The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel
DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take
ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame
audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was
the wall, and the one torn frame was a ring slot rewritten under the consumer's
read. Reference: WGC delivers at DWM/monitor cadence
(https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture)
and libyuv row-simple/fixed-point scaling
(https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule.
- DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage
math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame.
- 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of
serialized conversion/s; capacity sits just above the 60/s the pump can use.
- FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only
rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a
consumer still holds (session.LatestFrame survives conversions, dispatcher
preview lags) is never read-while-overwritten. Needs no consumer Release API.
- 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring
allocs — the device take is judgeable numerically.
Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings.
C4 (composite Epoch-cached downscale) deferred pending the device re-measure.
Local only, no push.
ty-1723/1726 device takes showed accelerated playback + audio tail cut-off:
a render overrun (~35ms vs the 16.6ms slot) SKIPPED the missed slots (slice 10's
freshness choice), so a 60fps-authoring pump wrote one frame per 35ms into a
60fps container — 1723: 697 frames/11.62s vs 11.84s audio; 1726: 163/2.72s vs
2.93s, video ending 0.21-0.24s early.
OBS never leaves a wall-time hole: the video thread emits one frame per tick and
a lagging producer DUPLICATES the newest frame ("lagged frames due to rendering
lag/stalls" — obs-output.c; "If the video frame queue is full, it will duplicate
the last frame" — docs.obsproject.com/backend-design). The pump's submit is now a
bounded catch-up over the missed slots (while now >= nextTick), fresh on the first,
repeated after — duration == wall, judder not fast-forward. Safe because Channel.
TryWrite never blocks (the take-9 smear was the blocking pipe-write; each emit is
nanoseconds). Burned frame index moved inside the loop: every emitted slot carries
its own +1 (also fixes the old unconditional pre-gate bump that gapped the judge
sequence on non-submitting fast-render iterations).
Good Dog test: Pump_Overrun_Renders_EmitsEverySlot_NotSkipped (60fps, 35ms render
cost, asserts >=0.65 of the wall slots emitted). 294/294 green, 0 warnings.
No push — web/A/V work is commit-local until greenlight.
The ~30Hz CapturePreviewAsync PNG poll capped real cadence at ~20Hz
(35–165ms full-HD encode+decode), so a 60fps widget still juddered at ~1/6
speed. Replaced polling with frame-driven capture of the composition
controller's root visual — the mechanism WebView2CompositionControl and
Flutter's webview_windows use (graphics_context.cc captures the root
surface_ visual via CreateGraphicsCaptureItemFromVisual; reference:
github.com/microsoft/Windows.UI.Composition.WinUI / flutter-internal
webview_windows). Frames now arrive at the renderer's own pace; capture
memory is epoch'd ring reuse + one crop-sized shared WriteableBitmap.
New Services/WebCaptureFrameSource.cs owns GraphicsCaptureItem + free-
threaded Direct3D11CaptureFramePool + session (Straight alpha readback,
per-frame FindContentBounds → CropBounds). WebView2Manager reworked around
per-session composition controllers + one UI-thread Compositor created via
the CoreMessaging CreateDispatcherQueueController P/Invoke (the 19041
projection lacks CreateOnCurrentThread); internal seam ctor
(Dispatcher, Func<string,IScreenCaptureSource>?) for hermetic tests.
CaptureScheduler.cs deleted; the three SetCaptureInterval cadence hooks
removed; InitWebView2() moved from MainWindow ctor to Loaded (a parent
HWND must exist for the composition controller); the hidden WebViewHostPanel
overlay deleted. TransparentBackgroundScript unchanged.
Tests: WebView2ManagerTests reworked — 4 control-size + scheduler tests
dropped, FindContentBounds tests moved to WebCaptureFrameSource, ONE
integration test (Frames_PublishCroppedPreview_And_CoalesceToLatest_CarryingCropBounds)
drives the seam with a FakeWebSource + background-STA DispatcherPump.
Suite 293/293, 0 warnings.
NOTE: composition path NOT yet verified on a device — the take is the
next step. Web work committed locally only (no push per standing rule).
Docs same-commit: ai.md Slice 14 + supersede marker on Slice 11, HANDOFF,
MyMistakes (CoreMessaging DQ + namespace-landmine recipe), TASK 17,
Controls/ViewModels/Services indexes.
Positive offsets still delay the whole mix via the delay line (lip-sync fix);
negative offsets now ARM once at StartLive and drop |N| ms off the pipe's write
head so audio events land earlier when audio runs BEHIND video. Slider relabeled
AUDIO SYNC, Min −500, locked while live/recording (IsEditMode). LayoutStore and
VM clamp to −500..500.
OBS reference for eat-the-head negative sync: https://obsproject.com/kb/obs-studio/buffering-time (negative sync values pull audio earlier by discarding buffered player audio).
Test: StartLive_NegativeOffset_AdvancesAudio_ByDroppingTheStreamHead (6x0.9 head
must be eaten before 0.2 bed reaches the wire).
AudioMixer.Start() begins mic/loopback capture at app startup to feed the
level meters, so both 2s ring buffers fill with pre-live audio. StartLive()
drained from the oldest tail sample, putting every recorded event ~2.2s late
in the audio track (confirmed by clap analysis + cross-correlation on two
takes: +2.11 to +2.22s).
Fix: AudioMixer.StartLive() now runs _micBuffer.Clear() + _loopbackBuffer.Clear()
immediately after the pipe starts, before the drain task runs. Recording now
begins at go-live; the <=10ms in-flight chunk evicted by Clear() is imperceptible.
Regression test: StartLive_DiscardsPreLiveBacklog_SoFirstAudioIsCurrent
saturates the loopback ring with stale 0.8 pre-live audio, then asserts the
wire carries fresh post-live 0.2 (max < 0.3).
Reference (external, per derivative-work rule): OBS 'Audio mixer' keeps its
buffers fed continuously and syncs the stream start timestamp at record time
rather than replaying pre-live capture; a go-live flush of the capture buffer
is the accepted pattern for live tools restarting a stream.
Docs: HANDOFF.md (fix shipped), MyMistakes.md (A/V sync measurement recipe).
silencedetect audio word times cross-correlated with per-frame video motion
(tblend=difference + signalstats) peaks at +22 frames: the 300ms TASK-22
offset plus ~70ms pipeline lag. 370ms reads as one full beat behind during
the ~2.3s pauses. Zeroing the offset targets the residual ~70ms.
- AudioSyncDelay.Configure reallocated/zeroed its buffer every ~10ms tick
(AudioMixer re-reads the UI setting each mix), so any non-zero sync offset
erased the just-written audio -> total silence. Now early-returns when the
delay samples are unchanged. Regression test proven both ways.
- SceneCompositor.BlitContentRaw defaulted cbW/cbH=0 when CropBounds is null
(regression from ed9d7c1) -> webcam blit to an empty rect = gray block.
Default to src.Width/Height. Regression test proven both ways.
- Truncated recordings: rawvideo mux stamps frames at declared 60fps by
arrival; a scene whose first layer is dynamic (hidden elements still count)
kills the bake cache -> full render ~35ms -> ~27fps submitted -> halved
file length. Static scenes bake once (246ms cold, then <1ms) -> 60fps,
full-length (probe + 13:53 take, 301/300 per 5s, 12.46s file from 12.3s
wall). FramePump.ProbeRender names the hot render path on slow frames.
The element-space raster builds on a TRANSPARENT base but BlitContentRaw's
partial-alpha branch applied the opaque-dst source-over blend: color premultiplied
by the sampled alpha, then alpha forced to 255. Pasting that raster saw a==255 and
straight-copied darkened ink over the scene — a recording box that the raw-bitmap
preview (correct alpha) never showed. Transparent margins and opaque content were
unaffected, which is why every take looped on the page/CSS while the capture was
transparent all along (15:51 dumps: alpha max 255, mean ~19, zero 57%).
BlitContentRaw now takes transparentDst; the raster call passes true and writes
straight color + straight alpha so the paste rows (BlendRowOpaque/Weighted) do the
real source-over onto the opaque master. Master paths byte-identical.
ONE integration test PasteCache_SemiTransparentLayer_RevealsBackdrop_NotOpaqueInk:
50%-blue over red reads (127,0,128) fixed vs (0,0,128) buggy — proven both ways
(verified by stashing the fix: fails before, passes after). Clean build, 0 warnings;
22/23 compositor-class tests pass, the sole failure the documented pre-existing
Composite_FullScene_MasterPixels pixel (1380,700).
Alpha-compositing model: standard source-over with producer-cached surfaces, the
OBS/libyuv paste model already cited in ai.md/MyMistakes (rawvideo recipe, row-blit
BLEND_NONE / straight-alpha branches); full story in MyMistakes (RESOLVED entry).
Per the good-dog rule: one integration test, memory updates (MyMistakes/ai.md/
HANDOFF) in the same commit.
The 15:34 take's box interior is the widget's OWN full-canvas opaque paint — a
black void with wide content strips at the top/bottom and a right-edge bar —
arriving AFTER the first-second blank-transparent dumps. A background-color-only
!important wipe (b4bba4b) can't touch it because CSS gradients/backdrops are
background-IMAGE, not background-color.
OBS's fix for this exact symptom is background-image:none (obsproject/obs-studio#6659:
"set the CSS for html and body to background: none !important"). Injection now wipes
background-image:none!important on html,body,html * alongside the color wipe; HTML
overlay art (<img>/DOM/CSS shapes) survives — that is browser-source semantics.
Also re-arms WidgetDumpRemaining=5 ~30s after nav so the next take dumps the real
document INSIDE the recording window (the previous dumps proved blank at +1s and
missed the later backdrop paint).
9/9 WebView2Manager tests, 0 warnings.
The real-widget dumps (12:54/13:50) proved the OLD inline
element.style.background='transparent' injection holds only while the page has
nothing to paint: the capture was alpha-transparent, yet the recording showed a
black opaque box over the whole element rect (1231,679 705x396) once the widget
connected and repainted a container background-COLOR — CapturePreviewAsync always
honors page CSS (MicrosoftEdge/WebView2Feedback specs/BackgroundColor.md), so any
page-painted background wins over DefaultBackgroundColor.
This is the OBS-solved class (all web-uri resources paint their own background):
browser sources use a Custom CSS override, and the decade-validated formula for
arbitrary pages is a pre-parse <style> with 'background-color: transparent
!important' — https://obsproject.com/forum/threads/translucent-transparent-browser-source.59549/
('body { background-color: rgba(0,0,0,0) !important }') plus the div-level variant
for stubborn widgets (woahtech.com OBS custom-CSS guide).
Injection is now an idempotent pre-parse style element wiping background-color on
html,body,html * with !important (outranks every page rule, runs before page parse
via AddScriptToExecuteOnDocumentCreatedAsync). Only background-COLOR is targeted —
background images and widget art survive. Regression test asserts the element-wide
!important form and that the losing inline form is gone.
9/9 WebView2Manager tests, 0 warnings.
The one-shot dump + alpha log was bound to the FIRST capture ever = the initial
about:blank placeholder (alpha=0, rgb=0, 5/5 sessions) — a blind instrument. The
pre-parse transparency injection (1a39b09) is intact but was NEVER verified
(MyMistakes '? VERIFY'), and the 2026-09-10 take still shows a black, opaque
block. Stop guessing: NavigationCompleted for a non-about:blank document now
arms WidgetDumpRemaining=5; the next 5 post-paint captures dump
%TEMP%\ytLive-web-<id>-w1..5.png + shared AlphaStats(min/max/mean/zero%) +
FindContentBounds to startup.log.
The stats line alone names the branch: zero%≈opaque ⇒ the injection did not hold
for this widget's CSS (fix: !important stylesheet / chroma-key); large zero% +
tight bounds ⇒ capture IS transparent and the black lives in the compositor
blend. No new test: WebView2 runtime is not instantiable in the suite and the
re-point is log-only; 8/8 WebView2 tests still pass.
The recorded audio is full-length silent AAC (−91dB, 1124 frames/23.95s) —
the named pipe carried ~9.1MB of zeros the entire take. The mixer's loop
ran, the pipe connected, ffmpeg read it, and the file is full-length
silence. Sources started fine ('Audio: using system default mic...' logged),
no failure callbacks fired, and the existing integration test
(Mix_WithFiltersDuckAndGain_Lands_On_AudioPipe) proves the loop→pipe path
carries real audio when fed — the fault is capture-side.
Added permanent per-5s live-loop telemetry to startup.log so the next take
names the exact stage without new code:
Audio live: pipe connected= dropped writes= micLevel= loopLevel= drained
N/N samples peakMix N (per 5s)
- AudioMixer.FillAndMix now returns (MicRms, MicDrained, LoopDrained) for
the accumulation; StartLive/StopLive log start/stop lines.
- NamedPipeAudioWriter.DroppedWrites: nonzero = audio dropped before ffmpeg
connected (names 'pipe never connected' stage).
Likely root cause: WASAPI loopback captures the default render endpoint —
if audio plays on a non-default device the recording is silently silent.
The fix requires device enumeration + selection (slice 13 candidate).
Suite 290/291 — same sole pre-existing compositor pixel failure.
The recording is 60fps but WebView2 capture was a blind 100ms DispatcherTimer = 10Hz;
each captured web frame repeated ~6x into the file caps web animation at the capture
rate, not the page's (user: 'the animation appears to be too slow').
- New CaptureScheduler (Services/CaptureScheduler.cs): per-session dispatcher timer
that DROPS a tick while a capture is in flight (latest-wins, never queues) — the
guard that makes a higher cadence safe: concurrent full-HD PNG CapturePreviewAsync
calls (~10-30ms each, slow per WebView2Feedback#20) would stack CPU and publish
stale-after-fresh. Effective cadence = max(interval, capture duration).
- Cadence: SetCaptureInterval(33) on record/stream start, (200) idle — applied via
MainViewModel.Streaming.Operations.cs.
- De-throttle the hidden page: shared CoreWebView2Environment created BEFORE
EnsureCoreWebView2Async with --disable-backgrounding-occluded-windows
--disable-renderer-backgrounding --disable-features=CalculateNativeWinOcclusion.
Off-screen WebView2 is a hidden page when the host window is unfocused/covered and
Chromium then parks rAF and clamps timers to ~1s (WebView2Feedback#1172/#3070,
Chrome-88 timer-throttling blog).
- Telemetry: first 30 captures per session log elapsed ms (PNG encode + decode) to
startup.log — that decides whether ~30Hz stays or drops to ~20Hz; the FramePump
drops frames (never time-lapses, slice 10) if UI-thread GC churn starves it.
- ONE test: CaptureScheduler_Drops_Ticks_While_Capture_InFlight_And_Resumes
(deterministic TCS-driven, no WebView2 runtime). Suite 290/291 — sole failure the
pre-existing compositor pixel test.
- Docs same-commit: ai.md slice 11, MyMistakes.md, HANDOFF.
References: https://github.com/MicrosoftEdge/WebView2Feedback/issues/1172https://github.com/MicrosoftEdge/WebView2Feedback/issues/3070https://github.com/MicrosoftEdge/WebView2Feedback/issues/20https://developer.chrome.com/blog/timer-throttling-in-chrome-88
slice 9 made the DURATION right but content still hiccuped; aggregates (301/300, uniform
file PTS) could not see it. Measured root cause: FfmpegEncoder.SubmitFrameAsync BLOCKED
on WriteAsync(8.3MB)+FlushAsync when ffmpeg lagged the pipe, and the burst while-loop
re-wrote that same stale composite per crossed slot — frozen runs.
OBS shape (derivative, wrapped pre-1.0): the encoder queue in libobs/obs-encoder.c —
encoder thread never couples back into the video thread; overflow = dropped data, never
a frozen producer. https://github.com/obsproject/obs-studio/blob/master/libobs/obs-encoder.c
- FfmpegEncoder: SubmitFrameAsync is now an enqueue (ArrayPool copy) into a bounded
Channel (cap 120) drained by its own task; drop-newest + count when full;
StopAsync flushes the queue then EOF (TryComplete). IFfmpegEncoder.DroppedFrames.
- FramePump: ONE fresh composite per iteration (burst loop deleted); worst-submit stat,
stall logger (>2x interval names the stage), dropped/stalls in stats.
- Burned-in 6-digit dot-matrix frame counter (white box, bottom-right) on every composite
— the clock-independent judge replacing the WSL ticker: +1/frame, jumps = counted drops.
- ONE new test Backpressure_QueueOverflow_DropsFrames_AndNeverBlocks (slow-sink fake:
submit never blocks, drops counted, stop flushes exactly submitted-minus-dropped).
- Full suite 290 tests, 289 pass — sole failure the pre-existing compositor pixel test.
- Docs same-commit: ai.md slice 10 (+ encoder/stop-note corrections), MyMistakes point 8,
HANDOFF.
Audio untouched (queued follow-up); web overlay still frozen pending timing closure.
Root cause (measured, not guessed): the take-3 'rebase on overrun' reset
nextTick to wall-now, erasing every missed slot. Pump delivered 215-219/300
per 5s (~43fps) but rawvideo carries no timestamps — ffmpeg muxes by frame
count at -framerate 60, so every recording played fast with honest-looking
stats. A second pacer, ffmpeg -re, throttled the demux separately
('Resumed reading ... after a lag' 0.79s->4.82s).
Fix follows libobs video-io.c (https://github.com/obsproject/obs-studio)
- the video thread never resets its deadline; one frame per interval slot,
a late render repeats content (judder), never skips time:
1. FramePump: count-based emission, while(now>=nextTick){Submit; nextTick+=I}
2. FfmpegArgs: -re removed - the pump is the pacer
Muxed duration is now frame-count/fps == wall time by construction. Covers
game background + webcam (one shared pump). 0 warnings; suite 288/289 -
the one failure (Composite_FullScene_MasterPixels 1380,700) also fails with
this change stashed: pre-existing, untouched, recorded as follow-up.
The widget page paints html/body opaque, so CapturePreviewAsync output had
zero alpha-0 margins — every take since bccdb48 built the crop on 'the
capture is transparent' and never verified it. WebView2 spec is explicit:
'WebView will always honor a webpage's background content' and
DefaultBackgroundColor only shows through pages with no background style
(MicrosoftEdge/WebView2Feedback specs/BackgroundColor.md). The late
ExecuteScriptAsync injection (NavStarting/NavCompleted) ran after page CSS
and lost the fight.
Fix via CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync — the
documented pre-parse hook ('before the HTML document has been parsed and
before any other script included by the HTML document is run'), the same
mechanism OBS user.css uses. Nav handlers kept as post-load re-assertion.
Restores the take-23/24 stride-correct crop path (my take-25 removal was
wrong — it regressed the bounding box into a solid black box).
Adds one-shot diagnostics: raw capture PNG -> %TEMP%/ytLive-web-<id>.png
plus alpha min/max/mean/%zero and FindContentBounds result in startup.log,
so the next run proves the capture is transparent instead of guessing.
Build 0 warnings, 289/289 tests pass. MyMistakes.md records the full chain.
User tested: 'didn't work at all, and broke additional crap.' Second failed remedy for
web-source-over-webcam dark composite -> AGENTS.md spin guard: no third guess; next attempt
requires external research (OBS/CEV browser-source transparency in recorded output, cited URL)
plus the blank-page falsifier. HANDOFF OPEN list updated: research-first step added, #11 noted
landed as d35823a.
Take 11 (c10ce06c) validated the off-UI architecture: typical frames land
work ~10ms + wait ~6.8ms = 16.7 exactly on the deadline; 212/300 best yet.
The ENTIRE remaining gap is periodic 35-65ms render spikes that WORSENED
across the take (189 -> 147) — the signature of gen2 GC pauses. Biggest
churner is structural: the screen capture minted a fresh ~8.3MB byte[] per
DWM frame (~500MB/s of LOH), a producer OBS never does (it owns fixed
surface pools).
- ScreenCaptureFrameSource: 4-deep buffer ring with size-matched slots (a
<=17ms consumer cannot be lapped at 60Hz) + reused downscale row scratch.
- VideoFrame.Epoch: monotonic per producer frame. The paste cache keys on
array IDENTITY, so recycled arrays MUST be distinguished — epoch joins the
PasteKey. Producers handing fresh arrays leave it 0 (key unchanged effect).
- Stats print 'gen2 +N' per 5s window: next take acquits or convicts GC
without another guess (rule: prove the stage).
- Test (the ONE): PasteCache_RecycledArrayWithNewEpoch_ReRasterizes_NotStaleHits
— same array, new content, bumped epoch; fails on the old key by
construction. 37/37 compositor/pump, clean build.
- Next suspect if gen2 stays hot: the 10Hz WebView2 capture (full-canvas PNG
decode + fresh arrays on the UI thread) — recorded, untouched.
Creator audio ask queued in the same working session (+40% post-mix master
gain before the -1dBFS limiter) lands as its own commit next.
Take 10 (59a02a5b, slice 6) finally produced a self-contradicting stat: render
22.4ms + submit 2.5 against a 16.7ms deadline, yet avg wait 10ms — a rebasing
pacer CANNOT sleep after a blown deadline. The wait was queue time: StartAsync
fires from a UI command handler, and async continuations re-capture the current
SynchronizationContext — the 'WPF-free, hermetic' frame pump had been rendering
ON THE DISPATCHER behind the live preview the entire starvation saga. OBS keeps
obs_graphics_thread/video_thread off-UI for exactly this reason (dedicated
threads; see docs.obsproject.com/backend-design 'Libobs Threads').
- FramePump: _pumpTask = Task.Run(() => PumpAsync(...)) — null context inside,
every continuation stays on the pool.
- Audited, not ignored, what that exposes: StaticPixelCache.Get now locks (pool
miss-decodes raced UI callers); ChatOverlayLayer.RenderFrame checks its cache
off-thread but marshals the rare raster MISS to the dispatcher (DrawingVisual
+ RenderTargetBitmap are UI-thread objects) and re-validates there; pump
events already marshal in the VM.
- GCLatencyMode.SustainedLowLatency for the pump's life (restored in finally).
- Stats gained 'worst render Xms' — bimodal averages hid per-tick spikes.
- Webcam routes through the paste cache (the IsOpaque bypass re-sampled ~156k
px every tick even between identical device frames).
ONE integration test: Pump_Produces_OffTheStartingContext — an inline-pumping
SynchronizationContext makes the old construction run the resolver on the
starting thread by capture; the loop must never. 70/70 per-class green, clean
build 0 warnings. Docs same commit (ai.md slice 7, TASKS take-11 gate,
MyMistakes #6, HANDOFF). take 11: ~300/300 + honest wait -> saga closed,
Unit B (two-line top bar spec, fully captured) starts.
Take 9's numbers were decisive: paste cache moved work to ~25ms/frame but the
period stayed ~37ms. The missing ~12ms per tick is Task.Delay rounding every
sub-tick request up to the Windows system-clock tick (~15.6ms default —
documented: learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay).
A frame finishing 3ms early requested 3ms and slept 15.6. Producer capped at
~27fps no matter how fast the compositor got — which is why two real render
fixes read as 'zero change' in playback. Game-loop/OBS canon for this
(stackoverflow.com/questions/5441464; learn.microsoft.com/en-us/windows/win32/
api/timeapi/nf-timeapi-timebeginperiod): raise the timer resolution for the
session, sleep only the bulk of the remainder, and SPIN the last ~2ms across
the deadline.
- FramePump: timeBeginPeriod(1) on entering the pump loop, timeEndPeriod(1) in
the finally; pacing = bulk _pacingDelay(ahead - 2ms) + bounded Thread.SpinWait
tail; blown deadlines rebase unchanged (never burst).
- Stats now report avg wait: render+submit+wait must equal the period — the
accounting is closed, no stage can hide in an unmeasured gap again.
- Webcam dropped its IsOpaque paste-cache bypass: it re-sampled ~156k px every
tick even between identical device frames; cached paste beats the sampler on
hits, costs the same on misses.
52/52 per-class green (pacing + pixel suites unchanged — output byte-stable),
clean build 0 warnings. Docs same commit (ai.md slice 6, TASKS.md take-10
gate, MyMistakes #3 + renumber, HANDOFF). User's top-bar spec remains next in
queue (Unit B) — re-sent many times, captured, no open questions.
The stamped build settled what slices 3-4 could not: chat cache works (resolve
~0.0ms) but render stayed 26-27ms -> 124-135/300. The cost was the compositor
re-rasterizing EVERY layer every tick: this Live scene re-samples chat (159k) +
web widget (271k) + image (95k) + cam (156k) ~ 680k px @ ~38ns — for layers
whose pixels do not change between chat/web/cam updates.
OBS shape: cache the surface, paste per tick. BlitCachedLayer rasterizes a
non-opaque layer ONCE into an element-space, transparent-based frame keyed by
(source-array identity, src W/H, ceil'd dst rect, round, mirror), then pastes:
integer position, row alpha-blend, opacity applied at paste. Producers hand out
fresh immutable arrays -> array-identity keys cannot serve stale content; dict
bounded (48, clears whole). Drag/opacity live in paste params, not keys, so
editing stops triggering resamples too. Opaque backdrop keeps the memcpy path;
the webcam keeps the direct path via its IsOpaque flag (revisit if take 9 is
borderline).
ONE integration test: PasteCache_RepeatRender_IsByteIdentical_And_ContentChange-
Propagates (byte-exact raster-vs-paste incl. round-clip margins, new-array
propagation); existing pixel suite guards sampler semantics. 85/85 across
compositor/pump/chat/capture/session classes, clean build 0 warnings. Also:
BuildStampTests.cs was written last commit but never staged — its own scope-check
slip, added here (the run had used the on-disk file; tracked now).
Docs same commit: ai.md slice 5 + stale 'general path only 130k' claim corrected,
TASKS.md take-9 gate, MyMistakes recipe (prove the stage; a fix that doesn't move
the stat wasn't the bottleneck), HANDOFF. take 9 expectation: 300/300, render
<= ~8ms -> saga closes, Unit B starts.
Take 6 measured render 35-41ms — WORSE than take 5's 25.5 — and the run could
not be attributed to a binary: exe mtime != build contents (incremental builds
serve stale exes; a source edit without rebuild is a silent old binary). Three
takes of a perf saga had been judged against builds nobody could prove.
- ytLive.csproj GenerateBuildStamp target: fresh GUID per compile (writes
obj/BuildStamp.g.cs -> Helpers/BuildStamp.Id/BuiltLocal). Deliberately defeats
incremental lies: every 'dotnet build' recompiles the app project.
- Wordmark shows the id as a superscript (TopBar.xaml, x:Static, 9px grey
BaselineAlignment=Superscript); startup.log records 'Build <id> (compiled
<time>)' so every take is cross-readable with the visible UI.
- FramePump stats split the tick: 'avg render Xms (resolve Y), avg submit Z' —
the resolver is timed separately (wrapper resolver on per-tick paths; bake
keeps the raw one) so take 7 names the hot half of 'render' with data.
- Fixed a latent transition-clock bug found on the way: lastTick now restarts
every frame (the branch rework had restarted it only during transitions,
letting a transition begun after idle complete instantly on its first Tick).
ONE integration test family: BuildStampTests (unit: shape) +
BuildStampDisplayTests (RealApp, namescoped FindName on TopBar proves the
wordmark SHOWS the id). 47/47 per-class green, clean build 0 warnings. Docs
same commit. User's top-bar/session spec (re-sent twice) + settled Q&A
decisions folded into HANDOFF Unit B — next work unit after take 7 verdict.
Take 5: render 58.9 -> 25.5ms (138/300, still ~2.2x). The blits were fixed; the
resolver was not: ResolveOutputFrame -> RenderChatBox ran a FULL WPF raster
(FormattedText + RenderTargetBitmap + CopyPixels + channel swap) EVERY tick
whenever the chat buffer was non-empty — and the buffer survives sessions, so
even a signed-out record-only take paid it. Established answer (OBS text
sources): re-render on change, blit the cache every tick.
ChatOverlayLayer: content version bumped from Messages.CollectionChanged
(covers adds, the 500-cap removal, the fade Clear from any caller) + a config
key (size + all Chat* appearance props); RenderFrame returns the cached
VideoFrame by identity until either changes (compositor only reads cached
frames). Conservative ordering (version latched BEFORE render) makes a
mid-render message re-render next tick, never serve stale.
ONE integration test: ChatOverlayLayerCacheTests (RealApp, real renderer):
Same() for unchanged inputs, NotSame() on message/config change, null on
empty. Full regression green (62 across touched classes), clean build 0
warnings. Accepted cost pending take 6: one ~15-25ms tick per arriving
message; if live-chat bursts sag n/300, next slice = debounced off-tick
re-render. Docs same commit: ai.md pipeline section, TASKS.md TASK 18,
MyMistakes recipe (raster-on-change + session-surviving-buffer trap),
HANDOFF (take 6 -> then Unit B, spec unchanged).
Two defects made the producer 17x slow (37s record -> 2.1s/127-frame file,
rawvideo stamps by arrival): FramePump slept the FULL interval after each
render (period = render+submit+interval) and SceneCompositor did per-pixel
float sampling + Math.Round blends over all 2.07M master pixels, scanning the
whole destination per overlay (258ms avg render vs 1.5ms submit).
Both solutions are established, not invented — researched before coding per
the derivative-work rule:
- deadline pacing: OBS libobs/media-io/video-io.c video_thread (nextTick +=
intervalTicks, sleep only the remainder, rebase on overrun, never burst)
- row blits: libyuv pattern (BSD-3, chromium.googlesource.com/libyuv/libyuv)
— 1:1 aligned identity fast path, per-pixel alpha branch, integer
fixed-point blend, overlay clipped to the intersection rect, skip the dead
black pre-fill when the backdrop covers
ONE integration test: Pump_Paces_To_The_Deadline_Compensating_Render_Cost
(lands after a fake-seam lesson: pacing fakes must await, not complete
synchronously, or the pump loop runs inline on StartAsync and hangs vstest).
Clean build 0 warnings; FramePumpTests 10/10, SceneCompositor/SceneGraph/
SocialBar/StretchMath 20/20. Docs same commit: ai.md pipeline section,
TASKS.md TASK 18 (webcam-in-output + rename modal verified from take 3),
MyMistakes recipe, HANDOFF rewritten. Take 4 pending on the user's machine.
All headless-testable TASK 21 decoder/mechanism slices shipped; the remaining UI
picker slice (AddMedia command + file dialog + Acquire/Release + Source.MediaIsLooping
-> IMediaFrameSource.Looping) is a GUI feature and is spec'd in HANDOFF for native
Windows build + verification. TASKS status + HANDOFF updated.
- IMediaFrameSource gains bool Looping.
- MediaVideoSource ctor takes Func<IDecodeProcess> processFactory instead of a
single IDecodeProcess: a System.Diagnostics.Process can't be re-Start()ed, so
each loop pass creates a fresh decoder. Decode wrapped in do-while(Looping):
restart on natural EOF instead of raising Completed.
- Production wiring (MainViewModel media factory): passes the process factory
AND FfmpegFrameRateProbe -- closes the slice-2b gap where production had no
probe and therefore no pacing.
- Tests: loop test (single frame re-emits across passes, Completed only when
loop cleared); fakes updated for the new interface member. Media tests 12/12,
build 0 warnings.
Wiring Source.MediaIsLooping into the flag needs a manager-level per-path loop
provider -> lands with the UI-picker (acquisition) slice.
Derivative reference: looping media by restarting decode on EOF, standard in
playback/overlay tooling (OBS media source repeat).
- MediaVideoSource takes optional IFrameRateProbe? + Func<TimeSpan,CancellationToken,Task>? delay
seams (default Task.Delay); probes FPS once in RunAsync, delays by 1/fps after each
emitted frame. No probe/unknown fps -> no pacing (ffmpeg pipe backpressure throttles).
- Test: MediaVideoSource_PacesFramesByProbedFps (fake probe returns 1000fps + recording
delay; one delay per frame ~= 1ms). Media tests 5/5, build 0 warnings.
Derivative reference: per-frame delay pacing of decoded output, standard in media playback.
- FfmpegFrameRateParser (pure): prefers avg_frame_rate= then r_frame_rate=,
rational N/N/M, unknown/0 -> null.
- IFrameRateProbe + FfmpegFrameRateProbe: derives sibling ffprobe.exe from the
located ffmpeg dir, reuses the IDecodeProcess seam for the ffprobe subprocess
text; null if ffprobe absent.
- FfmpegLocator now also extracts ffprobe.exe (ProbeFileName) from the pinned
archive, conditional so old caches without it degrade to no pacing.
- Tests: 6 pure parser units + 1 probe integration via fake locator/process;
FfmpegLocatorTests still green. 15/15, 0 warnings.
Pacing (probe->delay) is slice 2b. Derivative reference: standard ffprobe
avg_frame_rate probing used across OBS/media tooling.
Wire the MediaVideoSourceManager into the live pipeline (new partial
ViewModels/MainViewModel.Media.cs mirrors Webcam/Background partials):
readonly _mediaManager assigned in the core ctor (decode at 1920x1080 master,
dispatcher-coalesced preview), OnMediaPreviewBitmapChanged adopts the shared
bitmap onto every IsMediaSource with a matching MediaPath, a
Source { IsMediaSource, MediaPath } case in ResolveOutputFrame feeds
GetLatestFrame, and the manager is disposed on shutdown.
Source.DisplaySource now routes VideoImageSource for MediaSource (Source.cs),
so media previews/canvas show decoded video.
Compositor needs no change: BlitContent UniformToFill-scales any frame to the
element rect. Decoder/manager/seam derived-work mirrors ScreenCaptureManager
(citation in commit for slice-step-3).
Test (one unit for this change, mirroring the live-capture case):
Source_DisplaySource_IsVideoImageSource_WhenMediaSource. 14/14 media +
DisplaySource tests pass, build 0 warnings. Docs (TASKS/HANDOFF/ai.md) updated.
Adds the codec-agnostic decoder half of the media source. Spawns ffmpeg
with -f rawvideo -pix_fmt bgra (reusing the already-shipped ffmpeg via
IFfmpegLocator) and drains the raw BGRA stdout pipe into VideoFrames.
- Services/RawVideoFrameReader.cs: pure rawvideo BGRA stream -> frames
(partial reads kept across Feed; no ffmpeg needed to test)
- Services/IDecodeProcess.cs + FfmpegDecodeProcess.cs: binary-stdout
subprocess seam, mirror of the encoder's IEncoderProcess
- Services/MediaVideoSource.cs: owns the decode, raises FrameReady/Completed
- MediaVideoSourceTests: 3 pure reader + 1 integration (fake decode
process through the real source loop, frames in order)
Reference (external scan): ffmpeg rawvideo pipe decode is the canonical
codec-agnostic frame feeds pattern (ffmpeg docs -f rawvideo; how OBS/media
pipelines push frames to a compositor). Verified: 4/4 tests, build 0 warnings.
Native-FPS pacing + resolver/compositor wiring are the next slice.
- Add ElementKind (Static/Dynamic) to SceneElement base; Source/WebcamSceneConfig classify
- Services/SceneGraph.cs: owns the Scenes collection (ViewModel's Scenes delegates to it),
the element mutation surface (Add/Insert/Remove/Move, each invalidating the bake), and the
queries that were scattered LINQ (GetBackground/GetWebcam/GetChatBoxes/GetSplitPoint/IsStatic)
- SceneCompositor: split-aware BakeStaticBase + CompositeLayers + Render(.., staticBase, split);
builds/caches the static base below the split point in source-rect space
- FramePump: optional SceneGraph -> optimized bake+composite path; falls back to full render
- MainViewModel: routes element mutations through the graph; invalidates the bake on static
layout/opacity/visibility/useDefaultBackground changes and after background heal/Ensure
- Integration test SceneGraphTests.BakedStaticBase_WithDynamicLayer_CompositesCorrectly
Derivative survey (mandated): tried before writing — OBS does per-source opacity/visibility
caching and static-scene baking; this mirrors OBS's 'cached static source' optimization.
3 documented defensive deviations from the spec (ChatOverlayLayer stays decoupled;
background helpers stay VM-static for direct testability; full facade peel deferred post-1.0)
in TASKS.md + ai.md. Scope check passed; clean build 0 warnings.