ty-1723/1726 device takes showed accelerated playback + audio tail cut-off:
a render overrun (~35ms vs the 16.6ms slot) SKIPPED the missed slots (slice 10's
freshness choice), so a 60fps-authoring pump wrote one frame per 35ms into a
60fps container — 1723: 697 frames/11.62s vs 11.84s audio; 1726: 163/2.72s vs
2.93s, video ending 0.21-0.24s early.
OBS never leaves a wall-time hole: the video thread emits one frame per tick and
a lagging producer DUPLICATES the newest frame ("lagged frames due to rendering
lag/stalls" — obs-output.c; "If the video frame queue is full, it will duplicate
the last frame" — docs.obsproject.com/backend-design). The pump's submit is now a
bounded catch-up over the missed slots (while now >= nextTick), fresh on the first,
repeated after — duration == wall, judder not fast-forward. Safe because Channel.
TryWrite never blocks (the take-9 smear was the blocking pipe-write; each emit is
nanoseconds). Burned frame index moved inside the loop: every emitted slot carries
its own +1 (also fixes the old unconditional pre-gate bump that gapped the judge
sequence on non-submitting fast-render iterations).
Good Dog test: Pump_Overrun_Renders_EmitsEverySlot_NotSkipped (60fps, 35ms render
cost, asserts >=0.65 of the wall slots emitted). 294/294 green, 0 warnings.
No push — web/A/V work is commit-local until greenlight.
The ~30Hz CapturePreviewAsync PNG poll capped real cadence at ~20Hz
(35–165ms full-HD encode+decode), so a 60fps widget still juddered at ~1/6
speed. Replaced polling with frame-driven capture of the composition
controller's root visual — the mechanism WebView2CompositionControl and
Flutter's webview_windows use (graphics_context.cc captures the root
surface_ visual via CreateGraphicsCaptureItemFromVisual; reference:
github.com/microsoft/Windows.UI.Composition.WinUI / flutter-internal
webview_windows). Frames now arrive at the renderer's own pace; capture
memory is epoch'd ring reuse + one crop-sized shared WriteableBitmap.
New Services/WebCaptureFrameSource.cs owns GraphicsCaptureItem + free-
threaded Direct3D11CaptureFramePool + session (Straight alpha readback,
per-frame FindContentBounds → CropBounds). WebView2Manager reworked around
per-session composition controllers + one UI-thread Compositor created via
the CoreMessaging CreateDispatcherQueueController P/Invoke (the 19041
projection lacks CreateOnCurrentThread); internal seam ctor
(Dispatcher, Func<string,IScreenCaptureSource>?) for hermetic tests.
CaptureScheduler.cs deleted; the three SetCaptureInterval cadence hooks
removed; InitWebView2() moved from MainWindow ctor to Loaded (a parent
HWND must exist for the composition controller); the hidden WebViewHostPanel
overlay deleted. TransparentBackgroundScript unchanged.
Tests: WebView2ManagerTests reworked — 4 control-size + scheduler tests
dropped, FindContentBounds tests moved to WebCaptureFrameSource, ONE
integration test (Frames_PublishCroppedPreview_And_CoalesceToLatest_CarryingCropBounds)
drives the seam with a FakeWebSource + background-STA DispatcherPump.
Suite 293/293, 0 warnings.
NOTE: composition path NOT yet verified on a device — the take is the
next step. Web work committed locally only (no push per standing rule).
Docs same-commit: ai.md Slice 14 + supersede marker on Slice 11, HANDOFF,
MyMistakes (CoreMessaging DQ + namespace-landmine recipe), TASK 17,
Controls/ViewModels/Services indexes.
Positive offsets still delay the whole mix via the delay line (lip-sync fix);
negative offsets now ARM once at StartLive and drop |N| ms off the pipe's write
head so audio events land earlier when audio runs BEHIND video. Slider relabeled
AUDIO SYNC, Min −500, locked while live/recording (IsEditMode). LayoutStore and
VM clamp to −500..500.
OBS reference for eat-the-head negative sync: https://obsproject.com/kb/obs-studio/buffering-time (negative sync values pull audio earlier by discarding buffered player audio).
Test: StartLive_NegativeOffset_AdvancesAudio_ByDroppingTheStreamHead (6x0.9 head
must be eaten before 0.2 bed reaches the wire).
AudioMixer.Start() begins mic/loopback capture at app startup to feed the
level meters, so both 2s ring buffers fill with pre-live audio. StartLive()
drained from the oldest tail sample, putting every recorded event ~2.2s late
in the audio track (confirmed by clap analysis + cross-correlation on two
takes: +2.11 to +2.22s).
Fix: AudioMixer.StartLive() now runs _micBuffer.Clear() + _loopbackBuffer.Clear()
immediately after the pipe starts, before the drain task runs. Recording now
begins at go-live; the <=10ms in-flight chunk evicted by Clear() is imperceptible.
Regression test: StartLive_DiscardsPreLiveBacklog_SoFirstAudioIsCurrent
saturates the loopback ring with stale 0.8 pre-live audio, then asserts the
wire carries fresh post-live 0.2 (max < 0.3).
Reference (external, per derivative-work rule): OBS 'Audio mixer' keeps its
buffers fed continuously and syncs the stream start timestamp at record time
rather than replaying pre-live capture; a go-live flush of the capture buffer
is the accepted pattern for live tools restarting a stream.
Docs: HANDOFF.md (fix shipped), MyMistakes.md (A/V sync measurement recipe).
silencedetect audio word times cross-correlated with per-frame video motion
(tblend=difference + signalstats) peaks at +22 frames: the 300ms TASK-22
offset plus ~70ms pipeline lag. 370ms reads as one full beat behind during
the ~2.3s pauses. Zeroing the offset targets the residual ~70ms.
- AudioSyncDelay.Configure reallocated/zeroed its buffer every ~10ms tick
(AudioMixer re-reads the UI setting each mix), so any non-zero sync offset
erased the just-written audio -> total silence. Now early-returns when the
delay samples are unchanged. Regression test proven both ways.
- SceneCompositor.BlitContentRaw defaulted cbW/cbH=0 when CropBounds is null
(regression from ed9d7c1) -> webcam blit to an empty rect = gray block.
Default to src.Width/Height. Regression test proven both ways.
- Truncated recordings: rawvideo mux stamps frames at declared 60fps by
arrival; a scene whose first layer is dynamic (hidden elements still count)
kills the bake cache -> full render ~35ms -> ~27fps submitted -> halved
file length. Static scenes bake once (246ms cold, then <1ms) -> 60fps,
full-length (probe + 13:53 take, 301/300 per 5s, 12.46s file from 12.3s
wall). FramePump.ProbeRender names the hot render path on slow frames.
The element-space raster builds on a TRANSPARENT base but BlitContentRaw's
partial-alpha branch applied the opaque-dst source-over blend: color premultiplied
by the sampled alpha, then alpha forced to 255. Pasting that raster saw a==255 and
straight-copied darkened ink over the scene — a recording box that the raw-bitmap
preview (correct alpha) never showed. Transparent margins and opaque content were
unaffected, which is why every take looped on the page/CSS while the capture was
transparent all along (15:51 dumps: alpha max 255, mean ~19, zero 57%).
BlitContentRaw now takes transparentDst; the raster call passes true and writes
straight color + straight alpha so the paste rows (BlendRowOpaque/Weighted) do the
real source-over onto the opaque master. Master paths byte-identical.
ONE integration test PasteCache_SemiTransparentLayer_RevealsBackdrop_NotOpaqueInk:
50%-blue over red reads (127,0,128) fixed vs (0,0,128) buggy — proven both ways
(verified by stashing the fix: fails before, passes after). Clean build, 0 warnings;
22/23 compositor-class tests pass, the sole failure the documented pre-existing
Composite_FullScene_MasterPixels pixel (1380,700).
Alpha-compositing model: standard source-over with producer-cached surfaces, the
OBS/libyuv paste model already cited in ai.md/MyMistakes (rawvideo recipe, row-blit
BLEND_NONE / straight-alpha branches); full story in MyMistakes (RESOLVED entry).
Per the good-dog rule: one integration test, memory updates (MyMistakes/ai.md/
HANDOFF) in the same commit.
The 15:34 take's box interior is the widget's OWN full-canvas opaque paint — a
black void with wide content strips at the top/bottom and a right-edge bar —
arriving AFTER the first-second blank-transparent dumps. A background-color-only
!important wipe (b4bba4b) can't touch it because CSS gradients/backdrops are
background-IMAGE, not background-color.
OBS's fix for this exact symptom is background-image:none (obsproject/obs-studio#6659:
"set the CSS for html and body to background: none !important"). Injection now wipes
background-image:none!important on html,body,html * alongside the color wipe; HTML
overlay art (<img>/DOM/CSS shapes) survives — that is browser-source semantics.
Also re-arms WidgetDumpRemaining=5 ~30s after nav so the next take dumps the real
document INSIDE the recording window (the previous dumps proved blank at +1s and
missed the later backdrop paint).
9/9 WebView2Manager tests, 0 warnings.
The real-widget dumps (12:54/13:50) proved the OLD inline
element.style.background='transparent' injection holds only while the page has
nothing to paint: the capture was alpha-transparent, yet the recording showed a
black opaque box over the whole element rect (1231,679 705x396) once the widget
connected and repainted a container background-COLOR — CapturePreviewAsync always
honors page CSS (MicrosoftEdge/WebView2Feedback specs/BackgroundColor.md), so any
page-painted background wins over DefaultBackgroundColor.
This is the OBS-solved class (all web-uri resources paint their own background):
browser sources use a Custom CSS override, and the decade-validated formula for
arbitrary pages is a pre-parse <style> with 'background-color: transparent
!important' — https://obsproject.com/forum/threads/translucent-transparent-browser-source.59549/
('body { background-color: rgba(0,0,0,0) !important }') plus the div-level variant
for stubborn widgets (woahtech.com OBS custom-CSS guide).
Injection is now an idempotent pre-parse style element wiping background-color on
html,body,html * with !important (outranks every page rule, runs before page parse
via AddScriptToExecuteOnDocumentCreatedAsync). Only background-COLOR is targeted —
background images and widget art survive. Regression test asserts the element-wide
!important form and that the losing inline form is gone.
9/9 WebView2Manager tests, 0 warnings.
The one-shot dump + alpha log was bound to the FIRST capture ever = the initial
about:blank placeholder (alpha=0, rgb=0, 5/5 sessions) — a blind instrument. The
pre-parse transparency injection (1a39b09) is intact but was NEVER verified
(MyMistakes '? VERIFY'), and the 2026-09-10 take still shows a black, opaque
block. Stop guessing: NavigationCompleted for a non-about:blank document now
arms WidgetDumpRemaining=5; the next 5 post-paint captures dump
%TEMP%\ytLive-web-<id>-w1..5.png + shared AlphaStats(min/max/mean/zero%) +
FindContentBounds to startup.log.
The stats line alone names the branch: zero%≈opaque ⇒ the injection did not hold
for this widget's CSS (fix: !important stylesheet / chroma-key); large zero% +
tight bounds ⇒ capture IS transparent and the black lives in the compositor
blend. No new test: WebView2 runtime is not instantiable in the suite and the
re-point is log-only; 8/8 WebView2 tests still pass.
The recorded audio is full-length silent AAC (−91dB, 1124 frames/23.95s) —
the named pipe carried ~9.1MB of zeros the entire take. The mixer's loop
ran, the pipe connected, ffmpeg read it, and the file is full-length
silence. Sources started fine ('Audio: using system default mic...' logged),
no failure callbacks fired, and the existing integration test
(Mix_WithFiltersDuckAndGain_Lands_On_AudioPipe) proves the loop→pipe path
carries real audio when fed — the fault is capture-side.
Added permanent per-5s live-loop telemetry to startup.log so the next take
names the exact stage without new code:
Audio live: pipe connected= dropped writes= micLevel= loopLevel= drained
N/N samples peakMix N (per 5s)
- AudioMixer.FillAndMix now returns (MicRms, MicDrained, LoopDrained) for
the accumulation; StartLive/StopLive log start/stop lines.
- NamedPipeAudioWriter.DroppedWrites: nonzero = audio dropped before ffmpeg
connected (names 'pipe never connected' stage).
Likely root cause: WASAPI loopback captures the default render endpoint —
if audio plays on a non-default device the recording is silently silent.
The fix requires device enumeration + selection (slice 13 candidate).
Suite 290/291 — same sole pre-existing compositor pixel failure.
The recording is 60fps but WebView2 capture was a blind 100ms DispatcherTimer = 10Hz;
each captured web frame repeated ~6x into the file caps web animation at the capture
rate, not the page's (user: 'the animation appears to be too slow').
- New CaptureScheduler (Services/CaptureScheduler.cs): per-session dispatcher timer
that DROPS a tick while a capture is in flight (latest-wins, never queues) — the
guard that makes a higher cadence safe: concurrent full-HD PNG CapturePreviewAsync
calls (~10-30ms each, slow per WebView2Feedback#20) would stack CPU and publish
stale-after-fresh. Effective cadence = max(interval, capture duration).
- Cadence: SetCaptureInterval(33) on record/stream start, (200) idle — applied via
MainViewModel.Streaming.Operations.cs.
- De-throttle the hidden page: shared CoreWebView2Environment created BEFORE
EnsureCoreWebView2Async with --disable-backgrounding-occluded-windows
--disable-renderer-backgrounding --disable-features=CalculateNativeWinOcclusion.
Off-screen WebView2 is a hidden page when the host window is unfocused/covered and
Chromium then parks rAF and clamps timers to ~1s (WebView2Feedback#1172/#3070,
Chrome-88 timer-throttling blog).
- Telemetry: first 30 captures per session log elapsed ms (PNG encode + decode) to
startup.log — that decides whether ~30Hz stays or drops to ~20Hz; the FramePump
drops frames (never time-lapses, slice 10) if UI-thread GC churn starves it.
- ONE test: CaptureScheduler_Drops_Ticks_While_Capture_InFlight_And_Resumes
(deterministic TCS-driven, no WebView2 runtime). Suite 290/291 — sole failure the
pre-existing compositor pixel test.
- Docs same-commit: ai.md slice 11, MyMistakes.md, HANDOFF.
References: https://github.com/MicrosoftEdge/WebView2Feedback/issues/1172https://github.com/MicrosoftEdge/WebView2Feedback/issues/3070https://github.com/MicrosoftEdge/WebView2Feedback/issues/20https://developer.chrome.com/blog/timer-throttling-in-chrome-88
slice 9 made the DURATION right but content still hiccuped; aggregates (301/300, uniform
file PTS) could not see it. Measured root cause: FfmpegEncoder.SubmitFrameAsync BLOCKED
on WriteAsync(8.3MB)+FlushAsync when ffmpeg lagged the pipe, and the burst while-loop
re-wrote that same stale composite per crossed slot — frozen runs.
OBS shape (derivative, wrapped pre-1.0): the encoder queue in libobs/obs-encoder.c —
encoder thread never couples back into the video thread; overflow = dropped data, never
a frozen producer. https://github.com/obsproject/obs-studio/blob/master/libobs/obs-encoder.c
- FfmpegEncoder: SubmitFrameAsync is now an enqueue (ArrayPool copy) into a bounded
Channel (cap 120) drained by its own task; drop-newest + count when full;
StopAsync flushes the queue then EOF (TryComplete). IFfmpegEncoder.DroppedFrames.
- FramePump: ONE fresh composite per iteration (burst loop deleted); worst-submit stat,
stall logger (>2x interval names the stage), dropped/stalls in stats.
- Burned-in 6-digit dot-matrix frame counter (white box, bottom-right) on every composite
— the clock-independent judge replacing the WSL ticker: +1/frame, jumps = counted drops.
- ONE new test Backpressure_QueueOverflow_DropsFrames_AndNeverBlocks (slow-sink fake:
submit never blocks, drops counted, stop flushes exactly submitted-minus-dropped).
- Full suite 290 tests, 289 pass — sole failure the pre-existing compositor pixel test.
- Docs same-commit: ai.md slice 10 (+ encoder/stop-note corrections), MyMistakes point 8,
HANDOFF.
Audio untouched (queued follow-up); web overlay still frozen pending timing closure.
Root cause (measured, not guessed): the take-3 'rebase on overrun' reset
nextTick to wall-now, erasing every missed slot. Pump delivered 215-219/300
per 5s (~43fps) but rawvideo carries no timestamps — ffmpeg muxes by frame
count at -framerate 60, so every recording played fast with honest-looking
stats. A second pacer, ffmpeg -re, throttled the demux separately
('Resumed reading ... after a lag' 0.79s->4.82s).
Fix follows libobs video-io.c (https://github.com/obsproject/obs-studio)
- the video thread never resets its deadline; one frame per interval slot,
a late render repeats content (judder), never skips time:
1. FramePump: count-based emission, while(now>=nextTick){Submit; nextTick+=I}
2. FfmpegArgs: -re removed - the pump is the pacer
Muxed duration is now frame-count/fps == wall time by construction. Covers
game background + webcam (one shared pump). 0 warnings; suite 288/289 -
the one failure (Composite_FullScene_MasterPixels 1380,700) also fails with
this change stashed: pre-existing, untouched, recorded as follow-up.
The widget page paints html/body opaque, so CapturePreviewAsync output had
zero alpha-0 margins — every take since bccdb48 built the crop on 'the
capture is transparent' and never verified it. WebView2 spec is explicit:
'WebView will always honor a webpage's background content' and
DefaultBackgroundColor only shows through pages with no background style
(MicrosoftEdge/WebView2Feedback specs/BackgroundColor.md). The late
ExecuteScriptAsync injection (NavStarting/NavCompleted) ran after page CSS
and lost the fight.
Fix via CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync — the
documented pre-parse hook ('before the HTML document has been parsed and
before any other script included by the HTML document is run'), the same
mechanism OBS user.css uses. Nav handlers kept as post-load re-assertion.
Restores the take-23/24 stride-correct crop path (my take-25 removal was
wrong — it regressed the bounding box into a solid black box).
Adds one-shot diagnostics: raw capture PNG -> %TEMP%/ytLive-web-<id>.png
plus alpha min/max/mean/%zero and FindContentBounds result in startup.log,
so the next run proves the capture is transparent instead of guessing.
Build 0 warnings, 289/289 tests pass. MyMistakes.md records the full chain.
Slice 10 from today-slices (d1126dd): UVC cameras default MediaFrameReader to their
FIRST media type (YUY2 at crippled fps — C920 720p = 5fps). A 60fps container
then shows a 5-20fps face cam as "every few frames cut out". Fix:
VideoDeviceController.SetMediaStreamPropertiesAsync negotiated to best MJPEG
>=640x360 @>=30fps (<=1280 wide) BEFORE creating the frame reader, with
device-default fallback. Plus a startup.log truth line of the actual agreed media
type so the next diagnosis starts from facts (no more re-deriving the fps).
Build 0 warnings, 289 tests.
Slice 10 from today-slices (d1126dd): the paste cache minted a fresh raster
array per new source frame (webcam ~30 keys/s ≈ 18MB/s LOH churn →
gen2 pauses that ate camera frames). Fix: _elementRasters — re-rasterize
INTO the element's existing array when geometry holds, so steady-state
allocation ≈ 0. The paste cache keys on (array+epoch) still protect correctness;
_elementRasters[element] tracks the element's current raster so stale keys can
never paste a mid-update array. MaxPasteEntries 48→256 (was a leak guard,
not an allocation stream). Regression test:
PasteCache_RingRecycledSources_ReRasterInPlace_WithoutAllocationChurn.
Build 0 warnings, 289 tests.
TASKS.md is now the index (status table, open items, research pointer).
33 files: 32 task files + 1 research facts file. The full take-saga
narrative and all design decisions are preserved verbatim; the catalog
makes the queue readable without opening every task body. Schema and
AGENTS.md updated to reflect the new layout.
User tested: 'didn't work at all, and broke additional crap.' Second failed remedy for
web-source-over-webcam dark composite -> AGENTS.md spin guard: no third guess; next attempt
requires external research (OBS/CEV browser-source transparency in recorded output, cited URL)
plus the blank-page falsifier. HANDOFF OPEN list updated: research-first step added, #11 noted
landed as d35823a.
take-19-2/take-20 'webcam black square under the web widget': the capture was
alpha-cropped (FindContentBounds) and that CROP was fed to the compositor, whose
UniformToFill zoomed the opaque content box to cover the whole element rect the
moment the widget drew any content (idle transparent page = full-frame crop = the
correct viewport, hence take-1-good/take-2-bad). OBS model verified: the page is a
fixed 1920x1080 canvas and the element rect is a viewport onto it — measure with the
alpha crop (preview/selection), hand the composite the FULL canvas on an 8-deep ring
+ Epoch (same identity discipline as camera/screen). Regression test:
Composite_FullCanvasWebSource_TransparentMarginsRevealWebcam (the reveal contract:
margin pixel = webcam, badge pixel = widget).
Roll-forward of today-slices 6fd1d9c onto the slice-8 base, two-loop hunk dropped.
Release counter (per-build GUID read as noise; +1 per commit from git rev-list,
baseline 241 -> #13, generated by GenerateBuildStamp; GUID demotes to startup.log).
Tests pinned to Label/#N >= 13; wordmark display test asserts the Label.
Flash fix (take 14 finding): consumer holds must never outlive depth x source
period — 4 slots at high refresh lap ~27ms vs a <=50ms compositor read, so a
recycled slot flashed its new frame over the lagged old one. All shared rings 4->8
(OBS/overlay precedent for ring discipline).
Camera producer now rotates an 8-deep ring + Epoch instead of a fresh ~3.7MB
array per device frame (110-220MB/s LOH churn); WebView2 capture reuses a canvas
scratch + 8-deep output ring + a reused WriteableBitmap instead of two fresh
arrays + a fresh bitmap per 10Hz tick. Paste cache stays identity-keyed (Epoch).
Take 11 (c10ce06c) validated the off-UI architecture: typical frames land
work ~10ms + wait ~6.8ms = 16.7 exactly on the deadline; 212/300 best yet.
The ENTIRE remaining gap is periodic 35-65ms render spikes that WORSENED
across the take (189 -> 147) — the signature of gen2 GC pauses. Biggest
churner is structural: the screen capture minted a fresh ~8.3MB byte[] per
DWM frame (~500MB/s of LOH), a producer OBS never does (it owns fixed
surface pools).
- ScreenCaptureFrameSource: 4-deep buffer ring with size-matched slots (a
<=17ms consumer cannot be lapped at 60Hz) + reused downscale row scratch.
- VideoFrame.Epoch: monotonic per producer frame. The paste cache keys on
array IDENTITY, so recycled arrays MUST be distinguished — epoch joins the
PasteKey. Producers handing fresh arrays leave it 0 (key unchanged effect).
- Stats print 'gen2 +N' per 5s window: next take acquits or convicts GC
without another guess (rule: prove the stage).
- Test (the ONE): PasteCache_RecycledArrayWithNewEpoch_ReRasterizes_NotStaleHits
— same array, new content, bumped epoch; fails on the old key by
construction. 37/37 compositor/pump, clean build.
- Next suspect if gen2 stays hot: the 10Hz WebView2 capture (full-canvas PNG
decode + fresh arrays on the UI thread) — recorded, untouched.
Creator audio ask queued in the same working session (+40% post-mix master
gain before the -1dBFS limiter) lands as its own commit next.
Take 10 (59a02a5b, slice 6) finally produced a self-contradicting stat: render
22.4ms + submit 2.5 against a 16.7ms deadline, yet avg wait 10ms — a rebasing
pacer CANNOT sleep after a blown deadline. The wait was queue time: StartAsync
fires from a UI command handler, and async continuations re-capture the current
SynchronizationContext — the 'WPF-free, hermetic' frame pump had been rendering
ON THE DISPATCHER behind the live preview the entire starvation saga. OBS keeps
obs_graphics_thread/video_thread off-UI for exactly this reason (dedicated
threads; see docs.obsproject.com/backend-design 'Libobs Threads').
- FramePump: _pumpTask = Task.Run(() => PumpAsync(...)) — null context inside,
every continuation stays on the pool.
- Audited, not ignored, what that exposes: StaticPixelCache.Get now locks (pool
miss-decodes raced UI callers); ChatOverlayLayer.RenderFrame checks its cache
off-thread but marshals the rare raster MISS to the dispatcher (DrawingVisual
+ RenderTargetBitmap are UI-thread objects) and re-validates there; pump
events already marshal in the VM.
- GCLatencyMode.SustainedLowLatency for the pump's life (restored in finally).
- Stats gained 'worst render Xms' — bimodal averages hid per-tick spikes.
- Webcam routes through the paste cache (the IsOpaque bypass re-sampled ~156k
px every tick even between identical device frames).
ONE integration test: Pump_Produces_OffTheStartingContext — an inline-pumping
SynchronizationContext makes the old construction run the resolver on the
starting thread by capture; the loop must never. 70/70 per-class green, clean
build 0 warnings. Docs same commit (ai.md slice 7, TASKS take-11 gate,
MyMistakes #6, HANDOFF). take 11: ~300/300 + honest wait -> saga closed,
Unit B (two-line top bar spec, fully captured) starts.
Take 9's numbers were decisive: paste cache moved work to ~25ms/frame but the
period stayed ~37ms. The missing ~12ms per tick is Task.Delay rounding every
sub-tick request up to the Windows system-clock tick (~15.6ms default —
documented: learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay).
A frame finishing 3ms early requested 3ms and slept 15.6. Producer capped at
~27fps no matter how fast the compositor got — which is why two real render
fixes read as 'zero change' in playback. Game-loop/OBS canon for this
(stackoverflow.com/questions/5441464; learn.microsoft.com/en-us/windows/win32/
api/timeapi/nf-timeapi-timebeginperiod): raise the timer resolution for the
session, sleep only the bulk of the remainder, and SPIN the last ~2ms across
the deadline.
- FramePump: timeBeginPeriod(1) on entering the pump loop, timeEndPeriod(1) in
the finally; pacing = bulk _pacingDelay(ahead - 2ms) + bounded Thread.SpinWait
tail; blown deadlines rebase unchanged (never burst).
- Stats now report avg wait: render+submit+wait must equal the period — the
accounting is closed, no stage can hide in an unmeasured gap again.
- Webcam dropped its IsOpaque paste-cache bypass: it re-sampled ~156k px every
tick even between identical device frames; cached paste beats the sampler on
hits, costs the same on misses.
52/52 per-class green (pacing + pixel suites unchanged — output byte-stable),
clean build 0 warnings. Docs same commit (ai.md slice 6, TASKS.md take-10
gate, MyMistakes #3 + renumber, HANDOFF). User's top-bar spec remains next in
queue (Unit B) — re-sent many times, captured, no open questions.
The stamped build settled what slices 3-4 could not: chat cache works (resolve
~0.0ms) but render stayed 26-27ms -> 124-135/300. The cost was the compositor
re-rasterizing EVERY layer every tick: this Live scene re-samples chat (159k) +
web widget (271k) + image (95k) + cam (156k) ~ 680k px @ ~38ns — for layers
whose pixels do not change between chat/web/cam updates.
OBS shape: cache the surface, paste per tick. BlitCachedLayer rasterizes a
non-opaque layer ONCE into an element-space, transparent-based frame keyed by
(source-array identity, src W/H, ceil'd dst rect, round, mirror), then pastes:
integer position, row alpha-blend, opacity applied at paste. Producers hand out
fresh immutable arrays -> array-identity keys cannot serve stale content; dict
bounded (48, clears whole). Drag/opacity live in paste params, not keys, so
editing stops triggering resamples too. Opaque backdrop keeps the memcpy path;
the webcam keeps the direct path via its IsOpaque flag (revisit if take 9 is
borderline).
ONE integration test: PasteCache_RepeatRender_IsByteIdentical_And_ContentChange-
Propagates (byte-exact raster-vs-paste incl. round-clip margins, new-array
propagation); existing pixel suite guards sampler semantics. 85/85 across
compositor/pump/chat/capture/session classes, clean build 0 warnings. Also:
BuildStampTests.cs was written last commit but never staged — its own scope-check
slip, added here (the run had used the on-disk file; tracked now).
Docs same commit: ai.md slice 5 + stale 'general path only 130k' claim corrected,
TASKS.md take-9 gate, MyMistakes recipe (prove the stage; a fix that doesn't move
the stat wasn't the bottleneck), HANDOFF. take 9 expectation: 300/300, render
<= ~8ms -> saga closes, Unit B starts.
Take 6 measured render 35-41ms — WORSE than take 5's 25.5 — and the run could
not be attributed to a binary: exe mtime != build contents (incremental builds
serve stale exes; a source edit without rebuild is a silent old binary). Three
takes of a perf saga had been judged against builds nobody could prove.
- ytLive.csproj GenerateBuildStamp target: fresh GUID per compile (writes
obj/BuildStamp.g.cs -> Helpers/BuildStamp.Id/BuiltLocal). Deliberately defeats
incremental lies: every 'dotnet build' recompiles the app project.
- Wordmark shows the id as a superscript (TopBar.xaml, x:Static, 9px grey
BaselineAlignment=Superscript); startup.log records 'Build <id> (compiled
<time>)' so every take is cross-readable with the visible UI.
- FramePump stats split the tick: 'avg render Xms (resolve Y), avg submit Z' —
the resolver is timed separately (wrapper resolver on per-tick paths; bake
keeps the raw one) so take 7 names the hot half of 'render' with data.
- Fixed a latent transition-clock bug found on the way: lastTick now restarts
every frame (the branch rework had restarted it only during transitions,
letting a transition begun after idle complete instantly on its first Tick).
ONE integration test family: BuildStampTests (unit: shape) +
BuildStampDisplayTests (RealApp, namescoped FindName on TopBar proves the
wordmark SHOWS the id). 47/47 per-class green, clean build 0 warnings. Docs
same commit. User's top-bar/session spec (re-sent twice) + settled Q&A
decisions folded into HANDOFF Unit B — next work unit after take 7 verdict.
Take 5: render 58.9 -> 25.5ms (138/300, still ~2.2x). The blits were fixed; the
resolver was not: ResolveOutputFrame -> RenderChatBox ran a FULL WPF raster
(FormattedText + RenderTargetBitmap + CopyPixels + channel swap) EVERY tick
whenever the chat buffer was non-empty — and the buffer survives sessions, so
even a signed-out record-only take paid it. Established answer (OBS text
sources): re-render on change, blit the cache every tick.
ChatOverlayLayer: content version bumped from Messages.CollectionChanged
(covers adds, the 500-cap removal, the fade Clear from any caller) + a config
key (size + all Chat* appearance props); RenderFrame returns the cached
VideoFrame by identity until either changes (compositor only reads cached
frames). Conservative ordering (version latched BEFORE render) makes a
mid-render message re-render next tick, never serve stale.
ONE integration test: ChatOverlayLayerCacheTests (RealApp, real renderer):
Same() for unchanged inputs, NotSame() on message/config change, null on
empty. Full regression green (62 across touched classes), clean build 0
warnings. Accepted cost pending take 6: one ~15-25ms tick per arriving
message; if live-chat bursts sag n/300, next slice = debounced off-tick
re-render. Docs same commit: ai.md pipeline section, TASKS.md TASK 18,
MyMistakes recipe (raster-on-change + session-surviving-buffer trap),
HANDOFF (take 6 -> then Unit B, spec unchanged).
Two defects made the producer 17x slow (37s record -> 2.1s/127-frame file,
rawvideo stamps by arrival): FramePump slept the FULL interval after each
render (period = render+submit+interval) and SceneCompositor did per-pixel
float sampling + Math.Round blends over all 2.07M master pixels, scanning the
whole destination per overlay (258ms avg render vs 1.5ms submit).
Both solutions are established, not invented — researched before coding per
the derivative-work rule:
- deadline pacing: OBS libobs/media-io/video-io.c video_thread (nextTick +=
intervalTicks, sleep only the remainder, rebase on overrun, never burst)
- row blits: libyuv pattern (BSD-3, chromium.googlesource.com/libyuv/libyuv)
— 1:1 aligned identity fast path, per-pixel alpha branch, integer
fixed-point blend, overlay clipped to the intersection rect, skip the dead
black pre-fill when the backdrop covers
ONE integration test: Pump_Paces_To_The_Deadline_Compensating_Render_Cost
(lands after a fake-seam lesson: pacing fakes must await, not complete
synchronously, or the pump loop runs inline on StartAsync and hangs vstest).
Clean build 0 warnings; FramePumpTests 10/10, SceneCompositor/SceneGraph/
SocialBar/StretchMath 20/20. Docs same commit: ai.md pipeline section,
TASKS.md TASK 18 (webcam-in-output + rename modal verified from take 3),
MyMistakes recipe, HANDOFF rewritten. Take 4 pending on the user's machine.