Creator: 'when I attempt to refresh my YPP page, I get an error about not being able to reach YouTube. Seriously?' Real log: channels.list failed (403) x3. Root cause #1 (the 403): channels.list?mine=true&part=statistics,auditDetails, contentDetails returns 403 insufficientPermissions when the token lacks the youtubepartner-channel-audit scope — which the auditDetails part ALONE requires, per the docs ('A request that retrieves the auditDetails part ... must provide an authorization token that contains the youtubepartner-channel-audit scope'). That scope is MCN partner tooling with a 2-week token-revocation rule; the app must not hold it. TASK-39's 'current scopes suffice, no re-consent' slice-1 claim was wrong for this part; mock-fake tests never touched the real API, so it shipped green and 403'd every refresh since 2026-09-22. https://developers.google.com/youtube/v3/docs/channels/list Fix: part=statistics,contentDetails only; standing flags removed from ChannelStatsService -> YppStatSnapshot surface -> YppTrackerViewModel -> drawer, replaced by an honest deep-link row ('Channel standing isn't exposed to YouTube apps — check the Earn page'). YppSnapshot standing columns stay (schema-stable, always false). channels.list failures now log the response BODY — the bare code could not name insufficientPermissions, which is what made this undiagnosable. Root cause #2 (found by the new Good Dog, masked by the 403): statistics come back as JSON STRINGS ('350'); raw GetInt64() throws. Tolerant ReadInt64 (ValueKind-first; JsonElement.TryGetInt64 THROWS on strings — type-in, not try-type). Good Dog: ChannelStatsServiceTests.CaptureCurrent_RequestsNoAuditDetails_AndStillParsesTheSnapshot (URL asserts no auditDetails + snapshot parses); YppPullOutTests fixture updated. Recipes for both 403/scope and statistics-strings entered in MyMistakes.md. Also shipped in the same commit (shared PreviewPane.xaml + ai.md): the audio-sync status dot removal from the #77 feedback round (creator: 'what is the point of the status light? Lose it') — IntToSyncBrushConverter deleted with it. verify.sh gate: 0 warnings, 316/316 pass, scope-check clean.
55 KiB
MyMistakes.md
Two jobs, distinguished by heading:
- Per-task failure log — updated before every commit touching that task: current iteration + why the last one failed. On task complete, committed AND pushed → truncate to this stub. A new task does NOT seed this file until its first failure.
- Recipes registry (DERIVED-SOLUTION RULE, see
AGENTS.md🔬) — the durable home for one-off derived solutions, recipes, and how-tos. The moment you work out a reusable solution, write it here in the same session. GREP THIS FILE FIRST when you hit a "I've done this before but have to figure it out again" wall. Recipe entries stay permanently (they are NOT truncated on task completion) — only the failure log truncates.
🔬 Recipes registry
YOUTUBE STATISTICS ARE JSON STRINGS; TryGetInt64 THROWS ON STRINGS (RECIPE)
YouTube Data API v3 returns statistics.subscriberCount / viewCount / videoCount as
JSON strings ("350"), not numbers. (2026-09-23: the YPP parse bug surfaced the
moment the auditDetails 403 stopped masking it.) Fix = a tolerant read that checks
ValueKind FIRST, then TryGetInt64 for Number, then long.TryParse(GetString())
for String. Trap within the fix: JsonElement.TryGetInt64 THROWS
InvalidOperationException on any non-Number token ("requires an element of type
'Number'") — it is try-type-in, not try-catch. Guard on value.ValueKind before
calling it; a leaked 403→parse chain means your "whole request failed" symptom can
paper over a second crash that only appears once the 403 is fixed (test the full
happy path, not just the error path).
CHANNELS.LIST auditDetails PART 403s THE WHOLE REQUEST WITHOUT A PARTNER SCOPE (RECIPE)
YouTube Data API v3 channels.list rejects the ENTIRE request with
403 insufficientPermissions if your part= list includes auditDetails but the
token lacks https://www.googleapis.com/auth/youtubepartner-channel-audit — the
docs' exact words: "A request that retrieves the auditDetails part for a channel
resource must provide an authorization token that contains the
youtubepartner-channel-audit scope". That scope is MCN content-partner tooling
(with a two-week token-revocation rule); normal-creator apps should NEVER ask for
it. 2026-09-23 real-log evidence: YPP refresh 403'd three times in a row while the
mock-fake tests stayed green ("current scopes suffice, no re-consent" was wrong).
Fixes: (1) drop auditDetails from part=; (2) log the response BODY — the bare
status code could not name insufficientPermissions, which is what made this
failure undiagnosable for days; (3) when a feature needs data no ordinary scope
grants, deep-link to the site instead of requesting the partner privilege.
TRANSITION(COMPLETE) RACES AUTOSTOP: PRE-FLIGHT lifeCycleStatus (RECIPE)
A blind liveBroadcasts.transition?broadcastStatus=complete POST can 403
invalidTransition even though the stream just ended normally — 2026-09-22 logged
it on EVERY session end. Cause: the broadcast's own state moves toward complete
via enableAutoStop/YouTube auto-complete; the transition method's errors doc
(https://developers.google.com/youtube/v3/live/docs/liveBroadcasts/transition)
shows invalidTransition = "can't transition from its current status". A blind
POST just races — read the broadcast's lifeCycleStatus first
(liveBroadcasts.list?part=status) and only POST complete from live/testing;
skip silently otherwise and let enableAutoStop finish it. Never throw on the
close-out either way.
YOUTUBE liveStreams LIST: healthStatus IS AN OBJECT, NOT A STRING (RECIPE)
YouTube Data API v3 liveStreams.list nests health under
status.healthStatus = {status, lastUpdateTimeSeconds, configurationIssues[]} —
the healthStatus element is an OBJECT, and configurationIssues[] lives
INSIDE it, not directly under status. Reading healthStatus.GetString() throws
System.Text.Json's The requested operation requires an element of type 'String', but the target element has type 'Object' — the exact log line seen 2026-09-22 on
every live health poll (two test sessions). The parse must read
healthStatus["status"] (and nested healthStatus["configurationIssues"]).
Something like REST-shape drift is a good reason to grep the API reference
(https://developers.google.com/youtube/v3/live/docs/liveStreams) before writing
parsers against a "remembered" shape — our own flat-string fixture was the wrong
assumption the whole time.
WINRT RESOURCE-ALLOCATION CALLS: WRAP PER-CALL, DEGRADE TO NEXT OPTION (RECIPE)
Creator callout (2026-09-15): WinRT/COM calls that allocate or start a resource — InitializeAsync,
CreateFrameReaderAsync, StartAsync, CreateReaderAsync & friends — are THROW-HEAVY. An unsupported
subtype/format, a device that just vanished, an access mode rejected mid-flight: these surface as
ArgumentException/E_INVALIDARG ("value does not fall within the expected range") or HRESULTs, NOT as
a returned status you can switch on. Relying on ONE outer catch to "handle failures" is not handling —
one rejected call inside a fallback ladder aborts the whole ladder and every untried option. Rule:
- Each allocation/start call inside a try/catch of its OWN, so a throw on candidate N falls through to candidate N+1 (log each rejection with its message/status; collect them for the final error string).
return/breakon success must be reached WITHOUT passing through afinallythat disposes the resource you just committed (classic reader/capture dispose-after-commit bug).- Unsubscribe + dispose the partial resource in the catch block when the subscription happened before the throwing call.
- The outer catch stays as the LAST-RESORT net for device-level errors, not the primary one.
- Same discipline applies to the frame-consumption side: teardown races reader threads (see HANDOFF crash follow-up) — a frame callback can't assume the pipeline is alive.
Applied in MediaCaptureFrameSource's reader-subtype ladder (2026-09-15, webcam-take fix).
"SAVE-RECORDING DIALOG IS CLIPPING CONTENT" — WHERE + HOW (RECIPE)
WHERE: the end-of-recording modal (creator: "when I end a recording the app throws up a
save-recording dialog with a default filename") is RenameRecordingDialog.xaml at the repo ROOT
(title "Rename Recording", x:Class="ytLive.RenameRecordingDialog"). It is NOT one of the
*Dialog.xaml files under Controls/ — grep for the window title, not the filename, when the
user names a dialog by what it does. Only fix dialogs the user actually named; do not "while I'm
here" bump sibling dialogs.
HOW: the dialog is ResizeMode="NoResize" with a fixed Height attribute; content client area
≈ Height − ~37px of window chrome. Sizing/margin changes pushed content down in Z-order until it
clipped — first the file-name textbox (230 too short), and after the +20% bump (→276) the creator
found the Cancel/Save button row had ALSO been obscured. Fixed Height to 304 (+10% more),
x:Name the last interactive control (SaveRecordingButton) so a test can measure it.
Verify (never eyeball-predict): RenameRecordingDialogSizingTests (RealApp host) instantiates
the real dialog, Show() + UpdateLayout(), then asserts the BOTTOM EDGE of each named control
(TransformToAncestor against the content root → TransformBounds(RenderSize).Bottom) is ≤ the
client ActualHeight. Every such dialog fix ships with that assertion for every control that was
reported obscured. (Client-area math means the button row needs MORE slack than the textbox: its
bottom sits deeper because of the Margin="0,20,0,0" above it.)
SPIN GUARD → RESOLVED — web overlay transparency + bounding box (RECIPE)
THE ONE ROOT CAUSE THAT EXPLAINS EVERY FAILED TAKE: WebView2's CapturePreviewAsync
produces an OPAQUE PNG. From the WebView2 spec (sender: MicrosoftEdge/WebView2Feedback
specs/BackgroundColor.md): "WebView will always honor a webpage's background content."
DefaultBackgroundColor = Transparent only shows through pages with NO background style —
the widget's own CSS paints html/body opaque. Every take below built on the false premise
"the capture has transparent margins, alpha=0"; it never did. FindContentBounds then had
no alpha-0 margins to find → wrong crop → black bounding box. The compositor blend saw
alpha=255 → black over webcam = "transparency broken". Same bug, three symptoms.
THE FIX (the OBS way, applied 2026-09-08): inject the transparency BEFORE the page
parses using CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync — documented to run
"before the HTML document has been parsed and before any other script included by the HTML
document is run" (learn.microsoft.com/dotnet/api/microsoft.web.webview2.core.corewebview2.addscripttoexecuteondocumentcreatedasync).
The old ExecuteScriptAsync on NavigationStarting/NavigationCompleted ran AFTER page
scripts/CSS → widget page repainted background opaque → lost the fight. OBS browser sources
do the same via a pre-parse user.css. Kept the nav handlers as a post-load re-assertion.
Take timeline (the honest record):
bccdb48(Aug 28): added FindContentBounds crop — correct idea (tight bbox, no dead space), but the capture was OPAQUE so the bbox math was built on nothing.5348b5c(Aug 28, 2 min later): reverted to full-frame no-crop — looked "good" for a full-bleed widget, but floating widgets regained dead space ("ghost boundary").- take-21 (
e002847) full canvas → shrunken/offset widget (UnifomToFill of whole canvas into a small element = lost resizing). - take-22 (
ed9d7c1) crop width used as canvas stride for buffer indexing → garbage. - take-23 (
f6802c7) stride fixed withsrc.Width; PasteKey lacked CropBounds → stale raster cache → stale crop. - take-24 (
081e4c1) CropBounds in PasteKey — STILL BROKEN because the SOURCE ALPHA WAS NEVER REAL. - take-25 (ME, this session — the user's "RE-INTRODUCING THE BOUNDING-BOX PROBLEM"): I removed FindContentBounds + the crop path entirely, betting full-canvas UniformToFill was the answer. It WASN'T — the source is opaque-black, so the element rendered as a SOLID BLACK BOX (the user's screenshot: "a black box in the lower right corner"). Killed the crop → dead space returned AND black box. The compositor math was ALREADY correct; gutting it was vandalism in response to a source-level bug.
Rules, self-inflicted:
- Instrument FIRST. This session added: first-capture PNG dump of the raw WebView2 PNG
(to
%TEMP%\ytLive-web-<id>.png) + alpha min/max/mean/%zero + FindContentBounds result logged to startup.log once per session. That's the diff between a five-take loop and a five-minute diagnosis. - When the same symptom loops across takes, the PREMISE is wrong, not the code — the capture being transparent was the load-bearing premise and it was never verified.
- Do not delete code paths that fix one axis (crop=bbox) while debugging another (source alpha). Revert scope creep; keep layer contributions separable.
2026-09-10 follow-up → NOW VERIFIED and hardened. The 12:54 and 13:50 sessions
showed the REAL widget document capturing TRANSPARENT (widget dumps w1..5 for
8d7234ec: alpha 100% zero, content 12×3; whole-file decodes of
%TEMP%\ytLive-web-*.png alpha max = 0) while the RECORDING kept showing a black,
opaque box over the element rect (crisp edges at the exact rect 1231,679 705×396 —
NOT the desktop showing through). Concluson: branch (a) — the OLD inline
element.style.background='transparent' injection only wins while the page has
nothing to paint; once the widget connects and paints its own container
background-COLOR, the capture goes opaque again → black box. The OBS-validated
answer (valid for arbitrary pages for a decade) is an injected pre-parse <style>
with !important beating every page rule:
https://obsproject.com/forum/threads/translucent-transparent-browser-source.59549/
(body { background-color: rgba(0,0,0,0) !important }) + the div-level variant for
stubborn widgets (woahtech.com OBS custom-CSS guide). Applied 2026-09-10 (b4bba4b):
injection appends a style element wiping background-color:transparent!important
on html,body,html *. Second half (2026-09-10, 15:34 take): a background-COLOR wipe
is NOT enough. The 15:34 interior ASCII shows a black void with bright content strips
at top/bottom + a right-edge bar — the widget's full-canvas CSS background-image
(gradient/backdrop) painted after connect. The OBS fix for that is
background: none !important / background-image:none!important
(obsproject/obs-studio#6659 — "set the CSS for html and body to background: none
!important"). Wipe both moving forward; overlay art is <img>/DOM and survives.
Also the dumps only covered +1s (blank-transparent); the widget's backdrop paint
arrives later, so the mid-recording dump set is now re-armed ~30s in. Verify take:
element rect shows scene bg behind the widget art (animation, no black void) → loop
closes. Self-inflicted again: the agent re-derived the whole transparency story
(recording pixel archaeology, compositor blend re-verification) instead of reading
this entry — the instrument said transparent because the capture had NOT been
repainted yet. Do NOT re-derive this story a third time.
2026-09-12 → RESOLVED — the box lived in the paste-cache RASTER, not the page.
The 15:51 facts (alpha max=255, mean ~19, zero 57%, white rounded panel, real
transparent margins) + preview-correct + recording-box meant the capture transparency was
REAL all along; the fracture sat in SceneCompositor.BlitContentRaw
(SceneCompositor.cs:419-550) — the sampler that builds the element-space paste-cache
raster, the ONE path the take loop never read in full. Its partial-alpha branch applied
the OPAQUE-dst blend onto a TRANSPARENT raster base: dst = (src*sa + dst*inv)/255
with dst black → color PREMULTIPLIED by sa, then dst[+3] = 255 — a 50%-alpha widget
pixel became darkened color + FULL alpha. At paste time BlendRowOpaque saw alpha 255 →
straight copy → the scene behind was overwritten by darkened ink. Fully-transparent
margins (alpha 0, continue) and fully-opaque content (alpha 255 branch) survived —
which is why every take showed a box while the dumps and the preview (raw WriteableBitmap,
unaffected by the compositor) stayed correct, and why the CSS-wipe fixes (b4bba4b /
0f72c53) were red herrings: they treated the page as the villain, but the capture was
transparent from the start. FIX: BlitContentRaw gained transparentDst=false;
the raster call site passes true and writes STRAIGHT color + straight alpha so the
paste rows (BlendRowOpaque/BlendRowWeighted) do the real source-over onto the opaque
master. The master paths are untouched (bit-identical). ONE test
PasteCache_SemiTransparentLayer_RevealsBackdrop_NotOpaqueInk fails on the old code with
exactly the bug encoded: 50%-blue over red reads (0,0,128) instead of (127,0,128) —
backdrop never shows through. Next: verify take (element rect shows the backdrop behind
the widget art), then push gate #1 clears.
2026-09-12 → COLLATERAL — the SAME transparency saga broke the webcam next.
Verify-take of the transparency fix showed a perfect flat gray rectangle where the webcam
should be (preview fine, recording flat — stdev <1 across the whole element, despite the
raw camera frame handed to the compositor sampling min=0/max=255 one line earlier in the
pipeline — proven by instrumenting BOTH ends before touching any code, not guessing).
Root cause: ed9d7c1 ("CropBounds metadata + Fill-style scaling... widget fills element
rect" — take-22 of the SAME transparency saga above) added cbX/cbY/cbW/cbH to
BlitContentRaw's general sampler but only assigned them inside the
src.CropBounds is {} cb branch; every CropBounds-LESS source (webcam, images — anything
but the web widget) fell through with cbW=cbH=0. Since sxCrop = sxNorm * cbW and
syCrop = syNorm * cbH, both were always 0, so sxCanvas/syCanvas collapsed to
cbX/cbY = (0,0) for every destination pixel — an entire scaled webcam sampled ONE
source corner pixel. The web widget (the only CropBounds-bearing source) was never
affected, which is exactly why the transparency fix's own test suite stayed green while
this broke. FIX: default cbW = src.Width, cbH = src.Height (cbX/cbY = 0) before
the branch, only overridden when CropBounds is actually present. ONE test
BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin (two-color split
source, scaled non-1:1 through the paste cache) fails on the old code — right-half pixel
reads the left half's color — and passes on the fix; proven both ways with a stash/build/
revert cycle, not by inspection alone.
Lesson: when a shared low-level sampler gains a new optional code path (CropBounds),
audit EVERY variable the new branch introduces for a safe default in the branch it did
NOT touch — an uninitialized-to-zero "crop region" silently means "sample only pixel
(0,0)", not "no crop." grep for this shape (var x = 0; followed by an if (cond) x = ...
with no else) whenever a conditional metadata field is added to pixel math.
Shrink / re-encode an image for the README (screenshots → small hero image)
Worked out 2026-08-29 (the recipe was NEVER recorded the first time it was done, so it had to be re-derived from scratch — that's the incident this entry exists to end).
Approach: a throwaway Windows-dotnet console app uses WPF's imaging stack
(System.Windows.Media.Imaging) — same framework the app runs on, zero NuGet
packages, high-quality downscale via TransformedBitmap. Screenshots compress
far smaller as JPEG than PNG (PNG 1400px = ~1.2MB; JPEG q82 1400px = ~188KB).
Recipe (run via the Windows dotnet host from WSL):
- Create
imgresize.csprojtargetingnet8.0-windowswith<UseWPF>true</UseWPF>(SDK controller). Put it in a Windows-visible temp path, e.g.C:\Users\gramp\AppData\Local\Temp\imgresize— NOT/tmp(Windows dotnet can't reach a Linux-only path reliably). Program.cs: loadBitmapImage(CacheOption=OnLoad→Freeze()), downscale withTransformedBitmap(src, new ScaleTransform(scale, scale))to max width (1400 for the README hero), encode withJpegBitmapEncoder { QualityLevel = 82 }, save.- Run:
"/mnt/c/Program Files/dotnet/dotnet.exe" run -c Release --project .
-- "C:\Users\gramp\Downloads\Screenshot 2026-08-29 075626.png"
"C:\Users\gramp\Documents\Code\projects\ytLive\docs\ytLlive-preview.jpg" 1400
- Point
README.mdat the.jpg(not.png).
Result: 3.2MB screenshot → 1400×794 → 188KB ytLlive-preview.jpg in docs/.
Measuring audio-video A/V sync from a recording (no ears needed) — RECIPE
Derived 2026-09-12/14 (the "audio cut off at the end" / "audio delayed" reports). You cannot "listen" to a take — measure it. Proven on two independent takes (talking+claps, clean-clap test) with a tight, matching result each time.
Method (FFmpeg probe + numpy, run from WSL):
- Click/hiss-SILENT test takes are the gold standard. Have the creator do a loud, single-frame-syncable event (one clap after ~10s of near-silence) — the video position of the hands-meet peak and the audio position of the transient are both unambiguous. That single event replaced counting words forever.
- Extract audio as raw mono PCM (
ffmpeg -vn -ac 1 -ar 48000 -f s16le) and compute a short window RMS envelope (5ms windows) in numpy. The clap is the global max (argmax); also print a coarse 100ms table — it shows the pre-clap artifacts (faint blips) vs the event (100× larger) vs true digital silence. - Video motion per frame —
-vf "tblend=all_mode=difference,signalstats,metadata=print"captured to stdout (NOTfile=inside the filter — the\\path escapes mangle; have PowerShell-style quoting bite;metadata=printwrites to ffmpeg's log, so redirect stdout to a file). ParseYAVGvalues; the clap is the frame with the motion spike (2.75 vs a ~0.1–0.5 animated-widget background). - Offset = audio_event_time − video_event_time; every positive second means audio is BEHIND video by that much. Cross-correlate the two full envelopes (60fps-resampled RMS vs YAVG, normalized) to double-check the single-event peak — the autocorr-style peak must be sharp (unique max at +133 frames, second-best ≈0.26).
- Sanity checks baked in: count frames (
nb_framesvs wall log) — if video is real-time you've ruled out the truncation bug as the cause; compare stream durations (audio > video by the lag is the SIGNATURE of a delayed audio tail, not truncation); checkdropped writes— dropped audio moves events EARLIER (opposite sign), so it can never explain a lag. A constant offset ≠ drift: fixed lag = buffering/backlog, growing offset = clock mismatch.
The ROOT CAUSE the recipe led to (record it so it's never re-derived): the mixer's
capture rings are fed from APP STARTUP (for the meters); StartLive never clears them,
so the drained audio trail begins ~a full ring-depth (2.0s) behind go-live. Rule: any
"live preview/drain" sink fed by a continuously-capturing buffer MUST clear the buffer
at go-live, or early output is stale backlog. The 2s ring + configured 300ms sync
delay + ~100ms pipeline = exactly the measured 2.11–2.22s. The lag ALSO scales with
"time since app launch" up to the ring depth — a take right after a restart reads as
~370ms while a fully-warmed app reads ~2.2s. Same code, wildly different numbers.
Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
Derived 2026-09-03 (take-3 diagnosis — the stats seam from 97ffc42 named the stage
in one line: 17/300 frames per 5s, avg render 258.1ms, avg submit 1.5ms).
Both halves were solved by OBS/libyuv long ago; do not re-derive:
-
Pacing is a DEADLINE, never a post-render sleep.
sleep(interval)after each frame makes the periodrender + submit + interval— the producer can hit ≤ half the declared rate even with a free render. OBS'svideo_thread(libobs/media-io/video-io.c) advances an absolutenextTick += intervalTicksand sleeps only the remainder. NEVER "rebase" the deadline to wall-now when it blows — the take-3 note I originally wrote here ("skip … the missed ticks (rebase)") was the PROVEN-wrong advice: resettingnextTickerases the missed slots, so a 43fps reality was authored into a 60fps container and every recording played ~1.4x fast (take 12/13, 2026-09-10). The rebase even kept the stats looking honest (render+wait == period, no loss) because a rebased frame is never "late". OBS keeps the counter MOVING and outputs ONE frame per interval tick — a late render shows as REPEATED footage (judder; duration stays correct), never a skip and never a wall-now reset. Critical with rawvideo: pts is stamped by ARRIVAL, so the muxed duration is pure frame count ÷ fps — the only way to make duration == wall time under ANY load is exactly one frame per interval slot. Corollary: never add a second pacer.ffmpeg -reon the rawvideo input throttles the pipe independently and fights the pump (its "Resumed reading … after a lag" grew 0.79s→4.82s in the same take). Removed; the pump IS the pacer. -
A 1080p frame is ~2.07M pixels — the hot path must be row-simple. Per-pixel
Math.Round+ float source-over in managed code costs ~100ns/px = the whole 258ms. libyuv's pattern (https://chromium.googlesource.com/libyuv/libyuv/): branch per pixel on source alpha (opaque → 4-byte copy, transparent → skip), integer fixed-point blend(s*a + d*(255-a) + 127)/255otherwise; and ALWAYS clip the loop to the intersection rect (our social-bar overlay scanned all 2M dst px for a 64px strip). A full-cover 1:1 blit also obsoletes the opaque-black pre-fill — skip dead writes. -
On Windows,
Task.Delayis a 15.6ms QUANTUM, not a timer. Any request under one system-clock tick sleeps a full tick (documented — learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay: "approximately 15 milliseconds on Windows systems"). A deadline pacer built on Task.Delay caps the producer at ~40fps-ish EVEN IF render is instant — take 9 proved the signature: work fell 26.5→22.4ms but the period sat at ~37ms (≈ one padded wait/frame), so two real optimizations read as "zero change". Frame-accurate loops (OBS/Chromium/game-loop canon — stackoverflow.com/questions/5441464) do:timeBeginPeriod(1)for the session (paired withtimeEndPeriod), sleep only the BULK of the remainder, SPIN the last ~2ms across the deadline. Diagnostic before touching the compositor again: period ≈ work + 15.6 → the SLEEP is the bug, not the work.Related (take 14, 2026-09-04): recycled ring buffers are a race you must SIZE, not just own. Deepening shared frames to kill GC churn (a fresh 8.3MB/tick array) hands out REUSED memory — the ring's depth × source period must EXCEED the worst consumer hold (compositor read + lagged UI preview copy), not just "a few frames". 4 slots at 144Hz capture laps in ~27ms vs a ≤50ms read: half a new screen frame flashed over an old one in the recording ("bits flashing over other bits"). Depth 8 everywhere (screen/camera/web output rings); the paste-cache Epoch still guards identity.
-
Hermetic pacing test: inject the delay seam to RECORD the requested TimeSpan and genuinely await it (
Task.Delay(d, ct)) — a fake that returnsTask.CompletedTasksynchronously makes the whole pump loop run onStartAsync's sync continuation and hang the test run (hit this 2026-09-03; the existing fakes all yield for exactly this reason). Assert the REQUESTED wait (< interval with a ≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines. -
Expensive content: raster on change, never on read (take 5, 2026-09-04). A source that updates once a minute (chat text!) must not full-rasterize (
FormattedText+RenderTargetBitmap+CopyPixels≈ 15-25ms) every compositor tick. OBS text sources re-render on property/message change; the per-tick pass blits the cache. Implement as: content version (collection-changed counter) + config key (size/appearance) → cached immutableVideoFramereturned by identity. Gotcha: buffers that SURVIVE sessions (the chat log) silently arm the per-tick cost even in flows that never touch the feature (signed-out record-only takes paid chat rendering!). -
An async loop started from a UI handler runs ON THE UI THREAD until you take it off.
awaitcontinuations re-capture the currentSynchronizationContext— the frame pump was started from a WPF command handler, so the "WPF-free, hermetic" compositor rendered and read capture state ON THE DISPATCHER, serialized behind the live preview itself, for the whole starvation saga. Thewaitstat caught it only when the numbers became self-contradictory (render 22 + wait 10 > any rebasing deadline — a blown deadline cannot sleep). OBS runsobs_graphics_thread/video_threadas dedicated threads for exactly this reason. Pattern:_task = Task.Run(() => Loop())(null context inside), then audit EVERY object the loop touches for UI affinity (RenderTargetBitmap / DrawingVisual / WriteableBitmap: marshal the work or the rare miss; plain locked byte[] lookups: fine) and pin it with a context test (Pump_Produces_OffTheStartingContext, inline-pumping SynchronizationContext that the old code failed by construction). Cost: takes 3–10. -
Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8). The chat raster fix was REAL but the composer blamed it for the residual slowness it did not own; two takes burned before the render/resolve split showed
resolve ≈ 0and pointed at the compositor pasting static layers per tick (BlitCachedLayerfinished the job OBS-style). Before shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck. -
An encoder refed from a real-time loop must ENQUEUE, never pipe-write in the loop (2026-09-10, take 15/16 — the "1...23...4...56..." smeared ticker). After slice 9 the recording played at the right DURATION but the content still hiccuped — and the aggregates (301/300, uniform file PTS, 15.6s wall vs 15.74s file) could NOT see it. The cause finally measured in
FfmpegEncoder.SubmitFrameAsync: theWriteAsync(8.3MB)+FlushAsyncto ffmpeg's stdin BLOCKS whenever the encoder lags the pipe, and the slice-9 burstwhile (now>=nextTick)then re-wrote that SAME stale composite for every slot that ticked past — frozen content runs. OBS's answering machinery is the encoder queue (libobs/obs-encoder.c): the encoder thread NEVER couples back into the video thread; overflow = dropped data, NEVER a frozen producer. Fixed as: boundedChannel<byte[]>(cap 120) + a dedicated drain task owning stdin,SubmitFrameAsync= copy-to-pool-array +TryWrite(drop-newest + count when full), stop flushes the queue then EOF. Rule: verify with a clock-independent judge. The WSL ticker that "proved" slice 9 has its own Host-timer jitter under Windows load — so this slice burns a dot-matrix_outputIndexinto the bottom-right of every composite; decoding the recording reads the honest sequence (+1/frame, jumps = counted drops) with no external clock involved. Take 17 must read +1/frame from that strip. (A whole-frame duplicate scan was tried and is UNRELIABLE here: the scene is always animating — session elapsed timer + REC pulse — so no two frames are ever byte-identical.)SUPERSEDED by slice 15 (2026-09-14): the last clause "pump emits ONE fresh composite per iteration (no burst re-write)" had it BACKWARDS for the deadline. Slice 10 chose a skip when the render overruns — an overrun's slots vanish from the file — and that AUTHORS ACCELERATED playback (my item-1 lesson already said it: "NEVER a skip … a late render shows as REPEATED footage"). Device takes proved it: one fresh frame per 35ms stall →
ty-20260914-1726= 163 video frames (2.72s) against 2.93s of audio, video ending 0.22s early ("audio speeds up then cuts off"). OBS's answer (docs.obsproject.com/backend-design: "If the video frame queue is full, it will duplicate the last frame";libobs/obs-output.ccounts "lagged frames due to rendering lag/stalls" — never a time-hole) is to fill each missed slot by DUPLICATING the newest frame; duration == wall, judder not fast-forward. That's what the pump's submit now does: a catch-up loop over the missed slots emitting this iteration's composite again, clamped to adeadlineNowcaptured once (bounded — the smear that justified slice 10 was the BLOCKING pipe-write re-copying during a long freeze; the queue makes each emit nanoseconds, so the burst is safe). The burned index moved inside the submit loop: EVERY emitted slot carries its own +1 (this also fixed the old unconditional_outputIndex++that gapped the sequence on non-submitting fast-render iterations). A/V sync must be re-clap-measured after this fix — the +0.6s audio-late reading on 1726 was confounded by the 1.1x acceleration.
Latest A/V sync numbers (2026-09-14, pre-slice-15 pacing): keep the recipe below honest on clap measures. ty-1723: 697 video frames (60fps) = 11.62s vs 11.84s audio. ty-1726 (clap): 163 frames = 2.72s vs 2.93s audio; clap audio env peak 1.655s vs video motion cluster 0.75–1.03s → single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late) — BOTH confounded by the acceleration; re-measure after slice 15.
Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
the freeze-audit recipe (RECIPE — re-deriving it cost this session):
The slice-15 build fixed pacing but the desktop layer of the recording was still
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
measurement said something different AND worse than the running render theory.
FramePump stall… worst render 33-36ms was real but MOOT: once the camera+desktop
take was decoded to raw frames, the desktop band was frozen 21s of 23.35s (90%)
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
CONVERSION was the wall: the monitor delivers at the 240Hz DWM cadence, the
source converts ONE frame at a time (_framePending latest-wins), and each
2560×1440→1920×1080 DownscaleBgra — naive double-per-pixel bilinear — cost
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
MediaCapture path) and audio were fine — exactly what the user reported.
The audit recipe (ffmpeg → raw gray → numpy):
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
python3 - <<EOF
import numpy as np
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
EOF
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
the compositor render stats led nowhere until this number existed. A per-row
split detector (cumsum of per-row diff-to-next minus diff-to-prev, argmax =
split row) then separated real mid-frame tears (score ≈ huge, both halves match
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
were the session UI, not tears.
The fix (this slice):
-
integer 8.8 fixed-point downscale, "shift only at the end" — the SAME math as
SceneCompositor.Bilinear(rounded both stages in one 16.8 scale) ported intoDownscaleBgra, dropping per-pixel doubles to row-walk integer ops. The capture ring already had the integer-bilinear lesson; the capture downscale itself was still the naive float twin of the 258ms disaster. -
throttle to the slot cadence (
MinConvertInterval = 10ms): the 240Hz arrival is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps pump can use. -
ring reuse-distance, not ownership (
FrameRingBuffer, redLine 4): the take-14 "depth × period" rule guards size; the slice-16 addition makes it structural — a slot is only rewritten ≥4 rents after its last hand-out, else a fresh buffer.session.LatestFramesurvives across conversions and the dispatcher preview copy lags, so "who released it" is unknowable without a consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742 tear (new-top/old-bottom midway) is that read-under-write closed. -
measure before trusting the inherited plan: the approved native-res capture + composite-side downscale (C1) was recast to "fix the downscale in place" — relocating a 30ms float downscale from the capture thread to the render thread and caching by Epoch only moves the same ~30ms cost into the slot budget. The measurement said the cost ITSELF was the enemy; keep the architecture, make the op fast.
Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a bug): Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new telemetry said it plainly:
conv avg 46-50ms, max ~61ms, skip busy 37-83— the GPU→CPU readback (CreateCopyFromSurfaceAsync), NOTDownscaleBgra, is ~45ms of that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight = readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two follow-on decisions fixed by evidence:- pool size is a clip, not a scale (near-miss). The "obvious" fix was shrinking the pool to the master size to read back less. Microsoft's screen-capture page forbids it: "the underlying Direct3D surface is always the size specified when creating … the Direct3D11CaptureFramePool. If content is larger than the frame, the contents are clipped." Shrinking to 1920×1080 would CROP a 1440p monitor, not scale — silently encode the desktop cut off. Readback must stay native; the lever is concurrency, not size. Rule: read the platform doc for the exact primitive before "fixing" the pool/format; scaling assumptions about capture APIs have been wrong twice now.
- overlap the readbacks + a monotonic publish gate. With up to 3 conversions in
flight, completions can land out of order; a slow OLDER readback finishing last
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
MonotonicGate(seq set viaInterlocked.Incrementbefore the copy, verified by compare-exchange publish) drops stale completions instead. Ring gets a lock because rents are now concurrent; downscale row-scratch became per-conversion locals. Lesson: keep a single "what is bound?" number per layer (telemetry line) before choosing between throughput and latency fixes — both previous slices picked the wrong slot ("render" vs "conversion") until the audit existed.
Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
integer bilinear) followed the same libyuv research
(https://chromium.googlesource.com/libyuv/libyuv/ — row.cc/scale.cc keep both
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
My first Bilinear shifted stage 1 back to 8-bit AND shifted the final result >>16 —
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
NOTHING (green webcam silently vanished from output; the pixel probes caught what
the eye in a 2x time-lapse would not). Rule: multi-stage fixed point shifts only
at the end; verify against a uniform-255 sample before believing it. Second trap:
a stale-byte sentinel test whose source pattern can generate the sentinel value
itself (0xAB was a legitimate x+y pixel) — pick the sentinel coprime/out-of-range
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
that HOLDS submitted frames now must snapshot them (Clone) once the producer
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
A web widget captured at 10Hz inside a 60fps recording plays at ~1/6 speed
Derived 2026-09-10 (the "widget animation too slow" report). The recording is 60fps and the
WebView2 capture loop was a blind 100ms DispatcherTimer = 10Hz — each captured web frame gets
repeated ~6× in the file, so whatever the page animates at, the OUTPUT is capped at the capture
cadence, not the page's. Two stacked throttles, both real:
- The capture rate is the hard ceiling. web-layer motion in the recording can never beat
CapturePreviewAsyncfrequency. But you can't just raise the timer: full-HD PNG capture costs ~10-30ms (WebView2Feedback#20: "CapturePreviewAsync … produces PNG/JPG and is very slow"), so concurrent captures stack CPU AND can publish stale-after-fresh. The mandatory shape is a latest-wins drop: at most one capture in flight per session; a tick during the in-flight window is DROPPED, never queued; effective cadence = max(interval, capture duration). Do this before anyone touches the interval constant. - Chromium throttles hidden pages. An off-screen WebView2 (we place it at (-5000,-5000)) is a
hidden page the moment the host window is unfocused or covered:
requestAnimationFrameparks and JS timers clamp to ~1s (WebView2Feedback#1172 — background-throttled rAF; #3070 — a WebView2 withVisibility.Collapsedslows timers to 1s; Chrome-88 blog — heavy timer throttling of hidden tabs). There is NO supported per-page opt-out (#5250 still open). The embedder answer is browser args on theCoreWebView2EnvironmentOptionscreated BEFOREEnsureCoreWebView2Async:--disable-backgrounding-occluded-windows --disable-renderer-backgrounding --disable-features=CalculateNativeWinOcclusion(the Electron/Streamlabs-class fix for occluded-window animation throttling). Share ONE environment across sessions (one browser process). - Measure before picking the cadence. Log the first ~30 captures' elapsed ms on the first recording run; PNG encode + WPF decode of 1920×1080 is the per-capture cost that decides whether ~30Hz is affordable or it must drop to ~20Hz. The FramePump drops frames (never time-lapses) if UI-thread GC churn starves it, so cost shows up as dropped-frame stats — read them.
Splitting a large file into partials — NEVER awk … > SRC while awking SRC
(2026-08-31, Commit D) Tried to split SocialsDialogViewModel.cs in one line:
{ awk '…' SRC; echo ""; awk '…2…' SRC; } > SRC. The first write truncated
SRC to 1 line, so the second awk read the already-truncated file → the whole
source was lost (1 line left). Recovered with git checkout -- SRC, then redid
it, but the same bug could have meant making it up from scratch.
Rule: when a cut needs N blocks from one source into N files, never write a
block back onto the source that the awks still read. Instead:
- Read the source once at the start into temp files (
mktemp -d, one file per block), with a$Dvariable you carry forward. - Verify block sizes (
wc -l) and brace balance (python3 -ccounting{/}) before touching any real file. - Then assemble each new file from
cat D/block …— never truncating the source until every read is done.
git status can't save you here if you don't notice until the file is gone —
git checkout -- <path> from the last commit is the recovery. Cheap insurance:
restore-then-retry, do it atomically from temp files the first time.
Verifying an ffmpeg decode contract from WSL (no real CLR needed)
(2026-08-31, TASK 21) When a change depends on ffmpeg producing output with an exact frame-size contract (rawvideo W×H×4 BGRA), you can prove the command + frame accounting here without any .NET process:
- Fetch a static Linux ffmpeg into
/tmp/opencode(no sudo needed):curl -sLO https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz && tar -xf … - Generate a tiny known clip:
ffmpeg -f lavfi -i "testsrc2=duration=1:size=640x360:rate=30" -pix_fmt yuv420p clip.mp4 - Decode with exactly the app's args:
-f rawvideo -pix_fmt bgra -vf scale=640:360 -an - Assert
total_bytes % (W*H*4) == 0(python3) → exact integer frames, no pad.
Why not a dotnet-spawned ffmpeg here: the only CLR on this box
(/home/gramps/bin/dotnet) is a Windows-bound shim — Process.Start resolves
paths to \\wsl.localhost\Debian\… and throws "not a valid application for this
OS" when handed a Linux ELF ffmpeg. So never plan to have dotnet exec a Linux
ffmpeg here; verify the contract with shell/python instead, and leave the
CLR→real-ffmpeg run to the native Windows suite.
WPF hit-test truth in tests: UIElement.InputHitTest, NOT VisualTreeHelper.HitTest
(2026-09-01, the RoundClip "known failure" post-mortem — a failure the map carried as "not a regression" for weeks without ever recording WHY.)
The trap: VisualTreeHelper.HitTest(window, pt) returned the window's
WebViewHostPanel overlay (IsHitTestVisible="False", Opacity=0, ZERO children — since the
2026-09-14 composition-capture redesign the overlay is DELETED, but the API lesson stands) for
EVERY point in the window — so a "corner is grabbable" assertion could never pass, and
it looked like a real interaction bug. The actual input pipeline (UIElement.InputHitTest,
what Mouse routing uses) correctly returned the element's Grid at elem-center/corner-in/
corner-exact and fell through to CanvasGrid just past the corner. The product was fine;
the TEST was probing an API that doesn't model input semantics.
Rule: any test asserting "where does a click land" uses window.InputHitTest(pt) +
IsDescendantOf — never VisualTreeHelper.HitTest.
Diagnosis recipe (how the lie was caught in ~3 probe cycles, no guessing): add a TEMP
probe [Fact] in the RealApp collection that hit-tests a spread of points
(elem-center / corner-in / corner-exact / corner-out / bg-center) and Assert.Fails with a
composed dump: per-point VTH hit + InputHitTest hit + ancestor chain (GetParent walk
with #Name) + panel properties (IsHitTestVisible/Opacity/children/actual size) +
TranslatePoint origins. Run the class alone, read the message, delete the probe.
Two sibling facts learned the same session (record-once):
- A UserControl owns its own XAML namescope — after extracting a region out of a window,
window.FindName("InnerPart")returns null; resolve the UserControl by its window-level name, thenpane.FindName("InnerPart"). And window-scope STYLES are invisible to a UserControl'sStaticResourceat parse time — move such styles toThemes/Controls.xaml(the app-scope rule exists for this). - Per-class
dotnet.exe vstestfrom WSL DOES execute the RealApp/MainWindowtests fine (they passed natively 2026-09-01) — only the FULL suite hangs (WASAPI startup). And the test process shares%APPDATA%\ytLlive\startup.logwith the real app: lines likecamera 'test-camera' failedare test noise, not DB state — to check pollution, query the DB directly (python3 sqlite3,SELECT DeviceId FROM Webcam), not the log.
A "known failure" label without a recorded cause = a bug on life support
(2026-09-01, the audio triple-take) One line — _delayedMix (nullable, added by TASK 22,
never initialized) dereferenced as delayed.Length — produced THREE symptoms that lived in
the map as two separate "pre-existing, do-not-chase" entries: (a)
Mix_HonorsProviderGains… "known failure", (b) AudioPipelineTests hangs when run at all
(a test reading a named pipe with no writer blocks — hung test ≠ flaky test, it's a starved
producer), (c) startup.log flooded "Audio live loop error" every 10ms (the mixer loop caught
and logged only ex.Message — stack thrown away).
Rules derived:
- NEVER label a test "known/pre-existing" without writing WHY (exception type + first app frame). An unexplained known-failure is deferred archaeology that hardens into fog.
- A test that waits on IPC + a producer whose output vanished are usually ONE bug — look for the producer before blaming the test.
- Catch-and-log-swallow of
ex.Messagehides root causes; log with stack (AppLog.Write(ex, ...)) and throttle (5s) instead of dropping or flooding. Bonus: the "failing" test encoded the MAP's contract (loopbackGain = GameAudioVolume); the code had drifted to unity on a disproven premise (loopback capture does NOT follow endpoint volume — creator's 20%-volume/pegged-meter observation killed it). The test was right all along — failing tests may be the last honest witnesses; interrogate, don't pardon.
Feature provenance: record WHO asked and WHY, in the task entry itself
(2026-08-31/09-01, the SYNC slider scare) TASK 22's lip-sync slider surfaced on the preview rail and the creator's reaction was "totally don't remember ordering that" — because the queue entry recorded WHAT shipped (a slider, 0-500ms, a converter class) but not WHO asked (the creator, explicitly, for OBS's delay-filter fix built natively). Eight days later his own request read like AI drift and nearly got deleted. Rule: the moment a creator-driven feature is queued or shipped, its entry carries a one-line provenance — who asked, what triggered it ("creator: OBS delay-filter lip-sync fix, native"). Features without attribution become roadmap orphans that get punted, removed, or re-litigated. Same disease as an unexplained "known failure" label — a fact recorded without its reason is a future argument.
An un-attributed build invalidated three takes of a perf saga — stamp the binary
(2026-09-04, takes 4–6) After each render-perf fix the creator "exed the code" and re-recorded, but
the exe timestamp ≠ binary contents (incremental builds reuse whatever compiles clean; a source edit
with no rebuild serves the OLD exe). Take 6 measured render WORSE than take 5 (35-41ms) and there was
no honest way to tell "the chat cache fix doesn't work" from "the fix was never running" — three
hours of diagnosis on an unattributable sample. Rule: if takes measure the app, EVERY build carries
an id and EVERY log line traces to it — GenerateBuildStamp (csproj) writes a fresh GUID per
compile (deliberately defeating incremental lies), the wordmark shows it as a superscript, startup.log
records Build <id> (compiled <time>). A perf claim without build attribution is a guess; ask for the
stamp BEFORE theorizing. (Also this session: a sentinel-byte test where the source pattern could
GENERATE the sentinel, and a fixed-point bilinear that shifted BOTH stages and silently drew nothing —
see the rawvideo recipe.)
Signed A/V sync: negative offset = eat the buffer head (OBS semantics), armed at go-live
(2026-09-14) The ring-backlog fix shrank A/V lag from ~2.2s to ~0.54s, and ~300ms of that was
baked into the DB (Audio.SyncOffsetMs=300) via a POSITIVE-only control. For audio running BEHIND
video you cannot push audio later — positive delay makes it worse. OBS's answer is a NEGATIVE offset
that eats the buffer head: drop the first |N| ms of the written stream so every audio event lands
|N| ms EARLIER relative to video. It only makes sense at the head of the stream, so it is armed once
at StartLive (positive stays live-reactive via the delay line). Rule: record the signed semantics
together — positive = delay (ahead), negative = eat-head advance (behind) — and lock the control
while live (IsEditMode) since a mid-stream advance flip is meaningless. Test recipe:
StartLive_NegativeOffset_AdvancesAudio_ByDroppingTheStreamHead (emit 6×0.9 into an 8-tick budget,
long 0.2 bed, wire must show only 0.2).
Composition video capture of WebView2: the CoreMessaging DQ recipe + the build gotchas
(2026-09-14, the ~20→60fps web-capture redesign — record-once so the next GPU-path work never
re-derives it. Everything that follows was already settled by OBS/Flutter webview_windows; the
cost this session was the 19041 projection's gaps.)
- No
DispatcherQueueController.CreateOnCurrentThread()on the 10.0.19041 projection (CS0117 — onlyCreateOnDedicatedThread+FromAbi(IntPtr)exist). P/InvokecoreMessaging.dll!CreateDispatcherQueueControllerwith a sequentialDispatcherQueueOptions { DwSize, ThreadType=2 (DQTYPE_THREAD_CURRENT), ApartmentType=2 (DQTAT_COM_STA) }, wrap viaDispatcherQueueController.FromAbi(ptr)(mirror ofCaptureInterop), THENnew Compositor(). One controller + one compositor per UI thread, created once. CoreWebView2CompositionControllerneeds a real parent HWND — there is no window in the MainWindow ctor, soInitWebView2()moved from the ctor toMainWindow_Loaded(source registration is guarded byContainsKey, so the layout-load path is safe).- Capture the ROOT visual, not the control.
GraphicsItemfor a composition controller comes fromGraphicsCaptureItem.CreateFromVisual(root)whererootis your own 1920×1080ContainerVisual(IsVisible=true) holding the controller'sRootVisualTargetas aRelativeSizeAdjustment=1,1child — exactly whatwebview_windows'sgraphics_context.ccdoes (CreateGraphicsCaptureItemFromVisualon the rootsurface_visual). Hide nothing, move nothing, poll nothing — the capture is frame-driven at the renderer's pace. - Straight alpha, never premultiplied. Read back with
BitmapAlphaMode.Straight; the compositor's blend is straight-alpha, soPremultipliedreadback wrecks corner anti-aliasing. - Namespace landmines that each cost a build cycle in
Services/:Compositorresolves to the repo's OWNytLive.Services.Compositornamespace — fully-qualifyWindows.UI.Composition.Compositor;CoreWebView2CompositionController.Close()is the disposal call (noDispose());Coloris ambiguous (System.DrawingvsSystem.Windows.Media) — qualifySystem.Drawing.Color.Transparent. - Testing a compositor-built bitmap without a runtime: the internal seam ctor
(Dispatcher, Func<string, IScreenCaptureSource>?)+ aFakeWebSource+ a real background-STADispatcherPump(borrowed from ScreenCaptureManagerTests) drives the whole session path hermetic. When byte-comparing a crop, slice the source with stride gaps (helperCropBytes) — a contiguous range silently spans rows.
Recipe verified on green suite + 0-warning build; the composition path itself still needs a device take (open item, HANDOFF).
A pipe-read test that samples one tick is a timing flake by construction
(2026-09-14, re-discovered) Mix_HonorsProviderGains_AndGameMute_KillsTheLoopback fails
sporadically in ISOLATION (3/3 on the clean tree) but passes in the full suite: it reads exactly one
tick's worth after muting, and if the pipe still carries pre-mute buffered bytes the read spikes at
0.4 instead of silence. Any caller of ReadFullyAsync on a live pipe must either close/drain the
pipe first or assert on a multi-tick window. Do not blame a sync change for this — verify the grain
against the clean tree in the same mode before touching the mixer.
Slice-18 follow-up (2026-09-15) — C4 composite cache: test fakes must mirror production's STABLE scene and per-tick purity
The C4 blit-on-change cache keys on a hash of the resolved inputs, INCLUDING per-element reference
identity (RuntimeHelpers.GetHashCode(element)). Two test-setup habits silently broke/starved it:
NewPump's default scene is fresh per tick (() => BackgroundScene()— fine for pacing tests) — churned the element refs, so the signature NEVER matched and the cache looked broken (301 "renders" instead of 1). Production hands a STABLEStagedScene. Rule: any FramePump test that asserts per-frame content or cache behavior must passscene: () => scenewith one Scene instance; the fresh-scene default is only for pacing/diagnostic tests.- A call-count resolver flip (
flip++ % 2) double-advances under a cache-aware pump — the tick resolves TWICE (signature pass + compositor pass). Key per-tick content on the burnedOutputIndex(stable until submit, which happens AFTER render) or onencoder.Frames.Countinstead. A stable-per-tick token, not a call counter, is the deterministic alternation. - New frames that must invalidate the cache need BOTH a fresh array AND a fresh Epoch (array
identity alone is unchanged for a mutated-in-place array;
Epochis the monotonic generation marker the compositor's paste key and the C4 signature share).