Files
LlamaCasty/MyMistakes.md
T
gramps 724af1499b fix: audio silence (idempotent sync-delay configure), webcam gray block, truncated videos
- AudioSyncDelay.Configure reallocated/zeroed its buffer every ~10ms tick
  (AudioMixer re-reads the UI setting each mix), so any non-zero sync offset
  erased the just-written audio -> total silence. Now early-returns when the
  delay samples are unchanged. Regression test proven both ways.
- SceneCompositor.BlitContentRaw defaulted cbW/cbH=0 when CropBounds is null
  (regression from ed9d7c1) -> webcam blit to an empty rect = gray block.
  Default to src.Width/Height. Regression test proven both ways.
- Truncated recordings: rawvideo mux stamps frames at declared 60fps by
  arrival; a scene whose first layer is dynamic (hidden elements still count)
  kills the bake cache -> full render ~35ms -> ~27fps submitted -> halved
  file length. Static scenes bake once (246ms cold, then <1ms) -> 60fps,
  full-length (probe + 13:53 take, 301/300 per 5s, 12.46s file from 12.3s
  wall). FramePump.ProbeRender names the hot render path on slow frames.
2026-09-12 13:56:55 -07:00

32 KiB
Raw Blame History

MyMistakes.md

Two jobs, distinguished by heading:

  1. Per-task failure log — updated before every commit touching that task: current iteration + why the last one failed. On task complete, committed AND pushed → truncate to this stub. A new task does NOT seed this file until its first failure.
  2. Recipes registry (DERIVED-SOLUTION RULE, see AGENTS.md 🔬) — the durable home for one-off derived solutions, recipes, and how-tos. The moment you work out a reusable solution, write it here in the same session. GREP THIS FILE FIRST when you hit a "I've done this before but have to figure it out again" wall. Recipe entries stay permanently (they are NOT truncated on task completion) — only the failure log truncates.

🔬 Recipes registry

SPIN GUARD → RESOLVED — web overlay transparency + bounding box (RECIPE)

THE ONE ROOT CAUSE THAT EXPLAINS EVERY FAILED TAKE: WebView2's CapturePreviewAsync produces an OPAQUE PNG. From the WebView2 spec (sender: MicrosoftEdge/WebView2Feedback specs/BackgroundColor.md): "WebView will always honor a webpage's background content." DefaultBackgroundColor = Transparent only shows through pages with NO background style — the widget's own CSS paints html/body opaque. Every take below built on the false premise "the capture has transparent margins, alpha=0"; it never did. FindContentBounds then had no alpha-0 margins to find → wrong crop → black bounding box. The compositor blend saw alpha=255 → black over webcam = "transparency broken". Same bug, three symptoms.

THE FIX (the OBS way, applied 2026-09-08): inject the transparency BEFORE the page parses using CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync — documented to run "before the HTML document has been parsed and before any other script included by the HTML document is run" (learn.microsoft.com/dotnet/api/microsoft.web.webview2.core.corewebview2.addscripttoexecuteondocumentcreatedasync). The old ExecuteScriptAsync on NavigationStarting/NavigationCompleted ran AFTER page scripts/CSS → widget page repainted background opaque → lost the fight. OBS browser sources do the same via a pre-parse user.css. Kept the nav handlers as a post-load re-assertion.

Take timeline (the honest record):

  • bccdb48 (Aug 28): added FindContentBounds crop — correct idea (tight bbox, no dead space), but the capture was OPAQUE so the bbox math was built on nothing.
  • 5348b5c (Aug 28, 2 min later): reverted to full-frame no-crop — looked "good" for a full-bleed widget, but floating widgets regained dead space ("ghost boundary").
  • take-21 (e002847) full canvas → shrunken/offset widget (UnifomToFill of whole canvas into a small element = lost resizing).
  • take-22 (ed9d7c1) crop width used as canvas stride for buffer indexing → garbage.
  • take-23 (f6802c7) stride fixed with src.Width; PasteKey lacked CropBounds → stale raster cache → stale crop.
  • take-24 (081e4c1) CropBounds in PasteKey — STILL BROKEN because the SOURCE ALPHA WAS NEVER REAL.
  • take-25 (ME, this session — the user's "RE-INTRODUCING THE BOUNDING-BOX PROBLEM"): I removed FindContentBounds + the crop path entirely, betting full-canvas UniformToFill was the answer. It WASN'T — the source is opaque-black, so the element rendered as a SOLID BLACK BOX (the user's screenshot: "a black box in the lower right corner"). Killed the crop → dead space returned AND black box. The compositor math was ALREADY correct; gutting it was vandalism in response to a source-level bug.

Rules, self-inflicted:

  1. Instrument FIRST. This session added: first-capture PNG dump of the raw WebView2 PNG (to %TEMP%\ytLive-web-<id>.png) + alpha min/max/mean/%zero + FindContentBounds result logged to startup.log once per session. That's the diff between a five-take loop and a five-minute diagnosis.
  2. When the same symptom loops across takes, the PREMISE is wrong, not the code — the capture being transparent was the load-bearing premise and it was never verified.
  3. Do not delete code paths that fix one axis (crop=bbox) while debugging another (source alpha). Revert scope creep; keep layer contributions separable.

2026-09-10 follow-up → NOW VERIFIED and hardened. The 12:54 and 13:50 sessions showed the REAL widget document capturing TRANSPARENT (widget dumps w1..5 for 8d7234ec: alpha 100% zero, content 12×3; whole-file decodes of %TEMP%\ytLive-web-*.png alpha max = 0) while the RECORDING kept showing a black, opaque box over the element rect (crisp edges at the exact rect 1231,679 705×396 — NOT the desktop showing through). Concluson: branch (a) — the OLD inline element.style.background='transparent' injection only wins while the page has nothing to paint; once the widget connects and paints its own container background-COLOR, the capture goes opaque again → black box. The OBS-validated answer (valid for arbitrary pages for a decade) is an injected pre-parse <style> with !important beating every page rule: https://obsproject.com/forum/threads/translucent-transparent-browser-source.59549/ (body { background-color: rgba(0,0,0,0) !important }) + the div-level variant for stubborn widgets (woahtech.com OBS custom-CSS guide). Applied 2026-09-10 (b4bba4b): injection appends a style element wiping background-color:transparent!important on html,body,html *. Second half (2026-09-10, 15:34 take): a background-COLOR wipe is NOT enough. The 15:34 interior ASCII shows a black void with bright content strips at top/bottom + a right-edge bar — the widget's full-canvas CSS background-image (gradient/backdrop) painted after connect. The OBS fix for that is background: none !important / background-image:none!important (obsproject/obs-studio#6659 — "set the CSS for html and body to background: none !important"). Wipe both moving forward; overlay art is <img>/DOM and survives. Also the dumps only covered +1s (blank-transparent); the widget's backdrop paint arrives later, so the mid-recording dump set is now re-armed ~30s in. Verify take: element rect shows scene bg behind the widget art (animation, no black void) → loop closes. Self-inflicted again: the agent re-derived the whole transparency story (recording pixel archaeology, compositor blend re-verification) instead of reading this entry — the instrument said transparent because the capture had NOT been repainted yet. Do NOT re-derive this story a third time.

2026-09-12 → RESOLVED — the box lived in the paste-cache RASTER, not the page. The 15:51 facts (alpha max=255, mean ~19, zero 57%, white rounded panel, real transparent margins) + preview-correct + recording-box meant the capture transparency was REAL all along; the fracture sat in SceneCompositor.BlitContentRaw (SceneCompositor.cs:419-550) — the sampler that builds the element-space paste-cache raster, the ONE path the take loop never read in full. Its partial-alpha branch applied the OPAQUE-dst blend onto a TRANSPARENT raster base: dst = (src*sa + dst*inv)/255 with dst black → color PREMULTIPLIED by sa, then dst[+3] = 255 — a 50%-alpha widget pixel became darkened color + FULL alpha. At paste time BlendRowOpaque saw alpha 255 → straight copy → the scene behind was overwritten by darkened ink. Fully-transparent margins (alpha 0, continue) and fully-opaque content (alpha 255 branch) survived — which is why every take showed a box while the dumps and the preview (raw WriteableBitmap, unaffected by the compositor) stayed correct, and why the CSS-wipe fixes (b4bba4b / 0f72c53) were red herrings: they treated the page as the villain, but the capture was transparent from the start. FIX: BlitContentRaw gained transparentDst=false; the raster call site passes true and writes STRAIGHT color + straight alpha so the paste rows (BlendRowOpaque/BlendRowWeighted) do the real source-over onto the opaque master. The master paths are untouched (bit-identical). ONE test PasteCache_SemiTransparentLayer_RevealsBackdrop_NotOpaqueInk fails on the old code with exactly the bug encoded: 50%-blue over red reads (0,0,128) instead of (127,0,128) — backdrop never shows through. Next: verify take (element rect shows the backdrop behind the widget art), then push gate #1 clears.

2026-09-12 → COLLATERAL — the SAME transparency saga broke the webcam next. Verify-take of the transparency fix showed a perfect flat gray rectangle where the webcam should be (preview fine, recording flat — stdev <1 across the whole element, despite the raw camera frame handed to the compositor sampling min=0/max=255 one line earlier in the pipeline — proven by instrumenting BOTH ends before touching any code, not guessing). Root cause: ed9d7c1 ("CropBounds metadata + Fill-style scaling... widget fills element rect" — take-22 of the SAME transparency saga above) added cbX/cbY/cbW/cbH to BlitContentRaw's general sampler but only assigned them inside the src.CropBounds is {} cb branch; every CropBounds-LESS source (webcam, images — anything but the web widget) fell through with cbW=cbH=0. Since sxCrop = sxNorm * cbW and syCrop = syNorm * cbH, both were always 0, so sxCanvas/syCanvas collapsed to cbX/cbY = (0,0) for every destination pixel — an entire scaled webcam sampled ONE source corner pixel. The web widget (the only CropBounds-bearing source) was never affected, which is exactly why the transparency fix's own test suite stayed green while this broke. FIX: default cbW = src.Width, cbH = src.Height (cbX/cbY = 0) before the branch, only overridden when CropBounds is actually present. ONE test BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin (two-color split source, scaled non-1:1 through the paste cache) fails on the old code — right-half pixel reads the left half's color — and passes on the fix; proven both ways with a stash/build/ revert cycle, not by inspection alone. Lesson: when a shared low-level sampler gains a new optional code path (CropBounds), audit EVERY variable the new branch introduces for a safe default in the branch it did NOT touch — an uninitialized-to-zero "crop region" silently means "sample only pixel (0,0)", not "no crop." grep for this shape (var x = 0; followed by an if (cond) x = ... with no else) whenever a conditional metadata field is added to pixel math.

Shrink / re-encode an image for the README (screenshots → small hero image)

Worked out 2026-08-29 (the recipe was NEVER recorded the first time it was done, so it had to be re-derived from scratch — that's the incident this entry exists to end).

Approach: a throwaway Windows-dotnet console app uses WPF's imaging stack (System.Windows.Media.Imaging) — same framework the app runs on, zero NuGet packages, high-quality downscale via TransformedBitmap. Screenshots compress far smaller as JPEG than PNG (PNG 1400px = ~1.2MB; JPEG q82 1400px = ~188KB).

Recipe (run via the Windows dotnet host from WSL):

  1. Create imgresize.csproj targeting net8.0-windows with <UseWPF>true</UseWPF> (SDK controller). Put it in a Windows-visible temp path, e.g. C:\Users\gramp\AppData\Local\Temp\imgresize — NOT /tmp (Windows dotnet can't reach a Linux-only path reliably).
  2. Program.cs: load BitmapImage (CacheOption=OnLoad → Freeze()), downscale with TransformedBitmap(src, new ScaleTransform(scale, scale)) to max width (1400 for the README hero), encode with JpegBitmapEncoder { QualityLevel = 82 }, save.
  3. Run:
"/mnt/c/Program Files/dotnet/dotnet.exe" run -c Release --project .
   -- "C:\Users\gramp\Downloads\Screenshot 2026-08-29 075626.png"
      "C:\Users\gramp\Documents\Code\projects\ytLive\docs\ytLlive-preview.jpg" 1400
  1. Point README.md at the .jpg (not .png).

Result: 3.2MB screenshot → 1400×794 → 188KB ytLlive-preview.jpg in docs/.


Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget

Derived 2026-09-03 (take-3 diagnosis — the stats seam from 97ffc42 named the stage in one line: 17/300 frames per 5s, avg render 258.1ms, avg submit 1.5ms). Both halves were solved by OBS/libyuv long ago; do not re-derive:

  1. Pacing is a DEADLINE, never a post-render sleep. sleep(interval) after each frame makes the period render + submit + interval — the producer can hit ≤ half the declared rate even with a free render. OBS's video_thread (libobs/media-io/video-io.c) advances an absolute nextTick += intervalTicks and sleeps only the remainder. NEVER "rebase" the deadline to wall-now when it blows — the take-3 note I originally wrote here ("skip … the missed ticks (rebase)") was the PROVEN-wrong advice: resetting nextTick erases the missed slots, so a 43fps reality was authored into a 60fps container and every recording played ~1.4x fast (take 12/13, 2026-09-10). The rebase even kept the stats looking honest (render+wait == period, no loss) because a rebased frame is never "late". OBS keeps the counter MOVING and outputs ONE frame per interval tick — a late render shows as REPEATED footage (judder; duration stays correct), never a skip and never a wall-now reset. Critical with rawvideo: pts is stamped by ARRIVAL, so the muxed duration is pure frame count ÷ fps — the only way to make duration == wall time under ANY load is exactly one frame per interval slot. Corollary: never add a second pacer. ffmpeg -re on the rawvideo input throttles the pipe independently and fights the pump (its "Resumed reading … after a lag" grew 0.79s→4.82s in the same take). Removed; the pump IS the pacer.

  2. A 1080p frame is ~2.07M pixels — the hot path must be row-simple. Per-pixel Math.Round + float source-over in managed code costs ~100ns/px = the whole 258ms. libyuv's pattern (https://chromium.googlesource.com/libyuv/libyuv/): branch per pixel on source alpha (opaque → 4-byte copy, transparent → skip), integer fixed-point blend (s*a + d*(255-a) + 127)/255 otherwise; and ALWAYS clip the loop to the intersection rect (our social-bar overlay scanned all 2M dst px for a 64px strip). A full-cover 1:1 blit also obsoletes the opaque-black pre-fill — skip dead writes.

  3. On Windows, Task.Delay is a 15.6ms QUANTUM, not a timer. Any request under one system-clock tick sleeps a full tick (documented — learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay: "approximately 15 milliseconds on Windows systems"). A deadline pacer built on Task.Delay caps the producer at ~40fps-ish EVEN IF render is instant — take 9 proved the signature: work fell 26.5→22.4ms but the period sat at ~37ms (≈ one padded wait/frame), so two real optimizations read as "zero change". Frame-accurate loops (OBS/Chromium/game-loop canon — stackoverflow.com/questions/5441464) do: timeBeginPeriod(1) for the session (paired with timeEndPeriod), sleep only the BULK of the remainder, SPIN the last ~2ms across the deadline. Diagnostic before touching the compositor again: period ≈ work + 15.6 → the SLEEP is the bug, not the work.

    Related (take 14, 2026-09-04): recycled ring buffers are a race you must SIZE, not just own. Deepening shared frames to kill GC churn (a fresh 8.3MB/tick array) hands out REUSED memory — the ring's depth × source period must EXCEED the worst consumer hold (compositor read + lagged UI preview copy), not just "a few frames". 4 slots at 144Hz capture laps in ~27ms vs a ≤50ms read: half a new screen frame flashed over an old one in the recording ("bits flashing over other bits"). Depth 8 everywhere (screen/camera/web output rings); the paste-cache Epoch still guards identity.

  4. Hermetic pacing test: inject the delay seam to RECORD the requested TimeSpan and genuinely await it (Task.Delay(d, ct)) — a fake that returns Task.CompletedTask synchronously makes the whole pump loop run on StartAsync's sync continuation and hang the test run (hit this 2026-09-03; the existing fakes all yield for exactly this reason). Assert the REQUESTED wait (< interval with a ≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.

  5. Expensive content: raster on change, never on read (take 5, 2026-09-04). A source that updates once a minute (chat text!) must not full-rasterize (FormattedText + RenderTargetBitmap + CopyPixels ≈ 15-25ms) every compositor tick. OBS text sources re-render on property/message change; the per-tick pass blits the cache. Implement as: content version (collection-changed counter) + config key (size/appearance) → cached immutable VideoFrame returned by identity. Gotcha: buffers that SURVIVE sessions (the chat log) silently arm the per-tick cost even in flows that never touch the feature (signed-out record-only takes paid chat rendering!).

  6. An async loop started from a UI handler runs ON THE UI THREAD until you take it off. await continuations re-capture the current SynchronizationContext — the frame pump was started from a WPF command handler, so the "WPF-free, hermetic" compositor rendered and read capture state ON THE DISPATCHER, serialized behind the live preview itself, for the whole starvation saga. The wait stat caught it only when the numbers became self-contradictory (render 22 + wait 10 > any rebasing deadline — a blown deadline cannot sleep). OBS runs obs_graphics_thread/video_thread as dedicated threads for exactly this reason. Pattern: _task = Task.Run(() => Loop()) (null context inside), then audit EVERY object the loop touches for UI affinity (RenderTargetBitmap / DrawingVisual / WriteableBitmap: marshal the work or the rare miss; plain locked byte[] lookups: fine) and pin it with a context test (Pump_Produces_OffTheStartingContext, inline-pumping SynchronizationContext that the old code failed by construction). Cost: takes 3–10.

  7. Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8). The chat raster fix was REAL but the composer blamed it for the residual slowness it did not own; two takes burned before the render/resolve split showed resolve ≈ 0 and pointed at the compositor pasting static layers per tick (BlitCachedLayer finished the job OBS-style). Before shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.

  8. An encoder refed from a real-time loop must ENQUEUE, never pipe-write in the loop (2026-09-10, take 15/16 — the "1...23...4...56..." smeared ticker). After slice 9 the recording played at the right DURATION but the content still hiccuped — and the aggregates (301/300, uniform file PTS, 15.6s wall vs 15.74s file) could NOT see it. The cause finally measured in FfmpegEncoder.SubmitFrameAsync: the WriteAsync(8.3MB)+FlushAsync to ffmpeg's stdin BLOCKS whenever the encoder lags the pipe, and the slice-9 burst while (now>=nextTick) then re-wrote that SAME stale composite for every slot that ticked past — frozen content runs. OBS's answering machinery is the encoder queue (libobs/obs-encoder.c): the encoder thread NEVER couples back into the video thread; overflow = dropped data, NEVER a frozen producer. Fixed as: bounded Channel<byte[]> (cap 120) + a dedicated drain task owning stdin, SubmitFrameAsync = copy-to-pool-array + TryWrite (drop-newest + count when full), pump emits ONE fresh composite per iteration (no burst re-write), stop flushes the queue then EOF. Rule: verify with a clock-independent judge. The WSL ticker that "proved" slice 9 has its own Host-timer jitter under Windows load — so this slice burns a dot-matrix _outputIndex into the bottom-right of every composite; decoding the recording reads the honest sequence (+1/frame, jumps = counted drops) with no external clock involved. Take 17 must read +1/frame from that strip. (A whole-frame duplicate scan was tried and is UNRELIABLE here: the scene is always animating — session elapsed timer + REC pulse — so no two frames are ever byte-identical.)

    Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again: render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy + integer bilinear) followed the same libyuv research (https://chromium.googlesource.com/libyuv/libyuv/ — row.cc/scale.cc keep both interpolation stages in ONE fixed-point scale; rounding constant only at the end). My first Bilinear shifted stage 1 back to 8-bit AND shifted the final result >>16 — double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew NOTHING (green webcam silently vanished from output; the pixel probes caught what the eye in a 2x time-lapse would not). Rule: multi-stage fixed point shifts only at the end; verify against a uniform-255 sample before believing it. Second trap: a stale-byte sentinel test whose source pattern can generate the sentinel value itself (0xAB was a legitimate x+y pixel) — pick the sentinel coprime/out-of-range to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder that HOLDS submitted frames now must snapshot them (Clone) once the producer legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.


A web widget captured at 10Hz inside a 60fps recording plays at ~1/6 speed

Derived 2026-09-10 (the "widget animation too slow" report). The recording is 60fps and the WebView2 capture loop was a blind 100ms DispatcherTimer = 10Hz — each captured web frame gets repeated ~6× in the file, so whatever the page animates at, the OUTPUT is capped at the capture cadence, not the page's. Two stacked throttles, both real:

  1. The capture rate is the hard ceiling. web-layer motion in the recording can never beat CapturePreviewAsync frequency. But you can't just raise the timer: full-HD PNG capture costs ~10-30ms (WebView2Feedback#20: "CapturePreviewAsync … produces PNG/JPG and is very slow"), so concurrent captures stack CPU AND can publish stale-after-fresh. The mandatory shape is a latest-wins drop: at most one capture in flight per session; a tick during the in-flight window is DROPPED, never queued; effective cadence = max(interval, capture duration). Do this before anyone touches the interval constant.
  2. Chromium throttles hidden pages. An off-screen WebView2 (we place it at (-5000,-5000)) is a hidden page the moment the host window is unfocused or covered: requestAnimationFrame parks and JS timers clamp to ~1s (WebView2Feedback#1172 — background-throttled rAF; #3070 — a WebView2 with Visibility.Collapsed slows timers to 1s; Chrome-88 blog — heavy timer throttling of hidden tabs). There is NO supported per-page opt-out (#5250 still open). The embedder answer is browser args on the CoreWebView2EnvironmentOptions created BEFORE EnsureCoreWebView2Async: --disable-backgrounding-occluded-windows --disable-renderer-backgrounding --disable-features=CalculateNativeWinOcclusion (the Electron/Streamlabs-class fix for occluded-window animation throttling). Share ONE environment across sessions (one browser process).
  3. Measure before picking the cadence. Log the first ~30 captures' elapsed ms on the first recording run; PNG encode + WPF decode of 1920×1080 is the per-capture cost that decides whether ~30Hz is affordable or it must drop to ~20Hz. The FramePump drops frames (never time-lapses) if UI-thread GC churn starves it, so cost shows up as dropped-frame stats — read them.

Splitting a large file into partials — NEVER awk … > SRC while awking SRC

(2026-08-31, Commit D) Tried to split SocialsDialogViewModel.cs in one line: { awk '…' SRC; echo ""; awk '…2…' SRC; } > SRC. The first write truncated SRC to 1 line, so the second awk read the already-truncated file → the whole source was lost (1 line left). Recovered with git checkout -- SRC, then redid it, but the same bug could have meant making it up from scratch.

Rule: when a cut needs N blocks from one source into N files, never write a block back onto the source that the awks still read. Instead:

  1. Read the source once at the start into temp files (mktemp -d, one file per block), with a $D variable you carry forward.
  2. Verify block sizes (wc -l) and brace balance (python3 -c counting {/}) before touching any real file.
  3. Then assemble each new file from cat D/block … — never truncating the source until every read is done.

git status can't save you here if you don't notice until the file is gone — git checkout -- <path> from the last commit is the recovery. Cheap insurance: restore-then-retry, do it atomically from temp files the first time.


Verifying an ffmpeg decode contract from WSL (no real CLR needed)

(2026-08-31, TASK 21) When a change depends on ffmpeg producing output with an exact frame-size contract (rawvideo W×H×4 BGRA), you can prove the command + frame accounting here without any .NET process:

  1. Fetch a static Linux ffmpeg into /tmp/opencode (no sudo needed): curl -sLO https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz && tar -xf …
  2. Generate a tiny known clip: ffmpeg -f lavfi -i "testsrc2=duration=1:size=640x360:rate=30" -pix_fmt yuv420p clip.mp4
  3. Decode with exactly the app's args: -f rawvideo -pix_fmt bgra -vf scale=640:360 -an
  4. Assert total_bytes % (W*H*4) == 0 (python3) → exact integer frames, no pad.

Why not a dotnet-spawned ffmpeg here: the only CLR on this box (/home/gramps/bin/dotnet) is a Windows-bound shim — Process.Start resolves paths to \\wsl.localhost\Debian\… and throws "not a valid application for this OS" when handed a Linux ELF ffmpeg. So never plan to have dotnet exec a Linux ffmpeg here; verify the contract with shell/python instead, and leave the CLR→real-ffmpeg run to the native Windows suite.

WPF hit-test truth in tests: UIElement.InputHitTest, NOT VisualTreeHelper.HitTest

(2026-09-01, the RoundClip "known failure" post-mortem — a failure the map carried as "not a regression" for weeks without ever recording WHY.)

The trap: VisualTreeHelper.HitTest(window, pt) returned the window's WebViewHostPanel overlay (IsHitTestVisible="False", Opacity=0, ZERO children) for EVERY point in the window — so a "corner is grabbable" assertion could never pass, and it looked like a real interaction bug. The actual input pipeline (UIElement.InputHitTest, what Mouse routing uses) correctly returned the element's Grid at elem-center/corner-in/ corner-exact and fell through to CanvasGrid just past the corner. The product was fine; the TEST was probing an API that doesn't model input semantics.

Rule: any test asserting "where does a click land" uses window.InputHitTest(pt) + IsDescendantOf — never VisualTreeHelper.HitTest.

Diagnosis recipe (how the lie was caught in ~3 probe cycles, no guessing): add a TEMP probe [Fact] in the RealApp collection that hit-tests a spread of points (elem-center / corner-in / corner-exact / corner-out / bg-center) and Assert.Fails with a composed dump: per-point VTH hit + InputHitTest hit + ancestor chain (GetParent walk with #Name) + panel properties (IsHitTestVisible/Opacity/children/actual size) + TranslatePoint origins. Run the class alone, read the message, delete the probe.

Two sibling facts learned the same session (record-once):

  1. A UserControl owns its own XAML namescope — after extracting a region out of a window, window.FindName("InnerPart") returns null; resolve the UserControl by its window-level name, then pane.FindName("InnerPart"). And window-scope STYLES are invisible to a UserControl's StaticResource at parse time — move such styles to Themes/Controls.xaml (the app-scope rule exists for this).
  2. Per-class dotnet.exe vstest from WSL DOES execute the RealApp/MainWindow tests fine (they passed natively 2026-09-01) — only the FULL suite hangs (WASAPI startup). And the test process shares %APPDATA%\ytLlive\startup.log with the real app: lines like camera 'test-camera' failed are test noise, not DB state — to check pollution, query the DB directly (python3 sqlite3, SELECT DeviceId FROM Webcam), not the log.

A "known failure" label without a recorded cause = a bug on life support

(2026-09-01, the audio triple-take) One line — _delayedMix (nullable, added by TASK 22, never initialized) dereferenced as delayed.Length — produced THREE symptoms that lived in the map as two separate "pre-existing, do-not-chase" entries: (a) Mix_HonorsProviderGains… "known failure", (b) AudioPipelineTests hangs when run at all (a test reading a named pipe with no writer blocks — hung test ≠ flaky test, it's a starved producer), (c) startup.log flooded "Audio live loop error" every 10ms (the mixer loop caught and logged only ex.Message — stack thrown away).

Rules derived:

  1. NEVER label a test "known/pre-existing" without writing WHY (exception type + first app frame). An unexplained known-failure is deferred archaeology that hardens into fog.
  2. A test that waits on IPC + a producer whose output vanished are usually ONE bug — look for the producer before blaming the test.
  3. Catch-and-log-swallow of ex.Message hides root causes; log with stack (AppLog.Write(ex, ...)) and throttle (5s) instead of dropping or flooding. Bonus: the "failing" test encoded the MAP's contract (loopbackGain = GameAudioVolume); the code had drifted to unity on a disproven premise (loopback capture does NOT follow endpoint volume — creator's 20%-volume/pegged-meter observation killed it). The test was right all along — failing tests may be the last honest witnesses; interrogate, don't pardon.

Feature provenance: record WHO asked and WHY, in the task entry itself

(2026-08-31/09-01, the SYNC slider scare) TASK 22's lip-sync slider surfaced on the preview rail and the creator's reaction was "totally don't remember ordering that" — because the queue entry recorded WHAT shipped (a slider, 0-500ms, a converter class) but not WHO asked (the creator, explicitly, for OBS's delay-filter fix built natively). Eight days later his own request read like AI drift and nearly got deleted. Rule: the moment a creator-driven feature is queued or shipped, its entry carries a one-line provenance — who asked, what triggered it ("creator: OBS delay-filter lip-sync fix, native"). Features without attribution become roadmap orphans that get punted, removed, or re-litigated. Same disease as an unexplained "known failure" label — a fact recorded without its reason is a future argument.

An un-attributed build invalidated three takes of a perf saga — stamp the binary

(2026-09-04, takes 4–6) After each render-perf fix the creator "exed the code" and re-recorded, but the exe timestamp ≠ binary contents (incremental builds reuse whatever compiles clean; a source edit with no rebuild serves the OLD exe). Take 6 measured render WORSE than take 5 (35-41ms) and there was no honest way to tell "the chat cache fix doesn't work" from "the fix was never running" — three hours of diagnosis on an unattributable sample. Rule: if takes measure the app, EVERY build carries an id and EVERY log line traces to it — GenerateBuildStamp (csproj) writes a fresh GUID per compile (deliberately defeating incremental lies), the wordmark shows it as a superscript, startup.log records Build <id> (compiled <time>). A perf claim without build attribution is a guess; ask for the stamp BEFORE theorizing. (Also this session: a sentinel-byte test where the source pattern could GENERATE the sentinel, and a fixed-point bilinear that shifted BOTH stages and silently drew nothing — see the rawvideo recipe.)