Files
gramps e7cded6879 docs: pricing ruled — $10/mo + $400 lifetime, and the 15% fee is verified for subscriptions
Reconcile the pending doc edits with the creator's actual ruling, and close the
revenue-share question that was blocking the monthly price.

Verified from the App Developer Agreement v8.10 PDF itself (downloaded and
text-extracted, not a search excerpt). Section 6(b) has only three tiers:
6(b)(i) 15% for Apps and their In-App Products *not listed in* 6(b)(iii);
6(b)(ii) 12% Games-only; 6(b)(iii) 30% for Xbox console apps/games, Xbox
non-subscription IAP, and Windows 8/Phone 8. LlamaCasty is a Windows PC App,
so 6(b)(i) governs BOTH tiers -- there is no subscription surcharge. The
agreement's own changelog (v8.0, Oct 26 2017) states it outright: "implement
the 85/15 revenue share for non-Game subscriptions."

  => $10/mo nets $8.50 (~$102/yr); $400 lifetime nets $340.

Also confirmed: the 15% applies after VAT/GST (Net Receipts definition),
payouts are monthly above a $50 threshold, and no better small-indie rate
exists in the standard terms.

Creator ruling recorded: $10/mo subscription or $400 lifetime, no annual, no
.99. This supersedes the 2026-09-21 one-time $29 -> $49 model.

Two obligations ADA 6(h) attaches to the recurring tier, recorded because they
change its risk profile and are the reason lifetime is the hedge: we must
fulfil the subscription for the entire period as marketed (on breach Microsoft
may refund the full amount plus taxes in its sole discretion), and raising the
price disables auto-renew -- so $10 is effectively locked for the product's life.

Stale facts fixed in the same change rather than appended:
- research-store-certification.md: IARC was 11.11.1/11.11.2 in one table and
  10.11.1 in another; corrected to 10.11.x.
- research-store-certification.md: policy 10.8.1/10.8.2 still asserted "Polar
  explicitly permitted" and "tick the third-party purchase box". Void now that
  Store IAP is the route and Polar is being torn down.
- task-48: the working YouTube demo account (10.3.1) was accidentally dropped
  from the certification list during the Store-IAP edit. Restored -- a reviewer
  cannot use the app without one.
- MyMistakes.md: the verified-fact-vs-decided-outcome lesson now closes its
  loop (the creator did rule the way the research pointed), and records the
  over-correction that followed.

Docs-only. No code, no build, no tests.
2026-09-27 15:02:25 -07:00

1180 lines
83 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# MyMistakes.md
> **Two jobs**, distinguished by heading:
>
> 1. **Per-task failure log** — updated before every commit touching that task:
> current iteration + why the last one failed. On task complete, committed AND
> pushed → truncate to this stub. A new task does NOT seed this file until its
> first failure.
> 2. **Recipes registry** (DERIVED-SOLUTION RULE, see `AGENTS.md` 🔬) — the durable
> home for one-off derived solutions, recipes, and how-tos. The moment you work
> out a reusable solution, write it here **in the same session**. GREP THIS FILE
> FIRST when you hit a "I've done this before but have to figure it out again"
> wall. Recipe entries stay permanently (they are NOT truncated on task
> completion) — only the failure log truncates.
## 🔬 Recipes registry
### YOUTUBE STATISTICS ARE JSON STRINGS; TryGetInt64 THROWS ON STRINGS (RECIPE)
YouTube Data API v3 returns `statistics.subscriberCount / viewCount / videoCount` as
JSON **strings** (`"350"`), not numbers. (2026-09-23: the YPP parse bug surfaced the
moment the auditDetails 403 stopped masking it.) Fix = a tolerant read that checks
`ValueKind` FIRST, then `TryGetInt64` for `Number`, then `long.TryParse(GetString())`
for `String`. Trap within the fix: **`JsonElement.TryGetInt64` THROWS
`InvalidOperationException` on any non-Number token** ("requires an element of type
'Number'") — it is try-type-in, not try-catch. Guard on `value.ValueKind` before
calling it; a leaked 403→parse chain means your "whole request failed" symptom can
paper over a second crash that only appears once the 403 is fixed (test the full
happy path, not just the error path).
### liveChatId LIVES IN SNIPPET AND ONLY EXISTS ONCE THE BROADCAST IS LIVE (RECIPE)
YouTube Data API v3 `liveBroadcasts`. `liveChatId` is in **`snippet.liveChatId`** —
`contentDetails` has NO such property. AND it is only populated once the broadcast is
**live** (the official `GetLiveChatId.java` sample lists `broadcastStatus=active`);
a fetch right after insert (lifecycleStatus `ready`) returns nothing. (2026-09-25: this
cost a whole investigation — the TEST-tab "Chat polling couldn't start — Mock Chat Input
is disabled" report — because the code read `part=contentDetails` + ran before the frame
pump pushed RTMP.) Fix = read part=snippet + bounded retry AFTER the encoder starts
(`GetBroadcastLiveChatIdAsync(broadcastId, maxAttempts=10, delayMs=2000)`). Debug lens:
missing liveChatId ≈ "broadcast not live yet", NOT an auth failure.
### liveChat/MESSAGES.INSERT BODY REQUIRES snippet.type (RECIPE)
YouTube Data API v3 `liveChat/messages.insert` rejects the body with
`400 MISSING_REQUIRED_FIELD` (`domain: youtube.api.v3.LiveChatMessageInsertResponse.Error`)
unless the snippet declares `snippet.type` = `textMessageEvent` (or `pollEvent`) alongside
`liveChatId` and `textMessageDetails.messageText` — the official insert reference lists
`type` as a required property. (2026-09-25: the TEST-tab "YouTube rejected the message
(error 400)" report — the liveChatId fix in TASK 44 had worked and the drawer's Mock Chat
Input was issuing a real insert, but the body omitted `type`, and the TASK 41 test only
asserted `liveChatId` + `messageText` were present so it stayed green while real YouTube
rejected every send. Fix = add `type = "textMessageEvent"` to the body + assert it in the
Good Dog test.) Debug lens: MISSING_REQUIRED_FIELD ≠ auth/scope — it means the request
body shape is wrong, and the false-green test is the classic trap: assertion on the
body was about WHAT WE SEND, so YouTube's required fields must be mirrored in the test.
### CHANNELS.LIST auditDetails PART 403s THE WHOLE REQUEST WITHOUT A PARTNER SCOPE (RECIPE)
YouTube Data API v3 `channels.list` rejects the ENTIRE request with
`403 insufficientPermissions` if your `part=` list includes `auditDetails` but the
token lacks `https://www.googleapis.com/auth/youtubepartner-channel-audit` — the
docs' exact words: "A request that retrieves the auditDetails part for a channel
resource must provide an authorization token that contains the
youtubepartner-channel-audit scope". That scope is MCN content-partner tooling
(with a two-week token-revocation rule); normal-creator apps should NEVER ask for
it. 2026-09-23 real-log evidence: YPP refresh 403'd three times in a row while the
mock-fake tests stayed green ("current scopes suffice, no re-consent" was wrong).
Fixes: (1) drop `auditDetails` from `part=`; (2) log the response BODY — the bare
status code could not name `insufficientPermissions`, which is what made this
failure undiagnosable for days; (3) when a feature needs data no ordinary scope
grants, deep-link to the site instead of requesting the partner privilege.
### TRANSITION(COMPLETE) RACES AUTOSTOP: PRE-FLIGHT lifeCycleStatus (RECIPE)
A blind `liveBroadcasts.transition?broadcastStatus=complete` POST can 403
`invalidTransition` even though the stream just ended normally — 2026-09-22 logged
it on EVERY session end. Cause: the broadcast's own state moves toward complete
via `enableAutoStop`/YouTube auto-complete; the transition method's errors doc
(https://developers.google.com/youtube/v3/live/docs/liveBroadcasts/transition)
shows `invalidTransition` = "can't transition from its current status". A blind
POST just races — read the broadcast's `lifeCycleStatus` first
(`liveBroadcasts.list?part=status`) and only POST complete from `live`/`testing`;
skip silently otherwise and let `enableAutoStop` finish it. Never throw on the
close-out either way.
### YOUTUBE liveStreams LIST: healthStatus IS AN OBJECT, NOT A STRING (RECIPE)
YouTube Data API v3 `liveStreams.list` nests health under
`status.healthStatus = {status, lastUpdateTimeSeconds, configurationIssues[]}` —
the `healthStatus` element is an **OBJECT**, and `configurationIssues[]` lives
INSIDE it, not directly under `status`. Reading `healthStatus.GetString()` throws
System.Text.Json's `The requested operation requires an element of type 'String',
but the target element has type 'Object'` — the exact log line seen 2026-09-22 on
every live health poll (two test sessions). The parse must read
`healthStatus["status"]` (and nested `healthStatus["configurationIssues"]`).
Something like REST-shape drift is a good reason to grep the API reference
(https://developers.google.com/youtube/v3/live/docs/liveStreams) before writing
parsers against a "remembered" shape — our own flat-string fixture was the wrong
assumption the whole time.
### WINRT RESOURCE-ALLOCATION CALLS: WRAP PER-CALL, DEGRADE TO NEXT OPTION (RECIPE)
Creator callout (2026-09-15): WinRT/COM calls that allocate or start a resource — `InitializeAsync`,
`CreateFrameReaderAsync`, `StartAsync`, `CreateReaderAsync` & friends — are THROW-HEAVY. An unsupported
subtype/format, a device that just vanished, an access mode rejected mid-flight: these surface as
`ArgumentException`/`E_INVALIDARG` ("value does not fall within the expected range") or HRESULTs, NOT as
a returned status you can switch on. Relying on ONE outer catch to "handle failures" is not handling —
one rejected call inside a fallback ladder aborts the whole ladder and every untried option. Rule:
- Each allocation/start call inside a try/catch of its OWN, so a throw on candidate N falls through to
candidate N+1 (log each rejection with its message/status; collect them for the final error string).
- `return`/`break` on success must be reached WITHOUT passing through a `finally` that disposes the
resource you just committed (classic reader/capture dispose-after-commit bug).
- Unsubscribe + dispose the partial resource in the catch block when the subscription happened before
the throwing call.
- The outer catch stays as the LAST-RESORT net for device-level errors, not the primary one.
- Same discipline applies to the frame-consumption side: teardown races reader threads (see HANDOFF
crash follow-up) — a frame callback can't assume the pipeline is alive.
Applied in `MediaCaptureFrameSource`'s reader-subtype ladder (2026-09-15, webcam-take fix).
### "SAVE-RECORDING DIALOG IS CLIPPING CONTENT" — WHERE + HOW (RECIPE)
**WHERE:** the end-of-recording modal (creator: "when I end a recording the app throws up a
save-recording dialog with a default filename") is **`RenameRecordingDialog.xaml` at the repo ROOT**
(title "Rename Recording", `x:Class="ytLive.RenameRecordingDialog"`). It is NOT one of the
`*Dialog.xaml` files under `Controls/` — grep for the window title, not the filename, when the
user names a dialog by what it does. Only fix dialogs the user actually named; do not "while I'm
here" bump sibling dialogs.
**HOW:** the dialog is `ResizeMode="NoResize"` with a fixed `Height` attribute; content client area
≈ `Height − ~37px` of window chrome. Sizing/margin changes pushed content down in Z-order until it
clipped — first the file-name textbox (230 too short), and after the +20% bump (→276) the creator
found the Cancel/Save button row had ALSO been obscured. Fixed `Height` to 304 (+10% more),
`x:Name` the last interactive control (`SaveRecordingButton`) so a test can measure it.
**Verify (never eyeball-predict):** `RenameRecordingDialogSizingTests` (RealApp host) instantiates
the real dialog, `Show()` + `UpdateLayout()`, then asserts the BOTTOM EDGE of each named control
(`TransformToAncestor` against the content root → `TransformBounds(RenderSize).Bottom`) is ≤ the
client `ActualHeight`. Every such dialog fix ships with that assertion for every control that was
reported obscured. (Client-area math means the button row needs MORE slack than the textbox: its
bottom sits deeper because of the `Margin="0,20,0,0"` above it.)
### SPIN GUARD → RESOLVED — web overlay transparency + bounding box (RECIPE)
**THE ONE ROOT CAUSE THAT EXPLAINS EVERY FAILED TAKE:** WebView2's `CapturePreviewAsync`
produces an **OPAQUE** PNG. From the WebView2 spec (sender: MicrosoftEdge/WebView2Feedback
`specs/BackgroundColor.md`): "WebView will always honor a webpage's background content."
`DefaultBackgroundColor = Transparent` only shows through pages with NO background style —
the widget's own CSS paints html/body opaque. Every take below built on the false premise
"the capture has transparent margins, alpha=0"; it never did. `FindContentBounds` then had
no alpha-0 margins to find → wrong crop → black bounding box. The compositor blend saw
alpha=255 → black over webcam = "transparency broken". Same bug, three symptoms.
**THE FIX (the OBS way, applied 2026-09-08):** inject the transparency BEFORE the page
parses using `CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync` — documented to run
"before the HTML document has been parsed and before any other script included by the HTML
document is run" (learn.microsoft.com/dotnet/api/microsoft.web.webview2.core.corewebview2.addscripttoexecuteondocumentcreatedasync).
The old `ExecuteScriptAsync` on NavigationStarting/NavigationCompleted ran AFTER page
scripts/CSS → widget page repainted background opaque → lost the fight. OBS browser sources
do the same via a pre-parse user.css. Kept the nav handlers as a post-load re-assertion.
**Take timeline (the honest record):**
- `bccdb48` (Aug 28): added FindContentBounds crop — correct idea (tight bbox, no dead
space), but the capture was OPAQUE so the bbox math was built on nothing.
- `5348b5c` (Aug 28, 2 min later): reverted to full-frame no-crop — looked "good" for a
full-bleed widget, but floating widgets regained dead space ("ghost boundary").
- take-21 (`e002847`) full canvas → shrunken/offset widget (UnifomToFill of whole canvas
into a small element = lost resizing).
- take-22 (`ed9d7c1`) crop width used as canvas stride for buffer indexing → garbage.
- take-23 (`f6802c7`) stride fixed with `src.Width`; PasteKey lacked CropBounds → stale
raster cache → stale crop.
- take-24 (`081e4c1`) CropBounds in PasteKey — STILL BROKEN because the SOURCE ALPHA WAS
NEVER REAL.
- take-25 (**ME, this session — the user's "RE-INTRODUCING THE BOUNDING-BOX PROBLEM"**):
I removed FindContentBounds + the crop path entirely, betting full-canvas UniformToFill
was the answer. It WASN'T — the source is opaque-black, so the element rendered as a
SOLID BLACK BOX (the user's screenshot: "a black box in the lower right corner"). Killed
the crop → dead space returned AND black box. The compositor math was ALREADY correct;
gutting it was vandalism in response to a source-level bug.
**Rules, self-inflicted:**
1. Instrument FIRST. This session added: first-capture PNG dump of the raw WebView2 PNG
(to `%TEMP%\ytLive-web-<id>.png`) + alpha min/max/mean/%zero + FindContentBounds result
logged to startup.log once per session. That's the diff between a five-take loop and a
five-minute diagnosis.
2. When the same symptom loops across takes, the PREMISE is wrong, not the code — the
capture being transparent was the load-bearing premise and it was never verified.
3. Do not delete code paths that fix one axis (crop=bbox) while debugging another
(source alpha). Revert scope creep; keep layer contributions separable.
**2026-09-10 follow-up → NOW VERIFIED and hardened.** The 12:54 and 13:50 sessions
showed the REAL widget document capturing TRANSPARENT (widget dumps `w1..5` for
`8d7234ec`: alpha 100% zero, content 12×3; whole-file decodes of
`%TEMP%\ytLive-web-*.png` alpha max = 0) while the RECORDING kept showing a black,
opaque box over the element rect (crisp edges at the exact rect 1231,679 705×396 —
NOT the desktop showing through). Concluson: branch (a) — the OLD inline
`element.style.background='transparent'` injection only wins while the page has
nothing to paint; once the widget connects and paints its own container
background-COLOR, the capture goes opaque again → black box. The OBS-validated
answer (valid for arbitrary pages for a decade) is an injected pre-parse `<style>`
with `!important` beating every page rule:
https://obsproject.com/forum/threads/translucent-transparent-browser-source.59549/
(`body { background-color: rgba(0,0,0,0) !important }`) + the div-level variant for
stubborn widgets (woahtech.com OBS custom-CSS guide). Applied 2026-09-10 (b4bba4b):
injection appends a style element wiping `background-color:transparent!important`
on `html,body,html *`. **Second half (2026-09-10, 15:34 take): a background-COLOR wipe
is NOT enough.** The 15:34 interior ASCII shows a black void with bright content strips
at top/bottom + a right-edge bar — the widget's full-canvas CSS **background-image**
(gradient/backdrop) painted after connect. The OBS fix for that is
`background: none !important` / `background-image:none!important`
(obsproject/obs-studio#6659 — "set the CSS for html and body to background: none
!important"). Wipe both moving forward; overlay art is `<img>`/DOM and survives.
Also the dumps only covered +1s (blank-transparent); the widget's backdrop paint
arrives later, so the mid-recording dump set is now re-armed ~30s in. Verify take:
element rect shows scene bg behind the widget art (animation, no black void) → loop
closes. Self-inflicted again: the agent re-derived the whole transparency story
(recording pixel archaeology, compositor blend re-verification) instead of reading
this entry — the instrument said transparent because the capture had NOT been
repainted yet. Do NOT re-derive this story a third time.
**2026-09-12 → RESOLVED — the box lived in the paste-cache RASTER, not the page.**
The 15:51 facts (alpha max=255, mean ~19, zero 57%, white rounded panel, real
transparent margins) + preview-correct + recording-box meant the capture transparency was
REAL all along; the fracture sat in `SceneCompositor.BlitContentRaw`
(SceneCompositor.cs:419-550) — the sampler that builds the element-space paste-cache
raster, the ONE path the take loop never read in full. Its partial-alpha branch applied
the OPAQUE-dst blend onto a TRANSPARENT raster base: `dst = (src*sa + dst*inv)/255`
with dst black → color PREMULTIPLIED by sa, then `dst[+3] = 255` — a 50%-alpha widget
pixel became darkened color + FULL alpha. At paste time `BlendRowOpaque` saw alpha 255 →
straight copy → the scene behind was overwritten by darkened ink. Fully-transparent
margins (alpha 0, `continue`) and fully-opaque content (alpha 255 branch) survived —
which is why every take showed a box while the dumps and the preview (raw WriteableBitmap,
unaffected by the compositor) stayed correct, and why the CSS-wipe fixes (`b4bba4b` /
`0f72c53`) were red herrings: they treated the page as the villain, but the capture was
transparent from the start. **FIX:** `BlitContentRaw` gained `transparentDst=false`;
the raster call site passes `true` and writes STRAIGHT color + straight alpha so the
paste rows (`BlendRowOpaque`/`BlendRowWeighted`) do the real source-over onto the opaque
master. The master paths are untouched (bit-identical). ONE test
`PasteCache_SemiTransparentLayer_RevealsBackdrop_NotOpaqueInk` fails on the old code with
exactly the bug encoded: 50%-blue over red reads `(0,0,128)` instead of `(127,0,128)` —
backdrop never shows through. Next: verify take (element rect shows the backdrop behind
the widget art), then push gate #1 clears.
**2026-09-12 → COLLATERAL — the SAME transparency saga broke the webcam next.**
Verify-take of the transparency fix showed a perfect flat gray rectangle where the webcam
should be (preview fine, recording flat — stdev <1 across the whole element, despite the
raw camera frame handed to the compositor sampling min=0/max=255 one line earlier in the
pipeline — proven by instrumenting BOTH ends before touching any code, not guessing).
Root cause: `ed9d7c1` ("CropBounds metadata + Fill-style scaling... widget fills element
rect" — take-22 of the SAME transparency saga above) added `cbX/cbY/cbW/cbH` to
`BlitContentRaw`'s general sampler but only assigned them inside the
`src.CropBounds is {} cb` branch; every CropBounds-LESS source (webcam, images — anything
but the web widget) fell through with `cbW=cbH=0`. Since `sxCrop = sxNorm * cbW` and
`syCrop = syNorm * cbH`, both were always 0, so `sxCanvas`/`syCanvas` collapsed to
`cbX`/`cbY` = (0,0) for every destination pixel — an entire scaled webcam sampled ONE
source corner pixel. The web widget (the only CropBounds-bearing source) was never
affected, which is exactly why the transparency fix's own test suite stayed green while
this broke. **FIX:** default `cbW = src.Width`, `cbH = src.Height` (cbX/cbY = 0) before
the branch, only overridden when CropBounds is actually present. ONE test
`BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin` (two-color split
source, scaled non-1:1 through the paste cache) fails on the old code — right-half pixel
reads the left half's color — and passes on the fix; proven both ways with a stash/build/
revert cycle, not by inspection alone.
**Lesson:** when a shared low-level sampler gains a new optional code path (CropBounds),
audit EVERY variable the new branch introduces for a safe default in the branch it did
NOT touch — an uninitialized-to-zero "crop region" silently means "sample only pixel
(0,0)", not "no crop." grep for this shape (`var x = 0;` followed by an `if (cond) x = ...`
with no `else`) whenever a conditional metadata field is added to pixel math.
### Shrink / re-encode an image for the README (screenshots → small hero image)
Worked out 2026-08-29 (the recipe was NEVER recorded the first time it was done, so
it had to be re-derived from scratch — that's the incident this entry exists to end).
**Approach:** a throwaway Windows-dotnet console app uses WPF's imaging stack
(`System.Windows.Media.Imaging`) — same framework the app runs on, zero NuGet
packages, high-quality downscale via `TransformedBitmap`. Screenshots compress
**far smaller as JPEG than PNG** (PNG 1400px = ~1.2MB; JPEG q82 1400px = ~188KB).
**Recipe (run via the Windows dotnet host from WSL):**
1. Create `imgresize.csproj` targeting `net8.0-windows` with `<UseWPF>true</UseWPF>`
(SDK controller). Put it in a Windows-visible temp path, e.g.
`C:\Users\gramp\AppData\Local\Temp\imgresize` — NOT `/tmp` (Windows dotnet can't
reach a Linux-only path reliably).
2. `Program.cs`: load `BitmapImage` (`CacheOption=OnLoad` → `Freeze()`), downscale
with `TransformedBitmap(src, new ScaleTransform(scale, scale))` to max width
(1400 for the README hero), encode with `JpegBitmapEncoder { QualityLevel = 82 }`,
save.
3. Run:
```bash
"/mnt/c/Program Files/dotnet/dotnet.exe" run -c Release --project .
-- "C:\Users\gramp\Downloads\Screenshot 2026-08-29 075626.png"
"C:\Users\gramp\Documents\Code\projects\ytLive\docs\ytLlive-preview.jpg" 1400
```
4. Point `README.md` at the `.jpg` (not `.png`).
**Result:** 3.2MB screenshot → 1400×794 → **188KB** `ytLlive-preview.jpg` in `docs/`.
---
### Measuring audio-video A/V sync from a recording (no ears needed) — RECIPE
Derived 2026-09-12/14 (the "audio cut off at the end" / "audio delayed" reports). You
cannot "listen" to a take — measure it. Proven on two independent takes (talking+claps,
clean-clap test) with a tight, matching result each time.
**Method (FFmpeg probe + numpy, run from WSL):**
1. **Click/hiss-SILENT test takes are the gold standard.** Have the creator do a
loud, single-frame-syncable event (one clap after ~10s of near-silence) — the
video position of the hands-meet peak and the audio position of the transient
are both unambiguous. That single event replaced counting words forever.
2. **Extract audio as raw mono PCM** (`ffmpeg -vn -ac 1 -ar 48000 -f s16le`) and
compute a short window RMS envelope (5ms windows) in numpy. The clap is the global
max (`argmax`); also print a coarse 100ms table — it shows the pre-clap artifacts
(faint blips) vs the event (100× larger) vs true digital silence.
3. **Video motion per frame** — `-vf "tblend=all_mode=difference,signalstats,metadata=print"`
captured to stdout (NOT `file=` inside the filter — the `\\` path escapes mangle;
have PowerShell-style quoting bite; `metadata=print` writes to ffmpeg's log, so
redirect stdout to a file). Parse `YAVG` values; the clap is the frame with the
motion spike (2.75 vs a ~0.1–0.5 animated-widget background).
4. **Offset = audio_event_time − video_event_time**; every positive second means audio
is BEHIND video by that much. Cross-correlate the two full envelopes (60fps-resampled
RMS vs YAVG, normalized) to double-check the single-event peak — the autocorr-style
peak must be sharp (unique max at +133 frames, second-best ≈0.26).
5. **Sanity checks baked in:** count frames (`nb_frames` vs wall log) — if video is
real-time you've ruled out the truncation bug as the cause; compare stream
durations (audio > video by the lag is the SIGNATURE of a delayed audio tail, not
truncation); check `dropped writes` — dropped audio moves events EARLIER (opposite
sign), so it can never explain a lag. **A constant offset ≠ drift:** fixed lag =
buffering/backlog, growing offset = clock mismatch.
**The ROOT CAUSE the recipe led to (record it so it's never re-derived):** the mixer's
capture rings are fed from APP STARTUP (for the meters); `StartLive` never clears them,
so the drained audio trail begins ~a full ring-depth (2.0s) behind go-live. **Rule: any
"live preview/drain" sink fed by a continuously-capturing buffer MUST clear the buffer
at go-live, or early output is stale backlog.** The 2s ring + configured 300ms sync
delay + ~100ms pipeline = exactly the measured 2.11–2.22s. The lag ALSO scales with
"time since app launch" up to the ring depth — a take right after a restart reads as
~370ms while a fully-warmed app reads ~2.2s. Same code, wildly different numbers.
### Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
Derived 2026-09-03 (take-3 diagnosis — the stats seam from `97ffc42` named the stage
in one line: `17/300 frames per 5s, avg render 258.1ms, avg submit 1.5ms`).
Both halves were solved by OBS/libyuv long ago; do not re-derive:
1. **Pacing is a DEADLINE, never a post-render sleep.** `sleep(interval)` after each
frame makes the period `render + submit + interval` — the producer can hit ≤ half
the declared rate even with a free render. OBS's `video_thread`
(`libobs/media-io/video-io.c`) advances an absolute `nextTick += intervalTicks` and
sleeps only the remainder. **NEVER "rebase" the deadline to wall-now when it blows** —
the take-3 note I originally wrote here ("skip … the missed ticks (rebase)") was the
PROVEN-wrong advice: resetting `nextTick` erases the missed slots, so a 43fps reality
was authored into a 60fps container and every recording played ~1.4x fast (take 12/13,
2026-09-10). The rebase even kept the stats looking honest (render+wait == period, no
loss) because a rebased frame is never "late". OBS keeps the counter MOVING and outputs
ONE frame per interval tick — a late render shows as REPEATED footage (judder; duration
stays correct), never a skip and never a wall-now reset. Critical with rawvideo:
pts is stamped by ARRIVAL, so the muxed duration is pure frame count ÷ fps — the only
way to make duration == wall time under ANY load is exactly one frame per interval slot.
Corollary: **never add a second pacer.** `ffmpeg -re` on the rawvideo input throttles
the pipe independently and fights the pump (its "Resumed reading … after a lag" grew
0.79s→4.82s in the same take). Removed; the pump IS the pacer.
2. **A 1080p frame is ~2.07M pixels — the hot path must be row-simple.** Per-pixel
`Math.Round` + float source-over in managed code costs ~100ns/px = the whole 258ms.
libyuv's pattern (https://chromium.googlesource.com/libyuv/libyuv/): branch per
pixel on source alpha (opaque → 4-byte copy, transparent → skip), integer
fixed-point blend `(s*a + d*(255-a) + 127)/255` otherwise; and ALWAYS clip the loop
to the intersection rect (our social-bar overlay scanned all 2M dst px for a 64px
strip). A full-cover 1:1 blit also obsoletes the opaque-black pre-fill — skip dead
writes.
3. **On Windows, `Task.Delay` is a 15.6ms QUANTUM, not a timer.** Any request under one
system-clock tick sleeps a full tick (documented — learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay:
"approximately 15 milliseconds on Windows systems"). A deadline pacer built on Task.Delay caps the
producer at ~40fps-ish EVEN IF render is instant — take 9 proved the signature: work fell 26.5→22.4ms
but the period sat at ~37ms (≈ one padded wait/frame), so two real optimizations read as "zero change".
Frame-accurate loops (OBS/Chromium/game-loop canon — stackoverflow.com/questions/5441464) do:
`timeBeginPeriod(1)` for the session (paired with `timeEndPeriod`), sleep only the BULK of the
remainder, SPIN the last ~2ms across the deadline. Diagnostic before touching the compositor again:
period ≈ work + 15.6 → the SLEEP is the bug, not the work.
Related (take 14, 2026-09-04): **recycled ring buffers are a race you must SIZE, not just own.**
Deepening shared frames to kill GC churn (a fresh 8.3MB/tick array) hands out REUSED memory — the
ring's depth × source period must EXCEED the worst consumer hold (compositor read + lagged UI
preview copy), not just "a few frames". 4 slots at 144Hz capture laps in ~27ms vs a ≤50ms read: half
a new screen frame flashed over an old one in the recording ("bits flashing over other bits").
Depth 8 everywhere (screen/camera/web output rings); the paste-cache Epoch still guards identity.
4. **Hermetic pacing test:** inject the delay seam to RECORD the requested TimeSpan and
genuinely await it (`Task.Delay(d, ct)`) — a fake that returns
`Task.CompletedTask` synchronously makes the whole pump loop run on `StartAsync`'s
sync continuation and hang the test run (hit this 2026-09-03; the existing fakes all
yield for exactly this reason). Assert the REQUESTED wait (< interval with a
≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.
5. **Expensive content: raster on change, never on read (take 5, 2026-09-04).** A source
that updates once a minute (chat text!) must not full-rasterize (`FormattedText` +
`RenderTargetBitmap` + `CopyPixels` ≈ 15-25ms) every compositor tick. OBS text sources
re-render on property/message change; the per-tick pass blits the cache. Implement as:
content version (collection-changed counter) + config key (size/appearance) → cached
immutable `VideoFrame` returned by identity. Gotcha: buffers that SURVIVE sessions
(the chat log) silently arm the per-tick cost even in flows that never touch the
feature (signed-out record-only takes paid chat rendering!).
6. **An async loop started from a UI handler runs ON THE UI THREAD until you take it off.**
`await` continuations re-capture the current `SynchronizationContext` — the frame pump was started
from a WPF command handler, so the "WPF-free, hermetic" compositor rendered and read capture state
ON THE DISPATCHER, serialized behind the live preview itself, for the whole starvation saga. The
`wait` stat caught it only when the numbers became self-contradictory (render 22 + wait 10 > any
rebasing deadline — a blown deadline cannot sleep). OBS runs `obs_graphics_thread`/`video_thread`
as dedicated threads for exactly this reason. Pattern: `_task = Task.Run(() => Loop())` (null
context inside), then audit EVERY object the loop touches for UI affinity (RenderTargetBitmap /
DrawingVisual / WriteableBitmap: marshal the work or the rare miss; plain locked byte[] lookups:
fine) and pin it with a context test (`Pump_Produces_OffTheStartingContext`, inline-pumping
SynchronizationContext that the old code failed by construction). Cost: takes 3–10.
7. **Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8).**
The chat raster fix was REAL but the composer blamed it for the residual slowness it did not
own; two takes burned before the render/resolve split showed `resolve ≈ 0` and pointed at the
compositor pasting static layers per tick (`BlitCachedLayer` finished the job OBS-style). Before
shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the
NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.
8. **An encoder refed from a real-time loop must ENQUEUE, never pipe-write in the loop
(2026-09-10, take 15/16 — the "1...23...4...56..." smeared ticker).** After slice 9 the
recording played at the right DURATION but the content still hiccuped — and the aggregates
(301/300, uniform file PTS, 15.6s wall vs 15.74s file) could NOT see it. The cause finally
measured in `FfmpegEncoder.SubmitFrameAsync`: the `WriteAsync(8.3MB)+FlushAsync` to ffmpeg's
stdin BLOCKS whenever the encoder lags the pipe, and the slice-9 burst `while (now>=nextTick)`
then re-wrote that SAME stale composite for every slot that ticked past — frozen content runs.
OBS's answering machinery is the encoder queue (`libobs/obs-encoder.c`): the encoder thread
NEVER couples back into the video thread; overflow = dropped data, NEVER a frozen producer.
Fixed as: bounded `Channel<byte[]>` (cap 120) + a dedicated drain task owning stdin,
`SubmitFrameAsync` = copy-to-pool-array + `TryWrite` (drop-newest + count when full), stop
flushes the queue then EOF. **Rule: verify with a clock-independent judge.** The WSL ticker that
"proved" slice 9 has its
own Host-timer jitter under Windows load — so this slice burns a dot-matrix `_outputIndex`
into the bottom-right of every composite; decoding the recording reads the honest sequence
(+1/frame, jumps = counted drops) with no external clock involved. Take 17 must read
+1/frame from that strip. (A whole-frame duplicate scan was tried and is
UNRELIABLE here: the scene is always animating — session elapsed timer + REC pulse — so
no two frames are ever byte-identical.)
**SUPERSEDED by slice 15 (2026-09-14):** the last clause "pump emits ONE fresh composite per
iteration (no burst re-write)" had it BACKWARDS for the deadline. Slice 10 chose a skip when
the render overruns — an overrun's slots vanish from the file — and that AUTHORS ACCELERATED
playback (my item-1 lesson already said it: "NEVER a skip … a late render shows as REPEATED
footage"). Device takes proved it: one fresh frame per 35ms stall → `ty-20260914-1726` = 163
video frames (2.72s) against 2.93s of audio, video ending 0.22s early ("audio speeds up then
cuts off"). OBS's answer (docs.obsproject.com/backend-design: "If the video frame queue is
full, it will duplicate the last frame"; `libobs/obs-output.c` counts "lagged frames due to
rendering lag/stalls" — never a time-hole) is to fill each missed slot by DUPLICATING the
newest frame; duration == wall, judder not fast-forward. That's what the pump's submit now
does: a catch-up loop over the missed slots emitting this iteration's composite again, clamped
to a `deadlineNow` captured once (bounded — the smear that justified slice 10 was the
BLOCKING pipe-write re-copying during a long freeze; the queue makes each emit nanoseconds,
so the burst is safe). The burned index moved inside the submit loop: EVERY emitted slot
carries its own +1 (this also fixed the old unconditional `_outputIndex++` that gapped the
sequence on non-submitting fast-render iterations). A/V sync must be re-clap-measured after
this fix — the +0.6s audio-late reading on 1726 was confounded by the 1.1x acceleration.
**Latest A/V sync numbers (2026-09-14, pre-slice-15 pacing):** keep the recipe below honest on
clap measures. ty-1723: 697 video frames (60fps) = 11.62s vs 11.84s audio. ty-1726 (clap): 163
frames = 2.72s vs 2.93s audio; clap audio env peak 1.655s vs video motion cluster 0.75–1.03s →
single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late) — BOTH confounded by
the acceleration; re-measure after slice 15.
**Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
the freeze-audit recipe (RECIPE — re-deriving it cost this session):**
The slice-15 build fixed pacing but the desktop layer of the recording was still
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
measurement said something different AND worse than the running render theory.
`FramePump stall… worst render 33-36ms` was real but MOOT: once the camera+desktop
take was decoded to raw frames, the **desktop band was frozen 21s of 23.35s (90%)**
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
CONVERSION was the wall: the monitor delivers at the **240Hz DWM cadence**, the
source converts ONE frame at a time (`_framePending` latest-wins), and each
2560×1440→1920×1080 `DownscaleBgra` — naive double-per-pixel bilinear — cost
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
MediaCapture path) and audio were fine — exactly what the user reported.
**The audit recipe (ffmpeg → raw gray → numpy):**
```
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
python3 - <<EOF
import numpy as np
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
EOF
```
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
the compositor render stats led nowhere until this number existed. A **per-row
split detector** (`cumsum` of per-row diff-to-next minus diff-to-prev, argmax =
split row) then separated real mid-frame tears (score ≈ huge, both halves match
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
were the session UI, not tears.
**The fix (this slice):**
- **integer 8.8 fixed-point downscale, "shift only at the end"** — the SAME math as
`SceneCompositor.Bilinear` (rounded both stages in one 16.8 scale) ported into
`DownscaleBgra`, dropping per-pixel doubles to row-walk integer ops. The capture
ring already had the integer-bilinear lesson; the capture downscale itself was
still the naive float twin of the 258ms disaster.
- **throttle to the slot cadence** (`MinConvertInterval = 10ms`): the 240Hz arrival
is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per
second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps
pump can use.
- **ring reuse-distance, not ownership** (`FrameRingBuffer`, redLine 4): the
take-14 "depth × period" rule guards size; the slice-16 addition makes it
structural — a slot is only rewritten ≥4 rents after its last hand-out, else a
fresh buffer. `session.LatestFrame` survives across conversions and the
dispatcher preview copy lags, so "who released it" is unknowable without a
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
tear (new-top/old-bottom midway) is that read-under-write closed.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
bug):**
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
follow-on decisions fixed by evidence:
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
the pool to the master size to read back less. Microsoft's screen-capture page
forbids it: "the underlying Direct3D surface is always the size specified when
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
not scale — silently encode the desktop cut off. Readback must stay native; the
lever is concurrency, not size. **Rule: read the platform doc for the exact
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
have been wrong twice now.**
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
flight, completions can land out of order; a slow OLDER readback finishing last
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
compare-exchange publish) drops stale completions instead. Ring gets a lock because
rents are now concurrent; downscale row-scratch became per-conversion locals.
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
choosing between throughput and latency fixes — both previous slices picked the
wrong slot ("render" vs "conversion") until the audit existed.
---
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
integer bilinear) followed the same libyuv research
(https://chromium.googlesource.com/libyuv/libyuv/ — `row.cc`/`scale.cc` keep both
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
My first `Bilinear` shifted stage 1 back to 8-bit AND shifted the final result >>16 —
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
NOTHING (green webcam silently vanished from output; the pixel probes caught what
the eye in a 2x time-lapse would not). **Rule: multi-stage fixed point shifts only
at the end; verify against a uniform-255 sample before believing it.** Second trap:
a stale-byte sentinel test whose source pattern can generate the sentinel value
itself (0xAB was a legitimate `x+y` pixel) — pick the sentinel coprime/out-of-range
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
that HOLDS submitted frames now must snapshot them (`Clone`) once the producer
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
---
### A web widget captured at 10Hz inside a 60fps recording plays at ~1/6 speed
Derived 2026-09-10 (the "widget animation too slow" report). The recording is 60fps and the
WebView2 capture loop was a blind 100ms `DispatcherTimer` = 10Hz — each captured web frame gets
repeated ~6× in the file, so whatever the page animates at, the OUTPUT is capped at the *capture*
cadence, not the page's. Two stacked throttles, both real:
1. **The capture rate is the hard ceiling.** web-layer motion in the recording can never beat
`CapturePreviewAsync` frequency. But you can't just raise the timer: full-HD PNG capture costs
~10-30ms (WebView2Feedback#20: "CapturePreviewAsync … produces PNG/JPG and is very slow"), so
concurrent captures stack CPU AND can publish stale-after-fresh. The mandatory shape is a
**latest-wins drop**: at most one capture in flight per session; a tick during the in-flight
window is DROPPED, never queued; effective cadence = max(interval, capture duration). Do this
before anyone touches the interval constant.
2. **Chromium throttles hidden pages.** An off-screen WebView2 (we place it at (-5000,-5000)) is a
hidden page the moment the host window is unfocused or covered: `requestAnimationFrame` parks and
JS timers clamp to ~1s (WebView2Feedback#1172 — background-throttled rAF; #3070 — a WebView2 with
`Visibility.Collapsed` slows timers to 1s; Chrome-88 blog — heavy timer throttling of hidden tabs).
There is NO supported per-page opt-out (#5250 still open). The embedder answer is browser args on
the `CoreWebView2EnvironmentOptions` created BEFORE `EnsureCoreWebView2Async`:
`--disable-backgrounding-occluded-windows --disable-renderer-backgrounding
--disable-features=CalculateNativeWinOcclusion` (the Electron/Streamlabs-class fix for
occluded-window animation throttling). Share ONE environment across sessions (one browser process).
3. **Measure before picking the cadence.** Log the first ~30 captures' elapsed ms on the first
recording run; PNG encode + WPF decode of 1920×1080 is the per-capture cost that decides whether
~30Hz is affordable or it must drop to ~20Hz. The FramePump drops frames (never time-lapses) if
UI-thread GC churn starves it, so cost shows up as dropped-frame stats — read them.
---
## Splitting a large file into partials — NEVER `awk … > SRC` while awking SRC
(2026-08-31, Commit D) Tried to split `SocialsDialogViewModel.cs` in one line:
`{ awk '…' SRC; echo ""; awk '…2…' SRC; } > SRC`. The **first write truncated
SRC to 1 line**, so the second `awk` read the already-truncated file → the whole
source was lost (1 line left). Recovered with `git checkout -- SRC`, then redid
it, but the same bug could have meant making it up from scratch.
**Rule:** when a cut needs N blocks from one source into N files, never write a
block back onto the source that the `awk`s still read. Instead:
1. Read the source **once** at the start into temp files (`mktemp -d`, one file
per block), with a `$D` variable you carry forward.
2. Verify block sizes (`wc -l`) and brace balance (`python3 -c` counting `{`/`}`)
before touching any real file.
3. Then assemble each new file from `cat D/block …` — never truncating the source
until every read is done.
`git status` can't save you here if you don't notice until the file is gone —
`git checkout -- <path>` from the last commit is the recovery. Cheap insurance:
restore-then-retry, do it atomically from temp files the first time.
---
## Verifying an ffmpeg decode contract from WSL (no real CLR needed)
(2026-08-31, TASK 21) When a change depends on ffmpeg producing output with an
exact frame-size contract (rawvideo W×H×4 BGRA), you can prove the **command +
frame accounting** here without any .NET process:
1. Fetch a **static Linux ffmpeg** into `/tmp/opencode` (no sudo needed):
`curl -sLO https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz
&& tar -xf …`
2. Generate a tiny known clip: `ffmpeg -f lavfi -i "testsrc2=duration=1:size=640x360:rate=30" -pix_fmt yuv420p clip.mp4`
3. Decode with **exactly the app's args**: `-f rawvideo -pix_fmt bgra -vf scale=640:360 -an`
4. Assert `total_bytes % (W*H*4) == 0` (python3) → exact integer frames, no pad.
**Why not a dotnet-spawned ffmpeg here:** the only CLR on this box
(`/home/gramps/bin/dotnet`) is a **Windows-bound shim** — `Process.Start` resolves
paths to `\\wsl.localhost\Debian\…` and throws "not a valid application for this
OS" when handed a Linux ELF ffmpeg. So never plan to have dotnet exec a Linux
ffmpeg here; verify the contract with shell/python instead, and leave the
CLR→real-ffmpeg run to the native Windows suite.
## WPF hit-test truth in tests: `UIElement.InputHitTest`, NOT `VisualTreeHelper.HitTest`
(2026-09-01, the RoundClip "known failure" post-mortem — a failure the map carried as
"not a regression" for weeks without ever recording WHY.)
**The trap:** `VisualTreeHelper.HitTest(window, pt)` returned the window's
`WebViewHostPanel` overlay (`IsHitTestVisible="False"`, `Opacity=0`, ZERO children — since the
2026-09-14 composition-capture redesign the overlay is DELETED, but the API lesson stands) for
EVERY point in the window — so a "corner is grabbable" assertion could never pass, and
it looked like a real interaction bug. The actual input pipeline (`UIElement.InputHitTest`,
what Mouse routing uses) correctly returned the element's Grid at elem-center/corner-in/
corner-exact and fell through to CanvasGrid just past the corner. The product was fine;
the TEST was probing an API that doesn't model input semantics.
**Rule:** any test asserting "where does a click land" uses `window.InputHitTest(pt)` +
`IsDescendantOf` — never `VisualTreeHelper.HitTest`.
**Diagnosis recipe (how the lie was caught in ~3 probe cycles, no guessing):** add a TEMP
probe `[Fact]` in the RealApp collection that hit-tests a spread of points
(elem-center / corner-in / corner-exact / corner-out / bg-center) and `Assert.Fail`s with a
composed dump: per-point VTH hit + `InputHitTest` hit + ancestor chain (`GetParent` walk
with `#Name`) + panel properties (`IsHitTestVisible/Opacity/children/actual size`) +
`TranslatePoint` origins. Run the class alone, read the message, delete the probe.
**Two sibling facts learned the same session (record-once):**
1. A UserControl owns its own XAML namescope — after extracting a region out of a window,
`window.FindName("InnerPart")` returns null; resolve the UserControl by its window-level
name, then `pane.FindName("InnerPart")`. And window-scope STYLES are invisible to a
UserControl's `StaticResource` at parse time — move such styles to `Themes/Controls.xaml`
(the app-scope rule exists for this).
2. Per-class `dotnet.exe vstest` from WSL DOES execute the RealApp/`MainWindow` tests fine
(they passed natively 2026-09-01) — only the FULL suite hangs (WASAPI startup). And the
test process shares `%APPDATA%\ytLlive\startup.log` with the real app: lines like
`camera 'test-camera' failed` are test noise, not DB state — to check pollution, query
the DB directly (`python3 sqlite3`, `SELECT DeviceId FROM Webcam`), not the log.
## A "known failure" label without a recorded cause = a bug on life support
(2026-09-01, the audio triple-take) One line — `_delayedMix` (nullable, added by TASK 22,
never initialized) dereferenced as `delayed.Length` — produced THREE symptoms that lived in
the map as two separate "pre-existing, do-not-chase" entries: (a)
`Mix_HonorsProviderGains…` "known failure", (b) `AudioPipelineTests` hangs when run at all
(a test reading a named pipe with no writer blocks — hung test ≠ flaky test, it's a starved
producer), (c) startup.log flooded "Audio live loop error" every 10ms (the mixer loop caught
and logged only `ex.Message` — stack thrown away).
Rules derived:
1. NEVER label a test "known/pre-existing" without writing WHY (exception type + first app
frame). An unexplained known-failure is deferred archaeology that hardens into fog.
2. A test that waits on IPC + a producer whose output vanished are usually ONE bug — look
for the producer before blaming the test.
3. Catch-and-log-swallow of `ex.Message` hides root causes; log with stack (`AppLog.Write(ex,
...)`) and throttle (5s) instead of dropping or flooding.
Bonus: the "failing" test encoded the MAP's contract (`loopbackGain = GameAudioVolume`); the
code had drifted to unity on a disproven premise (loopback capture does NOT follow endpoint
volume — creator's 20%-volume/pegged-meter observation killed it). The test was right all
along — failing tests may be the last honest witnesses; interrogate, don't pardon.
## Feature provenance: record WHO asked and WHY, in the task entry itself
(2026-08-31/09-01, the SYNC slider scare) TASK 22's lip-sync slider surfaced on the preview rail and
the creator's reaction was "totally don't remember ordering that" — because the queue entry recorded
WHAT shipped (a slider, 0-500ms, a converter class) but not WHO asked (the creator, explicitly, for
OBS's delay-filter fix built natively). Eight days later his own request read like AI drift and nearly
got deleted. Rule: the moment a creator-driven feature is queued or shipped, its entry carries a
one-line provenance — *who asked, what triggered it* ("creator: OBS delay-filter lip-sync fix, native").
Features without attribution become roadmap orphans that get punted, removed, or re-litigated. Same
disease as an unexplained "known failure" label — a fact recorded without its reason is a future
argument.
## An un-attributed build invalidated three takes of a perf saga — stamp the binary
(2026-09-04, takes 4–6) After each render-perf fix the creator "exed the code" and re-recorded, but
the exe timestamp ≠ binary contents (incremental builds reuse whatever compiles clean; a source edit
with no rebuild serves the OLD exe). Take 6 measured render WORSE than take 5 (35-41ms) and there was
no honest way to tell "the chat cache fix doesn't work" from "the fix was never running" — three
hours of diagnosis on an unattributable sample. Rule: if takes measure the app, EVERY build carries
an id and EVERY log line traces to it — `GenerateBuildStamp` (csproj) writes a fresh GUID per
compile (deliberately defeating incremental lies), the wordmark shows it as a superscript, startup.log
records `Build <id> (compiled <time>)`. A perf claim without build attribution is a guess; ask for the
stamp BEFORE theorizing. (Also this session: a sentinel-byte test where the source pattern could
GENERATE the sentinel, and a fixed-point bilinear that shifted BOTH stages and silently drew nothing —
see the rawvideo recipe.)
## Signed A/V sync: negative offset = eat the buffer head (OBS semantics), armed at go-live
(2026-09-14) The ring-backlog fix shrank A/V lag from ~2.2s to ~0.54s, and ~300ms of that was
baked into the DB (`Audio.SyncOffsetMs=300`) via a POSITIVE-only control. For audio running BEHIND
video you cannot push audio later — positive delay makes it worse. OBS's answer is a NEGATIVE offset
that eats the buffer head: drop the first |N| ms of the written stream so every audio event lands
|N| ms EARLIER relative to video. It only makes sense at the head of the stream, so it is armed once
at `StartLive` (positive stays live-reactive via the delay line). Rule: record the signed semantics
together — positive = delay (ahead), negative = eat-head advance (behind) — and lock the control
while live (`IsEditMode`) since a mid-stream advance flip is meaningless. Test recipe:
`StartLive_NegativeOffset_AdvancesAudio_ByDroppingTheStreamHead` (emit 6×0.9 into an 8-tick budget,
long 0.2 bed, wire must show only 0.2).
## Composition video capture of WebView2: the CoreMessaging DQ recipe + the build gotchas
(2026-09-14, the ~20→60fps web-capture redesign — record-once so the next GPU-path work never
re-derives it. Everything that follows was already settled by OBS/Flutter `webview_windows`; the
cost this session was the 19041 projection's gaps.)
1. **No `DispatcherQueueController.CreateOnCurrentThread()` on the 10.0.19041 projection**
(CS0117 — only `CreateOnDedicatedThread` + `FromAbi(IntPtr)` exist). P/Invoke
`coreMessaging.dll!CreateDispatcherQueueController` with a sequential
`DispatcherQueueOptions { DwSize, ThreadType=2 (DQTYPE_THREAD_CURRENT), ApartmentType=2 (DQTAT_COM_STA) }`,
wrap via `DispatcherQueueController.FromAbi(ptr)` (mirror of `CaptureInterop`), THEN `new Compositor()`.
One controller + one compositor per UI thread, created once.
2. **`CoreWebView2CompositionController` needs a real parent HWND** — there is no window in the
MainWindow ctor, so `InitWebView2()` moved from the ctor to `MainWindow_Loaded` (source
registration is guarded by `ContainsKey`, so the layout-load path is safe).
3. **Capture the ROOT visual, not the control.** `GraphicsItem` for a composition controller comes
from `GraphicsCaptureItem.CreateFromVisual(root)` where `root` is your own 1920×1080
`ContainerVisual` (`IsVisible=true`) holding the controller's `RootVisualTarget` as a
`RelativeSizeAdjustment=1,1` child — exactly what `webview_windows`'s `graphics_context.cc`
does (`CreateGraphicsCaptureItemFromVisual` on the root `surface_` visual). Hide nothing,
move nothing, poll nothing — the capture is frame-driven at the renderer's pace.
4. **Straight alpha, never premultiplied.** Read back with `BitmapAlphaMode.Straight`; the
compositor's blend is straight-alpha, so `Premultiplied` readback wrecks corner anti-aliasing.
5. **Namespace landmines that each cost a build cycle in `Services/`:** `Compositor` resolves to the
repo's OWN `ytLive.Services.Compositor` namespace — fully-qualify `Windows.UI.Composition.Compositor`;
`CoreWebView2CompositionController.Close()` is the disposal call (no `Dispose()`); `Color` is
ambiguous (`System.Drawing` vs `System.Windows.Media`) — qualify `System.Drawing.Color.Transparent`.
6. **Testing a compositor-built bitmap without a runtime:** the internal seam ctor
`(Dispatcher, Func<string, IScreenCaptureSource>?)` + a `FakeWebSource` + a real background-STA
`DispatcherPump` (borrowed from ScreenCaptureManagerTests) drives the whole session path hermetic.
When byte-comparing a crop, slice the source with stride gaps (helper `CropBytes`) — a contiguous
range silently spans rows.
Recipe verified on green suite + 0-warning build; the composition path itself still needs a device
take (open item, HANDOFF).
## A pipe-read test that samples one tick is a timing flake by construction
(2026-09-14, re-discovered) `Mix_HonorsProviderGains_AndGameMute_KillsTheLoopback` fails
sporadically in ISOLATION (3/3 on the clean tree) but passes in the full suite: it reads exactly one
tick's worth after muting, and if the pipe still carries pre-mute buffered bytes the read spikes at
0.4 instead of silence. Any caller of `ReadFullyAsync` on a live pipe must either close/drain the
pipe first or assert on a multi-tick window. Do not blame a sync change for this — verify the grain
against the clean tree in the same mode before touching the mixer.
---
### Slice-18 follow-up (2026-09-15) — C4 composite cache: test fakes must mirror production's STABLE scene and per-tick purity
The C4 blit-on-change cache keys on a hash of the resolved inputs, INCLUDING per-element reference
identity (`RuntimeHelpers.GetHashCode(element)`). Two test-setup habits silently broke/starved it:
1. **`NewPump`'s default scene is fresh per tick** (`() => BackgroundScene()` — fine for pacing
tests) — churned the element refs, so the signature NEVER matched and the cache looked broken
(301 "renders" instead of 1). Production hands a STABLE `StagedScene`. **Rule: any FramePump test
that asserts per-frame content or cache behavior must pass `scene: () => scene` with one Scene
instance; the fresh-scene default is only for pacing/diagnostic tests.**
2. **A call-count resolver flip (`flip++ % 2`) double-advances under a cache-aware pump** — the
tick resolves TWICE (signature pass + compositor pass). Key per-tick content on the burned
`OutputIndex` (stable until submit, which happens AFTER render) or on `encoder.Frames.Count`
instead. A stable-per-tick token, not a call counter, is the deterministic alternation.
3. New frames that must invalidate the cache need BOTH a fresh array AND a fresh Epoch (array
identity alone is unchanged for a mutated-in-place array; `Epoch` is the monotonic generation
marker the compositor's paste key and the C4 signature share).
---
### ALERT-CLIP DECODE PIPE (RECIPE — TASK 47, 2026-09-26) — read a per-play mp4 into loose frames + stereo 48k audio
Worked out for the alert-box video: a short mp4 played **per alert** (one-shot) and dropped at
drain, NOT the refcounted media-source manager. Recipe:
- **Two ffmpeg pipe outputs, one child:** `ffmpeg -loglevel error -i "{file}" -f rawvideo -pix_fmt
bgra -vf scale={W}:{H} - -f f32le -ac 2 -ar 48000 -` — `-loglevel error` keeps stderr quiet (the
other pipes' ffmpeg rivals read the whole FD so a chatty stderr deadlocks); do **not** drain
stderr, mirror the `MediaVideoSource` split. Video pipe read backs onto `RawVideoFrameReader`
(TASK 21) which already pace-throttles to rtc-frame-rate… earlier live pacing is VFR-driven.
- **Audio = a carry-buffer loop, never a per-read slice.** PCM16 bytes → float32 chunks changes the
per-frame byte count; the naive `ReadFullyAsync(count)` per chunk **drops samples that straddle a
pipe read boundary** (the first draft's bug). The fix: an in-loop buffer that carries leftovers —
read `count - pending` bytes, emit whole float chunks from `pending + fresh`, keep the remainder.
Pacer: 50ms audio chunks → `Thread.Sleep` the (chunk length − read wall time) remainder.
- **bgra from ffmpeg is OPAQUE** (alpha=255). Whoever fades must copy the frame and scale alpha
across the copy (straight-alpha); the decoder never touches alpha. That keeps the decoder pure
and the fade a layer concern.
- **Fake-decoder convention:** the integration test implements `IAlertClipDecoder` inline with
`EmitFrame/EmitAudio/End` helpers — Start/Stop/Dispose are recorded as int counters, frames and
audio are pushed from the test thread, EOF is a test-raised event. No ffmpeg, no threads, still
the full real-time path through the layer.
- Re-grep here before building a second "play a file into the mix" path — snippet above is the
whole non-trivial part.
### A component's drain must reset its OWN anchor state (TASK 47 post-mortem — the test caught what a no-frame path hid)
`AlertOverlayLayer.Advance` gained a clip branch for video (TASK 47). In the ANIMATION branch the
old code did `_current = null; _elapsed = 0;` before `AdvanceToNext()`; the new clip branch called
only `StopClip(); AdvanceToNext();` — so after the EOF fade-out the layer kept `_current` pointing
at the finished message: `IsPlaying` stayed true, `RenderFrame` re-rendered the alert through the
animation fallback, and the layer never returned to idle. The six-animation tests couldn't see it
(they only ever exercise the animation branch); the video test's explicit
`Assert.False(layer.IsPlaying)` after drain nailed it on the first run. Rule: **when a second
execution path is added to a state machine, the drain/precondition contract of the ORIGINAL path is
part of the spec** — assert the component's end-state on the new path, don't assume the old path's
cleanup. The forest for the trip: the fallback path (animation) MASKED the broken path because its
output still looked like "an alert rendering", exactly the flick I'd have shipped if the test
hadn't forced the null-frame idle state.
### A seam fed from the chat poller must marshal EVERY WPF-object write to the UI thread (AlertOverlayLayer crash, 2026-09-26)
First crash in the app's history (startup.log `11:22:01.301`), and the sims gave zero warning for
months. Test-tab alert buttons run on the UI thread, so `OnMessageReceived` → preview writes were
always UI-threaded there. The moment a REAL message round-tripped through the chat poller:
`System.ArgumentException: Must create DependencySource on same Thread as the DependencyObject` in
`DataBindEngine.ProcessCrossThreadRequests` — WriteableBitmap created on the MTA poller thread,
`alertBox.VideoImageSource` raises INPC (`DisplaySource`) and is WPF-bound, the binding engine
re-binds cross-thread, process dies.
The chat layer already had the rule (`ChatOverlayLayer.OnMessageReceived` wraps its body in
`Application.Current.Dispatcher.Invoke`); the alert layer's copy of the seam didn't. Rule now
recorded for every layer: **the chat poller is an MTA producer — any INPC-raising seam it feeds
must marshal every WPF-object write (WriteableBitmap, DP/INPC-raised bound sources) to the UI
thread; wrap the WHOLE ingest, not just one writer, so enqueue/timer/preview stay coherent on the
UI thread. Guard `Application.Current` null like AlertTickerRenderer.Render does (pure seams).**
Red/green proof pattern: call the seam from a raw `new Thread` (MTA) while the RealApp message
loop runs, return to the loop so the marshalled work can execute (never `Join` on the UI thread),
then assert on the UI thread that `source.VideoImageSource.Dispatcher` is the App's dispatcher.
Red pre-fix (it was the poller's), green post-fix.
Grep-ahead: a `DispatcherTimer` ctor guarded by `if (Application.Current != null)` marks a class
with UI-thread affinity — audit every production ingress point for the marshal when adding one.
---
## A decoder that "starts" but yields nothing is DEAD AIR — fall back on a no-frame grace, never trust silent catches (2026-09-26)
Second live alert failure of the day, a different void than the crash: creator's repro — On-Air →
test → remove and re-add the Stream Alerts layer → TEST event → procs in chat but the alert box
stayed transparent, and startup.log 11:44's `peakMix 0.105` proved the alert audio never reached the
mix (vs 11:21's 0.375–0.891). ZERO exceptions anywhere; every failure stage was silent BY DESIGN:
`UpdatePreview`'s catch was `Debug.WriteLine`-only (nothing in release), and `AlertClipDecoder.RunAsync`
had a bare `catch { }` around the whole run (fires `Completed` regardless). A decode run that fails
AFTER `BeginClip` succeeds leaves `_clip != null` with `_latestClipFrame == null` → `RenderClipFrame`
returns null forever → box transparent FOREVER, and the six-animation fallback only ever ran on
setup-time throws. Even frame pacing was silently off (no frame-rate probe → `frameDuration = 0` →
video dumps instantly while 50 ms audio chunks pace the whole clip).
The established answer (derivative: any video-player/booth failure budget has the same contract):
**an alert that plays is a promise — the moment data stops arriving, degrade, don't vanish.** Fixed
in `AlertOverlayLayer.Advance`: a started clip that produced neither a video frame NOR an audio chunk
inside a `NoFrameFallbackSeconds` (1.0 s) grace is torn down, `_elapsed` resets to 0, and the SAME
alert continues as the six-animation branch — never drained early, never blank. And the silence is
gone: `BeginClip` writes the box id / resolved path / size + setup failure; `RunAsync` now logs the
frame/chunk end-state AND the exception it swallowed. Rules for every future gamble:
(1) a `catch` that hides WHY the output vanished is a lie — log the exception and the counters;
(2) any producer that yields zero output inside a grace becomes a fallback trigger, not a wait;
(3) one integration test per rescue: `SilentDecoder_FallsBackToTheAnimationAfterTheNoFrameGrace`
pins transparent→fell-back→renders→drains.
The pacing half that this entry flagged came true the same day: the 12:35 session showed a FULLY
HEALTHY decoder (`frames=240 audioChunks=156 failed=False`, alert audio in the mix) and the creator
still saw "no video / you lost the video". The clip played at ~240× for a second then froze on the
last frame: `AlertClipDecoderFor` passed no `frameRateProbe`, so `RunVideoAsync`'s pace
(`if (frameDuration > 0)`) was dead code while the audio chunk-pacer ran real-time. LESSON, sharper
than the whole fallback saga: **a decoder that drops data faster than wall-clock looks identical to a
dead decoder on screen — "plays but you don't see it" is a PACING bug before it's a decode bug.**
A healthy pipeline still needs an explicit pace per stream; never ship a pipe whose `frameDuration`
can be `TimeSpan.Zero`. Fix: the alert factory passes `FfmpegFrameRateProbe` exactly like the media
path does (MainViewModel.cs:308). Verify-by-seam once, everywhere: the same fake-process +
fake-probe delay-recorder test now pins real-time video pacing
(`AlertClipDecoderTests.AlertClipDecoder_PacesVideoFramesToTheProbedFrameRate`). Rule added to the
list above: (4) if output arrives but the USER still reports no/motionless picture, check pacing
before decode — count the bytes/s against wall-clock, don't re-audit the codec.
**RULE (5) — the last hop is a hop. A producer that hands over a correct frame has still
delivered nothing.** The 2026-09-26 alert hunt burned four builds (89/90 + two more) proving
the video was PERFECT: real h264, 240 frames, bright pixels at every timestamp, byte-identical
to the creator's source file, opaque `a255` frames logged every second into both the preview
and the compositor resolver. The box was blank because `PreviewPane.xaml`'s per-element
`Image` is `Collapsed` unless a DataTrigger says otherwise, and nobody had ever written the
`IsAlertBox` one. Two lessons, both expensive:
- **"It decodes, it's handed over, it's still invisible" ⇒ stop auditing the producer and
audit the CONSUMER'S VISIBILITY.** Bytes existing is not pixels being drawn. Diff the two
consumers' behaviour: chat box rendered, alert box didn't — that asymmetry was the whole
clue, and every per-frame log said "healthy".
- **A test whose fakes replace the failing seam will pass forever.** `AlertLayerVideoTests`
drove `RenderFrame` with a fake decoder and was green through the entire outage; the decode
test used a fake process; nothing ever crossed the XAML. When a feature's output is
*displayed*, the test must build the real display path (`AlertBoxPreviewVisibilityTests`
drives the real `MainWindow`+`PreviewPane`). And when the symptom is "nothing appears",
the honest test asserts on the CONSUMER (is the Image Visible / are the canvas pixels
different), never on the producer's own counters.
- Corollary: never trust "the data arrived" as a root cause. Ask what draws it.
**RULE (6) — a per-element `Image` is the wrong home for a GLOBAL overlay, and an
output-only producer is invisible by construction.** Same hunt, one layer up
(2026-09-26, after the `IsAlertBox` fix): the creator confirmed the clip plays, then
asked "where is the scrolling text? I do not see it." The ticker was healthy the whole
time — `AlertTickerRenderer` produced a 1920x48 frame and `FramePump` blitted it at
(0,0) — but it existed ONLY as a frame-pump callback (`alertTicker: () =>
_alertLayer.AlertTickerFrame`). `PreviewPane.xaml` had no element for it, because the
strip is master-width, not a `Source`, so it can't ride a per-element `Image`. Two
lessons:
- **When a feature's output is global (full frame width, a fixed overlay slot), it needs
its own root element + its own VM-published `WriteableBitmap`, exactly like
`SocialBarElement` — not a `Source`.** A global overlay that only a background thread
consumes will reach the stream and never the creator's preview, and no amount of
producer-side logging will show it: there is no consumer to be unhealthy.
- **Ask "who displays this?" while designing, not after the report.** RULE (5) says
audit the consumer when a picture is missing. RULE (6) is the earlier lesson: the
consumer may not exist yet, and that is invisible from the producer side.
**RULE (7) — pace a marquee in READS PER ALERT, never in pixels per second.** The ticker
shipped at a fixed `PixelsPerSecond = 140`, which sounds reasonable and was wrong: one
full pass took ~17s, so the 10s default clip showed the announcement barely once,
entering from the right and never crossing. The number that matches the creator's intent
("readable three times") is **passes per alert**, divided into the alert's own length —
the same convention the incumbent uses (Streamlabs exposes "Alert Duration: choose how
long your alert stays on your stream" and "Text Delay", never a scroll-speed slider;
https://support.streamlabs.com/hc/en-us/articles/52499995174299-Setting-up-Your-Streamlabs-Alerts).
So: `readsPerAlert / alertSeconds` => a pass length => a speed derived from it. Corollary
that cost a test rewrite: **a seamless marquee can never go blank**, so "the frame went
empty" is NOT a valid way to count passes — count the pill's leading edge resetting to
the right edge instead. And phase-start the run half a frame in, or the first published
frame is blank (the strip sits exactly off the right edge) and the preview opens on
nothing.
**RULE (8) — a bound `ItemsSource` ComboBox can break an unrelated real-mouse test.**
`LeftPanel.xaml` gained a display selector written as
`ItemsSource="{Binding ...}" + DisplayMemberPath + SelectedValuePath`; it compiled, my
feature's tests passed, and it broke `LayerReorderPersistenceTests.RealMouseDrag_OnThe
LayerList_PersistsTheReorder` — a test that injects PHYSICAL mouse input
(`SetCursorPos` + `mouse_event`). A data-bound selection resolves during load and
re-measures the panel under the test's feet, so the real cursor lands on the wrong row.
Rewriting it as inline `<ComboBoxItem>`s (the shape the chat Font selector already uses
in the same panel) fixed it with no other change. LESSON: **when a UI-only change breaks a
real-input test, the control's data-binding style is the suspect, not the feature** — and
bisect by reverting ONE file (`git checkout Controls/LeftPanel.xaml`) before theorising.
Prefer the idiom already in the file you are editing.
---
## 2026-09-26 — A blocking wait on the shared `RealApp` STA thread hangs the ENTIRE suite silently
**Symptom:** `vstest` printed 4 lines and then nothing. No test ever completed, no timeout, no
stack trace. Ran it three times, plus after a full rebuild and a reboot. Looked exactly like a
broken test runner.
**It was not the runner.** Two things were stacked:
1. **The runner output was being swallowed by my own shell plumbing.** Running `vstest` in the
foreground with a pipe to `tail` produced *zero* bytes even when the run finished. Launching it
with `nohup … > log 2>&1 &` and polling the log worked every time. **Recipe: always background
a `vstest` run and poll the log file.** Never trust a foreground piped `vstest` in this shell.
2. **The real hang was my test.** `LayerReorderPersistenceTests` and friends share ONE STA thread
via the `RealApp` collection. I had written the frame-pump tests as
`_app.Run(() => { pump.StartAsync(ct).GetAwaiter().GetResult(); … })`. `Run` *occupies* that
single thread, so the pump had nowhere to run and the two deadlocked — permanently, with no
timeout to save it.
**The rule:** a test that drives the frame pump MUST be `async` and use the new
`RealAppHost.RunAsync(Func<Task>)`, with plain `await pump.StartAsync(...)`. `await` yields the
STA thread back to the dispatcher; `.GetAwaiter().GetResult()` holds it hostage. **Never put a
blocking wait inside `RealAppHost.Run`.** A collection-serialized UI thread turns a missing
`await` into a whole-suite hang, which is why this presented as "the runner is broken."
**Also:** isolate before fixing. Running just the suspect class proved the runner was fine and my
test was the hang. Three blind retries on the same command would have "confirmed" a false theory.
---
## 2026-09-26 — Two real bugs my own new tests caught, and the cadence trap
Both were found only because the tests asserted the creator's *stated numbers* rather than
"something was drawn":
1. **Interval cadence implemented as an absolute deadline.** I had `_sinceStart` accumulate forever
and compared it to `_untilNextFlash`, which I *replaced* with a new interval on each fire. That
makes the second value an absolute time, not an interval, so gaps **shrink every cycle**:
measured 6.7s, 3.5s, 2.8s instead of 30–60s. Fix: keep an absolute `_nextFlashAt` but set it to
`_sinceStart + interval` on each fire, so the interval runs *from the last flash*. General trap:
whenever a timer holds "when did I last do X" and "when next", decide explicitly whether the
period is relative to the last event or to the start.
2. **A fade-in that starts at exactly 0 opacity emits a blank first frame.** `OpacityAt(0) == 0`, so
`Frame` returned null on the tick right after `Show()`. Not wrong (it is the start of a fade) but
it wasted a tick and made "is a credit on screen right now?" answer false. Worth knowing before
writing the assertion — my test's arithmetic was wrong, not the renderer, and I nearly "fixed"
the renderer to match the bad test.
**The lesson:** pin the *numbers the creator asked for* (2s, 30–60s, fully-on-canvas). "It rendered
something" passes against almost any implementation and caught neither bug.
**Third trap, cost me a wasted verification:** to prove a `const` isn't shipped I grepped the
Release binary for it — and it was absent from the *Debug* binary too, because **consts are inlined
at compile time and never reach metadata**. The control (run it against Debug as well) is what
exposed that. Grep for a **field/type** (`InstanceVariable`) or use reflection; never grep for an
inlined literal, and always run the check against the build you expect it to FAIL in.
## 2026-09-26 — A falsy `ShowDialog()` fell through into the side effect: "Cancel" that saved
The recording save prompt let the creator name the file. It always had a **Cancel** button, and
Cancel always *appeared* to work — the dialog closed, the toast was skipped, nothing visibly went
wrong. The footage was saved anyway, under the default name.
```csharp
// WRONG — and it type-checked, read fine, and looked correct
string stem = autoStem;
var dialog = new RenameRecordingDialog(autoStem);
if (dialog.ShowDialog() == true && !string.IsNullOrEmpty(dialog.ResultFileName))
stem = dialog.ResultFileName;
var finalPath = UniquePath(Path.Combine(dir, stem + ".mp4"));
File.Move(startPath, finalPath); // <-- runs even when the creator hit Cancel
```
**The trap:** `ShowDialog()` returning `false` was collapsed into the *default* branch instead of
being treated as a distinct outcome. "Not what I typed" and "I don't want this" are different
answers, and the nullable stem made the first one look like a legitimate value of the second. The
default value of the accumulator absorbed the cancel.
**The rule, and the grep-able form of it:** when a modal's result gates a side effect, every branch
is explicit and the side effect sits **inside** the accepting branch:
```csharp
if (!dialog.ShowDialog() == true) { DoTheOppositeThing(); return; } // read the FALSE case first
DoTheThing(); // the effect is not a fall-through
```
A `DialogResult`/return value that only ever *modifies* a later value (rather than *permitting* the
effect) is the smell. `null` from a dialog is the same trap wearing a nullable.
**Why it survived so long:** the side effect is in a *different method* from the modal, and the modal
had its own passing test (`RenameRecordingDialogSizingTests`, which only checks layout). The test
suite was green because the bug was never in the tested unit. **A modal's own test proves nothing
about the decision the modal feeds** — cover the commit-or-discard decision itself, against real
files (`ytLive.Tests/RecordingSaveDialogTests.cs` drives the extracted
`MainViewModel.CompleteRecordingSave` and asserts Cancel leaves the directory EMPTY, not just "the
stem is unchanged").
**Creator's actual ruling, worth keeping verbatim:** "if I click cancel, that doesn't mean to accept
the default file name and save the file — that's what pressing enter should do. Cancel should cancel
the save option, discarding the recording." A Cancel button on a destructive-and-final dialog means
*discard*, not *accept default*. When a dialog is the last gate before data is committed, ask what
Cancel should destroy — don't infer "keep the default" just because the default already exists.
---
## 2026-09-27 — A policy number in my head is not a policy finding: check that the policy *applies*
Two confident, wrong claims in one Store-research pass, both of the same shape — I had a policy
number and a confident paraphrase attached to it, and never asked whether the policy's own scope
clause covered us.
**1. "Store policy 10.10 (Advertising) requires X"** — 10.10 does not apply to LlamaCasty at
all. Every clause in it is conditioned on **ads your product displays**. YouTube injects ads
*server-side after the RTMP handoff*; the app renders no ad, hosts no ad SDK, and makes no ad
decision. I spent a cycle designing around a rule that was never in force. Same error later with
**10.2.5** (Store-only installation): the current policy text scopes it to **game and Xbox
console products**, and the older v7.7 wording that said "all of your apps" has been narrowed —
citing the stale wording would have produced a confidently wrong conclusion in the *opposite*
direction (claiming we were locked to MSIX).
**2. "Buy an EV certificate — EV gets instant SmartScreen reputation."** This one was worse than
wrong, because it was already **written into `Distribution.md`** and would have cost $400+/yr.
Microsoft **removed the SmartScreen instant-bypass for EV certs in March 2024**. EV and OV now
accrue reputation identically. I recited a fact that had been dead for over two years, in a
document that already contained the recommendation.
**The rule, and the generalisation:** *a policy citation is a claim about scope, not just about
text.* Before acting on one, read the **applicability sentence** and answer three questions —
does it cover our product class, does it cover our distribution model, and is the wording the
**current** version? A policy number without a verified scope is a guess wearing a citation's
clothes, and the failure mode is invisible because the sentence still *sounds* authoritative.
Corollary: **a stale external fact inside a shipped document is worse than a missing one** — it
is trusted, acted on, and costed. When a doc makes a factual claim about the outside world
(certificates, platform behaviour, policy text), re-verify it on touch, not on memory. This is
the same class as the dead-Pin incident in `FfmpegLocator` (`ai.md` → FFmpeg locator): an
external thing changed, our record of it did not, and the mismatch was discovered by a user
rather than by a test. **Records of the outside world rot silently and are believed until they
cost money.**
---
## 2026-09-27 — I offered a mitigation I had not checked the mechanics of: "just refund anyone who asks"
While discussing a subscription price, I wrote that at `$99/yr` with zero marginal cost the
creator could **"just refund anyone who asks"** — presented as a free mitigation against the
churn/refund burden of a subscription. **That lever does not exist for Microsoft Store IAP.**
Refunds run through Microsoft, not the developer; the buyer requests them from the Store, and
the developer cannot unilaterally issue one. I proposed it before ever checking how Store
refund mechanics work, and it sounded confident because it was arithmetically true and
operationally impossible.
**The rule:** *a mitigation is not an answer until you know who has the authority to execute
it.* Before offering "we can just do X," ask **who presses the button** — and if the answer is
"Microsoft, or the user's bank, or a regulator," then it is not our mitigation and must not be
counted as one. The same sentence reappears in a subtler failure: the underlying instinct
(refunds *should* be cheap when marginal cost is zero) was sound; the *agency* was invented.
**The verification that came out of it is worth more than the mistake.** Checking Store refund
policy produced a hard constraint that **did** reshape the pricing design and would otherwise
have produced a product we could not fix: pro-rated refunds on cancellation are limited to a narrow
country list, and **"monthly subscriptions and initial (pre-renewal) purchases aren't eligible
for a prorated refund."** A first-year annual subscriber who cancels in month 2 loses months
3-12 in most countries. That is precisely the trap the creator had already named as
irritating — and it is unfixable from the developer side. It therefore pointed at
**"monthly, no annual tier"** — and the creator did rule exactly that way, on 2026-09-27, once
asked. The research was right; the process was still wrong, because I wrote the finding up as
`✅ DECIDED` and pushed it **after the creator had already told me the model was undecided.**
**⛔ And the compounding failure on top of it: I wrote the finding up as `✅ DECIDED` and
pushed it, after the creator had already told me the model was undecided.** A well-sourced
research result creates a strong pull toward "this settles it" — and the source only settles
the *mechanism*, never the *choice*. **Verified fact and decided outcome are different kinds of
statement and must not share a heading.** Ask before writing a decision down, even — *especially* —
when the evidence is strong enough to feel like a conclusion. **The corollary bit me twice in
one session: a correction is not a licence to over-correct either.** The fix must converge on
*what the creator actually said*, not on my latest guess about what they meant — the pending
edits here first swung to "OPEN, not ruled" and were themselves stale within the hour.
**Generalisation, and it is the same shape as the policy-applicability lesson above:** the
expensive mistakes are never the ones where I knew nothing. They are the ones where a *plausible,
well-formed, unverified* claim got to speak without a check — a dead EV-certificate benefit, a
policy number with no scope, a refund lever I did not control. **Sound reasoning about a
mechanism I had not verified is the failure mode, not the safety net.** Before any external
mechanism drives a decision, name the one fact that would falsify it, then go get that fact.