e7cded6879
Reconcile the pending doc edits with the creator's actual ruling, and close the revenue-share question that was blocking the monthly price. Verified from the App Developer Agreement v8.10 PDF itself (downloaded and text-extracted, not a search excerpt). Section 6(b) has only three tiers: 6(b)(i) 15% for Apps and their In-App Products *not listed in* 6(b)(iii); 6(b)(ii) 12% Games-only; 6(b)(iii) 30% for Xbox console apps/games, Xbox non-subscription IAP, and Windows 8/Phone 8. LlamaCasty is a Windows PC App, so 6(b)(i) governs BOTH tiers -- there is no subscription surcharge. The agreement's own changelog (v8.0, Oct 26 2017) states it outright: "implement the 85/15 revenue share for non-Game subscriptions." => $10/mo nets $8.50 (~$102/yr); $400 lifetime nets $340. Also confirmed: the 15% applies after VAT/GST (Net Receipts definition), payouts are monthly above a $50 threshold, and no better small-indie rate exists in the standard terms. Creator ruling recorded: $10/mo subscription or $400 lifetime, no annual, no .99. This supersedes the 2026-09-21 one-time $29 -> $49 model. Two obligations ADA 6(h) attaches to the recurring tier, recorded because they change its risk profile and are the reason lifetime is the hedge: we must fulfil the subscription for the entire period as marketed (on breach Microsoft may refund the full amount plus taxes in its sole discretion), and raising the price disables auto-renew -- so $10 is effectively locked for the product's life. Stale facts fixed in the same change rather than appended: - research-store-certification.md: IARC was 11.11.1/11.11.2 in one table and 10.11.1 in another; corrected to 10.11.x. - research-store-certification.md: policy 10.8.1/10.8.2 still asserted "Polar explicitly permitted" and "tick the third-party purchase box". Void now that Store IAP is the route and Polar is being torn down. - task-48: the working YouTube demo account (10.3.1) was accidentally dropped from the certification list during the Store-IAP edit. Restored -- a reviewer cannot use the app without one. - MyMistakes.md: the verified-fact-vs-decided-outcome lesson now closes its loop (the creator did rule the way the research pointed), and records the over-correction that followed. Docs-only. No code, no build, no tests.
1180 lines
83 KiB
Markdown
1180 lines
83 KiB
Markdown
# MyMistakes.md
|
||
|
||
> **Two jobs**, distinguished by heading:
|
||
>
|
||
> 1. **Per-task failure log** — updated before every commit touching that task:
|
||
> current iteration + why the last one failed. On task complete, committed AND
|
||
> pushed → truncate to this stub. A new task does NOT seed this file until its
|
||
> first failure.
|
||
> 2. **Recipes registry** (DERIVED-SOLUTION RULE, see `AGENTS.md` 🔬) — the durable
|
||
> home for one-off derived solutions, recipes, and how-tos. The moment you work
|
||
> out a reusable solution, write it here **in the same session**. GREP THIS FILE
|
||
> FIRST when you hit a "I've done this before but have to figure it out again"
|
||
> wall. Recipe entries stay permanently (they are NOT truncated on task
|
||
> completion) — only the failure log truncates.
|
||
|
||
## 🔬 Recipes registry
|
||
|
||
### YOUTUBE STATISTICS ARE JSON STRINGS; TryGetInt64 THROWS ON STRINGS (RECIPE)
|
||
|
||
YouTube Data API v3 returns `statistics.subscriberCount / viewCount / videoCount` as
|
||
JSON **strings** (`"350"`), not numbers. (2026-09-23: the YPP parse bug surfaced the
|
||
moment the auditDetails 403 stopped masking it.) Fix = a tolerant read that checks
|
||
`ValueKind` FIRST, then `TryGetInt64` for `Number`, then `long.TryParse(GetString())`
|
||
for `String`. Trap within the fix: **`JsonElement.TryGetInt64` THROWS
|
||
`InvalidOperationException` on any non-Number token** ("requires an element of type
|
||
'Number'") — it is try-type-in, not try-catch. Guard on `value.ValueKind` before
|
||
calling it; a leaked 403→parse chain means your "whole request failed" symptom can
|
||
paper over a second crash that only appears once the 403 is fixed (test the full
|
||
happy path, not just the error path).
|
||
|
||
### liveChatId LIVES IN SNIPPET AND ONLY EXISTS ONCE THE BROADCAST IS LIVE (RECIPE)
|
||
|
||
YouTube Data API v3 `liveBroadcasts`. `liveChatId` is in **`snippet.liveChatId`** —
|
||
`contentDetails` has NO such property. AND it is only populated once the broadcast is
|
||
**live** (the official `GetLiveChatId.java` sample lists `broadcastStatus=active`);
|
||
a fetch right after insert (lifecycleStatus `ready`) returns nothing. (2026-09-25: this
|
||
cost a whole investigation — the TEST-tab "Chat polling couldn't start — Mock Chat Input
|
||
is disabled" report — because the code read `part=contentDetails` + ran before the frame
|
||
pump pushed RTMP.) Fix = read part=snippet + bounded retry AFTER the encoder starts
|
||
(`GetBroadcastLiveChatIdAsync(broadcastId, maxAttempts=10, delayMs=2000)`). Debug lens:
|
||
missing liveChatId ≈ "broadcast not live yet", NOT an auth failure.
|
||
|
||
### liveChat/MESSAGES.INSERT BODY REQUIRES snippet.type (RECIPE)
|
||
|
||
YouTube Data API v3 `liveChat/messages.insert` rejects the body with
|
||
`400 MISSING_REQUIRED_FIELD` (`domain: youtube.api.v3.LiveChatMessageInsertResponse.Error`)
|
||
unless the snippet declares `snippet.type` = `textMessageEvent` (or `pollEvent`) alongside
|
||
`liveChatId` and `textMessageDetails.messageText` — the official insert reference lists
|
||
`type` as a required property. (2026-09-25: the TEST-tab "YouTube rejected the message
|
||
(error 400)" report — the liveChatId fix in TASK 44 had worked and the drawer's Mock Chat
|
||
Input was issuing a real insert, but the body omitted `type`, and the TASK 41 test only
|
||
asserted `liveChatId` + `messageText` were present so it stayed green while real YouTube
|
||
rejected every send. Fix = add `type = "textMessageEvent"` to the body + assert it in the
|
||
Good Dog test.) Debug lens: MISSING_REQUIRED_FIELD ≠ auth/scope — it means the request
|
||
body shape is wrong, and the false-green test is the classic trap: assertion on the
|
||
body was about WHAT WE SEND, so YouTube's required fields must be mirrored in the test.
|
||
|
||
### CHANNELS.LIST auditDetails PART 403s THE WHOLE REQUEST WITHOUT A PARTNER SCOPE (RECIPE)
|
||
|
||
YouTube Data API v3 `channels.list` rejects the ENTIRE request with
|
||
`403 insufficientPermissions` if your `part=` list includes `auditDetails` but the
|
||
token lacks `https://www.googleapis.com/auth/youtubepartner-channel-audit` — the
|
||
docs' exact words: "A request that retrieves the auditDetails part for a channel
|
||
resource must provide an authorization token that contains the
|
||
youtubepartner-channel-audit scope". That scope is MCN content-partner tooling
|
||
(with a two-week token-revocation rule); normal-creator apps should NEVER ask for
|
||
it. 2026-09-23 real-log evidence: YPP refresh 403'd three times in a row while the
|
||
mock-fake tests stayed green ("current scopes suffice, no re-consent" was wrong).
|
||
Fixes: (1) drop `auditDetails` from `part=`; (2) log the response BODY — the bare
|
||
status code could not name `insufficientPermissions`, which is what made this
|
||
failure undiagnosable for days; (3) when a feature needs data no ordinary scope
|
||
grants, deep-link to the site instead of requesting the partner privilege.
|
||
|
||
### TRANSITION(COMPLETE) RACES AUTOSTOP: PRE-FLIGHT lifeCycleStatus (RECIPE)
|
||
|
||
A blind `liveBroadcasts.transition?broadcastStatus=complete` POST can 403
|
||
`invalidTransition` even though the stream just ended normally — 2026-09-22 logged
|
||
it on EVERY session end. Cause: the broadcast's own state moves toward complete
|
||
via `enableAutoStop`/YouTube auto-complete; the transition method's errors doc
|
||
(https://developers.google.com/youtube/v3/live/docs/liveBroadcasts/transition)
|
||
shows `invalidTransition` = "can't transition from its current status". A blind
|
||
POST just races — read the broadcast's `lifeCycleStatus` first
|
||
(`liveBroadcasts.list?part=status`) and only POST complete from `live`/`testing`;
|
||
skip silently otherwise and let `enableAutoStop` finish it. Never throw on the
|
||
close-out either way.
|
||
|
||
### YOUTUBE liveStreams LIST: healthStatus IS AN OBJECT, NOT A STRING (RECIPE)
|
||
|
||
YouTube Data API v3 `liveStreams.list` nests health under
|
||
`status.healthStatus = {status, lastUpdateTimeSeconds, configurationIssues[]}` —
|
||
the `healthStatus` element is an **OBJECT**, and `configurationIssues[]` lives
|
||
INSIDE it, not directly under `status`. Reading `healthStatus.GetString()` throws
|
||
System.Text.Json's `The requested operation requires an element of type 'String',
|
||
but the target element has type 'Object'` — the exact log line seen 2026-09-22 on
|
||
every live health poll (two test sessions). The parse must read
|
||
`healthStatus["status"]` (and nested `healthStatus["configurationIssues"]`).
|
||
Something like REST-shape drift is a good reason to grep the API reference
|
||
(https://developers.google.com/youtube/v3/live/docs/liveStreams) before writing
|
||
parsers against a "remembered" shape — our own flat-string fixture was the wrong
|
||
assumption the whole time.
|
||
|
||
### WINRT RESOURCE-ALLOCATION CALLS: WRAP PER-CALL, DEGRADE TO NEXT OPTION (RECIPE)
|
||
|
||
Creator callout (2026-09-15): WinRT/COM calls that allocate or start a resource — `InitializeAsync`,
|
||
`CreateFrameReaderAsync`, `StartAsync`, `CreateReaderAsync` & friends — are THROW-HEAVY. An unsupported
|
||
subtype/format, a device that just vanished, an access mode rejected mid-flight: these surface as
|
||
`ArgumentException`/`E_INVALIDARG` ("value does not fall within the expected range") or HRESULTs, NOT as
|
||
a returned status you can switch on. Relying on ONE outer catch to "handle failures" is not handling —
|
||
one rejected call inside a fallback ladder aborts the whole ladder and every untried option. Rule:
|
||
|
||
- Each allocation/start call inside a try/catch of its OWN, so a throw on candidate N falls through to
|
||
candidate N+1 (log each rejection with its message/status; collect them for the final error string).
|
||
- `return`/`break` on success must be reached WITHOUT passing through a `finally` that disposes the
|
||
resource you just committed (classic reader/capture dispose-after-commit bug).
|
||
- Unsubscribe + dispose the partial resource in the catch block when the subscription happened before
|
||
the throwing call.
|
||
- The outer catch stays as the LAST-RESORT net for device-level errors, not the primary one.
|
||
- Same discipline applies to the frame-consumption side: teardown races reader threads (see HANDOFF
|
||
crash follow-up) — a frame callback can't assume the pipeline is alive.
|
||
|
||
Applied in `MediaCaptureFrameSource`'s reader-subtype ladder (2026-09-15, webcam-take fix).
|
||
|
||
### "SAVE-RECORDING DIALOG IS CLIPPING CONTENT" — WHERE + HOW (RECIPE)
|
||
|
||
**WHERE:** the end-of-recording modal (creator: "when I end a recording the app throws up a
|
||
save-recording dialog with a default filename") is **`RenameRecordingDialog.xaml` at the repo ROOT**
|
||
(title "Rename Recording", `x:Class="ytLive.RenameRecordingDialog"`). It is NOT one of the
|
||
`*Dialog.xaml` files under `Controls/` — grep for the window title, not the filename, when the
|
||
user names a dialog by what it does. Only fix dialogs the user actually named; do not "while I'm
|
||
here" bump sibling dialogs.
|
||
|
||
**HOW:** the dialog is `ResizeMode="NoResize"` with a fixed `Height` attribute; content client area
|
||
≈ `Height − ~37px` of window chrome. Sizing/margin changes pushed content down in Z-order until it
|
||
clipped — first the file-name textbox (230 too short), and after the +20% bump (→276) the creator
|
||
found the Cancel/Save button row had ALSO been obscured. Fixed `Height` to 304 (+10% more),
|
||
`x:Name` the last interactive control (`SaveRecordingButton`) so a test can measure it.
|
||
|
||
**Verify (never eyeball-predict):** `RenameRecordingDialogSizingTests` (RealApp host) instantiates
|
||
the real dialog, `Show()` + `UpdateLayout()`, then asserts the BOTTOM EDGE of each named control
|
||
(`TransformToAncestor` against the content root → `TransformBounds(RenderSize).Bottom`) is ≤ the
|
||
client `ActualHeight`. Every such dialog fix ships with that assertion for every control that was
|
||
reported obscured. (Client-area math means the button row needs MORE slack than the textbox: its
|
||
bottom sits deeper because of the `Margin="0,20,0,0"` above it.)
|
||
|
||
### SPIN GUARD → RESOLVED — web overlay transparency + bounding box (RECIPE)
|
||
|
||
**THE ONE ROOT CAUSE THAT EXPLAINS EVERY FAILED TAKE:** WebView2's `CapturePreviewAsync`
|
||
produces an **OPAQUE** PNG. From the WebView2 spec (sender: MicrosoftEdge/WebView2Feedback
|
||
`specs/BackgroundColor.md`): "WebView will always honor a webpage's background content."
|
||
`DefaultBackgroundColor = Transparent` only shows through pages with NO background style —
|
||
the widget's own CSS paints html/body opaque. Every take below built on the false premise
|
||
"the capture has transparent margins, alpha=0"; it never did. `FindContentBounds` then had
|
||
no alpha-0 margins to find → wrong crop → black bounding box. The compositor blend saw
|
||
alpha=255 → black over webcam = "transparency broken". Same bug, three symptoms.
|
||
|
||
**THE FIX (the OBS way, applied 2026-09-08):** inject the transparency BEFORE the page
|
||
parses using `CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync` — documented to run
|
||
"before the HTML document has been parsed and before any other script included by the HTML
|
||
document is run" (learn.microsoft.com/dotnet/api/microsoft.web.webview2.core.corewebview2.addscripttoexecuteondocumentcreatedasync).
|
||
The old `ExecuteScriptAsync` on NavigationStarting/NavigationCompleted ran AFTER page
|
||
scripts/CSS → widget page repainted background opaque → lost the fight. OBS browser sources
|
||
do the same via a pre-parse user.css. Kept the nav handlers as a post-load re-assertion.
|
||
|
||
**Take timeline (the honest record):**
|
||
- `bccdb48` (Aug 28): added FindContentBounds crop — correct idea (tight bbox, no dead
|
||
space), but the capture was OPAQUE so the bbox math was built on nothing.
|
||
- `5348b5c` (Aug 28, 2 min later): reverted to full-frame no-crop — looked "good" for a
|
||
full-bleed widget, but floating widgets regained dead space ("ghost boundary").
|
||
- take-21 (`e002847`) full canvas → shrunken/offset widget (UnifomToFill of whole canvas
|
||
into a small element = lost resizing).
|
||
- take-22 (`ed9d7c1`) crop width used as canvas stride for buffer indexing → garbage.
|
||
- take-23 (`f6802c7`) stride fixed with `src.Width`; PasteKey lacked CropBounds → stale
|
||
raster cache → stale crop.
|
||
- take-24 (`081e4c1`) CropBounds in PasteKey — STILL BROKEN because the SOURCE ALPHA WAS
|
||
NEVER REAL.
|
||
- take-25 (**ME, this session — the user's "RE-INTRODUCING THE BOUNDING-BOX PROBLEM"**):
|
||
I removed FindContentBounds + the crop path entirely, betting full-canvas UniformToFill
|
||
was the answer. It WASN'T — the source is opaque-black, so the element rendered as a
|
||
SOLID BLACK BOX (the user's screenshot: "a black box in the lower right corner"). Killed
|
||
the crop → dead space returned AND black box. The compositor math was ALREADY correct;
|
||
gutting it was vandalism in response to a source-level bug.
|
||
|
||
**Rules, self-inflicted:**
|
||
1. Instrument FIRST. This session added: first-capture PNG dump of the raw WebView2 PNG
|
||
(to `%TEMP%\ytLive-web-<id>.png`) + alpha min/max/mean/%zero + FindContentBounds result
|
||
logged to startup.log once per session. That's the diff between a five-take loop and a
|
||
five-minute diagnosis.
|
||
2. When the same symptom loops across takes, the PREMISE is wrong, not the code — the
|
||
capture being transparent was the load-bearing premise and it was never verified.
|
||
3. Do not delete code paths that fix one axis (crop=bbox) while debugging another
|
||
(source alpha). Revert scope creep; keep layer contributions separable.
|
||
|
||
**2026-09-10 follow-up → NOW VERIFIED and hardened.** The 12:54 and 13:50 sessions
|
||
showed the REAL widget document capturing TRANSPARENT (widget dumps `w1..5` for
|
||
`8d7234ec`: alpha 100% zero, content 12×3; whole-file decodes of
|
||
`%TEMP%\ytLive-web-*.png` alpha max = 0) while the RECORDING kept showing a black,
|
||
opaque box over the element rect (crisp edges at the exact rect 1231,679 705×396 —
|
||
NOT the desktop showing through). Concluson: branch (a) — the OLD inline
|
||
`element.style.background='transparent'` injection only wins while the page has
|
||
nothing to paint; once the widget connects and paints its own container
|
||
background-COLOR, the capture goes opaque again → black box. The OBS-validated
|
||
answer (valid for arbitrary pages for a decade) is an injected pre-parse `<style>`
|
||
with `!important` beating every page rule:
|
||
https://obsproject.com/forum/threads/translucent-transparent-browser-source.59549/
|
||
(`body { background-color: rgba(0,0,0,0) !important }`) + the div-level variant for
|
||
stubborn widgets (woahtech.com OBS custom-CSS guide). Applied 2026-09-10 (b4bba4b):
|
||
injection appends a style element wiping `background-color:transparent!important`
|
||
on `html,body,html *`. **Second half (2026-09-10, 15:34 take): a background-COLOR wipe
|
||
is NOT enough.** The 15:34 interior ASCII shows a black void with bright content strips
|
||
at top/bottom + a right-edge bar — the widget's full-canvas CSS **background-image**
|
||
(gradient/backdrop) painted after connect. The OBS fix for that is
|
||
`background: none !important` / `background-image:none!important`
|
||
(obsproject/obs-studio#6659 — "set the CSS for html and body to background: none
|
||
!important"). Wipe both moving forward; overlay art is `<img>`/DOM and survives.
|
||
Also the dumps only covered +1s (blank-transparent); the widget's backdrop paint
|
||
arrives later, so the mid-recording dump set is now re-armed ~30s in. Verify take:
|
||
element rect shows scene bg behind the widget art (animation, no black void) → loop
|
||
closes. Self-inflicted again: the agent re-derived the whole transparency story
|
||
(recording pixel archaeology, compositor blend re-verification) instead of reading
|
||
this entry — the instrument said transparent because the capture had NOT been
|
||
repainted yet. Do NOT re-derive this story a third time.
|
||
|
||
**2026-09-12 → RESOLVED — the box lived in the paste-cache RASTER, not the page.**
|
||
The 15:51 facts (alpha max=255, mean ~19, zero 57%, white rounded panel, real
|
||
transparent margins) + preview-correct + recording-box meant the capture transparency was
|
||
REAL all along; the fracture sat in `SceneCompositor.BlitContentRaw`
|
||
(SceneCompositor.cs:419-550) — the sampler that builds the element-space paste-cache
|
||
raster, the ONE path the take loop never read in full. Its partial-alpha branch applied
|
||
the OPAQUE-dst blend onto a TRANSPARENT raster base: `dst = (src*sa + dst*inv)/255`
|
||
with dst black → color PREMULTIPLIED by sa, then `dst[+3] = 255` — a 50%-alpha widget
|
||
pixel became darkened color + FULL alpha. At paste time `BlendRowOpaque` saw alpha 255 →
|
||
straight copy → the scene behind was overwritten by darkened ink. Fully-transparent
|
||
margins (alpha 0, `continue`) and fully-opaque content (alpha 255 branch) survived —
|
||
which is why every take showed a box while the dumps and the preview (raw WriteableBitmap,
|
||
unaffected by the compositor) stayed correct, and why the CSS-wipe fixes (`b4bba4b` /
|
||
`0f72c53`) were red herrings: they treated the page as the villain, but the capture was
|
||
transparent from the start. **FIX:** `BlitContentRaw` gained `transparentDst=false`;
|
||
the raster call site passes `true` and writes STRAIGHT color + straight alpha so the
|
||
paste rows (`BlendRowOpaque`/`BlendRowWeighted`) do the real source-over onto the opaque
|
||
master. The master paths are untouched (bit-identical). ONE test
|
||
`PasteCache_SemiTransparentLayer_RevealsBackdrop_NotOpaqueInk` fails on the old code with
|
||
exactly the bug encoded: 50%-blue over red reads `(0,0,128)` instead of `(127,0,128)` —
|
||
backdrop never shows through. Next: verify take (element rect shows the backdrop behind
|
||
the widget art), then push gate #1 clears.
|
||
|
||
**2026-09-12 → COLLATERAL — the SAME transparency saga broke the webcam next.**
|
||
Verify-take of the transparency fix showed a perfect flat gray rectangle where the webcam
|
||
should be (preview fine, recording flat — stdev <1 across the whole element, despite the
|
||
raw camera frame handed to the compositor sampling min=0/max=255 one line earlier in the
|
||
pipeline — proven by instrumenting BOTH ends before touching any code, not guessing).
|
||
Root cause: `ed9d7c1` ("CropBounds metadata + Fill-style scaling... widget fills element
|
||
rect" — take-22 of the SAME transparency saga above) added `cbX/cbY/cbW/cbH` to
|
||
`BlitContentRaw`'s general sampler but only assigned them inside the
|
||
`src.CropBounds is {} cb` branch; every CropBounds-LESS source (webcam, images — anything
|
||
but the web widget) fell through with `cbW=cbH=0`. Since `sxCrop = sxNorm * cbW` and
|
||
`syCrop = syNorm * cbH`, both were always 0, so `sxCanvas`/`syCanvas` collapsed to
|
||
`cbX`/`cbY` = (0,0) for every destination pixel — an entire scaled webcam sampled ONE
|
||
source corner pixel. The web widget (the only CropBounds-bearing source) was never
|
||
affected, which is exactly why the transparency fix's own test suite stayed green while
|
||
this broke. **FIX:** default `cbW = src.Width`, `cbH = src.Height` (cbX/cbY = 0) before
|
||
the branch, only overridden when CropBounds is actually present. ONE test
|
||
`BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin` (two-color split
|
||
source, scaled non-1:1 through the paste cache) fails on the old code — right-half pixel
|
||
reads the left half's color — and passes on the fix; proven both ways with a stash/build/
|
||
revert cycle, not by inspection alone.
|
||
**Lesson:** when a shared low-level sampler gains a new optional code path (CropBounds),
|
||
audit EVERY variable the new branch introduces for a safe default in the branch it did
|
||
NOT touch — an uninitialized-to-zero "crop region" silently means "sample only pixel
|
||
(0,0)", not "no crop." grep for this shape (`var x = 0;` followed by an `if (cond) x = ...`
|
||
with no `else`) whenever a conditional metadata field is added to pixel math.
|
||
|
||
### Shrink / re-encode an image for the README (screenshots → small hero image)
|
||
|
||
Worked out 2026-08-29 (the recipe was NEVER recorded the first time it was done, so
|
||
it had to be re-derived from scratch — that's the incident this entry exists to end).
|
||
|
||
**Approach:** a throwaway Windows-dotnet console app uses WPF's imaging stack
|
||
(`System.Windows.Media.Imaging`) — same framework the app runs on, zero NuGet
|
||
packages, high-quality downscale via `TransformedBitmap`. Screenshots compress
|
||
**far smaller as JPEG than PNG** (PNG 1400px = ~1.2MB; JPEG q82 1400px = ~188KB).
|
||
|
||
**Recipe (run via the Windows dotnet host from WSL):**
|
||
|
||
1. Create `imgresize.csproj` targeting `net8.0-windows` with `<UseWPF>true</UseWPF>`
|
||
(SDK controller). Put it in a Windows-visible temp path, e.g.
|
||
`C:\Users\gramp\AppData\Local\Temp\imgresize` — NOT `/tmp` (Windows dotnet can't
|
||
reach a Linux-only path reliably).
|
||
2. `Program.cs`: load `BitmapImage` (`CacheOption=OnLoad` → `Freeze()`), downscale
|
||
with `TransformedBitmap(src, new ScaleTransform(scale, scale))` to max width
|
||
(1400 for the README hero), encode with `JpegBitmapEncoder { QualityLevel = 82 }`,
|
||
save.
|
||
3. Run:
|
||
```bash
|
||
"/mnt/c/Program Files/dotnet/dotnet.exe" run -c Release --project .
|
||
-- "C:\Users\gramp\Downloads\Screenshot 2026-08-29 075626.png"
|
||
"C:\Users\gramp\Documents\Code\projects\ytLive\docs\ytLlive-preview.jpg" 1400
|
||
```
|
||
4. Point `README.md` at the `.jpg` (not `.png`).
|
||
|
||
**Result:** 3.2MB screenshot → 1400×794 → **188KB** `ytLlive-preview.jpg` in `docs/`.
|
||
|
||
---
|
||
|
||
### Measuring audio-video A/V sync from a recording (no ears needed) — RECIPE
|
||
|
||
Derived 2026-09-12/14 (the "audio cut off at the end" / "audio delayed" reports). You
|
||
cannot "listen" to a take — measure it. Proven on two independent takes (talking+claps,
|
||
clean-clap test) with a tight, matching result each time.
|
||
|
||
**Method (FFmpeg probe + numpy, run from WSL):**
|
||
|
||
1. **Click/hiss-SILENT test takes are the gold standard.** Have the creator do a
|
||
loud, single-frame-syncable event (one clap after ~10s of near-silence) — the
|
||
video position of the hands-meet peak and the audio position of the transient
|
||
are both unambiguous. That single event replaced counting words forever.
|
||
2. **Extract audio as raw mono PCM** (`ffmpeg -vn -ac 1 -ar 48000 -f s16le`) and
|
||
compute a short window RMS envelope (5ms windows) in numpy. The clap is the global
|
||
max (`argmax`); also print a coarse 100ms table — it shows the pre-clap artifacts
|
||
(faint blips) vs the event (100× larger) vs true digital silence.
|
||
3. **Video motion per frame** — `-vf "tblend=all_mode=difference,signalstats,metadata=print"`
|
||
captured to stdout (NOT `file=` inside the filter — the `\\` path escapes mangle;
|
||
have PowerShell-style quoting bite; `metadata=print` writes to ffmpeg's log, so
|
||
redirect stdout to a file). Parse `YAVG` values; the clap is the frame with the
|
||
motion spike (2.75 vs a ~0.1–0.5 animated-widget background).
|
||
4. **Offset = audio_event_time − video_event_time**; every positive second means audio
|
||
is BEHIND video by that much. Cross-correlate the two full envelopes (60fps-resampled
|
||
RMS vs YAVG, normalized) to double-check the single-event peak — the autocorr-style
|
||
peak must be sharp (unique max at +133 frames, second-best ≈0.26).
|
||
5. **Sanity checks baked in:** count frames (`nb_frames` vs wall log) — if video is
|
||
real-time you've ruled out the truncation bug as the cause; compare stream
|
||
durations (audio > video by the lag is the SIGNATURE of a delayed audio tail, not
|
||
truncation); check `dropped writes` — dropped audio moves events EARLIER (opposite
|
||
sign), so it can never explain a lag. **A constant offset ≠ drift:** fixed lag =
|
||
buffering/backlog, growing offset = clock mismatch.
|
||
|
||
**The ROOT CAUSE the recipe led to (record it so it's never re-derived):** the mixer's
|
||
capture rings are fed from APP STARTUP (for the meters); `StartLive` never clears them,
|
||
so the drained audio trail begins ~a full ring-depth (2.0s) behind go-live. **Rule: any
|
||
"live preview/drain" sink fed by a continuously-capturing buffer MUST clear the buffer
|
||
at go-live, or early output is stale backlog.** The 2s ring + configured 300ms sync
|
||
delay + ~100ms pipeline = exactly the measured 2.11–2.22s. The lag ALSO scales with
|
||
"time since app launch" up to the ring depth — a take right after a restart reads as
|
||
~370ms while a fully-warmed app reads ~2.2s. Same code, wildly different numbers.
|
||
|
||
### Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
|
||
|
||
Derived 2026-09-03 (take-3 diagnosis — the stats seam from `97ffc42` named the stage
|
||
in one line: `17/300 frames per 5s, avg render 258.1ms, avg submit 1.5ms`).
|
||
Both halves were solved by OBS/libyuv long ago; do not re-derive:
|
||
|
||
1. **Pacing is a DEADLINE, never a post-render sleep.** `sleep(interval)` after each
|
||
frame makes the period `render + submit + interval` — the producer can hit ≤ half
|
||
the declared rate even with a free render. OBS's `video_thread`
|
||
(`libobs/media-io/video-io.c`) advances an absolute `nextTick += intervalTicks` and
|
||
sleeps only the remainder. **NEVER "rebase" the deadline to wall-now when it blows** —
|
||
the take-3 note I originally wrote here ("skip … the missed ticks (rebase)") was the
|
||
PROVEN-wrong advice: resetting `nextTick` erases the missed slots, so a 43fps reality
|
||
was authored into a 60fps container and every recording played ~1.4x fast (take 12/13,
|
||
2026-09-10). The rebase even kept the stats looking honest (render+wait == period, no
|
||
loss) because a rebased frame is never "late". OBS keeps the counter MOVING and outputs
|
||
ONE frame per interval tick — a late render shows as REPEATED footage (judder; duration
|
||
stays correct), never a skip and never a wall-now reset. Critical with rawvideo:
|
||
pts is stamped by ARRIVAL, so the muxed duration is pure frame count ÷ fps — the only
|
||
way to make duration == wall time under ANY load is exactly one frame per interval slot.
|
||
Corollary: **never add a second pacer.** `ffmpeg -re` on the rawvideo input throttles
|
||
the pipe independently and fights the pump (its "Resumed reading … after a lag" grew
|
||
0.79s→4.82s in the same take). Removed; the pump IS the pacer.
|
||
2. **A 1080p frame is ~2.07M pixels — the hot path must be row-simple.** Per-pixel
|
||
`Math.Round` + float source-over in managed code costs ~100ns/px = the whole 258ms.
|
||
libyuv's pattern (https://chromium.googlesource.com/libyuv/libyuv/): branch per
|
||
pixel on source alpha (opaque → 4-byte copy, transparent → skip), integer
|
||
fixed-point blend `(s*a + d*(255-a) + 127)/255` otherwise; and ALWAYS clip the loop
|
||
to the intersection rect (our social-bar overlay scanned all 2M dst px for a 64px
|
||
strip). A full-cover 1:1 blit also obsoletes the opaque-black pre-fill — skip dead
|
||
writes.
|
||
3. **On Windows, `Task.Delay` is a 15.6ms QUANTUM, not a timer.** Any request under one
|
||
system-clock tick sleeps a full tick (documented — learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay:
|
||
"approximately 15 milliseconds on Windows systems"). A deadline pacer built on Task.Delay caps the
|
||
producer at ~40fps-ish EVEN IF render is instant — take 9 proved the signature: work fell 26.5→22.4ms
|
||
but the period sat at ~37ms (≈ one padded wait/frame), so two real optimizations read as "zero change".
|
||
Frame-accurate loops (OBS/Chromium/game-loop canon — stackoverflow.com/questions/5441464) do:
|
||
`timeBeginPeriod(1)` for the session (paired with `timeEndPeriod`), sleep only the BULK of the
|
||
remainder, SPIN the last ~2ms across the deadline. Diagnostic before touching the compositor again:
|
||
period ≈ work + 15.6 → the SLEEP is the bug, not the work.
|
||
|
||
Related (take 14, 2026-09-04): **recycled ring buffers are a race you must SIZE, not just own.**
|
||
Deepening shared frames to kill GC churn (a fresh 8.3MB/tick array) hands out REUSED memory — the
|
||
ring's depth × source period must EXCEED the worst consumer hold (compositor read + lagged UI
|
||
preview copy), not just "a few frames". 4 slots at 144Hz capture laps in ~27ms vs a ≤50ms read: half
|
||
a new screen frame flashed over an old one in the recording ("bits flashing over other bits").
|
||
Depth 8 everywhere (screen/camera/web output rings); the paste-cache Epoch still guards identity.
|
||
|
||
4. **Hermetic pacing test:** inject the delay seam to RECORD the requested TimeSpan and
|
||
genuinely await it (`Task.Delay(d, ct)`) — a fake that returns
|
||
`Task.CompletedTask` synchronously makes the whole pump loop run on `StartAsync`'s
|
||
sync continuation and hang the test run (hit this 2026-09-03; the existing fakes all
|
||
yield for exactly this reason). Assert the REQUESTED wait (< interval with a
|
||
≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.
|
||
5. **Expensive content: raster on change, never on read (take 5, 2026-09-04).** A source
|
||
that updates once a minute (chat text!) must not full-rasterize (`FormattedText` +
|
||
`RenderTargetBitmap` + `CopyPixels` ≈ 15-25ms) every compositor tick. OBS text sources
|
||
re-render on property/message change; the per-tick pass blits the cache. Implement as:
|
||
content version (collection-changed counter) + config key (size/appearance) → cached
|
||
immutable `VideoFrame` returned by identity. Gotcha: buffers that SURVIVE sessions
|
||
(the chat log) silently arm the per-tick cost even in flows that never touch the
|
||
feature (signed-out record-only takes paid chat rendering!).
|
||
|
||
6. **An async loop started from a UI handler runs ON THE UI THREAD until you take it off.**
|
||
`await` continuations re-capture the current `SynchronizationContext` — the frame pump was started
|
||
from a WPF command handler, so the "WPF-free, hermetic" compositor rendered and read capture state
|
||
ON THE DISPATCHER, serialized behind the live preview itself, for the whole starvation saga. The
|
||
`wait` stat caught it only when the numbers became self-contradictory (render 22 + wait 10 > any
|
||
rebasing deadline — a blown deadline cannot sleep). OBS runs `obs_graphics_thread`/`video_thread`
|
||
as dedicated threads for exactly this reason. Pattern: `_task = Task.Run(() => Loop())` (null
|
||
context inside), then audit EVERY object the loop touches for UI affinity (RenderTargetBitmap /
|
||
DrawingVisual / WriteableBitmap: marshal the work or the rare miss; plain locked byte[] lookups:
|
||
fine) and pin it with a context test (`Pump_Produces_OffTheStartingContext`, inline-pumping
|
||
SynchronizationContext that the old code failed by construction). Cost: takes 3–10.
|
||
|
||
7. **Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8).**
|
||
The chat raster fix was REAL but the composer blamed it for the residual slowness it did not
|
||
own; two takes burned before the render/resolve split showed `resolve ≈ 0` and pointed at the
|
||
compositor pasting static layers per tick (`BlitCachedLayer` finished the job OBS-style). Before
|
||
shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the
|
||
NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.
|
||
|
||
8. **An encoder refed from a real-time loop must ENQUEUE, never pipe-write in the loop
|
||
(2026-09-10, take 15/16 — the "1...23...4...56..." smeared ticker).** After slice 9 the
|
||
recording played at the right DURATION but the content still hiccuped — and the aggregates
|
||
(301/300, uniform file PTS, 15.6s wall vs 15.74s file) could NOT see it. The cause finally
|
||
measured in `FfmpegEncoder.SubmitFrameAsync`: the `WriteAsync(8.3MB)+FlushAsync` to ffmpeg's
|
||
stdin BLOCKS whenever the encoder lags the pipe, and the slice-9 burst `while (now>=nextTick)`
|
||
then re-wrote that SAME stale composite for every slot that ticked past — frozen content runs.
|
||
OBS's answering machinery is the encoder queue (`libobs/obs-encoder.c`): the encoder thread
|
||
NEVER couples back into the video thread; overflow = dropped data, NEVER a frozen producer.
|
||
Fixed as: bounded `Channel<byte[]>` (cap 120) + a dedicated drain task owning stdin,
|
||
`SubmitFrameAsync` = copy-to-pool-array + `TryWrite` (drop-newest + count when full), stop
|
||
flushes the queue then EOF. **Rule: verify with a clock-independent judge.** The WSL ticker that
|
||
"proved" slice 9 has its
|
||
own Host-timer jitter under Windows load — so this slice burns a dot-matrix `_outputIndex`
|
||
into the bottom-right of every composite; decoding the recording reads the honest sequence
|
||
(+1/frame, jumps = counted drops) with no external clock involved. Take 17 must read
|
||
+1/frame from that strip. (A whole-frame duplicate scan was tried and is
|
||
UNRELIABLE here: the scene is always animating — session elapsed timer + REC pulse — so
|
||
no two frames are ever byte-identical.)
|
||
|
||
**SUPERSEDED by slice 15 (2026-09-14):** the last clause "pump emits ONE fresh composite per
|
||
iteration (no burst re-write)" had it BACKWARDS for the deadline. Slice 10 chose a skip when
|
||
the render overruns — an overrun's slots vanish from the file — and that AUTHORS ACCELERATED
|
||
playback (my item-1 lesson already said it: "NEVER a skip … a late render shows as REPEATED
|
||
footage"). Device takes proved it: one fresh frame per 35ms stall → `ty-20260914-1726` = 163
|
||
video frames (2.72s) against 2.93s of audio, video ending 0.22s early ("audio speeds up then
|
||
cuts off"). OBS's answer (docs.obsproject.com/backend-design: "If the video frame queue is
|
||
full, it will duplicate the last frame"; `libobs/obs-output.c` counts "lagged frames due to
|
||
rendering lag/stalls" — never a time-hole) is to fill each missed slot by DUPLICATING the
|
||
newest frame; duration == wall, judder not fast-forward. That's what the pump's submit now
|
||
does: a catch-up loop over the missed slots emitting this iteration's composite again, clamped
|
||
to a `deadlineNow` captured once (bounded — the smear that justified slice 10 was the
|
||
BLOCKING pipe-write re-copying during a long freeze; the queue makes each emit nanoseconds,
|
||
so the burst is safe). The burned index moved inside the submit loop: EVERY emitted slot
|
||
carries its own +1 (this also fixed the old unconditional `_outputIndex++` that gapped the
|
||
sequence on non-submitting fast-render iterations). A/V sync must be re-clap-measured after
|
||
this fix — the +0.6s audio-late reading on 1726 was confounded by the 1.1x acceleration.
|
||
|
||
**Latest A/V sync numbers (2026-09-14, pre-slice-15 pacing):** keep the recipe below honest on
|
||
clap measures. ty-1723: 697 video frames (60fps) = 11.62s vs 11.84s audio. ty-1726 (clap): 163
|
||
frames = 2.72s vs 2.93s audio; clap audio env peak 1.655s vs video motion cluster 0.75–1.03s →
|
||
single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late) — BOTH confounded by
|
||
the acceleration; re-measure after slice 15.
|
||
|
||
|
||
**Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
|
||
the freeze-audit recipe (RECIPE — re-deriving it cost this session):**
|
||
The slice-15 build fixed pacing but the desktop layer of the recording was still
|
||
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
|
||
measurement said something different AND worse than the running render theory.
|
||
`FramePump stall… worst render 33-36ms` was real but MOOT: once the camera+desktop
|
||
take was decoded to raw frames, the **desktop band was frozen 21s of 23.35s (90%)**
|
||
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
|
||
CONVERSION was the wall: the monitor delivers at the **240Hz DWM cadence**, the
|
||
source converts ONE frame at a time (`_framePending` latest-wins), and each
|
||
2560×1440→1920×1080 `DownscaleBgra` — naive double-per-pixel bilinear — cost
|
||
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
|
||
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
|
||
MediaCapture path) and audio were fine — exactly what the user reported.
|
||
|
||
**The audit recipe (ffmpeg → raw gray → numpy):**
|
||
```
|
||
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
|
||
python3 - <<EOF
|
||
import numpy as np
|
||
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
|
||
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
|
||
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
|
||
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
|
||
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
|
||
EOF
|
||
```
|
||
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
|
||
the compositor render stats led nowhere until this number existed. A **per-row
|
||
split detector** (`cumsum` of per-row diff-to-next minus diff-to-prev, argmax =
|
||
split row) then separated real mid-frame tears (score ≈ huge, both halves match
|
||
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
|
||
were the session UI, not tears.
|
||
|
||
**The fix (this slice):**
|
||
- **integer 8.8 fixed-point downscale, "shift only at the end"** — the SAME math as
|
||
`SceneCompositor.Bilinear` (rounded both stages in one 16.8 scale) ported into
|
||
`DownscaleBgra`, dropping per-pixel doubles to row-walk integer ops. The capture
|
||
ring already had the integer-bilinear lesson; the capture downscale itself was
|
||
still the naive float twin of the 258ms disaster.
|
||
- **throttle to the slot cadence** (`MinConvertInterval = 10ms`): the 240Hz arrival
|
||
is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per
|
||
second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps
|
||
pump can use.
|
||
- **ring reuse-distance, not ownership** (`FrameRingBuffer`, redLine 4): the
|
||
take-14 "depth × period" rule guards size; the slice-16 addition makes it
|
||
structural — a slot is only rewritten ≥4 rents after its last hand-out, else a
|
||
fresh buffer. `session.LatestFrame` survives across conversions and the
|
||
dispatcher preview copy lags, so "who released it" is unknowable without a
|
||
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
|
||
tear (new-top/old-bottom midway) is that read-under-write closed.
|
||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||
relocating a 30ms float downscale from the capture thread to the render thread
|
||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||
op fast.
|
||
|
||
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
|
||
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
|
||
bug):**
|
||
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
|
||
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
|
||
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
|
||
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
|
||
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
|
||
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
|
||
follow-on decisions fixed by evidence:
|
||
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
|
||
the pool to the master size to read back less. Microsoft's screen-capture page
|
||
forbids it: "the underlying Direct3D surface is always the size specified when
|
||
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
|
||
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
|
||
not scale — silently encode the desktop cut off. Readback must stay native; the
|
||
lever is concurrency, not size. **Rule: read the platform doc for the exact
|
||
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
|
||
have been wrong twice now.**
|
||
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
|
||
flight, completions can land out of order; a slow OLDER readback finishing last
|
||
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
|
||
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
|
||
compare-exchange publish) drops stale completions instead. Ring gets a lock because
|
||
rents are now concurrent; downscale row-scratch became per-conversion locals.
|
||
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
|
||
choosing between throughput and latency fixes — both previous slices picked the
|
||
wrong slot ("render" vs "conversion") until the audit existed.
|
||
|
||
---
|
||
|
||
|
||
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
|
||
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
|
||
integer bilinear) followed the same libyuv research
|
||
(https://chromium.googlesource.com/libyuv/libyuv/ — `row.cc`/`scale.cc` keep both
|
||
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
|
||
My first `Bilinear` shifted stage 1 back to 8-bit AND shifted the final result >>16 —
|
||
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
|
||
NOTHING (green webcam silently vanished from output; the pixel probes caught what
|
||
the eye in a 2x time-lapse would not). **Rule: multi-stage fixed point shifts only
|
||
at the end; verify against a uniform-255 sample before believing it.** Second trap:
|
||
a stale-byte sentinel test whose source pattern can generate the sentinel value
|
||
itself (0xAB was a legitimate `x+y` pixel) — pick the sentinel coprime/out-of-range
|
||
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
|
||
that HOLDS submitted frames now must snapshot them (`Clone`) once the producer
|
||
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
|
||
|
||
---
|
||
|
||
|
||
### A web widget captured at 10Hz inside a 60fps recording plays at ~1/6 speed
|
||
|
||
Derived 2026-09-10 (the "widget animation too slow" report). The recording is 60fps and the
|
||
WebView2 capture loop was a blind 100ms `DispatcherTimer` = 10Hz — each captured web frame gets
|
||
repeated ~6× in the file, so whatever the page animates at, the OUTPUT is capped at the *capture*
|
||
cadence, not the page's. Two stacked throttles, both real:
|
||
|
||
1. **The capture rate is the hard ceiling.** web-layer motion in the recording can never beat
|
||
`CapturePreviewAsync` frequency. But you can't just raise the timer: full-HD PNG capture costs
|
||
~10-30ms (WebView2Feedback#20: "CapturePreviewAsync … produces PNG/JPG and is very slow"), so
|
||
concurrent captures stack CPU AND can publish stale-after-fresh. The mandatory shape is a
|
||
**latest-wins drop**: at most one capture in flight per session; a tick during the in-flight
|
||
window is DROPPED, never queued; effective cadence = max(interval, capture duration). Do this
|
||
before anyone touches the interval constant.
|
||
2. **Chromium throttles hidden pages.** An off-screen WebView2 (we place it at (-5000,-5000)) is a
|
||
hidden page the moment the host window is unfocused or covered: `requestAnimationFrame` parks and
|
||
JS timers clamp to ~1s (WebView2Feedback#1172 — background-throttled rAF; #3070 — a WebView2 with
|
||
`Visibility.Collapsed` slows timers to 1s; Chrome-88 blog — heavy timer throttling of hidden tabs).
|
||
There is NO supported per-page opt-out (#5250 still open). The embedder answer is browser args on
|
||
the `CoreWebView2EnvironmentOptions` created BEFORE `EnsureCoreWebView2Async`:
|
||
`--disable-backgrounding-occluded-windows --disable-renderer-backgrounding
|
||
--disable-features=CalculateNativeWinOcclusion` (the Electron/Streamlabs-class fix for
|
||
occluded-window animation throttling). Share ONE environment across sessions (one browser process).
|
||
3. **Measure before picking the cadence.** Log the first ~30 captures' elapsed ms on the first
|
||
recording run; PNG encode + WPF decode of 1920×1080 is the per-capture cost that decides whether
|
||
~30Hz is affordable or it must drop to ~20Hz. The FramePump drops frames (never time-lapses) if
|
||
UI-thread GC churn starves it, so cost shows up as dropped-frame stats — read them.
|
||
|
||
|
||
---
|
||
|
||
|
||
## Splitting a large file into partials — NEVER `awk … > SRC` while awking SRC
|
||
|
||
(2026-08-31, Commit D) Tried to split `SocialsDialogViewModel.cs` in one line:
|
||
`{ awk '…' SRC; echo ""; awk '…2…' SRC; } > SRC`. The **first write truncated
|
||
SRC to 1 line**, so the second `awk` read the already-truncated file → the whole
|
||
source was lost (1 line left). Recovered with `git checkout -- SRC`, then redid
|
||
it, but the same bug could have meant making it up from scratch.
|
||
|
||
**Rule:** when a cut needs N blocks from one source into N files, never write a
|
||
block back onto the source that the `awk`s still read. Instead:
|
||
1. Read the source **once** at the start into temp files (`mktemp -d`, one file
|
||
per block), with a `$D` variable you carry forward.
|
||
2. Verify block sizes (`wc -l`) and brace balance (`python3 -c` counting `{`/`}`)
|
||
before touching any real file.
|
||
3. Then assemble each new file from `cat D/block …` — never truncating the source
|
||
until every read is done.
|
||
|
||
`git status` can't save you here if you don't notice until the file is gone —
|
||
`git checkout -- <path>` from the last commit is the recovery. Cheap insurance:
|
||
restore-then-retry, do it atomically from temp files the first time.
|
||
|
||
---
|
||
|
||
## Verifying an ffmpeg decode contract from WSL (no real CLR needed)
|
||
|
||
(2026-08-31, TASK 21) When a change depends on ffmpeg producing output with an
|
||
exact frame-size contract (rawvideo W×H×4 BGRA), you can prove the **command +
|
||
frame accounting** here without any .NET process:
|
||
|
||
1. Fetch a **static Linux ffmpeg** into `/tmp/opencode` (no sudo needed):
|
||
`curl -sLO https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz
|
||
&& tar -xf …`
|
||
2. Generate a tiny known clip: `ffmpeg -f lavfi -i "testsrc2=duration=1:size=640x360:rate=30" -pix_fmt yuv420p clip.mp4`
|
||
3. Decode with **exactly the app's args**: `-f rawvideo -pix_fmt bgra -vf scale=640:360 -an`
|
||
4. Assert `total_bytes % (W*H*4) == 0` (python3) → exact integer frames, no pad.
|
||
|
||
**Why not a dotnet-spawned ffmpeg here:** the only CLR on this box
|
||
(`/home/gramps/bin/dotnet`) is a **Windows-bound shim** — `Process.Start` resolves
|
||
paths to `\\wsl.localhost\Debian\…` and throws "not a valid application for this
|
||
OS" when handed a Linux ELF ffmpeg. So never plan to have dotnet exec a Linux
|
||
ffmpeg here; verify the contract with shell/python instead, and leave the
|
||
CLR→real-ffmpeg run to the native Windows suite.
|
||
|
||
## WPF hit-test truth in tests: `UIElement.InputHitTest`, NOT `VisualTreeHelper.HitTest`
|
||
|
||
(2026-09-01, the RoundClip "known failure" post-mortem — a failure the map carried as
|
||
"not a regression" for weeks without ever recording WHY.)
|
||
|
||
**The trap:** `VisualTreeHelper.HitTest(window, pt)` returned the window's
|
||
`WebViewHostPanel` overlay (`IsHitTestVisible="False"`, `Opacity=0`, ZERO children — since the
|
||
2026-09-14 composition-capture redesign the overlay is DELETED, but the API lesson stands) for
|
||
EVERY point in the window — so a "corner is grabbable" assertion could never pass, and
|
||
it looked like a real interaction bug. The actual input pipeline (`UIElement.InputHitTest`,
|
||
what Mouse routing uses) correctly returned the element's Grid at elem-center/corner-in/
|
||
corner-exact and fell through to CanvasGrid just past the corner. The product was fine;
|
||
the TEST was probing an API that doesn't model input semantics.
|
||
|
||
**Rule:** any test asserting "where does a click land" uses `window.InputHitTest(pt)` +
|
||
`IsDescendantOf` — never `VisualTreeHelper.HitTest`.
|
||
|
||
**Diagnosis recipe (how the lie was caught in ~3 probe cycles, no guessing):** add a TEMP
|
||
probe `[Fact]` in the RealApp collection that hit-tests a spread of points
|
||
(elem-center / corner-in / corner-exact / corner-out / bg-center) and `Assert.Fail`s with a
|
||
composed dump: per-point VTH hit + `InputHitTest` hit + ancestor chain (`GetParent` walk
|
||
with `#Name`) + panel properties (`IsHitTestVisible/Opacity/children/actual size`) +
|
||
`TranslatePoint` origins. Run the class alone, read the message, delete the probe.
|
||
|
||
**Two sibling facts learned the same session (record-once):**
|
||
1. A UserControl owns its own XAML namescope — after extracting a region out of a window,
|
||
`window.FindName("InnerPart")` returns null; resolve the UserControl by its window-level
|
||
name, then `pane.FindName("InnerPart")`. And window-scope STYLES are invisible to a
|
||
UserControl's `StaticResource` at parse time — move such styles to `Themes/Controls.xaml`
|
||
(the app-scope rule exists for this).
|
||
2. Per-class `dotnet.exe vstest` from WSL DOES execute the RealApp/`MainWindow` tests fine
|
||
(they passed natively 2026-09-01) — only the FULL suite hangs (WASAPI startup). And the
|
||
test process shares `%APPDATA%\ytLlive\startup.log` with the real app: lines like
|
||
`camera 'test-camera' failed` are test noise, not DB state — to check pollution, query
|
||
the DB directly (`python3 sqlite3`, `SELECT DeviceId FROM Webcam`), not the log.
|
||
|
||
## A "known failure" label without a recorded cause = a bug on life support
|
||
|
||
(2026-09-01, the audio triple-take) One line — `_delayedMix` (nullable, added by TASK 22,
|
||
never initialized) dereferenced as `delayed.Length` — produced THREE symptoms that lived in
|
||
the map as two separate "pre-existing, do-not-chase" entries: (a)
|
||
`Mix_HonorsProviderGains…` "known failure", (b) `AudioPipelineTests` hangs when run at all
|
||
(a test reading a named pipe with no writer blocks — hung test ≠ flaky test, it's a starved
|
||
producer), (c) startup.log flooded "Audio live loop error" every 10ms (the mixer loop caught
|
||
and logged only `ex.Message` — stack thrown away).
|
||
|
||
Rules derived:
|
||
1. NEVER label a test "known/pre-existing" without writing WHY (exception type + first app
|
||
frame). An unexplained known-failure is deferred archaeology that hardens into fog.
|
||
2. A test that waits on IPC + a producer whose output vanished are usually ONE bug — look
|
||
for the producer before blaming the test.
|
||
3. Catch-and-log-swallow of `ex.Message` hides root causes; log with stack (`AppLog.Write(ex,
|
||
...)`) and throttle (5s) instead of dropping or flooding.
|
||
Bonus: the "failing" test encoded the MAP's contract (`loopbackGain = GameAudioVolume`); the
|
||
code had drifted to unity on a disproven premise (loopback capture does NOT follow endpoint
|
||
volume — creator's 20%-volume/pegged-meter observation killed it). The test was right all
|
||
along — failing tests may be the last honest witnesses; interrogate, don't pardon.
|
||
|
||
## Feature provenance: record WHO asked and WHY, in the task entry itself
|
||
|
||
(2026-08-31/09-01, the SYNC slider scare) TASK 22's lip-sync slider surfaced on the preview rail and
|
||
the creator's reaction was "totally don't remember ordering that" — because the queue entry recorded
|
||
WHAT shipped (a slider, 0-500ms, a converter class) but not WHO asked (the creator, explicitly, for
|
||
OBS's delay-filter fix built natively). Eight days later his own request read like AI drift and nearly
|
||
got deleted. Rule: the moment a creator-driven feature is queued or shipped, its entry carries a
|
||
one-line provenance — *who asked, what triggered it* ("creator: OBS delay-filter lip-sync fix, native").
|
||
Features without attribution become roadmap orphans that get punted, removed, or re-litigated. Same
|
||
disease as an unexplained "known failure" label — a fact recorded without its reason is a future
|
||
argument.
|
||
|
||
## An un-attributed build invalidated three takes of a perf saga — stamp the binary
|
||
|
||
(2026-09-04, takes 4–6) After each render-perf fix the creator "exed the code" and re-recorded, but
|
||
the exe timestamp ≠ binary contents (incremental builds reuse whatever compiles clean; a source edit
|
||
with no rebuild serves the OLD exe). Take 6 measured render WORSE than take 5 (35-41ms) and there was
|
||
no honest way to tell "the chat cache fix doesn't work" from "the fix was never running" — three
|
||
hours of diagnosis on an unattributable sample. Rule: if takes measure the app, EVERY build carries
|
||
an id and EVERY log line traces to it — `GenerateBuildStamp` (csproj) writes a fresh GUID per
|
||
compile (deliberately defeating incremental lies), the wordmark shows it as a superscript, startup.log
|
||
records `Build <id> (compiled <time>)`. A perf claim without build attribution is a guess; ask for the
|
||
stamp BEFORE theorizing. (Also this session: a sentinel-byte test where the source pattern could
|
||
GENERATE the sentinel, and a fixed-point bilinear that shifted BOTH stages and silently drew nothing —
|
||
see the rawvideo recipe.)
|
||
|
||
## Signed A/V sync: negative offset = eat the buffer head (OBS semantics), armed at go-live
|
||
|
||
(2026-09-14) The ring-backlog fix shrank A/V lag from ~2.2s to ~0.54s, and ~300ms of that was
|
||
baked into the DB (`Audio.SyncOffsetMs=300`) via a POSITIVE-only control. For audio running BEHIND
|
||
video you cannot push audio later — positive delay makes it worse. OBS's answer is a NEGATIVE offset
|
||
that eats the buffer head: drop the first |N| ms of the written stream so every audio event lands
|
||
|N| ms EARLIER relative to video. It only makes sense at the head of the stream, so it is armed once
|
||
at `StartLive` (positive stays live-reactive via the delay line). Rule: record the signed semantics
|
||
together — positive = delay (ahead), negative = eat-head advance (behind) — and lock the control
|
||
while live (`IsEditMode`) since a mid-stream advance flip is meaningless. Test recipe:
|
||
`StartLive_NegativeOffset_AdvancesAudio_ByDroppingTheStreamHead` (emit 6×0.9 into an 8-tick budget,
|
||
long 0.2 bed, wire must show only 0.2).
|
||
|
||
## Composition video capture of WebView2: the CoreMessaging DQ recipe + the build gotchas
|
||
|
||
(2026-09-14, the ~20→60fps web-capture redesign — record-once so the next GPU-path work never
|
||
re-derives it. Everything that follows was already settled by OBS/Flutter `webview_windows`; the
|
||
cost this session was the 19041 projection's gaps.)
|
||
|
||
1. **No `DispatcherQueueController.CreateOnCurrentThread()` on the 10.0.19041 projection**
|
||
(CS0117 — only `CreateOnDedicatedThread` + `FromAbi(IntPtr)` exist). P/Invoke
|
||
`coreMessaging.dll!CreateDispatcherQueueController` with a sequential
|
||
`DispatcherQueueOptions { DwSize, ThreadType=2 (DQTYPE_THREAD_CURRENT), ApartmentType=2 (DQTAT_COM_STA) }`,
|
||
wrap via `DispatcherQueueController.FromAbi(ptr)` (mirror of `CaptureInterop`), THEN `new Compositor()`.
|
||
One controller + one compositor per UI thread, created once.
|
||
2. **`CoreWebView2CompositionController` needs a real parent HWND** — there is no window in the
|
||
MainWindow ctor, so `InitWebView2()` moved from the ctor to `MainWindow_Loaded` (source
|
||
registration is guarded by `ContainsKey`, so the layout-load path is safe).
|
||
3. **Capture the ROOT visual, not the control.** `GraphicsItem` for a composition controller comes
|
||
from `GraphicsCaptureItem.CreateFromVisual(root)` where `root` is your own 1920×1080
|
||
`ContainerVisual` (`IsVisible=true`) holding the controller's `RootVisualTarget` as a
|
||
`RelativeSizeAdjustment=1,1` child — exactly what `webview_windows`'s `graphics_context.cc`
|
||
does (`CreateGraphicsCaptureItemFromVisual` on the root `surface_` visual). Hide nothing,
|
||
move nothing, poll nothing — the capture is frame-driven at the renderer's pace.
|
||
4. **Straight alpha, never premultiplied.** Read back with `BitmapAlphaMode.Straight`; the
|
||
compositor's blend is straight-alpha, so `Premultiplied` readback wrecks corner anti-aliasing.
|
||
5. **Namespace landmines that each cost a build cycle in `Services/`:** `Compositor` resolves to the
|
||
repo's OWN `ytLive.Services.Compositor` namespace — fully-qualify `Windows.UI.Composition.Compositor`;
|
||
`CoreWebView2CompositionController.Close()` is the disposal call (no `Dispose()`); `Color` is
|
||
ambiguous (`System.Drawing` vs `System.Windows.Media`) — qualify `System.Drawing.Color.Transparent`.
|
||
6. **Testing a compositor-built bitmap without a runtime:** the internal seam ctor
|
||
`(Dispatcher, Func<string, IScreenCaptureSource>?)` + a `FakeWebSource` + a real background-STA
|
||
`DispatcherPump` (borrowed from ScreenCaptureManagerTests) drives the whole session path hermetic.
|
||
When byte-comparing a crop, slice the source with stride gaps (helper `CropBytes`) — a contiguous
|
||
range silently spans rows.
|
||
|
||
Recipe verified on green suite + 0-warning build; the composition path itself still needs a device
|
||
take (open item, HANDOFF).
|
||
|
||
## A pipe-read test that samples one tick is a timing flake by construction
|
||
|
||
(2026-09-14, re-discovered) `Mix_HonorsProviderGains_AndGameMute_KillsTheLoopback` fails
|
||
sporadically in ISOLATION (3/3 on the clean tree) but passes in the full suite: it reads exactly one
|
||
tick's worth after muting, and if the pipe still carries pre-mute buffered bytes the read spikes at
|
||
0.4 instead of silence. Any caller of `ReadFullyAsync` on a live pipe must either close/drain the
|
||
pipe first or assert on a multi-tick window. Do not blame a sync change for this — verify the grain
|
||
against the clean tree in the same mode before touching the mixer.
|
||
|
||
---
|
||
|
||
### Slice-18 follow-up (2026-09-15) — C4 composite cache: test fakes must mirror production's STABLE scene and per-tick purity
|
||
|
||
The C4 blit-on-change cache keys on a hash of the resolved inputs, INCLUDING per-element reference
|
||
identity (`RuntimeHelpers.GetHashCode(element)`). Two test-setup habits silently broke/starved it:
|
||
|
||
1. **`NewPump`'s default scene is fresh per tick** (`() => BackgroundScene()` — fine for pacing
|
||
tests) — churned the element refs, so the signature NEVER matched and the cache looked broken
|
||
(301 "renders" instead of 1). Production hands a STABLE `StagedScene`. **Rule: any FramePump test
|
||
that asserts per-frame content or cache behavior must pass `scene: () => scene` with one Scene
|
||
instance; the fresh-scene default is only for pacing/diagnostic tests.**
|
||
2. **A call-count resolver flip (`flip++ % 2`) double-advances under a cache-aware pump** — the
|
||
tick resolves TWICE (signature pass + compositor pass). Key per-tick content on the burned
|
||
`OutputIndex` (stable until submit, which happens AFTER render) or on `encoder.Frames.Count`
|
||
instead. A stable-per-tick token, not a call counter, is the deterministic alternation.
|
||
3. New frames that must invalidate the cache need BOTH a fresh array AND a fresh Epoch (array
|
||
identity alone is unchanged for a mutated-in-place array; `Epoch` is the monotonic generation
|
||
marker the compositor's paste key and the C4 signature share).
|
||
|
||
---
|
||
|
||
### ALERT-CLIP DECODE PIPE (RECIPE — TASK 47, 2026-09-26) — read a per-play mp4 into loose frames + stereo 48k audio
|
||
|
||
Worked out for the alert-box video: a short mp4 played **per alert** (one-shot) and dropped at
|
||
drain, NOT the refcounted media-source manager. Recipe:
|
||
|
||
- **Two ffmpeg pipe outputs, one child:** `ffmpeg -loglevel error -i "{file}" -f rawvideo -pix_fmt
|
||
bgra -vf scale={W}:{H} - -f f32le -ac 2 -ar 48000 -` — `-loglevel error` keeps stderr quiet (the
|
||
other pipes' ffmpeg rivals read the whole FD so a chatty stderr deadlocks); do **not** drain
|
||
stderr, mirror the `MediaVideoSource` split. Video pipe read backs onto `RawVideoFrameReader`
|
||
(TASK 21) which already pace-throttles to rtc-frame-rate… earlier live pacing is VFR-driven.
|
||
- **Audio = a carry-buffer loop, never a per-read slice.** PCM16 bytes → float32 chunks changes the
|
||
per-frame byte count; the naive `ReadFullyAsync(count)` per chunk **drops samples that straddle a
|
||
pipe read boundary** (the first draft's bug). The fix: an in-loop buffer that carries leftovers —
|
||
read `count - pending` bytes, emit whole float chunks from `pending + fresh`, keep the remainder.
|
||
Pacer: 50ms audio chunks → `Thread.Sleep` the (chunk length − read wall time) remainder.
|
||
- **bgra from ffmpeg is OPAQUE** (alpha=255). Whoever fades must copy the frame and scale alpha
|
||
across the copy (straight-alpha); the decoder never touches alpha. That keeps the decoder pure
|
||
and the fade a layer concern.
|
||
- **Fake-decoder convention:** the integration test implements `IAlertClipDecoder` inline with
|
||
`EmitFrame/EmitAudio/End` helpers — Start/Stop/Dispose are recorded as int counters, frames and
|
||
audio are pushed from the test thread, EOF is a test-raised event. No ffmpeg, no threads, still
|
||
the full real-time path through the layer.
|
||
- Re-grep here before building a second "play a file into the mix" path — snippet above is the
|
||
whole non-trivial part.
|
||
|
||
### A component's drain must reset its OWN anchor state (TASK 47 post-mortem — the test caught what a no-frame path hid)
|
||
|
||
`AlertOverlayLayer.Advance` gained a clip branch for video (TASK 47). In the ANIMATION branch the
|
||
old code did `_current = null; _elapsed = 0;` before `AdvanceToNext()`; the new clip branch called
|
||
only `StopClip(); AdvanceToNext();` — so after the EOF fade-out the layer kept `_current` pointing
|
||
at the finished message: `IsPlaying` stayed true, `RenderFrame` re-rendered the alert through the
|
||
animation fallback, and the layer never returned to idle. The six-animation tests couldn't see it
|
||
(they only ever exercise the animation branch); the video test's explicit
|
||
`Assert.False(layer.IsPlaying)` after drain nailed it on the first run. Rule: **when a second
|
||
execution path is added to a state machine, the drain/precondition contract of the ORIGINAL path is
|
||
part of the spec** — assert the component's end-state on the new path, don't assume the old path's
|
||
cleanup. The forest for the trip: the fallback path (animation) MASKED the broken path because its
|
||
output still looked like "an alert rendering", exactly the flick I'd have shipped if the test
|
||
hadn't forced the null-frame idle state.
|
||
|
||
### A seam fed from the chat poller must marshal EVERY WPF-object write to the UI thread (AlertOverlayLayer crash, 2026-09-26)
|
||
|
||
First crash in the app's history (startup.log `11:22:01.301`), and the sims gave zero warning for
|
||
months. Test-tab alert buttons run on the UI thread, so `OnMessageReceived` → preview writes were
|
||
always UI-threaded there. The moment a REAL message round-tripped through the chat poller:
|
||
`System.ArgumentException: Must create DependencySource on same Thread as the DependencyObject` in
|
||
`DataBindEngine.ProcessCrossThreadRequests` — WriteableBitmap created on the MTA poller thread,
|
||
`alertBox.VideoImageSource` raises INPC (`DisplaySource`) and is WPF-bound, the binding engine
|
||
re-binds cross-thread, process dies.
|
||
|
||
The chat layer already had the rule (`ChatOverlayLayer.OnMessageReceived` wraps its body in
|
||
`Application.Current.Dispatcher.Invoke`); the alert layer's copy of the seam didn't. Rule now
|
||
recorded for every layer: **the chat poller is an MTA producer — any INPC-raising seam it feeds
|
||
must marshal every WPF-object write (WriteableBitmap, DP/INPC-raised bound sources) to the UI
|
||
thread; wrap the WHOLE ingest, not just one writer, so enqueue/timer/preview stay coherent on the
|
||
UI thread. Guard `Application.Current` null like AlertTickerRenderer.Render does (pure seams).**
|
||
|
||
Red/green proof pattern: call the seam from a raw `new Thread` (MTA) while the RealApp message
|
||
loop runs, return to the loop so the marshalled work can execute (never `Join` on the UI thread),
|
||
then assert on the UI thread that `source.VideoImageSource.Dispatcher` is the App's dispatcher.
|
||
Red pre-fix (it was the poller's), green post-fix.
|
||
|
||
Grep-ahead: a `DispatcherTimer` ctor guarded by `if (Application.Current != null)` marks a class
|
||
with UI-thread affinity — audit every production ingress point for the marshal when adding one.
|
||
|
||
---
|
||
|
||
## A decoder that "starts" but yields nothing is DEAD AIR — fall back on a no-frame grace, never trust silent catches (2026-09-26)
|
||
|
||
Second live alert failure of the day, a different void than the crash: creator's repro — On-Air →
|
||
test → remove and re-add the Stream Alerts layer → TEST event → procs in chat but the alert box
|
||
stayed transparent, and startup.log 11:44's `peakMix 0.105` proved the alert audio never reached the
|
||
mix (vs 11:21's 0.375–0.891). ZERO exceptions anywhere; every failure stage was silent BY DESIGN:
|
||
`UpdatePreview`'s catch was `Debug.WriteLine`-only (nothing in release), and `AlertClipDecoder.RunAsync`
|
||
had a bare `catch { }` around the whole run (fires `Completed` regardless). A decode run that fails
|
||
AFTER `BeginClip` succeeds leaves `_clip != null` with `_latestClipFrame == null` → `RenderClipFrame`
|
||
returns null forever → box transparent FOREVER, and the six-animation fallback only ever ran on
|
||
setup-time throws. Even frame pacing was silently off (no frame-rate probe → `frameDuration = 0` →
|
||
video dumps instantly while 50 ms audio chunks pace the whole clip).
|
||
|
||
The established answer (derivative: any video-player/booth failure budget has the same contract):
|
||
**an alert that plays is a promise — the moment data stops arriving, degrade, don't vanish.** Fixed
|
||
in `AlertOverlayLayer.Advance`: a started clip that produced neither a video frame NOR an audio chunk
|
||
inside a `NoFrameFallbackSeconds` (1.0 s) grace is torn down, `_elapsed` resets to 0, and the SAME
|
||
alert continues as the six-animation branch — never drained early, never blank. And the silence is
|
||
gone: `BeginClip` writes the box id / resolved path / size + setup failure; `RunAsync` now logs the
|
||
frame/chunk end-state AND the exception it swallowed. Rules for every future gamble:
|
||
(1) a `catch` that hides WHY the output vanished is a lie — log the exception and the counters;
|
||
(2) any producer that yields zero output inside a grace becomes a fallback trigger, not a wait;
|
||
(3) one integration test per rescue: `SilentDecoder_FallsBackToTheAnimationAfterTheNoFrameGrace`
|
||
pins transparent→fell-back→renders→drains.
|
||
|
||
The pacing half that this entry flagged came true the same day: the 12:35 session showed a FULLY
|
||
HEALTHY decoder (`frames=240 audioChunks=156 failed=False`, alert audio in the mix) and the creator
|
||
still saw "no video / you lost the video". The clip played at ~240× for a second then froze on the
|
||
last frame: `AlertClipDecoderFor` passed no `frameRateProbe`, so `RunVideoAsync`'s pace
|
||
(`if (frameDuration > 0)`) was dead code while the audio chunk-pacer ran real-time. LESSON, sharper
|
||
than the whole fallback saga: **a decoder that drops data faster than wall-clock looks identical to a
|
||
dead decoder on screen — "plays but you don't see it" is a PACING bug before it's a decode bug.**
|
||
A healthy pipeline still needs an explicit pace per stream; never ship a pipe whose `frameDuration`
|
||
can be `TimeSpan.Zero`. Fix: the alert factory passes `FfmpegFrameRateProbe` exactly like the media
|
||
path does (MainViewModel.cs:308). Verify-by-seam once, everywhere: the same fake-process +
|
||
fake-probe delay-recorder test now pins real-time video pacing
|
||
(`AlertClipDecoderTests.AlertClipDecoder_PacesVideoFramesToTheProbedFrameRate`). Rule added to the
|
||
list above: (4) if output arrives but the USER still reports no/motionless picture, check pacing
|
||
before decode — count the bytes/s against wall-clock, don't re-audit the codec.
|
||
|
||
**RULE (5) — the last hop is a hop. A producer that hands over a correct frame has still
|
||
delivered nothing.** The 2026-09-26 alert hunt burned four builds (89/90 + two more) proving
|
||
the video was PERFECT: real h264, 240 frames, bright pixels at every timestamp, byte-identical
|
||
to the creator's source file, opaque `a255` frames logged every second into both the preview
|
||
and the compositor resolver. The box was blank because `PreviewPane.xaml`'s per-element
|
||
`Image` is `Collapsed` unless a DataTrigger says otherwise, and nobody had ever written the
|
||
`IsAlertBox` one. Two lessons, both expensive:
|
||
- **"It decodes, it's handed over, it's still invisible" ⇒ stop auditing the producer and
|
||
audit the CONSUMER'S VISIBILITY.** Bytes existing is not pixels being drawn. Diff the two
|
||
consumers' behaviour: chat box rendered, alert box didn't — that asymmetry was the whole
|
||
clue, and every per-frame log said "healthy".
|
||
- **A test whose fakes replace the failing seam will pass forever.** `AlertLayerVideoTests`
|
||
drove `RenderFrame` with a fake decoder and was green through the entire outage; the decode
|
||
test used a fake process; nothing ever crossed the XAML. When a feature's output is
|
||
*displayed*, the test must build the real display path (`AlertBoxPreviewVisibilityTests`
|
||
drives the real `MainWindow`+`PreviewPane`). And when the symptom is "nothing appears",
|
||
the honest test asserts on the CONSUMER (is the Image Visible / are the canvas pixels
|
||
different), never on the producer's own counters.
|
||
- Corollary: never trust "the data arrived" as a root cause. Ask what draws it.
|
||
|
||
**RULE (6) — a per-element `Image` is the wrong home for a GLOBAL overlay, and an
|
||
output-only producer is invisible by construction.** Same hunt, one layer up
|
||
(2026-09-26, after the `IsAlertBox` fix): the creator confirmed the clip plays, then
|
||
asked "where is the scrolling text? I do not see it." The ticker was healthy the whole
|
||
time — `AlertTickerRenderer` produced a 1920x48 frame and `FramePump` blitted it at
|
||
(0,0) — but it existed ONLY as a frame-pump callback (`alertTicker: () =>
|
||
_alertLayer.AlertTickerFrame`). `PreviewPane.xaml` had no element for it, because the
|
||
strip is master-width, not a `Source`, so it can't ride a per-element `Image`. Two
|
||
lessons:
|
||
- **When a feature's output is global (full frame width, a fixed overlay slot), it needs
|
||
its own root element + its own VM-published `WriteableBitmap`, exactly like
|
||
`SocialBarElement` — not a `Source`.** A global overlay that only a background thread
|
||
consumes will reach the stream and never the creator's preview, and no amount of
|
||
producer-side logging will show it: there is no consumer to be unhealthy.
|
||
- **Ask "who displays this?" while designing, not after the report.** RULE (5) says
|
||
audit the consumer when a picture is missing. RULE (6) is the earlier lesson: the
|
||
consumer may not exist yet, and that is invisible from the producer side.
|
||
|
||
**RULE (7) — pace a marquee in READS PER ALERT, never in pixels per second.** The ticker
|
||
shipped at a fixed `PixelsPerSecond = 140`, which sounds reasonable and was wrong: one
|
||
full pass took ~17s, so the 10s default clip showed the announcement barely once,
|
||
entering from the right and never crossing. The number that matches the creator's intent
|
||
("readable three times") is **passes per alert**, divided into the alert's own length —
|
||
the same convention the incumbent uses (Streamlabs exposes "Alert Duration: choose how
|
||
long your alert stays on your stream" and "Text Delay", never a scroll-speed slider;
|
||
https://support.streamlabs.com/hc/en-us/articles/52499995174299-Setting-up-Your-Streamlabs-Alerts).
|
||
So: `readsPerAlert / alertSeconds` => a pass length => a speed derived from it. Corollary
|
||
that cost a test rewrite: **a seamless marquee can never go blank**, so "the frame went
|
||
empty" is NOT a valid way to count passes — count the pill's leading edge resetting to
|
||
the right edge instead. And phase-start the run half a frame in, or the first published
|
||
frame is blank (the strip sits exactly off the right edge) and the preview opens on
|
||
nothing.
|
||
|
||
**RULE (8) — a bound `ItemsSource` ComboBox can break an unrelated real-mouse test.**
|
||
`LeftPanel.xaml` gained a display selector written as
|
||
`ItemsSource="{Binding ...}" + DisplayMemberPath + SelectedValuePath`; it compiled, my
|
||
feature's tests passed, and it broke `LayerReorderPersistenceTests.RealMouseDrag_OnThe
|
||
LayerList_PersistsTheReorder` — a test that injects PHYSICAL mouse input
|
||
(`SetCursorPos` + `mouse_event`). A data-bound selection resolves during load and
|
||
re-measures the panel under the test's feet, so the real cursor lands on the wrong row.
|
||
Rewriting it as inline `<ComboBoxItem>`s (the shape the chat Font selector already uses
|
||
in the same panel) fixed it with no other change. LESSON: **when a UI-only change breaks a
|
||
real-input test, the control's data-binding style is the suspect, not the feature** — and
|
||
bisect by reverting ONE file (`git checkout Controls/LeftPanel.xaml`) before theorising.
|
||
Prefer the idiom already in the file you are editing.
|
||
|
||
---
|
||
|
||
## 2026-09-26 — A blocking wait on the shared `RealApp` STA thread hangs the ENTIRE suite silently
|
||
|
||
**Symptom:** `vstest` printed 4 lines and then nothing. No test ever completed, no timeout, no
|
||
stack trace. Ran it three times, plus after a full rebuild and a reboot. Looked exactly like a
|
||
broken test runner.
|
||
|
||
**It was not the runner.** Two things were stacked:
|
||
|
||
1. **The runner output was being swallowed by my own shell plumbing.** Running `vstest` in the
|
||
foreground with a pipe to `tail` produced *zero* bytes even when the run finished. Launching it
|
||
with `nohup … > log 2>&1 &` and polling the log worked every time. **Recipe: always background
|
||
a `vstest` run and poll the log file.** Never trust a foreground piped `vstest` in this shell.
|
||
2. **The real hang was my test.** `LayerReorderPersistenceTests` and friends share ONE STA thread
|
||
via the `RealApp` collection. I had written the frame-pump tests as
|
||
`_app.Run(() => { pump.StartAsync(ct).GetAwaiter().GetResult(); … })`. `Run` *occupies* that
|
||
single thread, so the pump had nowhere to run and the two deadlocked — permanently, with no
|
||
timeout to save it.
|
||
|
||
**The rule:** a test that drives the frame pump MUST be `async` and use the new
|
||
`RealAppHost.RunAsync(Func<Task>)`, with plain `await pump.StartAsync(...)`. `await` yields the
|
||
STA thread back to the dispatcher; `.GetAwaiter().GetResult()` holds it hostage. **Never put a
|
||
blocking wait inside `RealAppHost.Run`.** A collection-serialized UI thread turns a missing
|
||
`await` into a whole-suite hang, which is why this presented as "the runner is broken."
|
||
|
||
**Also:** isolate before fixing. Running just the suspect class proved the runner was fine and my
|
||
test was the hang. Three blind retries on the same command would have "confirmed" a false theory.
|
||
|
||
---
|
||
|
||
## 2026-09-26 — Two real bugs my own new tests caught, and the cadence trap
|
||
|
||
Both were found only because the tests asserted the creator's *stated numbers* rather than
|
||
"something was drawn":
|
||
|
||
1. **Interval cadence implemented as an absolute deadline.** I had `_sinceStart` accumulate forever
|
||
and compared it to `_untilNextFlash`, which I *replaced* with a new interval on each fire. That
|
||
makes the second value an absolute time, not an interval, so gaps **shrink every cycle**:
|
||
measured 6.7s, 3.5s, 2.8s instead of 30–60s. Fix: keep an absolute `_nextFlashAt` but set it to
|
||
`_sinceStart + interval` on each fire, so the interval runs *from the last flash*. General trap:
|
||
whenever a timer holds "when did I last do X" and "when next", decide explicitly whether the
|
||
period is relative to the last event or to the start.
|
||
2. **A fade-in that starts at exactly 0 opacity emits a blank first frame.** `OpacityAt(0) == 0`, so
|
||
`Frame` returned null on the tick right after `Show()`. Not wrong (it is the start of a fade) but
|
||
it wasted a tick and made "is a credit on screen right now?" answer false. Worth knowing before
|
||
writing the assertion — my test's arithmetic was wrong, not the renderer, and I nearly "fixed"
|
||
the renderer to match the bad test.
|
||
|
||
**The lesson:** pin the *numbers the creator asked for* (2s, 30–60s, fully-on-canvas). "It rendered
|
||
something" passes against almost any implementation and caught neither bug.
|
||
|
||
**Third trap, cost me a wasted verification:** to prove a `const` isn't shipped I grepped the
|
||
Release binary for it — and it was absent from the *Debug* binary too, because **consts are inlined
|
||
at compile time and never reach metadata**. The control (run it against Debug as well) is what
|
||
exposed that. Grep for a **field/type** (`InstanceVariable`) or use reflection; never grep for an
|
||
inlined literal, and always run the check against the build you expect it to FAIL in.
|
||
|
||
## 2026-09-26 — A falsy `ShowDialog()` fell through into the side effect: "Cancel" that saved
|
||
|
||
The recording save prompt let the creator name the file. It always had a **Cancel** button, and
|
||
Cancel always *appeared* to work — the dialog closed, the toast was skipped, nothing visibly went
|
||
wrong. The footage was saved anyway, under the default name.
|
||
|
||
```csharp
|
||
// WRONG — and it type-checked, read fine, and looked correct
|
||
string stem = autoStem;
|
||
var dialog = new RenameRecordingDialog(autoStem);
|
||
if (dialog.ShowDialog() == true && !string.IsNullOrEmpty(dialog.ResultFileName))
|
||
stem = dialog.ResultFileName;
|
||
|
||
var finalPath = UniquePath(Path.Combine(dir, stem + ".mp4"));
|
||
File.Move(startPath, finalPath); // <-- runs even when the creator hit Cancel
|
||
```
|
||
|
||
**The trap:** `ShowDialog()` returning `false` was collapsed into the *default* branch instead of
|
||
being treated as a distinct outcome. "Not what I typed" and "I don't want this" are different
|
||
answers, and the nullable stem made the first one look like a legitimate value of the second. The
|
||
default value of the accumulator absorbed the cancel.
|
||
|
||
**The rule, and the grep-able form of it:** when a modal's result gates a side effect, every branch
|
||
is explicit and the side effect sits **inside** the accepting branch:
|
||
|
||
```csharp
|
||
if (!dialog.ShowDialog() == true) { DoTheOppositeThing(); return; } // read the FALSE case first
|
||
DoTheThing(); // the effect is not a fall-through
|
||
```
|
||
|
||
A `DialogResult`/return value that only ever *modifies* a later value (rather than *permitting* the
|
||
effect) is the smell. `null` from a dialog is the same trap wearing a nullable.
|
||
|
||
**Why it survived so long:** the side effect is in a *different method* from the modal, and the modal
|
||
had its own passing test (`RenameRecordingDialogSizingTests`, which only checks layout). The test
|
||
suite was green because the bug was never in the tested unit. **A modal's own test proves nothing
|
||
about the decision the modal feeds** — cover the commit-or-discard decision itself, against real
|
||
files (`ytLive.Tests/RecordingSaveDialogTests.cs` drives the extracted
|
||
`MainViewModel.CompleteRecordingSave` and asserts Cancel leaves the directory EMPTY, not just "the
|
||
stem is unchanged").
|
||
|
||
**Creator's actual ruling, worth keeping verbatim:** "if I click cancel, that doesn't mean to accept
|
||
the default file name and save the file — that's what pressing enter should do. Cancel should cancel
|
||
the save option, discarding the recording." A Cancel button on a destructive-and-final dialog means
|
||
*discard*, not *accept default*. When a dialog is the last gate before data is committed, ask what
|
||
Cancel should destroy — don't infer "keep the default" just because the default already exists.
|
||
|
||
---
|
||
|
||
## 2026-09-27 — A policy number in my head is not a policy finding: check that the policy *applies*
|
||
|
||
Two confident, wrong claims in one Store-research pass, both of the same shape — I had a policy
|
||
number and a confident paraphrase attached to it, and never asked whether the policy's own scope
|
||
clause covered us.
|
||
|
||
**1. "Store policy 10.10 (Advertising) requires X"** — 10.10 does not apply to LlamaCasty at
|
||
all. Every clause in it is conditioned on **ads your product displays**. YouTube injects ads
|
||
*server-side after the RTMP handoff*; the app renders no ad, hosts no ad SDK, and makes no ad
|
||
decision. I spent a cycle designing around a rule that was never in force. Same error later with
|
||
**10.2.5** (Store-only installation): the current policy text scopes it to **game and Xbox
|
||
console products**, and the older v7.7 wording that said "all of your apps" has been narrowed —
|
||
citing the stale wording would have produced a confidently wrong conclusion in the *opposite*
|
||
direction (claiming we were locked to MSIX).
|
||
|
||
**2. "Buy an EV certificate — EV gets instant SmartScreen reputation."** This one was worse than
|
||
wrong, because it was already **written into `Distribution.md`** and would have cost $400+/yr.
|
||
Microsoft **removed the SmartScreen instant-bypass for EV certs in March 2024**. EV and OV now
|
||
accrue reputation identically. I recited a fact that had been dead for over two years, in a
|
||
document that already contained the recommendation.
|
||
|
||
**The rule, and the generalisation:** *a policy citation is a claim about scope, not just about
|
||
text.* Before acting on one, read the **applicability sentence** and answer three questions —
|
||
does it cover our product class, does it cover our distribution model, and is the wording the
|
||
**current** version? A policy number without a verified scope is a guess wearing a citation's
|
||
clothes, and the failure mode is invisible because the sentence still *sounds* authoritative.
|
||
|
||
Corollary: **a stale external fact inside a shipped document is worse than a missing one** — it
|
||
is trusted, acted on, and costed. When a doc makes a factual claim about the outside world
|
||
(certificates, platform behaviour, policy text), re-verify it on touch, not on memory. This is
|
||
the same class as the dead-Pin incident in `FfmpegLocator` (`ai.md` → FFmpeg locator): an
|
||
external thing changed, our record of it did not, and the mismatch was discovered by a user
|
||
rather than by a test. **Records of the outside world rot silently and are believed until they
|
||
cost money.**
|
||
|
||
---
|
||
|
||
## 2026-09-27 — I offered a mitigation I had not checked the mechanics of: "just refund anyone who asks"
|
||
|
||
While discussing a subscription price, I wrote that at `$99/yr` with zero marginal cost the
|
||
creator could **"just refund anyone who asks"** — presented as a free mitigation against the
|
||
churn/refund burden of a subscription. **That lever does not exist for Microsoft Store IAP.**
|
||
Refunds run through Microsoft, not the developer; the buyer requests them from the Store, and
|
||
the developer cannot unilaterally issue one. I proposed it before ever checking how Store
|
||
refund mechanics work, and it sounded confident because it was arithmetically true and
|
||
operationally impossible.
|
||
|
||
**The rule:** *a mitigation is not an answer until you know who has the authority to execute
|
||
it.* Before offering "we can just do X," ask **who presses the button** — and if the answer is
|
||
"Microsoft, or the user's bank, or a regulator," then it is not our mitigation and must not be
|
||
counted as one. The same sentence reappears in a subtler failure: the underlying instinct
|
||
(refunds *should* be cheap when marginal cost is zero) was sound; the *agency* was invented.
|
||
|
||
**The verification that came out of it is worth more than the mistake.** Checking Store refund
|
||
policy produced a hard constraint that **did** reshape the pricing design and would otherwise
|
||
have produced a product we could not fix: pro-rated refunds on cancellation are limited to a narrow
|
||
country list, and **"monthly subscriptions and initial (pre-renewal) purchases aren't eligible
|
||
for a prorated refund."** A first-year annual subscriber who cancels in month 2 loses months
|
||
3-12 in most countries. That is precisely the trap the creator had already named as
|
||
irritating — and it is unfixable from the developer side. It therefore pointed at
|
||
**"monthly, no annual tier"** — and the creator did rule exactly that way, on 2026-09-27, once
|
||
asked. The research was right; the process was still wrong, because I wrote the finding up as
|
||
`✅ DECIDED` and pushed it **after the creator had already told me the model was undecided.**
|
||
|
||
**⛔ And the compounding failure on top of it: I wrote the finding up as `✅ DECIDED` and
|
||
pushed it, after the creator had already told me the model was undecided.** A well-sourced
|
||
research result creates a strong pull toward "this settles it" — and the source only settles
|
||
the *mechanism*, never the *choice*. **Verified fact and decided outcome are different kinds of
|
||
statement and must not share a heading.** Ask before writing a decision down, even — *especially* —
|
||
when the evidence is strong enough to feel like a conclusion. **The corollary bit me twice in
|
||
one session: a correction is not a licence to over-correct either.** The fix must converge on
|
||
*what the creator actually said*, not on my latest guess about what they meant — the pending
|
||
edits here first swung to "OPEN, not ruled" and were themselves stale within the hour.
|
||
|
||
**Generalisation, and it is the same shape as the policy-applicability lesson above:** the
|
||
expensive mistakes are never the ones where I knew nothing. They are the ones where a *plausible,
|
||
well-formed, unverified* claim got to speak without a check — a dead EV-certificate benefit, a
|
||
policy number with no scope, a refund lever I did not control. **Sound reasoning about a
|
||
mechanism I had not verified is the failure mode, not the safety net.** Before any external
|
||
mechanism drives a decision, name the one fact that would falsify it, then go get that fact.
|