fix(webcam): reader-output-subtype ladder + record WinRT per-call wrap rule

A browser grabbing the C920 kills our reader startup: the camera stays in
MJPG 1920x1080@30 (another app's format) and our OBS-refusal to re-negotiate
under SharedReadOnly contention (`SetMediaStreamPropertiesAsync` throws
'file in use / CaptureMode is SharedReadOnly') leaves it there.
CreateFrameReaderAsync(.., Bgra8) + StartAsync then refuses with
OutputFormatNotSupported — the MJPG-active source only exposes NV12 at the
reader level (clue: startup.log 19:01/19:05 sessions, same hardware that
started YUY2 640x480 fine at 08:52). Fresh launches showed no webcam and
'Add Webcam' failed.

Reader creation is now a per-candidate ladder (ReaderSubtypeCandidates):
Bgra8 for uncompressed cameras (unchanged fast path); NV12 then the
source-default for MJPG cameras — converted in OnFrameArrived like any
non-BGRA frame. Each candidate is allocated AND started under its OWN
catch: WinRT answers an unsupported subtype with a throw (E_INVALIDARG),
not a status, so a single rejected format must degrade to the next
candidate instead of aborting acquisition (creator rule — see MyMistakes
WINRT resource-allocation recipe). Rejections are logged and collected
into the final error.

Good Dog: 4 unit tests lock the candidate ordering (MJPG never Bgra8,
case-insensitive, uncompressed keeps Bgra8 first, unknown/null -> Bgra8).
301/301 green, 0 warnings. [no push]
This commit is contained in:
2026-09-15 19:25:39 -07:00
parent 64a5a6d06f
commit 94a934ffa9
5 changed files with 223 additions and 64 deletions
+62 -49
View File
@@ -1,10 +1,10 @@
# HANDOFF — 2026-09-15 (slice-18 C4 composite cache committed locally — device verify next)
# HANDOFF — 2026-09-15 (webcam take fix 1 of 2 committed locally; slice-18 already local — verify both)
## Branch / Commit State
`main` HEAD = **slice-18 commit** (C4 blit-on-change composite cache — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-17 overlapping
capture readbacks, slice-16 capture conversion fix, slice-15 pacing fix (FramePump), `c01206f`
`main` HEAD = **webcam reader-ladder fix** (committed LOCALLY with this handoff, NOT pushed — web/A/V
work stays commit-local until greenlight). Below it: **slice-18 C4 composite cache** (also local, un-pushed),
slice-17 overlapping capture readbacks, slice-16 capture conversion fix, slice-15 pacing fix, `c01206f`
(composition capture), `b22d08e` (signed audio-sync, pushed). Working tree clean.
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
@@ -13,46 +13,52 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
## ✅ Committed locally — slice 18: C4 = FramePump blit-on-change composite cache
## ✅ Committed locally — webcam take fix 1 of 2: reader-output-subtype ladder
The ty-1841 take (slice-17 build) proved capture fixed (band ~20 fresh updates/s, no tears, pacing
clean) but the render is STILL the wall: **FramePump stall on EVERY iteration** (`totalMs 21-44`,
`render=full-render split=0 elements=6 dynamic=4`, worst render 166ms startup spike) — only ~22-28
composites/s. The SceneGraph split can't fix it: `GetSplitPoint` returns **0** because the
live-capture backdrop is element 0 and CANNOT be baked (a cached capture goes stale).
**Incident (2026-09-15):** browser (msedge) grabbed the C920 → app died silently (last log line
09:16:47 is the `MediaCapture.Failed` "device no longer present" rollback; then nothing — no
AppDomain/Dispatcher handler fired → native WMF death). Fresh launches then showed NO webcam:
`CameraManager: camera '…GLOBAL' failed: '… frame reader refused to start: OutputFormatNotSupported'`.
**What** (`Services/Encoder/FramePump.cs`, + new Good Dog test in `ytLive.Tests/FramePumpTests.cs`):
the full-render path now caches the last composite + its INPUT IDENTITY. `BuildFullRenderSignature`
mirrors the compositor's own resolution (same resolver seam: element ref + layout/visual bits +
resolved frame's array identity + Epoch + CropBounds + options + social bar) — unchanged identity →
ONE `Buffer.BlockCopy` (~3ms) instead of the full re-composite (~30ms); changed identity → re-render.
Cache buffer is a separate long-lived array, written pre-burn/pre-recycle (never the scratch pool).
Gated on the 1:1 config (the only deployed tier). Telemetry: `cache {renders}R/{hits}H` on the 5s
stats line + internal `CacheHits`/`CacheRenders`/`OutputIndex`.
**Diagnosis (startup.log, four sessions):** the camera's live media type is the variable. Under
msedge + NVIDIA Broadcast contention our `SetMediaStreamPropertiesAsync` refuses ("file is being used
by another process / CaptureMode is SharedReadOnly") so the camera keeps whoever's format: YUY2 640x480
@07:47-08:52 → reader started fine (starved of frames while they held it; worked at 08:52 when they
didn't); **MJPG 1920x1080 @19:01/19:05 → `CreateFrameReaderAsync(.., Bgra8)` + `StartAsync` refused
with OutputFormatNotSupported** (MS-documented: an MJPG-active source can't be read as Bgra8 by a
MediaFrameReader — it exposes NV12; the MJPG→BGRA converter isn't on the reader pipeline).
**Good Dog test:** `FullRenderCache_StaticInputs_RenderOnce_Then_Reuse_UntilInputChanges` — static
scene renders ONCE then hits (byte-identical above the burn strip), a new frame (new array + Epoch)
invalidates + propagates. Existing `Pump_Pools_...` test now passes a STABLE scene (like production)
and keys alternation on `OutputIndex` (the two resolver passes per tick double-advanced a call-count
flip). **297/297 green, app + tests build 0 warnings.** Scope-locked (2 code files + ai.md +
HANDOFF): `Services/Encoder/FramePump.cs`, `ytLive.Tests/FramePumpTests.cs`.
**Fix** (`Services/MediaCaptureFrameSource.cs`): reader creation is now a per-candidate ladder —
`ReaderSubtypeCandidates(activeSubtype)`: Bgra8 for uncompressed cameras (unchanged fast path);
**NV12 then source-default for MJPG** (converted in `OnFrameArrived`, which already handled non-BGRA).
Each candidate is allocated AND started under its **own catch** — WinRT answers an unsupported subtype
with a THROW (`E_INVALIDARG`), not a status; the throw degrades to the next candidate instead of
aborting acquisition (creator rule: WINRT resource-allocation calls wrap per-call; see
`MyMistakes.md`). Rejections are logged + joined into the final error.
## ⚠️ Open items (before PUSHABLE)
**Good Dog tests:** `ytLive.Tests/MediaCaptureFrameSourceTests.cs` x4 — MJPG never asked as Bgra8,
case-insensitive, uncompressed keeps Bgra8 fast path, unknown/null → Bgra8. **301/301 green, clean
build 0 warnings, scope-check passed** (2 code files + ai.md + HANDOFF + MyMistakes).
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-18
build. Judge numerically:
- startup.log telemetry: `cache` line shows hits dominating on TV holds (`e.g. cache 1R/250H`),
`avg render` drops toward the ~3ms BlockCopy, FramePump **stalls disappear** (the per-iteration
stall was the C4 signature).
- Decode + `/tmp/opencode/freeze_audit.py` / `band_timeline.py`: desktop-band fresh updates/s up
toward ~60 (was ~20 first-11s; the render cap was the bind), no mid-frame splits.
- ffprobe: video ≈ audio ≈ wall (pacing already healthy at slice 17 — unchanged expected).
- If the desktop layer STILL reads choppy after cache hits dominate every static hold, the residual
is the **24fps TV → 60fps container pulldown** (inherent 3:2-ish repeats; the 1841 gap histogram
was 89×2-slot + 84×3-slot holds) — that's content, not the pipeline; decide with the creator
whether it needs an adaptive cadence or is acceptable.
- **No push yet** — commit-locally-until-greenlight for web/A/V work. After the take verdict, also
re-measure the clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.
## ⚠️ Open items
- **Webcam take fix 2 of 2 — the 09:16 crash is UNFIXED (hard, untested):** hypothesis — the app died
natively (no managed log line) when `MediaCapture.Failed` fired mid-stream: RollbackSession runs
`_ = SafeStopAsync` fire-and-forget → `StopAsync` disposes reader+capture while the frame-reader
thread is mid-`TryAcquireLatestFrame`/`Marshal.Copy` (that call region sits OUTSIDE the frame's
try/catch). Managed/unwrapped failures there surface via AppDomain handler (absent → native AV).
NOT fixed in this change: no repro, no integration test, and a native WMF race isn't catchable --
deferred per spin-guard. Record if a second occurrence shows a pattern.
- **Can't re-add webcam (3rd symptom):** partly the same root cause (add → AcquireAsync fails → empty
chip + toast). Note: if ANOTHER scene still holds a WebcamSceneConfig, `_webcam != null` persists and
"Add → Webcam" stays disabled by the single-identity rule — use right-click "Show Webcam" instead.
- **Device re-verify (two things, same take session):**
1. slice-18 cache (unchanged verdict criteria): `cache NR/WH` hits dominate on TV holds, stalls gone,
band fresh updates/s toward ~60 → else residual 24↔60 pulldown, content decision.
2. webcam ladder: while msedge holds the camera, a fresh launch must show the webcam (or degrade to
a clear chip, never a silent box); after closing the browser it must come up at 1080p30 MJPEG.
- **No push yet** — after the take verdict, re-measure the clap offset (`/tmp/opencode/avsync.py`),
then decide push with the user.
## Open threads (carried)
@@ -64,27 +70,34 @@ HANDOFF): `Services/Encoder/FramePump.cs`, `ytLive.Tests/FramePumpTests.cs`.
## Landmines
- testhost shares startup.log — filter by time.
- testhost shares startup.log — filter by time; the app also writes fake-device failures
('test-camera', 'dev-physical-9') from the CameraManager unit tests.
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL double-slashes mangle) before rebuilds — a live
app process locks `ytLive.exe` and the apphost copy fails (MSB3021).
- Build/tests: **Windows dotnet host** (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings —
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`. Known audio flake:
`Mix_HonorsProviderGains_AndGameMute_KillsTheLoopback` (float precision) — re-run if it trips
alone; it is unrelated to camera work.
- Camera contention failure modes (startup.log): "no frames within 4s" = init OK but starved;
"OutputFormatNotSupported" = reader subtype below the MJPG-active camera; "file is being used by
another process / CaptureMode is SharedReadOnly" = our stream re-negotiation refused while others
hold the device. CameraConflictProbe lists named processes that are merely RUNNING (msedge,
NVIDIA Broadcast) — a suspect list, not handle evidence.
- FramePump tests that assert per-frame CONTENT must pass a STABLE scene (`() => scene`) — the
default NewPump scene is fresh-per-tick (ok for pacing tests, but it churns the C4 render
signature and hides the cache). Cache-sensitive assertions also can't use a call-count resolver
flip (the tick resolves twice: signature + render) — key alternation on `OutputIndex`.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and the **WGC-CLIP** + slice
blocks — grep before re-deriving.
- `MyMistakes.md` has the WINRT resource-allocation RECIPE (new), freeze-audit RECIPE, A/V sync
measurement recipe, deadline-pacing lessons, CoreMessaging DQ recipe, WGC-CLIP + slice blocks —
grep before re-deriving. ai.md has the webcam reader-ladder note.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
## Next step
Creator records a tv-show take on the slice-18 build → read the startup.log telemetry: shift-stall
frequency and the `cache NR/WH` line (hits must dominate on TV holds) + the band audit + ffprobe
durations. If the desktop now tracks ~60 updates/s and stalls are gone: re-measure the clap offset,
then decide push with the user. If the layer is still choppy on fully-static holds, the pulldown
(readme) is the residual and it's a content decision, not a pipeline bug.
Creator records a take on this build (backend test): (a) with the browser holding the camera — fresh
launch must show the webcam or a named chip, then (b) close the browser and re-add the webcam to the
live view — must come back at 1080p30 MJPEG; (c) the slice-18 cache verdicts from the same session
(`cache NR/WH` hits, stall frequency, band audit). If (b) works: measure clap offset, decide push.