perf(pump): slice 18 — C4 blit-on-change composite cache

ty-1841: capture fixed (band ~20 updates/s, no tears) but FramePump stalls
on EVERY iteration (totalMs 21-44, render=full-render split=0 elements=6
dynamic=4) — the render ceiling, ~22-28 composites/s, caps the desktop in a
60fps file. SceneGraph can't help: the live backdrop is element 0 and cannot
be baked (a cached capture goes stale), so the cache lives at the pump.

RenderFull wraps both full-render call sites: BuildFullRenderSignature hashes
the full input identity (options rect, social bar, per-element layout/visual
bits + the frame each element would resolve through the SAME resolver seam,
using array identity + Epoch + CropBounds); unchanged identity reuses the last
composite with one Buffer.BlockCopy (~3ms) instead of a ~30ms re-composite.
Cache buffer is a separate long-lived array, written pre-burn/pre-recycle
(the caller burns the frame counter and recycles scratch AFTER render).
Engagement gated on the 1:1 config (the only deployed tier). Telemetry
surfaces `cache NR/WH` on the 5s stats line.

Same shape as OBS (sources cache their surface, the scene blits on update) —
docs.obsproject.com/backend-design, the pattern this repo cites since the
2026-09-04 paste-cache slice.

Good Dog: FullRenderCache_StaticInputs_RenderOnce_Then_Reuse_UntilInputChanges
(static scene renders ONCE + byte-identical reuse; new frame+Epoch invalidates).
297/297 green, 0 warnings, verify.sh scope-locked (FramePump.cs, FramePumpTests.cs
+ ai.md/HANDOFF/MyMistakes). No push — device re-verify next.
This commit is contained in:
2026-09-15 07:57:09 -07:00
parent b37b8a30f9
commit 64a5a6d06f
5 changed files with 378 additions and 67 deletions
+49 -59
View File
@@ -1,11 +1,11 @@
# HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next)
# HANDOFF — 2026-09-15 (slice-18 C4 composite cache committed locally — device verify next)
## Branch / Commit State
`main` HEAD = **slice-17 commit** (overlapping capture readbacks — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture
conversion fix, slice-15 pacing fix (FramePump), `c01206f` (composition capture), `b22d08e`
(signed audio-sync, pushed). Working tree **clean**.
`main` HEAD = **slice-18 commit** (C4 blit-on-change composite cache — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-17 overlapping
capture readbacks, slice-16 capture conversion fix, slice-15 pacing fix (FramePump), `c01206f`
(composition capture), `b22d08e` (signed audio-sync, pushed). Working tree clean.
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
@@ -13,61 +13,46 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
## 🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall
## ✅ Committed locally — slice 18: C4 = FramePump blit-on-change composite cache
Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the
creator reports the desktop layer **still looks like missing frames / choppy vs live**. Decoded
`ty-20260914-1824-0000-2.mp4` to `/mnt/c/tmpout/f1824.raw` and audited:
- Desktop band: **6.8 content updates/s, 88% frozen**, one **4.85s freeze** at start. Worse,
not better, than 1742.
- BUT the new slice-16 telemetry proved the downscale fix WORKS: `conv avg 46-50ms, max ~61ms,
~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0`. The ring is
steady-state (no allocs); the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`) is ~45ms of
the conversion** on a 240Hz-HDR box sharing the GPU with the encoder. **The wall was never
the CPU downscale.**
- FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render
is the SECOND cap, ~30 unique composites/s).
- Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that
theory is now closed).
The ty-1841 take (slice-17 build) proved capture fixed (band ~20 fresh updates/s, no tears, pacing
clean) but the render is STILL the wall: **FramePump stall on EVERY iteration** (`totalMs 21-44`,
`render=full-render split=0 elements=6 dynamic=4`, worst render 166ms startup spike) — only ~22-28
composites/s. The SceneGraph split can't fix it: `GetSplitPoint` returns **0** because the
live-capture backdrop is element 0 and CANNOT be baked (a cached capture goes stale).
## 🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool
**What** (`Services/Encoder/FramePump.cs`, + new Good Dog test in `ytLive.Tests/FramePumpTests.cs`):
the full-render path now caches the last composite + its INPUT IDENTITY. `BuildFullRenderSignature`
mirrors the compositor's own resolution (same resolver seam: element ref + layout/visual bits +
resolved frame's array identity + Epoch + CropBounds + options + social bar) — unchanged identity →
ONE `Buffer.BlockCopy` (~3ms) instead of the full re-composite (~30ms); changed identity → re-render.
Cache buffer is a separate long-lived array, written pre-burn/pre-recycle (never the scratch pool).
Gated on the 1:1 config (the only deployed tier). Telemetry: `cache {renders}R/{hits}H` on the 5s
stats line + internal `CacheHits`/`CacheRenders`/`OutputIndex`.
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/MonotonicGate.cs` + locked
`Services/FrameRingBuffer.cs`):
1. **MaxConcurrentConversions = 3** readbacks in flight (was one-in-flight `_framePending`),
pool buffers 2 → **5** so in-flight frames fit.
2. **MonotonicLatest publish gate** (`MonotonicGate`, new internal): a completed readback is
published ONLY if its Epoch is strictly newer than the last published. Overlapping
conversions can finish out of order; a slow OLDER completion must never overwrite a newer
`LatestFrame` (backwards time hole = the mirror of the 1742 tear). Epoch via
`Interlocked.Increment`.
3. `FrameRingBuffer.Rent`/`ConsumeAllocations` now take `_lock` (rents are concurrent); the
10ms floor and downscale stay; `DownscaleBgra` row-scratch is per-conversion locals (no
shared `_row0/_row1`).
4. **RESEARCH near-miss (docs'ed, MyMistakes):** shrinking the pool to 1920×1080 would have
been wrong — Microsoft Docs (screen capture): *"the underlying Direct3D surface is always
the size specified … **clipped**"* to the frame. Readback stays native; the lever is
concurrency.
**Good Dog test:** `ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins`.
**296/296 green, app build 0 warnings.** Scope-lock files (6): `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `Services/MonotonicGate.cs` (new), `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`
+ docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). `SceneCompositor.cs`/`FramePump.cs`
NOT touched (C4 is the next slice, pending this re-measure).
**Good Dog test:** `FullRenderCache_StaticInputs_RenderOnce_Then_Reuse_UntilInputChanges` — static
scene renders ONCE then hits (byte-identical above the burn strip), a new frame (new array + Epoch)
invalidates + propagates. Existing `Pump_Pools_...` test now passes a STABLE scene (like production)
and keys alternation on `OutputIndex` (the two resolver passes per tick double-advanced a call-count
flip). **297/297 green, app + tests build 0 warnings.** Scope-locked (2 code files + ai.md +
HANDOFF): `Services/Encoder/FramePump.cs`, `ytLive.Tests/FramePumpTests.cs`.
## ⚠️ Open items (before PUSHABLE)
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-17
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-18
build. Judge numerically:
- startup.log telemetry: conversions/s should jump from ~17-20 to **≥ ~30-40**, `skip busy`
falling, `ring allocs` ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps).
- Decode + `/tmp/opencode/tear_audit.py`: desktop-band fresh updates/s up toward the render
cap (~30+), frozen % well under 50%, no genuine mid-frame splits.
- ffprobe: video ≈ audio ≈ wall.
- If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots →
**C4 slice** (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags,
the bind is GPU contention — re-measure before touching anything.
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
- startup.log telemetry: `cache` line shows hits dominating on TV holds (`e.g. cache 1R/250H`),
`avg render` drops toward the ~3ms BlockCopy, FramePump **stalls disappear** (the per-iteration
stall was the C4 signature).
- Decode + `/tmp/opencode/freeze_audit.py` / `band_timeline.py`: desktop-band fresh updates/s up
toward ~60 (was ~20 first-11s; the render cap was the bind), no mid-frame splits.
- ffprobe: video ≈ audio ≈ wall (pacing already healthy at slice 17 — unchanged expected).
- If the desktop layer STILL reads choppy after cache hits dominate every static hold, the residual
is the **24fps TV → 60fps container pulldown** (inherent 3:2-ish repeats; the 1841 gap histogram
was 89×2-slot + 84×3-slot holds) — that's content, not the pipeline; decide with the creator
whether it needs an adaptive cadence or is acceptable.
- **No push yet** — commit-locally-until-greenlight for web/A/V work. After the take verdict, also
re-measure the clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.
## Open threads (carried)
@@ -85,16 +70,21 @@ NOT touched (C4 is the next slice, pending this re-measure).
- Build/tests: **Windows dotnet host** (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings —
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
- FramePump tests that assert per-frame CONTENT must pass a STABLE scene (`() => scene`) — the
default NewPump scene is fresh-per-tick (ok for pacing tests, but it churns the C4 render
signature and hides the cache). Cache-sensitive assertions also can't use a call-count resolver
flip (the tick resolves twice: signature + render) — key alternation on `OutputIndex`.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and now the **WGC-CLIP** + two
slice blocks — grep before re-deriving.
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and the **WGC-CLIP** + slice
blocks — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
## Next step
Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line
(conversions/s ≥ ~30-40, `skip busy` falling) + `tear_audit.py` cadence + ffprobe durations. If
the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice
next, then re-measure clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.
Creator records a tv-show take on the slice-18 build → read the startup.log telemetry: shift-stall
frequency and the `cache NR/WH` line (hits must dominate on TV holds) + the band audit + ffprobe
durations. If the desktop now tracks ~60 updates/s and stalls are gone: re-measure the clap offset,
then decide push with the user. If the layer is still choppy on fully-static holds, the pulldown
(readme) is the residual and it's a content decision, not a pipeline bug.