perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)

Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv
~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s.
Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms
of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact:
pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays
native; the lever is concurrency.

- MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit)
- new MonotonicGate (Interlocked compare-exchange): stale OLDER completions
  are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear)
- FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is
  per-conversion locals
- Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green,
  0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point.

Local only, no push.
This commit is contained in:
2026-09-14 18:40:20 -07:00
parent 71932b9756
commit b37b8a30f9
7 changed files with 238 additions and 114 deletions
+59 -60
View File
@@ -1,11 +1,11 @@
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
# HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next)
## Branch / Commit State
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
audio-sync, pushed). Working tree **clean**.
`main` HEAD = **slice-17 commit** (overlapping capture readbacks — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture
conversion fix, slice-15 pacing fix (FramePump), `c01206f` (composition capture), `b22d08e`
(signed audio-sync, pushed). Working tree **clean**.
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
@@ -13,62 +13,60 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
## 🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
tear_audit.py` + `mix_check.py`):
Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the
creator reports the desktop layer **still looks like missing frames / choppy vs live**. Decoded
`ty-20260914-1824-0000-2.mp4` to `/mnt/c/tmpout/f1824.raw` and audited:
- Desktop band: **6.8 content updates/s, 88% frozen**, one **4.85s freeze** at start. Worse,
not better, than 1742.
- BUT the new slice-16 telemetry proved the downscale fix WORKS: `conv avg 46-50ms, max ~61ms,
~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0`. The ring is
steady-state (no allocs); the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`) is ~45ms of
the conversion** on a 240Hz-HDR box sharing the GPU with the encoder. **The wall was never
the CPU downscale.**
- FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render
is the SECOND cap, ~30 unique composites/s).
- Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that
theory is now closed).
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
was real but MOOT.
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
reported.
## 🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/MonotonicGate.cs` + locked
`Services/FrameRingBuffer.cs`):
1. **MaxConcurrentConversions = 3** readbacks in flight (was one-in-flight `_framePending`),
pool buffers 2 → **5** so in-flight frames fit.
2. **MonotonicLatest publish gate** (`MonotonicGate`, new internal): a completed readback is
published ONLY if its Epoch is strictly newer than the last published. Overlapping
conversions can finish out of order; a slow OLDER completion must never overwrite a newer
`LatestFrame` (backwards time hole = the mirror of the 1742 tear). Epoch via
`Interlocked.Increment`.
3. `FrameRingBuffer.Rent`/`ConsumeAllocations` now take `_lock` (rents are concurrent); the
10ms floor and downscale stay; `DownscaleBgra` row-scratch is per-conversion locals (no
shared `_row0/_row1`).
4. **RESEARCH near-miss (docs'ed, MyMistakes):** shrinking the pool to 1920×1080 would have
been wrong — Microsoft Docs (screen capture): *"the underlying Direct3D surface is always
the size specified … **clipped**"* to the frame. Readback stays native; the lever is
concurrency.
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
preview lags).
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
`ring allocs` — so the next device take is judged numerically.
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
lessons + the recast note), this HANDOFF.
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
budget. C4 (compositor optimization) stays conditional on the re-measure.
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
**Good Dog test:** `ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins`.
**296/296 green, app build 0 warnings.** Scope-lock files (6): `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `Services/MonotonicGate.cs` (new), `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`
+ docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). `SceneCompositor.cs`/`FramePump.cs`
NOT touched (C4 is the next slice, pending this re-measure).
## ⚠️ Open items (before PUSHABLE)
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
slice-16 build. Judge numerically:
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-17
build. Judge numerically:
- startup.log telemetry: conversions/s should jump from ~17-20 to **≥ ~30-40**, `skip busy`
falling, `ring allocs` ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps).
- Decode + `/tmp/opencode/tear_audit.py`: desktop-band fresh updates/s up toward the render
cap (~30+), frozen % well under 50%, no genuine mid-frame splits.
- ffprobe: video ≈ audio ≈ wall.
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
composite downscale / blit-on-change), still staying 60fps.
- If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots →
**C4 slice** (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags,
the bind is GPU contention — re-measure before touching anything.
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
## Open threads (carried)
@@ -77,7 +75,7 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
- Signed A/V sync: verify the negative (advance) direction on device.
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
- Focus-loss capture lag — closed as NOT the cause (1824: delivery healthy ~60-100/s).
## Landmines
@@ -88,14 +86,15 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
recipe** — grep before re-deriving.
- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and now the **WGC-CLIP** + two
slice blocks — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
## Next step
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.
Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line
(conversions/s ≥ ~30-40, `skip busy` falling) + `tear_audit.py` cadence + ffprobe durations. If
the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice
next, then re-measure clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.