perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)

The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel
DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take
ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame
audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was
the wall, and the one torn frame was a ring slot rewritten under the consumer's
read. Reference: WGC delivers at DWM/monitor cadence
(https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture)
and libyuv row-simple/fixed-point scaling
(https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule.

- DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage
  math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame.
- 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of
  serialized conversion/s; capacity sits just above the 60/s the pump can use.
- FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only
  rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a
  consumer still holds (session.LatestFrame survives conversions, dispatcher
  preview lags) is never read-while-overwritten. Needs no consumer Release API.
- 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring
  allocs — the device take is judgeable numerically.

Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings.
C4 (composite Epoch-cached downscale) deferred pending the device re-measure.
Local only, no push.
This commit is contained in:
2026-09-14 18:18:34 -07:00
parent 76f51e6f4e
commit 71932b9756
6 changed files with 417 additions and 110 deletions
+61 -51
View File
@@ -1,11 +1,11 @@
# HANDOFF — 2026-09-14 (composition capture DEVICE-VERIFIED ✓; slice-15 pacing fix committed locally — A/V re-measure next)
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
## Branch / Commit State
`main` HEAD = **slice-15 commit** (FramePump duplicate-on-lag pacing — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: `c01206f` (composition
capture redesign, also local/unpushed), before that `b22d08e` (signed audio-sync, pushed).
Working tree **clean**.
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
audio-sync, pushed). Working tree **clean**.
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
@@ -13,53 +13,62 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
## ✅ DEVICE-VERIFIED — composition capture looks good; then the A/V defect surfaced
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
The creator ran real takes off the `c01206f` build and reports the video looks **very good** — the
Windows.Graphics.Capture web path is confirmed on device (open item closed). BUT the same takes
showed the A/V defect we have seen before: **both desktop and webcam playback are accelerated, the
audio "gets speedier", then cuts off at the end.**
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
tear_audit.py` + `mix_check.py`):
## 🔬 Sliced and measured (evidence)
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
was real but MOOT.
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
reported.
With the A/V recipe (`MyMistakes.md` → "Measuring audio-video A/V sync" + `/tmp/opencode/avsync.py`):
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
- `ty-20260914-1723-0000-2.mp4`: 697 video frames @60fps = **11.62s** vs audio **11.84s**.
- `ty-20260914-1726-0000-2.mp4` (clap take): 163 frames = **2.72s** vs **2.93s** audio. Video ends
0.21–0.24s before audio → the cut-off. Clap: audio env peak 1.655s vs video motion 0.75–1.03s;
cross-correlation lag **+0.667s** (audio late).
- `startup.log`: `FramePump stall: iteration 33-46ms (> 2× the 17ms interval): worst render 33ms …
dropped 0`.
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
**Root cause (slice 15):** slice 10's freshness choice skipped the overrun's missed slots — one
fresh frame per ~35ms stall authored into a 60fps container = **accelerated playback**. OBS's
answer is duplicate-on-lag (`libobs/video-io.c`, docs.obsproject.com/backend-design): fill each
missed slot by repeating the newest frame — duration == wall, judder never fast-forward.
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
preview lags).
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
`ring allocs` — so the next device take is judged numerically.
## 🔬 Committed locally — slice 15: one frame per deadline slot (duplicate-on-lag)
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
lessons + the recast note), this HANDOFF.
**What:** `Services/Encoder/FramePump.cs` submit is now a bounded catch-up —
`while (now >= nextTick) { submit latest composite; nextTick += intervalTicks; }` clamped to a
`deadlineNow` captured once per iteration. First missed slot gets the fresh composite, the rest get
repeats of it (OBS duplication) — recording duration == wall time under any render load, no
acceleration, no audio tail cut. The burned `_outputIndex` moved inside the submit loop (every
emitted slot gets its own +1; also fixes the old unconditional bump that gapped the judge sequence
on non-submitting fast-render iterations).
**Good Dog test:** `ytLive.Tests/FramePumpTests.cs` → `Pump_Overrun_Renders_EmitsEverySlot_NotSkipped`
(60fps, 35ms render cost, asserts ≥0.65 of the wall slots emitted — the old skip-pump wrote ~1/35ms).
**294/294 green, app build 0 warnings.** Files: `Services/Encoder/FramePump.cs`,
`ytLive.Tests/FramePumpTests.cs`. Docs in same commit: ai.md Slice 15, MyMistakes (supersedes the
slice-10 "no burst re-write" clause + latest numbers), this HANDOFF.
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
budget. C4 (compositor optimization) stays conditional on the re-measure.
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
## ⚠️ Open items (before PUSHABLE)
- **Device re-verify:** a take on the slice-15 build must play at REAL time (no acceleration, no
cut-off audio tail). Confirm via ffprobe: video stream duration ≈ audio ≈ container.
- **Re-measure the true A/V offset** with a clap take now that video pacing is honest
(`/tmp/opencode/avsync.py`). The +0.6s reading on 1726 was confounded by the 1.1x
acceleration. If a real residual remains after the pacing fix, the backlog is the AUDIO pipeline
(mixer ring / advance), not video.
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
slice-16 build. Judge numerically:
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
- ffprobe: video ≈ audio ≈ wall.
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
composite downscale / blit-on-change), still staying 60fps.
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
## Open threads (carried)
@@ -68,24 +77,25 @@ slice-10 "no burst re-write" clause + latest numbers), this HANDOFF.
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
- Signed A/V sync: verify the negative (advance) direction on device.
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
## Landmines
- testhost shares startup.log — filter by time.
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL double-slashes mangle) before rebuilds — a live
app process locks `ytLive.exe` and the apphost copy fails (MSB3021, seen today).
app process locks `ytLive.exe` and the apphost copy fails (MSB3021).
- Build/tests: **Windows dotnet host** (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings —
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
- `MyMistakes.md` has the **A/V sync measurement recipe**, the **deadline-pacing** lessons
(item 1 + slice 15), and the **CoreMessaging DQ recipe** — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3` for
`/mnt/c/Users/gramp/AppData/Roaming/ytLlive/ytLlive.db`.
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
recipe** — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
## Next step
Have the creator record a clap take on the slice-15 build → ffprobe durations (video == audio ==
wall) + `/tmp/opencode/avsync.py` for the honest offset. If durations match, the acceleration
defect is closed; then decide push with the user, and attack any true audio-latency residual as its
own work unit.
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.