Files
LlamaCasty/HANDOFF.md
T
gramps b37b8a30f9 perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)
Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv
~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s.
Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms
of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact:
pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays
native; the lever is concurrency.

- MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit)
- new MonotonicGate (Interlocked compare-exchange): stale OLDER completions
  are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear)
- FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is
  per-conversion locals
- Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green,
  0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point.

Local only, no push.
2026-09-14 18:40:20 -07:00

6.3 KiB
Raw Blame History

HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next)

Branch / Commit State

main HEAD = slice-17 commit (overlapping capture readbacks — committed LOCALLY, NOT pushed; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture conversion fix, slice-15 pacing fix (FramePump), c01206f (composition capture), b22d08e (signed audio-sync, pushed). Working tree clean.

⚠️ Branding (2026-09-14, creator-corrected): product = llamacasty, internals = ytLive

The product is llamacasty; repo path, csproj AssemblyName/RootNamespace, DB/log paths (%APPDATA%\ytLlive\...), and most code names are the legacy ytLive/ytLlive. User-facing language says "llamacasty"; code/assembly/repo names stay ytLive. See ai.md → Brand.

🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall

Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the creator reports the desktop layer still looks like missing frames / choppy vs live. Decoded ty-20260914-1824-0000-2.mp4 to /mnt/c/tmpout/f1824.raw and audited:

  • Desktop band: 6.8 content updates/s, 88% frozen, one 4.85s freeze at start. Worse, not better, than 1742.
  • BUT the new slice-16 telemetry proved the downscale fix WORKS: conv avg 46-50ms, max ~61ms, ~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0. The ring is steady-state (no allocs); the GPU→CPU readback (CreateCopyFromSurfaceAsync) is ~45ms of the conversion on a 240Hz-HDR box sharing the GPU with the encoder. The wall was never the CPU downscale.
  • FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render is the SECOND cap, ~30 unique composites/s).
  • Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that theory is now closed).

🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool

What (Services/ScreenCaptureFrameSource.cs + new Services/MonotonicGate.cs + locked Services/FrameRingBuffer.cs):

  1. MaxConcurrentConversions = 3 readbacks in flight (was one-in-flight _framePending), pool buffers 2 → 5 so in-flight frames fit.
  2. MonotonicLatest publish gate (MonotonicGate, new internal): a completed readback is published ONLY if its Epoch is strictly newer than the last published. Overlapping conversions can finish out of order; a slow OLDER completion must never overwrite a newer LatestFrame (backwards time hole = the mirror of the 1742 tear). Epoch via Interlocked.Increment.
  3. FrameRingBuffer.Rent/ConsumeAllocations now take _lock (rents are concurrent); the 10ms floor and downscale stay; DownscaleBgra row-scratch is per-conversion locals (no shared _row0/_row1).
  4. RESEARCH near-miss (docs'ed, MyMistakes): shrinking the pool to 1920×1080 would have been wrong — Microsoft Docs (screen capture): "the underlying Direct3D surface is always the size specified … clipped" to the frame. Readback stays native; the lever is concurrency.

Good Dog test: ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins. 296/296 green, app build 0 warnings. Scope-lock files (6): Services/ScreenCaptureFrameSource.cs, Services/FrameRingBuffer.cs, Services/MonotonicGate.cs (new), ytLive.Tests/ScreenCaptureFrameSourceTests.cs

  • docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). SceneCompositor.cs/FramePump.cs NOT touched (C4 is the next slice, pending this re-measure).

⚠️ Open items (before PUSHABLE)

  • Device re-verify (next step): creator records the SAME tv-show scenario on the slice-17 build. Judge numerically:
    • startup.log telemetry: conversions/s should jump from ~17-20 to ≥ ~30-40, skip busy falling, ring allocs ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps).
    • Decode + /tmp/opencode/tear_audit.py: desktop-band fresh updates/s up toward the render cap (~30+), frozen % well under 50%, no genuine mid-frame splits.
    • ffprobe: video ≈ audio ≈ wall.
  • If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots → C4 slice (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags, the bind is GPU contention — re-measure before touching anything.
  • No push yet — commit-locally-until-greenlight for web/A/V work.

Open threads (carried)

  • Audio-silence verification — fixed (724af14); creator heard real audio.
  • Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
  • Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
  • Signed A/V sync: verify the negative (advance) direction on device.
  • Focus-loss capture lag — closed as NOT the cause (1824: delivery healthy ~60-100/s).

Landmines

  • testhost shares startup.log — filter by time.
  • cmd.exe /c "taskkill /F /IM ytLive.exe" (WSL double-slashes mangle) before rebuilds — a live app process locks ytLive.exe and the apphost copy fails (MSB3021).
  • Build/tests: Windows dotnet host (/mnt/c/Program Files/dotnet/dotnet.exe). 0 warnings — only ./scripts/verify.sh "<files>"'s clean build counts. Building ytLive.csproj alone does NOT rebuild ytLive.Tests.dll — run the Tests csproj before vstest.
  • ffmpeg/ffprobe: /mnt/c/Program Files/Krita (x64)/bin/ with Windows paths.
  • MyMistakes.md has the freeze-audit RECIPE, the A/V sync measurement recipe, the deadline-pacing lessons, the CoreMessaging DQ recipe, and now the WGC-CLIP + two slice blocks — grep before re-deriving.
  • sqlite3 at /home/gramps/android-sdk/platform-tools/sqlite3.
  • C:\tmpout is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.

Next step

Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line (conversions/s ≥ ~30-40, skip busy falling) + tear_audit.py cadence + ffprobe durations. If the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice next, then re-measure clap offset (/tmp/opencode/avsync.py), then decide push with the user.