Files
LlamaCasty/HANDOFF.md
T
gramps 5ba4d709ae docs: audio-sync offset measured (~367ms lag w/ 300 offset) -> set Audio.SyncOffsetMs=0
silencedetect audio word times cross-correlated with per-frame video motion
(tblend=difference + signalstats) peaks at +22 frames: the 300ms TASK-22
offset plus ~70ms pipeline lag. 370ms reads as one full beat behind during
the ~2.3s pauses. Zeroing the offset targets the residual ~70ms.
2026-09-12 14:04:58 -07:00

8.9 KiB
Raw Blame History

HANDOFF — 2026-09-12 (webcam FIXED; audio-silence FIXED; TRUNCATION ROOT-CAUSED + FIXED for static scenes)

Branch / Commit State

main HEAD = a62a283. Ahead of origin by 25 commits. NOT pushing — user ruling: no push until web-overlay transparency AND audio-silence are addressed.

Uncommitted (all belong to the current work, none committed yet):

M .gitignore
M HANDOFF.md                       (this file)
M MyMistakes.md                    (webcam collateral recipe added)
M Services/Audio/AudioSyncDelay.cs (audio-silence fix)
M Services/Compositor/SceneCompositor.cs (webcam cbW/cbH fix)
M Services/MediaCaptureFrameSource.cs (diagnostics REMOVED — clean)
M ytLive.Tests/AudioSyncDelayTests.cs (silence regression test, proven both ways)
M ytLive.Tests/SceneCompositorTests.cs (webcam regression test, proven both ways)
M Services/Encoder/FramePump.cs (slow-render probe, see TRUNCATION below)

Build is green (0 warnings); 295 tests pass (audio + compositor + pump suites).

Audio.SyncOffsetMs changed DB-side to 0 (see AUDIO section) — HANDOFF-only change, nothing new to compile.

✅ TRANSPARENCY — CLOSED (pushed gate #1)

User confirmed transparency is fixed (efe88b8 era). Push gate #1 clears.

✅ WEBCAM GRAY BLOCK — CLOSED

Root cause: SceneCompositor.BlitContentRaw defaulted cbW=cbH=0 when src.CropBounds is null — regression from ed9d7c1. Every webcam frame (no crop bounds) blitted to an empty 0×0 source rect → nothing drawn where the webcam should be. Fix: default cbW=src.Width, cbH=src.Height. Regression test BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin proven BOTH ways (fails on old code, passes on new). User confirmed correct webcam render in the latest recording. All temp diagnostics stripped from MediaCaptureFrameSource.cs and SceneCompositor.cs.

🟡 AUDIO SILENCE — ROOT-CAUSED AND FIXED, pending final user take

Symptom: mic/desktop levels showed real values in preview and in the per-5s Audio live: telemetry, but peakMix stayed exactly 0.000 and saved recordings were digital silence (-91 dB, verified with ffprobe -af volumedetect).

Root cause: AudioMixer.FillAndMix calls _syncDelay.Configure(...) on EVERY ~10ms tick (it re-reads the live UI setting each mix). The old AudioSyncDelay.Configure unconditionally reallocated and zeroed the delay buffer AND reset _writePos=0 on every call — discarding the just-written audio before the delay offset could ever read it back. With offset 0 the early passthrough masked it; any non-zero saved offset = total silence. The DB has Audio.SyncOffsetMs = 500 (left over from TASK 22 testing) → that's what triggered it.

Fix: Configure now early-returns when the delay samples haven't actually changed (still reallocates/flushes on a change, so the live slider still clicks rather than smearing). Regression test RepeatedConfigureWithSameDelay_EveryTick_StillDelivers_TheMarker mirrors the mixer's exact calling pattern and was proven BOTH ways (marker lost = silence on old code; marker survives → expected 960-sample shift on new).

Telemtery proof the fix works: latest take (ty-20260912-1135-0000-2.mp4, recorded WITH the fix, sync offset still 500): peakMix = 0.434–0.523 across all five 5s windows (was 0.000 before). Audio is NOT silence anymore.

Follow-up: sync offset Audio.SyncOffsetMs measured and SET TO 0 (2026-09-12, take ty-...-1358):

Measurement method tri-wins instead of clap deltas: silencedetect on the audio gave the count-word starts (one≈0, two@4.12, three@6.48 — audio PTS, both streams start at 0, no mux offset), and per-frame video motion (tblend=difference + signalstats YAVG) cross-correlated against the audio envelope peaks at shift +22 frames → audio LAGS video by ~367ms with the 300ms offset active → the pipeline's own lag is only ~70ms. The saved 300ms (TASK 22 leftover) is what made it read as "one word late" (370ms ≈ a third of the ~2.3s pause ≈ perceptually a full beat behind). Fix = offset 0. Expected residual ≤~70ms (imperceptible).

✅ TRUNCATED VIDEO — ROOT-CAUSED AND FIXED (for baked/static scenes)

Symptom: recordings ran ~1s short at first, then HALF-length: 22.8s wall → 10.48s file. Not a mux/-shortest cut — the pump simply produces frames slower than the declared 60fps.

Mechanism (proven): ffmpeg rawvideo timestamps frames by ARRIVAL at declared -r 60. Pump deadline pacing (OBS libobs pattern) yields each frame when its ~17ms slot lands; if the render stage for that tick takes longer than the slot, the pump can only do ~1 frame per 35ms → ~27–28fps of submitted frames → file length ≈ submittedFrames/60. So a take that should be 22.8s muxes to ~10.5s, and logs show avg render 35–38ms with 136/300 per 5s. FramePump had NO way to tell which compositor pass ate a slow frame.

Root cause of the 35ms: the scene was NOT actually fully static. Any dynamic element still in the scene — including hidden ones, which still count as Dynamic in SceneGraph.GetSplitPoint — forces the per-frame composite path. If it sits at split 0, GetBakedBase returns null and every tick is a FULL render (no baked base). The 12:59/13:39 takes logged resolve 1.0ms per frame — a live element was still resolving, confirming dynamic content was present.

The fix that matters (no code changed in the renderer): a truly static scene renders in <1ms. New slow-render probe (ProbeRender in FramePump.cs) names the path; the 13:53 take of a 1-element static scene logs render=fully-static split=1 dynamic=0 bakeMs=246 totalMs=246 on the FIRST frame (cold raster of the background at 1080p), then 0.8→0.0ms avg render, 301/300 frames per 5s, 0 stalls → 12.46s file from 12.29s wall, 735 frames @ 60fps. No truncation.

Rule for recording: every scene used for RECORDING should be ≥1 static element first (background image) with all web/cam/chat widgets AFTER it in layer order (split ≥ 1 → bake feasible) — a scene whose FIRST layer is dynamic at recording time will truncate. If a take with widgets comes back truncated again, the probe now points at whether it took fully-static (bake), or base+layers / full-render (composite) and how long.

Probe left in place (≈0 overhead: one Stopwatch, string built only when a render ≥20ms): it is the diagnostic that named this bug; keep it until the dynamic-scene path is proven fast too, then remove.

🔴 OPEN — user-reported "video seems truncated ~1s" (unconfirmed)

Recording ty-20260912-1135-0000-2.mp4: format duration 26.370s, video 26.167s (1570 frames @ 60fps), audio 26.370s. FramePump stats for the whole take show dropped 0 and 300/300 frames every 5s window. End-frames (n=1500 vs 1560 vs 1569) differ — not a frozen/repeat-last-frame truncation. Log shows recording window 11:35:01.386 → 11:35:27.562 = 26.2s of video: matches the file. No evidence of truncation found in data. Possibly user perceived the 500ms audio delay (audio trailing video start / leading video end) as truncation — worth re-checking with offset zeroed. If a real freeze/truncation reappears: pull last-3-frames diffs and confirm_frame timestamps next.

Other threads (paused)

  • Audio-silence verification — the FIXED take has real peakMix; final = user re-records with the finalized offset and confirms audible audio in playback.
  • Audio sync offset final value — set to 0 via DB (measured residual ~70ms). NEEDS user re-take to confirm (the running app may still hold 300 in memory — if unsynced on the re-take, use the UI slider to 0, or restart the app BEFORE recording so the 0 loads).
  • Truncation with DYNAMIC scenes — proven mechanism; static scenes now record full-length. The 13:58 take was a 6-element/4-dynamic scene (split=0 → full-render, ~34ms, 20 stalls/5s): fully full-length (10.98s) but JITTERY (judder from the 34ms hard frames). If a widget-heavy scene reads badly again, the ProbeRender log names the stage; the next move is getting a baked base by reordering static elements first or optimizing full-render.
  • Webcam MJPG missing — "MJPG negotiation refused (being used by another process)". Queued.
  • Web capture speed — ~10-14Hz effective. Revisit only on request.
  • Layer order — SortOrder not persisting on drag. Queued.

Landmines

  • testhost shares startup.log — filter by time.
  • cmd.exe /c "taskkill /F /IM ytLive.exe" (WSL form double-slashes mangle) before rebuilds.
  • Build/tests: Windows dotnet host (/mnt/c/Program Files/dotnet/dotnet.exe). 0 warnings.
  • ffmpeg/ffprobe: /mnt/c/Program Files/Krita (x64)/bin/ffmpeg.exe with Windows paths.
  • Audio.SyncOffsetMs persisted non-zero WILL re-trigger silence symptoms if Configure is ever called unconditionally again — the idempotent guard is the fix, keep it.
  • MyMistakes.md has the transparency RESOLVED entry and the new webcam collateral recipe — grep before any re-derivation.