Files
LlamaCasty/HANDOFF.md
T
gramps 5ba4d709ae docs: audio-sync offset measured (~367ms lag w/ 300 offset) -> set Audio.SyncOffsetMs=0
silencedetect audio word times cross-correlated with per-frame video motion
(tblend=difference + signalstats) peaks at +22 frames: the 300ms TASK-22
offset plus ~70ms pipeline lag. 370ms reads as one full beat behind during
the ~2.3s pauses. Zeroing the offset targets the residual ~70ms.
2026-09-12 14:04:58 -07:00

161 lines
8.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# HANDOFF — 2026-09-12 (webcam FIXED; audio-silence FIXED; TRUNCATION ROOT-CAUSED + FIXED for static scenes)
## Branch / Commit State
`main` HEAD = `a62a283`. Ahead of origin by 25 commits. **NOT pushing** — user
ruling: no push until web-overlay transparency AND audio-silence are addressed.
Uncommitted (all belong to the current work, none committed yet):
```
M .gitignore
M HANDOFF.md (this file)
M MyMistakes.md (webcam collateral recipe added)
M Services/Audio/AudioSyncDelay.cs (audio-silence fix)
M Services/Compositor/SceneCompositor.cs (webcam cbW/cbH fix)
M Services/MediaCaptureFrameSource.cs (diagnostics REMOVED — clean)
M ytLive.Tests/AudioSyncDelayTests.cs (silence regression test, proven both ways)
M ytLive.Tests/SceneCompositorTests.cs (webcam regression test, proven both ways)
M Services/Encoder/FramePump.cs (slow-render probe, see TRUNCATION below)
```
Build is green (0 warnings); 295 tests pass (audio + compositor + pump suites).
Audio.SyncOffsetMs changed DB-side to 0 (see AUDIO section) — HANDOFF-only change,
nothing new to compile.
## ✅ TRANSPARENCY — CLOSED (pushed gate #1)
User confirmed transparency is fixed (`efe88b8` era). Push gate #1 clears.
## ✅ WEBCAM GRAY BLOCK — CLOSED
Root cause: `SceneCompositor.BlitContentRaw` defaulted `cbW=cbH=0` when
`src.CropBounds` is null — regression from `ed9d7c1`. Every webcam frame (no crop
bounds) blitted to an empty 0×0 source rect → nothing drawn where the webcam
should be. Fix: default `cbW=src.Width`, `cbH=src.Height`. Regression test
`BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin` proven BOTH
ways (fails on old code, passes on new). User confirmed correct webcam render in
the latest recording. All temp diagnostics stripped from
`MediaCaptureFrameSource.cs` and `SceneCompositor.cs`.
## 🟡 AUDIO SILENCE — ROOT-CAUSED AND FIXED, pending final user take
**Symptom:** mic/desktop levels showed real values in preview and in the
per-5s `Audio live:` telemetry, but `peakMix` stayed exactly `0.000` and saved
recordings were digital silence (-91 dB, verified with `ffprobe -af
volumedetect`).
**Root cause:** `AudioMixer.FillAndMix` calls `_syncDelay.Configure(...)` on
EVERY ~10ms tick (it re-reads the live UI setting each mix). The old
`AudioSyncDelay.Configure` unconditionally reallocated and zeroed the delay
buffer AND reset `_writePos=0` on every call — discarding the just-written
audio before the delay offset could ever read it back. With offset 0 the early
passthrough masked it; any non-zero saved offset = total silence. The DB has
`Audio.SyncOffsetMs = 500` (left over from TASK 22 testing) → that's what
triggered it.
**Fix:** `Configure` now early-returns when the delay samples haven't actually
changed (still reallocates/flushes on a *change*, so the live slider still
clicks rather than smearing). Regression test
`RepeatedConfigureWithSameDelay_EveryTick_StillDelivers_TheMarker` mirrors the
mixer's exact calling pattern and was proven BOTH ways (marker lost = silence
on old code; marker survives → expected 960-sample shift on new).
**Telemtery proof the fix works:** latest take (ty-20260912-1135-0000-2.mp4,
recorded WITH the fix, sync offset still 500): `peakMix` = 0.434–0.523 across
all five 5s windows (was 0.000 before). **Audio is NOT silence anymore.**
**Follow-up:** sync offset `Audio.SyncOffsetMs` measured and SET TO **0** (2026-09-12, take ty-...-1358):
Measurement method tri-wins instead of clap deltas: `silencedetect` on the audio
gave the count-word starts (one≈0, two@4.12, three@6.48 — audio PTS, both streams
start at 0, no mux offset), and per-frame video motion (`tblend=difference` +
`signalstats YAVG`) cross-correlated against the audio envelope peaks at shift
**+22 frames → audio LAGS video by ~367ms with the 300ms offset active** → the
pipeline's own lag is only ~70ms. The saved 300ms (TASK 22 leftover) is what made
it read as "one word late" (370ms ≈ a third of the ~2.3s pause ≈ perceptually a
full beat behind). Fix = offset 0. Expected residual ≤~70ms (imperceptible).
## ✅ TRUNCATED VIDEO — ROOT-CAUSED AND FIXED (for baked/static scenes)
**Symptom:** recordings ran ~1s short at first, then HALF-length: 22.8s wall →
10.48s file. Not a mux/`-shortest` cut — the pump simply produces frames slower
than the declared 60fps.
**Mechanism (proven):** ffmpeg rawvideo timestamps frames by ARRIVAL at declared
`-r 60`. Pump deadline pacing (OBS libobs pattern) yields each frame when its
~17ms slot lands; if the render stage for that tick takes longer than the slot,
the pump can only do ~1 frame per 35ms → ~27–28fps of *submitted* frames → file
length ≈ submittedFrames/60. So a take that *should* be 22.8s muxes to ~10.5s,
and logs show `avg render 35–38ms` with 136/300 per 5s. FramePump had NO way to
tell which compositor pass ate a slow frame.
**Root cause of the 35ms:** the scene was NOT actually fully static. Any dynamic
element still in the scene — **including hidden ones**, which still count as
Dynamic in `SceneGraph.GetSplitPoint` — forces the per-frame composite path. If
it sits at split 0, `GetBakedBase` returns null and every tick is a FULL render
(no baked base). The 12:59/13:39 takes logged `resolve 1.0ms` per frame — a live
element was still resolving, confirming dynamic content was present.
**The fix that matters (no code changed in the renderer):** a truly static scene
renders in <1ms. New slow-render probe (`ProbeRender` in FramePump.cs) names the
path; the 13:53 take of a 1-element static scene logs
`render=fully-static split=1 dynamic=0 bakeMs=246 totalMs=246` on the FIRST
frame (cold raster of the background at 1080p), then `0.8→0.0ms` avg render,
`301/300` frames per 5s, 0 stalls → **12.46s file from 12.29s wall, 735 frames
@ 60fps. No truncation.**
**Rule for recording:** every scene used for RECORDING should be ≥1 static
element first (background image) with all web/cam/chat widgets AFTER it in
layer order (split ≥ 1 → bake feasible) — a scene whose FIRST layer is dynamic
at recording time will truncate. If a take with widgets comes back truncated
again, the probe now points at whether it took `fully-static` (bake), or
`base+layers` / `full-render` (composite) and how long.
**Probe left in place** (≈0 overhead: one Stopwatch, string built only when a
render ≥20ms): it is the diagnostic that named this bug; keep it until the
dynamic-scene path is proven fast too, then remove.
## 🔴 OPEN — user-reported "video seems truncated ~1s" (unconfirmed)
Recording ty-20260912-1135-0000-2.mp4: format duration 26.370s, video 26.167s
(1570 frames @ 60fps), audio 26.370s. FramePump stats for the whole take show
`dropped 0` and 300/300 frames every 5s window. End-frames (n=1500 vs 1560 vs
1569) differ — not a frozen/repeat-last-frame truncation. Log shows recording
window 11:35:01.386 → 11:35:27.562 = 26.2s of video: matches the file. No
evidence of truncation found in data. Possibly user perceived the 500ms audio
delay (audio trailing video start / leading video end) as truncation — worth
re-checking with offset zeroed. If a real freeze/truncation reappears: pull
last-3-frames diffs and confirm_frame timestamps next.
## Other threads (paused)
- **Audio-silence verification** — the FIXED take has real peakMix; final =
user re-records with the finalized offset and confirms audible audio in playback.
- **Audio sync offset final value** — set to 0 via DB (measured residual ~70ms).
NEEDS user re-take to confirm (the running app may still hold 300 in memory —
if unsynced on the re-take, use the UI slider to 0, or restart the app BEFORE
recording so the 0 loads).
- **Truncation with DYNAMIC scenes** — proven mechanism; static scenes now
record full-length. The 13:58 take was a 6-element/4-dynamic scene
(`split=0 → full-render`, ~34ms, 20 stalls/5s): fully full-length (10.98s)
but JITTERY (judder from the 34ms hard frames). If a widget-heavy scene reads
badly again, the ProbeRender log names the stage; the next move is getting a
baked base by reordering static elements first or optimizing full-render.
- **Webcam MJPG missing** — "MJPG negotiation refused (being used by another
process)". Queued.
- **Web capture speed** — ~10-14Hz effective. Revisit only on request.
- **Layer order** — SortOrder not persisting on drag. Queued.
## Landmines
- testhost shares startup.log — filter by time.
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL form double-slashes mangle) before rebuilds.
- Build/tests: Windows dotnet host (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/ffmpeg.exe` with Windows paths.
- `Audio.SyncOffsetMs` persisted non-zero WILL re-trigger silence symptoms if
`Configure` is ever called unconditionally again — the idempotent guard is the
fix, keep it.
- `MyMistakes.md` has the transparency RESOLVED entry and the new webcam
collateral recipe — grep before any re-derivation.