fix: clear pre-live capture ring backlog at StartLive (2.2s A/V sync lag)
AudioMixer.Start() begins mic/loopback capture at app startup to feed the level meters, so both 2s ring buffers fill with pre-live audio. StartLive() drained from the oldest tail sample, putting every recorded event ~2.2s late in the audio track (confirmed by clap analysis + cross-correlation on two takes: +2.11 to +2.22s). Fix: AudioMixer.StartLive() now runs _micBuffer.Clear() + _loopbackBuffer.Clear() immediately after the pipe starts, before the drain task runs. Recording now begins at go-live; the <=10ms in-flight chunk evicted by Clear() is imperceptible. Regression test: StartLive_DiscardsPreLiveBacklog_SoFirstAudioIsCurrent saturates the loopback ring with stale 0.8 pre-live audio, then asserts the wire carries fresh post-live 0.2 (max < 0.3). Reference (external, per derivative-work rule): OBS 'Audio mixer' keeps its buffers fed continuously and syncs the stream start timestamp at record time rather than replaying pre-live capture; a go-live flush of the capture buffer is the accepted pattern for live tools restarting a stream. Docs: HANDOFF.md (fix shipped), MyMistakes.md (A/V sync measurement recipe).
This commit is contained in:
+111
-129
@@ -1,152 +1,126 @@
|
|||||||
# HANDOFF — 2026-09-12 (webcam FIXED; audio-silence FIXED; TRUNCATION ROOT-CAUSED + FIXED for static scenes)
|
# HANDOFF — 2026-09-14 (AUDIO−VIDEO SYNC FIX SHIPPED: StartLive now clears the pre-live ring backlog)
|
||||||
|
|
||||||
## Branch / Commit State
|
## Branch / Commit State
|
||||||
|
|
||||||
`main` HEAD = `a62a283`. Ahead of origin by 25 commits. **NOT pushing** — user
|
`main` HEAD = `5ba4d70`. Ahead of origin by **27 commits**. Working tree **dirty**:
|
||||||
ruling: no push until web-overlay transparency AND audio-silence are addressed.
|
the sync fix is implemented + tests green below, HANDOFF/MyMistakes updated — NOT yet
|
||||||
|
committed. **NOT pushing** — the prior no-push ruling (transparency + audio-silence) is
|
||||||
|
clear on the first gate; the second was the sync issue, whose fix is now in the tree
|
||||||
|
but the user has not greenlit a push. On-board next: verify a take, then decide.
|
||||||
|
|
||||||
Uncommitted (all belong to the current work, none committed yet):
|
Heads-up for any reader: the previous HANDOFF (a62a283 era) described webcam/audio/
|
||||||
|
truncation fixes as **uncommitted** — that file was stale on arrival. They were
|
||||||
|
actually already committed:
|
||||||
|
|
||||||
```
|
- `724af14` fix: audio silence (idempotent `AudioSyncDelay.Configure`), webcam gray
|
||||||
M .gitignore
|
block (cbW/cbH default), truncated videos (SceneGraph split/static bake docs + FramePump probe)
|
||||||
M HANDOFF.md (this file)
|
- `5ba4d70` docs: audio-sync offset measured ~367ms w/ 300 offset → "set SyncOffsetMs=0"
|
||||||
M MyMistakes.md (webcam collateral recipe added)
|
|
||||||
M Services/Audio/AudioSyncDelay.cs (audio-silence fix)
|
|
||||||
M Services/Compositor/SceneCompositor.cs (webcam cbW/cbH fix)
|
|
||||||
M Services/MediaCaptureFrameSource.cs (diagnostics REMOVED — clean)
|
|
||||||
M ytLive.Tests/AudioSyncDelayTests.cs (silence regression test, proven both ways)
|
|
||||||
M ytLive.Tests/SceneCompositorTests.cs (webcam regression test, proven both ways)
|
|
||||||
M Services/Encoder/FramePump.cs (slow-render probe, see TRUNCATION below)
|
|
||||||
```
|
|
||||||
|
|
||||||
Build is green (0 warnings); 295 tests pass (audio + compositor + pump suites).
|
Build was green (0 warnings) and 295 tests passing at the prior session end; no code
|
||||||
|
changed this session (analysis only).
|
||||||
|
|
||||||
Audio.SyncOffsetMs changed DB-side to 0 (see AUDIO section) — HANDOFF-only change,
|
## ✅ SHIPPED — pre-live capture ring backlog fix (2026-09-14)
|
||||||
nothing new to compile.
|
|
||||||
|
|
||||||
## ✅ TRANSPARENCY — CLOSED (pushed gate #1)
|
**Root cause:** `AudioMixer.Start()` begins mic and loopback capture at **app startup**
|
||||||
|
for the level meters. The 2-second ring buffers (`_micBuffer` = 96k floats = 2.0s mono;
|
||||||
|
`_loopbackBuffer` = 192k floats = 2.0s stereo) fill continuously with pre-live audio.
|
||||||
|
`StartLive` → `LiveLoopAsync` drained from the ring's tail — the **oldest** sample —
|
||||||
|
so every recorded event landed ~2.0s late in the audio track (fixed lag, scaled with
|
||||||
|
time-since-app-launch, capped at ring depth).
|
||||||
|
|
||||||
User confirmed transparency is fixed (`efe88b8` era). Push gate #1 clears.
|
**Fix:** `AudioMixer.StartLive()` now runs `_micBuffer.Clear(); _loopbackBuffer.Clear();`
|
||||||
|
immediately after the pipe starts, before the drain task runs — the in-flight ≤10ms
|
||||||
|
chunk loss is imperceptible and correct ("recording begins at go-live").
|
||||||
|
|
||||||
## ✅ WEBCAM GRAY BLOCK — CLOSED
|
**Regression test (the ONE integration test for this change):**
|
||||||
|
`StartLive_DiscardsPreLiveBacklog_SoFirstAudioIsCurrent` in
|
||||||
|
`ytLive.Tests/AudioPipelineTests.cs` — saturates the loopback ring with stale 0.8
|
||||||
|
(simulated pre-live meters), StartLive, then emits fresh 0.2; asserts the wire carries
|
||||||
|
~0.2 (max < 0.3), failing loudly if the stale 0.8 backlog survived the clear.
|
||||||
|
|
||||||
Root cause: `SceneCompositor.BlitContentRaw` defaulted `cbW=cbH=0` when
|
**Test collateral:** the 3 existing pipe-emit tests pre-filled the rings BEFORE
|
||||||
`src.CropBounds` is null — regression from `ed9d7c1`. Every webcam frame (no crop
|
StartLive (relying on the bug as a reservoir). Their emits now land right after
|
||||||
bounds) blitted to an empty 0×0 source rect → nothing drawn where the webcam
|
StartLive (post-clear, pre-connect — the pipe drops pre-connect writes, so this keeps
|
||||||
should be. Fix: default `cbW=src.Width`, `cbH=src.Height`. Regression test
|
the ring primed for the client) with the 2026-09-14 comments updated. All passed.
|
||||||
`BlitContentRaw_NoCropBounds_SamplesAcrossFullSource_NotJustOrigin` proven BOTH
|
|
||||||
ways (fails on old code, passes on new). User confirmed correct webcam render in
|
|
||||||
the latest recording. All temp diagnostics stripped from
|
|
||||||
`MediaCaptureFrameSource.cs` and `SceneCompositor.cs`.
|
|
||||||
|
|
||||||
## 🟡 AUDIO SILENCE — ROOT-CAUSED AND FIXED, pending final user take
|
Verified: `AudioPipelineTests` 26/26 green. Room-native build 0 warnings. Full gate
|
||||||
|
pending (verify.sh clean build + full suite + scope check).
|
||||||
|
|
||||||
**Symptom:** mic/desktop levels showed real values in preview and in the
|
### Read/Write semantics confirmed (AudioRingBuffer.cs)
|
||||||
per-5s `Audio live:` telemetry, but `peakMix` stayed exactly `0.000` and saved
|
|
||||||
recordings were digital silence (-91 dB, verified with `ffprobe -af
|
|
||||||
volumedetect`).
|
|
||||||
|
|
||||||
**Root cause:** `AudioMixer.FillAndMix` calls `_syncDelay.Configure(...)` on
|
Thread-safe via `_sync`: Write evicts oldest when full; Read returns from
|
||||||
EVERY ~10ms tick (it re-reads the live UI setting each mix). The old
|
`tail = (_head - _count + C) % C` — when full `tail == _head`, so every drained sample
|
||||||
`AudioSyncDelay.Configure` unconditionally reallocated and zeroed the delay
|
is one lap behind the last write. `Clear()` is locked; it is the right seam.
|
||||||
buffer AND reset `_writePos=0` on every call — discarding the just-written
|
|
||||||
audio before the delay offset could ever read it back. With offset 0 the early
|
|
||||||
passthrough masked it; any non-zero saved offset = total silence. The DB has
|
|
||||||
`Audio.SyncOffsetMs = 500` (left over from TASK 22 testing) → that's what
|
|
||||||
triggered it.
|
|
||||||
|
|
||||||
**Fix:** `Configure` now early-returns when the delay samples haven't actually
|
**Root cause (confirmed two independent takes):** `AudioMixer.Start()` begins mic and
|
||||||
changed (still reallocates/flushes on a *change*, so the live slider still
|
loopback capture at **app startup** for the level meters. The 2-second ring buffers
|
||||||
clicks rather than smearing). Regression test
|
(`_micBuffer` = 96k floats = 2.0s mono; `_loopbackBuffer` = 192k floats = 2.0s stereo)
|
||||||
`RepeatedConfigureWithSameDelay_EveryTick_StillDelivers_TheMarker` mirrors the
|
fill continuously with pre-live audio. `StartLive` → `LiveLoopAsync` drains from the
|
||||||
mixer's exact calling pattern and was proven BOTH ways (marker lost = silence
|
ring's tail — the **oldest** sample, which is the moment-of-go-live sample minus the
|
||||||
on old code; marker survives → expected 960-sample shift on new).
|
full ring depth — so every recorded event appears ~2.0s late in the audio track. The
|
||||||
|
lag is **fixed** for the entire recording and scales linearly with "how long the app was
|
||||||
|
open before you hit record", capped at the 2.0s ring depth (plus the configured sync
|
||||||
|
offset plus ~0.1s encoder pipeline).
|
||||||
|
|
||||||
**Telemtery proof the fix works:** latest take (ty-20260912-1135-0000-2.mp4,
|
**Read/Write semantics confirmed** in `AudioRingBuffer.cs` (thread-safe via `_sync`):
|
||||||
recorded WITH the fix, sync offset still 500): `peakMix` = 0.434–0.523 across
|
Write evicts oldest samples when full (`_count == _buffer.Length`); Read returns from
|
||||||
all five 5s windows (was 0.000 before). **Audio is NOT silence anymore.**
|
`tail = (_head - _count + C) % C` — when full, `tail == _head`, meaning every drained
|
||||||
|
sample is exactly one lap behind what was last written. `Clear()` exists and is locked;
|
||||||
|
it is the right seam.
|
||||||
|
|
||||||
**Follow-up:** sync offset `Audio.SyncOffsetMs` measured and SET TO **0** (2026-09-12, take ty-...-1358):
|
**Only one Clear call exists** in the capture path: `RestartMic()` (AudioMixer.cs:132)
|
||||||
|
clears `_micBuffer` on device swap. **Neither buffer is cleared at `StartLive`** — this
|
||||||
|
is the bug.
|
||||||
|
|
||||||
Measurement method tri-wins instead of clap deltas: `silencedetect` on the audio
|
**Taking:** the lag also explains the earlier 13:58 take that measured only ~367ms
|
||||||
gave the count-word starts (one≈0, two@4.12, three@6.48 — audio PTS, both streams
|
with 300ms offset active: the app had just been restarted for that take, so the rings
|
||||||
start at 0, no mux offset), and per-frame video motion (`tblend=difference` +
|
were only partially pre-filled (lag = min(2.0s, time-since-app-launch)). Two takes
|
||||||
`signalstats YAVG`) cross-correlated against the audio envelope peaks at shift
|
hours later (rings fully saturated) reproduce the full ~2.0s backlog consistently.
|
||||||
**+22 frames → audio LAGS video by ~367ms with the 300ms offset active** → the
|
|
||||||
pipeline's own lag is only ~70ms. The saved 300ms (TASK 22 leftover) is what made
|
|
||||||
it read as "one word late" (370ms ≈ a third of the ~2.3s pause ≈ perceptually a
|
|
||||||
full beat behind). Fix = offset 0. Expected residual ≤~70ms (imperceptible).
|
|
||||||
|
|
||||||
## ✅ TRUNCATED VIDEO — ROOT-CAUSED AND FIXED (for baked/static scenes)
|
---
|
||||||
|
|
||||||
**Symptom:** recordings ran ~1s short at first, then HALF-length: 22.8s wall →
|
**Take 1** (ty-20260912-1923-0000-2.mp4, talking + 3 claps):
|
||||||
10.48s file. Not a mux/`-shortest` cut — the pump simply produces frames slower
|
- Audio 17.25s, video 17.05s (1023 frames), no frame drops — NOT truncated.
|
||||||
than the declared 60fps.
|
- Per-clap offsets: +2.16 / +2.18 / +2.15s (audio later). Cross-correlation peak:
|
||||||
|
**+133 frames (+2.217s)**, sharp and unique. Fixed offset, no drift.
|
||||||
|
- DB: `Audio.SyncOffsetMs = 300` (not 0 as the 5ba4d70 commit message claimed —
|
||||||
|
the write never actually landed).
|
||||||
|
|
||||||
**Mechanism (proven):** ffmpeg rawvideo timestamps frames by ARRIVAL at declared
|
**Take 2** (ty-20260914-1104-0000-2.mp4, clean clap test):
|
||||||
`-r 60`. Pump deadline pacing (OBS libobs pattern) yields each frame when its
|
- Audio 22.31s, video 22.10s (1326 frames), no frame drops.
|
||||||
~17ms slot lands; if the render stage for that tick takes longer than the slot,
|
- Audio clap peak at **14.04s**; video motion peak at **11.93s** (yavg 2.75 vs
|
||||||
the pump can only do ~1 frame per 35ms → ~27–28fps of *submitted* frames → file
|
background ~0.1–0.5). Offset: **+2.11s** — confirms the same mechanism.
|
||||||
length ≈ submittedFrames/60. So a take that *should* be 22.8s muxes to ~10.5s,
|
- Tiny pre-clap blips at 0.2/1.5/4.5/13.5s are faint artifacts (< 100 RMS);
|
||||||
and logs show `avg render 35–38ms` with 136/300 per 5s. FramePump had NO way to
|
the 11617-RMS clap is unambiguous.
|
||||||
tell which compositor pass ate a slow frame.
|
|
||||||
|
|
||||||
**Root cause of the 35ms:** the scene was NOT actually fully static. Any dynamic
|
**Theory of the "cut off at the end":** audio events land ~2s late; when the video
|
||||||
element still in the scene — **including hidden ones**, which still count as
|
ends the audio track is still playing the final 2s of real-time content, giving the
|
||||||
Dynamic in `SceneGraph.GetSplitPoint` — forces the per-frame composite path. If
|
impression of a cut. The audio track's slightly longer duration (17.25 vs 17.05;
|
||||||
it sits at split 0, `GetBakedBase` returns null and every tick is a FULL render
|
22.31 vs 22.10) is the tail of that delayed signal — real audio outlives the
|
||||||
(no baked base). The 12:59/13:39 takes logged `resolve 1.0ms` per frame — a live
|
pump-stopped video.
|
||||||
element was still resolving, confirming dynamic content was present.
|
|
||||||
|
|
||||||
**The fix that matters (no code changed in the renderer):** a truly static scene
|
**Theories RULED OUT by data:**
|
||||||
renders in <1ms. New slow-render probe (`ProbeRender` in FramePump.cs) names the
|
- ❌ Frame truncation / slow full-render — video is real-time; frame counts match
|
||||||
path; the 13:53 take of a 1-element static scene logs
|
wall-clock (FramePump log: 301/300 per 5s, dropped 0, avg render 14ms).
|
||||||
`render=fully-static split=1 dynamic=0 bakeMs=246 totalMs=246` on the FIRST
|
- ❌ Named-pipe-connect drop — would *shorten* audio (events earlier), opposite sign.
|
||||||
frame (cold raster of the background at 1080p), then `0.8→0.0ms` avg render,
|
Pipe connected near-instantly in both takes.
|
||||||
`301/300` frames per 5s, 0 stalls → **12.46s file from 12.29s wall, 735 frames
|
- ❌ AudioSyncDelay alone — capped at 0..500ms; 300ms is a subset, not the whole.
|
||||||
@ 60fps. No truncation.**
|
- ❌ Audio drift/resampler rate — offset is constant, not growing.
|
||||||
|
|
||||||
**Rule for recording:** every scene used for RECORDING should be ≥1 static
|
## Open threads
|
||||||
element first (background image) with all web/cam/chat widgets AFTER it in
|
|
||||||
layer order (split ≥ 1 → bake feasible) — a scene whose FIRST layer is dynamic
|
|
||||||
at recording time will truncate. If a take with widgets comes back truncated
|
|
||||||
again, the probe now points at whether it took `fully-static` (bake), or
|
|
||||||
`base+layers` / `full-render` (composite) and how long.
|
|
||||||
|
|
||||||
**Probe left in place** (≈0 overhead: one Stopwatch, string built only when a
|
- **Audio-silence verification** — FIXED code (idempotent `Configure`, committed
|
||||||
render ≥20ms): it is the diagnostic that named this bug; keep it until the
|
`724af14`); the 19:23 take had real peakMix. Final user confirmation still pending.
|
||||||
dynamic-scene path is proven fast too, then remove.
|
- **Audio sync offset DB value** — `Audio.SyncOffsetMs = 300` (the "set to 0" write
|
||||||
|
in 5ba4d70 never actually landed). Fix for the ~2s ring backlog is now SHIPPED (local,
|
||||||
## 🔴 OPEN — user-reported "video seems truncated ~1s" (unconfirmed)
|
uncommitted). Next: re-record the clap pattern, run the cross-correlation recipe → expect
|
||||||
|
residual ≈ 400ms (300ms configured sync + ~100ms pipeline); then optionally zero the
|
||||||
Recording ty-20260912-1135-0000-2.mp4: format duration 26.370s, video 26.167s
|
slider to hit ~100ms.
|
||||||
(1570 frames @ 60fps), audio 26.370s. FramePump stats for the whole take show
|
- **Truncation with DYNAMIC scenes** — proven mechanism + probe in place; static
|
||||||
`dropped 0` and 300/300 frames every 5s window. End-frames (n=1500 vs 1560 vs
|
scenes full-length. 19:23 take confirms dynamic scenes hold ~60fps (no truncation)
|
||||||
1569) differ — not a frozen/repeat-last-frame truncation. Log shows recording
|
but with 30ms stall cadence.
|
||||||
window 11:35:01.386 → 11:35:27.562 = 26.2s of video: matches the file. No
|
- **Webcam MJPG missing**, **web capture ~10–14Hz**, **layer SortOrder not persisting**
|
||||||
evidence of truncation found in data. Possibly user perceived the 500ms audio
|
— queued, unchanged.
|
||||||
delay (audio trailing video start / leading video end) as truncation — worth
|
|
||||||
re-checking with offset zeroed. If a real freeze/truncation reappears: pull
|
|
||||||
last-3-frames diffs and confirm_frame timestamps next.
|
|
||||||
|
|
||||||
## Other threads (paused)
|
|
||||||
|
|
||||||
- **Audio-silence verification** — the FIXED take has real peakMix; final =
|
|
||||||
user re-records with the finalized offset and confirms audible audio in playback.
|
|
||||||
- **Audio sync offset final value** — set to 0 via DB (measured residual ~70ms).
|
|
||||||
NEEDS user re-take to confirm (the running app may still hold 300 in memory —
|
|
||||||
if unsynced on the re-take, use the UI slider to 0, or restart the app BEFORE
|
|
||||||
recording so the 0 loads).
|
|
||||||
- **Truncation with DYNAMIC scenes** — proven mechanism; static scenes now
|
|
||||||
record full-length. The 13:58 take was a 6-element/4-dynamic scene
|
|
||||||
(`split=0 → full-render`, ~34ms, 20 stalls/5s): fully full-length (10.98s)
|
|
||||||
but JITTERY (judder from the 34ms hard frames). If a widget-heavy scene reads
|
|
||||||
badly again, the ProbeRender log names the stage; the next move is getting a
|
|
||||||
baked base by reordering static elements first or optimizing full-render.
|
|
||||||
- **Webcam MJPG missing** — "MJPG negotiation refused (being used by another
|
|
||||||
process)". Queued.
|
|
||||||
- **Web capture speed** — ~10-14Hz effective. Revisit only on request.
|
|
||||||
- **Layer order** — SortOrder not persisting on drag. Queued.
|
|
||||||
|
|
||||||
## Landmines
|
## Landmines
|
||||||
|
|
||||||
@@ -154,8 +128,16 @@ last-3-frames diffs and confirm_frame timestamps next.
|
|||||||
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL form double-slashes mangle) before rebuilds.
|
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL form double-slashes mangle) before rebuilds.
|
||||||
- Build/tests: Windows dotnet host (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings.
|
- Build/tests: Windows dotnet host (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings.
|
||||||
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/ffmpeg.exe` with Windows paths.
|
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/ffmpeg.exe` with Windows paths.
|
||||||
- `Audio.SyncOffsetMs` persisted non-zero WILL re-trigger silence symptoms if
|
- **`Audio.SyncOffsetMs` currently = 300 in the DB** (contradicts the "set to 0"
|
||||||
`Configure` is ever called unconditionally again — the idempotent guard is the
|
line in 5ba4d70; the write likely never happened). Non-zero offsets still depend on
|
||||||
fix, keep it.
|
the idempotent `Configure` guard to avoid the silence bug.
|
||||||
- `MyMistakes.md` has the transparency RESOLVED entry and the new webcam
|
- `MyMistakes.md` → Recipes registry now has the **audio/video sync measurement
|
||||||
collateral recipe — grep before any re-derivation.
|
recipe** (claps + cross-correlation) — grep before re-deriving.
|
||||||
|
|
||||||
|
## Next step
|
||||||
|
|
||||||
|
**Verify the shipped fix:** re-record with the same clap pattern and run the
|
||||||
|
cross-correlation recipe from `MyMistakes.md` → expect offset ≈ 400ms (300ms configured
|
||||||
|
sync + ~100ms pipeline) with a `SyncOffsetMs=300` DB value, or ~100ms with the slider
|
||||||
|
zeroed. Then commit this session's work (StartLive clear + regression test + docs) and
|
||||||
|
decide with the user whether the no-push ruling is lifted.
|
||||||
@@ -174,6 +174,47 @@ packages, high-quality downscale via `TransformedBitmap`. Screenshots compress
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
### Measuring audio-video A/V sync from a recording (no ears needed) — RECIPE
|
||||||
|
|
||||||
|
Derived 2026-09-12/14 (the "audio cut off at the end" / "audio delayed" reports). You
|
||||||
|
cannot "listen" to a take — measure it. Proven on two independent takes (talking+claps,
|
||||||
|
clean-clap test) with a tight, matching result each time.
|
||||||
|
|
||||||
|
**Method (FFmpeg probe + numpy, run from WSL):**
|
||||||
|
|
||||||
|
1. **Click/hiss-SILENT test takes are the gold standard.** Have the creator do a
|
||||||
|
loud, single-frame-syncable event (one clap after ~10s of near-silence) — the
|
||||||
|
video position of the hands-meet peak and the audio position of the transient
|
||||||
|
are both unambiguous. That single event replaced counting words forever.
|
||||||
|
2. **Extract audio as raw mono PCM** (`ffmpeg -vn -ac 1 -ar 48000 -f s16le`) and
|
||||||
|
compute a short window RMS envelope (5ms windows) in numpy. The clap is the global
|
||||||
|
max (`argmax`); also print a coarse 100ms table — it shows the pre-clap artifacts
|
||||||
|
(faint blips) vs the event (100× larger) vs true digital silence.
|
||||||
|
3. **Video motion per frame** — `-vf "tblend=all_mode=difference,signalstats,metadata=print"`
|
||||||
|
captured to stdout (NOT `file=` inside the filter — the `\\` path escapes mangle;
|
||||||
|
have PowerShell-style quoting bite; `metadata=print` writes to ffmpeg's log, so
|
||||||
|
redirect stdout to a file). Parse `YAVG` values; the clap is the frame with the
|
||||||
|
motion spike (2.75 vs a ~0.1–0.5 animated-widget background).
|
||||||
|
4. **Offset = audio_event_time − video_event_time**; every positive second means audio
|
||||||
|
is BEHIND video by that much. Cross-correlate the two full envelopes (60fps-resampled
|
||||||
|
RMS vs YAVG, normalized) to double-check the single-event peak — the autocorr-style
|
||||||
|
peak must be sharp (unique max at +133 frames, second-best ≈0.26).
|
||||||
|
5. **Sanity checks baked in:** count frames (`nb_frames` vs wall log) — if video is
|
||||||
|
real-time you've ruled out the truncation bug as the cause; compare stream
|
||||||
|
durations (audio > video by the lag is the SIGNATURE of a delayed audio tail, not
|
||||||
|
truncation); check `dropped writes` — dropped audio moves events EARLIER (opposite
|
||||||
|
sign), so it can never explain a lag. **A constant offset ≠ drift:** fixed lag =
|
||||||
|
buffering/backlog, growing offset = clock mismatch.
|
||||||
|
|
||||||
|
**The ROOT CAUSE the recipe led to (record it so it's never re-derived):** the mixer's
|
||||||
|
capture rings are fed from APP STARTUP (for the meters); `StartLive` never clears them,
|
||||||
|
so the drained audio trail begins ~a full ring-depth (2.0s) behind go-live. **Rule: any
|
||||||
|
"live preview/drain" sink fed by a continuously-capturing buffer MUST clear the buffer
|
||||||
|
at go-live, or early output is stale backlog.** The 2s ring + configured 300ms sync
|
||||||
|
delay + ~100ms pipeline = exactly the measured 2.11–2.22s. The lag ALSO scales with
|
||||||
|
"time since app launch" up to the ring depth — a take right after a restart reads as
|
||||||
|
~370ms while a fully-warmed app reads ~2.2s. Same code, wildly different numbers.
|
||||||
|
|
||||||
### Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
|
### Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
|
||||||
|
|
||||||
Derived 2026-09-03 (take-3 diagnosis — the stats seam from `97ffc42` named the stage
|
Derived 2026-09-03 (take-3 diagnosis — the stats seam from `97ffc42` named the stage
|
||||||
|
|||||||
@@ -159,6 +159,15 @@ public sealed class AudioMixer : IDisposable
|
|||||||
if (_liveCts != null)
|
if (_liveCts != null)
|
||||||
return;
|
return;
|
||||||
|
|
||||||
|
// Capture runs from app STARTUP for the meters, so both rings carry up
|
||||||
|
// to a full ring depth (~2s) of pre-live backlog by the time the user
|
||||||
|
// hits record; draining it made every recorded event land ~2s late in
|
||||||
|
// the audio track (2026-09-14, two confirmed takes). Discard it so the
|
||||||
|
// recording starts at go-live. The writers keep running concurrently;
|
||||||
|
// Clear under the ring's lock, so only the in-flight chunk is dropped.
|
||||||
|
_micBuffer.Clear();
|
||||||
|
_loopbackBuffer.Clear();
|
||||||
|
|
||||||
var cts = new CancellationTokenSource();
|
var cts = new CancellationTokenSource();
|
||||||
_liveCts = cts;
|
_liveCts = cts;
|
||||||
_pipe = new NamedPipeAudioWriter();
|
_pipe = new NamedPipeAudioWriter();
|
||||||
|
|||||||
@@ -295,9 +295,15 @@ public class AudioPipelineTests
|
|||||||
loopbackGain: () => 1.0,
|
loopbackGain: () => 1.0,
|
||||||
mixInterval: TimeSpan.FromMilliseconds(5));
|
mixInterval: TimeSpan.FromMilliseconds(5));
|
||||||
|
|
||||||
|
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
||||||
|
var connected = client.ConnectAsync();
|
||||||
|
mixer.StartLive(pipeName);
|
||||||
|
|
||||||
// Loud mic (filtered + compressed toward ~0.5) + loud loopback (ducked
|
// Loud mic (filtered + compressed toward ~0.5) + loud loopback (ducked
|
||||||
// from unity down toward 0.25). Pre-fill the ring buffers so the very
|
// from unity down toward 0.25). Emitted right after StartLive — the
|
||||||
// first tick written to the pipe already carries real audio.
|
// rings are cleared at go-live so pre-live backlog never reaches the
|
||||||
|
// wire (the 2026-09-14 ~2s sync fix); dripping into the connected pipe
|
||||||
|
// is dropped, so the ring stays primed for the client.
|
||||||
var micChunk = new float[240];
|
var micChunk = new float[240];
|
||||||
Array.Fill(micChunk, 0.5f);
|
Array.Fill(micChunk, 0.5f);
|
||||||
var loopChunk = new float[480];
|
var loopChunk = new float[480];
|
||||||
@@ -308,11 +314,7 @@ public class AudioPipelineTests
|
|||||||
loopback.Emit(new AudioSample(loopChunk, 48000, 2));
|
loopback.Emit(new AudioSample(loopChunk, 48000, 2));
|
||||||
}
|
}
|
||||||
|
|
||||||
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
|
||||||
var connected = client.ConnectAsync();
|
|
||||||
mixer.StartLive(pipeName);
|
|
||||||
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
||||||
|
|
||||||
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
||||||
mixer.StopLive();
|
mixer.StopLive();
|
||||||
|
|
||||||
@@ -352,20 +354,21 @@ public class AudioPipelineTests
|
|||||||
loopbackGain: provider.LoopbackGain,
|
loopbackGain: provider.LoopbackGain,
|
||||||
mixInterval: TimeSpan.FromMilliseconds(5));
|
mixInterval: TimeSpan.FromMilliseconds(5));
|
||||||
|
|
||||||
|
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
||||||
|
var connected = client.ConnectAsync();
|
||||||
|
mixer.StartLive(pipeName);
|
||||||
|
|
||||||
// Mic silent (so the ducker stays at unity) + a loud loopback bed that
|
// Mic silent (so the ducker stays at unity) + a loud loopback bed that
|
||||||
// scales to 0.8 * 0.5 = 0.4 on the wire. Pre-fill the ring buffer so the
|
// scales to 0.8 * 0.5 = 0.4 on the wire. Emitted right after StartLive —
|
||||||
// first frames already carry real audio.
|
// the 2026-09-14 fix clears pre-live ring backlog at go-live, so this is
|
||||||
|
// the fresh bed the client pulls from.
|
||||||
var loopChunk = new float[480];
|
var loopChunk = new float[480];
|
||||||
Array.Fill(loopChunk, 0.8f);
|
Array.Fill(loopChunk, 0.8f);
|
||||||
for (var i = 0; i < 200; i++)
|
for (var i = 0; i < 200; i++)
|
||||||
loopback.Emit(new AudioSample(loopChunk, 48000, 2));
|
loopback.Emit(new AudioSample(loopChunk, 48000, 2));
|
||||||
|
|
||||||
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
|
||||||
var connected = client.ConnectAsync();
|
|
||||||
mixer.StartLive(pipeName);
|
|
||||||
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
|
||||||
|
|
||||||
// Phase 1: unmuted at 0.5 → scaled but clearly audible, bounded.
|
// Phase 1: unmuted at 0.5 → scaled but clearly audible, bounded.
|
||||||
|
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
||||||
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
||||||
var floats = new float[bytes.Length / 4];
|
var floats = new float[bytes.Length / 4];
|
||||||
Buffer.BlockCopy(bytes, 0, floats, 0, bytes.Length);
|
Buffer.BlockCopy(bytes, 0, floats, 0, bytes.Length);
|
||||||
@@ -401,19 +404,20 @@ public class AudioPipelineTests
|
|||||||
loopbackGain: () => 1.0,
|
loopbackGain: () => 1.0,
|
||||||
mixInterval: TimeSpan.FromMilliseconds(5));
|
mixInterval: TimeSpan.FromMilliseconds(5));
|
||||||
|
|
||||||
|
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
||||||
|
var connected = client.ConnectAsync();
|
||||||
|
mixer.StartLive(pipeName);
|
||||||
|
|
||||||
// Mic silent (the ducker stays at unity) + a 0.95 loopback bed — hotter
|
// Mic silent (the ducker stays at unity) + a 0.95 loopback bed — hotter
|
||||||
// than the -1 dBFS ceiling, so every tick must be trimmed. Pre-fill the
|
// than the -1 dBFS ceiling, so every tick must be trimmed. Emitted right
|
||||||
// ring buffer so the first written frame is already loud.
|
// after StartLive — the 2026-09-14 fix clears pre-live ring backlog at
|
||||||
|
// go-live, so this is the fresh bed the client pulls from.
|
||||||
var loopChunk = new float[480];
|
var loopChunk = new float[480];
|
||||||
Array.Fill(loopChunk, 0.95f);
|
Array.Fill(loopChunk, 0.95f);
|
||||||
for (var i = 0; i < 200; i++)
|
for (var i = 0; i < 200; i++)
|
||||||
loopback.Emit(new AudioSample(loopChunk, 48000, 2));
|
loopback.Emit(new AudioSample(loopChunk, 48000, 2));
|
||||||
|
|
||||||
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
|
||||||
var connected = client.ConnectAsync();
|
|
||||||
mixer.StartLive(pipeName);
|
|
||||||
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
||||||
|
|
||||||
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
||||||
mixer.StopLive();
|
mixer.StopLive();
|
||||||
|
|
||||||
@@ -428,6 +432,52 @@ public class AudioPipelineTests
|
|||||||
Assert.True(floats.Any(f => MathF.Abs(f) > 0.5f), "the capped mix must still be audible");
|
Assert.True(floats.Any(f => MathF.Abs(f) > 0.5f), "the capped mix must still be audible");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
[Fact]
|
||||||
|
public async Task StartLive_DiscardsPreLiveBacklog_SoFirstAudioIsCurrent()
|
||||||
|
{
|
||||||
|
// THE integration test for the 2026-09-14 sync fix: capture runs from
|
||||||
|
// app STARTUP for the meters, so by the time the user hits record both
|
||||||
|
// rings are full of ~2s of pre-live audio. Draining that backlog put
|
||||||
|
// every recorded event ~2s late in the audio track. StartLive must
|
||||||
|
// discard it so the recording begins at go-live.
|
||||||
|
var pipeName = "ytllive_test_" + Guid.NewGuid().ToString("N");
|
||||||
|
var mic = new FakeSource();
|
||||||
|
var loopback = new FakeSource();
|
||||||
|
using var mixer = new AudioMixer(
|
||||||
|
mic, loopback,
|
||||||
|
log: null,
|
||||||
|
mixInterval: TimeSpan.FromMilliseconds(5));
|
||||||
|
|
||||||
|
// Pre-live: saturate the loopback ring (0.8 — a full ring depth) with
|
||||||
|
// the OLD code this is exactly the backlog that leaked onto the wire.
|
||||||
|
var stale = new float[480];
|
||||||
|
Array.Fill(stale, 0.8f);
|
||||||
|
for (var i = 0; i < 500; i++)
|
||||||
|
loopback.Emit(new AudioSample(stale, 48000, 2));
|
||||||
|
|
||||||
|
using var client = new NamedPipeClientStream(".", pipeName, PipeDirection.In, PipeOptions.Asynchronous);
|
||||||
|
var connected = client.ConnectAsync();
|
||||||
|
mixer.StartLive(pipeName);
|
||||||
|
|
||||||
|
// Post-live: fresh audio at 0.2 — the loopback path has no voice chain,
|
||||||
|
// unity gain, ducker silent, limiter quiet: the wire carries it nearly
|
||||||
|
// 1:1. If the stale 0.8 survived the clear, it would dominate.
|
||||||
|
var fresh = new float[480];
|
||||||
|
Array.Fill(fresh, 0.2f);
|
||||||
|
for (var i = 0; i < 50; i++)
|
||||||
|
loopback.Emit(new AudioSample(fresh, 48000, 2));
|
||||||
|
|
||||||
|
await connected.WaitAsync(TimeSpan.FromSeconds(5));
|
||||||
|
var bytes = await ReadFullyAsync(client, 240 * 2 * 4, TimeSpan.FromSeconds(5));
|
||||||
|
mixer.StopLive();
|
||||||
|
|
||||||
|
var floats = new float[bytes.Length / 4];
|
||||||
|
Buffer.BlockCopy(bytes, 0, floats, 0, bytes.Length);
|
||||||
|
Assert.True(floats.Length > 0, "expected wired audio");
|
||||||
|
Assert.Equal(0.2f, floats.Max(), 3);
|
||||||
|
Assert.True(floats.Max() < 0.3f, "stale pre-live backlog leaked onto the wire");
|
||||||
|
}
|
||||||
|
|
||||||
private static async Task<byte[]> ReadFullyAsync(NamedPipeClientStream client, int count, TimeSpan timeout)
|
private static async Task<byte[]> ReadFullyAsync(NamedPipeClientStream client, int count, TimeSpan timeout)
|
||||||
{
|
{
|
||||||
var ms = new MemoryStream();
|
var ms = new MemoryStream();
|
||||||
|
|||||||
Reference in New Issue
Block a user