docs: audio-sync offset measured (~367ms lag w/ 300 offset) -> set Audio.SyncOffsetMs=0
silencedetect audio word times cross-correlated with per-frame video motion (tblend=difference + signalstats) peaks at +22 frames: the 300ms TASK-22 offset plus ~70ms pipeline lag. 370ms reads as one full beat behind during the ~2.3s pauses. Zeroing the offset targets the residual ~70ms.
This commit is contained in:
+22
-10
@@ -21,6 +21,9 @@ M Services/Encoder/FramePump.cs (slow-render probe, see TRUNCATION below)
|
|||||||
|
|
||||||
Build is green (0 warnings); 295 tests pass (audio + compositor + pump suites).
|
Build is green (0 warnings); 295 tests pass (audio + compositor + pump suites).
|
||||||
|
|
||||||
|
Audio.SyncOffsetMs changed DB-side to 0 (see AUDIO section) — HANDOFF-only change,
|
||||||
|
nothing new to compile.
|
||||||
|
|
||||||
## ✅ TRANSPARENCY — CLOSED (pushed gate #1)
|
## ✅ TRANSPARENCY — CLOSED (pushed gate #1)
|
||||||
|
|
||||||
User confirmed transparency is fixed (`efe88b8` era). Push gate #1 clears.
|
User confirmed transparency is fixed (`efe88b8` era). Push gate #1 clears.
|
||||||
@@ -63,12 +66,16 @@ on old code; marker survives → expected 960-sample shift on new).
|
|||||||
recorded WITH the fix, sync offset still 500): `peakMix` = 0.434–0.523 across
|
recorded WITH the fix, sync offset still 500): `peakMix` = 0.434–0.523 across
|
||||||
all five 5s windows (was 0.000 before). **Audio is NOT silence anymore.**
|
all five 5s windows (was 0.000 before). **Audio is NOT silence anymore.**
|
||||||
|
|
||||||
**Follow-up:** sync offset `Audio.SyncOffsetMs` is still in the DB — currently
|
**Follow-up:** sync offset `Audio.SyncOffsetMs` measured and SET TO **0** (2026-09-12, take ty-...-1358):
|
||||||
**300** (flipped 500→0→500→300 during lip-sync calibration; user judged 0 worse
|
|
||||||
than 500, so it was bisected back up). NOT the recording's natural state. When the
|
Measurement method tri-wins instead of clap deltas: `silencedetect` on the audio
|
||||||
user is ready to finalize, measure on the classic method (clap visible in frame;
|
gave the count-word starts (one≈0, two@4.12, three@6.48 — audio PTS, both streams
|
||||||
read delta on the waveform — the auto-correlation approach failed to converge at
|
start at 0, no mux offset), and per-frame video motion (`tblend=difference` +
|
||||||
~10–14Hz webcam shutter jitter).
|
`signalstats YAVG`) cross-correlated against the audio envelope peaks at shift
|
||||||
|
**+22 frames → audio LAGS video by ~367ms with the 300ms offset active** → the
|
||||||
|
pipeline's own lag is only ~70ms. The saved 300ms (TASK 22 leftover) is what made
|
||||||
|
it read as "one word late" (370ms ≈ a third of the ~2.3s pause ≈ perceptually a
|
||||||
|
full beat behind). Fix = offset 0. Expected residual ≤~70ms (imperceptible).
|
||||||
|
|
||||||
## ✅ TRUNCATED VIDEO — ROOT-CAUSED AND FIXED (for baked/static scenes)
|
## ✅ TRUNCATED VIDEO — ROOT-CAUSED AND FIXED (for baked/static scenes)
|
||||||
|
|
||||||
@@ -126,11 +133,16 @@ last-3-frames diffs and confirm_frame timestamps next.
|
|||||||
|
|
||||||
- **Audio-silence verification** — the FIXED take has real peakMix; final =
|
- **Audio-silence verification** — the FIXED take has real peakMix; final =
|
||||||
user re-records with the finalized offset and confirms audible audio in playback.
|
user re-records with the finalized offset and confirms audible audio in playback.
|
||||||
- **Audio sync offset final value** — DB is at 300 (bisect point). Needs the
|
- **Audio sync offset final value** — set to 0 via DB (measured residual ~70ms).
|
||||||
clap/waveform measurement; do NOT rely on the 500 left over from TASK 22.
|
NEEDS user re-take to confirm (the running app may still hold 300 in memory —
|
||||||
|
if unsynced on the re-take, use the UI slider to 0, or restart the app BEFORE
|
||||||
|
recording so the 0 loads).
|
||||||
- **Truncation with DYNAMIC scenes** — proven mechanism; static scenes now
|
- **Truncation with DYNAMIC scenes** — proven mechanism; static scenes now
|
||||||
record full-length. If a widget-heavy take truncates again, read the
|
record full-length. The 13:58 take was a 6-element/4-dynamic scene
|
||||||
`ProbeRender` path/stage breakdown in startup.log, then optimize composite.
|
(`split=0 → full-render`, ~34ms, 20 stalls/5s): fully full-length (10.98s)
|
||||||
|
but JITTERY (judder from the 34ms hard frames). If a widget-heavy scene reads
|
||||||
|
badly again, the ProbeRender log names the stage; the next move is getting a
|
||||||
|
baked base by reordering static elements first or optimizing full-render.
|
||||||
- **Webcam MJPG missing** — "MJPG negotiation refused (being used by another
|
- **Webcam MJPG missing** — "MJPG negotiation refused (being used by another
|
||||||
process)". Queued.
|
process)". Queued.
|
||||||
- **Web capture speed** — ~10-14Hz effective. Revisit only on request.
|
- **Web capture speed** — ~10-14Hz effective. Revisit only on request.
|
||||||
|
|||||||
Reference in New Issue
Block a user