fix(rec): bounded encoder queue + drop policy + burned frame counter — the "1...23...4...56..." smeared-ticker take

slice 9 made the DURATION right but content still hiccuped; aggregates (301/300, uniform
file PTS) could not see it. Measured root cause: FfmpegEncoder.SubmitFrameAsync BLOCKED
on WriteAsync(8.3MB)+FlushAsync when ffmpeg lagged the pipe, and the burst while-loop
re-wrote that same stale composite per crossed slot — frozen runs.

OBS shape (derivative, wrapped pre-1.0): the encoder queue in libobs/obs-encoder.c —
encoder thread never couples back into the video thread; overflow = dropped data, never
a frozen producer. https://github.com/obsproject/obs-studio/blob/master/libobs/obs-encoder.c

- FfmpegEncoder: SubmitFrameAsync is now an enqueue (ArrayPool copy) into a bounded
  Channel (cap 120) drained by its own task; drop-newest + count when full;
  StopAsync flushes the queue then EOF (TryComplete). IFfmpegEncoder.DroppedFrames.
- FramePump: ONE fresh composite per iteration (burst loop deleted); worst-submit stat,
  stall logger (>2x interval names the stage), dropped/stalls in stats.
- Burned-in 6-digit dot-matrix frame counter (white box, bottom-right) on every composite
  — the clock-independent judge replacing the WSL ticker: +1/frame, jumps = counted drops.
- ONE new test Backpressure_QueueOverflow_DropsFrames_AndNeverBlocks (slow-sink fake:
  submit never blocks, drops counted, stop flushes exactly submitted-minus-dropped).
- Full suite 290 tests, 289 pass — sole failure the pre-existing compositor pixel test.
- Docs same-commit: ai.md slice 10 (+ encoder/stop-note corrections), MyMistakes point 8,
  HANDOFF.

Audio untouched (queued follow-up); web overlay still frozen pending timing closure.
This commit is contained in:
2026-09-10 09:32:55 -07:00
parent bd396e488c
commit c45cbc93b7
8 changed files with 414 additions and 79 deletions
+20
View File
@@ -177,6 +177,26 @@ Both halves were solved by OBS/libyuv long ago; do not re-derive:
shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the
NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.
8. **An encoder refed from a real-time loop must ENQUEUE, never pipe-write in the loop
(2026-09-10, take 15/16 — the "1...23...4...56..." smeared ticker).** After slice 9 the
recording played at the right DURATION but the content still hiccuped — and the aggregates
(301/300, uniform file PTS, 15.6s wall vs 15.74s file) could NOT see it. The cause finally
measured in `FfmpegEncoder.SubmitFrameAsync`: the `WriteAsync(8.3MB)+FlushAsync` to ffmpeg's
stdin BLOCKS whenever the encoder lags the pipe, and the slice-9 burst `while (now>=nextTick)`
then re-wrote that SAME stale composite for every slot that ticked past — frozen content runs.
OBS's answering machinery is the encoder queue (`libobs/obs-encoder.c`): the encoder thread
NEVER couples back into the video thread; overflow = dropped data, NEVER a frozen producer.
Fixed as: bounded `Channel<byte[]>` (cap 120) + a dedicated drain task owning stdin,
`SubmitFrameAsync` = copy-to-pool-array + `TryWrite` (drop-newest + count when full), pump
emits ONE fresh composite per iteration (no burst re-write), stop flushes the queue then EOF.
**Rule: verify with a clock-independent judge.** The WSL ticker that "proved" slice 9 has its
own Host-timer jitter under Windows load — so this slice burns a dot-matrix `_outputIndex`
into the bottom-right of every composite; decoding the recording reads the honest sequence
(+1/frame, jumps = counted drops) with no external clock involved. Take 17 must read
+1/frame from that strip. (A whole-frame duplicate scan was tried and is
UNRELIABLE here: the scene is always animating — session elapsed timer + REC pulse — so
no two frames are ever byte-identical.)
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +