perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)

The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel
DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take
ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame
audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was
the wall, and the one torn frame was a ring slot rewritten under the consumer's
read. Reference: WGC delivers at DWM/monitor cadence
(https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture)
and libyuv row-simple/fixed-point scaling
(https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule.

- DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage
  math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame.
- 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of
  serialized conversion/s; capacity sits just above the 60/s the pump can use.
- FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only
  rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a
  consumer still holds (session.LatestFrame survives conversions, dispatcher
  preview lags) is never read-while-overwritten. Needs no consumer Release API.
- 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring
  allocs — the device take is judgeable numerically.

Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings.
C4 (composite Epoch-cached downscale) deferred pending the device re-measure.
Local only, no push.
This commit is contained in:
2026-09-14 18:18:34 -07:00
parent 76f51e6f4e
commit 71932b9756
6 changed files with 417 additions and 110 deletions
+61
View File
@@ -342,6 +342,67 @@ single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late)
the acceleration; re-measure after slice 15.
**Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
the freeze-audit recipe (RECIPE — re-deriving it cost this session):**
The slice-15 build fixed pacing but the desktop layer of the recording was still
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
measurement said something different AND worse than the running render theory.
`FramePump stall… worst render 33-36ms` was real but MOOT: once the camera+desktop
take was decoded to raw frames, the **desktop band was frozen 21s of 23.35s (90%)**
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
CONVERSION was the wall: the monitor delivers at the **240Hz DWM cadence**, the
source converts ONE frame at a time (`_framePending` latest-wins), and each
2560×1440→1920×1080 `DownscaleBgra` — naive double-per-pixel bilinear — cost
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
MediaCapture path) and audio were fine — exactly what the user reported.
**The audit recipe (ffmpeg → raw gray → numpy):**
```
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
python3 - <<EOF
import numpy as np
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
EOF
```
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
the compositor render stats led nowhere until this number existed. A **per-row
split detector** (`cumsum` of per-row diff-to-next minus diff-to-prev, argmax =
split row) then separated real mid-frame tears (score ≈ huge, both halves match
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
were the session UI, not tears.
**The fix (this slice):**
- **integer 8.8 fixed-point downscale, "shift only at the end"** — the SAME math as
`SceneCompositor.Bilinear` (rounded both stages in one 16.8 scale) ported into
`DownscaleBgra`, dropping per-pixel doubles to row-walk integer ops. The capture
ring already had the integer-bilinear lesson; the capture downscale itself was
still the naive float twin of the 258ms disaster.
- **throttle to the slot cadence** (`MinConvertInterval = 10ms`): the 240Hz arrival
is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per
second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps
pump can use.
- **ring reuse-distance, not ownership** (`FrameRingBuffer`, redLine 4): the
take-14 "depth × period" rule guards size; the slice-16 addition makes it
structural — a slot is only rewritten ≥4 rents after its last hand-out, else a
fresh buffer. `session.LatestFrame` survives across conversions and the
dispatcher preview copy lags, so "who released it" is unknowable without a
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
tear (new-top/old-bottom midway) is that read-under-write closed.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
---
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
integer bilinear) followed the same libyuv research