perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)
The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was the wall, and the one torn frame was a ring slot rewritten under the consumer's read. Reference: WGC delivers at DWM/monitor cadence (https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture) and libyuv row-simple/fixed-point scaling (https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule. - DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame. - 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of serialized conversion/s; capacity sits just above the 60/s the pump can use. - FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a consumer still holds (session.LatestFrame survives conversions, dispatcher preview lags) is never read-while-overwritten. Needs no consumer Release API. - 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring allocs — the device take is judgeable numerically. Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings. C4 (composite Epoch-cached downscale) deferred pending the device re-measure. Local only, no push.
This commit is contained in:
@@ -342,6 +342,67 @@ single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late)
|
||||
the acceleration; re-measure after slice 15.
|
||||
|
||||
|
||||
**Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
|
||||
the freeze-audit recipe (RECIPE — re-deriving it cost this session):**
|
||||
The slice-15 build fixed pacing but the desktop layer of the recording was still
|
||||
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
|
||||
measurement said something different AND worse than the running render theory.
|
||||
`FramePump stall… worst render 33-36ms` was real but MOOT: once the camera+desktop
|
||||
take was decoded to raw frames, the **desktop band was frozen 21s of 23.35s (90%)**
|
||||
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
|
||||
CONVERSION was the wall: the monitor delivers at the **240Hz DWM cadence**, the
|
||||
source converts ONE frame at a time (`_framePending` latest-wins), and each
|
||||
2560×1440→1920×1080 `DownscaleBgra` — naive double-per-pixel bilinear — cost
|
||||
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
|
||||
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
|
||||
MediaCapture path) and audio were fine — exactly what the user reported.
|
||||
|
||||
**The audit recipe (ffmpeg → raw gray → numpy):**
|
||||
```
|
||||
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
|
||||
python3 - <<EOF
|
||||
import numpy as np
|
||||
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
|
||||
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
|
||||
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
|
||||
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
|
||||
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
|
||||
EOF
|
||||
```
|
||||
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
|
||||
the compositor render stats led nowhere until this number existed. A **per-row
|
||||
split detector** (`cumsum` of per-row diff-to-next minus diff-to-prev, argmax =
|
||||
split row) then separated real mid-frame tears (score ≈ huge, both halves match
|
||||
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
|
||||
were the session UI, not tears.
|
||||
|
||||
**The fix (this slice):**
|
||||
- **integer 8.8 fixed-point downscale, "shift only at the end"** — the SAME math as
|
||||
`SceneCompositor.Bilinear` (rounded both stages in one 16.8 scale) ported into
|
||||
`DownscaleBgra`, dropping per-pixel doubles to row-walk integer ops. The capture
|
||||
ring already had the integer-bilinear lesson; the capture downscale itself was
|
||||
still the naive float twin of the 258ms disaster.
|
||||
- **throttle to the slot cadence** (`MinConvertInterval = 10ms`): the 240Hz arrival
|
||||
is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per
|
||||
second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps
|
||||
pump can use.
|
||||
- **ring reuse-distance, not ownership** (`FrameRingBuffer`, redLine 4): the
|
||||
take-14 "depth × period" rule guards size; the slice-16 addition makes it
|
||||
structural — a slot is only rewritten ≥4 rents after its last hand-out, else a
|
||||
fresh buffer. `session.LatestFrame` survives across conversions and the
|
||||
dispatcher preview copy lags, so "who released it" is unknowable without a
|
||||
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
|
||||
tear (new-top/old-bottom midway) is that read-under-write closed.
|
||||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||||
relocating a 30ms float downscale from the capture thread to the render thread
|
||||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||
op fast.
|
||||
|
||||
---
|
||||
|
||||
|
||||
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
|
||||
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
|
||||
integer bilinear) followed the same libyuv research
|
||||
|
||||
Reference in New Issue
Block a user