perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)

The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel
DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take
ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame
audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was
the wall, and the one torn frame was a ring slot rewritten under the consumer's
read. Reference: WGC delivers at DWM/monitor cadence
(https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture)
and libyuv row-simple/fixed-point scaling
(https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule.

- DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage
  math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame.
- 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of
  serialized conversion/s; capacity sits just above the 60/s the pump can use.
- FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only
  rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a
  consumer still holds (session.LatestFrame survives conversions, dispatcher
  preview lags) is never read-while-overwritten. Needs no consumer Release API.
- 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring
  allocs — the device take is judgeable numerically.

Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings.
C4 (composite Epoch-cached downscale) deferred pending the device re-measure.
Local only, no push.
This commit is contained in:
2026-09-14 18:18:34 -07:00
parent 76f51e6f4e
commit 71932b9756
6 changed files with 417 additions and 110 deletions
+51 -5
View File
@@ -324,10 +324,20 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
`WindowsRuntimeMarshal.TryGetDataUnsafe` (the same CsWinRT-safe read the webcam path uses) — the
`IMemoryBufferByteAccess` ComImport cast threw `Invalid cast` on **every frame** under CsWinRT, which
flooded `startup.log` (~5 MB in a session) and burned CPU, so it is gone. Surfaces larger than the
1920×1080 master are downscaled bilinearly to the master (`DownscaleBgra`) before the copy, and
per-frame conversion failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected
content delivers black frames (OS limitation, documented). Frame pool
pauses while the app is minimized — capture keeps running, the pool just stops delivering.
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
buffer is only rewritten ≥4 rents after its last hand-out, else a fresh one is allocated — a frame a
consumer still holds (`session.LatestFrame` survives conversions; the dispatcher preview copy lags)
can never be read-while-overwritten (the 1742 tear). Rolling telemetry (frames/s, conv avg/max ms,
drops, ring allocs) is logged to `startup.log` every 2s while converting. Per-frame conversion
failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected content delivers black
frames (OS limitation, documented). Frame pool pauses while the app is minimized — capture keeps
running, the pool just stops delivering.
- **Ownership:** `ScreenCaptureManager` mirrors `CameraManager` — refcounted by target key, one shared
`WriteableBitmap` per key, dispatcher-coalesced latest-frame copies, `PreviewBitmapChanged`/`CaptureFailed`
events, `ReleaseAllAsync` on re-designation. `ScreenCaptureSourceFactory.Resolve(key)` parses the key into
@@ -353,7 +363,9 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
app is in the background (worse under a full-screen game on 24H2/26100), GPU readback
(`CreateCopyFromSurfaceAsync`) contends with the foreground game, the source's one-in-flight `_framePending`
gate drops frames during stalls, and the manager's `DispatcherPriority.Render` copies only run as fast as
WPF presents the window. Recorded 2026-08-13; no mitigation attempted yet (deferred by user decision).
WPF presents the window. Slice 16 shrank the per-conversion stall (fast downscale + 10ms cadence floor)
but the OS-level delivery throttle under focus loss remains. Recorded 2026-08-13; no mitigation attempted
yet (deferred by user decision).
- **GPU posture:** same as webcam — CPU frames, WPF hardware-presents; D3DImage GPU compositing deferred
to the encoder task.
- **Preview watermark:** the "Preview" placeholder hides while a background capture renders —
@@ -1019,6 +1031,40 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
suite **294/294 green, 0 warnings**. The +0.6s audio-late clap reading on 1726 was confounded by
the 1.1x acceleration — re-measure on device; if a real residual remains it is the audio pipeline.
No push (web/A/V work stays local).
- **Slice 16 — the desktop capture conversion was the bottleneck: fast downscale +
cadence throttle + reuse-distance ring (2026-09-14, device take ty-1742):** slice 15
fixed pacing but the 1742 desktop layer was still jerky/frozen with a horizontal
tear. Decoded to raw frames and audited: the **desktop band was frozen 21s of 23.35s
(90%)** — ~6.1 content updates/s, freeze runs up to 2.28-2.78s, dup-run max 90
frames (1.5s). The render stat (`worst render 33-36ms`) was real but MOOT: the
capture CONVERSION was the wall. The monitor delivers at the **240Hz DWM cadence**
(~4.2ms); with one-in-flight conversions and each 2560×1440→1920×1080
`DownscaleBgra` at ~150ms under load, `LatestFrame` updated a handful of times/s —
the desktop feed inside a 60fps file read ~90% frozen. (Webcam + audio were their
own paths — fine, matching the report.) Three changes, all in
`Services/ScreenCaptureFrameSource.cs` (+ `Services/FrameRingBuffer.cs`):
(1) `DownscaleBgra` is now **integer 8.8 fixed-point, "shift only at the end"** —
the exact two-stage math of `SceneCompositor.Bilinear` (MyMistakes take-4 rule),
dropping double-per-pixel to a row-walk of integer ops (the capture downscale had
stayed the naive float twin of the 258ms disaster); (2) a **10ms conversion floor**
(`MinConvertInterval`) so the 240Hz tail stops queuing ~150ms of serialized
conversion per second — the open edge sits just above the ~60/s the pump can use;
(3) the ring is a **reuse-distance pool** (`FrameRingBuffer`, depth 8, redLine 4):
a buffer is only rewritten ≥4 rents after its last hand-out else a fresh allocation,
replacing the blind round-robin — since `ScreenCaptureManager` keeps
`session.LatestFrame` across conversions and the dispatcher preview copy lags,
"who released the buffer" needs a consumer API that doesn't exist; a reuse
DISTANCE needs no cooperation (the 1742 new-top/old-bottom tear is that
read-under-write, structural now). Telemetry added: a startup.log line every 2s
(`frames/s, conv avg/max ms, skip busy/cadence, ring allocs`) so the device take
can be judged numerically. **Good Dog test**
`Ring_NoLap_ReusesOnlyAfterRedLineRents` (depth 3/redLine 4 exercises the red-line
skip; the 8/4 config settles at 8 buffers and never grows). Full suite **295/295
green, 0 warnings**. C4 (compositor Epoch-cached downscale / blit-on-change) was
DEFERRED: the slice-15 approved plan assumed a native-res relocation; the
measurement recast it — relocating a ~30ms float downscale to the render thread
just moves the same cost into the slot budget. Re-measure on device; if render
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was