perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)
The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was the wall, and the one torn frame was a ring slot rewritten under the consumer's read. Reference: WGC delivers at DWM/monitor cadence (https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture) and libyuv row-simple/fixed-point scaling (https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule. - DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame. - 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of serialized conversion/s; capacity sits just above the 60/s the pump can use. - FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a consumer still holds (session.LatestFrame survives conversions, dispatcher preview lags) is never read-while-overwritten. Needs no consumer Release API. - 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring allocs — the device take is judgeable numerically. Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings. C4 (composite Epoch-cached downscale) deferred pending the device re-measure. Local only, no push.
This commit is contained in:
@@ -324,10 +324,20 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
||||
`WindowsRuntimeMarshal.TryGetDataUnsafe` (the same CsWinRT-safe read the webcam path uses) — the
|
||||
`IMemoryBufferByteAccess` ComImport cast threw `Invalid cast` on **every frame** under CsWinRT, which
|
||||
flooded `startup.log` (~5 MB in a session) and burned CPU, so it is gone. Surfaces larger than the
|
||||
1920×1080 master are downscaled bilinearly to the master (`DownscaleBgra`) before the copy, and
|
||||
per-frame conversion failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected
|
||||
content delivers black frames (OS limitation, documented). Frame pool
|
||||
pauses while the app is minimized — capture keeps running, the pool just stops delivering.
|
||||
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
|
||||
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
|
||||
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
|
||||
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
|
||||
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
|
||||
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
|
||||
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
|
||||
buffer is only rewritten ≥4 rents after its last hand-out, else a fresh one is allocated — a frame a
|
||||
consumer still holds (`session.LatestFrame` survives conversions; the dispatcher preview copy lags)
|
||||
can never be read-while-overwritten (the 1742 tear). Rolling telemetry (frames/s, conv avg/max ms,
|
||||
drops, ring allocs) is logged to `startup.log` every 2s while converting. Per-frame conversion
|
||||
failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected content delivers black
|
||||
frames (OS limitation, documented). Frame pool pauses while the app is minimized — capture keeps
|
||||
running, the pool just stops delivering.
|
||||
- **Ownership:** `ScreenCaptureManager` mirrors `CameraManager` — refcounted by target key, one shared
|
||||
`WriteableBitmap` per key, dispatcher-coalesced latest-frame copies, `PreviewBitmapChanged`/`CaptureFailed`
|
||||
events, `ReleaseAllAsync` on re-designation. `ScreenCaptureSourceFactory.Resolve(key)` parses the key into
|
||||
@@ -353,7 +363,9 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
||||
app is in the background (worse under a full-screen game on 24H2/26100), GPU readback
|
||||
(`CreateCopyFromSurfaceAsync`) contends with the foreground game, the source's one-in-flight `_framePending`
|
||||
gate drops frames during stalls, and the manager's `DispatcherPriority.Render` copies only run as fast as
|
||||
WPF presents the window. Recorded 2026-08-13; no mitigation attempted yet (deferred by user decision).
|
||||
WPF presents the window. Slice 16 shrank the per-conversion stall (fast downscale + 10ms cadence floor)
|
||||
but the OS-level delivery throttle under focus loss remains. Recorded 2026-08-13; no mitigation attempted
|
||||
yet (deferred by user decision).
|
||||
- **GPU posture:** same as webcam — CPU frames, WPF hardware-presents; D3DImage GPU compositing deferred
|
||||
to the encoder task.
|
||||
- **Preview watermark:** the "Preview" placeholder hides while a background capture renders —
|
||||
@@ -1019,6 +1031,40 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
|
||||
suite **294/294 green, 0 warnings**. The +0.6s audio-late clap reading on 1726 was confounded by
|
||||
the 1.1x acceleration — re-measure on device; if a real residual remains it is the audio pipeline.
|
||||
No push (web/A/V work stays local).
|
||||
- **Slice 16 — the desktop capture conversion was the bottleneck: fast downscale +
|
||||
cadence throttle + reuse-distance ring (2026-09-14, device take ty-1742):** slice 15
|
||||
fixed pacing but the 1742 desktop layer was still jerky/frozen with a horizontal
|
||||
tear. Decoded to raw frames and audited: the **desktop band was frozen 21s of 23.35s
|
||||
(90%)** — ~6.1 content updates/s, freeze runs up to 2.28-2.78s, dup-run max 90
|
||||
frames (1.5s). The render stat (`worst render 33-36ms`) was real but MOOT: the
|
||||
capture CONVERSION was the wall. The monitor delivers at the **240Hz DWM cadence**
|
||||
(~4.2ms); with one-in-flight conversions and each 2560×1440→1920×1080
|
||||
`DownscaleBgra` at ~150ms under load, `LatestFrame` updated a handful of times/s —
|
||||
the desktop feed inside a 60fps file read ~90% frozen. (Webcam + audio were their
|
||||
own paths — fine, matching the report.) Three changes, all in
|
||||
`Services/ScreenCaptureFrameSource.cs` (+ `Services/FrameRingBuffer.cs`):
|
||||
(1) `DownscaleBgra` is now **integer 8.8 fixed-point, "shift only at the end"** —
|
||||
the exact two-stage math of `SceneCompositor.Bilinear` (MyMistakes take-4 rule),
|
||||
dropping double-per-pixel to a row-walk of integer ops (the capture downscale had
|
||||
stayed the naive float twin of the 258ms disaster); (2) a **10ms conversion floor**
|
||||
(`MinConvertInterval`) so the 240Hz tail stops queuing ~150ms of serialized
|
||||
conversion per second — the open edge sits just above the ~60/s the pump can use;
|
||||
(3) the ring is a **reuse-distance pool** (`FrameRingBuffer`, depth 8, redLine 4):
|
||||
a buffer is only rewritten ≥4 rents after its last hand-out else a fresh allocation,
|
||||
replacing the blind round-robin — since `ScreenCaptureManager` keeps
|
||||
`session.LatestFrame` across conversions and the dispatcher preview copy lags,
|
||||
"who released the buffer" needs a consumer API that doesn't exist; a reuse
|
||||
DISTANCE needs no cooperation (the 1742 new-top/old-bottom tear is that
|
||||
read-under-write, structural now). Telemetry added: a startup.log line every 2s
|
||||
(`frames/s, conv avg/max ms, skip busy/cadence, ring allocs`) so the device take
|
||||
can be judged numerically. **Good Dog test**
|
||||
`Ring_NoLap_ReusesOnlyAfterRedLineRents` (depth 3/redLine 4 exercises the red-line
|
||||
skip; the 8/4 config settles at 8 buffers and never grows). Full suite **295/295
|
||||
green, 0 warnings**. C4 (compositor Epoch-cached downscale / blit-on-change) was
|
||||
DEFERRED: the slice-15 approved plan assumed a native-res relocation; the
|
||||
measurement recast it — relocating a ~30ms float downscale to the render thread
|
||||
just moves the same cost into the slot budget. Re-measure on device; if render
|
||||
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
|
||||
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
|
||||
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
|
||||
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
|
||||
|
||||
Reference in New Issue
Block a user