perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)
Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv ~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s. Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact: pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays native; the lever is concurrency. - MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit) - new MonotonicGate (Interlocked compare-exchange): stale OLDER completions are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear) - FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is per-conversion locals - Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green, 0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point. Local only, no push.
This commit is contained in:
@@ -327,7 +327,10 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
||||
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
|
||||
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
|
||||
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
|
||||
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
|
||||
Conversions run as up to MaxConcurrentConversions (3) overlapping readbacks
|
||||
(slice 17: the OS readback, not the downscale, is the ~47ms wall — see Slice 17) with a
|
||||
monotonic LatestFrame publish gate (`MonotonicGate`: a slow OLDER completion can never
|
||||
overwrite a newer frame), and are spaced by a 10ms floor
|
||||
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
|
||||
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
|
||||
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
|
||||
@@ -1065,6 +1068,28 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
|
||||
measurement recast it — relocating a ~30ms float downscale to the render thread
|
||||
just moves the same cost into the slot budget. Re-measure on device; if render
|
||||
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
|
||||
- **Slice 17 — the OS readback was the real wall: overlapping conversions + monotonic
|
||||
publish + deeper pool (2026-09-14, device take ty-1824):** slice 16's downscale fix
|
||||
landed but the desktop was still ~90% frozen on the 1824 take (band 6.8/s updates, max
|
||||
freeze 4.85s). The new 2s telemetry was decisive: `conv avg 46-50ms max ~61ms` with
|
||||
`skip busy 37-83` — the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), not
|
||||
`DownscaleBgra`**, is the ~47ms wall (240Hz HDR compositing + encoder + 2 capture
|
||||
devices share the GPU); at one-in-flight that caps the desktop feed at ~17-20
|
||||
updates/s, which is the file's whole-frame ~7 content-moments/s. Two facts reshaped
|
||||
the fix: **(a)** shrinking the pool size does NOT scale the desktop (Microsoft docs:
|
||||
"If content is larger than the frame, the contents are **clipped**") — readback stays
|
||||
at native 2560×1440; **(b)** delivery is healthy (60-100 arrivals/s), so the lever is
|
||||
conversion throughput, not the pool size. Changes in `Services/ScreenCaptureFrameSource.cs`:
|
||||
conversions overlap up to **MaxConcurrentConversions = 3** (pool deepened to 5 buffers
|
||||
so in-flight frames fit), each completion publishes ONLY if its Epoch is strictly
|
||||
newer than the last published (`MonotonicGate` — a slow older completion must never
|
||||
overwrite a newer LatestFrame), Epoch increments via `Interlocked`, and the
|
||||
DownscaleBgra row-scratch became per-conversion locals (concurrent callers). The
|
||||
render side (33-42ms → ~30 unique composites/s) is the NEXT cap after capture speeds
|
||||
up — that's the C4 slice, queued right after this re-measure. **Good Dog test**
|
||||
`PublishGate_TryPublish_OnlyStrictlyNewerWins`. Full suite **296/296 green, 0
|
||||
warnings**. NOT YET DEVICE-VERIFIED; target: telemetry frames/s jumps ≥ ~30-40 and the
|
||||
band audit drops below ~50% frozen. No push.
|
||||
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
|
||||
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
|
||||
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
|
||||
|
||||
Reference in New Issue
Block a user