perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)

Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv
~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s.
Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms
of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact:
pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays
native; the lever is concurrency.

- MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit)
- new MonotonicGate (Interlocked compare-exchange): stale OLDER completions
  are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear)
- FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is
  per-conversion locals
- Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green,
  0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point.

Local only, no push.
This commit is contained in:
2026-09-14 18:40:20 -07:00
parent 71932b9756
commit b37b8a30f9
7 changed files with 238 additions and 114 deletions
+26 -1
View File
@@ -327,7 +327,10 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
Conversions run as up to MaxConcurrentConversions (3) overlapping readbacks
(slice 17: the OS readback, not the downscale, is the ~47ms wall — see Slice 17) with a
monotonic LatestFrame publish gate (`MonotonicGate`: a slow OLDER completion can never
overwrite a newer frame), and are spaced by a 10ms floor
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
@@ -1065,6 +1068,28 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
measurement recast it — relocating a ~30ms float downscale to the render thread
just moves the same cost into the slot budget. Re-measure on device; if render
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
- **Slice 17 — the OS readback was the real wall: overlapping conversions + monotonic
publish + deeper pool (2026-09-14, device take ty-1824):** slice 16's downscale fix
landed but the desktop was still ~90% frozen on the 1824 take (band 6.8/s updates, max
freeze 4.85s). The new 2s telemetry was decisive: `conv avg 46-50ms max ~61ms` with
`skip busy 37-83` — the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), not
`DownscaleBgra`**, is the ~47ms wall (240Hz HDR compositing + encoder + 2 capture
devices share the GPU); at one-in-flight that caps the desktop feed at ~17-20
updates/s, which is the file's whole-frame ~7 content-moments/s. Two facts reshaped
the fix: **(a)** shrinking the pool size does NOT scale the desktop (Microsoft docs:
"If content is larger than the frame, the contents are **clipped**") — readback stays
at native 2560×1440; **(b)** delivery is healthy (60-100 arrivals/s), so the lever is
conversion throughput, not the pool size. Changes in `Services/ScreenCaptureFrameSource.cs`:
conversions overlap up to **MaxConcurrentConversions = 3** (pool deepened to 5 buffers
so in-flight frames fit), each completion publishes ONLY if its Epoch is strictly
newer than the last published (`MonotonicGate` — a slow older completion must never
overwrite a newer LatestFrame), Epoch increments via `Interlocked`, and the
DownscaleBgra row-scratch became per-conversion locals (concurrent callers). The
render side (33-42ms → ~30 unique composites/s) is the NEXT cap after capture speeds
up — that's the C4 slice, queued right after this re-measure. **Good Dog test**
`PublishGate_TryPublish_OnlyStrictlyNewerWins`. Full suite **296/296 green, 0
warnings**. NOT YET DEVICE-VERIFIED; target: telemetry frames/s jumps ≥ ~30-40 and the
band audit drops below ~50% frozen. No push.
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was