perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)
Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv ~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s. Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact: pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays native; the lever is concurrency. - MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit) - new MonotonicGate (Interlocked compare-exchange): stale OLDER completions are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear) - FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is per-conversion locals - Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green, 0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point. Local only, no push.
This commit is contained in:
+35
-6
@@ -393,12 +393,41 @@ the acceleration; re-measure after slice 15.
|
||||
dispatcher preview copy lags, so "who released it" is unknowable without a
|
||||
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
|
||||
tear (new-top/old-bottom midway) is that read-under-write closed.
|
||||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||||
relocating a 30ms float downscale from the capture thread to the render thread
|
||||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||
op fast.
|
||||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||||
relocating a 30ms float downscale from the capture thread to the render thread
|
||||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||
op fast.
|
||||
|
||||
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
|
||||
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
|
||||
bug):**
|
||||
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
|
||||
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
|
||||
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
|
||||
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
|
||||
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
|
||||
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
|
||||
follow-on decisions fixed by evidence:
|
||||
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
|
||||
the pool to the master size to read back less. Microsoft's screen-capture page
|
||||
forbids it: "the underlying Direct3D surface is always the size specified when
|
||||
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
|
||||
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
|
||||
not scale — silently encode the desktop cut off. Readback must stay native; the
|
||||
lever is concurrency, not size. **Rule: read the platform doc for the exact
|
||||
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
|
||||
have been wrong twice now.**
|
||||
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
|
||||
flight, completions can land out of order; a slow OLDER readback finishing last
|
||||
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
|
||||
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
|
||||
compare-exchange publish) drops stale completions instead. Ring gets a lock because
|
||||
rents are now concurrent; downscale row-scratch became per-conversion locals.
|
||||
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
|
||||
choosing between throughput and latency fixes — both previous slices picked the
|
||||
wrong slot ("render" vs "conversion") until the audit existed.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user