perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)

Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv
~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s.
Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms
of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact:
pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays
native; the lever is concurrency.

- MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit)
- new MonotonicGate (Interlocked compare-exchange): stale OLDER completions
  are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear)
- FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is
  per-conversion locals
- Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green,
  0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point.

Local only, no push.
This commit is contained in:
2026-09-14 18:40:20 -07:00
parent 71932b9756
commit b37b8a30f9
7 changed files with 238 additions and 114 deletions
+35 -6
View File
@@ -393,12 +393,41 @@ the acceleration; re-measure after slice 15.
dispatcher preview copy lags, so "who released it" is unknowable without a
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
tear (new-top/old-bottom midway) is that read-under-write closed.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
bug):**
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
follow-on decisions fixed by evidence:
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
the pool to the master size to read back less. Microsoft's screen-capture page
forbids it: "the underlying Direct3D surface is always the size specified when
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
not scale — silently encode the desktop cut off. Readback must stay native; the
lever is concurrency, not size. **Rule: read the platform doc for the exact
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
have been wrong twice now.**
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
flight, completions can land out of order; a slow OLDER readback finishing last
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
compare-exchange publish) drops stale completions instead. Ring gets a lock because
rents are now concurrent; downscale row-scratch became per-conversion locals.
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
choosing between throughput and latency fixes — both previous slices picked the
wrong slot ("render" vs "conversion") until the audit existed.
---