perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)

Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv
~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s.
Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms
of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact:
pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays
native; the lever is concurrency.

- MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit)
- new MonotonicGate (Interlocked compare-exchange): stale OLDER completions
  are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear)
- FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is
  per-conversion locals
- Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green,
  0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point.

Local only, no push.
This commit is contained in:
2026-09-14 18:40:20 -07:00
parent 71932b9756
commit b37b8a30f9
7 changed files with 238 additions and 114 deletions
+59 -60
View File
@@ -1,11 +1,11 @@
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
# HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next)
## Branch / Commit State
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
audio-sync, pushed). Working tree **clean**.
`main` HEAD = **slice-17 commit** (overlapping capture readbacks — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture
conversion fix, slice-15 pacing fix (FramePump), `c01206f` (composition capture), `b22d08e`
(signed audio-sync, pushed). Working tree **clean**.
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
@@ -13,62 +13,60 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
## 🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
tear_audit.py` + `mix_check.py`):
Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the
creator reports the desktop layer **still looks like missing frames / choppy vs live**. Decoded
`ty-20260914-1824-0000-2.mp4` to `/mnt/c/tmpout/f1824.raw` and audited:
- Desktop band: **6.8 content updates/s, 88% frozen**, one **4.85s freeze** at start. Worse,
not better, than 1742.
- BUT the new slice-16 telemetry proved the downscale fix WORKS: `conv avg 46-50ms, max ~61ms,
~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0`. The ring is
steady-state (no allocs); the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`) is ~45ms of
the conversion** on a 240Hz-HDR box sharing the GPU with the encoder. **The wall was never
the CPU downscale.**
- FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render
is the SECOND cap, ~30 unique composites/s).
- Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that
theory is now closed).
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
was real but MOOT.
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
reported.
## 🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/MonotonicGate.cs` + locked
`Services/FrameRingBuffer.cs`):
1. **MaxConcurrentConversions = 3** readbacks in flight (was one-in-flight `_framePending`),
pool buffers 2 → **5** so in-flight frames fit.
2. **MonotonicLatest publish gate** (`MonotonicGate`, new internal): a completed readback is
published ONLY if its Epoch is strictly newer than the last published. Overlapping
conversions can finish out of order; a slow OLDER completion must never overwrite a newer
`LatestFrame` (backwards time hole = the mirror of the 1742 tear). Epoch via
`Interlocked.Increment`.
3. `FrameRingBuffer.Rent`/`ConsumeAllocations` now take `_lock` (rents are concurrent); the
10ms floor and downscale stay; `DownscaleBgra` row-scratch is per-conversion locals (no
shared `_row0/_row1`).
4. **RESEARCH near-miss (docs'ed, MyMistakes):** shrinking the pool to 1920×1080 would have
been wrong — Microsoft Docs (screen capture): *"the underlying Direct3D surface is always
the size specified … **clipped**"* to the frame. Readback stays native; the lever is
concurrency.
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
preview lags).
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
`ring allocs` — so the next device take is judged numerically.
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
lessons + the recast note), this HANDOFF.
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
budget. C4 (compositor optimization) stays conditional on the re-measure.
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
**Good Dog test:** `ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins`.
**296/296 green, app build 0 warnings.** Scope-lock files (6): `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `Services/MonotonicGate.cs` (new), `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`
+ docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). `SceneCompositor.cs`/`FramePump.cs`
NOT touched (C4 is the next slice, pending this re-measure).
## ⚠️ Open items (before PUSHABLE)
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
slice-16 build. Judge numerically:
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-17
build. Judge numerically:
- startup.log telemetry: conversions/s should jump from ~17-20 to **≥ ~30-40**, `skip busy`
falling, `ring allocs` ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps).
- Decode + `/tmp/opencode/tear_audit.py`: desktop-band fresh updates/s up toward the render
cap (~30+), frozen % well under 50%, no genuine mid-frame splits.
- ffprobe: video ≈ audio ≈ wall.
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
composite downscale / blit-on-change), still staying 60fps.
- If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots →
**C4 slice** (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags,
the bind is GPU contention — re-measure before touching anything.
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
## Open threads (carried)
@@ -77,7 +75,7 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
- Signed A/V sync: verify the negative (advance) direction on device.
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
- Focus-loss capture lag — closed as NOT the cause (1824: delivery healthy ~60-100/s).
## Landmines
@@ -88,14 +86,15 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
recipe** — grep before re-deriving.
- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and now the **WGC-CLIP** + two
slice blocks — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
## Next step
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.
Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line
(conversions/s ≥ ~30-40, `skip busy` falling) + `tear_audit.py` cadence + ffprobe durations. If
the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice
next, then re-measure clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.
+35 -6
View File
@@ -393,12 +393,41 @@ the acceleration; re-measure after slice 15.
dispatcher preview copy lags, so "who released it" is unknowable without a
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
tear (new-top/old-bottom midway) is that read-under-write closed.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
bug):**
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
follow-on decisions fixed by evidence:
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
the pool to the master size to read back less. Microsoft's screen-capture page
forbids it: "the underlying Direct3D surface is always the size specified when
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
not scale — silently encode the desktop cut off. Readback must stay native; the
lever is concurrency, not size. **Rule: read the platform doc for the exact
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
have been wrong twice now.**
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
flight, completions can land out of order; a slow OLDER readback finishing last
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
compare-exchange publish) drops stale completions instead. Ring gets a lock because
rents are now concurrent; downscale row-scratch became per-conversion locals.
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
choosing between throughput and latency fixes — both previous slices picked the
wrong slot ("render" vs "conversion") until the audit existed.
---
+32 -22
View File
@@ -25,6 +25,7 @@ internal sealed class FrameRingBuffer
private readonly byte[]?[] _slots;
private readonly long[] _lastHandout;
private readonly int _redLine;
private readonly object _lock = new();
private int _next;
private long _seq;
private long _allocations;
@@ -38,43 +39,52 @@ internal sealed class FrameRingBuffer
}
/// <summary>Buffers freshly allocated after the last <see cref="ConsumeAllocations"/>.</summary>
public long Allocations => _allocations;
public long Allocations
{
get { lock (_lock) return _allocations; }
}
/// <summary>Returns the allocation count since the last call and resets it.</summary>
public long ConsumeAllocations()
{
var count = _allocations;
_allocations = 0;
return count;
lock (_lock)
{
var count = _allocations;
_allocations = 0;
return count;
}
}
/// <summary>Hands out a scratch buffer of <paramref name="size"/> bytes.</summary>
public byte[] Rent(int size)
{
var seq = ++_seq;
for (var tries = 0; tries < _slots.Length; tries++)
lock (_lock)
{
var idx = (_next + tries) % _slots.Length;
// Red line: rewriting this slot could hit a frame a consumer still reads.
if (seq - _lastHandout[idx] < _redLine) continue;
_next = (idx + 1) % _slots.Length;
if (_slots[idx] is { Length: var len } buf && len == size)
var seq = ++_seq;
for (var tries = 0; tries < _slots.Length; tries++)
{
var idx = (_next + tries) % _slots.Length;
// Red line: rewriting this slot could hit a frame a consumer still reads.
if (seq - _lastHandout[idx] < _redLine) continue;
_next = (idx + 1) % _slots.Length;
if (_slots[idx] is { Length: var len } buf && len == size)
{
_lastHandout[idx] = seq;
return buf;
}
// Length mismatch (or never allocated): a fresh array, never an in-place
// overwrite — the previous loan's bytes stay valid for whoever holds it.
_allocations++;
var fresh = new byte[size];
_slots[idx] = fresh;
_lastHandout[idx] = seq;
return buf;
return fresh;
}
// Length mismatch (or never allocated): a fresh array, never an in-place
// overwrite — the previous loan's bytes stay valid for whoever holds it.
// Defensive: the whole ring is inside its red line — do not lap a loaned slot.
_allocations++;
var fresh = new byte[size];
_slots[idx] = fresh;
_lastHandout[idx] = seq;
return fresh;
return new byte[size];
}
// Defensive: the whole ring is inside its red line — do not lap a loaned slot.
_allocations++;
return new byte[size];
}
}
+30
View File
@@ -0,0 +1,30 @@
using System;
using System.Threading;
namespace ytLive.Services;
/// <summary>
/// An atomic "publish only if strictly newer" gate for the capture's latest-wins
/// hand-out. Screen capture conversions may now overlap (slice 17), so two
/// conversions can finish briefly out of order; without a monotonic gate the slower
/// (older) completion would overwrite the faster (newer) one's LatestFrame and the
/// compositor would render STALE content — a time hole in the other direction.
/// Sequence must be a strictly increasing per-source counter
/// (<c>Interlocked.Increment</c> in <c>ScreenCaptureFrameSource</c>).
/// </summary>
internal sealed class MonotonicGate
{
private long _last;
/// <summary>Returns true iff <paramref name="seq"/> is strictly newer than every
/// seq accepted so far (and records it).</summary>
public bool TryPublish(long seq)
{
while (true)
{
var current = Interlocked.Read(ref _last);
if (seq <= current) return false;
if (Interlocked.CompareExchange(ref _last, seq, current) == current) return true;
}
}
}
+40 -25
View File
@@ -2,6 +2,7 @@ using System;
using System.Diagnostics;
using System.Runtime.InteropServices;
using System.Runtime.InteropServices.WindowsRuntime;
using System.Threading;
using System.Threading.Tasks;
using Windows.Graphics;
using Windows.Graphics.Capture;
@@ -28,7 +29,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
private GraphicsCaptureSession? _session;
private SizeInt32 _poolSize;
private bool _started;
private bool _framePending;
private int _convertInFlight;
private DateTime _lastErrorLog = DateTime.MinValue;
// Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh
@@ -43,8 +44,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
// consumers (the compositor's paste cache) cannot false-hit a recycled array.
private readonly FrameRingBuffer _ring = new(8, redLine: 4);
private long _epoch;
private byte[]? _row0;
private byte[]? _row1;
private readonly MonotonicGate _publishGate = new();
// The composition master frame (see ai.md "Resolution tiers"): the background
// is an input layer, so we never hold a CPU frame bigger than the master.
@@ -54,14 +54,19 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
// A failing conversion must not re-flood the log at frame rate.
private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5);
// Conversion cadence (slice 16, 2026-09-14): the monitor delivers at the 240Hz DWM
// cadence (~4.2ms) while slots are 16.6ms. A 10ms floor between conversion starts
// keeps the open edge above the ~60 conversions/s the pump can actually use, so the
// 240Hz tail stops chewing a conversion thread that the 1742 take measured at
// ~150ms/frame (a 90%-frozen desktop, ~6 fresh frames/s). Drop counters and the
// rolling conversion stats feed the 2-second startup.log telemetry line.
// Conversion cadence (slice 16, 2026-09-14; slice 17, 2026-09-14): the monitor
// delivers at the 240Hz DWM cadence (~4.2ms) while slots are 16.6ms. A 10ms floor
// between conversion starts keeps the open edge above the ~60 conversions/s the pump
// can use. What actually bounded the desktop feed on takes 1742/1824 was the SERIAL
// GPU→CPU readback (`CreateCopyFromSurfaceAsync` ≈ 40-50ms of the ~47ms conversion
// on a 240Hz-HDR box shared with the encoder), capping captures at ~17-20/s.
// Slice 17 overlaps up to MaxConcurrentConversions readbacks (pool sized to
// accommodate in-flight frames) and publishes only monotonically newer frames
// (MonotonicGate — a slow older completion must never overwrite a newer LatestFrame).
private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10);
private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2);
private const int MaxConcurrentConversions = 3;
private const int PoolBufferCount = 5;
private DateTime _lastConvertAt = DateTime.MinValue;
private DateTime _telemetryFrom = DateTime.UtcNow;
private DateTime _lastTelemetry = DateTime.UtcNow;
@@ -107,7 +112,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
var item = _item;
var device = Direct3D11Helper.CreateDevice();
var framePool = Direct3D11CaptureFramePool.CreateFreeThreaded(
device, DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, item.Size);
device, DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, item.Size);
_poolSize = item.Size;
var session = framePool.CreateCaptureSession(item);
framePool.FrameArrived += OnFrameArrived;
@@ -124,7 +129,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
lock (_gate)
{
_started = false;
_framePending = false;
_convertInFlight = 0;
if (_framePool != null)
_framePool.FrameArrived -= OnFrameArrived;
_session?.Dispose();
@@ -151,14 +156,15 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
{
sender.Recreate(Direct3D11Helper.CreateDevice(),
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, frame.ContentSize);
_poolSize = frame.ContentSize;
}
// One conversion at a time, spaced by MinConvertInterval (slice 16): the
// 240Hz delivery otherwise queued a conversion every ~4.2ms and the
// 1742 take's ~150ms conversion pinned the desktop layer ~90% frozen.
if (_framePending)
// Up to MaxConcurrentConversions readbacks in flight (slice 17), spaced by
// MinConvertInterval (slice 16): the 240Hz delivery otherwise queued one
// conversion every ~4.2ms, and the serial ~47ms readback pinned the desktop
// layer to ~17 updates/s on the 1824 take.
if (_convertInFlight >= MaxConcurrentConversions)
{
_skippedBusy++;
frame.Dispose();
@@ -172,7 +178,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
return;
}
_lastConvertAt = now;
_framePending = true;
_convertInFlight++;
}
_ = ProcessFrameAsync(frame);
}
@@ -186,7 +192,14 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
using (var softwareBitmap = await SoftwareBitmap.CreateCopyFromSurfaceAsync(
frame.Surface, BitmapAlphaMode.Ignore))
{
FrameAvailable?.Invoke(CopyToVideoFrame(softwareBitmap));
// Monotonic sequencing across the overlapping conversions (slice 17): a
// completed readback is published only if strictly newer than the last
// one published — a slow older completion must never overwrite a newer
// LatestFrame (that would be a time hole in the other direction).
var epoch = Interlocked.Increment(ref _epoch);
var videoFrame = CopyToVideoFrame(softwareBitmap, epoch);
if (_publishGate.TryPublish(epoch))
FrameAvailable?.Invoke(videoFrame);
}
}
catch (Exception ex)
@@ -203,8 +216,8 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
lock (_gate)
{
EmitTelemetry(sw.ElapsedMilliseconds);
_convertInFlight--;
}
_framePending = false;
}
}
@@ -234,7 +247,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
_skippedCadence = 0;
}
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap)
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap, long epoch)
{
var sw = bitmap.PixelWidth;
var sh = bitmap.PixelHeight;
@@ -257,12 +270,12 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
// DWM delivers an opaque surface (alpha 255); bilinear keeps it 255.
var scaled = _ring.Rent(dw * dh * 4);
return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled))
{ IsOpaque = true, Epoch = ++_epoch };
{ IsOpaque = true, Epoch = epoch };
}
var pixels = _ring.Rent(count);
Marshal.Copy(data, pixels, 0, pixels.Length);
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = ++_epoch };
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = epoch };
}
// Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16,
@@ -275,9 +288,11 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
// row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output.
private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst)
{
// Row scratch is per-capture-thread and reused across frames (same churn lesson).
var row0 = _row0 != null && _row0.Length >= srcStride ? _row0 : (_row0 = new byte[srcStride]);
var row1 = _row1 != null && _row1.Length >= srcStride ? _row1 : (_row1 = new byte[srcStride]);
// Row scratch is per-conversion (overlapping conversions since slice 17 each
// bring their own — two 10KB arrays, no shared state). Size: a 2560-wide row
// pair, the largest the monitor path delivers before the downscale.
var row0 = new byte[srcStride];
var row1 = new byte[srcStride];
for (var y = 0; y < dh; y++)
{
+26 -1
View File
@@ -327,7 +327,10 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
Conversions run as up to MaxConcurrentConversions (3) overlapping readbacks
(slice 17: the OS readback, not the downscale, is the ~47ms wall — see Slice 17) with a
monotonic LatestFrame publish gate (`MonotonicGate`: a slow OLDER completion can never
overwrite a newer frame), and are spaced by a 10ms floor
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
@@ -1065,6 +1068,28 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
measurement recast it — relocating a ~30ms float downscale to the render thread
just moves the same cost into the slot budget. Re-measure on device; if render
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
- **Slice 17 — the OS readback was the real wall: overlapping conversions + monotonic
publish + deeper pool (2026-09-14, device take ty-1824):** slice 16's downscale fix
landed but the desktop was still ~90% frozen on the 1824 take (band 6.8/s updates, max
freeze 4.85s). The new 2s telemetry was decisive: `conv avg 46-50ms max ~61ms` with
`skip busy 37-83` — the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), not
`DownscaleBgra`**, is the ~47ms wall (240Hz HDR compositing + encoder + 2 capture
devices share the GPU); at one-in-flight that caps the desktop feed at ~17-20
updates/s, which is the file's whole-frame ~7 content-moments/s. Two facts reshaped
the fix: **(a)** shrinking the pool size does NOT scale the desktop (Microsoft docs:
"If content is larger than the frame, the contents are **clipped**") — readback stays
at native 2560×1440; **(b)** delivery is healthy (60-100 arrivals/s), so the lever is
conversion throughput, not the pool size. Changes in `Services/ScreenCaptureFrameSource.cs`:
conversions overlap up to **MaxConcurrentConversions = 3** (pool deepened to 5 buffers
so in-flight frames fit), each completion publishes ONLY if its Epoch is strictly
newer than the last published (`MonotonicGate` — a slow older completion must never
overwrite a newer LatestFrame), Epoch increments via `Interlocked`, and the
DownscaleBgra row-scratch became per-conversion locals (concurrent callers). The
render side (33-42ms → ~30 unique composites/s) is the NEXT cap after capture speeds
up — that's the C4 slice, queued right after this re-measure. **Good Dog test**
`PublishGate_TryPublish_OnlyStrictlyNewerWins`. Full suite **296/296 green, 0
warnings**. NOT YET DEVICE-VERIFIED; target: telemetry frames/s jumps ≥ ~30-40 and the
band audit drops below ~50% frozen. No push.
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
@@ -14,6 +14,22 @@ namespace ytLive.Tests;
/// </summary>
public class ScreenCaptureFrameSourceTests
{
[Fact]
public void PublishGate_TryPublish_OnlyStrictlyNewerWins()
{
// Slice 17: overlapping conversions can finish out of order; the monotonic
// gate must keep a slow OLDER completion from overwriting a newer LatestFrame.
var gate = new MonotonicGate();
Assert.True(gate.TryPublish(1));
Assert.True(gate.TryPublish(2));
Assert.False(gate.TryPublish(1)); // stale replay (older than 2)
Assert.True(gate.TryPublish(3));
Assert.False(gate.TryPublish(3)); // duplicate — never accepted twice
Assert.False(gate.TryPublish(2)); // late older completion
Assert.False(gate.TryPublish(long.MinValue));
Assert.True(gate.TryPublish(long.MaxValue));
}
[Fact]
public void Ring_NoLap_ReusesOnlyAfterRedLineRents()
{