perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)
Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv ~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s. Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact: pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays native; the lever is concurrency. - MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit) - new MonotonicGate (Interlocked compare-exchange): stale OLDER completions are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear) - FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is per-conversion locals - Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green, 0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point. Local only, no push.
This commit is contained in:
+59
-60
@@ -1,11 +1,11 @@
|
||||
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
|
||||
# HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next)
|
||||
|
||||
## Branch / Commit State
|
||||
|
||||
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
|
||||
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
|
||||
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
|
||||
audio-sync, pushed). Working tree **clean**.
|
||||
`main` HEAD = **slice-17 commit** (overlapping capture readbacks — committed LOCALLY, **NOT
|
||||
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture
|
||||
conversion fix, slice-15 pacing fix (FramePump), `c01206f` (composition capture), `b22d08e`
|
||||
(signed audio-sync, pushed). Working tree **clean**.
|
||||
|
||||
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
|
||||
|
||||
@@ -13,62 +13,60 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
|
||||
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
|
||||
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
|
||||
|
||||
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
|
||||
## 🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall
|
||||
|
||||
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
|
||||
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
|
||||
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
|
||||
tear_audit.py` + `mix_check.py`):
|
||||
Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the
|
||||
creator reports the desktop layer **still looks like missing frames / choppy vs live**. Decoded
|
||||
`ty-20260914-1824-0000-2.mp4` to `/mnt/c/tmpout/f1824.raw` and audited:
|
||||
- Desktop band: **6.8 content updates/s, 88% frozen**, one **4.85s freeze** at start. Worse,
|
||||
not better, than 1742.
|
||||
- BUT the new slice-16 telemetry proved the downscale fix WORKS: `conv avg 46-50ms, max ~61ms,
|
||||
~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0`. The ring is
|
||||
steady-state (no allocs); the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`) is ~45ms of
|
||||
the conversion** on a 240Hz-HDR box sharing the GPU with the encoder. **The wall was never
|
||||
the CPU downscale.**
|
||||
- FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render
|
||||
is the SECOND cap, ~30 unique composites/s).
|
||||
- Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that
|
||||
theory is now closed).
|
||||
|
||||
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
|
||||
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
|
||||
was real but MOOT.
|
||||
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
|
||||
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
|
||||
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
|
||||
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
|
||||
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
|
||||
reported.
|
||||
## 🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool
|
||||
|
||||
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
|
||||
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/MonotonicGate.cs` + locked
|
||||
`Services/FrameRingBuffer.cs`):
|
||||
1. **MaxConcurrentConversions = 3** readbacks in flight (was one-in-flight `_framePending`),
|
||||
pool buffers 2 → **5** so in-flight frames fit.
|
||||
2. **MonotonicLatest publish gate** (`MonotonicGate`, new internal): a completed readback is
|
||||
published ONLY if its Epoch is strictly newer than the last published. Overlapping
|
||||
conversions can finish out of order; a slow OLDER completion must never overwrite a newer
|
||||
`LatestFrame` (backwards time hole = the mirror of the 1742 tear). Epoch via
|
||||
`Interlocked.Increment`.
|
||||
3. `FrameRingBuffer.Rent`/`ConsumeAllocations` now take `_lock` (rents are concurrent); the
|
||||
10ms floor and downscale stay; `DownscaleBgra` row-scratch is per-conversion locals (no
|
||||
shared `_row0/_row1`).
|
||||
4. **RESEARCH near-miss (docs'ed, MyMistakes):** shrinking the pool to 1920×1080 would have
|
||||
been wrong — Microsoft Docs (screen capture): *"the underlying Direct3D surface is always
|
||||
the size specified … **clipped**"* to the frame. Readback stays native; the lever is
|
||||
concurrency.
|
||||
|
||||
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
|
||||
|
||||
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
|
||||
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
|
||||
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
|
||||
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
|
||||
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
|
||||
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
|
||||
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
|
||||
preview lags).
|
||||
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
|
||||
`ring allocs` — so the next device take is judged numerically.
|
||||
|
||||
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
|
||||
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
|
||||
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
|
||||
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
|
||||
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
|
||||
lessons + the recast note), this HANDOFF.
|
||||
|
||||
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
|
||||
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
|
||||
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
|
||||
budget. C4 (compositor optimization) stays conditional on the re-measure.
|
||||
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
|
||||
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
|
||||
**Good Dog test:** `ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins`.
|
||||
**296/296 green, app build 0 warnings.** Scope-lock files (6): `Services/ScreenCaptureFrameSource.cs`,
|
||||
`Services/FrameRingBuffer.cs`, `Services/MonotonicGate.cs` (new), `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`
|
||||
+ docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). `SceneCompositor.cs`/`FramePump.cs`
|
||||
NOT touched (C4 is the next slice, pending this re-measure).
|
||||
|
||||
## ⚠️ Open items (before PUSHABLE)
|
||||
|
||||
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
|
||||
slice-16 build. Judge numerically:
|
||||
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
|
||||
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
|
||||
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
|
||||
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-17
|
||||
build. Judge numerically:
|
||||
- startup.log telemetry: conversions/s should jump from ~17-20 to **≥ ~30-40**, `skip busy`
|
||||
falling, `ring allocs` ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps).
|
||||
- Decode + `/tmp/opencode/tear_audit.py`: desktop-band fresh updates/s up toward the render
|
||||
cap (~30+), frozen % well under 50%, no genuine mid-frame splits.
|
||||
- ffprobe: video ≈ audio ≈ wall.
|
||||
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
|
||||
composite downscale / blit-on-change), still staying 60fps.
|
||||
- If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots →
|
||||
**C4 slice** (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags,
|
||||
the bind is GPU contention — re-measure before touching anything.
|
||||
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
|
||||
|
||||
## Open threads (carried)
|
||||
@@ -77,7 +75,7 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
|
||||
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
|
||||
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
|
||||
- Signed A/V sync: verify the negative (advance) direction on device.
|
||||
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
|
||||
- Focus-loss capture lag — closed as NOT the cause (1824: delivery healthy ~60-100/s).
|
||||
|
||||
## Landmines
|
||||
|
||||
@@ -88,14 +86,15 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
|
||||
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
|
||||
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
|
||||
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
|
||||
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
|
||||
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
|
||||
recipe** — grep before re-deriving.
|
||||
- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the
|
||||
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and now the **WGC-CLIP** + two
|
||||
slice blocks — grep before re-deriving.
|
||||
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
|
||||
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
|
||||
|
||||
## Next step
|
||||
|
||||
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
|
||||
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
|
||||
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.
|
||||
Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line
|
||||
(conversions/s ≥ ~30-40, `skip busy` falling) + `tear_audit.py` cadence + ffprobe durations. If
|
||||
the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice
|
||||
next, then re-measure clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.
|
||||
+35
-6
@@ -393,12 +393,41 @@ the acceleration; re-measure after slice 15.
|
||||
dispatcher preview copy lags, so "who released it" is unknowable without a
|
||||
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
|
||||
tear (new-top/old-bottom midway) is that read-under-write closed.
|
||||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||||
relocating a 30ms float downscale from the capture thread to the render thread
|
||||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||
op fast.
|
||||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||||
relocating a 30ms float downscale from the capture thread to the render thread
|
||||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||
op fast.
|
||||
|
||||
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
|
||||
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
|
||||
bug):**
|
||||
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
|
||||
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
|
||||
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
|
||||
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
|
||||
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
|
||||
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
|
||||
follow-on decisions fixed by evidence:
|
||||
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
|
||||
the pool to the master size to read back less. Microsoft's screen-capture page
|
||||
forbids it: "the underlying Direct3D surface is always the size specified when
|
||||
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
|
||||
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
|
||||
not scale — silently encode the desktop cut off. Readback must stay native; the
|
||||
lever is concurrency, not size. **Rule: read the platform doc for the exact
|
||||
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
|
||||
have been wrong twice now.**
|
||||
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
|
||||
flight, completions can land out of order; a slow OLDER readback finishing last
|
||||
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
|
||||
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
|
||||
compare-exchange publish) drops stale completions instead. Ring gets a lock because
|
||||
rents are now concurrent; downscale row-scratch became per-conversion locals.
|
||||
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
|
||||
choosing between throughput and latency fixes — both previous slices picked the
|
||||
wrong slot ("render" vs "conversion") until the audit existed.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+32
-22
@@ -25,6 +25,7 @@ internal sealed class FrameRingBuffer
|
||||
private readonly byte[]?[] _slots;
|
||||
private readonly long[] _lastHandout;
|
||||
private readonly int _redLine;
|
||||
private readonly object _lock = new();
|
||||
private int _next;
|
||||
private long _seq;
|
||||
private long _allocations;
|
||||
@@ -38,43 +39,52 @@ internal sealed class FrameRingBuffer
|
||||
}
|
||||
|
||||
/// <summary>Buffers freshly allocated after the last <see cref="ConsumeAllocations"/>.</summary>
|
||||
public long Allocations => _allocations;
|
||||
public long Allocations
|
||||
{
|
||||
get { lock (_lock) return _allocations; }
|
||||
}
|
||||
|
||||
/// <summary>Returns the allocation count since the last call and resets it.</summary>
|
||||
public long ConsumeAllocations()
|
||||
{
|
||||
var count = _allocations;
|
||||
_allocations = 0;
|
||||
return count;
|
||||
lock (_lock)
|
||||
{
|
||||
var count = _allocations;
|
||||
_allocations = 0;
|
||||
return count;
|
||||
}
|
||||
}
|
||||
|
||||
/// <summary>Hands out a scratch buffer of <paramref name="size"/> bytes.</summary>
|
||||
public byte[] Rent(int size)
|
||||
{
|
||||
var seq = ++_seq;
|
||||
for (var tries = 0; tries < _slots.Length; tries++)
|
||||
lock (_lock)
|
||||
{
|
||||
var idx = (_next + tries) % _slots.Length;
|
||||
// Red line: rewriting this slot could hit a frame a consumer still reads.
|
||||
if (seq - _lastHandout[idx] < _redLine) continue;
|
||||
_next = (idx + 1) % _slots.Length;
|
||||
if (_slots[idx] is { Length: var len } buf && len == size)
|
||||
var seq = ++_seq;
|
||||
for (var tries = 0; tries < _slots.Length; tries++)
|
||||
{
|
||||
var idx = (_next + tries) % _slots.Length;
|
||||
// Red line: rewriting this slot could hit a frame a consumer still reads.
|
||||
if (seq - _lastHandout[idx] < _redLine) continue;
|
||||
_next = (idx + 1) % _slots.Length;
|
||||
if (_slots[idx] is { Length: var len } buf && len == size)
|
||||
{
|
||||
_lastHandout[idx] = seq;
|
||||
return buf;
|
||||
}
|
||||
|
||||
// Length mismatch (or never allocated): a fresh array, never an in-place
|
||||
// overwrite — the previous loan's bytes stay valid for whoever holds it.
|
||||
_allocations++;
|
||||
var fresh = new byte[size];
|
||||
_slots[idx] = fresh;
|
||||
_lastHandout[idx] = seq;
|
||||
return buf;
|
||||
return fresh;
|
||||
}
|
||||
|
||||
// Length mismatch (or never allocated): a fresh array, never an in-place
|
||||
// overwrite — the previous loan's bytes stay valid for whoever holds it.
|
||||
// Defensive: the whole ring is inside its red line — do not lap a loaned slot.
|
||||
_allocations++;
|
||||
var fresh = new byte[size];
|
||||
_slots[idx] = fresh;
|
||||
_lastHandout[idx] = seq;
|
||||
return fresh;
|
||||
return new byte[size];
|
||||
}
|
||||
|
||||
// Defensive: the whole ring is inside its red line — do not lap a loaned slot.
|
||||
_allocations++;
|
||||
return new byte[size];
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
using System;
|
||||
using System.Threading;
|
||||
|
||||
namespace ytLive.Services;
|
||||
|
||||
/// <summary>
|
||||
/// An atomic "publish only if strictly newer" gate for the capture's latest-wins
|
||||
/// hand-out. Screen capture conversions may now overlap (slice 17), so two
|
||||
/// conversions can finish briefly out of order; without a monotonic gate the slower
|
||||
/// (older) completion would overwrite the faster (newer) one's LatestFrame and the
|
||||
/// compositor would render STALE content — a time hole in the other direction.
|
||||
/// Sequence must be a strictly increasing per-source counter
|
||||
/// (<c>Interlocked.Increment</c> in <c>ScreenCaptureFrameSource</c>).
|
||||
/// </summary>
|
||||
internal sealed class MonotonicGate
|
||||
{
|
||||
private long _last;
|
||||
|
||||
/// <summary>Returns true iff <paramref name="seq"/> is strictly newer than every
|
||||
/// seq accepted so far (and records it).</summary>
|
||||
public bool TryPublish(long seq)
|
||||
{
|
||||
while (true)
|
||||
{
|
||||
var current = Interlocked.Read(ref _last);
|
||||
if (seq <= current) return false;
|
||||
if (Interlocked.CompareExchange(ref _last, seq, current) == current) return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2,6 +2,7 @@ using System;
|
||||
using System.Diagnostics;
|
||||
using System.Runtime.InteropServices;
|
||||
using System.Runtime.InteropServices.WindowsRuntime;
|
||||
using System.Threading;
|
||||
using System.Threading.Tasks;
|
||||
using Windows.Graphics;
|
||||
using Windows.Graphics.Capture;
|
||||
@@ -28,7 +29,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
private GraphicsCaptureSession? _session;
|
||||
private SizeInt32 _poolSize;
|
||||
private bool _started;
|
||||
private bool _framePending;
|
||||
private int _convertInFlight;
|
||||
private DateTime _lastErrorLog = DateTime.MinValue;
|
||||
|
||||
// Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh
|
||||
@@ -43,8 +44,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
// consumers (the compositor's paste cache) cannot false-hit a recycled array.
|
||||
private readonly FrameRingBuffer _ring = new(8, redLine: 4);
|
||||
private long _epoch;
|
||||
private byte[]? _row0;
|
||||
private byte[]? _row1;
|
||||
private readonly MonotonicGate _publishGate = new();
|
||||
|
||||
// The composition master frame (see ai.md "Resolution tiers"): the background
|
||||
// is an input layer, so we never hold a CPU frame bigger than the master.
|
||||
@@ -54,14 +54,19 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
// A failing conversion must not re-flood the log at frame rate.
|
||||
private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5);
|
||||
|
||||
// Conversion cadence (slice 16, 2026-09-14): the monitor delivers at the 240Hz DWM
|
||||
// cadence (~4.2ms) while slots are 16.6ms. A 10ms floor between conversion starts
|
||||
// keeps the open edge above the ~60 conversions/s the pump can actually use, so the
|
||||
// 240Hz tail stops chewing a conversion thread that the 1742 take measured at
|
||||
// ~150ms/frame (a 90%-frozen desktop, ~6 fresh frames/s). Drop counters and the
|
||||
// rolling conversion stats feed the 2-second startup.log telemetry line.
|
||||
// Conversion cadence (slice 16, 2026-09-14; slice 17, 2026-09-14): the monitor
|
||||
// delivers at the 240Hz DWM cadence (~4.2ms) while slots are 16.6ms. A 10ms floor
|
||||
// between conversion starts keeps the open edge above the ~60 conversions/s the pump
|
||||
// can use. What actually bounded the desktop feed on takes 1742/1824 was the SERIAL
|
||||
// GPU→CPU readback (`CreateCopyFromSurfaceAsync` ≈ 40-50ms of the ~47ms conversion
|
||||
// on a 240Hz-HDR box shared with the encoder), capping captures at ~17-20/s.
|
||||
// Slice 17 overlaps up to MaxConcurrentConversions readbacks (pool sized to
|
||||
// accommodate in-flight frames) and publishes only monotonically newer frames
|
||||
// (MonotonicGate — a slow older completion must never overwrite a newer LatestFrame).
|
||||
private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10);
|
||||
private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2);
|
||||
private const int MaxConcurrentConversions = 3;
|
||||
private const int PoolBufferCount = 5;
|
||||
private DateTime _lastConvertAt = DateTime.MinValue;
|
||||
private DateTime _telemetryFrom = DateTime.UtcNow;
|
||||
private DateTime _lastTelemetry = DateTime.UtcNow;
|
||||
@@ -107,7 +112,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
var item = _item;
|
||||
var device = Direct3D11Helper.CreateDevice();
|
||||
var framePool = Direct3D11CaptureFramePool.CreateFreeThreaded(
|
||||
device, DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, item.Size);
|
||||
device, DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, item.Size);
|
||||
_poolSize = item.Size;
|
||||
var session = framePool.CreateCaptureSession(item);
|
||||
framePool.FrameArrived += OnFrameArrived;
|
||||
@@ -124,7 +129,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
lock (_gate)
|
||||
{
|
||||
_started = false;
|
||||
_framePending = false;
|
||||
_convertInFlight = 0;
|
||||
if (_framePool != null)
|
||||
_framePool.FrameArrived -= OnFrameArrived;
|
||||
_session?.Dispose();
|
||||
@@ -151,14 +156,15 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
|
||||
{
|
||||
sender.Recreate(Direct3D11Helper.CreateDevice(),
|
||||
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
|
||||
DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, frame.ContentSize);
|
||||
_poolSize = frame.ContentSize;
|
||||
}
|
||||
|
||||
// One conversion at a time, spaced by MinConvertInterval (slice 16): the
|
||||
// 240Hz delivery otherwise queued a conversion every ~4.2ms and the
|
||||
// 1742 take's ~150ms conversion pinned the desktop layer ~90% frozen.
|
||||
if (_framePending)
|
||||
// Up to MaxConcurrentConversions readbacks in flight (slice 17), spaced by
|
||||
// MinConvertInterval (slice 16): the 240Hz delivery otherwise queued one
|
||||
// conversion every ~4.2ms, and the serial ~47ms readback pinned the desktop
|
||||
// layer to ~17 updates/s on the 1824 take.
|
||||
if (_convertInFlight >= MaxConcurrentConversions)
|
||||
{
|
||||
_skippedBusy++;
|
||||
frame.Dispose();
|
||||
@@ -172,7 +178,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
return;
|
||||
}
|
||||
_lastConvertAt = now;
|
||||
_framePending = true;
|
||||
_convertInFlight++;
|
||||
}
|
||||
_ = ProcessFrameAsync(frame);
|
||||
}
|
||||
@@ -186,7 +192,14 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
using (var softwareBitmap = await SoftwareBitmap.CreateCopyFromSurfaceAsync(
|
||||
frame.Surface, BitmapAlphaMode.Ignore))
|
||||
{
|
||||
FrameAvailable?.Invoke(CopyToVideoFrame(softwareBitmap));
|
||||
// Monotonic sequencing across the overlapping conversions (slice 17): a
|
||||
// completed readback is published only if strictly newer than the last
|
||||
// one published — a slow older completion must never overwrite a newer
|
||||
// LatestFrame (that would be a time hole in the other direction).
|
||||
var epoch = Interlocked.Increment(ref _epoch);
|
||||
var videoFrame = CopyToVideoFrame(softwareBitmap, epoch);
|
||||
if (_publishGate.TryPublish(epoch))
|
||||
FrameAvailable?.Invoke(videoFrame);
|
||||
}
|
||||
}
|
||||
catch (Exception ex)
|
||||
@@ -203,8 +216,8 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
lock (_gate)
|
||||
{
|
||||
EmitTelemetry(sw.ElapsedMilliseconds);
|
||||
_convertInFlight--;
|
||||
}
|
||||
_framePending = false;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -234,7 +247,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
_skippedCadence = 0;
|
||||
}
|
||||
|
||||
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap)
|
||||
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap, long epoch)
|
||||
{
|
||||
var sw = bitmap.PixelWidth;
|
||||
var sh = bitmap.PixelHeight;
|
||||
@@ -257,12 +270,12 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
// DWM delivers an opaque surface (alpha 255); bilinear keeps it 255.
|
||||
var scaled = _ring.Rent(dw * dh * 4);
|
||||
return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled))
|
||||
{ IsOpaque = true, Epoch = ++_epoch };
|
||||
{ IsOpaque = true, Epoch = epoch };
|
||||
}
|
||||
|
||||
var pixels = _ring.Rent(count);
|
||||
Marshal.Copy(data, pixels, 0, pixels.Length);
|
||||
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = ++_epoch };
|
||||
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = epoch };
|
||||
}
|
||||
|
||||
// Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16,
|
||||
@@ -275,9 +288,11 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
// row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output.
|
||||
private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst)
|
||||
{
|
||||
// Row scratch is per-capture-thread and reused across frames (same churn lesson).
|
||||
var row0 = _row0 != null && _row0.Length >= srcStride ? _row0 : (_row0 = new byte[srcStride]);
|
||||
var row1 = _row1 != null && _row1.Length >= srcStride ? _row1 : (_row1 = new byte[srcStride]);
|
||||
// Row scratch is per-conversion (overlapping conversions since slice 17 each
|
||||
// bring their own — two 10KB arrays, no shared state). Size: a 2560-wide row
|
||||
// pair, the largest the monitor path delivers before the downscale.
|
||||
var row0 = new byte[srcStride];
|
||||
var row1 = new byte[srcStride];
|
||||
|
||||
for (var y = 0; y < dh; y++)
|
||||
{
|
||||
|
||||
@@ -327,7 +327,10 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
||||
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
|
||||
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
|
||||
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
|
||||
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
|
||||
Conversions run as up to MaxConcurrentConversions (3) overlapping readbacks
|
||||
(slice 17: the OS readback, not the downscale, is the ~47ms wall — see Slice 17) with a
|
||||
monotonic LatestFrame publish gate (`MonotonicGate`: a slow OLDER completion can never
|
||||
overwrite a newer frame), and are spaced by a 10ms floor
|
||||
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
|
||||
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
|
||||
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
|
||||
@@ -1065,6 +1068,28 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
|
||||
measurement recast it — relocating a ~30ms float downscale to the render thread
|
||||
just moves the same cost into the slot budget. Re-measure on device; if render
|
||||
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
|
||||
- **Slice 17 — the OS readback was the real wall: overlapping conversions + monotonic
|
||||
publish + deeper pool (2026-09-14, device take ty-1824):** slice 16's downscale fix
|
||||
landed but the desktop was still ~90% frozen on the 1824 take (band 6.8/s updates, max
|
||||
freeze 4.85s). The new 2s telemetry was decisive: `conv avg 46-50ms max ~61ms` with
|
||||
`skip busy 37-83` — the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), not
|
||||
`DownscaleBgra`**, is the ~47ms wall (240Hz HDR compositing + encoder + 2 capture
|
||||
devices share the GPU); at one-in-flight that caps the desktop feed at ~17-20
|
||||
updates/s, which is the file's whole-frame ~7 content-moments/s. Two facts reshaped
|
||||
the fix: **(a)** shrinking the pool size does NOT scale the desktop (Microsoft docs:
|
||||
"If content is larger than the frame, the contents are **clipped**") — readback stays
|
||||
at native 2560×1440; **(b)** delivery is healthy (60-100 arrivals/s), so the lever is
|
||||
conversion throughput, not the pool size. Changes in `Services/ScreenCaptureFrameSource.cs`:
|
||||
conversions overlap up to **MaxConcurrentConversions = 3** (pool deepened to 5 buffers
|
||||
so in-flight frames fit), each completion publishes ONLY if its Epoch is strictly
|
||||
newer than the last published (`MonotonicGate` — a slow older completion must never
|
||||
overwrite a newer LatestFrame), Epoch increments via `Interlocked`, and the
|
||||
DownscaleBgra row-scratch became per-conversion locals (concurrent callers). The
|
||||
render side (33-42ms → ~30 unique composites/s) is the NEXT cap after capture speeds
|
||||
up — that's the C4 slice, queued right after this re-measure. **Good Dog test**
|
||||
`PublishGate_TryPublish_OnlyStrictlyNewerWins`. Full suite **296/296 green, 0
|
||||
warnings**. NOT YET DEVICE-VERIFIED; target: telemetry frames/s jumps ≥ ~30-40 and the
|
||||
band audit drops below ~50% frozen. No push.
|
||||
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
|
||||
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
|
||||
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
|
||||
|
||||
@@ -14,6 +14,22 @@ namespace ytLive.Tests;
|
||||
/// </summary>
|
||||
public class ScreenCaptureFrameSourceTests
|
||||
{
|
||||
[Fact]
|
||||
public void PublishGate_TryPublish_OnlyStrictlyNewerWins()
|
||||
{
|
||||
// Slice 17: overlapping conversions can finish out of order; the monotonic
|
||||
// gate must keep a slow OLDER completion from overwriting a newer LatestFrame.
|
||||
var gate = new MonotonicGate();
|
||||
Assert.True(gate.TryPublish(1));
|
||||
Assert.True(gate.TryPublish(2));
|
||||
Assert.False(gate.TryPublish(1)); // stale replay (older than 2)
|
||||
Assert.True(gate.TryPublish(3));
|
||||
Assert.False(gate.TryPublish(3)); // duplicate — never accepted twice
|
||||
Assert.False(gate.TryPublish(2)); // late older completion
|
||||
Assert.False(gate.TryPublish(long.MinValue));
|
||||
Assert.True(gate.TryPublish(long.MaxValue));
|
||||
}
|
||||
|
||||
[Fact]
|
||||
public void Ring_NoLap_ReusesOnlyAfterRedLineRents()
|
||||
{
|
||||
|
||||
Reference in New Issue
Block a user