perf(capture): overlap GPU readbacks with monotonic publish gate (slice 17)
Measure take ty-1824 on the slice-16 build: the downscale fix worked (conv ~47ms, ring allocs 0) but desktop band was still 88% frozen at 6.8 updates/s. Telemetry isolated the real wall — CreateCopyFromSurfaceAsync readback ~45ms of each conversion, serialized one-in-flight => ~17/s capture cap. Docs fact: pool-sized surfaces CLIP, not scale (Microsoft Learn), so readback stays native; the lever is concurrency. - MaxConcurrentConversions=3 with pool 2->5 buffers (in-flight frames fit) - new MonotonicGate (Interlocked compare-exchange): stale OLDER completions are dropped, never overwrite a newer LatestFrame (mirror of 1742 tear) - FrameRingBuffer.Rent/ConsumeAllocations now lock; downscale row scratch is per-conversion locals - Good Dog test PublishGate_TryPublish_OnlyStrictlyNewerWins; 296/296 green, 0 warnings; docs cited Microsoft screen-capture page + libyuv fixed-point. Local only, no push.
This commit is contained in:
+59
-60
@@ -1,11 +1,11 @@
|
|||||||
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
|
# HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next)
|
||||||
|
|
||||||
## Branch / Commit State
|
## Branch / Commit State
|
||||||
|
|
||||||
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
|
`main` HEAD = **slice-17 commit** (overlapping capture readbacks — committed LOCALLY, **NOT
|
||||||
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
|
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture
|
||||||
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
|
conversion fix, slice-15 pacing fix (FramePump), `c01206f` (composition capture), `b22d08e`
|
||||||
audio-sync, pushed). Working tree **clean**.
|
(signed audio-sync, pushed). Working tree **clean**.
|
||||||
|
|
||||||
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
|
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
|
||||||
|
|
||||||
@@ -13,62 +13,60 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
|
|||||||
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
|
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
|
||||||
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
|
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
|
||||||
|
|
||||||
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
|
## 🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall
|
||||||
|
|
||||||
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
|
Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the
|
||||||
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
|
creator reports the desktop layer **still looks like missing frames / choppy vs live**. Decoded
|
||||||
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
|
`ty-20260914-1824-0000-2.mp4` to `/mnt/c/tmpout/f1824.raw` and audited:
|
||||||
tear_audit.py` + `mix_check.py`):
|
- Desktop band: **6.8 content updates/s, 88% frozen**, one **4.85s freeze** at start. Worse,
|
||||||
|
not better, than 1742.
|
||||||
|
- BUT the new slice-16 telemetry proved the downscale fix WORKS: `conv avg 46-50ms, max ~61ms,
|
||||||
|
~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0`. The ring is
|
||||||
|
steady-state (no allocs); the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`) is ~45ms of
|
||||||
|
the conversion** on a 240Hz-HDR box sharing the GPU with the encoder. **The wall was never
|
||||||
|
the CPU downscale.**
|
||||||
|
- FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render
|
||||||
|
is the SECOND cap, ~30 unique composites/s).
|
||||||
|
- Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that
|
||||||
|
theory is now closed).
|
||||||
|
|
||||||
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
|
## 🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool
|
||||||
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
|
|
||||||
was real but MOOT.
|
|
||||||
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
|
|
||||||
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
|
|
||||||
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
|
|
||||||
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
|
|
||||||
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
|
|
||||||
reported.
|
|
||||||
|
|
||||||
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
|
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/MonotonicGate.cs` + locked
|
||||||
|
`Services/FrameRingBuffer.cs`):
|
||||||
|
1. **MaxConcurrentConversions = 3** readbacks in flight (was one-in-flight `_framePending`),
|
||||||
|
pool buffers 2 → **5** so in-flight frames fit.
|
||||||
|
2. **MonotonicLatest publish gate** (`MonotonicGate`, new internal): a completed readback is
|
||||||
|
published ONLY if its Epoch is strictly newer than the last published. Overlapping
|
||||||
|
conversions can finish out of order; a slow OLDER completion must never overwrite a newer
|
||||||
|
`LatestFrame` (backwards time hole = the mirror of the 1742 tear). Epoch via
|
||||||
|
`Interlocked.Increment`.
|
||||||
|
3. `FrameRingBuffer.Rent`/`ConsumeAllocations` now take `_lock` (rents are concurrent); the
|
||||||
|
10ms floor and downscale stay; `DownscaleBgra` row-scratch is per-conversion locals (no
|
||||||
|
shared `_row0/_row1`).
|
||||||
|
4. **RESEARCH near-miss (docs'ed, MyMistakes):** shrinking the pool to 1920×1080 would have
|
||||||
|
been wrong — Microsoft Docs (screen capture): *"the underlying Direct3D surface is always
|
||||||
|
the size specified … **clipped**"* to the frame. Readback stays native; the lever is
|
||||||
|
concurrency.
|
||||||
|
|
||||||
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
|
**Good Dog test:** `ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins`.
|
||||||
|
**296/296 green, app build 0 warnings.** Scope-lock files (6): `Services/ScreenCaptureFrameSource.cs`,
|
||||||
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
|
`Services/FrameRingBuffer.cs`, `Services/MonotonicGate.cs` (new), `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`
|
||||||
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
|
+ docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). `SceneCompositor.cs`/`FramePump.cs`
|
||||||
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
|
NOT touched (C4 is the next slice, pending this re-measure).
|
||||||
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
|
|
||||||
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
|
|
||||||
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
|
|
||||||
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
|
|
||||||
preview lags).
|
|
||||||
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
|
|
||||||
`ring allocs` — so the next device take is judged numerically.
|
|
||||||
|
|
||||||
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
|
|
||||||
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
|
|
||||||
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
|
|
||||||
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
|
|
||||||
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
|
|
||||||
lessons + the recast note), this HANDOFF.
|
|
||||||
|
|
||||||
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
|
|
||||||
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
|
|
||||||
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
|
|
||||||
budget. C4 (compositor optimization) stays conditional on the re-measure.
|
|
||||||
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
|
|
||||||
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
|
|
||||||
|
|
||||||
## ⚠️ Open items (before PUSHABLE)
|
## ⚠️ Open items (before PUSHABLE)
|
||||||
|
|
||||||
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
|
- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-17
|
||||||
slice-16 build. Judge numerically:
|
build. Judge numerically:
|
||||||
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
|
- startup.log telemetry: conversions/s should jump from ~17-20 to **≥ ~30-40**, `skip busy`
|
||||||
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
|
falling, `ring allocs` ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps).
|
||||||
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
|
- Decode + `/tmp/opencode/tear_audit.py`: desktop-band fresh updates/s up toward the render
|
||||||
|
cap (~30+), frozen % well under 50%, no genuine mid-frame splits.
|
||||||
- ffprobe: video ≈ audio ≈ wall.
|
- ffprobe: video ≈ audio ≈ wall.
|
||||||
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
|
- If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots →
|
||||||
composite downscale / blit-on-change), still staying 60fps.
|
**C4 slice** (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags,
|
||||||
|
the bind is GPU contention — re-measure before touching anything.
|
||||||
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
|
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
|
||||||
|
|
||||||
## Open threads (carried)
|
## Open threads (carried)
|
||||||
@@ -77,7 +75,7 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
|
|||||||
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
|
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
|
||||||
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
|
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
|
||||||
- Signed A/V sync: verify the negative (advance) direction on device.
|
- Signed A/V sync: verify the negative (advance) direction on device.
|
||||||
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
|
- Focus-loss capture lag — closed as NOT the cause (1824: delivery healthy ~60-100/s).
|
||||||
|
|
||||||
## Landmines
|
## Landmines
|
||||||
|
|
||||||
@@ -88,14 +86,15 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT
|
|||||||
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
|
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
|
||||||
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
|
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
|
||||||
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
|
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
|
||||||
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
|
- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the
|
||||||
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
|
**deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and now the **WGC-CLIP** + two
|
||||||
recipe** — grep before re-deriving.
|
slice blocks — grep before re-deriving.
|
||||||
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
|
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
|
||||||
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
|
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
|
||||||
|
|
||||||
## Next step
|
## Next step
|
||||||
|
|
||||||
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
|
Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line
|
||||||
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
|
(conversions/s ≥ ~30-40, `skip busy` falling) + `tear_audit.py` cadence + ffprobe durations. If
|
||||||
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.
|
the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice
|
||||||
|
next, then re-measure clap offset (`/tmp/opencode/avsync.py`), then decide push with the user.
|
||||||
@@ -400,6 +400,35 @@ the acceleration; re-measure after slice 15.
|
|||||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||||
op fast.
|
op fast.
|
||||||
|
|
||||||
|
**Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real
|
||||||
|
wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a
|
||||||
|
bug):**
|
||||||
|
Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was
|
||||||
|
still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new
|
||||||
|
telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the
|
||||||
|
**GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of
|
||||||
|
that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight =
|
||||||
|
readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two
|
||||||
|
follow-on decisions fixed by evidence:
|
||||||
|
- **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking
|
||||||
|
the pool to the master size to read back less. Microsoft's screen-capture page
|
||||||
|
forbids it: "the underlying Direct3D surface is always the size specified when
|
||||||
|
creating … the Direct3D11CaptureFramePool. If content is larger than the frame,
|
||||||
|
the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor,
|
||||||
|
not scale — silently encode the desktop cut off. Readback must stay native; the
|
||||||
|
lever is concurrency, not size. **Rule: read the platform doc for the exact
|
||||||
|
primitive before "fixing" the pool/format; scaling assumptions about capture APIs
|
||||||
|
have been wrong twice now.**
|
||||||
|
- **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in
|
||||||
|
flight, completions can land out of order; a slow OLDER readback finishing last
|
||||||
|
would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear).
|
||||||
|
`MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by
|
||||||
|
compare-exchange publish) drops stale completions instead. Ring gets a lock because
|
||||||
|
rents are now concurrent; downscale row-scratch became per-conversion locals.
|
||||||
|
Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before
|
||||||
|
choosing between throughput and latency fixes — both previous slices picked the
|
||||||
|
wrong slot ("render" vs "conversion") until the audit existed.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -25,6 +25,7 @@ internal sealed class FrameRingBuffer
|
|||||||
private readonly byte[]?[] _slots;
|
private readonly byte[]?[] _slots;
|
||||||
private readonly long[] _lastHandout;
|
private readonly long[] _lastHandout;
|
||||||
private readonly int _redLine;
|
private readonly int _redLine;
|
||||||
|
private readonly object _lock = new();
|
||||||
private int _next;
|
private int _next;
|
||||||
private long _seq;
|
private long _seq;
|
||||||
private long _allocations;
|
private long _allocations;
|
||||||
@@ -38,18 +39,26 @@ internal sealed class FrameRingBuffer
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// <summary>Buffers freshly allocated after the last <see cref="ConsumeAllocations"/>.</summary>
|
/// <summary>Buffers freshly allocated after the last <see cref="ConsumeAllocations"/>.</summary>
|
||||||
public long Allocations => _allocations;
|
public long Allocations
|
||||||
|
{
|
||||||
|
get { lock (_lock) return _allocations; }
|
||||||
|
}
|
||||||
|
|
||||||
/// <summary>Returns the allocation count since the last call and resets it.</summary>
|
/// <summary>Returns the allocation count since the last call and resets it.</summary>
|
||||||
public long ConsumeAllocations()
|
public long ConsumeAllocations()
|
||||||
|
{
|
||||||
|
lock (_lock)
|
||||||
{
|
{
|
||||||
var count = _allocations;
|
var count = _allocations;
|
||||||
_allocations = 0;
|
_allocations = 0;
|
||||||
return count;
|
return count;
|
||||||
}
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// <summary>Hands out a scratch buffer of <paramref name="size"/> bytes.</summary>
|
/// <summary>Hands out a scratch buffer of <paramref name="size"/> bytes.</summary>
|
||||||
public byte[] Rent(int size)
|
public byte[] Rent(int size)
|
||||||
|
{
|
||||||
|
lock (_lock)
|
||||||
{
|
{
|
||||||
var seq = ++_seq;
|
var seq = ++_seq;
|
||||||
for (var tries = 0; tries < _slots.Length; tries++)
|
for (var tries = 0; tries < _slots.Length; tries++)
|
||||||
@@ -78,3 +87,4 @@ internal sealed class FrameRingBuffer
|
|||||||
return new byte[size];
|
return new byte[size];
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
using System;
|
||||||
|
using System.Threading;
|
||||||
|
|
||||||
|
namespace ytLive.Services;
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// An atomic "publish only if strictly newer" gate for the capture's latest-wins
|
||||||
|
/// hand-out. Screen capture conversions may now overlap (slice 17), so two
|
||||||
|
/// conversions can finish briefly out of order; without a monotonic gate the slower
|
||||||
|
/// (older) completion would overwrite the faster (newer) one's LatestFrame and the
|
||||||
|
/// compositor would render STALE content — a time hole in the other direction.
|
||||||
|
/// Sequence must be a strictly increasing per-source counter
|
||||||
|
/// (<c>Interlocked.Increment</c> in <c>ScreenCaptureFrameSource</c>).
|
||||||
|
/// </summary>
|
||||||
|
internal sealed class MonotonicGate
|
||||||
|
{
|
||||||
|
private long _last;
|
||||||
|
|
||||||
|
/// <summary>Returns true iff <paramref name="seq"/> is strictly newer than every
|
||||||
|
/// seq accepted so far (and records it).</summary>
|
||||||
|
public bool TryPublish(long seq)
|
||||||
|
{
|
||||||
|
while (true)
|
||||||
|
{
|
||||||
|
var current = Interlocked.Read(ref _last);
|
||||||
|
if (seq <= current) return false;
|
||||||
|
if (Interlocked.CompareExchange(ref _last, seq, current) == current) return true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -2,6 +2,7 @@ using System;
|
|||||||
using System.Diagnostics;
|
using System.Diagnostics;
|
||||||
using System.Runtime.InteropServices;
|
using System.Runtime.InteropServices;
|
||||||
using System.Runtime.InteropServices.WindowsRuntime;
|
using System.Runtime.InteropServices.WindowsRuntime;
|
||||||
|
using System.Threading;
|
||||||
using System.Threading.Tasks;
|
using System.Threading.Tasks;
|
||||||
using Windows.Graphics;
|
using Windows.Graphics;
|
||||||
using Windows.Graphics.Capture;
|
using Windows.Graphics.Capture;
|
||||||
@@ -28,7 +29,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
private GraphicsCaptureSession? _session;
|
private GraphicsCaptureSession? _session;
|
||||||
private SizeInt32 _poolSize;
|
private SizeInt32 _poolSize;
|
||||||
private bool _started;
|
private bool _started;
|
||||||
private bool _framePending;
|
private int _convertInFlight;
|
||||||
private DateTime _lastErrorLog = DateTime.MinValue;
|
private DateTime _lastErrorLog = DateTime.MinValue;
|
||||||
|
|
||||||
// Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh
|
// Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh
|
||||||
@@ -43,8 +44,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
// consumers (the compositor's paste cache) cannot false-hit a recycled array.
|
// consumers (the compositor's paste cache) cannot false-hit a recycled array.
|
||||||
private readonly FrameRingBuffer _ring = new(8, redLine: 4);
|
private readonly FrameRingBuffer _ring = new(8, redLine: 4);
|
||||||
private long _epoch;
|
private long _epoch;
|
||||||
private byte[]? _row0;
|
private readonly MonotonicGate _publishGate = new();
|
||||||
private byte[]? _row1;
|
|
||||||
|
|
||||||
// The composition master frame (see ai.md "Resolution tiers"): the background
|
// The composition master frame (see ai.md "Resolution tiers"): the background
|
||||||
// is an input layer, so we never hold a CPU frame bigger than the master.
|
// is an input layer, so we never hold a CPU frame bigger than the master.
|
||||||
@@ -54,14 +54,19 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
// A failing conversion must not re-flood the log at frame rate.
|
// A failing conversion must not re-flood the log at frame rate.
|
||||||
private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5);
|
private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5);
|
||||||
|
|
||||||
// Conversion cadence (slice 16, 2026-09-14): the monitor delivers at the 240Hz DWM
|
// Conversion cadence (slice 16, 2026-09-14; slice 17, 2026-09-14): the monitor
|
||||||
// cadence (~4.2ms) while slots are 16.6ms. A 10ms floor between conversion starts
|
// delivers at the 240Hz DWM cadence (~4.2ms) while slots are 16.6ms. A 10ms floor
|
||||||
// keeps the open edge above the ~60 conversions/s the pump can actually use, so the
|
// between conversion starts keeps the open edge above the ~60 conversions/s the pump
|
||||||
// 240Hz tail stops chewing a conversion thread that the 1742 take measured at
|
// can use. What actually bounded the desktop feed on takes 1742/1824 was the SERIAL
|
||||||
// ~150ms/frame (a 90%-frozen desktop, ~6 fresh frames/s). Drop counters and the
|
// GPU→CPU readback (`CreateCopyFromSurfaceAsync` ≈ 40-50ms of the ~47ms conversion
|
||||||
// rolling conversion stats feed the 2-second startup.log telemetry line.
|
// on a 240Hz-HDR box shared with the encoder), capping captures at ~17-20/s.
|
||||||
|
// Slice 17 overlaps up to MaxConcurrentConversions readbacks (pool sized to
|
||||||
|
// accommodate in-flight frames) and publishes only monotonically newer frames
|
||||||
|
// (MonotonicGate — a slow older completion must never overwrite a newer LatestFrame).
|
||||||
private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10);
|
private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10);
|
||||||
private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2);
|
private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2);
|
||||||
|
private const int MaxConcurrentConversions = 3;
|
||||||
|
private const int PoolBufferCount = 5;
|
||||||
private DateTime _lastConvertAt = DateTime.MinValue;
|
private DateTime _lastConvertAt = DateTime.MinValue;
|
||||||
private DateTime _telemetryFrom = DateTime.UtcNow;
|
private DateTime _telemetryFrom = DateTime.UtcNow;
|
||||||
private DateTime _lastTelemetry = DateTime.UtcNow;
|
private DateTime _lastTelemetry = DateTime.UtcNow;
|
||||||
@@ -107,7 +112,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
var item = _item;
|
var item = _item;
|
||||||
var device = Direct3D11Helper.CreateDevice();
|
var device = Direct3D11Helper.CreateDevice();
|
||||||
var framePool = Direct3D11CaptureFramePool.CreateFreeThreaded(
|
var framePool = Direct3D11CaptureFramePool.CreateFreeThreaded(
|
||||||
device, DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, item.Size);
|
device, DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, item.Size);
|
||||||
_poolSize = item.Size;
|
_poolSize = item.Size;
|
||||||
var session = framePool.CreateCaptureSession(item);
|
var session = framePool.CreateCaptureSession(item);
|
||||||
framePool.FrameArrived += OnFrameArrived;
|
framePool.FrameArrived += OnFrameArrived;
|
||||||
@@ -124,7 +129,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
lock (_gate)
|
lock (_gate)
|
||||||
{
|
{
|
||||||
_started = false;
|
_started = false;
|
||||||
_framePending = false;
|
_convertInFlight = 0;
|
||||||
if (_framePool != null)
|
if (_framePool != null)
|
||||||
_framePool.FrameArrived -= OnFrameArrived;
|
_framePool.FrameArrived -= OnFrameArrived;
|
||||||
_session?.Dispose();
|
_session?.Dispose();
|
||||||
@@ -151,14 +156,15 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
|
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
|
||||||
{
|
{
|
||||||
sender.Recreate(Direct3D11Helper.CreateDevice(),
|
sender.Recreate(Direct3D11Helper.CreateDevice(),
|
||||||
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
|
DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, frame.ContentSize);
|
||||||
_poolSize = frame.ContentSize;
|
_poolSize = frame.ContentSize;
|
||||||
}
|
}
|
||||||
|
|
||||||
// One conversion at a time, spaced by MinConvertInterval (slice 16): the
|
// Up to MaxConcurrentConversions readbacks in flight (slice 17), spaced by
|
||||||
// 240Hz delivery otherwise queued a conversion every ~4.2ms and the
|
// MinConvertInterval (slice 16): the 240Hz delivery otherwise queued one
|
||||||
// 1742 take's ~150ms conversion pinned the desktop layer ~90% frozen.
|
// conversion every ~4.2ms, and the serial ~47ms readback pinned the desktop
|
||||||
if (_framePending)
|
// layer to ~17 updates/s on the 1824 take.
|
||||||
|
if (_convertInFlight >= MaxConcurrentConversions)
|
||||||
{
|
{
|
||||||
_skippedBusy++;
|
_skippedBusy++;
|
||||||
frame.Dispose();
|
frame.Dispose();
|
||||||
@@ -172,7 +178,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
_lastConvertAt = now;
|
_lastConvertAt = now;
|
||||||
_framePending = true;
|
_convertInFlight++;
|
||||||
}
|
}
|
||||||
_ = ProcessFrameAsync(frame);
|
_ = ProcessFrameAsync(frame);
|
||||||
}
|
}
|
||||||
@@ -186,7 +192,14 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
using (var softwareBitmap = await SoftwareBitmap.CreateCopyFromSurfaceAsync(
|
using (var softwareBitmap = await SoftwareBitmap.CreateCopyFromSurfaceAsync(
|
||||||
frame.Surface, BitmapAlphaMode.Ignore))
|
frame.Surface, BitmapAlphaMode.Ignore))
|
||||||
{
|
{
|
||||||
FrameAvailable?.Invoke(CopyToVideoFrame(softwareBitmap));
|
// Monotonic sequencing across the overlapping conversions (slice 17): a
|
||||||
|
// completed readback is published only if strictly newer than the last
|
||||||
|
// one published — a slow older completion must never overwrite a newer
|
||||||
|
// LatestFrame (that would be a time hole in the other direction).
|
||||||
|
var epoch = Interlocked.Increment(ref _epoch);
|
||||||
|
var videoFrame = CopyToVideoFrame(softwareBitmap, epoch);
|
||||||
|
if (_publishGate.TryPublish(epoch))
|
||||||
|
FrameAvailable?.Invoke(videoFrame);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
catch (Exception ex)
|
catch (Exception ex)
|
||||||
@@ -203,8 +216,8 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
lock (_gate)
|
lock (_gate)
|
||||||
{
|
{
|
||||||
EmitTelemetry(sw.ElapsedMilliseconds);
|
EmitTelemetry(sw.ElapsedMilliseconds);
|
||||||
|
_convertInFlight--;
|
||||||
}
|
}
|
||||||
_framePending = false;
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -234,7 +247,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
_skippedCadence = 0;
|
_skippedCadence = 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap)
|
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap, long epoch)
|
||||||
{
|
{
|
||||||
var sw = bitmap.PixelWidth;
|
var sw = bitmap.PixelWidth;
|
||||||
var sh = bitmap.PixelHeight;
|
var sh = bitmap.PixelHeight;
|
||||||
@@ -257,12 +270,12 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
// DWM delivers an opaque surface (alpha 255); bilinear keeps it 255.
|
// DWM delivers an opaque surface (alpha 255); bilinear keeps it 255.
|
||||||
var scaled = _ring.Rent(dw * dh * 4);
|
var scaled = _ring.Rent(dw * dh * 4);
|
||||||
return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled))
|
return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled))
|
||||||
{ IsOpaque = true, Epoch = ++_epoch };
|
{ IsOpaque = true, Epoch = epoch };
|
||||||
}
|
}
|
||||||
|
|
||||||
var pixels = _ring.Rent(count);
|
var pixels = _ring.Rent(count);
|
||||||
Marshal.Copy(data, pixels, 0, pixels.Length);
|
Marshal.Copy(data, pixels, 0, pixels.Length);
|
||||||
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = ++_epoch };
|
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = epoch };
|
||||||
}
|
}
|
||||||
|
|
||||||
// Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16,
|
// Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16,
|
||||||
@@ -275,9 +288,11 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
|||||||
// row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output.
|
// row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output.
|
||||||
private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst)
|
private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst)
|
||||||
{
|
{
|
||||||
// Row scratch is per-capture-thread and reused across frames (same churn lesson).
|
// Row scratch is per-conversion (overlapping conversions since slice 17 each
|
||||||
var row0 = _row0 != null && _row0.Length >= srcStride ? _row0 : (_row0 = new byte[srcStride]);
|
// bring their own — two 10KB arrays, no shared state). Size: a 2560-wide row
|
||||||
var row1 = _row1 != null && _row1.Length >= srcStride ? _row1 : (_row1 = new byte[srcStride]);
|
// pair, the largest the monitor path delivers before the downscale.
|
||||||
|
var row0 = new byte[srcStride];
|
||||||
|
var row1 = new byte[srcStride];
|
||||||
|
|
||||||
for (var y = 0; y < dh; y++)
|
for (var y = 0; y < dh; y++)
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -327,7 +327,10 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
|||||||
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
|
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
|
||||||
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
|
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
|
||||||
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
|
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
|
||||||
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
|
Conversions run as up to MaxConcurrentConversions (3) overlapping readbacks
|
||||||
|
(slice 17: the OS readback, not the downscale, is the ~47ms wall — see Slice 17) with a
|
||||||
|
monotonic LatestFrame publish gate (`MonotonicGate`: a slow OLDER completion can never
|
||||||
|
overwrite a newer frame), and are spaced by a 10ms floor
|
||||||
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
|
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
|
||||||
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
|
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
|
||||||
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
|
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
|
||||||
@@ -1065,6 +1068,28 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
|
|||||||
measurement recast it — relocating a ~30ms float downscale to the render thread
|
measurement recast it — relocating a ~30ms float downscale to the render thread
|
||||||
just moves the same cost into the slot budget. Re-measure on device; if render
|
just moves the same cost into the slot budget. Re-measure on device; if render
|
||||||
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
|
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
|
||||||
|
- **Slice 17 — the OS readback was the real wall: overlapping conversions + monotonic
|
||||||
|
publish + deeper pool (2026-09-14, device take ty-1824):** slice 16's downscale fix
|
||||||
|
landed but the desktop was still ~90% frozen on the 1824 take (band 6.8/s updates, max
|
||||||
|
freeze 4.85s). The new 2s telemetry was decisive: `conv avg 46-50ms max ~61ms` with
|
||||||
|
`skip busy 37-83` — the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), not
|
||||||
|
`DownscaleBgra`**, is the ~47ms wall (240Hz HDR compositing + encoder + 2 capture
|
||||||
|
devices share the GPU); at one-in-flight that caps the desktop feed at ~17-20
|
||||||
|
updates/s, which is the file's whole-frame ~7 content-moments/s. Two facts reshaped
|
||||||
|
the fix: **(a)** shrinking the pool size does NOT scale the desktop (Microsoft docs:
|
||||||
|
"If content is larger than the frame, the contents are **clipped**") — readback stays
|
||||||
|
at native 2560×1440; **(b)** delivery is healthy (60-100 arrivals/s), so the lever is
|
||||||
|
conversion throughput, not the pool size. Changes in `Services/ScreenCaptureFrameSource.cs`:
|
||||||
|
conversions overlap up to **MaxConcurrentConversions = 3** (pool deepened to 5 buffers
|
||||||
|
so in-flight frames fit), each completion publishes ONLY if its Epoch is strictly
|
||||||
|
newer than the last published (`MonotonicGate` — a slow older completion must never
|
||||||
|
overwrite a newer LatestFrame), Epoch increments via `Interlocked`, and the
|
||||||
|
DownscaleBgra row-scratch became per-conversion locals (concurrent callers). The
|
||||||
|
render side (33-42ms → ~30 unique composites/s) is the NEXT cap after capture speeds
|
||||||
|
up — that's the C4 slice, queued right after this re-measure. **Good Dog test**
|
||||||
|
`PublishGate_TryPublish_OnlyStrictlyNewerWins`. Full suite **296/296 green, 0
|
||||||
|
warnings**. NOT YET DEVICE-VERIFIED; target: telemetry frames/s jumps ≥ ~30-40 and the
|
||||||
|
band audit drops below ~50% frozen. No push.
|
||||||
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
|
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
|
||||||
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
|
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
|
||||||
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
|
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
|
||||||
|
|||||||
@@ -14,6 +14,22 @@ namespace ytLive.Tests;
|
|||||||
/// </summary>
|
/// </summary>
|
||||||
public class ScreenCaptureFrameSourceTests
|
public class ScreenCaptureFrameSourceTests
|
||||||
{
|
{
|
||||||
|
[Fact]
|
||||||
|
public void PublishGate_TryPublish_OnlyStrictlyNewerWins()
|
||||||
|
{
|
||||||
|
// Slice 17: overlapping conversions can finish out of order; the monotonic
|
||||||
|
// gate must keep a slow OLDER completion from overwriting a newer LatestFrame.
|
||||||
|
var gate = new MonotonicGate();
|
||||||
|
Assert.True(gate.TryPublish(1));
|
||||||
|
Assert.True(gate.TryPublish(2));
|
||||||
|
Assert.False(gate.TryPublish(1)); // stale replay (older than 2)
|
||||||
|
Assert.True(gate.TryPublish(3));
|
||||||
|
Assert.False(gate.TryPublish(3)); // duplicate — never accepted twice
|
||||||
|
Assert.False(gate.TryPublish(2)); // late older completion
|
||||||
|
Assert.False(gate.TryPublish(long.MinValue));
|
||||||
|
Assert.True(gate.TryPublish(long.MaxValue));
|
||||||
|
}
|
||||||
|
|
||||||
[Fact]
|
[Fact]
|
||||||
public void Ring_NoLap_ReusesOnlyAfterRedLineRents()
|
public void Ring_NoLap_ReusesOnlyAfterRedLineRents()
|
||||||
{
|
{
|
||||||
|
|||||||
Reference in New Issue
Block a user