diff --git a/HANDOFF.md b/HANDOFF.md index cf5f1cc..302d7d7 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -1,11 +1,11 @@ -# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next) +# HANDOFF — 2026-09-14 (slice-17 concurrent capture conversions committed locally — device verify next) ## Branch / Commit State -`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT -pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix -(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed -audio-sync, pushed). Working tree **clean**. +`main` HEAD = **slice-17 commit** (overlapping capture readbacks — committed LOCALLY, **NOT +pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-16 capture +conversion fix, slice-15 pacing fix (FramePump), `c01206f` (composition capture), `b22d08e` +(signed audio-sync, pushed). Working tree **clean**. ## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive @@ -13,62 +13,60 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`, (`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand. -## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742) +## 🔬 Slice-16 build was still too frozen — take ty-20260914-1824 says the READBACK is the wall -Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer -(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great. -Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/ -tear_audit.py` + `mix_check.py`): +Slice 15 fixed pacing (video 729 frames @60 = 12.15s ≈ audio 12.35s — pacing healthy). But the +creator reports the desktop layer **still looks like missing frames / choppy vs live**. Decoded +`ty-20260914-1824-0000-2.mp4` to `/mnt/c/tmpout/f1824.raw` and audited: +- Desktop band: **6.8 content updates/s, 88% frozen**, one **4.85s freeze** at start. Worse, + not better, than 1742. +- BUT the new slice-16 telemetry proved the downscale fix WORKS: `conv avg 46-50ms, max ~61ms, + ~16-20 conversions/s, skip busy 37-83 per 2s, skip cadence 0, ring allocs 0`. The ring is + steady-state (no allocs); the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`) is ~45ms of + the conversion** on a 240Hz-HDR box sharing the GPU with the encoder. **The wall was never + the CPU downscale.** +- FramePump worst render 165ms startup spike → 33-42ms sustained (whole-frame ~7.1/s ⇒ render + is the SECOND cap, ~30 unique composites/s). +- Delivery is healthy (60-100 arrivals/s) ⇒ focus-loss OS throttling is NOT the cause (that + theory is now closed). -- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze - runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`) - was real but MOOT. -- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM - cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p - `DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** → - `LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on - a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as - reported. +## 🔬 Committed locally — slice 17: overlapping readbacks + monotonic publish + deeper pool -## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring +**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/MonotonicGate.cs` + locked +`Services/FrameRingBuffer.cs`): +1. **MaxConcurrentConversions = 3** readbacks in flight (was one-in-flight `_framePending`), + pool buffers 2 → **5** so in-flight frames fit. +2. **MonotonicLatest publish gate** (`MonotonicGate`, new internal): a completed readback is + published ONLY if its Epoch is strictly newer than the last published. Overlapping + conversions can finish out of order; a slow OLDER completion must never overwrite a newer + `LatestFrame` (backwards time hole = the mirror of the 1742 tear). Epoch via + `Interlocked.Increment`. +3. `FrameRingBuffer.Rent`/`ConsumeAllocations` now take `_lock` (rents are concurrent); the + 10ms floor and downscale stay; `DownscaleBgra` row-scratch is per-conversion locals (no + shared `_row0/_row1`). +4. **RESEARCH near-miss (docs'ed, MyMistakes):** shrinking the pool to 1920×1080 would have + been wrong — Microsoft Docs (screen capture): *"the underlying Direct3D surface is always + the size specified … **clipped**"* to the frame. Readback stays native; the lever is + concurrency. -**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`): - -1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage - math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost. -2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of - serialized conversion/s; capacity sits just above the ~60/s the pump can use. -3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only - rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that - needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher - preview lags). -4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`, - `ring allocs` — so the next device take is judged numerically. - -**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents` -(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and -never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`, -`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same -commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE + -lessons + the recast note), this HANDOFF. - -**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side -Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement: -relocating a 30ms float downscale to the render thread just moves the same cost into the slot -budget. C4 (compositor optimization) stays conditional on the re-measure. -No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit -was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched. +**Good Dog test:** `ScreenCaptureFrameSourceTests.PublishGate_TryPublish_OnlyStrictlyNewerWins`. +**296/296 green, app build 0 warnings.** Scope-lock files (6): `Services/ScreenCaptureFrameSource.cs`, +`Services/FrameRingBuffer.cs`, `Services/MonotonicGate.cs` (new), `ytLive.Tests/ScreenCaptureFrameSourceTests.cs` ++ docs (ai.md Slice 17, MyMistakes slice-17 block, this HANDOFF). `SceneCompositor.cs`/`FramePump.cs` +NOT touched (C4 is the next slice, pending this re-measure). ## ⚠️ Open items (before PUSHABLE) -- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the - slice-16 build. Judge numerically: - - startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady). - - Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen - % in the single digits, no mid-frame split survivors (social-bar strip churn is fine). +- **Device re-verify (next step):** creator records the SAME tv-show scenario on the slice-17 + build. Judge numerically: + - startup.log telemetry: conversions/s should jump from ~17-20 to **≥ ~30-40**, `skip busy` + falling, `ring allocs` ≈ 0, conv avg still ~40-50ms (readback isn't free, it just overlaps). + - Decode + `/tmp/opencode/tear_audit.py`: desktop-band fresh updates/s up toward the render + cap (~30+), frozen % well under 50%, no genuine mid-frame splits. - ffprobe: video ≈ audio ≈ wall. -- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached - composite downscale / blit-on-change), still staying 60fps. +- If capture now feeds ≥ render's unique-composite rate and render still busts 16.6ms slots → + **C4 slice** (Epoch-cached composite / blit-on-change), still 60fps. If capture still lags, + the bind is GPU contention — re-measure before touching anything. - **No push yet** — commit-locally-until-greenlight for web/A/V work. ## Open threads (carried) @@ -77,7 +75,7 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT - Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued. - Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22). - Signed A/V sync: verify the negative (advance) direction on device. -- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open. +- Focus-loss capture lag — closed as NOT the cause (1824: delivery healthy ~60-100/s). ## Landmines @@ -88,14 +86,15 @@ was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT only `./scripts/verify.sh ""`'s clean build counts. Building `ytLive.csproj` alone does NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`. - ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths. -- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the - **A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ - recipe** — grep before re-deriving. +- `MyMistakes.md` has the **freeze-audit RECIPE**, the **A/V sync measurement recipe**, the + **deadline-pacing** lessons, the **CoreMessaging DQ recipe**, and now the **WGC-CLIP** + two + slice blocks — grep before re-deriving. - sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`. - `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo. ## Next step -Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line + -`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed; -re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user. \ No newline at end of file +Creator records a tv-show take on the slice-17 build → read the startup.log telemetry line +(conversions/s ≥ ~30-40, `skip busy` falling) + `tear_audit.py` cadence + ffprobe durations. If +the desktop now tracks the render cap (~30+ updates/s, <50% frozen, no splits): C4 render slice +next, then re-measure clap offset (`/tmp/opencode/avsync.py`), then decide push with the user. \ No newline at end of file diff --git a/MyMistakes.md b/MyMistakes.md index 0a7daa5..6aee0d4 100644 --- a/MyMistakes.md +++ b/MyMistakes.md @@ -393,12 +393,41 @@ the acceleration; re-measure after slice 15. dispatcher preview copy lags, so "who released it" is unknowable without a consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742 tear (new-top/old-bottom midway) is that read-under-write closed. - - **measure before trusting the inherited plan:** the approved native-res capture + - composite-side downscale (C1) was recast to "fix the downscale in place" — - relocating a 30ms float downscale from the capture thread to the render thread - and caching by Epoch only moves the same ~30ms cost into the slot budget. The - measurement said the cost ITSELF was the enemy; keep the architecture, make the - op fast. +- **measure before trusting the inherited plan:** the approved native-res capture + + composite-side downscale (C1) was recast to "fix the downscale in place" — + relocating a 30ms float downscale from the capture thread to the render thread + and caching by Epoch only moves the same ~30ms cost into the slot budget. The + measurement said the cost ITSELF was the enemy; keep the architecture, make the + op fast. + + **Slice 17 follow-up (2026-09-14) — the readback, not the downscale, was the real + wall; and the pool-size lever was a trap (RESEARCH fact — would have shipped a + bug):** + Slice 16's downscale fix landed (20-50ms conversion, healthy) yet take ty-1824 was + still ~90% frozen in the desktop band (6.8 updates/s, 4.85s max freeze). The new + telemetry said it plainly: `conv avg 46-50ms, max ~61ms, skip busy 37-83` — the + **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), NOT `DownscaleBgra`**, is ~45ms of + that conversion on a 240Hz-HDR box sharing the GPU with the encoder. One-in-flight = + readback-bound at ~17-20 conversions/s — the real cap the whole way down. Two + follow-on decisions fixed by evidence: + - **pool size is a clip, not a scale (near-miss).** The "obvious" fix was shrinking + the pool to the master size to read back less. Microsoft's screen-capture page + forbids it: "the underlying Direct3D surface is always the size specified when + creating … the Direct3D11CaptureFramePool. If content is larger than the frame, + the contents are **clipped**." Shrinking to 1920×1080 would CROP a 1440p monitor, + not scale — silently encode the desktop cut off. Readback must stay native; the + lever is concurrency, not size. **Rule: read the platform doc for the exact + primitive before "fixing" the pool/format; scaling assumptions about capture APIs + have been wrong twice now.** + - **overlap the readbacks + a monotonic publish gate.** With up to 3 conversions in + flight, completions can land out of order; a slow OLDER readback finishing last + would stomp a newer frame (a backwards time hole — the mirror of the 1742 tear). + `MonotonicGate` (seq set via `Interlocked.Increment` before the copy, verified by + compare-exchange publish) drops stale completions instead. Ring gets a lock because + rents are now concurrent; downscale row-scratch became per-conversion locals. + Lesson: keep a *single* "what is bound?" number per layer (telemetry line) before + choosing between throughput and latency fixes — both previous slices picked the + wrong slot ("render" vs "conversion") until the audit existed. --- diff --git a/Services/FrameRingBuffer.cs b/Services/FrameRingBuffer.cs index 1ae63b0..bfe019b 100644 --- a/Services/FrameRingBuffer.cs +++ b/Services/FrameRingBuffer.cs @@ -25,6 +25,7 @@ internal sealed class FrameRingBuffer private readonly byte[]?[] _slots; private readonly long[] _lastHandout; private readonly int _redLine; + private readonly object _lock = new(); private int _next; private long _seq; private long _allocations; @@ -38,43 +39,52 @@ internal sealed class FrameRingBuffer } /// Buffers freshly allocated after the last . - public long Allocations => _allocations; + public long Allocations + { + get { lock (_lock) return _allocations; } + } /// Returns the allocation count since the last call and resets it. public long ConsumeAllocations() { - var count = _allocations; - _allocations = 0; - return count; + lock (_lock) + { + var count = _allocations; + _allocations = 0; + return count; + } } /// Hands out a scratch buffer of bytes. public byte[] Rent(int size) { - var seq = ++_seq; - for (var tries = 0; tries < _slots.Length; tries++) + lock (_lock) { - var idx = (_next + tries) % _slots.Length; - // Red line: rewriting this slot could hit a frame a consumer still reads. - if (seq - _lastHandout[idx] < _redLine) continue; - _next = (idx + 1) % _slots.Length; - if (_slots[idx] is { Length: var len } buf && len == size) + var seq = ++_seq; + for (var tries = 0; tries < _slots.Length; tries++) { + var idx = (_next + tries) % _slots.Length; + // Red line: rewriting this slot could hit a frame a consumer still reads. + if (seq - _lastHandout[idx] < _redLine) continue; + _next = (idx + 1) % _slots.Length; + if (_slots[idx] is { Length: var len } buf && len == size) + { + _lastHandout[idx] = seq; + return buf; + } + + // Length mismatch (or never allocated): a fresh array, never an in-place + // overwrite — the previous loan's bytes stay valid for whoever holds it. + _allocations++; + var fresh = new byte[size]; + _slots[idx] = fresh; _lastHandout[idx] = seq; - return buf; + return fresh; } - // Length mismatch (or never allocated): a fresh array, never an in-place - // overwrite — the previous loan's bytes stay valid for whoever holds it. + // Defensive: the whole ring is inside its red line — do not lap a loaned slot. _allocations++; - var fresh = new byte[size]; - _slots[idx] = fresh; - _lastHandout[idx] = seq; - return fresh; + return new byte[size]; } - - // Defensive: the whole ring is inside its red line — do not lap a loaned slot. - _allocations++; - return new byte[size]; } } \ No newline at end of file diff --git a/Services/MonotonicGate.cs b/Services/MonotonicGate.cs new file mode 100644 index 0000000..be3b178 --- /dev/null +++ b/Services/MonotonicGate.cs @@ -0,0 +1,30 @@ +using System; +using System.Threading; + +namespace ytLive.Services; + +/// +/// An atomic "publish only if strictly newer" gate for the capture's latest-wins +/// hand-out. Screen capture conversions may now overlap (slice 17), so two +/// conversions can finish briefly out of order; without a monotonic gate the slower +/// (older) completion would overwrite the faster (newer) one's LatestFrame and the +/// compositor would render STALE content — a time hole in the other direction. +/// Sequence must be a strictly increasing per-source counter +/// (Interlocked.Increment in ScreenCaptureFrameSource). +/// +internal sealed class MonotonicGate +{ + private long _last; + + /// Returns true iff is strictly newer than every + /// seq accepted so far (and records it). + public bool TryPublish(long seq) + { + while (true) + { + var current = Interlocked.Read(ref _last); + if (seq <= current) return false; + if (Interlocked.CompareExchange(ref _last, seq, current) == current) return true; + } + } +} \ No newline at end of file diff --git a/Services/ScreenCaptureFrameSource.cs b/Services/ScreenCaptureFrameSource.cs index 6018404..97ded68 100644 --- a/Services/ScreenCaptureFrameSource.cs +++ b/Services/ScreenCaptureFrameSource.cs @@ -2,6 +2,7 @@ using System; using System.Diagnostics; using System.Runtime.InteropServices; using System.Runtime.InteropServices.WindowsRuntime; +using System.Threading; using System.Threading.Tasks; using Windows.Graphics; using Windows.Graphics.Capture; @@ -28,7 +29,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource private GraphicsCaptureSession? _session; private SizeInt32 _poolSize; private bool _started; - private bool _framePending; + private int _convertInFlight; private DateTime _lastErrorLog = DateTime.MinValue; // Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh @@ -43,8 +44,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource // consumers (the compositor's paste cache) cannot false-hit a recycled array. private readonly FrameRingBuffer _ring = new(8, redLine: 4); private long _epoch; - private byte[]? _row0; - private byte[]? _row1; + private readonly MonotonicGate _publishGate = new(); // The composition master frame (see ai.md "Resolution tiers"): the background // is an input layer, so we never hold a CPU frame bigger than the master. @@ -54,14 +54,19 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource // A failing conversion must not re-flood the log at frame rate. private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5); - // Conversion cadence (slice 16, 2026-09-14): the monitor delivers at the 240Hz DWM - // cadence (~4.2ms) while slots are 16.6ms. A 10ms floor between conversion starts - // keeps the open edge above the ~60 conversions/s the pump can actually use, so the - // 240Hz tail stops chewing a conversion thread that the 1742 take measured at - // ~150ms/frame (a 90%-frozen desktop, ~6 fresh frames/s). Drop counters and the - // rolling conversion stats feed the 2-second startup.log telemetry line. + // Conversion cadence (slice 16, 2026-09-14; slice 17, 2026-09-14): the monitor + // delivers at the 240Hz DWM cadence (~4.2ms) while slots are 16.6ms. A 10ms floor + // between conversion starts keeps the open edge above the ~60 conversions/s the pump + // can use. What actually bounded the desktop feed on takes 1742/1824 was the SERIAL + // GPU→CPU readback (`CreateCopyFromSurfaceAsync` ≈ 40-50ms of the ~47ms conversion + // on a 240Hz-HDR box shared with the encoder), capping captures at ~17-20/s. + // Slice 17 overlaps up to MaxConcurrentConversions readbacks (pool sized to + // accommodate in-flight frames) and publishes only monotonically newer frames + // (MonotonicGate — a slow older completion must never overwrite a newer LatestFrame). private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10); private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2); + private const int MaxConcurrentConversions = 3; + private const int PoolBufferCount = 5; private DateTime _lastConvertAt = DateTime.MinValue; private DateTime _telemetryFrom = DateTime.UtcNow; private DateTime _lastTelemetry = DateTime.UtcNow; @@ -107,7 +112,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource var item = _item; var device = Direct3D11Helper.CreateDevice(); var framePool = Direct3D11CaptureFramePool.CreateFreeThreaded( - device, DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, item.Size); + device, DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, item.Size); _poolSize = item.Size; var session = framePool.CreateCaptureSession(item); framePool.FrameArrived += OnFrameArrived; @@ -124,7 +129,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource lock (_gate) { _started = false; - _framePending = false; + _convertInFlight = 0; if (_framePool != null) _framePool.FrameArrived -= OnFrameArrived; _session?.Dispose(); @@ -151,14 +156,15 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height) { sender.Recreate(Direct3D11Helper.CreateDevice(), - DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize); + DirectXPixelFormat.B8G8R8A8UIntNormalized, PoolBufferCount, frame.ContentSize); _poolSize = frame.ContentSize; } - // One conversion at a time, spaced by MinConvertInterval (slice 16): the - // 240Hz delivery otherwise queued a conversion every ~4.2ms and the - // 1742 take's ~150ms conversion pinned the desktop layer ~90% frozen. - if (_framePending) + // Up to MaxConcurrentConversions readbacks in flight (slice 17), spaced by + // MinConvertInterval (slice 16): the 240Hz delivery otherwise queued one + // conversion every ~4.2ms, and the serial ~47ms readback pinned the desktop + // layer to ~17 updates/s on the 1824 take. + if (_convertInFlight >= MaxConcurrentConversions) { _skippedBusy++; frame.Dispose(); @@ -172,7 +178,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource return; } _lastConvertAt = now; - _framePending = true; + _convertInFlight++; } _ = ProcessFrameAsync(frame); } @@ -186,7 +192,14 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource using (var softwareBitmap = await SoftwareBitmap.CreateCopyFromSurfaceAsync( frame.Surface, BitmapAlphaMode.Ignore)) { - FrameAvailable?.Invoke(CopyToVideoFrame(softwareBitmap)); + // Monotonic sequencing across the overlapping conversions (slice 17): a + // completed readback is published only if strictly newer than the last + // one published — a slow older completion must never overwrite a newer + // LatestFrame (that would be a time hole in the other direction). + var epoch = Interlocked.Increment(ref _epoch); + var videoFrame = CopyToVideoFrame(softwareBitmap, epoch); + if (_publishGate.TryPublish(epoch)) + FrameAvailable?.Invoke(videoFrame); } } catch (Exception ex) @@ -203,8 +216,8 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource lock (_gate) { EmitTelemetry(sw.ElapsedMilliseconds); + _convertInFlight--; } - _framePending = false; } } @@ -234,7 +247,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource _skippedCadence = 0; } - private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap) + private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap, long epoch) { var sw = bitmap.PixelWidth; var sh = bitmap.PixelHeight; @@ -257,12 +270,12 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource // DWM delivers an opaque surface (alpha 255); bilinear keeps it 255. var scaled = _ring.Rent(dw * dh * 4); return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled)) - { IsOpaque = true, Epoch = ++_epoch }; + { IsOpaque = true, Epoch = epoch }; } var pixels = _ring.Rent(count); Marshal.Copy(data, pixels, 0, pixels.Length); - return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = ++_epoch }; + return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = epoch }; } // Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16, @@ -275,9 +288,11 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource // row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output. private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst) { - // Row scratch is per-capture-thread and reused across frames (same churn lesson). - var row0 = _row0 != null && _row0.Length >= srcStride ? _row0 : (_row0 = new byte[srcStride]); - var row1 = _row1 != null && _row1.Length >= srcStride ? _row1 : (_row1 = new byte[srcStride]); + // Row scratch is per-conversion (overlapping conversions since slice 17 each + // bring their own — two 10KB arrays, no shared state). Size: a 2560-wide row + // pair, the largest the monitor path delivers before the downscale. + var row0 = new byte[srcStride]; + var row1 = new byte[srcStride]; for (var y = 0; y < dh; y++) { diff --git a/ai.md b/ai.md index 8bdb59f..d8724ab 100644 --- a/ai.md +++ b/ai.md @@ -327,7 +327,10 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr 1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear — slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take). - Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor + Conversions run as up to MaxConcurrentConversions (3) overlapping readbacks + (slice 17: the OS readback, not the downscale, is the ~47ms wall — see Slice 17) with a + monotonic LatestFrame publish gate (`MonotonicGate`: a slow OLDER completion can never + overwrite a newer frame), and are spaced by a 10ms floor (`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry). Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a @@ -1065,6 +1068,28 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t measurement recast it — relocating a ~30ms float downscale to the render thread just moves the same cost into the slot budget. Re-measure on device; if render still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push. +- **Slice 17 — the OS readback was the real wall: overlapping conversions + monotonic + publish + deeper pool (2026-09-14, device take ty-1824):** slice 16's downscale fix + landed but the desktop was still ~90% frozen on the 1824 take (band 6.8/s updates, max + freeze 4.85s). The new 2s telemetry was decisive: `conv avg 46-50ms max ~61ms` with + `skip busy 37-83` — the **GPU→CPU readback (`CreateCopyFromSurfaceAsync`), not + `DownscaleBgra`**, is the ~47ms wall (240Hz HDR compositing + encoder + 2 capture + devices share the GPU); at one-in-flight that caps the desktop feed at ~17-20 + updates/s, which is the file's whole-frame ~7 content-moments/s. Two facts reshaped + the fix: **(a)** shrinking the pool size does NOT scale the desktop (Microsoft docs: + "If content is larger than the frame, the contents are **clipped**") — readback stays + at native 2560×1440; **(b)** delivery is healthy (60-100 arrivals/s), so the lever is + conversion throughput, not the pool size. Changes in `Services/ScreenCaptureFrameSource.cs`: + conversions overlap up to **MaxConcurrentConversions = 3** (pool deepened to 5 buffers + so in-flight frames fit), each completion publishes ONLY if its Epoch is strictly + newer than the last published (`MonotonicGate` — a slow older completion must never + overwrite a newer LatestFrame), Epoch increments via `Interlocked`, and the + DownscaleBgra row-scratch became per-conversion locals (concurrent callers). The + render side (33-42ms → ~30 unique composites/s) is the NEXT cap after capture speeds + up — that's the C4 slice, queued right after this re-measure. **Good Dog test** + `PublishGate_TryPublish_OnlyStrictlyNewerWins`. Full suite **296/296 green, 0 + warnings**. NOT YET DEVICE-VERIFIED; target: telemetry frames/s jumps ≥ ~30-40 and the + band audit drops below ~50% frozen. No push. - **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits; an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was diff --git a/ytLive.Tests/ScreenCaptureFrameSourceTests.cs b/ytLive.Tests/ScreenCaptureFrameSourceTests.cs index 33741b7..79da91a 100644 --- a/ytLive.Tests/ScreenCaptureFrameSourceTests.cs +++ b/ytLive.Tests/ScreenCaptureFrameSourceTests.cs @@ -14,6 +14,22 @@ namespace ytLive.Tests; /// public class ScreenCaptureFrameSourceTests { + [Fact] + public void PublishGate_TryPublish_OnlyStrictlyNewerWins() + { + // Slice 17: overlapping conversions can finish out of order; the monotonic + // gate must keep a slow OLDER completion from overwriting a newer LatestFrame. + var gate = new MonotonicGate(); + Assert.True(gate.TryPublish(1)); + Assert.True(gate.TryPublish(2)); + Assert.False(gate.TryPublish(1)); // stale replay (older than 2) + Assert.True(gate.TryPublish(3)); + Assert.False(gate.TryPublish(3)); // duplicate — never accepted twice + Assert.False(gate.TryPublish(2)); // late older completion + Assert.False(gate.TryPublish(long.MinValue)); + Assert.True(gate.TryPublish(long.MaxValue)); + } + [Fact] public void Ring_NoLap_ReusesOnlyAfterRedLineRents() {