perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)

The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel
DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take
ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame
audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was
the wall, and the one torn frame was a ring slot rewritten under the consumer's
read. Reference: WGC delivers at DWM/monitor cadence
(https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture)
and libyuv row-simple/fixed-point scaling
(https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule.

- DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage
  math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame.
- 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of
  serialized conversion/s; capacity sits just above the 60/s the pump can use.
- FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only
  rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a
  consumer still holds (session.LatestFrame survives conversions, dispatcher
  preview lags) is never read-while-overwritten. Needs no consumer Release API.
- 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring
  allocs — the device take is judgeable numerically.

Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings.
C4 (composite Epoch-cached downscale) deferred pending the device re-measure.
Local only, no push.
This commit is contained in:
2026-09-14 18:18:34 -07:00
parent 76f51e6f4e
commit 71932b9756
6 changed files with 417 additions and 110 deletions
+61 -51
View File
@@ -1,11 +1,11 @@
# HANDOFF — 2026-09-14 (composition capture DEVICE-VERIFIED ✓; slice-15 pacing fix committed locally — A/V re-measure next)
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
## Branch / Commit State
`main` HEAD = **slice-15 commit** (FramePump duplicate-on-lag pacing — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: `c01206f` (composition
capture redesign, also local/unpushed), before that `b22d08e` (signed audio-sync, pushed).
Working tree **clean**.
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
audio-sync, pushed). Working tree **clean**.
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
@@ -13,53 +13,62 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
## ✅ DEVICE-VERIFIED — composition capture looks good; then the A/V defect surfaced
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
The creator ran real takes off the `c01206f` build and reports the video looks **very good** — the
Windows.Graphics.Capture web path is confirmed on device (open item closed). BUT the same takes
showed the A/V defect we have seen before: **both desktop and webcam playback are accelerated, the
audio "gets speedier", then cuts off at the end.**
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
tear_audit.py` + `mix_check.py`):
## 🔬 Sliced and measured (evidence)
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
was real but MOOT.
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
reported.
With the A/V recipe (`MyMistakes.md` → "Measuring audio-video A/V sync" + `/tmp/opencode/avsync.py`):
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
- `ty-20260914-1723-0000-2.mp4`: 697 video frames @60fps = **11.62s** vs audio **11.84s**.
- `ty-20260914-1726-0000-2.mp4` (clap take): 163 frames = **2.72s** vs **2.93s** audio. Video ends
0.21–0.24s before audio → the cut-off. Clap: audio env peak 1.655s vs video motion 0.75–1.03s;
cross-correlation lag **+0.667s** (audio late).
- `startup.log`: `FramePump stall: iteration 33-46ms (> 2× the 17ms interval): worst render 33ms …
dropped 0`.
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
**Root cause (slice 15):** slice 10's freshness choice skipped the overrun's missed slots — one
fresh frame per ~35ms stall authored into a 60fps container = **accelerated playback**. OBS's
answer is duplicate-on-lag (`libobs/video-io.c`, docs.obsproject.com/backend-design): fill each
missed slot by repeating the newest frame — duration == wall, judder never fast-forward.
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
preview lags).
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
`ring allocs` — so the next device take is judged numerically.
## 🔬 Committed locally — slice 15: one frame per deadline slot (duplicate-on-lag)
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
lessons + the recast note), this HANDOFF.
**What:** `Services/Encoder/FramePump.cs` submit is now a bounded catch-up —
`while (now >= nextTick) { submit latest composite; nextTick += intervalTicks; }` clamped to a
`deadlineNow` captured once per iteration. First missed slot gets the fresh composite, the rest get
repeats of it (OBS duplication) — recording duration == wall time under any render load, no
acceleration, no audio tail cut. The burned `_outputIndex` moved inside the submit loop (every
emitted slot gets its own +1; also fixes the old unconditional bump that gapped the judge sequence
on non-submitting fast-render iterations).
**Good Dog test:** `ytLive.Tests/FramePumpTests.cs` → `Pump_Overrun_Renders_EmitsEverySlot_NotSkipped`
(60fps, 35ms render cost, asserts ≥0.65 of the wall slots emitted — the old skip-pump wrote ~1/35ms).
**294/294 green, app build 0 warnings.** Files: `Services/Encoder/FramePump.cs`,
`ytLive.Tests/FramePumpTests.cs`. Docs in same commit: ai.md Slice 15, MyMistakes (supersedes the
slice-10 "no burst re-write" clause + latest numbers), this HANDOFF.
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
budget. C4 (compositor optimization) stays conditional on the re-measure.
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
## ⚠️ Open items (before PUSHABLE)
- **Device re-verify:** a take on the slice-15 build must play at REAL time (no acceleration, no
cut-off audio tail). Confirm via ffprobe: video stream duration ≈ audio ≈ container.
- **Re-measure the true A/V offset** with a clap take now that video pacing is honest
(`/tmp/opencode/avsync.py`). The +0.6s reading on 1726 was confounded by the 1.1x
acceleration. If a real residual remains after the pacing fix, the backlog is the AUDIO pipeline
(mixer ring / advance), not video.
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
slice-16 build. Judge numerically:
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
- ffprobe: video ≈ audio ≈ wall.
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
composite downscale / blit-on-change), still staying 60fps.
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
## Open threads (carried)
@@ -68,24 +77,25 @@ slice-10 "no burst re-write" clause + latest numbers), this HANDOFF.
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
- Signed A/V sync: verify the negative (advance) direction on device.
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
## Landmines
- testhost shares startup.log — filter by time.
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL double-slashes mangle) before rebuilds — a live
app process locks `ytLive.exe` and the apphost copy fails (MSB3021, seen today).
app process locks `ytLive.exe` and the apphost copy fails (MSB3021).
- Build/tests: **Windows dotnet host** (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings —
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
- `MyMistakes.md` has the **A/V sync measurement recipe**, the **deadline-pacing** lessons
(item 1 + slice 15), and the **CoreMessaging DQ recipe** — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3` for
`/mnt/c/Users/gramp/AppData/Roaming/ytLlive/ytLlive.db`.
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
recipe** — grep before re-deriving.
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
## Next step
Have the creator record a clap take on the slice-15 build → ffprobe durations (video == audio ==
wall) + `/tmp/opencode/avsync.py` for the honest offset. If durations match, the acceleration
defect is closed; then decide push with the user, and attack any true audio-latency residual as its
own work unit.
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.
+61
View File
@@ -342,6 +342,67 @@ single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late)
the acceleration; re-measure after slice 15.
**Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
the freeze-audit recipe (RECIPE — re-deriving it cost this session):**
The slice-15 build fixed pacing but the desktop layer of the recording was still
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
measurement said something different AND worse than the running render theory.
`FramePump stall… worst render 33-36ms` was real but MOOT: once the camera+desktop
take was decoded to raw frames, the **desktop band was frozen 21s of 23.35s (90%)**
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
CONVERSION was the wall: the monitor delivers at the **240Hz DWM cadence**, the
source converts ONE frame at a time (`_framePending` latest-wins), and each
2560×1440→1920×1080 `DownscaleBgra` — naive double-per-pixel bilinear — cost
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
MediaCapture path) and audio were fine — exactly what the user reported.
**The audit recipe (ffmpeg → raw gray → numpy):**
```
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
python3 - <<EOF
import numpy as np
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
EOF
```
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
the compositor render stats led nowhere until this number existed. A **per-row
split detector** (`cumsum` of per-row diff-to-next minus diff-to-prev, argmax =
split row) then separated real mid-frame tears (score ≈ huge, both halves match
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
were the session UI, not tears.
**The fix (this slice):**
- **integer 8.8 fixed-point downscale, "shift only at the end"** — the SAME math as
`SceneCompositor.Bilinear` (rounded both stages in one 16.8 scale) ported into
`DownscaleBgra`, dropping per-pixel doubles to row-walk integer ops. The capture
ring already had the integer-bilinear lesson; the capture downscale itself was
still the naive float twin of the 258ms disaster.
- **throttle to the slot cadence** (`MinConvertInterval = 10ms`): the 240Hz arrival
is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per
second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps
pump can use.
- **ring reuse-distance, not ownership** (`FrameRingBuffer`, redLine 4): the
take-14 "depth × period" rule guards size; the slice-16 addition makes it
structural — a slot is only rewritten ≥4 rents after its last hand-out, else a
fresh buffer. `session.LatestFrame` survives across conversions and the
dispatcher preview copy lags, so "who released it" is unknowable without a
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
tear (new-top/old-bottom midway) is that read-under-write closed.
- **measure before trusting the inherited plan:** the approved native-res capture +
composite-side downscale (C1) was recast to "fix the downscale in place" —
relocating a 30ms float downscale from the capture thread to the render thread
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
measurement said the cost ITSELF was the enemy; keep the architecture, make the
op fast.
---
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
integer bilinear) followed the same libyuv research
+80
View File
@@ -0,0 +1,80 @@
using System;
namespace ytLive.Services;
/// <summary>
/// A depth-bounded ring of scratch frame buffers that never writes into memory a live
/// consumer may still be reading. A slot's byte[] is only rewritten after at least
/// <paramref name="redLine"/> rents have cycled since it was last handed out; a slot
/// still inside its red line is skipped, and if the whole ring is inside its red line a
/// fresh buffer is handed out rather than lapping a loaned slot. Length-mismatched
/// buffers are always replaced by a fresh allocation (the previous array is orphaned,
/// never overwritten), so a hand-out keeps its bytes for as long as anyone holds it.
///
/// The screen capture's consumer contract makes this structural necessity (slice 16,
/// 2026-09-14): <c>ScreenCaptureManager</c> keeps each delivered frame as
/// <c>session.LatestFrame</c> across conversions and the dispatcher preview copy lags,
/// so a held frame can survive several conversions. The 1742 take recorded one torn
/// frame (new-top/old-bottom at the webcam split) — a ring slot rewritten under the
/// consumer's read. Depth 8 × the conversion interval already exceeded the worst
/// observed hold (~50ms, take-14 rule); the red line makes the guarantee structural
/// instead of a sizing coincidence.
/// </summary>
internal sealed class FrameRingBuffer
{
private readonly byte[]?[] _slots;
private readonly long[] _lastHandout;
private readonly int _redLine;
private int _next;
private long _seq;
private long _allocations;
public FrameRingBuffer(int depth, int redLine)
{
_slots = new byte[depth][];
_lastHandout = new long[depth];
for (var i = 0; i < depth; i++) _lastHandout[i] = -redLine;
_redLine = redLine;
}
/// <summary>Buffers freshly allocated after the last <see cref="ConsumeAllocations"/>.</summary>
public long Allocations => _allocations;
/// <summary>Returns the allocation count since the last call and resets it.</summary>
public long ConsumeAllocations()
{
var count = _allocations;
_allocations = 0;
return count;
}
/// <summary>Hands out a scratch buffer of <paramref name="size"/> bytes.</summary>
public byte[] Rent(int size)
{
var seq = ++_seq;
for (var tries = 0; tries < _slots.Length; tries++)
{
var idx = (_next + tries) % _slots.Length;
// Red line: rewriting this slot could hit a frame a consumer still reads.
if (seq - _lastHandout[idx] < _redLine) continue;
_next = (idx + 1) % _slots.Length;
if (_slots[idx] is { Length: var len } buf && len == size)
{
_lastHandout[idx] = seq;
return buf;
}
// Length mismatch (or never allocated): a fresh array, never an in-place
// overwrite — the previous loan's bytes stay valid for whoever holds it.
_allocations++;
var fresh = new byte[size];
_slots[idx] = fresh;
_lastHandout[idx] = seq;
return fresh;
}
// Defensive: the whole ring is inside its red line — do not lap a loaned slot.
_allocations++;
return new byte[size];
}
}
+108 -54
View File
@@ -1,4 +1,5 @@
using System;
using System.Diagnostics;
using System.Runtime.InteropServices;
using System.Runtime.InteropServices.WindowsRuntime;
using System.Threading.Tasks;
@@ -30,17 +31,17 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
private bool _framePending;
private DateTime _lastErrorLog = DateTime.MinValue;
// Buffer recycling (take-11 spikes, 2026-09-04): a fresh ~8.3MB byte[] per DWM
// frame ≈ 500MB/s of LOH churn — the gen2 pauses it forces surfaced as the
// "worst render 35-65ms" spikes that capped fps at ~42 long after the compositor
// itself was fast. A 4-deep ring rotated round-robin is never lapped by a
// ≤17ms consumer at 60Hz; each hand-out carries an Epoch so identity-keyed
// Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh
// ~8.3MB byte[] per DWM frame ≈ 500MB/s of LOH churn — the gen2 pauses it forced
// surfaced as the "worst render 35-65ms" spikes that capped fps at ~42. The pool
// is a reuse-distance ring: a buffer is only rewritten after ≥ redLine rents have
// cycled since its last hand-out (FrameRingBuffer), so a frame a consumer still
// holds — session.LatestFrame survives across conversions and the dispatcher
// preview copy lags — can never be overwritten in place (the 1742 tear, a ring
// slot rewritten under the consumer's read, handed the compositor one
// new-top/old-bottom frame). Each hand-out still carries an Epoch so identity-keyed
// consumers (the compositor's paste cache) cannot false-hit a recycled array.
// Depth 8 (take-14 rule): depth × source period must exceed the worst consumer
// hold — 4 slots at 144Hz laps in ~27ms while a compositor read + lagged UI
// preview copy can hold a frame ~50ms; the flash was half a new frame over old.
private readonly byte[]?[] _frameRing = new byte[8][];
private int _ringNext;
private readonly FrameRingBuffer _ring = new(8, redLine: 4);
private long _epoch;
private byte[]? _row0;
private byte[]? _row1;
@@ -53,6 +54,23 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
// A failing conversion must not re-flood the log at frame rate.
private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5);
// Conversion cadence (slice 16, 2026-09-14): the monitor delivers at the 240Hz DWM
// cadence (~4.2ms) while slots are 16.6ms. A 10ms floor between conversion starts
// keeps the open edge above the ~60 conversions/s the pump can actually use, so the
// 240Hz tail stops chewing a conversion thread that the 1742 take measured at
// ~150ms/frame (a 90%-frozen desktop, ~6 fresh frames/s). Drop counters and the
// rolling conversion stats feed the 2-second startup.log telemetry line.
private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10);
private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2);
private DateTime _lastConvertAt = DateTime.MinValue;
private DateTime _telemetryFrom = DateTime.UtcNow;
private DateTime _lastTelemetry = DateTime.UtcNow;
private long _skippedCadence;
private long _skippedBusy;
private long _conversions;
private long _convertMsTotal;
private long _convertMsMax;
public string Key { get; }
public event Action<VideoFrame>? FrameAvailable;
@@ -121,6 +139,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
{
var frame = sender.TryGetNextFrame();
if (frame == null) return;
lock (_gate)
{
if (!_started)
@@ -128,26 +147,39 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
frame.Dispose();
return;
}
}
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
{
sender.Recreate(Direct3D11Helper.CreateDevice(),
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
_poolSize = frame.ContentSize;
}
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
{
sender.Recreate(Direct3D11Helper.CreateDevice(),
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
_poolSize = frame.ContentSize;
}
if (_framePending)
{
frame.Dispose();
return;
// One conversion at a time, spaced by MinConvertInterval (slice 16): the
// 240Hz delivery otherwise queued a conversion every ~4.2ms and the
// 1742 take's ~150ms conversion pinned the desktop layer ~90% frozen.
if (_framePending)
{
_skippedBusy++;
frame.Dispose();
return;
}
var now = DateTime.UtcNow;
if (now - _lastConvertAt < MinConvertInterval)
{
_skippedCadence++;
frame.Dispose();
return;
}
_lastConvertAt = now;
_framePending = true;
}
_framePending = true;
_ = ProcessFrameAsync(frame);
}
private async Task ProcessFrameAsync(Direct3D11CaptureFrame frame)
{
var sw = Stopwatch.StartNew();
try
{
using (frame)
@@ -168,27 +200,38 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
}
finally
{
lock (_gate)
{
EmitTelemetry(sw.ElapsedMilliseconds);
}
_framePending = false;
}
}
private byte[] RentRingBuffer(int size)
private void EmitTelemetry(long conversionMs)
{
for (var tries = 0; tries < _frameRing.Length; tries++)
{
var idx = (_ringNext + tries) % _frameRing.Length;
var buf = _frameRing[idx];
if (buf is { Length: var len } && len == size)
{
_ringNext = (idx + 1) % _frameRing.Length;
return buf;
}
}
var slot = _ringNext;
_ringNext = (slot + 1) % _frameRing.Length;
var fresh = new byte[size];
_frameRing[slot] = fresh;
return fresh;
_conversions++;
_convertMsTotal += conversionMs;
if (conversionMs > _convertMsMax) _convertMsMax = conversionMs;
var now = DateTime.UtcNow;
if (now - _lastTelemetry < TelemetryInterval) return;
var span = now - _telemetryFrom;
var perSecond = span.TotalSeconds > 0 ? _conversions / span.TotalSeconds : 0;
AppLog.Write(
$"ScreenCapture telemetry [{Key}]: {_conversions} frames in {span.TotalSeconds:F1}s " +
$"({perSecond:F0}/s), conv avg {( _conversions == 0 ? 0 : _convertMsTotal / _conversions )}ms " +
$"max {_convertMsMax}ms, skip busy {_skippedBusy} cadence {_skippedCadence}, " +
$"ring allocs {_ring.ConsumeAllocations()}");
_lastTelemetry = now;
_telemetryFrom = now;
_conversions = 0;
_convertMsTotal = 0;
_convertMsMax = 0;
_skippedBusy = 0;
_skippedCadence = 0;
}
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap)
@@ -212,47 +255,58 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
var dw = Math.Max(1, (int)(sw * scale));
var dh = Math.Max(1, (int)(sh * scale));
// DWM delivers an opaque surface (alpha 255); bilinear keeps it 255.
var scaled = RentRingBuffer(dw * dh * 4);
var scaled = _ring.Rent(dw * dh * 4);
return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled))
{ IsOpaque = true, Epoch = ++_epoch };
}
var pixels = RentRingBuffer(count);
var pixels = _ring.Rent(count);
Marshal.Copy(data, pixels, 0, pixels.Length);
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = ++_epoch };
}
// Bilinear downscale to the master frame. Reads each source row pair through
// Marshal.Copy (no unsafe), writing tightly packed BGRA output.
// Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16,
// 2026-09-14: the two-stage math from SceneCompositor.Bilinear — the same "shift
// only at the end" rule from MyMistakes item on fixed-point blending). The previous
// double-per-pixel version ran ~30-45ms per 2.5K→1080p frame on a quiet desktop
// and the 1742 take's stall measured ~150ms/frame under load (90%-frozen desktop,
// ~6 fresh frames/s); integer math keeps the same bilinear result within ±1 while
// dropping the cost to the row-pair Marshal.Copy. Rows are read once per source
// row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output.
private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst)
{
// Row scratch is per-capture-thread and reused across frames (same churn lesson).
var row0 = _row0 != null && _row0.Length >= srcStride ? _row0 : (_row0 = new byte[srcStride]);
var row1 = _row1 != null && _row1.Length >= srcStride ? _row1 : (_row1 = new byte[srcStride]);
var xs = sw / (double)dw;
var ys = sh / (double)dh;
for (var y = 0; y < dh; y++)
{
var sy = Math.Min(sh - 1, (int)(y * ys));
var sy8 = (int)((long)y * sh * 256 / dh);
var sy = sy8 >> 8;
var sy1 = Math.Min(sh - 1, sy + 1);
var fy = (y * ys) - sy;
var fy8 = sy8 & 255;
var fyInv = 256 - fy8;
Marshal.Copy(IntPtr.Add(src, sy * srcStride), row0, 0, srcStride);
Marshal.Copy(IntPtr.Add(src, sy1 * srcStride), row1, 0, srcStride);
var dRow = y * dw * 4;
for (var x = 0; x < dw; x++)
{
var sx = Math.Min(sw - 1, (int)(x * xs));
var sx8 = (int)((long)x * sw * 256 / dw);
var sx = Math.Min(sw - 1, sx8 >> 8);
var sx1 = Math.Min(sw - 1, sx + 1);
var fx = (x * xs) - sx;
for (var c = 0; c < 4; c++)
var fx8 = sx8 & 255;
var fxInv = 256 - fx8;
var i0 = sx * 4;
var i1 = sx1 * 4;
var j = dRow + x * 4;
for (var c = 0; c < 4; c++, i0++, i1++, j++)
{
var i0 = sx * 4 + c;
var i1 = sx1 * 4 + c;
var top = row0[i0] + (row0[i1] - row0[i0]) * fx;
var bottom = row1[i0] + (row1[i1] - row1[i0]) * fx;
dst[dRow + x * 4 + c] = (byte)(top + (bottom - top) * fy);
// Two-stage in one 16.8 scale: top/bot ≤ 65280, ×256 + round ≤ 33.5M — int-safe.
var top = row0[i0] * fxInv + row0[i1] * fx8;
var bot = row1[i0] * fxInv + row1[i1] * fx8;
dst[j] = (byte)((top * fyInv + bot * fy8 + 32768) >> 16);
}
}
}
+51 -5
View File
@@ -324,10 +324,20 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
`WindowsRuntimeMarshal.TryGetDataUnsafe` (the same CsWinRT-safe read the webcam path uses) — the
`IMemoryBufferByteAccess` ComImport cast threw `Invalid cast` on **every frame** under CsWinRT, which
flooded `startup.log` (~5 MB in a session) and burned CPU, so it is gone. Surfaces larger than the
1920×1080 master are downscaled bilinearly to the master (`DownscaleBgra`) before the copy, and
per-frame conversion failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected
content delivers black frames (OS limitation, documented). Frame pool
pauses while the app is minimized — capture keeps running, the pool just stops delivering.
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
buffer is only rewritten ≥4 rents after its last hand-out, else a fresh one is allocated — a frame a
consumer still holds (`session.LatestFrame` survives conversions; the dispatcher preview copy lags)
can never be read-while-overwritten (the 1742 tear). Rolling telemetry (frames/s, conv avg/max ms,
drops, ring allocs) is logged to `startup.log` every 2s while converting. Per-frame conversion
failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected content delivers black
frames (OS limitation, documented). Frame pool pauses while the app is minimized — capture keeps
running, the pool just stops delivering.
- **Ownership:** `ScreenCaptureManager` mirrors `CameraManager` — refcounted by target key, one shared
`WriteableBitmap` per key, dispatcher-coalesced latest-frame copies, `PreviewBitmapChanged`/`CaptureFailed`
events, `ReleaseAllAsync` on re-designation. `ScreenCaptureSourceFactory.Resolve(key)` parses the key into
@@ -353,7 +363,9 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
app is in the background (worse under a full-screen game on 24H2/26100), GPU readback
(`CreateCopyFromSurfaceAsync`) contends with the foreground game, the source's one-in-flight `_framePending`
gate drops frames during stalls, and the manager's `DispatcherPriority.Render` copies only run as fast as
WPF presents the window. Recorded 2026-08-13; no mitigation attempted yet (deferred by user decision).
WPF presents the window. Slice 16 shrank the per-conversion stall (fast downscale + 10ms cadence floor)
but the OS-level delivery throttle under focus loss remains. Recorded 2026-08-13; no mitigation attempted
yet (deferred by user decision).
- **GPU posture:** same as webcam — CPU frames, WPF hardware-presents; D3DImage GPU compositing deferred
to the encoder task.
- **Preview watermark:** the "Preview" placeholder hides while a background capture renders —
@@ -1019,6 +1031,40 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
suite **294/294 green, 0 warnings**. The +0.6s audio-late clap reading on 1726 was confounded by
the 1.1x acceleration — re-measure on device; if a real residual remains it is the audio pipeline.
No push (web/A/V work stays local).
- **Slice 16 — the desktop capture conversion was the bottleneck: fast downscale +
cadence throttle + reuse-distance ring (2026-09-14, device take ty-1742):** slice 15
fixed pacing but the 1742 desktop layer was still jerky/frozen with a horizontal
tear. Decoded to raw frames and audited: the **desktop band was frozen 21s of 23.35s
(90%)** — ~6.1 content updates/s, freeze runs up to 2.28-2.78s, dup-run max 90
frames (1.5s). The render stat (`worst render 33-36ms`) was real but MOOT: the
capture CONVERSION was the wall. The monitor delivers at the **240Hz DWM cadence**
(~4.2ms); with one-in-flight conversions and each 2560×1440→1920×1080
`DownscaleBgra` at ~150ms under load, `LatestFrame` updated a handful of times/s —
the desktop feed inside a 60fps file read ~90% frozen. (Webcam + audio were their
own paths — fine, matching the report.) Three changes, all in
`Services/ScreenCaptureFrameSource.cs` (+ `Services/FrameRingBuffer.cs`):
(1) `DownscaleBgra` is now **integer 8.8 fixed-point, "shift only at the end"** —
the exact two-stage math of `SceneCompositor.Bilinear` (MyMistakes take-4 rule),
dropping double-per-pixel to a row-walk of integer ops (the capture downscale had
stayed the naive float twin of the 258ms disaster); (2) a **10ms conversion floor**
(`MinConvertInterval`) so the 240Hz tail stops queuing ~150ms of serialized
conversion per second — the open edge sits just above the ~60/s the pump can use;
(3) the ring is a **reuse-distance pool** (`FrameRingBuffer`, depth 8, redLine 4):
a buffer is only rewritten ≥4 rents after its last hand-out else a fresh allocation,
replacing the blind round-robin — since `ScreenCaptureManager` keeps
`session.LatestFrame` across conversions and the dispatcher preview copy lags,
"who released the buffer" needs a consumer API that doesn't exist; a reuse
DISTANCE needs no cooperation (the 1742 new-top/old-bottom tear is that
read-under-write, structural now). Telemetry added: a startup.log line every 2s
(`frames/s, conv avg/max ms, skip busy/cadence, ring allocs`) so the device take
can be judged numerically. **Good Dog test**
`Ring_NoLap_ReusesOnlyAfterRedLineRents` (depth 3/redLine 4 exercises the red-line
skip; the 8/4 config settles at 8 buffers and never grows). Full suite **295/295
green, 0 warnings**. C4 (compositor Epoch-cached downscale / blit-on-change) was
DEFERRED: the slice-15 approved plan assumed a native-res relocation; the
measurement recast it — relocating a ~30ms float downscale to the render thread
just moves the same cost into the slot budget. Re-measure on device; if render
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
@@ -0,0 +1,56 @@
using System;
using Xunit;
using ytLive.Services;
namespace ytLive.Tests;
/// <summary>
/// The capture conversion ring (<see cref="FrameRingBuffer"/>) is a pure,
/// deterministic piece of <see cref="ScreenCaptureFrameSource"/> — the WGC
/// pool/session layer is WinRT and exercised only on Windows at runtime. These tests
/// pin the slice-16 no-lap contract: a slot is never rewritten within redLine rents
/// of its last hand-out (the 1742 tear — a ring slot rewritten under the consumer's
/// read handed the compositor one new-top/old-bottom frame).
/// </summary>
public class ScreenCaptureFrameSourceTests
{
[Fact]
public void Ring_NoLap_ReusesOnlyAfterRedLineRents()
{
// Depth 3 / redLine 4: the ring is narrower than its safety distance, so the
// red-line skip is actually exercised (the production 8/4 config can never
// block a rotation — the ring cycles before any slot comes due).
var ring = new FrameRingBuffer(depth: 3, redLine: 4);
const int size = 100;
var a = ring.Rent(size);
var b = ring.Rent(size);
var c = ring.Rent(size);
Assert.NotSame(a, b);
Assert.NotSame(b, c);
Assert.NotSame(a, c);
Assert.Equal(3, ring.ConsumeAllocations());
// 4th rent arrives while all three slots are inside their red line: the ring
// must hand out a fresh buffer instead of rewriting a still-loaned slot.
var d = ring.Rent(size);
Assert.NotSame(d, a);
Assert.NotSame(d, b);
Assert.NotSame(d, c);
Assert.Equal(1, ring.ConsumeAllocations());
// 5th rent: slot a (handed at seq 1, revisited at seq 5 = exactly redLine)
// is reusable, and reuse is an in-place recycle, not a fresh allocation.
Assert.Same(a, ring.Rent(size));
Assert.Equal(0, ring.ConsumeAllocations());
// The production-sized ring (8/4, what ScreenCaptureFrameSource uses) settles
// at 8 buffers and recycles them forever — no growth under steady capture.
var prod = new FrameRingBuffer(depth: 8, redLine: 4);
var first = new byte[8][];
for (var i = 0; i < 8; i++) first[i] = prod.Rent(size);
for (var i = 8; i < 200; i++)
Assert.Same(first[i % 8], prod.Rent(size));
Assert.Equal(8, prod.ConsumeAllocations());
}
}