perf(capture): fast integer downscale + 10ms cadence floor + reuse-distance ring (slice 16)
The 240Hz monitor delivery + one-in-flight conversions + naive double-per-pixel DownscaleBgra (~150ms/frame under load) froze the desktop layer 90% of take ty-1742 (6.1 fresh content updates/s, freeze runs to 2.8s; decoded raw-frame audit). The render stat (33-36ms) was real but moot — the capture CONVERSION was the wall, and the one torn frame was a ring slot rewritten under the consumer's read. Reference: WGC delivers at DWM/monitor cadence (https://learn.microsoft.com/en-us/windows/apps/develop/media-authoring-processing/screen-capture) and libyuv row-simple/fixed-point scaling (https://chromium.googlesource.com/libyuv/libyuv/) — the repo's own take-4 rule. - DownscaleBgra: integer 8.8 fixed-point, shift-only-at-the-end (same two-stage math as SceneCompositor.Bilinear). ~150ms -> ~5ms per 2.5K->1080p frame. - 10ms MinConvertInterval: the ~4.2ms 240Hz tail stopped queuing ~150ms of serialized conversion/s; capacity sits just above the 60/s the pump can use. - FrameRingBuffer (depth 8, redLine 4): reuse-DISTANCE ring — a buffer is only rewritten >=4 rents after its last hand-out else fresh-allocated, so a frame a consumer still holds (session.LatestFrame survives conversions, dispatcher preview lags) is never read-while-overwritten. Needs no consumer Release API. - 2s startup.log telemetry: frames/s, conv avg/max ms, skip busy/cadence, ring allocs — the device take is judgeable numerically. Good Dog test: Ring_NoLap_ReusesOnlyAfterRedLineRents. 295/295 green, 0 warnings. C4 (composite Epoch-cached downscale) deferred pending the device re-measure. Local only, no push.
This commit is contained in:
+61
-51
@@ -1,11 +1,11 @@
|
||||
# HANDOFF — 2026-09-14 (composition capture DEVICE-VERIFIED ✓; slice-15 pacing fix committed locally — A/V re-measure next)
|
||||
# HANDOFF — 2026-09-14 (slice-16 capture-conversion fix committed locally — device verify next)
|
||||
|
||||
## Branch / Commit State
|
||||
|
||||
`main` HEAD = **slice-15 commit** (FramePump duplicate-on-lag pacing — committed LOCALLY, **NOT
|
||||
pushed**; web/A/V work stays commit-local until greenlight). Before it: `c01206f` (composition
|
||||
capture redesign, also local/unpushed), before that `b22d08e` (signed audio-sync, pushed).
|
||||
Working tree **clean**.
|
||||
`main` HEAD = **slice-16 commit** (capture conversion bottleneck — committed LOCALLY, **NOT
|
||||
pushed**; web/A/V work stays commit-local until greenlight). Before it: slice-15 pacing fix
|
||||
(FramePump), before that `c01206f` (composition capture), before that `b22d08e` (signed
|
||||
audio-sync, pushed). Working tree **clean**.
|
||||
|
||||
## ⚠️ Branding (2026-09-14, creator-corrected): product = **llamacasty**, internals = ytLive
|
||||
|
||||
@@ -13,53 +13,62 @@ The product is **llamacasty**; repo path, csproj `AssemblyName`/`RootNamespace`,
|
||||
(`%APPDATA%\ytLlive\...`), and most code names are the legacy **ytLive/ytLlive**. User-facing
|
||||
language says "llamacasty"; code/assembly/repo names stay ytLive. See `ai.md` → Brand.
|
||||
|
||||
## ✅ DEVICE-VERIFIED — composition capture looks good; then the A/V defect surfaced
|
||||
## 🔬 Measured the desktop-capture complaint (take ty-20260914-1742)
|
||||
|
||||
The creator ran real takes off the `c01206f` build and reports the video looks **very good** — the
|
||||
Windows.Graphics.Capture web path is confirmed on device (open item closed). BUT the same takes
|
||||
showed the A/V defect we have seen before: **both desktop and webcam playback are accelerated, the
|
||||
audio "gets speedier", then cuts off at the end.**
|
||||
Slice 15 fixed pacing (video 23.35s ≈ audio 23.52s) but the creator reports the desktop layer
|
||||
(tv show) still **jerky / laggy / frozen with a horizontal tear**; webcam + audio are great.
|
||||
Decoded `ty-20260914-1742-0000-2.mp4` (1401 frames) to raw gray and audited (`/tmp/opencode/
|
||||
tear_audit.py` + `mix_check.py`):
|
||||
|
||||
## 🔬 Sliced and measured (evidence)
|
||||
- **Desktop band (rows 40-320) frozen 21s of 23.35s (90%)** — ~6.1 content updates/s, freeze
|
||||
runs up to **2.28-2.78s**, dup-run max 90 frames (1.5s). Render stat (`worst render 33-36ms`)
|
||||
was real but MOOT.
|
||||
- Root cause: the **capture CONVERSION** is the wall. Monitor delivers at the **240Hz DWM
|
||||
cadence** (~4.2ms); one-in-flight conversions (`_framePending`), and each 2560×1440→1080p
|
||||
`DownscaleBgra` (naive double-per-pixel) ≈ 30-45ms quiet / **~150ms under 240Hz-HDR load** →
|
||||
`LatestFrame` updated ~6-9×/s. The tear (one new-top/old-bottom frame) = read-under-write on
|
||||
a recycled ring buffer / DWM readback race. Webcam+audio are separate paths — fine, as
|
||||
reported.
|
||||
|
||||
With the A/V recipe (`MyMistakes.md` → "Measuring audio-video A/V sync" + `/tmp/opencode/avsync.py`):
|
||||
## 🔬 Committed locally — slice 16: fast downscale + cadence throttle + reuse-distance ring
|
||||
|
||||
- `ty-20260914-1723-0000-2.mp4`: 697 video frames @60fps = **11.62s** vs audio **11.84s**.
|
||||
- `ty-20260914-1726-0000-2.mp4` (clap take): 163 frames = **2.72s** vs **2.93s** audio. Video ends
|
||||
0.21–0.24s before audio → the cut-off. Clap: audio env peak 1.655s vs video motion 0.75–1.03s;
|
||||
cross-correlation lag **+0.667s** (audio late).
|
||||
- `startup.log`: `FramePump stall: iteration 33-46ms (> 2× the 17ms interval): worst render 33ms …
|
||||
dropped 0`.
|
||||
**What** (`Services/ScreenCaptureFrameSource.cs` + new `Services/FrameRingBuffer.cs`):
|
||||
|
||||
**Root cause (slice 15):** slice 10's freshness choice skipped the overrun's missed slots — one
|
||||
fresh frame per ~35ms stall authored into a 60fps container = **accelerated playback**. OBS's
|
||||
answer is duplicate-on-lag (`libobs/video-io.c`, docs.obsproject.com/backend-design): fill each
|
||||
missed slot by repeating the newest frame — duration == wall, judder never fast-forward.
|
||||
1. `DownscaleBgra` → **integer 8.8 fixed-point, "shift only at the end"** (the exact two-stage
|
||||
math of `SceneCompositor.Bilinear`). Kills the double-per-pixel float cost.
|
||||
2. **10ms conversion floor** (`MinConvertInterval`): the 240Hz tail stops queuing ~150ms of
|
||||
serialized conversion/s; capacity sits just above the ~60/s the pump can use.
|
||||
3. Ring → **`FrameRingBuffer` reuse-distance pool (depth 8, redLine 4)**: a buffer is only
|
||||
rewritten ≥4 rents after its last hand-out, else a fresh allocation. Structural no-lap that
|
||||
needs no consumer Release API (`session.LatestFrame` survives conversions; dispatcher
|
||||
preview lags).
|
||||
4. **Telemetry:** a startup.log line every 2s — frames/s, conv avg/max ms, `skip busy/cadence`,
|
||||
`ring allocs` — so the next device take is judged numerically.
|
||||
|
||||
## 🔬 Committed locally — slice 15: one frame per deadline slot (duplicate-on-lag)
|
||||
**Good Dog test:** `ScreenCaptureFrameSourceTests.Ring_NoLap_ReusesOnlyAfterRedLineRents`
|
||||
(depth 3/redLine 4 exercises the red-line skip → fresh hand-out; 8/4 settles at 8 buffers and
|
||||
never grows). **295/295 green, app build 0 warnings.** Files: `Services/ScreenCaptureFrameSource.cs`,
|
||||
`Services/FrameRingBuffer.cs`, `ytLive.Tests/ScreenCaptureFrameSourceTests.cs`. Docs in same
|
||||
commit: ai.md (Slice 16 + capture bullet + focus-loss clause), MyMistakes (freeze-audit RECIPE +
|
||||
lessons + the recast note), this HANDOFF.
|
||||
|
||||
**What:** `Services/Encoder/FramePump.cs` submit is now a bounded catch-up —
|
||||
`while (now >= nextTick) { submit latest composite; nextTick += intervalTicks; }` clamped to a
|
||||
`deadlineNow` captured once per iteration. First missed slot gets the fresh composite, the rest get
|
||||
repeats of it (OBS duplication) — recording duration == wall time under any render load, no
|
||||
acceleration, no audio tail cut. The burned `_outputIndex` moved inside the submit loop (every
|
||||
emitted slot gets its own +1; also fixes the old unconditional bump that gapped the judge sequence
|
||||
on non-submitting fast-render iterations).
|
||||
|
||||
**Good Dog test:** `ytLive.Tests/FramePumpTests.cs` → `Pump_Overrun_Renders_EmitsEverySlot_NotSkipped`
|
||||
(60fps, 35ms render cost, asserts ≥0.65 of the wall slots emitted — the old skip-pump wrote ~1/35ms).
|
||||
**294/294 green, app build 0 warnings.** Files: `Services/Encoder/FramePump.cs`,
|
||||
`ytLive.Tests/FramePumpTests.cs`. Docs in same commit: ai.md Slice 15, MyMistakes (supersedes the
|
||||
slice-10 "no burst re-write" clause + latest numbers), this HANDOFF.
|
||||
**Note — approved-plan recast:** the earlier C1 (native-res capture + composite-side
|
||||
Epoch-cached downscale) was recast to "fix the downscale in place" after the measurement:
|
||||
relocating a 30ms float downscale to the render thread just moves the same cost into the slot
|
||||
budget. C4 (compositor optimization) stays conditional on the re-measure.
|
||||
No FPS-tier change, no HDR work (creator decisions preserved). Scope-lock list for this commit
|
||||
was the 6 files above (+docs); `SceneCompositor.cs` and `FramePump.cs` were NOT touched.
|
||||
|
||||
## ⚠️ Open items (before PUSHABLE)
|
||||
|
||||
- **Device re-verify:** a take on the slice-15 build must play at REAL time (no acceleration, no
|
||||
cut-off audio tail). Confirm via ffprobe: video stream duration ≈ audio ≈ container.
|
||||
- **Re-measure the true A/V offset** with a clap take now that video pacing is honest
|
||||
(`/tmp/opencode/avsync.py`). The +0.6s reading on 1726 was confounded by the 1.1x
|
||||
acceleration. If a real residual remains after the pacing fix, the backlog is the AUDIO pipeline
|
||||
(mixer ring / advance), not video.
|
||||
- **Device re-verify (next step):** the creator records the SAME tv-show scenario on the
|
||||
slice-16 build. Judge numerically:
|
||||
- startup.log telemetry: conv avg ≤ ~8ms, frames/s ≥ ~60, `ring allocs` ≈ 0 (steady).
|
||||
- Decode + `/tmp/opencode/tear_audit.py`: ≥ ~55 content updates/s in the desktop band, frozen
|
||||
% in the single digits, no mid-frame split survivors (social-bar strip churn is fine).
|
||||
- ffprobe: video ≈ audio ≈ wall.
|
||||
- If render still busts the 16.6ms slot after capture feeds real updates → C4 (Epoch-cached
|
||||
composite downscale / blit-on-change), still staying 60fps.
|
||||
- **No push yet** — commit-locally-until-greenlight for web/A/V work.
|
||||
|
||||
## Open threads (carried)
|
||||
@@ -68,24 +77,25 @@ slice-10 "no burst re-write" clause + latest numbers), this HANDOFF.
|
||||
- Webcam MJPG missing / ~10–14Hz, layer SortOrder, truncation-with-dynamic-scenes — queued.
|
||||
- Sync control user-doc tutorial — REQUIRED before 1.0 (creator directive; TASK 22).
|
||||
- Signed A/V sync: verify the negative (advance) direction on device.
|
||||
- Focus-loss capture lag (OS-level delivery throttle) — deferred, still open.
|
||||
|
||||
## Landmines
|
||||
|
||||
- testhost shares startup.log — filter by time.
|
||||
- `cmd.exe /c "taskkill /F /IM ytLive.exe"` (WSL double-slashes mangle) before rebuilds — a live
|
||||
app process locks `ytLive.exe` and the apphost copy fails (MSB3021, seen today).
|
||||
app process locks `ytLive.exe` and the apphost copy fails (MSB3021).
|
||||
- Build/tests: **Windows dotnet host** (`/mnt/c/Program Files/dotnet/dotnet.exe`). 0 warnings —
|
||||
only `./scripts/verify.sh "<files>"`'s clean build counts. Building `ytLive.csproj` alone does
|
||||
NOT rebuild `ytLive.Tests.dll` — run the Tests csproj before `vstest`.
|
||||
- ffmpeg/ffprobe: `/mnt/c/Program Files/Krita (x64)/bin/` with Windows paths.
|
||||
- `MyMistakes.md` has the **A/V sync measurement recipe**, the **deadline-pacing** lessons
|
||||
(item 1 + slice 15), and the **CoreMessaging DQ recipe** — grep before re-deriving.
|
||||
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3` for
|
||||
`/mnt/c/Users/gramp/AppData/Roaming/ytLlive/ytLlive.db`.
|
||||
- `MyMistakes.md` has the **freeze-audit RECIPE** (ffmpeg→raw-gray→numpy band audit), the
|
||||
**A/V sync measurement recipe**, the **deadline-pacing** lessons, and the **CoreMessaging DQ
|
||||
recipe** — grep before re-deriving.
|
||||
- sqlite3 at `/home/gramps/android-sdk/platform-tools/sqlite3`.
|
||||
- `C:\tmpout` is for ffmpeg evidence artifacts (raw decodes / PNGs); keep them out of the repo.
|
||||
|
||||
## Next step
|
||||
|
||||
Have the creator record a clap take on the slice-15 build → ffprobe durations (video == audio ==
|
||||
wall) + `/tmp/opencode/avsync.py` for the honest offset. If durations match, the acceleration
|
||||
defect is closed; then decide push with the user, and attack any true audio-latency residual as its
|
||||
own work unit.
|
||||
Creator records a tv-show take on the slice-16 build → read the startup.log telemetry line +
|
||||
`tear_audit.py` cadence + ffprobe durations. If conv~5ms + ≤60 fresh + no splits: defect closed;
|
||||
re-measure the clap offset (`/tmp/opencode/avsync.py`); then decide push with the user.
|
||||
@@ -342,6 +342,67 @@ single-peak offset ≈ +0.64–0.89s, cross-correlation lag +0.667s (audio late)
|
||||
the acceleration; re-measure after slice 15.
|
||||
|
||||
|
||||
**Slice 16 (2026-09-14) — the desktop capture conversion was the bottleneck; here is
|
||||
the freeze-audit recipe (RECIPE — re-deriving it cost this session):**
|
||||
The slice-15 build fixed pacing but the desktop layer of the recording was still
|
||||
"jerky / laggy / frozen with a horizontal tear". Measure, don't guess — and the
|
||||
measurement said something different AND worse than the running render theory.
|
||||
`FramePump stall… worst render 33-36ms` was real but MOOT: once the camera+desktop
|
||||
take was decoded to raw frames, the **desktop band was frozen 21s of 23.35s (90%)**
|
||||
with ~6.1 content updates/s and freeze intervals up to 2.28-2.78s. The capture
|
||||
CONVERSION was the wall: the monitor delivers at the **240Hz DWM cadence**, the
|
||||
source converts ONE frame at a time (`_framePending` latest-wins), and each
|
||||
2560×1440→1920×1080 `DownscaleBgra` — naive double-per-pixel bilinear — cost
|
||||
~30-45ms quiet and ~150ms+ under load (GPU-copy contention on 240Hz HDR). Result:
|
||||
~6-9 fresh frames/s of DESKTOP content inside a 60fps file. The webcam (its own
|
||||
MediaCapture path) and audio were fine — exactly what the user reported.
|
||||
|
||||
**The audit recipe (ffmpeg → raw gray → numpy):**
|
||||
```
|
||||
ffmpeg -i ty-*.mp4 -pix_fmt gray -f rawvideo /mnt/c/tmpout/f.take.raw
|
||||
python3 - <<EOF
|
||||
import numpy as np
|
||||
fr = np.memmap("/mnt/c/tmpout/f.take.raw", np.uint8, mode="r").reshape(n,h,w)
|
||||
band = fr[:,40:320,20:620] # desktop band, skip title/social bars
|
||||
d = [np.abs(band[i].astype(int16)-band[i-1]).mean() for i in range(1,n)]
|
||||
thr = np.percentile(d,25) + 0.5*(np.percentile(d,97)-np.percentile(d,25))
|
||||
print(sum(x>thr for x in d)/ (n/60)) # fresh content-updates/s
|
||||
EOF
|
||||
```
|
||||
"fresh content-updates/s" in the DESKTOP band vs 60 slots is the bottleneck read;
|
||||
the compositor render stats led nowhere until this number existed. A **per-row
|
||||
split detector** (`cumsum` of per-row diff-to-next minus diff-to-prev, argmax =
|
||||
split row) then separated real mid-frame tears (score ≈ huge, both halves match
|
||||
neighbours) from bottom-strip social-bar churn — the pairs it flagged at 97-100%
|
||||
were the session UI, not tears.
|
||||
|
||||
**The fix (this slice):**
|
||||
- **integer 8.8 fixed-point downscale, "shift only at the end"** — the SAME math as
|
||||
`SceneCompositor.Bilinear` (rounded both stages in one 16.8 scale) ported into
|
||||
`DownscaleBgra`, dropping per-pixel doubles to row-walk integer ops. The capture
|
||||
ring already had the integer-bilinear lesson; the capture downscale itself was
|
||||
still the naive float twin of the 258ms disaster.
|
||||
- **throttle to the slot cadence** (`MinConvertInterval = 10ms`): the 240Hz arrival
|
||||
is ~4.2ms — accepting every delivery queues ~150ms of serialized conversion per
|
||||
second minimum; a 10ms floor caps the open edge just above the ~60/s the 60fps
|
||||
pump can use.
|
||||
- **ring reuse-distance, not ownership** (`FrameRingBuffer`, redLine 4): the
|
||||
take-14 "depth × period" rule guards size; the slice-16 addition makes it
|
||||
structural — a slot is only rewritten ≥4 rents after its last hand-out, else a
|
||||
fresh buffer. `session.LatestFrame` survives across conversions and the
|
||||
dispatcher preview copy lags, so "who released it" is unknowable without a
|
||||
consumer API; a reuse-DISTANCE contract needs no consumer cooperation. The 1742
|
||||
tear (new-top/old-bottom midway) is that read-under-write closed.
|
||||
- **measure before trusting the inherited plan:** the approved native-res capture +
|
||||
composite-side downscale (C1) was recast to "fix the downscale in place" —
|
||||
relocating a 30ms float downscale from the capture thread to the render thread
|
||||
and caching by Epoch only moves the same ~30ms cost into the slot budget. The
|
||||
measurement said the cost ITSELF was the enemy; keep the architecture, make the
|
||||
op fast.
|
||||
|
||||
---
|
||||
|
||||
|
||||
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
|
||||
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
|
||||
integer bilinear) followed the same libyuv research
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
using System;
|
||||
|
||||
namespace ytLive.Services;
|
||||
|
||||
/// <summary>
|
||||
/// A depth-bounded ring of scratch frame buffers that never writes into memory a live
|
||||
/// consumer may still be reading. A slot's byte[] is only rewritten after at least
|
||||
/// <paramref name="redLine"/> rents have cycled since it was last handed out; a slot
|
||||
/// still inside its red line is skipped, and if the whole ring is inside its red line a
|
||||
/// fresh buffer is handed out rather than lapping a loaned slot. Length-mismatched
|
||||
/// buffers are always replaced by a fresh allocation (the previous array is orphaned,
|
||||
/// never overwritten), so a hand-out keeps its bytes for as long as anyone holds it.
|
||||
///
|
||||
/// The screen capture's consumer contract makes this structural necessity (slice 16,
|
||||
/// 2026-09-14): <c>ScreenCaptureManager</c> keeps each delivered frame as
|
||||
/// <c>session.LatestFrame</c> across conversions and the dispatcher preview copy lags,
|
||||
/// so a held frame can survive several conversions. The 1742 take recorded one torn
|
||||
/// frame (new-top/old-bottom at the webcam split) — a ring slot rewritten under the
|
||||
/// consumer's read. Depth 8 × the conversion interval already exceeded the worst
|
||||
/// observed hold (~50ms, take-14 rule); the red line makes the guarantee structural
|
||||
/// instead of a sizing coincidence.
|
||||
/// </summary>
|
||||
internal sealed class FrameRingBuffer
|
||||
{
|
||||
private readonly byte[]?[] _slots;
|
||||
private readonly long[] _lastHandout;
|
||||
private readonly int _redLine;
|
||||
private int _next;
|
||||
private long _seq;
|
||||
private long _allocations;
|
||||
|
||||
public FrameRingBuffer(int depth, int redLine)
|
||||
{
|
||||
_slots = new byte[depth][];
|
||||
_lastHandout = new long[depth];
|
||||
for (var i = 0; i < depth; i++) _lastHandout[i] = -redLine;
|
||||
_redLine = redLine;
|
||||
}
|
||||
|
||||
/// <summary>Buffers freshly allocated after the last <see cref="ConsumeAllocations"/>.</summary>
|
||||
public long Allocations => _allocations;
|
||||
|
||||
/// <summary>Returns the allocation count since the last call and resets it.</summary>
|
||||
public long ConsumeAllocations()
|
||||
{
|
||||
var count = _allocations;
|
||||
_allocations = 0;
|
||||
return count;
|
||||
}
|
||||
|
||||
/// <summary>Hands out a scratch buffer of <paramref name="size"/> bytes.</summary>
|
||||
public byte[] Rent(int size)
|
||||
{
|
||||
var seq = ++_seq;
|
||||
for (var tries = 0; tries < _slots.Length; tries++)
|
||||
{
|
||||
var idx = (_next + tries) % _slots.Length;
|
||||
// Red line: rewriting this slot could hit a frame a consumer still reads.
|
||||
if (seq - _lastHandout[idx] < _redLine) continue;
|
||||
_next = (idx + 1) % _slots.Length;
|
||||
if (_slots[idx] is { Length: var len } buf && len == size)
|
||||
{
|
||||
_lastHandout[idx] = seq;
|
||||
return buf;
|
||||
}
|
||||
|
||||
// Length mismatch (or never allocated): a fresh array, never an in-place
|
||||
// overwrite — the previous loan's bytes stay valid for whoever holds it.
|
||||
_allocations++;
|
||||
var fresh = new byte[size];
|
||||
_slots[idx] = fresh;
|
||||
_lastHandout[idx] = seq;
|
||||
return fresh;
|
||||
}
|
||||
|
||||
// Defensive: the whole ring is inside its red line — do not lap a loaned slot.
|
||||
_allocations++;
|
||||
return new byte[size];
|
||||
}
|
||||
}
|
||||
@@ -1,4 +1,5 @@
|
||||
using System;
|
||||
using System.Diagnostics;
|
||||
using System.Runtime.InteropServices;
|
||||
using System.Runtime.InteropServices.WindowsRuntime;
|
||||
using System.Threading.Tasks;
|
||||
@@ -30,17 +31,17 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
private bool _framePending;
|
||||
private DateTime _lastErrorLog = DateTime.MinValue;
|
||||
|
||||
// Buffer recycling (take-11 spikes, 2026-09-04): a fresh ~8.3MB byte[] per DWM
|
||||
// frame ≈ 500MB/s of LOH churn — the gen2 pauses it forces surfaced as the
|
||||
// "worst render 35-65ms" spikes that capped fps at ~42 long after the compositor
|
||||
// itself was fast. A 4-deep ring rotated round-robin is never lapped by a
|
||||
// ≤17ms consumer at 60Hz; each hand-out carries an Epoch so identity-keyed
|
||||
// Buffer recycling (take-11 spikes, 2026-09-04; slice 16, 2026-09-14): a fresh
|
||||
// ~8.3MB byte[] per DWM frame ≈ 500MB/s of LOH churn — the gen2 pauses it forced
|
||||
// surfaced as the "worst render 35-65ms" spikes that capped fps at ~42. The pool
|
||||
// is a reuse-distance ring: a buffer is only rewritten after ≥ redLine rents have
|
||||
// cycled since its last hand-out (FrameRingBuffer), so a frame a consumer still
|
||||
// holds — session.LatestFrame survives across conversions and the dispatcher
|
||||
// preview copy lags — can never be overwritten in place (the 1742 tear, a ring
|
||||
// slot rewritten under the consumer's read, handed the compositor one
|
||||
// new-top/old-bottom frame). Each hand-out still carries an Epoch so identity-keyed
|
||||
// consumers (the compositor's paste cache) cannot false-hit a recycled array.
|
||||
// Depth 8 (take-14 rule): depth × source period must exceed the worst consumer
|
||||
// hold — 4 slots at 144Hz laps in ~27ms while a compositor read + lagged UI
|
||||
// preview copy can hold a frame ~50ms; the flash was half a new frame over old.
|
||||
private readonly byte[]?[] _frameRing = new byte[8][];
|
||||
private int _ringNext;
|
||||
private readonly FrameRingBuffer _ring = new(8, redLine: 4);
|
||||
private long _epoch;
|
||||
private byte[]? _row0;
|
||||
private byte[]? _row1;
|
||||
@@ -53,6 +54,23 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
// A failing conversion must not re-flood the log at frame rate.
|
||||
private static readonly TimeSpan ErrorLogThrottle = TimeSpan.FromSeconds(5);
|
||||
|
||||
// Conversion cadence (slice 16, 2026-09-14): the monitor delivers at the 240Hz DWM
|
||||
// cadence (~4.2ms) while slots are 16.6ms. A 10ms floor between conversion starts
|
||||
// keeps the open edge above the ~60 conversions/s the pump can actually use, so the
|
||||
// 240Hz tail stops chewing a conversion thread that the 1742 take measured at
|
||||
// ~150ms/frame (a 90%-frozen desktop, ~6 fresh frames/s). Drop counters and the
|
||||
// rolling conversion stats feed the 2-second startup.log telemetry line.
|
||||
private static readonly TimeSpan MinConvertInterval = TimeSpan.FromMilliseconds(10);
|
||||
private static readonly TimeSpan TelemetryInterval = TimeSpan.FromSeconds(2);
|
||||
private DateTime _lastConvertAt = DateTime.MinValue;
|
||||
private DateTime _telemetryFrom = DateTime.UtcNow;
|
||||
private DateTime _lastTelemetry = DateTime.UtcNow;
|
||||
private long _skippedCadence;
|
||||
private long _skippedBusy;
|
||||
private long _conversions;
|
||||
private long _convertMsTotal;
|
||||
private long _convertMsMax;
|
||||
|
||||
public string Key { get; }
|
||||
public event Action<VideoFrame>? FrameAvailable;
|
||||
|
||||
@@ -121,6 +139,7 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
{
|
||||
var frame = sender.TryGetNextFrame();
|
||||
if (frame == null) return;
|
||||
|
||||
lock (_gate)
|
||||
{
|
||||
if (!_started)
|
||||
@@ -128,26 +147,39 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
frame.Dispose();
|
||||
return;
|
||||
}
|
||||
}
|
||||
|
||||
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
|
||||
{
|
||||
sender.Recreate(Direct3D11Helper.CreateDevice(),
|
||||
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
|
||||
_poolSize = frame.ContentSize;
|
||||
}
|
||||
if (frame.ContentSize.Width != _poolSize.Width || frame.ContentSize.Height != _poolSize.Height)
|
||||
{
|
||||
sender.Recreate(Direct3D11Helper.CreateDevice(),
|
||||
DirectXPixelFormat.B8G8R8A8UIntNormalized, 2, frame.ContentSize);
|
||||
_poolSize = frame.ContentSize;
|
||||
}
|
||||
|
||||
if (_framePending)
|
||||
{
|
||||
frame.Dispose();
|
||||
return;
|
||||
// One conversion at a time, spaced by MinConvertInterval (slice 16): the
|
||||
// 240Hz delivery otherwise queued a conversion every ~4.2ms and the
|
||||
// 1742 take's ~150ms conversion pinned the desktop layer ~90% frozen.
|
||||
if (_framePending)
|
||||
{
|
||||
_skippedBusy++;
|
||||
frame.Dispose();
|
||||
return;
|
||||
}
|
||||
var now = DateTime.UtcNow;
|
||||
if (now - _lastConvertAt < MinConvertInterval)
|
||||
{
|
||||
_skippedCadence++;
|
||||
frame.Dispose();
|
||||
return;
|
||||
}
|
||||
_lastConvertAt = now;
|
||||
_framePending = true;
|
||||
}
|
||||
_framePending = true;
|
||||
_ = ProcessFrameAsync(frame);
|
||||
}
|
||||
|
||||
private async Task ProcessFrameAsync(Direct3D11CaptureFrame frame)
|
||||
{
|
||||
var sw = Stopwatch.StartNew();
|
||||
try
|
||||
{
|
||||
using (frame)
|
||||
@@ -168,27 +200,38 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
}
|
||||
finally
|
||||
{
|
||||
lock (_gate)
|
||||
{
|
||||
EmitTelemetry(sw.ElapsedMilliseconds);
|
||||
}
|
||||
_framePending = false;
|
||||
}
|
||||
}
|
||||
|
||||
private byte[] RentRingBuffer(int size)
|
||||
private void EmitTelemetry(long conversionMs)
|
||||
{
|
||||
for (var tries = 0; tries < _frameRing.Length; tries++)
|
||||
{
|
||||
var idx = (_ringNext + tries) % _frameRing.Length;
|
||||
var buf = _frameRing[idx];
|
||||
if (buf is { Length: var len } && len == size)
|
||||
{
|
||||
_ringNext = (idx + 1) % _frameRing.Length;
|
||||
return buf;
|
||||
}
|
||||
}
|
||||
var slot = _ringNext;
|
||||
_ringNext = (slot + 1) % _frameRing.Length;
|
||||
var fresh = new byte[size];
|
||||
_frameRing[slot] = fresh;
|
||||
return fresh;
|
||||
_conversions++;
|
||||
_convertMsTotal += conversionMs;
|
||||
if (conversionMs > _convertMsMax) _convertMsMax = conversionMs;
|
||||
|
||||
var now = DateTime.UtcNow;
|
||||
if (now - _lastTelemetry < TelemetryInterval) return;
|
||||
|
||||
var span = now - _telemetryFrom;
|
||||
var perSecond = span.TotalSeconds > 0 ? _conversions / span.TotalSeconds : 0;
|
||||
AppLog.Write(
|
||||
$"ScreenCapture telemetry [{Key}]: {_conversions} frames in {span.TotalSeconds:F1}s " +
|
||||
$"({perSecond:F0}/s), conv avg {( _conversions == 0 ? 0 : _convertMsTotal / _conversions )}ms " +
|
||||
$"max {_convertMsMax}ms, skip busy {_skippedBusy} cadence {_skippedCadence}, " +
|
||||
$"ring allocs {_ring.ConsumeAllocations()}");
|
||||
|
||||
_lastTelemetry = now;
|
||||
_telemetryFrom = now;
|
||||
_conversions = 0;
|
||||
_convertMsTotal = 0;
|
||||
_convertMsMax = 0;
|
||||
_skippedBusy = 0;
|
||||
_skippedCadence = 0;
|
||||
}
|
||||
|
||||
private VideoFrame CopyToVideoFrame(SoftwareBitmap bitmap)
|
||||
@@ -212,47 +255,58 @@ public sealed class ScreenCaptureFrameSource : IScreenCaptureSource
|
||||
var dw = Math.Max(1, (int)(sw * scale));
|
||||
var dh = Math.Max(1, (int)(sh * scale));
|
||||
// DWM delivers an opaque surface (alpha 255); bilinear keeps it 255.
|
||||
var scaled = RentRingBuffer(dw * dh * 4);
|
||||
var scaled = _ring.Rent(dw * dh * 4);
|
||||
return new VideoFrame(dw, dh, DownscaleBgra(data, sw, sh, srcStride, dw, dh, scaled))
|
||||
{ IsOpaque = true, Epoch = ++_epoch };
|
||||
}
|
||||
|
||||
var pixels = RentRingBuffer(count);
|
||||
var pixels = _ring.Rent(count);
|
||||
Marshal.Copy(data, pixels, 0, pixels.Length);
|
||||
return new VideoFrame(sw, sh, pixels) { IsOpaque = true, Epoch = ++_epoch };
|
||||
}
|
||||
|
||||
// Bilinear downscale to the master frame. Reads each source row pair through
|
||||
// Marshal.Copy (no unsafe), writing tightly packed BGRA output.
|
||||
// Integer 8.8 fixed-point bilinear downscale to the master frame (slice 16,
|
||||
// 2026-09-14: the two-stage math from SceneCompositor.Bilinear — the same "shift
|
||||
// only at the end" rule from MyMistakes item on fixed-point blending). The previous
|
||||
// double-per-pixel version ran ~30-45ms per 2.5K→1080p frame on a quiet desktop
|
||||
// and the 1742 take's stall measured ~150ms/frame under load (90%-frozen desktop,
|
||||
// ~6 fresh frames/s); integer math keeps the same bilinear result within ±1 while
|
||||
// dropping the cost to the row-pair Marshal.Copy. Rows are read once per source
|
||||
// row pair through Marshal.Copy (no unsafe), writing tightly packed BGRA output.
|
||||
private byte[] DownscaleBgra(IntPtr src, int sw, int sh, int srcStride, int dw, int dh, byte[] dst)
|
||||
{
|
||||
// Row scratch is per-capture-thread and reused across frames (same churn lesson).
|
||||
var row0 = _row0 != null && _row0.Length >= srcStride ? _row0 : (_row0 = new byte[srcStride]);
|
||||
var row1 = _row1 != null && _row1.Length >= srcStride ? _row1 : (_row1 = new byte[srcStride]);
|
||||
var xs = sw / (double)dw;
|
||||
var ys = sh / (double)dh;
|
||||
|
||||
for (var y = 0; y < dh; y++)
|
||||
{
|
||||
var sy = Math.Min(sh - 1, (int)(y * ys));
|
||||
var sy8 = (int)((long)y * sh * 256 / dh);
|
||||
var sy = sy8 >> 8;
|
||||
var sy1 = Math.Min(sh - 1, sy + 1);
|
||||
var fy = (y * ys) - sy;
|
||||
var fy8 = sy8 & 255;
|
||||
var fyInv = 256 - fy8;
|
||||
Marshal.Copy(IntPtr.Add(src, sy * srcStride), row0, 0, srcStride);
|
||||
Marshal.Copy(IntPtr.Add(src, sy1 * srcStride), row1, 0, srcStride);
|
||||
|
||||
var dRow = y * dw * 4;
|
||||
for (var x = 0; x < dw; x++)
|
||||
{
|
||||
var sx = Math.Min(sw - 1, (int)(x * xs));
|
||||
var sx8 = (int)((long)x * sw * 256 / dw);
|
||||
var sx = Math.Min(sw - 1, sx8 >> 8);
|
||||
var sx1 = Math.Min(sw - 1, sx + 1);
|
||||
var fx = (x * xs) - sx;
|
||||
for (var c = 0; c < 4; c++)
|
||||
var fx8 = sx8 & 255;
|
||||
var fxInv = 256 - fx8;
|
||||
|
||||
var i0 = sx * 4;
|
||||
var i1 = sx1 * 4;
|
||||
var j = dRow + x * 4;
|
||||
for (var c = 0; c < 4; c++, i0++, i1++, j++)
|
||||
{
|
||||
var i0 = sx * 4 + c;
|
||||
var i1 = sx1 * 4 + c;
|
||||
var top = row0[i0] + (row0[i1] - row0[i0]) * fx;
|
||||
var bottom = row1[i0] + (row1[i1] - row1[i0]) * fx;
|
||||
dst[dRow + x * 4 + c] = (byte)(top + (bottom - top) * fy);
|
||||
// Two-stage in one 16.8 scale: top/bot ≤ 65280, ×256 + round ≤ 33.5M — int-safe.
|
||||
var top = row0[i0] * fxInv + row0[i1] * fx8;
|
||||
var bot = row1[i0] * fxInv + row1[i1] * fx8;
|
||||
dst[j] = (byte)((top * fyInv + bot * fy8 + 32768) >> 16);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -324,10 +324,20 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
||||
`WindowsRuntimeMarshal.TryGetDataUnsafe` (the same CsWinRT-safe read the webcam path uses) — the
|
||||
`IMemoryBufferByteAccess` ComImport cast threw `Invalid cast` on **every frame** under CsWinRT, which
|
||||
flooded `startup.log` (~5 MB in a session) and burned CPU, so it is gone. Surfaces larger than the
|
||||
1920×1080 master are downscaled bilinearly to the master (`DownscaleBgra`) before the copy, and
|
||||
per-frame conversion failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected
|
||||
content delivers black frames (OS limitation, documented). Frame pool
|
||||
pauses while the app is minimized — capture keeps running, the pool just stops delivering.
|
||||
1920×1080 master are downscaled to the master (`DownscaleBgra`, integer 8.8 fixed-point bilinear —
|
||||
slice 16: the same two-stage math as `SceneCompositor.Bilinear`; the previous double-per-pixel
|
||||
version was ~30-45ms quiet / ~150ms under 240Hz-HDR load and froze the desktop layer ~90% of a take).
|
||||
Conversions are serialized one-at-a-time (`_framePending`) and spaced by a 10ms floor
|
||||
(`MinConvertInterval`): the monitor delivers at the **240Hz DWM cadence** (~4.2ms), far too fast for
|
||||
the ~60/s the pump can use, so the extra arrivals are dropped (`skip busy`/`skip cadence` telemetry).
|
||||
Hand-out buffers come from a **reuse-distance ring** (`FrameRingBuffer`, depth 8, redLine 4): a
|
||||
buffer is only rewritten ≥4 rents after its last hand-out, else a fresh one is allocated — a frame a
|
||||
consumer still holds (`session.LatestFrame` survives conversions; the dispatcher preview copy lags)
|
||||
can never be read-while-overwritten (the 1742 tear). Rolling telemetry (frames/s, conv avg/max ms,
|
||||
drops, ring allocs) is logged to `startup.log` every 2s while converting. Per-frame conversion
|
||||
failures are logged at most once per 5 s (`ErrorLogThrottle`). DRM-protected content delivers black
|
||||
frames (OS limitation, documented). Frame pool pauses while the app is minimized — capture keeps
|
||||
running, the pool just stops delivering.
|
||||
- **Ownership:** `ScreenCaptureManager` mirrors `CameraManager` — refcounted by target key, one shared
|
||||
`WriteableBitmap` per key, dispatcher-coalesced latest-frame copies, `PreviewBitmapChanged`/`CaptureFailed`
|
||||
events, `ReleaseAllAsync` on re-designation. `ScreenCaptureSourceFactory.Resolve(key)` parses the key into
|
||||
@@ -353,7 +363,9 @@ This replaces the old five-seeder cluster (`Seed{Starting,Brb,Ending,Chat}Backgr
|
||||
app is in the background (worse under a full-screen game on 24H2/26100), GPU readback
|
||||
(`CreateCopyFromSurfaceAsync`) contends with the foreground game, the source's one-in-flight `_framePending`
|
||||
gate drops frames during stalls, and the manager's `DispatcherPriority.Render` copies only run as fast as
|
||||
WPF presents the window. Recorded 2026-08-13; no mitigation attempted yet (deferred by user decision).
|
||||
WPF presents the window. Slice 16 shrank the per-conversion stall (fast downscale + 10ms cadence floor)
|
||||
but the OS-level delivery throttle under focus loss remains. Recorded 2026-08-13; no mitigation attempted
|
||||
yet (deferred by user decision).
|
||||
- **GPU posture:** same as webcam — CPU frames, WPF hardware-presents; D3DImage GPU compositing deferred
|
||||
to the encoder task.
|
||||
- **Preview watermark:** the "Preview" placeholder hides while a background capture renders —
|
||||
@@ -1019,6 +1031,40 @@ Full suite 290/291 passing, the sole failure the pre-existing compositor pixel t
|
||||
suite **294/294 green, 0 warnings**. The +0.6s audio-late clap reading on 1726 was confounded by
|
||||
the 1.1x acceleration — re-measure on device; if a real residual remains it is the audio pipeline.
|
||||
No push (web/A/V work stays local).
|
||||
- **Slice 16 — the desktop capture conversion was the bottleneck: fast downscale +
|
||||
cadence throttle + reuse-distance ring (2026-09-14, device take ty-1742):** slice 15
|
||||
fixed pacing but the 1742 desktop layer was still jerky/frozen with a horizontal
|
||||
tear. Decoded to raw frames and audited: the **desktop band was frozen 21s of 23.35s
|
||||
(90%)** — ~6.1 content updates/s, freeze runs up to 2.28-2.78s, dup-run max 90
|
||||
frames (1.5s). The render stat (`worst render 33-36ms`) was real but MOOT: the
|
||||
capture CONVERSION was the wall. The monitor delivers at the **240Hz DWM cadence**
|
||||
(~4.2ms); with one-in-flight conversions and each 2560×1440→1920×1080
|
||||
`DownscaleBgra` at ~150ms under load, `LatestFrame` updated a handful of times/s —
|
||||
the desktop feed inside a 60fps file read ~90% frozen. (Webcam + audio were their
|
||||
own paths — fine, matching the report.) Three changes, all in
|
||||
`Services/ScreenCaptureFrameSource.cs` (+ `Services/FrameRingBuffer.cs`):
|
||||
(1) `DownscaleBgra` is now **integer 8.8 fixed-point, "shift only at the end"** —
|
||||
the exact two-stage math of `SceneCompositor.Bilinear` (MyMistakes take-4 rule),
|
||||
dropping double-per-pixel to a row-walk of integer ops (the capture downscale had
|
||||
stayed the naive float twin of the 258ms disaster); (2) a **10ms conversion floor**
|
||||
(`MinConvertInterval`) so the 240Hz tail stops queuing ~150ms of serialized
|
||||
conversion per second — the open edge sits just above the ~60/s the pump can use;
|
||||
(3) the ring is a **reuse-distance pool** (`FrameRingBuffer`, depth 8, redLine 4):
|
||||
a buffer is only rewritten ≥4 rents after its last hand-out else a fresh allocation,
|
||||
replacing the blind round-robin — since `ScreenCaptureManager` keeps
|
||||
`session.LatestFrame` across conversions and the dispatcher preview copy lags,
|
||||
"who released the buffer" needs a consumer API that doesn't exist; a reuse
|
||||
DISTANCE needs no cooperation (the 1742 new-top/old-bottom tear is that
|
||||
read-under-write, structural now). Telemetry added: a startup.log line every 2s
|
||||
(`frames/s, conv avg/max ms, skip busy/cadence, ring allocs`) so the device take
|
||||
can be judged numerically. **Good Dog test**
|
||||
`Ring_NoLap_ReusesOnlyAfterRedLineRents` (depth 3/redLine 4 exercises the red-line
|
||||
skip; the 8/4 config settles at 8 buffers and never grows). Full suite **295/295
|
||||
green, 0 warnings**. C4 (compositor Epoch-cached downscale / blit-on-change) was
|
||||
DEFERRED: the slice-15 approved plan assumed a native-res relocation; the
|
||||
measurement recast it — relocating a ~30ms float downscale to the render thread
|
||||
just moves the same cost into the slot budget. Re-measure on device; if render
|
||||
still >16.6ms slots after capture feeds ≤60 real updates/s, add C4. No push.
|
||||
- **Stop ordering matters:** `StopAsync` stops the encoder — since slice 10 it FLUSHES the pending
|
||||
queue (`Channel.TryComplete` → drain writes the leftovers, closes stdin → EOF → ffmpeg finalizes+exits;
|
||||
an accepted frame is never lost) — **before** awaiting the loop. The old reverse-order deadlock was
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
using System;
|
||||
using Xunit;
|
||||
using ytLive.Services;
|
||||
|
||||
namespace ytLive.Tests;
|
||||
|
||||
/// <summary>
|
||||
/// The capture conversion ring (<see cref="FrameRingBuffer"/>) is a pure,
|
||||
/// deterministic piece of <see cref="ScreenCaptureFrameSource"/> — the WGC
|
||||
/// pool/session layer is WinRT and exercised only on Windows at runtime. These tests
|
||||
/// pin the slice-16 no-lap contract: a slot is never rewritten within redLine rents
|
||||
/// of its last hand-out (the 1742 tear — a ring slot rewritten under the consumer's
|
||||
/// read handed the compositor one new-top/old-bottom frame).
|
||||
/// </summary>
|
||||
public class ScreenCaptureFrameSourceTests
|
||||
{
|
||||
[Fact]
|
||||
public void Ring_NoLap_ReusesOnlyAfterRedLineRents()
|
||||
{
|
||||
// Depth 3 / redLine 4: the ring is narrower than its safety distance, so the
|
||||
// red-line skip is actually exercised (the production 8/4 config can never
|
||||
// block a rotation — the ring cycles before any slot comes due).
|
||||
var ring = new FrameRingBuffer(depth: 3, redLine: 4);
|
||||
const int size = 100;
|
||||
|
||||
var a = ring.Rent(size);
|
||||
var b = ring.Rent(size);
|
||||
var c = ring.Rent(size);
|
||||
Assert.NotSame(a, b);
|
||||
Assert.NotSame(b, c);
|
||||
Assert.NotSame(a, c);
|
||||
Assert.Equal(3, ring.ConsumeAllocations());
|
||||
|
||||
// 4th rent arrives while all three slots are inside their red line: the ring
|
||||
// must hand out a fresh buffer instead of rewriting a still-loaned slot.
|
||||
var d = ring.Rent(size);
|
||||
Assert.NotSame(d, a);
|
||||
Assert.NotSame(d, b);
|
||||
Assert.NotSame(d, c);
|
||||
Assert.Equal(1, ring.ConsumeAllocations());
|
||||
|
||||
// 5th rent: slot a (handed at seq 1, revisited at seq 5 = exactly redLine)
|
||||
// is reusable, and reuse is an in-place recycle, not a fresh allocation.
|
||||
Assert.Same(a, ring.Rent(size));
|
||||
Assert.Equal(0, ring.ConsumeAllocations());
|
||||
|
||||
// The production-sized ring (8/4, what ScreenCaptureFrameSource uses) settles
|
||||
// at 8 buffers and recycles them forever — no growth under steady capture.
|
||||
var prod = new FrameRingBuffer(depth: 8, redLine: 4);
|
||||
var first = new byte[8][];
|
||||
for (var i = 0; i < 8; i++) first[i] = prod.Rent(size);
|
||||
for (var i = 8; i < 200; i++)
|
||||
Assert.Same(first[i % 8], prod.Rent(size));
|
||||
Assert.Equal(8, prod.ConsumeAllocations());
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user