perf(render): take-4 slice 2 — IsOpaque memcpy, integer bilinear, pump scratch pool

Take 4: pacing held (sync perfect) but avg render stayed 58.9ms — 2M
managed row-walk iterations + a fresh 8.3MB buffer every tick (LOH churn
into GC stalls inside the render measurement).

- VideoFrame.IsOpaque: producer-contract flag (screen capture + webcam —
  DWM/MF fill alpha 255; media/chat/web/static NOT flagged). Full-cover
  aligned opaque backdrop = ONE Buffer.BlockCopy; black pre-fill skipped
  when it covers.
- General BlitContent: integer 8.8 fixed-point bilinear + blend, row
  invariants hoisted, no per-pixel division/Math.Round. Within ±1 of the
  float reference (pixel tests allow ±2). Research per derivative-work
  rule: libyuv row/scale kernels (chromium.googlesource.com/libyuv/libyuv).
- FramePump scratch pool (max 4, length-keyed, owned-by-reference):
  release strictly AFTER SubmitFrameAsync returns (stdin write copies);
  Contains-guard makes the transition Cut alias safe.
- Removed the dead per-tick fromScene render + fromSceneProvider seam —
  BlendFrame consumes TransitionService.FromFrame captured at Start; the
  pump's render fed nothing. MainViewModel call site updated (signature).

Bugs caught by the pixel probes pre-ship (recorded MyMistakes): first
Bilinear double-shifted both stages (solid-255 sampled to ~1 -> general
path drew nothing); sentinel 0xAB collided with an x+y pixel. FakeEncoder
snapshots submitted frames (mirrors real copy semantics under recycling).

ONE integration test: Pump_Pools_ScratchBuffers_Across_Frames_Without_
Stale_Pixels (alternating backdrops + repeated backing identity). Direct
pin: Composite_OpaqueFullCover_Backdrop_CopiesEveryPixel_Into_Scratch.
Clean build 0 warnings; 59/59 per-class + RealApp boot-smoke. take 5
verdict: expect avg render <= ~10ms, ~300/300 frames. Docs same commit:
ai.md pipeline section, TASKS.md TASK 18, HANDOFF rewritten (Unit B spec
+ settled decisions queued).
This commit is contained in:
2026-09-04 09:35:59 -07:00
parent 716a77f61a
commit 432adfdaef
12 changed files with 370 additions and 102 deletions
+17
View File
@@ -75,6 +75,23 @@ Both halves were solved by OBS/libyuv long ago; do not re-derive:
yield for exactly this reason). Assert the REQUESTED wait (< interval with a
≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
integer bilinear) followed the same libyuv research
(https://chromium.googlesource.com/libyuv/libyuv/ — `row.cc`/`scale.cc` keep both
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
My first `Bilinear` shifted stage 1 back to 8-bit AND shifted the final result >>16 —
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
NOTHING (green webcam silently vanished from output; the pixel probes caught what
the eye in a 2x time-lapse would not). **Rule: multi-stage fixed point shifts only
at the end; verify against a uniform-255 sample before believing it.** Second trap:
a stale-byte sentinel test whose source pattern can generate the sentinel value
itself (0xAB was a legitimate `x+y` pixel) — pick the sentinel coprime/out-of-range
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
that HOLDS submitted frames now must snapshot them (`Clone`) once the producer
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
---
## Splitting a large file into partials — NEVER `awk … > SRC` while awking SRC