perf(render): take-4 slice 2 — IsOpaque memcpy, integer bilinear, pump scratch pool
Take 4: pacing held (sync perfect) but avg render stayed 58.9ms — 2M managed row-walk iterations + a fresh 8.3MB buffer every tick (LOH churn into GC stalls inside the render measurement). - VideoFrame.IsOpaque: producer-contract flag (screen capture + webcam — DWM/MF fill alpha 255; media/chat/web/static NOT flagged). Full-cover aligned opaque backdrop = ONE Buffer.BlockCopy; black pre-fill skipped when it covers. - General BlitContent: integer 8.8 fixed-point bilinear + blend, row invariants hoisted, no per-pixel division/Math.Round. Within ±1 of the float reference (pixel tests allow ±2). Research per derivative-work rule: libyuv row/scale kernels (chromium.googlesource.com/libyuv/libyuv). - FramePump scratch pool (max 4, length-keyed, owned-by-reference): release strictly AFTER SubmitFrameAsync returns (stdin write copies); Contains-guard makes the transition Cut alias safe. - Removed the dead per-tick fromScene render + fromSceneProvider seam — BlendFrame consumes TransitionService.FromFrame captured at Start; the pump's render fed nothing. MainViewModel call site updated (signature). Bugs caught by the pixel probes pre-ship (recorded MyMistakes): first Bilinear double-shifted both stages (solid-255 sampled to ~1 -> general path drew nothing); sentinel 0xAB collided with an x+y pixel. FakeEncoder snapshots submitted frames (mirrors real copy semantics under recycling). ONE integration test: Pump_Pools_ScratchBuffers_Across_Frames_Without_ Stale_Pixels (alternating backdrops + repeated backing identity). Direct pin: Composite_OpaqueFullCover_Backdrop_CopiesEveryPixel_Into_Scratch. Clean build 0 warnings; 59/59 per-class + RealApp boot-smoke. take 5 verdict: expect avg render <= ~10ms, ~300/300 frames. Docs same commit: ai.md pipeline section, TASKS.md TASK 18, HANDOFF rewritten (Unit B spec + settled decisions queued).
This commit is contained in:
@@ -75,6 +75,23 @@ Both halves were solved by OBS/libyuv long ago; do not re-derive:
|
||||
yield for exactly this reason). Assert the REQUESTED wait (< interval with a
|
||||
≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.
|
||||
|
||||
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
|
||||
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
|
||||
integer bilinear) followed the same libyuv research
|
||||
(https://chromium.googlesource.com/libyuv/libyuv/ — `row.cc`/`scale.cc` keep both
|
||||
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
|
||||
My first `Bilinear` shifted stage 1 back to 8-bit AND shifted the final result >>16 —
|
||||
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
|
||||
NOTHING (green webcam silently vanished from output; the pixel probes caught what
|
||||
the eye in a 2x time-lapse would not). **Rule: multi-stage fixed point shifts only
|
||||
at the end; verify against a uniform-255 sample before believing it.** Second trap:
|
||||
a stale-byte sentinel test whose source pattern can generate the sentinel value
|
||||
itself (0xAB was a legitimate `x+y` pixel) — pick the sentinel coprime/out-of-range
|
||||
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
|
||||
that HOLDS submitted frames now must snapshot them (`Clone`) once the producer
|
||||
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Splitting a large file into partials — NEVER `awk … > SRC` while awking SRC
|
||||
|
||||
Reference in New Issue
Block a user