perf(compositor): slice 5 — paste cache for non-opaque layers (take-7/8 data)

The stamped build settled what slices 3-4 could not: chat cache works (resolve
~0.0ms) but render stayed 26-27ms -> 124-135/300. The cost was the compositor
re-rasterizing EVERY layer every tick: this Live scene re-samples chat (159k) +
web widget (271k) + image (95k) + cam (156k) ~ 680k px @ ~38ns — for layers
whose pixels do not change between chat/web/cam updates.

OBS shape: cache the surface, paste per tick. BlitCachedLayer rasterizes a
non-opaque layer ONCE into an element-space, transparent-based frame keyed by
(source-array identity, src W/H, ceil'd dst rect, round, mirror), then pastes:
integer position, row alpha-blend, opacity applied at paste. Producers hand out
fresh immutable arrays -> array-identity keys cannot serve stale content; dict
bounded (48, clears whole). Drag/opacity live in paste params, not keys, so
editing stops triggering resamples too. Opaque backdrop keeps the memcpy path;
the webcam keeps the direct path via its IsOpaque flag (revisit if take 9 is
borderline).

ONE integration test: PasteCache_RepeatRender_IsByteIdentical_And_ContentChange-
Propagates (byte-exact raster-vs-paste incl. round-clip margins, new-array
propagation); existing pixel suite guards sampler semantics. 85/85 across
compositor/pump/chat/capture/session classes, clean build 0 warnings. Also:
BuildStampTests.cs was written last commit but never staged — its own scope-check
slip, added here (the run had used the on-disk file; tracked now).

Docs same commit: ai.md slice 5 + stale 'general path only 130k' claim corrected,
TASKS.md take-9 gate, MyMistakes recipe (prove the stage; a fix that doesn't move
the stat wasn't the bottleneck), HANDOFF. take 9 expectation: 300/300, render
<= ~8ms -> saga closes, Unit B starts.
This commit is contained in:
2026-09-04 11:44:19 -07:00
parent 27bf74389d
commit 6af2026906
7 changed files with 258 additions and 16 deletions
+8
View File
@@ -83,6 +83,14 @@ Both halves were solved by OBS/libyuv long ago; do not re-derive:
(the chat log) silently arm the per-tick cost even in flows that never touch the
feature (signed-out record-only takes paid chat rendering!).
5. **Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8).**
The chat raster fix was REAL but the composer blamed it for the residual slowness it did not
own; two takes burned before the render/resolve split showed `resolve ≈ 0` and pointed at the
compositor pasting static layers per tick (`BlitCachedLayer` finished the job OBS-style). Before
shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the
NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
integer bilinear) followed the same libyuv research