Files
LlamaCasty/HANDOFF.md
T
gramps eb4c379b91 fix(pump): slice 7 — take the loop off the UI thread (the 'wait 10ms after render 22ms' contradiction resolved)
Take 10 (59a02a5b, slice 6) finally produced a self-contradicting stat: render
22.4ms + submit 2.5 against a 16.7ms deadline, yet avg wait 10ms — a rebasing
pacer CANNOT sleep after a blown deadline. The wait was queue time: StartAsync
fires from a UI command handler, and async continuations re-capture the current
SynchronizationContext — the 'WPF-free, hermetic' frame pump had been rendering
ON THE DISPATCHER behind the live preview the entire starvation saga. OBS keeps
obs_graphics_thread/video_thread off-UI for exactly this reason (dedicated
threads; see docs.obsproject.com/backend-design 'Libobs Threads').

- FramePump: _pumpTask = Task.Run(() => PumpAsync(...)) — null context inside,
  every continuation stays on the pool.
- Audited, not ignored, what that exposes: StaticPixelCache.Get now locks (pool
  miss-decodes raced UI callers); ChatOverlayLayer.RenderFrame checks its cache
  off-thread but marshals the rare raster MISS to the dispatcher (DrawingVisual
  + RenderTargetBitmap are UI-thread objects) and re-validates there; pump
  events already marshal in the VM.
- GCLatencyMode.SustainedLowLatency for the pump's life (restored in finally).
- Stats gained 'worst render Xms' — bimodal averages hid per-tick spikes.
- Webcam routes through the paste cache (the IsOpaque bypass re-sampled ~156k
  px every tick even between identical device frames).

ONE integration test: Pump_Produces_OffTheStartingContext — an inline-pumping
SynchronizationContext makes the old construction run the resolver on the
starting thread by capture; the loop must never. 70/70 per-class green, clean
build 0 warnings. Docs same commit (ai.md slice 7, TASKS take-11 gate,
MyMistakes #6, HANDOFF). take 11: ~300/300 + honest wait -> saga closed,
Unit B (two-line top bar spec, fully captured) starts.
2026-09-04 12:16:20 -07:00

115 lines
9.7 KiB
Markdown

# HANDOFF — Session State
## Branch / Commit State
**`main`** — Unit A (render starvation, slice 2) committed locally this session; push on the user's word
(his pattern: says "push" explicitly). Slice 1 (deadline pacing + row-blit, take-3 fix) is already on
`origin/main` as `716a77f`.
## What just happened (2026-09-04, Unit A slice 2)
Take 4 verdict: pacing HELD (no stall cliff, sync intact — user confirmed game+webcam in-sync) but render
stayed **58.9ms** (budget 16.7) → file still ~3.4x time-lapse. Root: the 2M-iteration managed row walk +
8.3MB fresh buffer every tick. Shipped:
1. **`VideoFrame.IsOpaque`** producer-contract flag — set ONLY by `ScreenCaptureFrameSource` +
`MediaCaptureFrameSource` (DWM/MF fill alpha 255). Full-canvas aligned blit of an opaque frame =
ONE `Buffer.BlockCopy`; black pre-fill skipped when the backdrop covers.
2. **Integer fixed-point bilinear** in `BlitContent` general path (webcam: round/mirror/scaled) — no
divisions, no `Math.Round`; ±1 of the float reference (tests allow ±2).
3. **Pump scratch pool** — `AcquireScratch`/`ReleaseScratch` (max 4, length-keyed, owned-by-reference so
bake-cache/social-bar/static-art arrays can never be captured). Release strictly AFTER
`SubmitFrameAsync` returns (stdin write copies). `FakeEncoder` now snapshots frames like the real
encoder (holds would race legitimate recycling).
4. **Dead code kill:** the per-tick `fromScene` render in the transition branch fed NOTHING
(`BlendFrame` uses `TransitionService.FromFrame` captured at `Start`) — removed with the
`fromSceneProvider` seam + `MainViewModel.cs` call site. Transition cost halves as a side effect.
Bugs caught by the pixel probes before shipping (see MyMistakes take-4 follow-ups): first `Bilinear`
double-shifted (both stages scaled → solid-255 sampled to ~1 → general path drew nothing); sentinel
0xAB collided with a legitimate `x+y` value; pacing-fake synchronous completion hangs vstest (known,
re-trod).
**Verification:** clean build 0 warnings (both projects); per-class vstest 59/59 (FramePump 11 incl. the
new pooling test, SceneCompositor incl. NEW `Composite_OpaqueFullCover...`, SceneGraph, SocialBar, StretchMath,
Camera/ScreenCapture/MediaVideoSource producers, WebcamOutputKey, SessionTeardown) + SourceNaming
RealApp boot-smoke. Scope-check passed.
## OPEN — next, in order
1. **Take 5 happened (2026-09-04 10:49):** `138/300 frames per 5s, avg render 25.5ms` — blits
fixed but the resolver's `RenderChatBox` full-rasterized the chat box EVERY tick whenever the
message buffer was non-empty (the buffer survives sessions — a signed-out record-only take paid
chat render cost!). **Slice 3 shipped same day:** `ChatOverlayLayer` content-versioned cache —
raster on message/config change, blit the cached frame every tick (OBS text-source pattern).
Tests: `ChatOverlayLayerCacheTests` (the ONE, RealApp) + full regression green (62 across
touched classes), clean build 0 warnings.
2. **Takes 7/8 + slice 5 (2026-09-04):** the stamp (f190587b) settled attribution — chat cache REAL
(`resolve ≈ 0`) but render stayed 26-27ms: the compositor re-rasterized every non-opaque layer
per tick (his Live scene: chat+web+image+cam ≈ 680k samples @ ~38ns). Slice 5: `BlitCachedLayer`
— one raster per (source-array, rect, round/mirror), paste at integer pos with opacity; static
layers now cost row-blends, only content changes resample; drag/opacity changes are paste params,
not cache keys. Tests: `PasteCache_...` + 85/85 across compositor/pump/chat/capture/session
classes; clean build 0 warnings. (Cam bypasses the cache via IsOpaque; revisit if take 9 is
borderline. Follow-ups unchanged: vertical-tier alloc, debounced chat re-render on bursts.)
3. **Take 9/10 + slices 6-7 (2026-09-04):** slice 6 broke the 15.6ms Task.Delay sleep quantum (timeBeginPeriod + bulk-sleep + 2ms spin tail + wait stat); take 10's `wait 10ms after render 22ms` then exposed the FINAL structural bug: the pump loop's await-continuations inherited the UI SynchronizationContext — the "WPF-free" producer had been rendering ON THE DISPATCHER, queued behind the live preview, the whole time (explains every 'zero change' complaint). Slice 7: Task.Run the loop (OBS pattern) + StaticPixelCache lock + chat raster-miss marshalled to dispatcher + SustainedLowLatency GC + `worst render` stat + webcam through paste cache. Test `Pump_Produces_OffTheStartingContext`. 70/70 green, clean build. paste cache moved render 26.5→22.4ms yet
period stayed ~37ms — the gap is Task.Delay's ~15.6ms Windows sleep quantum padding every sub-tick
wait. **This is why takes 7→9 read as "zero change" despite real wins: the sleep floor dominated.**
Fixed (media-app canon, cited in code + MyMistakes #3): timeBeginPeriod(1) for the pump's life
(paired End in finally), sleep the bulk, SPIN the last 2ms across the deadline; stats gained
`avg wait` so render+submit+wait ≈ period (accounting closed — nothing can hide). Webcam now
routes through the paste cache too (bypass re-sampled 156k px even between identical device
frames). 52/52 per-class green, clean build 0 warnings, committed this slice.
4. **Take 11 (user, ~30s record-only):** read the superscript; expect `≈300/300 frames, wait ≈ the true remainder, worst render` now visible. If ~300: playback must be honest 1x — saga CLOSED, Unit B starts. If still ~250-280 with worst-render spikes: raster-miss spikes (web capture ~ every second) — next slice is pre-rasterizing on content change rather than on first-tick-after-change (cache the miss behind a swap-in). If wait is STILL large: the context theory was wrong and I have egg to eat — re-instrument, don't guess.
~22, wait ~0-3` — honest 60fps IF render+submit ≤ ~16.7. If wait is near zero and n/300 sits at
~200, the remaining gap is pure render 22ms → next slice = per-phase compositor timing (the stats
can split blit phases the same way they split resolve; that is the honest path, not a guess).
If take 10 lands at ~300/300 → recording saga CLOSED.
5. **Unit B — the top bar + session logic (user spec 2026-09-04 re-sent twice + decisions settled in Q&A):**
- Two-line top bar. Line 1: center = REC + **LIVE** pills (text renamed from ON-AIR; pills become
mutually-exclusive RADIOS — record-OR-stream ruling), right = avatar + **Login/Logout** button
(no account status light). Line 2: centered primary **Start** (grayed while NO pill armed —
INVERTS the 2026-09-01 "unarmed Start records" rule; fix the map when landing) that becomes the
Stop/End button while active.
- Avatar right-click → **Change Account** (creator: "standard google thing ... on a portrait
right-click"). Login = `SignInCommand` direct (context menu on Start dies; "Choose Record Folder"
lives in gear → App Settings only).
- LIVE pill stays login-gated (`CanToggleOnAir` exists ✓ 3a.viii).
- REC+Start → `Microsoft.Win32.SaveFileDialog` (InitialDirectory = settings folder, default name
`ty-…-0000.mp4`, NATIVE overwrite prompt covers exists/validate, Enter confirms). Cancel →
abort + disarm pill (lit pill with no session is a lie). Up-front naming RETIRES the stop-time
rename modal (assumption stated; user's dialog answer was about Go-Live confirmation).
- LIVE+Start → **Go-Live dialog stays** as preflight: prefilled from Text-drawer `Broadcast.*`,
unfilled fields visibly prompted, explicit confirm → `PrepareAndStartLiveAsync` (user: "going
live is scary — confirmation allows back-out + testing up to go-live").
- Bottom-bar metrics init/maintain: `ResetHealth` + `HealthUpdated` exist — verify on take 6.
- F6 "start/end" hotkey routes through `HandleHotkey` — check it honors the new grayed-Start gate.
- Login button text: "Login" (disconnected, LIVE pill greyed) → "Logout" (connected, avatar
appears left of it). The account status LIGHT is deleted per spec 2a.
- Primary button: grayed "Start" when NO pill armed (inverts 2026-09-01 rule — fix the map in
the landing commit); enabled when either armed; REC path = file dialog flow; LIVE path =
go-live after dialog confirm; becomes the Stop/End face while active (user's "Stop button
never active" complaint gets a hermetic test pinning visibility+CanExecute).
- ONE integration test (hermetic): pills↔button state machine + record-path seam (an
`internal static Func<SaveFileDialog-ish prompt>` override seam mirroring `RegistrarOverride`
— never pop real dialogs in tests).
3. **Follow-ups recorded (don't fix opportunistically):** vertical-tier `BilinearScale` per-frame
alloc; cosmetic `FramePump: encoder stop failed: No process is associated` double-stop race;
1440p capture downscale alloc.
## Landmines
- testhost shares startup.log with the app — filter by time when triaging.
- Stale testhost/exe locks the DLL (MSB3027): `taskkill /F /IM testhost.exe` / `ytLive.exe` first.
- Do NOT run full-suite vstest (WASAPI hang, pre-existing); flow = clean build + per-class + scope-check.
- Pacing-seam fakes MUST await/yield (sync-completed task → pump runs inline on StartAsync → hang).
- `FakeEncoder` snapshots submitted frames — keep any new fake encoder honest about buffer recycling.
- Multi-stage fixed-point: shift only at the end (MyMistakes 2026-09-04).
- Real-`MainWindow` tests: `LayoutPathOverride` + temp DB mandatory; `VolumePushOverride` for volume.
- ffmpeg: month-end pinned build; `Startup.log` "Recording saved:" lines show the real final path.
## @ User note
Recording fix FIRST (his order), UX queue right behind — spec + settled decisions above, don't re-ask.
Keep responses SHORT; one integration test per change; commit every unit; push on his word only.