c45cbc93b7
slice 9 made the DURATION right but content still hiccuped; aggregates (301/300, uniform file PTS) could not see it. Measured root cause: FfmpegEncoder.SubmitFrameAsync BLOCKED on WriteAsync(8.3MB)+FlushAsync when ffmpeg lagged the pipe, and the burst while-loop re-wrote that same stale composite per crossed slot — frozen runs. OBS shape (derivative, wrapped pre-1.0): the encoder queue in libobs/obs-encoder.c — encoder thread never couples back into the video thread; overflow = dropped data, never a frozen producer. https://github.com/obsproject/obs-studio/blob/master/libobs/obs-encoder.c - FfmpegEncoder: SubmitFrameAsync is now an enqueue (ArrayPool copy) into a bounded Channel (cap 120) drained by its own task; drop-newest + count when full; StopAsync flushes the queue then EOF (TryComplete). IFfmpegEncoder.DroppedFrames. - FramePump: ONE fresh composite per iteration (burst loop deleted); worst-submit stat, stall logger (>2x interval names the stage), dropped/stalls in stats. - Burned-in 6-digit dot-matrix frame counter (white box, bottom-right) on every composite — the clock-independent judge replacing the WSL ticker: +1/frame, jumps = counted drops. - ONE new test Backpressure_QueueOverflow_DropsFrames_AndNeverBlocks (slow-sink fake: submit never blocks, drops counted, stop flushes exactly submitted-minus-dropped). - Full suite 290 tests, 289 pass — sole failure the pre-existing compositor pixel test. - Docs same-commit: ai.md slice 10 (+ encoder/stop-note corrections), MyMistakes point 8, HANDOFF. Audio untouched (queued follow-up); web overlay still frozen pending timing closure.
346 lines
23 KiB
Markdown
346 lines
23 KiB
Markdown
# MyMistakes.md
|
||
|
||
> **Two jobs**, distinguished by heading:
|
||
>
|
||
> 1. **Per-task failure log** — updated before every commit touching that task:
|
||
> current iteration + why the last one failed. On task complete, committed AND
|
||
> pushed → truncate to this stub. A new task does NOT seed this file until its
|
||
> first failure.
|
||
> 2. **Recipes registry** (DERIVED-SOLUTION RULE, see `AGENTS.md` 🔬) — the durable
|
||
> home for one-off derived solutions, recipes, and how-tos. The moment you work
|
||
> out a reusable solution, write it here **in the same session**. GREP THIS FILE
|
||
> FIRST when you hit a "I've done this before but have to figure it out again"
|
||
> wall. Recipe entries stay permanently (they are NOT truncated on task
|
||
> completion) — only the failure log truncates.
|
||
|
||
## 🔬 Recipes registry
|
||
|
||
### ⚠ SPIN GUARD TRIGGERED → RESOLVED (? VERIFY) — web overlay transparency + bounding box
|
||
|
||
**THE ONE ROOT CAUSE THAT EXPLAINS EVERY FAILED TAKE:** WebView2's `CapturePreviewAsync`
|
||
produces an **OPAQUE** PNG. From the WebView2 spec (sender: MicrosoftEdge/WebView2Feedback
|
||
`specs/BackgroundColor.md`): "WebView will always honor a webpage's background content."
|
||
`DefaultBackgroundColor = Transparent` only shows through pages with NO background style —
|
||
the widget's own CSS paints html/body opaque. Every take below built on the false premise
|
||
"the capture has transparent margins, alpha=0"; it never did. `FindContentBounds` then had
|
||
no alpha-0 margins to find → wrong crop → black bounding box. The compositor blend saw
|
||
alpha=255 → black over webcam = "transparency broken". Same bug, three symptoms.
|
||
|
||
**THE FIX (the OBS way, applied 2026-09-08):** inject the transparency BEFORE the page
|
||
parses using `CoreWebView2.AddScriptToExecuteOnDocumentCreatedAsync` — documented to run
|
||
"before the HTML document has been parsed and before any other script included by the HTML
|
||
document is run" (learn.microsoft.com/dotnet/api/microsoft.web.webview2.core.corewebview2.addscripttoexecuteondocumentcreatedasync).
|
||
The old `ExecuteScriptAsync` on NavigationStarting/NavigationCompleted ran AFTER page
|
||
scripts/CSS → widget page repainted background opaque → lost the fight. OBS browser sources
|
||
do the same via a pre-parse user.css. Kept the nav handlers as a post-load re-assertion.
|
||
|
||
**Take timeline (the honest record):**
|
||
- `bccdb48` (Aug 28): added FindContentBounds crop — correct idea (tight bbox, no dead
|
||
space), but the capture was OPAQUE so the bbox math was built on nothing.
|
||
- `5348b5c` (Aug 28, 2 min later): reverted to full-frame no-crop — looked "good" for a
|
||
full-bleed widget, but floating widgets regained dead space ("ghost boundary").
|
||
- take-21 (`e002847`) full canvas → shrunken/offset widget (UnifomToFill of whole canvas
|
||
into a small element = lost resizing).
|
||
- take-22 (`ed9d7c1`) crop width used as canvas stride for buffer indexing → garbage.
|
||
- take-23 (`f6802c7`) stride fixed with `src.Width`; PasteKey lacked CropBounds → stale
|
||
raster cache → stale crop.
|
||
- take-24 (`081e4c1`) CropBounds in PasteKey — STILL BROKEN because the SOURCE ALPHA WAS
|
||
NEVER REAL.
|
||
- take-25 (**ME, this session — the user's "RE-INTRODUCING THE BOUNDING-BOX PROBLEM"**):
|
||
I removed FindContentBounds + the crop path entirely, betting full-canvas UniformToFill
|
||
was the answer. It WASN'T — the source is opaque-black, so the element rendered as a
|
||
SOLID BLACK BOX (the user's screenshot: "a black box in the lower right corner"). Killed
|
||
the crop → dead space returned AND black box. The compositor math was ALREADY correct;
|
||
gutting it was vandalism in response to a source-level bug.
|
||
|
||
**Rules, self-inflicted:**
|
||
1. Instrument FIRST. This session added: first-capture PNG dump of the raw WebView2 PNG
|
||
(to `%TEMP%\ytLive-web-<id>.png`) + alpha min/max/mean/%zero + FindContentBounds result
|
||
logged to startup.log once per session. That's the diff between a five-take loop and a
|
||
five-minute diagnosis.
|
||
2. When the same symptom loops across takes, the PREMISE is wrong, not the code — the
|
||
capture being transparent was the load-bearing premise and it was never verified.
|
||
3. Do not delete code paths that fix one axis (crop=bbox) while debugging another
|
||
(source alpha). Revert scope creep; keep layer contributions separable.
|
||
|
||
### Shrink / re-encode an image for the README (screenshots → small hero image)
|
||
|
||
Worked out 2026-08-29 (the recipe was NEVER recorded the first time it was done, so
|
||
it had to be re-derived from scratch — that's the incident this entry exists to end).
|
||
|
||
**Approach:** a throwaway Windows-dotnet console app uses WPF's imaging stack
|
||
(`System.Windows.Media.Imaging`) — same framework the app runs on, zero NuGet
|
||
packages, high-quality downscale via `TransformedBitmap`. Screenshots compress
|
||
**far smaller as JPEG than PNG** (PNG 1400px = ~1.2MB; JPEG q82 1400px = ~188KB).
|
||
|
||
**Recipe (run via the Windows dotnet host from WSL):**
|
||
|
||
1. Create `imgresize.csproj` targeting `net8.0-windows` with `<UseWPF>true</UseWPF>`
|
||
(SDK controller). Put it in a Windows-visible temp path, e.g.
|
||
`C:\Users\gramp\AppData\Local\Temp\imgresize` — NOT `/tmp` (Windows dotnet can't
|
||
reach a Linux-only path reliably).
|
||
2. `Program.cs`: load `BitmapImage` (`CacheOption=OnLoad` → `Freeze()`), downscale
|
||
with `TransformedBitmap(src, new ScaleTransform(scale, scale))` to max width
|
||
(1400 for the README hero), encode with `JpegBitmapEncoder { QualityLevel = 82 }`,
|
||
save.
|
||
3. Run:
|
||
```bash
|
||
"/mnt/c/Program Files/dotnet/dotnet.exe" run -c Release --project .
|
||
-- "C:\Users\gramp\Downloads\Screenshot 2026-08-29 075626.png"
|
||
"C:\Users\gramp\Documents\Code\projects\ytLive\docs\ytLlive-preview.jpg" 1400
|
||
```
|
||
4. Point `README.md` at the `.jpg` (not `.png`).
|
||
|
||
**Result:** 3.2MB screenshot → 1400×794 → **188KB** `ytLlive-preview.jpg` in `docs/`.
|
||
|
||
---
|
||
|
||
### Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
|
||
|
||
Derived 2026-09-03 (take-3 diagnosis — the stats seam from `97ffc42` named the stage
|
||
in one line: `17/300 frames per 5s, avg render 258.1ms, avg submit 1.5ms`).
|
||
Both halves were solved by OBS/libyuv long ago; do not re-derive:
|
||
|
||
1. **Pacing is a DEADLINE, never a post-render sleep.** `sleep(interval)` after each
|
||
frame makes the period `render + submit + interval` — the producer can hit ≤ half
|
||
the declared rate even with a free render. OBS's `video_thread`
|
||
(`libobs/media-io/video-io.c`) advances an absolute `nextTick += intervalTicks` and
|
||
sleeps only the remainder. **NEVER "rebase" the deadline to wall-now when it blows** —
|
||
the take-3 note I originally wrote here ("skip … the missed ticks (rebase)") was the
|
||
PROVEN-wrong advice: resetting `nextTick` erases the missed slots, so a 43fps reality
|
||
was authored into a 60fps container and every recording played ~1.4x fast (take 12/13,
|
||
2026-09-10). The rebase even kept the stats looking honest (render+wait == period, no
|
||
loss) because a rebased frame is never "late". OBS keeps the counter MOVING and outputs
|
||
ONE frame per interval tick — a late render shows as REPEATED footage (judder; duration
|
||
stays correct), never a skip and never a wall-now reset. Critical with rawvideo:
|
||
pts is stamped by ARRIVAL, so the muxed duration is pure frame count ÷ fps — the only
|
||
way to make duration == wall time under ANY load is exactly one frame per interval slot.
|
||
Corollary: **never add a second pacer.** `ffmpeg -re` on the rawvideo input throttles
|
||
the pipe independently and fights the pump (its "Resumed reading … after a lag" grew
|
||
0.79s→4.82s in the same take). Removed; the pump IS the pacer.
|
||
2. **A 1080p frame is ~2.07M pixels — the hot path must be row-simple.** Per-pixel
|
||
`Math.Round` + float source-over in managed code costs ~100ns/px = the whole 258ms.
|
||
libyuv's pattern (https://chromium.googlesource.com/libyuv/libyuv/): branch per
|
||
pixel on source alpha (opaque → 4-byte copy, transparent → skip), integer
|
||
fixed-point blend `(s*a + d*(255-a) + 127)/255` otherwise; and ALWAYS clip the loop
|
||
to the intersection rect (our social-bar overlay scanned all 2M dst px for a 64px
|
||
strip). A full-cover 1:1 blit also obsoletes the opaque-black pre-fill — skip dead
|
||
writes.
|
||
3. **On Windows, `Task.Delay` is a 15.6ms QUANTUM, not a timer.** Any request under one
|
||
system-clock tick sleeps a full tick (documented — learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay:
|
||
"approximately 15 milliseconds on Windows systems"). A deadline pacer built on Task.Delay caps the
|
||
producer at ~40fps-ish EVEN IF render is instant — take 9 proved the signature: work fell 26.5→22.4ms
|
||
but the period sat at ~37ms (≈ one padded wait/frame), so two real optimizations read as "zero change".
|
||
Frame-accurate loops (OBS/Chromium/game-loop canon — stackoverflow.com/questions/5441464) do:
|
||
`timeBeginPeriod(1)` for the session (paired with `timeEndPeriod`), sleep only the BULK of the
|
||
remainder, SPIN the last ~2ms across the deadline. Diagnostic before touching the compositor again:
|
||
period ≈ work + 15.6 → the SLEEP is the bug, not the work.
|
||
|
||
Related (take 14, 2026-09-04): **recycled ring buffers are a race you must SIZE, not just own.**
|
||
Deepening shared frames to kill GC churn (a fresh 8.3MB/tick array) hands out REUSED memory — the
|
||
ring's depth × source period must EXCEED the worst consumer hold (compositor read + lagged UI
|
||
preview copy), not just "a few frames". 4 slots at 144Hz capture laps in ~27ms vs a ≤50ms read: half
|
||
a new screen frame flashed over an old one in the recording ("bits flashing over other bits").
|
||
Depth 8 everywhere (screen/camera/web output rings); the paste-cache Epoch still guards identity.
|
||
|
||
4. **Hermetic pacing test:** inject the delay seam to RECORD the requested TimeSpan and
|
||
genuinely await it (`Task.Delay(d, ct)`) — a fake that returns
|
||
`Task.CompletedTask` synchronously makes the whole pump loop run on `StartAsync`'s
|
||
sync continuation and hang the test run (hit this 2026-09-03; the existing fakes all
|
||
yield for exactly this reason). Assert the REQUESTED wait (< interval with a
|
||
≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.
|
||
5. **Expensive content: raster on change, never on read (take 5, 2026-09-04).** A source
|
||
that updates once a minute (chat text!) must not full-rasterize (`FormattedText` +
|
||
`RenderTargetBitmap` + `CopyPixels` ≈ 15-25ms) every compositor tick. OBS text sources
|
||
re-render on property/message change; the per-tick pass blits the cache. Implement as:
|
||
content version (collection-changed counter) + config key (size/appearance) → cached
|
||
immutable `VideoFrame` returned by identity. Gotcha: buffers that SURVIVE sessions
|
||
(the chat log) silently arm the per-tick cost even in flows that never touch the
|
||
feature (signed-out record-only takes paid chat rendering!).
|
||
|
||
6. **An async loop started from a UI handler runs ON THE UI THREAD until you take it off.**
|
||
`await` continuations re-capture the current `SynchronizationContext` — the frame pump was started
|
||
from a WPF command handler, so the "WPF-free, hermetic" compositor rendered and read capture state
|
||
ON THE DISPATCHER, serialized behind the live preview itself, for the whole starvation saga. The
|
||
`wait` stat caught it only when the numbers became self-contradictory (render 22 + wait 10 > any
|
||
rebasing deadline — a blown deadline cannot sleep). OBS runs `obs_graphics_thread`/`video_thread`
|
||
as dedicated threads for exactly this reason. Pattern: `_task = Task.Run(() => Loop())` (null
|
||
context inside), then audit EVERY object the loop touches for UI affinity (RenderTargetBitmap /
|
||
DrawingVisual / WriteableBitmap: marshal the work or the rare miss; plain locked byte[] lookups:
|
||
fine) and pin it with a context test (`Pump_Produces_OffTheStartingContext`, inline-pumping
|
||
SynchronizationContext that the old code failed by construction). Cost: takes 3–10.
|
||
|
||
7. **Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8).**
|
||
The chat raster fix was REAL but the composer blamed it for the residual slowness it did not
|
||
own; two takes burned before the render/resolve split showed `resolve ≈ 0` and pointed at the
|
||
compositor pasting static layers per tick (`BlitCachedLayer` finished the job OBS-style). Before
|
||
shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the
|
||
NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.
|
||
|
||
8. **An encoder refed from a real-time loop must ENQUEUE, never pipe-write in the loop
|
||
(2026-09-10, take 15/16 — the "1...23...4...56..." smeared ticker).** After slice 9 the
|
||
recording played at the right DURATION but the content still hiccuped — and the aggregates
|
||
(301/300, uniform file PTS, 15.6s wall vs 15.74s file) could NOT see it. The cause finally
|
||
measured in `FfmpegEncoder.SubmitFrameAsync`: the `WriteAsync(8.3MB)+FlushAsync` to ffmpeg's
|
||
stdin BLOCKS whenever the encoder lags the pipe, and the slice-9 burst `while (now>=nextTick)`
|
||
then re-wrote that SAME stale composite for every slot that ticked past — frozen content runs.
|
||
OBS's answering machinery is the encoder queue (`libobs/obs-encoder.c`): the encoder thread
|
||
NEVER couples back into the video thread; overflow = dropped data, NEVER a frozen producer.
|
||
Fixed as: bounded `Channel<byte[]>` (cap 120) + a dedicated drain task owning stdin,
|
||
`SubmitFrameAsync` = copy-to-pool-array + `TryWrite` (drop-newest + count when full), pump
|
||
emits ONE fresh composite per iteration (no burst re-write), stop flushes the queue then EOF.
|
||
**Rule: verify with a clock-independent judge.** The WSL ticker that "proved" slice 9 has its
|
||
own Host-timer jitter under Windows load — so this slice burns a dot-matrix `_outputIndex`
|
||
into the bottom-right of every composite; decoding the recording reads the honest sequence
|
||
(+1/frame, jumps = counted drops) with no external clock involved. Take 17 must read
|
||
+1/frame from that strip. (A whole-frame duplicate scan was tried and is
|
||
UNRELIABLE here: the scene is always animating — session elapsed timer + REC pulse — so
|
||
no two frames are ever byte-identical.)
|
||
|
||
|
||
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
|
||
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
|
||
integer bilinear) followed the same libyuv research
|
||
(https://chromium.googlesource.com/libyuv/libyuv/ — `row.cc`/`scale.cc` keep both
|
||
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
|
||
My first `Bilinear` shifted stage 1 back to 8-bit AND shifted the final result >>16 —
|
||
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
|
||
NOTHING (green webcam silently vanished from output; the pixel probes caught what
|
||
the eye in a 2x time-lapse would not). **Rule: multi-stage fixed point shifts only
|
||
at the end; verify against a uniform-255 sample before believing it.** Second trap:
|
||
a stale-byte sentinel test whose source pattern can generate the sentinel value
|
||
itself (0xAB was a legitimate `x+y` pixel) — pick the sentinel coprime/out-of-range
|
||
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
|
||
that HOLDS submitted frames now must snapshot them (`Clone`) once the producer
|
||
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
|
||
|
||
|
||
---
|
||
|
||
## Splitting a large file into partials — NEVER `awk … > SRC` while awking SRC
|
||
|
||
(2026-08-31, Commit D) Tried to split `SocialsDialogViewModel.cs` in one line:
|
||
`{ awk '…' SRC; echo ""; awk '…2…' SRC; } > SRC`. The **first write truncated
|
||
SRC to 1 line**, so the second `awk` read the already-truncated file → the whole
|
||
source was lost (1 line left). Recovered with `git checkout -- SRC`, then redid
|
||
it, but the same bug could have meant making it up from scratch.
|
||
|
||
**Rule:** when a cut needs N blocks from one source into N files, never write a
|
||
block back onto the source that the `awk`s still read. Instead:
|
||
1. Read the source **once** at the start into temp files (`mktemp -d`, one file
|
||
per block), with a `$D` variable you carry forward.
|
||
2. Verify block sizes (`wc -l`) and brace balance (`python3 -c` counting `{`/`}`)
|
||
before touching any real file.
|
||
3. Then assemble each new file from `cat D/block …` — never truncating the source
|
||
until every read is done.
|
||
|
||
`git status` can't save you here if you don't notice until the file is gone —
|
||
`git checkout -- <path>` from the last commit is the recovery. Cheap insurance:
|
||
restore-then-retry, do it atomically from temp files the first time.
|
||
|
||
---
|
||
|
||
## Verifying an ffmpeg decode contract from WSL (no real CLR needed)
|
||
|
||
(2026-08-31, TASK 21) When a change depends on ffmpeg producing output with an
|
||
exact frame-size contract (rawvideo W×H×4 BGRA), you can prove the **command +
|
||
frame accounting** here without any .NET process:
|
||
|
||
1. Fetch a **static Linux ffmpeg** into `/tmp/opencode` (no sudo needed):
|
||
`curl -sLO https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz
|
||
&& tar -xf …`
|
||
2. Generate a tiny known clip: `ffmpeg -f lavfi -i "testsrc2=duration=1:size=640x360:rate=30" -pix_fmt yuv420p clip.mp4`
|
||
3. Decode with **exactly the app's args**: `-f rawvideo -pix_fmt bgra -vf scale=640:360 -an`
|
||
4. Assert `total_bytes % (W*H*4) == 0` (python3) → exact integer frames, no pad.
|
||
|
||
**Why not a dotnet-spawned ffmpeg here:** the only CLR on this box
|
||
(`/home/gramps/bin/dotnet`) is a **Windows-bound shim** — `Process.Start` resolves
|
||
paths to `\\wsl.localhost\Debian\…` and throws "not a valid application for this
|
||
OS" when handed a Linux ELF ffmpeg. So never plan to have dotnet exec a Linux
|
||
ffmpeg here; verify the contract with shell/python instead, and leave the
|
||
CLR→real-ffmpeg run to the native Windows suite.
|
||
|
||
## WPF hit-test truth in tests: `UIElement.InputHitTest`, NOT `VisualTreeHelper.HitTest`
|
||
|
||
(2026-09-01, the RoundClip "known failure" post-mortem — a failure the map carried as
|
||
"not a regression" for weeks without ever recording WHY.)
|
||
|
||
**The trap:** `VisualTreeHelper.HitTest(window, pt)` returned the window's
|
||
`WebViewHostPanel` overlay (`IsHitTestVisible="False"`, `Opacity=0`, ZERO children) for
|
||
EVERY point in the window — so a "corner is grabbable" assertion could never pass, and
|
||
it looked like a real interaction bug. The actual input pipeline (`UIElement.InputHitTest`,
|
||
what Mouse routing uses) correctly returned the element's Grid at elem-center/corner-in/
|
||
corner-exact and fell through to CanvasGrid just past the corner. The product was fine;
|
||
the TEST was probing an API that doesn't model input semantics.
|
||
|
||
**Rule:** any test asserting "where does a click land" uses `window.InputHitTest(pt)` +
|
||
`IsDescendantOf` — never `VisualTreeHelper.HitTest`.
|
||
|
||
**Diagnosis recipe (how the lie was caught in ~3 probe cycles, no guessing):** add a TEMP
|
||
probe `[Fact]` in the RealApp collection that hit-tests a spread of points
|
||
(elem-center / corner-in / corner-exact / corner-out / bg-center) and `Assert.Fail`s with a
|
||
composed dump: per-point VTH hit + `InputHitTest` hit + ancestor chain (`GetParent` walk
|
||
with `#Name`) + panel properties (`IsHitTestVisible/Opacity/children/actual size`) +
|
||
`TranslatePoint` origins. Run the class alone, read the message, delete the probe.
|
||
|
||
**Two sibling facts learned the same session (record-once):**
|
||
1. A UserControl owns its own XAML namescope — after extracting a region out of a window,
|
||
`window.FindName("InnerPart")` returns null; resolve the UserControl by its window-level
|
||
name, then `pane.FindName("InnerPart")`. And window-scope STYLES are invisible to a
|
||
UserControl's `StaticResource` at parse time — move such styles to `Themes/Controls.xaml`
|
||
(the app-scope rule exists for this).
|
||
2. Per-class `dotnet.exe vstest` from WSL DOES execute the RealApp/`MainWindow` tests fine
|
||
(they passed natively 2026-09-01) — only the FULL suite hangs (WASAPI startup). And the
|
||
test process shares `%APPDATA%\ytLlive\startup.log` with the real app: lines like
|
||
`camera 'test-camera' failed` are test noise, not DB state — to check pollution, query
|
||
the DB directly (`python3 sqlite3`, `SELECT DeviceId FROM Webcam`), not the log.
|
||
|
||
## A "known failure" label without a recorded cause = a bug on life support
|
||
|
||
(2026-09-01, the audio triple-take) One line — `_delayedMix` (nullable, added by TASK 22,
|
||
never initialized) dereferenced as `delayed.Length` — produced THREE symptoms that lived in
|
||
the map as two separate "pre-existing, do-not-chase" entries: (a)
|
||
`Mix_HonorsProviderGains…` "known failure", (b) `AudioPipelineTests` hangs when run at all
|
||
(a test reading a named pipe with no writer blocks — hung test ≠ flaky test, it's a starved
|
||
producer), (c) startup.log flooded "Audio live loop error" every 10ms (the mixer loop caught
|
||
and logged only `ex.Message` — stack thrown away).
|
||
|
||
Rules derived:
|
||
1. NEVER label a test "known/pre-existing" without writing WHY (exception type + first app
|
||
frame). An unexplained known-failure is deferred archaeology that hardens into fog.
|
||
2. A test that waits on IPC + a producer whose output vanished are usually ONE bug — look
|
||
for the producer before blaming the test.
|
||
3. Catch-and-log-swallow of `ex.Message` hides root causes; log with stack (`AppLog.Write(ex,
|
||
...)`) and throttle (5s) instead of dropping or flooding.
|
||
Bonus: the "failing" test encoded the MAP's contract (`loopbackGain = GameAudioVolume`); the
|
||
code had drifted to unity on a disproven premise (loopback capture does NOT follow endpoint
|
||
volume — creator's 20%-volume/pegged-meter observation killed it). The test was right all
|
||
along — failing tests may be the last honest witnesses; interrogate, don't pardon.
|
||
|
||
## Feature provenance: record WHO asked and WHY, in the task entry itself
|
||
|
||
(2026-08-31/09-01, the SYNC slider scare) TASK 22's lip-sync slider surfaced on the preview rail and
|
||
the creator's reaction was "totally don't remember ordering that" — because the queue entry recorded
|
||
WHAT shipped (a slider, 0-500ms, a converter class) but not WHO asked (the creator, explicitly, for
|
||
OBS's delay-filter fix built natively). Eight days later his own request read like AI drift and nearly
|
||
got deleted. Rule: the moment a creator-driven feature is queued or shipped, its entry carries a
|
||
one-line provenance — *who asked, what triggered it* ("creator: OBS delay-filter lip-sync fix, native").
|
||
Features without attribution become roadmap orphans that get punted, removed, or re-litigated. Same
|
||
disease as an unexplained "known failure" label — a fact recorded without its reason is a future
|
||
argument.
|
||
|
||
## An un-attributed build invalidated three takes of a perf saga — stamp the binary
|
||
|
||
(2026-09-04, takes 4–6) After each render-perf fix the creator "exed the code" and re-recorded, but
|
||
the exe timestamp ≠ binary contents (incremental builds reuse whatever compiles clean; a source edit
|
||
with no rebuild serves the OLD exe). Take 6 measured render WORSE than take 5 (35-41ms) and there was
|
||
no honest way to tell "the chat cache fix doesn't work" from "the fix was never running" — three
|
||
hours of diagnosis on an unattributable sample. Rule: if takes measure the app, EVERY build carries
|
||
an id and EVERY log line traces to it — `GenerateBuildStamp` (csproj) writes a fresh GUID per
|
||
compile (deliberately defeating incremental lies), the wordmark shows it as a superscript, startup.log
|
||
records `Build <id> (compiled <time>)`. A perf claim without build attribution is a guess; ask for the
|
||
stamp BEFORE theorizing. (Also this session: a sentinel-byte test where the source pattern could
|
||
GENERATE the sentinel, and a fixed-point bilinear that shifted BOTH stages and silently drew nothing —
|
||
see the rawvideo recipe.)
|
||
|