0f441d825a
take-19-2/take-20 'webcam black square under the web widget': the capture was alpha-cropped (FindContentBounds) and that CROP was fed to the compositor, whose UniformToFill zoomed the opaque content box to cover the whole element rect the moment the widget drew any content (idle transparent page = full-frame crop = the correct viewport, hence take-1-good/take-2-bad). OBS model verified: the page is a fixed 1920x1080 canvas and the element rect is a viewport onto it — measure with the alpha crop (preview/selection), hand the composite the FULL canvas on an 8-deep ring + Epoch (same identity discipline as camera/screen). Regression test: Composite_FullCanvasWebSource_TransparentMarginsRevealWebcam (the reveal contract: margin pixel = webcam, badge pixel = widget).
301 lines
19 KiB
Markdown
301 lines
19 KiB
Markdown
# MyMistakes.md
|
||
|
||
> **Two jobs**, distinguished by heading:
|
||
>
|
||
> 1. **Per-task failure log** — updated before every commit touching that task:
|
||
> current iteration + why the last one failed. On task complete, committed AND
|
||
> pushed → truncate to this stub. A new task does NOT seed this file until its
|
||
> first failure.
|
||
> 2. **Recipes registry** (DERIVED-SOLUTION RULE, see `AGENTS.md` 🔬) — the durable
|
||
> home for one-off derived solutions, recipes, and how-tos. The moment you work
|
||
> out a reusable solution, write it here **in the same session**. GREP THIS FILE
|
||
> FIRST when you hit a "I've done this before but have to figure it out again"
|
||
> wall. Recipe entries stay permanently (they are NOT truncated on task
|
||
> completion) — only the failure log truncates.
|
||
|
||
## 🔬 Recipes registry
|
||
|
||
### Web-overlay transparency: measure with the alpha crop, COMPOSE with the full canvas
|
||
|
||
(2026-09-05, take-19-2/take-20 — "the webcam is a black square under the web widget".
|
||
|
||
**Symptom:** WebView2-based web overlays (StreamElements alert hosts etc.) rendered
|
||
with hard opaque boxes over the webcam in the RECORDING but the preview looked fine,
|
||
and the FIRST take after a rebuild was good but later takes went opaque — a
|
||
stateful-looking bug.
|
||
|
||
**Root cause:** the capture was alpha-CROPPED to the widget's content bounding box
|
||
(`FindContentBounds` — every non-zero-alpha pixel) and that CROP was handed to the
|
||
compositor, which `UniformToFill`-scales a source into the element rect. A crop has
|
||
NO transparent margins by construction, so the composite zoomed the opaque content
|
||
to cover the ENTIRE element rect (hard box!) whenever the widget drew anything. It
|
||
looked stateful because an idle/transparent page yields a full-frame crop ≈ the
|
||
correct full-canvas viewport (good), while any real content yields a tight crop +
|
||
zoom (bad).
|
||
|
||
**The rule (OBS browser-source model):** the page is a FIXED canvas (1920×1080 here);
|
||
the element rect is a VIEWPORT onto it. Measure/select with the alpha crop; hand the
|
||
composite the FULL canvas with its transparent background intact — `UniformToFill`
|
||
then maps the page region 1:1 into the element rect and the page's transparent
|
||
margins reveal the layers beneath (the webcam). When feeding a compositor that
|
||
scales sources to fill a rect, NEVER feed it a crop of your source unless the rect
|
||
is meant to re-frame it.
|
||
|
||
Recipe:
|
||
1. Keep `FindContentBounds` for the preview bitmap / selection bounding box.
|
||
2. Feed the compositor a full-canvas frame whose background is alpha 0, from an
|
||
8-deep ring (same identity/Epoch discipline as every producer — a single reused
|
||
scratch array tears under the consumer).
|
||
3. Pin the reveal contract with a compositor test: full-canvas transparent-margin
|
||
frame over a webcam rect → margin pixel = webcam color, opaque badge pixel =
|
||
widget color.
|
||
|
||
Worked out 2026-08-29 (the recipe was NEVER recorded the first time it was done, so
|
||
it had to be re-derived from scratch — that's the incident this entry exists to end).
|
||
|
||
**Approach:** a throwaway Windows-dotnet console app uses WPF's imaging stack
|
||
(`System.Windows.Media.Imaging`) — same framework the app runs on, zero NuGet
|
||
packages, high-quality downscale via `TransformedBitmap`. Screenshots compress
|
||
**far smaller as JPEG than PNG** (PNG 1400px = ~1.2MB; JPEG q82 1400px = ~188KB).
|
||
|
||
**Recipe (run via the Windows dotnet host from WSL):**
|
||
|
||
1. Create `imgresize.csproj` targeting `net8.0-windows` with `<UseWPF>true</UseWPF>`
|
||
(SDK controller). Put it in a Windows-visible temp path, e.g.
|
||
`C:\Users\gramp\AppData\Local\Temp\imgresize` — NOT `/tmp` (Windows dotnet can't
|
||
reach a Linux-only path reliably).
|
||
2. `Program.cs`: load `BitmapImage` (`CacheOption=OnLoad` → `Freeze()`), downscale
|
||
with `TransformedBitmap(src, new ScaleTransform(scale, scale))` to max width
|
||
(1400 for the README hero), encode with `JpegBitmapEncoder { QualityLevel = 82 }`,
|
||
save.
|
||
3. Run:
|
||
```bash
|
||
"/mnt/c/Program Files/dotnet/dotnet.exe" run -c Release --project .
|
||
-- "C:\Users\gramp\Downloads\Screenshot 2026-08-29 075626.png"
|
||
"C:\Users\gramp\Documents\Code\projects\ytLive\docs\ytLlive-preview.jpg" 1400
|
||
```
|
||
4. Point `README.md` at the `.jpg` (not `.png`).
|
||
|
||
**Result:** 3.2MB screenshot → 1400×794 → **188KB** `ytLlive-preview.jpg` in `docs/`.
|
||
|
||
---
|
||
|
||
### Feeding a rawvideo pipe at 60fps: deadline pacing + row-blit budget
|
||
|
||
Derived 2026-09-03 (take-3 diagnosis — the stats seam from `97ffc42` named the stage
|
||
in one line: `17/300 frames per 5s, avg render 258.1ms, avg submit 1.5ms`).
|
||
Both halves were solved by OBS/libyuv long ago; do not re-derive:
|
||
|
||
1. **Pacing is a DEADLINE, never a post-render sleep.** `sleep(interval)` after each
|
||
frame makes the period `render + submit + interval` — the producer can hit ≤ half
|
||
the declared rate even with a free render. OBS's `video_thread`
|
||
(`libobs/media-io/video-io.c`) advances an absolute `nextTick += intervalTicks` and
|
||
sleeps only the remainder; if the deadline blew, skip the wait AND the missed ticks
|
||
(rebase, no catch-up burst — a burst queues stale frames). Critical with rawvideo:
|
||
pts is stamped by ARRIVAL, so a starved producer silently time-lapses the file.
|
||
2. **A 1080p frame is ~2.07M pixels — the hot path must be row-simple.** Per-pixel
|
||
`Math.Round` + float source-over in managed code costs ~100ns/px = the whole 258ms.
|
||
libyuv's pattern (https://chromium.googlesource.com/libyuv/libyuv/): branch per
|
||
pixel on source alpha (opaque → 4-byte copy, transparent → skip), integer
|
||
fixed-point blend `(s*a + d*(255-a) + 127)/255` otherwise; and ALWAYS clip the loop
|
||
to the intersection rect (our social-bar overlay scanned all 2M dst px for a 64px
|
||
strip). A full-cover 1:1 blit also obsoletes the opaque-black pre-fill — skip dead
|
||
writes.
|
||
3. **On Windows, `Task.Delay` is a 15.6ms QUANTUM, not a timer.** Any request under one
|
||
system-clock tick sleeps a full tick (documented — learn.microsoft.com/en-us/dotnet/api/system.threading.tasks.task.delay:
|
||
"approximately 15 milliseconds on Windows systems"). A deadline pacer built on Task.Delay caps the
|
||
producer at ~40fps-ish EVEN IF render is instant — take 9 proved the signature: work fell 26.5→22.4ms
|
||
but the period sat at ~37ms (≈ one padded wait/frame), so two real optimizations read as "zero change".
|
||
Frame-accurate loops (OBS/Chromium/game-loop canon — stackoverflow.com/questions/5441464) do:
|
||
`timeBeginPeriod(1)` for the session (paired with `timeEndPeriod`), sleep only the BULK of the
|
||
remainder, SPIN the last ~2ms across the deadline. Diagnostic before touching the compositor again:
|
||
period ≈ work + 15.6 → the SLEEP is the bug, not the work.
|
||
|
||
Related (take 14, 2026-09-04): **recycled ring buffers are a race you must SIZE, not just own.**
|
||
Deepening shared frames to kill GC churn (a fresh 8.3MB/tick array) hands out REUSED memory — the
|
||
ring's depth × source period must EXCEED the worst consumer hold (compositor read + lagged UI
|
||
preview copy), not just "a few frames". 4 slots at 144Hz capture laps in ~27ms vs a ≤50ms read: half
|
||
a new screen frame flashed over an old one in the recording ("bits flashing over other bits").
|
||
Depth 8 everywhere (screen/camera/web output rings); the paste-cache Epoch still guards identity.
|
||
|
||
4. **Hermetic pacing test:** inject the delay seam to RECORD the requested TimeSpan and
|
||
genuinely await it (`Task.Delay(d, ct)`) — a fake that returns
|
||
`Task.CompletedTask` synchronously makes the whole pump loop run on `StartAsync`'s
|
||
sync continuation and hang the test run (hit this 2026-09-03; the existing fakes all
|
||
yield for exactly this reason). Assert the REQUESTED wait (< interval with a
|
||
≥cost-ms fake render) — never wall-clock rate, which flakes on loaded machines.
|
||
5. **Expensive content: raster on change, never on read (take 5, 2026-09-04).** A source
|
||
that updates once a minute (chat text!) must not full-rasterize (`FormattedText` +
|
||
`RenderTargetBitmap` + `CopyPixels` ≈ 15-25ms) every compositor tick. OBS text sources
|
||
re-render on property/message change; the per-tick pass blits the cache. Implement as:
|
||
content version (collection-changed counter) + config key (size/appearance) → cached
|
||
immutable `VideoFrame` returned by identity. Gotcha: buffers that SURVIVE sessions
|
||
(the chat log) silently arm the per-tick cost even in flows that never touch the
|
||
feature (signed-out record-only takes paid chat rendering!).
|
||
|
||
6. **An async loop started from a UI handler runs ON THE UI THREAD until you take it off.**
|
||
`await` continuations re-capture the current `SynchronizationContext` — the frame pump was started
|
||
from a WPF command handler, so the "WPF-free, hermetic" compositor rendered and read capture state
|
||
ON THE DISPATCHER, serialized behind the live preview itself, for the whole starvation saga. The
|
||
`wait` stat caught it only when the numbers became self-contradictory (render 22 + wait 10 > any
|
||
rebasing deadline — a blown deadline cannot sleep). OBS runs `obs_graphics_thread`/`video_thread`
|
||
as dedicated threads for exactly this reason. Pattern: `_task = Task.Run(() => Loop())` (null
|
||
context inside), then audit EVERY object the loop touches for UI affinity (RenderTargetBitmap /
|
||
DrawingVisual / WriteableBitmap: marshal the work or the rare miss; plain locked byte[] lookups:
|
||
fine) and pin it with a context test (`Pump_Produces_OffTheStartingContext`, inline-pumping
|
||
SynchronizationContext that the old code failed by construction). Cost: takes 3–10.
|
||
|
||
7. **Prove the stage, then the fix — and re-prove after every slice (2026-09-04, takes 6-8).**
|
||
The chat raster fix was REAL but the composer blamed it for the residual slowness it did not
|
||
own; two takes burned before the render/resolve split showed `resolve ≈ 0` and pointed at the
|
||
compositor pasting static layers per tick (`BlitCachedLayer` finished the job OBS-style). Before
|
||
shipping a perf fix: name the stage with a measurement, not a story; after shipping one, the
|
||
NEXT number must move — a fix that doesn't change the stat wasn't the bottleneck.
|
||
|
||
|
||
**Take-4 follow-ups (2026-09-04) — the symptom needed a second pass, so cite again:**
|
||
render was still 58.9ms after slice 1. Slice 2 (buffer pool + opaque-row memcpy +
|
||
integer bilinear) followed the same libyuv research
|
||
(https://chromium.googlesource.com/libyuv/libyuv/ — `row.cc`/`scale.cc` keep both
|
||
interpolation stages in ONE fixed-point scale; rounding constant only at the end).
|
||
My first `Bilinear` shifted stage 1 back to 8-bit AND shifted the final result >>16 —
|
||
double scaling turned solid-255 samples into ~1, i.e. the "fixed" general path drew
|
||
NOTHING (green webcam silently vanished from output; the pixel probes caught what
|
||
the eye in a 2x time-lapse would not). **Rule: multi-stage fixed point shifts only
|
||
at the end; verify against a uniform-255 sample before believing it.** Second trap:
|
||
a stale-byte sentinel test whose source pattern can generate the sentinel value
|
||
itself (0xAB was a legitimate `x+y` pixel) — pick the sentinel coprime/out-of-range
|
||
to every channel formula (0xFD: odd, not ×4, above the R max). Third: a fake encoder
|
||
that HOLDS submitted frames now must snapshot them (`Clone`) once the producer
|
||
legitimately recycles buffers — mirror the real consumer's copy semantics in the fake.
|
||
|
||
|
||
---
|
||
|
||
## Splitting a large file into partials — NEVER `awk … > SRC` while awking SRC
|
||
|
||
(2026-08-31, Commit D) Tried to split `SocialsDialogViewModel.cs` in one line:
|
||
`{ awk '…' SRC; echo ""; awk '…2…' SRC; } > SRC`. The **first write truncated
|
||
SRC to 1 line**, so the second `awk` read the already-truncated file → the whole
|
||
source was lost (1 line left). Recovered with `git checkout -- SRC`, then redid
|
||
it, but the same bug could have meant making it up from scratch.
|
||
|
||
**Rule:** when a cut needs N blocks from one source into N files, never write a
|
||
block back onto the source that the `awk`s still read. Instead:
|
||
1. Read the source **once** at the start into temp files (`mktemp -d`, one file
|
||
per block), with a `$D` variable you carry forward.
|
||
2. Verify block sizes (`wc -l`) and brace balance (`python3 -c` counting `{`/`}`)
|
||
before touching any real file.
|
||
3. Then assemble each new file from `cat D/block …` — never truncating the source
|
||
until every read is done.
|
||
|
||
`git status` can't save you here if you don't notice until the file is gone —
|
||
`git checkout -- <path>` from the last commit is the recovery. Cheap insurance:
|
||
restore-then-retry, do it atomically from temp files the first time.
|
||
|
||
---
|
||
|
||
## Verifying an ffmpeg decode contract from WSL (no real CLR needed)
|
||
|
||
(2026-08-31, TASK 21) When a change depends on ffmpeg producing output with an
|
||
exact frame-size contract (rawvideo W×H×4 BGRA), you can prove the **command +
|
||
frame accounting** here without any .NET process:
|
||
|
||
1. Fetch a **static Linux ffmpeg** into `/tmp/opencode` (no sudo needed):
|
||
`curl -sLO https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz
|
||
&& tar -xf …`
|
||
2. Generate a tiny known clip: `ffmpeg -f lavfi -i "testsrc2=duration=1:size=640x360:rate=30" -pix_fmt yuv420p clip.mp4`
|
||
3. Decode with **exactly the app's args**: `-f rawvideo -pix_fmt bgra -vf scale=640:360 -an`
|
||
4. Assert `total_bytes % (W*H*4) == 0` (python3) → exact integer frames, no pad.
|
||
|
||
**Why not a dotnet-spawned ffmpeg here:** the only CLR on this box
|
||
(`/home/gramps/bin/dotnet`) is a **Windows-bound shim** — `Process.Start` resolves
|
||
paths to `\\wsl.localhost\Debian\…` and throws "not a valid application for this
|
||
OS" when handed a Linux ELF ffmpeg. So never plan to have dotnet exec a Linux
|
||
ffmpeg here; verify the contract with shell/python instead, and leave the
|
||
CLR→real-ffmpeg run to the native Windows suite.
|
||
|
||
## WPF hit-test truth in tests: `UIElement.InputHitTest`, NOT `VisualTreeHelper.HitTest`
|
||
|
||
(2026-09-01, the RoundClip "known failure" post-mortem — a failure the map carried as
|
||
"not a regression" for weeks without ever recording WHY.)
|
||
|
||
**The trap:** `VisualTreeHelper.HitTest(window, pt)` returned the window's
|
||
`WebViewHostPanel` overlay (`IsHitTestVisible="False"`, `Opacity=0`, ZERO children) for
|
||
EVERY point in the window — so a "corner is grabbable" assertion could never pass, and
|
||
it looked like a real interaction bug. The actual input pipeline (`UIElement.InputHitTest`,
|
||
what Mouse routing uses) correctly returned the element's Grid at elem-center/corner-in/
|
||
corner-exact and fell through to CanvasGrid just past the corner. The product was fine;
|
||
the TEST was probing an API that doesn't model input semantics.
|
||
|
||
**Rule:** any test asserting "where does a click land" uses `window.InputHitTest(pt)` +
|
||
`IsDescendantOf` — never `VisualTreeHelper.HitTest`.
|
||
|
||
**Diagnosis recipe (how the lie was caught in ~3 probe cycles, no guessing):** add a TEMP
|
||
probe `[Fact]` in the RealApp collection that hit-tests a spread of points
|
||
(elem-center / corner-in / corner-exact / corner-out / bg-center) and `Assert.Fail`s with a
|
||
composed dump: per-point VTH hit + `InputHitTest` hit + ancestor chain (`GetParent` walk
|
||
with `#Name`) + panel properties (`IsHitTestVisible/Opacity/children/actual size`) +
|
||
`TranslatePoint` origins. Run the class alone, read the message, delete the probe.
|
||
|
||
**Two sibling facts learned the same session (record-once):**
|
||
1. A UserControl owns its own XAML namescope — after extracting a region out of a window,
|
||
`window.FindName("InnerPart")` returns null; resolve the UserControl by its window-level
|
||
name, then `pane.FindName("InnerPart")`. And window-scope STYLES are invisible to a
|
||
UserControl's `StaticResource` at parse time — move such styles to `Themes/Controls.xaml`
|
||
(the app-scope rule exists for this).
|
||
2. Per-class `dotnet.exe vstest` from WSL DOES execute the RealApp/`MainWindow` tests fine
|
||
(they passed natively 2026-09-01) — only the FULL suite hangs (WASAPI startup). And the
|
||
test process shares `%APPDATA%\ytLlive\startup.log` with the real app: lines like
|
||
`camera 'test-camera' failed` are test noise, not DB state — to check pollution, query
|
||
the DB directly (`python3 sqlite3`, `SELECT DeviceId FROM Webcam`), not the log.
|
||
|
||
## A "known failure" label without a recorded cause = a bug on life support
|
||
|
||
(2026-09-01, the audio triple-take) One line — `_delayedMix` (nullable, added by TASK 22,
|
||
never initialized) dereferenced as `delayed.Length` — produced THREE symptoms that lived in
|
||
the map as two separate "pre-existing, do-not-chase" entries: (a)
|
||
`Mix_HonorsProviderGains…` "known failure", (b) `AudioPipelineTests` hangs when run at all
|
||
(a test reading a named pipe with no writer blocks — hung test ≠ flaky test, it's a starved
|
||
producer), (c) startup.log flooded "Audio live loop error" every 10ms (the mixer loop caught
|
||
and logged only `ex.Message` — stack thrown away).
|
||
|
||
Rules derived:
|
||
1. NEVER label a test "known/pre-existing" without writing WHY (exception type + first app
|
||
frame). An unexplained known-failure is deferred archaeology that hardens into fog.
|
||
2. A test that waits on IPC + a producer whose output vanished are usually ONE bug — look
|
||
for the producer before blaming the test.
|
||
3. Catch-and-log-swallow of `ex.Message` hides root causes; log with stack (`AppLog.Write(ex,
|
||
...)`) and throttle (5s) instead of dropping or flooding.
|
||
Bonus: the "failing" test encoded the MAP's contract (`loopbackGain = GameAudioVolume`); the
|
||
code had drifted to unity on a disproven premise (loopback capture does NOT follow endpoint
|
||
volume — creator's 20%-volume/pegged-meter observation killed it). The test was right all
|
||
along — failing tests may be the last honest witnesses; interrogate, don't pardon.
|
||
|
||
## Feature provenance: record WHO asked and WHY, in the task entry itself
|
||
|
||
(2026-08-31/09-01, the SYNC slider scare) TASK 22's lip-sync slider surfaced on the preview rail and
|
||
the creator's reaction was "totally don't remember ordering that" — because the queue entry recorded
|
||
WHAT shipped (a slider, 0-500ms, a converter class) but not WHO asked (the creator, explicitly, for
|
||
OBS's delay-filter fix built natively). Eight days later his own request read like AI drift and nearly
|
||
got deleted. Rule: the moment a creator-driven feature is queued or shipped, its entry carries a
|
||
one-line provenance — *who asked, what triggered it* ("creator: OBS delay-filter lip-sync fix, native").
|
||
Features without attribution become roadmap orphans that get punted, removed, or re-litigated. Same
|
||
disease as an unexplained "known failure" label — a fact recorded without its reason is a future
|
||
argument.
|
||
|
||
## An un-attributed build invalidated three takes of a perf saga — stamp the binary
|
||
|
||
(2026-09-04, takes 4–6) After each render-perf fix the creator "exed the code" and re-recorded, but
|
||
the exe timestamp ≠ binary contents (incremental builds reuse whatever compiles clean; a source edit
|
||
with no rebuild serves the OLD exe). Take 6 measured render WORSE than take 5 (35-41ms) and there was
|
||
no honest way to tell "the chat cache fix doesn't work" from "the fix was never running" — three
|
||
hours of diagnosis on an unattributable sample. Rule: if takes measure the app, EVERY build carries
|
||
an id and EVERY log line traces to it — `GenerateBuildStamp` (csproj) writes a fresh GUID per
|
||
compile (deliberately defeating incremental lies), the wordmark shows it as a superscript, startup.log
|
||
records `Build <id> (compiled <time>)`. A perf claim without build attribution is a guess; ask for the
|
||
stamp BEFORE theorizing. (Also this session: a sentinel-byte test where the source pattern could
|
||
GENERATE the sentinel, and a fixed-point bilinear that shifted BOTH stages and silently drew nothing —
|
||
see the rawvideo recipe.)
|
||
|