Files
ytLlive/TASKS.md
T
gramps 18a21010bb Scene compositor (TASK 4 ship step 1): encoder master-frame scaffolding
Software compositor that renders a scene into the encoder's master VideoFrame,
mirroring the XAML preview minus editing chrome: backdrop -> background ->
elements (UniformToFill cover-crop, round clip, mirror, opacity, border) ->
branding flash.

NOTE FOR USERS: this change shows NO difference in the app's UI — it is pure
backend scaffolding laying the groundwork for live video capture/streaming.
The preview you see is unchanged.

- Services/Compositor/: SceneCompositor (Render(scene, frameFor resolver,
  flashFrame, CompositorOptions)), CompositorOptions (source rect + output
  size; 16:9 full master, vertical 607x1080 -> 1080x1920), StretchMath (pure
  UniformToFill + bilinear), StaticPixelCache (asset bytes -> BGRA8 frame)
- Frame sources injected via Func<SceneElement, VideoFrame?> resolver, so the
  compositor is pure, WPF-free, and hermetic to test (D3D11 upgrade behind the
  same seam later)
- SceneElement.TryGetBorderColor public (shared hex parse), stale
  MainViewModel comment fixed, pre-existing CS1998 in YouTubeAuthServiceTests
  cleaned up
- Tests: SceneCompositorTests integration (full scene + vertical tier + flash)
  + StretchMath units, docs updated (72 tests passing, 0 warnings)
2026-08-10 10:15:58 -07:00

338 lines
31 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ytLlive — Task List
> Task queue and authoritative research. Memory-map conventions: [`schema.md`](schema.md);
> architecture/decisions: [`ai.md`](ai.md). Update statuses here whenever a task moves.
## YouTube Live API — research facts (authoritative, v3 build)
Lifecycle: `created → ready → [testing] → live → complete` (transitional `liveStarting` / `testStarting`).
- **liveBroadcasts.insert** requires: `snippet.title`, `snippet.scheduledStartTime`, `status.privacyStatus`, `status.selfDeclaredMadeForKids` (COPPA).
- **liveStreams.insert** requires: `snippet.title`, `cdn.frameRate`, `cdn.ingestionType`, `cdn.resolution`. **None of the four (except title) can ever change after creation** — changing them means delete + recreate the stream. This is the hard constraint behind the quality grey-out.
- **Title / description / privacy**: editable at any time, including while live (`liveBroadcasts.update`, part=`snippet,status`).
- **contentDetails** (DVR, recordFromStart, monitorStream, embed, latency): editable only in `created` / `ready`.
- **Transition to live** only allowed when the bound stream's `status.streamStatus == active`.
### Two features that reshape the design
1. **enableAutoStart / enableAutoStop** — instant one-click go-live, no transition call. With `enableAutoStart=true` we never call `transition(live)`: the broadcast auto-goes-live the moment the encoder starts. Combined with `enableMonitorStream=false` (our preview pane replaces YouTube's monitor stream — the thing that forces a testing stage), the flow is **create → bind → Start Stream → encoder starts → YouTube brings it live**. No testing, no transition polling, no liveStarting stuck-state handling.
2. **`cdn.resolution=variable` / `cdn.frameRate=variable`** — free auto step-down. YouTube auto-detects what we send; since we ARE the encoder we can drop bitrate/resolution on the fly with zero API calls. Declaring an explicit resolution instead (e.g. 1080p) requires a new stream, which can't happen mid-broadcast. Variable is the enabler for the whole auto step-down feature.
### Compliance gotchas (maps perfectly to report-by-exception)
- `liveStreams.status.healthStatus`: `good | ok | bad | noData` plus `configurationIssues[]` with `type` + `severity` (`info|warning|error`). Literally built for report-by-exception — poll it, render nothing on good/ok, surface a banner only on warning/error. No need to invent our own health logic.
- Encoder must comply or YouTube flags it: keyframes ≤ 4s (`gopSizeLong`), closed GOP, H.264, audio AAC/MP3 @ 44.1/48kHz, mono/stereo only.
- Error codes to handle: `errorStreamInactive`, `invalidTransition`, `redundantTransition`, `liveStreamDeletionNotAllowed`, `liveStreamModificationNotAllowed`, `liveBroadcastBindingNotAllowed`.
### Tips we should take advantage of
- **Reusable streams** (`isReusable=true`): one stream per channel, cache its ingestion URL + stream name, reuse for every broadcast. No rebinding dance each go-live. This is exactly the manual-stream-key baseline.
- **Backup ingestion address**: YouTube provides a simultaneous-push backup — future hardening, not v1.
- **`recordFromStart` + `enableDvr` default true** → every live is auto-recorded and immediately replayable. Free VOD archive, matches the v0.2 recording goal.
- **`latencyPreference`: `normal | low | ultraLow`** — for homelab streamers talking to chat, `low` (or `ultraLow`, capped at 1080p) is a real feature.
- **Broadcast ID == Video ID** — one ID to track everything.
---
## TASK 1 — Initial Scaffold
**Goal:** Working C# / WPF project with MVVM architecture, dark-theme main window, and YouTube service stubs.
### Status: ✅ Done
- Models: Scene, Source, StreamConfig, StreamHealth, YouTubeChannel, ChatMessage
- Services: YouTubeAuthService (OAuth2), YouTubeStreamService (broadcast/health), YouTubeChatService (chat polling)
- MainViewModel: scene management, stream controls, chat
- MainWindow: scene/source panel, preview area, chat panel, status bar
- Clean build, 0 warnings (WSL + Windows)
---
## TASK 2 — YouTube OAuth2 Authentication
**Goal:** Fully working Google OAuth2 flow — user clicks "YouTube", browser opens, authorization callback lands, channel info is stored.
**Design constraint:** Sign-in must NEVER block core exploration. Users can build scenes, add sources, and audition the software without authenticating. But **going live requires authentication** — the "Start Stream" dialog is where the account sign-in lives, alongside all stream metadata.
**Two-state flow:** There is no separate "Connect" button. The top bar shows a single button — **Start Stream** when idle, **End Stream** when live. Clicking Start Stream opens one dialog that supplies everything: account (previously-saved account shown as default, with a Change Account action) + title/description/visibility.
**Live indicators (unmissable):** Top bar + window title bar flip red, a pulsing **● LIVE** badge with elapsed timer appears in the top bar, the preview area gets a red glow, and the taskbar icon shows a red overlay dot. Window title bar shows the stream title once validated.
### Requirements:
1. **Google Cloud OAuth credentials** — client ID + secret, **baked into `Helpers/OAuthCredentials.cs`** (desktop "Desktop app" OAuth client; loopback callback — no console redirect URI registration needed; creators never configure)
2. **Local HTTP listener**`HttpListener` on `http://localhost:PORT/oauth2/callback` to catch the redirect
3. **Browser launch** — open the authorization URL in the default browser
4. **Token persistence** — store access/refresh tokens securely (Windows DPAPI), reload on startup
5. **UI state** — account shown in the Start Stream dialog; "Change Account" action triggers re-auth
6. **Go Live gated on auth** — Start Stream dialog requires sign-in to enable the Start button; scene building works without it
### Tests:
- Mock token exchange response, verify channel info parsed
- Verify token refresh triggers when near expiry
- Verify credential load/save roundtrip
### Status: ✅ Complete
- ✅ Two-state Start/End Stream button, go-live dialog (account + title/description/visibility), red top bar, pulsing LIVE badge + elapsed timer, preview glow, taskbar red dot
- ✅ Real OAuth2 wiring — baked-in Google credentials (desktop client; loopback callback) + `YouTubeAuthService` complete: browser launch, `HttpListener` callback, token exchange, refresh, channel fetch
- ✅ Token persistence via Windows DPAPI (`Helpers/TokenStore.cs``%APPDATA%\ytLlive\ytLlive.auth`), best-effort reload + proactive refresh at startup, saved after every exchange/refresh
- ✅ Account sign-in/change surfaced in the GoLive dialog (saved account shown with "Change Account"; "Sign in to YouTube" when none; Start disabled until signed in)
-**End Livestream signs out** — a graceful end completes the session: `StopStream()` calls `YouTubeAuthService.ClearSession()` + `TokenStore.Clear()` + `IsConnected = false`, so the next Start Stream dialog requires a fresh sign-in. A crash never runs End, so the DPAPI token survives and the creator stays signed in. Resume/reconnect after a midstream crash is deliberately deferred to TASK 3: the socket can't be resumed (it dies with the process), so "resume" = fast reconnect with a saved broadcast ID/stream key within YouTube's disconnect-grace window; too slow and `enableAutoStop` ends the broadcast
- ✅ Tests in `ytLive.Tests` (xUnit, net8.0-windows): TokenStore DPAPI roundtrip/corrupt/missing/clear + mocked exchange channel-parse + refresh expiry bump + `ClearSession` — 7 passing
---
## TASK 3 — Capture Pipeline (Scenes/Sources)
**Goal:** Real video preview in the center panel — the minimal source set below, composited per scene.
### The Minimal Source Set (design decision — do not expand casually)
ytLlive is YouTube-only and 90% of users are casual. OBS's long source list is off-putting; we ship
the hot few and nothing esoteric. If a user needs more, they've graduated to OBS.
1. **Webcam** — the face cam. Non-negotiable.
2. **Screen** — the main event (game, slides, browser). One source; a picker chooses a monitor *or* a
window. (Window capture is absorbed here — no separate source type.)
3. **Background** — a full-canvas backdrop image. Fills the whole scene automatically, zero fiddling.
Kept separate from Image on purpose: same pixels, but this one needs no positioning.
4. **Image** — a floating graphic/logo overlay (watermark, badge, corner branding). Free-positioned.
5. **Text** — live text ("Starting soon", "Back in 5", handle, callout). Casual streamers live on this.
6. **Chat box** — YouTube live chat rendered *on* the stream so viewers read along in-video. YT-native.
7. **Alerts** — Super Chat / membership / subscribe pop-ins. The dopamine source. **The one big lift**
(Super Chat event streaming + on-stream rendering/animation); build after the six. **Also the one
paid feature** — see Monetization in `ai.md`.
Deliberately NOT supported: game capture, browser source, media playlist, VLC, color-key voodoo, MIDI.
### Source memory model (design decision)
- A scene has resources. Resources can be shared across scenes.
- A resource exists exactly once in memory, no matter how many scenes use it (a logo in five
scenes = one loaded bitmap).
- Every resource carries a **catalog of scenes**: one usage entry per scene it appears in, each
entry dictating that scene's use — placement (X/Y/Width/Height), opacity, z-order, enabled,
scale mode, crop.
- Usages are named `{resourceName}.{sceneName}` — whatever the user named the resource, dot, the
scene name: `logo.starting`, `logo.live`, `myPic.brb`. Not a hardcoded "logo".
- A webcam in two scenes = one capture session, two catalog entries.
- Refcount by catalog size: the last usage removed → the resource is disposed and evicted.
- The resource (not a per-scene node) owns everything `IDisposable`.
### Scene transitions (design decision)
Scene switching while live must never stutter. Supported types, most → least economical:
1. **Cut** — instant switch. The default. Zero cost.
2. **Fade** — short crossfade (~300ms).
3. **Move** — a simple, economical move transition, done to perfection and memory-efficient. The
smart streamer's bread and butter.
4. **Custom (media) transitions** — require media elements (video/stinger playback during the
transition). Heavier, but creators pay for these, so we support them. Their media follows the
same resource memory model: loaded once, catalogued by scene.
Preview shows the transition too (WYSIWYG). No wipes/slides/LUTs beyond the four above.
### Requirements:
1. **Screen** — Windows.Graphics.Capture (WinRT), enumerate displays/windows, picker
2. **Webcam** — MediaCapture (WinRT SDK projection) with device enumeration — ✅ **milestone 1 done**:
- TFM bumped to `net8.0-windows10.0.19041.0` (app **and** tests) so the WinRT projection resolves from the SDK reference packs — no NuGet package, no capability manifest (unpackaged desktop app)
- `MediaCaptureFrameSource` (CPU-first: `MemoryPreference = Cpu`, BGRA8 via `CreateFrameReaderAsync`), `MediaCaptureCameraEnumerator` (`DeviceInformation.FindAllAsync(DeviceClass.VideoCapture)`)
- `CameraManager`: refcounted by `DeviceId`, one shared `WriteableBitmap` app-wide, dispatcher-coalesced UI updates (~render rate, latest-frame drop), placeholder/`AppLog` + warning on failure
- `CameraPickerDialog` (mirror of `ReuseImageDialog`) — "Searching for cameras…" / list / "No cameras found" states
- One webcam app-wide: Add → Webcam greyed out once one exists ("it's already in your stream" tooltip); persisted `DeviceId` re-acquires after layout load
- Default placement 16:9 **480×270**, bottom-right, 32px margin; drag/resize/selection shared with Image sources
- **Clip shapes: Traditional + Round** (phone view dropped — the 9:16 phone output is the vertical output-crop tier); **mirror**; both persisted in the layout DB (schema v2) and toggled from the source chip
- **Background removal = milestone 2** (ONNX Runtime + DirectML, MediaPipe Selfie Segmentation) — not in this build
3. **Background / Image / Text** — static sources positioned/scaled/opacity
4. **Chat box** — rendered from the live chat poll (right panel is the same feed, raw)
5. **Scene compositing** — per-scene source layering (z-order = sources list order, top-to-bottom
back-to-front), preview rendered via D3DImage or MediaElement
6. **Branding flash** — the topmost full-frame "made with ytLlive!" layer at ~25% opacity, ~1s on /
300s off (see Monetization in `ai.md`), gated on `BrandFlashEnabled` + live/recording. Lives in the
preview compositor now (`BrandFlashLayer` in `MainWindow.xaml` CanvasGrid, driven by
`BrandFlashActive`/`BrandFlashTimer` in `MainViewModel`); the encoder output renders the same layer,
and v0.2 local recordings carry it too
7. **Drag/drop placement & reorder** — intuitive, visual (per design principle):
- **Preview:** click-drag a source in the center panel to reposition it; resize via handles
- **Scenes list:** drag rows to reorder scenes
- **Sources list:** drag rows to reorder sources (this *is* the z-order) — implemented
### Status: 🔶 In progress — milestone 1 (webcam) shipped; **schema v3 (Ship Branch A) shipped**: multi-scene webcam (singleton `Webcam` + per-scene `WebcamSceneConfig`), right-click border/context menu, static OSB-style borders, 50%-per-dimension webcam size cap, device-swap (`ReleaseAllAsync`); **schema v4**: round→rect restore persisted (`WebcamSceneConfig.RectWidth`/`RectHeight`) + one-time legacy-square 16:9 heal on load — 25 tests passing; **screen backdrop (ship task #1) shipped**: live desktop/game capture as a permanent, non-deletable bottom layer (`Source.IsBackdrop`, schema v5), auto-detecting the full-screen game at launch/focus (else the **primary display** — never assumed monitor 0) via `Win32FullScreenDetector` (now with `GetDisplays()`/`PrimaryMonitorIndex()` for the in-app display picker), content re-designated via the OS `GraphicsCapturePicker` ("Change Capture…") or the in-app "Capture Display" submenu, refcounted/shared capture sessions in `ScreenCaptureManager` mirroring `CameraManager` — 45 tests passing; **schema v6**: `Scene.HasBackdrop`, now **Live-only by policy** — the backdrop is enforced by scene name on every load (`EnforceBackdropPolicy`: Starting/BRB/Chat/Ending never carry one; the one-time v5→v6 backfill covers all four), the scene context-menu "Backdrop" checkbox is gone (policy owns the flag), preview watermark hides when the backdrop renders, capture changed to `WindowsRuntimeMarshal.TryGetDataUnsafe` (the CsWinRT-safe frame-read) + downscale to the 1920×1080 master + 5s-throttled error logging (was flooding `startup.log` with 5 MB of cast errors and burning CPU), round webcam no longer re-rasterizes an `ImageBrush` every frame (Image + EllipseGeometry clip) — the live-mode stutter fix; **the five-scene catalog (`SceneCatalog`)**: Starting/Live/BRB/Chat/Ending is the product — work with less, never more; the (+) button only shows when a canonical scene is missing and re-adds it (its menu lists only the missing ones); **webcam-after-session-start fix**: a webcam added to a scene after the camera was already running (e.g. Chat) previously rendered a transparent container — `CameraManager.GetPreviewBitmap` + propagation in `AddWebcamToActiveSceneAsync`/`ReacquireWebcam` now hands the running shared frames to any newly added `WebcamSceneConfig` — 62 tests passing; the **Chat scene's webcam size cap** is raised from 50%-per-dimension (960×540) to half the screen AREA (~1358×764 @16:9, `MaxWebcamWidthFor`/`MaxWebcamHeightFor` keyed by canonical name) so the viewer sees the creator better — 65 tests passing; **"Add Webcam" always opens the camera picker** (deleting one scene's webcam then re-adding used to resurrect the old camera when another scene still used it — `SwapWebcamIdentityAsync` now swaps the app-wide identity if a different camera is chosen, same path as "Change Webcam…"); scenes/sources UI (add/reorder/rename, image + background overlays with move/resize/opacity/reuse) built; **audio UX shipped (UI)**: the bottom-bar footer is now two lines (dropped/duration moved under bitrate/fps), with the mic's sound meter + mute button + volume slider grouped CENTERED on the footer's top line, beneath the preview panel (meter: 288px, muted slate track with ruler graduations + muted yellow/red zone tints, green→yellow→red fill; mute = speaker icon → red do-not-symbol when muted, and the slider and speaker stay in sync (volume 0 ⇔ muted — sliding off flips the speaker to muted, sliding up from 0 clears it); mic volume defaults to 80%, muting zeroes the meter and restores the prior volume on unmute (which flashes the meter to the restored position ~300ms before it returns to the live level); the meter is a READ-ONLY realtime level display (fill = live level × volume — volume is a gain on ambient noise; while the slider is dragged the bar previews the slider position and bounces back to the live level on release, which is 0 with no input — clicking the meter does nothing), clicking the MIC label opens a microphone picker whose chosen source shows left-justified inside the meter bar (fill at 75% opacity so the name + ruler markings show through); slim dimensional slider — gradient track/fill, gloss-sphere thumb; the old flat pink 18px-filled one is gone), everything else on line 2 (bitrate/fps/dropped/duration/health left, quality + gear right) — the creator's only audio control, desktop/game audio is automatic (KISS rule); the connected YouTube account's avatar/name shows in the top bar next to Start Stream (`SyncConnectedAccount`); the scenes list is content-height now (no dead space before SOURCES); window capture, compositing, encoding pending
---
## TASK 4 — RTMP Ingest to YouTube
**Goal:** Push encoded video to YouTube's RTMP ingest.
### Requirements:
1. **Encoding** — H.264 (hardware via NVENC/AMD, fallback x264) + AAC audio; **must comply**: keyframes ≤ 4s (gopSizeLong), closed GOP, AAC/MP3 @ 44.1/48kHz, mono/stereo only. **License posture (decided): GPL-free build** — NVENC (NVIDIA) / QSV (Intel) / AMF (AMD) + OpenH264 software fallback + built-in AAC; no libx264 (GPL contaminates a paid product). Output containers are identical either way (H.264+AAC in `.flv` for RTMP, `.mp4`/`.ts` for VOD) — the format is NOT the differentiator, the license and per-GPU quality are.
2. **RTMP push****FFmpeg subprocess (decided)**: app feeds raw frames via stdin, parses stderr for health; one battle-tested binary does encode + FLV mux + push + reconnect. **Binary distribution (decided): check-then-pull** — probe `where ffmpeg`/PATH at first go-live; if absent, download a **pinned** build (~30 MB, standard gyan.dev/BtB N — no custom minimal build) to `%APPDATA%\ytLlive\tools\ffmpeg.exe` and cache it, offline-friendly. Behind an `IFfmpegLocator` seam so tests fake it. Push goes to the cached reusable stream's ingestion URL
3. **Quality ladder** — the offered tiers, with **1080p60 @ 8 Mbps as the standard/default**:
- 720p30 @ 6 Mbps
- 720p60 @ 6 Mbps
- 1080p30 @ 8 Mbps
- **1080p60 @ 8 Mbps** (default — mainstream ceiling, GPU hardware-encoded so the gaming
machine never notices; upload headroom stays comfortable)
- Vertical 1080×1920 @ 60fps @ 8 Mbps (9:16 phone tier)
The composition master is always 1920×1080; a tier is an output rect + target resolution
(see `ai.md` "Resolution tiers"). Vertical output = the centered 607×1080 crop of the master
scaled to 1080×1920 (semi-crop preview is already implemented; the encoder applies the same rect).
1080p60 is the ceiling by design — "if you want 1440 or 4K or 8K → OBS is your solution"; the app
targets the most mainstream creator, not power users.
Ladder is sculpted by a **cached probe** (IP-only TCP vs public ingest host; no auth required).
Quality is greyed out while live because the declared resolution can't change mid-stream — but
with `variable`, we can **auto step-down** bitrate/resolution on the fly with zero API calls
(no stream recreation); 60fps presumes a hardware encoder — no hardware encoder → auto
fallback to 720p60/1080p30
4. **Stream key management** — reuse the cached reusable stream (one per channel) instead of creating a new one per go-live; prefill default YouTube ingest URL `rtmp://a.rtmp.youtube.com/live2`
5. **Health stats** — bitrate, FPS, dropped frames reported live in the bottom bar (encoder-side)
6. **One-click go live** — defaults that work out of the box
7. **Audio capture (feeds the meter — this task ships the wiring)** — WASAPI loopback (desktop/game at unity, zero UI — "it just is") + the picked mic (`MicSourceName` from the `MicPickerDialog`). The mic capture feeds `AudioLevel` so the realtime meter comes alive (today it reads 0 — the mixer feed is pending, see `ai.md` audio notes). AAC mono/stereo @ 48 kHz per the compliance rules.
### Status: 🔶 In progress — **ship step 1 (the output compositor) SHIPPED** (2026-08-10); encoder/RTMP/audio follow it
The pipeline chain the encoder needs doesn't exist yet: **scene compositing** (the master 1920×1080 frame
without the preview's editing chrome) → **audio capture** (WASAPI, feeds the meter) → **H.264+AAC encode**
**vertical-tier crop/scale****RTMP push****health stats** into the bottom bar. Nothing can encode
until a frame source exists, so the compositor is ship step 1.
#### Ship step 1 — Scene compositor (the frame source)
**Goal:** a pure-CPU software compositor producing the encoder's master frame (BGRA8, the `VideoFrame`
seam) from the scene model. The preview stays XAML (the editing view); the compositor is the **output
view** — WPF's `RenderTargetBitmap` can't be used (software-rendered + captures chrome). Two renderers
must agree, so the XAML (`MainWindow.xaml` CanvasGrid + element DataTemplate) is the contract.
**Decisions (locked 2026-08-10):** **Path A CPU blitter** — GPU effort belongs to NVENC (the encoder),
not composition; with an FFmpeg subprocess the master crosses a CPU readback to the pipe every frame
anyway, so GPU compositing buys ~nothing at this layer count (2-3 live layers; static layers
pre-composite once). A D3D11 compositor can replace this one later **behind the same seam** (the CPU
master buffer stays the contract). **Render the output rect directly**: compositor is constructed with
`CompositorOptions {SourceRectX/Y/W/H, OutputWidth, OutputHeight}`; 16:9 tiers = full 1920×1080 1:1;
vertical (9:16) = composite the centered 607×1080 crop then bilinear-upscale to 1080×1920. Reuses
`MainViewModel.OutputRectX/Y/W/H` (note `(1920607)/2 = 656.5` → align to integer pixels for output).
**Render spec (back → front, mirror the XAML exactly):**
1. Backdrop — the Live scene's `IsBackdrop` Source (`CaptureKey` → live frame), `UniformToFill`
full-frame (XAML's separate `BackdropImage` layer; the backdrop *element* renders nothing — its
DataTemplate Image is Collapsed for DisplayCapture).
2. Background — the scene's `Background` Source, `UniformToFill` full-frame (the `ActiveBackgroundImage`
layer, not per-element).
3. Elements in `Scene.Elements` order (back→front), skip `IsVisible=false`. What actually renders:
- `Source` Type `Image` → static asset, `UniformToFill` cover-crop into (X, Y, W, H)
- `WebcamSceneConfig` → latest frame by `DeviceId`: Traditional = `UniformToFill` rect; Round = circle
diameter `min(W,H)` (alpha 0 outside — true circle, not oval); mirror = horizontal flip around
element center (`MirrorScale`); opacity = per-pixel multiply (content + border); border = stroked
rect / centered circle at `RoundBorderSize`, width `BorderWidth`, alpha `BorderOpacity`
- `Background` / `IsBackdrop` / `TextOverlay` are NOT per-element (layers above; Text not shipped)
4. Branding flash — pre-rendered full-frame "made with ytLlive!" at 25% alpha when live +
`BrandFlashEnabled` + timer active. Passed in as a `VideoFrame?` (compositor core stays pure byte-math,
no WPF; likely a bundled asset rather than runtime text rendering).
5. NOT in output (preview chrome only): SelectionOverlay, DimRects, output-rect outline, badge, placeholder.
**New files (all in `Services/Compositor/`):**
- `SceneCompositor.cs` — `Render(Scene, frameFor: Func<SceneElement, VideoFrame?>, flashFrame:
VideoFrame?, CompositorOptions) → VideoFrame` (output-sized). The caller's `frameFor` resolver maps
each element to its frame (webcam → DeviceId, image → AssetId via `StaticPixelCache`, backdrop →
CaptureKey) — the compositor stays pure/hermetic/no WPF.
- `CompositorOptions.cs` — source-rect + output W×H.
- `StretchMath.cs` — `UniformToFill` cover-crop, ellipse mask, bilinear scale (pure, unit-tested).
- `StaticPixelCache.cs` — asset `byte[]` → cached BGRA `VideoFrame` (WPF `BitmapDecoder` + `CopyPixels`,
decode once per content hash).
**Test plan (Good Dog Rule — ONE integration test):** `SceneCompositorTests` — a scene with backdrop
(solid red fake frame) + round webcam (solid green) + image (solid blue) → render 16:9 master → assert
per-layer probe pixels (corner = backdrop color, element center = webcam color, outside the round clip =
backdrop color, mirrored element swaps left/right); a vertical-tier variant asserts 1080×1920 output +
crop fidelity. Focused unit tests on `StretchMath`. Tests push frames directly — no capture managers
involved (they wire in a later step).
**Same-PR housekeeping:** fix the stale comment `MainViewModel.cs:324` ("shown under the meter on line 2"
→ "shown left-justified INSIDE the meter bar" — `ai.md` is the authority); this task's requirements now
include the explicit audio-capture/meter wiring (#7 above).
**Out of scope (later ship steps):** FFmpeg locator + license posture (covered in requirements 1-2),
encoder + RTMP push, WASAPI audio capture (loopback + mic) feeding `AudioLevel`, wiring
`CameraManager`/`ScreenCaptureManager` into the frame pipeline, brand-flash timer wiring, health stats
(bitrate/FPS/dropped).
**Built (2026-08-10):** all four files shipped in `Services/Compositor/`, `SceneElement.TryGetBorderColor`
made public (shared hex parse with the compositor — no duplicated color parsing), the stale
`MainViewModel.cs:324` comment corrected, and the pre-existing CS1998 in `YouTubeAuthServiceTests`
cleaned up — build **0 warnings**. Tests: the `SceneCompositorTests` integration test (full-scene master
pixels, vertical tier, flash) + 4 `StretchMath` units — **72 passing**.
---
## TASK 5 — YouTube Live Stream Management
**Goal:** Create/bind broadcasts, monitor YouTube-side stream health — the v3 way.
### Design decisions (v3)
1. **One-click go-live** — `liveBroadcasts.insert` with `enableAutoStart=true`, `enableAutoStop=true`, `enableMonitorStream=false`, `selfDeclaredMadeForKids=false`, `latencyPreference=low`. No `transition(live)` call, no testing stage, no liveStarting polling. Encoder starts → YouTube brings it live by itself.
2. **Variable reusable stream** — `cdn.resolution=variable`, `cdn.frameRate=variable`, `isReusable=true`. Create once per channel, cache ingestion URL + stream name, reuse for every broadcast. Any quality tier works without recreation; auto step-down needs no API calls.
3. **Report-by-exception** — poll `liveStreams.list`; banner only on `healthStatus` warning/error issues (`configurationIssues[]`). Bottom strip = YouTube logo + green/red connection dot (clickable → opens the dialog).
4. **One dialog, three states** — `not connected` (sign-in) / `connected-offline` (all editable) / `live` (title + description + visibility editable; quality + account greyed out). Both entry points (Start Stream button + bottom strip) open it; prefilled from saved session profile.
5. **Live edits** — `liveBroadcasts.update` with part=`snippet,status` for title/description/privacy.
6. **End stream** — stop encoder → `transition(complete)`, with `enableAutoStop` as the safety net.
7. **Broadcast ID == Video ID** — one ID to track status, health, and the auto-created VOD (`recordFromStart` + `enableDvr`).
### Requirements:
1. **Broadcast creation** — title/description/privacy/scheduledStartTime via API, with the v3 flags above
2. **Reusable stream** — create once, cache + reuse; bind to broadcast
3. **Health monitoring** — poll `liveStreams.list` `healthStatus` + `configurationIssues[]`, surface banner only on warning/error
4. **Live chat** — poll `liveChat/messages`, render in right panel, support Super Chat + membership badges
5. **Error handling** — the YouTube error codes: `errorStreamInactive`, `invalidTransition`, `redundantTransition`, `liveStreamDeletionNotAllowed`, `liveStreamModificationNotAllowed`, `liveBroadcastBindingNotAllowed`
### Status: Not started
---
## TASK 6 — Layout Persistence (SQLite)
**Goal:** Scenes, sources, and asset bytes survive restarts; assets are always available.
### Design decisions
1. **SQLite database** (`Microsoft.Data.Sqlite`) at `%APPDATA%\ytLlive\ytLlive.db`; schema versioned
via `PRAGMA user_version`.
2. **Assets live in the DB, not on disk** — `Asset` table stores image bytes (BLOB) keyed by a
SHA-256 content hash (unique). Identical image content collapses to one row regardless of file
name — the 1:M resource memory model, enforced by the database. No file paths; deleting the
original file never breaks a scene.
3. **File-model save/open** — the active layout file is tracked (default is the AppData DB).
**Save Layout As… / Open Layout…** switch the active file; auto-save writes to whatever is active.
4. **Auto-save (invisible)** — ~1.5s debounce on scene add/remove/reorder/rename/hide, source
add/remove/reorder, and any source transform change; flush on window close.
5. **Schema** — `Scene` (Id, Name, IsHidden, IsChatScene, HasBackdrop, SortOrder), `Asset` (Id, Hash, Data,
PixelWidth, PixelHeight), `Source` (Id, SceneId FK cascade, AssetId FK, Type, Name, IsEnabled,
X/Y/Width/Height/Opacity, MonitorIndex, DeviceId, ClipShape, IsMirrored, SortOrder) —
`user_version` **6** (v1 → v2 = `ALTER TABLE` adds the two webcam columns; v3 = singleton
`Webcam` + per-scene `WebcamSceneConfig`; v4 = `WebcamSceneConfig.RectWidth`/`RectHeight` for the
round-to-rect restore; v5 = `Source.IsBackdrop` + `Source.CaptureKey` for the live-capture
backdrop; v6 = `Scene.HasBackdrop` — the backdrop is **Live-only by policy**
(one-time backfill turns Starting/BRB/Chat/Ending off and drops their backdrop
sources; `EnforceBackdropPolicy` re-normalizes every load). `WindowHandle` stays in-memory
(per-session). Save = transactional rewrite; orphaned assets pruned.
6. **Startup** — load the active file; seed the five canonical scenes
(Starting/Live/BRB/Chat/Ending, `SceneCatalog`) only when the DB is empty. The (+)
button re-adds a missing canonical scene and is hidden once all five are present;
adding beyond the five is rejected — work with less, never more.
### Status: ✅ Implemented
---
## Backlog (future versions)
- v0.2 — Recording to local file (recordings carry the branding flash — see TASK 3 / `ai.md` Monetization)
- v0.3 — Stream scheduling
- v0.4 — Multi-destination restreaming
- v0.5 — Stream clipping