45 KiB
ytLlive — Task List
Task queue and authoritative research. Memory-map conventions:
schema.md; architecture/decisions:ai.md. Update statuses here whenever a task moves.
Checklist markers — every task's Status list uses the same states:
- ✅ — completed (green check)
- ☐ — not completed / pending (empty box)
- ❌ — exception (blocked, known-issue, or deliberately excluded from this build)
YouTube Live API — research facts (authoritative, v3 build)
Lifecycle: created → ready → [testing] → live → complete (transitional liveStarting / testStarting).
- liveBroadcasts.insert requires:
snippet.title,snippet.scheduledStartTime,status.privacyStatus,status.selfDeclaredMadeForKids(COPPA). - liveStreams.insert requires:
snippet.title,cdn.frameRate,cdn.ingestionType,cdn.resolution. None of the four (except title) can ever change after creation — changing them means delete + recreate the stream. This is the hard constraint behind the quality grey-out. - Title / description / privacy: editable at any time, including while live (
liveBroadcasts.update, part=snippet,status). - contentDetails (DVR, recordFromStart, monitorStream, embed, latency): editable only in
created/ready. - Transition to live only allowed when the bound stream's
status.streamStatus == active.
Two features that reshape the design
- enableAutoStart / enableAutoStop — instant one-click go-live, no transition call. With
enableAutoStart=truewe never calltransition(live): the broadcast auto-goes-live the moment the encoder starts. Combined withenableMonitorStream=false(our preview pane replaces YouTube's monitor stream — the thing that forces a testing stage), the flow is create → bind → Start Stream → encoder starts → YouTube brings it live. No testing, no transition polling, no liveStarting stuck-state handling. cdn.resolution=variable/cdn.frameRate=variable— free auto step-down. YouTube auto-detects what we send; since we ARE the encoder we can drop bitrate/resolution on the fly with zero API calls. Declaring an explicit resolution instead (e.g. 1080p) requires a new stream, which can't happen mid-broadcast. Variable is the enabler for the whole auto step-down feature.
Compliance gotchas (maps perfectly to report-by-exception)
liveStreams.status.healthStatus:good | ok | bad | noDataplusconfigurationIssues[]withtype+severity(info|warning|error). Literally built for report-by-exception — poll it, render nothing on good/ok, surface a banner only on warning/error. No need to invent our own health logic.- Encoder must comply or YouTube flags it: keyframes ≤ 4s (
gopSizeLong), closed GOP, H.264, audio AAC/MP3 @ 44.1/48kHz, mono/stereo only. - Error codes to handle:
errorStreamInactive,invalidTransition,redundantTransition,liveStreamDeletionNotAllowed,liveStreamModificationNotAllowed,liveBroadcastBindingNotAllowed.
Tips we should take advantage of
- Reusable streams (
isReusable=true): one stream per channel, cache its ingestion URL + stream name, reuse for every broadcast. No rebinding dance each go-live. This is exactly the manual-stream-key baseline. - Backup ingestion address: YouTube provides a simultaneous-push backup — future hardening, not v1.
recordFromStart+enableDvrdefault true → every live is auto-recorded and immediately replayable. Free VOD archive, matches the v0.2 recording goal.latencyPreference:normal | low | ultraLow— for homelab streamers talking to chat,low(orultraLow, capped at 1080p) is a real feature.- Broadcast ID == Video ID — one ID to track everything.
TASK 1 — Initial Scaffold
Goal: Working C# / WPF project with MVVM architecture, dark-theme main window, and YouTube service stubs.
Status: ✅ Done
- ✅ Models: Scene, Source, StreamConfig, StreamHealth, YouTubeChannel, ChatMessage
- ✅ Services: YouTubeAuthService (OAuth2), YouTubeStreamService (broadcast/health), YouTubeChatService (chat polling)
- ✅ MainViewModel: scene management, stream controls, chat
- ✅ MainWindow: scene/source panel, preview area, chat panel, status bar
- ✅ Clean build, 0 warnings (WSL + Windows)
TASK 2 — YouTube OAuth2 Authentication
Goal: Fully working Google OAuth2 flow — user clicks "YouTube", browser opens, authorization callback lands, channel info is stored.
Status: ✅ Done
- ✅ Two-state Start/End Stream button, go-live dialog (account + title/description/visibility), red top bar, pulsing LIVE badge + elapsed timer, preview glow, taskbar red dot
- ✅ Real OAuth2 wiring — baked-in Google credentials (desktop client; loopback callback) +
YouTubeAuthServicecomplete: browser launch,HttpListenercallback, token exchange, refresh, channel fetch - ✅ Token persistence via Windows DPAPI (
Helpers/TokenStore.cs→%APPDATA%\ytLlive\ytLlive.auth), best-effort reload + proactive refresh at startup, saved after every exchange/refresh - ✅ Account sign-in/change surfaced in the GoLive dialog (saved account shown with "Change Account"; "Sign in to YouTube" when none; Start disabled until signed in)
- ✅ End Livestream signs out — a graceful end completes the session:
StopStream()callsYouTubeAuthService.ClearSession()+TokenStore.Clear()+IsConnected = false, so the next Start Stream dialog requires a fresh sign-in. A crash never runs End, so the DPAPI token survives and the creator stays signed in. Resume/reconnect after a midstream crash is deliberately deferred to TASK 3: the socket can't be resumed (it dies with the process), so "resume" = fast reconnect with a saved broadcast ID/stream key within YouTube's disconnect-grace window; too slow andenableAutoStopends the broadcast - ✅ Tests in
ytLive.Tests(xUnit, net8.0-windows): TokenStore DPAPI roundtrip/corrupt/missing/clear + mocked exchange channel-parse + refresh expiry bump +ClearSession— 7 passing
Design constraint: Sign-in must NEVER block core exploration. Users can build scenes, add sources, and audition the software without authenticating. But going live requires authentication — the "Start Stream" dialog is where the account sign-in lives, alongside all stream metadata.
Two-state flow: There is no separate "Connect" button. The top bar shows a single button — Start Stream when idle, End Stream when live. Clicking Start Stream opens one dialog that supplies everything: account (previously-saved account shown as default, with a Change Account action) + title/description/visibility.
Live indicators (unmissable): Top bar + window title bar flip red, a pulsing ● LIVE badge with elapsed timer appears in the top bar, the preview area gets a red glow, and the taskbar icon shows a red overlay dot. Window title bar shows the stream title once validated.
Requirements:
- Google Cloud OAuth credentials — client ID + secret, baked into
Helpers/OAuthCredentials.cs(desktop "Desktop app" OAuth client; loopback callback — no console redirect URI registration needed; creators never configure) - Local HTTP listener —
HttpListeneronhttp://localhost:PORT/oauth2/callbackto catch the redirect - Browser launch — open the authorization URL in the default browser
- Token persistence — store access/refresh tokens securely (Windows DPAPI), reload on startup
- UI state — account shown in the Start Stream dialog; "Change Account" action triggers re-auth
- Go Live gated on auth — Start Stream dialog requires sign-in to enable the Start button; scene building works without it
Tests:
- Mock token exchange response, verify channel info parsed
- Verify token refresh triggers when near expiry
- Verify credential load/save roundtrip
TASK 3 — Capture Pipeline (Scenes/Sources)
Goal: Real video preview in the center panel — the minimal source set below, composited per scene.
Status: 🔶 In progress
- ✅ Milestone 1 — webcam — MediaCapture (WinRT SDK projection) with device enumeration, CPU-first frame source, refcounted
CameraManager, picker dialog, clip shapes (Traditional + Round) + mirror, 480×270 default placement — schema v2 - ✅ Schema v3 (Ship Branch A) — multi-scene webcam (singleton
Webcam+ per-sceneWebcamSceneConfig), right-click border/context menu, static OBS-style borders, 50%-per-dimension webcam size cap, device-swap (ReleaseAllAsync) — 25 tests passing - ✅ Schema v4 — round→rect restore persisted (
WebcamSceneConfig.RectWidth/RectHeight) + one-time legacy-square 16:9 heal on load - ✅ Screen backdrop (ship task #1, schema v5) — live desktop/game capture as a permanent, non-deletable bottom layer (
Source.IsBackdrop), auto-detecting the full-screen game at launch/focus (else the primary display — never assumed monitor 0) viaWin32FullScreenDetector(now withGetDisplays()/PrimaryMonitorIndex()for the in-app display picker), content re-designated via the OSGraphicsCapturePicker("Change Capture…") or the in-app "Capture Display" submenu, refcounted/shared capture sessions inScreenCaptureManagermirroringCameraManager— 45 tests passing - ✅ Schema v6 — backdrop Live-only by policy —
Scene.HasBackdrop, enforced by scene name on every load (EnforceBackdropPolicy: Starting/BRB/Chat/Ending never carry one; the one-time v5→v6 backfill covers all four), the scene context-menu "Backdrop" checkbox is gone (policy owns the flag), preview watermark hides when the backdrop renders, capture changed toWindowsRuntimeMarshal.TryGetDataUnsafe(the CsWinRT-safe frame-read) + downscale to the 1920×1080 master + 5s-throttled error logging (was floodingstartup.logwith 5 MB of cast errors and burning CPU), round webcam no longer re-rasterizes anImageBrushevery frame (Image + EllipseGeometry clip) — the live-mode stutter fix - ✅ The five-scene catalog (
SceneCatalog) — Starting/Live/BRB/Chat/Ending is the product — work with less, never more; the (+) button only shows when a canonical scene is missing and re-adds it (its menu lists only the missing ones) — 62 tests passing - ✅ Webcam-after-session-start fix — a webcam added to a scene after the camera was already running (e.g. Chat) previously rendered a transparent container —
CameraManager.GetPreviewBitmap+ propagation inAddWebcamToActiveSceneAsync/ReacquireWebcamnow hands the running shared frames to any newly addedWebcamSceneConfig— 65 tests passing - ✅ Chat scene webcam size cap — raised from 50%-per-dimension (960×540) to half the screen AREA (~1358×764 @16:9,
MaxWebcamWidthFor/MaxWebcamHeightForkeyed by canonical name) so the viewer sees the creator better - ✅ "Add Webcam" always opens the camera picker — deleting one scene's webcam then re-adding used to resurrect the old camera when another scene still used it —
SwapWebcamIdentityAsyncnow swaps the app-wide identity if a different camera is chosen, same path as "Change Webcam…" - ✅ Webcam resource validation + first-frame proof —
MediaCaptureFrameSourcevalidates post-init (VideoDeviceId match, stream properties ≥1,reader.StartAsync()status read + throws on non-Success); subscribescapture.Failed+CameraStreamStateChanged→SourceFailedevent on the seam; fallback ladder (VideoPreview → VideoRecord).CameraManager.AcquireAsyncrequires first-frame proof (4s timeout): returns true only after a real frame arrives — silent empty box impossible.MainViewModelsubscribesCameraFailed→ redWebcamErrorchip in preview + MessageBox names suspect apps (CameraConflictProbe). 19041 SDK projection gaps:Exclusive/DeviceLostnot projected;CameraStreamState.Failedcompared by(int)2. 81 tests passing - ✅ Scenes/sources UI — add/reorder/rename, image + background overlays with move/resize/opacity/reuse
- ✅ Audio UX shipped (UI) — the bottom-bar footer is now two lines (dropped/duration moved under bitrate/fps), with the mic's sound meter + mute button + volume slider grouped CENTERED on the footer's top line, beneath the preview panel (meter: 288px, muted slate track with ruler graduations + muted yellow/red zone tints, green→yellow→red fill; mute = speaker icon → red do-not-symbol when muted, and the slider and speaker stay in sync (volume 0 ⇔ muted — sliding off flips the speaker to muted, sliding up from 0 clears it); mic volume defaults to 80%, muting zeroes the meter and restores the prior volume on unmute (which flashes the meter to the restored position ~300ms before it returns to the live level); the meter is a READ-ONLY realtime level display (fill = live level × volume — volume is a gain on ambient noise; while the slider is dragged the bar previews the slider position and bounces back to the live level on release, which is 0 with no input — clicking the meter does nothing), clicking the MIC label opens a microphone picker whose chosen source shows left-justified inside the meter bar (fill at 75% opacity so the name + ruler markings show through); slim dimensional slider — gradient track/fill, gloss-sphere thumb; the old flat pink 18px-filled one is gone), everything else on line 2 (bitrate/fps/dropped/duration/health left, quality + gear right) — the creator's only audio control, desktop/game audio is automatic (KISS rule)
- ✅ The connected YouTube account's avatar/name shows in the top bar next to Start Stream (
SyncConnectedAccount); the scenes list is content-height now (no dead space before SOURCES) - ✅ Social bar v2 (six-slot dialog, sign-in gate, real logos) — global bar layer (never a Source, no sources-list row), content-sized, centered, GREEN glow when ON, top/bottom snap-drag (default BOTTOM, persisted
SocialBarPosition; drag clamps to 0/1040, tie→bottom). Footer Social button gets a state dot (green=ON). Dialog "Social Media Site Promotion" (SocialsDialog+ViewModels/SocialsDialogViewModel, WPF-free + injectedISocialValidator/sign-in/sign-out fakes): ON/OFF bar switch (SocialsConfig.BarEnabled, schema v8), 6 fixed slots — row 1 always YouTube (signed-in → channel handle; signed-out → sign-in gate → OAuth; delete → confirm sign-out, mirrorsStopStream), row 2 free, rows 3-6 lock icons on freemium (Premium seam: all six). Validation:DetectService(URL domain / fediverse@user@domain/ bare→Website) → asyncISocialValidatoron confirm/Save; valid snaps to text + real service logo (bundled SVG path data viaLogoDataFor, Simple Icons CC0 — initials badges gone); invalid → red do-not, stays editable, Save blocked. LCR justify dropped (BarJustifyunread), per-scene toggle dropped (Scene.HasSocialBarback-compat). Post-test fixes (2026-08-12): footer label "Socials" (not "Social"); fediverse@user@domainvalidates — the full handle is the identity end-to-end (DetectService/CanonicalUrlFor/HttpSocialValidatorbuildhttps://domain/@user, no domain loss); Cancel is a hard stop —ISocialValidator.LookupAsynctakes aCancellationToken, dialog VM owns a CTS, Cancel/X/Save abort in-flight lookups (HTTP request killed, canceled continuations never touch slot state), andConfirmEditskips re-submitting identical text (LostFocus on dismiss never re-fires a lookup).HttpSocialValidatornow has real tests (fakeHttpMessageHandler). 105 tests passing. Post-test fixes (2026-08-12, round 2): fediverse@user@domainno longer shows a generic chain — it resolves to the instance's actual software via nodeinfo (/.well-known/nodeinfo→software.name;SocialService.Fediverseenum member +SocialEntry.FediverseSoftwarepersisted in a newSocialEntry.Softwarecolumn, schema migration by column-presence) and renders that software's bundled logo (LogoDataForFediverse: mastodon/peertube/pixelfed/misskey/lemmy/pleroma/firefish, generic fediverse honeycomb fallback). Dialog row-2 edit/trash icons were too dark —IconButtonstyle gainsForeground=#d0d0d0; trash overrides#e94560(app red). 112 tests passing. Post-test fixes (2026-08-12, round 3): a fediverse handle whose identity domain is itself a redirect (e.g. YunoHost default-app subdomains —@user@llamachile.tubewhere the mastodon instance lives atmastodon.llamachile.tube) now still resolves its software: nodeinfo on the identity domain is SSO-blocked, soHttpSocialValidatorfollows the bare roothttps://domain/302 to the real instance host and re-runs the nodeinfo lookup there. - ☐ Window capture — absorbed into the Screen picker (no separate source type); dedicated window-as-source work is pending
- ☐ Scene compositing — the D3DImage/MediaElement preview compositor (this task's requirement 5; the output compositor ships as TASK 4 ship step 1)
- ☐ Text source — live text ("Starting soon", "Back in 5", handle, callout)
- ☐ Chat box — YouTube live chat rendered on the stream so viewers read along in-video
- ❌ Background removal (milestone 2) — ONNX Runtime + DirectML, MediaPipe Selfie Segmentation — deliberately NOT in this build
- ☐ Alerts — Super Chat / membership / subscribe pop-ins; build after the six; the one paid feature (see Monetization in
ai.md)
The Minimal Source Set (design decision — do not expand casually)
ytLlive is YouTube-only and 90% of users are casual. OBS's long source list is off-putting; we ship the hot few and nothing esoteric. If a user needs more, they've graduated to OBS.
- Webcam — the face cam. Non-negotiable.
- Screen — the main event (game, slides, browser). One source; a picker chooses a monitor or a window. (Window capture is absorbed here — no separate source type.)
- Background — a full-canvas backdrop image. Fills the whole scene automatically, zero fiddling. Kept separate from Image on purpose: same pixels, but this one needs no positioning.
- Image — a floating graphic/logo overlay (watermark, badge, corner branding). Free-positioned.
- Text — live text ("Starting soon", "Back in 5", handle, callout). Casual streamers live on this.
- Chat box — YouTube live chat rendered on the stream so viewers read along in-video. YT-native.
- Alerts — Super Chat / membership / subscribe pop-ins. The dopamine source. The one big lift
(Super Chat event streaming + on-stream rendering/animation); build after the six. Also the one
paid feature — see Monetization in
ai.md.
Deliberately NOT supported: game capture, browser source, media playlist, VLC, color-key voodoo, MIDI.
Source memory model (design decision)
- A scene has resources. Resources can be shared across scenes.
- A resource exists exactly once in memory, no matter how many scenes use it (a logo in five scenes = one loaded bitmap).
- Every resource carries a catalog of scenes: one usage entry per scene it appears in, each entry dictating that scene's use — placement (X/Y/Width/Height), opacity, z-order, enabled, scale mode, crop.
- Usages are named
{resourceName}.{sceneName}— whatever the user named the resource, dot, the scene name:logo.starting,logo.live,myPic.brb. Not a hardcoded "logo". - A webcam in two scenes = one capture session, two catalog entries.
- Refcount by catalog size: the last usage removed → the resource is disposed and evicted.
- The resource (not a per-scene node) owns everything
IDisposable.
Scene transitions (design decision)
Scene switching while live must never stutter. Supported types, most → least economical:
- Cut — instant switch. The default. Zero cost.
- Fade — short crossfade (~300ms).
- Move — a simple, economical move transition, done to perfection and memory-efficient. The smart streamer's bread and butter.
- Custom (media) transitions — require media elements (video/stinger playback during the transition). Heavier, but creators pay for these, so we support them. Their media follows the same resource memory model: loaded once, catalogued by scene.
Preview shows the transition too (WYSIWYG). No wipes/slides/LUTs beyond the four above.
Requirements:
- Screen — Windows.Graphics.Capture (WinRT), enumerate displays/windows, picker
- Webcam — MediaCapture (WinRT SDK projection) with device enumeration — ✅ milestone 1 done:
- TFM bumped to
net8.0-windows10.0.19041.0(app and tests) so the WinRT projection resolves from the SDK reference packs — no NuGet package, no capability manifest (unpackaged desktop app) MediaCaptureFrameSource(CPU-first:MemoryPreference = Cpu, BGRA8 viaCreateFrameReaderAsync),MediaCaptureCameraEnumerator(DeviceInformation.FindAllAsync(DeviceClass.VideoCapture))CameraManager: refcounted byDeviceId, one sharedWriteableBitmapapp-wide, dispatcher-coalesced UI updates (~render rate, latest-frame drop), placeholder/AppLog+ warning on failureCameraPickerDialog(mirror ofReuseImageDialog) — "Searching for cameras…" / list / "No cameras found" states- One webcam app-wide: Add → Webcam greyed out once one exists ("it's already in your stream" tooltip); persisted
DeviceIdre-acquires after layout load - Default placement 16:9 480×270, bottom-right, 32px margin; drag/resize/selection shared with Image sources
- Clip shapes: Traditional + Round (phone view dropped — the 9:16 phone output is the vertical output-crop tier); mirror; both persisted in the layout DB (schema v2) and toggled from the source chip
- Background removal = milestone 2 (ONNX Runtime + DirectML, MediaPipe Selfie Segmentation) — not in this build
- TFM bumped to
- Background / Image / Text — static sources positioned/scaled/opacity
- Chat box — rendered from the live chat poll (right panel is the same feed, raw)
- Scene compositing — per-scene source layering (z-order = sources list order, top-to-bottom back-to-front), preview rendered via D3DImage or MediaElement
- Branding flash — the topmost full-frame "made with ytLlive!" layer at ~25% opacity, ~1s on /
300s off (see Monetization in
ai.md), gated onBrandFlashEnabled+ live/recording. Lives in the preview compositor now (BrandFlashLayerinMainWindow.xamlCanvasGrid, driven byBrandFlashActive/BrandFlashTimerinMainViewModel); the encoder output renders the same layer, and v0.2 local recordings carry it too - Drag/drop placement & reorder — intuitive, visual (per design principle):
- Preview: click-drag a source in the center panel to reposition it; resize via handles
- Scenes list: drag rows to reorder scenes
- Sources list: drag rows to reorder sources (this is the z-order) — implemented
TASK 4 — RTMP Ingest to YouTube
Goal: Push encoded video to YouTube's RTMP ingest.
Status: 🔶 In progress
- ✅ Ship step 1 — the output compositor SHIPPED (2026-08-10)
- ✅ Ship step 2 — the FFmpeg locator SHIPPED (2026-08-10)
- ☐ Encoder + RTMP push — the FFmpeg subprocess: frames via stdin, stderr health parsing, FLV mux + push to the cached reusable stream's ingestion URL
- ☐ WASAPI audio capture — loopback (desktop/game) + the picked mic feeding
AudioLevelso the realtime meter comes alive (req 7) - ☐ Frame-pipeline wiring —
CameraManager/ScreenCaptureManager→ compositor resolver → encoder - ☐ Health stats — bitrate, FPS, dropped frames reported live in the bottom bar (req 5)
- ☐ One-click go live + private-only enforcement — Go Live always creates/updates the broadcast with
privacyStatus = "private"+ PRIVATE badge (req 8, test-verifiable)
The pipeline chain the encoder needs doesn't exist yet: scene compositing (the master 1920×1080 frame
without the preview's editing chrome) → audio capture (WASAPI, feeds the meter) → H.264+AAC encode
→ vertical-tier crop/scale → RTMP push → health stats into the bottom bar. Nothing can encode
until a frame source exists, so the compositor is ship step 1. The pipeline is
CameraManager + ScreenCaptureManager → compositor resolver → compositor → encoder → RTMP.
Requirements:
- Encoding — H.264 (hardware via NVENC/AMD, fallback x264) + AAC audio; must comply: keyframes ≤ 4s (gopSizeLong), closed GOP, AAC/MP3 @ 44.1/48kHz, mono/stereo only. License posture (decided): GPL-free build — NVENC (NVIDIA) / QSV (Intel) / AMF (AMD) + OpenH264 software fallback + built-in AAC; no libx264 (GPL contaminates a paid product). Output containers are identical either way (H.264+AAC in
.flvfor RTMP,.mp4/.tsfor VOD) — the format is NOT the differentiator, the license and per-GPU quality are. License guardrails (never violate — seeai.md→ "Licensing — do not violate"): only BtbNlgpl/lgpl-sharedbuilds; never GPL (gyan.dev) ornonfree(fdk-aac); never static for distribution (LGPL §6 relink material); never link FFmpeg into the app; never dropTHIRD-PARTY-NOTICES.txtfrom the app/About screen. - RTMP push — FFmpeg subprocess (decided): app feeds raw frames via stdin, parses stderr for health; one battle-tested binary does encode + FLV mux + push + reconnect. Binary distribution (decided): check-then-pull — probe
where ffmpeg/PATH at first go-live; if absent, download a pinned build (BtbN LGPL-shared win64 zip, ~75 MB — gyan.dev's builds are GPLv3 and ship libx264, which violates the license posture; BtbN's LGPL variant drops x264/x265 while keeping NVENC/QSV/AMF + libopenh264 + native AAC) to%APPDATA%\ytLlive\tools\ffmpeg.exe(extractffmpeg.exeplus thelibav*.dllfamily) and cache it, offline-friendly. Behind anIFfmpegLocatorseam so tests fake it (ship step 2, below). Push goes to the cached reusable stream's ingestion URL - Quality ladder — the offered tiers, with 1080p60 @ 8 Mbps as the standard/default:
- 720p30 @ 6 Mbps
- 720p60 @ 6 Mbps
- 1080p30 @ 8 Mbps
- 1080p60 @ 8 Mbps (default — mainstream ceiling, GPU hardware-encoded so the gaming machine never notices; upload headroom stays comfortable)
- Vertical 1080×1920 @ 60fps @ 8 Mbps (9:16 phone tier)
The composition master is always 1920×1080; a tier is an output rect + target resolution
(see
ai.md"Resolution tiers"). Vertical output = the centered 607×1080 crop of the master scaled to 1080×1920 (semi-crop preview is already implemented; the encoder applies the same rect). 1080p60 is the ceiling by design — "if you want 1440 or 4K or 8K → OBS is your solution"; the app targets the most mainstream creator, not power users. Ladder is sculpted by a cached probe (IP-only TCP vs public ingest host; no auth required). Quality is greyed out while live because the declared resolution can't change mid-stream — but withvariable, we can auto step-down bitrate/resolution on the fly with zero API calls (no stream recreation); 60fps presumes a hardware encoder — no hardware encoder → auto fallback to 720p60/1080p30
- Stream key management — reuse the cached reusable stream (one per channel) instead of creating a new one per go-live; prefill default YouTube ingest URL
rtmp://a.rtmp.youtube.com/live2 - Health stats — bitrate, FPS, dropped frames reported live in the bottom bar (encoder-side)
- One-click go live — defaults that work out of the box
- Audio capture (feeds the meter — this task ships the wiring) — WASAPI loopback (desktop/game at unity, zero UI — "it just is") + the picked mic (
MicSourceNamefrom theMicPickerDialog). The mic capture feedsAudioLevelso the realtime meter comes alive (today it reads 0 — the mixer feed is pending, seeai.mdaudio notes). AAC mono/stereo @ 48 kHz per the compliance rules. - Private-only go live until v1 (reputation guard, decided 2026-08-10) — until the v1 release, go-live is locked to private streams only so a software error can never publish something public/unlisted that damages the creator's reputation. RTMP push itself has no privacy — privacy lives on the YouTube live broadcast object, which this app already controls via its OAuth API calls. So the lock is purely API-side: the Go Live flow always creates/updates the broadcast with
privacyStatus = "private"and a guard refuses to set anything else (same spirit as the Live-only backdrop policy). The UI shows a clear "PRIVATE" badge next to the stream state so the creator always knows who can see them. Enforcement must be verifiable in the auth-service tests (fake the broadcast-insert/update call, assertprivacyStatusis forced to private). - v1 release gate: bundle the full license texts (decided 2026-08-10) —
THIRD-PARTY-NOTICES.txtcurrently links the canonical license texts rather than embedding them. At the v1 (GA) release, the full texts of every license it names (LGPL v2.1+, BSD-2-Clause, MIT, Apache-2.0) MUST be bundled alongside it (shipped in the app output, e.g. alicenses/folder next to the notices file, still reachable from the About screen). This is a release blocker for v1, not a task to queue early — do it in the release pass. The repo should treat this like the private-only go-live gate: a checkbox that cannot silently lapse.
Ship step 1 — Scene compositor (the frame source)
Goal: a pure-CPU software compositor producing the encoder's master frame (BGRA8, the VideoFrame
seam) from the scene model. The preview stays XAML (the editing view); the compositor is the output
view — WPF's RenderTargetBitmap can't be used (software-rendered + captures chrome). Two renderers
must agree, so the XAML (MainWindow.xaml CanvasGrid + element DataTemplate) is the contract.
Decisions (locked 2026-08-10): Path A CPU blitter — GPU effort belongs to NVENC (the encoder),
not composition; with an FFmpeg subprocess the master crosses a CPU readback to the pipe every frame
anyway, so GPU compositing buys ~nothing at this layer count (2-3 live layers; static layers
pre-composite once). A D3D11 compositor can replace this one later behind the same seam (the CPU
master buffer stays the contract). Render the output rect directly: compositor is constructed with
CompositorOptions {SourceRectX/Y/W/H, OutputWidth, OutputHeight}; 16:9 tiers = full 1920×1080 1:1;
vertical (9:16) = composite the centered 607×1080 crop then bilinear-upscale to 1080×1920. Reuses
MainViewModel.OutputRectX/Y/W/H (note (1920−607)/2 = 656.5 → align to integer pixels for output).
Render spec (back → front, mirror the XAML exactly):
- Backdrop — the Live scene's
IsBackdropSource (CaptureKey→ live frame),UniformToFillfull-frame (XAML's separateBackdropImagelayer; the backdrop element renders nothing — its DataTemplate Image is Collapsed for DisplayCapture). - Background — the scene's
BackgroundSource,UniformToFillfull-frame (theActiveBackgroundImagelayer, not per-element). - Elements in
Scene.Elementsorder (back→front), skipIsVisible=false. What actually renders:SourceTypeImage→ static asset,UniformToFillcover-crop into (X, Y, W, H)WebcamSceneConfig→ latest frame byDeviceId: Traditional =UniformToFillrect; Round = circle diametermin(W,H)(alpha 0 outside — true circle, not oval); mirror = horizontal flip around element center (MirrorScale); opacity = per-pixel multiply (content + border); border = stroked rect / centered circle atRoundBorderSize, widthBorderWidth, alphaBorderOpacityBackground/IsBackdrop/TextOverlayare NOT per-element (layers above; Text not shipped)
- Branding flash — pre-rendered full-frame "made with ytLlive!" at 25% alpha when live +
BrandFlashEnabled+ timer active. Passed in as aVideoFrame?(compositor core stays pure byte-math, no WPF; likely a bundled asset rather than runtime text rendering). - NOT in output (preview chrome only): SelectionOverlay, DimRects, output-rect outline, badge, placeholder.
New files (all in Services/Compositor/):
SceneCompositor.cs—Render(Scene, frameFor: Func<SceneElement, VideoFrame?>, flashFrame: VideoFrame?, CompositorOptions) → VideoFrame(output-sized). The caller'sframeForresolver maps each element to its frame (webcam → DeviceId, image → AssetId viaStaticPixelCache, backdrop → CaptureKey) — the compositor stays pure/hermetic/no WPF.CompositorOptions.cs— source-rect + output W×H.StretchMath.cs—UniformToFillcover-crop, ellipse mask, bilinear scale (pure, unit-tested).StaticPixelCache.cs— assetbyte[]→ cached BGRAVideoFrame(WPFBitmapDecoder+CopyPixels, decode once per content hash).
Test plan (Good Dog Rule — ONE integration test): SceneCompositorTests — a scene with backdrop
(solid red fake frame) + round webcam (solid green) + image (solid blue) → render 16:9 master → assert
per-layer probe pixels (corner = backdrop color, element center = webcam color, outside the round clip =
backdrop color, mirrored element swaps left/right); a vertical-tier variant asserts 1080×1920 output +
crop fidelity. Focused unit tests on StretchMath. Tests push frames directly — no capture managers
involved (they wire in a later step).
Same-PR housekeeping: fix the stale comment MainViewModel.cs:324 ("shown under the meter on line 2"
→ "shown left-justified INSIDE the meter bar" — ai.md is the authority); this task's requirements now
include the explicit audio-capture/meter wiring (#7 above).
Out of scope (later ship steps): FFmpeg locator + license posture (covered in requirements 1-2),
encoder + RTMP push, WASAPI audio capture (loopback + mic) feeding AudioLevel, wiring
CameraManager/ScreenCaptureManager into the frame pipeline, brand-flash timer wiring, health stats
(bitrate/FPS/dropped).
Built (2026-08-10): all four files shipped in Services/Compositor/, SceneElement.TryGetBorderColor
made public (shared hex parse with the compositor — no duplicated color parsing), the stale
MainViewModel.cs:324 comment corrected, and the pre-existing CS1998 in YouTubeAuthServiceTests
cleaned up — build 0 warnings. Tests: the SceneCompositorTests integration test (full-scene master
pixels, vertical tier, flash) + 4 StretchMath units — 72 passing.
Ship step 2 — FFmpeg locator (the encoder's binary)
Goal: resolve a usable ffmpeg.exe on demand (the encoder's one external dependency), never shipping
a binary in the repo. Returns an absolute path; downloads only when neither PATH nor the local cache
provides one.
Decisions (locked 2026-08-10):
- BtbN LGPL-shared win64 build — not gyan.dev (gyan's "essentials" is GPLv3 and ships libx264, which
violates requirement 1's license posture) and not the static lgpl build: LGPLv2.1 §6 wants
relinkable object files for static linking, but the shared (dynamic-DLL) variant sidesteps that —
compliance is "license text + source offer + unmodified binaries" (see
THIRD-PARTY-NOTICES.txtandai.md→ Licensing). Drops libx264/libx265 while keeping NVENC/QSV/AMF, libopenh264 (the LGPL-legal H.264 software fallback) and native AAC — exactly the requirement-1 encoder profile. - Pinned URL —
https://github.com/BtbN/FFmpeg-Builds/releases/download/autobuild-2026-08-09-13-03/ffmpeg-master-latest-win64-lgpl-shared.zip(~75 MB zip — earlier "~30 MB" estimate corrected). A dated autobuild tag is immutable; BtbN retention keeps the last 14 daily builds + each month-end build for 2 years, so a cold cache after retention expiry 404s — a logged, recoverable failure (the seam throws; the encoder step surfaces it). Once cached, the URL is never touched again. The pin is a singleconst, bumpable in one place — and must always stay on the shared variant (nevergpl,nonfree, or static; see ai.md Licensing). - Check-then-pull order — (1) PATH probe (the user's own install wins), (2) cached
%APPDATA%\ytLlive\tools\ffmpeg.exe, (3) download + extract. Extractffmpeg.exeplus thelibav*.dllfamily (the shared build's bin/ folder; Windows resolves the DLLs from the exe's own directory) into a staging dir then move into place — a crash never leaves a corrupt or partial cache. - Seam —
IFfmpegLocator.LocateAsync(CancellationToken): search dirs, tools dir, and the downloader (Func<string, CancellationToken, Task<byte[]>>) are constructor-injected with production defaults, so tests fake the network (feeding a real in-memory zip) and never touch disk outside a temp dir.
New files (all in Services/Encoder/):
IFfmpegLocator.cs— the seam.FfmpegLocator.cs— the impl (PATH probe → cache → pull+extract exe + DLLs), failures logged viaAppLog.THIRD-PARTY-NOTICES.txt(repo root) — the LGPL/BSD/MIT notices + source offer, copied to the build output and surfaced via the top-bar About button (MainWindowcode-behind, opens the file in the OS viewer).
Test plan: the hermetic integration test drives the full decision ladder against a temp tools dir and
a fake downloader returning a real in-memory zip (.../bin/ffmpeg.exe entry): PATH hit wins without
downloading, cache hit skips the network, cold cache downloads → extracts → ffmpeg.exe lands in the
tools dir, and a second call serves the cache (downloader invoked exactly once). Focused unit tests:
shared-build DLLs extract alongside the exe, empty zip throws, missing entry throws, empty download
throws, downloader failure propagates, zero-byte cache is refreshed.
Same-PR housekeeping: requirement 2's stale binary facts corrected in this plan (~30 MB → ~75 MB zip;
"gyan.dev/BtB N" → BtbN LGPL-shared only, with the why); the "never do" licensing guardrails recorded in
ai.md so the reasoning survives.
Out of scope (later ship steps): the FFmpeg subprocess encoder (frames in via stdin, stderr health parsing), RTMP push, WASAPI audio capture, the frame-pipeline wiring, health stats.
Built (2026-08-10): IFfmpegLocator + FfmpegLocator shipped in Services/Encoder/, pinned to the
lgpl-shared build autobuild-2026-08-09-13-03 (extracts ffmpeg.exe + the libav*.dll family via a
staging dir). THIRD-PARTY-NOTICES.txt (repo root) ships to the build output and is surfaced by a new
top-bar About button; the "never do" licensing guardrails are recorded in ai.md — build 0 warnings.
Tests: the hermetic FfmpegLocatorTests integration test (PATH → cache → download decision ladder with a
fake downloader serving a real in-memory zip) + edge/unit cases (shared-build DLL extraction, zero-byte
cache refresh, empty payload, missing zip entry, downloader failure) — 78 passing.
TASK 5 — YouTube Live Stream Management
Goal: Create/bind broadcasts, monitor YouTube-side stream health — the v3 way.
Status: ⏳ Not started
- ☐ Broadcast creation — title/description/privacy/scheduledStartTime via API, with the v3 flags above
- ☐ Reusable stream — create once, cache + reuse; bind to broadcast
- ☐ Health monitoring — poll
liveStreams.listhealthStatus+configurationIssues[], surface banner only on warning/error - ☐ Live chat — poll
liveChat/messages, render in right panel, support Super Chat + membership badges - ☐ Error handling — the YouTube error codes:
errorStreamInactive,invalidTransition,redundantTransition,liveStreamDeletionNotAllowed,liveStreamModificationNotAllowed,liveBroadcastBindingNotAllowed
Design decisions (v3)
- One-click go-live —
liveBroadcasts.insertwithenableAutoStart=true,enableAutoStop=true,enableMonitorStream=false,selfDeclaredMadeForKids=false,latencyPreference=low. Notransition(live)call, no testing stage, no liveStarting polling. Encoder starts → YouTube brings it live by itself. - Variable reusable stream —
cdn.resolution=variable,cdn.frameRate=variable,isReusable=true. Create once per channel, cache ingestion URL + stream name, reuse for every broadcast. Any quality tier works without recreation; auto step-down needs no API calls. - Report-by-exception — poll
liveStreams.list; banner only onhealthStatuswarning/error issues (configurationIssues[]). Bottom strip = YouTube logo + green/red connection dot (clickable → opens the dialog). - One dialog, three states —
not connected(sign-in) /connected-offline(all editable) /live(title + description + visibility editable; quality + account greyed out). Both entry points (Start Stream button + bottom strip) open it; prefilled from saved session profile. - Live edits —
liveBroadcasts.updatewith part=snippet,statusfor title/description/privacy. - End stream — stop encoder →
transition(complete), withenableAutoStopas the safety net. - Broadcast ID == Video ID — one ID to track status, health, and the auto-created VOD (
recordFromStart+enableDvr).
Requirements:
- Broadcast creation — title/description/privacy/scheduledStartTime via API, with the v3 flags above
- Reusable stream — create once, cache + reuse; bind to broadcast
- Health monitoring — poll
liveStreams.listhealthStatus+configurationIssues[], surface banner only on warning/error - Live chat — poll
liveChat/messages, render in right panel, support Super Chat + membership badges - Error handling — the YouTube error codes:
errorStreamInactive,invalidTransition,redundantTransition,liveStreamDeletionNotAllowed,liveStreamModificationNotAllowed,liveBroadcastBindingNotAllowed
TASK 6 — Layout Persistence (SQLite)
Goal: Scenes, sources, and asset bytes survive restarts; assets are always available.
Status: ✅ Done
- ✅ SQLite database (
Microsoft.Data.Sqlite) at%APPDATA%\ytLlive\ytLlive.db; schema versioned viaPRAGMA user_version(currently v6) - ✅ Assets live in the DB (BLOB keyed by SHA-256 content hash), never file paths — deleting the original file never breaks a scene
- ✅ File-model save/open — the active layout file is tracked (default is the AppData DB); Save Layout As… / Open Layout… switch the active file; auto-save writes to whatever is active
- ✅ Auto-save (invisible) — ~1.5s debounce on scene/source add/remove/reorder/rename/hide + any source transform change; flush on window close
- ✅ Startup — load the active file; seed the five canonical scenes only when the DB is empty; (+) re-adds a missing canonical scene and is hidden once all five are present; adding beyond the five is rejected
- ✅ Schema v1 → v6 — webcam columns (v2), singleton
Webcam+ per-sceneWebcamSceneConfig(v3),RectWidth/RectHeightround-to-rect restore (v4),Source.IsBackdrop+Source.CaptureKey(v5),Scene.HasBackdrop— backdrop Live-only by policy (v6, one-time backfill +EnforceBackdropPolicyon every load);WindowHandlestays in-memory (per-session); save = transactional rewrite; orphaned assets pruned
Design decisions
- SQLite database (
Microsoft.Data.Sqlite) at%APPDATA%\ytLlive\ytLlive.db; schema versioned viaPRAGMA user_version. - Assets live in the DB, not on disk —
Assettable stores image bytes (BLOB) keyed by a SHA-256 content hash (unique). Identical image content collapses to one row regardless of file name — the 1:M resource memory model, enforced by the database. No file paths; deleting the original file never breaks a scene. - File-model save/open — the active layout file is tracked (default is the AppData DB). Save Layout As… / Open Layout… switch the active file; auto-save writes to whatever is active.
- Auto-save (invisible) — ~1.5s debounce on scene add/remove/reorder/rename/hide, source add/remove/reorder, and any source transform change; flush on window close.
- Schema —
Scene(Id, Name, IsHidden, IsChatScene, HasBackdrop, SortOrder),Asset(Id, Hash, Data, PixelWidth, PixelHeight),Source(Id, SceneId FK cascade, AssetId FK, Type, Name, IsEnabled, X/Y/Width/Height/Opacity, MonitorIndex, DeviceId, ClipShape, IsMirrored, SortOrder) —user_version6 (v1 → v2 =ALTER TABLEadds the two webcam columns; v3 = singletonWebcam+ per-sceneWebcamSceneConfig; v4 =WebcamSceneConfig.RectWidth/RectHeightfor the round-to-rect restore; v5 =Source.IsBackdrop+Source.CaptureKeyfor the live-capture backdrop; v6 =Scene.HasBackdrop— the backdrop is Live-only by policy (one-time backfill turns Starting/BRB/Chat/Ending off and drops their backdrop sources;EnforceBackdropPolicyre-normalizes every load).WindowHandlestays in-memory (per-session). Save = transactional rewrite; orphaned assets pruned. - Startup — load the active file; seed the five canonical scenes
(Starting/Live/BRB/Chat/Ending,
SceneCatalog) only when the DB is empty. The (+) button re-adds a missing canonical scene and is hidden once all five are present; adding beyond the five is rejected — work with less, never more.
Backlog (future versions)
- v0.2 — Recording to local file (recordings carry the branding flash — see TASK 3 /
ai.mdMonetization) - v0.3 — Stream scheduling
- v0.4 — Multi-destination restreaming
- v0.5 — Stream clipping