Files
LlamaCasty/TASKS/task-32-stream-resilience.md
T
gramps 87509bcf99 docs: restructure TASKS.md into a catalog — one file per task in TASKS/
TASKS.md is now the index (status table, open items, research pointer).
33 files: 32 task files + 1 research facts file. The full take-saga
narrative and all design decisions are preserved verbatim; the catalog
makes the queue readable without opening every task body. Schema and
AGENTS.md updated to reflect the new layout.
2026-09-05 16:31:46 -07:00

45 lines
3.3 KiB
Markdown

# TASK 32 — Stream resilience (queued 2026-09-01 — merges the orphaned TASK 2 resume clause into the pills/signs architecture)
> Catalog: [`TASKS.md`](../TASKS.md) — status and requirements live here.
**Goal:** a dropped stream heals itself when physically possible, tells the truth when it isn't, and
gets the creator back on air in one click. The ON-AIR sign + primary button in the top bar are the
surface — they already mean "reality" (pills = intent, lights = reality, TASK 18/30); reality just
isn't two-state. Resumption of a completed broadcast is explicitly **out of scope** (impossible via
the API — creator ruling 2026-09-01): no lying in the copy.
**Slices (one integration test each, in order):**
1. ☐ **Blip retry** — app-side encoder restart: push dies while live (ProcessFailed / stdin
backpressure death) → bounded-backoff relaunch (2s/5s/10s, N attempts) against the cached
constant reusable-stream ingest URL. Success = push running again while the broadcast is still
live; viewers never know. Fixes the map's stale "ffmpeg does reconnect" claim (corrected
2026-09-01: reconnect flags are input-side only). Test: fake process fails once then starts →
pump recovers without surfacing to the VM.
2. ☐ **Measured grace + the sign** — while the push is dead, escalate the existing health poll
(30s → ~5s). ON-AIR dot goes **amber "RECONNECTING"** with a mm:ss countdown to the observed
YouTube cutoff; the grace length is a **measured, pinned constant** — probe it on a real private
test stream, cite the observation (spin-guard rule; do NOT invent a number). Poll reports the
broadcast `complete` → dot **red "OFF AIR"**, primary button relabels **"● Back on air"** and
pulses. Test: fake status provider drives the dot through live → reconnecting → cut-off.
3. ☐ **One-click Back on air** — the relabeled button fires the existing go-live path with saved
`Broadcast.*` + cached stream, dialog skipped: new broadcast, new VOD (accepted cost, stated
in-product). Test: seeded saved form + fake stream service → click creates + pushes, no dialog.
4. ☐ **Crash-safe recording** — fragmented MP4 on the record block (`-movflags
frag_keyframe+empty_moov+default_base_moof`): plain MP4's trailing moov atom means a power loss
kills the WHOLE file. Rename-on-stop unchanged. Closes the TASK 18 "MKV as an option" design-note
remnant (fragmented MP4 chosen over MKV: single container, YouTube-upload-native — verify player/
editor tolerance, cite the research). Test: `FfmpegArgs` emits the movflags only on the record block.
5. ☐ **Pre-flight** — before broadcast insert, one TCP probe to the ingest host:1935 (~2s timeout).
Unreachable → Error toast, no insert, no ghost broadcast ("check VPN/network"). The "forgot the
VPN" story, solved as prevention instead of recovery. Test: fake probe failure aborts go-live before any API call.
### Design decisions
- **The sign is the status, the button is the action** — no new chrome, no log lines, no spinners.
- **Honest windows:** countdown = observed grace; at zero we say OFF AIR, we don't pretend to reconnect.
- **Why us > incumbents:** constant reusable stream + persisted form make retry/restart cheap; OBS
shows a reconnecting log line and Streamlabs silently dies.