Files
LlamaCasty/TASKS/task-32-stream-resilience.md
T
gramps 87509bcf99 docs: restructure TASKS.md into a catalog — one file per task in TASKS/
TASKS.md is now the index (status table, open items, research pointer).
33 files: 32 task files + 1 research facts file. The full take-saga
narrative and all design decisions are preserved verbatim; the catalog
makes the queue readable without opening every task body. Schema and
AGENTS.md updated to reflect the new layout.
2026-09-05 16:31:46 -07:00

3.3 KiB

TASK 32 — Stream resilience (queued 2026-09-01 — merges the orphaned TASK 2 resume clause into the pills/signs architecture)

Catalog: TASKS.md — status and requirements live here.

Goal: a dropped stream heals itself when physically possible, tells the truth when it isn't, and gets the creator back on air in one click. The ON-AIR sign + primary button in the top bar are the surface — they already mean "reality" (pills = intent, lights = reality, TASK 18/30); reality just isn't two-state. Resumption of a completed broadcast is explicitly out of scope (impossible via the API — creator ruling 2026-09-01): no lying in the copy.

Slices (one integration test each, in order):

  1. ☐ Blip retry — app-side encoder restart: push dies while live (ProcessFailed / stdin backpressure death) → bounded-backoff relaunch (2s/5s/10s, N attempts) against the cached constant reusable-stream ingest URL. Success = push running again while the broadcast is still live; viewers never know. Fixes the map's stale "ffmpeg does reconnect" claim (corrected 2026-09-01: reconnect flags are input-side only). Test: fake process fails once then starts → pump recovers without surfacing to the VM.
  2. ☐ Measured grace + the sign — while the push is dead, escalate the existing health poll (30s → ~5s). ON-AIR dot goes amber "RECONNECTING" with a mm:ss countdown to the observed YouTube cutoff; the grace length is a measured, pinned constant — probe it on a real private test stream, cite the observation (spin-guard rule; do NOT invent a number). Poll reports the broadcast complete → dot red "OFF AIR", primary button relabels "● Back on air" and pulses. Test: fake status provider drives the dot through live → reconnecting → cut-off.
  3. ☐ One-click Back on air — the relabeled button fires the existing go-live path with saved Broadcast.* + cached stream, dialog skipped: new broadcast, new VOD (accepted cost, stated in-product). Test: seeded saved form + fake stream service → click creates + pushes, no dialog.
  4. ☐ Crash-safe recording — fragmented MP4 on the record block (-movflags frag_keyframe+empty_moov+default_base_moof): plain MP4's trailing moov atom means a power loss kills the WHOLE file. Rename-on-stop unchanged. Closes the TASK 18 "MKV as an option" design-note remnant (fragmented MP4 chosen over MKV: single container, YouTube-upload-native — verify player/ editor tolerance, cite the research). Test: FfmpegArgs emits the movflags only on the record block.
  5. ☐ Pre-flight — before broadcast insert, one TCP probe to the ingest host:1935 (~2s timeout). Unreachable → Error toast, no insert, no ghost broadcast ("check VPN/network"). The "forgot the VPN" story, solved as prevention instead of recovery. Test: fake probe failure aborts go-live before any API call.

Design decisions

  • The sign is the status, the button is the action — no new chrome, no log lines, no spinners.
  • Honest windows: countdown = observed grace; at zero we say OFF AIR, we don't pretend to reconnect.
  • Why us > incumbents: constant reusable stream + persisted form make retry/restart cheap; OBS shows a reconnecting log line and Streamlabs silently dies.