Splits pull+transcribe (now tower-side, eager) from extract+embed
(stays on the dev scheduler). Three machine-coordination pieces land
together because they reference each other:
- v3 migration adds sources.claimed_by + claimed_at — observability +
stale-claim recovery columns. The actual race-safety primitive is
`SELECT ... FOR UPDATE SKIP LOCKED` in the new claim helper, so two
machines can poll the queue without doubling work.
- src/second_brain/claim.py owns the claim dance (claim_next_source,
release_claim, reap_stale_claims). Both stage gates filter by
source_type so the tower never grabs articles and the dev side never
grabs videos.
- src/second_brain/transcribe.py wraps faster-whisper (lazy-imported so
it stays out of the dev install). resolve_settings() reads env >
[whisper] block > defaults, falling back to int8 on cpu / float16 on
cuda when compute_type is unspecified. Default model large-v3.
- src/second_brain/scheduler/transcribe_worker.py is the long-running
poll loop. Reads pipeline_settings every iteration so the dashboard's
enable/window/max-items/max-video-length take effect within one
cycle. Reaps stale claims at startup. SIGTERM-clean. DB-unreachable
backs off with a log line; never crash-loops.
- adapters/youtube.py drops the torch-whisper transcribe path; pull
stays. Removes openai-whisper from the default deps and gates
faster-whisper behind a new `tower` extra (uv sync --extra tower).
- main.py: new `second-brain transcribe-worker` (--once for ad-hoc).
`process` now article-only on the pull side but still picks up
TRANSCRIBED of any source_type for the extract step.
Live-verified: migration applies clean, transcribe-worker --once
honors transcription_enabled=false gate.