Tower worker now persists faster-whisper's segment-level output
(start/end/text + word-level timing when available) alongside the
existing joined `transcript_text`. The text column stays the canonical
input the extractor reads — this is additive.
Changes:
- alembic v4: sources.transcript_segments JSONB NULL. JSONB rather than
JSON so future equality/containment queries are indexable without a
re-migration. Same lovebug-no-CREATE-on-petalbrain guard as prior
migrations.
- ORM model: Optional[list] mapped to JSONB (postgresql dialect).
- transcribe.py:
- Always pass word_timestamps=True to faster-whisper.transcribe.
- New segment_to_dict() flattens the upstream NamedTuple-shaped
Segment/Word into JSON-safe plain dicts so the JSONB write doesn't
drag faster-whisper into any reader.
- Per-word defensive conversion: a single malformed word can't drop
the surrounding segment.
- transcribe_worker._advance: after a successful transcribe, persist
segments into source.transcript_segments inside a try/except. If the
JSONB write fails (oversize row, malformed dict, etc.) we log a
warning and still commit transcript_text + status=TRANSCRIBED — the
pipeline never crashes over the additive index.
- Tests: three new unit tests against fake Segment/Word objects cover
the happy path (word entries serialise), the no-words case
(`segment.words is None` → empty list), and the malformed-word skip.
json.dumps(d) asserts JSONB-binding compatibility.
Live-verified: migration applied clean against petalbrain (`\d sources`
shows transcript_segments jsonb); ORM round-trip writes and reads the
sample payload identically. GPU large-v3 word-timestamp behaviour is
unchanged from upstream — only the tower can validate that hot path.
Leftover ffmpeg-era imports from the earlier draft of the CPU transcribe
smoke test; the final numpy-array-bypass version doesn't need them.
Ruff autofix.
deploy/tower/:
- second-brain-transcribe.service — systemd unit. User=herbyadmin,
Type=simple, After=/Wants= wg-quick@wg-lan.service so the WG tunnel
must come up first. Restart=always with a StartLimitBurst guard.
- second-brain-transcribe.env.example — env file template documenting
the SECOND_BRAIN_DATABASE_URL form for db.wg.herbylab.dev (10.99.0.1)
and the optional WHISPER_* overrides.
- README.md — EndeavourOS install steps (nvidia/cuda/cudnn, ffmpeg, uv
+ tower extra, model pre-warm), WG topology reference, validation
checklist for what to confirm once the tunnel is live, and a
follow-ups section flagging the local-disk → NAS media migration as
out-of-scope-for-this-round.
Tests:
- tests/test_claim.py — live-DB race test. Two threads call
claim_next_source against a single PULLED video row; SKIP LOCKED
must give exactly one of them the row, the other gets None. Also
asserts the claimed_by/at columns land + release nulls them.
Auto-skips when no SECOND_BRAIN_DATABASE_URL is set.
- tests/test_transcribe.py — pure-Python coverage of resolve_settings
(cpu→int8, cuda→float16, env-over-block) and write_srt; plus a CPU
smoke test that synthesizes a numpy audio array and runs the `tiny`
model on cpu/int8 (auto-skipped when faster-whisper isn't installed,
i.e. on the dev side without --extra tower).