Commit Graph

6 Commits

Author SHA1 Message Date
Travis Herbranson
d055d1798d transcripts: capture segment-level output into a new JSONB column
Tower worker now persists faster-whisper's segment-level output
(start/end/text + word-level timing when available) alongside the
existing joined `transcript_text`. The text column stays the canonical
input the extractor reads — this is additive.

Changes:

- alembic v4: sources.transcript_segments JSONB NULL. JSONB rather than
  JSON so future equality/containment queries are indexable without a
  re-migration. Same lovebug-no-CREATE-on-petalbrain guard as prior
  migrations.

- ORM model: Optional[list] mapped to JSONB (postgresql dialect).

- transcribe.py:
  - Always pass word_timestamps=True to faster-whisper.transcribe.
  - New segment_to_dict() flattens the upstream NamedTuple-shaped
    Segment/Word into JSON-safe plain dicts so the JSONB write doesn't
    drag faster-whisper into any reader.
  - Per-word defensive conversion: a single malformed word can't drop
    the surrounding segment.

- transcribe_worker._advance: after a successful transcribe, persist
  segments into source.transcript_segments inside a try/except. If the
  JSONB write fails (oversize row, malformed dict, etc.) we log a
  warning and still commit transcript_text + status=TRANSCRIBED — the
  pipeline never crashes over the additive index.

- Tests: three new unit tests against fake Segment/Word objects cover
  the happy path (word entries serialise), the no-words case
  (`segment.words is None` → empty list), and the malformed-word skip.
  json.dumps(d) asserts JSONB-binding compatibility.

Live-verified: migration applied clean against petalbrain (`\d sources`
shows transcript_segments jsonb); ORM round-trip writes and reads the
sample payload identically. GPU large-v3 word-timestamp behaviour is
unchanged from upstream — only the tower can validate that hot path.
2026-05-25 13:40:49 -04:00
Travis Herbranson
f9feb0b393 gauntlet: fix lint — drop unused subprocess/shutil/tempfile/Path imports
Leftover ffmpeg-era imports from the earlier draft of the CPU transcribe
smoke test; the final numpy-array-bypass version doesn't need them.
Ruff autofix.
2026-05-25 10:34:54 -04:00
Travis Herbranson
913d4f0335 deploy/tower + claim race + transcribe smoke tests
deploy/tower/:
- second-brain-transcribe.service — systemd unit. User=herbyadmin,
  Type=simple, After=/Wants= wg-quick@wg-lan.service so the WG tunnel
  must come up first. Restart=always with a StartLimitBurst guard.
- second-brain-transcribe.env.example — env file template documenting
  the SECOND_BRAIN_DATABASE_URL form for db.wg.herbylab.dev (10.99.0.1)
  and the optional WHISPER_* overrides.
- README.md — EndeavourOS install steps (nvidia/cuda/cudnn, ffmpeg, uv
  + tower extra, model pre-warm), WG topology reference, validation
  checklist for what to confirm once the tunnel is live, and a
  follow-ups section flagging the local-disk → NAS media migration as
  out-of-scope-for-this-round.

Tests:
- tests/test_claim.py — live-DB race test. Two threads call
  claim_next_source against a single PULLED video row; SKIP LOCKED
  must give exactly one of them the row, the other gets None. Also
  asserts the claimed_by/at columns land + release nulls them.
  Auto-skips when no SECOND_BRAIN_DATABASE_URL is set.
- tests/test_transcribe.py — pure-Python coverage of resolve_settings
  (cpu→int8, cuda→float16, env-over-block) and write_srt; plus a CPU
  smoke test that synthesizes a numpy audio array and runs the `tiny`
  model on cpu/int8 (auto-skipped when faster-whisper isn't installed,
  i.e. on the dev side without --extra tower).
2026-05-25 10:20:46 -04:00
Travis Herbranson
f44ac2ff7c test: smoke test reads embedding model from config, not hardcoded literal
a-review (gemini, full migration diff) flagged that the smoke test had
the embedding model name baked into the SQL count query, decoupling it
from the application's configured value. Read it back from
`config.embeddings.model` (env override still wins) so the test stays
valid if the default ever moves off nomic-embed-text.
2026-05-24 22:58:30 -04:00
Travis Herbranson
2e07d7a8fe alembic: skip CREATE SCHEMA when present + add embedding smoke test
The runtime role (lovebug) doesn't have CREATE on the petalbrain database
even though it owns the second_brain schema, so a bare
`CREATE SCHEMA IF NOT EXISTS` errors out with permission denied. Gate
the bootstrap on a pg_namespace lookup so we only attempt the create
when the schema is genuinely missing — operators bootstrap it once as
postgres superuser, alembic just respects it afterward.

The smoke test exercises the full Postgres + embedding path against a
live DB + Ollama (autoskipped otherwise): writes an extraction, embeds
the summary, asserts public.embeddings has the expected row count, and
re-embeds to verify the delete-before-insert idempotency.
2026-05-24 22:49:06 -04:00
Travis Herbranson
250ca9fd2f project init 2026-05-22 19:08:22 -04:00