Splits pull+transcribe (now tower-side, eager) from extract+embed
(stays on the dev scheduler). Three machine-coordination pieces land
together because they reference each other:
- v3 migration adds sources.claimed_by + claimed_at — observability +
stale-claim recovery columns. The actual race-safety primitive is
`SELECT ... FOR UPDATE SKIP LOCKED` in the new claim helper, so two
machines can poll the queue without doubling work.
- src/second_brain/claim.py owns the claim dance (claim_next_source,
release_claim, reap_stale_claims). Both stage gates filter by
source_type so the tower never grabs articles and the dev side never
grabs videos.
- src/second_brain/transcribe.py wraps faster-whisper (lazy-imported so
it stays out of the dev install). resolve_settings() reads env >
[whisper] block > defaults, falling back to int8 on cpu / float16 on
cuda when compute_type is unspecified. Default model large-v3.
- src/second_brain/scheduler/transcribe_worker.py is the long-running
poll loop. Reads pipeline_settings every iteration so the dashboard's
enable/window/max-items/max-video-length take effect within one
cycle. Reaps stale claims at startup. SIGTERM-clean. DB-unreachable
backs off with a log line; never crash-loops.
- adapters/youtube.py drops the torch-whisper transcribe path; pull
stays. Removes openai-whisper from the default deps and gates
faster-whisper behind a new `tower` extra (uv sync --extra tower).
- main.py: new `second-brain transcribe-worker` (--once for ad-hoc).
`process` now article-only on the pull side but still picks up
TRANSCRIBED of any source_type for the extract step.
Live-verified: migration applies clean, transcribe-worker --once
honors transcription_enabled=false gate.
Swap the SQLite backing store for petalbrain Postgres + pgvector, modeled
on vault-mcp. All second-brain relational tables now live in the
`second_brain` schema (owned by the lovebug role); embeddings are written
to the shared public.embeddings table.
Locked design decisions (per Travis):
- DB: existing petalbrain Postgres, second_brain schema, lovebug role.
- Connection: containerized homelab-postgres:5432, plain psycopg_pool
(min=1/max=10), no PgBouncer.
- ORM stays SQLAlchemy; int autoincrement PKs + naive UTC DateTime.
- Embeddings: reuse shared public.embeddings keyed by
(source_schema='second_brain', source_table='extractions', source_id,
model='nomic-embed-text'). Summaries only for this round.
- Pipeline: chunk_text → Ollama nomic-embed-text → delete-before-insert
upsert, with graceful degradation (no DB / no Ollama → log + skip).
- Alembic stands up second-brain's own schema; public.embeddings stays
out-of-band.
- File-based wiki compiler is unchanged.
No SQLite data import — starting clean.
This commit is the scaffolding only; `alembic upgrade head` and a smoke
test of the embedding path are the next checkpoint.
Run extraction under the Max OAuth subscription via `claude -p` instead
of the per-token Anthropic API. The new src/second_brain/llm/claude_cli.py
spawns the CLI in a hermetic tempdir so the host project's CLAUDE.md,
hooks, MCP config, and settings don't leak into the prompt. Uses
--json-schema with LLMExtraction.model_json_schema() so the CLI guarantees
valid structured output — replaces the brittle markdown-fence stripping
in the old engine. The Anthropic SDK is preserved as an optional "api"
backend selectable via config.
While in here, fix a handful of blockers that the smoke test surfaced:
- scheduler filtered ANALYZED instead of TRANSCRIBED, so it never
actually advanced any sources
- process command read sources in a closed session, raising
DetachedInstanceError before any work happened
- config.prompts_dir walked one parent too many, resolving outside
the project and forcing the fallback prompt for every domain
- compiler called git rev-parse against a vault that was never
git-init'd; now auto-inits with an empty seed commit and skips empty
commits cleanly
- datetime.utcnow() deprecated in 3.12+ — single utcnow() helper in
models.py keeps naive UTC semantics so no DB migration is needed
- sess.query(...).get() deprecated in SA 2.x → sess.get(...)
- dead `import anthropic` removed from compiler
Smoke test (article → process → accept → compile) succeeds end-to-end
with ANTHROPIC_API_KEY unset. a-review run saved under reviews/.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>