The prior version was written at the initial Postgres migration. Since
then the pipeline has split across two machines, the web app shipped
(dashboard, queue, settings editor, add form, playlist fan-out, the
deploy/web container), the tower transcribe-worker landed with the
faster-whisper / CUDA-12 pin, the JSONB transcript_segments column
arrived, and several gotchas surfaced that are worth documenting once
so future-me doesn't relearn them.
Rewritten to be skimmable orientation, not a novel:
- Lead with the two-machine pipeline diagram + claim mechanism.
- Storage section covers schema, JSONB segments, the shared
public.embeddings table, and pipeline_settings.
- Web app section: routes, where it's deployed, brain.herbylab.dev,
no-auth caveat.
- Transcription section points at deploy/tower for the install detail
rather than duplicating it.
- Five load-bearing gotchas grouped explicitly:
(a) lovebug-can't-CREATE-SCHEMA → postgres superuser bootstrap once.
(b) Arch CUDA-13 vs ctranslate2 CUDA-12 → wheels pinned in venv +
systemd ExecStart wrapper.
(c) pg_hba needs both 172.19.0.0/16 (docker bridge) AND 10.99.0.0/24
(WG); we've lost each at least once.
(d) Both postgres port binds are load-bearing; don't consolidate.
(e) Traefik file-provider on the separate VM — no docker-label
routing in compose.yml.
- Code-level conventions (naive UTC, int PKs, re-query in session,
TemplateResponse signature, best-effort embedding, wrapper-vs-shell
LD_LIBRARY_PATH) kept as a checklist under the gotchas.
References the deploy READMEs rather than duplicating their content.
15 KiB
second-brain
Personal knowledge pipeline. Two machines coordinate purely through the shared Postgres queue: the dev side (herbys-dev) does article pull
- all extraction + embedding, and the tower (EndeavourOS, RTX 3080)
does video pull + transcription on the GPU. A FastAPI/Jinja/HTMX web
app fronts the queue, settings, and add-to-queue form; the file-based
Obsidian wiki is downstream of
acceptedextractions.
Two-machine pipeline
add (CLI or POST /sources/add — single video, article,
or YouTube playlist URL → fan-out into N rows)
│
▼
┌────────── second_brain.sources (the queue) ──────────┐
│ │
│ PENDING ── (article)──► dev `second-brain process`
│ trafilatura pull │
│ PENDING ── (video) ──► tower transcribe-worker │
│ yt-dlp pull → PULLED │
│ PULLED ── (video) ──► tower transcribe-worker │
│ faster-whisper │
│ both paths land at ──► TRANSCRIBED │
│ │
│ TRANSCRIBED ─────────► dev `second-brain schedule` │
│ Claude CLI extract + │
│ Ollama embed │
│ ──► ANALYZED ──► accept in web UI │
│ ──► ACCEPTED ──► `compile` writes vault│
│ ──► PUBLISHED │
│ │
│ FAILED is the terminal error state │
└───────────────────────────────────────────────────────┘
Source claiming — sources.claimed_by / claimed_at columns +
SELECT ... FOR UPDATE SKIP LOCKED keep the dev and tower workers from
grabbing the same row. Both stages filter by source_type:
- tower transcribe-worker:
status IN (PENDING, PULLED) AND source_type = 'video' - dev
process(article pull):status = PENDING AND source_type = 'article' - dev
schedule(extract+embed):status = TRANSCRIBED(any type)
A 30-minute stale-claim reaper runs at worker startup so a crashed machine doesn't dam the queue.
CLI commands
second-brain add <url>— queue a source. YouTubeplaylist?list=…URLs auto-expand into one row per video (cap 50, thewatch?v=…&list=…form stays a single video — seeis_youtube_playlist_urlfor the rule).second-brain process— dev-side: pull pending articles, extract- embed any TRANSCRIBED source. Skips videos (tower's job).
second-brain schedule— dev-side overnight extract+embed loop, honouringpipeline_settings.extraction_enabled/ active-hours / max-items.second-brain transcribe-worker— tower-side daemon.--oncefor ad-hoc runs. Honourspipeline_settings.transcription_*knobs.second-brain serve— local dev launch of the web UI (port 8000). Production runs as a container (seedeploy/web/).second-brain compile— wiki compiler: ACCEPTED → markdown invault_path, git-committed. File-based Obsidian vault only; we do not dual-write into petalbrain.wiki.
Storage
Postgres: petalbrain DB on homelab-postgres (pgvector/pg17, docker
container). Connection role: lovebug. Dev reaches it at
127.0.0.1:5433; tower reaches it over WireGuard at
db.wg.herbylab.dev:5432; the web container reaches it at
homelab-postgres:5432 on the homelab docker bridge. Keep both port
binds — the host-side 5433 publish and the docker-net 5432 service.
Schema second_brain (Alembic-managed in alembic/):
sources— the queue. Int SERIAL PK. Naive UTCtimestamp(without time zone) — nevertimestamptz.transcript_textis the joined full text the extractor reads;transcript_segmentsis JSONB with the faster-whisper per-segment / per-word breakdown.claimed_by/claimed_atfor the queue dance.extractions— one row per source. JSON columns for the structured output (summary, key_points, entities, claims, …).pipeline_settings— single-row config (PK pinned to id=1 with a CHECK constraint). What the dashboard edits and workers read each iteration: transcription_enabled / active_hours / max_gpu_jobs / max_video_length / max_items_per_run plus the equivalent extraction_* fields. DB values override settings.toml.wiki_pages— bookkeeping for the file-based wiki compiler.
Embeddings live in public.embeddings — the shared pgvector
table owned by postgres and consumed by vault-mcp / ob1 / second-brain.
We do not manage this in Alembic. Rows are keyed by
(source_schema='second_brain', source_table='extractions', source_id, embedding_model='nomic-embed-text'),
with delete-before-insert semantics on the 4-tuple so a re-embed
never leaves orphans. We embed extraction summaries only for now —
not key_points, not transcript chunks.
Web app
FastAPI + Jinja2 + HTMX, in src/second_brain/web/. Production lives in
a slim Python 3.12 container (see deploy/web/) on the homelab docker
network, reaching Postgres + Ollama by service DNS. Published at
10.0.21.207:8080 and fronted by the Traefik VM (10.0.11.20,
file-provider — no docker-label routing) as
https://brain.herbylab.dev.
Routes:
/— sources list (filterable by domain + status)/queue— in-flight sources (PENDING / PULLED / TRANSCRIBED) plus the add-to-queue form/dashboard— status counts, recent activity, settings snapshot, plus the add-to-queue form/sources/{id}— extraction detail + accept/reject HTMX actions/settings— full HTMX editor forpipeline_settingsPOST /sources/add— shared endpoint. Detects YouTube playlist URLs and fans out via yt-dlpextract_flat. ReadsHX-Current-URLto dispatch the response partial (_queue_body.htmlvs_dashboard_body.html) so each page refreshes its own region inline.
No auth. UI mutates state (accept/reject sources, edit settings) but
has no SSO / CSRF / rate-limiting. The Traefik route is currently
LAN-only by Travis's call — see deploy/web/README.md §SECURITY.
Transcription
src/second_brain/transcribe.py wraps faster-whisper (CTranslate2
backend). Defaults: model large-v3, device cuda, compute_type
float16. CPU fallback drops to int8. word_timestamps=True is
always passed; the segment dicts persisted to
sources.transcript_segments carry per-word {start, end, word, probability}.
The tower worker (src/second_brain/scheduler/transcribe_worker.py)
runs as a systemd service — see deploy/tower/README.md for the full
install runbook. ExecStart goes through
deploy/tower/run-transcribe-worker.sh which prepends the venv's
CUDA-12 lib dirs to LD_LIBRARY_PATH (see gotcha (b) below).
Extractor backends
[extractor].backend in config/settings.toml:
cli(default) —src/second_brain/llm/claude_cli.pyshells out toclaude -p --output-format json --model … --json-schema …. Runs under the Max OAuth subscription. Hermetic spawn: throwawaytempfile.TemporaryDirectory,ANTHROPIC_API_KEY/_AUTH_TOKENstripped from env so it uses OAuth, and--no-session-persistence --disable-slash-commands --tools "" --setting-sources ""so host CLAUDE.md / hooks / agents don't leak in.api—anthropicSDK, lazy-imported. RequiresANTHROPIC_API_KEY. Install withuv sync --extra api.
Schema-validated CLI output lands at payload["structured_output"], not
payload["result"]. The wrapper checks both.
Environment variables
SECOND_BRAIN_DATABASE_URL— Postgres DSN. Falls back toHERBYLAB_DATABASE_URL. Usepostgresql+psycopg://for SQLAlchemy; the embeddings + claim modules strip the+psycopgprefix for raw psycopg.OLLAMA_URL— overrides[embeddings].ollama_url. Defaulthttp://ollama:11434(homelab docker net); host-side runs usehttp://127.0.0.1:11434.EMBEDDING_MODEL— defaultnomic-embed-text(768-d).WHISPER_MODEL/WHISPER_DEVICE/WHISPER_COMPUTE_TYPE— override the[whisper]block on the tower.ANTHROPIC_API_KEY— only forextractor.backend = "api".SECOND_BRAIN_CONFIG— settings file path override.SECOND_BRAIN_WORKER_ID— stamped intosources.claimed_by.SECOND_BRAIN_DB_POOL_MIN/_MAX— SQLAlchemy pool sizing.SECOND_BRAIN_LIVE_NETWORK_TESTS=1— opt-in for the yt-dlp playlist test that hits real YouTube.
Domains
development / content / business / homelab — drives the
per-domain prompt template loaded by the context assembler.
Key files
src/second_brain/models.py— all ORM in thesecond_brainschema.src/second_brain/sources_service.py—add_source,add_playlist,is_youtube_playlist_url,expand_youtube_playlist. CLI + web both call into here so there's no behavior drift between entry points.src/second_brain/claim.py— the SKIP LOCKED dance + stale reaper.src/second_brain/settings_store.py— single-rowpipeline_settingsget/update + the active-window math the workers share.src/second_brain/transcribe.py— faster-whisper wrapper + segment serialiser.src/second_brain/scheduler/runner.py— dev extract scheduler.src/second_brain/scheduler/transcribe_worker.py— tower poll loop.src/second_brain/embeddings/__init__.py— chunk → Ollama → upsert intopublic.embeddings; best-effort with graceful degradation.alembic/+alembic.ini— schema migrations for thesecond_brainschema only.public.embeddingsis not ours to migrate.deploy/web/— Dockerfile + compose.yml + README for the web container. Travis applies the Knot DNS A-record + Traefik file-provider router from the README §2.deploy/tower/— systemd unit +run-transcribe-worker.shwrapper + env template + EndeavourOS install runbook.prompts/<domain>.md— extraction prompt templates.config/settings.toml—[database],[extractor],[embeddings],[scheduler],[whisper].
Load-bearing gotchas
These have all bitten us; don't relearn them.
-
(a) Schema bootstrap requires postgres superuser.
lovebugdoes NOT haveCREATEon thepetalbraindatabase, even though it owns thesecond_brainschema. The first migration needs a one-timeCREATE SCHEMA second_brain AUTHORIZATION lovebug; GRANT ALL PRIVILEGES ON SCHEMA second_brain TO lovebug;aspostgres.alembic/env.pyguardsCREATE SCHEMAbehind apg_namespacelookup so subsequent runs aslovebugdon't trip the permission check. -
(b) Arch ships CUDA 13, ctranslate2 needs CUDA 12. EndeavourOS's
pacman -S cuda cudnninstallslibcublas.so.13/libcudnn.so.10, whichctranslate24.7.x cannot load — it wants the .12 / .9 ABI. We pinnvidia-cublas-cu12andnvidia-cudnn-cu12>=9,<10in thetoweroptional-dependency group, and the systemd ExecStart callsdeploy/tower/run-transcribe-worker.shwhich resolves the venv's wheel-bundled.sodirectories at runtime and prepends them toLD_LIBRARY_PATH. The wrapper uses the namespace packages'__path__(not__file__, which isNonefor PEP 420 packages). systemd does not inherit shell exports. -
(c) pg_hba needs BOTH subnets. Two host rules are load-bearing:
host all all 172.19.0.0/16 scram-sha-256 # homelab docker bridge host all all 10.99.0.0/24 scram-sha-256 # WireGuard (tower)Dropping the 172.19/16 rule locks out the web container + vault-mcp + ob1-* + trellis-mcp (all on the homelab net). Dropping the 10.99/24 rule locks out the tower. Both have been removed once each during unrelated postgres edits — the symptom is
FATAL: no pg_hba.conf entry for host "…"in service logs. -
(d) Two postgres port binds, both load-bearing.
homelab-postgrespublishes127.0.0.1:5433 → 5432for dev-host access AND is reachable onhomelab-postgres:5432over the docker net. Tower hits it over WG atdb.wg.herbylab.dev:5432. Don't consolidate; each path serves a different consumer set. -
(e) Traefik is file-provider on a separate VM (10.0.11.20). It cannot read this host's docker labels. The web
compose.ymltherefore publishes back to a fixed host port (10.0.21.207:8080) that Traefik reaches over LAN, and intentionally carries NOtraefik.*labels. Router config lives in Traefik's dynamic-config YAML — seedeploy/web/README.md§2b for the snippet. -
Code-level conventions, since they bit us early:
- Naive UTC timestamps. Use
from second_brain.models import utcnow— neverdatetime.utcnow()(deprecated) ordatetime.now()(local). Postgres columns aretimestamp without time zone. - Int autoincrement PKs. Don't migrate to bigint-identity or UUID; the existing schema shape is locked.
- Re-query inside the session. SQLAlchemy ORM attribute access
after the session closes raises
DetachedInstanceError. Web routes and CLI loops collect IDs first, then re-fetch inside the session block; dict snapshots for templates are built insidewith db.session(). - Starlette TemplateResponse signature.
templates.TemplateResponse(request, name, context)— the old(name, context)form raisesTypeError: unhashable type: 'dict'. Do not put"request": requestinto the context dict. - Embedding is best-effort. The extract path flushes the
relational
extractionsrow first, then callsembed_extraction(). DB / Ollama failures log and return None; pipeline state stays canonical in the relational + file vault. - Worker LD_LIBRARY_PATH. See (b). When pre-warming or debugging
faster-whisper manually, run via
deploy/tower/run-transcribe-worker.shor replicate the path probe in your shell.
- Naive UTC timestamps. Use
In-flight / parked
- Vault location — recommended
/opt/projects/second-brain-vault/as a peer dir to the project. Not yet confirmed. - Media + subtitles on NAS — currently land on the tower's local SSD. Move to NAS once we have a stable mount path.
- SSO — Authentik forward-auth middleware in Traefik is the obvious next step before opening the UI past LAN.
- Dev-side extraction systemd timer — currently invoked manually; no systemd unit yet by Travis's call.
Validation
Per the user's global CLAUDE.md, run zero-check after any code
change. The pipeline's at ~/zero-check-pipeline/validate/run-all.sh "$(pwd)". Run a-review after non-trivial diffs (>50 lines):
python3 ~/.claude/skills/a-review/run_review.py. Test suite is uv run pytest; live DB + Ollama paths auto-skip when their endpoints
aren't configured.