second-brain/CLAUDE.md
Travis Herbranson 93f71ebb14 docs: refresh CLAUDE.md to current shipped state
The prior version was written at the initial Postgres migration. Since
then the pipeline has split across two machines, the web app shipped
(dashboard, queue, settings editor, add form, playlist fan-out, the
deploy/web container), the tower transcribe-worker landed with the
faster-whisper / CUDA-12 pin, the JSONB transcript_segments column
arrived, and several gotchas surfaced that are worth documenting once
so future-me doesn't relearn them.

Rewritten to be skimmable orientation, not a novel:

- Lead with the two-machine pipeline diagram + claim mechanism.
- Storage section covers schema, JSONB segments, the shared
  public.embeddings table, and pipeline_settings.
- Web app section: routes, where it's deployed, brain.herbylab.dev,
  no-auth caveat.
- Transcription section points at deploy/tower for the install detail
  rather than duplicating it.
- Five load-bearing gotchas grouped explicitly:
  (a) lovebug-can't-CREATE-SCHEMA → postgres superuser bootstrap once.
  (b) Arch CUDA-13 vs ctranslate2 CUDA-12 → wheels pinned in venv +
      systemd ExecStart wrapper.
  (c) pg_hba needs both 172.19.0.0/16 (docker bridge) AND 10.99.0.0/24
      (WG); we've lost each at least once.
  (d) Both postgres port binds are load-bearing; don't consolidate.
  (e) Traefik file-provider on the separate VM — no docker-label
      routing in compose.yml.
- Code-level conventions (naive UTC, int PKs, re-query in session,
  TemplateResponse signature, best-effort embedding, wrapper-vs-shell
  LD_LIBRARY_PATH) kept as a checklist under the gotchas.

References the deploy READMEs rather than duplicating their content.
2026-05-25 15:46:20 -04:00

15 KiB

second-brain

Personal knowledge pipeline. Two machines coordinate purely through the shared Postgres queue: the dev side (herbys-dev) does article pull

  • all extraction + embedding, and the tower (EndeavourOS, RTX 3080) does video pull + transcription on the GPU. A FastAPI/Jinja/HTMX web app fronts the queue, settings, and add-to-queue form; the file-based Obsidian wiki is downstream of accepted extractions.

Two-machine pipeline

                add (CLI or POST /sources/add — single video, article,
                     or YouTube playlist URL → fan-out into N rows)
                                  │
                                  ▼
   ┌────────── second_brain.sources (the queue) ──────────┐
   │                                                       │
   │   PENDING  ── (article)──►  dev `second-brain process`
   │                              trafilatura pull         │
   │   PENDING  ── (video)  ──►  tower transcribe-worker   │
   │                              yt-dlp pull → PULLED     │
   │   PULLED   ── (video)  ──►  tower transcribe-worker   │
   │                              faster-whisper           │
   │              both paths land at  ──►  TRANSCRIBED     │
   │                                                       │
   │   TRANSCRIBED ─────────►  dev `second-brain schedule` │
   │                              Claude CLI extract +     │
   │                              Ollama embed             │
   │              ──►  ANALYZED  ──►  accept in web UI     │
   │              ──►  ACCEPTED  ──►  `compile` writes vault│
   │              ──►  PUBLISHED                           │
   │                                                       │
   │   FAILED is the terminal error state                  │
   └───────────────────────────────────────────────────────┘

Source claimingsources.claimed_by / claimed_at columns + SELECT ... FOR UPDATE SKIP LOCKED keep the dev and tower workers from grabbing the same row. Both stages filter by source_type:

  • tower transcribe-worker: status IN (PENDING, PULLED) AND source_type = 'video'
  • dev process (article pull): status = PENDING AND source_type = 'article'
  • dev schedule (extract+embed): status = TRANSCRIBED (any type)

A 30-minute stale-claim reaper runs at worker startup so a crashed machine doesn't dam the queue.

CLI commands

  • second-brain add <url> — queue a source. YouTube playlist?list=… URLs auto-expand into one row per video (cap 50, the watch?v=…&list=… form stays a single video — see is_youtube_playlist_url for the rule).
  • second-brain process — dev-side: pull pending articles, extract
    • embed any TRANSCRIBED source. Skips videos (tower's job).
  • second-brain schedule — dev-side overnight extract+embed loop, honouring pipeline_settings.extraction_enabled / active-hours / max-items.
  • second-brain transcribe-worker — tower-side daemon. --once for ad-hoc runs. Honours pipeline_settings.transcription_* knobs.
  • second-brain serve — local dev launch of the web UI (port 8000). Production runs as a container (see deploy/web/).
  • second-brain compile — wiki compiler: ACCEPTED → markdown in vault_path, git-committed. File-based Obsidian vault only; we do not dual-write into petalbrain.wiki.

Storage

Postgres: petalbrain DB on homelab-postgres (pgvector/pg17, docker container). Connection role: lovebug. Dev reaches it at 127.0.0.1:5433; tower reaches it over WireGuard at db.wg.herbylab.dev:5432; the web container reaches it at homelab-postgres:5432 on the homelab docker bridge. Keep both port binds — the host-side 5433 publish and the docker-net 5432 service.

Schema second_brain (Alembic-managed in alembic/):

  • sources — the queue. Int SERIAL PK. Naive UTC timestamp (without time zone) — never timestamptz. transcript_text is the joined full text the extractor reads; transcript_segments is JSONB with the faster-whisper per-segment / per-word breakdown. claimed_by / claimed_at for the queue dance.
  • extractions — one row per source. JSON columns for the structured output (summary, key_points, entities, claims, …).
  • pipeline_settingssingle-row config (PK pinned to id=1 with a CHECK constraint). What the dashboard edits and workers read each iteration: transcription_enabled / active_hours / max_gpu_jobs / max_video_length / max_items_per_run plus the equivalent extraction_* fields. DB values override settings.toml.
  • wiki_pages — bookkeeping for the file-based wiki compiler.

Embeddings live in public.embeddings — the shared pgvector table owned by postgres and consumed by vault-mcp / ob1 / second-brain. We do not manage this in Alembic. Rows are keyed by (source_schema='second_brain', source_table='extractions', source_id, embedding_model='nomic-embed-text'), with delete-before-insert semantics on the 4-tuple so a re-embed never leaves orphans. We embed extraction summaries only for now — not key_points, not transcript chunks.

Web app

FastAPI + Jinja2 + HTMX, in src/second_brain/web/. Production lives in a slim Python 3.12 container (see deploy/web/) on the homelab docker network, reaching Postgres + Ollama by service DNS. Published at 10.0.21.207:8080 and fronted by the Traefik VM (10.0.11.20, file-provider — no docker-label routing) as https://brain.herbylab.dev.

Routes:

  • / — sources list (filterable by domain + status)
  • /queue — in-flight sources (PENDING / PULLED / TRANSCRIBED) plus the add-to-queue form
  • /dashboard — status counts, recent activity, settings snapshot, plus the add-to-queue form
  • /sources/{id} — extraction detail + accept/reject HTMX actions
  • /settings — full HTMX editor for pipeline_settings
  • POST /sources/add — shared endpoint. Detects YouTube playlist URLs and fans out via yt-dlp extract_flat. Reads HX-Current-URL to dispatch the response partial (_queue_body.html vs _dashboard_body.html) so each page refreshes its own region inline.

No auth. UI mutates state (accept/reject sources, edit settings) but has no SSO / CSRF / rate-limiting. The Traefik route is currently LAN-only by Travis's call — see deploy/web/README.md §SECURITY.

Transcription

src/second_brain/transcribe.py wraps faster-whisper (CTranslate2 backend). Defaults: model large-v3, device cuda, compute_type float16. CPU fallback drops to int8. word_timestamps=True is always passed; the segment dicts persisted to sources.transcript_segments carry per-word {start, end, word, probability}.

The tower worker (src/second_brain/scheduler/transcribe_worker.py) runs as a systemd service — see deploy/tower/README.md for the full install runbook. ExecStart goes through deploy/tower/run-transcribe-worker.sh which prepends the venv's CUDA-12 lib dirs to LD_LIBRARY_PATH (see gotcha (b) below).

Extractor backends

[extractor].backend in config/settings.toml:

  • cli (default) — src/second_brain/llm/claude_cli.py shells out to claude -p --output-format json --model … --json-schema …. Runs under the Max OAuth subscription. Hermetic spawn: throwaway tempfile.TemporaryDirectory, ANTHROPIC_API_KEY / _AUTH_TOKEN stripped from env so it uses OAuth, and --no-session-persistence --disable-slash-commands --tools "" --setting-sources "" so host CLAUDE.md / hooks / agents don't leak in.
  • apianthropic SDK, lazy-imported. Requires ANTHROPIC_API_KEY. Install with uv sync --extra api.

Schema-validated CLI output lands at payload["structured_output"], not payload["result"]. The wrapper checks both.

Environment variables

  • SECOND_BRAIN_DATABASE_URL — Postgres DSN. Falls back to HERBYLAB_DATABASE_URL. Use postgresql+psycopg:// for SQLAlchemy; the embeddings + claim modules strip the +psycopg prefix for raw psycopg.
  • OLLAMA_URL — overrides [embeddings].ollama_url. Default http://ollama:11434 (homelab docker net); host-side runs use http://127.0.0.1:11434.
  • EMBEDDING_MODEL — default nomic-embed-text (768-d).
  • WHISPER_MODEL / WHISPER_DEVICE / WHISPER_COMPUTE_TYPE — override the [whisper] block on the tower.
  • ANTHROPIC_API_KEY — only for extractor.backend = "api".
  • SECOND_BRAIN_CONFIG — settings file path override.
  • SECOND_BRAIN_WORKER_ID — stamped into sources.claimed_by.
  • SECOND_BRAIN_DB_POOL_MIN / _MAX — SQLAlchemy pool sizing.
  • SECOND_BRAIN_LIVE_NETWORK_TESTS=1 — opt-in for the yt-dlp playlist test that hits real YouTube.

Domains

development / content / business / homelab — drives the per-domain prompt template loaded by the context assembler.

Key files

  • src/second_brain/models.py — all ORM in the second_brain schema.
  • src/second_brain/sources_service.pyadd_source, add_playlist, is_youtube_playlist_url, expand_youtube_playlist. CLI + web both call into here so there's no behavior drift between entry points.
  • src/second_brain/claim.py — the SKIP LOCKED dance + stale reaper.
  • src/second_brain/settings_store.py — single-row pipeline_settings get/update + the active-window math the workers share.
  • src/second_brain/transcribe.py — faster-whisper wrapper + segment serialiser.
  • src/second_brain/scheduler/runner.py — dev extract scheduler.
  • src/second_brain/scheduler/transcribe_worker.py — tower poll loop.
  • src/second_brain/embeddings/__init__.py — chunk → Ollama → upsert into public.embeddings; best-effort with graceful degradation.
  • alembic/ + alembic.ini — schema migrations for the second_brain schema only. public.embeddings is not ours to migrate.
  • deploy/web/ — Dockerfile + compose.yml + README for the web container. Travis applies the Knot DNS A-record + Traefik file-provider router from the README §2.
  • deploy/tower/ — systemd unit + run-transcribe-worker.sh wrapper + env template + EndeavourOS install runbook.
  • prompts/<domain>.md — extraction prompt templates.
  • config/settings.toml[database], [extractor], [embeddings], [scheduler], [whisper].

Load-bearing gotchas

These have all bitten us; don't relearn them.

  • (a) Schema bootstrap requires postgres superuser. lovebug does NOT have CREATE on the petalbrain database, even though it owns the second_brain schema. The first migration needs a one-time CREATE SCHEMA second_brain AUTHORIZATION lovebug; GRANT ALL PRIVILEGES ON SCHEMA second_brain TO lovebug; as postgres. alembic/env.py guards CREATE SCHEMA behind a pg_namespace lookup so subsequent runs as lovebug don't trip the permission check.

  • (b) Arch ships CUDA 13, ctranslate2 needs CUDA 12. EndeavourOS's pacman -S cuda cudnn installs libcublas.so.13 / libcudnn.so.10, which ctranslate2 4.7.x cannot load — it wants the .12 / .9 ABI. We pin nvidia-cublas-cu12 and nvidia-cudnn-cu12>=9,<10 in the tower optional-dependency group, and the systemd ExecStart calls deploy/tower/run-transcribe-worker.sh which resolves the venv's wheel-bundled .so directories at runtime and prepends them to LD_LIBRARY_PATH. The wrapper uses the namespace packages' __path__ (not __file__, which is None for PEP 420 packages). systemd does not inherit shell exports.

  • (c) pg_hba needs BOTH subnets. Two host rules are load-bearing:

    host all all 172.19.0.0/16  scram-sha-256   # homelab docker bridge
    host all all 10.99.0.0/24   scram-sha-256   # WireGuard (tower)
    

    Dropping the 172.19/16 rule locks out the web container + vault-mcp + ob1-* + trellis-mcp (all on the homelab net). Dropping the 10.99/24 rule locks out the tower. Both have been removed once each during unrelated postgres edits — the symptom is FATAL: no pg_hba.conf entry for host "…" in service logs.

  • (d) Two postgres port binds, both load-bearing. homelab-postgres publishes 127.0.0.1:5433 → 5432 for dev-host access AND is reachable on homelab-postgres:5432 over the docker net. Tower hits it over WG at db.wg.herbylab.dev:5432. Don't consolidate; each path serves a different consumer set.

  • (e) Traefik is file-provider on a separate VM (10.0.11.20). It cannot read this host's docker labels. The web compose.yml therefore publishes back to a fixed host port (10.0.21.207:8080) that Traefik reaches over LAN, and intentionally carries NO traefik.* labels. Router config lives in Traefik's dynamic-config YAML — see deploy/web/README.md §2b for the snippet.

  • Code-level conventions, since they bit us early:

    • Naive UTC timestamps. Use from second_brain.models import utcnow — never datetime.utcnow() (deprecated) or datetime.now() (local). Postgres columns are timestamp without time zone.
    • Int autoincrement PKs. Don't migrate to bigint-identity or UUID; the existing schema shape is locked.
    • Re-query inside the session. SQLAlchemy ORM attribute access after the session closes raises DetachedInstanceError. Web routes and CLI loops collect IDs first, then re-fetch inside the session block; dict snapshots for templates are built inside with db.session().
    • Starlette TemplateResponse signature. templates.TemplateResponse(request, name, context) — the old (name, context) form raises TypeError: unhashable type: 'dict'. Do not put "request": request into the context dict.
    • Embedding is best-effort. The extract path flushes the relational extractions row first, then calls embed_extraction(). DB / Ollama failures log and return None; pipeline state stays canonical in the relational + file vault.
    • Worker LD_LIBRARY_PATH. See (b). When pre-warming or debugging faster-whisper manually, run via deploy/tower/run-transcribe-worker.sh or replicate the path probe in your shell.

In-flight / parked

  • Vault location — recommended /opt/projects/second-brain-vault/ as a peer dir to the project. Not yet confirmed.
  • Media + subtitles on NAS — currently land on the tower's local SSD. Move to NAS once we have a stable mount path.
  • SSO — Authentik forward-auth middleware in Traefik is the obvious next step before opening the UI past LAN.
  • Dev-side extraction systemd timer — currently invoked manually; no systemd unit yet by Travis's call.

Validation

Per the user's global CLAUDE.md, run zero-check after any code change. The pipeline's at ~/zero-check-pipeline/validate/run-all.sh "$(pwd)". Run a-review after non-trivial diffs (>50 lines): python3 ~/.claude/skills/a-review/run_review.py. Test suite is uv run pytest; live DB + Ollama paths auto-skip when their endpoints aren't configured.