repo for storing the second brain project
Go to file
Travis Herbranson 93f71ebb14 docs: refresh CLAUDE.md to current shipped state
The prior version was written at the initial Postgres migration. Since
then the pipeline has split across two machines, the web app shipped
(dashboard, queue, settings editor, add form, playlist fan-out, the
deploy/web container), the tower transcribe-worker landed with the
faster-whisper / CUDA-12 pin, the JSONB transcript_segments column
arrived, and several gotchas surfaced that are worth documenting once
so future-me doesn't relearn them.

Rewritten to be skimmable orientation, not a novel:

- Lead with the two-machine pipeline diagram + claim mechanism.
- Storage section covers schema, JSONB segments, the shared
  public.embeddings table, and pipeline_settings.
- Web app section: routes, where it's deployed, brain.herbylab.dev,
  no-auth caveat.
- Transcription section points at deploy/tower for the install detail
  rather than duplicating it.
- Five load-bearing gotchas grouped explicitly:
  (a) lovebug-can't-CREATE-SCHEMA → postgres superuser bootstrap once.
  (b) Arch CUDA-13 vs ctranslate2 CUDA-12 → wheels pinned in venv +
      systemd ExecStart wrapper.
  (c) pg_hba needs both 172.19.0.0/16 (docker bridge) AND 10.99.0.0/24
      (WG); we've lost each at least once.
  (d) Both postgres port binds are load-bearing; don't consolidate.
  (e) Traefik file-provider on the separate VM — no docker-label
      routing in compose.yml.
- Code-level conventions (naive UTC, int PKs, re-query in session,
  TemplateResponse signature, best-effort embedding, wrapper-vs-shell
  LD_LIBRARY_PATH) kept as a checklist under the gotchas.

References the deploy READMEs rather than duplicating their content.
2026-05-25 15:46:20 -04:00
alembic transcripts: capture segment-level output into a new JSONB column 2026-05-25 13:40:49 -04:00
config postgres migration: schema, models, embeddings, alembic 2026-05-24 22:46:48 -04:00
deploy Merge branch 'claude/laughing-shirley-52f4a6': queue-page add form + tower CUDA-12 pin 2026-05-25 15:43:47 -04:00
prompts project init 2026-05-22 19:08:22 -04:00
reviews docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
src/second_brain web: add-to-queue form on /queue too, with per-page HTMX dispatch 2026-05-25 14:21:02 -04:00
tests playlists: queue-time YouTube fan-out via yt-dlp extract_flat 2026-05-25 13:59:35 -04:00
.gitignore add .gitignore 2026-05-24 21:10:53 -04:00
alembic.ini postgres migration: schema, models, embeddings, alembic 2026-05-24 22:46:48 -04:00
CLAUDE.md docs: refresh CLAUDE.md to current shipped state 2026-05-25 15:46:20 -04:00
postgres-migration-planning.md docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
pyproject.toml tower: pin CUDA-12 wheels + LD_LIBRARY_PATH wrapper for the systemd service 2026-05-25 15:39:13 -04:00
README.md updates to project files 2026-05-24 22:23:59 -04:00
uv.lock tower: pin CUDA-12 wheels + LD_LIBRARY_PATH wrapper for the systemd service 2026-05-25 15:39:13 -04:00

second-brain

A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).

What it does

  1. Pull — download YouTube videos or fetch web articles
  2. Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
  3. Extract — single-shot Claude call produces structured notes per source
  4. Review — web UI to accept or reject extractions
  5. Compile — wiki compiler folds accepted extractions into a living Obsidian vault

Quick start

# Install dependencies
uv sync

# Copy and edit config
cp config/settings.toml config/settings.local.toml

# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...

# Run the pipeline
uv run second-brain process

# Review in the web UI
uv run second-brain serve

# Compile to wiki
uv run second-brain compile

Domains

Extractions are tagged by domain so the right prompt template is used:

Domain Focus
development Software, programming, systems
content Content creation, video production
business Entrepreneurship, marketing, ops
homelab Self-hosted infra, networking, DevOps

Extractor backends

Two backends, selectable via [extractor].backend in config/settings.toml:

  • cli (default) — shells out to claude -p --output-format json --json-schema .... Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwaway tempfile.TemporaryDirectory and strips ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN from the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.
  • api — uses the anthropic Python SDK, requires ANTHROPIC_API_KEY. Install with uv sync --extra api (the SDK is an optional dependency).

Requirements

  • Python 3.12+
  • Either the claude CLI on PATH (default backend) or ANTHROPIC_API_KEY (api backend)
  • ffmpeg (for Whisper audio extraction)

In-flight work

  • Postgres + pgvector migration — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in postgres-migration-planning.md; nothing has been migrated yet.