repo for storing the second brain project
The prior version was written at the initial Postgres migration. Since
then the pipeline has split across two machines, the web app shipped
(dashboard, queue, settings editor, add form, playlist fan-out, the
deploy/web container), the tower transcribe-worker landed with the
faster-whisper / CUDA-12 pin, the JSONB transcript_segments column
arrived, and several gotchas surfaced that are worth documenting once
so future-me doesn't relearn them.
Rewritten to be skimmable orientation, not a novel:
- Lead with the two-machine pipeline diagram + claim mechanism.
- Storage section covers schema, JSONB segments, the shared
public.embeddings table, and pipeline_settings.
- Web app section: routes, where it's deployed, brain.herbylab.dev,
no-auth caveat.
- Transcription section points at deploy/tower for the install detail
rather than duplicating it.
- Five load-bearing gotchas grouped explicitly:
(a) lovebug-can't-CREATE-SCHEMA → postgres superuser bootstrap once.
(b) Arch CUDA-13 vs ctranslate2 CUDA-12 → wheels pinned in venv +
systemd ExecStart wrapper.
(c) pg_hba needs both 172.19.0.0/16 (docker bridge) AND 10.99.0.0/24
(WG); we've lost each at least once.
(d) Both postgres port binds are load-bearing; don't consolidate.
(e) Traefik file-provider on the separate VM — no docker-label
routing in compose.yml.
- Code-level conventions (naive UTC, int PKs, re-query in session,
TemplateResponse signature, best-effort embedding, wrapper-vs-shell
LD_LIBRARY_PATH) kept as a checklist under the gotchas.
References the deploy READMEs rather than duplicating their content.
|
||
|---|---|---|
| alembic | ||
| config | ||
| deploy | ||
| prompts | ||
| reviews | ||
| src/second_brain | ||
| tests | ||
| .gitignore | ||
| alembic.ini | ||
| CLAUDE.md | ||
| postgres-migration-planning.md | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
second-brain
A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).
What it does
- Pull — download YouTube videos or fetch web articles
- Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
- Extract — single-shot Claude call produces structured notes per source
- Review — web UI to accept or reject extractions
- Compile — wiki compiler folds accepted extractions into a living Obsidian vault
Quick start
# Install dependencies
uv sync
# Copy and edit config
cp config/settings.toml config/settings.local.toml
# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...
# Run the pipeline
uv run second-brain process
# Review in the web UI
uv run second-brain serve
# Compile to wiki
uv run second-brain compile
Domains
Extractions are tagged by domain so the right prompt template is used:
| Domain | Focus |
|---|---|
development |
Software, programming, systems |
content |
Content creation, video production |
business |
Entrepreneurship, marketing, ops |
homelab |
Self-hosted infra, networking, DevOps |
Extractor backends
Two backends, selectable via [extractor].backend in config/settings.toml:
cli(default) — shells out toclaude -p --output-format json --json-schema .... Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwawaytempfile.TemporaryDirectoryand stripsANTHROPIC_API_KEY/ANTHROPIC_AUTH_TOKENfrom the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.api— uses theanthropicPython SDK, requiresANTHROPIC_API_KEY. Install withuv sync --extra api(the SDK is an optional dependency).
Requirements
- Python 3.12+
- Either the
claudeCLI on PATH (default backend) orANTHROPIC_API_KEY(api backend) - ffmpeg (for Whisper audio extraction)
In-flight work
- Postgres + pgvector migration — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in
postgres-migration-planning.md; nothing has been migrated yet.