repo for storing the second brain project
a-review (gemini, full migration diff) flagged that the smoke test had the embedding model name baked into the SQL count query, decoupling it from the application's configured value. Read it back from `config.embeddings.model` (env override still wins) so the test stays valid if the default ever moves off nomic-embed-text. |
||
|---|---|---|
| alembic | ||
| config | ||
| prompts | ||
| reviews | ||
| src/second_brain | ||
| tests | ||
| .gitignore | ||
| alembic.ini | ||
| CLAUDE.md | ||
| postgres-migration-planning.md | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
second-brain
A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).
What it does
- Pull — download YouTube videos or fetch web articles
- Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
- Extract — single-shot Claude call produces structured notes per source
- Review — web UI to accept or reject extractions
- Compile — wiki compiler folds accepted extractions into a living Obsidian vault
Quick start
# Install dependencies
uv sync
# Copy and edit config
cp config/settings.toml config/settings.local.toml
# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...
# Run the pipeline
uv run second-brain process
# Review in the web UI
uv run second-brain serve
# Compile to wiki
uv run second-brain compile
Domains
Extractions are tagged by domain so the right prompt template is used:
| Domain | Focus |
|---|---|
development |
Software, programming, systems |
content |
Content creation, video production |
business |
Entrepreneurship, marketing, ops |
homelab |
Self-hosted infra, networking, DevOps |
Extractor backends
Two backends, selectable via [extractor].backend in config/settings.toml:
cli(default) — shells out toclaude -p --output-format json --json-schema .... Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwawaytempfile.TemporaryDirectoryand stripsANTHROPIC_API_KEY/ANTHROPIC_AUTH_TOKENfrom the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.api— uses theanthropicPython SDK, requiresANTHROPIC_API_KEY. Install withuv sync --extra api(the SDK is an optional dependency).
Requirements
- Python 3.12+
- Either the
claudeCLI on PATH (default backend) orANTHROPIC_API_KEY(api backend) - ffmpeg (for Whisper audio extraction)
In-flight work
- Postgres + pgvector migration — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in
postgres-migration-planning.md; nothing has been migrated yet.