second-brain/README.md
2026-05-24 22:23:59 -04:00

62 lines
2.3 KiB
Markdown

# second-brain
A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).
## What it does
1. **Pull** — download YouTube videos or fetch web articles
2. **Transcribe** — extract transcripts via Whisper (videos) or trafilatura (articles)
3. **Extract** — single-shot Claude call produces structured notes per source
4. **Review** — web UI to accept or reject extractions
5. **Compile** — wiki compiler folds accepted extractions into a living Obsidian vault
## Quick start
```bash
# Install dependencies
uv sync
# Copy and edit config
cp config/settings.toml config/settings.local.toml
# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...
# Run the pipeline
uv run second-brain process
# Review in the web UI
uv run second-brain serve
# Compile to wiki
uv run second-brain compile
```
## Domains
Extractions are tagged by domain so the right prompt template is used:
| Domain | Focus |
|--------|-------|
| `development` | Software, programming, systems |
| `content` | Content creation, video production |
| `business` | Entrepreneurship, marketing, ops |
| `homelab` | Self-hosted infra, networking, DevOps |
## Extractor backends
Two backends, selectable via `[extractor].backend` in `config/settings.toml`:
- **`cli`** (default) — shells out to `claude -p --output-format json --json-schema ...`. Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwaway `tempfile.TemporaryDirectory` and strips `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN` from the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.
- **`api`** — uses the `anthropic` Python SDK, requires `ANTHROPIC_API_KEY`. Install with `uv sync --extra api` (the SDK is an optional dependency).
## Requirements
- Python 3.12+
- Either the `claude` CLI on PATH (default backend) **or** `ANTHROPIC_API_KEY` (api backend)
- ffmpeg (for Whisper audio extraction)
## In-flight work
- **Postgres + pgvector migration** — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in `postgres-migration-planning.md`; nothing has been migrated yet.