repo for storing the second brain project
Run extraction under the Max OAuth subscription via `claude -p` instead of the per-token Anthropic API. The new src/second_brain/llm/claude_cli.py spawns the CLI in a hermetic tempdir so the host project's CLAUDE.md, hooks, MCP config, and settings don't leak into the prompt. Uses --json-schema with LLMExtraction.model_json_schema() so the CLI guarantees valid structured output — replaces the brittle markdown-fence stripping in the old engine. The Anthropic SDK is preserved as an optional "api" backend selectable via config. While in here, fix a handful of blockers that the smoke test surfaced: - scheduler filtered ANALYZED instead of TRANSCRIBED, so it never actually advanced any sources - process command read sources in a closed session, raising DetachedInstanceError before any work happened - config.prompts_dir walked one parent too many, resolving outside the project and forcing the fallback prompt for every domain - compiler called git rev-parse against a vault that was never git-init'd; now auto-inits with an empty seed commit and skips empty commits cleanly - datetime.utcnow() deprecated in 3.12+ — single utcnow() helper in models.py keeps naive UTC semantics so no DB migration is needed - sess.query(...).get() deprecated in SA 2.x → sess.get(...) - dead `import anthropic` removed from compiler Smoke test (article → process → accept → compile) succeeds end-to-end with ANTHROPIC_API_KEY unset. a-review run saved under reviews/. Co-Authored-By: Claude Opus 4 <noreply@anthropic.com> |
||
|---|---|---|
| config | ||
| prompts | ||
| reviews | ||
| src/second_brain | ||
| tests | ||
| CLAUDE.md | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
second-brain
A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).
What it does
- Pull — download YouTube videos or fetch web articles
- Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
- Extract — single-shot Claude call produces structured notes per source
- Review — web UI to accept or reject extractions
- Compile — wiki compiler folds accepted extractions into a living Obsidian vault
Quick start
# Install dependencies
uv sync
# Copy and edit config
cp config/settings.toml config/settings.local.toml
# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...
# Run the pipeline
uv run second-brain process
# Review in the web UI
uv run second-brain serve
# Compile to wiki
uv run second-brain compile
Domains
Extractions are tagged by domain so the right prompt template is used:
| Domain | Focus |
|---|---|
development |
Software, programming, systems |
content |
Content creation, video production |
business |
Entrepreneurship, marketing, ops |
homelab |
Self-hosted infra, networking, DevOps |
Requirements
- Python 3.12+
ANTHROPIC_API_KEYenvironment variable- ffmpeg (for Whisper audio extraction)