repo for storing the second brain project
Pasting a YouTube playlist URL into either entry point now expands into one source row per video. Single-video URLs and non-YouTube URLs keep their existing behaviour untouched. Service module additions: - is_youtube_playlist_url(url): strict detector. Only `/playlist?list=…` on a known YouTube host (youtube.com / m / music / no-www) counts. A `watch?v=…&list=…` URL is ambiguous (user usually pasted a single video that happens to sit inside a playlist) and intentionally falls through to single-add. To fan out, paste the canonical playlist URL. - expand_youtube_playlist(url, *, max_items=50): yt-dlp with extract_flat=True, playlistend=max_items, skip_download. Builds a canonical https://www.youtube.com/watch?v={id} URL per entry and silently drops placeholders for private/removed videos. - add_playlist(sess, *, url, domain, focus, max_items, expander=None): loops expansion entries through add_source so URL validation, source_type detection, and the UNIQUE dedupe path stay identical to the single-add flow. Per-entry titles win over any caller-supplied title (a single playlist title would be wrong for N videos). Partial failures don't abort the batch — failed entries are tallied with up-to-10 (url, reason) tuples for the flash. `expander` is an injection seam for tests so the suite never hits YouTube unless explicitly opted in. - DEFAULT_PLAYLIST_MAX_ITEMS = 50 — shared ceiling, no throttle change. CLI: `second-brain add <playlist-url>` auto-detects and reports `expanded / added / duplicates / failed`. No new flag needed. Web: POST /sources/add same detection. _playlist_flash() builds the HTMX flash — "Queued N videos (M duplicates skipped, F failed)" with sensible plural forms and graceful omission of zero counters. Tests: - 15 pure-Python detection cases (positives + negatives, including the ambiguous watch?v=…&list=… rule). - 3 DB-backed add_playlist tests (with a mocked expander, so no network): count aggregation across new + pre-seeded duplicates, bad-entry tolerance, and the empty-playlist case. - 1 opt-in live-network test gated on SECOND_BRAIN_LIVE_NETWORK_TESTS=1 exercising expand_youtube_playlist against a real public playlist. Live-verified end to end: - web POST of a real 13-entry public playlist queued 13 video rows with titles, flash showed "Queued 13 videos". - re-POST returned "Queued 0 videos (13 duplicates skipped)". - watch?v=…&list=… correctly stayed a single-add. - CLI parity confirmed against the same playlist. |
||
|---|---|---|
| alembic | ||
| config | ||
| deploy | ||
| prompts | ||
| reviews | ||
| src/second_brain | ||
| tests | ||
| .gitignore | ||
| alembic.ini | ||
| CLAUDE.md | ||
| postgres-migration-planning.md | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
second-brain
A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).
What it does
- Pull — download YouTube videos or fetch web articles
- Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
- Extract — single-shot Claude call produces structured notes per source
- Review — web UI to accept or reject extractions
- Compile — wiki compiler folds accepted extractions into a living Obsidian vault
Quick start
# Install dependencies
uv sync
# Copy and edit config
cp config/settings.toml config/settings.local.toml
# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...
# Run the pipeline
uv run second-brain process
# Review in the web UI
uv run second-brain serve
# Compile to wiki
uv run second-brain compile
Domains
Extractions are tagged by domain so the right prompt template is used:
| Domain | Focus |
|---|---|
development |
Software, programming, systems |
content |
Content creation, video production |
business |
Entrepreneurship, marketing, ops |
homelab |
Self-hosted infra, networking, DevOps |
Extractor backends
Two backends, selectable via [extractor].backend in config/settings.toml:
cli(default) — shells out toclaude -p --output-format json --json-schema .... Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwawaytempfile.TemporaryDirectoryand stripsANTHROPIC_API_KEY/ANTHROPIC_AUTH_TOKENfrom the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.api— uses theanthropicPython SDK, requiresANTHROPIC_API_KEY. Install withuv sync --extra api(the SDK is an optional dependency).
Requirements
- Python 3.12+
- Either the
claudeCLI on PATH (default backend) orANTHROPIC_API_KEY(api backend) - ffmpeg (for Whisper audio extraction)
In-flight work
- Postgres + pgvector migration — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in
postgres-migration-planning.md; nothing has been migrated yet.