repo for storing the second brain project
Go to file
Travis Herbranson a475893403 playlists: queue-time YouTube fan-out via yt-dlp extract_flat
Pasting a YouTube playlist URL into either entry point now expands
into one source row per video. Single-video URLs and non-YouTube URLs
keep their existing behaviour untouched.

Service module additions:
- is_youtube_playlist_url(url): strict detector. Only `/playlist?list=…`
  on a known YouTube host (youtube.com / m / music / no-www) counts.
  A `watch?v=…&list=…` URL is ambiguous (user usually pasted a single
  video that happens to sit inside a playlist) and intentionally falls
  through to single-add. To fan out, paste the canonical playlist URL.
- expand_youtube_playlist(url, *, max_items=50): yt-dlp with
  extract_flat=True, playlistend=max_items, skip_download. Builds a
  canonical https://www.youtube.com/watch?v={id} URL per entry and
  silently drops placeholders for private/removed videos.
- add_playlist(sess, *, url, domain, focus, max_items, expander=None):
  loops expansion entries through add_source so URL validation,
  source_type detection, and the UNIQUE dedupe path stay identical to
  the single-add flow. Per-entry titles win over any caller-supplied
  title (a single playlist title would be wrong for N videos). Partial
  failures don't abort the batch — failed entries are tallied with
  up-to-10 (url, reason) tuples for the flash. `expander` is an
  injection seam for tests so the suite never hits YouTube unless
  explicitly opted in.
- DEFAULT_PLAYLIST_MAX_ITEMS = 50 — shared ceiling, no throttle change.

CLI: `second-brain add <playlist-url>` auto-detects and reports
`expanded / added / duplicates / failed`. No new flag needed.

Web: POST /sources/add same detection. _playlist_flash() builds the
HTMX flash — "Queued N videos (M duplicates skipped, F failed)"
with sensible plural forms and graceful omission of zero counters.

Tests:
- 15 pure-Python detection cases (positives + negatives, including the
  ambiguous watch?v=…&list=… rule).
- 3 DB-backed add_playlist tests (with a mocked expander, so no
  network): count aggregation across new + pre-seeded duplicates,
  bad-entry tolerance, and the empty-playlist case.
- 1 opt-in live-network test gated on SECOND_BRAIN_LIVE_NETWORK_TESTS=1
  exercising expand_youtube_playlist against a real public playlist.

Live-verified end to end:
- web POST of a real 13-entry public playlist queued 13 video rows
  with titles, flash showed "Queued 13 videos".
- re-POST returned "Queued 0 videos (13 duplicates skipped)".
- watch?v=…&list=… correctly stayed a single-add.
- CLI parity confirmed against the same playlist.
2026-05-25 13:59:35 -04:00
alembic transcripts: capture segment-level output into a new JSONB column 2026-05-25 13:40:49 -04:00
config postgres migration: schema, models, embeddings, alembic 2026-05-24 22:46:48 -04:00
deploy deploy/web: containerized FastAPI UI on the homelab docker network 2026-05-25 11:48:37 -04:00
prompts project init 2026-05-22 19:08:22 -04:00
reviews docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
src/second_brain playlists: queue-time YouTube fan-out via yt-dlp extract_flat 2026-05-25 13:59:35 -04:00
tests playlists: queue-time YouTube fan-out via yt-dlp extract_flat 2026-05-25 13:59:35 -04:00
.gitignore add .gitignore 2026-05-24 21:10:53 -04:00
alembic.ini postgres migration: schema, models, embeddings, alembic 2026-05-24 22:46:48 -04:00
CLAUDE.md docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
postgres-migration-planning.md docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
pyproject.toml pipeline split: tower transcribe worker + claim queue + faster-whisper 2026-05-25 10:11:37 -04:00
README.md updates to project files 2026-05-24 22:23:59 -04:00
uv.lock pipeline split: tower transcribe worker + claim queue + faster-whisper 2026-05-25 10:11:37 -04:00

second-brain

A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).

What it does

  1. Pull — download YouTube videos or fetch web articles
  2. Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
  3. Extract — single-shot Claude call produces structured notes per source
  4. Review — web UI to accept or reject extractions
  5. Compile — wiki compiler folds accepted extractions into a living Obsidian vault

Quick start

# Install dependencies
uv sync

# Copy and edit config
cp config/settings.toml config/settings.local.toml

# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...

# Run the pipeline
uv run second-brain process

# Review in the web UI
uv run second-brain serve

# Compile to wiki
uv run second-brain compile

Domains

Extractions are tagged by domain so the right prompt template is used:

Domain Focus
development Software, programming, systems
content Content creation, video production
business Entrepreneurship, marketing, ops
homelab Self-hosted infra, networking, DevOps

Extractor backends

Two backends, selectable via [extractor].backend in config/settings.toml:

  • cli (default) — shells out to claude -p --output-format json --json-schema .... Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwaway tempfile.TemporaryDirectory and strips ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN from the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.
  • api — uses the anthropic Python SDK, requires ANTHROPIC_API_KEY. Install with uv sync --extra api (the SDK is an optional dependency).

Requirements

  • Python 3.12+
  • Either the claude CLI on PATH (default backend) or ANTHROPIC_API_KEY (api backend)
  • ffmpeg (for Whisper audio extraction)

In-flight work

  • Postgres + pgvector migration — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in postgres-migration-planning.md; nothing has been migrated yet.