repo for storing the second brain project
Go to file
Travis Herbranson 597c83649c deploy/web: containerized FastAPI UI on the homelab docker network
Dockerfile, compose.yml, env template, and runbook for the second-brain
web container. Targets herbys-dev (10.0.21.207); Traefik (file-provider
on 10.0.11.20) reaches the published host port at 10.0.21.207:8080.

Image:
- python:3.12-slim + the official uv binary copied from the upstream
  image, plus apt-installed git + ca-certificates so the Gitea VCS pin
  for embedding-chunking resolves at build time.
- uv sync --frozen --no-dev --no-install-project, source copy, then a
  second uv sync to install the project itself. Two-step so the lock
  install layer caches independently of source edits.
- No ffmpeg / claude CLI / faster-whisper — web role doesn't need any
  of them. Extraction runs on the dev host's CLI; transcription on the
  tower.
- Drops to uid 1000 (`app`) before CMD. Uvicorn binds 0.0.0.0:8000
  inside the container, with --proxy-headers + --forwarded-allow-ips=*
  so Traefik's X-Forwarded-* survive.

compose.yml:
- Joins the existing external `homelab` bridge network so the container
  reaches homelab-postgres:5432 and ollama:11434 by service DNS.
- Publishes the uvicorn port at 10.0.21.207:8080 (LAN IP bound, not
  0.0.0.0) for Traefik on the separate VM to reach. NO traefik.* labels
  — file-provider Traefik can't read them.
- env_file: .env (0600, gitignored) — SECOND_BRAIN_DATABASE_URL points
  at homelab-postgres:5432 (containerised), NOT the host's 127.0.0.1:5433
  port-map.
- restart: unless-stopped.

README.md:
- Build + bring-up commands.
- SECURITY note: no SSO / no CSRF / mutating endpoints — Traefik route
  must be LAN-only for now (Travis's call).
- The two infra steps Travis applies himself, with ready-to-paste
  snippets:
  - Knot DNS: brain.herbylab.dev → 10.0.11.20.
  - Traefik dynamic config: file-provider router + service block.
- Verification checklist for both before-and-after-DNS states.

Live-verified on herbys-dev: container Up, dashboard returns 200 with
real status counts from petalbrain, settings save round-trip works,
no tracebacks in logs.
2026-05-25 11:48:37 -04:00
alembic pipeline split: tower transcribe worker + claim queue + faster-whisper 2026-05-25 10:11:37 -04:00
config postgres migration: schema, models, embeddings, alembic 2026-05-24 22:46:48 -04:00
deploy deploy/web: containerized FastAPI UI on the homelab docker network 2026-05-25 11:48:37 -04:00
prompts project init 2026-05-22 19:08:22 -04:00
reviews docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
src/second_brain pipeline split: tower transcribe worker + claim queue + faster-whisper 2026-05-25 10:11:37 -04:00
tests gauntlet: fix lint — drop unused subprocess/shutil/tempfile/Path imports 2026-05-25 10:34:54 -04:00
.gitignore add .gitignore 2026-05-24 21:10:53 -04:00
alembic.ini postgres migration: schema, models, embeddings, alembic 2026-05-24 22:46:48 -04:00
CLAUDE.md docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
postgres-migration-planning.md docs: update CLAUDE.md + migration plan to reflect shipped Postgres design 2026-05-24 22:57:31 -04:00
pyproject.toml pipeline split: tower transcribe worker + claim queue + faster-whisper 2026-05-25 10:11:37 -04:00
README.md updates to project files 2026-05-24 22:23:59 -04:00
uv.lock pipeline split: tower transcribe worker + claim queue + faster-whisper 2026-05-25 10:11:37 -04:00

second-brain

A personal knowledge system built on a video/article extraction pipeline and an LLM-maintained wiki (inspired by the Karpathy LLM Wiki pattern).

What it does

  1. Pull — download YouTube videos or fetch web articles
  2. Transcribe — extract transcripts via Whisper (videos) or trafilatura (articles)
  3. Extract — single-shot Claude call produces structured notes per source
  4. Review — web UI to accept or reject extractions
  5. Compile — wiki compiler folds accepted extractions into a living Obsidian vault

Quick start

# Install dependencies
uv sync

# Copy and edit config
cp config/settings.toml config/settings.local.toml

# Queue a source
uv run second-brain add https://www.youtube.com/watch?v=...

# Run the pipeline
uv run second-brain process

# Review in the web UI
uv run second-brain serve

# Compile to wiki
uv run second-brain compile

Domains

Extractions are tagged by domain so the right prompt template is used:

Domain Focus
development Software, programming, systems
content Content creation, video production
business Entrepreneurship, marketing, ops
homelab Self-hosted infra, networking, DevOps

Extractor backends

Two backends, selectable via [extractor].backend in config/settings.toml:

  • cli (default) — shells out to claude -p --output-format json --json-schema .... Runs under your Max OAuth subscription, no per-token billing. The wrapper spawns the CLI in a throwaway tempfile.TemporaryDirectory and strips ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN from the env so host CLAUDE.md, hooks, and settings don't leak in, and the subscription is used instead of the API key.
  • api — uses the anthropic Python SDK, requires ANTHROPIC_API_KEY. Install with uv sync --extra api (the SDK is an optional dependency).

Requirements

  • Python 3.12+
  • Either the claude CLI on PATH (default backend) or ANTHROPIC_API_KEY (api backend)
  • ffmpeg (for Whisper audio extraction)

In-flight work

  • Postgres + pgvector migration — SQLite is the current backing store but the project is moving to Travis's Postgres instance. Planning questions live in postgres-migration-planning.md; nothing has been migrated yet.