Pasting a YouTube playlist URL into either entry point now expands
into one source row per video. Single-video URLs and non-YouTube URLs
keep their existing behaviour untouched.
Service module additions:
- is_youtube_playlist_url(url): strict detector. Only `/playlist?list=…`
on a known YouTube host (youtube.com / m / music / no-www) counts.
A `watch?v=…&list=…` URL is ambiguous (user usually pasted a single
video that happens to sit inside a playlist) and intentionally falls
through to single-add. To fan out, paste the canonical playlist URL.
- expand_youtube_playlist(url, *, max_items=50): yt-dlp with
extract_flat=True, playlistend=max_items, skip_download. Builds a
canonical https://www.youtube.com/watch?v={id} URL per entry and
silently drops placeholders for private/removed videos.
- add_playlist(sess, *, url, domain, focus, max_items, expander=None):
loops expansion entries through add_source so URL validation,
source_type detection, and the UNIQUE dedupe path stay identical to
the single-add flow. Per-entry titles win over any caller-supplied
title (a single playlist title would be wrong for N videos). Partial
failures don't abort the batch — failed entries are tallied with
up-to-10 (url, reason) tuples for the flash. `expander` is an
injection seam for tests so the suite never hits YouTube unless
explicitly opted in.
- DEFAULT_PLAYLIST_MAX_ITEMS = 50 — shared ceiling, no throttle change.
CLI: `second-brain add <playlist-url>` auto-detects and reports
`expanded / added / duplicates / failed`. No new flag needed.
Web: POST /sources/add same detection. _playlist_flash() builds the
HTMX flash — "Queued N videos (M duplicates skipped, F failed)"
with sensible plural forms and graceful omission of zero counters.
Tests:
- 15 pure-Python detection cases (positives + negatives, including the
ambiguous watch?v=…&list=… rule).
- 3 DB-backed add_playlist tests (with a mocked expander, so no
network): count aggregation across new + pre-seeded duplicates,
bad-entry tolerance, and the empty-playlist case.
- 1 opt-in live-network test gated on SECOND_BRAIN_LIVE_NETWORK_TESTS=1
exercising expand_youtube_playlist against a real public playlist.
Live-verified end to end:
- web POST of a real 13-entry public playlist queued 13 video rows
with titles, flash showed "Queued 13 videos".
- re-POST returned "Queued 0 videos (13 duplicates skipped)".
- watch?v=…&list=… correctly stayed a single-add.
- CLI parity confirmed against the same playlist.
Extracts the `second-brain add` CLI's queueing logic into a service
module (sources_service.add_source) so the CLI and the new web POST
route share the same validation, dedupe, and source_type heuristic —
no behaviour drift between the two entry points.
UI:
- The dashboard body (Pipeline counts + Settings snapshot + Recent
activity) moves into _dashboard_body.html, wrapped in
`<div id="dashboard-body">`. HTMX targets that id for swap.
- A new "Add to queue" section sits at the top of the partial: URL
(required, type=url), Domain (select matching DOMAINS), Title and
Focus (both optional, match the CLI flags). hx-post=/sources/add,
hx-target=#dashboard-body, hx-swap=outerHTML — same pattern as the
settings save form.
Route:
- POST /sources/add calls sources_service.add_source inside a single
transaction, then re-renders _dashboard_body.html so the pipeline
counts and recent-activity list update in place. Flash slots above
the form report:
✓ Queued <type>: <url> on a fresh add
✓ Already queued (id=…, status=…) on a dupe (matches CLI text)
✗ <validation message> on bad URL / unknown domain
CLI:
- `second-brain add` now delegates to the same service. Field values
for the success echo are captured inside the session block so a
post-commit detached-instance access can't fail. Bad input exits 2
with a clear stderr message instead of raising.
Live-verified through the web container on herbys-dev: dashboard
renders the form, happy add lands a row, dup-detect matches the CLI
phrasing, two validation errors surface as red flashes, swap target
id survives across swaps, no tracebacks in uvicorn logs.
Splits pull+transcribe (now tower-side, eager) from extract+embed
(stays on the dev scheduler). Three machine-coordination pieces land
together because they reference each other:
- v3 migration adds sources.claimed_by + claimed_at — observability +
stale-claim recovery columns. The actual race-safety primitive is
`SELECT ... FOR UPDATE SKIP LOCKED` in the new claim helper, so two
machines can poll the queue without doubling work.
- src/second_brain/claim.py owns the claim dance (claim_next_source,
release_claim, reap_stale_claims). Both stage gates filter by
source_type so the tower never grabs articles and the dev side never
grabs videos.
- src/second_brain/transcribe.py wraps faster-whisper (lazy-imported so
it stays out of the dev install). resolve_settings() reads env >
[whisper] block > defaults, falling back to int8 on cpu / float16 on
cuda when compute_type is unspecified. Default model large-v3.
- src/second_brain/scheduler/transcribe_worker.py is the long-running
poll loop. Reads pipeline_settings every iteration so the dashboard's
enable/window/max-items/max-video-length take effect within one
cycle. Reaps stale claims at startup. SIGTERM-clean. DB-unreachable
backs off with a log line; never crash-loops.
- adapters/youtube.py drops the torch-whisper transcribe path; pull
stays. Removes openai-whisper from the default deps and gates
faster-whisper behind a new `tower` extra (uv sync --extra tower).
- main.py: new `second-brain transcribe-worker` (--once for ad-hoc).
`process` now article-only on the pull side but still picks up
TRANSCRIBED of any source_type for the extract step.
Live-verified: migration applies clean, transcribe-worker --once
honors transcription_enabled=false gate.
Ruff autofix removed pre-existing F401 unused imports across adapters,
context, compiler, extractor, scheduler, web, and main. Also reordered
models.py so its SQLAlchemy imports sit at the top of the file
(E402 was triggered by the inline `utcnow` helper definition).
Added a [dependency-groups] dev block (pytest, pytest-cov, ruff) so the
zero-check test gauntlet — which runs `pytest --cov --cov-report=term-missing`
— can resolve its plugins without a manual `uv pip install pytest-cov`.
Swap the SQLite backing store for petalbrain Postgres + pgvector, modeled
on vault-mcp. All second-brain relational tables now live in the
`second_brain` schema (owned by the lovebug role); embeddings are written
to the shared public.embeddings table.
Locked design decisions (per Travis):
- DB: existing petalbrain Postgres, second_brain schema, lovebug role.
- Connection: containerized homelab-postgres:5432, plain psycopg_pool
(min=1/max=10), no PgBouncer.
- ORM stays SQLAlchemy; int autoincrement PKs + naive UTC DateTime.
- Embeddings: reuse shared public.embeddings keyed by
(source_schema='second_brain', source_table='extractions', source_id,
model='nomic-embed-text'). Summaries only for this round.
- Pipeline: chunk_text → Ollama nomic-embed-text → delete-before-insert
upsert, with graceful degradation (no DB / no Ollama → log + skip).
- Alembic stands up second-brain's own schema; public.embeddings stays
out-of-band.
- File-based wiki compiler is unchanged.
No SQLite data import — starting clean.
This commit is the scaffolding only; `alembic upgrade head` and a smoke
test of the embedding path are the next checkpoint.
Run extraction under the Max OAuth subscription via `claude -p` instead
of the per-token Anthropic API. The new src/second_brain/llm/claude_cli.py
spawns the CLI in a hermetic tempdir so the host project's CLAUDE.md,
hooks, MCP config, and settings don't leak into the prompt. Uses
--json-schema with LLMExtraction.model_json_schema() so the CLI guarantees
valid structured output — replaces the brittle markdown-fence stripping
in the old engine. The Anthropic SDK is preserved as an optional "api"
backend selectable via config.
While in here, fix a handful of blockers that the smoke test surfaced:
- scheduler filtered ANALYZED instead of TRANSCRIBED, so it never
actually advanced any sources
- process command read sources in a closed session, raising
DetachedInstanceError before any work happened
- config.prompts_dir walked one parent too many, resolving outside
the project and forcing the fallback prompt for every domain
- compiler called git rev-parse against a vault that was never
git-init'd; now auto-inits with an empty seed commit and skips empty
commits cleanly
- datetime.utcnow() deprecated in 3.12+ — single utcnow() helper in
models.py keeps naive UTC semantics so no DB migration is needed
- sess.query(...).get() deprecated in SA 2.x → sess.get(...)
- dead `import anthropic` removed from compiler
Smoke test (article → process → accept → compile) succeeds end-to-end
with ANTHROPIC_API_KEY unset. a-review run saved under reviews/.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>