Pasting a YouTube playlist URL into either entry point now expands
into one source row per video. Single-video URLs and non-YouTube URLs
keep their existing behaviour untouched.
Service module additions:
- is_youtube_playlist_url(url): strict detector. Only `/playlist?list=…`
on a known YouTube host (youtube.com / m / music / no-www) counts.
A `watch?v=…&list=…` URL is ambiguous (user usually pasted a single
video that happens to sit inside a playlist) and intentionally falls
through to single-add. To fan out, paste the canonical playlist URL.
- expand_youtube_playlist(url, *, max_items=50): yt-dlp with
extract_flat=True, playlistend=max_items, skip_download. Builds a
canonical https://www.youtube.com/watch?v={id} URL per entry and
silently drops placeholders for private/removed videos.
- add_playlist(sess, *, url, domain, focus, max_items, expander=None):
loops expansion entries through add_source so URL validation,
source_type detection, and the UNIQUE dedupe path stay identical to
the single-add flow. Per-entry titles win over any caller-supplied
title (a single playlist title would be wrong for N videos). Partial
failures don't abort the batch — failed entries are tallied with
up-to-10 (url, reason) tuples for the flash. `expander` is an
injection seam for tests so the suite never hits YouTube unless
explicitly opted in.
- DEFAULT_PLAYLIST_MAX_ITEMS = 50 — shared ceiling, no throttle change.
CLI: `second-brain add <playlist-url>` auto-detects and reports
`expanded / added / duplicates / failed`. No new flag needed.
Web: POST /sources/add same detection. _playlist_flash() builds the
HTMX flash — "Queued N videos (M duplicates skipped, F failed)"
with sensible plural forms and graceful omission of zero counters.
Tests:
- 15 pure-Python detection cases (positives + negatives, including the
ambiguous watch?v=…&list=… rule).
- 3 DB-backed add_playlist tests (with a mocked expander, so no
network): count aggregation across new + pre-seeded duplicates,
bad-entry tolerance, and the empty-playlist case.
- 1 opt-in live-network test gated on SECOND_BRAIN_LIVE_NETWORK_TESTS=1
exercising expand_youtube_playlist against a real public playlist.
Live-verified end to end:
- web POST of a real 13-entry public playlist queued 13 video rows
with titles, flash showed "Queued 13 videos".
- re-POST returned "Queued 0 videos (13 duplicates skipped)".
- watch?v=…&list=… correctly stayed a single-add.
- CLI parity confirmed against the same playlist.
Tower worker now persists faster-whisper's segment-level output
(start/end/text + word-level timing when available) alongside the
existing joined `transcript_text`. The text column stays the canonical
input the extractor reads — this is additive.
Changes:
- alembic v4: sources.transcript_segments JSONB NULL. JSONB rather than
JSON so future equality/containment queries are indexable without a
re-migration. Same lovebug-no-CREATE-on-petalbrain guard as prior
migrations.
- ORM model: Optional[list] mapped to JSONB (postgresql dialect).
- transcribe.py:
- Always pass word_timestamps=True to faster-whisper.transcribe.
- New segment_to_dict() flattens the upstream NamedTuple-shaped
Segment/Word into JSON-safe plain dicts so the JSONB write doesn't
drag faster-whisper into any reader.
- Per-word defensive conversion: a single malformed word can't drop
the surrounding segment.
- transcribe_worker._advance: after a successful transcribe, persist
segments into source.transcript_segments inside a try/except. If the
JSONB write fails (oversize row, malformed dict, etc.) we log a
warning and still commit transcript_text + status=TRANSCRIBED — the
pipeline never crashes over the additive index.
- Tests: three new unit tests against fake Segment/Word objects cover
the happy path (word entries serialise), the no-words case
(`segment.words is None` → empty list), and the malformed-word skip.
json.dumps(d) asserts JSONB-binding compatibility.
Live-verified: migration applied clean against petalbrain (`\d sources`
shows transcript_segments jsonb); ORM round-trip writes and reads the
sample payload identically. GPU large-v3 word-timestamp behaviour is
unchanged from upstream — only the tower can validate that hot path.
Leftover ffmpeg-era imports from the earlier draft of the CPU transcribe
smoke test; the final numpy-array-bypass version doesn't need them.
Ruff autofix.
deploy/tower/:
- second-brain-transcribe.service — systemd unit. User=herbyadmin,
Type=simple, After=/Wants= wg-quick@wg-lan.service so the WG tunnel
must come up first. Restart=always with a StartLimitBurst guard.
- second-brain-transcribe.env.example — env file template documenting
the SECOND_BRAIN_DATABASE_URL form for db.wg.herbylab.dev (10.99.0.1)
and the optional WHISPER_* overrides.
- README.md — EndeavourOS install steps (nvidia/cuda/cudnn, ffmpeg, uv
+ tower extra, model pre-warm), WG topology reference, validation
checklist for what to confirm once the tunnel is live, and a
follow-ups section flagging the local-disk → NAS media migration as
out-of-scope-for-this-round.
Tests:
- tests/test_claim.py — live-DB race test. Two threads call
claim_next_source against a single PULLED video row; SKIP LOCKED
must give exactly one of them the row, the other gets None. Also
asserts the claimed_by/at columns land + release nulls them.
Auto-skips when no SECOND_BRAIN_DATABASE_URL is set.
- tests/test_transcribe.py — pure-Python coverage of resolve_settings
(cpu→int8, cuda→float16, env-over-block) and write_srt; plus a CPU
smoke test that synthesizes a numpy audio array and runs the `tiny`
model on cpu/int8 (auto-skipped when faster-whisper isn't installed,
i.e. on the dev side without --extra tower).
a-review (gemini, full migration diff) flagged that the smoke test had
the embedding model name baked into the SQL count query, decoupling it
from the application's configured value. Read it back from
`config.embeddings.model` (env override still wins) so the test stays
valid if the default ever moves off nomic-embed-text.
The runtime role (lovebug) doesn't have CREATE on the petalbrain database
even though it owns the second_brain schema, so a bare
`CREATE SCHEMA IF NOT EXISTS` errors out with permission denied. Gate
the bootstrap on a pg_namespace lookup so we only attempt the create
when the schema is genuinely missing — operators bootstrap it once as
postgres superuser, alembic just respects it afterward.
The smoke test exercises the full Postgres + embedding path against a
live DB + Ollama (autoskipped otherwise): writes an extraction, embeds
the summary, asserts public.embeddings has the expected row count, and
re-embeds to verify the delete-before-insert idempotency.