Gate was source_type == 'video', but the service-layer reset already
accepts any source with transcript_text — articles store the
trafilatura-extracted body there, so they're equally valid candidates
for an extract-only retry. Caught by a-review.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds per-record edit mode and retry to the source detail page, plus a
CLI `second-brain retry <id>` for parity.
Edit mode exposes domain (select), title, focus on /sources/{id}.
URL stays read-only — it's the UNIQUE dedupe key and the embeddings
identity tuple. Save goes through a new service helper
`update_source_metadata` so the web route stays a thin wrapper.
Retry exposes two modes, both available on any record:
- full: status -> PENDING (re-pull + re-process)
- extract: status -> TRANSCRIBED (re-extract only, requires an
existing transcript; offered when source_type=video and a
transcript is present, to avoid the expensive re-download path)
Both modes clear claimed_by/claimed_at and error_message so the queue
dance picks it up cleanly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Travis's original ask put the form on the "queue page" but it landed
only on /dashboard. Putting it on /queue too without duplicating any
logic:
- _add_to_queue_form.html — extracted the form markup into a shared
partial. Callers set `{% set swap_target = "#<id>-body" %}` before
including it; that's the only knob that differs between the two
pages. Same POST endpoint, same field set, same flash behaviour.
- _dashboard_body.html — now {% include %}s the shared partial with
swap_target="#dashboard-body". Net change: a six-line include
replacing the inline form.
- _queue_body.html (new) — `<div id="queue-body">` wrapping the shared
form (swap_target="#queue-body") + the in-flight source list (same
card markup as the Sources view, with the wrap-fix already applied).
Empty state copy matches the queue context.
- queue.html (new) — extends base.html, page-titles "Queue", and
includes _queue_body.html. /queue no longer reuses index.html.
- /queue route — renders queue.html via a new _build_queue_context
helper. Same in-flight filter as before (PENDING/PULLED/TRANSCRIBED).
- POST /sources/add — reads HX-Current-URL (HTMX sends it on every
request; falls back to Referer) and picks the response partial:
/queue → _queue_body.html, anything else → _dashboard_body.html. One
endpoint, one service call, two render branches — neither side
reaches into the other's state.
Live-verified through the rebuilt container:
- GET /queue (LAN + brain.herbylab.dev) — form present.
- GET /dashboard — form still present (unchanged behaviour).
- GET / — form still absent (Sources view stays clean).
- POST from /queue → response carries id="queue-body" (and not
dashboard-body); the freshly-added row appears in the swapped list.
- POST from /dashboard → response carries id="dashboard-body".
- Flash scenarios from /queue: queued / duplicate / playlist 13 videos
/ validation error — all four render correctly.
Pasting a YouTube playlist URL into either entry point now expands
into one source row per video. Single-video URLs and non-YouTube URLs
keep their existing behaviour untouched.
Service module additions:
- is_youtube_playlist_url(url): strict detector. Only `/playlist?list=…`
on a known YouTube host (youtube.com / m / music / no-www) counts.
A `watch?v=…&list=…` URL is ambiguous (user usually pasted a single
video that happens to sit inside a playlist) and intentionally falls
through to single-add. To fan out, paste the canonical playlist URL.
- expand_youtube_playlist(url, *, max_items=50): yt-dlp with
extract_flat=True, playlistend=max_items, skip_download. Builds a
canonical https://www.youtube.com/watch?v={id} URL per entry and
silently drops placeholders for private/removed videos.
- add_playlist(sess, *, url, domain, focus, max_items, expander=None):
loops expansion entries through add_source so URL validation,
source_type detection, and the UNIQUE dedupe path stay identical to
the single-add flow. Per-entry titles win over any caller-supplied
title (a single playlist title would be wrong for N videos). Partial
failures don't abort the batch — failed entries are tallied with
up-to-10 (url, reason) tuples for the flash. `expander` is an
injection seam for tests so the suite never hits YouTube unless
explicitly opted in.
- DEFAULT_PLAYLIST_MAX_ITEMS = 50 — shared ceiling, no throttle change.
CLI: `second-brain add <playlist-url>` auto-detects and reports
`expanded / added / duplicates / failed`. No new flag needed.
Web: POST /sources/add same detection. _playlist_flash() builds the
HTMX flash — "Queued N videos (M duplicates skipped, F failed)"
with sensible plural forms and graceful omission of zero counters.
Tests:
- 15 pure-Python detection cases (positives + negatives, including the
ambiguous watch?v=…&list=… rule).
- 3 DB-backed add_playlist tests (with a mocked expander, so no
network): count aggregation across new + pre-seeded duplicates,
bad-entry tolerance, and the empty-playlist case.
- 1 opt-in live-network test gated on SECOND_BRAIN_LIVE_NETWORK_TESTS=1
exercising expand_youtube_playlist against a real public playlist.
Live-verified end to end:
- web POST of a real 13-entry public playlist queued 13 video rows
with titles, flash showed "Queued 13 videos".
- re-POST returned "Queued 0 videos (13 duplicates skipped)".
- watch?v=…&list=… correctly stayed a single-add.
- CLI parity confirmed against the same playlist.
Tower worker now persists faster-whisper's segment-level output
(start/end/text + word-level timing when available) alongside the
existing joined `transcript_text`. The text column stays the canonical
input the extractor reads — this is additive.
Changes:
- alembic v4: sources.transcript_segments JSONB NULL. JSONB rather than
JSON so future equality/containment queries are indexable without a
re-migration. Same lovebug-no-CREATE-on-petalbrain guard as prior
migrations.
- ORM model: Optional[list] mapped to JSONB (postgresql dialect).
- transcribe.py:
- Always pass word_timestamps=True to faster-whisper.transcribe.
- New segment_to_dict() flattens the upstream NamedTuple-shaped
Segment/Word into JSON-safe plain dicts so the JSONB write doesn't
drag faster-whisper into any reader.
- Per-word defensive conversion: a single malformed word can't drop
the surrounding segment.
- transcribe_worker._advance: after a successful transcribe, persist
segments into source.transcript_segments inside a try/except. If the
JSONB write fails (oversize row, malformed dict, etc.) we log a
warning and still commit transcript_text + status=TRANSCRIBED — the
pipeline never crashes over the additive index.
- Tests: three new unit tests against fake Segment/Word objects cover
the happy path (word entries serialise), the no-words case
(`segment.words is None` → empty list), and the malformed-word skip.
json.dumps(d) asserts JSONB-binding compatibility.
Live-verified: migration applied clean against petalbrain (`\d sources`
shows transcript_segments jsonb); ORM round-trip writes and reads the
sample payload identically. GPU large-v3 word-timestamp behaviour is
unchanged from upstream — only the tower can validate that hot path.
Travis prefers the cards to auto-resize to their content rather than
ellipsis-clip the title + URL into a single line. Two places use the
same card markup — both updated:
- src/second_brain/web/templates/index.html — main sources list.
- src/second_brain/web/templates/_dashboard_body.html — recent activity
on the dashboard.
Changes are minimal and identical on both:
- title h3/p: drop `truncate`, add `break-words` so normal long titles
wrap at word boundaries.
- URL p: drop `truncate`, add `break-all` because URLs are typically
one unbreakable token and `break-words` alone wouldn't split them.
- `min-w-0 flex-1` on the left container kept — it lets the flex child
shrink so the right column's status / domain / timestamp badges stay
pinned and never get pushed off-screen by a long URL.
- The optional Focus line also gets `break-words` defensively.
Verified live on the rebuilt web container:
- /dashboard: 0 truncate hits in the response, RepoWise URL renders
intact with break-all applied.
- /?status=analyzed: 0 truncate hits, title + URL both carry the
new wrap classes.
Extracts the `second-brain add` CLI's queueing logic into a service
module (sources_service.add_source) so the CLI and the new web POST
route share the same validation, dedupe, and source_type heuristic —
no behaviour drift between the two entry points.
UI:
- The dashboard body (Pipeline counts + Settings snapshot + Recent
activity) moves into _dashboard_body.html, wrapped in
`<div id="dashboard-body">`. HTMX targets that id for swap.
- A new "Add to queue" section sits at the top of the partial: URL
(required, type=url), Domain (select matching DOMAINS), Title and
Focus (both optional, match the CLI flags). hx-post=/sources/add,
hx-target=#dashboard-body, hx-swap=outerHTML — same pattern as the
settings save form.
Route:
- POST /sources/add calls sources_service.add_source inside a single
transaction, then re-renders _dashboard_body.html so the pipeline
counts and recent-activity list update in place. Flash slots above
the form report:
✓ Queued <type>: <url> on a fresh add
✓ Already queued (id=…, status=…) on a dupe (matches CLI text)
✗ <validation message> on bad URL / unknown domain
CLI:
- `second-brain add` now delegates to the same service. Field values
for the success echo are captured inside the session block so a
post-commit detached-instance access can't fail. Bad input exits 2
with a clear stderr message instead of raising.
Live-verified through the web container on herbys-dev: dashboard
renders the form, happy add lands a row, dup-detect matches the CLI
phrasing, two validation errors surface as red flashes, swap target
id survives across swaps, no tracebacks in uvicorn logs.
Splits pull+transcribe (now tower-side, eager) from extract+embed
(stays on the dev scheduler). Three machine-coordination pieces land
together because they reference each other:
- v3 migration adds sources.claimed_by + claimed_at — observability +
stale-claim recovery columns. The actual race-safety primitive is
`SELECT ... FOR UPDATE SKIP LOCKED` in the new claim helper, so two
machines can poll the queue without doubling work.
- src/second_brain/claim.py owns the claim dance (claim_next_source,
release_claim, reap_stale_claims). Both stage gates filter by
source_type so the tower never grabs articles and the dev side never
grabs videos.
- src/second_brain/transcribe.py wraps faster-whisper (lazy-imported so
it stays out of the dev install). resolve_settings() reads env >
[whisper] block > defaults, falling back to int8 on cpu / float16 on
cuda when compute_type is unspecified. Default model large-v3.
- src/second_brain/scheduler/transcribe_worker.py is the long-running
poll loop. Reads pipeline_settings every iteration so the dashboard's
enable/window/max-items/max-video-length take effect within one
cycle. Reaps stale claims at startup. SIGTERM-clean. DB-unreachable
backs off with a log line; never crash-loops.
- adapters/youtube.py drops the torch-whisper transcribe path; pull
stays. Removes openai-whisper from the default deps and gates
faster-whisper behind a new `tower` extra (uv sync --extra tower).
- main.py: new `second-brain transcribe-worker` (--once for ad-hoc).
`process` now article-only on the pull side but still picks up
TRANSCRIBED of any source_type for the extract step.
Live-verified: migration applies clean, transcribe-worker --once
honors transcription_enabled=false gate.
The dev-side scheduler now snapshots the singleton pipeline_settings row
at the top of run() and applies three DB-driven gates:
- extraction_enabled = false → skip the entire run with a single log line.
Lets Travis pause extraction from the dashboard without touching the
config file or restarting anything.
- extraction_active_hours_start/_end → fast-fail if outside the window
AND (when set) override the static settings.toml [scheduler].window_*
bounds so the inner loop also respects the dashboard's choice. Either
side being NULL means "no constraint on that side" — matches the form's
"leave blank = always active" semantics.
- extraction_max_items_per_run → hard cap on processed items per
invocation. 0/null means unlimited (existing behavior).
Snapshot semantics are deliberate: a mid-run toggle doesn't half-apply,
same way max_calls_per_hour is tracked locally. Workers re-read on the
next invocation.
Adds three new routes on the FastAPI web app:
- GET /dashboard — pipeline counts by status (per-status drill-down
via filtered /?status=… links), recent activity
(latest 10 sources), and a settings snapshot
card with an "edit →" shortcut.
- GET /settings — full editing form for pipeline_settings.
- POST /settings/save — HTMX endpoint that validates the partial form,
upserts the row, and returns the rerendered
card with a "Saved." or error flash. The card
is its own _settings_card.html partial so the
HTMX swap targets only the form region.
Form parsing lives in settings_store helpers (parse_time_or_none,
parse_int_or_none) — keeps the route thin and the empty-string-to-NULL
normalisation in one place. Validation rejects negative counts and
sub-1 GPU concurrency.
The base template grew a Dashboard + Settings nav so the new routes are
reachable without typing URLs. Added python-multipart to deps because
FastAPI's Form(...) raises at app import without it.
Introduces second_brain.pipeline_settings — the single-row config row
the upcoming web dashboard edits and the workers read at the start of
each run. Pinned to id=1 by a CHECK constraint so upserts-by-PK keep
the table singleton, and the migration seeds the row with the table's
column defaults via INSERT ... ON CONFLICT DO NOTHING.
Two consumer groups carved out:
- transcription_* fields persist now; future tower-side worker reads them.
- extraction_* fields will be read by the existing scheduler in the next
commit, which is the actual behavior change Travis cares about today.
The settings_store helper centralises get/update + form-parsing
(time-of-day, int-or-none) and active-window math so the routes and the
scheduler don't reimplement them.
Without this, every `second-brain process` run prints
couldn't stop thread 'pool-1-worker-N' within 5.0 seconds
during interpreter shutdown — psycopg_pool's background workers don't
get a clean stop signal before Python's threading shutdown deadline.
Registering close_pool via atexit the first time we open the pool fixes
it without changing any API. `close_pool` is already idempotent, so the
explicit teardown path in the smoke test (which calls it directly) and
the atexit path coexist safely.
Ruff autofix removed pre-existing F401 unused imports across adapters,
context, compiler, extractor, scheduler, web, and main. Also reordered
models.py so its SQLAlchemy imports sit at the top of the file
(E402 was triggered by the inline `utcnow` helper definition).
Added a [dependency-groups] dev block (pytest, pytest-cov, ruff) so the
zero-check test gauntlet — which runs `pytest --cov --cov-report=term-missing`
— can resolve its plugins without a manual `uv pip install pytest-cov`.
Swap the SQLite backing store for petalbrain Postgres + pgvector, modeled
on vault-mcp. All second-brain relational tables now live in the
`second_brain` schema (owned by the lovebug role); embeddings are written
to the shared public.embeddings table.
Locked design decisions (per Travis):
- DB: existing petalbrain Postgres, second_brain schema, lovebug role.
- Connection: containerized homelab-postgres:5432, plain psycopg_pool
(min=1/max=10), no PgBouncer.
- ORM stays SQLAlchemy; int autoincrement PKs + naive UTC DateTime.
- Embeddings: reuse shared public.embeddings keyed by
(source_schema='second_brain', source_table='extractions', source_id,
model='nomic-embed-text'). Summaries only for this round.
- Pipeline: chunk_text → Ollama nomic-embed-text → delete-before-insert
upsert, with graceful degradation (no DB / no Ollama → log + skip).
- Alembic stands up second-brain's own schema; public.embeddings stays
out-of-band.
- File-based wiki compiler is unchanged.
No SQLite data import — starting clean.
This commit is the scaffolding only; `alembic upgrade head` and a smoke
test of the embedding path are the next checkpoint.
Run extraction under the Max OAuth subscription via `claude -p` instead
of the per-token Anthropic API. The new src/second_brain/llm/claude_cli.py
spawns the CLI in a hermetic tempdir so the host project's CLAUDE.md,
hooks, MCP config, and settings don't leak into the prompt. Uses
--json-schema with LLMExtraction.model_json_schema() so the CLI guarantees
valid structured output — replaces the brittle markdown-fence stripping
in the old engine. The Anthropic SDK is preserved as an optional "api"
backend selectable via config.
While in here, fix a handful of blockers that the smoke test surfaced:
- scheduler filtered ANALYZED instead of TRANSCRIBED, so it never
actually advanced any sources
- process command read sources in a closed session, raising
DetachedInstanceError before any work happened
- config.prompts_dir walked one parent too many, resolving outside
the project and forcing the fallback prompt for every domain
- compiler called git rev-parse against a vault that was never
git-init'd; now auto-inits with an empty seed commit and skips empty
commits cleanly
- datetime.utcnow() deprecated in 3.12+ — single utcnow() helper in
models.py keeps naive UTC semantics so no DB migration is needed
- sess.query(...).get() deprecated in SA 2.x → sess.get(...)
- dead `import anthropic` removed from compiler
Smoke test (article → process → accept → compile) succeeds end-to-end
with ANTHROPIC_API_KEY unset. a-review run saved under reviews/.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>