Commit Graph

17 Commits

Author SHA1 Message Date
Travis Herbranson
22f4a0915f web: show "Re-extract only" for articles too
Gate was source_type == 'video', but the service-layer reset already
accepts any source with transcript_text — articles store the
trafilatura-extracted body there, so they're equally valid candidates
for an extract-only retry. Caught by a-review.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 16:15:39 -04:00
Travis Herbranson
5131f3ccf6 web+cli: edit-mode for source metadata + retry/re-run actions
Adds per-record edit mode and retry to the source detail page, plus a
CLI `second-brain retry <id>` for parity.

Edit mode exposes domain (select), title, focus on /sources/{id}.
URL stays read-only — it's the UNIQUE dedupe key and the embeddings
identity tuple. Save goes through a new service helper
`update_source_metadata` so the web route stays a thin wrapper.

Retry exposes two modes, both available on any record:
  - full: status -> PENDING (re-pull + re-process)
  - extract: status -> TRANSCRIBED (re-extract only, requires an
    existing transcript; offered when source_type=video and a
    transcript is present, to avoid the expensive re-download path)

Both modes clear claimed_by/claimed_at and error_message so the queue
dance picks it up cleanly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 16:08:26 -04:00
Travis Herbranson
2b40157a0d web: add-to-queue form on /queue too, with per-page HTMX dispatch
Travis's original ask put the form on the "queue page" but it landed
only on /dashboard. Putting it on /queue too without duplicating any
logic:

- _add_to_queue_form.html — extracted the form markup into a shared
  partial. Callers set `{% set swap_target = "#<id>-body" %}` before
  including it; that's the only knob that differs between the two
  pages. Same POST endpoint, same field set, same flash behaviour.

- _dashboard_body.html — now {% include %}s the shared partial with
  swap_target="#dashboard-body". Net change: a six-line include
  replacing the inline form.

- _queue_body.html (new) — `<div id="queue-body">` wrapping the shared
  form (swap_target="#queue-body") + the in-flight source list (same
  card markup as the Sources view, with the wrap-fix already applied).
  Empty state copy matches the queue context.

- queue.html (new) — extends base.html, page-titles "Queue", and
  includes _queue_body.html. /queue no longer reuses index.html.

- /queue route — renders queue.html via a new _build_queue_context
  helper. Same in-flight filter as before (PENDING/PULLED/TRANSCRIBED).

- POST /sources/add — reads HX-Current-URL (HTMX sends it on every
  request; falls back to Referer) and picks the response partial:
  /queue → _queue_body.html, anything else → _dashboard_body.html. One
  endpoint, one service call, two render branches — neither side
  reaches into the other's state.

Live-verified through the rebuilt container:
- GET /queue (LAN + brain.herbylab.dev) — form present.
- GET /dashboard — form still present (unchanged behaviour).
- GET / — form still absent (Sources view stays clean).
- POST from /queue → response carries id="queue-body" (and not
  dashboard-body); the freshly-added row appears in the swapped list.
- POST from /dashboard → response carries id="dashboard-body".
- Flash scenarios from /queue: queued / duplicate / playlist 13 videos
  / validation error — all four render correctly.
2026-05-25 14:21:02 -04:00
Travis Herbranson
a475893403 playlists: queue-time YouTube fan-out via yt-dlp extract_flat
Pasting a YouTube playlist URL into either entry point now expands
into one source row per video. Single-video URLs and non-YouTube URLs
keep their existing behaviour untouched.

Service module additions:
- is_youtube_playlist_url(url): strict detector. Only `/playlist?list=…`
  on a known YouTube host (youtube.com / m / music / no-www) counts.
  A `watch?v=…&list=…` URL is ambiguous (user usually pasted a single
  video that happens to sit inside a playlist) and intentionally falls
  through to single-add. To fan out, paste the canonical playlist URL.
- expand_youtube_playlist(url, *, max_items=50): yt-dlp with
  extract_flat=True, playlistend=max_items, skip_download. Builds a
  canonical https://www.youtube.com/watch?v={id} URL per entry and
  silently drops placeholders for private/removed videos.
- add_playlist(sess, *, url, domain, focus, max_items, expander=None):
  loops expansion entries through add_source so URL validation,
  source_type detection, and the UNIQUE dedupe path stay identical to
  the single-add flow. Per-entry titles win over any caller-supplied
  title (a single playlist title would be wrong for N videos). Partial
  failures don't abort the batch — failed entries are tallied with
  up-to-10 (url, reason) tuples for the flash. `expander` is an
  injection seam for tests so the suite never hits YouTube unless
  explicitly opted in.
- DEFAULT_PLAYLIST_MAX_ITEMS = 50 — shared ceiling, no throttle change.

CLI: `second-brain add <playlist-url>` auto-detects and reports
`expanded / added / duplicates / failed`. No new flag needed.

Web: POST /sources/add same detection. _playlist_flash() builds the
HTMX flash — "Queued N videos (M duplicates skipped, F failed)"
with sensible plural forms and graceful omission of zero counters.

Tests:
- 15 pure-Python detection cases (positives + negatives, including the
  ambiguous watch?v=…&list=… rule).
- 3 DB-backed add_playlist tests (with a mocked expander, so no
  network): count aggregation across new + pre-seeded duplicates,
  bad-entry tolerance, and the empty-playlist case.
- 1 opt-in live-network test gated on SECOND_BRAIN_LIVE_NETWORK_TESTS=1
  exercising expand_youtube_playlist against a real public playlist.

Live-verified end to end:
- web POST of a real 13-entry public playlist queued 13 video rows
  with titles, flash showed "Queued 13 videos".
- re-POST returned "Queued 0 videos (13 duplicates skipped)".
- watch?v=…&list=… correctly stayed a single-add.
- CLI parity confirmed against the same playlist.
2026-05-25 13:59:35 -04:00
Travis Herbranson
d055d1798d transcripts: capture segment-level output into a new JSONB column
Tower worker now persists faster-whisper's segment-level output
(start/end/text + word-level timing when available) alongside the
existing joined `transcript_text`. The text column stays the canonical
input the extractor reads — this is additive.

Changes:

- alembic v4: sources.transcript_segments JSONB NULL. JSONB rather than
  JSON so future equality/containment queries are indexable without a
  re-migration. Same lovebug-no-CREATE-on-petalbrain guard as prior
  migrations.

- ORM model: Optional[list] mapped to JSONB (postgresql dialect).

- transcribe.py:
  - Always pass word_timestamps=True to faster-whisper.transcribe.
  - New segment_to_dict() flattens the upstream NamedTuple-shaped
    Segment/Word into JSON-safe plain dicts so the JSONB write doesn't
    drag faster-whisper into any reader.
  - Per-word defensive conversion: a single malformed word can't drop
    the surrounding segment.

- transcribe_worker._advance: after a successful transcribe, persist
  segments into source.transcript_segments inside a try/except. If the
  JSONB write fails (oversize row, malformed dict, etc.) we log a
  warning and still commit transcript_text + status=TRANSCRIBED — the
  pipeline never crashes over the additive index.

- Tests: three new unit tests against fake Segment/Word objects cover
  the happy path (word entries serialise), the no-words case
  (`segment.words is None` → empty list), and the malformed-word skip.
  json.dumps(d) asserts JSONB-binding compatibility.

Live-verified: migration applied clean against petalbrain (`\d sources`
shows transcript_segments jsonb); ORM round-trip writes and reads the
sample payload identically. GPU large-v3 word-timestamp behaviour is
unchanged from upstream — only the tower can validate that hot path.
2026-05-25 13:40:49 -04:00
Travis Herbranson
29832b3595 web: source cards wrap long titles/URLs instead of truncating
Travis prefers the cards to auto-resize to their content rather than
ellipsis-clip the title + URL into a single line. Two places use the
same card markup — both updated:

- src/second_brain/web/templates/index.html — main sources list.
- src/second_brain/web/templates/_dashboard_body.html — recent activity
  on the dashboard.

Changes are minimal and identical on both:
- title h3/p: drop `truncate`, add `break-words` so normal long titles
  wrap at word boundaries.
- URL p: drop `truncate`, add `break-all` because URLs are typically
  one unbreakable token and `break-words` alone wouldn't split them.
- `min-w-0 flex-1` on the left container kept — it lets the flex child
  shrink so the right column's status / domain / timestamp badges stay
  pinned and never get pushed off-screen by a long URL.
- The optional Focus line also gets `break-words` defensively.

Verified live on the rebuilt web container:
- /dashboard: 0 truncate hits in the response, RepoWise URL renders
  intact with break-all applied.
- /?status=analyzed: 0 truncate hits, title + URL both carry the
  new wrap classes.
2026-05-25 13:22:27 -04:00
Travis Herbranson
956bf8d0c3 web: add-to-queue form on the dashboard (HTMX, parity with CLI add)
Extracts the `second-brain add` CLI's queueing logic into a service
module (sources_service.add_source) so the CLI and the new web POST
route share the same validation, dedupe, and source_type heuristic —
no behaviour drift between the two entry points.

UI:
- The dashboard body (Pipeline counts + Settings snapshot + Recent
  activity) moves into _dashboard_body.html, wrapped in
  `<div id="dashboard-body">`. HTMX targets that id for swap.
- A new "Add to queue" section sits at the top of the partial: URL
  (required, type=url), Domain (select matching DOMAINS), Title and
  Focus (both optional, match the CLI flags). hx-post=/sources/add,
  hx-target=#dashboard-body, hx-swap=outerHTML — same pattern as the
  settings save form.

Route:
- POST /sources/add calls sources_service.add_source inside a single
  transaction, then re-renders _dashboard_body.html so the pipeline
  counts and recent-activity list update in place. Flash slots above
  the form report:
    ✓ Queued <type>: <url>            on a fresh add
    ✓ Already queued (id=…, status=…) on a dupe (matches CLI text)
    ✗ <validation message>            on bad URL / unknown domain

CLI:
- `second-brain add` now delegates to the same service. Field values
  for the success echo are captured inside the session block so a
  post-commit detached-instance access can't fail. Bad input exits 2
  with a clear stderr message instead of raising.

Live-verified through the web container on herbys-dev: dashboard
renders the form, happy add lands a row, dup-detect matches the CLI
phrasing, two validation errors surface as red flashes, swap target
id survives across swaps, no tracebacks in uvicorn logs.
2026-05-25 12:59:02 -04:00
Travis Herbranson
33ae5a6f7d pipeline split: tower transcribe worker + claim queue + faster-whisper
Splits pull+transcribe (now tower-side, eager) from extract+embed
(stays on the dev scheduler). Three machine-coordination pieces land
together because they reference each other:

- v3 migration adds sources.claimed_by + claimed_at — observability +
  stale-claim recovery columns. The actual race-safety primitive is
  `SELECT ... FOR UPDATE SKIP LOCKED` in the new claim helper, so two
  machines can poll the queue without doubling work.

- src/second_brain/claim.py owns the claim dance (claim_next_source,
  release_claim, reap_stale_claims). Both stage gates filter by
  source_type so the tower never grabs articles and the dev side never
  grabs videos.

- src/second_brain/transcribe.py wraps faster-whisper (lazy-imported so
  it stays out of the dev install). resolve_settings() reads env >
  [whisper] block > defaults, falling back to int8 on cpu / float16 on
  cuda when compute_type is unspecified. Default model large-v3.

- src/second_brain/scheduler/transcribe_worker.py is the long-running
  poll loop. Reads pipeline_settings every iteration so the dashboard's
  enable/window/max-items/max-video-length take effect within one
  cycle. Reaps stale claims at startup. SIGTERM-clean. DB-unreachable
  backs off with a log line; never crash-loops.

- adapters/youtube.py drops the torch-whisper transcribe path; pull
  stays. Removes openai-whisper from the default deps and gates
  faster-whisper behind a new `tower` extra (uv sync --extra tower).

- main.py: new `second-brain transcribe-worker` (--once for ad-hoc).
  `process` now article-only on the pull side but still picks up
  TRANSCRIBED of any source_type for the extract step.

Live-verified: migration applies clean, transcribe-worker --once
honors transcription_enabled=false gate.
2026-05-25 10:11:37 -04:00
Travis Herbranson
7eafdf3e03 scheduler: honor pipeline_settings (extraction_enabled, window, max_items)
The dev-side scheduler now snapshots the singleton pipeline_settings row
at the top of run() and applies three DB-driven gates:

- extraction_enabled = false → skip the entire run with a single log line.
  Lets Travis pause extraction from the dashboard without touching the
  config file or restarting anything.
- extraction_active_hours_start/_end → fast-fail if outside the window
  AND (when set) override the static settings.toml [scheduler].window_*
  bounds so the inner loop also respects the dashboard's choice. Either
  side being NULL means "no constraint on that side" — matches the form's
  "leave blank = always active" semantics.
- extraction_max_items_per_run → hard cap on processed items per
  invocation. 0/null means unlimited (existing behavior).

Snapshot semantics are deliberate: a mid-run toggle doesn't half-apply,
same way max_calls_per_hour is tracked locally. Workers re-read on the
next invocation.
2026-05-25 08:17:44 -04:00
Travis Herbranson
c92a40bce1 web: dashboard + settings editor (HTMX save) + pipeline observability
Adds three new routes on the FastAPI web app:
- GET  /dashboard        — pipeline counts by status (per-status drill-down
                           via filtered /?status=… links), recent activity
                           (latest 10 sources), and a settings snapshot
                           card with an "edit →" shortcut.
- GET  /settings         — full editing form for pipeline_settings.
- POST /settings/save    — HTMX endpoint that validates the partial form,
                           upserts the row, and returns the rerendered
                           card with a "Saved." or error flash. The card
                           is its own _settings_card.html partial so the
                           HTMX swap targets only the form region.

Form parsing lives in settings_store helpers (parse_time_or_none,
parse_int_or_none) — keeps the route thin and the empty-string-to-NULL
normalisation in one place. Validation rejects negative counts and
sub-1 GPU concurrency.

The base template grew a Dashboard + Settings nav so the new routes are
reachable without typing URLs. Added python-multipart to deps because
FastAPI's Form(...) raises at app import without it.
2026-05-25 08:15:29 -04:00
Travis Herbranson
09efbb5965 settings: v2 migration + ORM + settings_store helper
Introduces second_brain.pipeline_settings — the single-row config row
the upcoming web dashboard edits and the workers read at the start of
each run. Pinned to id=1 by a CHECK constraint so upserts-by-PK keep
the table singleton, and the migration seeds the row with the table's
column defaults via INSERT ... ON CONFLICT DO NOTHING.

Two consumer groups carved out:
- transcription_* fields persist now; future tower-side worker reads them.
- extraction_* fields will be read by the existing scheduler in the next
  commit, which is the actual behavior change Travis cares about today.

The settings_store helper centralises get/update + form-parsing
(time-of-day, int-or-none) and active-window math so the routes and the
scheduler don't reimplement them.
2026-05-25 08:11:53 -04:00
Travis Herbranson
0d9ebc9cc7 embeddings: register pool close() at interpreter exit
Without this, every `second-brain process` run prints
  couldn't stop thread 'pool-1-worker-N' within 5.0 seconds
during interpreter shutdown — psycopg_pool's background workers don't
get a clean stop signal before Python's threading shutdown deadline.

Registering close_pool via atexit the first time we open the pool fixes
it without changing any API. `close_pool` is already idempotent, so the
explicit teardown path in the smoke test (which calls it directly) and
the atexit path coexist safely.
2026-05-24 23:33:13 -04:00
Travis Herbranson
a2d46a7c6e gauntlet: fix lint — drop unused imports + reorder models.py + add dev deps
Ruff autofix removed pre-existing F401 unused imports across adapters,
context, compiler, extractor, scheduler, web, and main. Also reordered
models.py so its SQLAlchemy imports sit at the top of the file
(E402 was triggered by the inline `utcnow` helper definition).

Added a [dependency-groups] dev block (pytest, pytest-cov, ruff) so the
zero-check test gauntlet — which runs `pytest --cov --cov-report=term-missing`
— can resolve its plugins without a manual `uv pip install pytest-cov`.
2026-05-24 22:52:08 -04:00
Travis Herbranson
ceeae77e7d postgres migration: schema, models, embeddings, alembic
Swap the SQLite backing store for petalbrain Postgres + pgvector, modeled
on vault-mcp. All second-brain relational tables now live in the
`second_brain` schema (owned by the lovebug role); embeddings are written
to the shared public.embeddings table.

Locked design decisions (per Travis):
- DB: existing petalbrain Postgres, second_brain schema, lovebug role.
- Connection: containerized homelab-postgres:5432, plain psycopg_pool
  (min=1/max=10), no PgBouncer.
- ORM stays SQLAlchemy; int autoincrement PKs + naive UTC DateTime.
- Embeddings: reuse shared public.embeddings keyed by
  (source_schema='second_brain', source_table='extractions', source_id,
  model='nomic-embed-text'). Summaries only for this round.
- Pipeline: chunk_text → Ollama nomic-embed-text → delete-before-insert
  upsert, with graceful degradation (no DB / no Ollama → log + skip).
- Alembic stands up second-brain's own schema; public.embeddings stays
  out-of-band.
- File-based wiki compiler is unchanged.

No SQLite data import — starting clean.

This commit is the scaffolding only; `alembic upgrade head` and a smoke
test of the embedding path are the next checkpoint.
2026-05-24 22:46:48 -04:00
Travis Herbranson
3b269838d9 update to fix website error 2026-05-24 21:38:52 -04:00
Travis Herbranson
c4793f3a7f switch extractor to Claude CLI backend + pipeline fixes
Run extraction under the Max OAuth subscription via `claude -p` instead
of the per-token Anthropic API. The new src/second_brain/llm/claude_cli.py
spawns the CLI in a hermetic tempdir so the host project's CLAUDE.md,
hooks, MCP config, and settings don't leak into the prompt. Uses
--json-schema with LLMExtraction.model_json_schema() so the CLI guarantees
valid structured output — replaces the brittle markdown-fence stripping
in the old engine. The Anthropic SDK is preserved as an optional "api"
backend selectable via config.

While in here, fix a handful of blockers that the smoke test surfaced:
- scheduler filtered ANALYZED instead of TRANSCRIBED, so it never
  actually advanced any sources
- process command read sources in a closed session, raising
  DetachedInstanceError before any work happened
- config.prompts_dir walked one parent too many, resolving outside
  the project and forcing the fallback prompt for every domain
- compiler called git rev-parse against a vault that was never
  git-init'd; now auto-inits with an empty seed commit and skips empty
  commits cleanly
- datetime.utcnow() deprecated in 3.12+ — single utcnow() helper in
  models.py keeps naive UTC semantics so no DB migration is needed
- sess.query(...).get() deprecated in SA 2.x → sess.get(...)
- dead `import anthropic` removed from compiler

Smoke test (article → process → accept → compile) succeeds end-to-end
with ANTHROPIC_API_KEY unset. a-review run saved under reviews/.

Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2026-05-24 21:09:20 -04:00
Travis Herbranson
250ca9fd2f project init 2026-05-22 19:08:22 -04:00