Commit Graph

3 Commits

Author SHA1 Message Date
Travis Herbranson
b2e2359651 tower: pin CUDA-12 wheels + LD_LIBRARY_PATH wrapper for the systemd service
Arch / EndeavourOS now ships CUDA 13 (libcublas.so.13); ctranslate2
4.7.2 (which faster-whisper rides) wants CUDA 12 + cuDNN 9 and won't
load against the system libs. Travis got it working ad-hoc with a
shell export of LD_LIBRARY_PATH + manual pip install, but systemd
doesn't inherit either, so the next reboot would refire the
libcublas.so.12 load error.

Making it permanent + reproducible:

- pyproject `tower` extra now pins the CUDA-12 runtime as pip wheels
  alongside faster-whisper:
    nvidia-cublas-cu12; sys_platform == 'linux'
    nvidia-cudnn-cu12>=9,<10; sys_platform == 'linux'
  uv.lock resolves nvidia-cublas-cu12 12.9.2.10 + nvidia-cudnn-cu12
  9.22.0.52. Dev side (no --extra tower) stays clean — verified by a
  no-extra `uv sync` followed by `uv pip list | grep nvidia` returning
  empty.

- deploy/tower/run-transcribe-worker.sh (new, +x): computes the venv's
  CUDA-12 lib dirs at runtime via `uv run python` (resolving
  nvidia.cublas / nvidia.cudnn through __path__ — they're PEP 420
  namespace packages with no __file__), prepends them to
  LD_LIBRARY_PATH, then execs `uv run second-brain transcribe-worker`.
  No hard-coded python3.XX path so it survives Python upgrades. If the
  wheels aren't installed it aborts with a clear "uv sync --extra
  tower" hint instead of a silent libcublas load failure deep inside
  ctranslate2.

- second-brain-transcribe.service: ExecStart now points at the
  wrapper. Also moves StartLimitIntervalSec / StartLimitBurst from
  [Service] into [Unit] where modern systemd expects them
  (systemd-analyze verify previously flagged the misplaced keys as
  silently ignored). Restart=always, EnvironmentFile, After=/Wants=
  wg-quick@wg-lan.service, User=herbyadmin all unchanged.

- second-brain-transcribe.env.example: trimmed to just
  SECOND_BRAIN_DATABASE_URL with the placeholder spelled out, plus a
  clear pointer to the ready-to-scp env file generated on herbys-dev
  at /opt/backups/postgres-consolidation/second-brain-transcribe.env
  (mode 0600, regeneratable from credentials.env without ever echoing
  the password). The committed example never carries a real secret.

- deploy/tower/README.md: documents the CUDA-13-vs-CUDA-12 gotcha
  upfront ("don't `pacman -S cuda cudnn`"), the wrapper-based
  ExecStart, the scp-from-dev EnvironmentFile recipe with the
  password-regen one-liner, and the EnvironmentFile-vs-shell-export
  note.

Verified locally on dev (no GPU):
- uv.lock resolves with the new tower deps.
- A throwaway venv installed with the same `nvidia-cublas-cu12
  nvidia-cudnn-cu12>=9,<10` pins produces lib dirs containing
  libcublas.so.12 and libcudnn.so.9 via the wrapper's path probe.
- systemd-analyze verify is clean except the expected
  "/opt/projects/... not executable on this host" warning (the
  wrapper exists only in the tower's checkout).
- 27 passed / 2 skipped in pytest; zero-check 5/5.

GPU large-v3 + live service start under systemd remain tower-only
validation steps.
2026-05-25 15:39:13 -04:00
Travis Herbranson
597c83649c deploy/web: containerized FastAPI UI on the homelab docker network
Dockerfile, compose.yml, env template, and runbook for the second-brain
web container. Targets herbys-dev (10.0.21.207); Traefik (file-provider
on 10.0.11.20) reaches the published host port at 10.0.21.207:8080.

Image:
- python:3.12-slim + the official uv binary copied from the upstream
  image, plus apt-installed git + ca-certificates so the Gitea VCS pin
  for embedding-chunking resolves at build time.
- uv sync --frozen --no-dev --no-install-project, source copy, then a
  second uv sync to install the project itself. Two-step so the lock
  install layer caches independently of source edits.
- No ffmpeg / claude CLI / faster-whisper — web role doesn't need any
  of them. Extraction runs on the dev host's CLI; transcription on the
  tower.
- Drops to uid 1000 (`app`) before CMD. Uvicorn binds 0.0.0.0:8000
  inside the container, with --proxy-headers + --forwarded-allow-ips=*
  so Traefik's X-Forwarded-* survive.

compose.yml:
- Joins the existing external `homelab` bridge network so the container
  reaches homelab-postgres:5432 and ollama:11434 by service DNS.
- Publishes the uvicorn port at 10.0.21.207:8080 (LAN IP bound, not
  0.0.0.0) for Traefik on the separate VM to reach. NO traefik.* labels
  — file-provider Traefik can't read them.
- env_file: .env (0600, gitignored) — SECOND_BRAIN_DATABASE_URL points
  at homelab-postgres:5432 (containerised), NOT the host's 127.0.0.1:5433
  port-map.
- restart: unless-stopped.

README.md:
- Build + bring-up commands.
- SECURITY note: no SSO / no CSRF / mutating endpoints — Traefik route
  must be LAN-only for now (Travis's call).
- The two infra steps Travis applies himself, with ready-to-paste
  snippets:
  - Knot DNS: brain.herbylab.dev → 10.0.11.20.
  - Traefik dynamic config: file-provider router + service block.
- Verification checklist for both before-and-after-DNS states.

Live-verified on herbys-dev: container Up, dashboard returns 200 with
real status counts from petalbrain, settings save round-trip works,
no tracebacks in logs.
2026-05-25 11:48:37 -04:00
Travis Herbranson
913d4f0335 deploy/tower + claim race + transcribe smoke tests
deploy/tower/:
- second-brain-transcribe.service — systemd unit. User=herbyadmin,
  Type=simple, After=/Wants= wg-quick@wg-lan.service so the WG tunnel
  must come up first. Restart=always with a StartLimitBurst guard.
- second-brain-transcribe.env.example — env file template documenting
  the SECOND_BRAIN_DATABASE_URL form for db.wg.herbylab.dev (10.99.0.1)
  and the optional WHISPER_* overrides.
- README.md — EndeavourOS install steps (nvidia/cuda/cudnn, ffmpeg, uv
  + tower extra, model pre-warm), WG topology reference, validation
  checklist for what to confirm once the tunnel is live, and a
  follow-ups section flagging the local-disk → NAS media migration as
  out-of-scope-for-this-round.

Tests:
- tests/test_claim.py — live-DB race test. Two threads call
  claim_next_source against a single PULLED video row; SKIP LOCKED
  must give exactly one of them the row, the other gets None. Also
  asserts the claimed_by/at columns land + release nulls them.
  Auto-skips when no SECOND_BRAIN_DATABASE_URL is set.
- tests/test_transcribe.py — pure-Python coverage of resolve_settings
  (cpu→int8, cuda→float16, env-over-block) and write_srt; plus a CPU
  smoke test that synthesizes a numpy audio array and runs the `tiny`
  model on cpu/int8 (auto-skipped when faster-whisper isn't installed,
  i.e. on the dev side without --extra tower).
2026-05-25 10:20:46 -04:00