Two commits since the last merge:
- web: add-to-queue form on /queue too, with per-page HTMX dispatch.
Shared form partial reused by /dashboard and /queue; the POST
/sources/add endpoint dispatches the response partial by inspecting
HX-Current-URL so each page's own list/counts refresh inline. Same
service call, same flash behavior.
- tower: pin CUDA-12 wheels + LD_LIBRARY_PATH wrapper for the systemd
service. ctranslate2 4.7.2 needs CUDA 12 + cuDNN 9 but Arch ships
CUDA 13, so we pin nvidia-cublas-cu12 + nvidia-cudnn-cu12>=9,<10 in
the `tower` extra and add deploy/tower/run-transcribe-worker.sh,
which resolves the venv's lib dirs at runtime and prepends them to
LD_LIBRARY_PATH before exec-ing the worker. systemd doesn't inherit
shell exports, so this is what makes the service survive reboots
and Python upgrades. ExecStart= now points at the wrapper, and the
StartLimit* directives moved into [Unit] where modern systemd
expects them.
Travis's tower-side corrections on main (`User=trucktrav`,
`gitea-pbs-trucktrav` clone alias, tower WG IP 10.99.0.3) auto-merge
cleanly against these branch edits — different lines in the same files.
Arch / EndeavourOS now ships CUDA 13 (libcublas.so.13); ctranslate2
4.7.2 (which faster-whisper rides) wants CUDA 12 + cuDNN 9 and won't
load against the system libs. Travis got it working ad-hoc with a
shell export of LD_LIBRARY_PATH + manual pip install, but systemd
doesn't inherit either, so the next reboot would refire the
libcublas.so.12 load error.
Making it permanent + reproducible:
- pyproject `tower` extra now pins the CUDA-12 runtime as pip wheels
alongside faster-whisper:
nvidia-cublas-cu12; sys_platform == 'linux'
nvidia-cudnn-cu12>=9,<10; sys_platform == 'linux'
uv.lock resolves nvidia-cublas-cu12 12.9.2.10 + nvidia-cudnn-cu12
9.22.0.52. Dev side (no --extra tower) stays clean — verified by a
no-extra `uv sync` followed by `uv pip list | grep nvidia` returning
empty.
- deploy/tower/run-transcribe-worker.sh (new, +x): computes the venv's
CUDA-12 lib dirs at runtime via `uv run python` (resolving
nvidia.cublas / nvidia.cudnn through __path__ — they're PEP 420
namespace packages with no __file__), prepends them to
LD_LIBRARY_PATH, then execs `uv run second-brain transcribe-worker`.
No hard-coded python3.XX path so it survives Python upgrades. If the
wheels aren't installed it aborts with a clear "uv sync --extra
tower" hint instead of a silent libcublas load failure deep inside
ctranslate2.
- second-brain-transcribe.service: ExecStart now points at the
wrapper. Also moves StartLimitIntervalSec / StartLimitBurst from
[Service] into [Unit] where modern systemd expects them
(systemd-analyze verify previously flagged the misplaced keys as
silently ignored). Restart=always, EnvironmentFile, After=/Wants=
wg-quick@wg-lan.service, User=herbyadmin all unchanged.
- second-brain-transcribe.env.example: trimmed to just
SECOND_BRAIN_DATABASE_URL with the placeholder spelled out, plus a
clear pointer to the ready-to-scp env file generated on herbys-dev
at /opt/backups/postgres-consolidation/second-brain-transcribe.env
(mode 0600, regeneratable from credentials.env without ever echoing
the password). The committed example never carries a real secret.
- deploy/tower/README.md: documents the CUDA-13-vs-CUDA-12 gotcha
upfront ("don't `pacman -S cuda cudnn`"), the wrapper-based
ExecStart, the scp-from-dev EnvironmentFile recipe with the
password-regen one-liner, and the EnvironmentFile-vs-shell-export
note.
Verified locally on dev (no GPU):
- uv.lock resolves with the new tower deps.
- A throwaway venv installed with the same `nvidia-cublas-cu12
nvidia-cudnn-cu12>=9,<10` pins produces lib dirs containing
libcublas.so.12 and libcudnn.so.9 via the wrapper's path probe.
- systemd-analyze verify is clean except the expected
"/opt/projects/... not executable on this host" warning (the
wrapper exists only in the tower's checkout).
- 27 passed / 2 skipped in pytest; zero-check 5/5.
GPU large-v3 + live service start under systemd remain tower-only
validation steps.
deploy/tower/:
- second-brain-transcribe.service — systemd unit. User=herbyadmin,
Type=simple, After=/Wants= wg-quick@wg-lan.service so the WG tunnel
must come up first. Restart=always with a StartLimitBurst guard.
- second-brain-transcribe.env.example — env file template documenting
the SECOND_BRAIN_DATABASE_URL form for db.wg.herbylab.dev (10.99.0.1)
and the optional WHISPER_* overrides.
- README.md — EndeavourOS install steps (nvidia/cuda/cudnn, ffmpeg, uv
+ tower extra, model pre-warm), WG topology reference, validation
checklist for what to confirm once the tunnel is live, and a
follow-ups section flagging the local-disk → NAS media migration as
out-of-scope-for-this-round.
Tests:
- tests/test_claim.py — live-DB race test. Two threads call
claim_next_source against a single PULLED video row; SKIP LOCKED
must give exactly one of them the row, the other gets None. Also
asserts the claimed_by/at columns land + release nulls them.
Auto-skips when no SECOND_BRAIN_DATABASE_URL is set.
- tests/test_transcribe.py — pure-Python coverage of resolve_settings
(cpu→int8, cuda→float16, env-over-block) and write_srt; plus a CPU
smoke test that synthesizes a numpy audio array and runs the `tiny`
model on cpu/int8 (auto-skipped when faster-whisper isn't installed,
i.e. on the dev side without --extra tower).