Arch / EndeavourOS now ships CUDA 13 (libcublas.so.13); ctranslate2
4.7.2 (which faster-whisper rides) wants CUDA 12 + cuDNN 9 and won't
load against the system libs. Travis got it working ad-hoc with a
shell export of LD_LIBRARY_PATH + manual pip install, but systemd
doesn't inherit either, so the next reboot would refire the
libcublas.so.12 load error.
Making it permanent + reproducible:
- pyproject `tower` extra now pins the CUDA-12 runtime as pip wheels
alongside faster-whisper:
nvidia-cublas-cu12; sys_platform == 'linux'
nvidia-cudnn-cu12>=9,<10; sys_platform == 'linux'
uv.lock resolves nvidia-cublas-cu12 12.9.2.10 + nvidia-cudnn-cu12
9.22.0.52. Dev side (no --extra tower) stays clean — verified by a
no-extra `uv sync` followed by `uv pip list | grep nvidia` returning
empty.
- deploy/tower/run-transcribe-worker.sh (new, +x): computes the venv's
CUDA-12 lib dirs at runtime via `uv run python` (resolving
nvidia.cublas / nvidia.cudnn through __path__ — they're PEP 420
namespace packages with no __file__), prepends them to
LD_LIBRARY_PATH, then execs `uv run second-brain transcribe-worker`.
No hard-coded python3.XX path so it survives Python upgrades. If the
wheels aren't installed it aborts with a clear "uv sync --extra
tower" hint instead of a silent libcublas load failure deep inside
ctranslate2.
- second-brain-transcribe.service: ExecStart now points at the
wrapper. Also moves StartLimitIntervalSec / StartLimitBurst from
[Service] into [Unit] where modern systemd expects them
(systemd-analyze verify previously flagged the misplaced keys as
silently ignored). Restart=always, EnvironmentFile, After=/Wants=
wg-quick@wg-lan.service, User=herbyadmin all unchanged.
- second-brain-transcribe.env.example: trimmed to just
SECOND_BRAIN_DATABASE_URL with the placeholder spelled out, plus a
clear pointer to the ready-to-scp env file generated on herbys-dev
at /opt/backups/postgres-consolidation/second-brain-transcribe.env
(mode 0600, regeneratable from credentials.env without ever echoing
the password). The committed example never carries a real secret.
- deploy/tower/README.md: documents the CUDA-13-vs-CUDA-12 gotcha
upfront ("don't `pacman -S cuda cudnn`"), the wrapper-based
ExecStart, the scp-from-dev EnvironmentFile recipe with the
password-regen one-liner, and the EnvironmentFile-vs-shell-export
note.
Verified locally on dev (no GPU):
- uv.lock resolves with the new tower deps.
- A throwaway venv installed with the same `nvidia-cublas-cu12
nvidia-cudnn-cu12>=9,<10` pins produces lib dirs containing
libcublas.so.12 and libcudnn.so.9 via the wrapper's path probe.
- systemd-analyze verify is clean except the expected
"/opt/projects/... not executable on this host" warning (the
wrapper exists only in the tower's checkout).
- 27 passed / 2 skipped in pytest; zero-check 5/5.
GPU large-v3 + live service start under systemd remain tower-only
validation steps.
51 lines
1.9 KiB
Desktop File
51 lines
1.9 KiB
Desktop File
[Unit]
|
|
Description=second-brain transcribe worker (tower-side, faster-whisper on GPU)
|
|
Documentation=https://gitea.plantbasedsoutherner.com/petal-power/second-brain
|
|
# The worker reaches petalbrain over WireGuard, so the tunnel must be up first.
|
|
# `Wants=` is the soft form — if the tunnel disappears later the worker logs
|
|
# DB-unreachable warnings and retries, it doesn't crash-loop.
|
|
After=network-online.target wg-quick@wg-lan.service
|
|
Wants=network-online.target wg-quick@wg-lan.service
|
|
|
|
# Crash loop guard — systemd will give up after this if the unit keeps
|
|
# dying. The worker itself logs+backs-off on DB outages so this only
|
|
# trips on genuine failure. (StartLimit* directives live in [Unit] on
|
|
# modern systemd; placing them in [Service] is a silent no-op.)
|
|
StartLimitIntervalSec=300
|
|
StartLimitBurst=5
|
|
|
|
[Service]
|
|
Type=simple
|
|
User=herbyadmin
|
|
Group=herbyadmin
|
|
|
|
# Adjust to wherever you `git clone`d the repo on the tower.
|
|
WorkingDirectory=/opt/projects/homelab/second-brain
|
|
|
|
# Env (SECOND_BRAIN_DATABASE_URL, optional knobs). 0600, owned by herbyadmin.
|
|
EnvironmentFile=/etc/default/second-brain-transcribe
|
|
|
|
# The wrapper resolves the venv's CUDA-12 lib dirs (nvidia-cublas-cu12 /
|
|
# nvidia-cudnn-cu12 pulled in by `uv sync --extra tower`), prepends them
|
|
# to LD_LIBRARY_PATH, then execs the worker via `uv run`. systemd does
|
|
# NOT inherit a shell's LD_LIBRARY_PATH, so doing this inline in the
|
|
# unit would require a brittle hard-coded `python3.XX` path and break
|
|
# on Python upgrades.
|
|
ExecStart=/opt/projects/homelab/second-brain/deploy/tower/run-transcribe-worker.sh
|
|
|
|
Restart=always
|
|
RestartSec=10
|
|
|
|
# A few sandboxing nice-to-haves. Loose on purpose — the worker needs
|
|
# /dev/nvidia* for CUDA and writes to user-owned media/subtitles dirs.
|
|
ProtectSystem=full
|
|
ProtectHome=no
|
|
NoNewPrivileges=true
|
|
|
|
# Stdout/stderr → journald.
|
|
StandardOutput=journal
|
|
StandardError=journal
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|