second-brain/pyproject.toml
Travis Herbranson b2e2359651 tower: pin CUDA-12 wheels + LD_LIBRARY_PATH wrapper for the systemd service
Arch / EndeavourOS now ships CUDA 13 (libcublas.so.13); ctranslate2
4.7.2 (which faster-whisper rides) wants CUDA 12 + cuDNN 9 and won't
load against the system libs. Travis got it working ad-hoc with a
shell export of LD_LIBRARY_PATH + manual pip install, but systemd
doesn't inherit either, so the next reboot would refire the
libcublas.so.12 load error.

Making it permanent + reproducible:

- pyproject `tower` extra now pins the CUDA-12 runtime as pip wheels
  alongside faster-whisper:
    nvidia-cublas-cu12; sys_platform == 'linux'
    nvidia-cudnn-cu12>=9,<10; sys_platform == 'linux'
  uv.lock resolves nvidia-cublas-cu12 12.9.2.10 + nvidia-cudnn-cu12
  9.22.0.52. Dev side (no --extra tower) stays clean — verified by a
  no-extra `uv sync` followed by `uv pip list | grep nvidia` returning
  empty.

- deploy/tower/run-transcribe-worker.sh (new, +x): computes the venv's
  CUDA-12 lib dirs at runtime via `uv run python` (resolving
  nvidia.cublas / nvidia.cudnn through __path__ — they're PEP 420
  namespace packages with no __file__), prepends them to
  LD_LIBRARY_PATH, then execs `uv run second-brain transcribe-worker`.
  No hard-coded python3.XX path so it survives Python upgrades. If the
  wheels aren't installed it aborts with a clear "uv sync --extra
  tower" hint instead of a silent libcublas load failure deep inside
  ctranslate2.

- second-brain-transcribe.service: ExecStart now points at the
  wrapper. Also moves StartLimitIntervalSec / StartLimitBurst from
  [Service] into [Unit] where modern systemd expects them
  (systemd-analyze verify previously flagged the misplaced keys as
  silently ignored). Restart=always, EnvironmentFile, After=/Wants=
  wg-quick@wg-lan.service, User=herbyadmin all unchanged.

- second-brain-transcribe.env.example: trimmed to just
  SECOND_BRAIN_DATABASE_URL with the placeholder spelled out, plus a
  clear pointer to the ready-to-scp env file generated on herbys-dev
  at /opt/backups/postgres-consolidation/second-brain-transcribe.env
  (mode 0600, regeneratable from credentials.env without ever echoing
  the password). The committed example never carries a real secret.

- deploy/tower/README.md: documents the CUDA-13-vs-CUDA-12 gotcha
  upfront ("don't `pacman -S cuda cudnn`"), the wrapper-based
  ExecStart, the scp-from-dev EnvironmentFile recipe with the
  password-regen one-liner, and the EnvironmentFile-vs-shell-export
  note.

Verified locally on dev (no GPU):
- uv.lock resolves with the new tower deps.
- A throwaway venv installed with the same `nvidia-cublas-cu12
  nvidia-cudnn-cu12>=9,<10` pins produces lib dirs containing
  libcublas.so.12 and libcudnn.so.9 via the wrapper's path probe.
- systemd-analyze verify is clean except the expected
  "/opt/projects/... not executable on this host" warning (the
  wrapper exists only in the tower's checkout).
- 27 passed / 2 skipped in pytest; zero-check 5/5.

GPU large-v3 + live service start under systemd remain tower-only
validation steps.
2026-05-25 15:39:13 -04:00

66 lines
2.1 KiB
TOML

[project]
name = "second-brain"
version = "0.1.0"
description = "Personal knowledge system: video/article extraction pipeline + LLM-maintained wiki"
requires-python = ">=3.12"
dependencies = [
"alembic>=1.13.0",
"click>=8.1.0",
"embedding-chunking",
"fastapi>=0.115.0",
"httpx>=0.28.0",
"jinja2>=3.1.0",
"psycopg[binary]>=3.2",
"psycopg-pool>=3.2",
"pydantic>=2.0.0",
"python-dotenv>=1.0.0",
"python-multipart>=0.0.20",
"sqlalchemy>=2.0.0",
"tomli>=2.0.0",
"trafilatura>=1.9.0",
"uvicorn[standard]>=0.30.0",
"yt-dlp>=2024.1.0",
]
[project.optional-dependencies]
# Only required when extractor.backend = "api" in settings.toml.
# The default "cli" backend shells out to the Claude Code CLI under your
# Max OAuth subscription and has no Python-side Anthropic dependency.
api = ["anthropic>=0.40.0"]
# Tower-only: faster-whisper (CTranslate2) + the matched CUDA-12 runtime
# pulled in as wheels. Arch/EndeavourOS now ships CUDA 13 (libcublas.so.13)
# which ctranslate2 4.7+ cannot load — it wants libcublas.so.12 + cuDNN 9.
# Pinning the CUDA-12 runtime via pip wheels in the venv sidesteps the
# system CUDA mismatch entirely. sys_platform marker keeps these out of
# any non-Linux resolve (the nvidia wheels are Linux-only anyway, but the
# marker makes that explicit). Install with `uv sync --extra tower`; do
# not install on dev.
tower = [
"faster-whisper>=1.0.0",
"nvidia-cublas-cu12; sys_platform == 'linux'",
"nvidia-cudnn-cu12>=9,<10; sys_platform == 'linux'",
]
[dependency-groups]
dev = [
"pytest>=8.0",
"pytest-cov>=5.0",
"ruff>=0.8",
]
[tool.uv.sources]
# Shared chunker that vault-mcp / ob1-enricher / ob1-reflector also pin.
# Same source so the 512/64 chunking stays in lockstep across services.
embedding-chunking = { git = "https://gitea.plantbasedsoutherner.com/petal-power/embedding-chunking.git", tag = "v0.1.0" }
[project.scripts]
second-brain = "second_brain.main:cli"
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["src/second_brain"]