wiki-vault/Sources/Dev/pbs-api-db-connection-fix.md
Travis Herbranson 34a268d8dc migration: copy 62 notes from pbs-projects and homelab-projects
Two-layer structure: Sources (raw notes) + Wiki (compile output)
Four domains: Dev (40), Venture (3), Homelab (23), Reference (0)
Includes CLAUDE.md spec, index pages at all levels, compile log

Co-Authored-By: Lovebug <lovebug@herbylab.dev>
2026-05-08 17:21:13 -04:00

2.9 KiB

created path project status tags type updated
2026-04-15 Sources/Dev pbs-api-db-connection-fix active
pbs
pbs-api
flask
mysql
sqlalchemy
docker
session-notes 2026-04-15

PBS API — DB Connection Pool Fix

Problem

The pbs-api container was periodically appearing unavailable (content hub inaccessible). The container was not crashing — it was being marked unhealthy by Docker, causing Traefik to pull it from rotation.

Root cause: SQLAlchemy connection pool holding idle MySQL connections that MySQL had already closed server-side. On next use, the dead connection produces a broken pipe error. If this occurs during a healthcheck window, the health endpoint fails 3x and Traefik stops routing traffic.

Confirmed via logs:

  • pymysql.err.OperationalError: (2006, "MySQL server has gone away (BrokenPipeError(32, 'Broken pipe'))")

Container state confirmed clean — exit code 0, no OOM kill, restart policy unless-stopped.

Fix

Update get_engine() in the database module to add connection pool resilience for MySQL connections.

Change: get_engine() function

Locate the get_engine() function (currently uses bare create_engine(url, echo=False) for all DB types) and replace with the following:

def get_engine() -> Engine:
    global _engine
    if _engine is None:
        url = get_database_url()
        if url.startswith("sqlite"):
            from pathlib import Path
            db_path = Path(url.replace("sqlite:///", ""))
            db_path.parent.mkdir(parents=True, exist_ok=True)
            _engine = create_engine(url, echo=False)
        else:
            _engine = create_engine(
                url,
                echo=False,
                pool_pre_ping=True,   # test connection before use,
transparently reconnects dead ones
                pool_recycle=1800,    # proactively recycle
connections every 30 min
            )
    return _engine

Why this works

  • pool_pre_ping=True — issues a lightweight SELECT 1 before each query. If the connection is dead, SQLAlchemy silently replaces it. This is the self-healing mechanism.
  • pool_recycle=1800 — proactively replaces connections older than 30 minutes, before MySQL's wait_timeout closes them server-side.
  • SQLite path is unchanged — pooling does not apply.

Outcome

With pool_pre_ping in place, dead connections are detected and replaced transparently before they can cause a request failure or health check miss. No container restart required.

Notes

  • No changes needed to the /api/health endpoint — it is a simple 200 and does not touch the DB, which is correct.
  • Docker's unless-stopped restart policy does not auto-restart on healthcheck failure — this is by design and does not need to change given the root cause is now addressed.
  • Autoheal (docker-autoheal) was considered and deferred — not needed with pool_pre_ping in place.