2.9 KiB
| project | type | status | path | tags | created | updated | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| pbs-api-db-connection-fix | session-notes | active | Tech/Sessions |
|
2026-04-15 | 2026-04-15 |
PBS API — DB Connection Pool Fix
Problem
The pbs-api container was periodically appearing unavailable (content hub inaccessible). The container was not crashing — it was being marked unhealthy by Docker, causing Traefik to pull it from rotation.
Root cause: SQLAlchemy connection pool holding idle MySQL connections that MySQL had already closed server-side. On next use, the dead connection produces a broken pipe error. If this occurs during a healthcheck window, the health endpoint fails 3x and Traefik stops routing traffic.
Confirmed via logs:
pymysql.err.OperationalError: (2006, "MySQL server has gone away (BrokenPipeError(32, 'Broken pipe'))")
Container state confirmed clean — exit code 0, no OOM kill, restart
policy unless-stopped.
Fix
Update get_engine() in the database module to add connection pool
resilience for MySQL connections.
Change: get_engine() function
Locate the get_engine() function (currently uses bare
create_engine(url, echo=False) for all DB types) and replace with
the following:
def get_engine() -> Engine:
global _engine
if _engine is None:
url = get_database_url()
if url.startswith("sqlite"):
from pathlib import Path
db_path = Path(url.replace("sqlite:///", ""))
db_path.parent.mkdir(parents=True, exist_ok=True)
_engine = create_engine(url, echo=False)
else:
_engine = create_engine(
url,
echo=False,
pool_pre_ping=True, # test connection before use,
transparently reconnects dead ones
pool_recycle=1800, # proactively recycle
connections every 30 min
)
return _engine
Why this works
pool_pre_ping=True— issues a lightweightSELECT 1before each query. If the connection is dead, SQLAlchemy silently replaces it. This is the self-healing mechanism.pool_recycle=1800— proactively replaces connections older than 30 minutes, before MySQL'swait_timeoutcloses them server-side.- SQLite path is unchanged — pooling does not apply.
Outcome
With pool_pre_ping in place, dead connections are detected and
replaced transparently before they can cause a request failure or
health check miss. No container restart required.
Notes
- No changes needed to the
/api/healthendpoint — it is a simple 200 and does not touch the DB, which is correct. - Docker's
unless-stoppedrestart policy does not auto-restart on healthcheck failure — this is by design and does not need to change given the root cause is now addressed. - Autoheal (docker-autoheal) was considered and deferred — not needed with pool_pre_ping in place.