Create pbs-api-db-connection-fix.md via n8n
This commit is contained in:
parent
953aaa1cef
commit
2b4905f8fd
94
Tech/Sessions/pbs-api-db-connection-fix.md
Normal file
94
Tech/Sessions/pbs-api-db-connection-fix.md
Normal file
@ -0,0 +1,94 @@
|
||||
---
|
||||
project: pbs-api-db-connection-fix
|
||||
type: session-notes
|
||||
status: active
|
||||
path: Tech/Sessions
|
||||
tags:
|
||||
- pbs
|
||||
- pbs-api
|
||||
- flask
|
||||
- mysql
|
||||
- sqlalchemy
|
||||
- docker
|
||||
created: 2026-04-15
|
||||
updated: 2026-04-15
|
||||
---
|
||||
|
||||
# PBS API — DB Connection Pool Fix
|
||||
|
||||
## Problem
|
||||
|
||||
The pbs-api container was periodically appearing unavailable (content
|
||||
hub inaccessible). The container was not crashing — it was being
|
||||
marked unhealthy by Docker, causing Traefik to pull it from rotation.
|
||||
|
||||
Root cause: SQLAlchemy connection pool holding idle MySQL connections
|
||||
that MySQL had already closed server-side. On next use, the dead
|
||||
connection produces a broken pipe error. If this occurs during a
|
||||
healthcheck window, the health endpoint fails 3x and Traefik stops
|
||||
routing traffic.
|
||||
|
||||
Confirmed via logs:
|
||||
- `pymysql.err.OperationalError: (2006, "MySQL server has gone away
|
||||
(BrokenPipeError(32, 'Broken pipe'))")`
|
||||
|
||||
Container state confirmed clean — exit code 0, no OOM kill, restart
|
||||
policy `unless-stopped`.
|
||||
|
||||
## Fix
|
||||
|
||||
Update `get_engine()` in the database module to add connection pool
|
||||
resilience for MySQL connections.
|
||||
|
||||
### Change: `get_engine()` function
|
||||
|
||||
Locate the `get_engine()` function (currently uses bare
|
||||
`create_engine(url, echo=False)` for all DB types) and replace with
|
||||
the following:
|
||||
|
||||
```python
|
||||
def get_engine() -> Engine:
|
||||
global _engine
|
||||
if _engine is None:
|
||||
url = get_database_url()
|
||||
if url.startswith("sqlite"):
|
||||
from pathlib import Path
|
||||
db_path = Path(url.replace("sqlite:///", ""))
|
||||
db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
_engine = create_engine(url, echo=False)
|
||||
else:
|
||||
_engine = create_engine(
|
||||
url,
|
||||
echo=False,
|
||||
pool_pre_ping=True, # test connection before use,
|
||||
transparently reconnects dead ones
|
||||
pool_recycle=1800, # proactively recycle
|
||||
connections every 30 min
|
||||
)
|
||||
return _engine
|
||||
```
|
||||
|
||||
### Why this works
|
||||
|
||||
- `pool_pre_ping=True` — issues a lightweight `SELECT 1` before each
|
||||
query. If the connection is dead, SQLAlchemy silently replaces it.
|
||||
This is the self-healing mechanism.
|
||||
- `pool_recycle=1800` — proactively replaces connections older than 30
|
||||
minutes, before MySQL's `wait_timeout` closes them server-side.
|
||||
- SQLite path is unchanged — pooling does not apply.
|
||||
|
||||
## Outcome
|
||||
|
||||
With `pool_pre_ping` in place, dead connections are detected and
|
||||
replaced transparently before they can cause a request failure or
|
||||
health check miss. No container restart required.
|
||||
|
||||
## Notes
|
||||
|
||||
- No changes needed to the `/api/health` endpoint — it is a simple 200
|
||||
and does not touch the DB, which is correct.
|
||||
- Docker's `unless-stopped` restart policy does not auto-restart on
|
||||
healthcheck failure — this is by design and does not need to change
|
||||
given the root cause is now addressed.
|
||||
- Autoheal (docker-autoheal) was considered and deferred — not needed
|
||||
with pool_pre_ping in place.
|
||||
Loading…
Reference in New Issue
Block a user