pbs-projects/PBS/Tech/Projects/pbs-infrastructure-intelligence.md

481 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
project: pbs-infrastructure-intelligence
type: project-plan
status: active
path: PBS/Tech/Projects
tags:
- pbs
- monitoring
- docker
- go
- python
- mysql
- n8n
- streamlit
- security
created: 2026-04-16
updated: 2026-04-16
---
# PBS Infrastructure Intelligence (PBSII)
## Goal
Build a self-healing, intelligent monitoring system for the PBS
infrastructure. Four interconnected layers: a metrics collector, a
visualization dashboard, a threshold monitor, and an AI-powered healer.
## Problem Being Solved
- DoS attacks cause server thrashing with no early warning
- Container memory leaks go undetected until crashes occur
- Existing monitoring (Uptime Kuma) only detects hard-down states, not
degradation trends
- Manual log grepping is unsustainable as the site grows
- No visibility into what traffic reaches the server vs what Cloudflare
blocks
- Apache worker exhaustion causes memory crashes with no forensic data
## Architecture Overview
```
Layer 1: Collector (Go, systemd)
→ JSONL file (30s interval)
→ MySQL sync (every few minutes)
→ HTTP API (/metrics/current, /metrics/history, /health)
Log Parser (Python, n8n-scheduled)
→ docker logs --since
→ MySQL pbs_security_events + pbs_security_summary
Layer 2: Dashboard (Streamlit)
→ Reads MySQL
→ System metrics + security events + healer activity
Layer 3: Monitor (n8n, every 5 min)
→ Evaluates threshold rules against MySQL data
→ Fires warnings (Google Chat) or invokes healer
Layer 4: Healer (n8n + Claude API)
→ Gathers diagnostic context from all layers
→ Claude API reasons about the problem
→ Executes allowed actions or escalates
→ Logs all decisions to pbs_healer_actions
```
## Layer 1: Collector (Go Microservice)
### Purpose
Collect system and container metrics. Write to disk first, sync to MySQL
second. Serve current state via HTTP. No opinions, just data.
### Technology
- Language: Go (standard library first)
- Deployment: Single static binary, systemd service on host (not
containerized)
- Why systemd: Survives Docker daemon restarts, no Docker socket mount
needed for host metrics, resource-limited via unit file
### Metrics Collected
- Docker memory per container (replaces existing bash cron)
- Host total memory usage
- Host swap usage
- Host CPU usage
- Apache active worker count
### Collection Architecture
- Two goroutines in a single binary:
- **Collector goroutine:** Gathers metrics every 30 seconds, appends JSON
line to `/var/log/pbs-monitor/metrics.jsonl`. Cannot be slowed by MySQL.
- **Sync goroutine:** Reads new JSONL entries every few minutes, bulk
INSERTs to MySQL. If MySQL is unavailable, retries next cycle. File still
has the data.
- Log rotation via standard logrotate config
### HTTP API Endpoints
- `GET /metrics/current` — latest metrics snapshot
- `GET /metrics/history?range=1h` — recent history from MySQL
- `GET /health` — service health check
- Binds to localhost only; n8n reaches via Docker host bridge IP
### Data Retention
- Raw 30-second samples: 30 days in `pbs_system_metrics`
- Hourly rollup: kept indefinitely in `pbs_system_metrics_hourly`
- n8n handles both the hourly aggregation job and the 30-day cleanup job
### Systemd Unit File
```ini
[Unit]
Description=PBS Monitor - system metrics collector
After=network.target docker.service
[Service]
Type=simple
ExecStart=/usr/local/bin/pbs-monitor
Restart=on-failure
RestartSec=5s
User=pbs-monitor
Group=pbs-monitor
# Resource limits - blast radius containment
MemoryMax=256M
TasksMax=100
CPUQuota=20%
# Security hardening
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/var/log/pbs-monitor
[Install]
WantedBy=multi-user.target
```
### Ansible Deployment
- New Ansible role for systemd service deployment
- Tasks: copy binary, create config, create system user (in docker group),
install unit file, install logrotate config, enable and start service
- Follows existing staging-first deployment pattern
## Log Parser (Python)
### Purpose
Parse WordPress Docker logs for security events and request patterns. Feeds
both the dashboard and Layer 3 alert rules.
### Technology
- Language: Python (UV + venv, PyCharm Professional)
- Orchestration: n8n schedules execution every 15 minutes
- Log source: `docker logs wordpress --since` command (no container changes
required)
### Incremental Parsing
- Query `MAX(logged_at)` from `pbs_security_events` to determine where to
start
- Use `--since` flag on `docker logs` to avoid reprocessing
- Self-healing on restart — just picks up from last stored timestamp
### Fields Extracted
- `logged_at` — timestamp
- `ip_address` — source IP
- `method` — GET/POST
- `uri` — request path
- `status_code` — HTTP response code
- `response_size` — bytes sent
- `user_agent` — client user agent
- `event_type` — derived category (e.g., `login_attempt`, `404`, `normal`)
- Note: `response_time_ms` deferred — requires custom Apache LogFormat, not
needed for v1
### MySQL Tables
#### pbs_security_events (raw log entries)
```sql
CREATE TABLE pbs_security_events (
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
logged_at DATETIME NOT NULL,
ip_address VARCHAR(45) NOT NULL,
method VARCHAR(10),
uri TEXT,
status_code SMALLINT,
response_size INT,
user_agent TEXT,
event_type VARCHAR(50),
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
INDEX idx_logged_at (logged_at),
INDEX idx_ip (ip_address),
INDEX idx_uri (uri(100))
);
```
#### pbs_security_summary (hourly aggregates)
```sql
CREATE TABLE pbs_security_summary (
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
hour_bucket DATETIME NOT NULL,
total_requests INT,
unique_ips INT,
login_attempts INT,
login_page_loads INT,
status_2xx_count INT,
status_3xx_count INT,
status_4xx_count INT,
status_404_count INT,
status_5xx_count INT,
top_ip VARCHAR(45),
top_ip_count INT,
avg_requests_per_minute DECIMAL(10,2),
peak_requests_per_minute INT,
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
UNIQUE INDEX idx_hour (hour_bucket)
);
```
- Hourly rollup performed by n8n workflow (separate from the parser)
- Raw events retention: grow unbounded initially, add cleanup policy when
volume is understood
- Summary table retention: indefinite
## Layer 2: Dashboard (Streamlit)
### Purpose
Visualize system health and security data. Read-only, human-facing.
### Technology
- Streamlit, Python, reads from MySQL
- Charts via Plotly or Altair (time series)
### Pages
#### Page 1: System Metrics
- Time series graphs: Docker memory per container, host memory, swap, CPU,
Apache workers
- Horizontal threshold lines showing Layer 3 warning and healer trigger
levels
- Time range selector: 24h (default), 7 days, 30 days
#### Page 2: Security / Log Events
- wp-login.php attempts over time
- Top attacking IPs (table)
- Status code distribution
- Request volume trends (avg and peak requests/min)
#### Page 3: Healer Activity
- Log of healer invocations: trigger, decision, action, outcome
- Timeline view of actions taken
- Deferred actions and their resolution (executed, auto-cancelled,
overridden)
## Layer 3: Monitor (n8n)
### Purpose
Early warning system. Detects degradation trends before they become
outages. Evaluates rules, fires alerts or invokes the healer.
### Mechanism
- n8n workflow, scheduled every 5 minutes
- Queries MySQL for recent metrics
- Evaluates threshold rules
- Works alongside Uptime Kuma (not replacing it for now)
### Starter Rules
| Rule | Threshold | Action |
|---|---|---|
| Host RAM usage | >85% sustained 10 min | Google Chat warning |
| Host RAM usage | >95% sustained 5 min | Invoke healer |
| Swap usage | >50% of total swap | Google Chat warning |
| Container memory | >90% of cap sustained 5 min | Google Chat warning |
| Container memory | >98% of cap sustained 3 min | Invoke healer |
| Host CPU | >80% sustained 15 min | Google Chat warning (info only) |
| Apache workers | >80% of max sustained 10 min | Google Chat warning |
| wp-login.php POSTs | >30 in 5-min window | Google Chat warning |
"Sustained" = all samples in the window exceed the threshold. Prevents
single-spike false positives.
### Alert Tiers
- **Warning** — Google Chat notification, no action taken
- **Invoke Healer** — hands off to Layer 4 with trigger context
## Layer 4: Healer (n8n + Claude API)
### Purpose
Take corrective action when Layer 3 invokes it. Reason about the problem,
execute a safe response, report results.
### Mechanism
- n8n workflow triggered by Layer 3 or Uptime Kuma webhooks
- Portainer REST API for container operations
- Claude API for diagnostic reasoning
### Available Actions
| Action | Timing | Safety Rails |
|---|---|---|
| Container restart (non-protected) | Immediate | Max 1/hour, 3/24h per
container |
| WordPress restart | Deferred to quiet hours | Same caps + admin override
window |
| MySQL restart | Deferred to quiet hours | Same caps + admin override
window |
| Server reboot | Deferred to quiet hours | Max 1/24h + admin override
window |
| Alert only | Immediate | Always an option |
| Escalate | Immediate | When limits exceeded |
**Protected containers** (deferred to quiet hours only): WordPress, MySQL
### Quiet Hours
- Window: 02000500 ET
- Deferred actions are standing orders, not scheduled events
- During quiet hours, n8n checks: is there a pending action? Is it still
pending (not overridden)? Do the original conditions still exist?
- If conditions resolved → auto-cancel with Google Chat notification
- If overridden → mark as overridden
### Deferred Action Flow
1. Healer decides a protected container or server needs restart
2. Sets flag in `pbs_healer_actions`: `deferred_action`, `deferred_status =
'pending'`
3. Google Chat notification with action ID and reasoning
4. During quiet hours, n8n evaluates:
- Pending action exists? → Check if overridden → Check if conditions
persist → Execute or auto-cancel
5. Status transitions: `pending``executed` | `auto_cancelled` |
`overridden`
### Override Mechanism
#### V1
- **pbs-hub web page:** Shows pending healer actions, click to
cancel/approve
- **Google Chat webhook:** Text command with action ID (e.g., "override
1234", "cancel 1234")
- Both hit the same pbs-hub API endpoint
- pbs-hub API also available for future integrations
#### V2 (future)
- Google Chat interactive cards with approve/cancel/defer buttons
### Iteration Limits (Anti-Loop Protection)
- Per-container cooldown: 1 hour between automatic restarts
- Per-container daily cap: 3 automatic restarts per 24 hours
- Global healer cap: 10 healer-initiated actions per 24 hours
- Cooldown on escalation: 6 hours of no auto-action for that container
after escalation
- If any limit hit → escalate instead of acting
### Healer Decision Flow
1. Receive trigger from Layer 3 or Uptime Kuma
2. Check iteration limits → if exceeded, escalate and stop
3. Gather context:
- Current metrics from pbs-monitor HTTP API
- Recent log events from pbs_security_events
- Container inspect + recent logs from Portainer API
4. Send structured context to Claude API
5. Claude returns: recommended action, reasoning, confidence
6. If action is allowed AND confidence is high → execute (or defer if
protected)
7. Else → escalate with Claude's reasoning
8. Log decision + action + outcome to `pbs_healer_actions`
9. Google Chat notification with: what happened, what healer did (or
didn't), Claude's reasoning
### Claude API Prompt Shape
- Structured input: trigger reason, current metrics snapshot, 30-min
history, recent logs, container inspect
- Structured output: diagnosis, recommended action (enum:
`restart_container`, `defer_restart`, `reboot_server`, `no_action`,
`escalate`), reasoning, confidence level
### pbs_healer_actions Table
```sql
CREATE TABLE pbs_healer_actions (
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
triggered_at DATETIME NOT NULL,
trigger_source VARCHAR(50),
trigger_rule VARCHAR(100),
container_name VARCHAR(100),
metrics_snapshot JSON,
claude_reasoning TEXT,
action_recommended VARCHAR(50),
action_taken VARCHAR(50),
action_result TEXT,
deferred_action VARCHAR(50),
deferred_status VARCHAR(20) DEFAULT NULL,
iteration_limit_hit BOOLEAN DEFAULT FALSE,
escalated BOOLEAN DEFAULT FALSE,
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
INDEX idx_triggered (triggered_at),
INDEX idx_container (container_name),
INDEX idx_deferred (deferred_status)
);
```
### Starting Conservatism
- First period (target ~30 days): **dry-run mode** — healer makes
decisions, logs them, sends Google Chat with "I would have done X" but does
not execute
- After validation: flip to live mode, container restarts first
- Server reboot: enabled after container restart trust is established
- System reboot scheduling added only after override mechanism is validated
## Existing Infrastructure Context
- Docker/Traefik/Linode production stack
- Portainer running with REST API available
- Uptime Kuma operational for up/down checks and Google Chat alerts
- `pbs_automation` MySQL database exists with separate permissions
- n8n operational, Claude API integration is a known pattern
- Existing bash cron (60s, Docker memory only, flat file) — replaced by
Layer 1
- Crowdsec security hardening is a separate project (not duplicated here)
## Implementation Order
- [ ] Phase 1: Layer 1 — Go collector (systemd, JSONL, MySQL sync, HTTP API)
- [ ] Phase 2: Log Parser — Python WordPress log parser + MySQL schema
- [ ] Phase 3: Layer 2 — Streamlit dashboard (system metrics + security +
thresholds)
- [ ] Phase 4: Layer 3 — n8n monitor workflow (threshold rules, Google Chat
alerts)
- [ ] Phase 5: Layer 4 — n8n healer workflow (dry-run mode, Claude API
integration)
- [ ] Phase 6: Layer 4 live — enable live actions, override page in pbs-hub
- [ ] Phase 7: Deferred actions — quiet hours logic, WordPress/MySQL/server
restart scheduling
## Open Questions
- [ ] Go project structure — standard layout or flat? Decide during Phase 1
- [ ] Docker multi-stage build for Go binary — needed for CI, decide during
Phase 1
- [ ] MySQL schema for `pbs_system_metrics` and `pbs_system_metrics_hourly`
— design during Phase 1
- [ ] Apache max_workers value — needed to calculate percentage for Layer 3
rules
- [ ] Portainer API auth — how is it currently secured? Needed for Layer 4
- [ ] Claude API prompt engineering — detailed prompt design during Phase 5
- [ ] pbs-hub override page routing — how does it fit into existing pbs-hub
Flask app?
- [ ] Google Chat webhook for override commands — new webhook or extend
existing?
- [ ] Streamlit deployment — containerized or host-level? Decide during
Phase 3
...sent from Jenny & Travis