481 lines
15 KiB
Markdown
481 lines
15 KiB
Markdown
---
|
||
project: pbs-infrastructure-intelligence
|
||
type: project-plan
|
||
status: active
|
||
path: PBS/Tech/Projects
|
||
tags:
|
||
- pbs
|
||
- monitoring
|
||
- docker
|
||
- go
|
||
- python
|
||
- mysql
|
||
- n8n
|
||
- streamlit
|
||
- security
|
||
created: 2026-04-16
|
||
updated: 2026-04-16
|
||
---
|
||
|
||
# PBS Infrastructure Intelligence (PBSII)
|
||
|
||
## Goal
|
||
|
||
Build a self-healing, intelligent monitoring system for the PBS
|
||
infrastructure. Four interconnected layers: a metrics collector, a
|
||
visualization dashboard, a threshold monitor, and an AI-powered healer.
|
||
|
||
## Problem Being Solved
|
||
|
||
- DoS attacks cause server thrashing with no early warning
|
||
- Container memory leaks go undetected until crashes occur
|
||
- Existing monitoring (Uptime Kuma) only detects hard-down states, not
|
||
degradation trends
|
||
- Manual log grepping is unsustainable as the site grows
|
||
- No visibility into what traffic reaches the server vs what Cloudflare
|
||
blocks
|
||
- Apache worker exhaustion causes memory crashes with no forensic data
|
||
|
||
## Architecture Overview
|
||
|
||
```
|
||
Layer 1: Collector (Go, systemd)
|
||
→ JSONL file (30s interval)
|
||
→ MySQL sync (every few minutes)
|
||
→ HTTP API (/metrics/current, /metrics/history, /health)
|
||
|
||
Log Parser (Python, n8n-scheduled)
|
||
→ docker logs --since
|
||
→ MySQL pbs_security_events + pbs_security_summary
|
||
|
||
Layer 2: Dashboard (Streamlit)
|
||
→ Reads MySQL
|
||
→ System metrics + security events + healer activity
|
||
|
||
Layer 3: Monitor (n8n, every 5 min)
|
||
→ Evaluates threshold rules against MySQL data
|
||
→ Fires warnings (Google Chat) or invokes healer
|
||
|
||
Layer 4: Healer (n8n + Claude API)
|
||
→ Gathers diagnostic context from all layers
|
||
→ Claude API reasons about the problem
|
||
→ Executes allowed actions or escalates
|
||
→ Logs all decisions to pbs_healer_actions
|
||
```
|
||
|
||
## Layer 1: Collector (Go Microservice)
|
||
|
||
### Purpose
|
||
|
||
Collect system and container metrics. Write to disk first, sync to MySQL
|
||
second. Serve current state via HTTP. No opinions, just data.
|
||
|
||
### Technology
|
||
|
||
- Language: Go (standard library first)
|
||
- Deployment: Single static binary, systemd service on host (not
|
||
containerized)
|
||
- Why systemd: Survives Docker daemon restarts, no Docker socket mount
|
||
needed for host metrics, resource-limited via unit file
|
||
|
||
### Metrics Collected
|
||
|
||
- Docker memory per container (replaces existing bash cron)
|
||
- Host total memory usage
|
||
- Host swap usage
|
||
- Host CPU usage
|
||
- Apache active worker count
|
||
|
||
### Collection Architecture
|
||
|
||
- Two goroutines in a single binary:
|
||
- **Collector goroutine:** Gathers metrics every 30 seconds, appends JSON
|
||
line to `/var/log/pbs-monitor/metrics.jsonl`. Cannot be slowed by MySQL.
|
||
- **Sync goroutine:** Reads new JSONL entries every few minutes, bulk
|
||
INSERTs to MySQL. If MySQL is unavailable, retries next cycle. File still
|
||
has the data.
|
||
- Log rotation via standard logrotate config
|
||
|
||
### HTTP API Endpoints
|
||
|
||
- `GET /metrics/current` — latest metrics snapshot
|
||
- `GET /metrics/history?range=1h` — recent history from MySQL
|
||
- `GET /health` — service health check
|
||
- Binds to localhost only; n8n reaches via Docker host bridge IP
|
||
|
||
### Data Retention
|
||
|
||
- Raw 30-second samples: 30 days in `pbs_system_metrics`
|
||
- Hourly rollup: kept indefinitely in `pbs_system_metrics_hourly`
|
||
- n8n handles both the hourly aggregation job and the 30-day cleanup job
|
||
|
||
### Systemd Unit File
|
||
|
||
```ini
|
||
[Unit]
|
||
Description=PBS Monitor - system metrics collector
|
||
After=network.target docker.service
|
||
|
||
[Service]
|
||
Type=simple
|
||
ExecStart=/usr/local/bin/pbs-monitor
|
||
Restart=on-failure
|
||
RestartSec=5s
|
||
User=pbs-monitor
|
||
Group=pbs-monitor
|
||
|
||
# Resource limits - blast radius containment
|
||
MemoryMax=256M
|
||
TasksMax=100
|
||
CPUQuota=20%
|
||
|
||
# Security hardening
|
||
NoNewPrivileges=true
|
||
ProtectSystem=strict
|
||
ProtectHome=true
|
||
ReadWritePaths=/var/log/pbs-monitor
|
||
|
||
[Install]
|
||
WantedBy=multi-user.target
|
||
```
|
||
|
||
### Ansible Deployment
|
||
|
||
- New Ansible role for systemd service deployment
|
||
- Tasks: copy binary, create config, create system user (in docker group),
|
||
install unit file, install logrotate config, enable and start service
|
||
- Follows existing staging-first deployment pattern
|
||
|
||
## Log Parser (Python)
|
||
|
||
### Purpose
|
||
|
||
Parse WordPress Docker logs for security events and request patterns. Feeds
|
||
both the dashboard and Layer 3 alert rules.
|
||
|
||
### Technology
|
||
|
||
- Language: Python (UV + venv, PyCharm Professional)
|
||
- Orchestration: n8n schedules execution every 15 minutes
|
||
- Log source: `docker logs wordpress --since` command (no container changes
|
||
required)
|
||
|
||
### Incremental Parsing
|
||
|
||
- Query `MAX(logged_at)` from `pbs_security_events` to determine where to
|
||
start
|
||
- Use `--since` flag on `docker logs` to avoid reprocessing
|
||
- Self-healing on restart — just picks up from last stored timestamp
|
||
|
||
### Fields Extracted
|
||
|
||
- `logged_at` — timestamp
|
||
- `ip_address` — source IP
|
||
- `method` — GET/POST
|
||
- `uri` — request path
|
||
- `status_code` — HTTP response code
|
||
- `response_size` — bytes sent
|
||
- `user_agent` — client user agent
|
||
- `event_type` — derived category (e.g., `login_attempt`, `404`, `normal`)
|
||
- Note: `response_time_ms` deferred — requires custom Apache LogFormat, not
|
||
needed for v1
|
||
|
||
### MySQL Tables
|
||
|
||
#### pbs_security_events (raw log entries)
|
||
|
||
```sql
|
||
CREATE TABLE pbs_security_events (
|
||
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
|
||
logged_at DATETIME NOT NULL,
|
||
ip_address VARCHAR(45) NOT NULL,
|
||
method VARCHAR(10),
|
||
uri TEXT,
|
||
status_code SMALLINT,
|
||
response_size INT,
|
||
user_agent TEXT,
|
||
event_type VARCHAR(50),
|
||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||
INDEX idx_logged_at (logged_at),
|
||
INDEX idx_ip (ip_address),
|
||
INDEX idx_uri (uri(100))
|
||
);
|
||
```
|
||
|
||
#### pbs_security_summary (hourly aggregates)
|
||
|
||
```sql
|
||
CREATE TABLE pbs_security_summary (
|
||
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
|
||
hour_bucket DATETIME NOT NULL,
|
||
total_requests INT,
|
||
unique_ips INT,
|
||
login_attempts INT,
|
||
login_page_loads INT,
|
||
status_2xx_count INT,
|
||
status_3xx_count INT,
|
||
status_4xx_count INT,
|
||
status_404_count INT,
|
||
status_5xx_count INT,
|
||
top_ip VARCHAR(45),
|
||
top_ip_count INT,
|
||
avg_requests_per_minute DECIMAL(10,2),
|
||
peak_requests_per_minute INT,
|
||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||
UNIQUE INDEX idx_hour (hour_bucket)
|
||
);
|
||
```
|
||
|
||
- Hourly rollup performed by n8n workflow (separate from the parser)
|
||
- Raw events retention: grow unbounded initially, add cleanup policy when
|
||
volume is understood
|
||
- Summary table retention: indefinite
|
||
|
||
## Layer 2: Dashboard (Streamlit)
|
||
|
||
### Purpose
|
||
|
||
Visualize system health and security data. Read-only, human-facing.
|
||
|
||
### Technology
|
||
|
||
- Streamlit, Python, reads from MySQL
|
||
- Charts via Plotly or Altair (time series)
|
||
|
||
### Pages
|
||
|
||
#### Page 1: System Metrics
|
||
|
||
- Time series graphs: Docker memory per container, host memory, swap, CPU,
|
||
Apache workers
|
||
- Horizontal threshold lines showing Layer 3 warning and healer trigger
|
||
levels
|
||
- Time range selector: 24h (default), 7 days, 30 days
|
||
|
||
#### Page 2: Security / Log Events
|
||
|
||
- wp-login.php attempts over time
|
||
- Top attacking IPs (table)
|
||
- Status code distribution
|
||
- Request volume trends (avg and peak requests/min)
|
||
|
||
#### Page 3: Healer Activity
|
||
|
||
- Log of healer invocations: trigger, decision, action, outcome
|
||
- Timeline view of actions taken
|
||
- Deferred actions and their resolution (executed, auto-cancelled,
|
||
overridden)
|
||
|
||
## Layer 3: Monitor (n8n)
|
||
|
||
### Purpose
|
||
|
||
Early warning system. Detects degradation trends before they become
|
||
outages. Evaluates rules, fires alerts or invokes the healer.
|
||
|
||
### Mechanism
|
||
|
||
- n8n workflow, scheduled every 5 minutes
|
||
- Queries MySQL for recent metrics
|
||
- Evaluates threshold rules
|
||
- Works alongside Uptime Kuma (not replacing it for now)
|
||
|
||
### Starter Rules
|
||
|
||
| Rule | Threshold | Action |
|
||
|---|---|---|
|
||
| Host RAM usage | >85% sustained 10 min | Google Chat warning |
|
||
| Host RAM usage | >95% sustained 5 min | Invoke healer |
|
||
| Swap usage | >50% of total swap | Google Chat warning |
|
||
| Container memory | >90% of cap sustained 5 min | Google Chat warning |
|
||
| Container memory | >98% of cap sustained 3 min | Invoke healer |
|
||
| Host CPU | >80% sustained 15 min | Google Chat warning (info only) |
|
||
| Apache workers | >80% of max sustained 10 min | Google Chat warning |
|
||
| wp-login.php POSTs | >30 in 5-min window | Google Chat warning |
|
||
|
||
"Sustained" = all samples in the window exceed the threshold. Prevents
|
||
single-spike false positives.
|
||
|
||
### Alert Tiers
|
||
|
||
- **Warning** — Google Chat notification, no action taken
|
||
- **Invoke Healer** — hands off to Layer 4 with trigger context
|
||
|
||
## Layer 4: Healer (n8n + Claude API)
|
||
|
||
### Purpose
|
||
|
||
Take corrective action when Layer 3 invokes it. Reason about the problem,
|
||
execute a safe response, report results.
|
||
|
||
### Mechanism
|
||
|
||
- n8n workflow triggered by Layer 3 or Uptime Kuma webhooks
|
||
- Portainer REST API for container operations
|
||
- Claude API for diagnostic reasoning
|
||
|
||
### Available Actions
|
||
|
||
| Action | Timing | Safety Rails |
|
||
|---|---|---|
|
||
| Container restart (non-protected) | Immediate | Max 1/hour, 3/24h per
|
||
container |
|
||
| WordPress restart | Deferred to quiet hours | Same caps + admin override
|
||
window |
|
||
| MySQL restart | Deferred to quiet hours | Same caps + admin override
|
||
window |
|
||
| Server reboot | Deferred to quiet hours | Max 1/24h + admin override
|
||
window |
|
||
| Alert only | Immediate | Always an option |
|
||
| Escalate | Immediate | When limits exceeded |
|
||
|
||
**Protected containers** (deferred to quiet hours only): WordPress, MySQL
|
||
|
||
### Quiet Hours
|
||
|
||
- Window: 0200–0500 ET
|
||
- Deferred actions are standing orders, not scheduled events
|
||
- During quiet hours, n8n checks: is there a pending action? Is it still
|
||
pending (not overridden)? Do the original conditions still exist?
|
||
- If conditions resolved → auto-cancel with Google Chat notification
|
||
- If overridden → mark as overridden
|
||
|
||
### Deferred Action Flow
|
||
|
||
1. Healer decides a protected container or server needs restart
|
||
2. Sets flag in `pbs_healer_actions`: `deferred_action`, `deferred_status =
|
||
'pending'`
|
||
3. Google Chat notification with action ID and reasoning
|
||
4. During quiet hours, n8n evaluates:
|
||
- Pending action exists? → Check if overridden → Check if conditions
|
||
persist → Execute or auto-cancel
|
||
5. Status transitions: `pending` → `executed` | `auto_cancelled` |
|
||
`overridden`
|
||
|
||
### Override Mechanism
|
||
|
||
#### V1
|
||
|
||
- **pbs-hub web page:** Shows pending healer actions, click to
|
||
cancel/approve
|
||
- **Google Chat webhook:** Text command with action ID (e.g., "override
|
||
1234", "cancel 1234")
|
||
- Both hit the same pbs-hub API endpoint
|
||
- pbs-hub API also available for future integrations
|
||
|
||
#### V2 (future)
|
||
|
||
- Google Chat interactive cards with approve/cancel/defer buttons
|
||
|
||
### Iteration Limits (Anti-Loop Protection)
|
||
|
||
- Per-container cooldown: 1 hour between automatic restarts
|
||
- Per-container daily cap: 3 automatic restarts per 24 hours
|
||
- Global healer cap: 10 healer-initiated actions per 24 hours
|
||
- Cooldown on escalation: 6 hours of no auto-action for that container
|
||
after escalation
|
||
- If any limit hit → escalate instead of acting
|
||
|
||
### Healer Decision Flow
|
||
|
||
1. Receive trigger from Layer 3 or Uptime Kuma
|
||
2. Check iteration limits → if exceeded, escalate and stop
|
||
3. Gather context:
|
||
- Current metrics from pbs-monitor HTTP API
|
||
- Recent log events from pbs_security_events
|
||
- Container inspect + recent logs from Portainer API
|
||
4. Send structured context to Claude API
|
||
5. Claude returns: recommended action, reasoning, confidence
|
||
6. If action is allowed AND confidence is high → execute (or defer if
|
||
protected)
|
||
7. Else → escalate with Claude's reasoning
|
||
8. Log decision + action + outcome to `pbs_healer_actions`
|
||
9. Google Chat notification with: what happened, what healer did (or
|
||
didn't), Claude's reasoning
|
||
|
||
### Claude API Prompt Shape
|
||
|
||
- Structured input: trigger reason, current metrics snapshot, 30-min
|
||
history, recent logs, container inspect
|
||
- Structured output: diagnosis, recommended action (enum:
|
||
`restart_container`, `defer_restart`, `reboot_server`, `no_action`,
|
||
`escalate`), reasoning, confidence level
|
||
|
||
### pbs_healer_actions Table
|
||
|
||
```sql
|
||
CREATE TABLE pbs_healer_actions (
|
||
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
|
||
triggered_at DATETIME NOT NULL,
|
||
trigger_source VARCHAR(50),
|
||
trigger_rule VARCHAR(100),
|
||
container_name VARCHAR(100),
|
||
metrics_snapshot JSON,
|
||
claude_reasoning TEXT,
|
||
action_recommended VARCHAR(50),
|
||
action_taken VARCHAR(50),
|
||
action_result TEXT,
|
||
deferred_action VARCHAR(50),
|
||
deferred_status VARCHAR(20) DEFAULT NULL,
|
||
iteration_limit_hit BOOLEAN DEFAULT FALSE,
|
||
escalated BOOLEAN DEFAULT FALSE,
|
||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||
INDEX idx_triggered (triggered_at),
|
||
INDEX idx_container (container_name),
|
||
INDEX idx_deferred (deferred_status)
|
||
);
|
||
```
|
||
|
||
### Starting Conservatism
|
||
|
||
- First period (target ~30 days): **dry-run mode** — healer makes
|
||
decisions, logs them, sends Google Chat with "I would have done X" but does
|
||
not execute
|
||
- After validation: flip to live mode, container restarts first
|
||
- Server reboot: enabled after container restart trust is established
|
||
- System reboot scheduling added only after override mechanism is validated
|
||
|
||
## Existing Infrastructure Context
|
||
|
||
- Docker/Traefik/Linode production stack
|
||
- Portainer running with REST API available
|
||
- Uptime Kuma operational for up/down checks and Google Chat alerts
|
||
- `pbs_automation` MySQL database exists with separate permissions
|
||
- n8n operational, Claude API integration is a known pattern
|
||
- Existing bash cron (60s, Docker memory only, flat file) — replaced by
|
||
Layer 1
|
||
- Crowdsec security hardening is a separate project (not duplicated here)
|
||
|
||
## Implementation Order
|
||
|
||
- [ ] Phase 1: Layer 1 — Go collector (systemd, JSONL, MySQL sync, HTTP API)
|
||
- [ ] Phase 2: Log Parser — Python WordPress log parser + MySQL schema
|
||
- [ ] Phase 3: Layer 2 — Streamlit dashboard (system metrics + security +
|
||
thresholds)
|
||
- [ ] Phase 4: Layer 3 — n8n monitor workflow (threshold rules, Google Chat
|
||
alerts)
|
||
- [ ] Phase 5: Layer 4 — n8n healer workflow (dry-run mode, Claude API
|
||
integration)
|
||
- [ ] Phase 6: Layer 4 live — enable live actions, override page in pbs-hub
|
||
- [ ] Phase 7: Deferred actions — quiet hours logic, WordPress/MySQL/server
|
||
restart scheduling
|
||
|
||
## Open Questions
|
||
|
||
- [ ] Go project structure — standard layout or flat? Decide during Phase 1
|
||
- [ ] Docker multi-stage build for Go binary — needed for CI, decide during
|
||
Phase 1
|
||
- [ ] MySQL schema for `pbs_system_metrics` and `pbs_system_metrics_hourly`
|
||
— design during Phase 1
|
||
- [ ] Apache max_workers value — needed to calculate percentage for Layer 3
|
||
rules
|
||
- [ ] Portainer API auth — how is it currently secured? Needed for Layer 4
|
||
- [ ] Claude API prompt engineering — detailed prompt design during Phase 5
|
||
- [ ] pbs-hub override page routing — how does it fit into existing pbs-hub
|
||
Flask app?
|
||
- [ ] Google Chat webhook for override commands — new webhook or extend
|
||
existing?
|
||
- [ ] Streamlit deployment — containerized or host-level? Decide during
|
||
Phase 3
|
||
|
||
...sent from Jenny & Travis |