diff --git a/PBS/Tech/Projects/pbs-infrastructure-intelligence.md b/PBS/Tech/Projects/pbs-infrastructure-intelligence.md new file mode 100644 index 0000000..72a9d7c --- /dev/null +++ b/PBS/Tech/Projects/pbs-infrastructure-intelligence.md @@ -0,0 +1,481 @@ +--- +project: pbs-infrastructure-intelligence +type: project-plan +status: active +path: PBS/Tech/Projects +tags: + - pbs + - monitoring + - docker + - go + - python + - mysql + - n8n + - streamlit + - security +created: 2026-04-16 +updated: 2026-04-16 +--- + +# PBS Infrastructure Intelligence (PBSII) + +## Goal + +Build a self-healing, intelligent monitoring system for the PBS +infrastructure. Four interconnected layers: a metrics collector, a +visualization dashboard, a threshold monitor, and an AI-powered healer. + +## Problem Being Solved + +- DoS attacks cause server thrashing with no early warning +- Container memory leaks go undetected until crashes occur +- Existing monitoring (Uptime Kuma) only detects hard-down states, not +degradation trends +- Manual log grepping is unsustainable as the site grows +- No visibility into what traffic reaches the server vs what Cloudflare +blocks +- Apache worker exhaustion causes memory crashes with no forensic data + +## Architecture Overview + +``` +Layer 1: Collector (Go, systemd) + → JSONL file (30s interval) + → MySQL sync (every few minutes) + → HTTP API (/metrics/current, /metrics/history, /health) + +Log Parser (Python, n8n-scheduled) + → docker logs --since + → MySQL pbs_security_events + pbs_security_summary + +Layer 2: Dashboard (Streamlit) + → Reads MySQL + → System metrics + security events + healer activity + +Layer 3: Monitor (n8n, every 5 min) + → Evaluates threshold rules against MySQL data + → Fires warnings (Google Chat) or invokes healer + +Layer 4: Healer (n8n + Claude API) + → Gathers diagnostic context from all layers + → Claude API reasons about the problem + → Executes allowed actions or escalates + → Logs all decisions to pbs_healer_actions +``` + +## Layer 1: Collector (Go Microservice) + +### Purpose + +Collect system and container metrics. Write to disk first, sync to MySQL +second. Serve current state via HTTP. No opinions, just data. + +### Technology + +- Language: Go (standard library first) +- Deployment: Single static binary, systemd service on host (not +containerized) +- Why systemd: Survives Docker daemon restarts, no Docker socket mount +needed for host metrics, resource-limited via unit file + +### Metrics Collected + +- Docker memory per container (replaces existing bash cron) +- Host total memory usage +- Host swap usage +- Host CPU usage +- Apache active worker count + +### Collection Architecture + +- Two goroutines in a single binary: + - **Collector goroutine:** Gathers metrics every 30 seconds, appends JSON +line to `/var/log/pbs-monitor/metrics.jsonl`. Cannot be slowed by MySQL. + - **Sync goroutine:** Reads new JSONL entries every few minutes, bulk +INSERTs to MySQL. If MySQL is unavailable, retries next cycle. File still +has the data. +- Log rotation via standard logrotate config + +### HTTP API Endpoints + +- `GET /metrics/current` — latest metrics snapshot +- `GET /metrics/history?range=1h` — recent history from MySQL +- `GET /health` — service health check +- Binds to localhost only; n8n reaches via Docker host bridge IP + +### Data Retention + +- Raw 30-second samples: 30 days in `pbs_system_metrics` +- Hourly rollup: kept indefinitely in `pbs_system_metrics_hourly` +- n8n handles both the hourly aggregation job and the 30-day cleanup job + +### Systemd Unit File + +```ini +[Unit] +Description=PBS Monitor - system metrics collector +After=network.target docker.service + +[Service] +Type=simple +ExecStart=/usr/local/bin/pbs-monitor +Restart=on-failure +RestartSec=5s +User=pbs-monitor +Group=pbs-monitor + +# Resource limits - blast radius containment +MemoryMax=256M +TasksMax=100 +CPUQuota=20% + +# Security hardening +NoNewPrivileges=true +ProtectSystem=strict +ProtectHome=true +ReadWritePaths=/var/log/pbs-monitor + +[Install] +WantedBy=multi-user.target +``` + +### Ansible Deployment + +- New Ansible role for systemd service deployment +- Tasks: copy binary, create config, create system user (in docker group), +install unit file, install logrotate config, enable and start service +- Follows existing staging-first deployment pattern + +## Log Parser (Python) + +### Purpose + +Parse WordPress Docker logs for security events and request patterns. Feeds +both the dashboard and Layer 3 alert rules. + +### Technology + +- Language: Python (UV + venv, PyCharm Professional) +- Orchestration: n8n schedules execution every 15 minutes +- Log source: `docker logs wordpress --since` command (no container changes +required) + +### Incremental Parsing + +- Query `MAX(logged_at)` from `pbs_security_events` to determine where to +start +- Use `--since` flag on `docker logs` to avoid reprocessing +- Self-healing on restart — just picks up from last stored timestamp + +### Fields Extracted + +- `logged_at` — timestamp +- `ip_address` — source IP +- `method` — GET/POST +- `uri` — request path +- `status_code` — HTTP response code +- `response_size` — bytes sent +- `user_agent` — client user agent +- `event_type` — derived category (e.g., `login_attempt`, `404`, `normal`) +- Note: `response_time_ms` deferred — requires custom Apache LogFormat, not +needed for v1 + +### MySQL Tables + +#### pbs_security_events (raw log entries) + +```sql +CREATE TABLE pbs_security_events ( + id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY, + logged_at DATETIME NOT NULL, + ip_address VARCHAR(45) NOT NULL, + method VARCHAR(10), + uri TEXT, + status_code SMALLINT, + response_size INT, + user_agent TEXT, + event_type VARCHAR(50), + created_at DATETIME DEFAULT CURRENT_TIMESTAMP, + INDEX idx_logged_at (logged_at), + INDEX idx_ip (ip_address), + INDEX idx_uri (uri(100)) +); +``` + +#### pbs_security_summary (hourly aggregates) + +```sql +CREATE TABLE pbs_security_summary ( + id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY, + hour_bucket DATETIME NOT NULL, + total_requests INT, + unique_ips INT, + login_attempts INT, + login_page_loads INT, + status_2xx_count INT, + status_3xx_count INT, + status_4xx_count INT, + status_404_count INT, + status_5xx_count INT, + top_ip VARCHAR(45), + top_ip_count INT, + avg_requests_per_minute DECIMAL(10,2), + peak_requests_per_minute INT, + created_at DATETIME DEFAULT CURRENT_TIMESTAMP, + UNIQUE INDEX idx_hour (hour_bucket) +); +``` + +- Hourly rollup performed by n8n workflow (separate from the parser) +- Raw events retention: grow unbounded initially, add cleanup policy when +volume is understood +- Summary table retention: indefinite + +## Layer 2: Dashboard (Streamlit) + +### Purpose + +Visualize system health and security data. Read-only, human-facing. + +### Technology + +- Streamlit, Python, reads from MySQL +- Charts via Plotly or Altair (time series) + +### Pages + +#### Page 1: System Metrics + +- Time series graphs: Docker memory per container, host memory, swap, CPU, +Apache workers +- Horizontal threshold lines showing Layer 3 warning and healer trigger +levels +- Time range selector: 24h (default), 7 days, 30 days + +#### Page 2: Security / Log Events + +- wp-login.php attempts over time +- Top attacking IPs (table) +- Status code distribution +- Request volume trends (avg and peak requests/min) + +#### Page 3: Healer Activity + +- Log of healer invocations: trigger, decision, action, outcome +- Timeline view of actions taken +- Deferred actions and their resolution (executed, auto-cancelled, +overridden) + +## Layer 3: Monitor (n8n) + +### Purpose + +Early warning system. Detects degradation trends before they become +outages. Evaluates rules, fires alerts or invokes the healer. + +### Mechanism + +- n8n workflow, scheduled every 5 minutes +- Queries MySQL for recent metrics +- Evaluates threshold rules +- Works alongside Uptime Kuma (not replacing it for now) + +### Starter Rules + +| Rule | Threshold | Action | +|---|---|---| +| Host RAM usage | >85% sustained 10 min | Google Chat warning | +| Host RAM usage | >95% sustained 5 min | Invoke healer | +| Swap usage | >50% of total swap | Google Chat warning | +| Container memory | >90% of cap sustained 5 min | Google Chat warning | +| Container memory | >98% of cap sustained 3 min | Invoke healer | +| Host CPU | >80% sustained 15 min | Google Chat warning (info only) | +| Apache workers | >80% of max sustained 10 min | Google Chat warning | +| wp-login.php POSTs | >30 in 5-min window | Google Chat warning | + +"Sustained" = all samples in the window exceed the threshold. Prevents +single-spike false positives. + +### Alert Tiers + +- **Warning** — Google Chat notification, no action taken +- **Invoke Healer** — hands off to Layer 4 with trigger context + +## Layer 4: Healer (n8n + Claude API) + +### Purpose + +Take corrective action when Layer 3 invokes it. Reason about the problem, +execute a safe response, report results. + +### Mechanism + +- n8n workflow triggered by Layer 3 or Uptime Kuma webhooks +- Portainer REST API for container operations +- Claude API for diagnostic reasoning + +### Available Actions + +| Action | Timing | Safety Rails | +|---|---|---| +| Container restart (non-protected) | Immediate | Max 1/hour, 3/24h per +container | +| WordPress restart | Deferred to quiet hours | Same caps + admin override +window | +| MySQL restart | Deferred to quiet hours | Same caps + admin override +window | +| Server reboot | Deferred to quiet hours | Max 1/24h + admin override +window | +| Alert only | Immediate | Always an option | +| Escalate | Immediate | When limits exceeded | + +**Protected containers** (deferred to quiet hours only): WordPress, MySQL + +### Quiet Hours + +- Window: 0200–0500 ET +- Deferred actions are standing orders, not scheduled events +- During quiet hours, n8n checks: is there a pending action? Is it still +pending (not overridden)? Do the original conditions still exist? +- If conditions resolved → auto-cancel with Google Chat notification +- If overridden → mark as overridden + +### Deferred Action Flow + +1. Healer decides a protected container or server needs restart +2. Sets flag in `pbs_healer_actions`: `deferred_action`, `deferred_status = +'pending'` +3. Google Chat notification with action ID and reasoning +4. During quiet hours, n8n evaluates: + - Pending action exists? → Check if overridden → Check if conditions +persist → Execute or auto-cancel +5. Status transitions: `pending` → `executed` | `auto_cancelled` | +`overridden` + +### Override Mechanism + +#### V1 + +- **pbs-hub web page:** Shows pending healer actions, click to +cancel/approve +- **Google Chat webhook:** Text command with action ID (e.g., "override +1234", "cancel 1234") +- Both hit the same pbs-hub API endpoint +- pbs-hub API also available for future integrations + +#### V2 (future) + +- Google Chat interactive cards with approve/cancel/defer buttons + +### Iteration Limits (Anti-Loop Protection) + +- Per-container cooldown: 1 hour between automatic restarts +- Per-container daily cap: 3 automatic restarts per 24 hours +- Global healer cap: 10 healer-initiated actions per 24 hours +- Cooldown on escalation: 6 hours of no auto-action for that container +after escalation +- If any limit hit → escalate instead of acting + +### Healer Decision Flow + +1. Receive trigger from Layer 3 or Uptime Kuma +2. Check iteration limits → if exceeded, escalate and stop +3. Gather context: + - Current metrics from pbs-monitor HTTP API + - Recent log events from pbs_security_events + - Container inspect + recent logs from Portainer API +4. Send structured context to Claude API +5. Claude returns: recommended action, reasoning, confidence +6. If action is allowed AND confidence is high → execute (or defer if +protected) +7. Else → escalate with Claude's reasoning +8. Log decision + action + outcome to `pbs_healer_actions` +9. Google Chat notification with: what happened, what healer did (or +didn't), Claude's reasoning + +### Claude API Prompt Shape + +- Structured input: trigger reason, current metrics snapshot, 30-min +history, recent logs, container inspect +- Structured output: diagnosis, recommended action (enum: +`restart_container`, `defer_restart`, `reboot_server`, `no_action`, +`escalate`), reasoning, confidence level + +### pbs_healer_actions Table + +```sql +CREATE TABLE pbs_healer_actions ( + id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY, + triggered_at DATETIME NOT NULL, + trigger_source VARCHAR(50), + trigger_rule VARCHAR(100), + container_name VARCHAR(100), + metrics_snapshot JSON, + claude_reasoning TEXT, + action_recommended VARCHAR(50), + action_taken VARCHAR(50), + action_result TEXT, + deferred_action VARCHAR(50), + deferred_status VARCHAR(20) DEFAULT NULL, + iteration_limit_hit BOOLEAN DEFAULT FALSE, + escalated BOOLEAN DEFAULT FALSE, + created_at DATETIME DEFAULT CURRENT_TIMESTAMP, + INDEX idx_triggered (triggered_at), + INDEX idx_container (container_name), + INDEX idx_deferred (deferred_status) +); +``` + +### Starting Conservatism + +- First period (target ~30 days): **dry-run mode** — healer makes +decisions, logs them, sends Google Chat with "I would have done X" but does +not execute +- After validation: flip to live mode, container restarts first +- Server reboot: enabled after container restart trust is established +- System reboot scheduling added only after override mechanism is validated + +## Existing Infrastructure Context + +- Docker/Traefik/Linode production stack +- Portainer running with REST API available +- Uptime Kuma operational for up/down checks and Google Chat alerts +- `pbs_automation` MySQL database exists with separate permissions +- n8n operational, Claude API integration is a known pattern +- Existing bash cron (60s, Docker memory only, flat file) — replaced by +Layer 1 +- Crowdsec security hardening is a separate project (not duplicated here) + +## Implementation Order + +- [ ] Phase 1: Layer 1 — Go collector (systemd, JSONL, MySQL sync, HTTP API) +- [ ] Phase 2: Log Parser — Python WordPress log parser + MySQL schema +- [ ] Phase 3: Layer 2 — Streamlit dashboard (system metrics + security + +thresholds) +- [ ] Phase 4: Layer 3 — n8n monitor workflow (threshold rules, Google Chat +alerts) +- [ ] Phase 5: Layer 4 — n8n healer workflow (dry-run mode, Claude API +integration) +- [ ] Phase 6: Layer 4 live — enable live actions, override page in pbs-hub +- [ ] Phase 7: Deferred actions — quiet hours logic, WordPress/MySQL/server +restart scheduling + +## Open Questions + +- [ ] Go project structure — standard layout or flat? Decide during Phase 1 +- [ ] Docker multi-stage build for Go binary — needed for CI, decide during +Phase 1 +- [ ] MySQL schema for `pbs_system_metrics` and `pbs_system_metrics_hourly` +— design during Phase 1 +- [ ] Apache max_workers value — needed to calculate percentage for Layer 3 +rules +- [ ] Portainer API auth — how is it currently secured? Needed for Layer 4 +- [ ] Claude API prompt engineering — detailed prompt design during Phase 5 +- [ ] pbs-hub override page routing — how does it fit into existing pbs-hub +Flask app? +- [ ] Google Chat webhook for override commands — new webhook or extend +existing? +- [ ] Streamlit deployment — containerized or host-level? Decide during +Phase 3 + +...sent from Jenny & Travis \ No newline at end of file