--- created: 2026-04-16 path: Sources/Homelab project: pbs-infrastructure-intelligence status: active tags: - pbs - monitoring - docker - go - python - mysql - n8n - streamlit - security type: project-plan updated: 2026-04-16 --- # PBS Infrastructure Intelligence (PBSII) ## Goal Build a self-healing, intelligent monitoring system for the PBS infrastructure. Four interconnected layers: a metrics collector, a visualization dashboard, a threshold monitor, and an AI-powered healer. ## Problem Being Solved - DoS attacks cause server thrashing with no early warning - Container memory leaks go undetected until crashes occur - Existing monitoring (Uptime Kuma) only detects hard-down states, not degradation trends - Manual log grepping is unsustainable as the site grows - No visibility into what traffic reaches the server vs what Cloudflare blocks - Apache worker exhaustion causes memory crashes with no forensic data ## Architecture Overview ``` Layer 1: Collector (Go, systemd) → JSONL file (30s interval) → MySQL sync (every few minutes) → HTTP API (/metrics/current, /metrics/history, /health) Log Parser (Python, n8n-scheduled) → docker logs --since → MySQL pbs_security_events + pbs_security_summary Layer 2: Dashboard (Streamlit) → Reads MySQL → System metrics + security events + healer activity Layer 3: Monitor (n8n, every 5 min) → Evaluates threshold rules against MySQL data → Fires warnings (Google Chat) or invokes healer Layer 4: Healer (n8n + Claude API) → Gathers diagnostic context from all layers → Claude API reasons about the problem → Executes allowed actions or escalates → Logs all decisions to pbs_healer_actions ``` ## Layer 1: Collector (Go Microservice) ### Purpose Collect system and container metrics. Write to disk first, sync to MySQL second. Serve current state via HTTP. No opinions, just data. ### Technology - Language: Go (standard library first) - Deployment: Single static binary, systemd service on host (not containerized) - Why systemd: Survives Docker daemon restarts, no Docker socket mount needed for host metrics, resource-limited via unit file ### Metrics Collected - Docker memory per container (replaces existing bash cron) - Host total memory usage - Host swap usage - Host CPU usage - Apache active worker count ### Collection Architecture - Two goroutines in a single binary: - **Collector goroutine:** Gathers metrics every 30 seconds, appends JSON line to `/var/log/pbs-monitor/metrics.jsonl`. Cannot be slowed by MySQL. - **Sync goroutine:** Reads new JSONL entries every few minutes, bulk INSERTs to MySQL. If MySQL is unavailable, retries next cycle. File still has the data. - Log rotation via standard logrotate config ### HTTP API Endpoints - `GET /metrics/current` — latest metrics snapshot - `GET /metrics/history?range=1h` — recent history from MySQL - `GET /health` — service health check - Binds to localhost only; n8n reaches via Docker host bridge IP ### Data Retention - Raw 30-second samples: 30 days in `pbs_system_metrics` - Hourly rollup: kept indefinitely in `pbs_system_metrics_hourly` - n8n handles both the hourly aggregation job and the 30-day cleanup job ### Systemd Unit File ```ini [Unit] Description=PBS Monitor - system metrics collector After=network.target docker.service [Service] Type=simple ExecStart=/usr/local/bin/pbs-monitor Restart=on-failure RestartSec=5s User=pbs-monitor Group=pbs-monitor # Resource limits - blast radius containment MemoryMax=256M TasksMax=100 CPUQuota=20% # Security hardening NoNewPrivileges=true ProtectSystem=strict ProtectHome=true ReadWritePaths=/var/log/pbs-monitor [Install] WantedBy=multi-user.target ``` ### Ansible Deployment - New Ansible role for systemd service deployment - Tasks: copy binary, create config, create system user (in docker group), install unit file, install logrotate config, enable and start service - Follows existing staging-first deployment pattern ## Log Parser (Python) ### Purpose Parse WordPress Docker logs for security events and request patterns. Feeds both the dashboard and Layer 3 alert rules. ### Technology - Language: Python (UV + venv, PyCharm Professional) - Orchestration: n8n schedules execution every 15 minutes - Log source: `docker logs wordpress --since` command (no container changes required) ### Incremental Parsing - Query `MAX(logged_at)` from `pbs_security_events` to determine where to start - Use `--since` flag on `docker logs` to avoid reprocessing - Self-healing on restart — just picks up from last stored timestamp ### Fields Extracted - `logged_at` — timestamp - `ip_address` — source IP - `method` — GET/POST - `uri` — request path - `status_code` — HTTP response code - `response_size` — bytes sent - `user_agent` — client user agent - `event_type` — derived category (e.g., `login_attempt`, `404`, `normal`) - Note: `response_time_ms` deferred — requires custom Apache LogFormat, not needed for v1 ### MySQL Tables #### pbs_security_events (raw log entries) ```sql CREATE TABLE pbs_security_events ( id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY, logged_at DATETIME NOT NULL, ip_address VARCHAR(45) NOT NULL, method VARCHAR(10), uri TEXT, status_code SMALLINT, response_size INT, user_agent TEXT, event_type VARCHAR(50), created_at DATETIME DEFAULT CURRENT_TIMESTAMP, INDEX idx_logged_at (logged_at), INDEX idx_ip (ip_address), INDEX idx_uri (uri(100)) ); ``` #### pbs_security_summary (hourly aggregates) ```sql CREATE TABLE pbs_security_summary ( id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY, hour_bucket DATETIME NOT NULL, total_requests INT, unique_ips INT, login_attempts INT, login_page_loads INT, status_2xx_count INT, status_3xx_count INT, status_4xx_count INT, status_404_count INT, status_5xx_count INT, top_ip VARCHAR(45), top_ip_count INT, avg_requests_per_minute DECIMAL(10,2), peak_requests_per_minute INT, created_at DATETIME DEFAULT CURRENT_TIMESTAMP, UNIQUE INDEX idx_hour (hour_bucket) ); ``` - Hourly rollup performed by n8n workflow (separate from the parser) - Raw events retention: grow unbounded initially, add cleanup policy when volume is understood - Summary table retention: indefinite ## Layer 2: Dashboard (Streamlit) ### Purpose Visualize system health and security data. Read-only, human-facing. ### Technology - Streamlit, Python, reads from MySQL - Charts via Plotly or Altair (time series) ### Pages #### Page 1: System Metrics - Time series graphs: Docker memory per container, host memory, swap, CPU, Apache workers - Horizontal threshold lines showing Layer 3 warning and healer trigger levels - Time range selector: 24h (default), 7 days, 30 days #### Page 2: Security / Log Events - wp-login.php attempts over time - Top attacking IPs (table) - Status code distribution - Request volume trends (avg and peak requests/min) #### Page 3: Healer Activity - Log of healer invocations: trigger, decision, action, outcome - Timeline view of actions taken - Deferred actions and their resolution (executed, auto-cancelled, overridden) ## Layer 3: Monitor (n8n) ### Purpose Early warning system. Detects degradation trends before they become outages. Evaluates rules, fires alerts or invokes the healer. ### Mechanism - n8n workflow, scheduled every 5 minutes - Queries MySQL for recent metrics - Evaluates threshold rules - Works alongside Uptime Kuma (not replacing it for now) ### Starter Rules | Rule | Threshold | Action | |---|---|---| | Host RAM usage | >85% sustained 10 min | Google Chat warning | | Host RAM usage | >95% sustained 5 min | Invoke healer | | Swap usage | >50% of total swap | Google Chat warning | | Container memory | >90% of cap sustained 5 min | Google Chat warning | | Container memory | >98% of cap sustained 3 min | Invoke healer | | Host CPU | >80% sustained 15 min | Google Chat warning (info only) | | Apache workers | >80% of max sustained 10 min | Google Chat warning | | wp-login.php POSTs | >30 in 5-min window | Google Chat warning | "Sustained" = all samples in the window exceed the threshold. Prevents single-spike false positives. ### Alert Tiers - **Warning** — Google Chat notification, no action taken - **Invoke Healer** — hands off to Layer 4 with trigger context ## Layer 4: Healer (n8n + Claude API) ### Purpose Take corrective action when Layer 3 invokes it. Reason about the problem, execute a safe response, report results. ### Mechanism - n8n workflow triggered by Layer 3 or Uptime Kuma webhooks - Portainer REST API for container operations - Claude API for diagnostic reasoning ### Available Actions | Action | Timing | Safety Rails | |---|---|---| | Container restart (non-protected) | Immediate | Max 1/hour, 3/24h per container | | WordPress restart | Deferred to quiet hours | Same caps + admin override window | | MySQL restart | Deferred to quiet hours | Same caps + admin override window | | Server reboot | Deferred to quiet hours | Max 1/24h + admin override window | | Alert only | Immediate | Always an option | | Escalate | Immediate | When limits exceeded | **Protected containers** (deferred to quiet hours only): WordPress, MySQL ### Quiet Hours - Window: 0200–0500 ET - Deferred actions are standing orders, not scheduled events - During quiet hours, n8n checks: is there a pending action? Is it still pending (not overridden)? Do the original conditions still exist? - If conditions resolved → auto-cancel with Google Chat notification - If overridden → mark as overridden ### Deferred Action Flow 1. Healer decides a protected container or server needs restart 2. Sets flag in `pbs_healer_actions`: `deferred_action`, `deferred_status = 'pending'` 3. Google Chat notification with action ID and reasoning 4. During quiet hours, n8n evaluates: - Pending action exists? → Check if overridden → Check if conditions persist → Execute or auto-cancel 5. Status transitions: `pending` → `executed` | `auto_cancelled` | `overridden` ### Override Mechanism #### V1 - **pbs-hub web page:** Shows pending healer actions, click to cancel/approve - **Google Chat webhook:** Text command with action ID (e.g., "override 1234", "cancel 1234") - Both hit the same pbs-hub API endpoint - pbs-hub API also available for future integrations #### V2 (future) - Google Chat interactive cards with approve/cancel/defer buttons ### Iteration Limits (Anti-Loop Protection) - Per-container cooldown: 1 hour between automatic restarts - Per-container daily cap: 3 automatic restarts per 24 hours - Global healer cap: 10 healer-initiated actions per 24 hours - Cooldown on escalation: 6 hours of no auto-action for that container after escalation - If any limit hit → escalate instead of acting ### Healer Decision Flow 1. Receive trigger from Layer 3 or Uptime Kuma 2. Check iteration limits → if exceeded, escalate and stop 3. Gather context: - Current metrics from pbs-monitor HTTP API - Recent log events from pbs_security_events - Container inspect + recent logs from Portainer API 4. Send structured context to Claude API 5. Claude returns: recommended action, reasoning, confidence 6. If action is allowed AND confidence is high → execute (or defer if protected) 7. Else → escalate with Claude's reasoning 8. Log decision + action + outcome to `pbs_healer_actions` 9. Google Chat notification with: what happened, what healer did (or didn't), Claude's reasoning ### Claude API Prompt Shape - Structured input: trigger reason, current metrics snapshot, 30-min history, recent logs, container inspect - Structured output: diagnosis, recommended action (enum: `restart_container`, `defer_restart`, `reboot_server`, `no_action`, `escalate`), reasoning, confidence level ### pbs_healer_actions Table ```sql CREATE TABLE pbs_healer_actions ( id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY, triggered_at DATETIME NOT NULL, trigger_source VARCHAR(50), trigger_rule VARCHAR(100), container_name VARCHAR(100), metrics_snapshot JSON, claude_reasoning TEXT, action_recommended VARCHAR(50), action_taken VARCHAR(50), action_result TEXT, deferred_action VARCHAR(50), deferred_status VARCHAR(20) DEFAULT NULL, iteration_limit_hit BOOLEAN DEFAULT FALSE, escalated BOOLEAN DEFAULT FALSE, created_at DATETIME DEFAULT CURRENT_TIMESTAMP, INDEX idx_triggered (triggered_at), INDEX idx_container (container_name), INDEX idx_deferred (deferred_status) ); ``` ### Starting Conservatism - First period (target ~30 days): **dry-run mode** — healer makes decisions, logs them, sends Google Chat with "I would have done X" but does not execute - After validation: flip to live mode, container restarts first - Server reboot: enabled after container restart trust is established - System reboot scheduling added only after override mechanism is validated ## Existing Infrastructure Context - Docker/Traefik/Linode production stack - Portainer running with REST API available - Uptime Kuma operational for up/down checks and Google Chat alerts - `pbs_automation` MySQL database exists with separate permissions - n8n operational, Claude API integration is a known pattern - Existing bash cron (60s, Docker memory only, flat file) — replaced by Layer 1 - Crowdsec security hardening is a separate project (not duplicated here) ## Implementation Order - [ ] Phase 1: Layer 1 — Go collector (systemd, JSONL, MySQL sync, HTTP API) - [ ] Phase 2: Log Parser — Python WordPress log parser + MySQL schema - [ ] Phase 3: Layer 2 — Streamlit dashboard (system metrics + security + thresholds) - [ ] Phase 4: Layer 3 — n8n monitor workflow (threshold rules, Google Chat alerts) - [ ] Phase 5: Layer 4 — n8n healer workflow (dry-run mode, Claude API integration) - [ ] Phase 6: Layer 4 live — enable live actions, override page in pbs-hub - [ ] Phase 7: Deferred actions — quiet hours logic, WordPress/MySQL/server restart scheduling ## Open Questions - [ ] Go project structure — standard layout or flat? Decide during Phase 1 - [ ] Docker multi-stage build for Go binary — needed for CI, decide during Phase 1 - [ ] MySQL schema for `pbs_system_metrics` and `pbs_system_metrics_hourly` — design during Phase 1 - [ ] Apache max_workers value — needed to calculate percentage for Layer 3 rules - [ ] Portainer API auth — how is it currently secured? Needed for Layer 4 - [ ] Claude API prompt engineering — detailed prompt design during Phase 5 - [ ] pbs-hub override page routing — how does it fit into existing pbs-hub Flask app? - [ ] Google Chat webhook for override commands — new webhook or extend existing? - [ ] Streamlit deployment — containerized or host-level? Decide during Phase 3 ...sent from Jenny & Travis