15 KiB
| project | type | status | path | tags | created | updated | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| pbs-infrastructure-intelligence | project-plan | active | PBS/Tech/Projects |
|
2026-04-16 | 2026-04-16 |
PBS Infrastructure Intelligence (PBSII)
Goal
Build a self-healing, intelligent monitoring system for the PBS infrastructure. Four interconnected layers: a metrics collector, a visualization dashboard, a threshold monitor, and an AI-powered healer.
Problem Being Solved
- DoS attacks cause server thrashing with no early warning
- Container memory leaks go undetected until crashes occur
- Existing monitoring (Uptime Kuma) only detects hard-down states, not degradation trends
- Manual log grepping is unsustainable as the site grows
- No visibility into what traffic reaches the server vs what Cloudflare blocks
- Apache worker exhaustion causes memory crashes with no forensic data
Architecture Overview
Layer 1: Collector (Go, systemd)
→ JSONL file (30s interval)
→ MySQL sync (every few minutes)
→ HTTP API (/metrics/current, /metrics/history, /health)
Log Parser (Python, n8n-scheduled)
→ docker logs --since
→ MySQL pbs_security_events + pbs_security_summary
Layer 2: Dashboard (Streamlit)
→ Reads MySQL
→ System metrics + security events + healer activity
Layer 3: Monitor (n8n, every 5 min)
→ Evaluates threshold rules against MySQL data
→ Fires warnings (Google Chat) or invokes healer
Layer 4: Healer (n8n + Claude API)
→ Gathers diagnostic context from all layers
→ Claude API reasons about the problem
→ Executes allowed actions or escalates
→ Logs all decisions to pbs_healer_actions
Layer 1: Collector (Go Microservice)
Purpose
Collect system and container metrics. Write to disk first, sync to MySQL second. Serve current state via HTTP. No opinions, just data.
Technology
- Language: Go (standard library first)
- Deployment: Single static binary, systemd service on host (not containerized)
- Why systemd: Survives Docker daemon restarts, no Docker socket mount needed for host metrics, resource-limited via unit file
Metrics Collected
- Docker memory per container (replaces existing bash cron)
- Host total memory usage
- Host swap usage
- Host CPU usage
- Apache active worker count
Collection Architecture
- Two goroutines in a single binary:
- Collector goroutine: Gathers metrics every 30 seconds, appends JSON
line to
/var/log/pbs-monitor/metrics.jsonl. Cannot be slowed by MySQL. - Sync goroutine: Reads new JSONL entries every few minutes, bulk INSERTs to MySQL. If MySQL is unavailable, retries next cycle. File still has the data.
- Collector goroutine: Gathers metrics every 30 seconds, appends JSON
line to
- Log rotation via standard logrotate config
HTTP API Endpoints
GET /metrics/current— latest metrics snapshotGET /metrics/history?range=1h— recent history from MySQLGET /health— service health check- Binds to localhost only; n8n reaches via Docker host bridge IP
Data Retention
- Raw 30-second samples: 30 days in
pbs_system_metrics - Hourly rollup: kept indefinitely in
pbs_system_metrics_hourly - n8n handles both the hourly aggregation job and the 30-day cleanup job
Systemd Unit File
[Unit]
Description=PBS Monitor - system metrics collector
After=network.target docker.service
[Service]
Type=simple
ExecStart=/usr/local/bin/pbs-monitor
Restart=on-failure
RestartSec=5s
User=pbs-monitor
Group=pbs-monitor
# Resource limits - blast radius containment
MemoryMax=256M
TasksMax=100
CPUQuota=20%
# Security hardening
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/var/log/pbs-monitor
[Install]
WantedBy=multi-user.target
Ansible Deployment
- New Ansible role for systemd service deployment
- Tasks: copy binary, create config, create system user (in docker group), install unit file, install logrotate config, enable and start service
- Follows existing staging-first deployment pattern
Log Parser (Python)
Purpose
Parse WordPress Docker logs for security events and request patterns. Feeds both the dashboard and Layer 3 alert rules.
Technology
- Language: Python (UV + venv, PyCharm Professional)
- Orchestration: n8n schedules execution every 15 minutes
- Log source:
docker logs wordpress --sincecommand (no container changes required)
Incremental Parsing
- Query
MAX(logged_at)frompbs_security_eventsto determine where to start - Use
--sinceflag ondocker logsto avoid reprocessing - Self-healing on restart — just picks up from last stored timestamp
Fields Extracted
logged_at— timestampip_address— source IPmethod— GET/POSTuri— request pathstatus_code— HTTP response coderesponse_size— bytes sentuser_agent— client user agentevent_type— derived category (e.g.,login_attempt,404,normal)- Note:
response_time_msdeferred — requires custom Apache LogFormat, not needed for v1
MySQL Tables
pbs_security_events (raw log entries)
CREATE TABLE pbs_security_events (
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
logged_at DATETIME NOT NULL,
ip_address VARCHAR(45) NOT NULL,
method VARCHAR(10),
uri TEXT,
status_code SMALLINT,
response_size INT,
user_agent TEXT,
event_type VARCHAR(50),
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
INDEX idx_logged_at (logged_at),
INDEX idx_ip (ip_address),
INDEX idx_uri (uri(100))
);
pbs_security_summary (hourly aggregates)
CREATE TABLE pbs_security_summary (
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
hour_bucket DATETIME NOT NULL,
total_requests INT,
unique_ips INT,
login_attempts INT,
login_page_loads INT,
status_2xx_count INT,
status_3xx_count INT,
status_4xx_count INT,
status_404_count INT,
status_5xx_count INT,
top_ip VARCHAR(45),
top_ip_count INT,
avg_requests_per_minute DECIMAL(10,2),
peak_requests_per_minute INT,
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
UNIQUE INDEX idx_hour (hour_bucket)
);
- Hourly rollup performed by n8n workflow (separate from the parser)
- Raw events retention: grow unbounded initially, add cleanup policy when volume is understood
- Summary table retention: indefinite
Layer 2: Dashboard (Streamlit)
Purpose
Visualize system health and security data. Read-only, human-facing.
Technology
- Streamlit, Python, reads from MySQL
- Charts via Plotly or Altair (time series)
Pages
Page 1: System Metrics
- Time series graphs: Docker memory per container, host memory, swap, CPU, Apache workers
- Horizontal threshold lines showing Layer 3 warning and healer trigger levels
- Time range selector: 24h (default), 7 days, 30 days
Page 2: Security / Log Events
- wp-login.php attempts over time
- Top attacking IPs (table)
- Status code distribution
- Request volume trends (avg and peak requests/min)
Page 3: Healer Activity
- Log of healer invocations: trigger, decision, action, outcome
- Timeline view of actions taken
- Deferred actions and their resolution (executed, auto-cancelled, overridden)
Layer 3: Monitor (n8n)
Purpose
Early warning system. Detects degradation trends before they become outages. Evaluates rules, fires alerts or invokes the healer.
Mechanism
- n8n workflow, scheduled every 5 minutes
- Queries MySQL for recent metrics
- Evaluates threshold rules
- Works alongside Uptime Kuma (not replacing it for now)
Starter Rules
| Rule | Threshold | Action |
|---|---|---|
| Host RAM usage | >85% sustained 10 min | Google Chat warning |
| Host RAM usage | >95% sustained 5 min | Invoke healer |
| Swap usage | >50% of total swap | Google Chat warning |
| Container memory | >90% of cap sustained 5 min | Google Chat warning |
| Container memory | >98% of cap sustained 3 min | Invoke healer |
| Host CPU | >80% sustained 15 min | Google Chat warning (info only) |
| Apache workers | >80% of max sustained 10 min | Google Chat warning |
| wp-login.php POSTs | >30 in 5-min window | Google Chat warning |
"Sustained" = all samples in the window exceed the threshold. Prevents single-spike false positives.
Alert Tiers
- Warning — Google Chat notification, no action taken
- Invoke Healer — hands off to Layer 4 with trigger context
Layer 4: Healer (n8n + Claude API)
Purpose
Take corrective action when Layer 3 invokes it. Reason about the problem, execute a safe response, report results.
Mechanism
- n8n workflow triggered by Layer 3 or Uptime Kuma webhooks
- Portainer REST API for container operations
- Claude API for diagnostic reasoning
Available Actions
| Action | Timing | Safety Rails |
|---|---|---|
| Container restart (non-protected) | Immediate | Max 1/hour, 3/24h per |
| container | ||
| WordPress restart | Deferred to quiet hours | Same caps + admin override |
| window | ||
| MySQL restart | Deferred to quiet hours | Same caps + admin override |
| window | ||
| Server reboot | Deferred to quiet hours | Max 1/24h + admin override |
| window | ||
| Alert only | Immediate | Always an option |
| Escalate | Immediate | When limits exceeded |
Protected containers (deferred to quiet hours only): WordPress, MySQL
Quiet Hours
- Window: 0200–0500 ET
- Deferred actions are standing orders, not scheduled events
- During quiet hours, n8n checks: is there a pending action? Is it still pending (not overridden)? Do the original conditions still exist?
- If conditions resolved → auto-cancel with Google Chat notification
- If overridden → mark as overridden
Deferred Action Flow
- Healer decides a protected container or server needs restart
- Sets flag in
pbs_healer_actions:deferred_action,deferred_status = 'pending' - Google Chat notification with action ID and reasoning
- During quiet hours, n8n evaluates:
- Pending action exists? → Check if overridden → Check if conditions persist → Execute or auto-cancel
- Status transitions:
pending→executed|auto_cancelled|overridden
Override Mechanism
V1
- pbs-hub web page: Shows pending healer actions, click to cancel/approve
- Google Chat webhook: Text command with action ID (e.g., "override 1234", "cancel 1234")
- Both hit the same pbs-hub API endpoint
- pbs-hub API also available for future integrations
V2 (future)
- Google Chat interactive cards with approve/cancel/defer buttons
Iteration Limits (Anti-Loop Protection)
- Per-container cooldown: 1 hour between automatic restarts
- Per-container daily cap: 3 automatic restarts per 24 hours
- Global healer cap: 10 healer-initiated actions per 24 hours
- Cooldown on escalation: 6 hours of no auto-action for that container after escalation
- If any limit hit → escalate instead of acting
Healer Decision Flow
- Receive trigger from Layer 3 or Uptime Kuma
- Check iteration limits → if exceeded, escalate and stop
- Gather context:
- Current metrics from pbs-monitor HTTP API
- Recent log events from pbs_security_events
- Container inspect + recent logs from Portainer API
- Send structured context to Claude API
- Claude returns: recommended action, reasoning, confidence
- If action is allowed AND confidence is high → execute (or defer if protected)
- Else → escalate with Claude's reasoning
- Log decision + action + outcome to
pbs_healer_actions - Google Chat notification with: what happened, what healer did (or didn't), Claude's reasoning
Claude API Prompt Shape
- Structured input: trigger reason, current metrics snapshot, 30-min history, recent logs, container inspect
- Structured output: diagnosis, recommended action (enum:
restart_container,defer_restart,reboot_server,no_action,escalate), reasoning, confidence level
pbs_healer_actions Table
CREATE TABLE pbs_healer_actions (
id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
triggered_at DATETIME NOT NULL,
trigger_source VARCHAR(50),
trigger_rule VARCHAR(100),
container_name VARCHAR(100),
metrics_snapshot JSON,
claude_reasoning TEXT,
action_recommended VARCHAR(50),
action_taken VARCHAR(50),
action_result TEXT,
deferred_action VARCHAR(50),
deferred_status VARCHAR(20) DEFAULT NULL,
iteration_limit_hit BOOLEAN DEFAULT FALSE,
escalated BOOLEAN DEFAULT FALSE,
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
INDEX idx_triggered (triggered_at),
INDEX idx_container (container_name),
INDEX idx_deferred (deferred_status)
);
Starting Conservatism
- First period (target ~30 days): dry-run mode — healer makes decisions, logs them, sends Google Chat with "I would have done X" but does not execute
- After validation: flip to live mode, container restarts first
- Server reboot: enabled after container restart trust is established
- System reboot scheduling added only after override mechanism is validated
Existing Infrastructure Context
- Docker/Traefik/Linode production stack
- Portainer running with REST API available
- Uptime Kuma operational for up/down checks and Google Chat alerts
pbs_automationMySQL database exists with separate permissions- n8n operational, Claude API integration is a known pattern
- Existing bash cron (60s, Docker memory only, flat file) — replaced by Layer 1
- Crowdsec security hardening is a separate project (not duplicated here)
Implementation Order
- Phase 1: Layer 1 — Go collector (systemd, JSONL, MySQL sync, HTTP API)
- Phase 2: Log Parser — Python WordPress log parser + MySQL schema
- Phase 3: Layer 2 — Streamlit dashboard (system metrics + security + thresholds)
- Phase 4: Layer 3 — n8n monitor workflow (threshold rules, Google Chat alerts)
- Phase 5: Layer 4 — n8n healer workflow (dry-run mode, Claude API integration)
- Phase 6: Layer 4 live — enable live actions, override page in pbs-hub
- Phase 7: Deferred actions — quiet hours logic, WordPress/MySQL/server restart scheduling
Open Questions
- Go project structure — standard layout or flat? Decide during Phase 1
- Docker multi-stage build for Go binary — needed for CI, decide during Phase 1
- MySQL schema for
pbs_system_metricsandpbs_system_metrics_hourly— design during Phase 1 - Apache max_workers value — needed to calculate percentage for Layer 3 rules
- Portainer API auth — how is it currently secured? Needed for Layer 4
- Claude API prompt engineering — detailed prompt design during Phase 5
- pbs-hub override page routing — how does it fit into existing pbs-hub Flask app?
- Google Chat webhook for override commands — new webhook or extend existing?
- Streamlit deployment — containerized or host-level? Decide during Phase 3
...sent from Jenny & Travis