wiki-vault/Sources/Dev/herbys-dev-setup.md
Travis Herbranson 34a268d8dc migration: copy 62 notes from pbs-projects and homelab-projects
Two-layer structure: Sources (raw notes) + Wiki (compile output)
Four domains: Dev (40), Venture (3), Homelab (23), Reference (0)
Includes CLAUDE.md spec, index pages at all levels, compile log

Co-Authored-By: Lovebug <lovebug@herbylab.dev>
2026-05-08 17:21:13 -04:00

241 lines
9.0 KiB
Markdown

---
created: 2026-04-22
path: Sources/Dev
project: herbys-dev-setup
status: active
tags:
- homelab
- proxmox
- ansible
- dev-environment
- backups
- tailscale
type: project-plan
updated: 2026-04-22
---
## herbys-dev — Primary Dev Environment
Promote `us-test-authy` (VMID 100 on PVE, ZFS pool `vmstore`) from test VM
to the primary development environment. Rename to `herbys-dev`, harden it,
establish real backups, and codify the entire config as an idempotent
Ansible playbook in a new repo.
The `ladybug` name stays with the Claude Desktop + Dispatch instance
running on the host — not the host itself.
### Why
The Manjaro tower pulls ~1500W when active. VM on PVE is accessible
anywhere via Tailscale, persistent across sessions, and cheap to run 24/7.
Keeping the tower asleep unless GPU work (Stable Diffusion, Ollama, Immich
ML) is actually needed.
Current state has gaps: VM is named for its original test purpose, no
formal backup strategy beyond PVE snapshots, VNC exposed on LAN with
password-only auth, no config-as-code story.
### Goals
- Rename to `herbys-dev` everywhere (hostname, Tailscale, DNS, configs)
- Idempotent Ansible playbook that can rebuild the entire dev env from a
fresh Ubuntu VM
- Nightly file-level backups of entire `/home/herbyadmin` with offsite copy
- Weekly Proxmox VM backups with sensible retention
- Tailscale-only VNC access (no LAN exposure)
- SSH available on LAN + Tailscale
- Lightweight disk/health alerting
- Dotfiles repo created and sourced by the Ansible playbook
### Scope — what this plan covers
- Rename + DNS + Tailscale migration
- New Ansible repo with idempotent playbook
- Backup strategy (file-level + VM-level)
- Security hardening (VNC, SSH, unattended upgrades)
- Baseline tooling install
- Dotfiles repo creation
- Lightweight monitoring (disk space alert)
### Out of scope (future projects)
- Full Prometheus/Grafana observability
- Migration of other VMs to the same Ansible pattern
- Claude Desktop / Dispatch reconfiguration beyond what the rename requires
### Decisions made
- **Host name**: `herbys-dev`
- **Ansible repo**: new, separate from `trucktrav/ansible-drive`
- **Playbook must be idempotent** — safe to re-run anytime
- **Secrets**: Ansible Vault for anything sensitive; no plaintext tokens in
repo
- **File-level backup scope**: entire `/home/herbyadmin` with sensible
excludes (`.venv`, `node_modules`, `.cache`, `__pycache__`, build artifacts)
- **Backup frequency**: file-level nightly; VM-level weekly full + daily
incremental (pending retention window decision)
- **VNC**: Tailscale-bound only, remove LAN exposure
- **SSH**: allowed on LAN + Tailscale
- **Dotfiles**: include in this project, playbook pulls and symlinks from
the repo
- **Shell config**: existing setup to be pulled from current machines — not
creating from scratch
### Open decisions (answer during implementation)
- Proxmox backup target: shelved 1TB spinner re-added to PVE, or NAS over
NFS/SMB?
- File-level backup tool: Restic (leaning this way for rclone backend) or
Borg?
- Offsite destination: `tjgdrive` Google Drive remote, or different target?
- Backup retention windows: how many daily / weekly / monthly to keep?
- Rename approach: in-place (change hostname, Tailscale name, update DNS)
vs. fresh VM from snapshot with new name and data migration
- Ansible run mode: push from control node, or pull-mode via `ansible-pull`
cron?
- Baseline tooling list beyond the obvious (uv, Go, Node, git, tmux,
Neovim, Claude Code, ripgrep, fd, bat, jq, htop) — anything else to bake in?
- Disk alert delivery: local cron + email, or ping into existing
notification channel?
### Tasks
#### Phase 1: Preparation
- [ ] Take a pre-work Proxmox snapshot of current VM 100
- [ ] Document current state of `us-test-authy`: installed packages,
running services, `/home` contents, systemd user units, Claude Code config,
VNC config, OpenClaw state
- [ ] Decide rename approach (in-place vs fresh VM)
- [ ] Decide Proxmox backup target (spinner vs NAS)
- [ ] Decide file-level backup tool (Restic vs Borg)
- [ ] Decide offsite backup destination and retention policy
#### Phase 2: Ansible repo setup
- [ ] Create new GitHub repo `trucktrav/herbys-dev-ansible` (or similar)
- [ ] Initialize with inventory, playbook skeleton, roles directory
- [ ] Set up Ansible Vault for secrets
- [ ] Write first idempotent task as baseline (e.g., hostname set)
- [ ] Test run against current VM to verify no destructive behavior
#### Phase 3: Dotfiles repo
- [ ] Create new GitHub repo `trucktrav/dotfiles`
- [ ] Pull existing shell config from current setup
- [ ] Add git config, tmux config, Neovim config, SSH config template
- [ ] Document what's in there in a README
#### Phase 4: Ansible roles
- [ ] Role: `common` — hostname, timezone, locale, unattended-upgrades
- [ ] Role: `baseline-tools` — uv, Go, Node (via fnm or nvm), git, tmux,
Neovim, ripgrep, fd, bat, jq, htop, Claude Code
- [ ] Role: `dotfiles` — clone and symlink from dotfiles repo
- [ ] Role: `ssh-hardening` — key-only auth, sensible sshd_config
- [ ] Role: `tailscale` — ensure Tailscale running with key expiry disabled
- [ ] Role: `vnc` — XFCE + TigerVNC bound to Tailscale interface only,
systemd user unit for persistence
- [ ] Role: `claude-desktop` — install via patrickjaja apt repo, Dispatch +
Cowork service units, loginctl linger
- [ ] Role: `backups` — install backup tool, configure repo, systemd timer
for nightly runs
- [ ] Role: `monitoring` — disk space check cron, alert delivery
- [ ] Role: `firewall` — minimal ufw config (allow SSH LAN + Tailscale, VNC
Tailscale only)
#### Phase 5: Backups
- [ ] Configure Proxmox-level backup job (weekly full + daily incremental,
chosen target)
- [ ] Set up file-level backup tool (Restic or Borg)
- [ ] Test restore path end-to-end: delete a file, restore from backup,
verify integrity
- [ ] Document restore procedure in the Ansible repo README
#### Phase 6: Rename migration
- [ ] Take fresh snapshot immediately before rename
- [ ] Change Linux hostname to `herbys-dev`
- [ ] Update Tailscale node name
- [ ] Update Pi-hole local DNS record (if one exists)
- [ ] Update `/etc/hosts` on other machines (`truckslaptop`, `manjaro1`,
`pve`)
- [ ] Update SSH client config on other machines
- [ ] Update any hardcoded references in Claude Code configs, systemd
units, existing project `CLAUDE.md` files
- [ ] Update the `zero-check` skill if it references the old hostname
- [ ] Verify Dispatch still pairs correctly after rename
- [ ] Verify OpenClaw gateway still routes correctly over Tailscale
#### Phase 7: Hardening
- [ ] Migrate VNC to Tailscale-only binding
- [ ] Confirm no lingering LAN exposure on port scan
- [ ] Enable unattended-upgrades for security patches
- [ ] Verify SSH key-only auth (no password fallback)
- [ ] Review running services, disable anything unused
#### Phase 8: Validation
- [ ] Re-run full Ansible playbook against live VM — confirm zero changes
reported (true idempotency)
- [ ] Simulate fresh-rebuild scenario: clone fresh Ubuntu VM, run playbook
end-to-end, verify identical state
- [ ] Trigger a backup manually and verify offsite copy exists
- [ ] Trigger restore test on a throwaway file
- [ ] Trigger disk alert manually (e.g., create a test file to push past
threshold) and verify delivery
#### Phase 9: Documentation
- [ ] Write `MACHINE.md` on herbys-dev listing what's installed, where data
lives, what's backed up
- [ ] Update Obsidian memory / project notes: `us-test-authy`
`herbys-dev`, role change from test to primary dev
- [ ] Update any references in other project plans
(project-scaffolding-skill plan currently says "Manjaro tower is primary
dev" — needs correction)
- [ ] README in Ansible repo: what the playbook does, how to run it, how to
add a new role
### Testing strategy
- Every Ansible role must be safe to run twice in a row with zero changes
on the second run
- Backup + restore must be validated before the old `us-test-authy`
snapshot is deleted
- Rename migration must include rollback plan (revert snapshot) in case
Dispatch, OpenClaw, or Tailscale break
### Success criteria
- Playbook runs green against a fresh Ubuntu VM and produces an identical
dev environment
- Backups running nightly with verified restore path
- Zero LAN-exposed services beyond SSH
- Host renamed everywhere it matters
- Dotfiles repo populated and Ansible-managed
- Machine documented
### References
- Current VM: `us-test-authy`, VMID 100, ZFS pool `vmstore`
- Existing Ansible pattern: `trucktrav/ansible-drive` (for structure
reference only, not extended)
- Related project: `zero-check` skill at `~/.claude/skills/zero-check/`
- Related project: `project-scaffolding-skill` plan (filed separately)
- Key infra facts: Tailscale mesh connects `truckslaptop`, `manjaro1`,
`pve`, current VM; Pi-hole + Unbound at `10.0.11.50`
### Next session
1. Answer the open decisions above (backup targets, tool choice, rename
approach)
2. Create the new Ansible repo and initialize skeleton
3. Create dotfiles repo and seed with current config
4. Start with the `common` role and test idempotency before scaling to the
full set
...sent from Jenny & Travis