diff --git a/Tech/Sessions/ai-memory-architecture-research.md b/Tech/Sessions/ai-memory-architecture-research.md new file mode 100644 index 0000000..744fcf8 --- /dev/null +++ b/Tech/Sessions/ai-memory-architecture-research.md @@ -0,0 +1,142 @@ +--- +project: ai-memory-architecture-research +type: session-notes +status: active +path: Tech/Sessions +tags: + - homelab + - ai + - memory + - architecture + - research +created: 2026-05-04 +updated: 2026-05-04 +--- +# AI Memory Architecture Research + +Session exploring the landscape of AI agent harnesses, model memory +mechanisms, and persistent context layers. Started as a question about +two products and turned into the architectural backbone for several +follow-on projects. + +## Outcomes + +- Distinguished harnesses from models as separate layers in the AI +tooling stack. Harnesses are free or cheap; models are where the cost +lives. Swapping harnesses doesn't require migrating data or model +choice. +- Identified OB1 (by Nate B. Jones) as the right adoption target for a +personal AI memory layer, validating the self-hosted +MCP-server-plus-database pattern as the approach to project memory at +scale. +- Mapped the conceptual gap between current production AI memory +(transcript replay plus prefix caching) and the research frontier +(soft prompts, persona embeddings, weight-level continual learning). +Confirmed that frontier-API models will never expose substrate access; +open-weight local models do. +- Reframed the original "build a harness" instinct as actually "build +the memory layer." Most existing harnesses are good; what's missing is +the persistent project memory underneath them. + +## Topics covered + +### Poolside + +Frontier coding-focused AI lab, founded 2023 by Jason Warner +(ex-GitHub CTO) and Eiso Kant. Up to ~$12B valuation after Nvidia's +recent investment. Released the Laguna model family — including Laguna +XS.2, a 33B-total/3B-activated MoE under Apache 2.0 open weights that +runs on consumer hardware via Ollama. Their thesis: code is the best +beachhead to AGI because it forces long-horizon reasoning. + +Relevance: Laguna XS.2 is a candidate model for local agentic coding +workloads on the RTX 3080 and slots cleanly into the Phase 2 OpenClaw +plan. + +### Pi.dev + +A minimal terminal coding agent harness by Mario Zechner (badlogic), +MIT-licensed. Notable design choices: no MCP, no sub-agents, no plan +mode, no built-in todos — extensions instead of features. Supports 15+ +providers including Ollama. Tree-structured sessions with branching +and gist sharing. Same category as Claude Code and Aider, different +philosophy. + +Relevance: Useful as a reference architecture for understanding what a +harness is and isn't. Worth a look as a daily driver for local Ollama +work, separate from sanctioned `claude -p` for scripted Claude usage. + +### OpenBrain (OB1) + +Personal AI memory layer by Nate B. Jones. Self-hosted +Postgres-plus-Supabase stack with MCP server, semantic search via +pgvector, auto-extracted metadata, multiple capture sources (MCP, +REST, Slack webhook, Obsidian import). Provider-agnostic — any +MCP-compatible client can read and write through the same database. + +This is the convergence point of the conversation. Independent +implementation of the same architectural insight reached through +reasoning: AI memory should be a separate, owned, queryable layer that +any AI client can read and write through a standard interface, not a +feature locked inside one vendor's product. + +Adoption is the right call rather than building from scratch. The +technical-project extension layer (project state, decision logs, +session notes ingestion) is the gap to fill, and contributing it back +is plausible. + +## Why this matters + +The cliff-of-projects problem — what happens when there are 100 +projects instead of 5 and Dispatch sessions can't carry coherent +project memory — is solved at the architecture level by this category +of solution. Not by a better harness, not by more discipline, but by a +memory layer that lives alongside whatever tools come and go. + +The research-frontier digression on soft prompts and weight-level +memory established the boundary: production AI memory is RAG-shaped +(text retrieved by vector similarity), not weight-shaped or +activation-shaped. Frontier APIs will never expose the substrate. That +clarifies what is and isn't worth chasing — pursue the practical +memory layer (OB1), recognize the exotic stuff (persona embeddings, +hypernetworks) as research-frontier and not currently buildable on +closed-weight models. + +## Key learnings + +- Harnesses and models are independent layers. Picking a harness and +picking a model are separate decisions. +- "Memory" in current AI is the harness keeping a transcript and +replaying it. The model itself is stateless. Prefix caching is a cost +optimization, not a memory mechanism. +- Frontier API access stops at "send text, get text back." No weight +access, no activation hooks, no soft-prompt injection. Open-weight +local models are where substrate experimentation lives. +- Anthropic actively blocks third-party harnesses from using +subscription OAuth (April 2026 enforcement). API key with workspace +spend caps is the only sanctioned path for Claude access from +non-Anthropic tools. `claude -p` headless mode is the loophole — +invoking the official binary from a script is fine. +- The "build my own harness" instinct was actually about needing a +shared memory layer underneath any harness. Different problem, +different scope, much smaller project. + +## Follow-on projects + +- [ ] Deploy OB1 on homelab — Supabase stack on Proxmox, MCP server +reachable via Tailscale +- [ ] Wire Dispatch pre-compact hook to capture transcripts into OB1 via n8n +- [ ] Extend existing markdown summary n8n flow to also POST captures into OB1 +- [ ] Build technical-project extension for OB1 (project state schema, +decision log) +- [ ] Evaluate Laguna XS.2 on Ollama as local coding model +- [ ] Authelia in front of OB1 web surfaces (Studio UI), GoTrue stays +for application identity + +## References + +- Poolside Laguna: open weights via Ollama +- Pi.dev: github.com/badlogic/pi-mono +- OB1: github.com/NateBJones-Projects/OB1 +- srnichols OpenBrain (alternative implementation): lighter, plain +Postgres + pgvector, no Supabase dependency \ No newline at end of file