Create ai-memory-architecture-research.md via n8n
This commit is contained in:
parent
57f9fd6d7f
commit
53ab72a6ed
142
Tech/Sessions/ai-memory-architecture-research.md
Normal file
142
Tech/Sessions/ai-memory-architecture-research.md
Normal file
@ -0,0 +1,142 @@
|
|||||||
|
---
|
||||||
|
project: ai-memory-architecture-research
|
||||||
|
type: session-notes
|
||||||
|
status: active
|
||||||
|
path: Tech/Sessions
|
||||||
|
tags:
|
||||||
|
- homelab
|
||||||
|
- ai
|
||||||
|
- memory
|
||||||
|
- architecture
|
||||||
|
- research
|
||||||
|
created: 2026-05-04
|
||||||
|
updated: 2026-05-04
|
||||||
|
---
|
||||||
|
# AI Memory Architecture Research
|
||||||
|
|
||||||
|
Session exploring the landscape of AI agent harnesses, model memory
|
||||||
|
mechanisms, and persistent context layers. Started as a question about
|
||||||
|
two products and turned into the architectural backbone for several
|
||||||
|
follow-on projects.
|
||||||
|
|
||||||
|
## Outcomes
|
||||||
|
|
||||||
|
- Distinguished harnesses from models as separate layers in the AI
|
||||||
|
tooling stack. Harnesses are free or cheap; models are where the cost
|
||||||
|
lives. Swapping harnesses doesn't require migrating data or model
|
||||||
|
choice.
|
||||||
|
- Identified OB1 (by Nate B. Jones) as the right adoption target for a
|
||||||
|
personal AI memory layer, validating the self-hosted
|
||||||
|
MCP-server-plus-database pattern as the approach to project memory at
|
||||||
|
scale.
|
||||||
|
- Mapped the conceptual gap between current production AI memory
|
||||||
|
(transcript replay plus prefix caching) and the research frontier
|
||||||
|
(soft prompts, persona embeddings, weight-level continual learning).
|
||||||
|
Confirmed that frontier-API models will never expose substrate access;
|
||||||
|
open-weight local models do.
|
||||||
|
- Reframed the original "build a harness" instinct as actually "build
|
||||||
|
the memory layer." Most existing harnesses are good; what's missing is
|
||||||
|
the persistent project memory underneath them.
|
||||||
|
|
||||||
|
## Topics covered
|
||||||
|
|
||||||
|
### Poolside
|
||||||
|
|
||||||
|
Frontier coding-focused AI lab, founded 2023 by Jason Warner
|
||||||
|
(ex-GitHub CTO) and Eiso Kant. Up to ~$12B valuation after Nvidia's
|
||||||
|
recent investment. Released the Laguna model family — including Laguna
|
||||||
|
XS.2, a 33B-total/3B-activated MoE under Apache 2.0 open weights that
|
||||||
|
runs on consumer hardware via Ollama. Their thesis: code is the best
|
||||||
|
beachhead to AGI because it forces long-horizon reasoning.
|
||||||
|
|
||||||
|
Relevance: Laguna XS.2 is a candidate model for local agentic coding
|
||||||
|
workloads on the RTX 3080 and slots cleanly into the Phase 2 OpenClaw
|
||||||
|
plan.
|
||||||
|
|
||||||
|
### Pi.dev
|
||||||
|
|
||||||
|
A minimal terminal coding agent harness by Mario Zechner (badlogic),
|
||||||
|
MIT-licensed. Notable design choices: no MCP, no sub-agents, no plan
|
||||||
|
mode, no built-in todos — extensions instead of features. Supports 15+
|
||||||
|
providers including Ollama. Tree-structured sessions with branching
|
||||||
|
and gist sharing. Same category as Claude Code and Aider, different
|
||||||
|
philosophy.
|
||||||
|
|
||||||
|
Relevance: Useful as a reference architecture for understanding what a
|
||||||
|
harness is and isn't. Worth a look as a daily driver for local Ollama
|
||||||
|
work, separate from sanctioned `claude -p` for scripted Claude usage.
|
||||||
|
|
||||||
|
### OpenBrain (OB1)
|
||||||
|
|
||||||
|
Personal AI memory layer by Nate B. Jones. Self-hosted
|
||||||
|
Postgres-plus-Supabase stack with MCP server, semantic search via
|
||||||
|
pgvector, auto-extracted metadata, multiple capture sources (MCP,
|
||||||
|
REST, Slack webhook, Obsidian import). Provider-agnostic — any
|
||||||
|
MCP-compatible client can read and write through the same database.
|
||||||
|
|
||||||
|
This is the convergence point of the conversation. Independent
|
||||||
|
implementation of the same architectural insight reached through
|
||||||
|
reasoning: AI memory should be a separate, owned, queryable layer that
|
||||||
|
any AI client can read and write through a standard interface, not a
|
||||||
|
feature locked inside one vendor's product.
|
||||||
|
|
||||||
|
Adoption is the right call rather than building from scratch. The
|
||||||
|
technical-project extension layer (project state, decision logs,
|
||||||
|
session notes ingestion) is the gap to fill, and contributing it back
|
||||||
|
is plausible.
|
||||||
|
|
||||||
|
## Why this matters
|
||||||
|
|
||||||
|
The cliff-of-projects problem — what happens when there are 100
|
||||||
|
projects instead of 5 and Dispatch sessions can't carry coherent
|
||||||
|
project memory — is solved at the architecture level by this category
|
||||||
|
of solution. Not by a better harness, not by more discipline, but by a
|
||||||
|
memory layer that lives alongside whatever tools come and go.
|
||||||
|
|
||||||
|
The research-frontier digression on soft prompts and weight-level
|
||||||
|
memory established the boundary: production AI memory is RAG-shaped
|
||||||
|
(text retrieved by vector similarity), not weight-shaped or
|
||||||
|
activation-shaped. Frontier APIs will never expose the substrate. That
|
||||||
|
clarifies what is and isn't worth chasing — pursue the practical
|
||||||
|
memory layer (OB1), recognize the exotic stuff (persona embeddings,
|
||||||
|
hypernetworks) as research-frontier and not currently buildable on
|
||||||
|
closed-weight models.
|
||||||
|
|
||||||
|
## Key learnings
|
||||||
|
|
||||||
|
- Harnesses and models are independent layers. Picking a harness and
|
||||||
|
picking a model are separate decisions.
|
||||||
|
- "Memory" in current AI is the harness keeping a transcript and
|
||||||
|
replaying it. The model itself is stateless. Prefix caching is a cost
|
||||||
|
optimization, not a memory mechanism.
|
||||||
|
- Frontier API access stops at "send text, get text back." No weight
|
||||||
|
access, no activation hooks, no soft-prompt injection. Open-weight
|
||||||
|
local models are where substrate experimentation lives.
|
||||||
|
- Anthropic actively blocks third-party harnesses from using
|
||||||
|
subscription OAuth (April 2026 enforcement). API key with workspace
|
||||||
|
spend caps is the only sanctioned path for Claude access from
|
||||||
|
non-Anthropic tools. `claude -p` headless mode is the loophole —
|
||||||
|
invoking the official binary from a script is fine.
|
||||||
|
- The "build my own harness" instinct was actually about needing a
|
||||||
|
shared memory layer underneath any harness. Different problem,
|
||||||
|
different scope, much smaller project.
|
||||||
|
|
||||||
|
## Follow-on projects
|
||||||
|
|
||||||
|
- [ ] Deploy OB1 on homelab — Supabase stack on Proxmox, MCP server
|
||||||
|
reachable via Tailscale
|
||||||
|
- [ ] Wire Dispatch pre-compact hook to capture transcripts into OB1 via n8n
|
||||||
|
- [ ] Extend existing markdown summary n8n flow to also POST captures into OB1
|
||||||
|
- [ ] Build technical-project extension for OB1 (project state schema,
|
||||||
|
decision log)
|
||||||
|
- [ ] Evaluate Laguna XS.2 on Ollama as local coding model
|
||||||
|
- [ ] Authelia in front of OB1 web surfaces (Studio UI), GoTrue stays
|
||||||
|
for application identity
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- Poolside Laguna: open weights via Ollama
|
||||||
|
- Pi.dev: github.com/badlogic/pi-mono
|
||||||
|
- OB1: github.com/NateBJones-Projects/OB1
|
||||||
|
- srnichols OpenBrain (alternative implementation): lighter, plain
|
||||||
|
Postgres + pgvector, no Supabase dependency
|
||||||
Loading…
Reference in New Issue
Block a user