wiki-vault/Sources/Homelab/vault-semantic-search.md
2026-05-17 02:45:11 +00:00

226 lines
8.5 KiB
Markdown

---
created: '2026-05-17'
path: Sources/Homelab
project: vault-semantic-search
tags:
- mcp
- pgvector
- embeddings
- wiki
type: project-plan
---
# Vault Semantic Search
## Goal
Add a `vault:search` MCP tool that takes a query string (plus optional
metadata filters) and returns vault notes ranked by semantic similarity.
The tool runs entirely against the `wiki` schema and reuses the embedding
infrastructure built by the Postgres consolidation plan.
Concrete use case: "search the vault for context" returns relevant notes
without needing to remember slugs or tag names. Semantic search replaces
the current "list everything, filter client-side" pattern.
This plan is scoped to vault search only. A future unified memory search
spanning vault + thoughts + reflections is anticipated via a `scope`
parameter, but is deferred until vault search is in use and shaping
requirements are clearer.
## Locked Decisions
### Tool surface
- Name: `vault:search`
- Read-only. No writes, no side effects.
- Single MCP tool — no separate "by tag" or "by date" variants. All
filtering is parameterized on this one tool.
### Inputs
- `query` (required, string) — natural-language search query
- `tags` (optional, list of strings) — restrict to notes carrying any of
these tags
- `domain` (optional, string) — one of the schema's domain values
(homelab, dev, venture, reference)
- `status` (optional, list of strings) — one or more lifecycle statuses
- `note_type` (optional, string) — restrict to project-plan,
session-notes, entity, topic-landing, or synthesis
- `date_range` (optional, object with `start` and `end`) — filter by
`updated_at` timestamp range
- `limit` (optional, int, default 10) — max results to return
Filters compose via AND. Vector similarity ranks within the filtered
candidate set.
### Output shape (minimal)
```
[
{
"slug": "ob1-deployment",
"title": "OB1 Main Deployment",
"path": "Sources/Homelab/ob1-deployment.md",
"score": 0.82
},
...
]
```
Score is the cosine similarity (or pgvector distance converted to
similarity — implementation detail). Caller uses `vault:get_project`
to fetch full content for any result they want to read.
Keeping content fetch separate from search keeps the search response
cheap and lets callers decide when full body content is worth the round
trip.
### Embedding behavior
- Query embedding runs **server-side** inside the vault MCP server.
- The server uses the **same embedding model** that produced the stored
vectors in `public.embeddings` (recorded in the `embedding_model`
column).
- Mismatches are not silently tolerated — if the configured query model
doesn't match the model recorded on the rows being searched, the
search returns an error rather than incorrect rankings.
- Embedding model selection is inherited from the consolidation plan;
this plan doesn't pick or change models.
### Scope
- Searches the `wiki` schema only. Both Sources and Wiki layers are
searchable; filter via `note_type` if a caller wants Sources-only or
Wiki-only.
- Does not search `ob1.thoughts` or `ob1.reflections`. That's a separate
tool (or this tool extended with a `scope` parameter) in a future plan.
### Backend mechanics
- Vector similarity query against `public.embeddings` filtered by
`source_schema = 'wiki'`.
- Filters applied via WHERE clauses joining `public.embeddings` to
`wiki.notes` (or whatever the canonical wiki content table is named).
- HNSW index on `public.embeddings.embedding` is the consolidation
plan's responsibility; this plan reuses it.
## Open Items
- [ ] Final SQL query shape — JOIN pattern between `public.embeddings`
and the wiki content table depends on exact column names from the
consolidation plan. Resolve at implementation time.
- [ ] How to surface "no results" vs "search failed" — empty array
versus error response. Likely empty array for "no results" and error
for "model mismatch / query embed failed."
- [ ] Pagination — `limit` is in scope; offset / cursor pagination is
not. Add later if it's needed.
- [ ] Caching — query embedding cost is small but non-zero. No caching
in v1. If usage patterns show repeated queries, add later.
## Phases
### Phase 0 — Verify dependencies
- [ ] Consolidation plan's Phase 5 (vault indexing) is complete.
`public.embeddings` has rows where `source_schema = 'wiki'` covering
the full vault corpus.
- [ ] HNSW index exists on `public.embeddings.embedding`.
- [ ] Embedding model is documented and reachable from the vault MCP
server's host.
- [ ] The vault MCP user has read access to `public.embeddings`.
### Phase 1 — Tool implementation
- [ ] Add `search` function to the vault MCP server.
- [ ] Wire to the configured embedding model for query embedding.
- [ ] Build the parameterized SQL query (vector search + filters).
- [ ] Define the response shape per the locked decisions.
- [ ] Add input validation (filter values exist in schema, date format
correct, etc.).
### Phase 2 — Verification
- [ ] Smoke test: search for a topic that's definitely in the vault
("MCP", "Traefik") — confirm relevant results appear in top 5.
- [ ] Negative test: search for nonsense gibberish — confirm low scores
or empty results, not garbage matches.
- [ ] Filter test: search with `tags=["mcp"]` — confirm results are
restricted correctly.
- [ ] Combined filter test: query + tag + date range — confirm AND
semantics work.
- [ ] Model-mismatch test: temporarily configure a different embedding
model, confirm tool errors rather than returning bad rankings.
### Phase 3 — Documentation
- [ ] Add usage examples to the vault MCP server's README.
- [ ] Note the tool in CLAUDE.md so agents know to use it.
- [ ] Capture lessons learned in a session note.
## Notes
### Why this isn't part of the consolidation plan
The consolidation plan builds the substrate — embeddings table, vault
indexing, HNSW index. This plan builds the user-facing tool that
consumes the substrate. Splitting them keeps the consolidation plan
focused on infrastructure and lets the search tool ship with its own
small, focused scope.
### Why minimal output instead of full content
Most searches are exploratory — "what do I have on X" — and only a
fraction of returned hits will actually be read. Returning bodies in
the search response means every search pays for content fetching that
mostly isn't used. Keeping search lightweight and giving callers a
separate `get_project` path matches the actual usage pattern.
### Why a single tool with parameters instead of separate tools
A single parameterized tool (`vault:search` with optional tags, domain,
etc.) is more flexible than many tools (`vault:search_by_tag`,
`vault:search_recent`, etc.) and matches how natural-language queries
combine criteria. "Find recent context-window notes tagged claude-code"
is one tool call with three filters, not three tool calls composed
manually.
### Why server-side embedding is the only correct choice
Each embedding model produces vectors in its own geometry. Vectors from
different models aren't comparable, and a model mismatch produces
silently wrong rankings — not an error, just bad results. The server
owns the stored vectors and knows which model produced them, so it's
the right place to ensure the query uses the same model. Letting
clients embed queries themselves would require every client to know
which model is correct and would fail invisibly when wrong.
### Future extension: scope parameter
Anticipated shape for the unified memory search that follows:
```
vault:search(query, scope="vault" | "memory" | "all", ...)
```
- `scope="vault"` — current behavior, wiki schema only (default)
- `scope="memory"` — ob1.thoughts and ob1.reflections only
- `scope="all"` — everything
Implementation extends the WHERE clause to include additional
`source_schema` values. The output shape grows a `source` field so
callers can tell vault hits from memory hits. Deferred to a separate
plan once real vault-search usage informs the design.
### Flag for Lovebug
The Postgres consolidation plan (`postgres-consolidation-reflection-layer`,
saved 2026-05-15) was written via the now-deprecated `save_note` path.
The vault markdown file exists at
`Sources/Homelab/postgres-consolidation-ob1-reflection-layer.md`, but
the Postgres row does not — `vault:get_project` returns "not found" for
that slug. The plan needs to be re-registered via `vault:create_artifact`
so it shows up in `list_projects`, can have its status updated, and
participates in the Trello card projection. Not blocking for this
search plan, but worth handling before the consolidation plan goes
into execution tracking.