embedding-chunking/README.md
Travis Herbranson 176c04cd6c Initial commit: embedding-chunking v0.1.0
Shared 512-token chunking module for the OB1/vault embedding stack.
Wraps tiktoken cl100k_base; consumed by vault_mcp, ob1-enricher,
ob1-reflector before passing to Ollama (nomic-embed-text, 768-d).

Extracted from /opt/projects/homelab/shared/python/embedding_chunking
to its own repo so consumers can pin via git ref instead of editable
path source, unblocking docker builds outside the homelab tree.
2026-05-16 13:07:23 -04:00

949 B

embedding-chunking

Shared 512-token chunking for the OB1 / vault embedding stack.

Wraps tiktoken (cl100k_base) so all four embedders (vault_mcp, ob1-mcp, ob1-enricher, ob1-reflector) split source text the same way before passing to Ollama (nomic-embed-text, 768-d).

cl100k_base is not nomic's exact tokenizer but close enough for chunking decisions, and far faster than round-tripping every token count through Ollama.

Use

from embedding_chunking import chunk_text

for chunk in chunk_text(document):
    vec = embed_text(chunk.text)
    # upsert with chunk.index, chunk.text, chunk.token_count

Consumers must delete-before-insert keyed on (source_schema, source_table, source_id, embedding_model) before re-embedding, otherwise a shrinking source leaves orphan chunks that silently corrupt retrieval.

Constants

  • CHUNK_SIZE = 512
  • CHUNK_OVERLAP = 64

Step is CHUNK_SIZE - CHUNK_OVERLAP = 448.