wiki-vault/Sources/Dev/second-brain.md
Travis Herbranson 34a268d8dc migration: copy 62 notes from pbs-projects and homelab-projects
Two-layer structure: Sources (raw notes) + Wiki (compile output)
Four domains: Dev (40), Venture (3), Homelab (23), Reference (0)
Includes CLAUDE.md spec, index pages at all levels, compile log

Co-Authored-By: Lovebug <lovebug@herbylab.dev>
2026-05-08 17:21:13 -04:00

7.5 KiB
Raw Permalink Blame History

created path project status tags type updated
2026-05-04 Sources/Dev second-brain active
homelab
llm
obsidian
knowledge-management
claude-code
project-plan 2026-05-04

Second Brain — Project Plan

Concept

A personal knowledge system combining a video/article extraction pipeline with an LLM-maintained wiki (Karpathy LLM Wiki pattern). Sources get pulled, transcribed/extracted, reviewed, and compiled into an Obsidian-resident wiki that grows and self-maintains over time.

Phase 1 success: working knowledge extractor + working LLM wiki, end-to-end on one vertical slice, generalizable to all four domains.

Phase 2+ (out of scope for now): the "second brain" layer — static user-context, active concerns, recent decisions, open questions, wiki-aware extraction. Stretch goal, fun experiment, not a Phase 1 gate.

Architecture (3 layers)

  1. Wiki — long-term structured memory. Markdown files in Obsidian vault, LLM-maintained. Karpathy three-layer pattern: raw sources (immutable) → wiki (LLM-generated) → schema/config.
  2. Context assembler — module that builds the prompt bundle for each extraction. Phase 1 is thin (domain template + focus field + transcript) but built as a real module from day one so Phase 2 layers slot in cleanly.
  3. Extractor — stateless single-shot LLM call per source. Pure function: bundle in, structured extraction out.

Pipeline

pull → transcribe → extract → review → publish (LLM-compiled to wiki)

Existing tool already has pull/transcribe/extract and a Flask web UI with list/detail/accept. Extending it, not rewriting.

Decisions locked

Scope

  • Phase 1: extractor + wiki, all manual review
  • Phase 2 deferred: user-context, dynamic concerns, wiki-aware extraction

Sources & domains

  • Sources: web pages, videos
  • Domains: development, content, business, homelab

Scheduler

  • Rate-limit smoother, not a batch processor
  • Spaces N items across an overnight window (e.g., 10pm6am)
  • Token-aware budgeting preferred over pure time spacing (track tokens per job)
  • Failures retry with backoff, dead-letter after N attempts
  • Idempotent — re-running a finished URL is a no-op

LLM access

  • Claude Code Python SDK (subscription-based, not API)
  • Single-shot calls — no agent loop, no tool use, no session continuity
  • --bare mode equivalent for reproducibility
  • Subscription auth on herbydev — needs verification before scheduler goes live

Data model

  • Per-job domain field (one of four)
  • Per-job focus field (optional, free text — evaluative framing hint)
  • Status states: pending → pulled → transcribed → analyzed → accepted → published
  • Clean slate — no legacy data to migrate

Extraction output schema (contract between extractor and compiler)

yaml source: url, title, type, duration_or_length, published_at, ingested_at domain: focus: summary: key_points: [...] entities: [...] # candidate wiki entities claims: [...] # discrete claims tied to key_points open_questions: [...] contradictions: [...]

Wiki publish step

  • Option 3: LLM-compiled (full Karpathy pattern, not templated, not manual)
  • Compiler reads existing wiki, decides what to update/create, writes changes
  • Auto-accept on compiler writes (no diff/dry-run gate in Phase 1)
  • Git-as-recovery: commit before each compile run, commit after
    • Pre-write commit: pre-compile:
    • Post-write commit: compile: with reasoning in body (entities touched, created, flagged)
    • git log becomes the audit trail; git reset is the rollback

Wiki page types

  • Default to LLM Wiki spec — include all types (entity, concept, comparison, synthesis)
  • Skinny down later after running it

Source provenance

  • Default to LLM Wiki spec

Compiler decision authority

  • Default to LLM Wiki spec

Open items

Wiki folder structure

  • Review LLM Wiki reference projects before deciding flat vs. nested
  • Lean: may flatten (lots of folders not helpful)
  • Decision needed before compiler is wired to vault

Reference implementations to review

  • Ar9av/obsidian-wiki — skill-files for any agent, points at Obsidian vault, provenance tagging
  • lucasastorian/llmwiki — local app + MCP, indexes folder, Claude writes pages
  • nvk/llm-wiki — Claude Code plugin, parallel multi-agent research
  • Karpathy original gist — abstract spec

Verification before scheduler runs

  • Subscription auth in non-interactive cron context on herbydev

Build approach

Vertical slice first. Pick one domain (likely homelab) + one source type (likely YouTube). Build pull → transcribe → extract → review → publish end-to-end. Generalizing to other domains is mostly config; adding article ingestion is one new adapter.

Handoff to dev

This plan covers the what and why. Dev session on herbydev handles the where and how-to-wire-it-up:

  • Repo layout under ~/dev/
  • Integration with existing extraction tool (lift working pieces, fresh codebase)
  • Obsidian vault path on herbydev
  • Scheduler host/process structure (LXC, systemd timer, etc.)
  • Existing extractor code review and selective port

Phase 1 explicitly out of scope

  • Static user-context file
  • Active concerns / recent decisions / open questions context layers
  • Wiki-aware extraction (extractor doesn't read existing pages)
  • Auto-accept on extractions (always lands in analyzed for review)
  • LLM-maintained context
  • Diff/dry-run mode on compiler writes

Notes & learning

  • LLM Wiki pattern (Karpathy, April 2026): three layers — immutable sources, LLM-maintained wiki, schema/config file. Knowledge compiled once, kept current — not re-derived per query like RAG.
  • The "feeling of memory" in LLM systems is harness-driven, not model-driven. Model is stateless; harness shapes the bundle. Phase 2 "second brain" work is fundamentally context-assembler design, not extractor design.
  • SDK exercise: Claude Code Python SDK chosen partly for learning value. Uses ~10% of its surface for Phase 1 (single-shot calls), but that's fine — it's the right glue.
  • Real intellectual work in Phase 2 lives in the context assembler, not the extractor or the wiki. That's where the second-brain feeling actually comes from.

Tasks

Pre-build

  • Review LLM Wiki reference repos (30 min)
  • Decide wiki folder structure (flat vs. nested)
  • Verify subscription auth in cron on herbydev

Build (vertical slice)

  • Repo layout + project skeleton (dev session)
  • Port working pieces from existing extractor (dev session)
  • Context assembler module (thin, but real)
  • Extractor with Claude Code Python SDK
  • Per-domain prompt template files
  • Extraction output schema validation
  • Queue/scheduler with token-aware spacing
  • Web UI: domain filter, status filter, queue view, accept/publish actions
  • Wiki compiler module (reads vault, writes vault, two commits per run)
  • End-to-end test on one homelab YouTube video

Generalize

  • Add remaining three domains
  • Add web article adapter
  • Run for a week, observe, tune prompts

Phase 1 done when

  • All four domains have working prompt templates
  • Both source types ingest cleanly
  • Scheduler runs unattended overnight without intervention
  • Wiki has 20+ entity pages from real sources, with cross-references
  • Git history shows clean compile commits with usable audit trail

...sent from Jenny & Travis