A host-agnostic zettelkasten memory library. A flat corpus of atomic
Markdown notes — one thought per note, own words, plain-markdown links, YAML
frontmatter with a uuid. Judgment happens at write time: an injected LLM
decides whether a turn is worth a note and drafts it; recall is full-text
search. No Hermes / agent.* import anywhere in the package.
The Hermes plugin hermes-zk-memory is one thin adapter over this library
(a MemoryProvider that wires an injected StructuredLLM to Hermes'
auxiliary-task forced-tool-call machinery). This repo is the standalone
library the plugin wraps.
Memory is the mechanism by which past experience changes future behavior. It is not retrieval. Retrieval is a read. Memory is a write-then-use cycle: something happens, it is persisted, later it changes what the agent does. Perfect recall with bad writes is bad memory.
This library exists to excel at that cycle. The caretaker of this repo
owns the how — the mechanisms that write, manage, and use a corpus of
atomic notes. The sibling zk-memory-eval (../zk-memory-eval/eval.md)
owns the how we measure — whether a change actually moved write, manage,
or use, or we are just talking about it. eval.md is the living thesis;
update it when we learn something about memory. Do not fork the thesis
here.
Three phases. Every feature maps to one (or is not a memory feature):
- Write — decide what to persist. Selection, compression, update.
Here:
retain_turn/integrate/ingest(agent-voiced study of sources). Good write is selective, accurate, merge-aware, noise-resistant. - Manage — keep the store truthful over time. Belief revision,
eviction, consolidation, conflict resolution.
Here:
tend_writes,split_note, append-only merge. Still thin — notes mostly accumulate; nothing yet prunes or revises a superseded belief. That gap is load-bearing, not a footnote. - Use — persisted state changes later actions, not just later
answers. Here:
search/readfeeding the next turn. Coupling is still indirect (the host decides what to do with hits). Good use is action-coupled, latency-aware, abstention-capable.
Good is not "high recall on a benchmark." A change that does not move write quality, store truth, or action-coupling is not a win — ship it only if eval can say so. Tokens/query and p95 latency travel with every accuracy number; a number without cost is not a measurement.
- Host-agnostic core.
zk_memoryhas zero knowledge of Hermes,agent.*, or any LLM provider. It is embeddable by a plugin, a notebook, a script, another agent. - Write-time judgment. The LLM decides, per turn, whether anything is worth retaining and drafts it — not a raw transcript log, not recall-time-ranking of everything.
- Append-only, collision-safe writes.
mergeappends a dated fragment under a corpus-wide flock;writerefuses to overwrite. A bad merge can at worst add a wrong fragment, never destroy content. - One code path for volitional and automatic recall. The same corpus
operations back both the tool surface and the
retain_*motions.
The thing worth protecting:
- Atomic notes over transcripts. One thought per note means the corpus stays navigable and linkable, and merge decisions stay local.
- The concept / entity_update split.
entity_updateis a temporal or attribute-level fact that would be a useless orphan as its own note — it belongs appended to an existing entity note. Conflating the two kinds ruins the corpus. - Merge only into the same entity, prefer create. A wrong merge pollutes
an existing note; a missed merge is just slight duplication — the safer
failure. A
merge_target_refnot among the fetched hits is never trusted. rgfallback. Search never hard-fails without lancedb.- No free-form semantic tags. The only tag a note carries is
kind(concept/entity_update/decision) — a closed set that drives one structural decision (whether a note may merge). Free-form / LLM-invented tags are not allowed: a tag is semantic only if some consumer acts on it, and nothing here reads arbitrary tags — so they'd be dead frontmatter, a second crude copy of what the note's own content + links already say. Maintaining two copies is where the noise comes from. If a future feature genuinely consumes a tag (grouping, browse, recall filter), add it as a bounded, closed vocabulary with that consumer — never "tags for future use."
zk_memory/
__init__.py # exports Memory + module-level corpus functions
memory.py # Memory(root, llm=None, tracer=None) — the embeddable object
corpus.py # list/search/read/write/merge/tend (all take an explicit root)
integrate.py # the careful-write spine: decide_merge_target + integrate (functional pipeline)
indexing.py # IndexProvider + EmbeddingProvider protocols; Rg/LanceDB/Auto/Vector providers; registry
fts.py # LanceDB FTS engine (optional; search falls back to rg)
retain.py # retain_turn / retain_messages / process_candidate (composes over integrate)
split.py # the de-merge spine: decide_split_fragments + split_note (Z12)
judge.py # StructuredLLM protocol + distill/merge/split prompts & schemas
probe.py # trace(event, root, **fields) -> <root>/.zk/trace.jsonl
sidecar.py # the .zk/ sidecar paths (lock, index/, index/vector/, trace.jsonl, ingest/) — single source of truth
ingest.py # agent-driven source study: start/next/accept/skip/status (never authors prose)
cli/ # thin CLI: search / read / write / merge / tend / list / retain / integrate / split / tend-writes / split-candidates / ingest
cli/llm.py # CLI StructuredLLM over an OpenAI-compatible chat endpoint (optional httpx)
tests/ # corpus ops, probe, judge (StructuredLLM stubs), retain, indexing, integrate, split, ingest, cli
Install extras (all optional; the core degrades gracefully without them):
zk-memory[lancedb] full-text recall · zk-memory[faiss] vector recall
(embedder is caller-supplied) · zk-memory[cli-llm] LLM-backed CLI commands ·
zk-memory[test] pytest · zk-memory[markitdown] convert PDF/DOCX/HTML/… to
markdown before ingest. Combine as zk-memory[lancedb,faiss].
Module-level functions take an explicit root — for embedders that just want
files:
from zk_memory import search, read, write, merge, tend, list_notes
search("judy", root)Memory binds a root (and optionally an LLM and a tracer):
from zk_memory import Memory
m = Memory(root=Path("./zk"), llm=my_llm) # llm optional
m.search("judy", limit=8)
m.retain_turn(user, assistant) # distill -> merge|create
m.retain_messages(messages) # same pipeline over a batchNo LLM → retain_* returns empty / no-ops; corpus ops still work.
The library never imports openai / anthropic / litellm /
agent.auxiliary_client. judge.StructuredLLM is the contract:
class StructuredLLM(Protocol):
def __call__(self, messages: list[dict[str, str]], *,
schema: dict, name: str) -> dict | None: ...judge.py owns the prompts, JSON schemas, and orchestration; it calls
llm(messages, schema=..., name=...). The Hermes adapter implements this
with the auxiliary-task forced-tool-call path (so live retain behavior is
unchanged); a notebook implements it with whatever JSON mode it has.
retain_turn(user, assistant) (and retain_messages(messages) for a
compaction batch):
- Distill — one call, sees only the transcript, zero corpus visibility.
Splits it into candidates tagged
concept(an evergreen idea),entity_update(a temporal/attribute fact that belongs on an existing note), ordecision(a commitment/choice made — recorded as an authoritative, dated, recallable fact with choice/alternatives/rationale). - Per candidate —
searchits topic (no LLM). No hits → straight to create. Hits → fetch full bodies and make one comparison call across all of them (judge_merge) deciding merge-into-existing vs. create. Decisions skip this step entirely (never merge) and always become a standalone dated zettel.
The write path is deliberately cheap (arbitrary writes, mechanical
safety only). Quality is bought later by the gardener pass:
Memory.tend_writes() walks the most recent writes (by mtime) as the
highest-priority candidates and reconciles each — merges a duplicate into
an existing note (append-only fold) and retires it to .archive/
(reversible, never deleted), or appends [label](slug.md) out-links so
the graph grows. decision notes never merge. This is distinct from
tend (linlink structure hygiene: repair/check/mint/robustify). Capture fast,
integrate later; recency is the priority.
3. Write — merge (append-only) or write (new note).
The merge-or-create judgment is a single functional pipeline that three entry points compose over, so no path re-implements it:
decide_merge_target(root, *, content, topic, kind, llm, index, limit,
exclude_ref, exclude_path) -> uuid | None
integrate(root, *, content, topic, kind, llm, ...) -> {action, target?, path?, uuid?}
decide_merge_target is the pure decision: search (no LLM) → fetch full
bodies → drop the note itself (gardener case) → decisions never merge →
one judge_merge across all hits → verify the returned ref is among the
fetched uuids (a hallucinated ref is never honored). integrate wraps it
with the write: append-only corpus.merge, or corpus.write (building a
decision body from choice/rationale). Callers compose over it:
retain.process_candidate— capture-time flavor (create; returns label).tend._reconcile_note— gardener flavor (keep + link, or fold + archive).Memory.integrate(...)— the public careful write: a caller hands an atomic memory and gets merge-or-create with full verification. Requires anllm(returns{action:"error"}without one).
The inverse of merge — restores atomicity when an entity note has grown into a biography. Mirrors the merge pair:
decide_split_fragments(root, *, ref, llm, max_fragments=4)
-> {split, parent_summary, fragments} | {split: False} (pure decision)
split_note(root, *, ref, llm, ...) -> {action, parent, children} (decision -> write)
decide_split_fragments is the pure decision: the split judge (prefer-not-to-
split, at least as conservative as the merge judge) returns a summary parent
- atomic children, capped at 4 (schema + defensive truncation).
split_noteperforms the write: a new summary parent, new atomic children (decisions stay standalonedecisionzettels with choice/rationale), links between them, and the original biography retired to.archive/(reversible, never deleted).
Memory.split_note(ref)— the volitional entry point: the caller names the note to split. Requires anllm.split_candidates(root, top=)— the mechanical sweep: surfaces notes that need splitting by descending file size (no LLM). This is the sole authorization to split during gardening.- The gardener splits —
tend_writes(split_sweep=N)runs the sweep and de-merges the top N surfaced notes. The gardener splits only notes that came from the sweep; it must never decide on its own, mid-pass, that a note should be split. (split_noteis the shared split primitive.) - Parent/child merge guard —
decide_merge_targetnever merges into a note that already has a parent/child relation with the candidate, so a split artifact isn't folded back into its own biography.
Converting a large source tree is studying, not import. There is no
purely mechanical conversion, even with a supplied LLM. The notes
belong to the agent (or being) who writes them. Non-markdown files
(PDF, DOCX, HTML, …) are first converted to markdown via markitdown
(zk-memory[markitdown]) and cached under .zk/ingest/converted/;
the agent still studies the markdown and writes the thought.
ingest start --from DIR # inventory; per-doc state under <root>/.zk/ingest/
ingest next # surface one heading-aware window, then stop
ingest accept … # agent's prose → write or --merge
ingest skip … # this window holds no thought
ingest status [--json] # check-in (pending/done/notes/last-event age)
next shows: spine (headings mechanical; claim is the agent's), heading
path, section text, captured[] from this document, live-web hits.
accept is the only write. --keep leaves the window open for another
thought. The library never drafts a body.
A window is a reading unit, not a note. Process one document in reading order so later sections see notes this document just wrote.
- Corpus discipline. Flat
YYYYMMDD-slug.md; uuid minted vialinlink(never hand-written), with an own-uuid fallback when linlink is absent. Plain-markdown links[label](slug.md).tendis the link-integrity gate (check / repair / mint / robustify). It walks up from the corpus root to findlinlink.tomland runs linlink from that directory — so an adopted corpus in a subdir (genesis/zk) still resolveslin:citations.darnlinkis the old name; do not call it. - State footprint — one
.zk/sidecar inside the corpus. Everything the library plants (beside the notes) lives under<root>/.zk/, never beside it in the parent: the merge lock (.zk/lock), the LanceDB FTS index (.zk/index/), the FAISS pin (.zk/index/vector/— model + dims, must not mix embedders), the diagnostic trace (.zk/trace.jsonl), and ingest job state (.zk/ingest/). Retired notes go to<root>/.archive/(reversible, never deleted). So "adopt = point at the corpus dir" is self-contained — nothing leaks into the parent directory. Seezk_memory/sidecar.py(single source of path truth). - Diagnostics.
probe.trace(event, root, **fields)logs at INFO and appends one JSONL line to<root>/.zk/trace.jsonl. Never raises; a trace failure must never break the retain it describes. - Shared / multi-host corpora (e.g. a NAS every agent reads and writes).
Use the
rgsearch backend (Memory(backend="rg"),search(..., backend="rg"), or envZK_MEMORY_BACKEND=rg) — the LanceDB index is single-writer and unsafe to share. Writes are collision-safe and merges are append-only, so concurrent writers degrade gracefully;flockis best-effort only across hosts (the O_APPEND append is the real atomicity). Passsource=(host/agent name, or envZK_MEMORY_SOURCE) towrite/merge/retain_*for attribution. Givetend/check/repair/mint/robustifyto one caretaker host, never concurrent across hosts. - Recall is a pluggable engine (
indexing.IndexProvider).corpus.searchandMemoryresolve it three ways, in precedence: an injectedindex=/Memory(index=...)provider object (the DI seam — embedders bring their own remote/vector/custom engine), abackend=name (built-insauto/rg/fts, or anyregister_backend(name, provider)-ed name), else theZK_MEMORY_BACKENDenv var ("auto"). The chosen provider is threaded through the whole recall path —Memory.search,retain_*, andtend_writes— so a shared corpus configured forrgnever silently touches lancedb during retain/merge (a real bug before the abstraction). lancedb is a build-time extra (zk-memory[lancedb]); "other" providers are necessarily caller-supplied at runtime, hence the DI seam. Recall never hard-fails:auto/rgfall back to ripgrep,fts-only returns [] when lancedb is absent. - Vector recall is a DI seam, not a backend string.
indexing.VectorProvider(zk-memory[faiss]) is a FAISSIndexProviderthat needs an injectedEmbeddingProvider(the vector counterpart toStructuredLLM— the library never imports a provider SDK). It's injected asMemory(index=VectorProvider(embedder)), never abackend=name, because the embedder is caller-supplied. The index is pinned under<root>/.zk/index/vector/(meta.jsonrecordsmodel+dims,index.faissholds the vectors). Unchanged notes are not re-embedded. A later search with a different model or dimension raisesEmbedderMismatchunlessrebuild_index=True(replace the pin). A missing faiss or embedder degrades to [] (never hard-fails). CLI:--embed-model/--embed-base/--embed-key(orZK_MEMORY_EMBED_*/OPENROUTER_API_KEY); when set, search and ingestnextuse hybrid (vector + lexical). - Tests.
pytestin the repo root. Judge tests useStructuredLLMstubs — never fake OpenAI clients. To force search down thergfallback, install a fakezk_memory.ftswhoserun_ftsraisesImportError, or passbackend="rg". - Release discipline — cut versions for the consumer, not for yourself.
Callers pin by tag (e.g.
hermes-zk-memorypinszk-memory @ ...@vX.Y.Z), so a feature is only real to a consumer once it's cut into a release they can pin. The standing rule: when a feature lands onmainthat a consumer should reach, cut a new version promptly — bumppyproject.toml+__init__.py.__version__together (minor for additive features, patch for fixes) and tagvX.Y.Z. Don't let a pile of landed features accumulate behind an old tag — a consumer pinned to that tag silently misses them. What we do is land the feature and cut the release; re-pinning a consumer (e.g. the plugin) to the new tag is that repo's own caretaker's job, not ours to go do.
- House git (fleet_git, two modes). Mode 1 — active iteration: work on
main, commit small, revertable batches frequently, and never accumulate uncommitted work (commit roughly every 30 min). Push to a branch and open a PR when a change is ready; sync with origin/main before starting and before a PR. Mode 2 — parallel feature work: use linked worktrees underzk-memory.wt/<branch>/, mechanized bygit wt-new/git wt-rm— never hand-rungit worktree add. Branchdocs/<x>/fix/<x>/feat/<x>/task/<x>→ folderdocs--<x>; the.wt/namespace is flat (one worktree per branch, never category subfolders). This repo has no prior mainline — the first commit is main. Clean end-state (contract): no stale worktrees, no leftover local branches beyondmain, mainline at origin tip, primary clone clean. End every session by restoring it — or recording the deliberate deviation. - Custody. AGENTS.md +
skills/are the custody mechanism. No.agent/folder, no inhabit, no formal handoff. - AGENTS.md is the single source of truth.
README.mdis a symlink to this file for GitHub; there is no separate human doc to keep in sync. - Reversible-first. Prefer changes that are easy to revert; never leave the repo worse than you found it.
- All functionality and configuration is reachable via CLI arguments.
Nothing is library-only. Every
Memorymethod and corpus function has azk-memory <command>counterpart, and every knob (root, backend, LLM endpoint, split cap, kinds, source) is a CLI argument or env var — not a code change. LLM-backed commands take--llm-model/--llm-base/--llm-key(orZK_MEMORY_LLM_*/OPENROUTER_*env). Embeddings take--embed-model/--embed-base/--embed-key(orZK_MEMORY_EMBED_*/OPENROUTER_API_KEY). The library itself stays provider-free; the CLI'scli/llm.pyis the one place an HTTP client is imported (lazily, via thecli-llmextra).
zk-memory-eval (../zk-memory-eval) is the measurement sibling. This
repo is the how; that repo is the how we know. Its eval.md is the
living thesis (what memory is, what good means, how we measure). Do not
copy ambitions or scores here — change them there. A feature lands here
only if eval can say it moved write, manage, or use.
hermes-zk-memory (a separate repo, also in witt3rd) is the Hermes
MemoryProvider wrapper. Its __init__.py constructs Memory with an
adapter LLM, owns the tool text formatting / threading / config / auxiliary
task registration, and delegates every substantive op to this library. If
you find yourself reimplementing merge-or-create in the plugin, push it down
here instead (P9).
WHY-NOT.md answers "why don't you use mem0 / LangMem / Khoj / QMD / ...?"
for this library, anchored to PRINCIPLES.md. Maintenance rule: add a
line when a system is proposed; update when one materially changes. See
skills/why-not/ for the discipline.