Documentation

Cognitive Memory

A memory service for agents. It stores durable facts, decides what to put in front of a model each turn, and tells you why. The retrieval path is deterministic by design, and the reasoning behind every threshold below is written down because the numbers are only defensible with it.

The model

Most “memory” for agents is a list of strings and a similarity search. That is enough to make a demo feel like recall, and it fails in three specific ways that show up in production.

It cannot hold a contradiction.

Two stored claims that cannot both be true are, in a flat list, indistinguishable from two unrelated facts. There is nowhere to put “these disagree” — so one of them quietly wins, and the agent proceeds on a premise the user has already contradicted.

It has no notion of cost.

A similarity search returns k results, and k is chosen by whoever wrote the call. As the store grows, so does the prompt, and the material included falls into two groups: things that were relevant, and things that merely looked like they might be. The second group is what context rot actually measures.

It cannot say why something matched.

A cosine score is not an explanation. When a user asks why their agent believes something, the honest answer is a floating point number, which is not an answer.

This service is built around those three gaps. It stores tiers, so cost per turn is a decision rather than an accident. It stores tensions as first-class rows. And every injection is recorded with the rule that produced it, so the dashboard can show the exact block a model is about to receive.

Four tiers

A tier is a cost decision. L0 and L1 are written into the prompt every turn; L2 and L3 cost nothing until something promotes or recalls them.

TierNameHoldsCost
L0PinnedTensions, self-model guardrails, correction noticesFull body, always
L1Hot cacheNewly learned and pre-staged factsIndex line, body on trigger
L2Warm storeCandidates scored against each turnNothing until promoted
L3Cold archiveEverything else, still recallableNothing until recalled

New statements land in L1 immediately. Waiting for a promotion pass added a turn of latency, which meant a fact you taught on one turn was still missing from the very next prompt — the most visible possible way for memory to look broken.

L1 is capped by the token budget. When the cap bites, the report says truncated: true rather than dropping the tail in silence, and the least recently accessed entries are the ones that lose their place.

What gets learned

Learning runs deterministic patterns first, then optionally a model. The order is the point: a model is a good judge and a poor witness, because it can refuse, hedge, or return nothing, and a fact the user plainly stated then never gets learned at all.

URLs

Unambiguous. A host with a dot in it is captured whole, so a value like internal-hbr-2291.pineapple.example is never truncated at the first period.

Stated requirements

always / never / must / should / make sure to — kept in your own words, because paraphrasing once turned “Never force push” into “requires: force push”.

Assignments

“X is Y” where Y is identifier-shaped. Function words and filler are rejected, so “this is fine” never becomes a memory.

Model extraction

Optional, on top. Project facts, preferences and constraints the patterns cannot see — with the user's statement treated as the signal, not the assistant's confidence.

Three things are always refused. Instructions about how to behave in this conversation — “do not verify this against the repo” — are not project facts. A turn whose user message contains a question is a lookup, not a lesson, so learning is skipped: extracting from recall turns stored the assistant’s own answers back as memories, which duplicated facts and evicted the real ones. And anything interaction-scoped is returned as a rejection with a reason rather than dropped.

Reconciliation

Two statements can say the same thing. Storing both is how a memory layer becomes unsearchable inside a week, so a candidate is compared against what is already held and folded in — under one rule that governs every merge:

A merge must keep every distinctive token of both sides. A replacement missing one is refused, and both entries are kept.

Distinctive tokens are the ones that carry a fact’s identity — identifiers, numbers, codes. Function words and generic nouns (“file”, “name”, “project”) are excluded because they recur in every restatement and hide real differences. The test is deliberately biased towards reporting a loss: a false positive costs one duplicate row, which is recoverable, while a false negative deletes a fact permanently.

So “Always run migrations against staging” and “Always run migrations against staging, never production” do not merge. The second says strictly more, and dropping the qualifier is exactly the failure this rule exists to prevent.

With a model configured, the ADD / MERGE / REPLACE / REJECT decision is made once per turn for all candidates together, and any failure falls back to keeping both.

What goes into the prompt

POST /v1/context returns the block to prepend to a system prompt, plus the entries that produced it and what each cost.

indexgist lineThe default. What nearly every memory costs.
triggerfull bodyA concrete identifier absent from the transcript matched this memory.
tensionfull bodyAn unresolved contradiction, with the question to ask.
guardrailfull bodyA domain this agent has been unreliable in.

The trigger is the interesting one, and it is deterministic. Pull identifiers out of the message — URLs, dotted hosts, paths, SCREAMING_SNAKE, camelCase, long kebab-case, hex-ish codes — and keep only those absent from the visible transcript. If the caller named something concrete the model cannot already see, and a memory mentions it, that memory’s body is included. No model is asked whether that matters.

A synchronous fast gate runs first on the raw message and catches explicit corrections — “actually, we switched to Postgres”, “stop using that”, “that’s wrong”. When it fires, a premise-correction notice goes into the block, because a user visibly changing their mind is the single most reliable signal that the agent’s assumption is stale.

Recall

Ranking is token overlap, normalised by the smaller side, with paraphrases of one fact collapsed to their best-scoring instance. No model and no embedding service is involved.

That is a deliberate trade. A semantic index earns its keep on fuzzy paraphrase over very large corpora. The register an agent actually stores is URLs, build ids, ports, paths and conventions — where exact tokens are both faster and more precise, and where an embedding adds latency, cost, and a failure mode at the one moment you cannot afford one: mid-conversation.

The practical consequence is that recall quality is a property of the deployment rather than of the provider. A question that shares no words with a memory will not match it, which is why the index is pre-staged into every prompt as well: a memory you never had to ask for is a memory that cannot be missed.

An empty result returns empty: true. The SDK turns that into “if you were not told, say so rather than guessing”, because a silent miss is how a language model invents a fact it was never given.

Knowledge tensions

A tension is two claims that cannot both be true, stored as a pair with an actionable question. It is pinned into every context build until resolved, and it is the one thing that is always injected in full regardless of budget.

### Active Knowledge Tensions (Contradictions)
- [CRITICAL] “We deploy on Fridays” conflicts with “We never deploy on Fridays”.
  Ask: Which is it?

Resolving one keeps the resolution, including the reusable pattern it revealed. Deleting it would mean rediscovering the same contradiction next month.

The self-model

Reliability per domain, tracked as a moving average over outcomes you record. Any active domain below 75% is rendered into every prompt as a guardrail listing its known failure patterns and what has worked.

One detail worth stating plainly, because the obvious formula gets it wrong: the average carries a prior of two samples. Dividing by the sample count makes the first outcome fully replace the starting estimate, so a single failed task takes a domain from 0.8 to 0 — and the domain then trips the guardrail threshold forever. With the prior, one failure is a strong signal and ten are conclusive, which is what a reliability number should mean.

Nothing here is inferred. Without outcomes recorded through POST /v1/self-model/outcome, the model stays at its priors and no guardrail ever fires, which is the most common reason a self-model looks like it is not working.

API

Every endpoint is JSON. Authenticate with Authorization: Bearer <key>.

MethodPathDoesNeeds
POST/v1/contextBuild the prompt block. Returns text, entries with reasons, and the token cost.memories:read
POST/v1/recallDeterministic ranked lookup. Returns empty:true when nothing matched.memories:read
POST/v1/turnsLearn from a completed turn.memories:write
POST/v1/memoriesStore facts outright. Restatements are folded in.memories:write
GET/v1/memoriesList what is held, with tier counts.memories:read
GET/v1/memories/:idOne memory in full.memories:read
PATCH/v1/memories/:idMove between tiers.memories:write
DELETE/v1/memories/:idForget one memory.memories:write
GET/v1/tensionsContradictions, filterable by status.memories:read
POST/v1/tensionsRecord a contradiction.memories:write
POST/v1/tensions/:idResolve one, keeping the pattern.memories:write
GET/v1/self-modelReliability per domain.memories:read
POST/v1/self-model/outcomeRecord how a domain went.memories:write
GET/v1/statsCounts, weak domains, active tensions.stats:read
GET/api/v1/healthLiveness, limits, extractor mode. No key required.—
GET/v1/keysList your keys by prefix.session
POST/v1/keysMint a key. The secret is returned once.session
DELETE/v1/keys/:idRevoke a key.session

Errors carry a tag, a message, and whatever is actionable: requiredScope on a 403, issues on a 400. A caller can therefore tell a wrong key from a wrong scope, and a malformed body from an outage, without parsing prose.

Credentials

Two kinds of credential, deliberately separated.

A session identifies a person. Better Auth owns users, sessions and organisations; it is what the dashboard uses, and the only thing that can create an organisation or mint a key.

An API key identifies an agent. Keys are opaque, scoped, individually revocable, and stored as a sha256 hash — the secret is shown exactly once. A key resolves to one organisation and a set of scopes, and can only read and write that organisation’s memory.

The separation is the security model: a leaked agent key cannot mint a new key, cannot change its own scopes, and cannot create an organisation. Rotation happens from a signed-in session.

sha256 rather than a slow key derivation, on purpose. Password hashing exists to make guessing a human password expensive; these secrets carry 256 bits of entropy, so there is no search to slow down, and a deliberately slow hash would add latency to every request to protect against an attack that cannot happen. The public prefix means a presented key is one indexed lookup and one constant-time comparison.