Memory and continual learning for agents

Your agent forgets everything between sessions. This is the layer that stops it.

Cognitive Memory keeps the durable facts an agent is told — build ids, hosts, conventions, constraints — and hands back the relevant ones each turn. An index by default. Full bodies only where something earned them. No model in the retrieval path, so recall quality does not change when your provider does.

Self-hosted SQLite · REST API · TypeScript SDK · any model, any harness

what the agent actually receives
## Memory

### Memory index — established earlier in this project
One line per remembered item. Ask for the full item when a line is not enough.
- staging build ID is ZQ7X4M2K (deployment)
- file naming: kebab-case (naming)
- internal staging host is internal-hbr-2291.example (infra)

### Active Knowledge Tensions (Contradictions)
- [CRITICAL] “We deploy on Fridays” conflicts with “We never deploy on Fridays”.
  Ask: Which is it?

The loop

Five steps, and only one of them needs a model

A statement arrives in a conversation. It is checked against what is already known, folded in or kept apart, filed in a tier, and then either indexed or written out in full — depending on what the message in front of it actually referred to.

  1. 01

    Capture

    Deterministic patterns first: URLs, assignments, stated requirements. No model, no refusal, no silent loss.

  2. 02

    Reconcile

    Restatements are folded into what is held. A merge that would drop a qualifier is refused and both are kept.

  3. 03

    File

    Into one of four tiers. New facts land hot so the next prompt already knows them.

  4. 04

    Plan

    Build the block: a gist line for everything, a full body only for triggers, tensions and guardrails.

  5. 05

    Inject

    Into the system prompt, inside a token budget that reports when it truncated.

What this is

Cognitive memory, not a notes table

A store of facts is easy to build and quietly useless on its own. What an agent actually needs is somewhere to put the things a flat list of strings cannot represent.

It keeps contradictions

Two statements that cannot both be true are stored as a pair, pinned into every prompt with the question to ask, until somebody resolves them. This is the difference between an agent that asks and one that silently picks a side.

A store of facts cannot do this. A fact store has nowhere to put 'these disagree'.

It knows what it is bad at

Record how each domain went and the service tracks reliability as a moving average. Below 75%, that domain is rendered into every prompt as a guardrail, so a weak area gets explicit attention instead of a confident guess.

Nothing here is inferred. Without outcomes recorded, the self-model stays at its priors and no guardrail ever fires.

It refuses to invent

Instructions about how to behave in one conversation are not durable project facts, and are rejected with a reason. A recall that matches nothing returns an empty result, so the agent is told to admit ignorance instead of filling the gap.

The rejection is returned, not swallowed. A client that sent ten statements and got three back is told which seven did not land.

It shows its working

Every injection is recorded with the reason it was included and what it cost. The dashboard previews the exact block your agent is about to receive, for any message you type.

There is no published benchmark for pre-inject versus on-demand retrieval, so the only honest way to tune the tradeoff is to watch it.

Tiers

Four tiers, one budget

Tiers are not decoration. They decide what a memory costs per turn, and therefore how much of it you can afford to keep.

L0Pinned

Unresolved contradictions, the self-model, and any correction notice for the message being answered.

When the agent must not guess, or must be careful in a domain it has been failing in.

Full body, every turn
L1Hot cache

Facts pre-staged for the work in front of it. A new fact lands here immediately.

Default home for everything you tell the service.

One index line, plus a body when a trigger fires
L2Warm store

Indexed candidates, scored against each turn and promoted if they earn it.

Facts that are true but not currently relevant.

Nothing until promoted
L3Cold archive

Everything else, still searchable, still recallable on demand.

History. Retrieved by a question, never by the context builder.

Nothing until recalled

Why an index

Injecting everything is how memory becomes a liability

Research on context rot finds accuracy degrading with input length, and the damage comes from topically-related distractors rather than from structure. A small, high-signal index with on-demand bodies keeps recall cheap without filling the window with material that is usually irrelevant.

ReasonIncludedWhat earned it
indexA single gist line.The default, and what almost every memory costs. 200 memories index for ~1.3k tokens.
triggerThe full body.You named something concrete the transcript does not already contain — a build id, a host, a path — and a memory mentions it. No model decides this.
tensionThe full body.Two stored claims contradict each other. A contradiction is a question to ask, not trivia to skim.
guardrailThe full body.A domain this agent has been measurably unreliable in.
—Nothing.Above the index budget. The report says so rather than quietly dropping the tail.

measured, 200 memories

Index + triggers1,320 tok
Every body3,643 tok

Same 200 facts. 77% fewer tokens, and the window keeps its room for the actual task.

Reproduce with pnpm measure.

5.0 ms

recall p50

Deterministic ranked lookup over 200 stored memories. No model in the path, so it does not move when a provider does.

2.9 ms

context build p50

The whole prompt block — index, triggers, tensions, guardrails — assembled before the model is called.

77%

fewer tokens

Indexing 200 memories costs 1.3k tokens. Injecting all 200 bodies costs 3.6k. Bodies are spent only where something earned them.

0

models required

Deterministic extraction runs first. A deployment with no model key still learns plainly-stated facts.

Use cases

Built for agents that have to be trusted twice

The same failure everywhere: an agent that was right yesterday is confidently wrong today, and nobody can tell which. Memory is what turns a session into a relationship with the work.

Coding agents

Hold the build id, the deploy command, the naming convention, the thing that broke last month. Get them back in the one line that matters.

Support agents

Remember what this customer was told, what was actually true, and which of the two is now contradicted.

Research assistants

Keep a running set of findings and the open questions, with contradictions surfaced rather than averaged away.

Internal assistants

Company facts with a source, an owner, and a date — and a self-model of which topics the assistant is weak on.

Long-running work

A service the agent can call for the hundredth session. Memory survives deploys, restarts, and model swaps.

Multi-tenant products

Memory as a credentialed storage layer: your own database, your own keys, revocable per integration.

Quickstart

One call before the model, one after

The whole integration is two endpoints. If you wire only these two, memory works: the block goes in before the model runs, and the finished exchange comes back to be learned from.

TypeScript
import { createClient, runTurn } from "@astracollab/cogmem"

const memory = createClient({
  apiKey: process.env.COGNITIVE_MEMORY_KEY!,
  baseUrl: "http://localhost:3000",
})

// once per turn, and the loop is correct by construction
const { context, learning } = await runTurn(
  memory,
  {
    userMessage,
    run: (ctx) => callYourModel(ctx, userMessage),
  },
  { domain: "database" }
)
Any HTTP client
# what the agent should know before it answers
curl -X POST localhost:3000/api/v1/context \
  -H "Authorization: Bearer $COGNITIVE_MEMORY_KEY" \
  -H "content-type: application/json" \
  -d '{"userMessage":"deploy ZQ7X4M2K to staging"}'

# what it should remember from the finished turn
curl -X POST localhost:3000/api/v1/turns \
  -H "Authorization: Bearer $COGNITIVE_MEMORY_KEY" \
  -H "content-type: application/json" \
  -d '{"userMessage":"...","assistantResponse":"..."}'

The SDK is ~3kB gzipped, ESM-first, and depends on one small fetch wrapper. Full reference →

Questions

The ones worth answering honestly

Including the two that decide whether this is the right tool for you: it is not a vector database, and it does not need a model.

Is this a vector database with extra steps?+

No, and the difference is load-bearing. Ranking here is deterministic token overlap with identifier matching, so recall behaves identically on every run and on every model. A semantic index earns its keep on fuzzy paraphrase over very large corpora; for the register an agent actually stores — URLs, build ids, ports, conventions — exact tokens are both faster and more precise, and you can explain why a memory matched.

Do I need a model configured?+

No. Deterministic pattern extraction runs first, so a plainly-stated fact is captured even with no provider configured at all — health reports `extractor: rules-only` so you know which mode you are in. Add a key and turns are also summarised and contradictions are detected, with a typed output contract validated by the effect/ai layer rather than by parsing prose.

What happens when the agent is wrong?+

Nothing is treated as gospel. A restatement is folded into what is already held only when the merge keeps every distinctive token, so a qualifier like 'never production' cannot be quietly dropped. When a rewrite would lose information the service keeps both instead, because a duplicate costs one row and a lost fact is gone for good.

How do I stop it storing junk?+

Three ways. Interaction-scoped instructions are rejected with a reason. Questions are treated as lookups rather than lessons, so a recall turn does not store the assistant’s own answer back. And the store is tenant-scoped with a real delete: PATCH the tier to correct cheaply, DELETE when a stored fact is wrong enough that keeping it would keep poisoning recall.

How is this different from a markdown file in the repo?+

A file is read in full on every turn, so its cost grows with everything you have ever learned — and an agent has to be trusted to maintain it. Here the store is tiered and budgeted, so the per-turn cost is bounded, the contents are queryable, contradictions and weak domains are represented rather than flattened into prose, and the agent never authors its own memory.

Who can read my memory?+

Only the organisation the key belongs to. Keys are opaque, scoped, individually revocable, and stored as a hash — the secret is shown once. Minting a key requires a signed-in session rather than an existing key, so a leaked agent credential cannot escalate itself, and every read is filtered by organisation id in one auditable place.

Where does the data live?+

In your SQLite file, or your PostgreSQL-compatible deployment of it. There is no telemetry about your memory content, and the service is a single process with a single database file — which is the point: you can read every byte your agent has been told.

Give your agent something it can still remember on Friday.

Create an account, make an organisation, and mint a key. The secret is shown once, because only a hash is stored.