Yesterday a customer told the support agent they are allergic to latex. Today they write "order the same as last time", and the agent asks which gloves. The agent did not forget; it never remembered. Everything the model knows is in its context, assembled by your code on each request.

The model remembers nothing

Every model call is independent. Memory is three layers: the request context, session memory for the current conversation, and long-term memory that outlives it.

Session memory

A list of messages stored by conversation id and resent on every step. Long conversations get compressed: recent messages verbatim, older ones replaced by a summary that keeps numbers, amounts and decisions. Store the state outside the process so a restart does not drop the conversation.

Long-term memory: facts, not transcripts

Store facts extracted from conversations, each with a source and date, and replace outdated facts rather than silently deleting them.

{
  "user_id": "u-881",
  "facts": [
    { "text": "Latex allergy, nitrile gloves only", "source": "chat-2026-10-04", "kind": "preference" }
  ]
}

Getting memory in at the right moment

Fetch by structure (always include delivery preferences when ordering) and by meaning via embeddings, with a cap on count. Give the agent a "remember" tool for explicit requests.

What must never be stored

Card numbers, ID documents, passwords and one-time codes are scrubbed before extraction. Users can see and delete their memory, including embeddings. Search is always filtered by user id in the store, never by asking the model.

In short

  • Memory is code that builds context per request: context, session, long-term.
  • Compress long sessions; keep state outside the process.
  • Long-term memory holds facts with source and date.
  • Retrieve by kind and by meaning, filtered by user in storage.
  • Never store secrets; let users see and delete their memory.