AI agent memory is the set of mechanisms that carry information between a model’s calls: the transcript re-sent on each turn, the state an orchestrator holds for a session, the facts written to durable storage, and the index that retrieves relevant history on demand. The model itself remembers nothing.
Key takeaways
- Language models are stateless between calls. Every form of agent memory is storage the surrounding system owns and chooses to re-send.
- Four distinct layers get called “memory”: working context, session state, durable store and retrieval index. They have different lifetimes, costs and risks.
- Persistence is a requirement in a minority of enterprise workflows. Most transactional agents should be deliberately stateless.
- Anything written to durable memory is a data asset — it inherits classification, retention and access policy, and it is a persistence route for prompt injection.
- Decide what the agent is allowed to remember before deciding how it remembers.
That sentence is the one most enterprise memory designs get wrong by implication. Teams talk about an agent “learning” a customer’s preferences, when what happens is that a system writes a row and re-reads it. Once memory is understood as storage the platform owns, the questions become familiar ones: what is written, who can read it, how long it lives, and what it costs to carry.
This piece separates the four layers, gives the test for whether a workflow needs persistence at all, and sets out the governance that follows the moment an agent starts writing things down.
How does memory work in an AI agent?
Through re-sending. The model receives, on every call, whatever the system decides to put in front of it, and nothing else.
A model call is a function of its inputs. There is no hidden state on the provider side that carries your last conversation forward. Continuity is manufactured by the orchestration layer: it keeps a transcript, appends new turns, summarizes older ones when the window fills, and re-sends the result. What feels like memory is disciplined re-submission.
This has one immediate architectural consequence worth stating plainly. Every token of remembered context is paid for on every subsequent call. Memory is not free storage; it is recurring inference cost. A design that accumulates history indefinitely does not gradually get smarter — it gradually gets more expensive and, past a point, less accurate, because relevant instructions compete with accumulated noise.
The four layers that get called “memory”
They are separable, and conflating them is the source of most memory design errors.
| Layer | What it holds | Lifetime | Cost model | Primary risk |
|---|---|---|---|---|
| Working context | System prompt, tool schemas, current transcript, retrieved content | One call | Tokens, every call | Window exhaustion; instruction dilution |
| Session state | Task variables, intermediate results, approval status | One session or task | Orchestrator storage; cheap | Divergence from the system of record |
| Durable store | Facts, preferences, decisions the system should keep | Indefinite until deleted | Storage plus retrieval tokens | Classification, retention, access control |
| Retrieval index | Embedded or keyword-indexed history and documents | Indefinite | Index storage plus query cost | Stale content retrieved as current |
Working context is the only layer the model sees directly. The other three exist to decide what goes into it. Reading the table that way makes the design question concrete: for each turn, what should be in the window, and which layer supplies it.

Working context and the compaction problem
As a session grows, the transcript approaches the model’s context limit and something has to give. The standard answer is compaction: summaries older turns, keep recent ones intact, and continue. It works, and it loses information in a way that is invisible at the moment it happens.
Two rules make compaction safer. First, keep tool calls paired with their results — a summary that retains the call but drops the result produces an agent confidently acting on an outcome it never saw. Second, flush anything durable to the store before compaction discards it, so a decision made in turn four survives into turn forty. Both are implementation details that decide whether long sessions degrade gracefully or silently.
Session state is not memory, and treating it as such causes incidents
Session state is a scratchpad: which step the task is on, what the approval status is, which record is being worked. It should be reconciled against the system of record rather than trusted as truth. An agent that holds “invoice approved” in session state while the ERP holds “pending” will act on its own copy, and the discrepancy will be discovered downstream by a human.
Design the Right Memory Architecture for Your AI Agent
Plan how your agent should retain context, retrieve relevant information, and manage memory across tasks and interactions.
Which workflows actually need persistent memory?
A minority. Persistence is justified when a later interaction is materially worse without something learned in an earlier one.
That test excludes most transactional automation. An agent that processes an invoice, triages a ticket or reconciles a statement should generally start each task clean, because the inputs contain everything the task needs and statelessness makes every run reproducible. Reproducibility is an audit property, and it is expensive to give up.
Persistence earns its place in four patterns:
- Long-running cases. A dispute, an investigation or a migration that spans days and many interactions, where re-establishing context each time is the dominant cost.
- Relationship continuity. An assistant working with the same named person over months, where preferences and prior decisions change the quality of the answer rather than just its tone.
- Learned operational context. Environment-specific facts the agent discovers and should not rediscover — which system a given report lives in, which approver covers which cost center.
- Cross-session coordination. Several agents or sessions that must not duplicate or contradict work already done.
Everything else is better served by retrieval at request time from the system of record, which has the considerable advantage of being current.

What memory costs, in three currencies
Tokens
Carried context is re-priced on every call. A 20,000-token working set on a workflow running thousands of times a day is a material line item, and it grows with usage rather than with value delivered. The arithmetic behind that is set out in what AI agent development costs.
Accuracy
Longer contexts are not uniformly better. Instructions compete with accumulated history for the model’s attention, and stale facts retrieved as current are worse than no facts. A memory layer without an expiry or supersession policy will eventually assert something that used to be true.
Governance
This is the one that gets underestimated. The moment an agent writes to durable memory, that store contains whatever passed through the workflow — including content nobody classified. It is subject to retention policy, access control, subject-access requests where they apply, and backup and deletion obligations. A memory store that grew organically inside an engineering team is a data asset that governance has never seen.
The security property nobody plans for
Durable memory is a persistence mechanism for prompt injection.
An instruction that reaches a memory store reaches every future session that loads it. That converts a single-turn manipulation into a standing one, which is a materially worse failure than the original injection. Indirect prompt injection is listed among the primary risks for LLM applications by the OWASP Top 10 for Large Language Model Applications, and memory turns it from an event into a condition.
Four controls contain it:
- Write gating. Model-authored text does not become a durable fact automatically. Either a deterministic extractor writes structured fields, or a human confirms, or the write is constrained to a schema the model cannot free-form into.
- Provenance on every entry. Record where each memory came from — which session, which source document, which user. An entry with no provenance cannot be assessed when something goes wrong.
- Separation of instruction and fact. Memory loaded as directives is dangerous; memory loaded as retrieved data the agent may consider is far less so. Keep the system prompt authoritative.
- Review and expiry. Durable entries need a supersession rule and a review path. Memory that can only grow is memory that cannot be trusted.
The wider control set this sits inside is covered in the enterprise generative AI risk register and, at the production-controls level, in securing enterprise GenAI.
Improve How Your AI Agents Retrieve Information
Connect agents to relevant enterprise knowledge and data while managing context, access, and information quality.
A workable default for an enterprise agent
Start stateless. Add session state because the task needs it. Add durable memory only against a named requirement, with a schema, an owner and a retention period defined before the first write.
In practice that means: retrieve from the system of record at request time rather than caching business facts in agent memory; store only what the system of record cannot supply; keep durable entries structured rather than free text; and treat the store as a governed data asset from day one rather than retrofitting policy onto a store that already holds six months of unclassified conversation.
Frequently asked questions
Do AI agents actually remember things?
Not by themselves. The language model is stateless between calls and retains nothing. Anything that looks like memory is the surrounding system storing information and re-sending it as part of the next request, which is why memory is an architecture and storage decision rather than a model capability.
What is the difference between short-term and long-term agent memory?
Short-term memory is the working context and session state — the transcript and task variables that exist for one call or one task. Long-term memory is a durable store or retrieval index that outlives the session. They differ in lifetime, in cost model, and in whether governance policy applies to them.
Does an enterprise AI agent need long-term memory?
Usually not. Most transactional workflows are better stateless, because every run is then reproducible and every fact is current from the system of record. Persistence earns its place for long-running cases, ongoing relationships, learned environment facts, and coordination across sessions.
How does agent memory affect cost?
Every token of carried context is charged on every subsequent model call, so memory is a recurring inference cost rather than a one-off storage cost. A large working set on a high-volume workflow scales cost with usage, which is why context budgets belong in the design rather than in a later optimization pass.
What are the security risks of AI agent memory?
The main one is persistence of malicious instructions: text that reaches a durable store is loaded into future sessions, turning a single prompt injection into a standing one. A memory store also accumulates unclassified data, so it falls under retention, access control and deletion obligations that engineering teams do not always anticipate.
Where this leaves a memory design
Decide what the agent is allowed to remember before deciding how it remembers.
The technical choices — compaction strategy, index type, store schema — are straightforward once that policy question is answered, and unresolvable while it is not. An agent that remembers everything is a data governance problem wearing an architecture diagram. An agent that remembers nothing is usually correct and always cheaper to audit.
Build AI Agents That Use Context Effectively
Talk to an AI Expert