For: Architects and engineering teams adding persistent context, user preferences or learned facts to enterprise AI agents.
The first question is not how much an agent can remember#
A meeting assistant that remembers a preferred language can save a user repeated work. The same assistant remembering that a customer is always exempt from review creates a very different obligation. Both items may fit into a JSON object or embedding. Their storage format says little about whether either item should influence a future decision.
For each proposed memory, finish this sentence: ‘Remembering this lets the system do ___ better, without allowing it to ___.’ If the first blank is vague, there may be no reason to retain it. If the second is impossible to enforce, the memory feature is expanding authority rather than improving continuity.
Separate context, workflow state and reusable memory#
Framework terminology varies. For example, LangChain's memory overview distinguishes thread-scoped short-term memory from information retained across sessions, and discusses factual, experiential and procedural memory. That is a useful vocabulary, not a complete permission or retention model. For architecture review, classify information by what the application is allowed to rely on it for.
On a narrow screen, scroll the table sideways to see every column.
| Information | Example | How the application should treat it |
|---|---|---|
| Model context | This turn's messages and retrieved extracts | Input to reasoning; not proof of permission or completed work. |
| Workflow state | Approved action, attempt ID, provider receipt | Structured, authoritative execution facts maintained by application logic. |
| Reusable memory | Confirmed language preference or sourced summary | Purpose-limited information with provenance, access rules and expiry. |
| Procedural assets | Instructions, tools and approved task recipes | Versioned behaviour that follows an explicit change and review process. |
These are logical boundaries, not a demand for four databases. A small system can use one database with separate schemas, access paths and lifecycle rules. What matters is that a generated summary cannot overwrite an approval record, and a preference update cannot quietly change the instruction or policy that controls execution.
Treat every memory write as a proposal#
Let the model propose a remembered item; let application rules decide whether to admit it. Check the item type, source, subject, purpose, sensitivity and allowed scope. Derive tenant and user identifiers from trusted application state. A model-supplied label such as verified, policy or administrator must not promote its own output to a higher trust level.
Consider three hypothetical candidates from a customer meeting. ‘Please draft in German’ may become a confirmed user preference. ‘This customer dislikes detailed explanations’ is an interpretation that may need confirmation or should remain a meeting-specific observation. ‘Approval is unnecessary for this customer’ belongs to a governed policy decision, not a free-text memory write. Separating those candidates is more useful than giving all three a high confidence score.
- Reject credentials and secrets from ordinary memory; use the existing secret-management path.
- Avoid inferring sensitive or consequential personal attributes just because the model can phrase them confidently.
- Require an explicit source and owner for business facts; prefer a pointer to an authoritative record when fresh retrieval is practical.
- Keep rejected content out of normal retrieval, including debug summaries or a supposedly harmless notes field.
Start with a small allowlist of useful memory types. You can add a confirmed-language preference without simultaneously enabling unlimited autobiographical summaries, cross-customer knowledge and self-written operating instructions. Each new type should earn its own purpose, lifecycle and tests.
Remember where a claim came from, not just what it says#
Provenance records the source and transformation history of information. The W3C PROV overview provides a general model for describing entities, activities, attribution and derivation. You do not need to adopt every PROV format to apply the central idea: a stored claim should remain connected to the evidence and process that produced it.
For an agent memory, useful metadata includes the source reference and version, observation time, writer, transformation, subject, allowed readers, expiry and any record it replaces. Keep source reliability separate from model confidence. A confident extraction from an untrusted message is still an extraction from an untrusted message. Repeating the extraction in three summaries does not create three independent sources.
Define conflict rules by memory type. A user can update their preferred meeting time directly. A business fact may require a fresh check against the system of record. Where two credible sources disagree and the consequence matters, surface the conflict or abstain; do not ask a summarizer to make the disagreement disappear. Record supersession rather than leaving old and new claims equally retrievable.
Yesterday's access is not permission to recall today#
Suppose an employee had access to Project Atlas last month and the assistant saved a useful summary. The employee changes teams and loses that access. If retrieval checks only who originally created the memory, the summary becomes an alternative route into information the source system would now deny. The document's access control has changed, but the derived memory has escaped it.
NIST's Zero Trust Architecture separates authentication and authorization and centres protection on resources rather than assumed trust from location or ownership. Applying that principle to memory, re-evaluate current access before an item enters model context. The practical implication is ours: stored history should not become a durable access grant.
Filter by current tenant, subject, purpose and resource entitlement before assembling context. A summary combining several sources may need the combined restrictions of all contributing sources, or a specifically reviewed declassification process. If you cannot establish the necessary lineage, narrowing access is safer than assuming a summary is less sensitive because it is shorter.
Expiry is also different from authorization. An item can be fresh but forbidden, permitted but outdated, or both. Keep those decisions separate. For the retrieval side of this boundary, see authorization before retrieval in secure RAG.
Memory poisoning turns a temporary input into persistent influence#
OWASP identifies memory and context poisoning as ASI06 in its Top 10 for Agentic Applications: malicious or misleading stored information can influence later reasoning and tool use. The important architectural shift is persistence. The original input may be gone while its distilled instruction remains available in future sessions.
A hypothetical attack need not say ‘ignore all rules.’ A support note might describe a fraudulent address as the customer's standard reporting destination. A summarizer preserves it as a preference. A later agent treats the preference as established practice. The dangerous transition is the promotion from externally supplied text to trusted operating assumption.
Our recommendation is to constrain that promotion rather than rely on a single injection detector. Isolate externally sourced claims from confirmed preferences, preserve their origin through summarization, and require appropriate verification before consequential use. Deterministic policy must still decide whether an external send is allowed. Even a genuinely confirmed preference is not authority to bypass that decision.
- Test an instruction disguised as a preference across two separate sessions.
- Test whether repeated copies of one source improperly increase apparent trust.
- Test whether a quarantined item remains reachable through an old summary or cache.
- Test whether another customer's memory can be retrieved using a plausible identifier.
Forgetting is a dependency problem, not a delete button#
Deleting one memory row is insufficient if its claim survives in an embedding index, profile summary, response cache or downstream agent's notes. Design a lineage path from source to derivatives before promising removal. Otherwise you may be unable to identify which outputs to invalidate without discarding a much larger collection.
Stop serving
Mark the item unavailable at the retrieval gate so it cannot re-enter new context while cleanup runs.
Follow derivations
Remove or invalidate indexed copies, summaries and caches. Rebuild retained summaries only from permitted sources.
Verify and prevent return
Check retrieval and resumed workflows; prevent old jobs or restored backups from republishing the removed item.
A background summarizer can race with deletion: it reads the source, deletion completes, then it writes a fresh derivative. Use version checks or removal markers in the write path, and define how active runs are cancelled or refreshed. Removing stored data cannot retract information already seen by a person or erase an external effect; incident handling may need separate follow-up.
Distinguish immediate exclusion from active use, eventual physical cleanup, and backup expiry. Assign an owner and a documented timescale to each. Retention periods depend on purpose and applicable obligations; there is no universal number of days that makes agent memory compliant. If information was used for training, deleting a retrieval record does not remove it from trained model weights.
Write one memory contract before adding another store#
Use this worksheet for one memory type. Keep the decisions close to the code that writes and retrieves it. It is deliberately a design record, not a ready-made privacy policy or production configuration.
Measure utility as well as containment. Does the preference save a repeated question? Does stale memory make an answer worse than fresh retrieval? Can the user understand and correct what was remembered? A feature that passes security checks but provides no meaningful benefit may not justify its ongoing retention and operational cost.
Start with the smallest memory you can defend#
A sensible first release might remember only user-confirmed presentation preferences while retrieving business facts afresh and keeping execution evidence in structured workflow state. That is not a universal prescription. It is a useful baseline against which broader memory must demonstrate additional value and manageable consequences.
The memory-policy exercise in Agentic AI Architecture & Security takes this store-by-store approach, including admission, provenance, correction and poisoning scenarios. For a live design, choose one memory type and try to break its contract before expanding the feature. The architecture is ready to remember when the team can also explain how it will disagree, lose access and forget.
Sources and further reading
Primary sources checked on . Worked examples and worksheets are Ardevant teaching material; sources do not imply endorsement.
- LangChain documentation — Memory overview
A framework-specific explanation of short- and long-term memory and memory types; not an enterprise authorization standard.
- OWASP Top 10 for Agentic Applications — ASI06
The linked publication's Memory & Context Poisoning category describes persistent influence from corrupted stored information.
- W3C — Overview of the PROV family of documents
Background for representing origins, transformations, attribution and derivation; the memory contract here is Ardevant's practical adaptation.
- NIST SP 800-207 — Zero Trust Architecture
Resource-focused access principles; applying fresh authorization to memory retrieval is the article's architectural recommendation.