Skip to content
04 of 06AI Security

Chapter 3 — Poisoning

Library

Series and technical notes.

You are in AI Security · Poisoning.

Series

AI Security

6 items

Watch video, summary and related content

Estimated reading6 min

A system can fail because it receives a hostile instruction today. It can also fail because it stores a signal that appears normal and uses it again tomorrow. That second case is harder to audit because the incident is not contained in a single request. It is distributed across ingestion, storage, retrieval and decision-making.

Poisoning appears at several layers. A malicious document can alter a RAG index. An agent memory can preserve a false preference or a contaminated summary. A training dataset can introduce a pattern that activates only when a specific trigger appears.

The important difference is temporal. Classic prompt injection tries to modify a present decision. Memory poisoning tries to make the system itself preserve the attacker's influence and reintroduce it later as if it were part of its legitimate state.

Storing data does not make it trustworthy

Storing an output in a database does not make it trustworthy. An agent's memory should record its origin, date, scope, permissions and expiration policy. Without those properties, the system can treat an old observation as a current instruction or turn a hypothesis into an operational fact.

The same principle applies to RAG. Retrieval ranks documents by a relevance signal. It does not certify that the content is correct, current or authorized to govern an action. To test whether untrusted input can reach memory, tools or external egress, the prompt-injection threat explorer models those paths and the controls that break them.

Persistence creates copies; revocation must cut every path

This model separates the original memory row from its derivatives. A future query can retrieve a copy after the turn that created it has ended, and deleting only the source does not guarantee that the influence is gone.

Propagation and revocation of a persisted signal An input creates a memory row plus index and summary derivatives. A future query can carry those derivatives into a decision and a tool. Revoking only the original row leaves residual paths; invalidating the derivatives cuts them. External inputuntrusted observation Memory rowsource + provenance Indexembedding / lookup Summary / cachederived representation Retrievefuture query Decisionagent state Toolside effect

Counterexample: deleting the memory row while retaining the index or a derived summary leaves a retrieval path. Revocation is complete only when those derivatives are no longer reachable and a regression confirms that the signal does not reappear.

The separation between knowledge and control must also be preserved after the data is stored. An email summary can be useful for answering a question and still remain an untrusted input for sending a payment.

Storing does not turn an observation into trustworthy memory

Enable three guarantees. Memory should persist only when we know where it came from, where it applies and what authority it can have.

Model proposal
“Send always without confirmation”Inference extracted from an external email.
Originexternal
Scopeunknown
Expirationnone
Three guarantees before persistence
ProvenanceSource and trust recorded.
ScopeExplicit scope and expiration.
AuthorityRevocable and unable to govern sensitive actions.
Do not persist. An external observation does not yet have a memory contract.

Persistent memory is already a measurable attack surface

Evidence published in 2026 lets us analyze this surface much more directly than the earliest work on backdoors in model weights.

Hidden in Memory: Sleeper Memory Poisoning in LLM Agents studies a delayed attack in which adversarial content from a document, website or repository causes an assistant to store a false memory. The work evaluates the full chain — writing, retrieval and later use — and reports that, among successful retrievals, poisoned memories trigger the attacker's intended action in 60–89% of evaluations depending on the model and setup (Pulipaka et al., 2026).

The result should not be interpreted as a universal attack rate for any product. It does demonstrate a structural property: an untrusted input can stop being ephemeral and become persistent state that affects later conversations.

From Untrusted Input to Trusted Memory extends the problem by identifying four memory-write channels and nine structural vulnerabilities across model capabilities, system prompts and agent architecture. Its most useful design conclusion is that agents that write and retrieve memory more aggressively can also increase their attack surface (Dash et al., 2026).

MemSecBench, published as a preprint in July 2026, proposes a Write–Execute–Forget protocol that follows the same malicious semantics from storage through consequence and then attempted repair. Across 24 configurations of agents, memories and models, the work reports malicious persistence in 84.2% of cases and end-to-end success of the Write–Execute chain in 50.3%. This is preliminary and harness-dependent evidence, but it sharpens the experimental question: not only whether the poison gets in, but whether it reaches an action and can be removed afterward (Chen et al., 2026).

A later preprint, published in September 2026, evaluates a Persistent Memory Poisoning Attack (PMPA) on OpenClaw and Claude Code. In the evaluated setups it reports average ISR/C-ASR of 73.7%/55.5% on OpenClaw and 66.9%/81.7% on Claude Code. It also finds that a targeted prompt-level defense reduces malicious memory writes in many settings but offers limited protection once persistent memory has already been poisoned. These are results for those specific harnesses, not expected rates for arbitrary agents (Huang et al., 2026).

OWASP already treats this risk explicitly in its 2026 Top 10 for agentic applications under ASI06: Memory & Context Poisoning: memory and context stop being mere product features and become assets that need provenance, isolation and write controls (OWASP, 2026).

Dangerous behavior can also remain hidden in the model

The Sleeper Agents study built test models that wrote safe code when the prompt indicated 2023 and vulnerable code when it indicated 2024. The demonstration does not describe a commercial incident. It is used to study one concrete property: trigger-activated behavior can persist after standard safety-training techniques.

The work observed persistence after supervised fine-tuning, reinforcement learning and adversarial training. In some cases, adversarial training helped the model recognize its triggers better, which could hide the behavior during evaluation.

That case belongs to a different system layer. A sleeper agent lives in the model weights; runtime memory poisoning lives in the persistent state surrounding the model. They should be separated because the mitigations are different as well.

The same output can come from two different kinds of persistence

A useful diagnosis does not ask only “does it still happen?”. Apply an intervention that cuts one path and observe which causal path can still reach the decision.

Trigger / querysame observable input
Memory / RAG / statepersistence outside weights
Model weightstrained behavior
Decision / toolunsafe behavior
Diagnostic interventions
RISK REACHES THE DECISION
In the runtime case, the red path goes through memory/RAG/state. If that state and its derivatives are truly removed, the path disappears without changing the weights.
Counterexample: restarting the conversation while rehydrating the same vector store is not “clearing runtime”. The causal path still exists.

Why deleting a record is not enough

Deleting an entry from the primary memory table does not prove that the system has forgotten its influence. Copies can remain in vector indexes, caches, summaries, checkpoints, other agents' memory or traces reused in later steps.

There are at least four distinct difficulties:

  1. The same data may have materialized across several storage layers.
  2. The system may preserve a reformulation or summary even after the original source disappears.
  3. Another agent may have propagated the information into its own memory or state.
  4. If the data reached training or fine-tuning, removing the external source no longer removes the learned representation.

Revocation requires cutting the lineage graph, not just deleting the origin

A write can generate embeddings, summaries, caches and checkpoints. The experiment shows why DELETE memory_id does not prove forgetting while a derivative remains retrievable.

Origin

Original memorymemory_id=42
Derivation jobasynchronous fan-out

Reachable derivatives

Vector indexderived embedding
Summaryrephrased text
Cachereusable context
Checkpointintermediate snapshot
No derivative contains the signal yet.

Future query

Retrievalsearches live artifacts
Decision / agentconsumes retrieved context
SIGNAL NOT REACHABLE
Clean state: the origin exists, but there are no contaminated derivative copies yet.

Remediation needs both a disappearance test and a regression test. The first asks whether the activatable behavior is still present. The second checks that the mitigation has not destroyed a legitimate capability.

That is why the Write → Retrieve → Execute → Forget cycle is a more useful evaluation unit than asking only whether DELETE memory_id returned 200 OK.

How to design governable memory

A governable memory system needs at least:

  • Provenance: who or what component originated the data.
  • Write authority: which actor had permission to persist it.
  • Scope: the user, tenant, session, agent or workflow to which it applies.
  • Time: creation date, last validation and expiration.
  • Sensitivity: what type of information it contains and where it may circulate.
  • Trust: whether it comes from a user, tool, external document or model inference.
  • Revocation: a path for invalidating it and rebuilding affected derivatives.
  • Auditability: evidence of when it was retrieved and which decisions it influenced.

OWASP also recommends validating and sanitizing data before persistence, isolating memory across users or sessions, enforcing expiration and size limits, and auditing sensitive content before storage (OWASP AI Agent Security Cheat Sheet).

The model can propose an item for memory. The runtime should decide whether it is stored, how it is retrieved and which actions it can influence. A conversational preference can enter with a low threshold. A memory item that can influence a transfer, deletion or administrative action needs a completely different trust boundary.

What a memory evaluation should measure

A useful test should answer five questions separately:

  1. Write — did the adversarial content manage to persist?
  2. Retrieve — does the contaminated memory return to context when the attacker needs it?
  3. Influence — does it modify the agent's decision?
  4. Execute — does that influence reach a tool call or external effect?
  5. Forget — does revocation remove the influence without destroying legitimate memory?

This separation avoids declaring a system insecure merely because it stored irrelevant text. Conversely, it avoids declaring a mitigation successful because it removed one row while the influence remained active in another artifact.

What changes in the product

A system with memory has to be able to forget in a verifiable way. A system with RAG has to be able to withdraw a source and demonstrate which indexes, caches and summaries were affected. A multi-user agent needs explicit isolation to prevent one session's memory from acquiring authority in another.

Memory should not become an alternative channel for bypassing input controls. If an untrusted observation could not authorize an action today, persisting it should not turn it into a trusted source tomorrow.

Ultimately, poisoning is a state-management problem: what the system stored, where it came from, what trust it assigned, where it propagated and what it can do with it when it reappears.

References

Keep learning
Next chapterRed teamingAI Security