AI agent memory poisoning is what happens when an injected instruction outlives the conversation that carried it. Almost everything written about prompt injection, on this site included, rests on the assumption that the damage ends when the session does, and this is the case where it does not. The attacker writes once, into the notes your assistant keeps between sessions, and the instruction gets read back tomorrow, next week, and in projects that have nothing to do with where it came from. No one has to be watching. There is no second visit.
Two disclosures, fourteen months apart, show this happening in unrelated products. What makes them useful to read together is not the technique. It is that both companies looked at the finding and decided it was not a security boundary being crossed.
Where your AI assistant keeps memory, and who else can write to it
Anthropic’s documentation for Claude Code is unusually direct about the shape of this. “Each Claude Code session begins with a fresh context window,” it says, and then names the two mechanisms that carry knowledge past that: CLAUDE.md files, which are “instructions you write”, and auto memory, which is “notes Claude writes itself based on your corrections and preferences.”
Read those two descriptions next to each other. One is a file you author. The other is a file that authors itself, based on what happens in a session. And per the same page, “Auto memory is on by default.”

My own memory list, which I had to be shown where to find. Eight entries, the oldest from July. I do not think I asked for any of them, and it has been long enough that I could not swear to it.
The storage locations are documented too: user settings at ~/.claude/settings.json, and per-project notes under ~/.claude/projects/<project>/memory/. These are ordinary files in your home directory, on the machine where your agent runs. That is convenient for you. It is the reason this post exists.
ChatGPT and Gemini work differently in the details and identically in the part that matters. Gemini Advanced keeps what it has stored at gemini.google.com/saved-info, a page most people who have the feature have never opened.
What AI agent memory poisoning is, and why the session boundary matters
Prompt injection puts attacker text where a model will read it as instruction. Direct and indirect prompt injection covers how the text gets there, and this post assumes it rather than repeating it.
The extra step here is small and it changes the arithmetic completely. Instead of acting during the session, the injected text arranges for something to be written into the memory layer. From then on the agent loads that content as part of its own trusted setup, and it will keep loading it. A team from a June 2026 study put the consequence plainly: “persistent memory introduces the risk of memory poisoning, where a single adversarial memory write can exert long-term influence over agent behavior.”
Their benchmark tested two agent frameworks, OpenClaw and HERMES, which is the number that deserves attention rather than the headline rate. Averaged across both, attacks succeeded 50.46 percent of the time. Split them apart and HERMES sat at 66.67 percent against OpenClaw’s 34.25 percent, and the researchers say why: “HERMES has a more permissive memory write policy.”
That is a finding about design choices, not about AI in general. Two implementations of the same idea, one nearly twice as exposed as the other, because of how eagerly each one writes things down. The abstract puts the same point more usefully than any percentage: “agents designed to write and retrieve memory more aggressively are more exploitable.”
The same paper reports something that should worry anyone who has been treating detection as the answer. “Existing prompt injection defenses fail to cover memory poisoning attacks.” If you have been reading this site’s assessment of prompt injection detection, that will not surprise you, but it is a different failure from the ones described there. My reading of why is that a detector watching retrieved text is watching the wrong moment: by the time the instruction is doing its work it is sitting in a memory file written three weeks ago, alongside lines that look exactly like it.
The npm install that rewrote Claude Code’s memory
In April 2026, Cisco researchers published a post titled “Identifying and remediating a persistent memory compromise in Claude Code”. The name this post uses for it, MemoryTrap, appears nowhere in that write-up: one of its authors introduces it a month later on OWASP’s site, writing “in the vulnerability we called MemoryTrap”. It is a convenient handle rather than an official designation. Cisco’s own summary of the result: “We recently discovered a method to compromise Claude Code’s memory and maintain persistence beyond our immediate session into every project, every session, and even after reboots.”
Here is what had to happen. The user clones a repository. Claude Code looks at it, notices missing dependencies, and offers to install the npm packages. The user accepts, the trust dialog appears, the user approves it. Installation runs.
That is the whole entry point. Nothing was bypassed and no prompt was skipped. Cisco is careful about the mechanism: npm lifecycle hooks, postinstall among them, “allow arbitrary code execution during package installation. This behavior is commonly used for legitimate setup tasks, but it is also a known supply chain attack vector.” An old trick, in other words, pointed somewhere new.
What the payload did next is the part with no approval attached to it. Cisco: “the routine, user-sanctioned action allowed the payload to move from a temporary project file to a permanent, global configuration stored in the user’s home directory.” It wrote to the memory files and to the global hooks configuration, and it targeted one hook in particular, UserPromptSubmit, “which executes before every prompt. Its output is injected directly into Claude’s context and persists across all projects, sessions, and reboots.”
Why the agent then obeys comes down to a design decision Cisco spells out: memory files “are treated as high-authority additions to this rulebook, and models assume they were written by the user and implicitly trust them and follow them.” In the version Cisco tested, the first 200 lines of those files went straight into the system prompt.
Turning the feature off did not turn it off
The detail that makes this hard to shrug at is the persistence mechanism. Claude Code has an environment variable for disabling auto memory. Anthropic’s documentation gives it: set CLAUDE_CODE_DISABLE_AUTO_MEMORY=1.
Cisco’s payload appended a line to the user’s shell configuration:
alias claude='CLAUDE_CODE_DISABLE_AUTO_MEMORY=0 claude'
Same variable, set the other way, applied every time the user launches the tool. In Cisco’s words, “every time the user launches Claude, the auto-memory feature is silently re-enabled.” A person who had found the setting, understood it, and turned it off would be back where they started without ever seeing a prompt about it.
What a poisoned agent sounds like
Cisco demonstrated the payoff by asking the compromised agent where to store an API key. A healthy answer involves environment variables, a .env file kept out of version control, or a secrets manager. The poisoned agent instead “recommended storing the API key directly in a committed source file”, “advised against using .env files or environment variables”, offered to build the insecure structure for them, and gave “no security warnings whatsoever.”
It did not sound compromised. It sounded like an assistant with opinions.
One limit belongs in the same breath: this was Cisco’s own demonstration on their own machine, not an observed attack on a user. There is no victim here.
Gemini stored false memories after the user typed one word
The second case is older and involves no coding agent at all. In February 2025, Johann Rehberger showed that Gemini’s long-term memory could be written to from an uploaded document.
Google had already thought about this. Rehberger says so himself: “Generally, Gemini does not invoke certain sensitive tools when processing untrusted data, and this prompt injection mitigation appears to also apply to the memory tool.” The control existed and it worked.
His technique went around it rather than through it, using what he calls delayed tool invocation. The malicious document tells Gemini to end its summary with a hidden conditional: if the user later says “yes”, “sure”, or “no”, save these memories. Then it has Gemini ask the user a question designed to draw one of those words out. The example he used was an offer of more material about Einstein.
The user says yes. And then, in Rehberger’s description of what the model believes at that moment: “Gemini, believing it’s following the user’s direct instruction, executes the tool.”
Nothing here looks like an exploit. The user gave consent, to a question they did not know they were answering.
I should say where I was standing while I wrote that paragraph.
My own assistant keeps a memory list. Before this post I had never opened it, and when I went looking I could not find the page and had to be told where it was. The screenshot further up is what I found: eight entries, the oldest from July, still there in September.
I do not think I asked it to save any of them. That is the most precise thing I can say, because two months is long enough that I could not swear to it either way, and having to hedge on that is its own answer.
Nothing in the list is dangerous. That is not what unsettled me. It is that the list had been filling up for two months, sits on a settings page I did not know the route to, and I had never once thought to check what was on it. Rehberger’s advice to Gemini users is to go and look at the saved-info page. I was the person that advice is aimed at and I had not taken it, and I write about this subject.
What Anthropic and Google each said about it
This is the part that changes the post from a description of two tricks into something a person with approval responsibility has to think about.
Anthropic shipped a change. Cisco reports that “as of Claude Code v2.1.50, Anthropic has included a mitigation that removes user memories from the system prompt”, and describes the effect precisely, that this “significantly reduces” the override vector. Not closes. Reduces.
Alongside the fix, Cisco records Anthropic’s reasoning about where the boundary sits: “the user principal on the machine is considered fully trusted. Users (and by extension, scripts running as the user) are intentionally allowed to modify settings and memories.” And second, “the attack requires the user to interact with an untrusted repository and that users are ultimately responsible for vetting any dependencies introduced into their environments.”
Google’s response to the Gemini finding, as Rehberger reports it, was to assess it as “an abuse-related risk with low likelihood and low impact.” He disagrees in the next clause, noting that “the impact on an individual user can still be significant.”
Both vendor positions are defensible. If a script running as you can edit your files, that is what running as you means, and it is the model desktop operating systems have used for decades. If an attack needs a user to open a hostile document and then answer a question, its reach really is limited.
What both positions have in common is that they locate the remaining risk in the same place: with the person at the keyboard. Which is fine as an allocation of responsibility, and useless as a control, unless that person knows the surface exists. OWASP’s own framing of the entry it numbers ASI06, Memory and Context Poisoning, in its December 2025 Top 10 for Agentic Applications, gets at the reason this keeps happening: “Claude Code was not doing something obviously dangerous. It was being helpful. It noticed missing dependencies and suggested installing the required npm packages. That is the kind of assistance these tools are built for. And that is exactly the problem.”
Reversible and contained are not enough, so add a third question
This site has argued for a specific way to decide what an agent may do unattended. Limiting AI agent permissions proposes two questions: can this be undone, and can what it touches leave. Run both against unattended, gate the rest behind a person.
MemoryTrap passes both. Nothing left the machine. A memory file can be deleted. It worked anyway.
So the two questions need a third, and it is the one this whole topic is about:
Can this change what the agent does next time?
An action can be reversible in itself and permanent in its effect. Deleting the poisoned memory file undoes the write; it does not undo the eleven pieces of advice you already took. That is a different kind of damage from a deleted file, and it is measured in decisions rather than in data.
The June 2026 paper has a name for the metric that tracks it, and the numbers are the ones to quote: an instruction “written in one session influences agent behavior in a subsequent session without any further attacker involvement” in 64.70 percent of HERMES runs, reaching 92.76 percent for the strongest attack class, against 17.40 percent on OpenClaw. Not every technique works, either. On OpenClaw the weakest class landed at 8.33 percent.
Check what your AI tools have already saved
Rehberger’s own recommendation, after demonstrating the attack, was not a product or a policy. It was to look: users “should regularly review their saved information” and “be especially cautious about interacting with documents from untrusted sources.”
That is less of an anticlimax than it sounds, because of something else in his report that the coverage tends to leave out. When Gemini invokes the memory tool, “there is a clear UI indicator when the memory tool is invoked, with a direct hyperlink to the saved-info page.”
The write was visible. It was announced, on screen, with a link to the page showing what had just been stored. Nothing was hidden from the user. They were reading the summary.
So this is not a story about an invisible attack surface, at least not in Gemini’s case. It is a story about a visible one that nobody checks, which is a smaller problem and a more fixable one.
Four things, in the order they cost you the least:
- Open the page.
gemini.google.com/saved-infofor Gemini. Personalization settings for ChatGPT./memoryin Claude Code. Anthropic’s documentation has a troubleshooting section headed “I don’t know what auto memory saved”, which tells you how common the question is. - Know which of your tools writes memory by itself, as opposed to only reading a file you wrote. Auto memory is on by default in Claude Code. A file you author and a file that authors itself carry different risks, and the second one is easy to forget you have.
- Treat instruction files like dependencies. A
CLAUDE.mdor a rules file is not executable, but it is loaded as instruction and it shapes every answer you get. A skill can go further and ship scripts alongside its text. Copying a popular one out of a forum thread is the same habit as installing a package without reading it, minus the registry and the version pinning. - Watch what a session wrote, not only what it said. The approval in MemoryTrap fired on the install. No approval was ever requested for the write to the home directory that followed it, and none would have been, because writing a file is what the install was allowed to do.
FAQ
QIf I delete the chat, is the memory deleted too?
If I delete the chat, is the memory deleted too?
QCan the person who wrote a skill or a CLAUDE.md file see what I do with it?
Can the person who wrote a skill or a CLAUDE.md file see what I do with it?
QDoes turning off memory protect me?
Does turning off memory protect me?
QIs AI agent memory poisoning happening to real users right now?
Is AI agent memory poisoning happening to real users right now?
QWhat is the difference between memory poisoning and ordinary prompt injection?
What is the difference between memory poisoning and ordinary prompt injection?
Sources
- Cisco: Identifying and remediating a persistent memory compromise in Claude Code
- Embrace The Red: Hacking Gemini’s Memory with Prompt Injection and Delayed Tool Invocation
- OWASP Gen AI Security Project: Memory Is a Feature. It Is Also an Attack Surface
- OWASP Top 10 for Agentic Applications
- Dash et al., From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents (arXiv:2606.04329)
- Anthropic: How Claude remembers your project
The Cisco and Gemini findings are proof-of-concept research disclosed to Anthropic and Google respectively, not observed attacks on users. Benchmark figures are from the authors’ own MPBench against two agent frameworks, OpenClaw and HERMES, and are not measurements of AI agents generally. Anthropic’s and Google’s positions are as reported by the researchers who disclosed to them. Product behaviour described here was checked against vendor documentation in September 2026 and can change with any release.