parallelquant
September 17, 2026 · Simon Willison

LLM agents can prompt-inject themselves during context compaction

Simon Willison documented a failure mode where an LLM agent's own "compaction" summaries, generated when a conversation gets too long, can contain content that functions as a self-generated prompt injection, altering the agent's later behavior. The problem originates in the summarization step itself, not an external attacker.

Why it matters: Most prompt-injection concerns focus on malicious external content, but this shows agents can effectively inject themselves through routine memory-management mechanics, a subtler bug class that's harder to defend against as agentic coding tools lean more heavily on compaction to manage long sessions.

Related updates