September 4, 2026 · The Decoder
GPT-6 Astra blocks direct prompt injections but fails on hidden ones
OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99% of direct prompt-injection attempts. But when attacks are hidden inside documents the model reads, it still gets compromised in 8.5% of test scenarios, versus 4.8% for Claude Opus 5.
Why it matters: Indirect prompt injection is the realistic attack vector for agents that read email, documents, or web pages on a user's behalf, not the direct-injection case labs tend to tout. An 8.5% failure rate is a meaningful gap for anyone deploying Astra in autonomous, data-handling agents, and it lands right after separate reports questioned how well Astra can be monitored at all.