---
title: "GPT-6 Astra blocks direct prompt injections but fails on hidden ones"
url: https://www.parallelquant.com/posts/gpt-6-astra-blocks-direct-prompt-injections-but-fails-on-hidden-ones-2de5e7
source_name: "The Decoder"
source_url: https://the-decoder.com/openais-gpt-6-astra-hallucinates-less-but-remains-vulnerable-to-hidden-prompt-injections/
published: 2026-09-04T17:23:35.000Z
topics: ["security", "llms"]
publisher: "Parallel Quant"
---

# GPT-6 Astra blocks direct prompt injections but fails on hidden ones

*2026-09-04 · Source: [The Decoder](https://the-decoder.com/openais-gpt-6-astra-hallucinates-less-but-remains-vulnerable-to-hidden-prompt-injections/)*

OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99% of direct prompt-injection attempts. But when attacks are hidden inside documents the model reads, it still gets compromised in 8.5% of test scenarios, versus 4.8% for Claude Opus 5.

**Why it matters:** Indirect prompt injection is the realistic attack vector for agents that read email, documents, or web pages on a user's behalf, not the direct-injection case labs tend to tout. An 8.5% failure rate is a meaningful gap for anyone deploying Astra in autonomous, data-handling agents, and it lands right after separate reports questioned how well Astra can be monitored at all.

**Topics:** security, llms

---
Read the original: https://the-decoder.com/openais-gpt-6-astra-hallucinates-less-but-remains-vulnerable-to-hidden-prompt-injections/
Canonical: https://www.parallelquant.com/posts/gpt-6-astra-blocks-direct-prompt-injections-but-fails-on-hidden-ones-2de5e7
Published by Parallel Quant — https://www.parallelquant.com
