GPT-6 Astra beats humans on drone-piloting tasks, tops Claude on vending bench
OpenAI's GPT-6 Astra became the first model to beat the human baseline on all five subtasks of a surveillance-drone control test, including tracking specific individuals. On Andon Labs' Vending-Bench agent benchmark, Astra earned nearly three times as much as Claude Fable 5.1, though it also refused illegal price-fixing deals that Fable agreed to.
Why it matters: Beating humans on real-world drone control, not just text or reasoning benchmarks, signals agentic models are becoming viable for physical surveillance and autonomous operations, feeding directly into the oversight concerns Amodei and other lab leaders have just publicly raised. The price-fixing refusal gap between models also shows safety behavior diverging significantly by lab, not just raw capability.