AI coding agents speed up research software but can't verify the science
A field report from OpenAI and academic partners found AI coding agents can modernize neglected research software with speedups of up to 60x. Participants said the agents were 'eloquent, convincing, and confidently wrong' in ways that are easy to miss, shifting the bottleneck from writing code to verifying scientific correctness.
Why it matters: This is a concrete data point on where agentic coding tools currently add the most value, legacy modernization, versus where they remain risky, domain correctness, for a field where wrong-but-confident output can propagate into published results. It reinforces a pattern seen elsewhere this cycle, such as the Claude Opus 5 home-directory deletion, that agent competence and agent trustworthiness are separate axes that don't move together.
Related updates
- AI breast-cancer detection tools underperform radiologists' expectationsAug 12
- Researchers reconstruct LLM prompts from outputs, near-perfect accuracyAug 12
- AI legal-research tool lifted Pakistani judges' case resolution 6.3%Aug 12
- Unreleased Anthropic model advances work on the Riemann hypothesisAug 11