Prime Intellect open-sources an agent harness that beats a human baseline
Prime Intellect released Prime Agent, an open-source coding and research harness that treats sub-agents as function calls inside a persistent IPython kernel and lets the agent edit its own prompts, skills, and memory mid-run. Running on Opus 5, it reportedly scored 95.5% Best@1 on ARC-AGI-3, above the cited human expert baseline of 95.4%.
Why it matters: Open-sourcing a harness that edits its own scaffolding mid-task, rather than relying purely on a bigger underlying model, suggests some of the next capability gains will come from agent architecture. Beating a human baseline on ARC-AGI-3, a benchmark designed to resist memorization, is also a concrete data point in the ongoing debate over how close current systems are to general reasoning ability.