September 6, 2026 · MarkTechPost
Meta FAIR's AI judges cut wasted research compute
Meta FAIR, Oxford, and UCL built AI Research Preference Models (RPMs) - frozen large language model (LLM) judges that rank 15 candidate experiments and select just one to run, instead of running all of them. On the AIRS-Bench benchmark this raised the average normalized score from 0.684 to 0.729, and reached the baseline's 24-hour result in about 15 hours.
Why it matters: As AI research agents increasingly propose more experiments than available GPU budgets allow, tools that triage which experiments are worth running could meaningfully speed up the AI research loop itself by cutting wasted compute. This fits a broader trend of building AI systems that accelerate AI research rather than just apply it.