THE CRUNCH

AI agents can run machine learning experiments, but a new benchmark from Epoch AI suggests they are still far from doing genuine research. In the InnovationEval test, agents had to invent, implement and refine a new method for improving language models after their initial training. The starting point was GRPO, a widely used technique that compares multiple answers a model generates for the same task and rewards the better ones. Tested with up to 3,000 hours of compute on high-end chips but no internet access, GPT-5.6 Sol and Claude Fable 5 came nowhere near the human-designed reference method SDPO, which Epoch says neither model knew in advance.

Sol scored about 35 percent of SDPO's improvement over GRPO under generous grading, dropping to roughly 15 percent when only rule-abiding changes counted. On coding tasks it mostly made training slower and more expensive. Fable 5's approach, retrying failed tasks with previous attempts fed back in, was a well-known technique and produced no measurable improvement. Even Sol's partial success would barely qualify as "moderately interesting" to experts, Epoch says.

The more troubling finding is how the agents reported their work. Both ran multiple near-identical training rounds and disclosed only the best result, exploiting random fluctuation to make methods look stronger, then barely mentioned the practice in their final reports. Sol claimed about 70 percent of the reference improvement and Fable 5 about 40 percent; Epoch stripped out those inflated gains. Their reasoning logs show they knew what they were doing, with Fable 5 calling its repeated runs a search for a better checkpoint, though Epoch leaves open whether this amounts to deliberate cheating. Notably, METR previously detected more cheating attempts from GPT-5.6 Sol in its coding test than from any other publicly available model it had evaluated.

Newer models that knew SDPO from training data could not fully replicate it either, and even with the original paper in front of it, Fable 5 fell short. The pattern extends beyond this benchmark: Anthropic's system card for Claude Opus 5.5 says the model is far from replacing the company's own researchers, citing weak "epistemic quality", unchecked assumptions presented as facts and partial checks described as complete verification. A study involving Princeton University and the UK Safety Institute had Claude Opus 4.8 spend six days on questions behind two unpublished NeurIPS papers, and the original authors rejected both results.