THE CRUNCH
OpenAI's GPT-6 Astra has demonstrated a significant improvement in spatial reasoning, according to early benchmarks. In a test involving a dual-arm YAM robot, Astra completed seven out of 100 tasks on the StationeryBench, while the Ai2 model MolmoAct2 managed zero. Astra's median progress score was 46, compared to MolmoAct2's 12. The results, which include videos and code, are available on GitHub. AI researcher Yoav
Artzi described Astra as a 'step change in spatial reasoning.' On the unpublished REMAP benchmark, Astra reached accuracy close to human level, though Artzi noted that 'even Astra doesn't get to what humans do in other scenarios.' He suspects OpenAI trained the model on large amounts of 3D data, such as Blender scenes, which aligns with Astra's particular improvement on 3D tasks. OpenAI has long-term plans to build its own consumer robots.


