OpenAI‘s GPT-6 Astra model outperformed Ai2’s MolmoAct2 on StationeryBench, a robotics evaluation involving five desk-based tasks. Both systems controlled identical dual-arm YAM robots across 200 trials. Astra successfully finished seven out of 100 attempts, achieving a median progress score of 46. MolmoAct2 completed zero tasks and scored 12. The full results, videos, and code are available on GitHub. OpenAI intends to use this technology for future consumer robots.
Researchers attribute the improvement to training on large volumes of 3D data, likely including Blender scenes. Yoav Artzi from Cornell and Google DeepMind describes the result as a step change in spatial reasoning. However, Artzi notes that the model still falls short of human capabilities in many scenarios. The REMAP benchmark, which remains unpublished, shows Astra reaching accuracy close to human levels.
* Astra finished 7 of 100 tasks while MolmoAct2 finished none
* Median progress scores were 46 for Astra and 12 for MolmoAct2
* All evaluation materials are hosted on GitHub




