OpenAI states that GPT-5.6 Sol achieves 38.3 percent on the ARC-AGI-3 benchmark, surpassing Anthropic‘s Opus 5 score of 30.2 percent. This claim relies on the Responses API with Retained Reasoning and Compaction settings, features not available in the official test harness where GPT-5.6 Sol previously scored 7.8 percent. François Chollet, co-founder of the ARC Prize, acknowledged that OpenAI’s configuration offers an advantage by retaining context and reasoning steps that the standard environment discards. He argued that as long as these specific settings and associated costs are clearly reported, the disparity does not invalidate the comparison. The dispute centres on whether the official test environment unfairly penalises OpenAI due to a lack of the same API features that Anthropic utilises. Chollet suggested that different providers using different tools creates a potential parity issue but remains acceptable if transparency is maintained. This situation highlights the growing difficulty in comparing large language models when technical infrastructure varies significantly between competitors.
- Official ARC-AGI-3 scores use a standardised approach without provider-specific settings
- OpenAI’s Responses API allows context summarisation rather than truncation
- Retained Reasoning keeps the model’s chain of thought between steps




