OpenAI states that GPT-5.6 Sol achieves a 38.3 percent score on the ARC-AGI-3 benchmark, surpassing Anthropic’s Opus 5 which recorded 30.2 percent. This comparison relies entirely on OpenAI’s custom test harness rather than the standard evaluation environment. The company runs the model through its Responses API using specific settings like Retained Reasoning and Compaction to maintain context. Without these adjustments, GPT-5.6 Sol scores just 7.8 percent in the official harness where reasoning chains are discarded after each action.
The dispute highlights how technical infrastructure influences reported performance metrics. ARC-AGI-3 was designed to isolate pure model capability by preventing external aids from skewing results. Opus 5 met the benchmark constraints while running inside a similar proprietary environment for Claude Code, suggesting its potential is higher under standardised conditions. OpenAI’s approach demonstrates that system configuration can artificially inflate scores.
* GPT-5.6 Sol scores 7.8 percent in the official harness
* Opus 5 scored 30.2 percent under identical constraints
* ARC-AGI-3 tests pure model performance without external aids




