OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

AI Summary
OpenAI's GPT-5.6 Sol achieved a score of 38.3 percent on the ARC-AGI-3 benchmark using its own API features, surpassing Opus 5. However, the model scored only 7.8 percent in the official test setup, raising concerns about the fairness of the comparison due to possible outdated API usage by ARC Prize.
From the source
OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instead of the official test setup, where the model landed at 7.8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the comparison with Opus 5. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings appeared first on The Decoder.
The full text couldn't be loaded here (the source may require a subscription).
View original at The Decoder