Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

AI Summary
Anthropic's Claude Opus 5 achieved a score of 30.2 percent on the ARC-AGI-3 benchmark, significantly outperforming GPT-5.6 Sol's previous record of 7.8 percent. The performance indicates enhanced logical reasoning capabilities, as Opus 5 independently formulated reflection equations, a first for AI models.
From the source
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning. The article Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence appeared first on The Decoder.
The full text couldn't be loaded here (the source may require a subscription).
View original at The Decoder