GPT-5.6 Sol and Claude Opus 5 Just Cleared ARC Prize's Solved Threshold
ARC-AGI-2 was built specifically to resist the kind of memorization and pattern matching that made earlier benchmarks like MMLU easy to game over time. As of July 2026, according to figures from the aggregator BenchLM, GPT-5.6 Sol leads the benchmark at 92.5% and Claude Opus 5 follows at 90.4%. Those specific percentages come from a secondary tracker rather than a figure independently reverified against ARC Prize's own live leaderboard, so treat the exact decimal points as approximate rather than official.
What is not in dispute is the threshold both models have now crossed. ARC Prize organizers set 85% earlier in 2026 as the bar for calling the benchmark solved. Two frontier models clearing that bar within a few months of it being set is forcing the benchmark's own organizers to reconsider what a passing score should actually mean.
That reconsideration is the real story here, not the leaderboard order. ARC-AGI-2 was designed as a harder successor to a benchmark that got gamed, and if it is already being cleared this fast, the question is whether 85% ever measured genuine reasoning the way it was meant to, or whether it measured a difficulty level that increasingly sophisticated pattern matching against the test's own structure can now reach too. Benchmark thresholds set to be hard have a track record of aging faster than expected once frontier labs start optimizing against them directly, and ARC-AGI-2 looks like it is following that pattern rather than breaking it.