Grok 4.6 Matches GPT-5.6 Sol on Paper, Trails It on Actual Agentic Work
xAI released Grok 4.6 on August 12, 2026, holding pricing flat at $2 per million input tokens and $6 per million output tokens, rising to $4 and $12 past a 200K token context window. On the Artificial Analysis Intelligence Index, Grok 4.6 matched GPT-5.6 Sol.
That parity doesn't hold on Terminal-Bench v3.0, a benchmark built specifically to test agentic, tool-using, terminal-based task completion rather than general knowledge or reasoning. According to a developer community writeup rather than xAI's own benchmarks page, Grok 4.6 trailed GPT-5.6 Sol by roughly 8.6 points there. That figure is worth treating as directional rather than official until xAI or a more established benchmark tracker confirms it independently.
The gap matters more than the headline number does. An intelligence index score is a reasonable proxy for how a model performs on knowledge and reasoning questions, but it says very little about whether a model can reliably chain tool calls, execute terminal commands, and complete a multi-step agentic task without losing the thread. Those are increasingly the tasks companies are actually buying these models to do.
If you're choosing a model specifically for agent-building or coding-automation work, matching on a general intelligence index isn't a substitute for checking a benchmark that actually measures tool use. Grok 4.6's pricing is competitive and its general capability looks solid, but the Terminal-Bench gap is the number to check before betting an agent pipeline on it.