@xeophon: Pacing the frontier
Summary
Vals AI evaluated Grok 4.7, finding it ranks #24 on the Vals Index with a score of 54.2%, down from Grok 4.6, but shows improvements in legal and medical domains.
View Cached Full Text
Cached at: 09/22/26, 05:48 AM
Pacing the frontier
Vals AI (@ValsAI): We evaluated Grok 4.7 across the Vals benchmark suite. It ranks #24 on the Vals Index at 54.2%, down 5.0 points from Grok 4.6 (#14, 59.2%), but still ahead of Grok 4.5 (#30, 51.5%). Grok 4.7 improves the most on legal and medical work.
Similar Articles
Grok 4.6
xAI releases Grok 4.6, a frontier model focused on long-running agents and ambitious interactive/visual work, matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index and available in Cursor and Grok Build.
SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier alongside GPT-5.6 Sol and Claude models, with strong agentic performance at lower cost.
@elonmusk: Grok 4.6 is #1 on healthcare questions
Grok 4.6 has achieved the top position on MedAgentBench, outperforming GPT-5.6 Sol and Grok 4.5 in real-world healthcare AI tasks.
@gregpr07: Grok 4.6 is within 1 point of Opus 5 at ~40% lower cost. Across 7 runs, it solved 105/106 hard browser tasks. A little …
Grok 4.6 achieves near-Opus 5 performance at a 40% lower cost, solving 105/106 hard browser tasks, indicating potential for further advancement with reinforcement learning.
@rohanpaul_ai: Legal work saw one of Grok 4.7’s biggest jumps. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while GPT-5…
Grok 4.7 shows significant improvements in legal work and terminal tasks, outperforming competitors on benchmarks like Harvey Legal Agent Benchmark and Terminal-Bench 4.0, while keeping token prices unchanged.