Artificial Analysis updates its Intelligence Index to version 4.3

Reddit r/singularity News

Summary

Artificial Analysis has updated its Intelligence Index to version 4.3, incorporating new benchmarks like Terminal-Bench v4.0 and AutomationBench-AA to better evaluate AI model performance and cost-efficiency.

No content available
Original Article
View Cached Full Text

Cached at: 09/07/26, 09:08 PM

# Announcing the Artificial Analysis Intelligence Index v4.3 Source: [https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3) ## Results ### [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index) Artificial Analysis Intelligence Index v4\.3 incorporates 10 evaluations: AA\-Briefcase, GDPval\-AA v2, AutomationBench\-AA, Terminal\-Bench v4\.0, SciCode, Humanity's Last Exam, GDP\.pdf, CritPt, AA\-Omniscience, AA\-LCR v1\.1 Artificial Analysis Intelligence Index v4\.3includes:AA\-Briefcase, GDPval\-AA v2, AutomationBench\-AA, Terminal\-Bench v4\.0, SciCode, Humanity's Last Exam, GDP\.pdf, CritPt, AA\-Omniscience, AA\-LCR v1\.1\. See[Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking)for further details, including a breakdown of each evaluation and how we run them\. - **Claude Fable 5\.1 and GPT\-6 Astra lead the Intelligence Index:**Both Claude Fable 5\.1 \(max with fallback\) and GPT\-6 Astra \(max\) score 53 on Intelligence Index v4\.3, followed by Claude Opus 5 \(max, 51\), Claude Fable 5 \(with fallback, 50\), Muse Spark 1\.3 \(max, 48\) and GPT\-5\.6 Sol \(max, 47\)\. - **GLM\-5\.3 and Kimi K3 continue to lead open weights models:**GLM\-5\.3 Flash \(42\) is the third strongest open weights model, followed by Qwen3\.8 2\.4T A95B \(40\) and DeepSeek V4 Pro 0813 \(max, 36\)\. - **4 labs occupy the Intelligence vs\. Cost per Task Pareto frontier:**OpenAI occupies the majority of the cost\-efficiency frontier, with all five reasoning efforts of the recently released GPT\-6 Astra offering the lowest Cost per Task at their respective levels of intelligence\. MiMo\-V2\.5\-Pro \(26\), GLM\-5\.3\-Flash \(42\) and Claude Fable 5\.1 \(xhigh, max, 53\) round out the rest of the frontier\. ## Changelog Upgraded ### Terminal\-Bench: v2\.1 to v4 Completing our upgrade to the latest version of Terminal\-Bench\. Replaced ### 𝜏³\-Banking with AutomationBench\-AA Our implementation of Zapier’s business workflow automation benchmark\. We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5\. Each change in v4\.2 and v4\.3 stands on its own merits and brings the index closer to real\-world problem solving, adds more private test sets to prevent gaming, and reduces saturation\. Intelligence Index v4\.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested\. Because we use a held\-out test set for AutomationBench in collaboration with Zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%\. Category weights are unchanged from v4\.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%\. [Read the full methodology changelog](https://artificialanalysis.ai/methodology/intelligence-benchmarking#version-history) ## Evaluation updates ### Terminal\-Bench v4\.0 **We are upgrading from Terminal\-Bench v2\.1 to v4\.0, which features harder tasks from a terminal\.**Terminal\-Bench v4\.0 tests whether an agent can complete complex work through the terminal, across software, machine learning, science, operations, security, hardware, and media\. The update recalibrates compute and time allowances and improves task instructions, environments, and verification\. We run all 66 tasks three times and report average pass@1\. GPT\-6 Astra \(max\) scores 59\.1%, compared with 52\.0% for Claude Fable 5\.1 \(max with fallback\) and 49\.0% for Claude Opus 5 \(max\)\. Astra is 19\.2 percentage points ahead of GPT\-5\.6 Sol \(max\), which scores 39\.9%\. [Evaluation details](https://artificialanalysis.ai/evaluations/terminalbench-v4-0)### Terminal\-Bench v4\.0: Score Independently benchmarked by Artificial Analysis ### AutomationBench\-AA **We are replacing 𝜏³\-Banking with AutomationBench\-AA, featuring broader business workflows across applications\.** In collaboration with Zapier, we run the held\-out test set of 657 tasks, using the v1\.0\.6 version of the benchmark\. We call our implementation AutomationBench\-AA because we award partial credit for completed objectives, with any guardrail violation reducing the task’s score to zero\. Its 657 tasks span Finance, HR, Marketing, Operations, Sales, and Support\. Agents work across simulated business applications and discover the relevant APIs to complete each task\. For each task, we measure the share of objectives completed\. Any guardrail violation gives that task a score of zero\. ‘Score’ averages these task scores across all 657 workflows and is the metric used in the Intelligence Index\. ‘Tasks Completed’ separately reports the share of workflows where every objective is completed without a guardrail violation\. GPT\-6 Astra \(max\) scores 68\.5%, compared with 66\.7% for Grok 4\.6 \(high\) and 62\.2% for GLM\-5\.3 \(max\)\. Astra \(max\) completes every objective without a guardrail violation on 41\.6% of workflows, compared with 32\.1% for Claude Fable 5\.1 \(max with fallback\) and 28\.3% for Claude Opus 5 \(max\)\. Completing every objective while respecting all guardrails remains harder than completing part of a workflow\. [Evaluation details](https://artificialanalysis.ai/evaluations/automationbench-aa)### AutomationBench\-AA: Score Share of task objectives completed with no guardrail violations · Higher is better · Benchmark developed by Zapier · Independently benchmarked by Artificial Analysis AutomationBench\-AA reports two measures\. The Score is the share of task objectives a model completes, where any task with a guardrail violation scores zero\. Tasks Completed is the share of tasks completed in full, with every objective met and no guardrail violations\. Higher is better for both\. ## Cost **Similar Intelligence Index scores can come at very different costs\.**GPT\-6 Astra \(max\) and Claude Fable 5\.1 \(max with fallback\) both score 53, but their average cost per Intelligence Index task is $3\.26 and $7\.63 respectively \- 57% lower for Astra\. GLM\-5\.3\-Flash and GPT\-5\.6 Terra \(max\) both score 42 on the Intelligence Index, but GLM\-5\.3\-Flash costs just 18% as much per task \($0\.25 versus $1\.40\)\. At a lower price point, GPT\-5\.6 Luna \(max\) scores 38 at $0\.18 per task\. ### Intelligence Index vs\. Cost per Intelligence Index Task Artificial Analysis Intelligence Index · Weighted average cost \(USD\) per Artificial Analysis Intelligence Index task Weighted average cost per Intelligence Index task\. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight\. Artificial Analysis Intelligence Index v4\.3includes:AA\-Briefcase, GDPval\-AA v2, AutomationBench\-AA, Terminal\-Bench v4\.0, SciCode, Humanity's Last Exam, GDP\.pdf, CritPt, AA\-Omniscience, AA\-LCR v1\.1\. See[Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking)for further details, including a breakdown of each evaluation and how we run them\. ## Full results and methodology **The overall scores reflect different strengths\.**Claude Fable 5\.1 scores higher on AA\-Briefcase and SciCode, while GPT\-6 Astra scores higher on Terminal\-Bench v4\.0 and AutomationBench\-AA Score\. Category contributions to the Intelligence Index remain at Agents: 30%, Coding: 20%, General: 30%, and Scientific Reasoning: 20%\. Terminal\-Bench 4\.0 keeps the same weighting as Terminal\-Bench 2\.1, and AutomationBench\-AA replaces 𝜏³\-Banking at its 5% weighting\. Evaluations with private questions or answers account for 45% of the Intelligence Index v4\.3 weighting, up from 40% in v4\.2\. Intelligence Index v4\.3 \- evaluations, private test sets, and weightsCategoryEvaluationPrivate test setWeightAgents30%AA\-BriefcaseYes15%GDPval\-AA v2No10%AutomationBench\-AA\(new in index\)Yes5%Coding20%Terminal\-Bench v4\.0\(new in index\)No10%SciCodeNo10%General30%AA\-Omniscience \- AccuracyYes10%AA\-Omniscience \- Non\-hallucinationYes5%GDP\.pdfNo10%AA\-LCR v1\.1No5%Scientific Reasoning20%HLE \(Humanity’s Last Exam\)No10%CritPtYes10%Total100%

Similar Articles

My issue with Artificial Analysis's 'intelligence index'

Reddit r/LocalLLaMA

The article criticizes Artificial Analysis's intelligence index, claiming that a sudden v4.1.1 update reweighted metrics to downgrade the open-source Qwen 3.8 Max below Anthropic's Claude Opus, suggesting bias or sponsorship influence.

GLM5.3 Artificial Analysis Benchmarks

Reddit r/LocalLLaMA

This article presents a detailed benchmark analysis of the GLM-5.3 AI model, evaluating its intelligence and performance across multiple tests by Artificial Analysis.

New benchmark dropped

Reddit r/singularity

A new benchmark has been released, likely for evaluating AI or software performance.

Artificial Intelligence Index Report 2026

Hugging Face Daily Papers

The ninth edition of the AI Index report analyzes the gap between AI advancement and societal preparedness, featuring new assessments of reasoning, safety, economic value, labor effects, and dedicated chapters on AI in science and medicine.