Artificial Analysis v4.3: the new AI index test arrives

The artificial intelligence index that a large part of the market uses to decide which model to buy changed instruments. Artificial Analysis published version v4.3 of the Intelligence Index on 2026-09-07, with Terminal-Bench moving from v2.1 to v4.0, AutomationBench-AA taking the 5% weight tau3-Banking held, and evaluations with private tasks or answers rising from 40% to 45% of the total. When the ruler changes, the same model's score stops being the same number. A score without the index version beside it is a number without a ruler, and that is how it circulates in sales proposals, in procurement spreadsheets and in committee slides.
Why it matters
Three of the v4.3 changes are methodological, not cosmetic. Terminal-Bench was rewritten from v2.1 to v4.0, which means the set of terminal tasks changed in content; tau3-Banking left the index and AutomationBench-AA entered, inheriting the 5% weight; and the block of evaluations with private tasks or answers gained five percentage points.
None of that changes how the model behaves.
It changes what the number measures.
The practical effect is direct for the buyer. If your routing policy was calibrated on a v4.2 score, it was calibrated on another instrument. A model that climbed three points between the previous reading and the current one may have climbed on its own merit, on an easier task in the new benchmark harness, or on a different weight composition. Artificial Analysis publishes the version alongside the number for exactly this reason; whoever repeats the number without the version is the one breaking the series.
Comparability is the asset an evaluation base takes years to build and a rebase gives back in a day. It is not a defect of v4.3. It is the cost of fixing an instrument that fell out of date, and the bill arrives for whoever treated the score as a constant.
What exactly changed in index v4.3
The 2026-09-14 reading of the primary source records four changes to the instrument. Terminal-Bench moves to a major version, tau3-Banking leaves the composition and AutomationBench-AA takes the 5% weight that was its. And the aggregate weight of evaluations with private tasks or answers rises from 40% to 45%.
The fourth item on the list is the one that most disturbs the nature of the measurement, and it is worth opening in two parts. AutomationBench-AA is new and enters with a private task set.
| Benchmark | Previous version | v4.3 | Nature of the change |
|---|---|---|---|
| Terminal-Bench | v2.1 | v4.0 | Major version: the set of terminal tasks is rewritten |
| tau3-Banking | present, 5% weight | leaves the index | AutomationBench-AA takes the 5% weight that was its |
| AutomationBench-AA | did not exist | new, 5% weight | Agentic workflow automation benchmark, built with Zapier, 657 items in a held-out private task set, version 1.0.6 |
| Evaluations with private tasks or answers | 40% | 45% | Larger aggregate weight, a consequence of the private task set entering |
Each row deserves a sentence. Terminal-Bench measures execution in a command line environment, and the move to v4.0 swaps the content of the test, not the model taking the test. tau3-Banking still exists as an evaluation; what happened is that it stopped composing this index, and AutomationBench-AA inherited its slice. AutomationBench-AA charges workflow automation across six domains: Finance, HR, Marketing, Operations, Sales and Support. And the larger private weight is the arithmetic consequence of a private task set having entered in place of a public benchmark.
A held-out private task set, and why it changes the measurement
A held-out private task set is a test set that does not circulate publicly. The gain is measuring generalization instead of measuring familiarity. When the set is public, a lab can train against it, and the score rises from overfitting to the test, not from capability.
The detail that sustains that gain is operational: the set appears neither in training material nor in an optimization bench. With a private task set of 657 items, that route is closed. It is the same logic that already appeared when three agentic evaluations changed how models are chosen for agents, and it is worth recording what it does not solve.
AutomationBench-AA was built in collaboration with Zapier and the tasks are real world workflow automation. Zapier owns the set, and that is part of the instrument's design, not sponsorship of the result: whoever writes the tasks needed real workflows, and Zapier operates those workflows in production. Artificial Analysis maintains the evaluation and reports the numbers. These are separate roles, and the distinction matters for anyone reading the result with healthy skepticism.
The rigor of the design shows in a detail that usually goes unnoticed. The benchmark reports two distinct measures. The Score is the average fraction of objectives completed per task, and that is the metric that enters the Intelligence Index; any guardrail violation zeroes that task, which is a design choice and not an execution detail, because it dumps an entire piece of work that progressed down to zero. Tasks Completed is the fraction of workflows where all objectives were completed without any guardrail violation. The source states that completing everything while respecting the guardrails remains harder than completing part of the workflow. Reading the two measures as if they were one is the classic error of whoever quotes a benchmark from memory: the difference between them is exactly what the evaluation wants to measure.
There is an honest limit to this kind of set. A private task set reduces optimization against the test, but it does not eliminate the risk of the set aging. 657 tasks describe today's workflows well and will describe 2027's worse. No held-out benchmark is permanent, which is one more reason to cite the version alongside the number.
What the benchmark swap does to the historical series
A v4.2 score and a v4.3 score of the same model do not share the same ruler. The question the buyer needs to ask changed from "which is the best model" to "best relative to which instrument", and it is not rhetorical: it is the difference between a re-anchored decision and an inherited one.
Two disturbed benchmarks and a larger block of private weight are enough to invalidate the direct comparison between the two readings, even if the model name is identical in both rows of the table. That was exactly the problem the piece on the v4.2 rebase and the decision to choose again raised, and v4.3 did not solve it: it repeated it somewhere else in the instrument.
The point deserves to be said without hedging: a score is a third-party measurement at a dated version, never verified production truth. It was produced by an instrument with a date, under a weight composition that is published, on tasks you did not watch run against your workload. That does not disqualify the index. It is, today, the most consistent public reference for positioning models, and the alternative is deciding in the dark. What it is not, is a certificate about your case.
The governance consequence is boring and necessary. Every internal document that cites a score needs to carry the index version alongside it, the way it already carries the date of the dollar quote. A requirement that says "model above 50 on the Intelligence Index" ages badly and will be reinterpreted by someone in three months. A requirement that says "above 50 on v4.3, review scheduled for the next version" survives an audit.
What the new index shows, and what it does not decide
The v4.3 reading of 2026-09-14 puts Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) tied at the top, at 53. Next come Claude Opus 5 (max) at 51, Claude Fable 5 (with fallback) at 50, Muse Spark 1.3 (max) at 48 and GPT-5.6 Sol (max) at 47.
None of that is permanent. It is a dated reading, not a ranking, and the next version of the index can reorder that list without any of the models having changed. The top of the index, in fact, has already been read as structure rather than as a table: that is what produced the snapshot of the end of the model duopoly, which still holds as a market portrait and still says nothing about the instrument that produced the numbers.
On Terminal-Bench v4.0, GPT-6 Astra (max) scores 59.1% pass@1 over 66 tasks run three times each, against 52.0% for Claude Fable 5.1 and 49.0% for Claude Opus 5. The 7.1 point gap between first and second is large for a benchmark of the same kind, and it is worth remembering that it was measured on the new instrument: the distance is not comparable with any distance published on v2.1.
The top two on the index have the same score and costs per task that look nothing alike. GPT-6 Astra (max) appears at US$ 3.26 per task and Claude Fable 5.1 (max with fallback) at US$ 7.63, a difference of 57% with an identical index score. Further down, GLM-5.3-Flash and GPT-5.6 Terra (max) tie at 42 at US$ 0.25 and US$ 1.40 per task, respectively: GLM-5.3-Flash costs 18% of the pair. These are visible consequences of an instrument that now weights more evaluation with a private task, and the reading that matters is about the instrument: the index cost per task is the average of a task set that is not yours, so it serves to raise the routing hypothesis, never to close it.
What this changes for anyone routing models
The operational consequence is a procedure, and it is short. Whoever routes models should treat an index rebase the way they would treat a schema change: identify what broke, revalidate what matters and move on with an adjusted policy, without redoing the entire vendor analysis.
The work is not redoing the entire vendor analysis. It is re-anchoring the decisions that rested on the score, and the evaluation guide with a stable ruler remains the starting point for anyone redoing the math.
- Write the index version beside every score that circulates internally. Tender requirement, purchase justification, vendor comparison. A score without a version is an orphan number, and it will be cited out of context by someone with no way to know where it came from.
- Treat the v4.2 series as closed, not as a trend. Do not calculate the variation between the previous and current score of the same model as if it were a gain or loss of capability. If the difference matters to the decision, the path is rerunning on the new instrument, not subtracting the two readings.
- Re-test with your own traffic before re-deciding the route. A private task set with 657 items measures workflow automation in general. Your workflow is not in that set. The parallel test, with the same prompt across several models, settles in hours what the score debate does not settle in weeks.
- Write the policy per task, with fallback and a ceiling. A single score does not sustain a routing decision. What sustains it is a route rule per key for each task class, with fallback configured for when the primary provider fails and a spend ceiling per key, per project, so that the route change does not become a surprise at month close.
- Keep the audit trail of every call. When the next rebase arrives, and it will, the real execution history is worth more than any public score comparison. It is the only data that answers "how did the model behave on my workload, under my policy".
- Reschedule the policy review for the next index version. v4.3 is today's ruler. Review dates tied to a version number expire on their own and do not depend on anyone remembering.
This is where the Nexforce Router enters, and it enters for the boring reason. None of those six actions requires switching model providers. All of them require a layer sitting between the application and the providers, and it is the same layer the criteria for evaluating an LLM gateway describe before contracting. The Router is an LLM gateway with intelligent routing by cost, performance, latency and context, parallel testing of one prompt across several models, real-time model ranking by performance and price, automatic failover and configurable fallback, route rules per key and a spend ceiling per key or per project. It keeps central observability and an audit trail of every call, and it switches models without reintegration, which is what makes it possible to re-decide a route without opening an engineering project.
The honesty point: the Router does not decide what index v4.3 means for your case, and it cannot. It lets you measure. The next time Artificial Analysis rewrites the instrument, whoever has the parallel test in place re-anchors the policy in an afternoon. Whoever has the score pasted into the spreadsheet will rediscover the difference between citing a number and deciding based on one.
Frequently asked questions about the Artificial Analysis index v4.3
What changed in the artificial intelligence index in v4.3? Artificial Analysis published v4.3 on 2026-09-07 with four changes: Terminal-Bench moves from v2.1 to v4.0, tau3-Banking leaves the index, AutomationBench-AA enters and takes the 5% weight that was tau3-Banking's, and evaluations with private tasks or answers rise from 40% to 45%.
What is AutomationBench-AA? It is an agentic workflow automation benchmark, at version 1.0.6, built by Artificial Analysis in collaboration with Zapier on a held-out private task set of 657 items, across Finance, HR, Marketing, Operations, Sales and Support. Any guardrail violation zeroes the task, and it takes the 5% weight that was tau3-Banking's.
Why are scores from different index versions not comparable? Because the instrument changed. Two benchmarks were swapped or rewritten and the weight of evaluations with private tasks or answers rose from 40% to 45%. The same model measured on v4.2 and on v4.3 was not measured by the same ruler, so the difference between the two readings mixes a change of model with a change of test.
Is a private task set reliable? It is more resistant to optimization against the test than a public set, because the content does not circulate and does not enter a training bench. Reliable does not mean definitive: a held-out private task set of 657 items describes today's workflows well and ages like any set. The correct reading is that of a dated instrument, with the version cited.
What do I do with the index score in my routing policy? Use it as a starting point, never as a final answer. Re-test with your own traffic in a parallel test, write the rule per task, configure fallback and a spend ceiling, and keep the audit trail of every call. With each new version of the index, re-anchor the policy with your own data instead of recalculating the difference between scores.
References and Further Reading
- Artificial Analysis, "Announcing the Artificial Analysis Intelligence Index v4.3", published on 2026-09-07. artificialanalysis.ai
- Zapier, the company's benchmarks page, which confirms the held-out task set owned by Zapier and used by AutomationBench-AA. zapier.com/benchmarks
- AI model ranking reset: what changes in the choice, the sibling piece on the v4.2 rebase and the route decision. nexforce.ai
- Intelligence Index 2026: the end of the AI model duopoly, the structural snapshot of who sits at the top of the index. nexforce.ai
- Three agentic evaluations change how to choose models for agents, on what to do when the evaluation set changes. nexforce.ai
- LLM benchmark: how to evaluate and choose the right model, the evaluation guide with a stable ruler. nexforce.ai
- How to evaluate and choose an LLM gateway, the criteria that matter before contracting the routing layer. nexforce.ai
What to watch in the next version
v4.3 is today's ruler and will be replaced, probably sooner than most internal routing policies expect. The signal to follow is not any specific model's position: it is the methodology note Artificial Analysis publishes with each version. As long as it keeps opening the benchmark weights and the size of the private sets, the number stays auditable. On the day that note comes shorter, that is when the routing decision gets harder, because the ruler starts changing without public explanation. And then nobody knows what changed.
For anyone running models in production, the practical conclusion fits in one line: treat the index as a dated instrument, not as a verdict. Cite the version, re-test with your own traffic, and keep routing in a layer where switching models costs a configuration and not a project. The first three weeks after a rebase are the most expensive for whoever decided on the old score and did not know.
Models enter and leave the top of the index at a frequency that no longer surprises anyone. What changes slowly is the instrument, and that is why v4.3 deserves more attention than any model's position on the list. When the ruler changes, whoever had the number memorized needs to relearn how to measure.

Accelerate your company'sbusiness and operational efficiency
We design the technology of tomorrow to boost your business operational scale
Talk to a SpecialistRelated articles

Google Launches Gemini 3.8 Live and Extended Thinking: Parallel Voice Reasoning
Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking featuring parallel reasoning and asynchronous tool execution during continuous voice dialogue.
Read more
TypeSafe launches Jev: the model that generates no text
TypeSafe launched Jev, a model that abandons text generation and returns calibrated probability decisions. Why it feeds LLM routing instead of replacing it.
Read more
Anthropic urges slowing AI; Trump and Beijing refuse
On September 14, 2026, Trump and Beijing rejected the plan to slow down the AI frontier. With no coordination, model routing becomes a compliance choice.
Read more