Cost per Task: How an AI Routing Decision Gets Made

Artificial Analysis published version v4.3 of its Intelligence Index on 2026-09-07, and the top of the table turned awkward for anyone who has to decide. Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) both scored 53 points. A tie. One of them costs $3.26 per task. The other costs $7.63.
A tied score is the worst possible setup for a routing decision, because it is not a tie. It is a problem wearing the costume of a non-problem. Anyone reading only the score column concludes that either model will do and picks on some other ground: the model the team already knows, the one the vendor pushed harder, the one that showed up first in the internal evaluation. Any criterion works once the number everyone trusted stops discriminating.
Cost per task is not that number by accident. It is the only one that answers the question the score left open, and it answers in dollars: what it costs to solve one task from the evaluation set, with the model delivering the quality the job demands. This text walks the decision in five steps, from the quality floor to the route audit, and it shows where the published data ends and where the measurement only your own workload can produce begins.
Why the intelligence score no longer decides the route
An intelligence score measures capability inside a standardized set of tasks. When two models measure the same number, that number stops discriminating, and the criterion left standing is the cost of solving the task. The score goes from being the goal of the decision to being its filter.
The trap sits in the gap between measured capability and required capability. The index does not measure your problem, it measures a compendium of problems Artificial Analysis chose, and its value is precisely that it is standardized enough to compare different labs. Standardization has a price: the task where your product makes or loses money may not be represented there with the weight it carries in your real volume.
GPT-6 Astra is the clearest case of the moment. The model launch was already covered here with pricing, benchmarks and the Critical safety rating, in GPT-6 Astra: pricing, benchmarks and OpenAI safety, on 2026-09-09. What was missing was the arithmetic. A model that ties at the top of the index and costs 57% less per task than its tie partner is not a matter of taste, it is a margin decision waiting for someone to sign off on it.
What the v4.3 index measures
The Intelligence Index v4.3 is a composite that now includes version v4.0 of Terminal-Bench. It replaced the banking evaluation with AutomationBench-AA, built with Zapier over 657 private, held-out tasks. As a result, the weight of evaluations using private tasks or responses rose from 40% to 45%, and a single guardrail violation zeroes the task.
Private testing matters for one simple reason: a held-out set does not leak into the next model's training, so it measures capability instead of memory. Every number in this piece carries the version and the date for that reason. The index has a history of re-basing. The reset already forced teams to re-decide routes in AI model ranking reset, and a cost per task without a version next to it is a number with no declared expiry date.
What cost per task is and why it separates what the score ties
Cost per task is what it costs to solve one task from the index set, in dollars, calculated from the price of input tokens, cache reads, cache writes, reasoning and output, divided by the number of tasks and weighted by each evaluation's share of the index. In practice it is the unit that turns a scoreboard into a budget, because it converts measured capability into cost per unit of delivered work.
The v4.3 index publishes that number alongside the score, and the 2026-09-07 reading delivers two pairs that tell the whole story. The first is an expensive tie. The second is a cheap one.
| Pair | Score (max) | Cost per task | Difference |
|---|---|---|---|
| GPT-6 Astra (max) | 53 | $3.26 | reference for the expensive pair |
| Claude Fable 5.1 (max with fallback) | 53 | $7.63 | 57% more expensive at the same score |
| GLM-5.3-Flash | 42 | $0.25 | reference for the budget pair |
| GPT-5.6 Terra (max) | 42 | $1.40 | 18% of the cost; 5.6x is a derived ratio, not measured by the source |
The budget pair teaches more than the expensive one, because 42 points is still a respectable score and $0.25 per task is another order of magnitude. A team that only reads the top of the table never finds that line. It is the same kind of asymmetry that already carried the economic argument for routing in AI model costs in 2026, now with a decision unit in place of a spread per token.
The tie on the index is not the only quality signal available, and the second axis shows that the expensive pair is not a technical tie in disguise. On Terminal-Bench v4.0, measured across 66 tasks at 3 runs each, GPT-6 Astra (max) scores 59.1% pass@1. Claude Fable 5.1 sits at 52.0% and Claude Opus 5 at 49.0%. The top two have strengths on different axes, and Artificial Analysis itself records that Fable 5.1 leads on AA-Briefcase and SciCode while Astra leads on Terminal-Bench v4.0 and AutomationBench-AA. The composite score treats as equal two models that win in different places, which is why it does not decide.
Step 1: set the acceptable quality floor
Before comparing price, the task defines which score is enough. The acceptable quality floor is the minimum capability the task family tolerates without generating rework, and it comes from the cost of the error. Not from the ranking. Written once, it turns the score into a filter and frees cost per task to decide among the ones that passed.
That floor is a business decision disguised as a technical parameter. A pipeline that produces a draft for human review tolerates a 42 point model. A route that writes straight into the client's system does not. The question is not which model is better. It is how much error the step absorbs before the cheap option turns expensive, and that answer lives with whoever owns the process, not with whoever wrote the prompt.
Starting from the ranking inverts the order and locks the decision. The floor has to be written before anyone looks at the table. Once the scoreboard enters the room it becomes the goal, and a team with the wrong goal picks the most expensive model with conviction.
Step 2: read the index's cost per task with its date
A published cost per task is read with three things in hand: the index version, the publication date and the definition of the column. Without them the number becomes a supermarket shelf price, comparable to anything and accountable for nothing. The 2026-09-07 reading is the portrait valid until the next version.
The most common error is treating published cost per task as an invoice forecast. It is not. It is a measurement taken on a fixed set of tasks, under an execution policy that is not yours, and it serves to compare models against each other, not to estimate your month. Anyone who needs the bridge between a third party's measurement and their own budget has already run into that difference, and the documented path is in How to measure LLM provider performance.
The second error is reading the column in isolation, without the score beside it. A cost per task of $0.25 looks unbeatable until someone remembers it solves a 42 point task. The third error is ignoring the version: a v4.2 number compared against a v4.3 number is a comparison between different rulers, and it is worth exactly what comparing two measurements of different things is worth.
Step 3: measure your cost per task with your own traffic
The published number defines the shortlist. The number that decides the route is yours, measured in a parallel test, with the same prompt sent to two or three models that passed the floor and the answers scored against the criterion for that step. This is where most of the work of the decision lives.
A parallel test is not a game of comparing answers in a chat interface. It is instrumentation: the same input on both sides, the degradation rate counted, the accumulated cost added up. What almost nobody accounts for is what happens after the first attempt. Retries and fallback do not appear in the published cost per task and they are the slice that makes a cheap model expensive when it fails midway through a long task. Routing by request complexity already solves part of this, and the criterion is worth reading in Routing by complexity, but complexity is not cost: the first estimates effort, the second measures the result.
The honest count of your cost per task includes four lines that the published measurement does not charge for. Input tokens in long context, which grow with the conversation history. Retries and the second call nobody budgets. Fallback, when the primary model going down pushes traffic to a more expensive one. And the task abandoned midway, which consumed tokens and delivered no result, the line that tends to be the largest and the least visible. And the retry that enters the count needs a limit: without exponential backoff with a ceiling and a circuit breaker, aggressive fallback becomes the problem itself, because every client re-executes at the same time and reissues the most expensive call instead of avoiding it.
Step 4: write the routing policy with fallback and a spend cap
Five steps, and this is the one that demands a written decision. The routing policy names the default model, the fallback model, the condition that authorizes escalation and the spend cap that ends the argument. Without all four on paper, the route becomes an oral convention that changes the day whoever upheld it leaves.
The consolidated procedure, in order:
- Set the acceptable quality floor per task family, based on the cost of the error each step tolerates.
- Read cost per task from index v4.3 with version and date, crossed with the score. The output is a list of two to three models.
- Measure the real cost per task of each finalist in a parallel test, with your own traffic, counting retries, fallback, long context and abandoned tasks.
- Write the routing policy with a default model, a fallback model, an escalation condition and a spend cap per key or per project.
- Audit. Re-measure whenever a new model enters the shortlist, and before that on the cadence defined in the previous step.
Without a gateway layer in between, none of those four decisions survives the first day on which switching models requires re-integrating the application. Nexforce Router is that layer: model selection by cost, performance and context, parallel testing with one prompt across several models, real-time model ranking, configurable fallback with automatic failover, routing rules per key, a spend cap per key or per project, an audit trail for every call and model switching without re-integration.
The spend cap deserves the paragraph it almost never gets, because it is the only item in the policy that works without trust. The team can disagree about the floor, the default and the escalation condition. The cap does not ask for agreement. It halts the route when consumption crosses the defined limit, and that halt is the difference between a month of predictable cost and an email from finance asking what happened in October.
Step 5: audit the result against the decision
Promoting a model to default is a hypothesis, and a hypothesis without verification becomes corporate folklore within two weeks. Auditing means comparing the cost per task measured after the change against what the decision predicted, and checking whether quality held. If cost fell and degradation rose, the route got cheaper and worse. Nobody noticed.
The audit has one central question and two supporting ones. The central one: did cost per task fall by the proportion the Step 3 measurement predicted? The supporting ones: did the retry rate change, and where did fallback fire in the last period. A fallback that fires every week stopped being contingency and became part of the route, which changes the arithmetic and deserves a re-decision. The cost of operating that layer, which is small but not zero, is laid out in Operating cost of an AI gateway in production.
Cadence matters more than the sophistication of the report. One re-measurement per quarter, plus a forced reading whenever Artificial Analysis publishes a new version of the index, covers the interval in which the market moves. New models enter the shortlist every week. A route revalidated every six months is a bet, not a policy.
When the scoreboard still decides
There are four cases where the index score takes charge again, and one of them is a trap. The first is the task at the edge of measured capability: with no slack in the floor, two points separate working from not working. There, cost per task is the dependent variable. The second is the audit. There the choice needs published criteria.
The third is the very low volume task. If the route solves two hundred tasks a month, cost per task moves cents. The decision should go to quality or to reducing operational risk. Automating a low impact decision is spending governance on someone who does not need it. The fourth case is the limitation of the unit itself: the cost per task the index publishes is an average over a set of tasks that is not yours.
Published cost per task is an institution's measurement, on a fixed set of tasks, under an execution policy that is not yours. It does not include long context, retries, fallback, abandoned tasks or the human review that comes after the answer. Treating that number as the December invoice is the mirror image of the mistake made by the team that decides on the scoreboard alone.
FAQ
The five questions below close the points that decide a route in practice: what the unit means, what to do when two models tie on the index, why the published number is not your invoice, how to measure your own workload, and where the spend cap enters the policy.
What is cost per task in AI models? It is the dollar cost of solving one task from an evaluation set, adding up the tokens the model consumes to deliver the answer at the required quality. In the Intelligence Index v4.3, published on 2026-09-07, cost per task and intelligence score appear side by side, model by model.
If two models tie on the index, which one do you pick? Cost per task decides, after setting the quality floor the task tolerates. GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) tied at 53 points in the 2026-09-07 reading, at $3.26 against $7.63 per task. An identical score separates nothing; the cost separates.
Is the published cost per task what I will pay? No. It is an institution's measurement on a fixed set of tasks, under an execution policy that is not yours. It compares models against each other, and it does not estimate your invoice. The number that decides your route comes out of a parallel test on your own traffic, counting retries, fallback and abandoned tasks.
How do I measure the cost per task of my own workload? By running the same input through two or three models that passed the floor, in a parallel test, and adding the cost of every attempt up to the acceptable answer, not just the first. Your real task composition is the divisor. A standardized set built by another company never represents your volume.
Where does the spend cap enter the routing policy? As the only item that does not depend on consensus. The policy names a default model, a fallback, an escalation condition and a cap per key or per project. The cap halts consumption when it crosses the limit, and it is what turns the routing decision into predictable cost instead of a declared intention.
References and Further Reading
- Artificial Analysis, Intelligence Index v4.3, published 2026-09-07: artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
- Artificial Analysis, GPT-6 Astra benchmarking article: artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- Nexforce, routing pillar: Model Router as middleware
- Nexforce, GPT-6 Astra launch: GPT-6 Astra: pricing, benchmarks and OpenAI safety
- Nexforce, index re-base: AI model ranking reset
- Nexforce, the economic argument for routing: AI model costs in 2026
- Nexforce, cost comparison by model: LLM cost comparison in 2026
- Nexforce Router, the routing layer: nexforce.ai/router
The decision that stays
Next Monday, the decision moves. The scoreboard leaves the decider's chair and takes the filter's. The team writes the quality floor, reads cost per task from the index with version and date, measures the real number against the finalists and closes the policy with default, fallback and cap. Then it audits.
The 53 point tie between GPT-6 Astra and Claude Fable 5.1 will be remembered as the day the index stopped answering. The 57% difference in cost per task between two models the scoreboard treats as identical is the answer the arithmetic was waiting for. It only appears to whoever decided to look at the cost before the ranking.

Save up to 50% in creditswith a single smart API
Connect your operations to our AI Router and optimize the consumption of multiple LLMs
Free TrialRelated articles

Caller identity in agent and tool traffic: who the gateway sees
When two teams share one agent, the gateway recognizes the credential and not the caller. Caller identity separates quota, access, and trace per business unit.
Read more
How to decide your LLM route with real traffic evidence
A five-step method to decide your LLM route with real traffic evidence: shadow on live traffic, blind judging, statistical criteria, and promotion only after measurement.
Read more
LLM Routing: What to Do When the Token Price Changes
Two models repriced in opposite directions inside the same three-week window: a 33% output cut at the top model and a 371% output increase at the cheapest. What that does to the cost of a fixed route, and how routing absorbs the move.
Read more