AI model ranking: what the rebase changes for model choice

The AI model ranking reset in one stroke: a methodological rebase of the leading artificial intelligence index, published on 2026-09-04, rescaled every score. The top dropped from 63 to 57, eight of the ten leaders left the group, and no number from the previous version compares with a number from the new one. The model choice starts over.
The 63 that topped the table was Claude Opus 5 max, in the reading of 2026-08-17. The 57 that opens it now is Claude Fable 5.1 max with fallback, in the reading of 2026-09-07, measured by Artificial Analysis. The whole leaderboard moved. Between the two numbers sits a rebase: it swaps the ruler without asking permission, and every leaderboard saved in a deck, a spreadsheet or a policy loses its basis for comparison.
For the buyer, the reset is not a laboratory curiosity. It is an order to redecide. The ruler that guided the choice became another one, and the question returns to "which route serves each task now". That is exactly the work the model router pillar has tracked since its first post.
Key findings
Six findings summarize the Artificial Analysis v4.2 rebase, published on 2026-09-04 and measured in the reading of 2026-09-07: the top of the index fell from 63 to 57, eight of the ten former leaders left the top 10, eight labs now sit above 50, the evaluation set grew to 10, and prices moved in both directions over the same period.
Each line below stands alone as a citation, with source and date.
- The top of the artificial intelligence index fell from 63 to 57: Claude Opus 5 max scored 63 in the reading of 2026-08-17 and Claude Fable 5.1 max with fallback scores 57 in the reading of 2026-09-07 (Artificial Analysis measurement).
- 8 of the 10 former top-placed models left the top 10 in the v4.2 rebase of 2026-09-04.
- Claude Opus 5 max dropped from #1 to #6, from 63 to 54; GPT-5.6 Sol max left the top 10, from 61 to 51; Qwen3.8 Max went from 58 to 47.
None of these lines claims a model got worse. It claims the ruler changed underneath everyone, and the shift between the two readings measures how much work the rebase delivered onto the decision maker's table.
- In the reading of 2026-09-07, 8 labs sit above 50: Anthropic, OpenAI, Moonshot AI, Meta, Google, DeepSeek, SpaceXAI and Alibaba.
- The evaluation set grew to 10: GDP.pdf (Surge AI), CritPt, SciCode and AA-Omniscience joined, and the old Long Context Reasoning became AA-LCR v1.1.
- Prices moved in both directions over the same period: GPT-5.6 Sol max output became 33% cheaper and DeepSeek V4 Flash 0731 max output became 371% more expensive, from the reading of 2026-08-17 to the reading of 2026-09-07.
The ruler is the finding.
The set grew to 10 in v4.2, and the rest is consequence: another set, another score, another position on the board. Comparing leaderboards from different versions is the most expensive mistake of the week.
Who entered and who left the AI model ranking?
The v4.2 rebase reorganized the top 10 from top to bottom: 8 of the 10 former leaders left, and Claude Fable 5.1 entered at #1, #2 and #4, GPT-6 Astra at #3, #5 and #8, and Muse Spark 1.3, from Meta, at #10. Grok 4.6 fell from #6 to #17 and Kimi K3 from #7 to #18.
Whoever left did so through rescaling. No production measurement said the work of those models got worse; the entire earlier ruler was rescaled into the v4.2. The table organizes the movement.
| Model | Before the rebase | After, reading of 2026-09-07 |
|---|---|---|
| Claude Opus 5 max | 63, #1 (reading of 2026-08-17) | 54, #6 |
| GPT-5.6 Sol max | 61, top 10 | 51, out of the top 10 |
| Qwen3.8 Max | 58, top 10 | 47 |
| Grok 4.6 | #6 | #17 |
| Kimi K3 | #7 | #18 |
| Claude Fable 5.1 | out of the top 10 | 57, #1 (max with fallback); variants at #2 and #4 |
| GPT-6 Astra | out of the top 10 | #3, #5 and #8 |
| Muse Spark 1.3, from Meta | out of the top 10 | #10 |
Artificial Analysis measurement: the "after" column is v4.2, rebase of 2026-09-04, in the reading of 2026-09-07; the "before" column is the earlier ruler, its top dated in the reading of 2026-08-17.
That leaves a geography: 8 labs above 50 in the reading of 2026-09-07 is a compressed frontier, with new names at the top. A compressed leaderboard does not decide a route by itself, and the reading of when a small model approaches the frontier leader already covered that squeeze.
How does the index rebase work?
A rebase changes the methodology of the intelligence index and rescales every published score to the new ruler. In v4.2, rebase of 2026-09-04, the set grew to 10 evaluations and every published number was rescaled at once. That is why a score from the earlier version and a v4.2 score do not compare: they do not measure the same thing.
Scores are relative to the ruler that produced them. When the set and the methodology change, the numbers lose their common denominator, and Artificial Analysis rescales everything into the new version. Within one version, comparison holds. Across versions, it does not.
Every index figure cited in this piece is a third party measurement, published by Artificial Analysis in the v4.2 rebase article in dated form, and no score here is verified production performance. The v4.2 is different. The verification that replaces trust in the leaderboard is the own-traffic retest, in the procedure ahead.
For a policy that consumes a leaderboard, a score without a version is a loose number. Requiring the version next to every score that enters a rule is a job for scale routing middleware, not for an analyst's memory.
What happens to saved comparisons?
Every saved comparison between scores from different versions lost its basis in the rebase of 2026-09-04: a selection report, a board deck and a routing policy written on an old score now mix two rulers in the same document. The number did not break; the cross-version comparison is what stopped existing.
The failure mode has a name: double ruler. The report compares the 61 of an old reading with the 51 of the reading of 2026-09-07 and concludes the model aged. No measurement said that. A number read on the wrong ruler is not a bad number. It is a number that does not exist.
| Saved document | What it said | What it says now |
|---|---|---|
| Selection report | "Model A: 61; model B: 58, choose A" | Different rulers |
| Board deck | "The index top is 63" | The top is 57 in another index version, so cite the version with the number |
| Routing policy by score | "Route X requires score above 55" | A threshold without a version does not execute |
| LLM benchmark spreadsheet | Score columns from different versions stacked in the same file | Separate by index version before any average, sum or internal ranking |
What survives after the rebase of 2026-09-04 is the comparison internal to the version, because inside the reading of 2026-09-07 the numbers talk to each other and outside it the comparison disappears. And the damage is smaller for whoever already decided by cost before score, the ground of the LLM benchmark for CFOs. The ruler that does not wobble is cost per task.
Rewriting a routing threshold on the current ruler, with a retest and a log, is model policy governance: whoever already has it in the gateway spends an afternoon, whoever improvises in a spreadsheet spends a quarter.
Did prices follow the leaderboard?
There was no single direction: between the reading of 2026-08-17 and the reading of 2026-09-07, GPT-5.6 Sol max output became 33% cheaper and DeepSeek V4 Flash 0731 max output became 371% more expensive. Price and leaderboard are different rulers, and the routing decision consumes both, each with its own dated reading.
The two moves live in the same window.
| Model (max) | Input, reading of 2026-08-17 | Output, reading of 2026-08-17 | Input, reading of 2026-09-07 | Output, reading of 2026-09-07 |
|---|---|---|---|---|
| GPT-5.6 Sol | 5.00 USD | 30.00 USD | 4.00 USD | 20.00 USD |
| DeepSeek V4 Flash 0731 | 0.14 USD | 0.28 USD | 0.44 USD | 1.32 USD |
USD per million tokens: each column anchored to its own reading, 2026-08-17 and 2026-09-07. Source: Artificial Analysis.
Price is also a leaderboard.
There is no reading that causally connects price and score, and this piece does not invent one: they are a ruler and a label, each with its own date. What the buyer does with that is an old house rule from the token price collapse: routing cost is recalculated per task, never inherited from last month.
What enters the leaderboard now?
The index evaluation set grew to 10. GDP.pdf (Surge AI), CritPt, SciCode and AA-Omniscience enter; the old Long Context Reasoning becomes AA-LCR v1.1; and the MLCR-AA, built with Wisedocs, debuted outside the set. The leaderboard changed because the definition of artificial intelligence measured by the index changed.
| Evaluation | Status in the v4.2 set |
|---|---|
| GDP.pdf (Surge AI) | new to the set |
| CritPt | new to the set |
| SciCode | new to the set |
| AA-Omniscience | new to the set |
| AA-LCR v1.1 | renamed (was Long Context Reasoning) |
| MLCR-AA, with Wisedocs | debuted outside the set |
The team's mental average of "which is the smartest model" was computed on the old set. The v4.2 list, with the full method, is in the rebase article from Artificial Analysis. The set lasts; the readings pass. Whoever documents the set version next to the number sleeps better.
How do you redecide the route when the whole leaderboard resets?
The redecision has four steps: retest with your own traffic, redecide the route per task, rewrite the policy with fallback and a spend cap, and audit the result. The leaderboard guides, it does not decide: the verification is the own-traffic retest, one prompt across several models.
The old question of how to choose an AI model changes shape when the ruler resets. You choose a route, not a champion. The procedure fits into one meeting and holds for the next ruler too. The detail of evaluating endpoints and routing policy before closing each route already has its own post on the blog.
Step 1: own-traffic retest. The leaderboard is a third party measurement in a dated version; your own load is the measurement that decides. The parallel test of one prompt on several models on the Nexforce Router gives the reading the leaderboard cannot.
Step 2: redecision per task. Each route answers what the task asks in cost per result, task performance and context, because the overall score does not pick a route.
Step 3: policy with fallback and a spend cap. Routing rules by key, automatic failover when the path fails, and a spend cap per key or project, so the redecision does not become an open invoice.
Step 4: audit of the result. A trace for every call and central observability: the decision stays logged, and the next ruler starts with history, not memory.
Latency enters as a routing criterion, not as praise: when the request has a deadline, it is the gateway that picks the path that meets the deadline.
That is exactly the job of the Nexforce Router, the in-house gateway and routing layer: smart routing by cost, performance and context, real time model ranking by performance and price, automatic failover and configurable fallback, routing rules by key, spend cap by key or project, an audit trace for every call, central observability, and model swaps without reintegration.
Frequently asked questions
Four questions cover the essentials of the rebase for whoever decides a route: what compares across versions, what the reset says about quality, how price enters the redecision, and what to do on the next ruler. The answers hold for any index version, including v4.2.
Can I compare today's score with the score I saved in August?
No. The 63 at the top, in the reading of 2026-08-17, and the 57, in the reading of 2026-09-07, came from different rulers, and that is why scores from different versions do not compare in a report, a spreadsheet or a policy. A valid comparison is one that happens inside the same index version, with the version logged next to the number.
Does the rebase say the models got worse?
No. The rebase of 2026-09-04 changed the methodology and rescaled every score; no reading claims a worsening or an improvement. The v4.2 is different. The verification that replaces trust in the leaderboard is the own-traffic retest, in which the in-house Nexforce Router runs one prompt across several models.
Does price follow score?
No, and the rebase window proved it in both directions: GPT-5.6 Sol max output fell from 30.00 (reading of 2026-08-17) to 20.00 USD (reading of 2026-09-07), 33% cheaper, and DeepSeek V4 Flash 0731 max output rose from 0.28 (reading of 2026-08-17) to 1.32 USD (reading of 2026-09-07), 371% more expensive.
How often does the leaderboard reset again?
Index readings come out on their own dates, 2026-08-17 and 2026-09-07, and a rebase, which swaps the ruler, happens when the methodology changes, on 2026-09-04, so the defense is not predicting the date. It is logging version and reading next to every score and keeping the routing policy ready to redecide without reintegration.
<!-- The questions above feed the page's FAQPage schema. -->References and Further Reading
The numbers in this piece come from one dated primary source, the Artificial Analysis v4.2 rebase article published on 2026-09-04 with a reading of 2026-09-07, and the internal posts below complement the procedure.
One primary source, four complements.
Primary source: Artificial Analysis, Intelligence Index v4.2, the rebase article of 2026-09-04, reading of 2026-09-07.
The stable ruler, for evaluating from scratch, is in the LLM benchmark for CFOs.
The price dynamic has its own post in the token price collapse.
The execution mechanism, from endpoint to routing rule, is in endpoint evaluation and routing policy.
And the frontier compression appears when a small model approaches the leader.
The index version is part of the number
The citation instruction is one and no score travels without its version and reading, because a bare 63 means nothing and a 63 in the reading of 2026-08-17 means something. The update cadence follows the index. New readings reopen the comparison within the version. The next rebase restarts the procedure from zero.
The work of the week is to log version and reading next to every score that enters a document, spreadsheet or policy. The payoff arrives the day the ruler changes again: whoever knows which ruler produced each number redecides in an afternoon, and whoever mixes versions starts over from zero.
And when the next ruler arrives, the middleware that swaps models without reintegration is the difference between adjusting a policy and rebuilding an integration.
The leaderboard will reset again. That is the only commitment an artificial intelligence index signs. Between one reset and the next, the work is the work of this piece: ruler logged, route retested, cost per task. The leaderboard guides. The route decides.

Save up to 50% in creditswith a single smart API
Connect your operations to our AI Router and optimize the consumption of multiple LLMs
Free TrialRelated articles

Durable execution for AI agents: the engine lives in code
Long-running AI agents fail in the middle, and the difference between redoing and resuming decides cost and trust. Durable execution written in the code itself, with a checkpoint at every step, beats the dedicated orchestration engine for most B2B agent workloads.
Read more
MCP gateway: control plane for AI agent tool traffic
How the MCP protocol reshapes AI agent integration and why tool traffic governance requires a centralized gateway for security, auditing, and cost control.
Read more
Managing Context and Capacity Limits in Multiple AI Models
How to run several AI models while honoring each one's context window and rate limits, routing every task to the model whose window and throughput peak actually fit.
Read more