Skip to main content

Quasar 438B: the European 438B model enters the routing leaderboard

Camila Duarte
Camila DuarteSeptember 8, 202610 min. read
Quasar 438B: the European 438B model enters the routing leaderboard

What happened

Multiverse Computing launched Quasar on September 2, 2026, its first large reasoning model, native in English and Spanish. The leaderboard that feeds AI model routing gained a new point: index 43 on the Artificial Analysis Intelligence Index v4.1.1, read on 2026-09-07, per the official announcement. It is 500 tokens in 15.3 seconds.

The verified launch numbers

Quasar 438B is a 438B-class reasoning model built for enterprise agents and code, with planning, tool use, and large context, available through the CompactifAI API and with no public price at the time of publication. The numbers below were published by the vendor itself, citing Artificial Analysis; they count as a dated citation, read on 2026-09-07.

The index 43 is a composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. In the published comparison, the model passes Mistral Medium 3.5 (30), NVIDIA Nemotron 3 Ultra (38), and Inkling (42). The field leader remains Claude Opus 5, at 63.

Speed repositions the leaderboard. Quasar 438B returns 500 tokens with thinking included in 15.3 seconds. In the source's comparison, only three models are faster: Nemotron 3.5 Lightning (index 24, at 9.4s), Gemini 3.5 Flash-Lite (37, at 10.8s), and Gemini 3.7 Flash (56, at 11.5s). Of the models that score above Quasar 438B, only two respond in under 25 seconds. Inkling, the immediate neighbor on the index, returns the same 500 tokens in 48.3s.

Those two facts together are the point: the models above Quasar 438B on the index are, with two exceptions, slow on the published metric.

On long context, the AA-LCR of 75.0 tied Grok 4.6 (high) and finished less than one point below Claude Opus 5 (75.7) and Qwen3.8 2.4T A95B (75.3), on the same reading. On Terminal-Bench v2.1, the model scored 69.3, the result the source itself flags as the largest improvement margin; category leadership belongs to Claude Opus 5, at 89.1.

Two cautions before using these numbers. First: they are the vendor's numbers, citing Artificial Analysis, not an independent test by this blog. Second: the announcement page carries an internal inconsistency in the NVIDIA Nemotron 3 Ultra index, 38 in the leaderboard paragraph and 36 in the caption of one figure; this post adopts 38, the value from the index paragraph.

The bilingualism is official: English and Spanish. On Portuguese, the source says nothing.

Why a European model changes route decisions

Because the frontier between quality and latency gained a point that did not exist before September 2, 2026. Fast points already existed in the published comparison: Gemini 3.7 Flash (56, at 11.5s) and Mistral Medium 3.5 (30, at 18.8s). What was missing was the pairing: AA-LCR of 75.0, tied at the top of the published reading, 0.7 from Claude Opus 5, with 15.3s per 500 tokens with thinking, plus European jurisdiction and native English and Spanish. A European AI model with reasoning, 43 on the index, and 15.3 seconds per 500 tokens widens the set of viable routes for anyone routing per task.

For the CTO and the CFO who pay per token consumed, the useful question was never which model is the best in the world. It is which route serves this task inside the latency and cost budget. That is the thesis the post LLM Benchmarks for CFOs: cost, not technical score already defended with other data. Today's delta: the set of routes that decision evaluates has grown, with a point that combines reasoning, long context, and low latency. The economics behind that budget is what the post The token price collapse and the real cost of AI mapped: the price fell, the total bill did not.

Flash-latency reasoning shows up in production as operations, not as a demo: interactive agent loops, where every iteration pays for the thinking, and long bilingual document analysis, where English and Spanish arrive mixed in the same flow. In a loop with 20 responses of 500 tokens, the difference between 48.3s and 15.3s is 16 minutes versus 5, on the published metric alone, before counting the rest of the pipeline.

Inkling and Quasar sit in the same neighborhood on the index, 42 and 43. On latency, they live on different continents. It is 48.3s against 15.3s. Whoever compares leaderboards without the wait column buys a route too expensive for the loop that has to run.

One detail remains, and it only appears with the version in hand: 43 is the largest result among European models in the published comparison (v4.1.1, read on 2026-09-07). For the buyer evaluating jurisdiction and data residency, there is now a European candidate with a competitive score. The source does not declare where the CompactifAI API serves inference. A European vendor does not imply an EU serving region or EU data residency by default. Verifying serving regions and processor terms before the contract is the buyer's job. Serving region and data residency are contract terms to verify, not consequences of a vendor's nationality. Jurisdiction is not a dimension the gateway applies on its own; it is a buyer criterion, and the route now exists.

inline-01.png

What changes in practice

The route criterion for long-context reasoning workloads no longer has a single default that is expensive in waiting time. Before September 2, no point in the published comparison carried AA-LCR 75.0 with 15.3s per 500 tokens with thinking, and the best European on the leaderboard was at 30. After the launch, there is a bilingual point with index 43, AA-LCR 75.0, and 15.3 seconds to evaluate as a candidate.

Route criterionBefore September 2After the launch
Latency near the frontierFast points existed (Gemini 3.7 Flash: 56 at 11.5s; Mistral Medium 3.5: 30 at 18.8s), but not the AA-LCR 75.0 pairing with 15.3sA point at 15.3s per 500 tokens with thinking, AA-LCR 75.0, and index 43
Language coverageNo European reasoning point with declared native English and Spanish on the published v4.1.1 leaderboardEnglish and Spanish in the same reasoning model
European jurisdictionBest European in the published comparison: Mistral Medium 3.5 (30)A European route the buyer can evaluate
DecisionChoice by aggregate index, with latency as a side effectChoice per task, with an explicit latency and cost budget

Table numbers: Multiverse Computing citing Artificial Analysis, Intelligence Index v4.1.1, read on 2026-09-07.

The operational difference sits in the last row: before, latency was a side effect of a choice made by the aggregate index; after, the latency and cost budget is named first, and the index is evaluated against it.

The table does not say the new route wins everything. It says the route enters the set. For the task that demands the absolute top of the leaderboard, 63 from Claude Opus 5 on the 2026-09-07 reading, nothing changed. What changed is the band where 43 is enough, and that band is wide in production: most of a company's internal tasks are not frontier benchmarks.

What to do now

Five actions fit this week, before the next leaderboard moves the numbers. They apply to teams already routing per task and to anyone still choosing a model by the aggregate index. The deadline is the news itself. The 2026-09-07 reading ages fast.

  1. Model latency and cost per task, not the aggregate index alone. A 43 at 15.3s beats a 42 at 48.3s on most interactive workloads, even though the leaderboard treats the two as neighbors. The cost framing is in How to reduce LLM inference costs.
  2. Include the European route as a candidate for bilingual long-context workloads and agentic work, with a pilot measuring cost per task and perceived latency, not the score alone.
  3. Pin the index version in the model dossiers. v4.1.1 and the v4.2 rebase of 2026-09-04 are not comparable with each other; a citation without a version contaminates the decision.
  4. Pilot before changing the default route. Index and latency are the vendor's numbers, citing Artificial Analysis. Your workload decides, with real traffic and an explicit latency and cost budget.
  5. Keep model choice decoupled from the application. A gateway with per-task rules and failover lets the new route enter testing without replatforming; that is the design of Corporate AI gateway: LLM routing.

Frequently asked questions about Quasar 438B

Short answers to the questions that arrive before the analysis. The numbers cited carry the index version and the read date.

How much does Quasar 438B cost?

There is no public price at the time of publication. The September 2, 2026 announcement confirms availability through the CompactifAI API and carries no price table. A number quoted by a third party before an official pricing page is a guess, and the cost per token of your operation is defined in the contract and by the volume.

Does Quasar 438B have open weights?

The source does not mention open weights, and speculating about undeclared plans adds nothing. The criterion for choosing between open weights and hosted models is what the post Open weights vs hosted models: selection and governance details; what this launch changes is the route offering, not the distribution regime.

What is Intelligence Index v4.1.1?

It is the Artificial Analysis composite that sums nine evaluations, from GDPval-AA v2 to AA-LCR. The version matters. On 2026-09-04 there was a rebase to v4.2, with every score redeclared. Quasar 438B's 43 is from v4.1.1, read on 2026-09-07, and it cannot be placed side by side with v4.2 numbers.

Does Nexforce Router route to Quasar 438B today?

No, and this post does not claim it. The dimensions the Router applies in model selection are cost, performance, latency, and context; a new route enters the way any candidate does, with a pilot before the default route changes. For bilingual long-context workloads, the model now belongs in the set worth evaluating.

Is Quasar 438B good for Portuguese workloads?

The announcement declares English and Spanish as native languages and does not mention Portuguese. There is no published data on pt-BR performance, so the honest answer is that it cannot be affirmed. Teams with Portuguese workloads should treat the route as bilingual until there is a test of their own or a vendor statement.

References and Further Reading

Primary sources and supporting reading for the route decision. The bulletin that surfaced the event is a discovery channel, not a citation; what it reported enters here through the original source. The numbers in this post are the 2026-09-07 reading of the September 2, 2026 announcement, which cites Artificial Analysis, and they count as a dated citation, not as independent testing. Start with the primary source.

The European route on the leaderboard: what to watch next

Europe entered the routing leaderboard with a route that combines reasoning, long context, and low latency, and the leaderboard rewards whoever treats each new point as a testable candidate, not as a headline. The next test named by the source itself is Terminal-Bench: today's 69.3 sits 19.8 points from category leadership, and that is where the improvement margin holds up or does not, on the 2026-09-07 reading.

The leaderboard will move again.

For anyone operating routes in production, the value of the launch is not swapping everything for a new route. It is the ability to test it without replatforming: keeping model choice decoupled from the application is what a gateway with per-task rules offers, and that is the design of Nexforce Router. The new route does not change your application's contract. It changes the set that routing evaluates, and the set got better for anyone who speaks English and Spanish and pays the thinking bill.

Event: September 2, 2026. Reads and verification: 2026-09-07.

Nexforce

Accelerate your company'sbusiness and operational efficiency

We design the technology of tomorrow to boost your business operational scale

Talk to a Specialist

Related articles