Skip to main content

Token Prices Drop, But AI Costs Keep Rising

Rafael Torres
Rafael TorresAugust 9, 20265 min. read
Token Prices Drop, But AI Costs Keep Rising

A company can reduce its price per token and still increase its AI cost. The keyword "AI cost" only becomes intelligible when the equation stops looking at the isolated rate and starts tracking the workload, the completed task, the route, and the cash actually disbursed.

The price of a unit fell. The basket of goods changed.

AI cost is a multiplication with a terrible sense of humor: price per token times tokens consumed, plus the work that the architecture hides until close. A cheaper request can encourage more requests, larger contexts, new features, and automatic retries. The price per token does not need to rise for the bill to grow.

The article AI Token Price Collapse, But AI Costs Are Rising, published by the GTM Newsletter, serves as a secondary source and starting point. It does not serve as primary proof of any single number. The thesis of this article is narrower: CTOs and CFOs must manage effective cost per workload, with volume, route quality, context, retries, operations, currency, and contract in the same frame.

Why Can the Price Per Token Drop While AI Costs Rise?

AI costs rise when the reduction in the price per token is smaller than the increase in work sent to the system or the costs required to complete that work. The relationship is not a universal law. It appears when the product expands usage, the application loads more context, the operation repeats requests, or the contract adds expenses outside the rate card.

The price per token answers a narrow question: how much does it cost to process a unit on that route, under that contract, and in that currency. The bill answers another: how much work the company sent, how many times it sent it, and how many attempts were needed.

The difference stops being academic when a summarization function starts running on every interaction. Context grows. A second attempt is enabled for slow requests. Each decision may make sense in isolation, but the bill observes the whole.

The old efficiency paradox appears once again: when the toll gets cheaper, more people take the road.

A lower price per token reduces the bill if volume, product behavior, route, and contract remain stable. That is the condition. Without measurement, the company calls it savings what may simply be an expansion of purchasing.

What the Price Per Token Does Not Reveal About AI Cost

The price per token does not reveal effective volume, the composition between input and output, context size, repeated requests, operations, or procurement in foreign currency. To compare workloads, the useful unit is cost per completed task. Price per token is one dimension of the equation, not the whole equation.

A team may celebrate cheaper input tokens while ignoring that each request now carries the full conversation history. Another may raise the output limit to avoid truncated responses. The rate fell. The package got larger.

The budget needs to separate five layers:

  1. Price per token: contracted rate for input and output tokens.
  2. Volume: requests, tokens, and completed tasks in the period.
  3. Request quality: success, retries, fallback, and discarded requests.
  4. Operations: observability, caching, queues, context processing, and incidents.
  5. Effective cost: procurement, exchange rate, applicable taxes, fees, and credits captured or lost.

The third layer tends to be silent. A request that fails and is repeated shows up as consumption. The dashboard does not ask whether the first response served the user. The budget should.

The CFO-oriented LLM benchmark puts technical score and economics in the same conversation. This article advances along a different track: the best rate fails when the system does not measure the wasted work around it.

LayerFinancial questionPossible distortion
Price per tokenHow much does the unit cost?The rate drops, but context grows
VolumeHow much is consumed?A function fires requests on every interaction
Request qualityHow much does completion cost?Retries become normal volume
OperationsHow much does maintenance cost?Logs and queues stay outside the comparison
Effective costHow much leaves the bank?Currency, taxes, and contract alter the value

The table is an antidote to the spreadsheet that contains only the "price per million" column.

When Does a Drop in Price Per Token Expand Usage?

A drop in price per token expands usage when it reduces the perceived cost of a task and frees product, engineering, or operations to send more work to the model. The effect depends on pent-up demand, application design, and user behavior. It must be measured as a hypothesis, never treated as destiny.

A product that used AI only at ticket closure may start using AI at triage, search, draft, and audit. The rate remains lower. The consumption points multiply.

The same happens within a single request. Larger contexts reduce pre-processing outside the model but transfer the work to inference. The team may accept this trade when accuracy or development speed justifies the expense. The budget needs to record the choice.

Expansion usually shows up through four signals: requests per user, tokens per request, automated requests, and users or use cases. No signal proves causality on its own.

Attribution requires a baseline with date, route, workload, and volume before the change. Without that baseline, "consumption grew" is an observation, not an explanation.

The technical figure in this article organizes that attribution without inventing market values: baseline, classification, route, context, retries, operations, and effective cost appear as steps in the same equation.

inline-01.png

Where Do Retries, Context, and Routing Eat the Budget?

Retries, context, and routing raise costs when the company pays for work that does not improve the completed task. A retry may recover availability. A larger context may improve the response. A cheaper route may fail on the wrong workload. The diagnosis depends on outcome, not on the isolated rate.

Retries are the easiest case to underestimate. A retry policy protects against transient failures but needs a limit, classification, and telemetry. Without these, a timeout can generate two or three charges, depending on the retry policy and the provider's billing behavior.

Fallback also has a cost. Migrating a request to another route may preserve the service, but the company needs to record how many requests took the secondary path, for what reason, and with what result. Failover is an insurance policy. Insurance policies need claims accounted for.

Context is another form of invisible inflation. Instructions, history, and documents may be re-sent in successive requests because the application did not define compression, caching, or chunk selection. The user sees one response. The company pays for the entire library.

Price-based routing can produce false savings. When a latency-sensitive request goes to a slow route, timeouts, retries, and abandonment appear. The lowest-rate route was chosen; the most expensive task was manufactured.

The Nexforce Router page describes routing by cost, performance, and latency, automatic fallback, budget limits, analytics, and local payment. The point here is not to attribute to the product a promise of automatic savings. It is to make the routing decision observable and auditable.

Why Does Effective Cost Demand a Brazilian Condition?

In CLIENT mode, the direct procurement of a foreign supplier is a hypothesis for analysis, not a ready-made rate. The qualification of the operation, the contractual separation, the beneficiary's jurisdiction, the municipality, and the buyer's tax regime define which line items enter the effective cost. The distributor or reseller mode follows a different logic, and its interpretations do not carry over to the end customer.

A rate in foreign currency is not a universal cash-outflow rate. In Brazil, the direct buyer must classify the operation before calculating any tax. In Solução de Consulta Cosit No. 191, of March 23, 2017, the Receita Federal addressed, based on the facts presented by the interested party, authorizations for SaaS use and access and classified that specific operation as a technical service for IRRF and CIDE purposes. This does not turn every SaaS into an automatic tax category. A pure, segregated license without technology transfer follows the exception provided in art. 2, paragraph 1-A, of Law 10.168/2000. Classification beats the shortcut.

Classification first.

For CLIENT mode, the validation map starts with RIR/2018, arts. 767 and 786, for IRRF and gross-up; Law 10.168/2000, art. 2, paragraphs 1-A and 2, for CIDE; Law 10.865/2004, arts. 7 and 8, for PIS/COFINS-Importation; LC 116/2003, arts. 1, paragraph 1, and 6, paragraph 2, I, for ISS; and Decree 6.306/2007, art. 15-B, XXIV, for IOF-exchange. Application depends on the contract and the rule in force.

This section describes the regime in force as of August 8, 2026. The transition under Lei Complementar 214/2025 extinguishes PIS/COFINS in 2027 and replaces their incidence with CBS; ISS is gradually reduced from 2029 to 2032 and ceases to exist in 2033. The full CBS/IBS rate still depends on the applicable definition. Any future percentage must be identified as a planning assumption, not as a rate currently in force.

This article does not fix rates nor presume that an AI procurement has a single tax treatment. The nature of the operation, the contract, the supplier, and the contracting entity must be examined before any effective cost calculation. The official source cited in the references serves for normative consultation and does not replace case-specific analysis.

The Brazilian equation must record the entity, supplier, nature of service, contract, currency, exchange date, any withholding, and validated tax treatment. No invoice should receive a fixed multiplier by editorial shortcut. The classification depends on the operation's premises, and internal product references do not replace the public source nor the tax validation of the concrete contract.

Which Framework Measures Real AI Cost?

The correct framework tracks each workload from request to completed task. It records consumption, route, quality, and procurement cost under the same identifier. This way, the company separates usage growth, route degradation, operational waste, and currency effects instead of attributing every increase to the model price.

The operation can start with a simple unit: project, key, workload, or period. What matters is keeping the unit stable enough to compare before and after.

Start with the completed task.

The ruler must be stable before the next price-per-token change. Without it, the team measures noise and calls the noise savings. When the definition includes success, route, context, retries, and currency, the comparison stops rewarding the nominal rate and starts showing the cost that actually follows the workload.

  1. Define the completed task. A generated response is not necessarily a resolved task.
  2. Record the route. Model, provider, region, policy version, and fallback reason enter the same event.
  3. Separate input and output. Context tokens and response tokens behave differently.
  4. Count retries. Initial attempt, repetition, timeout, error, and final result remain visible.
  5. Add operations. Caching, observability, queues, and auxiliary processing receive a cost center or traceable estimate.
  6. Convert to effective cost. The procurement currency and the reporting currency appear together, with date and assumptions.
  7. Compare by workload. Price per token is a dimension. Cost per completed task is the decision.

The framework does not need to start perfect. It needs to start before the next price-per-token change.

What Routing Solves, and What It Does Not Solve

Routing solves the lack of control over which request goes to which model, under which rules, and with which result. It does not turn wasted volume into efficiency nor fix a poor definition of success. Its role is to make cost, performance, availability, and limits observable operational decisions.

Routing matters.

An LLM gateway can distribute requests and choose routes by cost, latency, or performance. It can also apply spending limits, log requests, and trigger failover, when these capabilities are documented in the adopted solution. The economic value appears when the policy is explicit and the company can compare the route outcome with the completed task.

Low-complexity requests do not need to consume the same route as a task requiring broad context. A latency-sensitive path must not be optimized for nominal price if the delay triggers repetition. A project that hit its limit needs to stop, alert, or change policy.

The Router must not be used as an excuse to erase complexity. A spending limit without task measurement can cut a critical workload. Failover without quality analysis can preserve availability while degrading the response.

Mature architecture treats routing as economic policy applied in real time. The question stops being "which model is cheaper?" and becomes "which route delivers this task within the accepted budget and service level?"

Is the Argument for Price Per Token Wrong?

The best counterargument says that if the price per token falls and the workload remains the same, total cost falls. Under those premises, the argument is correct. The divergence begins when "same workload" is merely a hypothesis, without a baseline for volume, context, retries, route, input-output distribution, and contractual conditions.

The steelman deserves to be preserved. An operation with fixed volume, the same success rate, the same context, and the same procurement captures the drop in price per token. Denying this turns a good thesis into cheerleading.

The math does work.

The conclusion changes when the basket purchased changes: the team may unlock a feature previously blocked by budget, replace one request with three steps, or attach more context to reduce retrieval, while the finance team continues comparing only the nominal price per token published on the provider's rate card.

Whoever can prove stability should capture the reduction. Whoever cannot prove stability should treat the drop as an opportunity to investigate the workload. The counterargument wins in the controlled environment; it loses when the company confuses price per token with quantity of work purchased.

Frequently Asked Questions About Price Per Token and AI Cost

The short answer is that the price per token informs the rate, while AI cost measures the completed work and everything required to complete it. The difference includes volume, context, retries, route, operations, currency, and contract. The company that does not measure these layers knows the list price but does not know its real expense.

Does a cheap token always reduce the bill?

No. The price per token reduces the request rate when the route and contract are the same. The bill rises when volume, context, responses, retries, or number of tasks grow faster than the price reduction. The only way to confirm savings is to compare the baseline with the subsequent workload.

Which metric should go into the budget?

The primary metric should be effective cost per workload or per completed task. Price per token remains useful for negotiating, simulating, and comparing routes, but it does not show repeated requests, excessive context, operations, currency, or contractual conditions. The budget must keep these dimensions separate before aggregating them.

Does usage expansion happen with every price drop?

No. Expansion depends on pent-up demand, the product, the users, and engineering decisions. The analysis must compare requests per user, tokens per request, tasks, features, and success rate before and after the change. Without a baseline, the growth may have another cause.

Are retries always waste?

No. Retries can recover a transient failure and preserve an important task. Waste appears when there is no limit, classification, or telemetry to distinguish the attempt that solved the problem from the repetition that merely generated another charge. The correct metric combines number of retries, reason, outcome, and cost.

Does the exchange rate change the price per token?

The exchange rate does not change the price per token contracted in foreign currency, but it changes the effective cost for those who report in local currency. The impact depends on the contract, settlement date, spread, and, in CLIENT mode, the tax qualification of the operation and the buyer's regime. A Brazilian reference must not be automatically applied to another operation.

When does routing truly help?

Routing helps when the company has different workloads, cost or performance rules, a need for failover, and little visibility over requests. It makes the policy explicit and auditable. It does not replace the definition of a completed task, the baseline, quality measurement, or contract validation.

References and Further Reading

The Equation That Survives the Price

The AI budget needs to survive a price per token that changes. It requires measuring the workload, recording the route, counting retries, and separating the price per token from the effective cost. This discipline reveals when a price reduction became a consumption expansion, when a failure produced extra requests, and when foreign-currency procurement altered the local outlay.

The company that measures only the token asks how much a part costs. The company that measures the workload discovers how much the entire machine costs, including the screw that falls to the floor every time a retry fires.

Price per token is negotiation. Effective cost is management.

Nexforce

Save up to 50% in creditswith a single smart API

Connect your operations to our AI Router and optimize the consumption of multiple LLMs

Free Trial

Related articles