AI Economics 12 min read By GreyBath Technology

The AI Cost Paradox: Cheaper Models, Bigger AI Bills

AI inference keeps getting cheaper, yet agents, frontier training, power and wider adoption can push total spending higher. Here is what businesses should measure.

The AI Cost Paradox: Cheaper Models, Bigger AI Bills

AI models are getting cheaper to run. AI budgets are still getting larger. Both trends are real, and the reason is easier to see once we stop treating “AI cost” as one number.

A company can pay less for a million tokens and spend more overall because it is processing more documents, giving agents longer jobs and adding AI to workflows that did not exist a year earlier. Outside the API bill, model developers are funding training clusters, data centres, memory, power connections and research teams.

The practical unit for a business is not the token. It is a completed, checked outcome: one support case resolved, one invoice processed or one qualified lead handed to sales.

Fixed-capability intelligence is becoming cheaper. Reliable automation still has a full operating cost.

Five costs that should not be confused

Most arguments about AI economics mix numbers from different layers. Separating them prevents a cheap API request from being compared with a billion-dollar data-centre project.

  1. Price per token. The developer's published API charge for model input and output.
  2. Provider inference cost. The chips, memory, electricity, cooling, networking, spare capacity and serving software used to generate a response. Providers rarely disclose this figure in full.
  3. Cost per completed task. Every model call, search, tool, retry, verification step and human check needed to finish useful work.
  4. Training cost. The compute, data, researchers, evaluation and experimentation required to create or materially improve a model.
  5. Infrastructure cost. Land, buildings, accelerators, power connections, cooling and backup systems that may support many models over several years.

The third figure is the one most business cases miss. Customers do not buy tokens for their own sake. They buy an answer or an action they can trust.

Why fixed-capability inference keeps getting cheaper

Stanford's 2025 AI Index estimated that the cost of reaching GPT-3.5-level performance on the MMLU benchmark fell from about $20 per million tokens in November 2022 to $0.07 by October 2024. That is a reduction of more than 280 times for a comparable capability level. [6]

Decline in the cost of reaching GPT-3.5 level performance from 2022 to 2024
The cost of reaching a fixed benchmark level fell from about $20 to $0.07 per million tokens.

Epoch AI's 2026 estimate is more useful for planning than the most dramatic historical comparisons: fixed-capability inference has recently become roughly five to ten times cheaper per year, with considerable variation by task. [7]

No single breakthrough produced that curve. New accelerators process more tokens per watt; smaller models now handle work that once required a frontier model; quantisation reduces memory; distillation transfers useful behaviour; and batching, caching, routing and better serving software keep expensive hardware busier.

The August 2026 price ladder

Published API prices still span a wide range. The figures below are US dollars per one million tokens and should be checked again before procurement because providers change prices, discounts and caching rules.

ModelInputOutput
OpenAI GPT-5.6 Luna$0.20$1.20
Google Gemini 3.7 Flash$0.75$3.75
OpenAI GPT-5.6 Terra$2.00$12.00
Anthropic Claude Sonnet 5$2.00$10.00
OpenAI GPT-5.6 Sol$5.00$30.00
Input and output prices for major AI model APIs in August 2026
Headline price is only one variable; output, retries, tools and tokenisation also affect the result.

OpenAI, Anthropic and Google apply different tokenisers, reasoning settings, context limits and cache charges. [1] [2] [3] A lower-priced model that needs three attempts may cost more than a stronger model that succeeds once.

A cheap query does not reveal provider profitability

Public information is not sufficient to claim that every paid API request loses money. A well-utilised service can earn a positive margin on an individual request while the company remains unprofitable after free usage, research, training, unused capacity, depreciation, safety work, support and financing.

List prices also reflect competition. OpenAI reduced prices in July 2026, Anthropic retained an introductory Sonnet price, and Google published a future price change for Gemini. [1] [2] [3] Those decisions are commercial strategy as well as engineering arithmetic.

Why agents can cost more even when tokens get cheaper

A chatbot normally receives one request and returns one answer. An agent may plan, search, read company documents, call an API, inspect the response, retry a failed step and ask a stronger model to verify the final result. Some of that activity is never visible to the user.

Gartner expects provider inference cost per agentic workflow to increase more than fivefold through 2028. Goldman Sachs describes one user request expanding into 10, 20 or even 50 model operations and forecasts a 24-fold rise in total token consumption from 2026 to 2030. [8] [9]

Cost comparison between a simple chatbot request and a multi-step AI agent
An agent's planning, tools, verification and retries can turn one request into many model calls.

A calculation using the listed OpenAI prices

Consider a simple chat using 2,000 input and 1,000 output tokens, then an agent using 100,000 total input and 20,000 output tokens across its full run.

WorkloadGPT-5.6 LunaGPT-5.6 Sol
Simple chat request$0.0016$0.04
Complex agent workflow$0.044$1.10

Search APIs, databases, orchestration, monitoring and human review sit outside those model totals. The finished Sol workflow costs about 27.5 times the simple Sol request even though both use the same price card.

Cheaper supply also creates more demand. Teams add AI to more products, process larger documents and automate work that was previously uneconomic. Economists recognise this as the Jevons effect: efficiency can increase total consumption rather than reduce it.

Training, data centres and power move on a different curve

Serving yesterday's capability is getting cheaper. Building the frontier remains capital intensive. Stanford estimated GPT-4 training compute near $78 million and Gemini Ultra near $191 million. Epoch AI projects that the largest training runs could exceed $1 billion by 2027 if historical scaling continues. [10] [11]

Estimated growth in frontier AI model training costs
Frontier training estimates have moved from tens of millions towards billion-dollar runs.

These estimates are ranges, not invoices. Stanford's 2026 review found that developers are disclosing less about datasets, parameter counts, training code and compute use, which makes outside cost estimates harder. [12]

The bill extends beyond accelerators. Frontier work needs high-bandwidth memory, fast networking, storage, datasets, reinforcement-learning environments, researchers, evaluation and thousands of smaller experiments. Epoch AI estimates that high-bandwidth memory now represents about 63% of AI-accelerator component cost, up from 52% in early 2024. [13]

A data centre is not one training run

Epoch AI estimates roughly $38 billion of upfront capital for a typical one-gigawatt AI data centre and about $0.9 billion in annual operating expense. [14] That facility can train, test and serve multiple models for years. Treating its entire cost as the price of one model exaggerates the comparison.

The slower part of the stack

The IEA's 2026 outlook puts global data-centre electricity use near 485 terawatt-hours in 2025 and about 950 terawatt-hours in 2030, with AI-focused facilities potentially tripling their demand. [22] Epoch AI estimates that the largest individual training runs could seek four to sixteen gigawatts by 2030, while warning that the upper end may be physically unattainable. [16]

Projected growth in global data-centre electricity demand through 2030
Data-centre electricity demand may almost double between 2025 and 2030.

A model can change in months. Transmission lines, substations, transformers and cooling systems take years. The IEA estimates that roughly 20% of planned data-centre projects could be delayed without action on grid constraints. [17]

What one query uses

Microsoft Research measured a median of about 0.31 watt-hours for an optimised frontier-scale query. A long reasoning request producing around 5,000 output tokens used roughly thirteen times more. [18] A normal prompt is small; billions of prompts and persistent agents are not. Workload shape matters more than a single viral estimate.

Bubble risk or structural boom?

AI is a stack of businesses rather than one investment. Chips, data centres, model providers, API platforms and workflow applications have different margins and failure modes.

Nvidia reported $75.2 billion of data-centre revenue in the first quarter of its 2027 financial year, up 92% year over year, with gross margin near 75%. [19] Gartner expects inference infrastructure spending of $23.3 billion in 2026 to exceed training infrastructure spending of $19 billion. [5] Production demand and supplier revenue are measurable.

Goldman Sachs' baseline still points to about $765 billion of annual AI capital expenditure in 2026, rising towards $1.6 trillion in 2031. [4] Returns will not be evenly distributed. Capacity can arrive before customers, grid delays can leave hardware idle, and subsidised model or application pricing may not survive competition.

Productivity evidence is similarly uneven. Stanford's 2026 AI Index cites gains around 14% to 15% in customer support and 26% in software development, with weaker results where work was hard to monitor or errors needed heavy review. [21] A workflow costing ₹100 and saving ₹1,000 has a sound business case. A ₹10 workflow producing unreliable work does not.

AI is a structural technology boom with speculative pockets, not a single yes-or-no bubble.

What pricing may look like by 2030

Forecasts should be treated as planning ranges. Four directions have stronger support than a detailed prediction for every market segment.

  1. Fixed-capability AI becomes cheaper. Gartner expects provider inference cost for a one-trillion-parameter model to be more than 90% lower in 2030 than in 2025. [20] Providers may retain part of that saving for reliability, security and margin.
  2. Frontier intelligence retains a premium. The industry will keep creating a new top tier with longer context, more reasoning, better tools and reserved capacity, even as older capability moves down the price curve.
  3. Agent workflows remain expensive in the near term. Longer jobs, premium decisions, tools and verification can outweigh reductions in token price.
  4. Business pricing shifts towards outcomes. Tokens remain useful for infrastructure accounting, while customers increasingly pay for cases resolved, documents processed, leads qualified, time saved or service guarantees.

Some prices will move against the broad trend. Google, for example, published a scheduled increase for Gemini 3.7 Flash from January 2027. [3] Budgets need room for provider changes as well as efficiency gains.

What Indian businesses should measure

The cheapest token price is rarely enough to choose an architecture. Salary structure, error risk, customer expectations, language, exchange rates, tax and integration effort all influence the completed-task cost.

1. Price the successful outcome

Track cost per support case resolved, invoice processed, accurate product description, appointment booked or qualified sales lead. Include human correction and the cost of errors.

2. Route work instead of using one model everywhere

A small model can classify and draft, a mid-tier model can handle normal work, and a frontier model can be reserved for difficult cases. High-risk decisions still need a person who is accountable for the result.

Cost-efficient multi-model AI architecture with routing and human review
Route each request to the least expensive capable model, then verify according to risk.

3. Put boundaries around agents

Set maximum steps, tokens, retries, searches, tool calls, time and spend per task. Retrieve only the documents needed for the current job; sending the whole history or document library on every run wastes money and can increase privacy risk.

4. Test Indian operating conditions

Hindi, Marathi, Gujarati, Tamil, Bengali and other languages can tokenise differently across models. Test real customer conversations rather than extrapolating from English. Add USD-INR movement, GST treatment, cloud charges and payment costs to the budget.

5. Keep the workflow portable

Separate business logic from the provider, cache repeated instructions, batch work that is not urgent and keep interfaces replaceable. Compare the automated workflow with the full human process: turnaround time, missed follow-ups, service quality and capacity matter alongside salary.

GreyBath architecture principle: use the least expensive model that can complete the task reliably, limit its freedom, verify according to business risk and record the cost of the outcome. This is an engineering framework, not a claim based on invented project savings.

Cheap intelligence still needs disciplined economics

Model efficiency is improving quickly, while agents, frontier research and physical infrastructure expand the amount of AI work being attempted. The result is a divided market: older capability becomes abundant, frontier reliability stays premium, and total enterprise spending grows as more workflows become viable.

Durable value is likely to sit with businesses that connect models to useful distribution, clean data and real workflows, then deliver reliability at a controlled serving cost. Model size alone is not a business moat.

For buyers, the decision is concrete: measure the verified task, keep the system portable and spend on stronger models only when their quality changes the result.

Frequently Asked Questions

What are AI inference costs?

They are the costs of running a trained model to process input and generate a response: accelerator time, memory, electricity, cooling, networking and serving software. A business workflow may add tools, data services, monitoring and human review.

Why can AI agents cost more than chatbots?

An agent may plan, retrieve documents, call tools, inspect results and retry. One user request can therefore create many model calls and a much larger input and output total.

Is AI a bubble?

AI has real adoption, supplier revenue and productivity gains, alongside speculative valuations and projects that may not earn their capital cost. Risk differs across chips, infrastructure, models and applications.

How should a business control AI cost?

Measure verified outcomes, route simple work to smaller models, limit agent steps and retries, cache repeated context, batch non-urgent work and retain human review where errors are expensive.

Build a cost-efficient AI workflow

GreyBath Technology can help design a practical multi-model system for your website, CRM, portal or internal operations.

Discuss your workflow

References

  1. OpenAI, Advancing the price-performance frontier with GPT-5.6.
  2. Anthropic, Introducing Claude Sonnet 5 and API pricing update.
  3. Google AI for Developers, Gemini Developer API pricing.
  4. Goldman Sachs, The Assumptions Shaping the Scale of the AI Build-Out, May 2026.
  5. Gartner, Worldwide AI-Optimized IaaS Spending Forecast, August 2026.
  6. Stanford HAI, AI Index 2025: State of AI in 10 Charts.
  7. Epoch AI, How persistent is the inference cost burden?, February 2026.
  8. Gartner, Agentic Workflow Inference Cost Forecast, August 2026.
  9. Goldman Sachs, AI Agents Forecast to Boost Tech Cash Flow as Usage Soars, May 2026.
  10. Stanford HAI, AI Index: State of AI in 13 Charts.
  11. Epoch AI, How much does it cost to train frontier AI models?
  12. Stanford HAI, Inside the AI Index: 12 Takeaways from the 2026 Report.
  13. Epoch AI, AI chip component cost shares, May 2026.
  14. Epoch AI, The Finances of AI: Data and Research.
  15. International Energy Agency, Energy demand from AI.
  16. Epoch AI, Power demands of frontier AI training.
  17. International Energy Agency, Energy and AI executive summary.
  18. Microsoft Research, Energy use of AI inference, efficiency pathways, and test-time scaling, Joule, 2026.
  19. NVIDIA, First Quarter Fiscal 2027 Financial Results.
  20. Gartner, 2030 LLM Inference Cost Forecast, March 2026.
  21. Stanford HAI, Economy: The 2026 AI Index Report.
  22. International Energy Agency, Key Questions on Energy and AI, 2026.

Editorial note: provider prices, infrastructure forecasts and market conditions change. Verify current pricing before making procurement or budget decisions.

More practical thinking from GreyBath across design, engineering and growth.

All articles

Think fast. Build faster - launch a human-centered product that performs.