Introduction
July 2026 compressed a year of AI pricing news into three weeks. SpaceXAI (formerly xAI) launched Grok 4.5 on July 8, 2026, at $2 per million input tokens and $6 per million output tokens. OpenAI shipped its GPT-5.6 family on July 9, then cut GPT-5.6 Luna by 80% on July 30. In between, Anthropic released Claude Opus 5 on July 24 at $5 and $25 per million tokens, the same rates as Opus 4.8.
Grok 4.5’s low launch price, GPT Luna’s 80% cut, and Opus 5’s price hold are three versions of the same move. Each lowers the cost of intelligence, yet only OpenAI’s counts as a straight price cut. Anthropic instead held its price while claiming near-Fable 5 capability at half of Fable’s $10/$50 rate. The labs are no longer competing on leaderboard rank alone. They are competing on intelligence per dollar.
So, which frontier AI model gives businesses the most intelligence for their money? The correct answer depends on the workload, because the cheapest token is rarely the cheapest finished task. This article covers the forces behind the AI price war in 2026 and explains the Claude Opus 5 vs GPT-5.6 vs Grok 4.5 pricing picture, and how to find the cheapest model for your work.
What is Driving the AI Price War in 2026?
Three main factors are driving the AI price war in 2026:
- Lower costs per query
- Buyers demanding proof of return
- Budget rivals competing at the low end of the market
All three are pushing prices down, and they kind of came together in July. This led to an unusually large number of pricing changes in a single month.
Improving model economics comes first. OpenAI attributed its July 30 cuts to efficiency gains made during GPT-5.6’s development. They claimed that the model helped make itself cheaper to serve by optimizing its own production code. It also shows that falling serving costs give labs room to cut prices without giving up margin.
The continuous pressure from buyers is another thing that moves the needle. According to CNBC, enterprises have grown reluctant to deploy expensive models without a clear picture of the return on their investment. Anthropic tried to answer this by adjusting the product, rather than the price tag. Axios reported that they created an effort dial on Opus 5 to allow users to trade capability for cost per call.
Lastly, open-weight models, like GLM 5.2 and Moonshot’s Kimi K3, keep squeezing the value argument for frontier pricing. Google added to this pressure in the same July window with Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both these models are built around lower model operating costs. Each of these releases improves the tradeoff between intelligence and cost. In simple terms, open-weight models give you more capability for the same price or similar capability for less.
Claude Opus 5 vs GPT-5.6 vs Grok 4.5 at a Glance
The table below compares the three model families on the numbers that decide a purchase. All prices are API rates per million tokens, verified against provider announcements on July 31, 2026.
|
Factor |
Claude Opus 5 |
GPT-5.6 |
Grok 4.5 |
|
Provider |
Anthropic |
OpenAI |
SpaceXAI (formerly xAI) |
|
Released |
July 24, 2026 |
July 9, 2026 |
July 8, 2026 |
|
Positioning |
Everyday frontier model, near Fable 5 at half the price |
Three tiers: Sol (flagship), Terra (balanced), Luna (fast and cheap) |
Value flagship for coding and agents |
|
Input price |
$5 |
Sol $5 / Terra $2 / Luna $0.20 |
$2 |
|
Output price |
$25 |
Sol $30 / Terra $12 / Luna $1.20 |
$6 |
|
Context window |
1M tokens |
1.05M tokens |
500K tokens |
|
Intelligence Index (Artificial Analysis) |
61 (#1) |
Sol 59 / Terra 55 / Luna 51 |
54 |
|
Best fit |
Complex coding, agents, and knowledge work |
Tiered routing, high-volume work on Luna |
Cost-sensitive agentic and coding workloads |
Pricing verified July 31, 2026. Intelligence Index scores from Artificial Analysis Index v4.1.
The table already shows the strategy split. Anthropic holds price at the top, OpenAI spreads three tiers underneath it, and SpaceXAI undercuts the flagships with one model.
Claude Opus 5 Pricing: Unchanged at $5/$25
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens on the Claude API. It is the same as Opus 4.8, as Anthropic held the price while raising capability. According to the launch announcement, they positioned Opus 5 as approaching Fable 5’s frontier intelligence at half the price. Independent scoring supports the claim. Artificial Analysis places Opus 5 first on its Intelligence Index at 61, one point above Fable 5, at 26% lower cost per task.
The standard API price is $5 per million input tokens and $25 per million output tokens. Batch processing cuts both rates in half, to $2.50/$12.50, for jobs that do not need to run immediately.
Cached inputs are even cheaper at $0.50 per million tokens, a 90% discount. Anthropic also offers a Fast mode that runs about 2.5 times faster for twice the price, at $10/$50.
For people using Claude through the app, the Pro plan starts at $20 per month. For businesses running Claude via the API, the API pricing is the more important metric.
One more billing detail deserves attention before migration. Opus 5 runs adaptive thinking by default, and thinking tokens bill as output at $25 per million. Because of this, a flat sticker price can still produce a higher invoice if the model reasons longer on the same prompts. The full Claude Opus 5 breakdown covers the benchmarks and the migration details.
To make the rates concrete, a call with a 100,000-token prompt and a 10,000-token response costs $0.75 at standard rates. Batching the same call drops it to $0.38, and caching a stable prefix pushes it lower again.

GPT-5.6 Pricing: Luna, Terra, and Sol
GPT-5.6 pricing splits across three tiers, and two of them just became far cheaper. OpenAI launched the family on July 9, 2026, with Sol at $5/$30, Terra at $2.50/$15, and Luna at $1/$6 per million tokens. Three weeks later, on July 30, the company cut Luna by 80% and Terra by 20% while leaving Sol untouched.
|
GPT-5.6 variant |
Launch price (July 9) |
Price after July 30 cut |
Change |
|
Luna |
$1 / $6 |
$0.20 / $1.20 |
80% cut |
|
Terra |
$2.50 / $15 |
$2 / $12 |
20% cut |
|
Sol |
$5 / $30 |
$5 / $30 |
Unchanged |
|
Sol Fast mode |
Not offered |
$10 / $60 |
New, 2.5x throughput |
Input/output rates per million tokens, per OpenAI announcements.
The GPT-5.6 Luna price is the headline. At $0.20 input and $1.20 output, Luna’s combined $1.40 per million tokens undercuts Google’s Gemini 3.5 Flash-Lite at $2.80 combined and Gemini 3.6 Flash at $9, according to VentureBeat’s comparison. Luna also scores 51 on the Artificial Analysis Intelligence Index, which puts a frontier-family model at commodity rates.
Terra now sits at $2/$12 for balanced production work, while Sol stays at $5/$30 as the flagship. Instead of cutting Sol, OpenAI added a Fast mode at $10/$60 that delivers up to 2.5 times the throughput of standard processing. The move mirrors Anthropic’s Fast mode: the premium buys latency, not intelligence.
For deployment economics, the cut changes routing math more than model choice. OpenAI recommends Sol for the hardest reasoning, Terra for steady production tasks, and Luna for high-volume routine calls, and the price gap between those tiers has widened further. A workload that was borderline on Luna at $1 input is economical at $0.20.

Grok 4.5 API Pricing
Grok 4.5 API pricing is $2 per million input tokens and $6 per million output tokens, with cached input at $0.50 per million and a 500,000-token context window. SpaceXAI, the SpaceX division formed from xAI in the February 2026 merger, launched the model on July 8, 2026. The prices it offers are below both rival flagships from day one.
However, the headline rate hides several extras worth watching. Priority processing doubles token rates to $4/$12. Similarly, live web and X search are billed separately at $5 per 1,000 tool calls. SpaceXAI is also not offering any batch discount for Grok 4.5. Higher input rates also apply above 200,000 prompt tokens, so long-context jobs cost more than the sticker suggests.
Even with those extras, the case for Grok 4.5 rests on price per point. Artificial Analysis scored 54 on its Intelligence Index, fourth at launch behind only Fable 5, GPT-5.5, and Opus 4.8. Its reported cost was around $0.31 per Index task.
Since then, Opus 5, GPT-5.6 Sol, and Kimi K3 have all landed above it. This leaves Grok 4.5 below the top tier on capability and far below it on cost. This trade of rank for cost is the story we told in how Grok 4.5 turned the AI race into a pricing battle. Three weeks of market response have confirmed our predictions.

Which Is the Cheapest Frontier AI Model?
GPT-5.6 Luna is the cheapest frontier AI model by token price in July 2026. It is priced at $0.20 input and $1.20 output per million tokens. Among the flagship tiers, Grok 4.5 is cheapest at $2/$6, ahead of Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30. Cheapest input tokens and cheapest output tokens both point to Luna, while the cheapest flagship points to Grok 4.5. However, this answer only holds for one definition of cheap.
Things tend to change a little when we talk about the cheapest mode for a specific task. It depends on the input/output mix, since Opus 5 beats Sol on output-heavy work ($25 versus $30 per million) while matching it on input. This means that if you need a certain level of quality, the cheapest model may be different again. A cheaper model is not a good deal if it cannot meet the quality you need.
In practice, the useful question is which model is cheapest for my workload. The next two sections show why.
The Problem With Comparing Token Prices
Token prices are only one input to the final bill. Real workloads differ on input/output ratio, prompt length, reasoning volume, retries, tool calls, caching, latency, and how much human review the output needs. Each factor can flip a comparison that looks settled on paper.
The output side of the bill diverges first. Grok 4.5 charges three times its input rate for output, while most rivals charge five to six times. This is the gap that favors Grok on output-heavy generation. Reasoning volume compounds the same side of the bill, since Opus 5 and the GPT-5.6 family both bill thinking tokens as output. More thinking means a bigger invoice at identical rates.
The input side has its own levers. Opus 5 reads cached prefixes at $0.50 per million (with a new 512-token minimum), Grok 4.5 matches that $0.50 rate, and GPT-5.6 discounts cached input by 90%. Prompt length is the other input lever. This is the one that teams can control directly, and our prompt engineering guide shows where to trim it.

In the above example, the model with 10 times the token price ships the work at roughly a third of the total cost. Gaps that wide are realistic on hard tasks and unrealistic on easy ones, which is the point. The flip depends on your workload, so measure it on your workload.
What Does Price Per Intelligence Mean?
Price per intelligence is the cost required to reach a defined level of useful task performance. Instead of asking what a million tokens costs, it asks what a correct answer, a merged pull request, or a resolved support ticket costs. It also includes the cost of retries and human correction time.
A public version of the idea exists. Artificial Analysis publishes cost-per-task figures alongside its Intelligence Index. This is how Grok 4.5 can trail the leaders on raw score and still look strong. Its 54 points at roughly $0.31 per Index task are a different proposition from a higher score at several times that cost. The same lens explains why Anthropic markets Opus 5 on cost per task instead of rank alone.
Businesses should build their own benchmark on the tasks they actually run. Public benchmarks grade exam performance and say nothing about your invoice. A benchmark built on your own work replaces that exam score with numbers you can act on. Six metrics do most of that work.
- Cost per completed task: Count only the tasks that finished to your standard. Failed attempts still cost money, so their tokens stay in the total while the attempts stay out of the count.
- Cost per accurate answer: Divide total spend by the answers that survived checking. A model that's cheap per call and wrong a third of the time isn't cheap.
- Cost per production-ready code change: Measure the change that passed review and shipped, not the diff the model produced. Draft code that needs a rewrite gets paid for twice.
- Cost per resolved support case: Resolved means the customer stopped writing back. A ticket that closes and reopens a day later belongs in the failure column.
- Cost per research report: Include the searches, the reruns, and the sources someone had to verify by hand. Those three usually cost more than the final generation did.
- Minutes of human correction per output: Track this next to every figure above, because it never shows up on the invoice. A cheap model that needs 20 minutes of editing costs more than an expensive one that needs two.
One formula covers all of them.
Price per intelligence = (API spend + human correction cost) ÷ tasks completed to your quality bar.
A model wins on this measure by charging less per token, by failing less often, or both. This is the axis Anthropic, OpenAI, and SpaceXAI now compete on, and it explains July’s pricing moves better than any leaderboard does.
Claude Opus 5 vs GPT-5.6 vs Grok 4.5: Performance vs Price
The performance picture splits by workload, and the source of each number matters. The independent Artificial Analysis Coding Agent Index ranks GPT-5.6 Sol first, with a score of 80 on OpenAI’s Codex harness. It also has a lower cost per task than Fable 5 and Opus 4.8.
Grok 4.5 scores 76 in its Grok Build harness, on par with GPT-5.5 at a small fraction of the token spend. Anthropic’s own testing, which is vendor data rather than independent evaluation, puts Opus 5 within 0.5% of Fable 5’s peak CursorBench score at half the cost per task.
On agentic knowledge work, Artificial Analysis now ranks Opus 5 first on its AA-Briefcase benchmark. It is nearly 150 Elo ahead of Fable 5 at 20% lower cost per task. Grok 4.5 ranks fourth on GDPval-AA v2, Artificial Analysis’s test of economically valuable work, between Opus 4.8 and GLM 5.2.
On raw intelligence, the July 2026 ladder on Index v4.1 puts Opus 5 on top by a single point.
- Opus 5: 61
- Fable 5: 60
- GPT-5.6 Sol: 59
- Kimi K3: 57
- Opus 4.8: 56
- GPT-5.6 Terra: 55
- Grok 4.5: 54
- GPT-5.6 Luna: 51
A 10-point spread separates Opus 5 from GPT-5.6 Luna, while input prices across the same eight models differ by 25 times. So, the capability band is narrow and the price band is wide. That combination is what starts a price war.
The top three sit inside two points of each other, which is close enough that a rerun could reorder them. Buyers who can't separate Opus 5, Fable 5, and Sol on capability will separate them on price instead. The pressure lands hardest at the bottom of the ladder, where Luna at 51 has the least capability to defend whatever it charges.
In the long context, Opus 5 includes its full 1 million-token window at the standard rate. GPT-5.6 lists a 1.05 million-token window, but bills prompts above 272,000 input tokens at double the input rate, and Grok 4.5 caps at 500,000 tokens with a surcharge above 200,000. Long-document pipelines should price their real prompt sizes instead of the advertised window.
Note: Benchmark scores travel badly, and a model that tops its own vendor’s harness may not top yours. Treat these numbers as a screening filter rather than a verdict.
Which Model Offers the Best Value?
No universal winner exists, so the best value differs by workload. For complex coding and long-horizon agents, Claude Opus 5 offers the strongest value right now. It tops the independent intelligence and agentic-work rankings at half of Fable 5’s price. Also, it includes the full 1M context as standard. GPT-5.6 Sol is the counterargument for Codex-committed teams, with the top Coding Agent Index score at $5/$30.
For high-volume routine work, such as classification, extraction, summarization, and support triage, GPT-5.6 Luna is the clear value at $0.20/$1.20. Nothing else in the frontier families comes close on cost per call at that volume. This makes Luna the natural engine for AI automation pipelines.
For cost-sensitive agentic workloads that still need real capability, Grok 4.5 earns its slot. An Index score of 54 at $2/$6 buys most of the leaders’ capability for far less money. The caveats are the 200,000-token surcharge line and the separate tool-call fees, so budget for both.
AI Model Pricing 2026: What Happens Next?
Expect the price war to continue, because none of its three drivers are easing things down. Inference keeps getting cheaper, buyers keep demanding a return, and open-weight models keep improving. The following two second-order effects matter more for planning than the next cut.
First, model routing becomes the default architecture. When Luna handles routine calls at $1.40 combined and Opus 5 handles hard ones at $30 combined, a router that classifies task difficulty pays for itself quickly.
Both vendors are also building for it. OpenAI publishes three-tier routing guidance, and TechCrunch reported that Anthropic now supports switching models mid-task. Building that routing layer is where AI integration work concentrates in the coming months of 2026.
Second, multi-model setups become insurance rather than optimization. Fable 5's 19-day government suspension in June showed how fast a single-model dependency turns into an outage. Open-weight fallbacks, like GLM 5.2, make a second lane affordable.
However, open weights are not cheap to run by default. Artificial Analysis notes that Kimi K3 costs more per task than Opus 4.8. Each of those tasks also takes it nearly an hour.
Beyond those two effects, expect commoditization at the bottom and differentiation at the top. Routine intelligence keeps falling toward commodity pricing, while the labs charge premiums for latency, autonomy, and reliability. The twin $10/$50 and $10/$60 Fast modes are the clearest sign of that premium.
How Should Businesses Choose an AI Model?
Treat model selection as a measurement exercise rather than a shopping decision. The same process applies whether you are picking a chat model or scoping a full generative AI build:
- Define the workload as concrete tasks and monthly volumes.
- Set the quality bar: what counts as an acceptable output, and who checks it.
- Measure token usage per task on realistic prompts, including thinking tokens.
- Shortlist two or three models at different price points, such as Opus 5, Sol, and Grok 4.5 for agentic work.
- Run 50 to 100 real tasks through each model.
- Count completed tasks, retries, and tool calls.
- Time the human correction each model’s failures require.
- Compute cost per completed task (correction time included).
- Check latency, rate limits, and compliance terms.
- Pick per workload, and revisit quarterly, because July 2026 proved that prices move monthly.
Most teams skip steps 5 through 8 for lack of time, and these are the steps that decide the outcome. A pricing table takes an afternoon to read, but a cost-per-task evaluation is the only comparison that holds up in production.
Claude Opus 5 vs GPT-5.6 vs Grok 4.5: Final Verdict
Claude Opus 5 wins on capability per dollar, GPT-5.6 Luna wins on raw price, and Grok 4.5 wins on price per intelligence point in between. GPT-5.6 Sol stays the pick for teams committed to Codex. It’s critical to understand that each vendor is pricing for a different segment instead of racing for the same crown. This is why no honest comparison produces one winner.
The larger shift will outlast this month’s numbers. The real AI price war is moving from price per token to value per successful outcome. The effort dials, model tiers, and Fast modes that shipped in July are all built for that fight. Buyers should do the same. Measure cost per completed task and route work across models.
Get Help Choosing the Right Model
Most teams comparing these three models never get past the pricing table, because a real cost-per-task evaluation takes engineering time they don’t have. This evaluation is the core of our AI consulting and strategy work.
Our AI agent development team also builds the routing layers that let you use each model where it is cheapest. If you want help pressure-testing your setup before the next price move, book a free consultation. We’ll walk through your workloads together to see how we can get maximum gains for your business.
Conclusion
A new model will land next week, or the week after. July alone produced four frontier releases and one 80% price cut, and August is unlikely to be quieter. Chasing each release is a treadmill, and businesses that switch models on each announcement pay for the churn in re-testing and broken edge cases.
Your regular tasks don’t need the latest and greatest. They need the model that fits. The one that clears your quality bar at the lowest cost per completed task, whatever its leaderboard position. A model that did your work well on July 23 still does it well today, and no launch announcement changes that.
Price moves are the exception worth acting on. When a cut is large enough to change your cost per completed task, re-run the evaluation and switch if the numbers say so. Luna’s 80% drop just crossed that bar for high-volume work. Until then, stability beats novelty.
Book a Free 30-Minute Meeting
Discover how our services can support your goals — no strings attached. Schedule your free 30-minute consultation today and let's explore the possibilities.
Book a Free Call