Introduction
All three clouds charge $2 per million input tokens for their mid-tier model, and $10 to $12 for output. The rates have converged so cloud choice barely moves the bill now. The defining factor is the model choice, which moves it by 25 times or more. The same application can cost $39 or $975 a month on identical traffic.
Model pricing is now the dominant line item for most production AI applications. All three hyperscalers publish per-token rates and leave the costly extras in footnotes.
Every figure below comes from the providers’ own pricing pages, checked in the first week of August 2026. All are US dollar list prices, excluding tax, on standard US-region tiers. This article covers what each cloud charges per token, what it charges around the token, and how to work out which provider is most feasible for you.
What Do AI Models Cost on AWS, Azure, and Google Cloud?
Token pricing across the three clouds runs from about $0.04 per million input tokens to $5.50. Output rates run from $0.08 to $33.00.
|
Cloud |
Tier |
Model |
Input per 1M |
Output per 1M |
|---|---|---|---|---|
|
AWS Bedrock |
Cheap |
Amazon Nova Micro |
$0.035 |
$0.14 |
|
AWS Bedrock |
Mid |
Claude Sonnet 5 |
$2.00 |
$10.00 |
|
AWS Bedrock |
Flagship |
GPT-5.6 Sol, in-region |
$5.50 |
$33.00 |
|
Azure Foundry |
Cheap |
GPT-5.6 Luna |
$0.20 |
$1.20 |
|
Azure Foundry |
Mid |
GPT-5.6 Terra |
$2.00 |
$12.00 |
|
Azure Foundry |
Flagship |
GPT-5.6 Sol |
$5.00 |
$30.00 |
|
Google Cloud |
Cheap |
Gemini 3.1 Flash-Lite |
$0.25 |
$1.50 |
|
Google Cloud |
Mid |
Gemini 3.1 Pro Preview |
$2.00 |
$12.00 |
|
Google Cloud |
Flagship |
Gemini 3.1 Pro, long context |
$4.00 |
$18.00 |
Two patterns matter more than any individual number. The mid-tier has converged on $2 and $12 across all three clouds, so the rate no longer separates providers. The vertical gap inside each cloud stays close to 25 times.
Every model on the list charges four to eight times more for output than for input. This means any verbose feature spends on the expensive side of the meter, and trimming response length is easier than switching models.
How AWS, Azure, and GCP Structure AI Model Pricing?
Each provider sells the same tokens through a different purchasing structure, and the structure moves the effective rate more than the headline does.
Bedrock lists Standard, Flex, Priority, and Reserved tiers. Standard is the default rate, Flex sits 50% below it, and Priority carries a 75% premium. Batch inference runs at the same 50% discount as Flex. Provisioned Throughput comes from the account team rather than the price list.
Azure prices by deployment type rather than by tier. Every model deploys as Global, Data Zone, or Regional. Microsoft publishes Global rates and quotes Data Zone and Priority Processing through sales. Batch returns completions within 24 hours for 50% off Global Standard. Provisioned Throughput Units reserve capacity at an hourly rate whether you use it or not.
Google Cloud publishes three tiers per model. Priority costs 1.8 times standard, and Flex and Batch both cost half. Non-global endpoints carry a 10% premium, in force for Gemini 3 models, from July 1, 2026.
As a result, two teams on the same published rate can land three times apart on effective cost. One runs everything on Priority against a regional endpoint. The other batches half its traffic globally, while neither of them changed models.

Amazon Bedrock Pricing in 2026
Bedrock’s advantage is catalog breadth, and the rate card reflects it. Amazon Bedrock’s pricing page now lists models from more than 15 providers, so cross-vendor routing works inside one API.
Claude Sonnet 5 carries a promotional rate of $2.00 and $10.00 per million tokens through August 31, 2026. The standard $3.00 and $15.00 rates apply after that, so budget with the higher figure.
Bedrock also serves OpenAI’s models for in-region inference in US East and Ohio. AWS prices them at parity with OpenAI’s data residency tier, which works out to 10% above the standard rates.
|
Bedrock model |
Input per 1M |
Output per 1M |
Context |
|---|---|---|---|
|
GPT-5.6 Sol |
$5.50 |
$33.00 |
272K |
|
GPT-5.6 Terra |
$2.20 |
$13.20 |
272K |
|
GPT-5.6 Luna |
$0.22 |
$1.32 |
272K |
|
GPT-5.6 Sol |
$11.00 |
$49.50 |
1M |
|
GPT-5.6 Terra |
$4.40 |
$19.80 |
1M |
|
GPT-5.6 Luna |
$0.44 |
$1.98 |
1M |
The million-token context window doubles the input rate and lifts output by half. GPT-5.6 Sol goes from $5.50 and $33.00 to $11.00 and $49.50. Teams enable long context by default without checking whether requests need it, then pay for capacity nobody uses.
Bedrock gets cheap on the open-weight side. DeepSeek v3.2 costs $0.62 and $1.85. Mistral Large 3 costs $0.50 and $1.50, GLM 4.7 Flash costs $0.07 and $0.40, and Gemma 3 4B costs $0.04 and $0.08. The four models handle classification, data extraction, and routing at a fraction of frontier rates.
Azure and Microsoft Foundry Pricing in 2026
Azure sells models through Microsoft Foundry, and the rate card starts from OpenAI’s published list. OpenAI cut Terra by 20% and Luna by 80% on July 30, 2026, while Sol held. Foundry’s published Global rates now reflect the cut.
|
Model |
Input per 1M |
Output per 1M |
|---|---|---|
|
GPT-5.6 Sol |
$5.00 |
$30.00 |
|
GPT-5.6 Terra |
$2.00 |
$12.00 |
|
GPT-5.6 Luna |
$0.20 |
$1.20 |
The Luna cut widened Azure’s spread between the cheapest and most expensive GPT-5.6 tier from five times to 25 times. Tier routing is now the highest-return optimization on Azure.
Verifying the Azure figures takes an extra step. Azure’s OpenAI pricing page renders rates through a region and currency selector rather than a static table, so the numbers never appear in a saved copy. Check your own region and deployment type first.
Azure’s real pricing story is rarely the list. Organizations with an Enterprise Agreement negotiate against committed spend, and the discount usually beats anything published here. If your company has a Microsoft contract, start there.
Google Cloud Gemini Pricing in 2026
Google publishes the cleanest rate card of the three. Every tier and context band is laid out on its generative AI pricing page.
|
Model |
Input per 1M |
Output per 1M |
|---|---|---|
|
Gemini 3.1 Pro Preview |
$2.00 |
$12.00 |
|
Gemini 3.6 Flash |
$1.50 |
$7.50 |
|
Gemini 3.5 Flash |
$1.50 |
$9.00 |
|
Gemini 3 Flash Preview |
$0.50 |
$3.00 |
|
Gemini 3.5 Flash-Lite |
$0.30 |
$2.50 |
|
Gemini 3.1 Flash-Lite |
$0.25 |
$1.50 |
|
Gemini 2.5 Flash-Lite |
$0.10 |
$0.40 |
Two footnotes change budgets more than the headline rates. Pro models reprice above 200K input tokens, where Gemini 3.1 Pro moves to $4.00 and $18.00. The higher rate applies to the whole request, not just the excess.
Tuned endpoints are the second surprise. Google charges 1.5 times the base rate to serve a fine-tuned Gemini 3 model, so customization carries a permanent inference premium.
The context window is Google’s real pricing advantage. Every tier down to Flash-Lite supports a very large context, which sometimes removes the need for a retrieval pipeline. Sending a whole document to a cheap model can cost less than building one. Generative AI solutions designed around a big context window have fewer moving parts.
Four Discounts Worth Using
Four discounts cut token costs materially, and all three clouds offer a version of each.
|
Mechanism |
AWS Bedrock |
Azure Foundry |
Google Cloud |
|---|---|---|---|
|
Batch or asynchronous |
50% off |
50% off Global Standard |
50% off |
|
Flex or deferred |
50% off |
Included in Batch |
50% off |
|
Cached input reads |
Up to 90% off |
90% off uncached input |
90% off uncached input |
|
Priority or fast |
75% premium |
Priority processing available |
1.8 times standard |
Caching is the discount most teams under-use. Take an application sending 1,500 input tokens per request, 1,000 of them a fixed system prompt. At 2 million requests a month on a $2 mid-tier model, the uncached input bill is $6,000. Caching the fixed portion cuts it to $2,400. The saving is $3,600 a month for a change to one API parameter.
Cache writes have one catch. On GPT-5.6 and later models, writing to the cache costs 1.25 times the normal input rate, so a prefix used once or twice loses money. Caching pays on heavy reuse.
Batch is the easiest saving on the list and the most neglected. Any workload not needing a real-time answer qualifies, from document processing to overnight summarization. All three clouds charge half price.
Charges Beyond the Token Rate
Most teams model for the token rates so they are part of the calculations. The rest of the cost arrives as a surprise, and this is where things can get messy.
- Managed retrieval: Bedrock Knowledge Bases cost $5.00 per GB of raw data per month, plus $1.00 per 1,000 retrieval calls. AWS’s worked example reaches $350 a month at 50 GB and 100,000 queries, and storage accrues whether anyone searches or not.
- Agentic retrieval: Multi-hop retrieval on Bedrock costs $4.00 per 1,000 calls on top of the underlying retrievals. AI agents making two hops per query pay $6.00 per 1,000.
- Grounding: Google includes 5,000 grounding queries a month at no charge across Gemini 3 models, then charges $14 per 1,000 against Search or Maps. Grounding against your own data costs $2.50 per 1,000.
- Web search: Amazon Bedrock AgentCore Web Search costs $7 per 1,000 queries.
- Guardrails: Bedrock content filters and denied-topic filters each cost $0.15 per 1,000 text units of up to 1,000 characters. An AI chatbot running both on 1,000 queries an hour pays about $650 a month, on AWS’s own worked example.
- Model routing: Bedrock’s Intelligent Prompt Routing costs $1.00 per 1,000 requests, or $300 a month at 10,000 requests a day.
None of these charges is unreasonable alone. Together they routinely add 30% to 50% to a token bill, and none appear in the comparison tables teams use.
What AI Model Costs Look Like on a Real Monthly Bill?
Assume 1,500 input tokens and 400 output tokens per request, typical for a retrieval-backed assistant. The table below prices three traffic levels against three model tiers. The figures are illustrative, not quotes.
|
Monthly traffic |
Cheap tier |
Mid tier |
Flagship |
|---|---|---|---|
|
50,000 requests |
$39 to $49 |
$350 to $390 |
$975 |
|
2 million requests |
$1,560 to $1,950 |
$14,000 to $15,600 |
$39,000 |
|
20 million requests |
$15,600 to $19,500 |
$140,000 to $156,000 |
$390,000 |
The cheap-tier column uses GPT-5.6 Luna and Gemini 3.1 Flash-Lite. The mid-tier uses Claude Sonnet 5 and Gemini 3.1 Pro. The flagship column uses GPT-5.6 Sol.
Read the rows rather than the columns. At 2 million requests a month, moving from flagship to cheap tier saves $37,000. Moving between clouds at a fixed tier saves a few hundred dollars. Provider comparison is the wrong optimization to run first.
The second lesson is in the bottom row. A mid-tier default at enterprise volume costs $150,000 a month. Route half the traffic to a cheap tier and the saving is about $65,000. Routing at scale is architecture.
When Self-Hosting Beats Per-Token Pricing
Self-hosting starts winning at high sustained volume and loses badly below it. The breakeven is easier to calculate than most teams assume.
A managed API bills per token and scales to zero, so a quiet weekend costs nothing. A self-hosted model bills per hour whether traffic arrives or not. On AWS, a real-time SageMaker endpoint on ml.g5.2xlarge costs about $1.52 an hour, or roughly $1,110 a month running continuously. SageMaker AI pricing puts the managed premium at around 25% over bare EC2.
The same $1,110 buys around 740 million output tokens at Gemini 3.1 Flash-Lite rates. Very few applications generate 740 million output tokens a month, which is why managed APIs win.
Utilization decides everything above the breakeven. An always-on endpoint at 5% utilization costs 20 times more per request than the same endpoint at full load. The invoice looks identical. Measure your traffic distribution, not your peak.
Bedrock’s Custom Model Import sits between the two. It charges $0.05718 per Custom Model Unit per minute, billed only on invocation, and scales to zero after five idle minutes. A Llama 3.1 70B model needs eight units, so about $27 an hour while serving and nothing at idle.

Which Cloud is Cheapest for AI Models in 2026?
No provider wins outright, and the honest answer depends on the shape of the workload.
|
Workload |
Usually cheapest |
Reason |
|---|---|---|
|
High-Volume Simple Tasks |
Azure |
GPT-5.6 Luna at $0.20 and $1.20 is the lowest published frontier-family rate |
|
Long-Document Processing |
Google Cloud |
Large context at Flash-Lite prices can replace a retrieval pipeline |
|
Open-Weight And Specialist Models |
AWS |
The widest catalog, with several models under $0.10 per million input tokens |
|
Batch And Offline Work |
Any of the three |
All three discount by exactly 50% |
|
Regulated Workloads |
Azure |
Data Zone and Regional deployments, priced above Global but below the alternatives |
|
Retrieval-Heavy Applications |
Depends on design |
Managed retrieval charges vary more than token rates |
|
Existing Enterprise Agreements |
Whichever holds the commitment |
Negotiated discounts beat list-price arbitrage |
Region, deployment type, and commitment level all flip these answers. So does where your data already lives. Egress usually costs more than the token saving is worth. Any serious AI strategy prices the whole path rather than the model alone.
How To Reduce AI Model Costs Across AWS, Azure, And GCP
Seven changes are worth making, roughly in order of return on effort.
- Route by task complexity: Send classification and extraction to a cheap tier. Reserve the flagship for work with a measurable need.
- Cache fixed prefixes: System prompts, few-shot examples, and static context all qualify for the 90% discount on cached reads.
- Batch everything non-interactive: The same model costs half price, and the only code change is the API call.
- Cut retrieved context: Returning 20 mediocre chunks instead of three good ones multiplies the input bill by seven and usually produces worse answers.
- Constrain output length: Output costs four to eight times input, so an instruction to be concise pays for itself immediately.
- Check your context window setting: On Bedrock, the million-token window doubles input rates whether or not requests need it.
- Measure cost per successful task: Cost per request rewards cheap failures. Cost per completed task rewards the outcome you actually want.
Most of the list is configuration rather than engineering. Points two, three, five, and six take an afternoon and cut a token bill by a third.
Conclusion
Stop comparing providers and start comparing tiers. The gap between AWS, Azure, and Google Cloud on the same model is about 10%. Inside any one of them, the cheapest tier to flagship is 25 times or more. This makes it very clear that model routing returns far more than provider selection.
Everything above assumes the application sends sensible prompts to a sensibly chosen model, and the assumption fails often. The inflated AI bills we review trace back to one default nobody revisited. A prototype flagship never gets re-tested. A retrieval system returns too much context. A real-time call runs where a batch job would do. All three are AI integration problems, and they’re cheaper to fix than to keep paying for.
For a second opinion on where your model spend is going, book a free consultation with our team. We’ll look at your traffic, your prompts, and your invoice, then say what we’d change first.
Book a Free 30-Minute Meeting
Discover how our services can support your goals — no strings attached. Schedule your free 30-minute consultation today and let's explore the possibilities.
Book a Free Call