AI pricing

AI model API pricing comparison

51 current models from 11 providers, priced per 1 million tokens from each official pricing page (2026-10-08). Type your monthly volume to compare what it costs on every model.
ProviderCached inContext
Nova Micro

Bedrock on-demand standard tier, US pricing

Amazon (Bedrock)$0.035$0.140——$0.630
Command R7B

trial keys are free but rate limited and not for production

Cohere$0.037$0.150—128K$0.675
Nova Lite

Bedrock on-demand standard tier, US pricing

Amazon (Bedrock)$0.060$0.240——$1.08
GPT-OSS 20B

Groq-hosted open model, ~1000 tokens/sec

Groqopen model$0.075$0.300—131K$1.35
Qwen3.8 FlashTogether AIopen model$0.090$0.282—1M$1.46
Ministral 3 8B

batch 50% off

Mistral$0.150$0.150$0.015—$1.80
Claude Haiku 5.5

prompts up to 100K tokens; over 100K: $0.50 in, $2.50 out, $0.05 cached. Batch 50% off

Anthropic$0.100$0.500$0.010—$2.00
GPT-6 Luna

long context: $0.20 in / $0.75 out; batch 50% off

OpenAI$0.100$0.500$0.010—$2.00
Mistral Small 4

batch 50% off

Mistral$0.150$0.600$0.015—$2.70
GPT-OSS 120B

Groq-hosted open model, ~500 tokens/sec

Groqopen model$0.150$0.600—131K$2.70
GPT-OSS 120BTogether AIopen model$0.150$0.600—131K$2.70
OpenAI GPT OSS 120B

serverless standard mode; batch 50% off

Fireworks AIopen model$0.150$0.600$0.015—$2.70
Codestral

batch 50% off

Mistral$0.300$0.900$0.030—$4.80
DeepSeek V4.1 Flash

peak-hour rates (cache miss); off-peak half price: $0.15 in / $0.60 out / $0.003 cached

DeepSeek$0.300$1.20$0.0061M$5.40
DeepSeek V4.1 Flash

serverless standard mode; batch 50% off

Fireworks AIopen model$0.300$1.20$0.006—$5.40
Gemini 3.1 Flash-Lite

paid tier, text/image/video input; audio $0.50 in

Google$0.250$1.50$0.025—$5.50
Gemini 3.5 Flash-Lite

paid tier, text/image/video input; audio input differs

Google$0.300$2.50$0.030—$8.00
Gemini 2.5 Flash

paid tier, text/image/video input; audio $1.00 in

Google$0.300$2.50$0.030—$8.00
Mistral Large 3

batch 50% off

Mistral$0.500$1.50$0.050—$8.00
Nova 2 Lite

Bedrock global cross-region standard tier; geo/in-region $0.33 in / $2.75 out

Amazon (Bedrock)$0.300$2.50——$8.00
Mistral Large 4

listed as a sale price (original $1.36 in / $4.18 out); batch 50% off

Mistral$0.680$2.09$0.070—$10.98
Llama 3.3 70B Instruct TurboTogether AIopen model$1.04$1.04—131K$12.48
Grok Build 0.1

prompts under 200K tokens; long-context rates apply to all tokens once prompt reaches 200K: $2 in / $4 out

xAI$1.00$2.00$0.200256K$14.00
Nova Pro

Bedrock on-demand standard tier, US pricing

Amazon (Bedrock)$0.800$3.20——$14.40
Gemini 3.8 Flash

paid tier; introductory price through Dec 31, 2026, then $1.50 in / $7.50 out / $0.15 cached; batch 50% off

Google$0.750$3.75$0.075—$15.00
Qwen3.8 27B

preview model, may be discontinued at short notice

Groqopen model$0.800$4.00—131K$16.00
Grok 4.3

prompts under 200K tokens; long-context rates apply to all tokens once prompt reaches 200K: $2.50 in / $5 out; batch 20% off

xAI$1.25$2.50$0.2001M$17.50
Claude Haiku 4.5

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$1.00$5.00$0.100—$20.00
DeepSeek V4 Pro

peak-hour rates (cache miss); off-peak half price: $0.66 in / $1.98 out / $0.022 cached

DeepSeek$1.32$3.96$0.0441M$21.12
DeepSeek V4 Pro 0813Together AIopen model$1.32$3.96$0.1301M$21.12
GLM-5.3Together AIopen model$1.40$4.40$0.2601M$22.80
GLM 5.3

serverless standard mode; batch 50% off

Fireworks AIopen model$1.40$4.40$0.260—$22.80
Mistral Medium 3.5

batch 50% off

Mistral$1.50$7.50$0.150—$30.00
Grok 4.7

prompts under 200K tokens; long-context rates apply to all tokens once prompt reaches 200K: $4 in / $12 out

xAI$2.00$6.00$0.500500K$32.00
Grok 4.6

prompts under 200K tokens; long-context rates apply to all tokens once prompt reaches 200K: $4 in / $12 out

xAI$2.00$6.00$0.500500K$32.00
Qwen 3.8 Max

serverless standard mode; batch 50% off

Fireworks AIopen model$2.00$6.00$0.250—$32.00
Gemini 2.5 Pro

prompts up to 200K tokens. Over 200K: $2.50 in / $15 out

Google$1.25$10.00$0.125—$32.50
Nova 2 Pro (Preview)

preview; Bedrock global cross-region standard tier

Amazon (Bedrock)$1.25$10.00——$32.50
Gemini 3.5 Flash

paid tier; cache storage billed extra per hour

Google$1.50$9.00$0.150—$33.00
Claude Sonnet 5.5

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$2.00$10.00$0.1001M$40.00
Claude Sonnet 5

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$2.00$10.00$0.2001M$40.00
GPT-6.1 Sol

long context: $4 in / $15 out; batch 50% off

OpenAI$2.00$10.00$0.100—$40.00
Gemini 3.1 Pro Preview

preview; prompts up to 200K tokens. Over 200K: $4 in / $18 out

Google$2.00$12.00$0.200—$44.00
Command A

trial keys are free but rate limited and not for production

Cohere$2.50$10.00—256K$45.00
GPT-5.3 Codex

Codex model; fast mode costs 2x

OpenAI$1.75$14.00$0.175—$45.50
Claude Sonnet 4.6

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$3.00$15.00$0.3001M$60.00
Kimi K3

serverless standard mode; batch 50% off

Fireworks AIopen model$3.00$15.00$0.300—$60.00
Claude Opus 5.5

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$4.00$20.00$0.2001M$80.00
Claude Opus 5

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$5.00$25.00$0.5001M$100
Claude Fable 5.1

batch 50% off; cached = cache hits (5m write 1.25x input)

Anthropic$10.00$50.00$0.2501M$200
GPT-6 Astra

short context: $20 in / $75 out for long context; batch 50% off

OpenAI$10.00$50.00$1.00—$200

Need benchmarks and specs too? LLM Models compares models in depth. Building with them? See AI dev tools and open-source ChatGPT alternatives.

FAQ

Which AI model API is cheapest?+

By list price per 1M tokens, Nova Micro (Amazon (Bedrock)) is the cheapest on this page: $0.035 input and $0.14 output. Cheaper is not always better: compare quality on your own prompts, and use cached input and batch discounts where providers offer them.

How are these prices measured?+

US dollars per 1 million tokens, standard (non-batch) tier, from each provider's official pricing page on 2026-10-08. Where a price depends on prompt length or time of day, the base tier is shown and the note explains the rest.

What does the "Your month" column mean?+

It multiplies each model’s input and output prices by the monthly token volumes you type above the table, so you can compare what the same workload costs on every model.