Prompt Token Cost Calculator

Paste your prompt, set output tokens, and instantly compare API costs across GPT, Claude, Gemini, and Llama — with caching and batch pricing.

Cost Calculator
Quick fill:
Input
0
Output
500
Total
500
Chars
0
375 words
For scale table

Embed v3 English512K

Cohere

$0.00
In $0.00Out $0.00

Rerank v34K

Cohere

$0.00
In $0.00Out $0.00

text-embedding-3-large8K

OpenAI

$0.00
In $0.00Out $0.00

Mistral Embed8K

Mistral AI

$0.00
In $0.00Out $0.00

text-embedding-3-small8K

OpenAI

$0.00
In $0.00Out $0.00
Cheapest

Embed English v34K

Cohere

$0.00
In $0.00Out $0.00

Snowflake Arctic4K

Snowflake

$0.000009
In $0.00Out $0.000009

Ministral 3B128K

Mistral AI

$0.00002
In $0.00Out $0.00002

Gemma 3 4B128K

Google

$0.00002
In $0.00Out $0.00002

Phi-4 Mini16K

Microsoft

$0.000025
In $0.00Out $0.000025

Llama 3.2 3B128K

Meta

$0.00003
In $0.00Out $0.00003

Qwen 3 7B128K

Alibaba

$0.00003
In $0.00Out $0.00003

Llama 3.1 8B128K

Meta

$0.00004
In $0.00Out $0.00004

Llama 3.1 8B (Groq)128K

Groq

$0.00004
In $0.00Out $0.00004

Ministral 8B128K

Mistral AI

$0.00005
In $0.00Out $0.00005

Production Scale — 1,000 calls/day

ModelProviderPer CallDaily CostMonthly (×30d)Annual (×365d)
Embed v3 EnglishCohere$0.00$0.00$0.00$0.00
Rerank v3Cohere$0.00$0.00$0.00$0.00
text-embedding-3-largeOpenAI$0.00$0.00$0.00$0.00
Mistral EmbedMistral AI$0.00$0.00$0.00$0.00
text-embedding-3-smallOpenAI$0.00$0.00$0.00$0.00
Embed English v3Cohere$0.00$0.00$0.00$0.00
Snowflake ArcticSnowflake$0.000009$0.01$0.27$3.29
Ministral 3BMistral AI$0.00002$0.02$0.60$7.30
Gemma 3 4BGoogle$0.00002$0.02$0.60$7.30
Phi-4 MiniMicrosoft$0.000025$0.03$0.75$9.13
Llama 3.2 3BMeta$0.00003$0.03$0.90$10.95
Qwen 3 7BAlibaba$0.00003$0.03$0.90$10.95
Llama 3.1 8BMeta$0.00004$0.04$1.20$14.60
Llama 3.1 8B (Groq)Groq$0.00004$0.04$1.20$14.60
Ministral 8BMistral AI$0.00005$0.05$1.50$18.25
StableLM 2 12BStability AI$0.00005$0.05$1.50$18.25
Mistral 7BMistral AI$0.00005$0.05$1.50$18.25
Llama 3.1 8B (Cerebras)Cerebras$0.00005$0.05$1.50$18.25
Gemma 3 12BGoogle$0.00006$0.06$1.80$21.90
Amazon Nova MicroAmazon$0.00007$0.07$2.10$25.55
Phi-4Microsoft$0.00007$0.07$2.10$25.55
Qwen 3 14BAlibaba$0.000075$0.08$2.25$27.38
DeepSeek R1 Distill Qwen 7BDeepSeek$0.000075$0.08$2.25$27.38
Pixtral 12BMistral AI$0.000075$0.08$2.25$27.38
Llama 3.2 11B VisionMeta$0.00009$0.09$2.70$32.85
Qwen 2.5 Coder 32BAlibaba$0.0001$0.10$3.00$36.50
Mixtral 8×7BMistral AI$0.00012$0.12$3.60$43.80
Amazon Nova LiteAmazon$0.00012$0.12$3.60$43.80
Mixtral 8×7B (Groq)Groq$0.00012$0.12$3.60$43.80
Gemma 3 27BGoogle$0.000135$0.14$4.05$49.28
Gemma 3 27B (Groq)Groq$0.000135$0.14$4.05$49.28
DeepSeek V2.5DeepSeek$0.00014$0.14$4.20$51.10
Gemini 2.0 Flash-LiteGoogle$0.00015$0.15$4.50$54.75
Mistral Small 3Mistral AI$0.00015$0.15$4.50$54.75
Gemini 1.5 FlashGoogle$0.00015$0.15$4.50$54.75
Gemini 2.5 Flash-LiteGoogle$0.0002$0.20$6.00$73.00
Gemini 2.0 FlashGoogle$0.0002$0.20$6.00$73.00
Llama 3.3 70BMeta$0.0002$0.20$6.00$73.00
Qwen 3 32BAlibaba$0.0002$0.20$6.00$73.00
Jamba 1.5 MiniAI21 Labs$0.0002$0.20$6.00$73.00
Grok 3 minixAI$0.00025$0.25$7.50$91.25
DeepSeek R1 Distill Qwen 32BDeepSeek$0.00025$0.25$7.50$91.25
Llama 3.1 70B TurboMeta$0.00027$0.27$8.10$98.55
Llama 4 ScoutMeta$0.000295$0.30$8.85$107.68
GPT-4o miniOpenAI$0.0003$0.30$9.00$109.50
Command RCohere$0.0003$0.30$9.00$109.50
Qwen 2.5 72BAlibaba$0.0003$0.30$9.00$109.50
Llama 3.3 70B (Cerebras)Cerebras$0.0003$0.30$9.00$109.50
Llama 3.3 70B (Groq)Groq$0.000395$0.40$11.85$144.18
Llama 4 MaverickMeta$0.000425$0.43$12.75$155.13
DeepSeek R1 Distill Llama 70BDeepSeek$0.00044$0.44$13.20$160.60
Llama 3.1 70BMeta$0.00044$0.44$13.20$160.60
Qwen 3 72BAlibaba$0.00045$0.45$13.50$164.25
QwQ 32BAlibaba$0.00045$0.45$13.50$164.25
Llama 3.2 90B VisionMeta$0.00045$0.45$13.50$164.25
CodestralMistral AI$0.00045$0.45$13.50$164.25
FireFunction v2Fireworks AI$0.00045$0.45$13.50$164.25
SonarPerplexity$0.0005$0.50$15.00$182.50
DeepSeek V3DeepSeek$0.00055$0.55$16.50$200.75
Mixtral 8×22BMistral AI$0.0006$0.60$18.00$219.00
Sky-T1 32BNovaSky$0.0006$0.60$18.00$219.00
GPT-5 miniOpenAI$0.001$1.00$30.00$365.00
DeepSeek R1DeepSeek$0.001095$1.10$32.85$399.68
Gemini 2.5 FlashGoogle$0.00125$1.25$37.50$456.25
Gemini 3 FlashGoogle$0.0015$1.50$45.00$547.50
Llama 3.1 405BMeta$0.0015$1.50$45.00$547.50
Yi Large01.AI$0.0015$1.50$45.00$547.50
Amazon Nova ProAmazon$0.0016$1.60$48.00$584.00
Claude Haiku 3.5Anthropic$0.002$2.00$60.00$730.00
Claude 3.5 HaikuAnthropic$0.002$2.00$60.00$730.00
Grok 3 Mini FastxAI$0.002$2.00$60.00$730.00
o4-miniOpenAI$0.0022$2.20$66.00$803.00
o1-miniOpenAI$0.0022$2.20$66.00$803.00
o3-miniOpenAI$0.0022$2.20$66.00$803.00
Claude Haiku 4.5Anthropic$0.0025$2.50$75.00$912.50
Gemini 1.5 ProGoogle$0.0025$2.50$75.00$912.50
Mistral Large 2Mistral AI$0.003$3.00$90.00$1,095.00
Mistral Large 2411Mistral AI$0.003$3.00$90.00$1,095.00
Jamba 1.5 LargeAI21 Labs$0.004$4.00$120.00$1,460.00
GPT-4oOpenAI$0.005$5.00$150.00$1,825.00
Gemini 2.5 ProGoogle$0.005$5.00$150.00$1,825.00
Grok 2xAI$0.005$5.00$150.00$1,825.00
Command R+Cohere$0.005$5.00$150.00$1,825.00
Palmyra X4Writer$0.005$5.00$150.00$1,825.00
Inflection 3 PiInflection$0.005$5.00$150.00$1,825.00
GPT-5.4OpenAI$0.0075$7.50$225.00$2,737.50
Claude Sonnet 5Anthropic$0.0075$7.50$225.00$2,737.50
Claude Sonnet 4Anthropic$0.0075$7.50$225.00$2,737.50
Gemini 3 ProGoogle$0.0075$7.50$225.00$2,737.50
Grok 3xAI$0.0075$7.50$225.00$2,737.50
Sonar ProPerplexity$0.0075$7.50$225.00$2,737.50
Palmyra X5Writer$0.01$12.00$360.00$4,380.00
Claude Opus 5Anthropic$0.01$12.50$375.00$4,562.50
o3OpenAI$0.02$20.00$600.00$7,300.00
GPT-Image-1OpenAI$0.02$20.00$600.00$7,300.00
o1OpenAI$0.03$30.00$900.00$10,950.00
Claude Opus 4Anthropic$0.04$37.50$1,125.00$13,687.50
GPT-4.5OpenAI$0.08$75.00$2,250.00$27,375.00
Multi-turn projection (conversation cost grows as history accumulates)
10
10-turn conversation on Embed v3 English: $0.00 total (× a single call) · 22,500 input tokens · 5,000 output tokens. Full conversation simulator →
Building an AI agent or multi-step loop?
Agent costs grow quadratically with step count — your system prompt is re-sent every step. Use the Agent Cost Calculator for full modelling. Agent Cost Calculator →
How to cut this cost
Enable prompt cachingCache the system prompt prefix. Breaks even at 2 calls per TTL window.
Up to 90% on input
Use Batch APIFor non-realtime workloads (classification, summarization, data extraction).
50% all costs
Route to cheaper modelRun evals first — most classification tasks don't need a flagship model.
80–95% overall
Shorten system promptEvery input token is paid on every call. Audit for redundancy and verbose instructions.
5–30% on input

How to Use

Paste your prompt

Type or paste your system message, user prompt, or full conversation. Token count updates instantly in your browser — nothing is uploaded.

Set expected output

Enter the average output tokens the model will generate. Default is 500 tokens (≈ 375 words). Adjust to match your typical response length.

Filter, sort & compare

Use provider tabs to focus on one vendor. Toggle "Cheapest first" to rank models by cost. Enable "Batch API" to see 50%-off async pricing. Toggle caching to simulate cached input.

Project production cost

Enter your daily API call volume — the scale table shows projected daily and monthly spend (×30 days) per model so you can budget before committing.

How LLM API pricing works

Every major LLM provider bills on a pay-per-token model. You pay separately for input tokens (your prompt, system message, conversation history) and output tokens (the model's generated response). Output tokens are typically priced 3–10× higher than input because generating each token requires a full sequential forward pass — whereas input tokens are read in parallel in a single pass.

The formula: Total cost = (input tokens × input rate) + (output tokens × output rate). Rates are expressed per million tokens ($/1M). A 1,000-token prompt at $3/1M costs $0.003 — small per call, but 10,000 calls/day at $0.01 each is $3,000/month. Use the scale table above to see how costs compound at your volume.

Looking to just count tokens without cost analysis? Try our AI Token Calculator — it shows words, characters, and a token-to-cost reference table alongside the count.

2026 LLM Pricing Reference

Current pricing per 1 million tokens. Batch pricing is 50% of standard for async workloads. Always verify with your provider before committing to a budget.

ModelProviderContextInput / 1MOutput / 1MOut/In RatioType
GPT-5.4OpenAI1M$2.50$15.006.0×Standard
GPT-5.4 (Batch)OpenAI1M$1.25$7.506.0×Batch
GPT-5 miniOpenAI400K$0.250$2.008.0×Standard
GPT-5 mini (Batch)OpenAI400K$0.125$1.008.0×Batch
GPT-4oOpenAI128K$2.50$10.004.0×Standard
GPT-4o miniOpenAI128K$0.150$0.6004.0×Standard
o3OpenAI200K$10.00$40.004.0×Standard
o4-miniOpenAI200K$1.10$4.404.0×Standard
Claude Opus 5Anthropic200K$5.00$25.005.0×Standard
Claude Sonnet 5Anthropic200K$3.00$15.005.0×Standard
Claude Haiku 4.5Anthropic200K$1.00$5.005.0×Standard
Claude Haiku 4.5 (Batch)Anthropic200K$0.500$2.505.0×Batch
Claude Haiku 3.5Anthropic200K$0.800$4.005.0×Standard
Gemini 3 FlashGoogle1M$0.500$3.006.0×Standard
Gemini 2.5 ProGoogle2M$1.25$10.008.0×Standard
Gemini 2.5 FlashGoogle1M$0.300$2.508.3×Standard
Gemini 2.5 Flash-LiteGoogle1M$0.100$0.4004.0×Standard
Gemini 2.0 FlashGoogle1M$0.100$0.4004.0×Standard
Gemini 2.0 Flash-LiteGoogle1M$0.075$0.3004.0×Standard
Llama 4 MaverickMeta1M$0.270$0.8503.1×Standard
Llama 4 ScoutMeta128K$0.180$0.5903.3×Standard
Llama 3.3 70BMeta128K$0.230$0.4001.7×Standard
Llama 3.1 405BMeta128K$1.00$3.003.0×Standard
Llama 3.1 8BMeta128K$0.050$0.0801.6×Standard
DeepSeek V3DeepSeek64K$0.270$1.104.1×Standard
DeepSeek R1DeepSeek64K$0.550$2.194.0×Standard
Mistral Large 2Mistral AI128K$2.00$6.003.0×Standard
Mistral Small 3Mistral AI128K$0.100$0.3003.0×Standard
Grok 3 minixAI131K$0.300$0.5001.7×Standard
Command RCohere128K$0.150$0.6004.0×Standard
Embed v3 EnglishCohere512K$0.100$0.0000.0×Standard
Rerank v3Cohere4K$2.00$0.0000.0×Standard
text-embedding-3-largeOpenAI8K$0.130$0.0000.0×Standard
Gemma 3 12BGoogle128K$0.120$0.1201.0×Standard
Ministral 8BMistral AI128K$0.100$0.1001.0×Standard
Ministral 3BMistral AI128K$0.040$0.0401.0×Standard
Mistral EmbedMistral AI8K$0.100$0.0000.0×Standard
Qwen 3 72BAlibaba128K$0.900$0.9001.0×Standard
Qwen 3 32BAlibaba128K$0.400$0.4001.0×Standard
QwQ 32BAlibaba128K$0.900$0.9001.0×Standard
Qwen 3 14BAlibaba128K$0.150$0.1501.0×Standard
Llama 3.2 3BMeta128K$0.060$0.0601.0×Standard
DeepSeek R1 Distill Llama 70BDeepSeek128K$0.880$0.8801.0×Standard
o1OpenAI200K$15.00$60.004.0×Standard
o1-miniOpenAI128K$1.10$4.404.0×Standard
o3-miniOpenAI200K$1.10$4.404.0×Standard
text-embedding-3-smallOpenAI8K$0.020$0.0000.0×Standard
Claude Opus 4Anthropic200K$15.00$75.005.0×Standard
Claude Sonnet 4Anthropic200K$3.00$15.005.0×Standard
Claude 3.5 HaikuAnthropic200K$0.800$4.005.0×Standard
Gemini 3 ProGoogle2M$2.50$15.006.0×Standard
Gemma 3 27BGoogle128K$0.270$0.2701.0×Standard
Gemma 3 4BGoogle128K$0.040$0.0401.0×Standard
Llama 3.1 70BMeta128K$0.680$0.8801.3×Standard
Llama 3.2 11B VisionMeta128K$0.180$0.1801.0×Standard
Llama 3.2 90B VisionMeta128K$0.900$0.9001.0×Standard
Mistral Large 2411Mistral AI128K$2.00$6.003.0×Standard
Mixtral 8×22BMistral AI64K$1.20$1.201.0×Standard
Mixtral 8×7BMistral AI32K$0.240$0.2401.0×Standard
CodestralMistral AI256K$0.300$0.9003.0×Standard
DeepSeek R1 Distill Qwen 32BDeepSeek128K$0.500$0.5001.0×Standard
DeepSeek V2.5DeepSeek64K$0.140$0.2802.0×Standard
Grok 3xAI131K$3.00$15.005.0×Standard
Grok 2xAI131K$2.00$10.005.0×Standard
Command R+Cohere128K$2.50$10.004.0×Standard
Embed English v3Cohere4K$0.100$0.0000.0×Standard
Qwen 3 7BAlibaba128K$0.060$0.0601.0×Standard
Qwen 2.5 72BAlibaba128K$0.600$0.6001.0×Standard
Amazon Nova ProAmazon300K$0.800$3.204.0×Standard
Amazon Nova LiteAmazon300K$0.060$0.2404.0×Standard
Amazon Nova MicroAmazon128K$0.035$0.1404.0×Standard
Jamba 1.5 LargeAI21 Labs256K$2.00$8.004.0×Standard
Jamba 1.5 MiniAI21 Labs256K$0.200$0.4002.0×Standard
Sonar ProPerplexity200K$3.00$15.005.0×Standard
SonarPerplexity128K$1.00$1.001.0×Standard
Palmyra X5Writer1M$6.00$24.004.0×Standard
Palmyra X4Writer128K$2.50$10.004.0×Standard
Snowflake ArcticSnowflake4K$0.006$0.0183.0×Standard
Yi Large01.AI32K$3.00$3.001.0×Standard
FireFunction v2Fireworks AI8K$0.900$0.9001.0×Standard
Inflection 3 PiInflection8K$2.50$10.004.0×Standard
Llama 3.3 70B (Groq)Groq128K$0.590$0.7901.3×Standard
Llama 3.1 8B (Groq)Groq128K$0.050$0.0801.6×Standard
Mixtral 8×7B (Groq)Groq32K$0.240$0.2401.0×Standard
Gemma 3 27B (Groq)Groq128K$0.270$0.2701.0×Standard
Qwen 2.5 Coder 32BAlibaba128K$0.200$0.2001.0×Standard
DeepSeek R1 Distill Qwen 7BDeepSeek128K$0.150$0.1501.0×Standard
Phi-4Microsoft16K$0.070$0.1402.0×Standard
Phi-4 MiniMicrosoft16K$0.025$0.0502.0×Standard
StableLM 2 12BStability AI4K$0.100$0.1001.0×Standard
Sky-T1 32BNovaSky32K$1.20$1.201.0×Standard
GPT-4o (Batch)OpenAI128K$1.25$5.004.0×Batch
GPT-4o mini (Batch)OpenAI128K$0.075$0.3004.0×Batch
Claude Sonnet 5 (Batch)Anthropic200K$1.50$7.505.0×Batch
Gemini 2.5 Flash (Batch)Google1M$0.150$1.258.3×Batch
GPT-4.5OpenAI128K$75.00$150.002.0×Standard
GPT-Image-1OpenAI128K$5.00$40.008.0×Standard
Gemini 1.5 ProGoogle2M$1.25$5.004.0×Standard
Gemini 1.5 FlashGoogle1M$0.075$0.3004.0×Standard
Llama 3.1 70B TurboMeta128K$0.540$0.5401.0×Standard
Mistral 7BMistral AI32K$0.100$0.1001.0×Standard
Pixtral 12BMistral AI128K$0.150$0.1501.0×Standard
Grok 3 Mini FastxAI131K$0.600$4.006.7×Standard
Llama 3.3 70B (Cerebras)Cerebras128K$0.600$0.6001.0×Standard
Llama 3.1 8B (Cerebras)Cerebras128K$0.100$0.1001.0×Standard

Cost Optimization Strategies

Practical techniques to reduce API spend without sacrificing output quality.

StrategyTypical SavingsNotes
Prompt CachingUp to 90% on inputCache static system prompts or documents. Supported by OpenAI and Anthropic. Toggle in calculator above.
Batch API50% overallOpenAI Batch and Anthropic Batches process requests within 24 hours at 50% discount. Enable in calculator above.
Shorten system prompt10–40% on inputEvery request sends the system prompt. A 500-token reduction × 10K calls/day = 5M fewer input tokens/day.
Limit max_tokens20–60% on outputSet max_tokens to the minimum your use case needs. Output is typically your most expensive line item.
Route to smaller model80–95% overallUse Haiku / Flash-Lite for classification, extraction, and summarization. Reserve Opus / Pro for complex reasoning.
Compress conversation history30–70% on inputSummarize older turns instead of appending the full chat history to every new request.

Savings percentages are estimates. Actual results depend on your specific prompts, output length, and provider plan. Always measure with your real workload.

FAQ

Have more questions? Contact us

Last verified: 2026-08-20 · methodology · data sources