Prompt Token Cost Calculator
Paste your prompt, set output tokens, and instantly compare API costs across GPT, Claude, Gemini, and Llama — with caching and batch pricing.
Embed v3 English512K
Cohere
Rerank v34K
Cohere
text-embedding-3-large8K
OpenAI
Mistral Embed8K
Mistral AI
text-embedding-3-small8K
OpenAI
Embed English v34K
Cohere
Snowflake Arctic4K
Snowflake
Ministral 3B128K
Mistral AI
Gemma 3 4B128K
Phi-4 Mini16K
Microsoft
Llama 3.2 3B128K
Meta
Qwen 3 7B128K
Alibaba
Llama 3.1 8B128K
Meta
Llama 3.1 8B (Groq)128K
Groq
Ministral 8B128K
Mistral AI
Production Scale — 1,000 calls/day
| Model | Provider | Per Call | Daily Cost | Monthly (×30d) | Annual (×365d) |
|---|---|---|---|---|---|
| Embed v3 English | Cohere | $0.00 | $0.00 | $0.00 | $0.00 |
| Rerank v3 | Cohere | $0.00 | $0.00 | $0.00 | $0.00 |
| text-embedding-3-large | OpenAI | $0.00 | $0.00 | $0.00 | $0.00 |
| Mistral Embed | Mistral AI | $0.00 | $0.00 | $0.00 | $0.00 |
| text-embedding-3-small | OpenAI | $0.00 | $0.00 | $0.00 | $0.00 |
| Embed English v3 | Cohere | $0.00 | $0.00 | $0.00 | $0.00 |
| Snowflake Arctic | Snowflake | $0.000009 | $0.01 | $0.27 | $3.29 |
| Ministral 3B | Mistral AI | $0.00002 | $0.02 | $0.60 | $7.30 |
| Gemma 3 4B | $0.00002 | $0.02 | $0.60 | $7.30 | |
| Phi-4 Mini | Microsoft | $0.000025 | $0.03 | $0.75 | $9.13 |
| Llama 3.2 3B | Meta | $0.00003 | $0.03 | $0.90 | $10.95 |
| Qwen 3 7B | Alibaba | $0.00003 | $0.03 | $0.90 | $10.95 |
| Llama 3.1 8B | Meta | $0.00004 | $0.04 | $1.20 | $14.60 |
| Llama 3.1 8B (Groq) | Groq | $0.00004 | $0.04 | $1.20 | $14.60 |
| Ministral 8B | Mistral AI | $0.00005 | $0.05 | $1.50 | $18.25 |
| StableLM 2 12B | Stability AI | $0.00005 | $0.05 | $1.50 | $18.25 |
| Mistral 7B | Mistral AI | $0.00005 | $0.05 | $1.50 | $18.25 |
| Llama 3.1 8B (Cerebras) | Cerebras | $0.00005 | $0.05 | $1.50 | $18.25 |
| Gemma 3 12B | $0.00006 | $0.06 | $1.80 | $21.90 | |
| Amazon Nova Micro | Amazon | $0.00007 | $0.07 | $2.10 | $25.55 |
| Phi-4 | Microsoft | $0.00007 | $0.07 | $2.10 | $25.55 |
| Qwen 3 14B | Alibaba | $0.000075 | $0.08 | $2.25 | $27.38 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | $0.000075 | $0.08 | $2.25 | $27.38 |
| Pixtral 12B | Mistral AI | $0.000075 | $0.08 | $2.25 | $27.38 |
| Llama 3.2 11B Vision | Meta | $0.00009 | $0.09 | $2.70 | $32.85 |
| Qwen 2.5 Coder 32B | Alibaba | $0.0001 | $0.10 | $3.00 | $36.50 |
| Mixtral 8×7B | Mistral AI | $0.00012 | $0.12 | $3.60 | $43.80 |
| Amazon Nova Lite | Amazon | $0.00012 | $0.12 | $3.60 | $43.80 |
| Mixtral 8×7B (Groq) | Groq | $0.00012 | $0.12 | $3.60 | $43.80 |
| Gemma 3 27B | $0.000135 | $0.14 | $4.05 | $49.28 | |
| Gemma 3 27B (Groq) | Groq | $0.000135 | $0.14 | $4.05 | $49.28 |
| DeepSeek V2.5 | DeepSeek | $0.00014 | $0.14 | $4.20 | $51.10 |
| Gemini 2.0 Flash-Lite | $0.00015 | $0.15 | $4.50 | $54.75 | |
| Mistral Small 3 | Mistral AI | $0.00015 | $0.15 | $4.50 | $54.75 |
| Gemini 1.5 Flash | $0.00015 | $0.15 | $4.50 | $54.75 | |
| Gemini 2.5 Flash-Lite | $0.0002 | $0.20 | $6.00 | $73.00 | |
| Gemini 2.0 Flash | $0.0002 | $0.20 | $6.00 | $73.00 | |
| Llama 3.3 70B | Meta | $0.0002 | $0.20 | $6.00 | $73.00 |
| Qwen 3 32B | Alibaba | $0.0002 | $0.20 | $6.00 | $73.00 |
| Jamba 1.5 Mini | AI21 Labs | $0.0002 | $0.20 | $6.00 | $73.00 |
| Grok 3 mini | xAI | $0.00025 | $0.25 | $7.50 | $91.25 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | $0.00025 | $0.25 | $7.50 | $91.25 |
| Llama 3.1 70B Turbo | Meta | $0.00027 | $0.27 | $8.10 | $98.55 |
| Llama 4 Scout | Meta | $0.000295 | $0.30 | $8.85 | $107.68 |
| GPT-4o mini | OpenAI | $0.0003 | $0.30 | $9.00 | $109.50 |
| Command R | Cohere | $0.0003 | $0.30 | $9.00 | $109.50 |
| Qwen 2.5 72B | Alibaba | $0.0003 | $0.30 | $9.00 | $109.50 |
| Llama 3.3 70B (Cerebras) | Cerebras | $0.0003 | $0.30 | $9.00 | $109.50 |
| Llama 3.3 70B (Groq) | Groq | $0.000395 | $0.40 | $11.85 | $144.18 |
| Llama 4 Maverick | Meta | $0.000425 | $0.43 | $12.75 | $155.13 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | $0.00044 | $0.44 | $13.20 | $160.60 |
| Llama 3.1 70B | Meta | $0.00044 | $0.44 | $13.20 | $160.60 |
| Qwen 3 72B | Alibaba | $0.00045 | $0.45 | $13.50 | $164.25 |
| QwQ 32B | Alibaba | $0.00045 | $0.45 | $13.50 | $164.25 |
| Llama 3.2 90B Vision | Meta | $0.00045 | $0.45 | $13.50 | $164.25 |
| Codestral | Mistral AI | $0.00045 | $0.45 | $13.50 | $164.25 |
| FireFunction v2 | Fireworks AI | $0.00045 | $0.45 | $13.50 | $164.25 |
| Sonar | Perplexity | $0.0005 | $0.50 | $15.00 | $182.50 |
| DeepSeek V3 | DeepSeek | $0.00055 | $0.55 | $16.50 | $200.75 |
| Mixtral 8×22B | Mistral AI | $0.0006 | $0.60 | $18.00 | $219.00 |
| Sky-T1 32B | NovaSky | $0.0006 | $0.60 | $18.00 | $219.00 |
| GPT-5 mini | OpenAI | $0.001 | $1.00 | $30.00 | $365.00 |
| DeepSeek R1 | DeepSeek | $0.001095 | $1.10 | $32.85 | $399.68 |
| Gemini 2.5 Flash | $0.00125 | $1.25 | $37.50 | $456.25 | |
| Gemini 3 Flash | $0.0015 | $1.50 | $45.00 | $547.50 | |
| Llama 3.1 405B | Meta | $0.0015 | $1.50 | $45.00 | $547.50 |
| Yi Large | 01.AI | $0.0015 | $1.50 | $45.00 | $547.50 |
| Amazon Nova Pro | Amazon | $0.0016 | $1.60 | $48.00 | $584.00 |
| Claude Haiku 3.5 | Anthropic | $0.002 | $2.00 | $60.00 | $730.00 |
| Claude 3.5 Haiku | Anthropic | $0.002 | $2.00 | $60.00 | $730.00 |
| Grok 3 Mini Fast | xAI | $0.002 | $2.00 | $60.00 | $730.00 |
| o4-mini | OpenAI | $0.0022 | $2.20 | $66.00 | $803.00 |
| o1-mini | OpenAI | $0.0022 | $2.20 | $66.00 | $803.00 |
| o3-mini | OpenAI | $0.0022 | $2.20 | $66.00 | $803.00 |
| Claude Haiku 4.5 | Anthropic | $0.0025 | $2.50 | $75.00 | $912.50 |
| Gemini 1.5 Pro | $0.0025 | $2.50 | $75.00 | $912.50 | |
| Mistral Large 2 | Mistral AI | $0.003 | $3.00 | $90.00 | $1,095.00 |
| Mistral Large 2411 | Mistral AI | $0.003 | $3.00 | $90.00 | $1,095.00 |
| Jamba 1.5 Large | AI21 Labs | $0.004 | $4.00 | $120.00 | $1,460.00 |
| GPT-4o | OpenAI | $0.005 | $5.00 | $150.00 | $1,825.00 |
| Gemini 2.5 Pro | $0.005 | $5.00 | $150.00 | $1,825.00 | |
| Grok 2 | xAI | $0.005 | $5.00 | $150.00 | $1,825.00 |
| Command R+ | Cohere | $0.005 | $5.00 | $150.00 | $1,825.00 |
| Palmyra X4 | Writer | $0.005 | $5.00 | $150.00 | $1,825.00 |
| Inflection 3 Pi | Inflection | $0.005 | $5.00 | $150.00 | $1,825.00 |
| GPT-5.4 | OpenAI | $0.0075 | $7.50 | $225.00 | $2,737.50 |
| Claude Sonnet 5 | Anthropic | $0.0075 | $7.50 | $225.00 | $2,737.50 |
| Claude Sonnet 4 | Anthropic | $0.0075 | $7.50 | $225.00 | $2,737.50 |
| Gemini 3 Pro | $0.0075 | $7.50 | $225.00 | $2,737.50 | |
| Grok 3 | xAI | $0.0075 | $7.50 | $225.00 | $2,737.50 |
| Sonar Pro | Perplexity | $0.0075 | $7.50 | $225.00 | $2,737.50 |
| Palmyra X5 | Writer | $0.01 | $12.00 | $360.00 | $4,380.00 |
| Claude Opus 5 | Anthropic | $0.01 | $12.50 | $375.00 | $4,562.50 |
| o3 | OpenAI | $0.02 | $20.00 | $600.00 | $7,300.00 |
| GPT-Image-1 | OpenAI | $0.02 | $20.00 | $600.00 | $7,300.00 |
| o1 | OpenAI | $0.03 | $30.00 | $900.00 | $10,950.00 |
| Claude Opus 4 | Anthropic | $0.04 | $37.50 | $1,125.00 | $13,687.50 |
| GPT-4.5 | OpenAI | $0.08 | $75.00 | $2,250.00 | $27,375.00 |
How to Use
Paste your prompt
Type or paste your system message, user prompt, or full conversation. Token count updates instantly in your browser — nothing is uploaded.
Set expected output
Enter the average output tokens the model will generate. Default is 500 tokens (≈ 375 words). Adjust to match your typical response length.
Filter, sort & compare
Use provider tabs to focus on one vendor. Toggle "Cheapest first" to rank models by cost. Enable "Batch API" to see 50%-off async pricing. Toggle caching to simulate cached input.
Project production cost
Enter your daily API call volume — the scale table shows projected daily and monthly spend (×30 days) per model so you can budget before committing.
How LLM API pricing works
Every major LLM provider bills on a pay-per-token model. You pay separately for input tokens (your prompt, system message, conversation history) and output tokens (the model's generated response). Output tokens are typically priced 3–10× higher than input because generating each token requires a full sequential forward pass — whereas input tokens are read in parallel in a single pass.
The formula: Total cost = (input tokens × input rate) + (output tokens × output rate). Rates are expressed per million tokens ($/1M). A 1,000-token prompt at $3/1M costs $0.003 — small per call, but 10,000 calls/day at $0.01 each is $3,000/month. Use the scale table above to see how costs compound at your volume.
Looking to just count tokens without cost analysis? Try our AI Token Calculator — it shows words, characters, and a token-to-cost reference table alongside the count.
2026 LLM Pricing Reference
Current pricing per 1 million tokens. Batch pricing is 50% of standard for async workloads. Always verify with your provider before committing to a budget.
| Model | Provider | Context | Input / 1M | Output / 1M | Out/In Ratio | Type |
|---|---|---|---|---|---|---|
| GPT-5.4 | OpenAI | 1M | $2.50 | $15.00 | 6.0× | Standard |
| GPT-5.4 (Batch) | OpenAI | 1M | $1.25 | $7.50 | 6.0× | Batch |
| GPT-5 mini | OpenAI | 400K | $0.250 | $2.00 | 8.0× | Standard |
| GPT-5 mini (Batch) | OpenAI | 400K | $0.125 | $1.00 | 8.0× | Batch |
| GPT-4o | OpenAI | 128K | $2.50 | $10.00 | 4.0× | Standard |
| GPT-4o mini | OpenAI | 128K | $0.150 | $0.600 | 4.0× | Standard |
| o3 | OpenAI | 200K | $10.00 | $40.00 | 4.0× | Standard |
| o4-mini | OpenAI | 200K | $1.10 | $4.40 | 4.0× | Standard |
| Claude Opus 5 | Anthropic | 200K | $5.00 | $25.00 | 5.0× | Standard |
| Claude Sonnet 5 | Anthropic | 200K | $3.00 | $15.00 | 5.0× | Standard |
| Claude Haiku 4.5 | Anthropic | 200K | $1.00 | $5.00 | 5.0× | Standard |
| Claude Haiku 4.5 (Batch) | Anthropic | 200K | $0.500 | $2.50 | 5.0× | Batch |
| Claude Haiku 3.5 | Anthropic | 200K | $0.800 | $4.00 | 5.0× | Standard |
| Gemini 3 Flash | 1M | $0.500 | $3.00 | 6.0× | Standard | |
| Gemini 2.5 Pro | 2M | $1.25 | $10.00 | 8.0× | Standard | |
| Gemini 2.5 Flash | 1M | $0.300 | $2.50 | 8.3× | Standard | |
| Gemini 2.5 Flash-Lite | 1M | $0.100 | $0.400 | 4.0× | Standard | |
| Gemini 2.0 Flash | 1M | $0.100 | $0.400 | 4.0× | Standard | |
| Gemini 2.0 Flash-Lite | 1M | $0.075 | $0.300 | 4.0× | Standard | |
| Llama 4 Maverick | Meta | 1M | $0.270 | $0.850 | 3.1× | Standard |
| Llama 4 Scout | Meta | 128K | $0.180 | $0.590 | 3.3× | Standard |
| Llama 3.3 70B | Meta | 128K | $0.230 | $0.400 | 1.7× | Standard |
| Llama 3.1 405B | Meta | 128K | $1.00 | $3.00 | 3.0× | Standard |
| Llama 3.1 8B | Meta | 128K | $0.050 | $0.080 | 1.6× | Standard |
| DeepSeek V3 | DeepSeek | 64K | $0.270 | $1.10 | 4.1× | Standard |
| DeepSeek R1 | DeepSeek | 64K | $0.550 | $2.19 | 4.0× | Standard |
| Mistral Large 2 | Mistral AI | 128K | $2.00 | $6.00 | 3.0× | Standard |
| Mistral Small 3 | Mistral AI | 128K | $0.100 | $0.300 | 3.0× | Standard |
| Grok 3 mini | xAI | 131K | $0.300 | $0.500 | 1.7× | Standard |
| Command R | Cohere | 128K | $0.150 | $0.600 | 4.0× | Standard |
| Embed v3 English | Cohere | 512K | $0.100 | $0.000 | 0.0× | Standard |
| Rerank v3 | Cohere | 4K | $2.00 | $0.000 | 0.0× | Standard |
| text-embedding-3-large | OpenAI | 8K | $0.130 | $0.000 | 0.0× | Standard |
| Gemma 3 12B | 128K | $0.120 | $0.120 | 1.0× | Standard | |
| Ministral 8B | Mistral AI | 128K | $0.100 | $0.100 | 1.0× | Standard |
| Ministral 3B | Mistral AI | 128K | $0.040 | $0.040 | 1.0× | Standard |
| Mistral Embed | Mistral AI | 8K | $0.100 | $0.000 | 0.0× | Standard |
| Qwen 3 72B | Alibaba | 128K | $0.900 | $0.900 | 1.0× | Standard |
| Qwen 3 32B | Alibaba | 128K | $0.400 | $0.400 | 1.0× | Standard |
| QwQ 32B | Alibaba | 128K | $0.900 | $0.900 | 1.0× | Standard |
| Qwen 3 14B | Alibaba | 128K | $0.150 | $0.150 | 1.0× | Standard |
| Llama 3.2 3B | Meta | 128K | $0.060 | $0.060 | 1.0× | Standard |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 128K | $0.880 | $0.880 | 1.0× | Standard |
| o1 | OpenAI | 200K | $15.00 | $60.00 | 4.0× | Standard |
| o1-mini | OpenAI | 128K | $1.10 | $4.40 | 4.0× | Standard |
| o3-mini | OpenAI | 200K | $1.10 | $4.40 | 4.0× | Standard |
| text-embedding-3-small | OpenAI | 8K | $0.020 | $0.000 | 0.0× | Standard |
| Claude Opus 4 | Anthropic | 200K | $15.00 | $75.00 | 5.0× | Standard |
| Claude Sonnet 4 | Anthropic | 200K | $3.00 | $15.00 | 5.0× | Standard |
| Claude 3.5 Haiku | Anthropic | 200K | $0.800 | $4.00 | 5.0× | Standard |
| Gemini 3 Pro | 2M | $2.50 | $15.00 | 6.0× | Standard | |
| Gemma 3 27B | 128K | $0.270 | $0.270 | 1.0× | Standard | |
| Gemma 3 4B | 128K | $0.040 | $0.040 | 1.0× | Standard | |
| Llama 3.1 70B | Meta | 128K | $0.680 | $0.880 | 1.3× | Standard |
| Llama 3.2 11B Vision | Meta | 128K | $0.180 | $0.180 | 1.0× | Standard |
| Llama 3.2 90B Vision | Meta | 128K | $0.900 | $0.900 | 1.0× | Standard |
| Mistral Large 2411 | Mistral AI | 128K | $2.00 | $6.00 | 3.0× | Standard |
| Mixtral 8×22B | Mistral AI | 64K | $1.20 | $1.20 | 1.0× | Standard |
| Mixtral 8×7B | Mistral AI | 32K | $0.240 | $0.240 | 1.0× | Standard |
| Codestral | Mistral AI | 256K | $0.300 | $0.900 | 3.0× | Standard |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 128K | $0.500 | $0.500 | 1.0× | Standard |
| DeepSeek V2.5 | DeepSeek | 64K | $0.140 | $0.280 | 2.0× | Standard |
| Grok 3 | xAI | 131K | $3.00 | $15.00 | 5.0× | Standard |
| Grok 2 | xAI | 131K | $2.00 | $10.00 | 5.0× | Standard |
| Command R+ | Cohere | 128K | $2.50 | $10.00 | 4.0× | Standard |
| Embed English v3 | Cohere | 4K | $0.100 | $0.000 | 0.0× | Standard |
| Qwen 3 7B | Alibaba | 128K | $0.060 | $0.060 | 1.0× | Standard |
| Qwen 2.5 72B | Alibaba | 128K | $0.600 | $0.600 | 1.0× | Standard |
| Amazon Nova Pro | Amazon | 300K | $0.800 | $3.20 | 4.0× | Standard |
| Amazon Nova Lite | Amazon | 300K | $0.060 | $0.240 | 4.0× | Standard |
| Amazon Nova Micro | Amazon | 128K | $0.035 | $0.140 | 4.0× | Standard |
| Jamba 1.5 Large | AI21 Labs | 256K | $2.00 | $8.00 | 4.0× | Standard |
| Jamba 1.5 Mini | AI21 Labs | 256K | $0.200 | $0.400 | 2.0× | Standard |
| Sonar Pro | Perplexity | 200K | $3.00 | $15.00 | 5.0× | Standard |
| Sonar | Perplexity | 128K | $1.00 | $1.00 | 1.0× | Standard |
| Palmyra X5 | Writer | 1M | $6.00 | $24.00 | 4.0× | Standard |
| Palmyra X4 | Writer | 128K | $2.50 | $10.00 | 4.0× | Standard |
| Snowflake Arctic | Snowflake | 4K | $0.006 | $0.018 | 3.0× | Standard |
| Yi Large | 01.AI | 32K | $3.00 | $3.00 | 1.0× | Standard |
| FireFunction v2 | Fireworks AI | 8K | $0.900 | $0.900 | 1.0× | Standard |
| Inflection 3 Pi | Inflection | 8K | $2.50 | $10.00 | 4.0× | Standard |
| Llama 3.3 70B (Groq) | Groq | 128K | $0.590 | $0.790 | 1.3× | Standard |
| Llama 3.1 8B (Groq) | Groq | 128K | $0.050 | $0.080 | 1.6× | Standard |
| Mixtral 8×7B (Groq) | Groq | 32K | $0.240 | $0.240 | 1.0× | Standard |
| Gemma 3 27B (Groq) | Groq | 128K | $0.270 | $0.270 | 1.0× | Standard |
| Qwen 2.5 Coder 32B | Alibaba | 128K | $0.200 | $0.200 | 1.0× | Standard |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 128K | $0.150 | $0.150 | 1.0× | Standard |
| Phi-4 | Microsoft | 16K | $0.070 | $0.140 | 2.0× | Standard |
| Phi-4 Mini | Microsoft | 16K | $0.025 | $0.050 | 2.0× | Standard |
| StableLM 2 12B | Stability AI | 4K | $0.100 | $0.100 | 1.0× | Standard |
| Sky-T1 32B | NovaSky | 32K | $1.20 | $1.20 | 1.0× | Standard |
| GPT-4o (Batch) | OpenAI | 128K | $1.25 | $5.00 | 4.0× | Batch |
| GPT-4o mini (Batch) | OpenAI | 128K | $0.075 | $0.300 | 4.0× | Batch |
| Claude Sonnet 5 (Batch) | Anthropic | 200K | $1.50 | $7.50 | 5.0× | Batch |
| Gemini 2.5 Flash (Batch) | 1M | $0.150 | $1.25 | 8.3× | Batch | |
| GPT-4.5 | OpenAI | 128K | $75.00 | $150.00 | 2.0× | Standard |
| GPT-Image-1 | OpenAI | 128K | $5.00 | $40.00 | 8.0× | Standard |
| Gemini 1.5 Pro | 2M | $1.25 | $5.00 | 4.0× | Standard | |
| Gemini 1.5 Flash | 1M | $0.075 | $0.300 | 4.0× | Standard | |
| Llama 3.1 70B Turbo | Meta | 128K | $0.540 | $0.540 | 1.0× | Standard |
| Mistral 7B | Mistral AI | 32K | $0.100 | $0.100 | 1.0× | Standard |
| Pixtral 12B | Mistral AI | 128K | $0.150 | $0.150 | 1.0× | Standard |
| Grok 3 Mini Fast | xAI | 131K | $0.600 | $4.00 | 6.7× | Standard |
| Llama 3.3 70B (Cerebras) | Cerebras | 128K | $0.600 | $0.600 | 1.0× | Standard |
| Llama 3.1 8B (Cerebras) | Cerebras | 128K | $0.100 | $0.100 | 1.0× | Standard |
Cost Optimization Strategies
Practical techniques to reduce API spend without sacrificing output quality.
| Strategy | Typical Savings | Notes |
|---|---|---|
| Prompt Caching | Up to 90% on input | Cache static system prompts or documents. Supported by OpenAI and Anthropic. Toggle in calculator above. |
| Batch API | 50% overall | OpenAI Batch and Anthropic Batches process requests within 24 hours at 50% discount. Enable in calculator above. |
| Shorten system prompt | 10–40% on input | Every request sends the system prompt. A 500-token reduction × 10K calls/day = 5M fewer input tokens/day. |
| Limit max_tokens | 20–60% on output | Set max_tokens to the minimum your use case needs. Output is typically your most expensive line item. |
| Route to smaller model | 80–95% overall | Use Haiku / Flash-Lite for classification, extraction, and summarization. Reserve Opus / Pro for complex reasoning. |
| Compress conversation history | 30–70% on input | Summarize older turns instead of appending the full chat history to every new request. |
Savings percentages are estimates. Actual results depend on your specific prompts, output length, and provider plan. Always measure with your real workload.
FAQ
Have more questions? Contact us
Last verified: 2026-08-20 · methodology · data sources