Cost Calculators → Cost Scenarios
Semantic Search — Embedding Cost at Scale
Embedding cost for building and querying a semantic search index over 1M documents.
Low estimate
$500/mo
Mid estimate
$750/mo
High estimate
$1,200/mo
Assumptions
- Initial embedding: 1M documents × 500 tokens avg = 500M tokens
- Re-embedding: 5% of corpus updated monthly = 25M tokens
- Query embedding: 100,000 queries/month × 50 tokens = 5M tokens
- Vector DB: Pinecone Serverless, 1M 1536-dim vectors
- LLM generation: 100k queries × 3k context × $3/M + 100k × 400 out × $15/M
Base model: Claude Sonnet 5 — $3/M input · $15/M output · 2026-08-19
Cost breakdown by component
LLM generation (dominant)82%
Vector DB storage/queries12%
Query embedding4%
Re-embedding2%
How to cut this cost by 50%+
Reduce retrieved chunks
Retrieving 5 chunks instead of 10 halves context cost — test quality first
30-40% savings
Cache frequently-queried docs
Common queries repeat — cache the RAG context prefix
20-30% savings
Calculate for your exact workload
These are estimates based on typical patterns. Enter your specific token counts and volume for a precise calculation.
Frequently Asked Questions
Is embedding or generation more expensive in a RAG system?
Generation is almost always the dominant cost — typically 80-90% of total RAG cost. Embedding 1M documents with OpenAI text-embedding-3-small costs ~$20 one-time. Monthly LLM generation on those same documents at 100k queries/month costs 10-50× more. Optimizing generation (fewer chunks, caching) has far more impact than embedding cost.
Pricing verified 2026-08-20. All figures are estimates with stated assumptions. See methodology.