Cost Calculators → Cost Scenarios

Semantic Search — Embedding Cost at Scale

Embedding cost for building and querying a semantic search index over 1M documents.

Low estimate
$500/mo
Mid estimate
$750/mo
High estimate
$1,200/mo

Assumptions

  • Initial embedding: 1M documents × 500 tokens avg = 500M tokens
  • Re-embedding: 5% of corpus updated monthly = 25M tokens
  • Query embedding: 100,000 queries/month × 50 tokens = 5M tokens
  • Vector DB: Pinecone Serverless, 1M 1536-dim vectors
  • LLM generation: 100k queries × 3k context × $3/M + 100k × 400 out × $15/M

Base model: Claude Sonnet 5 — $3/M input · $15/M output · 2026-08-19

Cost breakdown by component

LLM generation (dominant)82%
Vector DB storage/queries12%
Query embedding4%
Re-embedding2%

How to cut this cost by 50%+

Reduce retrieved chunks
Retrieving 5 chunks instead of 10 halves context cost — test quality first
30-40% savings
Cache frequently-queried docs
Common queries repeat — cache the RAG context prefix
20-30% savings
Calculate for your exact workload

These are estimates based on typical patterns. Enter your specific token counts and volume for a precise calculation.

Open calculator →prompt cost calculatorprompt caching calculator

Frequently Asked Questions

Is embedding or generation more expensive in a RAG system?

Generation is almost always the dominant cost — typically 80-90% of total RAG cost. Embedding 1M documents with OpenAI text-embedding-3-small costs ~$20 one-time. Monthly LLM generation on those same documents at 100k queries/month costs 10-50× more. Optimizing generation (fewer chunks, caching) has far more impact than embedding cost.

Pricing verified 2026-08-20. All figures are estimates with stated assumptions. See methodology.