Budget routing
Cheapest LLM API for your workload
The cheapest model depends on cache ratio, output length, and traffic — not headline list prices. Start with a preset, then scan the ranked table for the lowest monthly estimate.
Calculate my API costRelated: Llm Pricing For Rag · Openai vs Anthropic
Calculator
Start with your workload
Step 1
Review your assumptions
These are estimates, not measured usage. Edit the inputs or choose a preset.
Edit assumptions · 50,000 requests/month · 800 tokens/request · 82% input · 55% of input cached
50K requests / mo
800 tokens avg
Per request: 656 in · 144 out · 361 cached
Step 2
Lowest estimated cost
Top 5 models for this workload (unreviewed and conditional tariffs excluded)
General intelligence is not a task-specific quality guarantee. Test shortlisted models on your own examples. Source reviewed — exact offer checked on its provider page. Needs review — missing, stale or conflicting evidence; excluded from default recommendations. Estimates cover input/cache-read/output tokens only. Cache writes, storage, tools, search fees, taxes and unconfigured billing conditions are additional. Include billed reasoning tokens in your output estimate.
Selected model: cost breakdown
Estimate
Lowest estimated cost
AWS Bedrock · Amazon Nova Micro (standard)
From monthly cost (USD)
From $1.68
Based on: 50K requests · 656 input tokens/request · 144 output tokens/request · 55% of input from cache
Token costs only. Tool calls, search, storage, regional pricing, commitment discounts, and taxes may be excluded.
Explore all priced models
Step 3
All priced models
Jump to any of 62 models for this workload
Pick a model to focus the estimate above, or browse the ranked list.
How costs are calculated
Prices from ai-provider-pricing-validated.json, validated Wednesday 7th October 2026. Confirm on official provider pages before billing decisions. Dataset contains 95 offer records, including archived entries and offers that require additional billing inputs or price evidence. Only reviewed offers appear in this calculator. See excluded offers and reasons
Model Cost Comparison · Built by Lazige · Methodology
How we calculate cost
Monthly estimate = (input tokens × input $/MTok) + (cached tokens × cached $/MTok when published) + (output tokens × output $/MTok), scaled to your message volume. See the methodology for validation sources and update cadence.
Topic scenarios
Cheapest LLM API for your workload
The cheapest model depends on cache ratio, output length, and traffic — not headline list prices. Start with a preset, then scan the ranked table for the lowest monthly estimate.
“Start with a preset that matches product shape, then adjust volume after a week of real traffic.”
“Compare supported cache scenarios for RAG and agents, then validate the assumptions against actual usage.”
“Use the ranked list to shortlist two to three models, then validate quality on your eval set.”
“Re-run when traffic grows or vendors publish new tiers; check the dataset validation date and official rates.”