Cache-aware math

Cached input pricing comparison

RAG and agent apps reuse long prompts — cached-input pricing often matters more than output rates. Adjust the cache slider to see which models reward retrieval-heavy workloads.

Calculate my API cost

Related: Llm Pricing For Rag · Ai Api Cost Per Month

Calculator

Start with your workload

Step 1

Review your assumptions

These are estimates, not measured usage. Edit the inputs or choose a preset.

Synced Oct 8, 2026
Edit assumptions · 1,000 requests/month · 1,000 tokens/request · 75% input · 65% of input cached

1K requests / mo

1K tokens avg

Step 2

Lowest estimated cost

Top 5 models for this workload (unreviewed and conditional tariffs excluded)

General intelligence is not a task-specific quality guarantee. Test shortlisted models on your own examples. Source reviewed — exact offer checked on its provider page. Needs review — missing, stale or conflicting evidence; excluded from default recommendations. Estimates cover input/cache-read/output tokens only. Cache writes, storage, tools, search fees, taxes and unconfigured billing conditions are additional. Include billed reasoning tokens in your output estimate.

Selected model: cost breakdown

Estimate

Lowest estimated cost

AWS Bedrock · Amazon Nova Micro (standard)

From monthly cost (USD)

From $0.05

Based on: 1K requests · 750 input tokens/request · 250 output tokens/request · 65% of input from cache

Input $0.01 Cached 0.49M tok Output $0.04
Total tokens1.00M
Per 1k requests$0.05
Rank#1

Token costs only. Tool calls, search, storage, regional pricing, commitment discounts, and taxes may be excluded.

Explore all priced models

Step 3

All priced models

Jump to any of 62 models for this workload

Pick a model to focus the estimate above, or browse the ranked list.

How costs are calculated

Prices from ai-provider-pricing-validated.json, validated Thursday 8th October 2026. Confirm on official provider pages before billing decisions. Dataset contains 95 offer records, including archived entries and offers that require additional billing inputs or price evidence. Only reviewed offers appear in this calculator. See excluded offers and reasons

Embed this calculator

Embed

Want this calculator on your site?

Copy the iframe snippet below and paste it into any page, doc, or WordPress Custom HTML block.

<iframe src="https://modelcostcomparison.com/embed/ai-api-pricing-calculator?ref=topic-cached-input-pricing-comparison" width="100%" height="980" style="border:0;border-radius:12px;overflow:hidden" loading="lazy" referrerpolicy="strict-origin-when-cross-origin" title="AI API Pricing Calculator by Model Cost Comparison"></iframe>

Model Cost Comparison · Built by Lazige · Methodology

How we calculate cost

Monthly estimate = (input tokens × input $/MTok) + (cached tokens × cached $/MTok when published) + (output tokens × output $/MTok), scaled to your message volume. See the methodology for validation sources and update cadence.

Topic scenarios

Cached input pricing comparison

RAG and agent apps reuse long prompts — cached-input pricing often matters more than output rates. Adjust the cache slider to see which models reward retrieval-heavy workloads.

“Start with a preset that matches product shape, then adjust volume after a week of real traffic.”
Planning baselinerag preset
“Compare supported cache scenarios for RAG and agents, then validate the assumptions against actual usage.”
Cache-sensitive appsAdvanced token assumptions
“Use the ranked list to shortlist two to three models, then validate quality on your eval set.”
Model shortlistLowest estimated cost cards
“Re-run when traffic grows or vendors publish new tiers; check the dataset validation date and official rates.”
Quarterly reviewCheck source freshness