Instructions

How to use the AI API cost calculator

Compare 62 models across 11 providers in four steps. Set one workload, read the monthly estimate, rank every model, and tune cache-aware token math. Pricing reference updated Wednesday 7th October 2026.

Quick start (under 2 minutes)

  1. 01

    Set your workload

  2. 02

    Review lowest estimated cost

  3. 03

    Optional: focus a model

  4. 04

    Fine-tune token mix (optional)

Workload presets

Presets fill messages, tokens, and token-mix defaults. Click one on the calculator, then adjust if your traffic differs.

  • Support bot

    High volume, short replies

    50,000 / mo · 800 avg

    Customer support, FAQ bots, ticket triage

  • Code assistant

    Medium volume, long context

    5,000 / mo · 4,000 avg

    IDE copilots, PR review, refactors

  • RAG / search

    Retrieval-heavy prompts

    10,000 / mo · 2,500 avg

    Document Q&A, knowledge bases, search-augmented apps

  • Agent / tools

    Multi-step, mixed I/O

    2,000 / mo · 8,000 avg

    Tool-calling agents, workflows, multi-turn reasoning

01

Set your workload

Start with a preset or your own volume. Recommendations appear immediately — no model pick required.

  • Requests per month = API calls or chat turns (one user request + model reply).
  • Tokens per request = prompt + completion size. Use your analytics average, or start with a preset.
  • The Support bot preset loads by default so you see ranked costs on first paint.

02

Review lowest estimated cost

Scan the top recommendations ranked by monthly token cost for the same workload.

  • Up to five cards show name, provider, monthly estimate, and fit notes.
  • Unreviewed offers and conditional Batch/Flex/off-peak discounts are excluded from default recommendations.
  • Source reviewed means the exact API offer has current primary evidence. Needs review means it is not eligible for default recommendations.

03

Optional: focus a model

Select a recommendation (or a row from View all models) to inspect the estimate detail panel.

  • See assumptions (requests, tokens, input %, cache %) plus input/cached/output breakdown.
  • Clear focus to return to recommendations-only view with a From $X/mo headline.
  • Expand View all models anytime to search and filter the full ranked list.

04

Fine-tune token mix (optional)

Open Advanced token mix when RAG, agents, or long-context apps need a more accurate split.

  • Input vs output ratio — share of tokens sent as prompt/context vs model reply.
  • Cached input ratio — share of prompt tokens served from provider cache (lower rate when available).
  • Per-request token summary updates live so you can sanity-check the math.

Embed on your site

Add the calculator to docs, internal tools, or landing pages with the embed widget. It loads the same live pricing engine as the main site.

Compare providers and explore pricing

Already considering a provider? Compare its rates, explore alternatives or read a workload guide.

FAQ

Common questions about using the calculator

ready to compare?

Run your workload through 62 models now

Same baseline, every provider — updated Wednesday 7th October 2026.