What Claude, GPT and Gemini actually cost to run
Every rate below was checked against the vendor's own current pricing page, not recalled or aggregated. By Causal Labs engineering team.
Methodology · Every rate in this index was checked directly against the vendor's own current pricing documentation — Anthropic, OpenAI and Google's official pages, not a third-party aggregator or recalled from memory. Each entry links to the exact source page. Workload costs are computed from stated token counts using each vendor's published per-token rate — no rounding beyond standard currency display. Model pricing changes frequently; treat the verified-on date as an expiry, not a guarantee, and confirm against the source before budgeting a production workload off these numbers.
Rates, per million tokens
| Vendor | Model | Input | Output | Context | Source |
|---|---|---|---|---|---|
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | 1M tokens | Vendor docs |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1M tokens | Vendor docs |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 1M tokens | Vendor docs |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | Not confirmed | Vendor docs |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | Not confirmed | Vendor docs |
| OpenAI | GPT-4o mini | $0.15 | $0.60 | Not confirmed | Vendor docs |
| Gemini 3.6 Flash | $1.50 | $7.50 | Not confirmed | Vendor docs | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Not confirmed | Vendor docs |
Claude Sonnet 5: Introductory rate through 2026-08-31; standard rate afterward is $3 / $15 per MTok.
GPT-5.6 Terra: Balanced everyday-work tier, released 2026-07-09.
GPT-5.6 Luna: Fastest, lowest-cost GPT-5.6 tier.
Gemini 3.6 Flash: Current Flash-tier default, shipped 2026-07-21.
Real workload costs
What each model actually costs per month against three common automation patterns — not a single generic “cost per 1K tokens” that hides how volume changes the picture.
10,000 invoice/receipt extractions / month
~1,500 input tokens (document + prompt), ~300 output tokens (structured JSON) each.
| Model | Monthly cost |
|---|---|
| GPT-4o mini | $4.05 |
| GPT-5.6 Luna | $6.60 |
| Gemini 3.5 Flash-Lite | $12.00 |
| Claude Haiku 4.5 | $30.00 |
| Gemini 3.6 Flash | $45.00 |
| Claude Sonnet 5 | $60.00 |
| GPT-5.6 Terra | $66.00 |
| Claude Opus 5 | $150.00 |
5,000 customer-support draft replies / month
~800 input tokens (ticket + context), ~250 output tokens (draft reply) each.
| Model | Monthly cost |
|---|---|
| GPT-4o mini | $1.35 |
| GPT-5.6 Luna | $2.30 |
| Gemini 3.5 Flash-Lite | $4.33 |
| Claude Haiku 4.5 | $10.25 |
| Gemini 3.6 Flash | $15.38 |
| Claude Sonnet 5 | $20.50 |
| GPT-5.6 Terra | $23.00 |
| Claude Opus 5 | $51.25 |
50,000 lead-enrichment lookups / month
High-volume, low-complexity: ~400 input tokens, ~100 output tokens each.
| Model | Monthly cost |
|---|---|
| GPT-4o mini | $6.00 |
| GPT-5.6 Luna | $10.00 |
| Gemini 3.5 Flash-Lite | $18.50 |
| Claude Haiku 4.5 | $45.00 |
| Gemini 3.6 Flash | $67.50 |
| Claude Sonnet 5 | $90.00 |
| GPT-5.6 Terra | $100.00 |
| Claude Opus 5 | $225.00 |
Model your own workload
Enter your own request volume and token counts in the free calculator — it uses these same editable rates.
Open the AI API Cost Calculator