Skip to content
CAUSAL LABS
Research · verified 2026-08-09

What Claude, GPT and Gemini actually cost to run

Every rate below was checked against the vendor's own current pricing page, not recalled or aggregated. By Causal Labs engineering team.

Methodology · Every rate in this index was checked directly against the vendor's own current pricing documentation — Anthropic, OpenAI and Google's official pages, not a third-party aggregator or recalled from memory. Each entry links to the exact source page. Workload costs are computed from stated token counts using each vendor's published per-token rate — no rounding beyond standard currency display. Model pricing changes frequently; treat the verified-on date as an expiry, not a guarantee, and confirm against the source before budgeting a production workload off these numbers.

Rates, per million tokens

VendorModelInputOutputContextSource
AnthropicClaude Opus 5$5.00$25.001M tokensVendor docs
AnthropicClaude Sonnet 5$2.00$10.001M tokensVendor docs
AnthropicClaude Haiku 4.5$1.00$5.001M tokensVendor docs
OpenAIGPT-5.6 Terra$2.00$12.00Not confirmedVendor docs
OpenAIGPT-5.6 Luna$0.20$1.20Not confirmedVendor docs
OpenAIGPT-4o mini$0.15$0.60Not confirmedVendor docs
GoogleGemini 3.6 Flash$1.50$7.50Not confirmedVendor docs
GoogleGemini 3.5 Flash-Lite$0.30$2.50Not confirmedVendor docs

Claude Sonnet 5: Introductory rate through 2026-08-31; standard rate afterward is $3 / $15 per MTok.

GPT-5.6 Terra: Balanced everyday-work tier, released 2026-07-09.

GPT-5.6 Luna: Fastest, lowest-cost GPT-5.6 tier.

Gemini 3.6 Flash: Current Flash-tier default, shipped 2026-07-21.

Real workload costs

What each model actually costs per month against three common automation patterns — not a single generic “cost per 1K tokens” that hides how volume changes the picture.

10,000 invoice/receipt extractions / month

~1,500 input tokens (document + prompt), ~300 output tokens (structured JSON) each.

ModelMonthly cost
GPT-4o mini$4.05
GPT-5.6 Luna$6.60
Gemini 3.5 Flash-Lite$12.00
Claude Haiku 4.5$30.00
Gemini 3.6 Flash$45.00
Claude Sonnet 5$60.00
GPT-5.6 Terra$66.00
Claude Opus 5$150.00

5,000 customer-support draft replies / month

~800 input tokens (ticket + context), ~250 output tokens (draft reply) each.

ModelMonthly cost
GPT-4o mini$1.35
GPT-5.6 Luna$2.30
Gemini 3.5 Flash-Lite$4.33
Claude Haiku 4.5$10.25
Gemini 3.6 Flash$15.38
Claude Sonnet 5$20.50
GPT-5.6 Terra$23.00
Claude Opus 5$51.25

50,000 lead-enrichment lookups / month

High-volume, low-complexity: ~400 input tokens, ~100 output tokens each.

ModelMonthly cost
GPT-4o mini$6.00
GPT-5.6 Luna$10.00
Gemini 3.5 Flash-Lite$18.50
Claude Haiku 4.5$45.00
Gemini 3.6 Flash$67.50
Claude Sonnet 5$90.00
GPT-5.6 Terra$100.00
Claude Opus 5$225.00

Model your own workload

Enter your own request volume and token counts in the free calculator — it uses these same editable rates.

Open the AI API Cost Calculator