Explore AI
What it actually costs
Live prices for 444 models, translated into something you can picture: one long document summarised.
Prices fetched 2026-09-22 today
“Task” = a 10,000-token document in, a 1,000-token answer out. That is one long PDF summarised, or a substantial code review.
| Model | Context | In /M | Out /M | One task |
|---|---|---|---|---|
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai | 262k | free | free | free |
| Nex AGI: Nex-N2.5-Mini (free)nex-agi | 262k | free | free | free |
| Nex AGI: Nex-N2.5-Pro (free)nex-agi | 262k | free | free | free |
| inclusionAI: Ling 3.0 Flash Sante (free)inclusionai | 262k | free | free | free |
| inclusionAI: Ling 3.0 Flash Fin (free)inclusionai | 262k | free | free | free |
| Qwen: Qwen3.8 27B (free)qwen | 262k | free | free | free |
| Dots Studio: Dots3-Note Preview (free)dots-studio | 512k | free | free | free |
| LiquidAI: LFM2.5-2.6B (free)liquid | 66k | free | free | free |
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia | 1000k | free | free | free |
| Thinking Machines: Inkling Small (free)thinkingmachines | 1049k | free | free | free |
| Poolside: Laguna S 2.1 (free)poolside | 262k | free | free | free |
| Thinking Machines: Inkling (free)thinkingmachines | 1049k | free | free | free |
| Poolside: Laguna XS 2.1 (free)poolside | 262k | free | free | free |
| Cohere: North Mini Code (free)cohere | 256k | free | free | free |
| Z.ai: GLM 5.2 (free)z-ai | 33k | free | free | free |
| NVIDIA: Nemotron 3.5 Content Safety (free)nvidia | 128k | free | free | free |
| NVIDIA: Nemotron 3 Ultra (free)nvidia | 1000k | free | free | free |
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia | 256k | free | free | free |
| Google: Gemma 4 26B A4B (free)google | 262k | free | free | free |
| Google: Gemma 4 31B (free)google | 262k | free | free | free |
| Google: Lyria 3 Pro Previewgoogle | 1049k | free | free | free |
| Google: Lyria 3 Clip Previewgoogle | 1049k | free | free | free |
| NVIDIA: Nemotron 3 Super (free)nvidia | 262k | free | free | free |
| Free Models Routeropenrouter | 200k | free | free | free |
| IBM: Granite 4.0 Microibm-granite | 131k | $0.02 | $0.11 | $0.0003 |
| Sao10K: Llama 3 8B Lunarissao10k | 8k | $0.04 | $0.05 | $0.0004 |
| OpenAI: GPT-4.1 Nano (batch)openai | 1048k | $0.05 | $0.20 | $0.0007 |
| IBM: Granite 4.2 8Bibm-granite | 131k | $0.06 | $0.25 | $0.0008 |
| Qwen: Qwen3 Coder 30B A3B Instructqwen | 262k | $0.07 | $0.28 | $0.0010 |
| Z.ai: GLM 5.3 Flash (batch)z-ai | 1049k | $0.07 | $0.25 | $0.0010 |
| Qwen: Qwen3 32Bqwen | 131k | $0.08 | $0.28 | $0.0011 |
| Google: Gemma 4 26B A4B google | 262k | $0.09 | $0.30 | $0.0012 |
| Qwen: Qwen2.5 7B Instructqwen | 33k | $0.10 | $0.20 | $0.0012 |
| DeepSeek: DeepSeek V4 Flash Latest~deepseek | 1311k | $0.03 | $1.00 | $0.0013 |
| Mistral: Voxtral Small 24B 2507mistralai | 33k | $0.10 | $0.30 | $0.0013 |
| OpenAI: GPT-5.6 Luna (batch)openai | 1050k | $0.10 | $0.60 | $0.0016 |
| Tencent: Hy3tencent | 262k | $0.13 | $0.53 | $0.0018 |
| Qwen: Qwen3 Coder Nextqwen | 262k | $0.12 | $0.80 | $0.0020 |
| OpenAI: gpt-oss-120bopenai | 131k | $0.15 | $0.60 | $0.0021 |
| Tencent: Hy3 previewtencent | 262k | $0.18 | $0.60 | $0.0024 |
| Google: Gemini 3.5 Flash Lite (batch)google | 1049k | $0.15 | $1.25 | $0.0027 |
| OpenAI: GPT Luna Latest~openai | 1050k | $0.20 | $1.20 | $0.0032 |
| Inception: Mercury 2inception | 128k | $0.25 | $0.75 | $0.0033 |
| Qwen: Qwen Plus 0728qwen | 1000k | $0.26 | $0.78 | $0.0034 |
| Qwen: Qwen3 VL 235B A22B Instructqwen | 262k | $0.21 | $1.90 | $0.0040 |
| Qwen2.5 72B Instructqwen | 33k | $0.36 | $0.40 | $0.0040 |
| MiniMax: MiniMax M2.1minimax | 205k | $0.30 | $1.20 | $0.0042 |
| Qwen: Qwen3 VL 30B A3B Thinkingqwen | 262k | $0.20 | $2.40 | $0.0044 |
| Meta: Llama 3.1 70B Instructmeta-llama | 131k | $0.40 | $0.40 | $0.0044 |
| OpenAI: GPT-5 Miniopenai | 400k | $0.25 | $2.00 | $0.0045 |
| Google: Gemini 3.5 Flash Litegoogle | 1049k | $0.30 | $2.50 | $0.0055 |
| Google: Gemini 2.5 Flashgoogle | 1049k | $0.30 | $2.50 | $0.0055 |
| Thinking Machines: Inkling Smallthinkingmachines | 1049k | $0.45 | $1.20 | $0.0057 |
| Mistral: Devstral 2 2512mistralai | 262k | $0.40 | $2.00 | $0.0060 |
| OpenAI: o3 Mini (batch)openai | 200k | $0.55 | $2.20 | $0.0077 |
| Google: Gemini 3 Flash Previewgoogle | 1049k | $0.50 | $3.00 | $0.0080 |
| MoonshotAI: Kimi K2 0905moonshotai | 262k | $0.60 | $2.50 | $0.0085 |
| DeepSeek: DeepSeek V4 Pro 0813 (batch)deepseek | 1049k | $0.66 | $1.98 | $0.0086 |
| Sao10K: Llama 3.1 Euryale 70B v2.2sao10k | 131k | $0.85 | $0.85 | $0.0094 |
| AionLabs: Aion-RP 1.0 (8B)aion-labs | 33k | $0.80 | $1.60 | $0.0096 |
| MoonshotAI: Kimi K2.7 Codemoonshotai | 262k | $0.71 | $3.30 | $0.01 |
| Google: Gemini 2.5 Pro (batch)google | 1049k | $0.63 | $5.00 | $0.01 |
| OpenAI: GPT Mini Latest~openai | 400k | $0.75 | $4.50 | $0.01 |
| SpaceXAI: Grok 4.3 (batch)x-ai | 1000k | $1.00 | $2.00 | $0.01 |
| OpenAI: GPT-4.1 (batch)openai | 1048k | $1.00 | $4.00 | $0.01 |
| Thinking Machines: Inklingthinkingmachines | 1049k | $1.00 | $4.05 | $0.01 |
| SpaceXAI: Grok 4.3x-ai | 1000k | $1.25 | $2.50 | $0.02 |
| OpenAI: o3 Mini Highopenai | 200k | $1.10 | $4.40 | $0.02 |
| OpenAI: GPT-5openai | 400k | $1.25 | $10.00 | $0.02 |
| MoonshotAI: Kimi Latest~moonshotai | 1049k | $1.50 | $7.50 | $0.02 |
| Qwen: Qwen3.8 2.4T A95Bqwen | 1049k | $2.00 | $6.00 | $0.03 |
| OpenAI: o3openai | 200k | $2.00 | $8.00 | $0.03 |
| Anthropic: Claude Sonnet 5anthropic | 1000k | $2.00 | $10.00 | $0.03 |
| OpenAI: GPT-5.2-Codexopenai | 400k | $1.75 | $14.00 | $0.03 |
| Paretounbiased | 262k | $2.50 | $7.50 | $0.03 |
| OpenAI: GPT Audioopenai | 128k | $2.50 | $10.00 | $0.04 |
| Cohere: Command R+ (08-2024)cohere | 128k | $2.50 | $10.00 | $0.04 |
| Anthropic: Claude Sonnet 4.5anthropic | 1000k | $3.00 | $15.00 | $0.05 |
| OpenAI: GPT-6 Astra Pro (batch)openai | 1050k | $5.00 | $25.00 | $0.08 |
| OpenAI: GPT-5.5openai | 1050k | $5.00 | $30.00 | $0.08 |
| OpenAI: GPT-5 Pro (batch)openai | 400k | $7.50 | $60.00 | $0.14 |
| Anthropic: Claude Fable Latest~anthropic | 1000k | $10.00 | $50.00 | $0.15 |
| OpenAI: GPT-5 Proopenai | 400k | $15.00 | $120.00 | $0.27 |
| OpenAI: GPT-4openai | 8k | $30.00 | $60.00 | $0.36 |
The number that should change your mind
Input prices across the models on this page span about 8,824 times. That is not a small difference in quality for a small difference in price — it is a difference of three orders of magnitude, for models that all answer the same kind of question.
Output costs several times more than input on nearly every model. If you are generating long answers rather than reading long documents, your bill is dominated by output tokens. Ask for less, and pay less.
What actually drives your bill
- Output length, not input length, on most tasks. A terse instruction that produces a 200-token answer costs a fraction of a chatty one.
- Re-sending context. In a long conversation, the whole history is charged again on every turn. Long chats are quadratically expensive, which is why harnesses manage context carefully.
- Choosing a frontier model for a routine job. The cheapest paid models here are perfectly adequate for classification, extraction and formatting — tasks that make up most of the volume in real systems.
How to buy credits
Every route differs, and this is where people waste money:
- Directly from a lab — an account and a card per provider. Best price, worst convenience, one bill per vendor.
- An aggregator — one account and one credit balance reaching hundreds of models. Slightly more expensive per token, and worth it while you are deciding.
- Cloud platforms — the model is one line item on an existing cloud bill, which suits companies that already have one.
- Free tiers — real, and usually paid for with your data. Never send anything confidential through one.
- Prepaid credits — most providers let you buy credit rather than attach a card. Do that first: it caps the damage from a runaway loop.
Buy the smallest credit pack a provider offers and set a spend limit. An agent stuck in a retry loop can burn a surprising amount in an hour, and the invoice arrives later.