Compare Claude, GPT, Gemini, DeepSeek, Groq, and Together AI pricing. Anthropic, OpenAI and Google prices verified Sep 5, 2026; other providers May 2026.
Cost Calculator
| Model | Provider | Input / 1M | Output / 1M | Context | Cost / request ↑ |
|---|---|---|---|---|---|
| Llama 3.1 8B (Groq) | Groq | $0.05 | $0.08 | 128K | $0.0001cheapest |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | 1M | $0.0003 |
| Llama 4 Scout (Groq) | Groq | $0.11 | $0.34 | 128K | $0.0003 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | $0.0003 | |
| Mistral Small 4 | Mistral | $0.15 | $0.60 | 128K | $0.0004 |
| Qwen3 32B (Groq) | Groq | $0.29 | $0.59 | 131K | $0.0006 |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | 1M | $0.0008 |
| DeepSeek V4-Pro | DeepSeek | $0.43 | $0.87 | 1M | $0.0009 |
| Llama 3.3 70B (Groq) | Groq | $0.59 | $0.79 | 128K | $0.0010 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | $0.0010 | |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | 128K | $0.0013 |
| Llama 3.3 70B (Together) | Together AI | $0.88 | $0.88 | 128K | $0.0013 |
| DeepSeek V3.1 (Together) | Together AI | $0.60 | $1.70 | 128K | $0.0014 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | $0.0015 | |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1M | $0.0015 | |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1M | $0.0026 | |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1M | $0.0026 | |
| Gemini 3.6 Flash | $0.75 | $3.75 | 1M | $0.0026 | |
| GPT-5.4 mini | OpenAI | $0.75 | $4.50 | 400K | $0.0030 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | 200K | $0.0035 |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | 128K | $0.0053 |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1M | $0.0060 | |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1M | $0.0063 | |
| DeepSeek R1-0528 (Together) | Together AI | $3.00 | $7.00 | 128K | $0.0065 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | 1M | $0.0070 |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | 1M | $0.0080 |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M | $0.0080 | |
| GPT-5.3-Codex | OpenAI | $1.75 | $14.00 | 400K | $0.0088 |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | 1M | $0.010 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | 1M | $0.011 |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 | 1M | $0.014 |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | 1M | $0.018 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | 1M | $0.018 |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | 1M | $0.018 |
| Claude Opus 4.6 | Anthropic | $5.00 | $25.00 | 1M | $0.018 |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | 1M | $0.020 |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | 1M | $0.035 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | 1M | $0.035 |
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | 1M | $0.035 |
* DeepSeek V4-Flash: Cache hit: $0.003/1M
* GPT-5.6 Luna: Fast everyday model. Cached input: $0.02/1M
* DeepSeek V4-Pro: Cache hit: $0.0036/1M
* Llama 3.3 70B (Groq): Ultra-fast inference
* Gemini 3.1 Flash-Lite: Stable. Cache: $0.025/1M
* Mistral Large 3: Open-weight flagship
* Gemini 3.5 Flash-Lite: Most cost-efficient GA tier. Cache: $0.03/1M
* Gemini 3.8 Flash: New, GA Sep 2, 2026 — Google's 'most intelligent Flash model'. Introductory thru Dec 31, 2026 — $1.50/$7.50 standard from Jan 1, 2027. Cache: $0.075/1M
* Gemini 3.7 Flash: GA Aug 13, 2026, now previous-generation Flash. Introductory thru Dec 31, 2026 — $1.50/$7.50 standard from Jan 1, 2027. Cache: $0.075/1M
* Gemini 3.6 Flash: Cut to the same introductory rate as 3.7/3.8 (was $1.50/$7.50 on Aug 25) — $1.50/$7.50 standard from Jan 1, 2027. Cache: $0.075/1M
* GPT-5.4 mini: Cached input: $0.075/1M
* Gemini 3.5 Flash: Stable. Cache: $0.15/1M
* Gemini 2.5 Pro: $2.50/$15 for >200K
* Claude Sonnet 5: Standard price — the planned Sep 1, 2026 hike to $3/$15 was cancelled, $2/$10 stays. Cache hit: $0.20/1M
* GPT-5.6 Terra: Balanced high-volume model. Cached input: $0.20/1M
* Gemini 3.1 Pro: Preview. $4/$18 for >200K tokens
* GPT-5.3-Codex: Agentic coding model. Feb 2026
* GPT-5.4: Cached input: $0.25/1M
* GPT-5.6 Sol: Agentic workhorse one tier below GPT-6 Astra. Promotional price thru at least Nov 21, 2026 (was $5/$30). Cached input: $0.40/1M
* Claude Opus 5: Current Opus flagship, Jul 24, 2026. Cache hit: $0.50/1M
* Claude Opus 4.8: Previous Opus flagship, same price. Cache hit: $0.50/1M
* GPT-5.5: Previous flagship. Cached input: $0.50/1M
* Claude Fable 5.1: New flagship, Sep 1, 2026 — adaptive thinking always on, 128K max output. Cache hit: $0.25/1M (2.5% of input; every other Claude model is 10%). Claude Mythos 5.1 costs the same, limited availability
* Claude Fable 5: Previous Fable flagship, now on Anthropic's legacy list (still available). Cache hit: $1/1M
* GPT-6 Astra: New flagship, Sep 3, 2026 — OpenAI's 'most capable model, built for the hardest end-to-end work'. Cached input: $1/1M; long-context requests $20/$75
Prices from official provider docs. Last verified June 29, 2026 (Anthropic, OpenAI, Google: Sep 5, 2026). Cached input pricing not shown.
Pro tip
Using AI APIs for code review? Git AutoReview runs Claude, Gemini, and GPT in parallel with BYOK — you control costs, we handle the orchestration.
AI providers charge based on the number of tokens processed. A token is roughly 4 characters or 0.75 words in English. Input tokens (your prompt) are cheaper than output tokens (the model's response) because generation requires more compute. Most providers list prices per 1 million tokens.
For a typical code review, you send 2,000-5,000 input tokens (the PR diff plus system prompt) and receive 500-1,500 output tokens (review comments). At Claude Sonnet 4.6 rates, that is about $0.01-$0.03 per review. Running 100 reviews per day costs $1-$3 — far less than the developer time saved.
For code review specifically, Claude Opus 4.7 catches the deepest architectural issues but costs more. Gemini 2.5 Pro offers strong value with its 1M token context at $1.25/1M input, though the newer Gemini 3.1 Pro brings better performance at $2.00/1M input. GPT-5.4 nano at $0.20/$1.25 is the budget option for high-volume reviews.
Git AutoReview lets you run multiple models in parallel and compare findings. This catches more issues than any single model because different models have different strengths — Claude excels at logic bugs, GPT at security patterns, Gemini at understanding large codebases.
Context window determines how much code the model can see at once. For reviewing a small PR (under 500 lines), any model works. For large PRs or when you need the model to understand surrounding code, the 1M token context of Claude, GPT-5.4, and Gemini 2.5/3.1 is essential.
DeepSeek's 128K context is enough for most individual file reviews but may truncate large monorepo diffs. Groq's speed makes it ideal for quick checks where latency matters more than context size.
If your system prompt stays the same across requests (which it does for code review), cached input pricing saves 50-90%. Anthropic caches prompt prefixes at $0.30/1M (vs $3.00 standard) for Sonnet 4.6. DeepSeek caches at $0.003/1M. This makes repeated reviews significantly cheaper.
All prices on this page are sourced from official provider pricing pages. We verify against platform.claude.com, developers.openai.com, ai.google.dev, api-docs.deepseek.com, groq.com, and together.ai directly. Prices are updated bi-weekly. Last verified: June 29, 2026 (Anthropic, OpenAI, Google: Sep 5, 2026).
Most AI providers charge per token — a token is roughly 4 characters or 0.75 words in English. Pricing is split between input tokens (your prompt) and output tokens (the model's response). Prices are listed per 1 million tokens.
On this table the floor is Gemini 2.5 Flash-Lite at $0.10/$0.40 and DeepSeek V4-Flash at $0.14/$0.28 per 1M tokens. For review quality per dollar, the Gemini 3.x Flash line at $0.75/$3.75 (introductory through December 31, 2026) is hard to beat. Git AutoReview supports BYOK, so you pay whichever provider you pick directly, with your own key.
Input tokens are the text you send to the model (your prompt, context, instructions). Output tokens are the text the model generates in response. Output tokens are typically 3-5x more expensive than input tokens because generation requires more compute.
A typical PR review uses 2,000-5,000 input tokens (diff plus context) and 500-1,500 output tokens (comments). On Claude Sonnet 5 at $2/$10 that is roughly $0.01-$0.03 per review; on Gemini 3.x Flash or DeepSeek, under a cent. Git AutoReview is BYOK, so that per-review API cost is what you pay on top of the subscription — typically $2-5 a month for a working developer.
BYOK (Bring Your Own Key) means you use your own API keys and pay the provider directly. This gives you full control over costs and means your code goes directly to the AI provider, not through a middleman. Git AutoReview supports BYOK on all plans including Free.
Most current flagships run 1 million tokens: Claude Fable 5.1, Opus 5 and Sonnet 5, GPT-6 Astra and the GPT-5.6 family, and Gemini from 3.8 Flash down to 2.5 Pro. On Anthropic's current tokenizer that is roughly 555,000 words in one request — enough for a mid-sized codebase.
AI API prices typically drop every 3-6 months as providers optimize their infrastructure, and the September 2026 sweep alone brought a new flagship from each of Anthropic, OpenAI and Google. We re-verify this table against the official pricing pages every two weeks; the Anthropic, OpenAI and Google rows were last checked on September 5, 2026.
Yes. Most providers discount cached or repeated prompt tokens. Anthropic charges 10% of the input price on a cache hit (2.5% on Claude Fable 5.1), OpenAI charges 10% across the GPT-5.x and GPT-6 lines, and DeepSeek's cache hits sit near $0.003-$0.004 per 1M tokens. This is especially useful for code review, where the system prompt is the same every time.
Yes. Git AutoReview runs Claude, Gemini, and GPT in parallel and merges duplicate findings. This catches more issues than any single model. With BYOK, you control the cost of each model independently.
Claude Fable 5.1, the new flagship from September 1, 2026: $10/1M input, $50/1M output, with cache hits at just $0.25/1M. Claude Opus 5: $5/1M input, $25/1M output. Claude Sonnet 5: $2/1M input, $10/1M output — the planned September hike to $3/$15 was cancelled. Claude Sonnet 4.6: $3/1M input, $15/1M output. Claude Haiku 4.5: $1/1M input, $5/1M output. Claude Fable 5 stays available at the same $10/$50 as 5.1, now on Anthropic's legacy list.
GPT-6 Astra, released September 3, 2026: $10/1M input, $50/1M output. GPT-5.6 Sol: $4/1M input, $20/1M output at a promotional price that holds at least through November 21, 2026. GPT-5.6 Terra: $2/$12. GPT-5.6 Luna: $0.20/$1.20. GPT-5.5: $5/$30. GPT-5.4: $2.50/$15, GPT-5.4 mini: $0.75/$4.50, GPT-5.4 nano: $0.20/$1.25. Cached input costs 10% of the input price across the GPT-5.x and GPT-6 lines.
Gemini 3.8 Flash, generally available since September 2, 2026, and its 3.7 and 3.6 siblings all cost $0.75/1M input, $3.75/1M output through December 31, 2026, then $1.50/$7.50. Gemini 3.1 Pro: $2/1M input, $12/1M output (preview; $4/$18 above 200K tokens). Gemini 2.5 Pro: $1.25/1M input, $10/1M output ($2.50/$15 above 200K). Gemini 2.5 Flash: $0.30/$2.50. Gemini 2.5 Flash-Lite: $0.10/$0.40.
Developer Toolkit by Git AutoReview
Free tools for developers. AI code review for teams.