Updated 5 August 2026
Comparison of 8 large language models — 4 proprietary and 4 open-weight. Capability is the Artificial Analysis Intelligence Index, one eval suite run by an independent third party across every model here. Context windows, output limits and prices are read from each vendor’s own documentation, dated 2026-08-22. We do not assign scores of our own.
| Model | Provider | Context | Input $/1M | Output $/1M | Vision | Open Source | Best For |
|---|---|---|---|---|---|---|---|
| Claude Opus 5Frontier | Anthropic | 1M tokens | $5.00 | $25.00 | Complex analysis | ||
| GPT-5.6 SolFrontier | OpenAI | 1.05M tokens | $4.00 | $20.00 | Production AI applications | ||
| Claude Sonnet 5Frontier | Anthropic | 1M tokens | $3.00 | $15.00 | Customer-facing chatbots | ||
| Gemini 3.1 Pro PreviewFrontier | Google DeepMind | 1M tokens | $2.00 | $12.00 | Multimodal workflows | ||
| DeepSeek V4 FlashOpen Source | DeepSeek | 1M tokens | $0.22 | $0.66 | On-premise and air-gapped deployment | ||
| Qwen3.5 397BOpen Source | Alibaba Cloud | 128K tokens | Self-host | Self-host | Multilingual open-weight deployments | ||
| Mistral Large 3Open Source | Mistral AI | 128K tokens | Self-host | Self-host | European data residency and GDPR-constrained deployments where jurisdiction outranks raw capability | ||
| Llama 4 MaverickOpen Source | Meta | 1M tokens | Self-host | Self-host | Existing Llama deployments and fine-tuning pipelines already built around its tooling |
List prices read from each vendor’s own pricing page on 5 August 2026. Claude Sonnet 5 shows its durable $3.00 / $15.00 rate; an introductory $2.00 / $10.00 applies through 31 August 2026. Gemini 3.1 Pro is $2.00 / $12.00 for prompts at or under 200K tokens and $4.00 / $18.00 above. Open-weight models can also be rented from third-party hosts (Together AI, Fireworks, Groq) at their own rates.
We do not score these models ourselves. Every capability figure on this page is the Artificial Analysis Intelligence Index — a single composite eval suite run by an independent third party across 177 models, so the numbers are comparable to each other. Every spec and price is read from the vendor's own documentation, with the date we read it. Where a vendor does not publish a figure we print 'Not published' rather than an estimate. An earlier version of this page assigned its own 1-10 scores across eight dimensions; those were removed on 2026-08-05 because no vendor publishes such a score and we could not substantiate ours.
Intelligence Index
Artificial Analysis' composite score, 0-100, run identically across every model listed. Read 2026-08-05.
Specs
Context window and max output, taken from vendor documentation. 'Not published' means exactly that.
Pricing
List price per 1M tokens from the vendor's own pricing page. Introductory rates are labelled with their end date.
Licensing
Whether weights are published and under what terms — MIT, Apache 2.0, and bespoke community licences are not interchangeable.
Anthropic · Jul 2026 · Undisclosed · Undisclosed
Artificial Analysis Intelligence Index (adaptive reasoning, max effort), read 2026-08-05. Source
Strengths
Weaknesses
OpenAI · Jul 2026 · Undisclosed · Undisclosed
Artificial Analysis Intelligence Index (max effort), read 2026-08-05. Source
Strengths
Weaknesses
Anthropic · Jun 2026 · Undisclosed · Undisclosed
Artificial Analysis Intelligence Index (max effort), read 2026-08-05. Source
Strengths
Weaknesses
Google DeepMind · Jan 2026 · Multimodal MoE Transformer · Undisclosed
Artificial Analysis Intelligence Index, read 2026-08-05. Source
Strengths
Weaknesses
DeepSeek · Jul 2026 · MoE with Compressed Sparse Attention · 284B total (13B active per token)
Artificial Analysis Intelligence Index (max effort), read 2026-08-05. Source
Strengths
Weaknesses
Alibaba Cloud · Feb 2026 · MoE Transformer · 397B total (17B active per token)
Artificial Analysis Intelligence Index, read 2026-08-05. Source
Strengths
Weaknesses
Mistral AI · Jan 2026 · MoE Transformer · 675B MoE
Artificial Analysis Intelligence Index, read 2026-08-05. Source
Strengths
Weaknesses
Meta · Feb 2026 · MoE Transformer (128 experts, top-1 routing) · 400B total (17B active per token)
Artificial Analysis Intelligence Index, read 2026-08-05. Source
Strengths
Weaknesses
| Feature | Opus 5 | GPT-5.6 Sol | Sonnet 5 | Gemini 3.1 Pro | DeepSeek V4 Flash | Qwen3.5 397B | Mistral Large 3 | Llama 4 Maverick |
|---|---|---|---|---|---|---|---|---|
| Intelligence Index | 61 | 59 | 53 | 46 | 50 | 34 | 16 | 14 |
| Context Window | 1M | 1.05M | 1M | 1M | 1M | 128K | 128K | 1M |
| Max Output Tokens | 128,000 | 128,000 | 128,000 | 64,000 | Not published | Not published | Not published | Not published |
| Input Price (per 1M tokens) | $5.00 | $4.00 | $3.00 | $2.00 | $0.22 | Free (self-host) | Free (self-host) | Free (self-host) |
| Output Price (per 1M tokens) | $25.00 | $20.00 | $15.00 | $12.00 | $0.66 | Free (self-host) | Free (self-host) | Free (self-host) |
| Multimodal (Vision) | Yes | Yes | Yes | Yes (+ video, audio) | No | Yes | No | Yes |
| Tool / Function Calling | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Open Weights | No | No | No | No | Yes (MIT) | Yes (Apache 2.0) | Yes (Apache 2.0) | Yes (Community License) |
| Self-Hosting | No | No | No | No | Yes | Yes | Yes | Yes |
Our recommendation for each scenario, and the reason. Where the call rests on a measured figure we say which one; where it rests on preference we say that instead.
24/7 AI agent handling customer inquiries, FAQs, and ticket routing
Sonnet 5 carries an index of 53 at $3/$15 with the same 1M context as Opus 5 — the best capability-per-dollar of the proprietary models, and support volume is where per-token cost compounds fastest.
Writing, reviewing, and debugging production code across multiple languages
Opus 5 leads the independent index at 61 and Anthropic positions it specifically for agentic coding. This page previously named GPT as the coding leader; at 61 vs 59 that is no longer the ordering the data supports.
Analyzing long contracts, medical records, or regulatory filings
Highest measured intelligence, 1M context, 128K output, and knowledge reliable through May 2026 — the most recent of any model here, which matters when the documents reference recent rules.
Processing images, video, audio, and text in unified pipelines
Gemini is the only model listed that understands video and audio natively alongside text and images in a single API call. Its index of 46 is the lowest of the frontier group, so use it for the modalities, not the reasoning.
High-volume inference where cost per token is the primary constraint
DeepSeek V4 Flash scores 50 at $0.22/$0.66 — higher than Gemini 3.1 Pro at roughly a third of the input price. Nothing else listed is close on capability per dollar.
Deployments requiring full data control with no external API calls
MIT licence, open weights, 1M context, and an index of 50 against Llama 4 Maverick's 14. This page previously recommended Llama 4 Maverick here; the independent index does not support that and it has been corrected.
Creating and translating content across multiple languages and markets
Google publishes 92.6% on MMMLU for Gemini 3.1 Pro on its own model card — an actual multilingual benchmark figure rather than an impression. Qwen remains the strongest open-weight option for Asian languages.
Blog posts, ad copy, email campaigns, and brand content
Writing quality is the one dimension here with no clean public benchmark, so this is a preference stated as a preference: in our own use Claude holds tone and brand voice better than the alternatives. Treat it as opinion, not measurement.
EU-based deployments requiring GDPR compliance and local hosting
Mistral is Paris-headquartered, Apache 2.0, and available on European infrastructure. Its index of 16 is low — this recommendation is about jurisdiction, and you are trading real capability for it.
Claude Opus 5, GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.1 Pro
Best when: You need top-tier quality, don't want to manage infrastructure, and API costs are acceptable for your volume.
DeepSeek V4 Flash, Qwen3.5 397B, Mistral Large 3, Llama 4 Maverick
Best when: You need on-premise deployment, have high token volumes, require custom training, or need data residency compliance.
For most business AI implementations, we recommend Claude Sonnet 5 as the default. It scores 53 on the independent index at $3/$15 per 1M tokens, with the same 1M token context as Opus 5 — the best capability per dollar of the proprietary models, which is what matters once volume is real.
For complex analysis, legal documents, agentic coding, or research-heavy workflows, upgrade to Claude Opus 5 (index 61, the highest measured here) or GPT-5.6 Sol (59, with the largest context at 1.05M tokens).
For on-premise, air-gapped, or budget-dominated work we deploy DeepSeek V4 Flash — open weights under an MIT licence, 1M token context, and an index of 50 at $0.14/$0.28 per 1M tokens. It measures higher than Gemini 3.1 Pro at a fraction of the cost.
This page previously recommended Llama 4 Maverick for on-premise deployment. On the independent index it scores 14, against DeepSeek V4 Flash’s 50, so that recommendation has been withdrawn. Llama remains a reasonable choice where you have already built tooling and fine-tunes around it. Where EU data residency is the binding constraint, Mistral Large 3 is still the answer — but at an index of 16 you are trading real capability for jurisdiction, and you should know that is the trade you are making.
On the Artificial Analysis Intelligence Index — one eval suite run identically across 177 models — Claude Opus 5 leads at 61, followed by GPT-5.6 Sol at 59 and Claude Sonnet 5 at 53 (read 2026-08-05). But 'best' depends on the constraint: for capability per dollar, DeepSeek V4 Flash scores 50 at $0.22/$0.66 per 1M tokens; for video and audio, Gemini 3.1 Pro is the only option listed that handles them natively; for open weights with permissive terms, DeepSeek V4 Flash is MIT-licensed.
Per 1M tokens, list prices read from the vendors: DeepSeek V4 Flash $0.22 in / $0.66 out (off-peak); Gemini 3.1 Pro Preview $2.00 / $12.00 (prompts at or under 200K tokens; $4.00 / $18.00 above); Claude Sonnet 5 $3.00 / $15.00; Claude Opus 5 $5.00 / $25.00; GPT-5.6 Sol $4.00 / $20.00. Open-weight models such as Qwen3.5, Mistral Large 3 and Llama 4 Maverick have no API cost when self-hosted — you pay for compute instead.
The gap is narrower than it was. DeepSeek V4 Flash is open weights under an MIT licence and scores 50 on the independent index — above Gemini 3.1 Pro at 46 — with a 1M token context. That makes on-premise a real option rather than a compromise. The remaining reasons to pay for a proprietary model are the top of the index (Opus 5 at 61), native video and audio (Gemini), and not having to run inference infrastructure at all.
GPT-5.6 Sol at 1.05M tokens, narrowly ahead of Claude Opus 5, Claude Sonnet 5, Gemini 3.1 Pro Preview, DeepSeek V4 Flash and Llama 4 Maverick at 1M each. A million tokens is roughly 555,000 words. Context is no longer a differentiator between frontier models the way it was in 2025 — max output is: Claude and GPT-5.6 publish 128K, Gemini publishes 64K, and several open-weight models publish none.
Claude Sonnet 5. It scores 53 on the independent index at $3/$15 per 1M tokens with a 1M token context for large knowledge bases — the strongest capability-per-dollar of the proprietary models, which matters most in support workloads where volume is high. GPT-5.6 Sol is the alternative if you need the wider plugin ecosystem, and GPT-5.6 Terra at $2/$12 is worth pricing if cost dominates.
Yes — DeepSeek V4 Flash (MIT), Qwen3.5 397B (Apache 2.0), Mistral Large 3 (Apache 2.0) and Llama 4 Maverick (Llama 4 Community License) all publish weights. DeepSeek V4 Flash activates only 13B of 284B parameters per token, which is what makes self-hosting practical at its capability level. Claude, GPT-5.6 and Gemini are API-only.
We start from the constraint, not the leaderboard. If data cannot leave the building, that is DeepSeek V4 Flash self-hosted. If cost per token dominates, the same model on its hosted API. If the work is hard reasoning or agentic coding, Claude Opus 5. If it is a high-volume customer-facing chatbot, Claude Sonnet 5. If it involves video or audio, Gemini 3.1 Pro. Every engagement starts with a discovery phase, and we re-check the index before recommending — the ordering on this page has already changed once.
In a Mixture of Experts model only a fraction of the total parameters activate for any given token. DeepSeek V4 Flash has 284B total parameters but routes to just 13B per token; Llama 4 Maverick is 400B total and 17B active; Qwen3.5 397B is 17B active. The effect is that a large model can be served at the speed and cost of a much smaller one, which is why the open-weight tier is dominated by MoE designs. The frontier labs do not disclose whether their current models use it.
Head-to-head: OpenAI vs Anthropic for business AI
Best workflow automation platform compared
AI coding assistants for developer teams
Voice AI platforms for business automation
AI image generation for marketing teams
AI video generation platforms compared
Ready?
Not sure which model fits your business? We'll recommend the optimal AI stack during a free discovery call.
Aria picks up instantly · 24/7 · Free assessment · 30-day guarantee