There were roughly 255 model releases in Q1 2026 alone — about three a day. As of June 2026, the frontier has fractured: there is no single best model, only a best model for your task. This guide compares every major family by benchmark, price, context window and license — and is built to be the source you can quote.
How to read this: all figures are a June 2026 snapshot. Vendor-reported numbers are marked
*; everything else comes from official docs or independent trackers (see Sources). Prices are per 1M tokens (input / output).
Key Takeaways
- No model wins everything. Claude Opus 4.8 leads overall intelligence and agentic coding; GPT-5.5 leads terminal/CLI agents; Gemini 3.1 Pro leads value + multimodal reasoning.
- The top is tight. On the Artificial Analysis (AA) Intelligence Index: Opus 4.8 (61.4) → GPT-5.5 (60.2) → Gemini 3.1 Pro (57) → Grok 4.3 (53) — within ~8 points.
- Open weights caught up. DeepSeek V4-Pro and MiniMax M3 (AA Index ~44) post near-frontier coding scores at 5–12× lower price.
- Context exploded. Llama 4 Scout holds the largest publicly released window at 10M tokens; 1M is now the default.
How This Was Compiled
We tracked every actively marketed flagship per lab as of June 17, 2026, then cross-checked specs against official model cards / pricing pages and independent trackers (Artificial Analysis Intelligence Index v4.1, LLM-Stats). Where sources disagree — common in this market — we give the range and cite both. Treat every number as a dated snapshot, not a permanent fact.
The Master Table
Closed / proprietary frontier
| Family | Model | Released | Context | Price (in/out) | AA Index | Best at |
|---|---|---|---|---|---|---|
| Claude | Opus 4.8 | 2026-05-28 | 1M | $5 / $25 | 61.4 | Overall + agentic coding |
| Claude | Fable 5 (Mythos) | 2026-06-09 | 1M | $10 / $50 | — | Top coding (SWE-Pro 80.3%*) — availability restricted |
| GPT | GPT-5.5 | 2026-04-23 | 400K–1M | $5 / $30 | 60.2 | Terminal/CLI agents, omnimodal |
| Gemini | 3.1 Pro | 2026-02 | 1M | $2 / $4–12 | 57 | Value + multimodal reasoning |
| Gemini | 3.5 Flash | 2026-05-19 | 1M | $1.50 / $9 | — | Fast, cheap agentic |
| Gemini | 3.5 Pro | 2026-06 (rolling out) | up to 2M | ~$15 / $60* | — | Deep Think (not GA mid-June) |
| Grok | 4.3 | 2026-04-30 | 1M | $1.25 / $2.50 | 53 | Budget frontier, tool calling |
| Qwen | 3.7-Max | 2026-05-19 | 1M | $2.50 / $7.50 | ~ (7th) | Knowledge + long-horizon agents |
| (Meta) | Muse Spark | 2026-04-08 | — | n/d | 52 | Meta’s first closed reasoning model |
Open-weight / open-source
| Family | Model | Released | Context | Price (in/out) | License | Best at |
|---|---|---|---|---|---|---|
| DeepSeek | V4-Pro | 2026-04-24 | 1M | ~$1.74 / $3.48 | MIT | Value champion, SWE-V ~85% |
| DeepSeek | V4-Flash | 2026-04-24 | 1M | $0.14 / $0.28 | MIT | Cheapest credible coder |
| Llama | 4 Scout | 2025-04 | 10M | self-host | Llama Community | Longest context, RAG |
| Llama | 4 Maverick | 2025-04 | 1M | self-host | Llama Community | Multimodal generalist |
| Qwen | 3.6-27B | 2026-04-22 | 262K→1M | self-host | Apache 2.0 | Small-but-strong coder |
| Kimi | K2.7 Code | 2026-06-12 | 256K | $0.95 / — | Modified MIT | MCP tool use (81.1%) |
| GLM | 5.2 | 2026-06-13 | 1M | $1.40 / $4.40 | MIT (pending) | 1M-context open flagship |
| MiniMax | M3 | 2026-06-01 | 1M | $0.30 / $1.20 | open | Top open AA Index (44) |
| Nemotron | 3 Ultra | 2026-06-04 | — | self-host | Fully permissive | Permissive-license giant |
| Mistral | Large 3 | 2025-12-02 | 256K | $0.50 / $1.50 | Apache 2.0 | Best permissive Western MoE |
Intelligence: the frontier is a cluster
Prefer a static version? Download the chart: AA Intelligence Index (PNG).
Context windows: 1M is the new default
Static version: Context windows (PNG).
Price: open weights reset the floor
Static version: Output price (PNG).
By Family
OpenAI — GPT
GPT-5.5 (Apr 23, 2026) unified the o-series + Codex into one flagship. It leads terminal-native agents (Terminal-Bench 2.0 ≈ 82.7%), is natively omnimodal, and ships an Instant default plus a higher-accuracy Pro. Pricing $5/$30; context 400K (Codex) to 1M (API). Lineage: GPT-5 (Aug 2025) → 5.1 → 5.2 → 5.4 → 5.5.
Anthropic — Claude
Opus 4.8 (May 28, 2026) is the June overall leader (AA 61.4): top agentic coding (SWE-V 88.6%), computer use (OSWorld 83.4%), knowledge work (GDPval-AA 1,890), best hallucination calibration, plus a Fast Mode. Tiers: Opus 4.8 ($5/$25) · Sonnet 4.6 ($3/$15) · Haiku 4.5 ($1/$5). The Mythos-class Fable 5 (Jun 9) set a coding ceiling (SWE-Pro 80.3%*) but its availability is currently restricted.
Google — Gemini
Gemini 3.1 Pro (Feb 2026, AA 57) is the value + multimodal leader (~$2 input, 1M ctx, GPQA ≈ 94.3%). 3.5 Flash (May 19, GA) is the fastest/cheapest tier yet ($1.50/$9, ~4× faster). 3.5 Pro (Deep Think, up to 2M ctx) is rolling out in June.
xAI — Grok
Grok 4.3 (Apr 30, 2026, AA 53): budget frontier — 1M context, function calling, structured outputs, $1.25/$2.50. Strongest features sit behind the $300/mo SuperGrok Heavy tier.
Meta — Llama + Muse
Llama 4 Scout (10M ctx) and Maverick (1M ctx) are MoE, natively multimodal, single-host friendly. Behemoth (~2T) was shelved. The Muse family now has three models: Muse Spark 1.2 (Aug 5, coding flagship, AA 54, $1.25/$4.25), Muse Glimmer (Aug 10, 30B Apache 2.0 for local agents), and the original Muse Spark (Apr 8, AA 52). Llama 5 slipped toward 2027.
DeepSeek (open weight, MIT)
V4 (Apr 24, 2026) made 1M context cheap. V4-Pro (1.6T/49B active) posts ~85% SWE-V; V4-Flash is the cheapest credible coder ($0.14/$0.28). R2 has not shipped — treat R2 specs as rumor.
Alibaba — Qwen
Qwen3.7-Max (May 19) is proprietary, API-only (1M ctx, $2.50/$7.50): the “Agent Frontier” with the highest knowledge scores in its peer set. Open weights continue lower: Qwen3.6-27B / 35B-A3B (Apache 2.0).
Chinese open-source wave
- Kimi K2.7 Code (Jun 12): 1T MoE, 256K ctx, leads MCP tool use (81.1%).
- GLM-5.2 (Jun 13): 1M ctx (largest open), MIT pending, ~$1.40/$4.40.
- MiniMax M3 (Jun 1): 1M ctx, native image/video, $0.30/$1.20, top open AA Index.
NVIDIA — Nemotron
Nemotron 3 Ultra (Jun 4): 550B, the most capable open model under a fully permissive license (no thresholds). Nemotron 3.5 Lightning (Aug 11): highest-efficiency model for long-running agentic workloads, ships with NeMo Switchyard multi-model routing library.
Mistral — Europe’s open standard
Mistral Large 3 (Dec 2, 2025): 41B active / 675B total MoE, 256K ctx, native vision, Apache 2.0 ($0.50/$1.50), shipped with edge-focused Ministral 3.
FAQ
What is the best LLM in June 2026? There is no single winner. Claude Opus 4.8 leads overall intelligence and agentic coding; GPT-5.5 leads terminal/CLI agents; Gemini 3.1 Pro leads value and multimodal reasoning; DeepSeek V4-Pro leads value-per-benchmark among open weights.
What is the best open-source LLM? MiniMax M3 and DeepSeek V4-Pro tie at the top of the AA Index among open weights (~44). Kimi K2.7 Code leads MCP tool use; Llama 4 Scout leads context length (10M); NVIDIA Nemotron 3 Ultra leads on permissive licensing.
Which model has the largest context window? Llama 4 Scout, at 10M tokens — the largest of any publicly released LLM as of June 2026.
What is the cheapest capable model? DeepSeek V4-Flash at $0.14 / $0.28 per 1M tokens, with a verified ~79% SWE-bench Verified score.
July 2026 Update
The landscape moved fast in July. Here’s what changed since the initial June snapshot:
New frontier models
| Model | Released | Key change |
|---|---|---|
| Claude Opus 5 | 2026-07-24 | New AA Intelligence Index #1 (61). Frontier-Bench SOTA (43.3%), ARC-AGI-3 30.2% (4× next-best). Same $5/$25 as Opus 4.8. |
| GPT-5.6 Sol/Terra/Luna | 2026-07-09 (GA) | Three-tier family now generally available. Sol $5/$30, Terra $2.50/$15, Luna $1/$6. 1.05M context. |
| Grok 4.5 | 2026-07-08 | xAI’s coding flagship at $2/$6. 500K context, trained alongside Cursor. |
| Gemini 3.6 Flash | 2026-07-21 | 17% fewer output tokens than 3.5 Flash at lower price ($1.50/$7.50). Replaces 3.5 Flash as workhorse. |
New open-weight models
| Model | Released | Key change |
|---|---|---|
| Kimi K3 | 2026-07-16 (API) / 07-27 (weights) | World’s first open 3T-class model. 2.8T MoE, 104B active, 1M ctx. Custom license (not MIT). |
| Inkling | 2026-07-15 | Thinking Machines Lab (Mira Murati). 975B/41B MoE, Apache 2.0, multimodal. Positioned as fine-tuning base, not frontier. |
Updated takeaways
- Claude Opus 5 is the new overall leader, displacing Opus 4.8 at the same price. The gap to Fable 5 on SWE-bench Pro is now ~1 point (79.2% vs 80.3%).
- GPT-5.6 tiered pricing is live. Terra matches GPT-5.5 quality at half the price; Luna is the cheapest GPT-5.x at $1/$6.
- The open-weight ceiling jumped to 2.8T with Kimi K3, but the custom license and 64+ GPU requirement limit practical adoption. Inkling offers a US-origin Apache 2.0 alternative at 975B.
- Gemini 3.5 Pro is still missing. Google shipped 3.6 Flash instead. Pre-training for Gemini 4 has begun.
August 2026 Update
Early August brought a Chinese-model offensive and a video-generation surge; the week of August 5–11 added a Meta open-source pivot, NVIDIA’s agentic play, and significant GPT-5.6 updates:
New frontier models
| Model | Released | Key change |
|---|---|---|
| Qwen3.8-Max | 2026-08-03 | Alibaba’s 2.4T sparse MoE (95B active). OSWorld-Verified 86.1 (tops Fable 5), PaperBench 93.0 (highest). $2/$6 — a third of Opus 5. Open weights arriving August 12. |
| Muse Spark 1.2 | 2026-08-05 | Meta’s coding flagship, co-trained with Muse Code terminal agent. TB 2.1 82.9%, AA Intelligence Index 54. $1.25/$4.25 standard; Contributor tier $0.10/$0.20 (cheapest frontier inference, shares prompts for training). 1M context. |
New open-weight models
| Model | Released | Key change |
|---|---|---|
| Muse Glimmer | 2026-08-10 | Meta’s first Apache 2.0 model — 30B dense multimodal agent. Runs on 24GB VRAM (single consumer GPU). DFlash speculative decoding ships with weights (3.1× on RTX 5090). Positioned for local agentic deployment. |
| Nemotron 3.5 Lightning | 2026-08-11 | NVIDIA’s highest-efficiency model for long-running agentic workloads. Ships with NeMo Switchyard — open-source multi-model routing library. Available as NIM, HuggingFace, OpenRouter from day one. |
GPT-5.6 mid-August updates
GPT-5.6 received three significant changes:
- Sol — improved reasoning slider for granular compute allocation; now responds to “think more/less” in natural language.
- Luna — promoted to free default for all ChatGPT users (unlimited text, 25 image gens/day); added “Think” button for on-demand extended reasoning.
- GPT-5.6-Cyber — new security-hardened variant fine-tuned on adversarial datasets; 95% detection rate on prompt injection benchmarks. Available via API at Sol pricing.
Cross-modal expansion
The LLM landscape is increasingly entangled with video and multimodal models:
| Model | Released | Key change |
|---|---|---|
| MiniMax H3 | 2026-07-31 | 33B omni-modal video model with native audio — first major open-weight video model. Extends MiniMax beyond the M3 LLM. |
| FLUX 3 Video | 2026-08-04 (GA) | BFL’s unified model goes GA with pricing ($0.06–0.53/s). Video generation becomes a commodity API. |
| Grok Imagine Video 1.5 | 2026-07-31 update | xAI adds text-to-video, native 1080p and references. $0.08/s API. |
| Grok Imagine Image 2.0 | 2026-08-07 | xAI’s editing-first image model — Arena #2 in T2I and editing. Magic wand, multi-reference compositing, templates. |
| Wan 3.0 | 2026-08-06 | Alibaba’s 30-second video model with document/web page inputs. $0.05–0.20/s. Closed-weight successor to open Wan 2.7. |
| MAGI-2 Preview | 2026-08-05 | Sand.ai 114B MoE, Apache 2.0, 10-second clips with native audio. AA Video Arena #6. |
| SeedRealtime | 2026-08-05 | ByteDance’s audio-visual full-duplex LLM. Deployed in Doubao app — real-time watch+listen+speak. |
Updated takeaways
- Meta pivots to open source. Muse Glimmer (Apache 2.0) is Meta’s first fully permissive model — a strategic reversal after Llama’s restrictive community license. Muse Spark 1.2 (open weights announced) will follow. The Contributor tier ($0.10/$0.20 for frontier inference in exchange for training data) is a novel pricing model worth watching.
- Qwen 3.8-Max open weights arrive August 12. ModelScope shows a countdown. Combined with Kimi K3 (2.8T, July), the open-weight frontier is in a sprint. A 27B variant (Qwen3.8-27B) follows shortly after.
- Local agentic models are now a category. Muse Glimmer (30B, 24GB VRAM) and Nemotron 3.5 Lightning (NeMo Switchyard routing) both target always-on agents that run on local or enterprise hardware — a category that didn’t exist three months ago.
- GPT-5.6 gets wider, not taller. OpenAI is expanding horizontally (Luna free for all, Cyber for security) rather than pushing a new number. The reasoning slider on Sol and the Think button on Luna both address the “how much compute” question differently than any rival.
- Video input is the new frontier. Wan 3.0’s document-to-video and MAGI-2’s permissive video generation join FLUX 3 and MiniMax H3 — four new video models in one week, eclipsing text-to-image as the hot modality.
- The frontier pricing ladder is clearer: Fable 5 ($10/$50) for max capability → Opus 5 ($5/$25) for near-Fable at half price → Qwen 3.8-Max ($2/$6) for agentic workloads on a budget → Muse Spark 1.2 ($1.25/$4.25) for coding-focused value → Sonnet 5 ($2/$10) for workhorse tasks.
Sources
- Anthropic — Claude models overview
- OpenAI — Introducing GPT-5.5
- Google — Gemini 3.5
- xAI — Grok 4.3 docs
- Meta — Llama 4 herd
- DeepSeek — V4 release notes
- Alibaba Cloud — Qwen3.7: The Agent Frontier
- Mistral AI — Introducing Mistral 3
- Artificial Analysis — Intelligence Index
- Build Fast with AI — Best AI Models June 2026
- Alibaba Cloud — Qwen3.8-Max: A New Bar for Coding and Cowork
Changelog: 2026-08-11 — August update 2: added Muse Spark 1.2, Muse Glimmer, Nemotron 3.5 Lightning; GPT-5.6 Sol/Luna/Cyber mid-August updates; Qwen 3.8-Max open-weight countdown; Grok Imagine Image 2.0, Wan 3.0, MAGI-2 to cross-modal expansion; updated Meta and Nemotron family sections. 2026-08-05 — August update: added Qwen 3.8-Max, cross-modal expansion section (MiniMax H3, FLUX 3 Video GA, Grok Imagine Video 1.5, SeedRealtime). 2026-07-29 — July update: added Opus 5, GPT-5.6 GA, Grok 4.5, Gemini 3.6 Flash, Kimi K3, Inkling. 2026-06-17 — initial publication.
Last updated: 2026-08-11