There were roughly 255 model releases in Q1 2026 alone — about three a day. As of June 2026, the frontier has fractured: there is no single best model, only a best model for your task. This guide compares every major family by benchmark, price, context window and license — and is built to be the source you can quote.

How to read this: all figures are a June 2026 snapshot. Vendor-reported numbers are marked *; everything else comes from official docs or independent trackers (see Sources). Prices are per 1M tokens (input / output).

Key Takeaways

  • No model wins everything. Claude Opus 4.8 leads overall intelligence and agentic coding; GPT-5.5 leads terminal/CLI agents; Gemini 3.1 Pro leads value + multimodal reasoning.
  • The top is tight. On the Artificial Analysis (AA) Intelligence Index: Opus 4.8 (61.4) → GPT-5.5 (60.2) → Gemini 3.1 Pro (57) → Grok 4.3 (53) — within ~8 points.
  • Open weights caught up. DeepSeek V4-Pro and MiniMax M3 (AA Index ~44) post near-frontier coding scores at 5–12× lower price.
  • Context exploded. Llama 4 Scout holds the largest publicly released window at 10M tokens; 1M is now the default.

How This Was Compiled

We tracked every actively marketed flagship per lab as of June 17, 2026, then cross-checked specs against official model cards / pricing pages and independent trackers (Artificial Analysis Intelligence Index v4.1, LLM-Stats). Where sources disagree — common in this market — we give the range and cite both. Treat every number as a dated snapshot, not a permanent fact.

The Master Table

Closed / proprietary frontier

FamilyModelReleasedContextPrice (in/out)AA IndexBest at
ClaudeOpus 4.82026-05-281M$5 / $2561.4Overall + agentic coding
ClaudeFable 5 (Mythos)2026-06-091M$10 / $50Top coding (SWE-Pro 80.3%*) — availability restricted
GPTGPT-5.52026-04-23400K–1M$5 / $3060.2Terminal/CLI agents, omnimodal
Gemini3.1 Pro2026-021M$2 / $4–1257Value + multimodal reasoning
Gemini3.5 Flash2026-05-191M$1.50 / $9Fast, cheap agentic
Gemini3.5 Pro2026-06 (rolling out)up to 2M~$15 / $60*Deep Think (not GA mid-June)
Grok4.32026-04-301M$1.25 / $2.5053Budget frontier, tool calling
Qwen3.7-Max2026-05-191M$2.50 / $7.50~ (7th)Knowledge + long-horizon agents
(Meta)Muse Spark2026-04-08n/d52Meta’s first closed reasoning model

Open-weight / open-source

FamilyModelReleasedContextPrice (in/out)LicenseBest at
DeepSeekV4-Pro2026-04-241M~$1.74 / $3.48MITValue champion, SWE-V ~85%
DeepSeekV4-Flash2026-04-241M$0.14 / $0.28MITCheapest credible coder
Llama4 Scout2025-0410Mself-hostLlama CommunityLongest context, RAG
Llama4 Maverick2025-041Mself-hostLlama CommunityMultimodal generalist
Qwen3.6-27B2026-04-22262K→1Mself-hostApache 2.0Small-but-strong coder
KimiK2.7 Code2026-06-12256K$0.95 / —Modified MITMCP tool use (81.1%)
GLM5.22026-06-131M$1.40 / $4.40MIT (pending)1M-context open flagship
MiniMaxM32026-06-011M$0.30 / $1.20openTop open AA Index (44)
Nemotron3 Ultra2026-06-04self-hostFully permissivePermissive-license giant
MistralLarge 32025-12-02256K$0.50 / $1.50Apache 2.0Best permissive Western MoE

Intelligence: the frontier is a cluster

Artificial Analysis Intelligence Index (June 2026)
Claude Opus 4.861.4GPT-5.560.2Gemini 3.1 Pro57Grok 4.353Muse Spark52DeepSeek V4-Pro44MiniMax M344
ClosedOpen
The top four closed models sit within ~8 points; open weights trail by ~16.
Source: Artificial Analysis Intelligence Index v4.1, as of June 2026.

Prefer a static version? Download the chart: AA Intelligence Index (PNG).

Context windows: 1M is the new default

LLM Context Windows (June 2026)
Llama 4 Scout10M tokensGemini 3.5 Pro2M tokensOpus 4.81M tokensGPT-5.5 (API)1M tokensDeepSeek V41M tokensGLM-5.21M tokensGPT-5.5 (Codex)400K tokensMistral Large 3256K tokensHaiku 4.5200K tokens
Llama 4 Scout's 10M-token window is 10× the new 1M-token default.
Source: official model cards, as of June 2026. Log scale.

Static version: Context windows (PNG).

Price: open weights reset the floor

Output Price ($/1M tokens, June 2026)
DeepSeek V4-Flash$0.28MiniMax M3$1.2Mistral Large 3$1.5Grok 4.3$2.5DeepSeek V4-Pro$3.48Qwen3.7-Max$7.5Gemini 3.5 Flash$9Claude Opus 4.8$25GPT-5.5$30Claude Fable 5$50
OpenClosed
DeepSeek V4-Flash ($0.28) is ~180× cheaper on output than Claude Fable 5 ($50).
Source: official API pricing pages, as of June 2026.

Static version: Output price (PNG).

By Family

OpenAI — GPT

GPT-5.5 (Apr 23, 2026) unified the o-series + Codex into one flagship. It leads terminal-native agents (Terminal-Bench 2.0 ≈ 82.7%), is natively omnimodal, and ships an Instant default plus a higher-accuracy Pro. Pricing $5/$30; context 400K (Codex) to 1M (API). Lineage: GPT-5 (Aug 2025) → 5.1 → 5.2 → 5.4 → 5.5.

Anthropic — Claude

Opus 4.8 (May 28, 2026) is the June overall leader (AA 61.4): top agentic coding (SWE-V 88.6%), computer use (OSWorld 83.4%), knowledge work (GDPval-AA 1,890), best hallucination calibration, plus a Fast Mode. Tiers: Opus 4.8 ($5/$25) · Sonnet 4.6 ($3/$15) · Haiku 4.5 ($1/$5). The Mythos-class Fable 5 (Jun 9) set a coding ceiling (SWE-Pro 80.3%*) but its availability is currently restricted.

Google — Gemini

Gemini 3.1 Pro (Feb 2026, AA 57) is the value + multimodal leader (~$2 input, 1M ctx, GPQA ≈ 94.3%). 3.5 Flash (May 19, GA) is the fastest/cheapest tier yet ($1.50/$9, ~4× faster). 3.5 Pro (Deep Think, up to 2M ctx) is rolling out in June.

xAI — Grok

Grok 4.3 (Apr 30, 2026, AA 53): budget frontier — 1M context, function calling, structured outputs, $1.25/$2.50. Strongest features sit behind the $300/mo SuperGrok Heavy tier.

Meta — Llama + Muse

Llama 4 Scout (10M ctx) and Maverick (1M ctx) are MoE, natively multimodal, single-host friendly. Behemoth (~2T) was shelved. The Muse family now has three models: Muse Spark 1.2 (Aug 5, coding flagship, AA 54, $1.25/$4.25), Muse Glimmer (Aug 10, 30B Apache 2.0 for local agents), and the original Muse Spark (Apr 8, AA 52). Llama 5 slipped toward 2027.

DeepSeek (open weight, MIT)

V4 (Apr 24, 2026) made 1M context cheap. V4-Pro (1.6T/49B active) posts ~85% SWE-V; V4-Flash is the cheapest credible coder ($0.14/$0.28). R2 has not shipped — treat R2 specs as rumor.

Alibaba — Qwen

Qwen3.7-Max (May 19) is proprietary, API-only (1M ctx, $2.50/$7.50): the “Agent Frontier” with the highest knowledge scores in its peer set. Open weights continue lower: Qwen3.6-27B / 35B-A3B (Apache 2.0).

Chinese open-source wave

  • Kimi K2.7 Code (Jun 12): 1T MoE, 256K ctx, leads MCP tool use (81.1%).
  • GLM-5.2 (Jun 13): 1M ctx (largest open), MIT pending, ~$1.40/$4.40.
  • MiniMax M3 (Jun 1): 1M ctx, native image/video, $0.30/$1.20, top open AA Index.

NVIDIA — Nemotron

Nemotron 3 Ultra (Jun 4): 550B, the most capable open model under a fully permissive license (no thresholds). Nemotron 3.5 Lightning (Aug 11): highest-efficiency model for long-running agentic workloads, ships with NeMo Switchyard multi-model routing library.

Mistral — Europe’s open standard

Mistral Large 3 (Dec 2, 2025): 41B active / 675B total MoE, 256K ctx, native vision, Apache 2.0 ($0.50/$1.50), shipped with edge-focused Ministral 3.

FAQ

What is the best LLM in June 2026? There is no single winner. Claude Opus 4.8 leads overall intelligence and agentic coding; GPT-5.5 leads terminal/CLI agents; Gemini 3.1 Pro leads value and multimodal reasoning; DeepSeek V4-Pro leads value-per-benchmark among open weights.

What is the best open-source LLM? MiniMax M3 and DeepSeek V4-Pro tie at the top of the AA Index among open weights (~44). Kimi K2.7 Code leads MCP tool use; Llama 4 Scout leads context length (10M); NVIDIA Nemotron 3 Ultra leads on permissive licensing.

Which model has the largest context window? Llama 4 Scout, at 10M tokens — the largest of any publicly released LLM as of June 2026.

What is the cheapest capable model? DeepSeek V4-Flash at $0.14 / $0.28 per 1M tokens, with a verified ~79% SWE-bench Verified score.

July 2026 Update

The landscape moved fast in July. Here’s what changed since the initial June snapshot:

New frontier models

ModelReleasedKey change
Claude Opus 52026-07-24New AA Intelligence Index #1 (61). Frontier-Bench SOTA (43.3%), ARC-AGI-3 30.2% (4× next-best). Same $5/$25 as Opus 4.8.
GPT-5.6 Sol/Terra/Luna2026-07-09 (GA)Three-tier family now generally available. Sol $5/$30, Terra $2.50/$15, Luna $1/$6. 1.05M context.
Grok 4.52026-07-08xAI’s coding flagship at $2/$6. 500K context, trained alongside Cursor.
Gemini 3.6 Flash2026-07-2117% fewer output tokens than 3.5 Flash at lower price ($1.50/$7.50). Replaces 3.5 Flash as workhorse.

New open-weight models

ModelReleasedKey change
Kimi K32026-07-16 (API) / 07-27 (weights)World’s first open 3T-class model. 2.8T MoE, 104B active, 1M ctx. Custom license (not MIT).
Inkling2026-07-15Thinking Machines Lab (Mira Murati). 975B/41B MoE, Apache 2.0, multimodal. Positioned as fine-tuning base, not frontier.

Updated takeaways

  • Claude Opus 5 is the new overall leader, displacing Opus 4.8 at the same price. The gap to Fable 5 on SWE-bench Pro is now ~1 point (79.2% vs 80.3%).
  • GPT-5.6 tiered pricing is live. Terra matches GPT-5.5 quality at half the price; Luna is the cheapest GPT-5.x at $1/$6.
  • The open-weight ceiling jumped to 2.8T with Kimi K3, but the custom license and 64+ GPU requirement limit practical adoption. Inkling offers a US-origin Apache 2.0 alternative at 975B.
  • Gemini 3.5 Pro is still missing. Google shipped 3.6 Flash instead. Pre-training for Gemini 4 has begun.

August 2026 Update

Early August brought a Chinese-model offensive and a video-generation surge; the week of August 5–11 added a Meta open-source pivot, NVIDIA’s agentic play, and significant GPT-5.6 updates:

New frontier models

ModelReleasedKey change
Qwen3.8-Max2026-08-03Alibaba’s 2.4T sparse MoE (95B active). OSWorld-Verified 86.1 (tops Fable 5), PaperBench 93.0 (highest). $2/$6 — a third of Opus 5. Open weights arriving August 12.
Muse Spark 1.22026-08-05Meta’s coding flagship, co-trained with Muse Code terminal agent. TB 2.1 82.9%, AA Intelligence Index 54. $1.25/$4.25 standard; Contributor tier $0.10/$0.20 (cheapest frontier inference, shares prompts for training). 1M context.

New open-weight models

ModelReleasedKey change
Muse Glimmer2026-08-10Meta’s first Apache 2.0 model — 30B dense multimodal agent. Runs on 24GB VRAM (single consumer GPU). DFlash speculative decoding ships with weights (3.1× on RTX 5090). Positioned for local agentic deployment.
Nemotron 3.5 Lightning2026-08-11NVIDIA’s highest-efficiency model for long-running agentic workloads. Ships with NeMo Switchyard — open-source multi-model routing library. Available as NIM, HuggingFace, OpenRouter from day one.

GPT-5.6 mid-August updates

GPT-5.6 received three significant changes:

  • Sol — improved reasoning slider for granular compute allocation; now responds to “think more/less” in natural language.
  • Luna — promoted to free default for all ChatGPT users (unlimited text, 25 image gens/day); added “Think” button for on-demand extended reasoning.
  • GPT-5.6-Cyber — new security-hardened variant fine-tuned on adversarial datasets; 95% detection rate on prompt injection benchmarks. Available via API at Sol pricing.

Cross-modal expansion

The LLM landscape is increasingly entangled with video and multimodal models:

ModelReleasedKey change
MiniMax H32026-07-3133B omni-modal video model with native audio — first major open-weight video model. Extends MiniMax beyond the M3 LLM.
FLUX 3 Video2026-08-04 (GA)BFL’s unified model goes GA with pricing ($0.06–0.53/s). Video generation becomes a commodity API.
Grok Imagine Video 1.52026-07-31 updatexAI adds text-to-video, native 1080p and references. $0.08/s API.
Grok Imagine Image 2.02026-08-07xAI’s editing-first image model — Arena #2 in T2I and editing. Magic wand, multi-reference compositing, templates.
Wan 3.02026-08-06Alibaba’s 30-second video model with document/web page inputs. $0.05–0.20/s. Closed-weight successor to open Wan 2.7.
MAGI-2 Preview2026-08-05Sand.ai 114B MoE, Apache 2.0, 10-second clips with native audio. AA Video Arena #6.
SeedRealtime2026-08-05ByteDance’s audio-visual full-duplex LLM. Deployed in Doubao app — real-time watch+listen+speak.

Updated takeaways

  • Meta pivots to open source. Muse Glimmer (Apache 2.0) is Meta’s first fully permissive model — a strategic reversal after Llama’s restrictive community license. Muse Spark 1.2 (open weights announced) will follow. The Contributor tier ($0.10/$0.20 for frontier inference in exchange for training data) is a novel pricing model worth watching.
  • Qwen 3.8-Max open weights arrive August 12. ModelScope shows a countdown. Combined with Kimi K3 (2.8T, July), the open-weight frontier is in a sprint. A 27B variant (Qwen3.8-27B) follows shortly after.
  • Local agentic models are now a category. Muse Glimmer (30B, 24GB VRAM) and Nemotron 3.5 Lightning (NeMo Switchyard routing) both target always-on agents that run on local or enterprise hardware — a category that didn’t exist three months ago.
  • GPT-5.6 gets wider, not taller. OpenAI is expanding horizontally (Luna free for all, Cyber for security) rather than pushing a new number. The reasoning slider on Sol and the Think button on Luna both address the “how much compute” question differently than any rival.
  • Video input is the new frontier. Wan 3.0’s document-to-video and MAGI-2’s permissive video generation join FLUX 3 and MiniMax H3 — four new video models in one week, eclipsing text-to-image as the hot modality.
  • The frontier pricing ladder is clearer: Fable 5 ($10/$50) for max capability → Opus 5 ($5/$25) for near-Fable at half price → Qwen 3.8-Max ($2/$6) for agentic workloads on a budget → Muse Spark 1.2 ($1.25/$4.25) for coding-focused value → Sonnet 5 ($2/$10) for workhorse tasks.

Sources

  1. Anthropic — Claude models overview
  2. OpenAI — Introducing GPT-5.5
  3. Google — Gemini 3.5
  4. xAI — Grok 4.3 docs
  5. Meta — Llama 4 herd
  6. DeepSeek — V4 release notes
  7. Alibaba Cloud — Qwen3.7: The Agent Frontier
  8. Mistral AI — Introducing Mistral 3
  9. Artificial Analysis — Intelligence Index
  10. Build Fast with AI — Best AI Models June 2026
  11. Alibaba Cloud — Qwen3.8-Max: A New Bar for Coding and Cowork

Changelog: 2026-08-11 — August update 2: added Muse Spark 1.2, Muse Glimmer, Nemotron 3.5 Lightning; GPT-5.6 Sol/Luna/Cyber mid-August updates; Qwen 3.8-Max open-weight countdown; Grok Imagine Image 2.0, Wan 3.0, MAGI-2 to cross-modal expansion; updated Meta and Nemotron family sections. 2026-08-05 — August update: added Qwen 3.8-Max, cross-modal expansion section (MiniMax H3, FLUX 3 Video GA, Grok Imagine Video 1.5, SeedRealtime). 2026-07-29 — July update: added Opus 5, GPT-5.6 GA, Grok 4.5, Gemini 3.6 Flash, Kimi K3, Inkling. 2026-06-17 — initial publication.


Last updated: 2026-08-11