Text & reasoning

Large Language Models

Frontier and open-weight language models for chat, reasoning, coding and agents — compared by price, context window and benchmark.

Meta open

Muse Glimmer

Meta's first Apache 2.0 model — a 30B dense multimodal agent optimized for local deployment on consumer GPUs, with DFlash speculative decoding for fast inference.

Free (open weights) / hosted via partners
Meta

Muse Spark 1.2

Meta's coding-focused flagship — co-trained with Muse Code terminal agent, TB 2.1 82.9%, $1.25/$4.25 standard pricing, 1M context, open weights announced.

$1.25 / $4.25 per 1M tokens (input / output)
Alibaba

Qwen3.8-Max

Alibaba's new flagship — 2.4T-parameter sparse MoE (95B active), multimodal, top scores on agentic/research benchmarks, $2/$6 pricing, with open weights promised.

$2 / $6 per 1M tokens (input / output)
Anthropic

Claude Opus 5

Anthropic's July 2026 flagship — near-Fable intelligence at Opus pricing ($5/$25), SOTA on Frontier-Bench, ARC-AGI-3 and OSWorld 2.0, with thinking on by default.

$5 / $25 per 1M tokens (input / output)
Google

Gemini 3.6 Flash

Google's July 2026 workhorse model — 17% fewer output tokens than 3.5 Flash, cheaper at $1.50/$7.50, optimized for agentic coding and knowledge work.

$1.50 / $7.50 per 1M tokens (input / output)
Moonshot AI open

Kimi K3

The world's first open 3T-class frontier model — 2.8T parameters (104B active MoE), 1M context, native vision, open weights under a revenue-tiered custom license.

$3 / $15 per 1M tokens (input / output)
xAI

Grok 4.5

xAI's July 2026 coding-and-agent flagship — 500K context, $2/$6 per 1M, default in Grok Build and Cursor, trained alongside Cursor.

$2 / $6 per 1M tokens (input / output)
Anthropic

Claude Sonnet 5

Anthropic's most agentic Sonnet yet — near-Opus quality for coding, agents and knowledge work at $2/$10 introductory pricing, and the new default model on claude.ai.

$2 / $10 per 1M tokens (intro, through Aug 31) · $3 / $15 standard
OpenAI

GPT-5.6

OpenAI's GPT-5.6 family — flagship Sol, balanced Terra and fast Luna — with max/ultra reasoning modes and tiered pricing. Generally available since July 9, 2026.

Sol $5/$30 · Terra $2.50/$15 · Luna $1/$6 per 1M
Sakana AI

Sakana Fugu

Tokyo-based Sakana AI's June 2026 multi-agent orchestration system — shipped as a single OpenAI-compatible model that dynamically routes each task to a swappable pool of frontier and open models.

Fugu Ultra API $5 / $30 per 1M tokens (input / output)
Z.ai open

GLM-5.2

Z.ai's open-weight long-horizon flagship — a 753B-param MoE with a stable 1M context under an MIT license, beating GPT-5.5 on several coding benchmarks.

~$1.40 / $4.40 per 1M tokens (input / output)
Moonshot AI open

Kimi K2.7 Code

Moonshot's open-weight agentic coding flagship — a 1T-param MoE (32B active) with 256K context and native image/video input, thinking ~30% leaner than K2.6.

~$0.95 / $4.00 per 1M tokens (input / output)
Anthropic

Claude Fable 5

Anthropic's first Mythos-class model for general use — state-of-the-art on nearly every benchmark, with safety classifiers that fall back to Opus 4.8 on sensitive topics. Suspended by US export controls June 12–30, back globally since July 1.

$10 / $50 per 1M tokens (input / output)
Anthropic

Claude Opus 4.8

The June 2026 overall intelligence leader and top model for agentic coding, with the most careful hallucination calibration of any frontier model.

$5 / $25 per 1M tokens (input / output)
DeepSeek open

DeepSeek V4-Pro

The open-weight value champion — near-frontier coding (SWE-V ~85%) under an MIT license with a 1M-token context at a fraction of closed-model prices.

~$1.74 / $3.48 per 1M tokens (input / output)
OpenAI

GPT-5.5

OpenAI's unified reasoning-and-chat flagship — the strongest model for terminal-native agentic work, natively omnimodal across text, image, audio and video.

$5 / $30 per 1M tokens (input / output)
Google

Gemini 3.1 Pro

The highest-value closed frontier model — top-tier reasoning and native multimodality at roughly $2 per 1M input tokens with a 1M-token context.

$2 / $4–12 per 1M tokens (input / output)
NVIDIA open

Nemotron 3.5 Lightning

NVIDIA's highest-efficiency open model for long-running agentic AI workloads — built by the Nemotron Coalition, available with NeMo Switchyard for intelligent multi-model routing.

Free (open weights) / hosted via partners
Thinking Machines Lab open

Inkling

Mira Murati's first model — a 975B/41B-active open-weight MoE with Apache 2.0 license, native text/image/audio, 1M context, designed as a base for fine-tuning.

$1.87 / $4.68 per 1M tokens (64K ctx) · $3.74 / $9.36 (256K ctx)
MiniMax open

MiniMax M3

The first open-weight model to combine frontier coding, a 1M-token context and native multimodality — a 428B MoE (22B active) built on sparse attention.

~$0.60 / $2.40 per 1M tokens (input / output)
DeepReinforce open

Ornith-1.0

Ornith-1.0 is DeepReinforce's open-source, self-improving LLM family for agentic coding — from a 9B edge model to a 397B MoE that rivals Claude Opus 4.7 on SWE-Bench.

Open weight (free to self-host)
Alibaba

Qwen3.7-Max

Alibaba's proprietary "Agent Frontier" flagship — top knowledge scores and long-horizon autonomy across a 1M-token context.

Pay-as-you-go via Alibaba Cloud Model Studio
xAI

Grok 4.3

xAI's budget-frontier model — strong agentic tool calling and a self-reported low hallucination rate inside a 1M-token window at $1.25 / $2.50 per 1M.

$1.25 / $2.50 per 1M tokens (input / output)