Text & reasoning
Large Language Models
Frontier and open-weight language models for chat, reasoning, coding and agents — compared by price, context window and benchmark.
Muse Glimmer
Meta's first Apache 2.0 model — a 30B dense multimodal agent optimized for local deployment on consumer GPUs, with DFlash speculative decoding for fast inference.
Muse Spark 1.2
Meta's coding-focused flagship — co-trained with Muse Code terminal agent, TB 2.1 82.9%, $1.25/$4.25 standard pricing, 1M context, open weights announced.
Qwen3.8-Max
Alibaba's new flagship — 2.4T-parameter sparse MoE (95B active), multimodal, top scores on agentic/research benchmarks, $2/$6 pricing, with open weights promised.
Claude Opus 5
Anthropic's July 2026 flagship — near-Fable intelligence at Opus pricing ($5/$25), SOTA on Frontier-Bench, ARC-AGI-3 and OSWorld 2.0, with thinking on by default.
Gemini 3.6 Flash
Google's July 2026 workhorse model — 17% fewer output tokens than 3.5 Flash, cheaper at $1.50/$7.50, optimized for agentic coding and knowledge work.
Kimi K3
The world's first open 3T-class frontier model — 2.8T parameters (104B active MoE), 1M context, native vision, open weights under a revenue-tiered custom license.
Grok 4.5
xAI's July 2026 coding-and-agent flagship — 500K context, $2/$6 per 1M, default in Grok Build and Cursor, trained alongside Cursor.
Claude Sonnet 5
Anthropic's most agentic Sonnet yet — near-Opus quality for coding, agents and knowledge work at $2/$10 introductory pricing, and the new default model on claude.ai.
GPT-5.6
OpenAI's GPT-5.6 family — flagship Sol, balanced Terra and fast Luna — with max/ultra reasoning modes and tiered pricing. Generally available since July 9, 2026.
Sakana Fugu
Tokyo-based Sakana AI's June 2026 multi-agent orchestration system — shipped as a single OpenAI-compatible model that dynamically routes each task to a swappable pool of frontier and open models.
GLM-5.2
Z.ai's open-weight long-horizon flagship — a 753B-param MoE with a stable 1M context under an MIT license, beating GPT-5.5 on several coding benchmarks.
Kimi K2.7 Code
Moonshot's open-weight agentic coding flagship — a 1T-param MoE (32B active) with 256K context and native image/video input, thinking ~30% leaner than K2.6.
Claude Fable 5
Anthropic's first Mythos-class model for general use — state-of-the-art on nearly every benchmark, with safety classifiers that fall back to Opus 4.8 on sensitive topics. Suspended by US export controls June 12–30, back globally since July 1.
Claude Opus 4.8
The June 2026 overall intelligence leader and top model for agentic coding, with the most careful hallucination calibration of any frontier model.
DeepSeek V4-Pro
The open-weight value champion — near-frontier coding (SWE-V ~85%) under an MIT license with a 1M-token context at a fraction of closed-model prices.
GPT-5.5
OpenAI's unified reasoning-and-chat flagship — the strongest model for terminal-native agentic work, natively omnimodal across text, image, audio and video.
Gemini 3.1 Pro
The highest-value closed frontier model — top-tier reasoning and native multimodality at roughly $2 per 1M input tokens with a 1M-token context.
Nemotron 3.5 Lightning
NVIDIA's highest-efficiency open model for long-running agentic AI workloads — built by the Nemotron Coalition, available with NeMo Switchyard for intelligent multi-model routing.
Inkling
Mira Murati's first model — a 975B/41B-active open-weight MoE with Apache 2.0 license, native text/image/audio, 1M context, designed as a base for fine-tuning.
MiniMax M3
The first open-weight model to combine frontier coding, a 1M-token context and native multimodality — a 428B MoE (22B active) built on sparse attention.
Ornith-1.0
Ornith-1.0 is DeepReinforce's open-source, self-improving LLM family for agentic coding — from a 9B edge model to a 397B MoE that rivals Claude Opus 4.7 on SWE-Bench.
Qwen3.7-Max
Alibaba's proprietary "Agent Frontier" flagship — top knowledge scores and long-horizon autonomy across a 1M-token context.
Grok 4.3
xAI's budget-frontier model — strong agentic tool calling and a self-reported low hallucination rate inside a 1M-token window at $1.25 / $2.50 per 1M.