Directory
AI Models
The models worth knowing across five modalities — each with sourced specs, current pricing and an honest verdict. Filter by type below.
Muse Glimmer
Meta's first Apache 2.0 model — a 30B dense multimodal agent optimized for local deployment on consumer GPUs, with DFlash speculative decoding for fast inference.
Muse Spark 1.2
Meta's coding-focused flagship — co-trained with Muse Code terminal agent, TB 2.1 82.9%, $1.25/$4.25 standard pricing, 1M context, open weights announced.
Qwen3.8-Max
Alibaba's new flagship — 2.4T-parameter sparse MoE (95B active), multimodal, top scores on agentic/research benchmarks, $2/$6 pricing, with open weights promised.
Claude Opus 5
Anthropic's July 2026 flagship — near-Fable intelligence at Opus pricing ($5/$25), SOTA on Frontier-Bench, ARC-AGI-3 and OSWorld 2.0, with thinking on by default.
Gemini 3.6 Flash
Google's July 2026 workhorse model — 17% fewer output tokens than 3.5 Flash, cheaper at $1.50/$7.50, optimized for agentic coding and knowledge work.
Kimi K3
The world's first open 3T-class frontier model — 2.8T parameters (104B active MoE), 1M context, native vision, open weights under a revenue-tiered custom license.
Grok 4.5
xAI's July 2026 coding-and-agent flagship — 500K context, $2/$6 per 1M, default in Grok Build and Cursor, trained alongside Cursor.
Claude Sonnet 5
Anthropic's most agentic Sonnet yet — near-Opus quality for coding, agents and knowledge work at $2/$10 introductory pricing, and the new default model on claude.ai.
GPT-5.6
OpenAI's GPT-5.6 family — flagship Sol, balanced Terra and fast Luna — with max/ultra reasoning modes and tiered pricing. Generally available since July 9, 2026.
Sakana Fugu
Tokyo-based Sakana AI's June 2026 multi-agent orchestration system — shipped as a single OpenAI-compatible model that dynamically routes each task to a swappable pool of frontier and open models.
GLM-5.2
Z.ai's open-weight long-horizon flagship — a 753B-param MoE with a stable 1M context under an MIT license, beating GPT-5.5 on several coding benchmarks.
Kimi K2.7 Code
Moonshot's open-weight agentic coding flagship — a 1T-param MoE (32B active) with 256K context and native image/video input, thinking ~30% leaner than K2.6.
Claude Fable 5
Anthropic's first Mythos-class model for general use — state-of-the-art on nearly every benchmark, with safety classifiers that fall back to Opus 4.8 on sensitive topics. Suspended by US export controls June 12–30, back globally since July 1.
Claude Opus 4.8
The June 2026 overall intelligence leader and top model for agentic coding, with the most careful hallucination calibration of any frontier model.
DeepSeek V4-Pro
The open-weight value champion — near-frontier coding (SWE-V ~85%) under an MIT license with a 1M-token context at a fraction of closed-model prices.
GPT-5.5
OpenAI's unified reasoning-and-chat flagship — the strongest model for terminal-native agentic work, natively omnimodal across text, image, audio and video.
Gemini 3.1 Pro
The highest-value closed frontier model — top-tier reasoning and native multimodality at roughly $2 per 1M input tokens with a 1M-token context.
Nemotron 3.5 Lightning
NVIDIA's highest-efficiency open model for long-running agentic AI workloads — built by the Nemotron Coalition, available with NeMo Switchyard for intelligent multi-model routing.
Inkling
Mira Murati's first model — a 975B/41B-active open-weight MoE with Apache 2.0 license, native text/image/audio, 1M context, designed as a base for fine-tuning.
MiniMax M3
The first open-weight model to combine frontier coding, a 1M-token context and native multimodality — a 428B MoE (22B active) built on sparse attention.
Ornith-1.0
Ornith-1.0 is DeepReinforce's open-source, self-improving LLM family for agentic coding — from a 9B edge model to a 397B MoE that rivals Claude Opus 4.7 on SWE-Bench.
Qwen3.7-Max
Alibaba's proprietary "Agent Frontier" flagship — top knowledge scores and long-horizon autonomy across a 1M-token context.
Grok 4.3
xAI's budget-frontier model — strong agentic tool calling and a self-reported low hallucination rate inside a 1M-token window at $1.25 / $2.50 per 1M.
FLUX.2
The open-weight champion for photorealism, text rendering and developer workflows — top quality at a few cents per image via API or self-hosting.
Nano Banana 2
Google's Gemini 3 Pro Image model — 4K photorealism free inside the Gemini app, with near-perfect text rendering and conversational editing.
Midjourney v8.1
The industry leader for raw artistic quality and aesthetic polish — a subscription-only image generator beloved by designers and concept artists.
Grok Imagine Image 2.0
xAI's editing-first image model — Arena #2 in both T2I and image editing behind GPT-Image-2, with magic wand, segmentation, multi-reference compositing and templates.
GPT Image 2
OpenAI's image model — the strongest at prompt instruction-following and conversational editing, native to the ChatGPT and OpenAI API stack.
Boogu-Image 0.1
Apache-2.0 open-source image generation and editing family (Base / Turbo / Edit) with strong bilingual text rendering and near closed-source quality — trained on roughly 10× less data and fully self-hostable.
Recraft V4
The design-and-brand specialist — vector/SVG output, brand-style consistency and strong typography built for professional design work.
Ideogram 3.0
The specialist for typography and logo design — the most reliable model for correct, legible text inside generated images.
Imagen 4
Google's photorealism leader — best-in-class typography and people, available cheaply through Gemini Advanced or per-image on Vertex AI.
Stable Diffusion 3.5
The open-weights standard for local, private and customizable image generation — run it on your own GPU at no per-image cost, or via Stability's API.
MiniMax H3
The first major open-weight video model with native stereo audio — a 33B omni-modal transformer generating 4–15s clips at up to 2K resolution from text, image, video and audio references.
FLUX 3
BFL's unified multimodal model — video (up to 20s with native audio), image, and action prediction from one set of weights, built on Self-Flow. Video GA as of August 4, 2026.
Seedance 2.5
ByteDance's long-sequence video model — native 30-second single-shot clips, native 4K, up to 50 multimodal reference assets, phoneme-level lip-sync and one-pass synced audio.
PixVerse V6
The best free pick for short-form video — multi-shot character consistency, native audio and 20+ cinematic lens controls, with daily free credits.
Kling 3.0
A top value-for-performance video model — excellent human motion and face consistency, native 4K output and multilingual audio.
Veo 3.1
Google's cinematic video model with native audio — the quality leader for polished, sound-synced clips, billed per second on the Gemini API.
Runway Gen-4.5
The professional's video model — rich creative controls like motion brush and camera moves, built into a mature editorial toolset.
Wan 3.0
Alibaba's unified video generation model — 30-second single-pass clips, multimodal inputs including documents and web pages, $0.05–$0.20/sec API pricing.
MAGI-2 Preview
Sand.ai's open-source 114B MoE video model — generates 10-second clips with synchronized audio using just 6B active parameters per token. Apache 2.0.
Grok Imagine Video 1.5
xAI's dedicated video generation model — native 1080p, 6–15 seconds with synced audio, image/voice references, built on the Aurora autoregressive engine.
Wan 2.7
Alibaba's flagship video model — 1080p with native audio sync, first/last-frame and multi-scene controls, from the Wan series whose open-source lineage made it the community base.
Seedance 2.0
A multimodal-control video model built for e-commerce and reference-heavy jobs — strong image-to-video fit, motion physics and multi-asset input.
Hailuo 2.3
The budget speed champion of AI video — the lowest cost per second and fast turnaround, ideal for drafts and high-volume social content.
Sora 2
OpenAI's video model known for physics realism and synchronized audio, available through ChatGPT and the API.
Suno v5.5
The leading AI music generator for full, radio-ready songs — natural vocals, real song structure, voice cloning and 30+ genres from a single text prompt.
ACE-Step
The open-source full-song generator — free, self-hostable music with vocals that runs on a single consumer GPU.
ElevenLabs v3
The industry-standard text-to-speech model — the most expressive, emotionally rich voice synthesis, with 70+ languages and multi-speaker dialogue.
ElevenLabs Music v2
The commercially safest AI music generator — built with label and publisher partnerships, with section-by-section editing, a full API and studio-grade exports.
Lyria 3
Google DeepMind's music generation model — high-fidelity 48kHz audio up to three minutes, available through Vertex AI and the Gemini API.
MiniMax Speech 2.8
A top-ranked text-to-speech model — expressive, multilingual voice synthesis with fast voice cloning and a large built-in voice library.
Stable Audio 2.5
Stability's music-and-sound model — fast, commercially-cleared instrumental tracks and sound design with an enterprise-friendly API.
Udio v4
The audio-quality leader among AI music tools — 48kHz rendering with the cleanest instrument separation, for film and brand-grade music.
Meshy 5
The most reliable all-rounder for AI 3D — best-in-class PBR textures, fast iteration, and a deep plugin ecosystem for Blender and Unity.
Rodin Gen-2
The highest-fidelity AI 3D generator — production-ready geometry with clean quad topology, the top pick for hero assets and photorealism.
Hunyuan3D 2.1
The leading open-source 3D generator — fully open weights and training code with a production-ready PBR texture pipeline, self-hostable for free.
Hitem3D
The ultra-detailed specialist — high-resolution 3D models for miniatures, e-commerce and product visualization.
Luma Genie
A free, accessible text-to-3D generator from Luma AI — a fast, no-cost way to turn prompts into 3D models via web and Discord.
Sloyd AI
Parametric, game-ready 3D — instant low-poly assets with clean topology, tuned for real-time engines.
TRELLIS 2
Microsoft's open-source 3D generator known for high-quality Gaussian-splatting visuals and flexible output representations from images or text.
Tripo 3.0
The speed-and-value leader for game-ready 3D — clean low-poly topology and auto-rigging at the lowest cost per asset.