Text- & image-to-video

Video Models

Text- and image-to-video generators — cinematic quality, native audio and budget options, compared by cost per second.

MiniMax open

MiniMax H3

The first major open-weight video model with native stereo audio — a 33B omni-modal transformer generating 4–15s clips at up to 2K resolution from text, image, video and audio references.

API pricing varies by resolution; open weights available
Black Forest Labs

FLUX 3

BFL's unified multimodal model — video (up to 20s with native audio), image, and action prediction from one set of weights, built on Self-Flow. Video GA as of August 4, 2026.

$0.06–0.53 / second (varies by mode and resolution)
ByteDance

Seedance 2.5

ByteDance's long-sequence video model — native 30-second single-shot clips, native 4K, up to 50 multimodal reference assets, phoneme-level lip-sync and one-pass synced audio.

Not public yet (enterprise beta; general availability early July 2026)
PixVerse

PixVerse V6

The best free pick for short-form video — multi-shot character consistency, native audio and 20+ cinematic lens controls, with daily free credits.

Free daily credits · $10–$199 / month · API ~$0.025–$0.115 / sec
Kuaishou

Kling 3.0

A top value-for-performance video model — excellent human motion and face consistency, native 4K output and multilingual audio.

~$0.067 / sec (Standard) to ~$0.17 / sec (Pro)
Google

Veo 3.1

Google's cinematic video model with native audio — the quality leader for polished, sound-synced clips, billed per second on the Gemini API.

$0.40 / sec (1080p std, with audio) · $0.05–0.12 / sec (Fast/Lite)
Runway

Runway Gen-4.5

The professional's video model — rich creative controls like motion brush and camera moves, built into a mature editorial toolset.

Credit-based (~$0.10–0.25 / sec effective)
Alibaba

Wan 3.0

Alibaba's unified video generation model — 30-second single-pass clips, multimodal inputs including documents and web pages, $0.05–$0.20/sec API pricing.

$0.05–$0.20 per second (by resolution)
Sand.ai open

MAGI-2 Preview

Sand.ai's open-source 114B MoE video model — generates 10-second clips with synchronized audio using just 6B active parameters per token. Apache 2.0.

Free (open weights)
xAI

Grok Imagine Video 1.5

xAI's dedicated video generation model — native 1080p, 6–15 seconds with synced audio, image/voice references, built on the Aurora autoregressive engine.

$0.08 / second (API)
Alibaba

Wan 2.7

Alibaba's flagship video model — 1080p with native audio sync, first/last-frame and multi-scene controls, from the Wan series whose open-source lineage made it the community base.

~$6 / minute (cloud), a ~33% cut versus Wan 2.5/2.6
ByteDance

Seedance 2.0

A multimodal-control video model built for e-commerce and reference-heavy jobs — strong image-to-video fit, motion physics and multi-asset input.

~$0.022–0.14 / sec (varies by provider)
MiniMax

Hailuo 2.3

The budget speed champion of AI video — the lowest cost per second and fast turnaround, ideal for drafts and high-volume social content.

~$0.025 / sec
OpenAI

Sora 2

OpenAI's video model known for physics realism and synchronized audio, available through ChatGPT and the API.

Via ChatGPT Plus/Pro · ~$0.15 / sec (third-party API routes)