Video Models

FLUX 3

BFL's unified multimodal model — video (up to 20s with native audio), image, and action prediction from one set of weights, built on Self-Flow. Video GA as of August 4, 2026.

FLUX 3 is Black Forest Labs’ July 23, 2026 step from image generation into full multimodal visual intelligence. Where FLUX.2 was image-only, FLUX 3 is a single unified model that jointly learned from images, video and audio — generating video with sound, editing images, rendering readable text and even predicting robot actions, all from one set of weights.

The headline capability: 20-second video with natively synchronized audio in a single generation — a first for BFL and competitive with the longest clips from Veo 3.1 and Seedance 2.0. The model is built on Self-Flow, BFL’s approach to self-supervised flow matching across modalities, and is already being tested by Canva, Burda, Magnific, Krea and Picsart.

What’s available

FLUX 3 is rolling out in stages. As of August 5, 2026:

CapabilityAccessStatus
FLUX 3 Video (video + native audio)BFL API + select partnersGenerally Available (August 4, 2026)
FLUX 3 Action (robotics action prediction)Selected research & commercial partnersActive (began with mimic robotics)
FLUX 3 Image (synthesis + editing)API + private weight accessComing in the following weeks
FLUX 3 Dev (open-weight backbone for all modalities)Open weightsPlanned — no confirmed timeline

Pricing (FLUX 3 Video)

Published August 4, 2026. Audio is included in all tiers.

ModeDraft (HD)HD (720p)Full HD (1080p)
Text-to-video / Image-to-video$0.06/s$0.17/s$0.29/s
Video-to-video$0.41/s$0.53/s

Draft mode generates a lower-fidelity preview at $0.06/s — iterate on direction before committing to a full-quality render. A 20-second draft costs $1.20; a 5-second HD clip costs $0.85.

Video generation

FLUX 3 Video supports multiple generation modes:

  • Text-to-video — up to 20 seconds with synchronized audio
  • Image-to-video — start from a reference frame
  • Video-to-video — carry central elements forward
  • Keyframe-to-video — define specific moments and interpolate
  • Multi-shot clip chaining — agent-driven sequences linking individual clips into longer narratives
  • Multilingual dialogue generation

BFL highlights particular strength in human facial expressions, matching sounds to physical events (a door slamming sounds like a door, not generic ambience) and multilingual capabilities.

Competitive signals

BFL’s own Elo testing (published at GA) puts FLUX 3 Video at:

  • Text-to-video Elo: 1,135 (#1 in BFL’s test, ahead of Gemini Omni Flash, MiniMax H3 and Seedance 2.0)
  • Image-to-video Elo: 1,051 (#1 in BFL’s test)

Earlier early-access evaluations showed ties with Seedance 2.0 and Gemini Omni Flash at 52% preference, and leads against Luma Ray 3.2 (93%) and Runway Gen-4.5 (77%).

Treat BFL’s own Elo numbers as vendor-reported — independent evaluations like Artificial Analysis Video Arena will provide the definitive ranking.

Architecture: Self-Flow

FLUX 3 is built on Self-Flow, BFL’s self-supervised flow matching method for training one model to both generate and understand content simultaneously. A multimodal transformer uses dedicated components to convert images, video and audio into a shared latent space, then reconstructs outputs from that representation. The March 2026 research paper (Chefer, Esser et al.) laid the foundation.

FLUX-mimic and robotics

Perhaps the most surprising piece of the launch: FLUX-mimic extends the same architecture to predict robot actions from video observations. BFL partnered with mimic robotics for the initial deployment, and Audi is testing it for manipulation tasks. This is early-stage but signals BFL’s ambition to make FLUX 3 a general-purpose visual foundation model — not just a creative tool.

How it compares

  • vs FLUX.2: FLUX.2 is image-only; FLUX 3 adds video, audio and action prediction from one model. FLUX.2 [dev] open weights are available now; FLUX 3 Dev weights are planned but unscheduled.
  • vs Seedance 2.0 / Seedance 2.5: Seedance leads the AA Video Arena on T2V and I2V. FLUX 3 ties Seedance 2.0 in early preference tests and adds native audio + image + action from one model, but lacks public pricing and broad access.
  • vs MiniMax H3: H3 offers open weights (33B, 768p local) and a broad reference system. FLUX 3 claims higher Elo and generates longer clips (20s vs 15s) but has no open weights yet.
  • vs Grok Imagine Video 1.5: Imagine Video 1.5 offers native 1080p and voice references at $0.08/s. FLUX 3 supports longer clips and more generation modes but costs more at HD quality ($0.17/s).
  • vs Veo 3.1: Veo generates videos with original audio at $0.40/s. FLUX 3 matches on duration (20s) and native audio at lower cost ($0.17/s HD) and is now generally available.
  • vs Runway Gen-4.5: FLUX 3 led Gen-4.5 by 77% in preference tests. Gen-4.5 has professional control features (motion brush, camera moves) that FLUX 3 doesn’t yet match.

Best for

  • Text-to-video and image-to-video generation with native audio
  • Keyframe-to-video transitions for defined story moments
  • Multi-shot agent-driven clip chaining for longer sequences
  • Robotics action prediction (FLUX-mimic, tested at Audi)

Pros & cons

Strengths

  • Single unified model generating video, image, audio and action predictions — not a patchwork pipeline
  • Up to 20-second video with natively synchronized audio in a single generation
  • Strong early evaluations on human facial expressions, sound-event matching and multilingual dialogue
  • Open-weight FLUX 3 Dev backbone planned — continuing BFL's open tradition

Limitations

  • Video is GA but image generation not yet available
  • Image generation not yet available (coming in following weeks)
  • Benchmarks not yet publicly available — "full results and methodology alongside broader availability"
  • No confirmed timeline for FLUX 3 Dev open weights

Frequently asked questions

When was FLUX 3 released?

Black Forest Labs announced FLUX 3 in Early Access on July 23, 2026. FLUX 3 Video became generally available on August 4, 2026 with published pricing. FLUX 3 Image and the open-weight FLUX 3 Dev backbone will follow later.

What makes FLUX 3 different from FLUX.2?

FLUX.2 was image-only. FLUX 3 is a unified multimodal foundation model jointly trained across images, video and audio using Self-Flow (self-supervised flow matching). It generates video up to 20 seconds with native audio, while also supporting image generation and even action prediction for robotics — all from one set of weights.

How long can FLUX 3 videos be?

Up to 20 seconds with natively synchronized audio in a single generation. Generation modes include text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue and agent-driven multi-shot clip chaining.

Will FLUX 3 be open source?

BFL plans to release an open-weight FLUX 3 Dev backbone for image, video, audio and action prediction. No timeline has been confirmed. The staged rollout is Video/Action first, Image next, then Dev weights last.

What is FLUX-mimic?

FLUX-mimic is a video action model for robotics applications built on the same FLUX 3 architecture. It is being tested with selected research and commercial partners, beginning with mimic robotics, and has been tested at Audi for manipulation tasks.

Sources

Last updated: 2026-08-05 · Specs and pricing change fast — verify on the vendor's site before relying on them.