Video Models

Grok Imagine Video 1.5

xAI's dedicated video generation model — native 1080p, 6–15 seconds with synced audio, image/voice references, built on the Aurora autoregressive engine.

Grok Imagine Video 1.5 is xAI’s dedicated video generation model — not the Grok chatbot, but a standalone creative tool built on the Aurora autoregressive engine trained on 110,000 NVIDIA GB200 GPUs. First launched on May 31, 2026, it received a major update on July 31 adding text-to-video, native 1080p and image/voice references.

The standout feature: native 1080p at a time when most competitors cap at 720p (or charge a premium for upscaling), combined with image and voice references that let you maintain character and style consistency across clips. At $0.08 per second in the API, the pricing is competitive with mid-tier video models.

Specs

SpecValue
EngineAurora autoregressive
Resolution480p (draft), 720p, 1080p (native)
Frame rate24 FPS
Duration6–15 seconds
AudioNative, generated in same pass
Aspect ratios1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3
References1–7 images + optional voice reference
Generation speed5–30 seconds
API model IDgrok-imagine-video-1.5

Pricing

Rate
Output$0.08 / second

Available via the xAI API (grok-imagine-video-1.5). Consumer access through grok.com/imagine and iOS/Android apps — SuperGrok Heavy and Plus tiers.

Generation modes

  • Image-to-video — the original and strongest mode: provide a starting image, describe the motion
  • Text-to-video — describe the shot without a starting image (added July 31)
  • Reference-to-video — 1–7 reference images for character/style anchoring (added July 31)
  • Voice reference — clone a speaking style for lip-synced dialogue (available on request in API)

How it compares

  • vs Seedance 2.0: Seedance leads the AA Video Arena on blind preference and costs less (~$0.022–0.14/sec). Imagine Video 1.5 offers native 1080p and voice references that Seedance doesn’t.
  • vs Veo 3.1: Veo generates more cinematic output at $0.40/sec — 5× the price. Imagine Video 1.5 is faster and cheaper for iteration.
  • vs FLUX 3: FLUX 3 Video supports up to 20 seconds with a broader control surface. Imagine Video 1.5 is simpler and generally available today.
  • vs MiniMax H3: H3 offers open weights and a richer reference system (images + video + audio). Imagine Video 1.5 is faster, simpler to use, and has native 1080p.
  • vs Grok 4.5 (the LLM): Completely separate products. Grok 4.5 is a text/code LLM. Imagine Video 1.5 is a video generator. Same company, different models.

For the full video model landscape, see the video models directory.

Best for

  • Quick social and marketing video from product images
  • Character-consistent video series with image references
  • Voice-referenced clips for brand consistency
  • Text-to-video drafting and iteration at fast turnaround

Pros & cons

Strengths

  • Native 1080p generation — one of few video models offering full HD natively
  • Image and voice references for character/style consistency across clips
  • Fast generation (5–30 seconds per clip) with the Aurora autoregressive engine
  • Clean API at $0.08/sec — simple, predictable pricing

Limitations

  • No open weights — proprietary and API-only
  • 15-second maximum duration — shorter than FLUX 3 (20s) or Seedance 2.5 (30s)
  • US-first rollout for reference features — international availability lags
  • Separate product from Grok LLM — no integrated text reasoning in generation

Frequently asked questions

When was Grok Imagine Video 1.5 released?

xAI launched Grok Imagine Video 1.5 on May 31, 2026 (GA in the API on June 16). A major update on July 31, 2026 added text-to-video, native 1080p, and image/voice reference capabilities.

Is this the same as the Grok chatbot?

No. Grok Imagine Video 1.5 is a standalone video generation model, entirely separate from the Grok LLM chatbot. They share the xAI brand but serve different purposes — the video model generates clips, the chatbot handles text conversations.

How long can videos be?

6 to 15 seconds per clip, at 24 FPS. Supported resolutions are 480p (drafting) and 720p/1080p (output). Seven aspect ratios are available including 16:9, 9:16 and 1:1.

What are image and voice references?

You can provide 1–7 reference images that contribute people, objects, styles or settings to the generated video. Voice references let you clone a speaking style for lip-synced dialogue. These shipped on July 31, 2026 — initially US-only for SuperGrok Heavy/Plus, rolling out more broadly.

Sources

Last updated: 2026-08-05 · Specs and pricing change fast — verify on the vendor's site before relying on them.