Grok Imagine Video 1.5 is xAI’s dedicated video generation model — not the Grok chatbot, but a standalone creative tool built on the Aurora autoregressive engine trained on 110,000 NVIDIA GB200 GPUs. First launched on May 31, 2026, it received a major update on July 31 adding text-to-video, native 1080p and image/voice references.
The standout feature: native 1080p at a time when most competitors cap at 720p (or charge a premium for upscaling), combined with image and voice references that let you maintain character and style consistency across clips. At $0.08 per second in the API, the pricing is competitive with mid-tier video models.
Specs
| Spec | Value |
|---|---|
| Engine | Aurora autoregressive |
| Resolution | 480p (draft), 720p, 1080p (native) |
| Frame rate | 24 FPS |
| Duration | 6–15 seconds |
| Audio | Native, generated in same pass |
| Aspect ratios | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3 |
| References | 1–7 images + optional voice reference |
| Generation speed | 5–30 seconds |
| API model ID | grok-imagine-video-1.5 |
Pricing
| Rate | |
|---|---|
| Output | $0.08 / second |
Available via the xAI API (grok-imagine-video-1.5). Consumer access through grok.com/imagine and iOS/Android apps — SuperGrok Heavy and Plus tiers.
Generation modes
- Image-to-video — the original and strongest mode: provide a starting image, describe the motion
- Text-to-video — describe the shot without a starting image (added July 31)
- Reference-to-video — 1–7 reference images for character/style anchoring (added July 31)
- Voice reference — clone a speaking style for lip-synced dialogue (available on request in API)
How it compares
- vs Seedance 2.0: Seedance leads the AA Video Arena on blind preference and costs less (~$0.022–0.14/sec). Imagine Video 1.5 offers native 1080p and voice references that Seedance doesn’t.
- vs Veo 3.1: Veo generates more cinematic output at $0.40/sec — 5× the price. Imagine Video 1.5 is faster and cheaper for iteration.
- vs FLUX 3: FLUX 3 Video supports up to 20 seconds with a broader control surface. Imagine Video 1.5 is simpler and generally available today.
- vs MiniMax H3: H3 offers open weights and a richer reference system (images + video + audio). Imagine Video 1.5 is faster, simpler to use, and has native 1080p.
- vs Grok 4.5 (the LLM): Completely separate products. Grok 4.5 is a text/code LLM. Imagine Video 1.5 is a video generator. Same company, different models.
For the full video model landscape, see the video models directory.
Best for
- Quick social and marketing video from product images
- Character-consistent video series with image references
- Voice-referenced clips for brand consistency
- Text-to-video drafting and iteration at fast turnaround
Pros & cons
Strengths
- Native 1080p generation — one of few video models offering full HD natively
- Image and voice references for character/style consistency across clips
- Fast generation (5–30 seconds per clip) with the Aurora autoregressive engine
- Clean API at $0.08/sec — simple, predictable pricing
Limitations
- No open weights — proprietary and API-only
- 15-second maximum duration — shorter than FLUX 3 (20s) or Seedance 2.5 (30s)
- US-first rollout for reference features — international availability lags
- Separate product from Grok LLM — no integrated text reasoning in generation
Frequently asked questions
When was Grok Imagine Video 1.5 released?
xAI launched Grok Imagine Video 1.5 on May 31, 2026 (GA in the API on June 16). A major update on July 31, 2026 added text-to-video, native 1080p, and image/voice reference capabilities.
Is this the same as the Grok chatbot?
No. Grok Imagine Video 1.5 is a standalone video generation model, entirely separate from the Grok LLM chatbot. They share the xAI brand but serve different purposes — the video model generates clips, the chatbot handles text conversations.
How long can videos be?
6 to 15 seconds per clip, at 24 FPS. Supported resolutions are 480p (drafting) and 720p/1080p (output). Seven aspect ratios are available including 16:9, 9:16 and 1:1.
What are image and voice references?
You can provide 1–7 reference images that contribute people, objects, styles or settings to the generated video. Voice references let you clone a speaking style for lip-synced dialogue. These shipped on July 31, 2026 — initially US-only for SuperGrok Heavy/Plus, rolling out more broadly.