Xai (via Future Tools)

Imagine Video 1.5 with References

Brief

Imagine Video 1.5 adds image and voice references, native 1080p, and combined text-to-video/image-to-video workflows: a single character image and voice can be preserved across scenes, up to seven reference images per generation. Features are live on grok.com/imagine, iOS, Android, and via the xAI API model grok-imagine-video-1.5; voice refs require a request.

Why it matters

Imagine Video 1.5 (announcement dated 2026-07-31) now supports image and voice references, native 1080p output, and both text-to-video and image-to-video generation; initial rollout in the US for SuperGrok Heavy and SuperGrok Plus begins immediately and expands to all tiers over the next few days.

Key details

  • The xAI API exposes the model grok-imagine-video-1.5 with image references, text-to-video, and native 1080p; generations can include up to seven reference images, and voice-reference support is available on request (contact-sales).
Source evidence

Jul 31, 2026

Our best video model, now with text, image, and voice references — generating up to 1080p.

When we launched Imagine Video 1.5 last month, it was our best video model yet — better motion, better physics, and better audio. Today it goes further: image and voice references, video from a prompt alone, and native 1080p generation.

Image and voice references start today in the US for SuperGrok Heavy and SuperGrok Plus on grok.com/imagine and iOS, rolling out to all tiers over the next few days.

Describe the shot — no starting image needed. Text-to-video pairs our image generation with image-to-video. Native 1080p is now supported with text-to-video and image-to-video. Text-to-video and native 1080p are generally available on grok.com/imagine, iOS, and Android.

Pass in a character image and a voice reference, and both hold — the same face and the same voice in every scene.

Character

Voice

Each reference image locks one thing in place — a face, a product, a location. Keep a character and swap the scene, keep the scene and swap the character, or hold both and change only the action. Up to seven references per generation.

Image references, text-to-video, and native 1080p are live in the xAI API with our best video model, grok-imagine-video-1.5. Voice reference support is available on request.