Grok Imagine 1.5
xAI's flagship video generation model — cinematic motion, synchronized audio, and 15-second 1080p clips with privacy-first processing on Venice.
Generate videoGet API keyWhat is Grok Imagine 1.5?
Grok Imagine 1.5 is xAI's flagship video generation model, released in June 2026. It creates cinematic 15-second videos from text, images, or reference inputs, with synchronized audio, improved physics, and support for 1080p resolution. The model powers both text-to-video and image-to-video workflows with enhanced speed and realism.
Use Grok Imagine 1.5 privately on Venice
On Venice, Grok Imagine 1.5 runs with zero retention — your prompts, images, and generated videos are never stored or profiled. This means full creative sovereignty: you generate content privately, without surveillance or data reuse. The model’s fast tier ensures low-latency output, ideal for iterative creative work under strict privacy.
What can Grok Imagine 1.5 do?
- •Supports 15-second clips — 50% longer than the previous version — enabling more complete scenes in single-pass generation.
- •1080p resolution output, bringing AI video into broadcast-quality range for professional workflows.
- •Audio is generated in the same pass as video, with clear speech and precise sync to on-screen action.
- •Strong motion coherence and physics — fewer warps, more believable weight and momentum over time.
- •Fast generation tier cuts latency by ~40%, producing 6-second 720p clips in about 25 seconds.
- •Proprietary model — no open weights, so self-hosting or fine-tuning is not possible.
- •No end-to-end encryption or TEE protection on Venice, though prompts are not retained.
- •Reference-to-video variant limited to 480p and 720p — 1080p not supported in this mode.
- •Text-to-video requires a detailed prompt to match the fidelity of image-to-video inputs.
Sample outputs
Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.
Slow aerial drone shot gliding over a misty mountain valley at golden hour, sunlight piercing clouds onto a winding river, ancient pine forests on either side, ultra-smooth motion, professional color grading, atmospheric haze, 4K
Calm ocean waves rolling onto a black-sand beach at sunrise, golden light on wet sand, foam dissolving into the shore, a single silhouetted figure at the waterline, smooth continuous forward push, serene cinematic atmosphere
A slow tracking shot through a rain-soaked Tokyo street at night, neon reflecting in puddles, steam rising from a food cart, a person with a translucent umbrella, shallow depth of field, teal-and-orange grade, smooth steady camera
Grok Imagine 1.5 model variants
Grok Imagine 1.5 runs on Venice as 3 variants of the same underlying model — pick by what you're starting from: a written prompt, a still image, reference images, or an existing clip. Each variant is its own model id on the API; the generation quality is the same across the family.
| Variant | What it is | Clip lengths | Resolutions | Aspect ratios | Audio | Model ID |
|---|---|---|---|---|---|---|
| Text to Videoflagship | Generate a clip from a written prompt | 1s – 15s | 480p, 720p, 1080p | 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16 | grok-imagine-1-5-text-to-video-private | |
| Image to Video | Animate a still image into motion | 1s – 15s | 480p, 720p, 1080p | — | grok-imagine-1-5-image-to-video-private | |
| Reference to Video | Keep a subject consistent using reference images | 1s – 15s | 480p, 720p | 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16 | grok-imagine-1-5-reference-to-video-private |
Capability data comes straight from the Venice model API and refreshes with every catalog ingest. The specs and pricing on this page are captured from the flagship variant; pass the model id of the variant you want to the API.
Grok Imagine 1.5 Text to Video
Generate a clip from a written prompt. Supports clips of 1s – 15s, 480p, 720p, 1080p output, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16 aspect ratios, with native audio.
grok-imagine-1-5-text-to-video-privateGrok Imagine 1.5 Image to Video
Animate a still image into motion. Supports clips of 1s – 15s, 480p, 720p, 1080p output, with native audio.
grok-imagine-1-5-image-to-video-privateGrok Imagine 1.5 Reference to Video
Keep a subject consistent using reference images. Supports clips of 1s – 15s, 480p, 720p output, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16 aspect ratios, with native audio.
grok-imagine-1-5-reference-to-video-privateHow to use Grok Imagine 1.5 via API
Venice exposes this model through the REST API. Queue a generation with grok-imagine-1-5-text-to-video-private.
curl https://api.venice.ai/api/v1/video/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-imagine-1-5-text-to-video-private",
"prompt": "Aerial drone shot over a misty mountain valley at golden hour"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.Specifications
Pricing
Pay per clip on Venice — price scales with resolution and duration (1s–15s), from $0.09.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
Grok Imagine 1.5 vs alternatives
| Model | Max resolution | Strongest at | Open weights | Price (Venice) |
|---|---|---|---|---|
| Grok Imagine 1.5 | 1080p | Cinematic motion & audio sync | No | from $0.09 |
| Kling O3 Pro | 1080p | Text-to-video realism | No | from $0.46 |
| Wan 2.7 Enhanced | 1080p | Open-source flexibility | Yes | from $0.68 |
| Vidu Q3 | 1080p | Long-duration consistency | No | from $0.27 |
xAI's fastest, highest-fidelity video model with audio and 15-second clips.
What is Grok Imagine 1.5 good for?
- •Creating short-form social content with synchronized audio and cinematic motion.
- •Animating concept art or storyboards into dynamic previews for film and game dev.
- •Generating branded video clips from still assets with consistent character rendering.
- •Prototyping ad creatives using reference images for character and style continuity.
- •Producing fast-turnaround visual content for news, commentary, or educational media.
Prompting tips
- •Use natural language to describe camera moves, pacing, and sound design for richer output.
- •For image-to-video, start with high-resolution, well-lit source images for best fidelity.
- •Chain short clips together using consistent reference images to build longer scenes.
- •Specify timing cues like 'slow push-in' or 'sudden cut' to guide motion structure.
- •Include audio intent: 'wind howling', 'dialogue begins at 3s', or 'tense music builds'.
Version history
Initial release — 10s clips, 720p max, no audio.
CurrentCurrent — 15s clips, 1080p, audio, faster generation.
Frequently asked questions
Grok Imagine 1.5 is xAI's flagship video generation model, released in June 2026. It creates up to 15-second cinematic videos from text, images, or reference inputs, with synchronized audio, improved physics, and support for 1080p resolution. It powers both text-to-video and image-to-video workflows.
On Venice, pricing starts at $0.09 per clip for 480p at 1 second, scaling with resolution and duration. A 1080p 15-second clip costs $4.35. You pay per generation — no subscription required.
No. Grok Imagine 1.5 is a proprietary model developed by xAI. It is not open source, and weights are not available for self-hosting or fine-tuning.
Yes. The variant 'grok-imagine-1-5-text-to-video-private' generates video directly from a written prompt, supporting resolutions up to 1080p and clip lengths from 1 to 15 seconds.
Yes. The variant 'grok-imagine-1-5-image-to-video-private' animates a still image into motion, preserving lighting and detail while adding cinematic movement and sound.
Yes. The variant 'grok-imagine-1-5-reference-to-video-private' uses reference images to maintain subject consistency across motion, ideal for character-locked animation in 480p and 720p.
Grok Imagine 1.5 supports 480p, 720p, and 1080p resolutions across all variants. The reference-to-video mode does not support 1080p.
Yes. All variants generate synchronized audio — including sound effects, ambience, and dialogue — in the same pass as video, with clear speech and precise timing.
Grok Imagine 1.5 excels in cinematic motion, audio sync, and image-to-video fidelity, while Kling O3 Pro leads in text-to-video realism. Grok also offers longer clips (15s vs 10s) and faster generation, but Kling has stronger text rendering in scenes.
Related models
Run Grok Imagine 1.5 privately.
No prompt logging. No data used for training. Free to start — no credit card.
