
Text to Video
Type a prompt, get a cinematic clip with sound.
- Generates video from text alone, no image or footage needed.
- Native audio comes with the clip.
- Up to 30 seconds, 1080p, in the aspect ratio you choose.
State-of-the-art cinematic video generation.
Three ways in
Generate up to 30 seconds of cinematic video with native audio from text, an image, or up to 50 references. Available on Venice.

Type a prompt, get a cinematic clip with sound.

Turn a single image into motion.

Guide the video with your own references.
Specs
Features
Native audio-video generation means dialogue, ambience, and effects come with the clip. No separate sound pass.
Up to 30 seconds in a single generation for real storytelling, not just 5-second tests.
Tuned for film-grade motion and lighting.
Stronger multilingual instruction following, including directional camera language.
Reference-driven generation keeps characters and style locked, then edit and extend without starting over.
Generated on Venice with no training on your inputs.

Lumara Film Festival

Everything above was built for exactly this. Lumara is an independent film festival for the artists shaping the next era of cinema, presented by MoonPay and Venice. Seven awards, $100,000 in cash prizes and matching Venice credits. All visual footage is AI-generated, on any platform you choose.
Bring a story with a clear point of view and the nerve to take risks.
Submissions close September 15, 2026.
Anonymous on Venice.