Gemini Omni Flash
Google's fast, multimodal video generation model — creates and edits 10-second clips from text, images, or reference media with conversational control.
Generate videoGet API keyWhat is Gemini Omni Flash?
Gemini Omni Flash is Google's multimodal video generation model, introduced in May 2026 as the first in the Omni family. It creates up to 10-second video clips from text prompts, images, or reference media, and supports conversational editing. It excels at fast, high-quality video generation grounded in real-world knowledge.
Use Gemini Omni Flash privately on Venice
On Venice, Gemini Omni Flash runs under an anonymized privacy tier — your prompts are not stored, profiled, or used for training. You retain full control over your creative inputs without surveillance. This enables permissionless, uncensored experimentation with one of the most capable video models while preserving privacy and sovereignty.
What can Gemini Omni Flash do?
- •Fast, conversational video generation with strong multimodal grounding in real-world knowledge.
- •Supports multiple input types — text, image, and reference media for consistent subject animation.
- •Enables iterative, chat-based editing of generated clips for precise creative control.
- •Optimized for speed and cost-efficiency, ideal for rapid prototyping and high-volume use.
- •Available in multiple variants (text-to-video, image-to-video, reference-to-video) for flexible workflows.
- •Closed and proprietary — no open weights, so self-hosting or fine-tuning is not possible.
- •No native audio generation — clips are silent by default.
- •Maximum clip length capped at 10 seconds, limiting use for longer-form content.
- •Not fully uncensored — content policies apply as per Google's guidelines.
Sample outputs
Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.
Slow aerial drone shot gliding over a misty mountain valley at golden hour, sunlight piercing clouds onto a winding river, ancient pine forests on either side, ultra-smooth motion, professional color grading, atmospheric haze, 4K
Calm ocean waves rolling onto a black-sand beach at sunrise, golden light on wet sand, foam dissolving into the shore, a single silhouetted figure at the waterline, smooth continuous forward push, serene cinematic atmosphere
A slow tracking shot through a rain-soaked Tokyo street at night, neon reflecting in puddles, steam rising from a food cart, a person with a translucent umbrella, shallow depth of field, teal-and-orange grade, smooth steady camera
Gemini Omni Flash model variants
Gemini Omni Flash runs on Venice as 3 variants of the same underlying model — pick by what you're starting from: a written prompt, a still image, reference images, or an existing clip. Each variant is its own model id on the API; the generation quality is the same across the family.
| Variant | What it is | Clip lengths | Resolutions | Aspect ratios | Audio | Model ID |
|---|---|---|---|---|---|---|
| Text to Videoflagship | Generate a clip from a written prompt | 4s, 6s, 8s, 10s | — | 16:9, 9:16 | gemini-omni-flash-text-to-video | |
| Image to Video | Animate a still image into motion | 4s, 6s, 8s, 10s | — | 16:9, 9:16 | gemini-omni-flash-image-to-video | |
| Reference to Video | Keep a subject consistent using reference images | 4s, 6s, 8s, 10s | — | 16:9, 9:16 | gemini-omni-flash-reference-to-video |
Capability data comes straight from the Venice model API and refreshes with every catalog ingest. The specs and pricing on this page are captured from the flagship variant; pass the model id of the variant you want to the API.
Gemini Omni Flash Text to Video
Generate a clip from a written prompt. Supports clips of 4s, 6s, 8s, 10s, 16:9, 9:16 aspect ratios.
gemini-omni-flash-text-to-videoGemini Omni Flash Image to Video
Animate a still image into motion. Supports clips of 4s, 6s, 8s, 10s, 16:9, 9:16 aspect ratios.
gemini-omni-flash-image-to-videoGemini Omni Flash Reference to Video
Keep a subject consistent using reference images. Supports clips of 4s, 6s, 8s, 10s, 16:9, 9:16 aspect ratios.
gemini-omni-flash-reference-to-videoHow to use Gemini Omni Flash via API
Venice exposes this model through the REST API. Queue a generation with gemini-omni-flash-text-to-video.
curl https://api.venice.ai/api/v1/video/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-omni-flash-text-to-video",
"prompt": "Aerial drone shot over a misty mountain valley at golden hour"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.Specifications
Pricing
Pay per clip on Venice — price scales with resolution and duration (4s–10s), from $0.55.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
Gemini Omni Flash vs alternatives
| Model | Max resolution | Strongest at | Open weights | Price (Venice) |
|---|---|---|---|---|
| Gemini Omni Flash | 720p | Multimodal input & conversational editing | No | from $0.55 |
| Kling O3 Pro | 4K | High-res cinematic output | No | from $0.46 |
| Wan 2.7 Enhanced | 1080p | Photorealistic motion | Yes | from $0.68 |
| Vidu Q3 | 2K | Longer clips up to 30s | No | from $0.27 |
Excels at combining text, image, and reference inputs with natural language editing.
What is Gemini Omni Flash good for?
- •Rapid ideation and A/B testing for ads or social content using short video clips.
- •Animating still images or product shots into motion for e-commerce or marketing.
- •Creating consistent character or object animations using reference images.
- •Iterative video editing via natural language for non-technical creators.
- •Prototyping cinematic scenes or visual effects within a 10-second window.
Prompting tips
- •Use clear, descriptive language when specifying scene changes or object transformations.
- •Include reference images when consistency in subject appearance is critical.
- •Chain edit commands in conversation to build complex sequences step-by-step.
- •Specify aspect ratio (16:9 or 9:16) early to match target platform requirements.
Version history
CurrentInitial public preview release with text-to-video and editing capabilities.
Frequently asked questions
Gemini Omni Flash is Google's fast, multimodal video generation model, released in May 2026. It creates up to 10-second video clips from text, images, or reference media and supports conversational editing through natural language.
On Venice, pricing starts at $0.55 per 4-second clip, scaling with resolution and duration. You pay per clip, with no subscription required.
No. Gemini Omni Flash is a proprietary Google model — it is not free to use at scale and does not offer open weights. You cannot self-host or modify it.
No. Gemini Omni Flash is designed for video generation and editing, not agentic tool use. It does not execute code, call APIs, or interact with external systems.
Gemini Omni Flash generates video at 720p resolution (1280×720) by default, with support for 16:9 and 9:16 aspect ratios.
No. The model generates silent video clips. Audio must be added in post-production or through external tools.
Yes. The variant 'gemini-omni-flash-image-to-video' specifically enables animating still images into motion using AI.
Yes. The 'gemini-omni-flash-reference-to-video' variant allows you to maintain subject consistency by providing reference images during generation.
Gemini Omni Flash excels in multimodal input flexibility and conversational editing, while Kling O3 Pro leads in resolution and cinematic detail. Choose Omni Flash for iterative, mixed-input workflows; Kling for high-fidelity output.
Related models
Run Gemini Omni Flash privately.
No prompt logging. No data used for training. Free to start — no credit card.
