Flux 3 First Last Frame
FLUX 3 First Last Frame generates smooth video transitions between two provided keyframes, with native audio and up to 20-second clips in 1080p.
curl https://api.venice.ai/api/v1/video/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "flux-3-first-last-frame-to-video",
"prompt": "Aerial drone shot over a misty mountain valley at golden hour"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.Overview
What is Flux 3 First Last Frame
Flux 3 First Last Frame is an image-to-video model by Black Forest Labs that generates a coherent video sequence by interpolating between a defined start and end image. It supports native audio, clips from 5 to 20 seconds, and outputs in 720p or 1080p, all processed without storing user prompts on Venice.
Using it anonymously on Venice
On Venice, Flux 3 First Last Frame runs with anonymized privacy — your prompts and reference images are not stored or used for training. This means you retain full sovereignty over creative direction, with zero retention of inputs. It's permissionless access to a high-fidelity video interpolator, free from Big Tech surveillance.
Specifications
Datasheet
- Maker
- Black Forest Labs
- Modality
- Image-to-video (first and last frame interpolation)
- Open weights
- No — proprietary
- License
- Proprietary
- Clip lengths
- 5s – 20s in 5s steps
- Resolutions
- 720p, 1080p
- Mode
- image-to-video
- Audio
- Yes
- References
- Up to 2 images — start and end frames
- Reference video
- Not supported
- Released
- August 4, 2026
- Aspect ratios
- auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, 9:21
- Privacy on Venice
- Anonymized — prompts not stored
- Available on Venice since
- Aug 2026
Assessment
Strengths and limitations
- Precise control via keyframes: define exact start and end states for reliable interpolation.
- Native audio generation synchronized with video motion and events (e.g., footsteps, impacts).
- Supports natural scene transitions, typography, and multilingual dialogue with lip-syncing.
- High resolution output (1080p) and long durations (up to 20s) for professional-grade clips.
- Runs on Venice with anonymized privacy: no prompt storage or profiling.
- No open weights: cannot be self-hosted or fine-tuned.
- Limited to two reference images; no support for video, audio, or document references.
- No tool use, vision beyond keyframes, or web search capabilities.
- Not uncensored: content policies apply as per Black Forest Labs' guidelines.
Use cases
What it is good for
- 01Creating smooth product transformations (e.g., before/after visuals).
- 02Generating cinematic transitions between storyboard keyframes.
- 03Producing social media clips with synchronized voiceover and motion.
- 04Animating educational content with precise start and end states.
- 05Designing UI/UX prototypes with realistic state transitions.
Prompting
Getting better results
Describe the motion path clearly — how the subject moves, how lighting shifts, or how the camera pans.
Use concise, action-oriented language to define the interpolation (e.g., 'the camera dollies in smoothly').
Specify audio cues in the prompt (e.g., 'birds chirping softly, fading as wind picks up').
Ensure both keyframes are high quality and consistent in style to avoid visual artifacts.
Match aspect ratio to platform (e.g., 9:16 for TikTok, 16:9 for YouTube).
Use 1080p for final deliverables; 720p for fast iteration.
Samples
Sample outputs
Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.
Calm ocean waves rolling onto a black-sand beach at sunrise, golden light on wet sand, foam dissolving into the shore, a single silhouetted figure at the waterline, smooth continuous forward push, serene cinematic atmosphere
A slow tracking shot through a rain-soaked Tokyo street at night, neon reflecting in puddles, steam rising from a food cart, a person with a translucent umbrella, shallow depth of field, teal-and-orange grade, smooth steady camera
Alternatives
How it compares
| Model | Best for | Max duration | Max resolution | Native audio |
|---|---|---|---|---|
| Flux 3 First Last Frame | Keyframe-to-video interpolation | 20s | 1080p | Yes |
| Kling O3 Pro | Longer cinematic clips | 15s | 1080p | Yes |
| Wan 2.7 Enhanced | Open-source flexibility | 15s | 1080p | Yes |
| Vidu Q3 | High-fidelity text-to-video | 16s | 1080p | Yes |
Flux 3 First Last Frame is the right pick when you need precise control over video start and end states — ideal for product animations, storyboards, and transitions where visual consistency is critical.
API
Call it from your code
Venice exposes this model through the REST API. Queue a generation with the model id.
curl https://api.venice.ai/api/v1/video/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "flux-3-first-last-frame-to-video",
"prompt": "Aerial drone shot over a misty mountain valley at golden hour"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.Pricing
What it costs on Venice
Pay per clip on Venice — price scales with resolution and duration (5s–20s), from $0.94.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
FAQ
Frequently asked questions
Flux 3 First Last Frame is an AI video model by Black Forest Labs that generates a video by interpolating between two provided images — a start frame and an end frame. It creates smooth transitions up to 20 seconds long with native audio, running on Venice with full privacy.
On Venice, pricing starts at $0.94 for a 720p, 5-second clip, scaling with resolution and duration. A 1080p, 5-second clip costs $1.60. You pay per generation with no subscription required.
No, Flux 3 First Last Frame is not free to use at scale and is not open source. It is a proprietary model by Black Forest Labs, available via API on Venice and other platforms.
No. The model does not support tool use, web search, or external API calls. It operates solely on the two input images and the text prompt describing the transition.
It supports 720p and 1080p output resolutions, with aspect ratios including 16:9, 9:16, 1:1, and others. The video is generated natively at the selected resolution.
Yes. It generates native audio synchronized with the video — including ambient sounds, dialogue with lip-syncing, and event-based effects like impacts or footsteps — all created jointly with the visual output.
Choose Flux 3 First Last Frame for precise keyframe control and smooth interpolation between defined states. Choose Kling O3 Pro for longer, cinematic text-to-video generation without keyframe constraints.
Use Flux 3 First Last Frame anonymously
Venice does not store your prompts. Chat history stays in your browser.