Now on VeniceVideoAnonymous

Flux 3 First Last Frame

FLUX 3 First Last Frame generates smooth video transitions between two provided keyframes, with native audio and up to 20-second clips in 1080p.

For agents
curl https://api.venice.ai/api/v1/video/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "flux-3-first-last-frame-to-video",
    "prompt": "Aerial drone shot over a misty mountain valley at golden hour"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.
Model IDflux-3-first-last-frame-to-video
Maker
Black Forest Labs
Modality
Video + audio
Max duration
20 seconds
Max resolution
1080p

Overview

What is Flux 3 First Last Frame

Flux 3 First Last Frame is an image-to-video model by Black Forest Labs that generates a coherent video sequence by interpolating between a defined start and end image. It supports native audio, clips from 5 to 20 seconds, and outputs in 720p or 1080p, all processed without storing user prompts on Venice.

Using it anonymously on Venice

On Venice, Flux 3 First Last Frame runs with anonymized privacy — your prompts and reference images are not stored or used for training. This means you retain full sovereignty over creative direction, with zero retention of inputs. It's permissionless access to a high-fidelity video interpolator, free from Big Tech surveillance.

AnonymousNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Specifications

Datasheet

Maker
Black Forest Labs
Modality
Image-to-video (first and last frame interpolation)
Open weights
No — proprietary
License
Proprietary
Clip lengths
5s – 20s in 5s steps
Resolutions
720p, 1080p
Mode
image-to-video
Audio
Yes
References
Up to 2 images — start and end frames
Reference video
Not supported
Released
August 4, 2026
Aspect ratios
auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, 9:21
Privacy on Venice
Anonymized — prompts not stored
Available on Venice since
Aug 2026

Assessment

Strengths and limitations

Strengths
  • Precise control via keyframes: define exact start and end states for reliable interpolation.
  • Native audio generation synchronized with video motion and events (e.g., footsteps, impacts).
  • Supports natural scene transitions, typography, and multilingual dialogue with lip-syncing.
  • High resolution output (1080p) and long durations (up to 20s) for professional-grade clips.
  • Runs on Venice with anonymized privacy: no prompt storage or profiling.
Limitations
  • No open weights: cannot be self-hosted or fine-tuned.
  • Limited to two reference images; no support for video, audio, or document references.
  • No tool use, vision beyond keyframes, or web search capabilities.
  • Not uncensored: content policies apply as per Black Forest Labs' guidelines.

Use cases

What it is good for

  1. 01Creating smooth product transformations (e.g., before/after visuals).
  2. 02Generating cinematic transitions between storyboard keyframes.
  3. 03Producing social media clips with synchronized voiceover and motion.
  4. 04Animating educational content with precise start and end states.
  5. 05Designing UI/UX prototypes with realistic state transitions.

Prompting

Getting better results

Describe the motion path clearly — how the subject moves, how lighting shifts, or how the camera pans.

Use concise, action-oriented language to define the interpolation (e.g., 'the camera dollies in smoothly').

Specify audio cues in the prompt (e.g., 'birds chirping softly, fading as wind picks up').

Ensure both keyframes are high quality and consistent in style to avoid visual artifacts.

Match aspect ratio to platform (e.g., 9:16 for TikTok, 16:9 for YouTube).

Use 1080p for final deliverables; 720p for fast iteration.

Samples

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

Seamless loop

Calm ocean waves rolling onto a black-sand beach at sunrise, golden light on wet sand, foam dissolving into the shore, a single silhouetted figure at the waterline, smooth continuous forward push, serene cinematic atmosphere

Urban cinematic

A slow tracking shot through a rain-soaked Tokyo street at night, neon reflecting in puddles, steam rising from a food cart, a person with a translucent umbrella, shallow depth of field, teal-and-orange grade, smooth steady camera

Compare every video model on these prompts

Alternatives

How it compares

ModelBest forMax durationMax resolutionNative audio
Flux 3 First Last FrameKeyframe-to-video interpolation20s1080pYes
Kling O3 ProLonger cinematic clips15s1080pYes
Wan 2.7 EnhancedOpen-source flexibility15s1080pYes
Vidu Q3High-fidelity text-to-video16s1080pYes

Flux 3 First Last Frame is the right pick when you need precise control over video start and end states — ideal for product animations, storyboards, and transitions where visual consistency is critical.

API

Call it from your code

Venice exposes this model through the REST API. Queue a generation with the model id.

curl https://api.venice.ai/api/v1/video/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "flux-3-first-last-frame-to-video",
    "prompt": "Aerial drone shot over a misty mountain valley at golden hour"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.

Pricing

What it costs on Venice

Pay per clip on Venice — price scales with resolution and duration (5s–20s), from $0.94.

720p · 5s
$0.94
Per clip
1080p · 5s
$1.60
Per clip

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

FAQ

Frequently asked questions

Flux 3 First Last Frame is an AI video model by Black Forest Labs that generates a video by interpolating between two provided images — a start frame and an end frame. It creates smooth transitions up to 20 seconds long with native audio, running on Venice with full privacy.

On Venice, pricing starts at $0.94 for a 720p, 5-second clip, scaling with resolution and duration. A 1080p, 5-second clip costs $1.60. You pay per generation with no subscription required.

No, Flux 3 First Last Frame is not free to use at scale and is not open source. It is a proprietary model by Black Forest Labs, available via API on Venice and other platforms.

No. The model does not support tool use, web search, or external API calls. It operates solely on the two input images and the text prompt describing the transition.

It supports 720p and 1080p output resolutions, with aspect ratios including 16:9, 9:16, 1:1, and others. The video is generated natively at the selected resolution.

Yes. It generates native audio synchronized with the video — including ambient sounds, dialogue with lip-syncing, and event-based effects like impacts or footsteps — all created jointly with the visual output.

Choose Flux 3 First Last Frame for precise keyframe control and smooth interpolation between defined states. Choose Kling O3 Pro for longer, cinematic text-to-video generation without keyframe constraints.

Use Flux 3 First Last Frame anonymously

Venice does not store your prompts. Chat history stays in your browser.