ImageAnonymized

Wan 2.7

Alibaba's Wan 2.7 is a unified image and video generation model with open weights, native audio, and instruction-based editing — built for production workflows.

Maker
Alibaba
Modality
Image
Resolution
1080p
Open weights
Yes — Apache 2.0

Overview

What is Wan 2.7

Wan 2.7 is Alibaba's open-weight text-to-image and image-to-video model, released on April 1, 2026. It supports high-resolution image generation, 1080p video up to 15 seconds, native audio co-generation, and instruction-based video editing — all under an Apache 2.0 license for commercial use.

Running it privately on Venice

On Venice, Wan 2.7 runs with full prompt anonymity — your inputs are not stored or profiled. This enables private, permissionless creation using one of the few open-weight models capable of both high-quality image and video synthesis. You gain sovereign control over visual content without Big Tech surveillance.

AnonymizedNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Assessment

Strengths and limitations

Strengths
  • Open-weight under Apache 2.0: you can self-host, fine-tune, or integrate without licensing risk.
  • Unified model for both image and video generation, reducing pipeline fragmentation.
  • Supports instruction-based video editing: modify existing clips with natural language commands.
  • Native audio co-generation with phoneme-level lip sync and voice cloning capabilities.
  • Flexible input modes: text, image, or reference video to drive generation.
Limitations
  • Max clip length capped at 15 seconds: shorter than some cinematic use cases require.
  • 1080p resolution limit; no native 4K output like in Kling 3.0 or Sora 2.
  • No tool use, web search, or reasoning capabilities — strictly multimodal generation.
  • Still requires prompt engineering for consistent character or object persistence.

Samples

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

In-image textIn-image text

A retro travel poster with the bold headline "VENICE" in large condensed serif type, sunset color palette, clean layout

PhotorealismPhotorealism

Photorealistic close-up portrait of a weathered fisherman at golden hour, 85mm lens, shallow depth of field, natural skin texture

Instruction-followingInstruction-following

A small red cube balanced on top of a large glossy blue sphere, with a green cone to the right, plain light-grey studio background

Illustration styleIllustration style

Cozy watercolor illustration of a hillside village in autumn, warm tones, soft paper texture

Compare every image model on these prompts

Capabilities

What it supports

  • Text to image
  • Image to image

Specifications

Datasheet

Maker
Alibaba
Released
April 1, 2026
Modality
Text-to-image, image-to-video, reference-to-video
Max resolution
1080p
Clip length
Up to 15 seconds
Open weights
Yes — Apache 2.0
Aspect ratios
1:1, 3:2, 16:9, 21:9, 9:16, 2:3, 3:4, 4:5
Prompt limit
3,000 chars
Privacy on Venice
Anonymized — prompts not stored
Available on Venice since
Apr 2026
License
Apache 2.0

API

Call it from your code

Venice exposes this model through the REST API. Queue a generation with the model id.

curl https://api.venice.ai/api/v1/image/generate \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "wan-2-7-text-to-image",
    "prompt": "A serene mountain lake at dawn, photorealistic"
  }' --output image.png

Pricing

What it costs on Venice

Flat per-image pricing on Venice: $0.04 per generation.

Generation
$0.04
Per image
Upscale 2×
$0.02
Per image
Upscale 4×
$0.08
Per image

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

Alternatives

How it compares

ModelMax resolutionStrongest atOpen weightsPrice (Venice)
Wan 2.71080pFlexible video workflowsNo$0.04 / image
Grok Imagine High Quality (SOTA)4KCinematic qualityNofrom $0.06 / image
Chroma1080pSpeed & open weightsYes$0.01 / image
Lustify v81080pAdult content generationNo$0.01 / image

Open-weight advantage with editing and reference inputs.

Use cases

What it is good for

  1. 01Branded short-form video with consistent voice and visuals across clips.
  2. 02Social media content where native audio and lip sync improve engagement.
  3. 03Internal creative tools that require self-hosting or on-premise deployment.
  4. 04Video editing workflows driven by natural language instructions.
  5. 05Multimodal storytelling using image, text, and audio in a single pipeline.

Prompting

Getting better results

Use reference images or videos to lock in style, motion, or composition.

Specify audio requirements clearly: 'voiceover in Spanish', 'background music: ambient synth'.

For editing, describe the change precisely: 'zoom in on the character's face', 'replace the background with a forest'.

Leverage aspect ratio support for platform-specific formats: 9:16 for TikTok, 16:9 for YouTube.

Version history

Wan 2.6
2025-12

Predecessor with basic video generation.

Wan 2.7
2026-04

Current — adds open weights, audio, editing, reference inputs.

FAQ

Frequently asked questions

Wan 2.7 is Alibaba's open-weight multimodal model for text-to-image and image-to-video generation, released on April 1, 2026. It supports 1080p video up to 15 seconds, native audio, voice cloning, and instruction-based editing under an Apache 2.0 license.

On Venice, Wan 2.7 costs $0.04 per image generation. Upscaling is $0.02 for 2× and $0.08 for 4×. Video generation pricing is not included in this tier.

Wan 2.7 is not free, but it is open-weight under the Apache 2.0 license, meaning you can download, modify, and commercially use the model weights without restriction.

Yes. Wan 2.7 supports text-to-video, image-to-video, and reference-to-video generation with up to 15 seconds of 1080p output, including native audio and lip sync.

Yes. Wan 2.7 supports instruction-based video editing — you can upload a clip and apply natural language commands like 'zoom in' or 'change background' to modify it.

No. Wan 2.7 is a pure generation model and does not support tool use, web search, or reasoning. It focuses on high-quality image and video synthesis from text, image, or reference inputs.

Wan 2.7 excels in open-weight flexibility, video editing, and native audio, while GPT Image 2 leads in in-image text accuracy and 4K resolution. Choose Wan 2.7 for editable, audio-rich video; GPT Image 2 for precise, text-heavy visuals.

Wan 2.7 supports up to 1080p resolution for both images and videos, with aspect ratios including 1:1, 16:9, 9:16, and others optimized for social and cinematic formats.

Wan 2.7 is not uncensored — it includes standard content filters. However, its open-weight nature allows developers to deploy modified versions without filters in permissionless environments.

Run Wan 2.7 privately

No prompt logging. No data used for training.