HappyHorse 1.1
Alibaba's multimodal video model that generates 1080p clips with native audio and supports reference-driven subject consistency across text, image, and reference-to-video modes.
Generate videoGet API keyWhat is HappyHorse 1.1?
HappyHorse 1.1 is Alibaba's upgraded video generation model, released in June 2026. It creates 3–15 second clips at up to 1080p from text, images, or reference photos, with native audio and lip-sync. The 1.1 update improves motion expressiveness, subject consistency, and visual fidelity over its predecessor.
Use HappyHorse 1.1 privately on Venice
On Venice, HappyHorse 1.1 runs under an anonymized privacy tier — your prompts are not stored, profiled, or retained for training. You pay per clip with credits instead of a subscription, starting at $0.46 for a 720p 3-second generation, with no Big-Tech surveillance attached to your creative workflow.
What can HappyHorse 1.1 do?
- •Native audio generation with lip-sync in a single forward pass — no separate audio render needed.
- •Strong subject consistency across shots using reference images, enabling persistent characters and products.
- •Smooth, coherent motion modeling for complex action sequences.
- •Supports 1080p and 720p output with selectable durations from 3s to 15s.
- •Improved visual fidelity and prompt adherence over HappyHorse 1.0.
- •Closed and proprietary — no open weights, so self-hosting or fine-tuning is impossible.
- •Wan 2.7 is the open-weights alternative on Venice if you need an open pipeline.
- •Clip length capped at 15 seconds; not suited for long-form storytelling without external editing.
- •Pricing per clip scales with resolution and duration, so 1080p or longer clips cost more than 720p drafts.
- •No TEE or end-to-end encryption on Venice — privacy is anonymized but not hardware-isolated.
HappyHorse 1.1 model variants
HappyHorse 1.1 runs on Venice as 3 variants of the same underlying model — pick by what you're starting from: a written prompt, a still image, reference images, or an existing clip. Each variant is its own model id on the API; the generation quality is the same across the family.
| Variant | What it is | Clip lengths | Resolutions | Aspect ratios | Audio | Model ID |
|---|---|---|---|---|---|---|
| Text to Video | Generate a clip from a written prompt | 3s – 15s | 1080p, 720p | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5 | happyhorse-1-1-text-to-video | |
| Image to Videoflagship | Animate a still image into motion | 3s – 15s | 1080p, 720p | — | happyhorse-1-1-image-to-video | |
| Reference to Video | Keep a subject consistent using reference images | 3s – 15s | 1080p, 720p | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5 | happyhorse-1-1-reference-to-video |
Capability data comes straight from the Venice model API and refreshes with every catalog ingest. The specs and pricing on this page are captured from the flagship variant; pass the model id of the variant you want to the API.
HappyHorse 1.1 Text to Video
Generate a clip from a written prompt. Supports clips of 3s – 15s, 1080p, 720p output, 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5 aspect ratios, with native audio.
happyhorse-1-1-text-to-videoHappyHorse 1.1 Image to Video
Animate a still image into motion. Supports clips of 3s – 15s, 1080p, 720p output, with native audio.
happyhorse-1-1-image-to-videoHappyHorse 1.1 Reference to Video
Keep a subject consistent using reference images. Supports clips of 3s – 15s, 1080p, 720p output, 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5 aspect ratios, with native audio.
happyhorse-1-1-reference-to-videoHow to use HappyHorse 1.1 via API
Venice exposes this model through the REST API. Queue a generation with happyhorse-1-1-image-to-video.
curl https://api.venice.ai/api/v1/video/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "happyhorse-1-1-image-to-video",
"prompt": "Aerial drone shot over a misty mountain valley at golden hour"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/video/retrieve.
# Call /video/complete after downloading if needed.Specifications
Pricing
Pay per clip on Venice — price scales with resolution and duration (3s–15s), from $0.46.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
HappyHorse 1.1 vs alternatives
| Model | Max resolution | Audio | Open weights | Price (Venice) |
|---|---|---|---|---|
| HappyHorse 1.1 | 1080p | Yes | No | from $0.46 |
| Wan 2.7 | — | — | Yes | from $0.55 |
| Kling O3 Pro | — | — | No | from $0.46 |
| HappyHorse 1.0 | — | — | No | from $0.46 |
The only Venice-hosted model in this family with native audio, lip-sync, and up to 15s reference-to-video output.
What is HappyHorse 1.1 good for?
- •Short-form social media ads and product demos with consistent branding.
- •E-commerce video content featuring persistent characters or products across clips.
- •Dialogue scenes with synced lip movement and ambient sound.
- •Reference-driven cinematic shots where visual fidelity to source images matters.
- •Rapid prototyping of motion concepts from still images or text prompts.
Prompting tips
- •Use reference images to lock in subject consistency across generations.
- •Start with 720p 3s drafts to iterate cheaply, then scale to 1080p for final delivery.
- •Describe motion and camera direction explicitly to take advantage of the model's improved motion handling.
Version history
Predecessor with weaker motion and audio sync.
CurrentCurrent — improved motion, subject consistency, and native audio.
Frequently asked questions
HappyHorse 1.1 is Alibaba's upgraded AI video generation model, released in June 2026. It produces 3–15 second video clips at up to 1080p from text prompts, still images, or reference photos, with native audio and lip-sync.
On Venice you pay per clip: from $0.46 for a 720p 3-second video up to $0.59 for a 1080p 3-second clip, with longer durations costing more. There is no subscription required.
You can try it on Venice with credits; new accounts receive welcome credits that can cover initial generations. There is no free tier that offers unlimited use.
No. HappyHorse 1.1 is closed and proprietary to Alibaba. If you need open weights, Wan 2.7 is available on Venice as an open-source alternative.
Yes. The text-to-video variant (model id happyhorse-1-1-text-to-video) generates clips directly from written prompts, supporting the same 3–15s lengths and 1080p/720p resolutions as the rest of the family.
Yes. The reference-to-video variant (model id happyhorse-1-1-reference-to-video) accepts multiple reference images to maintain subject consistency across generated frames.
HappyHorse 1.1 is the better choice for native audio, lip-sync, and reference-driven character consistency in short cinematic clips. Wan 2.7 is the choice if you need open weights and the freedom to self-host or fine-tune outside Venice.
Venice processes HappyHorse 1.1 under an anonymized privacy tier — prompts are not stored or used for training. Note that it does not currently run inside a TEE or use end-to-end encryption, so privacy is enforced by policy-level zero retention rather than hardware isolation.
It outputs 720p and 1080p video in clips ranging from 3 seconds to 15 seconds, with selectable durations in one-second increments.
Related models
Run HappyHorse 1.1 privately.
No prompt logging. No data used for training. Free to start — no credit card.
