Seed Audio 1.0
Seed Audio 1.0 is ByteDance's all-in-one audio scene generator — produces multi-character dialogue, music, SFX, and ambience from one prompt with precise timing control.
Get API keyWhat is Seed Audio 1.0?
Seed Audio 1.0 is ByteDance's text-to-audio model that generates full-scene audio, including multi-character dialogue, sound effects, music, and ambience in one pass. Released in June 2026, it enables creators to produce broadcast-ready audio tracks from a single prompt without manual mixing or stitching.
Use Seed Audio 1.0 privately on Venice
On Venice, Seed Audio 1.0 runs under an anonymized privacy tier — your prompts are never stored or profiled. This means you can generate rich, cinematic audio scenes without surveillance, ideal for creators who value privacy and control. Since Venice bills per second with zero retention, experimentation stays private and affordable.
What can Seed Audio 1.0 do?
- •Generates multi-layer audio scenes — dialogue, SFX, music, and ambience — in a single pass, eliminating post-mixing.
- •Supports zero-shot voice cloning from text or reference audio (up to three clips), enabling consistent character voices.
- •Fine-grained timing control at 100 ms intervals for precise dialogue and sound cue placement.
- •Produces broadcast-ready, fully mixed audio ideal for podcasts, radio dramas, and video dubbing.
- •Available on Venice with zero retention — no storage of prompts or audio outputs.
- •Currently limited to English and Chinese language support, with no official multilingual expansion announced.
- •No open weights or self-hosting options — entirely proprietary and closed model.
- •Lacks streaming output; audio is generated synchronously and delivered as a complete file.
- •Reference image input for voice shaping remains unverified in official documentation.
Sample outputs
Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.
“An uplifting cinematic orchestral build with soaring strings, warm brass, and a hopeful resolution.”
“A mellow lo-fi hip-hop beat with a soft jazzy piano loop, vinyl crackle, and a relaxed late-night mood.”
How to use Seed Audio 1.0 via API
Venice exposes this model through the REST API. Queue a generation with seed-audio-1-0.
curl https://api.venice.ai/api/v1/audio/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seed-audio-1-0",
"prompt": "An uplifting cinematic orchestral build with soaring strings"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.Specifications
Pricing
Billed per second of audio on Venice: $0 per second.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
Seed Audio 1.0 vs alternatives
| Model | Max output length | Strongest at | Open weights | Price (Venice) |
|---|---|---|---|---|
| Seed Audio 1.0 | 120 seconds | Full-scene audio in one pass | No | $0 / sec |
| ElevenLabs Music | — | Background music generation | No | from $0.69 / track |
| Lyria 3 Pro | — | High-fidelity music generation | No | $0.10 / track |
| MiniMax Music 2.5 | — | Music with emotional tone control | No | $0.18 / track |
Generates multi-character dialogue, SFX, and music together — ideal for finished audio scenes.
What is Seed Audio 1.0 good for?
- •Creating immersive audio scenes for podcasts, audiobooks, and radio dramas with synchronized SFX and music.
- •Generating voiceovers with ambient context for marketing videos or documentaries.
- •Prototyping sound design for games or films without requiring a full DAW workflow.
- •Producing multilingual content in English and Chinese with consistent character voices.
- •Private audio creation for sensitive or confidential projects using Venice’s anonymized tier.
Prompting tips
- •Structure prompts as scene descriptions with character dialogue, emotional tone, and ambient cues.
- •Use timing tags (e.g., '[00:15]') to control when lines or effects occur within the 120-second window.
- •Include voice descriptions (e.g., 'calm female voice') or upload reference audio for character consistency.
- •Specify output format and sample rate in API calls if high fidelity (e.g., 48 kHz) is required.
Version history
CurrentInitial release — unified audio scene generation
Frequently asked questions
Seed Audio 1.0 is ByteDance's text-to-audio model that generates full-scene audio — including multi-character dialogue, sound effects, music, and ambience — from a single prompt. It produces broadcast-ready tracks up to 120 seconds long without requiring manual mixing.
On Venice, Seed Audio 1.0 is billed per second of output at $0 per second. This means you can use it at no cost while running on Venice’s anonymized, zero-retention infrastructure.
Seed Audio 1.0 is not open source — it is a proprietary model developed by ByteDance. However, it is available for free per-second usage on Venice, with no storage or profiling of your prompts.
Yes. Seed Audio 1.0 supports zero-shot voice cloning using text descriptions or reference audio clips (up to three). This allows for consistent character voices across scenes without additional training.
Seed Audio 1.0 supports multiple output formats including mp3, wav, pcm, and ogg_opus, with sample rates ranging from 8 kHz to 48 kHz, making it suitable for both web and broadcast applications.
Yes. That’s its core strength — it generates dialogue, music, SFX, and ambience in one unified pass, creating fully mixed, scene-level audio that’s ready to use without further editing.
Seed Audio 1.0 is better for complete audio scenes with dialogue and SFX, while ElevenLabs Music focuses solely on music generation. If you need integrated, multi-element audio, Seed Audio 1.0 is the stronger choice.
Currently, Seed Audio 1.0 supports English and Chinese. While some third-party sources suggest broader capabilities, official documentation confirms only these two languages at launch.
Seed Audio 1.0 can generate up to 120 seconds (2 minutes) of audio per request, with precise timing control down to 100 ms intervals for dialogue and sound cues.
Related models
Run Seed Audio 1.0 privately.
No prompt logging. No data used for training. Free to start — no credit card.
