AudioAnonymized

Seed Audio 1.0

Seed Audio 1.0 is ByteDance's all-in-one audio scene generator — produces multi-character dialogue, music, SFX, and ambience from one prompt with precise timing control.

Maker
ByteDance
Modality
Audio
License
Proprietary
Open weights
No — proprietary

Overview

What is Seed Audio 1.0

Seed Audio 1.0 is ByteDance's text-to-audio model that generates full-scene audio, including multi-character dialogue, sound effects, music, and ambience in one pass. Released in June 2026, it enables creators to produce broadcast-ready audio tracks from a single prompt without manual mixing or stitching.

Running it privately on Venice

On Venice, Seed Audio 1.0 runs under an anonymized privacy tier — your prompts are never stored or profiled. This means you can generate rich, cinematic audio scenes without surveillance, ideal for creators who value privacy and control. Since Venice bills per second with zero retention, experimentation stays private and affordable.

AnonymizedNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Assessment

Strengths and limitations

Strengths
  • Generates multi-layer audio scenes: dialogue, SFX, music, and ambience — in a single pass, eliminating post-mixing.
  • Supports zero-shot voice cloning from text or reference audio (up to three clips), enabling consistent character voices.
  • Fine-grained timing control at 100 ms intervals for precise dialogue and sound cue placement.
  • Produces broadcast-ready, fully mixed audio ideal for podcasts, radio dramas, and video dubbing.
  • Available on Venice with zero retention: no storage of prompts or audio outputs.
Limitations
  • Currently limited to English and Chinese language support, with no official multilingual expansion announced.
  • No open weights or self-hosting options: entirely proprietary and closed model.
  • Lacks streaming output; audio is generated synchronously and delivered as a complete file.
  • Reference image input for voice shaping remains unverified in official documentation.

Samples

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

Cinematic score

An uplifting cinematic orchestral build with soaring strings, warm brass, and a hopeful resolution.

Lo-fi beat

A mellow lo-fi hip-hop beat with a soft jazzy piano loop, vinyl crackle, and a relaxed late-night mood.

Compare every audio model on these prompts

Specifications

Datasheet

Maker
ByteDance
Released
June 29, 2026
Modality
Text-to-audio scene generation
Max output length
120 seconds
Timing precision
100 ms intervals
Output formats
mp3, wav, pcm, ogg_opus
Sample rates
8 kHz to 48 kHz
Open weights
No — proprietary
Privacy on Venice
Anonymized — prompts not stored
Available on Venice since
Jun 2026
License
Proprietary

API

Call it from your code

Venice exposes this model through the REST API. Queue a generation with the model id.

curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "seed-audio-1-0",
    "prompt": "An uplifting cinematic orchestral build with soaring strings"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.

Pricing

What it costs on Venice

Billed per second of audio on Venice: $0 per second.

Per second
$0
Per track

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

Alternatives

How it compares

ModelMax output lengthStrongest atOpen weightsPrice (Venice)
Seed Audio 1.0120 secondsFull-scene audio in one passNo$0 / sec
ElevenLabs MusicBackground music generationNofrom $0.69 / track
Lyria 3 ProHigh-fidelity music generationNo$0.10 / track
MiniMax Music 2.5Music with emotional tone controlNo$0.18 / track

Generates multi-character dialogue, SFX, and music together — ideal for finished audio scenes.

Use cases

What it is good for

  1. 01Creating immersive audio scenes for podcasts, audiobooks, and radio dramas with synchronized SFX and music.
  2. 02Generating voiceovers with ambient context for marketing videos or documentaries.
  3. 03Prototyping sound design for games or films without requiring a full DAW workflow.
  4. 04Producing multilingual content in English and Chinese with consistent character voices.
  5. 05Private audio creation for sensitive or confidential projects using Venice’s anonymized tier.

Prompting

Getting better results

Structure prompts as scene descriptions with character dialogue, emotional tone, and ambient cues.

Use timing tags (e.g., '[00:15]') to control when lines or effects occur within the 120-second window.

Include voice descriptions (e.g., 'calm female voice') or upload reference audio for character consistency.

Specify output format and sample rate in API calls if high fidelity (e.g., 48 kHz) is required.

Version history

Seed Audio 1.0
2026-06

Initial release — unified audio scene generation

FAQ

Frequently asked questions

Seed Audio 1.0 is ByteDance's text-to-audio model that generates full-scene audio — including multi-character dialogue, sound effects, music, and ambience — from a single prompt. It produces broadcast-ready tracks up to 120 seconds long without requiring manual mixing.

On Venice, Seed Audio 1.0 is billed per second of output at $0 per second. This means you can use it at no cost while running on Venice’s anonymized, zero-retention infrastructure.

Seed Audio 1.0 is not open source — it is a proprietary model developed by ByteDance. However, it is available for free per-second usage on Venice, with no storage or profiling of your prompts.

Yes. Seed Audio 1.0 supports zero-shot voice cloning using text descriptions or reference audio clips (up to three). This allows for consistent character voices across scenes without additional training.

Seed Audio 1.0 supports multiple output formats including mp3, wav, pcm, and ogg_opus, with sample rates ranging from 8 kHz to 48 kHz, making it suitable for both web and broadcast applications.

Yes. That’s its core strength — it generates dialogue, music, SFX, and ambience in one unified pass, creating fully mixed, scene-level audio that’s ready to use without further editing.

Seed Audio 1.0 is better for complete audio scenes with dialogue and SFX, while ElevenLabs Music focuses solely on music generation. If you need integrated, multi-element audio, Seed Audio 1.0 is the stronger choice.

Currently, Seed Audio 1.0 supports English and Chinese. While some third-party sources suggest broader capabilities, official documentation confirms only these two languages at launch.

Seed Audio 1.0 can generate up to 120 seconds (2 minutes) of audio per request, with precise timing control down to 100 ms intervals for dialogue and sound cues.

Run Seed Audio 1.0 privately

No prompt logging. No data used for training.