
ElevenLabs TTS v4
ElevenLabs' flagship expressive TTS model — context-aware emotional delivery, stackable expression tags, IPA control, and 90+ languages.
curl https://api.venice.ai/api/v1/audio/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs-tts-v4",
"prompt": "An uplifting cinematic orchestral build with soaring strings"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.Overview
What is ElevenLabs TTS v4
ElevenLabs TTS v4 is ElevenLabs' flagship text-to-speech model, released in September 2026 as the successor to v3. It delivers expressive, context-aware speech in more than 90 languages, supports stackable inline expression tags, IPA pronunciation control, and voice cloning from just 10 seconds of audio.
Using it anonymously on Venice
On Venice, ElevenLabs TTS v4 runs under the anonymized privacy tier — your prompts and scripts are not stored, profiled, or used for training, so sensitive narration, dialogue, or business audio never builds a personal history. You pay per character with credits instead of a subscription, and there's no account-level profiling tied to what you synthesize. Note that v4 is a closed, proprietary model, so private access via Venice is one of the few ways to use it without self-hosting trade-offs.
Specifications
Datasheet
- Maker
- ElevenLabs
- Modality
- Text-to-speech
- Open weights
- No — proprietary
- License
- Proprietary
- Prompt length
- Reportedly up to 10,000 characters per request
- Input images
- Not supported
- Released
- September 28, 2026
- Languages
- 90+ (up from 70+ in v3)
- Voice cloning
- From as little as 10 seconds of audio
- Expression control
- Inline tags, stackable and sequential; IPA phonetic control
- Privacy on Venice
- Anonymized — prompts not stored
- Available on Venice since
- Sep 2026
Assessment
Strengths and limitations
- Context-aware delivery: it reads the surrounding text to adjust tone, pacing, and emotion rather than narrating line by line.
- Stackable inline expression tags: you can chain multiple tags and the model follows the sequence, a major upgrade over v3.
- 90+ languages, with the largest reported quality jumps in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.
- Preserves speaker identity far better over long-form text, where earlier models tended to drift.
- Ranked first on the Artificial Analysis Speech Arena at launch, with the highest pronunciation-robustness score the benchmark has recorded.
- Closed and proprietary: no open weights, so self-hosting or fine-tuning is not possible.
- It is not the low-latency variant: for real-time voice agents, ElevenLabs TTS v4 Turbo is purpose-built with roughly 100ms latency.
- Its headline benchmark rankings come from third-party evaluators (Artificial Analysis), not independent reproduction.
- Per-character billing adds up on very long scripts — worth estimating cost before batch-rendering an entire book.
Use cases
What it is good for
- 01Audiobook and long-form narration where the voice must stay consistent across chapters.
- 02Character dialogue and performance for games, animation, and audio drama using stacked expression tags.
- 03Multilingual voiceovers — especially Japanese, Brazilian Portuguese, Mandarin, and Cantonese, where v4's quality gains are largest.
- 04Marketing and explainer audio where emotional delivery matters more than raw speed.
- 05Cloning a brand or creator voice from a short sample for repeatable, on-brand narration.
Prompting
Getting better results
Stack inline expression tags in sequence (e.g. a whisper tag followed by an excited tag) — v4 follows the chain rather than applying one flat tone.
Use IPA notation for names, acronyms, or loanwords the model habitually mispronounces.
Clone with a clean 10-second sample: quiet room, consistent distance from the mic, no background music.
Leave surrounding sentences in the input — v4 reads context to set expression, so isolated lines sound flatter.
Tag each speaker's lines individually in multi-character scripts so identity and tone stay separated.
For long scripts, render in consistent chunks with the same voice settings rather than one giant request, so pacing stays uniform.
Samples
Sample outputs
Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.
“An uplifting cinematic orchestral build with soaring strings, warm brass, and a hopeful resolution.”
“A mellow lo-fi hip-hop beat with a soft jazzy piano loop, vinyl crackle, and a relaxed late-night mood.”
Alternatives
How it compares
| Model | Best for | Voices | Open weights | Price (Venice) |
|---|---|---|---|---|
| ElevenLabs TTS v4 | Expressive narration & character work | 90+ languages | No | $0.09 / 1K chars |
| ElevenLabs TTS v4 Turbo | Real-time voice agents | 90+ languages | No | $0.05 / 1K chars |
| ElevenLabs TTS v3 | Proven expressive TTS | 70+ languages | No | $0.12 / 1K chars |
| ElevenLabs Multilingual v2 | Stable multilingual voiceovers | Multilingual | No | $0.12 / 1K chars |
Pick ElevenLabs TTS v4 when delivery quality is the product — narration, audiobooks, character performance — and you need emotional control across 90+ languages. For live, conversational voice agents, TTS v4 Turbo is the faster and cheaper pick.
API
Call it from your code
Venice exposes this model through the REST API. Queue a generation with the model id.
curl https://api.venice.ai/api/v1/audio/queue \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs-tts-v4",
"prompt": "An uplifting cinematic orchestral build with soaring strings"
}'
# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.Pricing
What it costs on Venice
Billed per character on Venice: $0.09 per 1,000 characters.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
FAQ
Frequently asked questions
ElevenLabs TTS v4 is ElevenLabs' flagship text-to-speech model, launched on September 28, 2026 as the successor to v3. It is built on a new architecture that delivers context-aware emotional speech, supports more than 90 languages, and enables voice cloning from as little as 10 seconds of audio.
Venice bills ElevenLabs TTS v4 per character: $0.09 per 1,000 characters, paid in credits with no subscription required. Cost scales linearly with script length, so a long narration project is easy to estimate before you render.
Neither. ElevenLabs TTS v4 is closed and proprietary — the weights are not published, so it cannot be self-hosted or fine-tuned. On Venice it is pay-per-character at $0.09 per 1,000 characters rather than subscription-only.
Use v4 for quality-critical, expressive output: narration, audiobooks, and character work where tone and pacing matter. Use v4 Turbo for real-time voice agents — it targets roughly 100ms latency for fluid conversation and costs $0.05 per 1,000 characters on Venice versus $0.09 for v4.
More than 90, up from roughly 70 in v3. ElevenLabs reports the biggest quality improvements in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, making v4 a strong choice for multilingual voiceover work.
Yes. The v4 architecture supports voice cloning from as little as 10 seconds of reference audio, and it preserves the cloned speaker's identity more reliably across long texts than previous generations.
v3 introduced inline tags to mark how a passage should be delivered. v4 expands this: you can stack multiple tags and the model follows their sequence, while also using the surrounding text's context to adjust expression automatically — reducing the amount of manual annotation needed.
Venice runs ElevenLabs TTS v4 under its anonymized privacy tier: prompts and scripts are not stored, profiled, or used for training. You can synthesize audio in the Venice app or via the Venice API without building a personal generation history tied to your identity.
Use ElevenLabs TTS v4 anonymously
Venice does not store your prompts. Chat history stays in your browser.