Now on VeniceAudioAnonymous

ElevenLabs TTS v4

ElevenLabs' flagship expressive TTS model — context-aware emotional delivery, stackable expression tags, IPA control, and 90+ languages.

For agents
curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "elevenlabs-tts-v4",
    "prompt": "An uplifting cinematic orchestral build with soaring strings"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.
Model IDelevenlabs-tts-v4
Maker
ElevenLabs
Modality
Audio
License
Proprietary
Open weights
No — proprietary

Overview

What is ElevenLabs TTS v4

ElevenLabs TTS v4 is ElevenLabs' flagship text-to-speech model, released in September 2026 as the successor to v3. It delivers expressive, context-aware speech in more than 90 languages, supports stackable inline expression tags, IPA pronunciation control, and voice cloning from just 10 seconds of audio.

Using it anonymously on Venice

On Venice, ElevenLabs TTS v4 runs under the anonymized privacy tier — your prompts and scripts are not stored, profiled, or used for training, so sensitive narration, dialogue, or business audio never builds a personal history. You pay per character with credits instead of a subscription, and there's no account-level profiling tied to what you synthesize. Note that v4 is a closed, proprietary model, so private access via Venice is one of the few ways to use it without self-hosting trade-offs.

AnonymousNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Specifications

Datasheet

Maker
ElevenLabs
Modality
Text-to-speech
Open weights
No — proprietary
License
Proprietary
Prompt length
Reportedly up to 10,000 characters per request
Input images
Not supported
Released
September 28, 2026
Languages
90+ (up from 70+ in v3)
Voice cloning
From as little as 10 seconds of audio
Expression control
Inline tags, stackable and sequential; IPA phonetic control
Privacy on Venice
Anonymized — prompts not stored
Available on Venice since
Sep 2026

Assessment

Strengths and limitations

Strengths
  • Context-aware delivery: it reads the surrounding text to adjust tone, pacing, and emotion rather than narrating line by line.
  • Stackable inline expression tags: you can chain multiple tags and the model follows the sequence, a major upgrade over v3.
  • 90+ languages, with the largest reported quality jumps in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.
  • Preserves speaker identity far better over long-form text, where earlier models tended to drift.
  • Ranked first on the Artificial Analysis Speech Arena at launch, with the highest pronunciation-robustness score the benchmark has recorded.
Limitations
  • Closed and proprietary: no open weights, so self-hosting or fine-tuning is not possible.
  • It is not the low-latency variant: for real-time voice agents, ElevenLabs TTS v4 Turbo is purpose-built with roughly 100ms latency.
  • Its headline benchmark rankings come from third-party evaluators (Artificial Analysis), not independent reproduction.
  • Per-character billing adds up on very long scripts — worth estimating cost before batch-rendering an entire book.

Use cases

What it is good for

  1. 01Audiobook and long-form narration where the voice must stay consistent across chapters.
  2. 02Character dialogue and performance for games, animation, and audio drama using stacked expression tags.
  3. 03Multilingual voiceovers — especially Japanese, Brazilian Portuguese, Mandarin, and Cantonese, where v4's quality gains are largest.
  4. 04Marketing and explainer audio where emotional delivery matters more than raw speed.
  5. 05Cloning a brand or creator voice from a short sample for repeatable, on-brand narration.

Prompting

Getting better results

Stack inline expression tags in sequence (e.g. a whisper tag followed by an excited tag) — v4 follows the chain rather than applying one flat tone.

Use IPA notation for names, acronyms, or loanwords the model habitually mispronounces.

Clone with a clean 10-second sample: quiet room, consistent distance from the mic, no background music.

Leave surrounding sentences in the input — v4 reads context to set expression, so isolated lines sound flatter.

Tag each speaker's lines individually in multi-character scripts so identity and tone stay separated.

For long scripts, render in consistent chunks with the same voice settings rather than one giant request, so pacing stays uniform.

Samples

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

Cinematic score

“An uplifting cinematic orchestral build with soaring strings, warm brass, and a hopeful resolution.”

Lo-fi beat

“A mellow lo-fi hip-hop beat with a soft jazzy piano loop, vinyl crackle, and a relaxed late-night mood.”

Compare every audio model on these prompts →

Alternatives

How it compares

ModelBest forVoicesOpen weightsPrice (Venice)
ElevenLabs TTS v4Expressive narration & character work90+ languagesNo$0.09 / 1K chars
ElevenLabs TTS v4 TurboReal-time voice agents90+ languagesNo$0.05 / 1K chars
ElevenLabs TTS v3Proven expressive TTS70+ languagesNo$0.12 / 1K chars
ElevenLabs Multilingual v2Stable multilingual voiceoversMultilingualNo$0.12 / 1K chars

Pick ElevenLabs TTS v4 when delivery quality is the product — narration, audiobooks, character performance — and you need emotional control across 90+ languages. For live, conversational voice agents, TTS v4 Turbo is the faster and cheaper pick.

API

Call it from your code

Venice exposes this model through the REST API. Queue a generation with the model id.

curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "elevenlabs-tts-v4",
    "prompt": "An uplifting cinematic orchestral build with soaring strings"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.

Pricing

What it costs on Venice

Billed per character on Venice: $0.09 per 1,000 characters.

Characters / 1K
$0.09
Per track

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

FAQ

Frequently asked questions

ElevenLabs TTS v4 is ElevenLabs' flagship text-to-speech model, launched on September 28, 2026 as the successor to v3. It is built on a new architecture that delivers context-aware emotional speech, supports more than 90 languages, and enables voice cloning from as little as 10 seconds of audio.

Venice bills ElevenLabs TTS v4 per character: $0.09 per 1,000 characters, paid in credits with no subscription required. Cost scales linearly with script length, so a long narration project is easy to estimate before you render.

Neither. ElevenLabs TTS v4 is closed and proprietary — the weights are not published, so it cannot be self-hosted or fine-tuned. On Venice it is pay-per-character at $0.09 per 1,000 characters rather than subscription-only.

Use v4 for quality-critical, expressive output: narration, audiobooks, and character work where tone and pacing matter. Use v4 Turbo for real-time voice agents — it targets roughly 100ms latency for fluid conversation and costs $0.05 per 1,000 characters on Venice versus $0.09 for v4.

More than 90, up from roughly 70 in v3. ElevenLabs reports the biggest quality improvements in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, making v4 a strong choice for multilingual voiceover work.

Yes. The v4 architecture supports voice cloning from as little as 10 seconds of reference audio, and it preserves the cloned speaker's identity more reliably across long texts than previous generations.

v3 introduced inline tags to mark how a passage should be delivered. v4 expands this: you can stack multiple tags and the model follows their sequence, while also using the surrounding text's context to adjust expression automatically — reducing the amount of manual annotation needed.

Venice runs ElevenLabs TTS v4 under its anonymized privacy tier: prompts and scripts are not stored, profiled, or used for training. You can synthesize audio in the Venice app or via the Venice API without building a personal generation history tied to your identity.

Use ElevenLabs TTS v4 anonymously

Venice does not store your prompts. Chat history stays in your browser.