Now on VeniceAudioAnonymous

ElevenLabs TTS v4 Turbo

ElevenLabs' low-latency speech model — real-time streaming, 90+ languages, and inline expression tags, built for voice agents.

For agents
curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "elevenlabs-tts-v4-turbo",
    "prompt": "An uplifting cinematic orchestral build with soaring strings"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.
Model IDelevenlabs-tts-v4-turbo
Maker
ElevenLabs
Modality
Audio
License
Proprietary
Open weights
No — proprietary

Overview

What is ElevenLabs TTS v4 Turbo

ElevenLabs TTS v4 Turbo is the real-time variant of ElevenLabs' v4 speech generation, released in September 2026 alongside the flagship v4. It converts text into lifelike speech with low latency for streaming and voice agents, supports more than 90 languages, and controls delivery through natural-language audio tags.

Using it anonymously on Venice

On Venice, ElevenLabs TTS v4 Turbo runs under the anonymized privacy tier: your prompts are not stored, profiled, or used for training — a sharp contrast to Big-Tech voice platforms that keep every transcript. You pay per character with credits instead of a subscription, and you can call it through the Venice API without building a personal generation history. Sovereignty over your scripts stays with you.

AnonymousNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Specifications

Datasheet

Maker
ElevenLabs
Modality
Text-to-speech with real-time streaming output
Open weights
No — proprietary
License
Proprietary
Modes
Not supported
Prompt length
10,000 characters per request (v4 family)
Input images
Not supported
Released
September 28, 2026
Languages
90+ (up from 70+ in v3)
Latency
Built for real-time voice agents and streaming
Privacy on Venice
Anonymized — prompts not stored
Available on Venice since
Sep 2026

Assessment

Strengths and limitations

Strengths
  • Purpose-built for real-time use: low-latency streaming makes it suitable for voice agents, live calls, and interactive applications.
  • Broad language coverage: more than 90 languages, with ElevenLabs reporting its biggest quality jumps in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
  • Expression control via inline natural-language audio tags such as [laughs] and [whispers], which can be stacked and followed in sequence.
  • Reads text with context awareness, adjusting pacing and emotion based on what a line means rather than just what it says.
  • Cheapest v4-generation speech model on Venice at $0.05 per 1,000 characters — a third of the price of v3 or Multilingual v2.
Limitations
  • Closed and proprietary: no open weights, so it cannot be self-hosted, audited, or fine-tuned.
  • SSML break tags are disabled; pauses and delivery must be directed through audio tags rather than markup.
  • For maximum expressiveness in narration and character performance, the flagship ElevenLabs TTS v4 (also on Venice) is the stronger pick — Turbo trades some of that for speed.
  • Venice's anonymized tier means prompts are not stored, but the model does not run in a TEE and traffic is not end-to-end encrypted.

Use cases

What it is good for

  1. 01Voice agents and interactive assistants that need fluid, low-latency spoken responses.
  2. 02Live narration and streaming audio generated from text as it arrives.
  3. 03Multilingual voiceovers across 90+ languages without switching models.
  4. 04Prototyping dialogue delivery with inline tags like [laughs] and [whispers] before committing to a full production run.
  5. 05Cost-sensitive bulk narration at $0.05 per 1,000 characters.

Prompting

Getting better results

Direct emotion inline with natural-language tags — write [whispers] or [laughs] directly in the script where the shift should happen.

Stack multiple tags in sequence (e.g. [sighs] then [excited]) — v4 follows tag order more reliably than v3.

Don't use SSML break tags; they're disabled. Describe pauses in plain language instead, e.g. "...a long pause...".

Give the model context about who is speaking and why — v4 reads with awareness of the speaker and preceding text, so framing lines improves delivery.

Keep drafts short while iterating: at $0.05 per 1,000 characters, testing on a paragraph costs a fraction of a full script.

Name the target language or dialect explicitly in the surrounding instructions when generating non-English speech to lock in pronunciation.

Samples

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

Cinematic score

“An uplifting cinematic orchestral build with soaring strings, warm brass, and a hopeful resolution.”

Lo-fi beat

“A mellow lo-fi hip-hop beat with a soft jazzy piano loop, vinyl crackle, and a relaxed late-night mood.”

Compare every audio model on these prompts →

Alternatives

How it compares

ModelBest forVoicesOpen weightsPrice (Venice)
ElevenLabs TTS v4 TurboReal-time agents & streamingVoice libraryNo$0.05 / 1K chars
ElevenLabs TTS v4Expressive narration & charactersVoice libraryNo$0.09 / 1K chars
ElevenLabs TTS v3Dramatic multi-speaker dialogueVoice libraryNo$0.12 / 1K chars
ElevenLabs Multilingual v2Stable long-form narrationVoice libraryNo$0.12 / 1K chars

ElevenLabs TTS v4 Turbo is the right pick when audio must respond in real time — voice agents, live streams, interactive apps — at the lowest per-character price in the v4 family. For produced narration where expressiveness matters more than latency, step up to ElevenLabs TTS v4.

API

Call it from your code

Venice exposes this model through the REST API. Queue a generation with the model id.

curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "elevenlabs-tts-v4-turbo",
    "prompt": "An uplifting cinematic orchestral build with soaring strings"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.

Pricing

What it costs on Venice

Billed per character on Venice: $0.05 per 1,000 characters.

Characters / 1K
$0.05
Per track

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

FAQ

Frequently asked questions

ElevenLabs TTS v4 Turbo is the real-time variant of ElevenLabs' v4 text-to-speech generation, released on September 28, 2026. It turns text into lifelike speech with low latency for streaming and voice agents, supports more than 90 languages, and controls delivery through inline audio tags rather than SSML.

Venice bills it per character: $0.05 per 1,000 characters, with no subscription required. That makes it the cheapest v4-generation speech model on Venice — compared with $0.09 for ElevenLabs TTS v4 and $0.12 for v3 or Multilingual v2.

It is neither free-software nor open-source: the model is proprietary to ElevenLabs with no open weights, so it cannot be self-hosted or fine-tuned. You can use it on Venice by paying per character in credits, and new Venice accounts include free daily usage to try it before spending credits.

Choose v4 Turbo for real-time workloads — voice agents, live calls, streaming — where latency matters most and cost per character is lowest. Choose the flagship v4 for produced narration and character performances, where its stronger expression control and voice-identity stability over long text justify the higher price.

The v4 generation supports more than 90 languages, up from roughly 70 in v3. ElevenLabs reports the largest quality improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

Not for pauses: SSML break tags are disabled in the v4 generation. Pauses and delivery are instead controlled through natural-language audio tags such as [laughs], [whispers] and [door slams], which can be stacked and are followed in sequence.

Yes. Venice runs it under the anonymized privacy tier — prompts are not stored, profiled, or used for training, and generations are tied to no personal account history. Note that the model does not run inside a trusted execution environment and traffic is not end-to-end encrypted.

The v4 architecture supports cloning a voice from as little as 10 seconds of audio, a capability ElevenLabs highlighted at launch. Availability of cloning through Venice's API may be limited to ElevenLabs' own platform features, so verify voice-clone support in the Venice interface before planning a project around it.

Use ElevenLabs TTS v4 Turbo anonymously

Venice does not store your prompts. Chat history stays in your browser.