xAI TTS v1
xAI's high-fidelity, low-latency text-to-speech API with expressive voices, inline speech tags, and enterprise-grade multilingual support — now on Venice with anonymized processing.
Get API keyWhat is xAI TTS v1?
xAI TTS v1 is xAI's text-to-speech model, released in April 2026 as part of the Grok Voice stack. It converts text into natural, expressive speech with sub-second latency, supporting 20+ languages and 80+ voices, with fine-grained control via speech tags and multiple audio formats for production use.
Use xAI TTS v1 privately on Venice
On Venice, xAI TTS v1 runs under an anonymized privacy tier — your prompts are never stored, profiled, or used for training. This means developers can generate speech at scale without compromising user privacy or building a traceable history, while still accessing the full expressiveness and low latency of xAI’s production-grade TTS.
What can xAI TTS v1 do?
- •Sub-second latency ideal for real-time applications like voice agents and customer support bots.
- •Supports 80+ expressive voices across 25+ languages with fine-grained control via inline speech tags (pauses, laughter, whispers, pitch, speed).
- •Multiple audio output formats including telephony-optimized μ-law and high-fidelity FLAC for diverse use cases.
- •Built on the same stack powering Tesla vehicles and Starlink customer support, ensuring enterprise-grade reliability.
- •Competitive per-character pricing makes it cost-effective for high-volume production workloads.
- •Proprietary and closed — no open weights, so self-hosting or fine-tuning is not possible.
- •No voice cloning support in the initial release, limiting personalization options.
- •Still in early adoption phase compared to mature TTS providers, so ecosystem tooling and community support are limited.
- •Limited public benchmarking data on long-form narration quality and emotional range.
Sample outputs
Generated on Venice with our standard prompt suite — the same scripts we run through every model of this type, so you can judge it like-for-like.
“On Venice, your prompts are processed privately and never stored, profiled, or used to train anyone's model.”
“Wait — so I can run a private voice model with zero data retention, and pay only for what I use? That's genuinely useful.”
“Three… two… one… liftoff! The rocket roared into the night sky as the crowd erupted in cheers.”
How to use xAI TTS v1 via API
Venice exposes an OpenAI-compatible API. Swap your base URL and call tts-xai-v1.
curl https://api.venice.ai/api/v1/audio/speech \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-xai-v1",
"input": "On Venice, your prompts are processed privately.",
"voice": "af_sky",
"response_format": "mp3"
}' --output speech.mp3Specifications
Pricing
Billed per character on Venice: $18.75 per 1M characters of synthesized speech.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
xAI TTS v1 vs alternatives
| Model | Price (Venice) | Voices | Open weights | Latency |
|---|---|---|---|---|
| xAI TTS v1 | $18.75 / 1M chars | 80+ | No | Sub-second |
| Inworld TTS-1.5 Max | $12.50 / 1M chars | — | No | — |
| Chatterbox HD (Resemble AI) | $50 / 1M chars | — | No | — |
| Kokoro Text to Speech | $3.50 / 1M chars | — | Yes | — |
xAI's expressive, low-latency TTS with broad language support and speech tags.
What is xAI TTS v1 good for?
- •Real-time voice agents for customer service, sales, and support.
- •Accessibility tools that convert text content to natural speech.
- •Podcast and audiobook narration with expressive delivery control.
- •Interactive voice response (IVR) systems and telephony applications.
- •Multilingual content localization with consistent, high-quality voices.
Prompting tips
- •Use inline speech tags like [pause] and [whisper] to control tone and pacing precisely.
- •Choose output format based on use case: Opus for streaming, WAV for post-production, μ-law for telephony.
- •Test multiple voices for emotional fit — some are optimized for warmth, others for clarity or authority.
- •Break long texts into smaller chunks for consistent latency and error resilience.
Version history
CurrentInitial public release
Frequently asked questions
xAI TTS v1 is xAI's text-to-speech model, released in April 2026 as part of the Grok Voice stack. It converts text into natural, expressive speech with sub-second latency, supporting 25+ languages and 80+ voices, with fine-grained delivery control via speech tags.
On Venice, xAI TTS v1 costs $18.75 per 1 million characters of synthesized speech. Pricing is transparent and usage-based, with no subscription required.
No. xAI TTS v1 is a proprietary model — it is not free for unlimited use and does not have open weights. It cannot be self-hosted or fine-tuned by users.
No, voice cloning is not supported in the initial release of xAI TTS v1. However, the model offers a rich library of 80+ production-ready voices across multiple languages.
xAI TTS v1 supports over 25 languages, including English, Chinese, Russian, Italian, and others, with language-optimized voices for natural pronunciation and intonation.
Yes. xAI TTS v1 supports inline speech tags for pauses, laughter, whispers, pitch shifts, and speed changes, allowing fine-grained expressive control over the generated speech.
xAI TTS v1 offers lower latency and broader language support at a lower price point, while ElevenLabs excels in emotional range and voice cloning. For real-time, multilingual production use, xAI has an edge; for creative storytelling, ElevenLabs may be preferred.
Related models
Run xAI TTS v1 privately.
No prompt logging. No data used for training. Free to start — no credit card.
