AudioAnonymized

MMAudio V2

MMAudio V2 is a 157M-parameter flow-matching audio model from UIUC and Sony AI researchers that generates synchronized sound from text or video, with a v2 checkpoint tuned for stronger real-world generalization.

Get API key

What is MMAudio V2?

MMAudio V2 is a 157M-parameter flow-matching audio model developed by UIUC and Sony AI researchers. It generates synchronized sound from text or video inputs using multimodal joint training, and the v2 checkpoint trades benchmark scores for stronger real-world generalization.

Use MMAudio V2 privately on Venice

On Venice, MMAudio V2 runs with zero retention — your prompts are anonymized and never stored for training. Audio generation costs $0 per second, so you can experiment permissionlessly without surveillance or a subscription.

Anonymized
No prompt training
TEE · hardware enclave
End-to-end encrypted

What can MMAudio V2 do?

Strengths
  • Extremely lightweight at 157M parameters, yet produces high-fidelity 44.1kHz audio via flow matching.
  • Multimodal joint training handles both text-to-audio and video-to-audio with a frame-level synchronization module.
  • Fast inferencethe paper reports just 1.23 seconds to generate an 8-second audio clip.
  • The v2 checkpoint generalizes better to new data than the original benchmark-tuned version, according to the authors.
Limitations
  • Closed weights on Venice — the hosted model cannot be self-hosted or fine-tuned through the platform.
  • The v2 checkpoint underperforms the original on standard benchmarks such as Fréchet distance.
  • Designed for short-form Foley and ambient sound rather than complex multi-track music production.

Sample outputs

Generated on Venice with our standard prompt suite — the same prompts we run through every model of this type, so you can judge it like-for-like.

Cinematic score

An uplifting cinematic orchestral build with soaring strings, warm brass, and a hopeful resolution.

Lo-fi beat

A mellow lo-fi hip-hop beat with a soft jazzy piano loop, vinyl crackle, and a relaxed late-night mood.

Compare every audio model on these prompts

How to use MMAudio V2 via API

Venice exposes this model through the REST API. Queue a generation with mmaudio-v2-text-to-audio.

curl https://api.venice.ai/api/v1/audio/queue \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mmaudio-v2-text-to-audio",
    "prompt": "An uplifting cinematic orchestral build with soaring strings"
  }'

# Use the returned queue_id with https://api.venice.ai/api/v1/audio/retrieve.
# Call /audio/complete after downloading if needed.

Specifications

MakerHo Kei Cheng et al. (UIUC, Sony AI, Sony Group Corporation)
ReleasedDecember 7, 2024 (research); v2 checkpoint updated 2025
ModalityText-to-audio, video-to-audio
ArchitectureFlow matching with conditional synchronization module
Parameters157M
Open weightsNo — proprietary
Privacy on VeniceAnonymized — prompts not stored
Available on Venice sinceFeb 2026

Pricing

Billed per second of audio on Venice: $0 per second.

Per second
$0

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

MMAudio V2 vs alternatives

ModelPrice (Venice)ContextOpen weights
MMAudio V2$0 / secNo
ACE-Step 1.5from $0.03 / trackNo
ElevenLabs Sound Effects v2$0 / secNo
ElevenLabs Musicfrom $0.69 / trackNo

Free per-second generation for text- or video-conditioned audio; v2 emphasizes generalization over benchmark scores.

What is MMAudio V2 good for?

  • Adding synchronized Foley and ambient audio to short video clips.
  • Rapid soundscape prototyping for game development and social media content.
  • Generating text-conditioned sound effects and atmospheric audio for creative projects.
  • Revitalizing historical footage with period-appropriate generated audio.

Prompting tips

  • Describe the scene, mood, and sound sources explicitly; the model uses detailed text guidance.
  • Start with short clips to verify sync and audio quality before generating longer sequences.
  • Iterate on prompt specificity — tighter descriptions increase adherence, looser ones yield more natural variation.

Version history

MMAudio
2024-12

Initial research release with 157M parameters and multimodal joint training.

MMAudio V2
2025

CurrentRecommended checkpoint that trades benchmark scores for improved generalization to new data.

Frequently asked questions

MMAudio V2 is a 157M-parameter flow-matching audio synthesis model developed by researchers at UIUC, Sony AI, and Sony Group Corporation. It generates synchronized audio from text or video inputs, and the v2 checkpoint is optimized for stronger generalization to real-world data.

On Venice, MMAudio V2 is free to use — you pay $0 per second of generated audio with no subscription required.

Yes. Venice bills MMAudio V2 at $0 per second, so you can generate audio without per-use charges.

The research code is available under an MIT license, but the weights hosted on Venice are proprietary and not available for download or local fine-tuning.

Use MMAudio V2 for free, per-second generation of soundscapes and Foley tied to text or video. Choose ACE-Step 1.5 if you need dedicated music track generation and prefer a flat per-track price.

The underlying research model supports video-to-audio synthesis, but the Venice endpoint is currently configured for text-to-audio generation.

Yes. Venice processes requests under an anonymized privacy tier, meaning your prompts are not stored, profiled, or used for training.

The model generates 44.1kHz audio using a flow-matching objective and outputs synchronized sound effects and ambient audio.

Related models

Run MMAudio V2 privately.

No prompt logging. No data used for training. Free to start — no credit card.

Room