Speech & music
Audio Models
Text-to-speech, voice and music generation models — from expressive narration to radio-ready songs and commercial-safe audio.
Suno v5.5
The leading AI music generator for full, radio-ready songs — natural vocals, real song structure, voice cloning and 30+ genres from a single text prompt.
ACE-Step
The open-source full-song generator — free, self-hostable music with vocals that runs on a single consumer GPU.
ElevenLabs v3
The industry-standard text-to-speech model — the most expressive, emotionally rich voice synthesis, with 70+ languages and multi-speaker dialogue.
ElevenLabs Music v2
The commercially safest AI music generator — built with label and publisher partnerships, with section-by-section editing, a full API and studio-grade exports.
Lyria 3
Google DeepMind's music generation model — high-fidelity 48kHz audio up to three minutes, available through Vertex AI and the Gemini API.
MiniMax Speech 2.8
A top-ranked text-to-speech model — expressive, multilingual voice synthesis with fast voice cloning and a large built-in voice library.
Stable Audio 2.5
Stability's music-and-sound model — fast, commercially-cleared instrumental tracks and sound design with an enterprise-friendly API.
Udio v4
The audio-quality leader among AI music tools — 48kHz rendering with the cleanest instrument separation, for film and brand-grade music.