Speech & music

Audio Models

Text-to-speech, voice and music generation models — from expressive narration to radio-ready songs and commercial-safe audio.

Suno

Suno v5.5

The leading AI music generator for full, radio-ready songs — natural vocals, real song structure, voice cloning and 30+ genres from a single text prompt.

Free tier · paid Pro and Premier plans
ACE Studio / StepFun open

ACE-Step

The open-source full-song generator — free, self-hostable music with vocals that runs on a single consumer GPU.

Free (open source, self-hosted)
ElevenLabs

ElevenLabs v3

The industry-standard text-to-speech model — the most expressive, emotionally rich voice synthesis, with 70+ languages and multi-speaker dialogue.

~$0.12 / minute (v3) · Flash/Turbo ~$0.06 / min
ElevenLabs

ElevenLabs Music v2

The commercially safest AI music generator — built with label and publisher partnerships, with section-by-section editing, a full API and studio-grade exports.

$11 / mo (30 tracks) to $99 / mo (Scale, API)
Google

Lyria 3

Google DeepMind's music generation model — high-fidelity 48kHz audio up to three minutes, available through Vertex AI and the Gemini API.

Via Vertex AI / Gemini API (usage-based)
MiniMax

MiniMax Speech 2.8

A top-ranked text-to-speech model — expressive, multilingual voice synthesis with fast voice cloning and a large built-in voice library.

Usage-based via MiniMax platform (varies)
Stability AI

Stable Audio 2.5

Stability's music-and-sound model — fast, commercially-cleared instrumental tracks and sound design with an enterprise-friendly API.

Via Stability AI subscription / API (usage-based)
Udio

Udio v4

The audio-quality leader among AI music tools — 48kHz rendering with the cleanest instrument separation, for film and brand-grade music.

~$10 / mo (Standard) · ~$30 / mo (Pro)