Skip to content

Audio AI covers four jobs: text-to-speech (a computer voice reads your text), voice cloning, speech-to-text (transcription), and music generation. Most products specialize in one or two of them.

ElevenLabs is the best-known name and spans TTS, cloning, transcription, and music in one platform. The enterprise clouds (Google, Microsoft, OpenAI) sell the same four jobs as developer APIs priced per character or per minute.

Open-weight models are strong in audio. OpenAI's Whisper is a free, MIT-licensed transcription model you can run on your own machine, Kokoro is a small TTS model light enough for a CPU, and Resemble's Chatterbox does zero-shot voice cloning under an MIT license.

Snapshot of 7 June 2026. 31 entries in this group, about 25 with a free option and 8 that can run on your own hardware. Prices and versions change often, and each name links to its official page. The full directory has search and filters across all categories.

How to Choose

Voiceover for a video:

ElevenLabs and Murf both have free tiers to test voices. Free tiers commonly exclude commercial use, so plan on a paid tier for published work.

Transcribe recordings free:

Whisper is open-weight and free to run locally. Otter's free tier covers 300 minutes a month if you prefer an app.

Generate music:

Suno and Udio produce full songs with vocals from a text prompt. Free tiers are non-commercial, and commercial rights come with the paid plans.

Build voice into an app:

Deepgram and AssemblyAI give real free starting credit (200 and 50 dollars) on fast transcription APIs. Cartesia's Sonic targets low-latency voice agents.

Clone a voice:

ElevenLabs and Resemble offer instant cloning on paid tiers. Chatterbox is the open-source route.

All 31 Voice, audio & music Entries

Tool Maker What it is Cost Where it runs Hardware
AssemblyAI AssemblyAI Transcription API with speaker labels and audio-intelligence add-ons. Free + paid Cloud Any device
Azure AI Speech Microsoft Neural text-to-speech, custom voice, and speech-to-text on Azure. Free + paid Cloud Any device
Cartesia (Sonic) Cartesia Low-latency real-time TTS plus speech-to-text for voice apps. Free + paid Cloud Any device
Chatterbox Resemble AI Open zero-shot voice cloning TTS with a low-latency Turbo variant. Open-weight Cloud or local GPU recommended
Deepgram (Nova-3) Deepgram Fast, accurate transcription API, plus Aura TTS and voice agents. Free + paid Cloud Any device
ElevenLabs ElevenLabs High-quality text-to-speech, voice cloning, dubbing, and sound effects. Free + paid Cloud Any device
ElevenLabs Music ElevenLabs Text-to-music generation inside the ElevenLabs platform. Free + paid Cloud Any device
ElevenLabs Scribe ElevenLabs High-accuracy transcription API with a realtime variant. Free + paid Cloud Any device
Google Cloud STT Google Streaming and batch transcription across 45+ languages. Free + paid Cloud Any device
Google Cloud TTS Google Enterprise text-to-speech with WaveNet, Neural2, and Chirp 3 HD voices. Free + paid Cloud Any device
Google Lyria Google Enterprise music-generation model. 48kHz stereo with SynthID watermarking. Paid Cloud Any device
Hume (Octave) Hume AI Emotionally expressive text-to-speech with prompt-controlled tone. Free + paid Cloud Any device
Kokoro hexgrad (community) Small, fast open text-to-speech model that runs almost anywhere. Open-weight Cloud or local CPU or modest GPU (82M)
Mubert Mubert Royalty-free generative music for creators, plus an API. Free + paid Cloud Any device
Murf AI Murf Studio-style voiceover with 200+ voices, translation, and dubbing. Free + paid Cloud Any device
MusicGen Meta Open text-to-music model (AudioCraft). Weights are non-commercial. Open-weight Cloud or local GPU recommended
NVIDIA Canary NVIDIA Open transcription and translation across 25 European languages. Open-weight Cloud or local NVIDIA GPU
NVIDIA Parakeet NVIDIA High-throughput open English transcription model. Open-weight Cloud or local NVIDIA GPU
OpenAI Realtime voice OpenAI End-to-end speech-in, speech-out model for low-latency voice agents. Paid Cloud Any device
OpenAI TTS OpenAI Developer text-to-speech API with steerable tone. Paid Cloud Any device
Otter.ai Otter.ai Meeting transcription with live notes, summaries, and speaker labels. Free + paid Cloud Any device
PlayHT / PlayAI PlayAI Text-to-speech and instant voice cloning for creators and agents. Free + paid Cloud Any device
Resemble AI Resemble AI Voice cloning, TTS, voice changer, plus deepfake detection. Paid Cloud Any device
Riffusion Riffusion Text-to-music app producing royalty-free full songs. Free + paid Cloud Any device
Speechify Speechify Consumer read-aloud text-to-speech across apps, plus an API. Free + paid Cloud or local Any device
Stable Audio Stability AI Text-to-audio for music and sound effects, enterprise-oriented. Paid Cloud Any device
Stable Audio Open Stability AI Open text-to-audio model for short samples and sound effects. Open-weight Cloud or local GPU recommended
Suno Suno Consumer text-to-song generator with vocals and instrumentation. Free + paid Cloud Any device
Udio Udio Text-to-music generator producing full songs with vocals. Free + paid Cloud Any device
WellSaid Labs WellSaid Enterprise text-to-speech for narration and corporate voiceover. Paid Cloud Any device
Whisper OpenAI Widely used open transcription model. Multilingual and robust. Open-weight Cloud or local GPU for large. CPU for small

Common Questions

Can I make AI music for free?

Yes. Suno's free tier gives 50 credits a day (about 10 songs) and Udio gives 10 a day plus a monthly allowance, both for non-commercial use. Releasing music commercially requires their paid plans.

What is the best free transcription?

Whisper, OpenAI's open-weight model, is free, multilingual, and runs on your own machine, with smaller variants that run on a CPU. Cloud APIs like Deepgram and AssemblyAI are faster at scale and give free starting credit.

Can I clone my own voice?

Yes, it is a standard product feature on ElevenLabs, Resemble, and PlayHT. Cloning another person's voice without consent breaks these platforms' terms of use, and Resemble ships deepfake detection and watermarking alongside its cloning tools for exactly that reason.