Audio AI covers four jobs: text-to-speech (a computer voice reads your text), voice cloning, speech-to-text (transcription), and music generation. Most products specialize in one or two of them.
ElevenLabs is the best-known name and spans TTS, cloning, transcription, and music in one platform. The enterprise clouds (Google, Microsoft, OpenAI) sell the same four jobs as developer APIs priced per character or per minute.
Open-weight models are strong in audio. OpenAI's Whisper is a free, MIT-licensed transcription model you can run on your own machine, Kokoro is a small TTS model light enough for a CPU, and Resemble's Chatterbox does zero-shot voice cloning under an MIT license.
Snapshot of 7 June 2026. 31 entries in this group, about 25 with a free option and 8 that can run on your own hardware. Prices and versions change often, and each name links to its official page. The full directory has search and filters across all categories.
How to Choose
Voiceover for a video:
ElevenLabs and Murf both have free tiers to test voices. Free tiers commonly exclude commercial use, so plan on a paid tier for published work.
Transcribe recordings free:
Whisper is open-weight and free to run locally. Otter's free tier covers 300 minutes a month if you prefer an app.
Generate music:
Suno and Udio produce full songs with vocals from a text prompt. Free tiers are non-commercial, and commercial rights come with the paid plans.
Build voice into an app:
Deepgram and AssemblyAI give real free starting credit (200 and 50 dollars) on fast transcription APIs. Cartesia's Sonic targets low-latency voice agents.
Clone a voice:
ElevenLabs and Resemble offer instant cloning on paid tiers. Chatterbox is the open-source route.
All 31 Voice, audio & music Entries
| Tool | Maker | What it is | Cost | Where it runs | Hardware |
|---|---|---|---|---|---|
| AssemblyAI | AssemblyAI | Transcription API with speaker labels and audio-intelligence add-ons. | Free + paid | Cloud | Any device |
| Azure AI Speech | Microsoft | Neural text-to-speech, custom voice, and speech-to-text on Azure. | Free + paid | Cloud | Any device |
| Cartesia (Sonic) | Cartesia | Low-latency real-time TTS plus speech-to-text for voice apps. | Free + paid | Cloud | Any device |
| Chatterbox | Resemble AI | Open zero-shot voice cloning TTS with a low-latency Turbo variant. | Open-weight | Cloud or local | GPU recommended |
| Deepgram (Nova-3) | Deepgram | Fast, accurate transcription API, plus Aura TTS and voice agents. | Free + paid | Cloud | Any device |
| ElevenLabs | ElevenLabs | High-quality text-to-speech, voice cloning, dubbing, and sound effects. | Free + paid | Cloud | Any device |
| ElevenLabs Music | ElevenLabs | Text-to-music generation inside the ElevenLabs platform. | Free + paid | Cloud | Any device |
| ElevenLabs Scribe | ElevenLabs | High-accuracy transcription API with a realtime variant. | Free + paid | Cloud | Any device |
| Google Cloud STT | Streaming and batch transcription across 45+ languages. | Free + paid | Cloud | Any device | |
| Google Cloud TTS | Enterprise text-to-speech with WaveNet, Neural2, and Chirp 3 HD voices. | Free + paid | Cloud | Any device | |
| Google Lyria | Enterprise music-generation model. 48kHz stereo with SynthID watermarking. | Paid | Cloud | Any device | |
| Hume (Octave) | Hume AI | Emotionally expressive text-to-speech with prompt-controlled tone. | Free + paid | Cloud | Any device |
| Kokoro | hexgrad (community) | Small, fast open text-to-speech model that runs almost anywhere. | Open-weight | Cloud or local | CPU or modest GPU (82M) |
| Mubert | Mubert | Royalty-free generative music for creators, plus an API. | Free + paid | Cloud | Any device |
| Murf AI | Murf | Studio-style voiceover with 200+ voices, translation, and dubbing. | Free + paid | Cloud | Any device |
| MusicGen | Meta | Open text-to-music model (AudioCraft). Weights are non-commercial. | Open-weight | Cloud or local | GPU recommended |
| NVIDIA Canary | NVIDIA | Open transcription and translation across 25 European languages. | Open-weight | Cloud or local | NVIDIA GPU |
| NVIDIA Parakeet | NVIDIA | High-throughput open English transcription model. | Open-weight | Cloud or local | NVIDIA GPU |
| OpenAI Realtime voice | OpenAI | End-to-end speech-in, speech-out model for low-latency voice agents. | Paid | Cloud | Any device |
| OpenAI TTS | OpenAI | Developer text-to-speech API with steerable tone. | Paid | Cloud | Any device |
| Otter.ai | Otter.ai | Meeting transcription with live notes, summaries, and speaker labels. | Free + paid | Cloud | Any device |
| PlayHT / PlayAI | PlayAI | Text-to-speech and instant voice cloning for creators and agents. | Free + paid | Cloud | Any device |
| Resemble AI | Resemble AI | Voice cloning, TTS, voice changer, plus deepfake detection. | Paid | Cloud | Any device |
| Riffusion | Riffusion | Text-to-music app producing royalty-free full songs. | Free + paid | Cloud | Any device |
| Speechify | Speechify | Consumer read-aloud text-to-speech across apps, plus an API. | Free + paid | Cloud or local | Any device |
| Stable Audio | Stability AI | Text-to-audio for music and sound effects, enterprise-oriented. | Paid | Cloud | Any device |
| Stable Audio Open | Stability AI | Open text-to-audio model for short samples and sound effects. | Open-weight | Cloud or local | GPU recommended |
| Suno | Suno | Consumer text-to-song generator with vocals and instrumentation. | Free + paid | Cloud | Any device |
| Udio | Udio | Text-to-music generator producing full songs with vocals. | Free + paid | Cloud | Any device |
| WellSaid Labs | WellSaid | Enterprise text-to-speech for narration and corporate voiceover. | Paid | Cloud | Any device |
| Whisper | OpenAI | Widely used open transcription model. Multilingual and robust. | Open-weight | Cloud or local | GPU for large. CPU for small |
Common Questions
Can I make AI music for free?
Yes. Suno's free tier gives 50 credits a day (about 10 songs) and Udio gives 10 a day plus a monthly allowance, both for non-commercial use. Releasing music commercially requires their paid plans.
What is the best free transcription?
Whisper, OpenAI's open-weight model, is free, multilingual, and runs on your own machine, with smaller variants that run on a CPU. Cloud APIs like Deepgram and AssemblyAI are faster at scale and give free starting credit.
Can I clone my own voice?
Yes, it is a standard product feature on ElevenLabs, Resemble, and PlayHT. Cloning another person's voice without consent breaks these platforms' terms of use, and Resemble ships deepfake detection and watermarking alongside its cloning tools for exactly that reason.
Keep Reading
Verified articles related to this page:
Other model pages: