AI Models · Audio

Best Audio AI Models

Models that understand or generate audio beyond plain speech TTS.

Models
7
Providers
4
Categories
35
Updated
2026-06

Best Audio AI Models (2026): GPT-4o, Gemini, ElevenLabs & More

7 models matched. Click any column to sort.

Capabilities
Gemini 2.5 ProGoogle78110 tok/s0.7s2M$1.25$5.00$2.19
GPT-4oOpenAI72110 tok/s0.4s128k$2.50$10.00$4.38
Gemini 1.5 ProGoogle6760 tok/s0.9s2M$1.25$5.00$2.19
Gemini 2.0 FlashGoogle64220 tok/s0.3s1M$0.10$0.40$0.18
Cartesia SonicCartesia0.09s
OpenAI TTS (HD)OpenAI0.5s
ElevenLabs Multilingual v2ElevenLabs0.4s

Showing 7 of 7 models. Click any column header to sort. Prices are USD per 1M tokens unless noted otherwise. Estimates marked with *.

Browse AI Models by category

Drill into a slice of the catalog — reasoning, coding, image, video, speech, embeddings, agents, open weights, and more.

By Provider

Frequently asked questions

Audio models either understand sound (speech-in, ambient audio) or generate it. Multimodal LLMs like GPT-4o and Gemini accept audio input; TTS models like ElevenLabs and Cartesia generate speech output.

Explore the full catalog

See every AI model in one place — intelligence, speed and price on a single sortable table.