AI Models · Audio
Best Audio AI Models
Models that understand or generate audio beyond plain speech TTS.
Best Audio AI Models (2026): GPT-4o, Gemini, ElevenLabs & More
7 models matched. Click any column to sort.
| Gemini 2.5 Pro | 78 | 110 tok/s | 0.7s | 2M | $1.25 | $5.00 | $2.19 | |
| GPT-4o | OpenAI | 72 | 110 tok/s | 0.4s | 128k | $2.50 | $10.00 | $4.38 |
| Gemini 1.5 Pro | 67 | 60 tok/s | 0.9s | 2M | $1.25 | $5.00 | $2.19 | |
| Gemini 2.0 Flash | 64 | 220 tok/s | 0.3s | 1M | $0.10 | $0.40 | $0.18 | |
| Cartesia Sonic | Cartesia | — | — | 0.09s | — | — | — | — |
| OpenAI TTS (HD) | OpenAI | — | — | 0.5s | — | — | — | — |
| ElevenLabs Multilingual v2 | ElevenLabs | — | — | 0.4s | — | — | — | — |
Showing 7 of 7 models. Click any column header to sort. Prices are USD per 1M tokens unless noted otherwise. Estimates marked with *.
Browse AI Models by category
Drill into a slice of the catalog — reasoning, coding, image, video, speech, embeddings, agents, open weights, and more.
By Use case
Frontier
10The most capable models from each major lab.
Reasoning Models
7Long chain-of-thought models built for hard math, code, and planning.
Coding Models
14Models specialised for software engineering and code generation.
Agents
13Models strong at tool use, multi-step workflows, and agentic systems.
Fast & Cheap
6Production-grade workhorses with the best speed and cost.
Multimodal
11Models that natively understand text, images, and beyond.
By Modality
By License
By Audience
By Provider
OpenAI
11All models from OpenAI — GPT, o-series and beyond.
Anthropic
4Claude family of models from Anthropic.
Gemini family of models from Google DeepMind.
Meta
3Open-weights Llama family from Meta AI.
Mistral
2Models from Mistral AI — EU-hosted, multilingual.
DeepSeek
2Frontier-class open-weights models from DeepSeek.
xAI
1Grok models from xAI.
Alibaba
1Qwen family of open-weights models from Alibaba.
Voyage AI
1Embedding models from Voyage AI.
Midjourney
1Midjourney text-to-image models.
Black Forest Labs
1FLUX image-generation models from Black Forest Labs.
Stability AI
1Stable Diffusion image models from Stability AI.
Runway
1Gen-series text-to-video models from Runway.
Kuaishou
1Kling text-to-video models from Kuaishou.
ElevenLabs
1Text-to-speech voice models from ElevenLabs.
Cartesia
1Sonic low-latency text-to-speech models from Cartesia.
Frequently asked questions
Explore the full catalog
See every AI model in one place — intelligence, speed and price on a single sortable table.