Deepgram
AI speech-to-text API built for developers with industry-leading speed and accuracy.
Quick take
Deepgram is the best speech-to-text API for most developers. Speed (sub-300ms streaming) and a strong accuracy record (4.7% WER on the earlier Nova-2) make it hard to beat; on pre-recorded price alone AssemblyAI Universal-2 ($0.15/hr) is cheaper. AssemblyAI is the closest competitor, with stronger LLM-powered features (summarization and Q&A via LLM Gateway) and a cheaper pre-recorded rate on Universal-2. Choose Deepgram for raw transcription speed and cost. Choose AssemblyAI for post-processing intelligence.
Deepgram scorecard
How we gradeFastest streaming speech API with strong developer value.
Overview
Deepgram is a speech-to-text API built for developers who need fast, accurate transcription at scale. Unlike consumer tools (Otter, Fathom), Deepgram is pure infrastructure: you send audio in, you get text back. The company raised a $130M Series C at a $1.3B valuation in January 2026, employs 101-250 people, and processes billions of minutes of audio per year. With 440 G2 reviews at 4.6/5 and 76% satisfaction, Deepgram is one of the most adopted speech APIs in the developer community. Its current models are Nova-3 and Flux, a conversational model with built-in turn detection.
Strengths & limitations
Strengths
- +Fastest streaming transcription available
- +Competitive accuracy
- +Good developer experience and docs
Limitations
- –Fewer features than AssemblyAI for intelligence
- –Custom models require enterprise plan
- –Pre-recorded Nova-3 ($0.0043/min) costs more than AssemblyAI Universal-2 ($0.15/hr)
Pricing
Pay As You Go: $0.0043/minute for Nova-3 (pre-recorded); Nova-3 streaming is $0.0077/minute regular, shown at a $0.0048/minute promotional price on October 3, 2026. Flux is $0.0077/minute regular. $200 free credit to start. Growth (pre-paid annual credits) saves up to 20%, and Enterprise offers custom models and on-premise deployment. No minimum commitment on Pay As You Go.
Who it's for
Developers building voice applications, transcription features, or real-time captioning. Teams that already have an audio capture pipeline and need the best accuracy-to-price ratio for the transcription step. If you need an all-in-one meeting recording solution (capture + transcription + summaries), look at Recall.ai + Deepgram as a stack, not Deepgram alone.
Verdict
A- · 84/100Deepgram is the best speech-to-text API for most developers. Speed (sub-300ms streaming) and a strong accuracy record (4.7% WER on the earlier Nova-2) make it hard to beat; on pre-recorded price alone AssemblyAI Universal-2 ($0.15/hr) is cheaper. AssemblyAI is the closest competitor, with stronger LLM-powered features (summarization and Q&A via LLM Gateway) and a cheaper pre-recorded rate on Universal-2. Choose Deepgram for raw transcription speed and cost. Choose AssemblyAI for post-processing intelligence.
Key features
- Real-time streaming transcription
- Pre-recorded audio API
- Custom model training
- Speaker diarization
- Language detection
What users say
Works well with different voices and accents, even with background noise or strong accents.
G2
Can sometimes struggle when the audio is very noisy or when multiple people speak over each other.
G2
Alternatives to Deepgram
OpenAI Whisper
Free (open source) / API $0.006/min (whisper-1, removed from the API Feb 26, 2027)
AssemblyAI
Free tier (up to 185 hrs pre-recorded) / Pay-as-you-go from $0.15/hr (Universal-2) or $0.21/hr (Universal-3.5 Pro) / Streaming from $0.15/hr
Rev
Free / Essentials $25.49 per seat/mo annual ($29.99 monthly) / Pro $47.99 annual ($59.99 monthly) / Unlimited custom / Human from $1.99/min
Speechmatics
Free trial / Pay-as-you-go / Enterprise custom