Speech-to-Text API
Built for production.Optimized for batch.
High-accuracy multilingual transcription with speaker diarization, word-level timestamps, and automatic language detection.
Designed for large-scale batch workloads.
Meet Zephyr
Upload a file or record a sample to see the Zephyr API in action.
Input
Language
Output
Options
Upload a file or record a sample to begin. Zephyr will detect the language, process your audio, and format the results. Your completed transcript will appear here.
Speaker diarization
DER measures missed or misassigned speaker turns. Lower DER means cleaner speaker labels for analytics, compliance, and conversation intelligence.
Datasets
- AMI-SDM
- NotSoFar
↓ Lower is better
SpeechRevolutions Zephyr
AssemblyAI Universal 3.5 Pro
ElevenLabs Scribe V2
Deepgram Nova-3
Soniox V5
Why Zephyr leads
Zephyr delivers the industry's lowest speaker diarization error rate, ranking #1 across every evaluated dataset. Fewer missed speakersand cleaner speaker attribution make it the ideal choice for meetings, podcasts, contact centers, and conversation intelligence.
View methodology →Pay for speed, not features.
Every request includes speaker diarization, word-level timestamps, and multilingual transcription,seamless code switching, keyterm prompting, punctuation, and multiple output formats. No premium add-ons.
Processing
Price
Delivery
ActionStandard Batch
High-priority processing with results delivered in seconds.
Price
$0.003/ min
Delivery
Seconds
Interactive APIs
Standard Batch
High-priority processing with results delivered in seconds.
Price
$0.003/ min
Delivery
Seconds
Interactive APIs
Economy Batch
Uses idle compute for the same accuracy at a lower cost.
Price
$0.0015/ min
Delivery
≤ 7 days
Large backlogs
Economy Batch
Uses idle compute for the same accuracy at a lower cost.
Price
$0.0015/ min
Delivery
≤ 7 days
Large backlogs
Need enterprise pricing?
Volume discounts · Dedicated concurrency · Custom SLAs
$10 free credit on signup
Frequently asked questions
We focus exclusively on batch transcription at 300× realtime with transparent per-minute pricing and no add-on fees for diarization, timestamps, or repunctuation. If you process large volumes and want fast batch turnaround at a lower cost, we're built for that.
Not at launch. Speech Revolutions is batch-only today — upload audio, track progress via SSE, download structured output. Live streaming is on the roadmap.
Standard runs on dedicated compute for fast turnaround (300× realtime). Economy uses idle capacity at ~$0.09/hr with a 7-day delivery SLA or explicit failure notification — ideal for large backlogs where cost matters more than speed.
No. We do not use customer audio to train speech recognition models. Uploaded files and transcripts are deleted within 30 minutes of job completion.
JSON (with word timestamps and speaker labels), TXT, SRT, VTT, DOCX, and PDF — all from a single API call.
Up to 100 MB via direct REST upload, or up to 20 GB through the SDK using presigned multipart uploads.