Speech-to-Text API

Built for production.Optimized for batch.

High-accuracy multilingual transcription with speaker diarization, word-level timestamps, and automatic language detection.

Designed for large-scale batch workloads.

300x RTF
long-form throughput
100+
Languages
20 GB
Maximum upload
$0.18 hr/audio
no extra add on costs

Meet Zephyr

Upload a file or record a sample to see the Zephyr API in action.

Input

or

Language

Output

Options

Speaker labels
Timestamps

Upload a file or record a sample to begin. Zephyr will detect the language, process your audio, and format the results. Your completed transcript will appear here.

Speaker diarization

DER measures missed or misassigned speaker turns. Lower DER means cleaner speaker labels for analytics, compliance, and conversation intelligence.

Datasets

  • AMI-SDM
  • NotSoFar
Lowest DER9.70% · 10.9%
25.3% · 23.3%
33.7% · 24.4%
30.9% · 28.4%
36.6% · 32.2%

AssemblyAI Universal 3.5 Pro

ElevenLabs Scribe V2

Deepgram Nova-3

Soniox V5

Why Zephyr leads

Zephyr delivers the industry's lowest speaker diarization error rate, ranking #1 across every evaluated dataset. Fewer missed speakersand cleaner speaker attribution make it the ideal choice for meetings, podcasts, contact centers, and conversation intelligence.

View methodology →

Pay for speed, not features.

Every request includes speaker diarization, word-level timestamps, and multilingual transcription,seamless code switching, keyterm prompting, punctuation, and multiple output formats. No premium add-ons.

Standard Batch

High-priority processing with results delivered in seconds.

Price

$0.003/ min

Delivery

Seconds

Interactive APIs

Economy Batch

Uses idle compute for the same accuracy at a lower cost.

Price

$0.0015/ min

Delivery

≤ 7 days

Large backlogs

Need enterprise pricing?

Volume discounts · Dedicated concurrency · Custom SLAs

Talk to our team

$10 free credit on signup

Frequently asked questions

We focus exclusively on batch transcription at 300× realtime with transparent per-minute pricing and no add-on fees for diarization, timestamps, or repunctuation. If you process large volumes and want fast batch turnaround at a lower cost, we're built for that.

Not at launch. Speech Revolutions is batch-only today — upload audio, track progress via SSE, download structured output. Live streaming is on the roadmap.

Standard runs on dedicated compute for fast turnaround (300× realtime). Economy uses idle capacity at ~$0.09/hr with a 7-day delivery SLA or explicit failure notification — ideal for large backlogs where cost matters more than speed.

No. We do not use customer audio to train speech recognition models. Uploaded files and transcripts are deleted within 30 minutes of job completion.

JSON (with word timestamps and speaker labels), TXT, SRT, VTT, DOCX, and PDF — all from a single API call.

Up to 100 MB via direct REST upload, or up to 20 GB through the SDK using presigned multipart uploads.