Speech-to-text pricing
One rate, billed by the audio minute, with every production feature included. No plans to compare, no per-feature surcharges, and nothing to negotiate before you can try it.
$0.003 / audio minute
$0.18 per audio hour · $10 of credit free on signup
Pay for speed, not features.
Every request includes speaker diarization, word-level timestamps, and multilingual transcription, seamless code switching, keyterm prompting, punctuation, and multiple output formats. No premium add-ons.
Processing
Price
Delivery
ActionStandard Batch
High-priority processing with results delivered in seconds.
Price
$0.003/ min
Delivery
Seconds
Interactive APIs
Standard Batch
High-priority processing with results delivered in seconds.
Price
$0.003/ min
Delivery
Seconds
Interactive APIs
Economy Batch
Uses idle compute for the same accuracy at a lower cost.
Price
$0.0015/ min
Delivery
≤ 24 hours
Large backlogs
Coming soon
Economy Batch
Uses idle compute for the same accuracy at a lower cost.
Price
$0.0015/ min
Delivery
≤ 24 hours
Large backlogs
Coming soon
Need enterprise pricing?
Volume discounts · Dedicated concurrency · Custom SLAs
$10 free credit on signup
What the price includes
Everything below is on by default and costs nothing extra. Where other providers price diarization, timestamps or keyword boosting as add-ons, the effective cost of an audio hour is the sum of those — which is the number worth comparing.
- Speaker diarization
- Who spoke when, as labelled turns. Charged for by most others.
- Word-level timestamps
- Start and end times per word, for captions and alignment.
- Punctuation and casing
- Readable text, not a lowercase stream.
- Automatic language detection
- 99+ languages, including switching mid-file.
- Six output formats
- JSON, TXT, SRT, VTT, DOCX and PDF from one call.
- Custom vocabulary
- Bias towards names, jargon and product terms.
- Webhooks and progress events
- Signed callbacks, plus live progress over SSE.
- Files up to 10 GB
- Through the SDKs, with multipart handled for you.
How billing works
Credits are prepaid and never expire. You buy them in the console, they are drawn down by the audio second, and the balance is visible there in real time — there is no invoice to wait for and no minimum commitment.
Billing is by the second of audio, not the minute, so a 90-second file costs ninety seconds. New accounts get $10 of credit, which is about 56 audio hours — enough to run a real backlog through before deciding.
Auto-recharge is optional and off by default. With it off, requests are refused once the balance reaches zero rather than running up a bill.
Frequently asked questions
We focus exclusively on batch transcription, with transparent per-minute pricing and no add-on fees for diarization, timestamps, or repunctuation. If you process large volumes and want fast batch turnaround at a lower cost, we're built for that.
Not at launch. Speech Revolutions is batch-only today — upload audio, track progress via SSE, download structured output. Live streaming is on the roadmap.
Standard runs on dedicated compute, and most files come back in seconds. Economy is not available yet: when it ships it will use idle capacity at ~$0.09/hr with a 24-hour delivery SLA or explicit failure notification, for large backlogs where cost matters more than speed.
No. We do not use customer audio to train speech recognition models. Uploaded files and transcripts are deleted within 30 minutes of job completion.
JSON (with word timestamps and speaker labels), TXT, SRT, VTT, DOCX, and PDF — all from a single API call.
Up to 200 MB via direct REST upload, or up to 10 GB through the SDK using presigned multipart uploads.