Speech-to-text pricing

One rate, billed by the audio minute, with every production feature included. No plans to compare, no per-feature surcharges, and nothing to negotiate before you can try it.

$0.003 / audio minute

$0.18 per audio hour · $10 of credit free on signup

Pay for speed, not features.

Every request includes speaker diarization, word-level timestamps, and multilingual transcription, seamless code switching, keyterm prompting, punctuation, and multiple output formats. No premium add-ons.

Standard Batch

High-priority processing with results delivered in seconds.

Price

$0.003/ min

Delivery

Seconds

Interactive APIs

Economy Batch

Uses idle compute for the same accuracy at a lower cost.

Price

$0.0015/ min

Delivery

≤ 24 hours

Large backlogs

Coming soon

Need enterprise pricing?

Volume discounts · Dedicated concurrency · Custom SLAs

Talk to our team

$10 free credit on signup

What the price includes

Everything below is on by default and costs nothing extra. Where other providers price diarization, timestamps or keyword boosting as add-ons, the effective cost of an audio hour is the sum of those — which is the number worth comparing.

Speaker diarization
Who spoke when, as labelled turns. Charged for by most others.
Word-level timestamps
Start and end times per word, for captions and alignment.
Punctuation and casing
Readable text, not a lowercase stream.
Automatic language detection
99+ languages, including switching mid-file.
Six output formats
JSON, TXT, SRT, VTT, DOCX and PDF from one call.
Custom vocabulary
Bias towards names, jargon and product terms.
Webhooks and progress events
Signed callbacks, plus live progress over SSE.
Files up to 10 GB
Through the SDKs, with multipart handled for you.

How billing works

Credits are prepaid and never expire. You buy them in the console, they are drawn down by the audio second, and the balance is visible there in real time — there is no invoice to wait for and no minimum commitment.

Billing is by the second of audio, not the minute, so a 90-second file costs ninety seconds. New accounts get $10 of credit, which is about 56 audio hours — enough to run a real backlog through before deciding.

Auto-recharge is optional and off by default. With it off, requests are refused once the balance reaches zero rather than running up a bill.

Frequently asked questions

We focus exclusively on batch transcription, with transparent per-minute pricing and no add-on fees for diarization, timestamps, or repunctuation. If you process large volumes and want fast batch turnaround at a lower cost, we're built for that.

Not at launch. Speech Revolutions is batch-only today — upload audio, track progress via SSE, download structured output. Live streaming is on the roadmap.

Standard runs on dedicated compute, and most files come back in seconds. Economy is not available yet: when it ships it will use idle capacity at ~$0.09/hr with a 24-hour delivery SLA or explicit failure notification, for large backlogs where cost matters more than speed.

No. We do not use customer audio to train speech recognition models. Uploaded files and transcripts are deleted within 30 minutes of job completion.

JSON (with word timestamps and speaker labels), TXT, SRT, VTT, DOCX, and PDF — all from a single API call.

Up to 200 MB via direct REST upload, or up to 10 GB through the SDK using presigned multipart uploads.