The recordings you already have
Drop in an interview, a lecture or a screen capture and get a transcript back. Long files are split at natural pauses and processed part by part, and everything you have transcribed stays searchable in one library.
Both apps do the same core thing well: press a hotkey, speak, and the text lands in whatever app you are in.
The difference is what happens to recordings you already have. SayToType transcribes audio and video files too — with speakers told apart, timestamps, and SRT or VTT subtitle export — which is the part Wispr Flow does not describe on its site.
Free to start, on Windows and macOS.
Checked against Wispr Flow's own site, not from memory.
| Feature | SayToType | Wispr Flow |
|---|---|---|
| Dictation into any app | Yes | Yes |
| Desktop | Windows, macOS | Windows, macOS |
| Mobile | iOS, Android (beta) | iOS, Android |
| Recognition languages | 47 | 100+ (their figure) |
| Transcribe audio and video files you already have | Yes | Not described on their site |
| Speakers told apart in a transcript | Yes | Not described on their site |
| SRT and VTT subtitle export | Yes | Not described on their site |
| Translate a finished transcript | Yes | Not described on their site |
| Use your own API key | OpenAI, Mistral AI, Deepgram, any OpenAI-compatible | Not described on their site |
| On-device offline dictation | macOS on Apple Silicon | Not described on their site |
| Free tier | Monthly allowance of transcription minutes, no card needed | 2,000 words a week on desktop |
| Paid plan | Regional pricing — see the pricing page | From $15 per user per month |
Wispr Flow details compared on 7 September 2026, from publicly available information on wisprflow.ai. Products change — check theirs before deciding.
Almost all of it is the same idea: the recording you already made is also text.
Drop in an interview, a lecture or a screen capture and get a transcript back. Long files are split at natural pauses and processed part by part, and everything you have transcribed stays searchable in one library.
Speakers are told apart and labelled automatically, so a meeting reads as a conversation. Rename a speaker everywhere at once, or reassign a single line where the recording was ambiguous.
Any transcript exports as SRT or VTT, with or without speaker names and timestamps. Translate it first and the timings carry over, so a second-language subtitle track costs a couple of clicks.
Connect OpenAI, Mistral AI, Deepgram or any OpenAI-compatible endpoint. Work run through your own key spends no minutes from your plan, which changes the arithmetic if you transcribe a lot.
Local modes run on-device Whisper models on an M-series Mac — six of them, two free — so dictation can happen with no connection and nothing leaving the machine.
A mode fixes the recognition language, an optional language to translate into as you speak, and a prompt that shapes the result, so dictating a commit message and dictating an email give different output.
Three cases where we would not argue.
Wispr Flow advertises over 100 recognition languages against SayToType’s 47. If the language you work in is not among ours, that decides it.
Their iOS and Android apps are generally available. SayToType dictates on both, but those apps are still in beta.
If you never work with recorded files, the whole second half of SayToType is weight you are not using, and the comparison comes down to preference.
Both apps are good at the thing they share. If dictation into any app is the whole job, you will be fine either way — try both and keep the one whose hotkey and output you prefer.
Superwhisper comes up in the same searches and overlaps more than it used to: it runs on macOS, Windows and iOS, transcribes uploaded audio and video files, and does its recognition on-device, which is its main draw. A page comparing it properly is worth writing rather than summarising in a paragraph, so this one deliberately stops here.
File transcription. SayToType turns audio and video recordings you already have into transcripts with speakers told apart, timestamps, search, translation, and SRT or VTT subtitle export. Wispr Flow’s site describes dictation and does not describe transcribing uploaded files.
Yes. Both insert text at the cursor in whatever app you are using, from a global hotkey, on Windows and macOS.
Wispr Flow, by its own figure of over 100. SayToType recognises 47, and can additionally translate into another language as you dictate or after a transcript is finished.
SayToType has local dictation modes on Macs with Apple Silicon that run on-device Whisper models with no connection. File transcription is online either way. Wispr Flow’s pricing page does not describe an offline or on-device mode.
It depends on where you are and how much you use. SayToType prices regionally through Paddle, so the pricing page is the only honest answer for your country. Both have a free tier: Wispr Flow caps its free plan at 2,000 words a week on desktop, SayToType gives every account a monthly allowance of transcription minutes shared between dictation and file transcription.
See plans and pricesDictation on Windows and macOS, plus the transcripts, speaker labels and subtitles for the files already on your disk.
Free to start — every account gets a monthly allowance of cloud transcription at no cost, and no card is needed to sign up.