SayToTypeVoice Typing & Audio Transcription

Dictate anywhere. Transcribe any recording.

Two tools in one app: press a hotkey and speak to write in any application, or drop in a recording and get a full transcript with speakers, timestamps and subtitles.

Speak instead of type. Read what you already recorded.

Voice typing on all four platforms. File transcription runs on Windows and macOS.

Voice typing in any app

Press a hotkey and speak. The finished text lands where your cursor is — browsers, messengers, documents, code editors or the terminal.

1

Pick your hotkey

Setup takes a minute. Choose the key combination that suits you, and dictation is one keypress away in every application you use.

2

Speak into the app you are already in

Advanced speech recognition turns your voice into accurate, editable text and inserts it straight into whatever you are working in. No copy-paste, no switching windows.

3

Set the language, the translation and the tone

A mode fixes the input language, the output language and a prompt that shapes the result — meeting notes, a formal email, or a reply translated as you speak it.

Turn recordings into text

Meetings, interviews, lectures, videos — audio and video files alike: MP3, WAV, M4A, AAC, FLAC, OGG, OPUS, MP4, MOV, WEBM and more.

File transcription runs in the desktop app for Windows and macOS.

Drop in a file, get a transcript

Long recordings are split at natural pauses and transcribed part by part, with progress you can watch and cancel at any point. Search across every transcript you have made, or inside a single one. More on transcribing video and audio files

Speakers told apart automatically

Speakers are told apart and labelled, so an interview or a meeting reads as a conversation instead of a wall of text. Rename a speaker everywhere, or hand a single line to a different one.

TXT, SRT and VTT out of the box

Copy the result, or export it as plain text, SRT or VTT subtitles — with or without speaker names and timestamps. Play the recording back with the spoken line highlighted and correct any line by hand. More on making SRT and VTT subtitles

Subtitles in a second language

A finished transcript can be translated into another language on a tab beside the original. Timestamps and speaker names carry over, so subtitling a video in a second language takes a couple of clicks.

Key Features

Everything you need to speak instead of type — and to read what you have already recorded

Fast Transcription

Voice to text in seconds — fast speech recognition with low latency.

47 Languages

Speech recognition in 47 languages, with automatic language detection.

Works in Any App

Voice input for browsers, messengers, documents, code editors and the terminal.

Automatic Pasting

Speak to type directly into any text field — the text appears where your cursor is.

Audio and Video Files

Transcribe MP3, WAV, M4A, AAC, FLAC, OGG, OPUS, MP4, MOV, WEBM and more.

Speaker Labels

Speakers are told apart and labelled, so a meeting reads as a conversation.

Playback and Editing

Play a recording back with the spoken line highlighted, and correct any line by hand.

Subtitle Export

Export transcripts as text or SRT/VTT subtitles, with speaker names and timestamps.

Transcript Translation

Translate a finished transcript — timestamps and speakers carry over to the subtitles.

Search

Search across every transcript you have made, or inside a single one.

Custom Hotkeys and Modes

Start dictation with one keypress, and create modes that fix the language, translate your speech or shape the result with a prompt.

Use Your Own API

Connect OpenAI, Mistral AI, Deepgram or any OpenAI-compatible provider.

Inside the App

Your recognition, your rules

Choose who does the recognition

Use SayToType Cloud, or connect OpenAI, Mistral AI, Deepgram or any OpenAI-compatible provider. A file transcribed through your own key spends no minutes from your plan.

Settings that stay out of the way

The app lives in the system tray: pick your microphone, decide what happens while you record, and it stays invisible until you press the hotkey.

Who it's for

Everyone who talks faster than they type — and everyone sitting on recordings nobody has time to listen to

Anyone working with AI

Prompts, instructions and follow-up questions are long, and far faster spoken than typed — a paragraph of context costs you a sentence of breath.

Anyone who lives in messengers

Slack, Teams, Google Chat, Discord: reply at the speed you would say it, without taking your hands off what you were doing.

Journalists and researchers

Hours of interviews become searchable text with speakers told apart, so a quote is a search away instead of a re-listen.

Students and professionals

Lectures and meeting calls turn into notes you can read, search and quote the same day they happened.

Podcasters and video makers

Subtitles in SRT or VTT, with timestamps that survive translation — captioning an episode in a second language is a couple of clicks.

Developers

Speak your code comments, commit messages and pull request reviews without leaving the editor or the terminal.

Frequently asked questions

What people ask before they download

Which audio and video formats can SayToType transcribe?

MP3, WAV, M4A, AAC, FLAC, OGG, OPUS, MP4, MOV, WEBM and more. Audio and video both work, and long recordings are split at natural pauses and transcribed part by part.

How many languages does the speech recognition support?

47. You can fix the language for a mode or let the app detect it, and a mode can also translate your speech into another language as you dictate.

Does voice typing work in any application?

Yes. Press the hotkey, speak, and the finished text is inserted where your cursor is — browsers, messengers, documents, code editors or the terminal. No copy-paste, no switching windows.

Does it tell speakers apart?

Yes. In a file transcription speakers are told apart and labelled, so an interview or a meeting reads as a conversation. You can rename a speaker everywhere, or hand a single line to a different one.

Can I get subtitles for a video?

Yes. Any transcript exports as SRT or VTT subtitles, with or without speaker names and timestamps. Translate the transcript first and the timestamps carry over, so a second-language subtitle track is a couple of clicks.

How subtitle export works

Can I use my own OpenAI or Deepgram API key?

Yes. Alongside SayToType Cloud you can connect OpenAI, Mistral AI, Deepgram or any OpenAI-compatible provider, and a file transcribed through your own key spends no minutes from your plan.

Which platforms does SayToType run on?

Voice typing and file transcription both run in the desktop app for Windows and macOS. Dictation is also available on iOS and Android, in beta.

Is there a free plan?

Yes. Every account gets a monthly allowance of cloud transcription at no cost, and dictation and file transcription draw on the same balance. Paid plans add more minutes, and your own API key costs no minutes at all.

See plans and prices

Ready to stop typing?

Voice typing in every app you use, and a transcript for everything you have already recorded

3-4x
Faster than typing
47
Languages supported
<2 sec
Processing time

Interface available in 15 languages:

Bulgarian · Czech · Dutch · English · French · German · Italian · Polish · Portuguese · Romanian · Russian · Slovak · Spanish · Swedish · Ukrainian