Transcribe video and audio to text

Drop in a recording and get a transcript you can actually read — speakers told apart, every line searchable and correctable.

SayToType transcribes audio and video files on your Windows or Mac desktop. Long recordings are split at natural pauses and transcribed part by part, so a two-hour interview is one file, not twelve.

Download SayToType

Free to start, on Windows and macOS.

From recording to transcript

Four steps, all of them inside the desktop app.

1

Add the file

Any audio or video file the app supports — a meeting recording, a lecture, an interview, a screen capture. Long files are split at natural pauses and transcribed part by part, with progress you can watch and cancel at any point.

2

Read it as a conversation

Speakers are told apart and labelled automatically, so an interview or a meeting reads as a dialogue instead of a wall of text. Rename a speaker everywhere at once, or hand a single line to a different one where the recording was ambiguous.

3

Correct what matters

Play the recording back with the spoken line highlighted as it goes, and edit any line by hand. Names, jargon and half-swallowed words are where a transcript needs a human, and this is where you fix them.

4

Keep it, search it, export it

Every transcript stays in the library. Search across all of them at once when you remember a phrase but not which recording it was in, or search inside a single long one. Copy the text out, or export it as a file.

What you get back

A transcript that is worth reading, not just a dump of words.

47 languages

Recognition covers the same 47 languages as the rest of the app, worked out from the recording itself. A finished transcript can then be translated into a language you pick.

Speaker labels

Who said what, worked out automatically and editable afterwards. This is what turns a raw recording into something you can quote from.

Timestamps throughout

Every line is anchored to its moment in the recording, so you can jump back to the audio for any sentence in the transcript.

Search across everything

One search box over every transcript you have made, and a second one inside the transcript you are reading.

Translation built in

A finished transcript can be translated into another language on a tab beside the original, with timestamps and speaker names carried over.

Your own API key

Alongside SayToType Cloud you can connect OpenAI, Mistral AI, Deepgram or any OpenAI-compatible provider. A file transcribed through your own key spends no minutes from your plan — you pay that provider at its own rates instead.

Formats it accepts

Audio and video alike — if it plays, it usually transcribes.

Audio

  • MP3
  • WAV
  • M4A
  • AAC
  • FLAC
  • OGG
  • OPUS

Video

  • MP4
  • MOV
  • WEBM

…and more. A video is read for its sound, so a screen recording transcribes exactly like a voice memo does.

Who transcribes with it

The same feature, four different reasons to want it.

Journalists and researchers

Interviews come back as a labelled conversation with timestamps, so quoting someone means finding the line rather than scrubbing the audio.

Students and lecturers

A recorded lecture becomes a searchable document. Look up the ten minutes that mattered instead of replaying ninety.

Teams that record meetings

A meeting recording turns into text you can skim, search and paste into notes, with each participant on their own labelled line.

Anyone with a backlog

Voice memos, old podcast episodes, screen recordings — anything already sitting on disk becomes text you can search.

Questions about file transcription

Which audio and video formats can SayToType transcribe?

MP3, WAV, M4A, AAC, FLAC, OGG, OPUS, MP4, MOV, WEBM and more. Audio and video both work, and long recordings are split at natural pauses and transcribed part by part.

How long a recording can I transcribe?

Long recordings are handled by splitting them at natural pauses and transcribing them piece by piece, with progress you can watch and cancel at any point. What limits you in practice is the transcription minutes on your plan, not the length of a single file.

Does it tell speakers apart?

Yes. Speakers are told apart and labelled, so an interview or a meeting reads as a conversation. You can rename a speaker everywhere, or hand a single line to a different one.

How many languages does it recognise?

47, the same set the rest of the app uses. The spoken language is worked out from the recording, and a finished transcript can then be translated into a language you choose, line by line, so timestamps and speaker names keep working on the translation.

Is my recording stored anywhere?

Not by us. The recording is processed in real time, and nothing — neither the audio nor the text — is kept on our servers, on free plans included. The transcript you get back lives only on your own device, in the app’s library.

Can I transcribe files on my phone?

Not yet. File transcription runs in the desktop app for Windows and macOS. The iOS and Android apps are dictation only, in beta.

Turn your recordings into text

SayToType runs on Windows and macOS, and transcribes the files you already have.

Free to start — every account gets 50 minutes of cloud transcription a month at no cost, and no card is needed to sign up.