Make subtitles from a video — SRT and VTT

Get a properly timed caption file out of a recording, ready to upload or drop onto a timeline.

SayToType times every line against the audio, lets you correct the ones that came out wrong, and exports a standard .srt or .vtt file — with or without speaker names and timestamps.

Download SayToType

Free to start, on Windows and macOS.

From video to caption file

Four steps, and only one of them needs your attention.

1

Add the video

MP4, MOV, WEBM and more, alongside plain audio files. A video is read for its sound, so a screen capture works exactly like camera footage. Long recordings are split at natural pauses and processed part by part, with progress you can watch and cancel.

2

Check the timings

Play the recording back with the current line highlighted as it goes. Where a caption starts a beat early or the recognition mangled a name, fix that line by hand before you export.

3

Choose what goes in the file

Export with or without speaker names, with or without timestamps. A two-person interview can ship with “Anna:” in front of every line, or as clean captions with no labels at all.

4

Add a second language, if you need one

Translate the finished transcript on a tab beside the original. Timestamps and speaker names carry over, so a second subtitle track is a couple of clicks rather than a second pass over the video.

SRT or VTT — which do you need?

Two formats, one rule of thumb.

Choose SRT when

You are uploading to a video platform or handing the file to an editor. It is the most widely accepted caption format there is, and the safe default when nobody has told you otherwise.

Choose VTT when

The captions are going onto a web page — VTT is the format an HTML5 video element expects for its text track.

Both are plain text. You can open either in any editor, read the timings, and see exactly what will appear on screen — which also means you are never locked out of your own captions.

What the export gives you

A file that behaves like a caption file should.

SRT and VTT, both standard

Plain, well-formed caption files. Nothing proprietary, nothing that needs SayToType installed to open on the other end.

Watch it before you export

Play the video inside the app with the captions running on it, full screen if you want, with the transcript and the translation beside it. What you see there is what the exported file will do.

Timings you can fix

Captions are only as good as their timing. Every line is editable against playback, with the spoken line highlighted as the audio runs.

Speaker names optional

Keep the labels for an interview or a panel, drop them for a single-voice narration. The same transcript exports either way.

A translated track from the same file

Translate the transcript and the timings come with it, so the second-language track lines up with the first without redoing the work.

47 languages recognised

Recognition covers 47 languages, worked out from the recording itself. Subtitling something you did not record yourself is the normal case, not the exception.

Who needs captions

Four jobs that all end in the same file.

Video editors

A timed caption file that drops onto the timeline, instead of typing captions against a scrub bar for an afternoon.

Creators publishing online

Most people watch with the sound off at least some of the time. Captions are the difference between being watched and being scrolled past.

Anyone localising a video

One recording, two subtitle tracks: the original and a translated one that keeps the same timings.

Teams making internal video

Training clips and recorded walkthroughs become searchable and accessible without a separate captioning step.

Questions about subtitle export

What is the difference between SRT and VTT?

Both are plain-text caption files that pair lines of dialogue with timings. SRT is the older and more widely accepted of the two and is the safe default for video platforms and editing software. VTT is the format used for the text track of an HTML5 video on a web page. SayToType exports either.

Can I edit the timings before exporting?

Yes. Play the recording back with the spoken line highlighted and correct any line by hand. What you export is what you last corrected.

Can I keep speaker names in the captions?

Yes, and you can leave them out. Export runs with or without speaker names and with or without timestamps, from the same transcript.

Can I subtitle a video in a different language from the audio?

Yes. Translate the finished transcript on a tab beside the original and export that. Timestamps and speaker names carry over, so the translated track keeps the timing of the original.

Which video files can I caption?

MP4, MOV, WEBM and more, plus audio files such as MP3, WAV, M4A, AAC, FLAC, OGG and OPUS. A video is read for its sound, so a screen recording works the same as camera footage.

Do I need to be online?

Yes, for file transcription. The recognition runs on a service — SayToType Cloud by default, or OpenAI, Mistral AI or Deepgram if you supply your own key — so the file is transcribed online.

Caption your next video

SayToType runs on Windows and macOS and exports SRT and VTT from any recording you give it.

Free to start — every account gets 50 minutes of cloud transcription a month at no cost, and no card is needed to sign up.