VTT to SRT converter
Convert a WebVTT .vtt file to SubRip .srt: paste it or open it, and the SRT appears as you go, with cues numbered, dots swapped for commas and the WebVTT-only parts removed. Convert a whole folder in one go. The work happens in this tab, so your captions are never uploaded.
The file is read with your browser’s FileReader and parsed in this tab. Nothing is uploaded — there is no server on the other end of this page.
Everything here runs on your device. Nothing you paste is uploaded.
What changes when WebVTT becomes SRT
WebVTT was designed for web video, so it carries things a browser needs: a signature line, comments, CSS, and settings that place each caption on screen. SubRip is older and plainer. It has no formal specification, just a convention every player agrees on: a number, a timing line, the caption text, a blank line. Converting from one to the other means keeping everything SRT can express and removing the rest. Here is each part of a WebVTT file and what becomes of it.
- The header. A WebVTT file must open with
WEBVTT, and the same line may carry a title (WEBVTT - Lecture 2). Some exporters add header lines under it, such asKind: captionsandLanguage: enin files downloaded from YouTube. SRT has no header, so all of it goes. - Timestamps. WebVTT puts a dot before the milliseconds and lets you leave out the hours:
00:04.500is valid. SRT uses a comma and always writes the hours, so the converter writes every time in full as00:00:04,500. The values themselves do not move by a millisecond. - Cue identifiers. In WebVTT the line above a timing is optional and can be any text:
intro,chapter-3, or the long unique ID that Microsoft Teams puts above every cue in a transcript. SRT expects a plain count, so cues are numbered 1, 2, 3 in file order and the WebVTT identifiers are discarded. - Cue settings. Anything after the end time (
vertical,line,position,size,alignandregion) tells a web player where to draw the caption. SRT has no equivalent, so these are dropped and the player puts every caption in its default place, usually bottom centre. Some players honour a code such as{\an8}at the start of an SRT caption to lift it to the top, but that is a convention borrowed from another format, and there is no faithful way to translateline:10%into it, so the converter does not guess. - NOTE, STYLE and REGION blocks. Notes are comments for whoever edits the file, STYLE blocks hold CSS for
::cue, and REGION blocks define areas of the frame for scrolling captions. None of them has anywhere to go in SRT, so all three are removed, taking any colours and fonts with them. - Tags inside the caption text. Bold, italic and underline (
<b>,<i>,<u>) are understood by most SRT players and are kept. WebVTT’s own tags, which are voice<v Priya>, class<c.yellow>, language<lang>, ruby annotations and word-timing tags such as<00:01.200>, are removed with Tidy the cue text switched on, and the words inside them are kept. - Character codes. WebVTT has to write an ampersand as
&and a less-than sign as<. SubRip has no escaping at all, so those codes are turned back into the characters they stand for. The exception is a tag written out as text, such as<i>in a caption about formatting: decoded, it would become real italics in the SRT, so it is left as it was.
A short WebVTT file with most of those features:
WEBVTT - Lecture 2
NOTE Checked against the recording on 12 March.
intro
00:00:01.000 --> 00:00:04.200 align:start line:85%
<v Priya>Welcome back to the second lecture.</v>
00:04.500 --> 00:08.000
Questions go in the chat & I will
take them at the end.And the SRT the converter writes from it:
1
00:00:01,000 --> 00:00:04,200
Welcome back to the second lecture.
2
00:00:04,500 --> 00:00:08,000
Questions go in the chat & I will
take them at the end.Notice what went with the voice tag. SRT has no field for a speaker’s name, so “Priya” disappears along with the tag that held it. If the name matters on screen, type it into the caption (PRIYA: Welcome back…) before you convert, or untick Tidy the cue text to keep every tag exactly as it was written and deal with them in your editor.
Where .vtt files come from
Most people with a VTT file did not choose the format. It is simply what the software they used hands back:
- Meeting recordings. With audio transcription turned on, Zoom saves the transcript of a cloud recording as a
.vttfile next to the video, and Microsoft Teams offers its transcripts as.vttor.docx, with each speaker’s name in a voice tag. - YouTube. If the video is yours, YouTube Studio lets you download a caption track as
.srt,.vttor.sbv, so you may be able to skip the conversion entirely. Automatic captions fetched with a download tool arrive as WebVTT in a scrolling layout, which has its own section below. - Speech-to-text tools. OpenAI’s Whisper writes
.vttalongside.srt,.txt,.tsvand.jsonunless you ask for one format, and many transcription services offer VTT because it drops straight into a web player. The MP4 to text and MP3 to text tools on this site can export either format directly. - Web video. WebVTT is the caption format of the HTML
<track>element, and HLS streams usually carry their subtitles as WebVTT, so caption files saved from web players and video platforms tend to be VTT.
The places that want SRT instead are just as common: video editors, the media player on a smart TV reading a subtitle file from a USB stick, and upload forms that list SRT first or accept nothing else. That mismatch is the whole reason this page exists.
YouTube automatic captions: the repeated-line problem
Captions that YouTube generates automatically are built to scroll. Each cue shows the previous line again above a new one, the new words carry a timing tag for every word so they can appear one at a time, and a cue lasting a fraction of a second sits between each pair to move the text up. In a browser that looks smooth. Converted line for line to SRT, it looks like every sentence is said twice, and the file is cluttered with tags such as <00:00:01.669><c> since</c>.
With Tidy the cue text on, the converter looks at the whole file. If at least half of the cues begin with the line the previous cue ended on, and a good share of those add a new line underneath it, it treats the file as rolling captions: the word-timing tags are stripped, the repeated lines and the fraction-of-a-second cues are dropped, and each line is written once, starting at the moment it first appeared. A note under the stats says how many repeated lines were removed.
Because the check is made across the whole file rather than one cue at a time, ordinary files rarely trip it. Two consecutive captions that both say “No.” are left alone, and so is a run of identical [applause] captions, since nothing scrolls. If a file is ever collapsed when it should not have been, the note under the stats will say so, and unticking Tidy the cue text brings every cue back.
These files often have a second oddity that stops most converters outright: the words of a cue sit one line below where they should, separated from their timing line by a line containing only a space. In a WebVTT file, text found on its own straight after an empty cue is attached to that cue rather than reported as an error. The exception is a block that looks like a damaged cue, one that starts with a cue number or contains what looks like a mistyped timing line: that is still reported with its line number, rather than turned into caption text.
Converting a folder of VTT files
A lecture series, a season of a podcast or a quarter’s worth of meeting recordings means dozens of files, and the Batch tab is built for that. Choose them all at once in the file picker. Each one is read, parsed and converted with the same settings, and listed with its own Copy and Download buttons. Download all saves them one after another, each keeping its name with the new extension: week-04-seminar.vtt becomes week-04-seminar.srt.
Every file in the list is converted from the copy already in memory, so switching the output to plain text, or turning tidying off, rewrites the whole batch instantly without reading anything again. An SRT file that has slipped into the selection is harmless: it is renumbered, its timings are rewritten in the standard layout, and it is saved alongside the rest. A file that cannot be parsed is marked as failed with the line that stopped it, and does not hold up the others.
When a VTT file will not convert
The error names a line number, and it is nearly always one of a few causes:
- The
WEBVTTline is missing. Without it the file is read as SRT. Timings written with dots are still accepted, but aNOTEorSTYLEblock then looks like a caption with no timing line and is reported as one. PutWEBVTTback on the first line. - A broken arrow. The separator must be
-->, with two hyphens. A hand-edited->or an en dash pasted from a word processor will not be recognised. - An end time before its start. Usually a typo in the hours or minutes after manual editing. The error quotes both times so you can see which is wrong.
- Mangled characters. WebVTT must be UTF-8. A file that was re-saved in another encoding arrives with accented letters turned into symbols; re-save it as UTF-8 from your text editor and open it again.
Some things that look wrong are fine. A fraction written with one or two digits, such as 00:01.5, is read as 500 milliseconds and written out in full. A missing blank line between two cues is tolerated. A caption that happens to contain an arrow (“from 12 --> 48”) is treated as text, not as a new cue.
Checking the SRT before you use it
Two quick checks catch nearly every problem. First, compare the first and last cues with the originals: same start times and the same number of captions, except for rolling captions, where the count is meant to fall. Second, play it. VLC and most desktop players load an .srt automatically when it sits in the same folder as the video and shares its name, so talk.mp4 and talk.srt together are enough. If captions you had positioned at the top of the frame now cover a name caption at the bottom, that is the dropped cue settings, and your editor is the place to move them back.
If you need the words rather than the captions, choose Plain text under Convert to, or use the SRT to text page, which explains the options for turning captions into a readable transcript. For the opposite direction, SRT to WebVTT, the subtitle converter opens with WebVTT selected.
Where VoiceSnap Pro fits
This converter is free and stays free. It exists because we are building VoiceSnap Pro, a dictation app for macOS and Windows: hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is in, whether that is a caption editor, a video description or the show notes you write after a recording. Every dictation is also saved to a searchable notes library. It will be a one-time $39 purchase, not a subscription.
Two honest caveats. It is a dictation app, not a file transcriber, so it will not turn a recording into subtitles; the MP4 to text tool does that in your browser. And it has not shipped yet, so there is nothing to download today. Join the waitlist and you will get one email when it is released.
Questions people ask
No. A file you open is read from your disk by the browser’s FileReader, the conversion is JavaScript running in this tab, and the SRT you download is built in the tab’s memory. The page sends no request containing your captions, and closing the tab discards everything.
No. Every start and end time keeps its exact millisecond value. Only the way it is written changes: SRT uses a comma before the milliseconds where WebVTT uses a dot, and SRT always writes the hours, so 00:04.500 becomes 00:00:04,500.
They are dropped. SRT has no way to position an individual caption, so the player decides where every caption goes, usually bottom centre. If a caption was moved to keep it clear of on-screen text, note which one and reposition it in your video editor after converting.
No. NOTE blocks are comments for whoever edits the file, and STYLE and REGION blocks hold CSS and layout areas for web players. SRT has nowhere to put any of them, so they are removed, and colours or fonts defined in a STYLE block are lost with them.
With Tidy the cue text ticked, which is the default, the tag is removed and the words inside it are kept. SRT has no field for a speaker’s name, so the name goes with the tag. If the names need to appear on screen, add them to the caption text before converting, or untick Tidy the cue text to leave every tag exactly as written.
Usually, yes. YouTube’s automatic captions scroll: each cue repeats the previous line above a new one, and the new words carry word-by-word timing tags. With Tidy the cue text ticked, the page recognises that pattern across the whole file, removes the timing tags and drops the repeated lines, so each line appears once, starting when it first appeared. A file without the pattern is left alone.
Yes. Switch to the batch tab and select as many .vtt files as you like. Each one becomes an SRT file with the same name, lecture-02.vtt becoming lecture-02.srt, and Download all saves them one after another. A file that will not parse is listed with the line number that stopped it, and the rest still convert.