SRT to text converter
Turn an SRT subtitle file into plain text: cue numbers and timing lines removed, each caption on its own line, formatting tags stripped. Keep a [mm:ss] timestamp at the start of each line if you want one. Paste the file or open it; it is converted in this tab and never uploaded.
The file is read with your browser’s FileReader and parsed in this tab. Nothing is uploaded — there is no server on the other end of this page.
Everything here runs on your device. Nothing you paste is uploaded.
What the conversion keeps and what it removes
An SRT file is a list of numbered captions. Each one has a cue number, a timing line saying when it appears and disappears, one or two lines of text, and a blank line to finish it. Only the text is the transcript, and getting it out cleanly means making a decision about each of the other parts.
- Cue numbers are removed. They only count the captions and say nothing about the words.
- Timing lines are removed. If you tick Keep a timestamp in front of each line, the start time of each caption survives as a short
[mm:ss]marker, and the end time is dropped. - Line breaks inside a caption become spaces. Subtitle editors wrap a caption onto a second line to keep it narrow on screen, and that wrap means nothing in running text, so each caption comes out as one line.
- The break between captions stays as a new line. The output has exactly one line per caption, in the order they appear in the file, and captions with no text are skipped.
- Formatting is removed with Tidy the cue text on, which is the default: italics and bold (
<i>,<b>), colour tags (<font color="…">) and positioning codes such as{\an8}, which some SRT files use to lift a caption to the top of the frame. Codes such as&become the character they stand for.
An SRT file going in:
1
00:00:00,500 --> 00:00:03,200
Right, let's pick up where
we left off last week.
2
00:00:03,600 --> 00:00:05,100
<i>[door closes]</i>
3
00:00:05,400 --> 00:00:09,000
The interview notes are in the shared folder,
under <i>Research</i>.The plain text coming out:
Right, let's pick up where we left off last week.
[door closes]
The interview notes are in the shared folder, under Research.Note what stayed. [door closes] is a sound description, the kind of line written for viewers who cannot hear the audio, and to the converter it is caption text like any other. The same goes for speaker labels such as PRIYA: and music symbols. What to do with them depends on what the text is for, and the tools further down this page handle each case.
With timestamps or without
Tick Keep a timestamp in front of each line and the same file comes out like this:
[00:00] Right, let's pick up where we left off last week.
[00:03] [door closes]
[00:05] The interview notes are in the shared folder, under Research.The marker is the caption’s start time, cut to the whole second. It switches to [hh:mm:ss] for captions that start after the first hour, so a long recording stays unambiguous. Keep timestamps when you need to get back to the recording from the text: pulling quotes from an interview, checking a disputed word against the audio, writing show notes that point to the moment a topic starts, or reviewing a lecture where you want to rewatch one explanation. They are also the honest choice when the text will be quoted, because anyone can check the words against the source.
Leave them off when the text is the destination. For reading, summarising, pasting into a document or counting words, a timestamp every few seconds is noise, and taking it out later is a chore. One subtlety if you do keep them: the time marks where a caption starts, not where a sentence starts. A sentence that begins halfway through one caption is stamped with that caption’s time.
Why the text still reads like subtitles
Captions are cut for the screen, not for the page. Netflix’s English style guide, for example, allows at most 42 characters per line and two lines per subtitle, and every caption has to stay up long enough to be read. So a sentence of any length is split across two or three captions, often at a comma or in the middle of a clause. Because this converter writes one line per caption, those splits survive: you get short lines that end mid-thought.
That is deliberate. Where a caption ends is information, and a converter that guessed at paragraphs would sometimes join things that belong apart, such as two speakers, or a sound description and the line after it. One line per caption keeps the output faithful to the file, and the rejoining is a separate, visible step you can check.
What subtitle text is good for, and where it falls short
When a caption file already exists, converting it is the quickest route to a transcript: the video you published with captions, a recorded lecture, a meeting that was transcribed as it happened. The text works well for show notes and summaries, for searching a long recording for the moment something was said, for pulling quotes, and for publishing a readable transcript next to the video. That helps people who would rather read than watch, and it puts the words on the page itself, where search engines can read them.
Keep its limits in mind. Subtitles are sometimes condensed to keep up with fast speech, and automatic captions mishear names and technical terms, so check any quotation against the recording before you publish it. Captions also describe what is heard, not what is shown: a transcript made from them will not mention a chart on screen or a name that only appears as text in the picture. If that matters, add it by hand at the point the timestamps show.
Cleaning the transcript up
Once the subtitles are plain text, two more passes turn them into something that reads like prose. Paste the result into the transcript cleaner first. Its unwrap pass joins a line onto the one above when the line above ended mid-sentence, which turns caption-length lines back into paragraphs. It can also strip speaker labels such as JOHN: or SPEAKER 1:, common markers such as [inaudible], [crosstalk] and [laughter], and the timestamps, if you kept them and changed your mind.
Then use the text cleaner for the finishing work. It can take filler words (um, uh, you know) out of a verbatim transcript, straighten or curl quotation marks, fix spacing around punctuation, and turn a file written entirely in capitals, which older broadcast captions often are, into sentence case. Each pass is a separate switch that reports what it changed. To see how long the finished text is, or how long it would take to read aloud, paste it into the word counter.
Automatic captions and VTT files
This page also reads WebVTT. The format is detected from the first line, so a file that starts with WEBVTT is handled as WebVTT: its header, NOTE and STYLE blocks and the positioning settings after each timing are skipped, and the caption text comes out the same way. That covers the transcripts Zoom and Microsoft Teams save as .vtt files. In Teams transcripts the speaker’s name sits in a voice tag, which tidying removes along with the name; untick Tidy the cue text if you want to keep the names and strip the tags yourself.
YouTube’s automatic captions need special handling. They scroll, so each caption repeats the line before it and adds a new one, and a converter that copies them line for line gives you every sentence twice. With tidying on, the converter recognises that pattern across the file and keeps each line once, stamped with the moment it first appeared, and a note under the stats says how many repeated lines were dropped. If you would rather go from WebVTT to SubRip than to text, the VTT to SRT converter applies the same clean-up.
A folder of subtitles at once
To turn a whole course or a podcast back catalogue into text, switch to the Batch tab and select every file. Each .srt becomes a .txt with the same name, so episode-12.srt becomes episode-12.txt, using whatever timestamp and tidying settings you have chosen. Changing a setting rewrites the whole list at once. A file that cannot be read is listed with the line number that stopped it, and the others still convert.
If you have the recording but no subtitle file, you need transcription rather than conversion. The MP3 to text and MP4 to text tools run the Whisper speech model in your browser and produce a transcript, with timestamps, from the audio itself.
Where VoiceSnap Pro fits
We make these tools because we are building VoiceSnap Pro, a dictation app for macOS and Windows. You hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is in. That suits the writing that follows a transcript, such as the summary, the show notes or the email about what was decided, which is usually quicker to say than to type. Every dictation is also saved to a searchable notes library. It will be a one-time $39 purchase, not a subscription.
It dictates what you say as you say it. It does not transcribe recordings or subtitle files, which is what the tools above are for. The app has not shipped yet, so there is nothing to download today. Join the waitlist and you will get one email when it is released.
Questions people ask
No. The file is read from your disk by the browser’s FileReader and converted by JavaScript running in this tab. The .txt you download is built in the tab’s memory, no request containing your subtitles is sent, and closing the tab discards everything.
Cue numbers and timing lines are removed. The lines inside a caption are joined with a space, so a two-line caption becomes one line of text, and each caption goes on its own line in the order it appears in the file. Captions with no text are skipped.
Yes. Tick Keep a timestamp in front of each line and every line starts with the caption’s start time, written as [mm:ss], or [hh:mm:ss] once the file passes the first hour. Times are cut to the whole second, and end times are not kept.
Because captions are cut to fit the screen and the time a viewer has to read them, so one sentence often runs over two or three captions. The text keeps one line per caption. To join those lines back into paragraphs, paste the result into the transcript cleaner, whose unwrap pass rejoins lines that end mid-sentence.
Yes, with Tidy the cue text ticked, which is the default. Italic, bold, underline and font colour tags are removed, and so are the positioning codes in curly brackets that some SRT files use to lift a caption to the top of the screen. Character codes, such as the one that stands for an ampersand, become the character itself. Untick it to keep the text exactly as written in the file.
Yes. The format is detected from the first line, so a file that starts with WEBVTT is read as WebVTT: its header, NOTE and STYLE blocks and cue settings are skipped, and the caption text comes out the same way. YouTube’s automatic captions, which repeat each line in the next cue, are recognised and every line is kept once.
They are part of the caption text, so they stay. The transcript cleaner can strip speaker labels such as JOHN: and common markers such as [inaudible] and [laughter]; for any other descriptions, find and delete them in your editor.