MP4 to text converter
Drop in an MP4 and get back a transcript you can edit, copy, or save as SRT or VTT captions. Your browser decodes only the audio track, and OpenAI's Whisper model transcribes it on your own device, so the video is never uploaded.
Drag an MP4 video here, or choose one below.
Only the soundtrack is used, and it is decoded by your browser. MOV and WebM videos and ordinary audio files work too, if your browser can decode them.
Choose a file first.
Your file is transcribed by the Whisper model running on your own device. The audio is never uploaded; the only downloads are the model files from Hugging Face and the transformers.js library that runs them, from jsDelivr, fetched when you first press Transcribe and cached by your browser.
How this MP4 to text converter works
A video file is two streams in one box: the picture and the sound, wrapped in a container. Transcription needs only the sound, so that is all this page touches.
- Decoding. The file is read from your disk into the tab, and your browser’s own media decoder pulls out the audio track. The video frames are never decoded or drawn.
- Preparing. The audio is mixed down to one channel and resampled to 16 kHz, the format Whisper expects.
- Transcribing. The samples go to a Web Worker — a background thread, so the page never freezes — running OpenAI’s open Whisper model through the transformers.js library. Long recordings are worked through in 30-second chunks, the window Whisper was built around, and every chunk comes back with timestamps.
- Output. The transcript appears as timestamped segments you can edit, then copy or download as plain text, SRT or WebVTT.
The transcriber fetches nothing when the page loads, and nothing while your file is decoded either. The library (transformers.js, pinned to version 4.3.0) and the model are downloaded only when you press Transcribe, and the browser keeps them afterwards, so it is a one-time download unless you clear the site’s data.
Only the audio track: what that means for a video
Because the picture is ignored, it makes no difference to the transcript whether the video is a talking head, a slide deck or a black screen. It also means anything that exists only as pixels is invisible to it: the text on a slide, a terminal in a screen recording, captions already burned into the frame. This is speech recognition, not text recognition, and it does not tell speakers apart either — a three-person meeting comes out as one continuous transcript.
MP4 is not the only thing it opens. Any file whose audio your browser can decode will work: Chrome decodes the sound from MP4, MOV and WebM video, and plain audio such as M4A or MP3 opens too. Safari cannot reliably decode the audio in a WebM file, so use Chrome, Edge or Firefox for those. A recording made with no audio track at all — a screen capture with the microphone and system sound both off — has nothing to decode, and the page reports that it could not decode the file rather than handing you an empty transcript.
Recordings worth turning into text
- Screen recordings. A walkthrough, a bug report, a product demo. The transcript becomes the written steps, the ticket or the description under the video. The macOS screenshot toolbar saves recordings as MOV and browser-based recorders usually produce WebM; both open here.
- Meeting recordings. A Zoom recording often comes with an audio-only M4A file as well as the MP4; current versions name it
audiofollowed by a string of digits. If you have one, transcribe the M4A: it is the same sound without the weight of the video. - Lectures and talks you recorded. These tend to be long and large, which is what the section on file size below is about.
- Social clips you made. Short, often vertical, and often started muted by the feed they land in, which is the whole case for captions. Transcribe the clip, export an SRT and add it in your editor or when you upload.
Turning a video into SRT or VTT captions
The timestamps are what make a video transcript more than a text file. Download SRT or VTT and each segment becomes a caption cue that appears when the words were spoken.
- SRT for a video editor. Premiere Pro, DaVinci Resolve and Final Cut Pro all import SRT onto a caption track. Final Cut Pro does not import WebVTT, so SRT is the safe choice whenever an editor is involved.
- VTT for your own website. The HTML
<track>element reads WebVTT, not SRT, so if you are embedding the video yourself this is the file to point it at. - Either for YouTube. YouTube accepts SRT and WebVTT when you upload subtitles in YouTube Studio. It also generates automatic captions itself for many languages; uploading your own file is for when you want to read and correct the words before viewers see them.
Correct the transcript here before you export, because a caption file puts every misheard name straight onto the screen. Every segment is editable, the downloads use your edited text, and clicking a segment’s time plays the recording from that moment so you can check it by ear. Where the model’s own timestamps broke down and the page had to estimate them, the time is marked ≈; nudge those cues in your editor.
Each segment becomes one cue, and Whisper decides where segments end by its own sense of a phrase: some run to two sentences, some stop mid-sentence. Press Enter inside a long segment to break it over two lines of the same cue, or split it into separate cues in your editor. Already have captions in one format and need the other? The SRT to VTT converter and the VTT to SRT converter switch between them without moving a single timing, and SRT to TXT turns a caption file back into readable text.
Large video files and the 200 MB limit
The page refuses any file over 200 MB, and with video that limit arrives sooner than you might expect. Before the browser’s decoder can find the audio, the whole file has to be read into the tab’s memory, picture included — and the picture is nearly all of the bytes. At the 8 Mbps YouTube recommends for a 1080p upload, the video stream alone is 60 MB a minute, so a 1080p recording reaches the limit a little over three minutes in. A soundtrack in AAC at 128 kbps, by contrast, is under 1 MB a minute.
So for anything long, extract the audio first and transcribe that. It takes a moment:
- On a Mac, open the video in QuickTime Player and choose File → Export As → Audio Only. You get an M4A file with an AAC audio track.
- Anywhere, with ffmpeg, run
ffmpeg -i talk.mp4 -vn -c:a copy talk.m4a.-vndrops the video and-c:a copykeeps the audio stream exactly as it is, without re-encoding, so it finishes in seconds. That works when the audio is AAC, as it is in most MP4 files. - From Zoom, use the audio-only
.m4afile if the recording came with one.
Then drop the M4A here: same page, same transcriber. The second limit is length. We recommend files under about 30 minutes, because decoding expands the compressed audio into uncompressed samples in memory before it is resampled, and transcription time grows with every minute. A longer file still runs, with a warning. For a two-hour lecture it is safer to split the audio and transcribe the parts in turn — ffmpeg -i talk.m4a -f segment -segment_time 1500 -c copy part%02d.m4a cuts it into 25-minute files without re-encoding.
Fast or Accurate, and how long it takes
Two sizes of Whisper are on offer. Fast is whisper-tiny, about a 41 MB download; Accurate is whisper-base, about 77 MB. Each is downloaded from Hugging Face the first time you use it, with progress shown, and cached by the browser after that. Fast is quickest and copes with clear speech in a quiet room. For captions other people will read, Accurate mishears fewer words and is worth the extra download.
Speed depends mostly on your machine. The page runs the model on your graphics processor through WebGPU when the browser supports it, and falls back to WebAssembly on the CPU when it does not, or when the graphics chip fails part-way; the Processor setting can also force the CPU. On the CPU, transcription is slower than real time on some machines — a ten-minute video can take more than ten minutes — so try Fast on a short clip first. Progress is shown as it goes, and Cancel stops it at any point, keeping whatever has been transcribed so far.
Leave the language on auto-detect or pick it from the list. Detection listens to the first 30-second chunk that contains any sound, so if a video opens with a music intro or someone speaking a different language, choose the language yourself; when the model is unsure of its guess, the page says so. The Translate to English switch uses Whisper’s own translation task to write English text from speech in another language, timestamps included — a quick route to rough English subtitles for a foreign-language clip. It only translates into English, and with models this small the result is a draft to check, not a subtitle to publish unread.
Check the transcript before it becomes captions
Whisper is good, not infallible, and the two small versions that fit in a browser tab make more mistakes than its larger ones. Three places to look first:
- Names and jargon. Surnames, product names and acronyms are where it slips most, and they are exactly the words a viewer notices.
- Music, noise and repetition. Over a music bed or background noise, Whisper can produce text nobody said — a phrase repeated over and over, or a stray sign-off line. The page skips any 30-second chunk that is silent throughout and highlights segments where it catches the model repeating itself, but it cannot catch every invention. Check the segments around intros, outros and gaps.
- Crosstalk. Where two people talk over each other, words can be dropped or run together, and with no speaker labels nothing tells you whose words survived. Click the segment’s time and listen again.
If what you need in the end is prose rather than captions, download the .txt: it is the words alone, without timestamps, with a new paragraph wherever the speaker paused for two seconds or more. For a transcript from somewhere else that arrived with timestamps and speaker labels baked in, the transcript cleaner strips them out.
What leaves your device
Your video does not. The audio track is decoded in the tab and transcribed in a worker in the same tab; neither the file nor its samples are sent anywhere. Two things are downloaded, and only after you press Transcribe: the transformers.js library and its WebAssembly runtime from jsDelivr, and the Whisper model from Hugging Face. You can check this yourself — open the browser’s network panel and run a transcription. Apart from the page’s own worker script, every request the transcriber makes goes to jsDelivr or Hugging Face, and every one of them is a download; your file and its audio are never sent.
One other service shows up in that panel: this site’s cookieless page-view counter, Fathom, which runs on every page. It counts the visit to the page and carries nothing from your file or its transcript.
Where VoiceSnap Pro fits
This converter is for recordings that already exist. VoiceSnap Pro is for the words you have not said yet. It is a voice-to-text dictation app for macOS and Windows: hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is already in — an email, a Slack message, a document. Every dictation is also saved to a searchable notes library. It is a one-time $39 purchase, not a subscription.
It does not transcribe video or audio files — that is what this page is for — and it has not shipped yet, so there is nothing to download today. Join the waitlist for a single email on release day. If you are weighing it against running Whisper yourself, VoiceSnap Pro vs Whisper sets out what each one does. For recordings that are already audio, the audio to text converter and MP3 to text use the same transcriber as this page.
Questions people ask
No. Your browser reads the file and decodes its audio track inside the tab, and Whisper transcribes it in a background worker in the same tab. The only things the transcriber downloads are the transcription library and its WebAssembly runtime, from jsDelivr, and the Whisper model, from Hugging Face, and it fetches them only after you press Transcribe. Your video and its audio are never sent anywhere.
Files over 200 MB are refused, and the limit counts the whole file, picture included, because the browser has to read all of it into memory before its decoder can find the audio. Extract the audio first and transcribe that instead: in QuickTime Player choose File, Export As, Audio Only, or run ffmpeg -i video.mp4 -vn -c:a copy audio.m4a. The audio track is usually a small fraction of the file.
Yes. Download SRT or VTT and each timestamped segment becomes a caption cue. YouTube accepts both formats when you upload subtitles. Premiere Pro, DaVinci Resolve and Final Cut Pro all import SRT, and Final Cut Pro does not import WebVTT, so SRT is the safer choice for an editor. Read the transcript through and correct misheard names before you export.
Any file whose audio your browser can decode. Chrome decodes the audio from MP4, MOV and WebM video, and audio files such as M4A, MP3 and WAV work too. Safari cannot reliably decode WebM, so open those in Chrome, Edge or Firefox. A video with no audio track at all has nothing to transcribe, and the page tells you it could not decode the file.
It depends on your hardware, the model you pick and the length of the video. The page runs Whisper on your graphics processor through WebGPU when your browser supports it, and otherwise on the CPU through WebAssembly, where transcription can be slower than real time on some machines. The first run also downloads the model, about 41 MB for Fast or 77 MB for Accurate, which the browser then keeps. Try a short clip on Fast first to see what your machine does.
Neither. Whisper turns speech into text with timestamps; it does not tell voices apart, so a meeting comes out as one continuous transcript without speaker names. Only the audio track is decoded, so slides, on-screen text and captions burned into the picture are ignored.
Yes. Turn on Translate to English and Whisper writes English text from the speech, still with timestamps, so the SRT and VTT downloads become English subtitles. Whisper only translates into English, and with the small models that run in a browser the result is a draft to check, not a finished subtitle.