MP3 to text converter
Choose an MP3 — a podcast episode, a lecture, an interview, a voice memo — and OpenAI's Whisper model transcribes it inside this tab. You get an editable, timestamped transcript to copy or save as TXT, SRT or VTT, and the audio never leaves your device.
Drag an MP3 here, or choose one below.
Any other audio or video file your browser can decode works as well: M4A, WAV, FLAC, or the soundtrack of an MP4.
Choose a file first.
Your file is transcribed by the Whisper model running on your own device. The audio is never uploaded; the only downloads are the model files from Hugging Face and the transformers.js library that runs them, from jsDelivr, fetched when you first press Transcribe and cached by your browser.
What happens to an MP3 when you press Transcribe
Your browser’s own MP3 decoder — the one that plays audio on any web page — turns the file into raw samples. The page averages the channels into one and resamples the result to 16 kHz, because that is the format Whisper expects: the model was trained on 16 kHz audio, and it turns every input into a log-Mel spectrogram before it reads a single word.
Those samples are handed to a Web Worker, which runs on its own thread, so the page keeps responding while the model works. Only when you start a transcription does the worker fetch what it needs: Hugging Face’s transformers.js library and the WebAssembly runtime it runs the model with, both pinned to exact versions and served from jsDelivr, and the Whisper model you chose, from Hugging Face. Opening this page downloads none of them. Your browser caches them, so the next file starts without the wait, unless the cache has been cleared in between.
If your browser exposes WebGPU, the model tries your graphics chip first. If WebGPU is missing or fails, the tool falls back to WebAssembly on the CPU, and the Processor setting lets you choose the CPU outright. The CPU is slower — on some machines slower than real time, so a ten-minute MP3 can take more than ten minutes. On either path, your file stays in the tab.
Fast or Accurate: which Whisper model to pick
Both choices are OpenAI’s open-weight Whisper, in quantised builds converted to run in a browser. The difference is size, and size is accuracy.
- Fast is Whisper tiny: 39 million parameters, about a 41 MB download. It is the right first try for a clear recording of one person at a steady pace — a narrated voice memo, a solo podcast, a talk given into a decent microphone.
- Accurate is Whisper base: 74 million parameters, about a 77 MB download. It is slower, and worth the wait when the audio works against you: two people talking over each other, a strong accent, a noisy room, specialist vocabulary, or a language other than English.
A sensible routine is to run Fast and read the first minute of output. If it is mostly right, let it finish and fix the stragglers in the editor. If whole phrases are wrong, switch to Accurate rather than editing your way through a bad draft.
Leave the language on automatic detection and Whisper identifies it from the first 30 seconds of the file that contain sound, then uses that language for the whole file. So choose it from the list yourself when the file starts with music, or with a greeting in a different language from the rest; if the model was unsure, the tool tells you which language it guessed. The Translate to English switch runs Whisper’s translate task, which writes English text from speech in another language in a single pass. English is the only target it offers, and from models this small the output is best read as the gist rather than a finished translation.
Does MP3 bitrate affect transcription accuracy?
Less than you might expect, down to a point, and the reason is that 16 kHz resample. A signal sampled at 16 kHz can only carry frequencies up to 8 kHz, so everything above that is discarded before Whisper hears anything. The shimmering treble an encoder preserved at 320 kbps goes in the bin either way.
An MP3 encoder saves space partly by filtering out high frequencies, and the lower the bitrate, the lower it sets the cut. LAME, the best-known open-source MP3 encoder, chooses its default low-pass filter from the bitrate and the number of channels. For a stereo constant-bitrate file it cuts at about 17 kHz at 128 kbps, 11 kHz at 64 kbps and 5.5 kHz at 32 kbps. A mono file spends every bit on one channel, so LAME sets its filter higher: about 17 kHz at 64 kbps and 8.3 kHz at 32 kbps. Spoken-word MP3s are often mono, so check which kind you have. Set those figures against the 8 kHz line:
- Stereo at 56 kbps and above, mono at 32 kbps and above: the filter sits above 8 kHz, so it removes nothing the model would have used. On bandwidth alone, a spoken-word podcast at 64 kbps is at no disadvantage against the same episode at 192 kbps.
- Stereo at 32 to 48 kbps, mono at 16 to 24 kbps: the cut falls between about 5.5 and 7.6 kHz, inside the band Whisper listens to, and the lower the bitrate, the more of the top of that band goes. The upper part of the band carries much of the hiss of an s and the softer f and th, so those consonants lose definition.
- Stereo at 24 kbps or less, mono at 8 kbps: the filter drops to between 2 and 4 kHz, about as narrow as a phone line or narrower, and compression artefacts become audible. Expect more mistakes.
Two things follow. Re-encoding a low-bitrate MP3 at a higher bitrate does not help: what the first encode removed is gone, and a second lossy pass can only add damage. And converting the MP3 to WAV before opening it here gains nothing, because decoding to uncompressed samples is the first thing this page does anyway. What does help is the next recording: sit closer to the microphone, find a quieter room, and check the level with the microphone test before a long session.
Podcasts, lectures, interviews and voice memos
The same model handles all four, but each kind of MP3 has its own trap.
- Podcast episodes tend to open with music and carry jingles, ad breaks and pauses between segments. That is where Whisper is weakest. A 2024 study of its transcripts, published as “Careless Whisper”, found it occasionally invents whole phrases nobody said, more often in audio with long stretches without speech. This tool skips any 30-second window that stays close to silent from start to finish, so a long stretch of true silence is passed over rather than transcribed; a pause inside a window that also holds speech still goes to the model, and so does music. Cut a long musical intro off with the audio trimmer first, and read anything transcribed over a music bed with suspicion.
- Lectures are long, usually one voice, often recorded from the back of a room with plenty of echo. Use Accurate, set the language rather than relying on auto-detect, and expect names and technical terms to need correcting.
- Interviews have two or more voices, and the transcript does not say who is speaking: Whisper transcribes words, not people. You add the names yourself; when you are not sure who said a line, click the time beside it and the recording plays from that point.
- Voice memos are often not MP3s at all — the iPhone Voice Memos app saves M4A, for example. That is fine: despite the name, this page accepts any audio or video file your browser can decode, so there is nothing to convert first.
Long MP3 files: chunks, memory and the 30-minute guideline
Whisper reads audio in 30-second windows. It was trained that way and cannot take a longer stretch at once, so a long MP3 is transcribed in 30-second chunks, one after another, and each piece of text keeps the timestamps of the audio it came from. A sentence that runs into the end of a chunk is probably cut mid-word, so the tool drops it and starts the next chunk where that sentence began. The transcript fills in as each chunk finishes, the progress bar shows how much of the file is done, and Cancel stops the job at any point while keeping what has been transcribed so far.
Two marks point to stretches that need a second listen. A time shown with ≈ means the model’s first pass over that stretch kept repeating itself, so the tool decoded it again without timestamps and estimated the times. The model struggled there, and the second pass can still drop or invent words, so play that stretch and compare. A line highlighted in amber still looked repetitive after the retries. Neither mark catches every mistake; they show where to start checking, not where to stop.
The file limit is 200 MB, which at 128 kbps is more than three hours of audio. The practical limit is lower, and memory sets it rather than the file. Before anything can be resampled, the browser decodes the entire MP3 into uncompressed 32-bit samples at its own playback rate, usually 44.1 or 48 kHz. A 30-minute stereo MP3 of around 29 MB becomes well over half a gigabyte of samples, and only after downmixing and resampling does it shrink to about 115 MB of 16 kHz mono. Add the model on top and a long file can exhaust a tab on a phone or an older laptop.
Hence the recommendation to keep files under about 30 minutes. The audio trimmer can cut a longer recording into parts, with one catch: it saves WAV, and a 30-minute stereo part comes to over 300 MB, which this page refuses. Press its 16 kHz mono preset before each export and the same part is about 55 MB. The trimmer also decodes the whole recording into memory, well over 2 GB for two hours of stereo, so split a recording that long in a desktop audio editor instead. Each part’s timestamps start from zero, so note where every part began before you join the transcripts.
Turning the transcript into subtitles or clean prose
Every line of the transcript is editable, and Copy and every download take the text as you have edited it, so the best time to fix names, figures and misheard words is while it is still on screen. Export before you transcribe again, though: a new run replaces your edits. Then pick the format for the job:
- .srt is the subtitle format most video editors and upload forms accept. Each timestamped segment becomes one cue.
- .vtt (WebVTT) is the subtitle format the HTML
<track>element uses for video on the web. If you need the other format later, the subtitle converter switches between the two. - .txt is the words on their own, with a paragraph break wherever two segments are two seconds or more apart — the starting point for notes, an article or show notes. Copy puts the same text on your clipboard.
Whisper’s segments follow the speech, not a subtitle style guide, so skim the cue lengths before you publish captions. Going the other way, if you already have subtitles for an episode and want prose, SRT to TXT strips the cue numbers and times out of an .srt file, and the transcript cleaner removes timestamps and speaker labels from any transcript and rejoins lines that were hard-wrapped. For a video rather than an audio file, MP4 to text runs the same model on the soundtrack.
What this MP3 to text converter will not do
- It will not label speakers. Everything comes out as one stream of segments with times.
- It will not guarantee a verbatim record. Whisper is good, the small models less so, and it can mishear or invent. Anything that will be quoted, published or relied on — a name, a number, a direct quotation — deserves a check against the audio.
- It will not keep anything. Close the tab and the transcript is gone, so copy or download it first.
Where VoiceSnap Pro fits
Converting an MP3 makes sense when the recording already exists. When the words have not been spoken yet, recording first and transcribing afterwards is a detour. VoiceSnap Pro is the desktop app we are building for that case. On a Mac or a Windows PC you hold a keyboard shortcut, say what you want to write and let go, and punctuated text lands in the field you were already typing in: a Slack reply, an email, a document. Each dictation is kept in a notes library you can search. The price will be a single payment of $39, with no subscription.
The app is still in development and cannot be installed yet. The waitlist takes an email address and sends one message, on the day it ships. In the meantime, the voice typing tool gives you live speech to text in the browser. It relies on your browser’s speech recognition service, which may process the audio on a server, so read its privacy note before you dictate anything private.
Questions people ask
No. The MP3 is read from your device by this browser tab, decoded there, and transcribed by a copy of Whisper running in a Web Worker on your own machine. The network traffic runs the other way: once you press Transcribe, jsDelivr supplies the transformers.js library with the WebAssembly runtime it needs, Hugging Face supplies the model files, and the browser caches both for next time. No request carries your audio out, which you can confirm in the Network tab of your browser's developer tools.
That depends on how long the MP3 is, which model you choose and what your computer can do. Before the first transcription the chosen model has to arrive, roughly 41 MB for Fast and 77 MB for Accurate, and later runs load it from the browser cache. With WebGPU the work goes to the graphics chip; without it, it runs on the CPU through WebAssembly, which on some machines cannot keep up with real time, so a ten-minute episode can take longer than ten minutes. A progress bar tracks how far through the file it has got and estimates the minutes remaining.
Only at the low end. Whisper resamples everything to 16 kHz, so it never uses anything above 8 kHz. The LAME encoder's default low-pass filter stays above 8 kHz for stereo files at 56 kbps and above, and for mono files at 32 kbps and above, because LAME sets a higher filter for mono. Those files lose nothing the model would have used. Below those bitrates the filter cuts into the band Whisper listens to (for a stereo file at 32 kbps it sits at about 5.5 kHz), and consonants such as s and f lose definition. Re-encoding the file at a higher bitrate cannot bring that detail back.
Yes. Every segment of the transcript carries a start and end time, and the .srt and .vtt downloads turn each segment into one subtitle cue. The .txt download gives you the words alone. Proofread names and numbers before you publish the subtitles.
The tool refuses any file over 200 MB, since the entire MP3 is decoded into memory before transcription starts. On long recordings memory usually runs short well before that cap, so files under about 30 minutes work best. Split a longer recording into parts and transcribe them one at a time.
Yes. Left on automatic detection, Whisper works out the language from the opening half-minute of the file that has sound in it and keeps that language to the end, so pick it from the list yourself if the recording starts with music or with someone speaking a different language. Turning on Translate to English runs Whisper's translate task instead, which writes English text directly from non-English speech. English is the only language it translates into.
Whisper occasionally invents text, and research on the model found this happens more often in audio with long stretches of no speech. This tool skips 30-second windows that stay close to silent throughout, but music is harder to guard against, so trim a long musical intro or outro before transcribing, and check any passage recorded over music against the audio.