Audio to text converter
Turn a recording into text with OpenAI's open Whisper model, running on your own device. Pick an audio or video file, press Transcribe, then edit the timestamped transcript and save it as TXT, SRT or VTT. The file is never uploaded.
Drag an audio or video file here, or choose one below.
MP3, M4A, WAV and FLAC open in every current browser; Ogg, Opus and WebM need Chrome, Edge or Firefox. Video files work too: only the soundtrack is used.
Choose a file first.
Your file is transcribed by the Whisper model running on your own device. The audio is never uploaded; the only downloads are the model files from Hugging Face and the transformers.js library that runs them, from jsDelivr, fetched when you first press Transcribe and cached by your browser.
How on-device transcription works
Many free audio-to-text sites upload your file to a server, run it through a speech model there and send the text back. This page does the same job without the upload. The model comes to your file instead of your file going to the model, and every step happens inside the browser tab you are looking at:
- Decode. Your browser reads the file from disk and decodes it with its own audio decoder, the same one it uses to play media.
- Prepare. The sound is mixed down to one channel and resampled to 16 kHz, the sample rate Whisper was trained on. That discards everything above 8 kHz, which the model never uses anyway.
- Load the model. A background worker loads transformers.js, Hugging Face’s library for running models in the browser, and the Whisper model you chose. The first time, both are downloaded; after that your browser keeps them in its cache.
- Transcribe. Whisper reads audio in 30-second windows, so the recording is fed through one window at a time. Each finished window adds its lines to the transcript, with a start and end time on every line.
Because the work happens in a background worker, the page stays responsive while it runs: you can scroll, read the first lines as they arrive, or press Cancel, which stops the model at once and keeps what has been transcribed so far. When a line runs into the last moments of a window it has probably been cut off mid-word, so it is dropped and the next window starts where that line began. That costs a little extra work and saves a garbled word at every 30-second boundary.
What Whisper is
Whisper is a speech recognition model that OpenAI released in September 2022, with the code and the trained weights published under the MIT licence. It was trained on 680,000 hours of audio collected from the web, covering many languages and several tasks at once: transcribing speech, translating speech into English, identifying the language being spoken, and predicting when each phrase starts and ends. One model does all of that, which is why a single download here covers transcription, translation and timestamps.
OpenAI published Whisper in several sizes. This page offers the two smallest, because they are the ones a browser can download and run comfortably. Fast is Whisper tiny, with 39 million parameters, about 41 MB as downloaded here. Accurate is Whisper base, with 74 million parameters, about 77 MB. Both are 8-bit quantised versions converted for the browser, which is what keeps the downloads that small. The larger Whisper models are more accurate still, but they run to hundreds of megabytes or more, which is a lot to ask of a free web page.
How accurate is it?
Honestly, it depends on the recording more than anything this page can control. One person speaking clearly, close to the microphone, in a quiet room, usually comes out with only the occasional slip. Accuracy drops with a distant or built-in laptop microphone, background music, echoey rooms, several people talking over each other, and strong accents. OpenAI’s own model card is candid that performance is uneven across languages, weaker in languages with less training data, and varies between accents and dialects of the same language.
There are two failure modes worth knowing about. The first is invention: Whisper sometimes writes a plausible phrase nobody said, most often over music, applause or a long silence, because it is partly predicting the next likely words. The second is looping, where the model repeats one word or sentence over and over. This page watches for loops: when a window’s text looks like one, that window is transcribed again without timestamps, a different path through the model that can break the loop, and its lines are given estimated times, marked with ≈. If the second attempt repeats itself too, or stops short, the window is cut at a pause near its middle, the first part is transcribed on its own and the next window starts at the cut. Lines that still repeat after all that are highlighted so you know to check them against the audio, and the Accurate model is the thing to try next.
A few things reliably help. Choose the language yourself if you know it. Trim long musical intros and outros with the audio trimmer before transcribing. Use Accurate for accented speech, poor recordings or anything you will publish. And proofread names, numbers and technical terms, which no speech model gets right every time. If you record yourself, run a microphone test first: a clean input fixes more errors than any model setting.
How long it takes
Speed depends on your hardware, the model and the length of the recording. Where your browser offers WebGPU, the model tries the graphics chip first and falls back to the CPU if that fails. On the CPU it runs through WebAssembly on a single core, and on some machines that is slower than real time: a ten-minute recording can take more than ten minutes. Accurate does more work per second of audio than Fast, so it is slower on the same machine. The Processor setting lets you force the CPU if the graphics-chip path misbehaves on your device.
The first run has a one-off cost on top: the model download (about 41 or 77 MB) plus about 6 MB for the library and its runtime. The progress panel shows the download, then how much of the recording has been transcribed and an estimate of the time left. When it finishes, the summary tells you how long the transcription took and whether it ran on the graphics chip or the CPU.
File types and limits
The page accepts anything your browser can decode. MP3, WAV, M4A (AAC) and FLAC open in every current browser. Ogg, Opus and WebM need Chrome, Edge or Firefox. Video files work too, because only the soundtrack is used, so a screen recording or a phone video can go straight in. For a page aimed at one format, see MP3 to text and MP4 to text; they run the same transcriber.
Two limits apply. Files over 200 MB are refused, because the browser decodes the whole file into memory before anything else happens. And recordings under about 30 minutes work best: a long file needs a lot of memory at every stage, and on a slow CPU it takes a long time. Before decoding, the page reads the recording’s length from the file’s header where the browser can, and asks before going on with anything over 30 minutes. Video files reach the 200 MB cap far sooner than audio of the same length. For long recordings, cut them into parts with the audio trimmer and transcribe the parts in turn.
TXT, SRT and VTT exports
Every line of the transcript is an editable box with its start time beside it. Click a time to hear that part of the recording, fix what the model misheard, and the copy button and downloads pick up your edits. There are three downloads:
- .txt is the words alone, with a paragraph break wherever the speaker paused for two seconds or more. Use it for notes, articles and meeting records.
- .srt (SubRip) turns each line into a numbered subtitle cue with a start and end time. YouTube and most video editors accept it.
- .vtt (WebVTT) holds the same cues in the format HTML5 video players read, for captions on your own site.
If you need to change format later, the SRT to VTT and VTT to SRT converters do it without re-running the model, and SRT to TXT strips a subtitle file back to plain text. To tidy a transcript for reading, the transcript cleaner strips timestamps and rejoins hard-wrapped lines.
Languages and translation
Whisper was trained on 99 languages, and the language list here offers all of them. By default the page detects the language from the first 30-second stretch that contains sound, then transcribes the whole file in that language. Detection is usually right, but it can be misled by a musical intro, a short greeting in another language, or a very quiet start, so choose the language yourself when you know it. If the model was unsure, the page says so.
Tick Translate to English and Whisper switches to its translate task: it listens in the source language and writes English. It is a single step inside the model, not a transcript passed through a separate translator, and it only goes into English. Translation quality follows transcription quality, so the same advice about clear audio applies twice over.
Privacy: what is downloaded and what is not
Your recording is never uploaded. The file is read from your disk by the browser, decoded in the tab and handed to a worker in the same tab. Nothing about it is sent to us or to anyone else, and there is no account to sign into. Closing the tab discards the audio and the transcript, so download or copy what you want to keep.
Three things do come over the network, and only once you press Transcribe: the transformers.js library and its WebAssembly runtime from jsDelivr, a public code CDN, and the Whisper model files from Hugging Face. They are software and model weights, the same for every visitor, and your browser keeps them in its cache so the next run does not fetch them again. If you clear this site’s data, the model is downloaded again the next time.
What it does not do
- No speaker labels. Whisper records what was said, not who said it. Lines break at pauses and sentence ends, not at changes of speaker.
- No word-by-word timing. Times are per line, which suits subtitles and navigation but not karaoke-style highlighting.
- One file at a time. There is no batch queue; transcribe a file, save the result, then load the next.
- No guarantee of accuracy. Treat the output as a fast first draft and proofread anything that matters.
Transcribing a recording is not the same as live dictation
This page works on audio that already exists: an interview, a lecture, a voice memo, a video. You get the whole transcript with timestamps once the model has read the file. Live dictation is the other direction: you speak and the words appear as you go. The speech to text tool on this site does that in the browser, but it works differently: it uses your browser’s built-in speech service rather than a model on your device, and in Chrome that service sends your audio to Google. Use this page for files and that one for talking, and read each page’s privacy note.
Where VoiceSnap Pro fits
Recording first and transcribing afterwards is the long way round when all you wanted was text in an email or a document. VoiceSnap Pro, the app we are building, is the short way. It is a voice-to-text dictation app for macOS and Windows: hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is already in — an email, a Slack message, a code editor. Every dictation is also saved to a searchable notes library. It is a one-time $39 purchase, not a subscription.
It has not shipped yet, so there is nothing to download today. Join the waitlist for an email on release day. Until then, this page and the speech to text tool are free to use as often as you like, and if you are weighing up running Whisper yourself, our VoiceSnap Pro vs Whisper comparison sets out the differences.
Questions people ask
No. Your browser decodes the file inside the tab, and the Whisper model transcribes it in a background worker in the same tab. The only things downloaded are the transformers.js library and its WebAssembly runtime, from jsDelivr, and the model files, from Hugging Face, and only after you press Transcribe. Your browser caches them, so later runs skip the download. The audio itself never leaves your device.
It depends far more on the recording than on the tool. One clear voice close to the microphone in a quiet room comes out well; distant microphones, background music, people talking over each other and strong accents all cost accuracy, and Whisper is noticeably stronger in widely spoken languages than in ones with little training data. It can also invent a phrase over music or long silences. The Accurate model makes fewer mistakes than Fast. Either way, check names, numbers and technical terms before you rely on the text.
It depends on your hardware, the model and the length of the file. The first run also downloads the model once: about 41 MB for Fast or about 77 MB for Accurate. Where the browser offers WebGPU the model tries the graphics chip first; otherwise it runs on the CPU, which on some machines is slower than real time. The progress bar shows how much has been transcribed and estimates the time left, and Cancel keeps whatever has been transcribed so far.
Anything your browser can decode. MP3, WAV, M4A and FLAC open in every current browser; Ogg, Opus and WebM need Chrome, Edge or Firefox. For a video file only the soundtrack is used. Files over 200 MB are refused, because the whole file is decoded into memory first, and recordings under about 30 minutes work best. Split anything longer into parts.
No. Whisper writes down what was said, not who said it, so the transcript has timestamps but no speaker labels. Lines break where the model finds a natural pause or the end of a sentence, which often but not always coincides with a change of speaker. Add the names yourself while you proofread.
Whisper was trained on 99 languages. Leave the language on Detect automatically and it is identified from the first 30 seconds that contain sound, or choose it from the list, which is safer when a recording opens with music or with a different language. Tick Translate to English and Whisper uses its translate task instead, writing English text from speech in another language. It translates into English only.
Yes. Every line of the transcript is an editable text box with its start time beside it, and clicking the time plays the recording from that point. Copy and the downloads use your edited text: .txt is the words alone, while .srt and .vtt turn each line into a subtitle cue with its start and end time.