Speech to text, online and free
A free online speech to text converter: press one button, talk, and your words appear as punctuated text you can copy or save. No signup, no upload, no waiting. It listens to your microphone live — it cannot transcribe a recording you already have.
This is live dictation from your microphone.
You press start, you talk, the words appear below. It is not a file transcriber — there is nothing to upload here, and it cannot turn an MP3, an M4A, a voice memo or a video you already recorded into text. If that is what you came for, this page cannot do it and you should look for an audio transcription service instead.
Command list
- “comma”,
- “period / full stop”.
- “question mark”?
- “exclamation mark”!
- “colon”:
- “semicolon”;
- “hyphen”-
- “dash”—
- “ellipsis”…
- “open quote / close quote”“ ”
- “open parenthesis / close parenthesis”( )
- “new line”line break
- “new paragraph”blank line
The trade-off is unavoidable: a sentence about “the reporting period” will end up with a full stop in the middle of it. Untick the box for that kind of dictation, or fix the one word afterwards.
- Words
- 0
- Characters
- 0
“Undo last phrase” takes back the last thing the microphone added, as many times as you press it. Editing the box by hand starts you fresh — the tool will not undo a phrase once you have changed the text around it.
Honest caveat: this page uses the browser’s own speech engine. In Chrome that engine streams your audio to Google’s servers for recognition — the page itself sends nothing anywhere, but the browser does. Keep genuinely sensitive dictation out of any browser-based recogniser.
This page stores nothing and sends nothing to us — but recognition is done by your browser's speech service, and in Chrome that means your audio goes to Google.
How to use this free online speech to text converter
- Open this page in Chrome, Edge, or another Chromium browser. Firefox has no speech recognition at all, and Safari’s version usually gives up after a few seconds.
- Pick your language from the dropdown. Set it before you start, not after.
- Press Start dictating and approve the microphone prompt. The browser only asks once per site, and only on an HTTPS page.
- Speak in whole sentences at a normal pace. Say “comma”, “full stop” or “new paragraph” where you want them and the tool writes the character rather than the word. Grey italic text under the box is the engine’s running guess; it firms up into the main box a second or two later.
- Got a sentence wrong? Undo last phrase takes back exactly what the microphone last added, as many times as you press it — no dragging a cursor through a paragraph to select it.
- Fix anything left over by typing directly in the box, then Copy text to paste it somewhere or Save .txt to keep a file.
Live dictation, not file transcription
This is the distinction that decides whether this page is any use to you, so it is worth being blunt about. Two completely different jobs get called “speech to text” online:
- Dictation — you talk, and text appears as you speak. That is what this page does.
- Transcription — you already have an MP3, an M4A, a voice memo, a Zoom recording or a video, and you want the words out of it. This page cannot do that. There is no upload button here because there is nothing on the other end to upload to: the recognition happens in your browser, and browsers only expose a live microphone stream, not a file decoder you can feed a recording into.
If you have a recording, you need a service that accepts file uploads, and you should expect to pay for anything longer than a few minutes. If you have something to say and an empty box to fill, you are in the right place. One workaround people try — playing the recording out loud into the microphone — reliably produces worse text than the original audio deserves, because the recogniser is now listening to a speaker, a room and your microphone rather than a voice.
Speaking your punctuation
The Web Speech API hands a page raw recognised words. Chrome adds some punctuation of its own in English, inconsistently and differently between versions, and there is no standard for saying “new paragraph” at all. So this page resolves the commands itself, in the page, after the words come back: “comma” becomes ,, “full stop” and “period” become ., “question mark” becomes ?, and “new paragraph” starts a fresh block. The next sentence gets a capital letter automatically. The full list is under the tool, behind “Command list”.
The trade-off is unavoidable and worth knowing before it surprises you: any tool that treats spoken words as commands will occasionally punctuate a sentence that was about punctuation. Dictate “the reporting period was strong” and you will get a full stop in the middle of it. Untick Spoken punctuation commands when that matters, and the words are kept verbatim.
Which browsers support online speech to text?
Browser dictation is built on the Web Speech API, and support for it is uneven in a way that catches people out. Here is the honest state of it:
- Google Chrome (desktop and Android) — full support. This is the reference implementation and the one everything else is measured against.
- Microsoft Edge, Brave, Arc, Opera, Vivaldi — supported, because they are all Chromium underneath.
- Safari — partial. The API exists behind the
webkitSpeechRecognitionprefix, but sessions end quickly, continuous mode is unreliable, and results often stop arriving with no error. - Firefox — not implemented. The tool will tell you so rather than silently doing nothing.
Two other conditions are easy to miss. The page must be served over HTTPS — browsers refuse microphone access on a plain http:// origin, with localhost as the only exception. And the tab has to stay in the foreground: switch away and most browsers suspend recognition within a few seconds.
Where your audio actually goes
This is the part most free dictation pages leave out. This page has no server component and never posts your transcript anywhere — but the recognition itself is done by the browser, and Chrome performs it in the cloud. Your microphone audio is streamed to Google’s speech service, converted to text there, and sent back to the tab. That is why online dictation stops working the moment your connection drops, and why nobody can honestly describe a browser-based recogniser as offline or fully private.
For a shopping list or a blog paragraph, fine. For a client’s medical history, a legal note, an unreleased product spec or anything under an NDA, decide deliberately rather than by default. Any browser-based recogniser has this property, no matter how the page is worded.
Getting a usable first draft
- Keep the microphone about a hand’s width from your mouth and slightly off to one side, so plosives (“p”, “b”) do not thump the capsule.
- Speak in complete sentences. Recognition uses surrounding words as context, so half-sentences and long pauses in the middle of a clause produce worse text.
- Do not stop to correct a single word. Keep going and clean up at the end — you will finish faster and the engine will make fewer mistakes.
- Names, product names, jargon and acronyms are where every general-purpose engine struggles. Expect to fix them by hand here; a desktop tool with a custom vocabulary is the real fix.
- A wired headset or any dedicated microphone beats a laptop’s built-in array, especially in a room with hard surfaces.
- Spoken drafts run long — most people speak around 150 words a minute and type at a fraction of that. Run the result through the text cleaner before you send it, and if you want to know your own speaking pace, the voice typing test measures it against a set passage.
What a browser tab cannot do
This page is a demonstration of speech to text, not a way of working. The text only appears in this one text box, in this one tab. Every dictation ends with the same four steps: select, copy, switch app, paste. That is fine once. It is not fine forty times a day, and it is why online dictation never becomes a habit for most people.
Everything else it lacks follows from the same limitation. There is no custom vocabulary for the names you say constantly. Filler words are transcribed faithfully, so every “um” and “you know” survives into the draft. There is no record of what you dictated last week. And the tab has to stay focused, so you cannot dictate into the thing you are actually looking at.
How VoiceSnap Pro handles the same job
VoiceSnap Pro is a dictation app for macOS and Windows built around removing that copy-paste step. You hold one keyboard shortcut anywhere on your machine, speak, and clean punctuated text appears in whatever field your cursor is already in — a Gmail reply, a Slack message, a Jira ticket, a commit message, a browser form. Filler words are stripped automatically, paragraph breaks are inserted where you paused, and you can teach it the names and acronyms you use. It handles 50+ languages, including switching language mid-sentence, and every dictation is also saved to a searchable notes library so last Tuesday’s recording is still findable. Nothing you dictate is used to train AI models, and it is a one-time $39 purchase rather than another monthly line item.
VoiceSnap Pro has not shipped yet. Join the waitlist and you’ll get one email on release day. In the meantime this page is genuinely useful, and so are the other free tools.
Questions people ask
Is this speech to text tool free, and do I need an account?
It is free, and there is no account, no signup, no watermark and no usage limit. The tool is a single page of JavaScript talking to the speech engine already inside your browser, so there is nothing for us to meter.
Can I upload an audio file and get a transcript?
No. This is live dictation from your microphone only. Browsers expose a live audio stream to the speech API and nothing else, so there is no way for a page like this one to accept an MP3, a WAV, an M4A or a video and return text. A recording needs a transcription service that does the work on a server.
How do I add punctuation?
Say it. “Comma”, “full stop”, “question mark”, “new line” and “new paragraph” are all recognised while the spoken-punctuation box is ticked, and the next sentence is capitalised for you. You can also just type the punctuation in afterwards — the box is an ordinary text field.
Why did it stop listening on its own?
Browsers end a recognition session after a stretch of silence. This page restarts it automatically, which is why long dictations keep working, but a very long silence or a backgrounded tab can still end it. Press Start again and carry on — the text you already have is untouched.
The microphone prompt never appeared.
Either permission was denied earlier, or the page is not on HTTPS. Click the padlock or microphone icon in the address bar, set the microphone to Allow, and reload. On macOS also check System Settings → Privacy & Security → Microphone and confirm your browser is listed and enabled.
Can I use voice to text online on a phone?
Chrome on Android works. iOS is a different story: every iOS browser is Safari underneath, so support is the partial, unreliable kind described above. On an iPhone the built-in keyboard dictation key is a better bet.
Does it work offline?
No. Recognition happens on a server, so a dropped connection stops the transcript. Desktop dictation apps that run the model locally are the answer if offline matters to you.
How accurate should I expect it to be?
On clear speech, a decent microphone and a quiet room, browser recognition is good enough that editing is faster than typing from scratch. Accuracy falls off sharply with background noise, several people talking, strong accents in a language variant you did not select, and any word the engine has never seen — which is most product names.
Can I save the transcript?
Yes — Save .txt writes the text to a plain text file with today’s date in the name, built in your browser from what is in the box. Nothing is uploaded to produce it. For subtitles rather than prose, the SRT to VTT converter handles caption files, and the transcript cleaner tidies exports from other tools.