Back to all tools
Free · runs in your browser

AI text cleaner

Paste anything a model wrote. This AI text cleaner shows you every invisible codepoint, em dash and curly quote the text is carrying — each one in place, with its Unicode number — and only then strips whatever you tell it to.

What is in there

Paste something above, or load the example, and every flagged character is counted here before anything is removed.

Audit view

0 flagged in context

Your text with every flagged codepoint swapped for a visible chip carrying its Unicode number. Nothing here has been removed yet.

The highlighted version of your text appears here.

This is paste hygiene, not detector evasion. A browser can find anomalous invisible codepoints and strip them; it cannot detect a statistical watermark, because that lives in the choice of words themselves and no amount of character surgery touches it.

Some of these characters are also meant to be there. A non-breaking space in “10 kg”, a soft hyphen in justified copy, a zero-width joiner inside an emoji — all legitimate. That is exactly why the audit shows you each one in context instead of stripping blind.

What to remove
Replace each em dash with

Parentheses only apply to a matched pair of dashes around an aside. A lone dash falls back to a comma.

Characters in
0
Characters out
0
Changes made
0

Everything here runs on your device. Nothing you paste is uploaded.

What an AI text cleaner should show you before it cleans

Every other AI text cleaner works the same way: you paste, it strips, it hands back a block of text. You have no idea what was in there. You cannot tell whether it removed four zero-width spaces or forty, whether the non-breaking space it deleted was junk from a model or the one you deliberately put in “10 kg”, or whether it touched anything at all — because the characters in question are invisible. That is the whole problem, and a silent stripper reproduces it.

So this one is built backwards. The first thing on screen after you paste is an audit: your text, unchanged, with every flagged codepoint replaced by a visible chip carrying its Unicode number, so a zero-width space becomes a readable U+200B ZWSP sitting exactly where it was. Above it is a summary line — four zero-width spaces, eleven em dashes, two non-breaking spaces. Only below that comes the cleaned version.

That ordering is not aesthetic. Some of these characters are legitimate: a non-breaking space between a number and its unit is correct typography, a soft hyphen in justified copy is a break point a designer put there, a zero-width joiner holds a family emoji together. Stripping blind quietly breaks all three. Seeing each one in context is what lets you decide.

Invisible characters remover: what each class actually breaks

“Invisible characters” sounds cosmetic until one costs you an afternoon.

Zero width space remover

A zero-width space (U+200B) has no width and no glyph, but it is still a character, so string comparison sees it. Paste a class name containing one into a stylesheet and the selector silently fails to match — valid CSS that applies to nothing, and you will read that line twenty times without seeing why. The same character means zero results in a search box, an undefined lookup on a JSON key, and a login failure you cannot reproduce in a password field.

They arrive from more places than chatbots: web pages insert them to control line breaking in long URLs, rich-text editors add them at formatting boundaries. This tool removes them with their relatives: the zero-width joiner and non-joiner, the word joiner (U+2060), the byte-order mark (U+FEFF) that heads anything exported from a Windows tool, the bidirectional controls that make a rendered line differ from the line a compiler sees, and the U+E0000 tag block — 128 codepoints that render as nothing and can hide an entire ASCII message inside one visible sentence.

Non-breaking spaces and the import that fails on row 4,000

A non-breaking space (U+00A0) looks exactly like a space and is not one. Split a line on " " and the field containing it does not split — that is a CSV import failing on one row in four thousand with an error about column counts. It is also why a value reading 1 234 will not parse as a number and why two identical-looking strings compare as unequal. Narrow non-breaking space (U+202F) does the same and is harder to spot. Both need a judgement, not a rule: the one in “10 kg” stays, the one sprinkled through a bullet list does not.

Soft hyphens that survive into a URL slug

A soft hyphen (U+00AD) renders only when a line actually breaks there; everywhere else it is invisible. Copy justified text out of a PDF and you drag soft hyphens with it into a slug generator, a search index or a filename, none of which can do anything sensible with an unprintable character. The word rollback with a soft hyphen after the second L is indistinguishable from the one without it until something tries to match it.

Remove em dashes from AI text

The em dash is the most recognisable habit of generated prose — usually unspaced, often twice in a sentence. If you want to remove em dashes from AI text, the question is not how to delete them but what stands in their place: the dash is doing grammatical work, and deleting it leaves a sentence with no joint. Three replacements, not interchangeable:

  • A spaced hyphen — word - word. The literal swap: same pause, works everywhere. The default, because it never changes your meaning.
  • A comma — word, word. Usually the best choice for prose. Most em dashes in generated text join a clause a comma joins just as well, and the sentence stops announcing itself.
  • Parentheses — word (aside) word. Applies only to a matched pair wrapping an aside. A lone dash falls back to a comma rather than producing an unbalanced bracket.

Spacing is repaired either way, so negotiable—the vendor does not become negotiablethe vendor, and a dash opening a line is treated as a bullet rather than gaining a leading space. Em dashes have their own switch, separate from the pass that straightens curly quotes, en dashes and the ellipsis character — visible characters, but multi-byte in UTF-8, so they turn into mojibake wherever something guesses the wrong encoding, and a curly quote in code is a syntax error your eyes read as valid.

ChatGPT text cleaner: the paste path that breaks

Used as a ChatGPT text cleaner, the sequence is almost always the same. The answer arrives formatted for a chat window — markdown headings, bold asterisks, bullet glyphs, curly quotes, em dashes, plus whatever zero-width characters the renderer inserted. You paste it somewhere with no markdown renderer, and the reader sees literal **bold** asterisks and a line beginning ###.

The markdown pass handles the visible half: heading hashes come off and the text stays, emphasis markers go and the words between them stay, fenced code lines go entirely, and bullet glyphs become the plain hyphen every plain-text field understands. The audit flags each marker first, so you can see whether a stray ** is formatting or something you meant to write. None of it is tuned to one vendor — Claude, Gemini and Copilot output has the same shape.

How to clean AI generated text without flattening it

A workable order of operations if you want to clean AI generated text and still have it sound like something you would sign your name to.

  • Read the summary first. Eleven em dashes in four paragraphs is a style problem; forty zero-width spaces is a copy-path problem, worth tracing to whichever editor inserts them.
  • Decide the em dash question deliberately. Replacing every dash with a comma is the fastest way to make generated prose stop announcing itself, and it is a real edit, so read the result. If a flagged character is one you put there, switch that category off and clean the rest.
  • Then edit for content. Character hygiene does not fix the real tell of generated writing, which is rhythm: every sentence the same length, every list three items. Our sentence counter shows the length distribution, and a flat one is what readers register long before a dash.

What this tool deliberately does not do

It does not defeat AI detectors, and it is not built to. A browser can find anomalous invisible codepoints and remove them, because those are discrete characters in a string. It cannot detect a statistical watermark: that lives in the choice of words themselves, across thousands of small decisions, and no amount of character surgery touches it. Detectors score how predictable the phrasing is, and this tool changes no words at all.

So treat it as paste hygiene: making text render correctly in a CMS, a code editor, a spreadsheet or a form. It also does not guess at single asterisks or underscores for italics, because a * b and snake_case_name are ordinary text far more often than they are markup.

How this differs from the general text cleaner

The two pages share their transform code, so they can never disagree about what a rule does. What differs is the surface: this page leads with the audit and has the em-dash modes, while the text cleaner has eleven passes for things AI output does not produce — hard-wrapped PDF lines, words split across a break by a hyphen, punctuation spacing, ALL-CAPS sentence casing. For recordings and captions, the transcript cleaner strips timestamps, speaker labels and filler words. To insert one of these characters rather than remove it, the alt codes reference has them with click-to-copy.

Where VoiceSnap Pro fits

VoiceSnap Pro is a voice-to-text dictation app for macOS and Windows. Hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is already in — the CMS box, the email, the pull request comment. It adds punctuation and paragraph breaks as it transcribes, strips filler words, and saves every dictation to a searchable notes library. It is a one-time $39 purchase, not a subscription.

The honest connection to this page: dictated text has none of the problems above. No markdown to strip, no zero-width characters from a chat renderer, no em dash habit, because the words went straight from your mouth into the field. Cleaning generated output is work you only do because the text detoured through a chat window first. VoiceSnap Pro has not shipped yet — join the waitlist for one email on release day, or try the browser voice typing tool to see how dictated text reads before you edit it.

Questions people ask

No. Everything here is JavaScript running in your tab: your text is never sent to a server, never written to disk and never seen by a model. There is no API call behind the audit or the cleaning, only string rules, and closing the tab discards everything. That is why you can safely paste something confidential into it.

No, and you should not use it for that. Removing invisible characters and em dashes changes nothing a detector measures. The honest use is making text paste correctly.

Because it cannot tell the difference and does not pretend to — a non-breaking space is the same codepoint whether a model emitted it or you typed it. That is why the audit shows each one in context. Anything with a common legitimate use also gets a note under the summary saying what that use is.

Not unless you ask. The tidy option collapses runs of spaces, strips trailing whitespace and reduces three or more blank lines to one — repair work for the gaps that removing a character leaves behind. It never joins lines or rewraps paragraphs.

Multi-codepoint emoji are read as whole characters, so a family emoji or a flag is never split down the middle. The joiners inside them are still flagged, because the same character is used to hide things — if the audit shows they are all inside emoji, leave the invisible pass off.

The counts and the cleaned output have no cap. The highlight view stops drawing chips after twenty thousand characters, because a chip per character across a manuscript would make the page crawl. A note appears when that happens; the numbers still cover everything.

Further readingChatGPT DictationA longer read on the VoiceSnap Pro blog.