What it fixes
Paste text and you get a clean copy right away, with a list of what was fixed. These are the characters that cause most paste problems:
| Character | Where it comes from | What it becomes |
|---|---|---|
| Zero-width space, joiner, word joiner | Web pages, chat apps, AI tools | Removed |
| Non-breaking and other special spaces | Word, web pages, PDFs | A normal space |
| Soft hyphen | Word and web pages that hyphenate | Removed |
| Byte order mark | Files saved by some Windows programs | Removed |
| Direction marks | Text mixed with right-to-left languages | Removed |
| Curly quotes and apostrophes (“ ” ‘ ’) | Word, Google Docs, Pages, phones | Straight “ and ’ |
| En and em dashes (– —), minus sign | Word, publishing tools | A hyphen - |
| Ellipsis (…) | Word, phones | Three dots … |
| Joined letters (fi, fl) | Text copied from PDFs | Separate letters fi, fl |
| Line separator, page break, Windows line ends | Word, PDFs, Windows files | A standard line break |
Remove extra spaces also takes away spaces at the end of lines and double spaces between words, but keeps the indentation at the start of a line, so code stays lined up.
Why it matters
Hidden characters cause problems that are hard to spot because the text looks fine:
- Excel and Google Sheets: a number with a non-breaking space after it is treated as text, so it won’t add up, and a lookup fails when one cell has a hidden character the other doesn’t.
- Code and config files: a curly quote or a zero-width space breaks a script or a JSON file, often with an error that points nowhere useful.
- Website editors and forms: soft hyphens and odd spaces show up as strange gaps or break words in the wrong places.
How it works
Each character is checked against its Unicode category. Format characters (like zero-width spaces and direction marks) and control characters are removed, except tabs and line breaks. One exception is kept on purpose: the zero-width joiner between two emoji, which is what joins them into one emoji such as a family. Special spaces become normal spaces, and typographic quotes, dashes, and dots become their plain versions. Finally the text is normalized (Unicode NFC), so an accent stored separately from its letter is joined back to it, which keeps searches and comparisons working.
Limits
Turning dashes into hyphens and quotes into straight ones is right for code, spreadsheets, and data, but it’s a style change in finished writing. Untick those boxes to keep them. Plain ASCII only is for systems that reject anything else; it removes emoji and characters from non-Latin alphabets.