Text Cleaner
Text copied from a PDF, an email or a word processor arrives with trailing spaces, mixed line endings, non-breaking spaces and zero-width characters that break searches and diffs. Each repair is a separate toggle, and Furtu lists what it changed and why.
CRLF, CR and LF all become LF.
Zero-width characters, soft hyphens, non-breaking spaces, directional marks and byte order marks. These are why a search for a visible word sometimes finds nothing.
Removes the markup and the contents of any script or style element.
Spaces and tabs at the end of each line.
Three or more newlines in a row become one blank line.
Within a line only. Turn this on carefully if the spacing is deliberate, as it is in ASCII art and hand-aligned tables.
How text cleaner works
- Paste the textStraight from a document, a PDF or a terminal.
- Choose the repairsSix toggles, each independent, and each reported only when it changed something.
- Read the changesThe notes name the code points involved, so a stubborn character can be tracked down.
What you get
Fix the invisible damage in pasted text, one step at a time. Everything happens inside this page: the file is read by your browser, transformed in memory and handed straight back to you as a download. There is no upload queue, no waiting for a server, and nothing left behind when you close the tab.
Supported formats
This tool works on text you paste or type, so there is no file format to worry about. Nothing you type is sent anywhere.
Limitations, stated up front
- HTML entity references are left as written rather than decoded.
- Nothing is repaired beyond whitespace, invisible characters and tags: encoding problems inside the bytes need the source fixed first.
Frequently asked questions
What are the invisible characters?
They are real code points that occupy no width: a soft hyphen, zero-width characters, directional marks, bidirectional overrides that can reverse how a line displays, and a byte order mark. A non-breaking space is also invisible, as a space. Each is named with its code point in the notes, so you can see which one was in your text.
Does stripping HTML decode the entities?
No. The tags go and the entities stay, as written. Decoding them means choosing a character set and accepting that a bare ampersand is not always an entity, and getting it wrong corrupts text that was fine. If you need the characters, the document tools make that decision explicitly.
Is it safe to collapse repeated spaces?
For prose, yes. For anything where spacing is structure — ASCII art, a hand-aligned table, a signature block, poetry — no, and that is why the toggle is off by default. The other five steps do not change what the text means.
Why is the trailing newline removed?
Because a stray newline at the end of a copied fragment is almost never intentional, and it breaks a byte-for-byte comparison with an expected value. It is a final tidy-up rather than one of the toggles, and it is named in the notes when it happens.
Is the text I am cleaning uploaded?
No. Cleaning is local, which is the point — the text people clean is usually text lifted from a paid report, a competitor page or a private document.