100% local — your data never leaves your browser

HTML Entities to Text — Read a Scraped Snippet

Decode HTML entities back to characters. Numeric references are all supported, and an unknown name stays visible so you can see what it was.

Instant Private Zero cookies

HTML Entities input

Text output

What this tool does

Every &…; reference becomes the character it names.

  • Numeric, decimal — é gives é.
  • Numeric, hexadecimal — é and é both give é.
  • Named — &,  , —, é and around sixty other common names.

An astral character arrives as one reference and leaves as one character: 😀 and 😀 both give the emoji.

An unknown name stays visible

♥ comes out as ♥.

HTML5 defines more than two thousand names. This tool carries the ones you actually meet, and leaves the rest untouched — because the alternatives are worse. Deleting an unrecognised reference loses data silently. Guessing at it invents a character that was never there. Leaving it visible tells you exactly what was not handled, and its numeric form always works.

Decoding runs once

&amp;lt; gives &lt;, not <.

That is the right answer, not a limitation. &amp;lt; is what an encoder produces when the source text was the literal four characters &lt;, so a single pass returns exactly that source text. Running the decoder again would turn escaped markup into live markup — the exact failure escaping exists to prevent.

The semicolon is required

&amp x is left alone. Browsers do decode some references without their terminator, for compatibility with documents written in the nineties, but which ones depends on the parser and the context. Requiring the semicolon means the result does not depend on whose parser you ask.

References that name no character

Three cases are returned as written rather than decoded:

  • &#0; — a NUL has no legal representation in HTML, and producing one truncates a C string, breaks a SQL insert and corrupts whatever reads the result;
  • &#xD800; through &#xDFFF; — surrogate halves are not characters, only an encoding artefact;
  • anything above &#1114111; — beyond the top of Unicode.

You see the reference you typed, which is the honest signal that it did not name anything.

Private by design

Everything runs locally in your browser with JavaScript. Your data is never uploaded, which makes the tool safe for sensitive content, and it keeps working offline.

Frequently asked questions

Why is &hearts; still in my output?
Because it is not in the list of names this tool decodes, and leaving it visible is better than deleting it or guessing. HTML5 defines over two thousand names; the common ones are covered, and anything else stays exactly as you typed it so you can see what was not handled. Its numeric form, `&#9829;`, decodes.
Why did &amp;lt; only become &lt; and not <?
Because decoding runs once. `&amp;lt;` is what an encoder produces from the literal text `&lt;`, so one pass gives that text back — which is the correct answer. Decoding twice would turn escaped markup into live markup, and that is precisely the bug that escaping exists to prevent.
Why was &#0; left alone?
Because there is no character to produce. A NUL has no legal representation in an HTML document, and emitting one would truncate a C string, break a SQL insert and corrupt whatever read the result. The same applies to a surrogate half and to anything above U+10FFFF: the reference is returned as written, so you can see it.

Related converters