Dev Toolbox · Windows 10 and 11

Escape and unescape HTML entities

Turn text into something you can drop into HTML without breaking the page, or turn a mess of &, " and ' back into plain words. Paste on one side, read the result on the other, all on your own PC.

Updated

How to escape or unescape HTML

  1. 1

    Open HTML entities

    Open the Dev Toolbox and choose HTML entities in the Data group.

  2. 2

    Escape or Unescape

    Set Direction to Escape to make text safe for HTML, or to Unescape to turn entities back into characters.

  3. 3

    Choose how far to go

    When escaping, switch on Also non-ASCII characters to write every accented letter, symbol and emoji as a numeric entity such as é.

  4. 4

    Copy the result

    Paste into Input; the Result follows as you type. Press Copy to take it.

The five characters that need escaping

HTML gives a few characters special jobs. Put them in a page as plain text and the browser reads them as markup instead. Escaping replaces each with an entity, a short code that the browser shows as the character itself:

CharacterBecomesBecause it
&&Starts every entity
<&lt;Opens a tag
>&gt;Closes a tag
"&quot;Ends an attribute value written in double quotes
'&#39;Ends an attribute value written in single quotes

The ampersand goes first for a reason: escape it after the others and you would escape your own new entities a second time. Octoolo keeps the order right. Writing the apostrophe as &#39; rather than &apos; is deliberate too: &apos; was not part of HTML 4, while a numeric entity works everywhere.

Escaping is how you stop broken pages and injected scripts

Any time text from somewhere else ends up inside HTML (a comment, a product name, a search term, an error message), it has to be escaped. Otherwise a product called Salt & Pepper <Large> loses its last word, because the browser takes <Large> for a tag. And a "name" that contains a <script> tag runs as code in every visitor's browser. That second case is cross-site scripting (XSS), one of the most common holes in web applications.

This tool is for the times you escape by hand: putting a code sample into a blog post or documentation, adding text to an HTML email template, writing a value into a static page, or checking what your template engine should produce. Two limits are worth knowing:

  • Escaping makes text safe between tags and inside quoted attributes. Text going into a script block, a style block or a link's address needs a different kind of encoding; for links, use URL encoding.
  • In your own code, let the framework do it. React, Vue, Angular and most server-side templates escape output by default, and escaping text that is already escaped is how &amp;amp; ends up on a live page.

Numeric entities for accents, symbols and emoji

With Also non-ASCII characters on, every character outside basic English (code points above 127) is written as its number: é becomes &#233;, € becomes &#8364; and 😀 becomes &#128512;. The page looks the same, but the file itself is pure ASCII.

A page saved and served as UTF-8 does not need this, and most of the web is UTF-8. It helps when text passes through something that mangles it: an old CMS or database set to Latin-1, an email system that turns é into é, a config value that only takes ASCII. If you have seen é or ’ where an accent or a curly apostrophe should be, that is UTF-8 read as the wrong encoding, and numeric entities sidestep it.

Unescaping: from entities back to text

Unescape understands every named entity in HTML, from &amp; and &nbsp; to &eacute;, &hellip; and &rarr;, and numeric ones in decimal (&#8217;) or hex (&#x2019;). It is what you need when text copied from a page's source, an RSS feed, a JSON API or a database export comes full of codes.

Tags in the input are left alone and come out as text, so unescaping never runs or removes markup. To strip tags as well, use Remove HTML tags in Text Studio. One thing to watch: &nbsp; becomes a non-breaking space, which looks like a space but is a different character, and searching, splitting and comparing treat it differently. Collapse extra spaces in Text Studio turns those back into ordinary spaces.

Questions, answered

Which characters must be escaped in HTML?

Between tags, & and < are the ones that truly matter, and > is escaped to match. Inside an attribute, the quote that wraps the value must be escaped too. Octoolo escapes all five every time, which is safe in both places.

What is the difference between &nbsp; and a normal space?

A non-breaking space keeps the words on either side on the same line, and HTML never merges it with its neighbors as it does with ordinary spaces. It is a different character, U+00A0, so code that splits text on spaces will not split on it.

Why does a web page show &amp; instead of &?

The text was escaped twice, once when it was stored and again when it was shown. Unescape it once here to see what it should be, then fix the step that escapes it the second time.

Should I use named or numeric entities?

Either works in every current browser. Named ones like &copy; are easier to read; numeric ones like &#169; exist for every character, including the many that have no name, such as most emoji.

Is escaping HTML the same as URL encoding?

No. HTML escaping makes text safe inside a page; URL encoding makes it safe inside a web address, with %26 for & instead of &amp;. A link inside an href can need both: URL-encode the values first, then escape the whole address.

HTML entities, and 16 more apps.

Download for Windows

7 days free, then from $3.99 a month for all 17 apps. Windows 10 and 11.