Guide

HTML entities: named, decimal and hex, and when you actually need them

An HTML entity is a way of writing a character using only ASCII. The degree sign can be written as its named entity, as its decimal reference, or as its hexadecimal reference. All three render identically. On a page served as UTF-8 you can also just type the character.

Last checked 2026-08-05.

The three forms

The same character written four ways
FormExampleNotes
The character°Works on any page served as UTF-8. Shortest and most readable.
Named entity°Only exists for a fixed list of characters. Readable, but not universal.
Decimal reference°Works for every Unicode character. The number is the code point in decimal.
Hex reference°Works for every Unicode character. The number is the code point in hexadecimal, matching how U+00B0 is written.

The five characters you must always escape

Everything else is a preference. These five are not, because leaving them raw changes how the browser parses your document or opens an injection hole.

  • The ampersand, written &. Escape it first, always, or your other escapes get mangled.
  • The less-than sign, written <. Raw, it starts a tag.
  • The greater-than sign, written >. Required inside attribute-heavy contexts and safest everywhere.
  • The double quote, written ". Required inside a double-quoted attribute value.
  • The single quote, written '. Required inside a single-quoted attribute value. The named form ' is valid HTML5 but not in older HTML, so the numeric form is the safe choice.

Named entities are a fixed, finite list

There is no named entity for most characters. The HTML standard defines a specific set, and it stopped growing long ago. That is why the degree sign, the arrows and the Greek letters have friendly names while the vast majority of Unicode has only numeric references.

This site shows the named entity on a symbol page only where a standard one genuinely exists. If a page shows the decimal and hex forms and no named form, that is the answer: there is no name for that character, and the numeric reference is what you should use.

Why your page probably needs no entities at all

If your document declares UTF-8 and is actually saved as UTF-8, you can paste the character directly into your markup and it will render. That is more readable in source, shorter over the wire and easier to search for. The declaration is a single tag in the head of the document, and it must appear within the first kilobyte.

Entities earn their keep in three situations: when the file has to survive a pipeline that mangles non-ASCII bytes, when you are generating markup inside a string in another language and want to avoid encoding questions entirely, and when the character is invisible. A non-breaking space written as a raw character is indistinguishable from a normal space in your editor; written as its named entity it is obvious.

Where entities go wrong in practice

The first failure is double escaping. Text that already contains an entity gets escaped again on its way through a template or a storage layer, so the ampersand becomes its own entity and the reader sees the literal characters of the escape rather than the character it stood for. The tell is a visible ampersand followed by a word and a semicolon on the page. The fix is upstream: escape once, at the point where text enters markup, and not again.

The second is a missing semicolon. Most browsers recover from it for a handful of legacy named entities and do not recover for anything else, which makes the behaviour look random. Always terminate the reference.

The third is an entity in the wrong context. Entities are parsed by the HTML parser, so they work in text content and in attribute values. They do not work inside a script block, inside a style block, inside a plain text email, or in a JSON payload. In each of those places the surrounding language has its own escape and the HTML one is just literal text.

The characters worth escaping even when you do not have to

Beyond the five that are mandatory, a small group is worth escaping simply because it is invisible or ambiguous in source. The non-breaking space is the clearest case: as a raw character it is indistinguishable from a normal space in an editor, and a stray one is a common cause of text that will not wrap where you expect. Written as its named entity it is obvious to the next person reading the file.

The same argument applies to the zero-width characters, to the narrow no-break space and to the soft hyphen. All of them are legitimate and useful, and all of them are impossible to see. Escaping them is a note to your future self rather than a technical requirement.

Everything else is a preference, and on a UTF-8 page the readable choice is usually the character itself. A line of markup containing an arrow is easier to scan, easier to search for and shorter over the wire than the same line containing a numeric reference.

Entities and accessibility

A screen reader never sees your entity. By the time the text reaches it, the HTML parser has already turned the reference into a character, and the character is announced by its Unicode name or by whatever pronunciation rule the reader applies. Writing a character as a named entity therefore changes nothing about how it is read aloud, which is worth knowing before anyone proposes escaping for accessibility reasons.

What does change the announcement is which character you chose. A multiplication sign is announced differently from the letter x, a minus sign differently from a hyphen, and a square root differently from a check mark. That is the accessibility argument for picking the semantically correct character, and it has nothing to do with entities at all.

The related escapes: CSS and JavaScript

HTML entities do not work outside HTML. In a CSS content property the same character is written as a backslash followed by the hexadecimal code point, and the escape must be terminated by a space or padded to six digits, or the next character gets absorbed into it. In JavaScript the escape is a backslash and a lowercase u followed by four hexadecimal digits, and for characters above U+FFFF the digits go in braces.

Every symbol page on this site prints all three next to each other with a copy button on each row, precisely because mixing them up is the most common way this goes wrong.

Frequently asked questions

Is there a named entity for every character?

No. The named list is fixed and covers a small fraction of Unicode. Decimal and hexadecimal numeric references work for every character.

Do I need entities on a UTF-8 page?

Only for the ampersand, the angle brackets and the quotes used inside attributes. Everything else can be the character itself.

Why does my entity show as literal text?

Almost always a missing semicolon, or the ampersand itself was escaped twice so the browser sees the literal text rather than a reference.

Copy the characters in this guide

Related category hubs

Other guides

Or skip the keystrokes: open the full symbol reference and click any glyph to copy it.