UTF-8 frequently asked questions

Short answers to the questions that come up when text is stored, transmitted, or displayed incorrectly.

Is UTF-8 the same thing as Unicode?

No. Unicode assigns code points to characters. UTF-8 is an encoding form that represents those code points as bytes. UTF-16 is another encoding form for the same Unicode text.

Does one character always take one byte in UTF-8?

No. A Unicode code point takes one to four bytes in UTF-8. Basic ASCII characters take one byte; ü takes two, € takes three, and 😀 takes four. A visible text symbol can consist of several code points, so byte count is not the same as visible character count.

Can I use ASCII text as UTF-8?

Yes. ASCII byte values are preserved in UTF-8. A file containing only ASCII bytes is valid UTF-8, although that alone cannot tell you whether a non-ASCII file was intended to use UTF-8.

How do I declare UTF-8 in HTML?

Save the document as UTF-8 and put <meta charset="utf-8"> near the beginning of its <head>. The server should send a compatible Content-Type header. A declaration does not change the bytes already stored in the file.

What is a UTF-8 BOM?

A UTF-8 byte order mark is the optional byte sequence EF BB BF at the start of a file. UTF-8 has no endianness, so the BOM is only an encoding signature. Some tools add or remove it; check what your receiving application expects. It is usually unnecessary for HTML served on the web.

Why does my text look like “Grüße”?

This kind of mojibake often happens when UTF-8 bytes are decoded as Windows-1252 and the resulting text is saved. Stop converting the file and keep the original bytes. Recovery may be possible if the original encoding and each conversion step are known, but there is no safe universal repair for every garbled string.

Does valid UTF-8 guarantee that the displayed text is correct?

No. Validity only means the byte sequence follows UTF-8 rules. If the bytes were already decoded using the wrong charset and saved again, they may be valid UTF-8 that encodes the wrong visible text.

Should I use HTML entities for accented letters?

Usually not. If the document is saved and served as UTF-8, write characters such as é directly. Use a character reference when useful for markup, for reserved syntax such as &amp;, or when an output format requires it. Character references are HTML syntax, not a replacement for the file encoding.

Is URL encoding the same as UTF-8 conversion?

No. URL percent-encoding represents bytes safely inside a URL. A web application commonly converts text to UTF-8 bytes and then percent-encodes those bytes. UTF-8 encoding and URL escaping are separate steps.

How can I see the bytes for a string?

Enter it in the UTF-8 byte inspector. It shows code points, UTF-8 byte sequences, byte count, and (when the browser supports it) grapheme-cluster count. It analyzes browser text, not the raw bytes of an uploaded file.

Where can I verify the details?

RFC 3629 defines UTF-8 byte sequences. The Unicode UTF FAQ covers encoding forms and BOMs; the W3C guide covers HTML declarations.