UTF-8 vs UTF-16

UTF-8 and UTF-16 encode the same Unicode code points using different byte sequences. Neither encoding changes which characters Unicode defines. The practical differences are compatibility, byte layout, and the systems that read or write the data.

Quick comparison

PropertyUTF-8UTF-16
Encoding unit8-bit byte16-bit code unit
Bytes per code point1–4Usually 2; 4 for supplementary code points, represented by a surrogate pair
ASCII textSame byte values as ASCIIEach ASCII code point takes a 16-bit code unit
Byte orderNo endianness choiceByte order matters; a BOM or an explicit label can identify it
Common web useDefault choice for HTML and web APIsUsed by some runtimes, file formats, and internal APIs

Example: the same text in both encodings

For Aü€😀, UTF-8 produces 41 C3 BC E2 82 AC F0 9F 98 80 (10 bytes). UTF-16 uses code units for each character: 0041 00FC 20AC D83D DE00. The exact byte order depends on UTF-16LE or UTF-16BE, and the emoji is represented by two code units. “One character equals two bytes” is therefore not a reliable rule for UTF-16.

UTF-8 is especially compact for text dominated by ASCII because each ASCII code point stays one byte. UTF-16 can use fewer bytes for some text dominated by code points in the Basic Multilingual Plane, but that is not a universal size or speed advantage. Measure your data and follow the format or protocol you are implementing.

Which should a website use?

For HTML served on the web, UTF-8 is the interoperable default. It preserves ASCII byte values, avoids UTF-16 byte-order issues, and is supported throughout current browsers and web infrastructure. Declare it in HTML with <meta charset="utf-8">, save the file as UTF-8, and align the HTTP Content-Type charset with the bytes.

UTF-16 can be appropriate when an existing file format, API, or runtime requires it. When exchanging UTF-16 data, specify whether bytes are little-endian or big-endian, and confirm how a byte order mark is handled. Do not choose UTF-16 only because some languages use many non-ASCII characters; storage depends on the actual text and surrounding format.

Encoding is separate from string operations

Both encodings can represent the same Unicode text, but a programming language may store strings internally using a different representation. Operations such as length, slicing, sorting, and cursor movement may count code units, code points, or grapheme clusters. Choose APIs designed for Unicode text instead of assuming one storage unit equals one visible character.

See the Unicode FAQ on UTF encodings and byte order marks and RFC 3629. To inspect individual UTF-8 sequences, try the byte inspector.