UTF-8 and Unicode references
These links point to standards bodies and language documentation that define or implement character encoding. Each entry includes what it is useful for, so you can go straight to the source you need.
Standards and definitions
- RFC 3629: UTF-8
The Internet Standard that defines valid UTF-8 byte sequences, ASCII compatibility, and security considerations.
- Unicode FAQ: UTF-8, UTF-16, UTF-32, and BOM
Explains Unicode encoding forms, supplementary characters, byte order marks, and common implementation questions.
- Unicode FAQ: normalization
Explains canonically equivalent strings and the NFC, NFD, NFKC, and NFKD normalization forms.
- W3C: declaring character encodings in HTML
Practical guidance on the HTML charset declaration and the HTTP response header.
- WHATWG Encoding Standard
Defines browser encoding labels and decoding behavior, including legacy web encodings.
Implementation documentation
- Python codecs
Python’s standard library for encoding, decoding, stream readers, and strict error handling.
- PHP mbstring
Multibyte string functions for detecting, converting, and processing text in PHP.
- MDN: TextEncoder
Browser API reference for encoding JavaScript strings as UTF-8 bytes.
- MDN: TextDecoder
Browser API reference for decoding byte arrays with a selected encoding and strict error behavior.
Use references with the right scope
A character encoding tells software how to map bytes to text. It does not define font coverage, translation, sort order, URL escaping, or database migration rules. When diagnosing a problem, consult the source that matches the layer you are changing: Unicode for code points and normalization, HTML guidance for web declarations, or your language’s documentation for conversion APIs.
For a worked walkthrough, see UTF-8 troubleshooting. To compare a string with its byte representation, open the byte inspector.