UTF-8 troubleshooting: a practical checklist
When text looks wrong, the cause is often a mismatch between bytes and the charset used to decode them. Check each boundary before changing stored data.
1. Keep an untouched copy
Before converting a source file or database, make a backup and work on a copy. Write down where the broken text appears: in the editor, a web page, a database export, an email, or a terminal. If it is wrong in one place but correct in another, that narrows down the layer to investigate.
2. Check the actual file bytes
Do not infer the encoding only from the file extension or the text currently shown by an editor. Inspect the file’s byte-level encoding and compare a known non-ASCII example, such as ä or €, against the source. The converter can decode a local text file when you select its known source encoding. It does not guess the encoding.
For a command-line conversion, specify both sides and write to a new file. For example, with GNU iconv and a file known to be Windows-1252:
iconv -f CP1252 -t UTF-8 legacy.txt > converted.txtCheck the result before replacing the original. Do not use this command just because the output contains strange symbols; first establish the source encoding.
3. Check HTML bytes and the HTTP response
For an HTML document, confirm that the editor saved the file as UTF-8 and include this declaration near the top of the <head>:
<meta charset="utf-8">Then inspect the network response in your browser’s developer tools. The response header should say Content-Type: text/html; charset=utf-8. The header, meta declaration, and file bytes should agree. The declaration does not transcode a Windows-1252 file into UTF-8.
4. Check databases at every connection boundary
For MySQL or MariaDB, check the database and table character sets, the connection character set, and the application’s input/output handling. For full Unicode coverage, prefer utf8mb4. Changing only a column or table does not repair text that was already stored incorrectly. Back up the database and test a migration on a copy before altering production data.
5. Check the full path, not just the page
Follow one known string from its source file or form submission through application code, database, response header, and browser. Inspect logs and CSV exports separately; command-line tools and spreadsheet software can apply their own encoding assumptions. Use the same sample at each step and note where it first changes.
Common symptoms and likely causes
| Symptom | What to inspect first |
|---|---|
| Accents are replaced by � | Whether the decoder received invalid UTF-8 or used the wrong declared charset |
| “ü” appears where “ü” belongs | A UTF-8 byte sequence decoded as Windows-1252 or a similar legacy encoding |
| Only one route or page is broken | That file’s bytes, response header, and template metadata |
| Text breaks after a database round trip | Connection charset, table charset, and whether the stored data was already garbled |
| Emoji becomes a question mark in storage | Database column and connection support for the full Unicode range, including utf8mb4 in MySQL |
For a quick byte-level view, use the UTF-8 inspector. For reference material, see the standards and implementation links.