UTF-8 troubleshooting: a practical checklist

When text looks wrong, the cause is often a mismatch between bytes and the charset used to decode them. Check each boundary before changing stored data.

1. Keep an untouched copy

Before converting a source file or database, make a backup and work on a copy. Write down where the broken text appears: in the editor, a web page, a database export, an email, or a terminal. If it is wrong in one place but correct in another, that narrows down the layer to investigate.

2. Check the actual file bytes

Do not infer the encoding only from the file extension or the text currently shown by an editor. Inspect the file’s byte-level encoding and compare a known non-ASCII example, such as ä or €, against the source. The converter can decode a local text file when you select its known source encoding. It does not guess the encoding.

For a command-line conversion, specify both sides and write to a new file. For example, with GNU iconv and a file known to be Windows-1252:

iconv -f CP1252 -t UTF-8 legacy.txt > converted.txt

Check the result before replacing the original. Do not use this command just because the output contains strange symbols; first establish the source encoding.

3. Check HTML bytes and the HTTP response

For an HTML document, confirm that the editor saved the file as UTF-8 and include this declaration near the top of the <head>:

<meta charset="utf-8">

Then inspect the network response in your browser’s developer tools. The response header should say Content-Type: text/html; charset=utf-8. The header, meta declaration, and file bytes should agree. The declaration does not transcode a Windows-1252 file into UTF-8.

4. Check databases at every connection boundary

For MySQL or MariaDB, check the database and table character sets, the connection character set, and the application’s input/output handling. For full Unicode coverage, prefer utf8mb4. Changing only a column or table does not repair text that was already stored incorrectly. Back up the database and test a migration on a copy before altering production data.

5. Check the full path, not just the page

Follow one known string from its source file or form submission through application code, database, response header, and browser. Inspect logs and CSV exports separately; command-line tools and spreadsheet software can apply their own encoding assumptions. Use the same sample at each step and note where it first changes.

Common symptoms and likely causes

SymptomWhat to inspect first
Accents are replaced by �Whether the decoder received invalid UTF-8 or used the wrong declared charset
“ü” appears where “ü” belongsA UTF-8 byte sequence decoded as Windows-1252 or a similar legacy encoding
Only one route or page is brokenThat file’s bytes, response header, and template metadata
Text breaks after a database round tripConnection charset, table charset, and whether the stored data was already garbled
Emoji becomes a question mark in storageDatabase column and connection support for the full Unicode range, including utf8mb4 in MySQL

For a quick byte-level view, use the UTF-8 inspector. For reference material, see the standards and implementation links.