Weekly ecommerce tips, deals & news.
A character encoding error happens when text saved in one encoding is read in another, so letters turn to garbage. An accented “é” shows up as “é”, and a curly apostrophe becomes “’”. In a WooCommerce store, it usually appears after a CSV import or export.
A character encoding error works like a message written in one secret code and decoded with another. Computers don’t store letters. Instead, they store numbers, and an encoding is the codebook that turns those numbers back into letters. When the writer and the reader use different codebooks, the letters come out wrong.
Plain English letters look the same in almost every encoding. Accented letters, currency symbols, and smart quotes are where the codebooks disagree. UTF-8 stores “é” as two bytes, while older Windows formats store it as one.
So the error shows up in two classic patterns:
WooCommerce stores and displays text as UTF-8, and its built-in product importer expects a UTF-8 CSV. Most encoding errors therefore start outside the store, in a spreadsheet. Older versions of Excel save a plain “CSV (Comma delimited)” file in a legacy Windows format, not UTF-8.
The common entry points look like this:
A character encoding error rarely stops an import. The importer happily writes “Crème” into the product title and reports success. That’s what makes it costly: nobody notices until a customer does.
Meanwhile, the broken text spreads. It lands in product pages, order emails, invoices, and any product data feed built from your catalog. On top of that, a shopper who searches “crème” won’t find a product stored as “Crème”.
Exports suffer the same way. An order file opened in Excel can turn a customer named “José” into “José”. Your bookkeeper then imports those names into accounting software, and the junk spreads again. Fixing it later means cleaning two systems instead of one file.
No public survey counts encoding errors in online stores. However, the data on encodings themselves explains why UTF-8 is the only safe default. W3Techs reports that UTF-8 is used by 99.1% of websites whose encoding it knows. Legacy formats survive only in small slivers, such as ISO-8859-1 at 0.8% and Windows-1252 at 0.2%.
The reason is coverage. The Unicode Standard now defines 159,801 characters, from accented letters to emoji. UTF-8 can write every one of them using 1 to 4 bytes per character, according to its official specification. A legacy single-byte format, by contrast, can only hold a tiny slice of them.
A character encoding error in practice usually looks like one quick spreadsheet edit that quietly damages hundreds of products. Here’s a hypothetical example. Imagine a small bakery-supply store that sells 600 products, many with French names.
The owner gets a new price list from a supplier. It covers 180 products, including “Crème Pâtissière Mix” and “Brûlée Torch”. The supplier’s file was saved in an old Windows format. The owner opens it in Excel, pastes in the new prices, and saves it as a plain CSV.
Next, the owner runs the WooCommerce importer with “update existing products” switched on. The importer matches each row by SKU and finishes with no errors. The owner closes the laptop, happy that the price update took five minutes.
Two days later, a customer emails a screenshot of “Cr�me P�tissi�re Mix”. It turns out 140 of the 180 titles contained an accent, and every one is now broken. Because the descriptions came from the same file, those broke too.
The damage goes further than looks. Site search for “crème” now returns nothing, so those products stop selling from search. Say those 140 products normally bring in $3,500 a week.
A week of broken titles could cost a good share of that. It also costs trust with every customer who saw the junk.
On top of that, the store’s shopping feed picks up the broken titles on its next run. So the ads now show the same garbled names to new shoppers.
Luckily, the owner exported the full catalog the week before. First, the owner reimports the titles and descriptions from that backup, matched by SKU. Then the owner reopens the supplier file and saves it with Excel’s “CSV UTF-8 (Comma delimited)” option.
Before importing again, the owner opens the file in a plain text editor and checks a few accented names. They look right, so the price-only import runs. Finally, the owner spot-checks five product pages and the next feed update.
After that, the store adds two habits. Every supplier file gets re-saved as UTF-8 before import. Also, every bulk edit starts with a fresh export, so there’s always a clean copy to restore.
You fix a character encoding error by restoring clean text and then making sure every file moves through UTF-8. Patching titles by hand works for five products, not five hundred. A clean backup and a corrected file do the job faster.
Visser Labs’ guide to importing WooCommerce products calls encoding the quietest failure in the whole process. Its Store Exporter Deluxe plugin writes a UTF-8 byte order mark by default. That mark helps spreadsheet apps detect the encoding. Plus, the plugin can export straight to XLS or XLSX, which skips the CSV guesswork entirely.
Microsoft’s own support page confirms the byte order mark point. A UTF-8 CSV opens normally in Excel if it was saved with one.
| What you’re comparing | Character Encoding Error | Failed Import |
|---|---|---|
| What goes wrong | Text is read with the wrong codebook | The upload stops, skips rows, or maps wrong |
| Does the import finish? | Usually yes, with no error shown | Often no, or only partly |
| What you see | Odd symbols inside otherwise correct products | Missing products, empty fields, or an error |
| Typical cause | A file saved in a legacy format | Mapping, delimiters, or server limits |
A character encoding error is one specific cause, while a failed import is the wider family of upload problems. An encoding error can cause a failed import, but it more often slips through as a “successful” one. So check the text itself after every import, not just the success message.
Your file is UTF-8, but the program opening it is reading it as an older Windows format. Each accented letter uses two bytes in UTF-8, so it shows up as two odd symbols. Reopen the file and choose UTF-8 as the encoding.
Go to File, then Save As, and pick “CSV UTF-8 (Comma delimited)” from the format list. The plain “CSV (Comma delimited)” option can save in a legacy format instead. Google Sheets exports CSV files as UTF-8 by default.
A BOM, or byte order mark, is a short invisible marker at the start of a file. It tells programs like Excel that the file is UTF-8. Without it, Excel may guess wrong when you double-click the file.
A character encoding error matters because it damages your catalog without telling you. Broken titles hurt search, ads, and customer trust all at once. Keeping files in UTF-8 and backing up before bulk edits keeps your product data clean as you grow.
Copyright © StoreOwnerTips.com. All Rights Reserved.