Store Owner Tips

Subscribe to our newsletter

Weekly ecommerce tips, deals & news.

Thank You, we'll be in touch soon.

Latest News

Character Encoding Error

A character encoding error happens when text saved in one encoding is read in another, so letters turn to garbage. An accented “é” shows up as “é”, and a curly apostrophe becomes “’”. In a WooCommerce store, it usually appears after a CSV import or export.


Key Takeaways

  • It’s a translation mix-up, not lost data: The bytes are usually intact, but the reader uses the wrong codebook.
  • It fails silently: The import often reports success while product titles fill up with odd symbols.
  • Spreadsheet apps cause most of it: Saving a CSV in a legacy format is the most common trigger.
  • UTF-8 everywhere is the fix: Save, import, and export every file as UTF-8, and check a sample before going live.

How Does a Character Encoding Error Work?

A character encoding error works like a message written in one secret code and decoded with another. Computers don’t store letters. Instead, they store numbers, and an encoding is the codebook that turns those numbers back into letters. When the writer and the reader use different codebooks, the letters come out wrong.

Why the Same Bytes Look Different

Plain English letters look the same in almost every encoding. Accented letters, currency symbols, and smart quotes are where the codebooks disagree. UTF-8 stores “é” as two bytes, while older Windows formats store it as one.

So the error shows up in two classic patterns:

  • UTF-8 read as an old Windows format: Each two-byte letter splits into two odd characters, so “Café” becomes “Café”.
  • An old Windows format read as UTF-8: The single byte isn’t valid UTF-8. You get a black diamond with a question mark (�).
  • Double encoding: A file gets converted twice, and the junk characters multiply with each round trip.

Where Encoding Errors Enter a WooCommerce Store

WooCommerce stores and displays text as UTF-8, and its built-in product importer expects a UTF-8 CSV. Most encoding errors therefore start outside the store, in a spreadsheet. Older versions of Excel save a plain “CSV (Comma delimited)” file in a legacy Windows format, not UTF-8.

The common entry points look like this:

  • Supplier price lists: A supplier sends a CSV saved in a regional format, and you import it as is.
  • Edit-and-reimport loops: You export products, open the file in Excel, save it, and then import it back.
  • Store moves: A store migration pulls data from an old database set to a different character set.
  • Opening exports by double-click: Excel may misread a UTF-8 file that lacks a byte order mark.

Why It Hurts More Than It Looks

A character encoding error rarely stops an import. The importer happily writes “Crème” into the product title and reports success. That’s what makes it costly: nobody notices until a customer does.

Meanwhile, the broken text spreads. It lands in product pages, order emails, invoices, and any product data feed built from your catalog. On top of that, a shopper who searches “crème” won’t find a product stored as “Crème”.

Exports suffer the same way. An order file opened in Excel can turn a customer named “José” into “José”. Your bookkeeper then imports those names into accounting software, and the junk spreads again. Fixing it later means cleaning two systems instead of one file.

What Do the Numbers Say About Character Encoding Errors?

No public survey counts encoding errors in online stores. However, the data on encodings themselves explains why UTF-8 is the only safe default. W3Techs reports that UTF-8 is used by 99.1% of websites whose encoding it knows. Legacy formats survive only in small slivers, such as ISO-8859-1 at 0.8% and Windows-1252 at 0.2%.

The reason is coverage. The Unicode Standard now defines 159,801 characters, from accented letters to emoji. UTF-8 can write every one of them using 1 to 4 bytes per character, according to its official specification. A legacy single-byte format, by contrast, can only hold a tiny slice of them.


What Does a Character Encoding Error Look Like in Practice?

A character encoding error in practice usually looks like one quick spreadsheet edit that quietly damages hundreds of products. Here’s a hypothetical example. Imagine a small bakery-supply store that sells 600 products, many with French names.

The Edit

The owner gets a new price list from a supplier. It covers 180 products, including “Crème Pâtissière Mix” and “Brûlée Torch”. The supplier’s file was saved in an old Windows format. The owner opens it in Excel, pastes in the new prices, and saves it as a plain CSV.

Next, the owner runs the WooCommerce importer with “update existing products” switched on. The importer matches each row by SKU and finishes with no errors. The owner closes the laptop, happy that the price update took five minutes.

The Fallout

Two days later, a customer emails a screenshot of “Cr�me P�tissi�re Mix”. It turns out 140 of the 180 titles contained an accent, and every one is now broken. Because the descriptions came from the same file, those broke too.

The damage goes further than looks. Site search for “crème” now returns nothing, so those products stop selling from search. Say those 140 products normally bring in $3,500 a week.

A week of broken titles could cost a good share of that. It also costs trust with every customer who saw the junk.

On top of that, the store’s shopping feed picks up the broken titles on its next run. So the ads now show the same garbled names to new shoppers.

The Fix

Luckily, the owner exported the full catalog the week before. First, the owner reimports the titles and descriptions from that backup, matched by SKU. Then the owner reopens the supplier file and saves it with Excel’s “CSV UTF-8 (Comma delimited)” option.

Before importing again, the owner opens the file in a plain text editor and checks a few accented names. They look right, so the price-only import runs. Finally, the owner spot-checks five product pages and the next feed update.

After that, the store adds two habits. Every supplier file gets re-saved as UTF-8 before import. Also, every bulk edit starts with a fresh export, so there’s always a clean copy to restore.


How Do You Fix a Character Encoding Error?

You fix a character encoding error by restoring clean text and then making sure every file moves through UTF-8. Patching titles by hand works for five products, not five hundred. A clean backup and a corrected file do the job faster.

  • Export before every bulk edit: A fresh export gives you a clean copy to restore from.
  • Save as UTF-8 on purpose: In Excel, choose “CSV UTF-8 (Comma delimited)”, not the plain CSV option.
  • Check in a text editor: A plain text editor shows the real characters, without a spreadsheet’s guesswork.
  • Open exports the safe way: In Excel, use Data, then Get Data, then From Text/CSV, and pick UTF-8.

Visser Labs’ guide to importing WooCommerce products calls encoding the quietest failure in the whole process. Its Store Exporter Deluxe plugin writes a UTF-8 byte order mark by default. That mark helps spreadsheet apps detect the encoding. Plus, the plugin can export straight to XLS or XLSX, which skips the CSV guesswork entirely.

Microsoft’s own support page confirms the byte order mark point. A UTF-8 CSV opens normally in Excel if it was saved with one.


What’s the Difference Between a Character Encoding Error and a Failed Import?

What you’re comparingCharacter Encoding ErrorFailed Import
What goes wrongText is read with the wrong codebookThe upload stops, skips rows, or maps wrong
Does the import finish?Usually yes, with no error shownOften no, or only partly
What you seeOdd symbols inside otherwise correct productsMissing products, empty fields, or an error
Typical causeA file saved in a legacy formatMapping, delimiters, or server limits

A character encoding error is one specific cause, while a failed import is the wider family of upload problems. An encoding error can cause a failed import, but it more often slips through as a “successful” one. So check the text itself after every import, not just the success message.


Frequently Asked Questions

Why does my CSV show weird characters like é?

Your file is UTF-8, but the program opening it is reading it as an older Windows format. Each accented letter uses two bytes in UTF-8, so it shows up as two odd symbols. Reopen the file and choose UTF-8 as the encoding.

How do I save a CSV as UTF-8 in Excel?

Go to File, then Save As, and pick “CSV UTF-8 (Comma delimited)” from the format list. The plain “CSV (Comma delimited)” option can save in a legacy format instead. Google Sheets exports CSV files as UTF-8 by default.

What is a BOM in a CSV file?

A BOM, or byte order mark, is a short invisible marker at the start of a file. It tells programs like Excel that the file is UTF-8. Without it, Excel may guess wrong when you double-click the file.


Why Does a Character Encoding Error Matter?

A character encoding error matters because it damages your catalog without telling you. Broken titles hurt search, ads, and customer trust all at once. Keeping files in UTF-8 and backing up before bulk edits keeps your product data clean as you grow.

Share article

Subscribe to our newsletter

Weekly ecommerce tips, deals & news.

Nice – You're in!

Copyright © StoreOwnerTips.com. All Rights Reserved.