Skip to main content

CSV Strange Characters: How to Fix Them

Nothing uploadedTested up to 150 MB

Why CSV text turns into mojibake, how to identify the likely source encoding, and when changing encoding will not repair a structurally broken CSV.

Published: 2026-08-28Updated: 2026-09-083 minFix the CSV encoding
Data workflow illustration for CSV Strange Characters: How to Fix Them

If José appears as José, the text is usually not missing. The bytes were decoded with the wrong character encoding. The useful question is therefore not “which symbol should I replace?”, but “which encoding was used when this file was written, and which encoding did the reader assume?”

Nablyx processes the CSV locally in your browser without uploading it, up to 150 MB, and writes a UTF-8 CSV named csv-corrected.csv using the selected delimiter, line ending, and BOM settings. Encoding repair cannot recreate characters already lost in the source or repair malformed CSV structure; docs/tool-audit/batch-a.md documents the first-party browser evidence method used to verify CSV processing boundaries.

Mojibake is a decoding mismatch

CSV describes rows and fields; character encoding describes how text becomes bytes. Those are separate layers. A file can have perfectly valid rows and columns while its text is unreadable.

Typical clues are repeatable substitutions rather than random damage: accented letters become short clusters such as é, punctuation becomes sequences such as ’, or a replacement character � appears where the reader could not decode a byte sequence.

Do not repeatedly save the damaged-looking version. Each save can turn a reversible decoding mistake into changed text.

Work out the likely source encoding before converting

Start with provenance. The application that created the CSV is usually a better clue than the characters on screen.

  • A recent web app, database export, or modern command-line tool will often produce UTF-8.
  • An older Windows desktop application may use a legacy Windows code page.
  • A file exported by a business system can inherit that system's locale and export settings.
  • A file that looks correct in one program but wrong in another strongly suggests the reader chose the wrong decoding.

If you still have access to the source system, inspect its export dialog or re-export a tiny sample containing names, punctuation, and non-ASCII text. That gives you known text to compare against the damaged version.

A useful test row

customer_id,name,city,note
A-17,José,Montréal,customer’s preferred café

Once the likely source encoding is known, decode the original file with that setting and write a new UTF-8 copy. Verify several meaningful values before replacing anything.

Nablyx uses a deliberately conservative detection order. A UTF-8, UTF-16LE, or UTF-16BE byte-order mark takes precedence. Without one, the tool first requires valid UTF-8; if that check fails, it falls back to Windows-1252 and reports that fallback so you can treat it as a lower-confidence guess rather than a certainty.

The repaired UTF-8 file can be larger than the source. Windows-1252 stores many accented characters in one byte, while the same characters need multiple bytes in UTF-8. A larger output after a successful repair is therefore not, by itself, evidence that rows were added or duplicated.

Decode the original CSV and save a UTF-8 copy

When changing encoding will not fix the file

Encoding repair cannot reconstruct CSV structure. If columns shift partway through the file, rows have different field counts, headers are duplicated, or a quote is left open, the problem is structural; use the broken CSV diagnosis instead.

What you seeLikely problem
Columns line up, but names or punctuation are garbledEncoding mismatch
Every line appears in one spreadsheet columnSeparator/import mismatch
Columns shift after a particular rowMalformed CSV structure
Only the first header has an odd invisible prefixPossible BOM handling issue

If the text has already been replaced with ?, empty boxes, or altered characters in the source file itself, changing the declared encoding cannot recreate the original characters. Recover from an earlier export or from the source system if possible.

Verify the repair with real text, not just the first row

Check values that are sensitive to encoding: accented names, curly punctuation, currency symbols, and any non-Latin text your dataset actually contains. Then open the corrected file in the application that will consume it. A repair is successful when the intended text survives that real import path, not merely when one preview looks better.

See the transformation

Broken text and mixed separators become one reliable CSV

Repair garbled text and inconsistent separators while leaving the actual cell values alone.

Input
name;city;note
José;München,West;"ready"
Chloë|Zürich;review
Result
name;city;note
José;"München,West";ready
Chloë;Zürich;review