CSV Strange Characters: How to Fix Them
Why CSV text turns into mojibake, how to identify the likely source encoding, and when changing encoding will not repair a structurally broken CSV.
If José appears as José, the text is usually not missing. The bytes were decoded with the wrong character encoding. The useful question is therefore not “which symbol should I replace?”, but “which encoding was used when this file was written, and which encoding did the reader assume?”
Nablyx processes the CSV locally in your browser without uploading it, up to 150 MB, and writes a UTF-8 CSV named csv-corrected.csv using the selected delimiter, line ending, and BOM settings. Encoding repair cannot recreate characters already lost in the source or repair malformed CSV structure; docs/tool-audit/batch-a.md documents the first-party browser evidence method used to verify CSV processing boundaries.
Mojibake is a decoding mismatch
CSV describes rows and fields; character encoding describes how text becomes bytes. Those are separate layers. A file can have perfectly valid rows and columns while its text is unreadable.
Typical clues are repeatable substitutions rather than random damage: accented letters become short clusters such as é, punctuation becomes sequences such as ’, or a replacement character � appears where the reader could not decode a byte sequence.
Do not repeatedly save the damaged-looking version. Each save can turn a reversible decoding mistake into changed text.
Work out the likely source encoding before converting
Start with provenance. The application that created the CSV is usually a better clue than the characters on screen.
- A recent web app, database export, or modern command-line tool will often produce UTF-8.
- An older Windows desktop application may use a legacy Windows code page.
- A file exported by a business system can inherit that system's locale and export settings.
- A file that looks correct in one program but wrong in another strongly suggests the reader chose the wrong decoding.
If you still have access to the source system, inspect its export dialog or re-export a tiny sample containing names, punctuation, and non-ASCII text. That gives you known text to compare against the damaged version.
A useful test row
customer_id,name,city,note
A-17,José,Montréal,customer’s preferred café
Once the likely source encoding is known, decode the original file with that setting and write a new UTF-8 copy. Verify several meaningful values before replacing anything.
Nablyx uses a deliberately conservative detection order. A UTF-8, UTF-16LE, or UTF-16BE byte-order mark takes precedence. Without one, the tool first requires valid UTF-8; if that check fails, it falls back to Windows-1252 and reports that fallback so you can treat it as a lower-confidence guess rather than a certainty.
The repaired UTF-8 file can be larger than the source. Windows-1252 stores many accented characters in one byte, while the same characters need multiple bytes in UTF-8. A larger output after a successful repair is therefore not, by itself, evidence that rows were added or duplicated.
Decode the original CSV and save a UTF-8 copyWhen changing encoding will not fix the file
Encoding repair cannot reconstruct CSV structure. If columns shift partway through the file, rows have different field counts, headers are duplicated, or a quote is left open, the problem is structural; use the broken CSV diagnosis instead.
| What you see | Likely problem |
|---|---|
| Columns line up, but names or punctuation are garbled | Encoding mismatch |
| Every line appears in one spreadsheet column | Separator/import mismatch |
| Columns shift after a particular row | Malformed CSV structure |
| Only the first header has an odd invisible prefix | Possible BOM handling issue |
If the text has already been replaced with ?, empty boxes, or altered characters in the source file itself, changing the declared encoding cannot recreate the original characters. Recover from an earlier export or from the source system if possible.
Verify the repair with real text, not just the first row
Check values that are sensitive to encoding: accented names, curly punctuation, currency symbols, and any non-Latin text your dataset actually contains. Then open the corrected file in the application that will consume it. A repair is successful when the intended text survives that real import path, not merely when one preview looks better.
See the transformation
Broken text and mixed separators become one reliable CSV
Repair garbled text and inconsistent separators while leaving the actual cell values alone.
Need to do this now?
Fix the CSV encodingRelated tools
Open the tool this guide is about, or explore a related one.
- Fix CSV encoding online — free, no uploadDetect CSV encoding, delimiter, and quoting problems, then download a normalized UTF-8 file without uploading it.
- Clean and fix CSV files — no uploadRepair common CSV problems including wrong delimiters, uneven rows, empty lines, whitespace, BOM markers, and duplicate headers.