BROWSER-BASED TABLE TOOL

CSV Encoding Converter

Choose a source encoding from original bytes and convert CSV to UTF-8 without rewriting its structure.

File contents are processed locally in your browser. Refreshing or closing clears the session. The website still requests static resources.

1 Choose original files2 Confirm settings3 Review and deliver
or drop files here

Choose files, then confirm the scope

Original inputs stay in page memory. Refreshing or closing clears them; changing language does not change data.

Synthetic examples · Normal / Boundary / Error
Page memory only · Cancellable · New copies · Refreshing or closing clears data

Recover a correct reading before changing the file

A CSV Encoding Converter changes how text is represented as bytes. It is useful when an export contains Chinese, Japanese, accented names, or other characters that appear incorrectly in the receiving application. Begin with the original file from the exporting system. A copy already opened under the wrong encoding and saved again can contain different characters, question marks, or replacement symbols. Changing the encoding of that damaged copy cannot reliably reconstruct information that is no longer present.

The workbench keeps the selected source bytes in page memory. Selecting a different source encoding always decodes those bytes again. It does not convert the current garbled preview into another encoding. This distinction gives you a controlled way to compare plausible readings while preserving your starting point. The original file on your device is never overwritten; downloading creates a separately named copy.

Choose and confirm each input

Use Choose files or drop CSV, TSV, or text files into the file area. Paste a table is also available, but pasted text has already been decoded by another application. It therefore cannot reveal the original encoding of a file. For an encoding investigation, selecting the original file is the stronger input. Each selected file has its own Source encoding, Delimiter, Quote character, and confirmation checkbox.

Choose a file in Current file, select a likely encoding, and use Preview original bytes. Open Original text and byte excerpts to compare candidates. The byte excerpt uses zero-based byte offsets. The decoded excerpt uses UTF-16 code units, which are positions in JavaScript text rather than positions in the source file. Neither excerpt is presented as an exact mapping between every byte and every visible character.

Read suggestions as evidence, not certainty

A recognizable byte order mark provides a useful encoding hint. Without that marker, valid UTF-8 is offered as a suggestion. Other supported decoders are shown for comparison. Strict decoding success means that the byte sequence is acceptable to that decoder; it does not establish that the resulting words are the intended words. Some legacy encodings accept many possible byte sequences, including sequences originating in a different encoding.

Compare portions containing meaningful non-English text, not just column numbers and punctuation. A file containing only ordinary ASCII characters offers little evidence for distinguishing several encodings. If you know which system created the file, check its export setting as well. The confirmation checkbox records that you reviewed this particular input; confirming one supplier's file does not confirm another supplier's encoding.

Supported input encodings and output choices

The tested source set includes UTF-8, UTF-8 with a leading BOM, UTF-16 little endian, UTF-16 big endian, Shift-JIS, GB18030, and Windows-1252. Browser support is required. An unavailable decoder or an invalid sequence produces a visible error rather than a fabricated successful conversion. UTF-16 endianness matters because the same pair of bytes can identify different characters depending on its order.

Output is UTF-8, with the Output UTF-8 with BOM checkbox controlling a leading encoding marker. The converter does not offer arbitrary legacy output encodings. The BOM is separate from the file's delimiters and field values. Switching that option must not rewrite commas, semicolons, whitespace, quoted fields, or embedded line breaks. A normal character appearing later in the content is not removed merely because it resembles an encoding marker.

Keep transcoding separate from CSV repair

Delimiter, Quote character, the header setting, and Expected fields are used to inspect the decoded structure. They do not instruct this tool to reserialize the file. A comma inside a correctly quoted field remains inside that field, and a newline inside a field remains a newline inside that field. Record separators also retain their original spelling, including CRLF where it was present.

A structurally malformed file can decode perfectly. For example, three column names followed by a four-field record is a structure problem even when every character is readable. The report lists structural issues separately. If you explicitly accept exporting the original structure, the converter can create a UTF-8 copy carrying those issues forward. Use the repair tool to make reviewed structural changes; do not interpret a successful encoding conversion as a repaired CSV.

A complete example

Suppose a source file contains an identifier written as 00123 and a quoted description containing Japanese text, an accented product name, a comma, and an embedded newline. Select the source encoding that produces the intended description and confirm the file's comma delimiter and double quotes. Run complete input, inspect the report, and verify the export. Decoding the resulting UTF-8 data file should produce the same description, identifier, quotes, and newlines.

If BOM is enabled, the data file begins with the UTF-8 marker before its content. With BOM disabled, that added marker is absent. The identifier remains the text 00123 in the CSV. This does not force a spreadsheet program to treat the field as text when you double-click the file; the receiving application's import options still control automatic number and date interpretation.

Understand strict errors and the lossy option

Strict decoding is the normal path. Invalid bytes stop the selected decoding and preserve the source. First try a different plausible source encoding or obtain a fresh export. The explicit invalid-byte replacement option permits a lossy conversion and is labelled accordingly. Replacement characters are not evidence that the file was repaired, and the report must not present that output as a lossless recovery.

A visible replacement symbol may also have been present in the original text. Finding that character alone cannot prove whether decoding just failed or a previous application already discarded information. If the original system can still export the data, regenerate the file with a declared encoding. Keep the original and the conversion report until you have checked representative values in the receiving workflow.

Review the complete result and download

Run complete input processes the selected files, rather than just the displayed excerpt. Verify export and show summary prepares the download, and Confirm download new copy saves a ZIP containing numbered converted files and an encoding report. Generic output names avoid overwriting a source file. The report records the chosen source encoding, UTF-8 output option, size, available record count, and structural issues.

Pure transcoding preserves text, including text that a spreadsheet application might interpret as a formula. It does not silently prefix those values with apostrophes. If you need formula protection, continue into a table operation and review its explicit export option. Structural uncertainty can prevent a precise logical-record total; that uncertainty is reported rather than replaced by a count of physical newlines.

Limits and continued work

Inputs are bounded by file size, combined input size, record count, column count, field length, and processing time. The current guards include ten MiB per input and thirty MiB across selected files. These are rejection boundaries, not promises that every file at those limits will process comfortably on every device. Cancel task stops the active worker. A changed input or setting makes the old result stale and disables delivery until another run.

After conversion, continue to structure repair or the existing cleaning workbench. Data stays in memory during supported navigation and language changes; it is not placed in the page address. Refreshing or closing clears the session. Static website assets still require network requests, but file contents are processed locally. Saved rules can contain filenames or field settings, so inspect a rule file before sharing it.

Frequently asked questions

Why does readable CSV text turn into garbled characters in Excel?

The receiving application may be using a different character encoding. Start with the original bytes, compare supported source encodings, and confirm one that produces the intended text. Export a UTF-8 copy and choose the BOM option for your destination. This cannot recover characters already replaced or lost in an earlier save.

Does a BOM fix every garbled CSV?

No. A BOM is an encoding marker, not a character recovery mechanism. It can help a receiving application choose UTF-8, but the source bytes still need to be decoded correctly. It cannot restore characters already replaced or discarded by another program.

Why do several encoding previews pass?

Several encodings may accept the same bytes. Use meaningful words, the exporting application’s setting, and the byte marker together. An accepted byte sequence only establishes decoder validity; the intended language and content still require your review.

Will this convert commas to semicolons?

No. Pure transcoding preserves the decoded text structure. Delimiter selection helps diagnose the file. Use a separate reviewed parse and export operation when you deliberately want another delimiter or record-separator style.

Can different files use different source encodings?

Yes. Select and confirm each file separately. A shared UTF-8 output choice does not imply a shared source encoding. The numbered output files and report let you check which choice was used for each input.