Diagnose a structure problem before cleaning cells
CSV repair addresses files that cannot yet be read as a reliable table. Typical symptoms include inconsistent field counts, an unclosed quoted field, or a delimiter choice that turns a whole record into one field. Trimming whitespace or removing duplicate records assumes that cell boundaries have already been established. Those operations cannot safely decide where a broken quoted record was meant to end.
This tool retains the original bytes and the original decoded text. It builds a candidate from explicit decisions and validates that candidate again. A repaired structure is a proposed interpretation of the file, not proof that missing business information has been recovered. Keep the source available and compare the candidate with the exporting system when the intended values are uncertain.
Establish encoding and parsing rules
Select a CSV, TSV, or text file, then choose Source encoding. Preview original bytes displays candidate readings and the selected decoding. If strict decoding fails, resolve that issue first; treating unreadable bytes as a quotation problem can lead to incorrect edits. The encoding converter is available as a separate step when several source files need to become UTF-8 copies.
Set Delimiter and Quote character to the rules used by the source. The default expected field count is suggested from the first logical record, but you can enter a count explicitly. Indicate whether the first record is a header. An incorrect delimiter can still produce a technically consistent one-column file, so also inspect the text and the expected business structure. Consistency alone does not prove that the selected dialect is correct.
Understand the location labels
A physical line is a line in the text file. A logical record can span several physical lines when a quoted value contains embedded newlines. The problem list shows physical-line ranges separately from logical-record positions. An ordinary field-count problem usually has a known record boundary, making a precise record position meaningful.
An unclosed quote can make the remaining content look like one long field. In that case the parser cannot establish the intended downstream record boundaries. The issue is labelled uncertain and includes the suspected text range. The tool does not claim an exact count of the missing records or remove later text to make a convenient success message. Original text remains available for review even when no complete recovery is possible.
Pad only an explicitly short record
For a record with a certain boundary and fewer fields than expected, choose Pad trailing empty fields. If the expected width is three and the record contains values 1 and 2, the candidate becomes a three-field record containing 1, 2, and an empty final field. The first two values do not change. The change log records the record position, original fields, resulting fields, and the selected reason.
Padding only addresses missing fields at the end. It cannot tell you whether the omitted business value originally belonged in the middle. If a supplier meant the short record to represent the first and third columns, appending an empty field would put the second value in the wrong business column. That situation needs a deliberate text correction or a fresh source export, not an automatic padding assumption.
Isolate extra fields without discarding their evidence
A record containing too many fields is offered for isolation. The tool does not automatically delete the last fields or concatenate them into another value. Either action could destroy a legitimate identifier, amount, or note. An extra delimiter might be an exporter error, but it might also be an unquoted comma that belongs inside a particular field; the parser cannot decide that business meaning for you.
Isolating a complete record excludes it from the candidate data file and retains its original fragment in the report. Delivery then requires the explicit partial-result option. A partial result is useful when you need the confirmed good portion for investigation, but it must not be passed downstream as a complete file. The downloaded filename and report identify that limited scope.
Make a precise text edit when you know the answer
Edit an explicit text fragment lets you identify a start offset, an exclusive end offset, a replacement, and an edit reason. Offsets refer to UTF-16 code units in the original decoded text, starting at zero. They are not source byte offsets, Unicode character counts, or physical line numbers. The panel shows the selected original fragment so you can check it before adding the edit.
Non-overlapping edits can be combined. The engine checks that every recorded original fragment still matches its specified position; a stale fragment is rejected. To insert a known missing quotation mark, use equal start and end offsets at the confirmed insertion position. Do not guess the location from a wrapped screen line. Undo settings and Redo settings let you revisit decisions without overwriting the source file.
A complete three-column example
Consider a file with the header id, name, note. One complete record contains only an identifier and a name, while another contains four fields. Confirm the comma delimiter, double quotes, three expected fields, and the header setting. Preview the problems. Choose padding for the short record only if the missing note really belongs at the end, and choose isolation for the extra-field record while you investigate its meaning.
Enable the explicitly labelled partial-result option and run the complete input. The candidate contains the valid records and the padded short record, while the four-field record remains in the excluded-fragment report. The parser reads the entire candidate again using the selected quotation rules. Quoted commas, escaped double quotes, embedded newlines, and final empty fields are verified as fields rather than counted with simple newline splitting.
Read the candidate and the report together
Repair output is reserialized using the selected delimiter and quotation character, with CRLF record separators and quoted empty fields. These are documented format changes. Unlike pure encoding conversion, structure repair may therefore change the spelling of valid CSV syntax while preserving the selected field values. Review the candidate preview and the before-and-after log together before delivery.
Verify export and show summary prepares the candidate package; Confirm download new copy saves it. The package contains the candidate CSV and a report with explicit edits, exclusions, unresolved fragments, and scope. The report can contain original sensitive values, so it belongs with the source during investigation. If you need a shareable sample, continue with the confirmed data and apply the separate masking tool before sharing.
Know when to stop and obtain a fresh export
If quoting is uncertain across a large part of the file, the best next step may be to correct the source exporter and generate a new file. The tool cannot infer missing rows or values from a malformed fragment. A repaired CSV can be syntactically valid and still contain incorrect business assignments. Check identifiers, totals, representative multiline notes, and any records that changed.
File size, record count, column count, field length, and task duration are bounded. Cancel task ends the current worker. Changing encoding, parsing settings, input, or repair decisions makes the previous result stale. The grid displays pages of records but the validation run uses the complete selected input. Switching interface language changes labels, not parsing rules or cell contents. Refreshing or closing clears the memory-only session.
Protect formula-like CSV text with an apostrophe controls a separate delivery transformation. The local repaired candidate remains unchanged; the downloaded copy and report show how many fields were protected, including legitimate negative numbers. Disable protection and explicitly acknowledge raw-text risk when exact original field text is required. Protection is not a guarantee of safe behavior in every spreadsheet application.
Frequently asked questions
What should I do when CSV rows have different numbers of columns?
Check the delimiter and quoting first, because an embedded separator may belong inside a quoted field. The repair workspace diagnoses structural issues and lets you review supported repairs or exclusions. Keep unresolved records in the report. Do not treat a partial export as a complete repaired copy of the source.
Does the repair tool delete broken records automatically?
No. Complete records with problems require an explicit isolation choice, and delivery of excluded content is labelled partial. Uncertain quotation ranges remain in the report. The source file is preserved even when the candidate contains only confirmed good records.
Why can’t an unclosed quote always be fixed automatically?
The intended closing position may be ambiguous. A newline could be part of a note or the start of another record. Adding a quotation mark in the wrong position can silently move values across columns. Supply a confirmed edit or regenerate the source.
What if the candidate has duplicate headers?
The candidate may have valid CSV structure while its headers still prevent a normal table import. Resolve empty or repeated column names in the appropriate header or mapping step. A syntactic success does not authorize automatic overwriting of columns with the same name.
Can I use the partial output in a normal workflow?
Yes, when that limited scope is acceptable and clearly communicated. Check the excluded ranges first, retain the report, and do not interpret downstream row counts or totals as representing the original complete file. The partial filename is an intentional reminder.