Learn the shape of an unfamiliar CSV
A CSV Data Profiler helps you inspect a file before choosing cleaning rules or making statistical claims. It reports how many records and fields were analyzed, which values are empty, how often values occur, and which columns appear constant or contain mixed text shapes. The profile describes observations. It does not decide that an unusual amount is a business error, assign an unexplained quality score, or rewrite a value merely because it resembles a date or number.
Profiling is especially useful when a supplier changes an export, a category vocabulary grows unexpectedly, or an unfamiliar table arrives without a schema. It can reveal questions worth investigating: why does one status dominate, why is an address column empty, or why are some quantities stored with a comma? Those observations are starting points for decisions. An optional field can legitimately be empty, and a constant country field can be entirely correct in a file covering one country.
Import and choose the population being described
Choose files reads UTF-8 CSV or TSV, while Paste a table accepts copied rows. Confirm parsing before interpreting the summary. The header choice determines whether the first record names the fields or participates in the data. Quoted delimiters and multiline cells are handled as structured CSV, so the number of logical records may differ from the number of physical lines in a text editor. Duplicate-looking field labels should be inspected in the import preview rather than merged by appearance.
Processing scope defaults to All records. First 20 records only is available for an exploratory sample, but its output is clearly marked as a sample and downloads include SAMPLE in the filename. A sample profile cannot establish full-file uniqueness, full error counts, or a complete category distribution. Later records can contain values absent from the beginning of the file. Switching to All records recalculates every metric; it does not extrapolate the sample percentages to the entire input.
Define empty values before comparing percentages
The empty definition distinguishes missing cells, present empty strings, and whitespace-only text. Missing cells and empty strings are included initially. Whitespace is not included unless selected. You may add literal marker values such as a source-specific dash, but NULL, N/A, zero, and false remain ordinary values by default. This policy applies consistently to empty counts, nonempty counts, distinct counts, and numeric participation. A changed policy creates a new profile rather than modifying the source.
The empty percentage uses the number of analyzed records as its denominator. With values A, A, B, and an empty string, one of four observations is empty, giving 25 percent. The nonempty count is three. There are two different nonempty values, and one of those values occurs only once. These metrics answer different questions; labeling all of them simply unique would make it easy to mistake a vocabulary size for the number of singly occurring observations.
Read distinct values and frequency distributions
Distinct counts the number of different nonempty values under strict source-text comparison. Singletons counts how many of those values appear exactly once. Case, surrounding whitespace, accents, and punctuation remain significant in this descriptive view. A and a are different values, and a trailing space is not silently removed. If those differences should be standardized, review them in the value-standardization tool and then profile the explicitly changed result.
Frequent values shows the most common items and their counts. Its preview limit can be set from one to one hundred per column. That display limit does not reduce the distinct-value calculation or change the population being counted. A high-cardinality identifier field may have thousands of different values while the frequency view shows only a small subset. Use source links to locate the observations behind a displayed value, and use the report's full distinct count when discussing the field's diversity.
Treat shape suggestions as suggestions
An empty column has no nonempty observations under the selected empty policy. A constant column has one distinct nonempty value, even if some observations are missing. Mixed-format identifies more than one broad text shape, such as numeric-looking text together with ordinary words. A date-shaped string is not automatically a valid calendar date. The profiler does not convert these suggestions into new column types or modify original values during analysis.
A column containing 00123 and 00124 may look numeric while actually holding identifiers. Adding those identifiers would usually have no meaningful interpretation. Conversely, a quantity column containing 10 and unknown contains a numeric-looking value and text, but the text might be a documented missing marker rather than a typographical error. Establish the data contract before adding validation or conversion. The purpose of a profile is to make those distinctions visible, not to manufacture confidence from formatting alone.
Confirm numeric interpretation before calculating
Use Add numeric statistics to select a field and configure its numeric convention. Confirm this numeric interpretation is required. Decimal and grouping separators, negative notation, currency tokens, and percentage semantics come from the same settings used by number normalization. Interface language does not choose these rules. An applied normalized column can supply its previously confirmed interpretation, but you should still review the selected field before treating its values as measurements.
Successful numeric observations contribute to minimum, maximum, mean, and quartiles. Empty values are omitted according to the empty policy. Invalid or unsupported numeric text is counted separately and appears in the issue-location report; it is never substituted with zero. The numeric count therefore may be smaller than the nonempty count. When discussing a mean, cite the valid numeric population and the invalid count together, especially if the missing or rejected observations might systematically differ from the included observations.
Understand the fixed quantile calculation
Quartiles use a fixed linear interpolation convention, often called type seven. Sort the valid numeric values, calculate the zero-based position h = (n − 1) × p, and interpolate between the values on either side of that position. The lower quartile uses p = 0.25, the median uses 0.5, and the upper quartile uses 0.75. This method is stated in the configuration report so a later library default cannot silently redefine the result.
For values zero and ten, the lower-quartile position is 0.25. Interpolating one quarter of the distance from zero to ten gives 2.5. The median is five and the upper quartile is 7.5. With a single numeric observation, all three quartiles equal that observation. With no valid numbers, numeric statistics remain blank rather than reporting a fabricated zero. Decimal arithmetic preserves supported input precision; a recurring mean is rounded to the documented maximum of forty fractional places.
Locate observations and choose an appropriate next step
After Run and review, Inspect column selects a field in the profile panel. Its statistics, common values, and empty or invalid-value locations appear together. Source links open the underlying records without exposing their contents in the page URL. A high-frequency category can be traced to its records, and an invalid numeric cell can be compared with adjacent identifiers or source information. The grid's search affects the view only and does not redefine the statistical population.
Use Standardize Categorical Data for a reviewed vocabulary change, CSV Number Format Converter for an explicit numeric representation change, or CSV Data Validator for a known requirement. Profiling by itself does not know which correction is justified. Apply to workflow keeps the current table available for the next tool without changing its cells. After an actual transformation, return to the profiler and calculate a new report to verify the intended effect on counts and distributions.
Export a reproducible report and interpret its limits
Export new copy can download Column profiles, Frequent values, Issue locations, or Settings. The settings record includes processing scope, analyzed record count, input identity, run time, numeric configuration, and the quantile convention. Keep that report with the output when another reviewer needs to understand the denominator or calculation policy. A column report without its population definition can be misleading even when every displayed number is arithmetically correct.
The shared exporter verifies generated CSV or TSV by reading it back. Formula-prefix protection and raw-text preservation are explicit choices. Large issue reports may exceed audit limits even when the imported table fits the input limit; split the work or reduce the scope, and keep the resulting scope labels. Cancel task stops active processing without deleting the input. Data lives in the page session, so refresh or closing clears it. No cloud file-processing service, automatic repair, quality certification, or universal performance guarantee is part of this profiler.
Frequently asked questions
How can I find missing values and frequently repeated values in a column?
Choose the columns to profile and inspect the missing-value counts, distinct values, and frequency information. Confirm whether the run covers the full input or a sample before drawing conclusions. Numeric summaries depend on confirmed number interpretation; a nonempty text field is not necessarily a valid numeric value.
What is the difference between distinct values and singletons?
Distinct is the vocabulary size among nonempty observations. Singletons is the number of values that occur exactly once. A, A, B contains two distinct values but only one singleton.
Can a sample prove that a column is unique?
No. Duplicate values can occur outside the sample. Use a complete validation run with a uniqueness rule before asserting that the full file has unique keys.
Does an inferred numeric shape change my values?
No. Shape suggestions are descriptive. Numeric statistics require an explicitly confirmed interpretation, and the original table remains unchanged.
Why is the numeric count below the nonempty count?
Some nonempty values failed the configured numeric interpretation. They are reported separately and excluded from numeric statistics rather than treated as zero.