BROWSER-BASED TABLE TOOL

CSV Test Data Generator

Generate synthetic CSV test data with configured fields, seeded values, and located anomalies. Work locally in your browser and review results before export.

File contents are processed locally in your browser. Refreshing or closing clears the session. The website still requests static resources.

1Import data—2Configure—3Review—4Export
0 files · 0 applied steps
Source data retained. Configure rules and run. No records have been changed or deleted.
0 columns · 0 records
#Source · Logical record
No records in this view. An empty result is valid.

Cell text

Kept in page memory · Continue across tools · Refreshing or closing clears the sessionCSV / TSV · UTF-8
Sample library · Normal / Boundary / Error

Clear this session?

Imported data, steps, and results will be released. Export any needed copies first. Original files are unaffected.

Create synthetic records without an input file

CSV Test Data Generator creates a table from a field configuration rather than a customer file. It is useful for testing imports, demonstrating a workflow, checking quoting behavior, or reproducing a failure involving deliberately incomplete data. The initial configuration produces a sequence identifier and an explicit TEST label. No external personal-profile service is contacted, and generated values are not presented as information about real people.

You can start immediately without choosing a file. Alternatively, use Choose files or Paste a table to supply a template containing only headers. Confirm file parsing verifies the template, then the generator creates one configurable field for each header. Templates with data records are rejected. This prevents a test-data operation from quietly treating an existing customer dataset as a source of example identities. Header names should still be reviewed before sharing a saved configuration.

Configure field names, order, and record count

Record count sets the number of generated records. Add field creates another field card, and the up and down buttons change the output order. Every field needs a unique, nonempty name. Remove field deletes that field from the configuration. Select Field type and fill in the corresponding settings before generation. A preview can help check a value's shape, but a completed full run is necessary before exporting the requested dataset.

Sequence produces consecutive integer identifiers from Starting ID and pads them to Minimum ID width when needed. The sign of a negative identifier stays before its zero padding. Fixed-length text uses lowercase letters and digits with a configured length. Allowed values draws from an explicit list, one value per line. Duplicate list entries do not create additional unique possibilities. Test label emits the fixed text you enter, making values such as TEST-ORDER or SYNTHETIC easy to recognize in a demonstration.

Integer and Decimal use an inclusive configured range. Decimal also takes a fixed number of decimal places. Ranges must be ordinary decimal text without exponent notation or thousands separators; endpoints must align to the requested precision. Supported numeric configurations are bounded to avoid rounding or huge allocations. Values with more than one hundred digits or more than 4,294,967,296 candidate positions are rejected. A narrow high-value range can still be useful for checking that another system preserves long numeric text.

Dates and reproducibility use explicit settings

Date range accepts valid calendar dates written as YYYY-MM-DD and includes both endpoints. Generation uses a fixed range, so opening the tool on a different day does not silently move the dates. Fixed date reference records an explicit reference date with the configuration. It does not shift the chosen date endpoints or refer to the current clock. Keeping both the endpoints and the reference in the saved settings makes the intended fixture easier to explain.

Random seed is text that initializes the deterministic sequence of random choices. To reproduce a dataset, preserve the seed, generator version, record count, field order, all field settings, anomaly rules, and fixed reference date. A seed alone is insufficient when a field is inserted or a range changes, because those changes can consume random choices differently. The settings report stores the generator version and complete configuration alongside the generated output.

Preview sample and estimate validates the requested normal configuration and shows at most twenty sample records. It intentionally omits anomaly injection. Anomalies are placed during the full run, where the complete set of records is available. Do not infer final missing-value positions from the preview. Run and review produces the requested count, applies the enabled anomaly rules, and creates the actual anomaly report. Repeating that full configuration produces the same values and locations for the same generator version.

Check uniqueness before allocating records

Unique normal values requires a field to have enough distinct values for the requested record count. Three unique records cannot be drawn from an allowed list containing only A and B. The generator blocks this configuration before generation instead of retrying indefinitely or quietly allowing a duplicate. The same capacity check applies to integer, decimal, date, and fixed-length-text fields. A constant test label has only one possible value and therefore cannot be unique across several records.

Uniqueness describes the normal values before deliberately injected anomalies. A later missing-value or duplicate-key injection can break it, and that is part of the test configuration. Sequence fields naturally provide distinct consecutive values. Duplicate-key injection requires a sequence or an explicitly unique normal field so that an additional duplicate has a clear meaning. If you simply need ordinary repeated categories, use Allowed values without uniqueness instead of describing their natural repeats as injected key errors.

Inject known errors and inspect actual locations

Inject anomalies is collapsed and empty by default. Add injection lets you choose a field, an anomaly type, and an exact count. Missing values replaces selected cells with empty strings. Duplicate keys rewrites additional records to an existing nonempty key from the same field. Long fields writes a configured number of X characters. Cell line breaks appends a newline and a TEST marker, which is useful for exercising CSV quoting and multiline parsing.

Injections for the same field use distinct cells, and a duplicate donor is reserved before other injections. This protects a requested missing-value count from being accidentally multiplied by copying a missing donor. If there are not enough independent cells for the requested rules, the configuration is blocked. Injections in different fields can still affect the same record. The Actual anomalies report lists logical record number, field, type, previous value, final value, and whether the final value matches the recorded injection.

Duplicate count means the number of additional records actually rewritten to an existing key. It does not mean the size of the resulting duplicate group. One donor plus three rewritten records forms a group involving four records. The summary separately reports rewritten extras, final duplicate extras, final duplicate groups, and involved records. When multiple fields receive duplicate injections, those final counts are per-field totals and may count the same business record in more than one field's grouping.

A complete fixture and an impossible configuration

Keep the default hundred records and add a Missing values injection to label with count five. Run and review should produce a hundred rows, with exactly five injected empty label cells and five corresponding report entries. The id values keep their configured width. Add a line-break injection to another available label position and inspect the result through the cell-detail view. Export and reimport that CSV to check that the newline stays inside one logical field rather than becoming an extra record.

For a deliberate configuration error, choose three records and one Allowed values field containing A and B with uniqueness enabled. Preview and full generation both reject the insufficient value space. The normal configuration is preserved for editing. Widen the list, reduce the count, or explicitly remove uniqueness if repetition is actually acceptable. This is a useful negative test because an apparently successful file with duplicated values would contradict the selected rule.

Export real data files and retain the test explanation

Export new copy offers CSV and TSV, the generation settings, the actual anomaly list, or a ZIP containing the result artifacts. Quotes, separators, non-Latin text, accents, emoji, and embedded newlines are serialized through the shared exporter and parsed back before download. Formula-prefix protection is explicit. If a test intentionally exercises formula-looking strings, select the raw-text option with its acknowledgement and document that export choice; enabling protection would change the fixture's actual text.

Use clearly marked test forms for contact-related fields, such as labels or reserved example-domain addresses that you deliberately supply. The generator does not promise that an arbitrary realistic-looking address is safe to contact. Apply to workflow can pass the synthetic table to fixed-seed sampling, splitting, or comparison. Save rules before sharing a reproducible fixture and check its field names and fixed values. No cloud backup is created. Refreshing or closing clears the in-memory configuration and result.

Size guards and cancellation

The generator checks record count, column count, total cells, individual field lengths, unique value capacity, and a conservative generated-text budget. Long-field and newline injections contribute to that projected budget. The configured character budget is an allocation guard, not a promise about final UTF-8 download bytes or a universally supported dataset size. Large Unicode values can occupy more bytes than their character count suggests. Device memory and browser behavior still affect what completes comfortably.

A blocked size check asks you to reduce the count, range, length, or injection scope; it does not truncate the final file. Cancel task terminates a running worker and leaves the prior valid result available. Editing the configuration makes that result stale and disables export until a new full run succeeds. The preview is intentionally bounded, and a successful twenty-record preview does not certify the performance of the largest allowed full configuration on every device.

Frequently asked questions

Can I generate test data with intentional missing values and duplicates?

Yes. Configure the supported field types and row count, then add the available anomaly settings. Review the generated output and anomaly report rather than assuming the small preview shows every injected issue. The data is synthetic, and a reproducible seed helps repeat a test without using real customer records.

Is generated data taken from real users?

No external personal-profile data is used. Values come from your configured sequences, ranges, lists, and labels. Do not enter real customer values into an allowed list if the fixture must remain synthetic.

Why can two allowed values not produce three unique records?

There are only two distinct possibilities. The generator reports insufficient capacity before creating the dataset, rather than silently violating uniqueness or retrying indefinitely.

Does the preview show the final anomaly positions?

No. It shows a small normal sample without injections. The complete run produces the actual anomaly list, including final values and logical record positions.

What does a duplicate injection count of three mean?

Three additional records are rewritten to an existing nonempty key. One donor plus those records involves four rows. Final duplicate groups and involved records are reported separately from actual rewritten extras.