BROWSER-BASED TABLE TOOL

Random Sample CSV Rows

Create a random CSV sample by count, percentage, or group with a reproducible seed. Work locally in your browser and review results before export.

File contents are processed locally in your browser. Refreshing or closing clears the session. The website still requests static resources.

1Import data—2Configure—3Review—4Export

Drop your table here

Drag and drop a CSV, TSV, or delimited text file

UTF-8 · Up to 10 MB per file

Start with your data

File content is processed in your browser, without uploading it.
Kept in page memory · Continue across tools · Refreshing or closing clears the sessionCSV / TSV · UTF-8
Sample library · Normal / Boundary / Error

Clear this session?

Imported data, steps, and results will be released. Export any needed copies first. Original files are unaffected.

Draw a reproducible subset of complete records

Random Sample CSV Rows selects whole source records for inspection, debugging, or a demonstration. You can request a fixed count, a percentage of the table, a fixed count within each group, or a total distributed proportionally across groups. The default never selects the same source record twice. Every chosen record keeps all its columns, original text, and source position, so the relationship between an identifier and its other fields remains intact.

A random sample is not automatically representative of every characteristic of the source population. A small category can be absent from an ordinary random draw, and a missing or biased source cannot be repaired by changing a seed. Proportional allocation follows the selected grouping fields only. Use the tool to create a transparent operational sample and inspect the plan; it does not certify a statistical research design or provide confidence intervals.

Confirm the table and its order

Choose files or Paste a table, then inspect Confirm file parsing. Select Parse and preview to verify the file delimiter, headers, quoted cells, and record boundaries. Confirm import checks the full file before sampling. Input is UTF-8 CSV, TSV, or delimited text, and all values remain text. A quoted multiline note belongs to one record and is selected together with that record rather than being treated as several independent lines.

The current record order is part of reproducibility. If you sort or filter the table before sampling, you have changed the input to the random procedure even when the set of visible values looks similar. Apply those preparation steps explicitly and retain their rules. A sample drawn from a workflow result uses that current result. It does not quietly return to the originally imported file or incorporate unrelated files elsewhere in the session.

Choose Sampling mode before entering the amount. Whole table draws from all available records as one pool. Fixed count per group treats each distinct group as its own pool and requires Count as the amount type. Proportional group allocation uses a requested total for the whole table and divides that total across groups according to their available record counts. Choose one or more Group columns for either grouped mode.

Counts and percentages have precise meanings

Count means a whole number of records. For Whole table, a request for ten produces ten distinct source records if at least ten are available. A request larger than the pool is blocked. The tool does not silently switch to sampling with replacement or reduce the request to whatever happened to be available. Zero is allowed and produces a valid empty sample, which can be useful for checking a downstream import path.

Percentage of full table uses a value from zero through one hundred and rounds the resulting count down. For example, ten percent of twenty-five records produces two records. The percentage always describes the full input for the total-count modes. It is not independently rounded inside every group. This distinction prevents several separate rounding decisions from accidentally producing more or fewer sampled records than the requested total.

Fixed count per group requests the same integer from every group. If one group contains fewer records, the default reports the shortage and blocks the run. You can explicitly choose Take every available record for undersized groups. That decision changes the final total, so review the plan rather than multiplying the number of groups by the requested count and assuming the product was achieved. The shortage option never adds duplicate source records to reach a quota.

Proportional allocation preserves the requested total

Proportional group allocation first computes each group's exact share of the requested total. It assigns the integer part of each share, then distributes the remaining records to groups with the largest fractional remainders. Ties use the groups' first-appearance order in the current input. This makes the rule deterministic and keeps the sum of quotas exactly equal to the requested total without making a separate rounding choice for each group.

Suppose three groups have thirty-four, thirty-three, and thirty-three records and the requested total is ten. Their shares are 3.4, 3.3, and 3.3. The initial quotas are three each, leaving one record. The first group receives that remainder, producing quotas of four, three, and three. The actual draw then selects distinct records inside each pool according to those quotas. The plan shows available, planned, and actual counts for every group.

Grouping uses a structured combination of the selected field values. It does not join key components with an ambiguous punctuation character. Empty group values form groups rather than being dropped, and missing cells remain distinguishable from present empty strings in the internal representation. Whitespace and literal strings such as NULL are not normalized away. Clean or normalize a grouping field through a separate explicit step if the source convention requires it.

Seed, algorithm, and input identity belong together

Random seed is an explicit text value. The same ordered input, configuration, seed, and algorithm version produce the same selection. Changing any of them can change the result. Save rules retains the seed and options, while Settings and reproducibility records the algorithm version, input record count, input-order convention, and a compact input signature. The signature is a diagnostic checksum, not a cryptographic proof that two files are identical.

The algorithm version identifies the seeded random generator and shuffle procedure used by this implementation. Keep the version information with a sample intended for a reproducible bug report. A future implementation can intentionally use a different procedure, and a seed alone would not explain why its selections differ. The source names and row identities remain in local reports; they are not put into a URL or shared automatically with another service.

Output order offers Source order and Draw order. Source order sorts the selected indices back into their original relative order, which often makes manual checking easier. Draw order keeps the order produced by the draws; in grouped modes, groups are visited in their first-appearance order. Changing output order changes the arrangement of the same selected records, not the quotas. It does not create an additional draw or allow replacement.

Inspect a complete sample and its provenance

For a hundred source records, choose Whole table, Count, amount ten, and a fixed seed such as review-1. Run and review should produce ten records with ten distinct source identities. Export the sample, run again without changing anything, and compare the output. The same records and order should recur. If the source contains two identical business rows, both can be selected because they occupy different source positions. This is not a promise of unique business content.

Preview sample and estimate shows a limited view of the planned selection. Full export still requires Run and review, which establishes the completed result in the workspace. Group sampling plan makes quota differences visible, and Source mapping records each selected record's origin. If the draw does not include a record you expected, first check whether it was part of the current input and whether its group had a quota; a seed is not a guarantee that any specific item will appear.

Export a sample without bundling the entire source

Export new copy lets you choose the sample, group plan, source mapping, or reproducibility settings. The ZIP option bundles these result artifacts and a manifest; it does not automatically include the entire original dataset. CSV and TSV serialization is verified by parsing the output again before download. Formula-prefix protection is explicit and reports any modified cells. Protected text may differ from raw source text, so retain the export settings with a sample used to reproduce a parser issue.

Source identifiers such as 00123 remain textual values in the file, although another program may infer numbers when it opens the download. Import identifier columns as text when needed. Current input, output, cell, field, report, and time limits still apply. Cancel task stops a pending worker, and changed rules make the old result stale. Apply to workflow passes the reviewed subset to another tool. Refreshing or closing clears the session; download the sample and its rule file before leaving if later reproduction matters.

Frequently asked questions

Can I randomly select a fixed number of rows without selecting the same row twice?

Yes. Request a count within the available pool; sampling does not select the same source record twice. Identical-looking rows can still appear if they are separate records in the source. Keep the input order, settings, and seed if you need to reproduce the same draw later.

How do I get the same sample next time?

Keep the input content and order, seed, configuration, and algorithm version. Retain the reproducibility report and any earlier sorting or filtering steps; the seed alone is not sufficient.

What happens if a group cannot supply its quota?

Fixed count per group blocks by default and reports availability. Explicitly choose the option to take all available records if a smaller sample from that group is acceptable.

Why can the sample contain identical rows?

Distinct source records can have identical contents. Sampling prevents selecting the same source identity twice, not selecting two different records that happen to contain the same business values.

Does proportional sampling guarantee representativeness?

No. It follows the chosen grouping distribution and exact quota rule. Other characteristics, source bias, and a small requested total can still make a sample unsuitable for a particular statistical conclusion.