Synthetic proof · Data QA sprint

Eight messy rows. Five actionable findings. One corrected artifact.

This small, reproducible example shows the output of a bounded CSV/JSON validation sprint for teams that need a clean handoff—not a vague “data looks fine” assurance.

Evidence boundary: this is self-directed synthetic data, not client work or a claim about production outcomes. No customer data, credentials, uploads, or external systems are used.
8
input rows

Order-like records created only for this sample.

5
ranked findings

Schema, duplicate, missing-field, date, and outlier checks.

6
clean output rows

Duplicate removed; invalid record quarantined.

Representative input

order_idcreated_atskuquantitystatus
ORD-10012026-07-20SKU-A2shipped
ORD-10012026-07-20SKU-A2duplicate
ORD-10032026-02-30SKU-C1pending
ORD-10042026-07-21(missing)3paid
ORD-10052026-07-21SKU-B999paid

Ranked findings and disposition

SeverityFindingEvidenceDisposition
HighDuplicate primary keyORD-1001 occurs twiceKeep first record; quarantine duplicate
HighInvalid calendar date2026-02-30Quarantine pending source correction
HighRequired value missingsku empty for ORD-1004Flag for source enrichment
MediumQuantity outlier999 exceeds sample rule of 100Retain but require review
MediumUnexpected enum valuecomplete outside approved status setMap only after owner confirmation

Smallest sellable sprint

USD 150 for one CSV or JSON export, up to 10,000 records and one agreed schema.

Includes duplicate/missing/type/date/outlier checks, a ranked report, corrected artifact where deterministic, and one review pass. No production integration.

What the buyer receives

  • Input/profile summary
  • Evidence-linked findings
  • Corrected and quarantined records
  • Validation rules and limitations
Request a free fit check

Describe the format, approximate row count, and validation goal. Do not email customer data, credentials, or the dataset itself.