YourSOTAGet a sample
Home/Insights/DATA QUALITY & PREPARATION

DATA QUALITY & PREPARATION

Fine-tuning dataset quality: a practical checklist for real project data

A dataset can contain readable files and still teach the wrong behavior. Before fine-tuning, check whether each example represents a task you care about, includes the right context and points to an output worth learning from.

THE PRACTICAL TAKEAWAY

Review examples at the task level. Measure useful coverage, explain missing context, and keep related project versions together when separating training from evaluation.

Define the behavior before judging the examples

‘Good data’ is a useful phrase only when you can explain the task. A requirements assistant may need to extract constraints from a brief. A design assistant may need to respond to a revision while preserving earlier requirements. Those tasks need different examples and different checks.

Write a success rubric before reviewing a sample. For extraction, ask whether required fields are present and supported by the source. For revision tasks, ask whether the new request is followed and earlier constraints remain intact. Specify acceptable abstention when the material does not support an answer.

Include common failure cases from your product. A dataset that addresses your repeated errors may be more useful than a larger collection that only repeats tasks the baseline already handles.

Check the instruction, context and output as one record

For each prepared example, inspect the human input, the context provided to the model and the target output together. Confirm the output belongs to the right project and file version. If the instruction references an attachment, make sure that attachment is included or clearly marked as unavailable.

Keep source identifiers alongside prepared text. They let a reviewer trace a claim back to the document instead of relying on an extraction pipeline’s confidence. Preserve units, document language and any conversion choices that change how the example should be interpreted.

Source files are useful starting material, but they do not automatically establish a supervised training pair. Extraction, pairing, annotation and review should be described as preparation steps with their own deliverables.

Distinguish a finished file from a reviewed answer

A final-looking filename is not enough to decide that an output should become a training target. Ask whether the file was reviewed, which requirements were checked and what evidence supports that status. Preserve corrections instead of silently replacing the original record.

Use clear labels such as ‘source artifact’, ‘prepared example’, ‘human-reviewed target’ and ‘review evidence unavailable’. An output can still be useful for retrieval or project understanding when it is unsuitable as an approved answer.

For multi-step examples, check the order of instructions and identify the relevant result after each change. Missing intermediate stages should be visible. The task-history guide shows how to inspect those links without assuming every project includes a full conversation.

Measure coverage and defects separately

Record what the sample covers: tasks, project types, languages, disciplines, formats and available histories. Then record defects such as broken files, missing dependencies, contradictory targets or incomplete instructions. Coverage gaps and file defects call for different decisions.

A small sample-review scorecard
CheckEvidence to record
UsabilityCan the file be opened or parsed in the agreed environment?
ContextAre referenced instructions, attachments and units available?
PairingDoes the output match the correct task and revision?
Review statusWhat supports using the output as a target?
CoverageWhich product failure cases are represented?
PreparationWhich steps remain before the records enter your pipeline?

Report both counts and denominators. ‘Eight missing attachments in forty inspected records’ is easier to assess than an unexplained quality percentage.

Review duplicates and split by related projects

Exact duplicate checks are a useful starting point. Also look for exports of the same model, renamed copies, repeated templates and adjacent document revisions. Decide whether a repetition adds useful behavior or merely increases the apparent size of the dataset.

A random file split can place a project brief in training and a closely related drawing in evaluation. For a project-based dataset, group related material before assigning splits. Keep the group identifiers and split decisions in a reproducible manifest.

Project separation reduces one obvious leakage route; it does not prove a clean evaluation. Review shared templates and reused content across projects too. Your evaluation should test the skill you want to learn, with enough independent tasks to expose meaningful errors.

Turn the review into a preparation and acceptance plan

After reviewing the sample, decide what you can use immediately, what needs preparation and what should be excluded. Attach the schema, review rubric and known limitations to the delivery specification. Agree on how issues are reported, corrected and checked again.

The Datasheets for Datasets paper proposes documentation of a dataset’s motivation, composition, collection and recommended uses. That is a useful starting point for asking a supplier concrete questions; your product-specific checks still need to be defined.

Finish with a measurable pilot. Track accepted examples, preparation effort and results on held-out tasks. A well-documented sample review helps you scope the next delivery, while performance in your pipeline determines whether to scale.

Useful for your team?Share on LinkedIn ↗

FROM THE GUIDE TO YOUR NEXT STEP

Want to inspect data quality before a larger purchase?

Request a YourSOTA sample around your target behavior. Discuss source coverage, preparation scope and the checks your team needs to run.

Request a sample & data brief