Define the behavior before judging the examples
‘Good data’ is a useful phrase only when you can explain the task. A requirements assistant may need to extract constraints from a brief. A design assistant may need to respond to a revision while preserving earlier requirements. Those tasks need different examples and different checks.
Write a success rubric before reviewing a sample. For extraction, ask whether required fields are present and supported by the source. For revision tasks, ask whether the new request is followed and earlier constraints remain intact. Specify acceptable abstention when the material does not support an answer.
Include common failure cases from your product. A dataset that addresses your repeated errors may be more useful than a larger collection that only repeats tasks the baseline already handles.
Check the instruction, context and output as one record
For each prepared example, inspect the human input, the context provided to the model and the target output together. Confirm the output belongs to the right project and file version. If the instruction references an attachment, make sure that attachment is included or clearly marked as unavailable.
Keep source identifiers alongside prepared text. They let a reviewer trace a claim back to the document instead of relying on an extraction pipeline’s confidence. Preserve units, document language and any conversion choices that change how the example should be interpreted.
Source files are useful starting material, but they do not automatically establish a supervised training pair. Extraction, pairing, annotation and review should be described as preparation steps with their own deliverables.
Distinguish a finished file from a reviewed answer
A final-looking filename is not enough to decide that an output should become a training target. Ask whether the file was reviewed, which requirements were checked and what evidence supports that status. Preserve corrections instead of silently replacing the original record.
Use clear labels such as ‘source artifact’, ‘prepared example’, ‘human-reviewed target’ and ‘review evidence unavailable’. An output can still be useful for retrieval or project understanding when it is unsuitable as an approved answer.
For multi-step examples, check the order of instructions and identify the relevant result after each change. Missing intermediate stages should be visible. The task-history guide shows how to inspect those links without assuming every project includes a full conversation.
Measure coverage and defects separately
Record what the sample covers: tasks, project types, languages, disciplines, formats and available histories. Then record defects such as broken files, missing dependencies, contradictory targets or incomplete instructions. Coverage gaps and file defects call for different decisions.
| Check | Evidence to record |
|---|---|
| Usability | Can the file be opened or parsed in the agreed environment? |
| Context | Are referenced instructions, attachments and units available? |
| Pairing | Does the output match the correct task and revision? |
| Review status | What supports using the output as a target? |
| Coverage | Which product failure cases are represented? |
| Preparation | Which steps remain before the records enter your pipeline? |
Report both counts and denominators. ‘Eight missing attachments in forty inspected records’ is easier to assess than an unexplained quality percentage.
Review duplicates and split by related projects
Exact duplicate checks are a useful starting point. Also look for exports of the same model, renamed copies, repeated templates and adjacent document revisions. Decide whether a repetition adds useful behavior or merely increases the apparent size of the dataset.
A random file split can place a project brief in training and a closely related drawing in evaluation. For a project-based dataset, group related material before assigning splits. Keep the group identifiers and split decisions in a reproducible manifest.
Project separation reduces one obvious leakage route; it does not prove a clean evaluation. Review shared templates and reused content across projects too. Your evaluation should test the skill you want to learn, with enough independent tasks to expose meaningful errors.
Turn the review into a preparation and acceptance plan
After reviewing the sample, decide what you can use immediately, what needs preparation and what should be excluded. Attach the schema, review rubric and known limitations to the delivery specification. Agree on how issues are reported, corrected and checked again.
The Datasheets for Datasets paper proposes documentation of a dataset’s motivation, composition, collection and recommended uses. That is a useful starting point for asking a supplier concrete questions; your product-specific checks still need to be defined.
Finish with a measurable pilot. Track accepted examples, preparation effort and results on held-out tasks. A well-documented sample review helps you scope the next delivery, while performance in your pipeline determines whether to scale.