YourSOTAGet a sample
Home/Insights/TASK DATA & EVALUATION

TASK DATA & EVALUATION

Multi-step instruction datasets: how to evaluate real task history

A finished file shows the outcome. A useful task history also shows what someone requested, what changed and which result followed each instruction. That structure helps you test whether a model can adapt to real work.

THE PRACTICAL TAKEAWAY

Inspect the links between requests, revisions and files. Keep review evidence separate, and evaluate on projects the model has not seen during training.

What a multi-step instruction dataset contains

In plain English: a person asks for something, adds changes, and the work is updated. The technical description is multi-step task instructions with iteration over the initial instruction.

A record may include the initial human input, ordered follow-up requests, constraints that still apply and the files produced at each stage. Conversation history can help explain why a change happened. If only a revised document survives, describe it as a document revision rather than implying that a full conversation exists.

A simple record structure to inspect

  1. Human inputThe original request, its source and the expected deliverable.
  2. ChangesLater instructions in order, with any retained messages or document revisions.
  3. Result filesThe output associated with each available stage, with a version identifier.
  4. Validation filesChecks or review records referring to the relevant result version.

This is a suggested review structure, not a claim that every project contains all four parts. A coverage field should make missing history, unknown linking and absent validation visible.

Use a small change request to test the links

Illustrative example

Initial request: “Create a room layout with a desk and a clear entry path.”

Follow-up: “Move the desk closer to the window. Keep the entry path clear.”

Result: a revised drawing or model.

Validation: an available check or review note about the revised layout.

Ask whether the file follows the new instruction and preserves the earlier constraint. Check that the reference output and review note belong to the same version. A collection of plausible-looking files is weaker evidence than a traceable relationship between the request and the work.

Separate production from validation

A result file is what was made. A validation file describes how it was assessed: a test output, review comment, issue report or acceptance record. Even then, identify the scope of the check. A geometry check does not establish that every design requirement was satisfied.

Where formal validation is absent, you can agree on expert review as part of a preparation or pilot scope. Record who reviewed the example, which criteria they applied and any unresolved questions. Do not quietly turn “file exists” into “task accepted.”

Hold out whole projects, not random near-duplicates

Related instructions, exports and revisions can reveal the answer across a random file split. Our recommended starting point is to keep all records from the same project on one side of the training/evaluation boundary, then inspect related templates and duplicate content separately.

Use a baseline with the same tools and access to context. Evaluate adherence to the latest instruction, preservation of earlier constraints, output usability and the ability to identify missing information. Report coverage and sample size alongside the scores. A small pilot supports a bounded conclusion about that pilot.

Turn real work into an evaluation protocol

The SWE-bench project is an example from software engineering: repository-level tasks are evaluated against concrete resolution criteria. Architecture and construction need their own task definitions and checking methods; software benchmark scores do not transfer automatically to building workflows.

Define the instruction, allowed context, tools, output target and scoring rules before running the pilot. Some checks can be automated; others need domain review. Keep those judgments distinct, and publish the limits of what your evaluation can establish.

For procurement, use the data buyer checklist. For choosing how to use the records, read RAG vs. fine-tuning.

Useful for your team?Share on LinkedIn ↗

FROM THE GUIDE TO YOUR NEXT STEP

Looking for instructions, revisions and the work they produced?

Tell YourSOTA the task history and review evidence you need. We will discuss available project coverage and a sample that makes the links and gaps explicit.

Request a sample & data brief