YourSOTAGet a sample
Home/Insights/MULTI-TURN TASKS & REVISIONS

MULTI-TURN TASKS & REVISIONS

Your AI follows the first instruction. Then the customer changes it.

The demo works until the customer says, “Change this one thing.” The assistant makes the change and breaks something that was already correct. For a product built around ongoing work, handling the second request can matter as much as handling the first.

THE PRACTICAL TAKEAWAY

Train and test the whole change: what was requested, what changed, which constraints remain active and whether the new result satisfies them. Check application memory before blaming the model.

Test a change, not just a fresh answer

Use an illustrative layout task. The first instruction asks for a seating arrangement with an unobstructed entry route. The next instruction changes the position of a desk. A successful revision moves the desk while retaining the active access requirement.

A second-turn answer that follows only “move the desk” is incomplete. A model that refuses every change also fails. The useful task is controlled modification: preserve what remains valid, update what has changed and ask when the instructions conflict.

This is why isolated prompt–answer examples can leave a gap. They may show how to produce an output without showing how to revise an existing result under a continuing set of requirements.

Did the model forget, or did the application omit the context?

Inspect the actual request sent to the model at the failing turn. Does it contain the relevant initial instruction, the current state and the latest change? Check truncation, conversation summaries and which document version your retrieval step supplied.

If the model never received the earlier condition, more instruction-tuning examples cannot restore that missing fact at inference time. Test context assembly or an explicit current-requirements record before changing training.

If the condition is present but ignored, examine model behavior and representative revision examples. Separate memory delivery from constraint handling so your experiment can identify which change helped.

Keep, change or clarify: give every requirement a status

A conversation is not a list of instructions that all remain valid forever. “Use six seats” followed by “Actually, make that four” changes the earlier count. “Move the desk, keeping everything else” preserves the other requirements.

Represent what remains active at each step. A simple task record can mark a condition as kept, replaced or unresolved and link that status to the relevant user input. This makes both annotation and evaluation easier to inspect.

Include cases where the request lacks enough information or conflicts with an active condition. The correct next action may be a clarification, rather than a guessed final output. Specify that behavior through the product’s rules and reviewed examples.

What should multi-turn training records contain?

For the skill you want to teach, request the initial human input, available starting files, ordered revisions, each corresponding result and any relevant review evidence. Connect those elements through stable identifiers.

Distinguish the user’s actual words from a summary written later. A reconstructed instruction can be useful, but label it as reconstructed. Do not turn an inferred reason for a design change into a claim that the client explicitly requested it.

The Multi-IF benchmark evaluates multi-turn instruction following and reported increasing failures across turns for its tested models. It motivates testing revisions explicitly; its results do not predict performance on your architecture or construction workflow.

Score the change and the requirements it must preserve

Checks for a successful revision
CheckWhat counts as success?
Requested changeThe new instruction is implemented.
Retained conditionsEarlier active requirements still hold.
Replaced conditionsSuperseded requirements do not reappear.
Missing informationThe assistant asks or escalates as the workflow requires.
Result versionThe delivered artifact corresponds to the current step.

Apply these checks to the output or artifact, not only to the assistant’s explanation. Saying that an entry route remains clear is different from producing a layout where it is clear.

Report failures by turn and constraint type. An overall conversation score can conceal the precise point where the system starts breaking prior work. Track reviewer corrections as well.

Real project history helps only when the links survive

A folder with several drawing versions is not automatically a multi-step instruction dataset. You need to know which change each version reflects and what evidence supports the mapping. Dates alone may not establish the full sequence.

Ask a supplier for a small traceable chain before ordering a larger preparation job. Inspect missing steps, ambiguous mappings and unavailable approvals. Exclude uncertain targets or document uncertainty according to the pilot’s requirements.

YourSOTA can discuss available project revisions and result files as source material. Whether a delivery includes complete correspondence, prepared task records or independent review is established for the proposed package. See our task-history guide for the source checks.

Do you need full conversations to improve revision handling?

You need enough reliable context to teach and evaluate the change. That can include full conversations, selected task steps or a reviewed current-state representation. Choose based on how your product operates; do not discard active constraints just to make records shorter.

Useful for your team?Share on LinkedIn ↗

FROM THE GUIDE TO YOUR NEXT STEP

Does your assistant break earlier requirements after a revision?

Share the type of change your product struggles with. YourSOTA can discuss source histories, linked revisions and corresponding result files where available, and define what needs separate task preparation.

Request a sample & data brief