Start with the mistake the customer has to fix
Write the failure as an observable event. “The model is weak at architecture” is difficult to test. “After the client changes the seating layout, the assistant removes the earlier access requirement” gives you an input, a failed behavior and a check.
Collect permitted examples of what the user supplied, what your system actually saw, what it returned and what a reviewer changed. Include the deployed prompt, retrieved material and relevant tool results. A transcript of the user’s request alone may omit the cause.
Group repeated mistakes before choosing the next intervention. One fluent but incorrect answer can come from several different failures. Your training budget should follow the diagnosis.
Is the problem knowledge, behavior or execution?
Use the following checks as hypotheses, not automatic prescriptions. A product can have more than one problem at once.
| Observed failure | Check first |
|---|---|
| Uses an outdated requirement | Was the current source available in the model’s context? |
| Ignores a requirement it received | Do training examples teach and evaluation checks measure that behavior? |
| Calls the right tool but returns the wrong file | Inspect version selection, tool output and application state. |
| Always answers when information is missing | Look for examples and product rules for clarification or escalation. |
If supplying the correct document fixes the mistake, test retrieval and context assembly first. If the information is present but the assistant repeatedly mishandles it, investigate behavior and task examples. A tool or file-selection bug needs an execution fix.
Check what your examples reward
A beautifully written target can teach the wrong behavior. Inspect whether it uses facts available at that point, selects the correct file version, preserves active constraints and stops when the task cannot be completed responsibly.
Check for contradictions: two records may respond differently to the same condition without explaining why. If a later approved drawing becomes the target for an earlier request, the example can quietly teach the model to use information it would not have at deployment.
The LIMA study showed strong instruction-tuning results with a small curated set in its experimental setting. It supports taking curation seriously; it does not establish a universal sample count or prove that a small dataset will solve your product’s task.
Compare training inputs with production inputs
Your model may train on clean, complete requests and serve messy fragments. It may learn from full conversations while your application supplies only the last message. It may expect one document format while customers upload another.
Compare a training record and a real inference request side by side. Inspect instructions, role boundaries, context length, available files and the representation of tool results. Check that the intended answer tokens are included in the training objective and that the deployed model or adapter is the version you tested.
These checks should precede speculative parameter changes. A deployment mismatch can make a useful training change look ineffective.
Run an experiment that can explain its result
Freeze a small, relevant evaluation set before editing the training data. Keep related project versions together so an evaluation drawing is not a near-duplicate of a training drawing. Include ordinary tasks and the costly failures you identified.
Compare the current system, a prompt or retrieval correction where relevant, and the tuned candidate under the same operating conditions. Change one major factor at a time. Record successes, serious errors, reviewer effort and performance on tasks that previously worked.
If results are inconsistent, inspect the actual examples and repeat the uncertain comparison where justified. “Better on average” should not hide a new failure in a business-critical task. Use our pilot guide to turn the experiment into a buying decision.
Ask for the missing task evidence, not another archive
Turn the diagnosed gap into a data request. Specify the human input, available context, expected action or result, review status and variation needed. For revision handling, request ordered changes and the corresponding outputs rather than unrelated finished files.
Ask the supplier which components exist in the source and which require reconstruction or annotation. A professional archive can contain useful material without containing a complete, ready-to-train conversation. Missing context should stay marked as missing.
At YourSOTA, sample scope starts with the task you want to test. The source package, preparation work and licensing scope are discussed explicitly; access to a growing corpus is not a promise that every failure has a ready-made example.
Should you add more data when fine-tuning does not work?
Only when the diagnosis points to missing or insufficient task coverage. More examples can help a genuine coverage gap; repeating the same incorrect targets, incomplete context or deployment mismatch does not address its cause. Decide what evidence the next batch must add before increasing volume.