HUMAN INPUT / INSTRUCTION
“Here is what I need.”
The original brief, requirements and constraints. The human input that starts the work.
HUMAN INPUT. REAL WORK. USEFUL DATA.
Human instructions. The changes. The results.
What people asked for. How they changed the task. What the team made.
Human input → Multi-step instructions → Result & validation files
License real work from architecture, construction and smart-building projects to train and test your AI.
Contract-based sourcing. Licensing for your intended AI use.

The instruction that starts the work.
Human-written briefs and requirements explain the goal, the constraints and the expected result.
See how the request changed.
Where history is retained, later messages and revised instructions show each change to the original task.
Open the work behind the instruction.
Models, drawings and documents show what was produced. Linking each result to its instruction is part of the agreed preparation.
See the checks and review notes.
Available checks, comments and acceptance records help assess the result. A finished file alone is not proof that it was validated.
History, file versions and review records are included where available. Each sample shows exactly what it contains.
01 / FOLLOW THE WHOLE TASK
A model needs more than a finished file. It needs to understand what the person wanted, what changed and how the work responded.
HUMAN INPUT / INSTRUCTION
The original brief, requirements and constraints. The human input that starts the work.
MULTI-STEP TASK INSTRUCTIONS
Follow-up messages and revised requirements, where retained. See how instructions change over several steps.
RESULT & REVIEW
Result files show what was made. Available validation files show what was checked or reviewed.
Multi-step task instructions with iteration over the initial instruction.
In plain English: a person asks for something, adds changes, and the work is updated.
A SIMPLE EXAMPLE
Illustrative workflow, not an actual conversation from the dataset.
“Create a room layout with space for a desk.”
“Keep the desk, but move it closer to the window.”
A revised drawing or model, linked to the task.
Available review notes explain what was checked.
02 / WHAT YOUR AI CAN LEARN
Select source files for your own pipeline, or agree on a package that links instructions to the work they produced.
Use human-written requirements to build tasks for domain-specific fine-tuning.
Where histories exist, test whether a model follows new instructions while keeping earlier context.
Link text with BIM geometry, CAD drawings and project documents.
Scope a pilot with reference files, review criteria and separate training and evaluation projects.
Task preparation, annotation and benchmark creation are separate scopes. Performance is measured in your pilot.
03 / REAL SOURCE MATERIAL
Technical briefs, BIM models, CAD drawings and project documents. More than 1 TB of industry data, with new projects added and targeted sourcing available to expand your coverage.
A single archive to illustrate the data.
More than a terabyte. New data added.
ONE EXAMPLE ARCHIVE
1 archive file · 10 source files inside · 4 formats
Includes human-written technical requirements, tracked edits to those requirements and BIM/CAD project files.
The documents in this example are in Russian. This example does not represent the full corpus or its language coverage.
Tracked document edits are present in this example. Conversation trails, linked file versions and completed validation reports are checked separately for each proposed package.
04 / KNOW WHAT YOU GET
Request a representative sample and a data brief. Your research, engineering and legal teams can review the same proposed package.
Project and file counts, measured volume, languages, formats and software compatibility.
Links between instructions and files. Available messages, revisions and review records. Any missing stages.
File integrity, duplicates, missing references and known limits. How training and evaluation projects are separated.
Native files or separately prepared records. Discuss extraction, linking, translation, annotation and JSONL or Parquet exports.
Source information, relevant permissions and license terms for your intended AI use.
Sample and pilot scope, delivery, acceptance criteria, pricing and an optional supply schedule.
Bring your task and evaluation criteria.
We will discuss the sample that fits.
05 / WHERE THE DATA COMES FROM
We obtain professional project data through direct agreements with data owners and agent-assisted acquisitions.
We discuss the source, relevant permissions and the license your team needs before production delivery.
Discuss licensingDiscuss the source party, contractual basis and relevant authorization, including the sourcing partner’s role where applicable.
Define training, fine-tuning, evaluation or commercial model use. Agree on duration, territories, recipients and any redistribution or derived-data terms.
Agree on review of personal information and third-party content, redaction, access, transfer and retention requirements.
Discuss warranties, remedies, liability, exclusivity and treatment of trained models after license expiry.
06 / FROM A SAMPLE TO A SUPPLY
A focused pilot, a corpus license or ongoing supply. Scope, volume, cadence and pricing are agreed individually.
Your model, target skills, languages and intended use.
Inspect the files, task coverage and data brief.
Agree on preparation and measure fit in your pipeline.
Finalize terms, then scope more projects and optional updates.
07 / A FEW USEFUL ANSWERS
The starting point is professional source material. Extraction, linking, task assembly, annotation and normalized training records are separately scoped. We agree on what will be delivered and how it will be assessed.
No single package structure applies to every source. Conversation history, instruction revisions and file versions can be included where retained and permitted. The sample brief identifies what is present and what is missing.
Result files are what the team made. Validation files are checks, review comments or acceptance records showing how the result was assessed. A finished file on its own does not prove it was checked or accepted.
The reviewed example contains Russian-language documents and IFC, RVT, DWG and DOCX files. Languages, formats and compatibility for your purchase are specified in its data brief. Translation and normalized exports can be discussed separately.
Our industry corpus contains more than 1 TB of data and continues to grow. The 55 MB archive shown here is one example. We can source additional projects through contractual agreements and rights acquisition. Measured counts and volume are provided for the proposed purchase. Expansion targets and timing are agreed before commitment.
Yes. Include your source and licensing requirements in the inquiry so we can discuss the acquisition route, supporting authorization information, restrictions and intended AI use before a production license.
Terms depend on the chosen data, rights, preparation work and supply schedule. A pilot, a corpus license and ongoing supply can be discussed separately. Exclusivity, warranties and any indemnities require explicit agreement.
We can scope records with an initial instruction, source references, changes in order, associated file versions and available review evidence. Missing stages and unknown review status should remain explicit. This is a target structure, not a claim that all source data is already converted.
08 / GUIDES FOR AI BUILDERS
Practical guides for founders and model teams: buy the right data, choose the right experiment and evaluate real task history.
DATA PROCUREMENT
A practical checklist for buying AI training data: sample quality, instruction coverage, licensing scope, preparation costs and a measurable pilot.
Read the guideMODEL STRATEGY
A founder-friendly guide to RAG vs. fine-tuning: diagnose missing knowledge or weak task behavior, choose the right data and test a focused pilot.
Read the guideTASK DATA & EVALUATION
Learn what to inspect in multi-step instruction datasets: human inputs, ordered revisions, result files, validation evidence and project-level evaluation splits.
Read the guideLET’S TALK ABOUT YOUR DATA
Tell us the task. We will discuss a sample, what it contains and the terms for using it.