Teaching machines to do real work: Part 2 of 4
How to design training exercises for AI agents
Why row counts lie
Post-training rewards a policy for doing better on the exercises it sees (part 1 of this series covers the loop (opens in a new tab)). The unit of model training/learning is the “situation”, not the “prompt”. Ten prompts about the same fictional coffee chain's store performance teach the model one thing about coffee chains.
The first job with any exercises is to define the root unit and count it. For decks it can be the scenario briefs. For spreadsheets it can be "family" (i.e., a whole build, a rebuild-one-sheet variant, a diagnose-and-repair variant, all on the same workbook). For documents the bank normally is flat, one row per root type.
Imbalance lives at the root
Earlier this year, I was handed a spreadsheet with a few thousand rows. Each row was an exercise: a brief written in a requester's voice, asking an AI agent to build or fix a slide deck. The sheet looked like a rich, varied training set. Sector column, audience column, chart types, constraints. Lots of everything. It was a few hundred scenarios.
Each scenario, a four-to-six-slide brief for a fictional company, had been sliced about ten ways: build the whole deck, build slide one, build slide two, add slide three to a deck that already has the others, fix a version of the deck that came back with faults. Every slice inherited the parent's sector, audience and vocabulary. The row count said thousands. The number of distinct situations the model would ever meet said hundreds.
The same instinct shows up wherever people generate scenarios rather than write them. Husain and Shankar's advice for simulating user queries (opens in a new tab) is to fix a handful of dimensions (task, persona, request clarity), enumerate the combinations, and make one model call per combination, because asking for everything in one shot collapses the diversity. A transplant is the file-based version: the dimensions are sector and audience, and the combination is filled by copying rather than generating.

Designing a balanced training exercise set
At row level the slide bank looked reasonably spread across sectors. At root level, three sectors held close to half of it. One commercially important sector had a single root. Customer-facing decks, the kind a sales team sends out, were only a handful.
Chart coverage had the same shape. Measured against a reference list of chart types, most were named at least once somewhere in the bank, which sounds like breadth. But the shape-built class, the diagrams a consultant draws by hand because no native chart does the job, rested on a handful of types.
None of this was visible at row level, because the derived rows copied their parent's labels and the copies looked like coverage. This post is about what I did with that, and what I would do differently. The principles are not specific to file format.
What does not work: down-weighting
The obvious cheap fix is to reduce the training weight of rows in saturated cells.
It is still worth doing, because it stops a saturated family from dominating any single batch. But it cannot fix root imbalance, because the roots are still the roots. You cannot down-weight your way out of not having sector specific scenarios. You must add the missing sector combinations. Set that expectation before you run the stage, and never zero a row for merely sitting in a crowded cell; a good exercise in a saturated sector is still a good exercise.
What does work: transplants
Writing new scenarios from scratch is slow and expensive, and every new brief needs its own starting files, its own answer key, and its own review. A few dozen twins took the thinnest sector from one root to a workable base and cut the empty cell share meaningfully, with zero new task designs.
The rules that made it defensible took longer to arrive at than the method. Here they are:
- Run a portability test before you touch anything. A scenario is portable when its visuals are not domain-bound and its vocabulary is light. A ski resort's catchment map is the task; re-skinning it into a bank is exactly the force-fit you set out to avoid. Process flows, org charts and network diagrams also travel badly.
- Match purpose to target. Performance updates port almost anywhere. Compliance briefings go to regulated sectors. Grant applications cannot go to banks.
- Audience pivots are cheaper than sector pivots. Keep the organization and the numbers, change only the framing. Same data, different reader, new cell filled. An internal negotiation-prep deck becomes the version put in front of the counterparty.
- Region pivots are their own kind. Keep numerals and date values identical, change currency words only through the dictionary, never touch a date format string, and give region twins stricter checks. Changing a region changes locale semantics: currency symbol, decimal separator, date order, fiscal-year start. For a spreadsheet task that can silently change the answer.
- Transplant before you enrich. Any enrichment applied to a source after twinning has to be mirrored to the twin by hand. Do the copying first.
- Split families. A twin shares its task layer with its source. This is the same contamination problem that rephrased benchmark samples (opens in a new tab) cause at the benchmark level, reproduced inside your own bank. If one lands in training and the other in the held-out evaluation set, the evaluation is contaminated. Add a split-family column before anyone builds a split, and size the holdout on families, not roots.
- File-dependent rows cannot be twinned cleanly. Add-a-slide, fix, rebuild and repair rows start from a file that belongs to the source organization. Create them for family completeness, park them in reserve, and list the regeneration backlog row by row. Where the bank already carries build attempts, re-running the build on the twin brief clears the backlog cheaply. Anthropic's evals guide (opens in a new tab) makes a similar point about task selection: the eval should look like the work.
What makes an exercise good
Assume a meaningful share of any bank is broken until someone has checked. When OpenAI had 93 developers annotate the original SWE-bench, they filtered out 68% of the samples (opens in a new tab): more than a third had underspecified problem statements and more than half had tests that would reject a valid solution. Having reviewed a few thousand briefs closely, the ones I would keep share five properties.
- Realistic: a professional would send this brief, in this form, to this reader.
- Clear and complete: a careful person could say what "done" means.
- Robust: there is no cheap way to satisfy the words while failing the client.
- Demanding in the ways real work is: numbers that must reconcile, objects that must be native and editable, and layout judgment.
- Useful as a tool, which is a different question from whether it is a fair test.
The weaknesses I saw most were small and fixable.
- Ambiguous time basis (calendar or fiscal quarters?).
- "Reconcile the figures" with no rule for which source wins.
- Style rules that conflict with a named template.
- Vocabulary a practitioner in that sector would never use, which is the tell of a lazy twin.
- Gameable asks like "the headline states the finding" with no check behind them, which is an invitation the model will accept.
Thinking Machines Lab audited a widely used text-to-SQL training set and found 61% of items had at least one error (opens in a new tab), including gold answers that were simply wrong. Those are two of the most-used datasets in the field. An internal bank nobody outside has seen is not going to be cleaner.
Spec-dense briefs, with exact inches and named layouts, are easy to grade and not how people write. Vague briefs are how people write and hard to grade. The answer is not to pick one. Put the same scenario at several points on a spec-density spectrum, and at the vague end add rubric items for "assumptions stated" and "reasonable default chosen."
Public benchmarks have been wrestling with the same tension. SWE-bench (opens in a new tab) got its realism by taking tasks from real repositories. OSWorld (opens in a new tab) built real computer environments and found humans at seventy-two percent and the best model at twelve.
OpenAI's BrowseComp (opens in a new tab) shows what an explicit difficulty gate looks like: a task was accepted only if existing models failed it, the author could not find the answer in five searches, and a second person could not solve it in ten minutes. A training bank does not need gates that strict, but it needs some.
What the bank was missing that no re-skin fixes
Every input file in the bank was generated in-house. Clean. No legacy macros, no external links, no merged headers, no mixed locales, no fifty-megabyte deck with embedded video. Real work is mostly inherited mess, and the model had never seen any.
The same fictional organizations also recur, so memorized structure can pass for skill. Proposals on the list: cap reuse per organization, quota task kinds by observed product demand rather than by what is easy to author, build a messy-input generator, and add non-editing exercises (explain this workbook, audit this deck, answer a question from this contract) since those are a large share of what people actually ask.
Cursor's account of building its model router (opens in a new tab) is a good model for the demand side: they built a task taxonomy from hundreds of thousands of real turns and used what users did next (moved on or corrected) as the label. That is where a task-kind quota should come from.
GDPval (opens in a new tab) commissioned experts across forty-four occupations to write tasks with real deliverables. Each of them chose realism over authoring convenience and paid for it. An internal training bank must make the same choice, and it is easy to drift the other way because nobody outside sees the bank.
----
Daljeet Saran (dsaran@motifplatforms.com) is the founder of Motif Platforms, where he builds AI and data systems for commercial teams. He currently spends most of his time on pre- and post-training data and evaluation for AI agents that do real knowledge work.
Part 1 of this series: Reward hacking explained: how one bad grader teaches AI a thousand bad lessons (opens in a new tab). Part 3: Nine ways a grader lies. Part 4: The correct-looking file.
Resources
- Anthropic. (2026, January 9). Demystifying evals for AI agents. Anthropic Engineering. (opens in a new tab)
- Husain, H., & Shankar, S. (2026, September 22). Advanced evals: How to find (and fix) hidden AI failures in your product. Lenny's Newsletter. (opens in a new tab)
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can language models resolve real-world GitHub issues? arXiv. (opens in a new tab)
- O'Keefe, C., & Volkov, Y. (2026, August 6). How Cursor Router chooses the right model for the task. Cursor. (opens in a new tab)
- OpenAI. (2024, August 13). Introducing SWE-bench Verified. (opens in a new tab)
- OpenAI. (2025a, April 10). BrowseComp: A benchmark for browsing agents. (opens in a new tab)
- OpenAI. (2025b, September 25). Measuring the performance of our models on real-world tasks [GDPval]. (opens in a new tab)
- Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., & Yu, T. (2024). OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv. (opens in a new tab)
- Yang, S., Chiang, W.-L., Zheng, L., Gonzalez, J. E., & Stoica, I. (2023). Rethinking benchmark and contamination for language models with rephrased samples. arXiv. (opens in a new tab)
- Zhu, Y., Jin, T., Choi, Y., & Kang, D. (2026, August 27). Putting task expertise into RL achieves state-of-the-art performance on text-to-SQL. Thinking Machines Lab. (opens in a new tab)