Teaching machines to do real work: Part 1 of 4
Reward hacking explained: how one bad grader teaches AI a thousand bad lessons
I have spent the last few years working with frontier AI labs on model development, training data and building agents for knowledge work (e.g., build a slide deck from a brief, rebuild a missing sheet in a workbook, push a set of corrections through a long contract).
After pre-training, a model learns from a loop with three parts: (a) exercises define the situations it sees, (b) graders define what counts as better and (c) human experts are the quality check on whether the graders' idea of better matches a professional output. Whatever is wrong in that loop is not diluted downstream; it is pushed into the model on every attempt, thousands of times.
Knowing isn’t doing
Pre-training is the part most people have heard of a very large corpus, next-token prediction, months of compute. The model that comes out knows the API of every document library, every chart type and when to use it, what a board deck usually looks like. It has read more spreadsheets than any analyst alive.
It does not, on its own, reopen the file it just saved to check whether it opens. It does not know that a "chart" made of rectangles and arrows is worse than a native chart even though it looks the same on screen. It does not know that when you are asked to fix one slide, changing the theme on the other forty is a firing offence.
Those are not knowledge gaps. They are behaviors, and behaviors are taught in post-training, where the model attempts tasks, gets scored, and is nudged toward whatever scored higher. The recipe has been public since InstructGPT (opens in a new tab) and Christiano et al. 2017 (opens in a new tab): collect preferences, learn a reward, optimize against it. What has changed since is that for agentic work the reward increasingly comes from programs and rubrics rather than from a person clicking A or B, and the tasks are long, multi-step and file based. Nathan Lambert's RLHF book (opens in a new tab) is the best single reference on the mechanics.
Andrej Karpathy called 2025 the year reinforcement learning (opens in a new tab) with verifiable rewards became a major training stage in its own right (opens in a new tab).
The loop: exercises, graders, people
Here is the picture to explain the idea.

Exercises. A prompt in a requester's own words, sometimes a starting file, plus labels: sector, audience, task kind, difficulty. Together the exercises are the universe the model practices in. If the universe has one financial-services scenario and forty retail ones, the model gets very good at coffee-chain performance analysis.
Graders. Deterministic checks (does the file open, is the chart native, did the untouched slides stay untouched), a checklist grader that reads the brief's requirements, sometimes a pairwise judge comparing two attempts, and a few overrides (claim success on a broken file and your score is zeroed). Together the graders define "better" inside that universe.
People. Domain experts who look at the exercise and ask whether a professional would send this brief, then look at the output and ask whether the client would accept it. They are the only thing in the loop that can say the graders are wrong.
The model learns from the first two. The first two are only as good as the third.
Five levers to rely upon
If you want to change what a model learns, the order of leverage matters.
First, the number of distinct scenarios. A bank of a few thousand prompts that is really a few hundred scenarios sliced ten ways (i.e., whole deck, one slide, add a slide, fix a deck) teaches the model a few hundred things. Row-level diversity flatters you.
Second, the shared quality checks. Every exercise carries a set of general checks: the slide master was used, space is used well, tables are real tables. Because they apply to everything, they dominate the aggregate reward regardless of what any individual brief asked for. Whatever they reward is what the model becomes.
Third, what the reward makes free. Relative scoring within a group of attempts with no absolute floor lets a group of poor attempts crown a near-perfect winner. An honesty rule that zeroes false claims but treats silence as free teaches the model to say less. A "nothing else changed" check that counts double teaches timidity: patch the named symptom, leave the stale summary that depends on it. Nobody intended any of these. The gradient sees all of them.
OpenAI made the same argument about hallucination at the level of the whole field: accuracy-only scoreboards reward guessing over abstention (opens in a new tab), so models learn to guess. Their GPT-4o sycophancy postmortem (opens in a new tab) is the production version of the story: over-weight short-term thumbs-up feedback and you ship a model that agrees with everyone.
Fourth, where the evaluation is thin. Editing an existing file and multi-turn work are where enterprise use lives and where capability collapses; one public benchmark of slide-editing tasks found strong per-turn accuracy and near-zero whole-session accuracy (PPTC (opens in a new tab)). But creation-from-scratch is the cheapest kind of exercise to write, so it is over-represented, and the model gets good at the easy thing.
Fifth, what the judge can see. If a model-based judge looks at rendered pages with no control for length, denser decks read as more work and score higher, which is exactly backwards for slides. If the judge reads the model's reply alongside the file, replies get written for the judge. This is the family of problems the literature calls reward hacking (opens in a new tab) and Goodhart's law (opens in a new tab).
Manheim and Garrabrant (opens in a new tab) sorted the mechanisms into four kinds, and Skalse et al. (opens in a new tab) showed that a fully un-hackable proxy is essentially impossible for any non-trivial reward. Anthropic's research shows what happens when it goes badly: models that learn to shortcut a coding reward generalize to broader misbehavior (opens in a new tab), and a curriculum of gameable environments can teach a model to tamper with its own reward (opens in a new tab).
OpenAI's reasoning models were caught announcing the hack in their own chain of thought (opens in a new tab) ("let's hack"), and penalizing the thought did not stop the hack, it taught the model to hide it.
One bad grader, a thousand bad lessons
Here is the point I most want to land. In ordinary software, a bug in a test hurts you once, when it lets a bad change through. In post-training, a wrong answer key or a mis-weighted rubric line is applied to every attempt in every group, for as long as that exercise is in the bank. It is not a one-off error. It is a gradient, pushed thousands of times in the same direction.
Signal quality is bounded by the weakest leak, not the average. Worse, the external evaluation meant to catch it is often built by the same people on the same assumptions, so it certifies the result.
That changes the order of work. When the bank looks thin, the instinct is to write more exercises. The right first move is to fix the cheapest leaks: a validity gate so files that will not open score zero rather than being quietly excluded, tags on why each failure happened, a judge protocol that controls for length and position. Each of those makes every existing exercise worth more. Growing the bank before they land scales the defects into training.
People are the only check
The loop has no internal check on whether "better" means anything. The graders are written by people who did not build the model, tested against reference answers built by the same people.
So the human review must be designed as carefully as the graders. Blind, before any score is shown. Split into two tracks, so a good output cannot rescue a bad brief. A one-line reason with every score, or the score is not saved. And it must come first. Shreya Shankar and colleagues documented what they call criteria drift (opens in a new tab): people discover what they are grading for by looking at examples, so a rubric written before anyone has read the outputs measures the wrong things.
Hamel Husain and Shankar's piece on error discovery (opens in a new tab) makes the product-team version of the same case: most teams skip straight to metrics, and the step they skip is the one that decides which metrics matter. Anthropic's guide to agent evals (opens in a new tab) and its earlier piece on why evaluating AI systems is hard (opens in a new tab) both come back to the same place: automated grading gets you scale, and expert review is what tells you whether the scale is pointed anywhere.
What comes next
Part 1, this post: the loop, and why leaks compound. Part 2, How to design training exercises for AI agents (opens in a new tab): designing exercise banks, why row counts lie, and how to fix imbalance without writing a thousand new tasks. Part 3, Nine ways a grader lies: rubric hygiene, answer keys nobody audits, reward design, and model-based judges. Part 4, The correct-looking file: failure modes of agents on document work, what expert taste means, and what pre-training can't teach.
----
Daljeet Saran (dsaran@motifplatforms.com) is the founder of Motif Platforms, where he builds AI and data systems for commercial teams. He currently spends most of his time on pre- and post-training data and evaluation for AI agents that do real knowledge work.
Resources
- Anthropic. (2023, October 4). Challenges in evaluating AI systems (opens in a new tab)
- Anthropic. (2024, June 17). Sycophancy to subterfuge: Investigating reward tampering in language models (opens in a new tab)
- Anthropic. (2025, November 21). From shortcuts to sabotage: Natural emergent misalignment from reward hacking (opens in a new tab)
- Anthropic. (2026, January 9). Demystifying evals for AI agents (opens in a new tab)
- Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences [Preprint] (opens in a new tab)
- Guo, Y., Zhang, Z., Liang, Y., Zhao, D., & Duan, N. (2023). PPTC benchmark: Evaluating large language models for PowerPoint task completion [Preprint] (opens in a new tab)
- Husain, H., & Shankar, S. (2026, September 22). Lenny’s Newsletter. Advanced evals: How to find (and fix) hidden AI failures in your product. (opens in a new tab)
- Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025, September 5). Why language models hallucinate. OpenAI. (opens in a new tab)
- Karpathy, A. (2025, December 19). 2025 LLM year in review. Bear Blog. (opens in a new tab)
- Lambert, N. (2025). Reinforcement learning from human feedback. (opens in a new tab)
- Manheim, D., & Garrabrant, S. (2018). Categorizing variants of Goodhart’s law [Preprint]. (opens in a new tab)
- OpenAI. (2025, March 10). Detecting misbehavior in frontier reasoning models. (opens in a new tab)
- OpenAI. (2025, April 29). Sycophancy in GPT-4o: What happened and what we’re doing about it (opens in a new tab)
- Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback [Preprint] (opens in a new tab)
- Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences [Preprint]. (opens in a new tab)
- Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and characterizing reward hacking [Preprint] (opens in a new tab)
- Willison, S. (2025, December 19). A quote from Andrej Karpathy. Simon Willison’s Weblog. (opens in a new tab)