Headhunted for a second frontier AI engagement, this time grading the work products that AI agents actually produce.
The first generation of model evaluation asked a simple question: is the answer right? That was tractable. A fact is a fact, a sum is a sum, and a wrong answer announces itself.
That question is now mostly obsolete. Frontier models no longer return answers. They return deliverables: a twelve-page memo, a board deck, a financial model with linked assumptions across four tabs. Grading a deliverable is a different discipline entirely. A document can be factually flawless and still be unusable, because the structure buries the recommendation, the formatting signals carelessness, or the whole thing reads like competent filler that nobody would sign their name to.
Judging that gap is not a technical skill. It is professional judgement, and it is exactly what frontier labs are now buying.
David Nagtzaam has been selected as an Evaluator on a large-scale evaluation programme run for a leading AI lab. The invitation came through direct headhunting rather than application, following earlier work on Mercor's Argentum and APEX frameworks and on OpenAI's reasoning initiative.
The specifics of the engagement are confidential. What can be said is that the work sits at the sharp end of a shift the industry has been building toward for two years. Mercor, which now works with frontier developers including OpenAI and Anthropic, has built its business on the bet that expert-driven evaluation defines the next phase of model development, and pays contracted specialists to test and score model performance. Its APEX benchmarks test models on the professional tasks knowledge workers actually do, across domains like investment banking, consulting and corporate law.
From Answers to Artifacts
Mercor describes the Artifacts Evaluator role publicly as reviewing, assessing and giving structured feedback on domain-specific documents, ensuring quality, accuracy and relevance. That flat description hides how hard the work is.
Scoring a model-produced document against a rubric means holding several judgements at once. Is the reasoning sound? Is the structure serving the reader or the writer? Does the formatting hold up under scrutiny? And the one that separates useful evaluators from the rest: would a professional in this field actually send this, or would they quietly rewrite it first?
Getting that consistently right, across hundreds of artifacts, at a standard reproducible enough to train on, is the bottleneck in frontier AI right now. Compute is abundant. Expert judgement is not.
Why This Matters for DECODE Clients
Evaluation work is the closest vantage point available on where these systems break. You see the failure modes before they are documented, and you see which capabilities are real versus demoed.
"Most organisations are still evaluating AI output the way they evaluate a search result," David notes. "Is it correct. But once a model is producing the deliverable, correctness is the floor, not the bar. The question becomes whether it holds up in front of a client, a regulator, or a board. Very few teams have built the judgement to answer that."
That perspective feeds directly into how DECODE advises on AI adoption: not which model scores highest on a public leaderboard, but which one produces work your organisation can actually put its name on.
Deploying AI on work that carries real consequences? Get in touch to discuss where these systems hold up and where they don't.