LLM Evaluation Built From Your Own Failures
Search for this and you get a taxonomy. Reference based and reference free, model evaluation against product evaluation, a catalogue of metrics with names like faithfulness and answer relevance, and a framework comparison at the bottom. The pages are written by the companies selling the evaluation platforms, they are technically correct, and they all share one assumption that quietly does not hold: that you know what you are measuring before you start.
You do not, and that is not a gap in your preparation. On a real integration the thing worth measuring is not knowable in advance, because it is determined by what your particular system does wrong, and your particular system has not run yet. Starting from a metric list means picking measurements before you have any evidence about which failures matter, and the predictable result is a dashboard of green numbers sitting above a support queue that disagrees with it.
The alternative is to build the evaluation set backwards, out of failures rather than out of metrics. It is slower to start and it is the only version that stays useful past the first month.
The benchmark problem, stated plainly
A benchmark tells you how a model performs on somebody else's distribution of inputs. That is genuinely useful information when you are choosing between models, which is a decision you make once and then mostly stop thinking about.
It tells you almost nothing about the system you built. Your prompt is not their prompt, your retrieval corpus is not their corpus, your users do not phrase things the way the benchmark authors did, and the failures that will actually cost you money are the ones specific to that combination. A model that scores well and then fails on the way your customers write dates is a system that is broken, and no public leaderboard was ever going to catch it.
So the benchmark answers the model selection question and then stops. Everything after that is your own. Model selection has its own inputs, and if that is the decision in front of you, the context window comparison and the token counter are the two things worth checking before a benchmark table, because both are properties of your inputs rather than of the model.
Where the cases come from
An evaluation set is only as good as the inputs in it, and the good inputs are already in your system. Four places produce them, and they produce them continuously, which is the property that matters.
The exception path. Every integration has one: the branch that fires when the model returns something the code could not use. Malformed output, a tool call with arguments that do not validate, a refusal, a timeout, a response that parsed but was empty. Most teams log these and then treat the log as an operational concern, something to page on and clear. It is the single richest source of evaluation cases you will ever have, because every entry is a real input that a real user sent, and it is already labelled as a failure by the code itself. Nothing needs annotating.
The corrections. Wherever a human touches the output before it goes out, the difference between what the model produced and what shipped is a labelled pair, free. This is the one source most teams have and never harvest, usually because the correction happens in a different tool from the one the logs live in.
The complaints. Support tickets about a wrong answer are cases where the failure passed every automated check you had. They are the most valuable and the hardest to get, since they require someone to connect a ticket back to a specific generation. A system with no path from a ticket to the input that caused it is a system that cannot learn from its worst failures, and that path is worth building before any metric is chosen.
The near misses. Outputs that were technically fine and that you would not want to defend. These need judgement and they are where the domain expert earns their place, since they are invisible to every automated check by definition.
Start with the exception path. It is already instrumented, it needs no annotation, and it will produce more cases in a fortnight than a brainstorm produces in a day.
The three lines that never appear on an invoice
There is a pattern in how integration work gets sold, and it is the same pattern the AI integration services post describes: the quote covers the parts that are visible in a demo and omits the parts that decide whether the thing survives. Evaluation has three of those omissions, and they are not optional extras.
Somebody owns the wrong answer. Not the infrastructure, the answer. When the model says something incorrect to a customer, there is a named person whose job includes noticing that and deciding what to do. Without that name the failures accumulate unexamined, because reading them is nobody's task. This is an organisational line rather than a technical one, which is exactly why it is the one most often skipped.
The drift check runs on a schedule. Model behaviour changes under you. Providers update models, you change a prompt, your retrieval corpus grows, your users start asking about a feature that did not exist last quarter. Any of those moves the behaviour without anything in your repository changing. A suite that runs only when someone remembers is a suite that is silently stale at the moment you most need it, so it runs on a schedule and it alerts when it moves.
The set gets curated. An evaluation set that only grows becomes a wall of cases nobody reads, and one that is never pruned keeps testing failure modes you fixed two quarters ago while missing the ones you have now. Curation is recurring work with an owner, not a task that completes.
None of the three is expensive. All three are the difference between an evaluation practice and a folder of test files.
Choosing the check, after you have the cases
Once the cases exist the measurement question becomes much easier, because you are no longer choosing a metric in the abstract. You are choosing how to detect a specific failure you have already seen, and most failures sort into three kinds.
Some are checkable by a program. Did it produce valid JSON, was the tool call well formed, did the cited document identifier exist, is the answer inside the allowed set. These are ordinary assertions, they cost nothing to run, and they should be exhausted before anything cleverer is considered. A surprising proportion of real failures live here.
Some need a model to judge them, because the property is semantic. Did the answer follow from the retrieved passage, does it contradict itself, does it answer the question that was asked. A model grading another model is a legitimate technique and it carries an obvious problem, which is that the judge has failure modes of its own. The discipline that makes it trustworthy is to check the judge against human labels on a sample before relying on it, and to recheck when the judge model changes, exactly as you would with any other measurement instrument.
Some need a person. Tone, domain correctness, whether an answer is defensible to a regulator or a customer. These are expensive, so they are spent on the near misses and on auditing the judge rather than on the cases a program can already settle.
The order is the point. Program first, model second, person third, and each layer only handles what the one before it could not.
What good looks like
A working evaluation practice is unglamorous and it has a recognisable shape. Cases arrive continuously from production rather than being written once. The suite runs on every prompt change and on a schedule besides. Failures have an owner who reads them. The set is pruned as often as it is extended. Nobody quotes a benchmark score in a status update, because everyone understands it describes a different system.
The version that fails is also recognisable. It was built in a week from a metric list, it scored well immediately, and nobody has looked at it since, because looking at it has never once told anybody something they did not know. Green numbers above a support queue that disagrees.
The difference between the two is not the tooling. It is whether the cases came from a catalogue or from your own failures.