Machine Learning Consulting and the LLM Default
Search for machine learning consulting and every result is a directory. Ranked lists of the best firms, refreshed with a new month in the title, each one a grid of logos sorted by size and industry. None of them describes the work, and none of them raises the question that decides most of the budget: whether your problem needs a model trained on your own history, or a prompt against somebody else's.
That question used to answer itself. A prediction problem was a prediction problem and the only route to it was to build a model. The default has since moved, and it has moved all the way: a great many projects now begin with an API call to a language model, and the decision to do that is rarely made explicitly. It is simply where the conversation starts.
Sometimes that is right and the alternative would be absurd. Often it is the expensive answer to a problem that was solved more cheaply, more predictably and more defensibly long before anyone had heard of a transformer. Telling those apart is most of what a competent engagement does in the first fortnight, and it is worth being able to do it yourself before you commission anything.
What still makes something a classical ML problem
The shape is recognisable and it has very little to do with the subject matter. Four properties travel together, and when all four are present you are looking at a problem that wants a trained model.
The decision repeats, at volume, in an automated path. Scoring every transaction, ranking every search result, forecasting demand per item per store per week. The repetition is what makes a small gain per decision worth anything, and it is also what makes a per-call cost matter.
The output is bounded. A probability, a class from a known set, a number in a range, a ranking. Not prose, not a plan, not an explanation. The answer has a type, and the type is narrow.
Your own history is the signal. The reason to train on your data is that the pattern lives in your data: your customers, your fraud, your churn, your seasonality. A model that has read the internet knows nothing in particular about any of it.
And the quality of the answer is measurable against outcomes that actually happened. You know who churned. You know which transactions were disputed. That labelled history is the asset, and it is also the thing most organisations undervalue, because it looks like old records rather than like a model.
Churn, propensity, demand forecasting, fraud scoring, credit risk, pricing, routing, anomaly detection on telemetry, recommendation and ranking all sit squarely here. So does most of what gets called predictive maintenance. The subject is varied and the shape is identical.
Why reaching for a language model here is the expensive answer
Not wrong in every case, and expensive along several axes at once, most of which are invisible in a prototype.
Cost scales with volume and with payload. A trained model's cost is nearly all paid up front, in the building, and the marginal cost of a prediction rounds to nothing. A language model inverts that: almost nothing up front and a bill attached to every decision, forever, growing with the length of what you send it. On a decision that fires a handful of times a day the difference is irrelevant. On a decision that fires per request, per row, per event, it is the dominant line in the budget. If you want that in your own numbers rather than in the abstract, the LLM cost calculator and the token counter turn a prompt and a volume into a figure, and the AWS cost calculator does the same for the hosting a trained model would need.
Latency sits in a different range. An inline decision inside a request has a budget measured in milliseconds, and a generated answer is not usually delivered in that budget. Pushing the call out of the request path solves the latency and introduces asynchrony, which is a design change rather than a configuration one.
Calibration is the one that gets missed. Plenty of these problems do not want a label, they want a well behaved score: a probability you can threshold, move, and reason about. Risk, pricing, routing and anything with an approval queue all work that way, because the threshold is a business decision that gets tuned after launch. A model asked for a confidence will produce a number that looks like a probability and does not behave like one, and that gap is invisible in testing and expensive in production.
Stability matters more than people expect. A trained model you host is a fixed artefact: the same input gives the same output next quarter. A hosted language model is a moving dependency, and a provider's update is not an event you control. For a regulated decision, or any decision somebody may later have to explain, a fixed artefact with a recorded training set and a measurable error profile is a materially stronger position.
And supervision is cheap on this shape of problem. You have labels, so you have a loss function, and you can measure a change honestly before shipping it. Evaluating a generative system is a project in itself, which we have written about separately. Giving that up on a problem where you did not have to is a real loss.
Where a language model is clearly the right tool
The honest version of this argument has to say where the default is correct, because it frequently is.
Unstructured input with no labels is the obvious case. Contracts, tickets, call transcripts, scanned forms, free text fields nobody ever validated. Extracting structure from those used to be a bespoke modelling project per document type and is now largely a prompt, and that change is genuine.
Open-ended output is the second. Anything where the answer is a summary, a draft, a reply or an explanation has no bounded output space to train against, so the classical framing does not apply at all.
A long tail of categories you could never label is the third. Thousands of classes with a handful of examples each is a bad training problem and a reasonable zero-shot one.
And cold starts belong here. With no labelled history, a language model gets a working system in front of users in days, which is also how you start generating the labels you lacked. That is a sound sequence, provided somebody wrote down that the second phase exists.
The combination that usually wins
The framing of either-or is mostly false, and the arrangement that holds up in production uses both for what each is good at.
Use the language model as a feature extractor and a trained model as the decision maker. Ticket text becomes a handful of structured fields; a gradient boosted tree trained on your own outcomes decides what happens next. The expensive flexible component runs once per document, the cheap calibrated component runs on every decision, and the thing you have to defend to an auditor is the small one.
Use it to bootstrap and then distil. Label a corpus with the large model, train a small supervised model on those labels, serve the small one. You pay the generative cost once rather than per request, and you end up with the fixed artefact.
Keep it for the exceptions. A trained model handles the overwhelming bulk of cases at effectively zero marginal cost, and the residue that falls outside its confidence goes to the expensive path. Most volume never touches the expensive path, which is the entire point.
Either way, the plumbing questions are the same ones any integration has to answer, and they are the part of the quote that decides whether the thing survives a year of real traffic. We worked through those in what AI integration services do not quote, and the question of which actions a system may take unattended is sorted in types of AI agents.
What a classical engagement actually contains
Modelling is the small part, and it is the part the capability grids describe. The rest is where the time goes.
Defining the label is first and it is harder than it sounds. What exactly is churn, measured on what window, and is the thing you can predict the thing you can act on. A technically excellent model of the wrong target is a complete waste and looks like a success in every metric.
Then leakage, which is the most common way a project fails late. A feature that is only populated after the outcome has happened will produce a spectacular offline result and nothing in production. Finding those requires somebody who understands how each column came to exist, which is domain work rather than modelling work.
Then a baseline, honestly built. Last year's value, the segment average, the rule the operations team already uses in their heads. A surprising share of models never beat it, and the ones that do are much easier to justify once the comparison exists.
Then the gap between offline and online, which is a matter of serving the same features at prediction time that you trained on, with the same definitions, against data that arrives late and sometimes not at all.
Then the threshold, which is not a modelling decision. Where you cut a score is a trade between what a missed case costs and what a false alarm costs, and those are business numbers that belong to the business. A delivery that hands over a model without that conversation has handed over half of something.
Then monitoring for drift. A trained model degrades quietly as the world moves away from its training set. No exception, no alert, just predictions that are slowly less useful, which is cheap to instrument at the start and awkward to retrofit.
What to ask the firms on the list
The directories rank by size and by logo, which tells you nothing you need. These questions separate the providers who have done this from the ones who have read about it.
Ask what evidence would make them recommend against a model. A provider with no such answer is selling a model regardless of your problem.
Ask what the baseline is and how they will measure against it. Then ask what result would count as a failure, before anything is built.
Ask how the label is defined and who defines it. If that answer is yours alone to give, the engagement is thinner than it looks.
Ask where inference will run, what it costs per decision at your volume, and what happens to that figure if volume doubles.
Ask who owns the artefact, the features and the training set at the end, and whether you can retrain without them.
Ask what they will leave behind for monitoring, and what threshold triggers a retrain or a rollback.
None of those are about which model is used. All of them decide whether the thing is still in service next year, and a firm that has delivered this before will be glad to be asked rather than pushed.