How to Choose Between AI Consulting Firms
Search for AI consulting firms and every result is a ranked list of them. The lists are largely accurate and they are also nearly useless for choosing, because they answer a question you did not ask. You do not need to know who the best AI consultancy is in the abstract. You need to know which kind of company can do the thing in front of you, and that is decidable from your own situation before you speak to anybody.
The firms on those lists are not competing with each other in any meaningful sense. They sort into three groups that sell different products, produce different artefacts, and fail in different ways. Sorting yourself first removes most of the list.
Three kinds of company share one label
Global strategy consultancies. The large names, with AI practices attached to an existing relationship with your board. What they sell is a decision made defensible: a portfolio view, a comparison of options with risk attached to each, and a recommendation your executives can act on without having to defend the reasoning themselves. What they produce is documents and a plan. Delivery is usually a separate engagement, often subcontracted, and the gap between the strategy and something running is where most of the disappointment in this category lives.
Engineering and build shops. Larger delivery organisations and offshore development firms, selling capacity. What they sell is people who will build what you specify. They are genuinely good at that, and the failure mode is that they will build what you specify, including when the specification is the thing that was wrong. A build shop handed an underdefined problem produces working software that solves it badly, on time.
Independents and small specialist firms. One person or a handful, usually engineers who do the strategy because the two are not separable at small scale. What they sell is judgement plus hands. The failure mode is capacity: an independent cannot staff a programme across three business units, and one who says otherwise is describing a subcontracting arrangement.
None of these is better. They are answers to different questions, and the expensive mistake is buying one while having the problem that another solves.
Which one you need follows from where you are stuck
Three questions sort it faster than any comparison table.
Is the disagreement about what to build, or about how to build it? If your organisation does not agree on which process is the painful one, no amount of engineering capacity helps, and hiring it means paying builders to wait. That is a strategy problem and it wants the strategy answer, at whatever scale matches your company.
Has anybody written down what working would mean? Not a goal, a test: a statement both sides could check afterwards to agree whether it happened. If nobody has, a build shop is the wrong next call, because the specification they need does not exist yet and they will not refuse the work for that reason.
How many teams does this touch? One workflow in one department is independent territory. A programme crossing several business units with its own governance needs an organisation that can staff it, and that is a real constraint rather than a snobbery about size.
What a first engagement should produce
Whichever tier you land on, the first piece of work should produce decisions rather than software, and it should be short enough that the decisions are still cheap to change.
At the end of it you should hold a shortlist of candidate workflows with an honest note on each about why it is or is not a fit, one chosen workflow, and a written definition of what working means for it. That last item is the one that gets skipped, and it is the one that later decides whether anybody can tell if the thing succeeded.
A proposal that names a model, a framework or a vendor before it names a workflow has the order backwards. The model is close to the last decision rather than the first, and it is the decision most likely to be revised after launch anyway. We set out what that first engagement contains in more detail in a piece on what a small business AI consultant actually does, and the shape holds at larger sizes too.
Do the running-cost arithmetic before you take anybody's estimate
Every proposal quotes the build. Almost none of them quote what the thing costs to run, and the running cost is the half that lasts.
Usage of a hosted model is billed per token, with the text going in and the text coming out priced on separate lines rather than pooled. That structure means the cost of a feature depends on which direction the text is travelling rather than on how much of it there is, which is why estimates handed to you by a firm proposing the work are worth checking rather than accepting. How the billing is actually put together, and where to read live rates, is covered in our piece on API pricing.
Then compute it yourself. Take one realistic request from the workflow you are considering, count what it genuinely sends with a token counter, and put that through the LLM cost calculator at the volume you actually expect. It takes a few minutes, and a firm whose estimate does not survive that check has told you something useful about the rest of their estimates.
Questions that separate the three tiers quickly
Ask all of these of any firm, regardless of tier. The answers place them faster than their case studies do.
- Which single workflow would we start with, and why that one rather than the others?
- What does working mean for it, stated so that we could both agree afterwards whether it happened?
- Who does the build, you or somebody you subcontract, and who maintains it after handover?
- What does this cost per month to run at our current volume, and at twice it?
- What happens when the model is wrong, and where does the wrong answer stop before it reaches a customer?
- What do we own at the end, and can somebody else maintain it?
The third and the last are the ones that surface the difference. A strategy firm will be candid that delivery is somebody else's; a build shop will answer the cost and maintenance questions precisely and the workflow question vaguely; an independent will answer all of them and should be pressed on capacity instead.
You can do the first part without hiring anybody
For a lot of organisations the honest answer is that the readiness is not there yet, and that is a real finding rather than a failure. Data spread across systems nobody can query, no agreement internally on which process hurts, or a workload too small to repay any build are all reasons to wait, and none of them are improved by buying a strategy document about them.
The AI readiness assessment on this site walks the same ground a scoping engagement covers, and it costs nothing. If it comes back saying you are not ready, that is the same answer a good consultant would give you, arrived at sooner. If it comes back with a workflow worth doing, you now have a specific thing to take to a firm rather than a general intention, and specific briefs are quoted far more honestly than general ones.