Before You Hire an AI Software Development Company
Search this and every result is a sales page. Hire our developers, our engineers are vetted, our team ships in weeks, logos arranged by industry. They are selling the same thing in different fonts, and none of them answers the question you were actually holding when you typed it, which is whether you need to hire anyone.
That question has a real answer and it is not always no. It turns on one distinction: whether what you want is a model call or a system. A model call is a request to a provider that comes back with text. A system is everything that has to exist around that request before the answer can be trusted, delivered to the right person, and relied on when the provider is slow. The first is something your existing team can build this quarter. The second is a software project, and software projects are what development companies are for.
The trouble is that the two look identical in a demo. Both produce an answer on a screen. The difference only appears later, under real traffic, with real permissions and real consequences for a wrong answer, which is why so many engagements are scoped against the first and priced against the second, or the reverse.
The line, drawn at four specific places
We wrote up the four parts of an integration in a piece on what AI integration services do not quote. They are the same four here, used differently: there, they are what a proposal should cover, and here they are the test for whether you have a project at all. Walk through them honestly about your own case and the answer usually falls out.
Where the context comes from. If the model can do the job with what the user types plus a fixed instruction, there is no retrieval problem and you do not have a system. If it needs to know things about your business, something has to find those things, keep them current, and decide which of them is relevant to this request. That is an indexing and freshness problem, it is ordinary engineering, and it is where answer quality actually lives. The moment this appears, you have crossed the line.
Who is allowed to see it. If every user of the thing may see everything it can reach, you have no permissions problem. If they may not, the enforcement has to happen at retrieval rather than in an instruction to the model, because the instruction and the content share one channel and anyone typing into the box is writing in that channel too. Permission-aware retrieval is not hard and it is not optional, and it is the thing most often missing from a demo where one person is logged in as themselves.
What happens when nothing answers. Model providers tend to fail slowly rather than loudly: the request hangs. A toy can hang. Anything a customer or a colleague depends on needs a defined behaviour for that case, including whether a retry can repeat a side effect that already happened. If your thing writes a record, sends an email or moves money, this is not a detail you add later.
How a wrong answer gets caught. A wrong answer that reaches a customer is the failure that costs the most and the one that is invisible in testing, because during testing you are reading every output. In production nobody is. Catching it means a way of measuring output quality that is built from your own failures rather than from a benchmark, which is a body of work in itself and the subject of our piece on LLM evaluation.
Zero or one of those four and you have a feature. Three or four and you have a system, with a surface area that keeps existing after launch.
What a system costs that a feature does not
The part that surprises people is not the build. It is that a system has an owner, forever. Retrieval drifts as the source documents change. Permissions change when someone leaves. A provider deprecates a model and the prompt that was tuned against it behaves differently. None of that is a defect and all of it is work that arrives whether or not anybody budgeted for it.
Infrastructure is the visible half of this and the easier half to estimate. If your plan involves running anything of your own rather than only calling a provider, the AWS cost calculator here is a faster way to find the shape of that number than a vendor conversation is, and it is worth having before the conversation rather than after.
The invisible half is the attention. Somebody has to notice that the answers got worse, and noticing is harder than fixing. That is a staffing question rather than a procurement one, and it is the one a statement of work will not resolve for you.
So when is hiring the right answer
Three cases, and they are narrower than the search results imply.
The first is that you have a genuine system by the four tests above and no team with the time to build it. This is the straightforward case and it is what these companies are for. What you are buying is throughput, and the thing to check is whether the proposal addresses all four parts or quietly leaves two of them with you.
The second is that you have the team but not the particular experience, and the work is on a deadline where learning it in public is expensive. Here what you want is not a dedicated team but a smaller engagement that leaves your own people able to run it afterwards, which is a different shape of contract and one you have to ask for explicitly, because it is not the default.
The third is that you do not yet know which of these you are in. That is the most common position and it is badly served by the vendor pages, since every one of them resolves the ambiguity in the direction of a larger engagement. It is the case a small-business AI consultant is genuinely for: a short, bounded piece of work whose only deliverable is a decision.
And the case for not hiring: one of the four parts, a contained audience, and nothing irreversible happening downstream of a wrong answer. Build it, keep it small, and find out what it does under real use before you commission anything larger. The information you get from a month of real traffic is worth more than any discovery phase, and it is cheaper.
Questions worth asking before the demo
A demo is designed to show the part that already works. These are the questions that reach the rest of it, and they work on any vendor regardless of how the page is worded.
Ask where the context comes from, and specifically how long it takes for a change in a source document to reach an answer. A vague reply here means the retrieval design does not exist yet.
Ask how permissions are enforced. If the answer involves telling the model what it must not reveal, that is the wrong answer and it is a common one.
Ask what happens when the provider is slow rather than down, and whether a retry can repeat an action that already completed. This is the question that separates people who have run one of these from people who have built one.
Ask how they will know the output has got worse months after launch. If the answer is a dashboard of generic metrics, ask which of your failures those metrics would have caught.
Ask what you own at the end, and who maintains it. The distinction between a deliverable and a dependency is the thing most easily left ambiguous and the most expensive to discover later.
None of these are trick questions and a good partner will welcome all five, because they are the questions their own engineers are already asking internally. A vendor who cannot answer them is not necessarily bad at building software. They are describing a feature, and you came here to find out whether you have one.