Skip to content
Insights · Enterprise

When you need an AI development partner, and when you don't

Five conditions make an outside AI partner the wrong hire, and all five are visible before you sign. Then the narrower set of cases where outside help genuinely pays — written by a firm that sells it.

Scroll
EnterpriseSep 10, 20267 min readBy Salman Naqvi, Founder & CEO
When you need an AI development partner, and when you don't

Five conditions make hiring an outside AI partner the wrong decision, and they are worth naming before the reasons to hire one. The problem is not actually an AI problem. Something already on the market solves it. There is no data worth retrieving. Nobody inside can say what a correct answer looks like. Or the organisation cannot yet operate what it would receive. Any one of those turns an engagement into expensive theatre, and every one is visible before a contract is signed. Outside help does earn its cost, in a narrower band than the market admits — and that band is the second half of this article.

The conflict of interest is total, so it belongs at the top rather than in a footnote. We sell AI development services, so every case below where the answer is hire nobody is revenue we are arguing ourselves out of. It gets published because the alternative costs more on both sides: an engagement built on a problem that was never an AI problem ends with a system that works, is not used, and cannot be referenced afterwards.

The problem is not an AI problem more often than the discourse allows, and the tells are specific. There is a rule, written down somewhere, and the only difficulty is that nobody has encoded it. Or the bottleneck is a permission nobody will grant, a process step with no owner, or a field entered by hand because no system owns it — in which case a model guesses from exactly the information the human had. A language model is a probabilistic component; putting one where a query, a form and a scheduled job would do adds an evaluation burden, a per-token bill and a new class of failure to a problem that had none. The people who would maintain it are already sceptical. In Stack Overflow's 2025 Developer Survey, ranking what attracts them to a technology, respondents put AI integration or AI Agent capabilities ninth out of ten — behind an easy-to-use API, a robust API, a reputation for quality, reliability, and manageable costs (Stack Overflow, 2025 Developer Survey: Work).

Something you can already buy solves it in more cases every quarter, because the vendors of the software you rent are shipping these capabilities into their own products. Eurostat's 2025 figures map where: among EU enterprises using AI, 34.70% used it for marketing or sales and 31.05% for the organisation of business administration processes or management — the two commonest purposes by a distance — while logistics came last at 6.08% (Eurostat, Use of artificial intelligence in enterprises, 2025 data). Copy drafting, meeting summaries, ticket triage: those arrive inside a subscription you already pay for. The test is the one in bespoke or off-the-shelf — would a customer ever switch to a competitor over this component? If not, rent it.

There is no data worth retrieving more often than anyone discovers before kickoff, because retrieval projects are sold on the assumption that an organisation's knowledge lives in its documents. Frequently it lives in people. What lives in the documents is eleven versions of the same policy with no way to tell which is current, and a wiki last edited by someone who has left. Retrieval over that corpus does not fail loudly — it answers confidently from the wrong version, which is worse than no system, because the wrong answer now carries the platform's authority. The diagnostic takes ten minutes: name your three most-referenced internal documents, ask two people to produce the current version of each, and see whether you get the same six files. If not, the work in front of you is curation and ownership, and only your own people can do it (RAG is a data problem is the long version).

Nobody internally can say what a correct answer looks like is the condition that quietly ends the most engagements, because it is invisible during a demo and fatal after one. No vendor can define correct on your behalf: what counts as correct in a claims workflow, a clinical summary or a pricing recommendation is a statement about your business and your liability. And the failure it produces is the one practitioners report most — Stack Overflow's 2025 survey found 66% of developers naming "AI solutions that are almost right, but not quite" as their single biggest frustration, and 46% actively distrusting the accuracy of AI output against 33% who trust it (Stack Overflow, 2025 Developer Survey: AI). Almost-right is only detectable against a standard. If you cannot assemble fifty real examples with agreed answers, no evaluation suite can be built — and without one you are buying a demo with a support contract. The evaluation suite sets out what that artefact contains.

The organisation cannot yet operate what it would receive is the least discussed condition and the most measurable, because there is public evidence for how badly institutions estimate their own delivery capacity. The US Government Accountability Office reported that the Office of Management and Budget "now requires agencies to deliver useable functionality every 6 months," that 22 agencies reported 64 percent of their software development projects would do so, and that GAO's review of seven departments found approximately half actually reported delivering functionality every six months — the difference traced partly to "a lack of support for reported delivery" (US Government Accountability Office, Information Technology Reform: Agencies Need to Increase Their Use of Incremental Development Practices, GAO-16-469, August 2016). Fourteen points of optimism inside organisations with a mandate and an auditor; a company with neither should assume its gap is wider. So the operating questions belong before the build: who runs this on the Tuesday after the team leaves, who is paged when it produces a wrong answer at scale, and which budget line pays for the inference. The handover you must receive lists the artefacts that decide the answer.

Now the other half — and the source that argues hardest against hiring makes the strongest case for it. Eurostat reports 19.95% of EU enterprises used AI technologies in 2025: 17% of small enterprises, 30.36% of medium and 55.03% of large. Among enterprises that considered AI and did not adopt it, the commonest reason by a wide margin was a lack of relevant expertise, at 70.89%, then lack of clarity about the legal consequences (52.52%) and concerns about violation of data protection and privacy (48.83%). The *least* common reason, at 20.68%, was that the technologies were not useful for the enterprise (Eurostat, Use of artificial intelligence in enterprises, 2025 data). That ordering is the market in one figure: one company in five walked away because it was not useful to them; seven in ten because they did not have the people.

It pays when the skill is needed intensely now and thinly forever. Evaluation harnesses, retrieval quality work, guardrail design, cost and latency engineering against a provider's API — each absorbs a quarter of concentrated attention, then needs a fraction of a person indefinitely. Hiring for that shape means a search you have no time to run and a role that becomes boring to the person you finally hired. The counter-case is argued in in-house AI team vs AI development agency: if the system is the product and will be built on for a decade, hire, and hire early.

It pays when the hard part is the integration rather than the model. Connecting to a CRM is a week of work. What follows is identity — a service account with far more reach than the person who asked it for anything — plus permission inheritance into a retrieval index, freshness against a system of record, idempotent writes so a retry does not issue a second refund, and an audit trail that can bound the damage of a mistake nobody noticed for a month. That is where the schedule actually goes (what breaks when AI meets a system of record is the inventory), and it is cheapest done first: the student information system we built for Concordia Colleges had per-branch data isolation and audit trails from the first sprint of a 22-month engagement, not fitted underneath 150-plus live branches afterwards.

It pays when something already exists and has stalled. A pilot that demos beautifully and will not ship is a diagnosable condition rather than a failure of ambition — usually missing evaluations, missing guardrails, an integration scoped as a week, or a security review nobody planned for. Outside help is efficient here precisely because the work is bounded and the diagnosis is a known list. Why AI pilots fail to reach production is that list, and the pilot-to-production rescue is the engagement shaped around it, with the scope fixed before it starts rather than discovered inside it.

It pays when what you need is a decision rather than a build, and this is the case buyers skip most often, because it looks like a smaller purchase than it is. If you cannot yet tell which of the five conditions applies to you, the cheapest thing to buy is the answer. Our own first engagement is deliberately that shape: a two-week AI readiness assessment mapping systems, data and security posture against three to five candidate use cases, handing back a roadmap that is yours whether or not you build with us. A firm that only earns money by building has no mechanism for telling you not to build — the same reason pricing publishes fixed deliverables and no rate card.

Three questions separate a partner from a supplier of hours. Which of these five conditions applies to us? A firm that cannot name one has not read your situation, or has decided agreement closes faster than accuracy. What would you need from us before an evaluation suite could exist? The right answer includes a quantity of work that is yours rather than theirs. Who operates this after you leave? An answer that terminates in a retainer is describing a dependency, not a handover.

None of this makes the decision easy; it makes it decidable. The one-sentence version: hire an outside partner when the constraint is expertise you need intensely now and thinly later, integration into systems you do not control, or a build that stalled at a diagnosable point — and hire nobody when the problem is a rule nobody wrote down, a product you could rent, a corpus nobody maintains, a definition of correct that does not exist, or an operating capability you have not built yet. If you do need help, the specialist criteria are in how to choose an AI agent development company, and the general version — how to test any firm, this one included — is in the criteria for choosing a top saas development company.

Find this useful? Tell Google to show you more of it.

Let's put AI to work in your business.

A 30-minute call. You bring the workflow or the roadmap — we'll tell you what's feasible, what it costs, and what we'd build first.

Book a call