Skip to content
Insights · Engineering

What a discovery phase should produce

Six artefacts separate a discovery worth its fee from a proposal with a fee attached — and the last of them is a recommendation the author is willing to make against their own commercial interest.

Scroll
EngineeringSep 15, 20268 min readBy Salman Naqvi, Founder & CEO
What a discovery phase should produce

A discovery phase is the only part of an AI engagement whose entire output is a document, which makes it the easiest thing in this industry to sell and the hardest to grade. Six artefacts settle it. A measured baseline for the task under consideration — how it is done today, how long one instance takes, how often it has to be redone. A written definition of a wrong answer, produced by your people rather than the vendor's. A data and permission map: what exists, who may see which part of it, and which system is authoritative when two disagree. A scored shortlist in which the non-AI alternative has been costed beside the AI one. A cost and latency model expressed per completed task rather than per thousand tokens. And a go/no-go recommendation the author is willing to make against their own commercial interest. If what lands instead is a deck of recommendations only the firm that wrote it can execute, you did not buy a discovery. You bought a proposal with a fee attached.

The conflict is total, so it belongs here rather than in a footnote: we sell this phase. Our first engagement with most clients is a two-week AI readiness assessment at a fixed fee agreed before it starts, and its roadmap is yours whether or not you build with us. That page also publishes who it is not for, and one line matters more than anything in the sales copy: it is not for organisations unwilling to share real systems and data access for two weeks. A discovery without access is a workshop. Everything below is a standard we can be held to as easily as the firm you are comparing us with.

The phase resists grading for a structural reason rather than a moral one. Every other phase produces something that either runs or does not. Discovery produces judgement, delivered as prose, usually by the firm bidding for the build that follows — so one document serves as an honest assessment and as a sales instrument, and nothing in its format says which you are holding. The six artefacts are useful precisely because each is falsifiable by someone who was not in the room. A baseline is a number somebody measured or it is not.

The measured baseline is the first artefact and the one most often missing, because measuring how the work is done today is slow, unglamorous, and visibly not building anything. It is also the only thing that will later let anybody say whether the system helped. The standards body names the difficulty rather than waving it away: NIST's AI Risk Management Framework lists what it calls the human baseline among the challenges of measuring AI risk, observing that risk management of systems "intended to augment or replace human activity, for example decision making, requires some form of baseline metrics for comparison", and that this "is difficult to systematize" because AI systems carry out different tasks, and perform them differently, than humans do (NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023). Difficult is not the same as optional. A discovery that skips it has removed the only evidence that could settle the argument you will be having in month nine.

Four numbers per candidate task are enough, and none requires a vendor to collect. How many times does this happen in a month. How long does one instance take, end to end, including the waiting. How often does it produce something that has to be redone. And who does it now, at what grade. If no system records any of that, count by hand for a fortnight — every later claim of improvement is measured against this number or against nothing.

The written definition of a wrong answer is the second artefact, and the discovery's job is not to write it. It is to extract it from you and get it signed by someone who can be held to it. What counts as wrong in your business is a business judgement: a refund that should have escalated, a clause the summary dropped, a record that must not surface in a search result for this user. No vendor can supply that, and one never made to ask will supply it anyway, by accident, in code. What a good discovery produces here is short — a page of rules, a set of real cases where two competent people disagreed about the right answer, and the name of the person who breaks the tie. The evaluation suite is what that page becomes once there is a system to run it against, and the substance behind AI evaluation and observability.

The data and permission map is the third artefact and the one most likely to change the plan while changing it is still cheap. Three questions per source. Does it exist in a form a machine can read, or as a habit in somebody's head. Who is entitled to see each part of it, and is that entitlement enforced by a system or by an understanding. And which source is authoritative when two disagree, because they will. The third question reprices projects quietly: a retrieval layer over two systems that contradict each other does not fail loudly, it answers confidently from whichever one it happened to read. RAG is a data problem covers the corpus half of this and what breaks when AI meets a system of record the write half.

The scored shortlist is the fourth, and it is honest only if the non-AI option is priced in the same column. The framework asks for exactly that, listing among the resources to be accounted for "viable non-AI alternative systems, approaches, or methods", and asking that "potential costs, including non-monetary costs, which result from expected or realized AI errors or system functionality and trustworthiness" be examined and documented. Read that second clause as the line item nobody puts in a business case: the cost of being wrong at volume, in refunds, in rework, in the hour a person spends not trusting an answer. A shortlist on which every candidate scores well is a menu, and a menu is what a sales exercise produces, because striking a row costs the author revenue.

A discovery also has to say which technology it is recommending, because the category is a basket whose members have almost nothing in common commercially. Eurostat's 2025 figures show no dominant member at all: among EU enterprises with ten or more people, the most used AI technology was text mining at 11.75%, followed by technologies generating pictures, video or audio at 9.55% and natural language generation or speech synthesis at 8.76%, down to technologies letting machines move autonomously by observing their surroundings at 1.39% (Eurostat, Use of artificial intelligence in enterprises, data extracted December 2025, from a survey of 157,000 of the EU's 1.53 million enterprises). Eight technologies, eight sets of data requirements, eight failure modes. A document recommending that you adopt AI without naming which of these, against which task, has narrowed nothing.

The same survey is a useful corrective on where the market actually is. Among EU enterprises not using any AI technology in 2025, 14.21% had ever considered using one — 36.54% of large enterprises against 12.65% of small ones. Consideration is the stage at which a discovery gets bought, and most of the market has not reached it. Worth holding on to when a proposal implies you are late: nearly everybody is at the beginning.

The fifth artefact is a cost and latency model per completed task, and it is the one buyers accept in the wrong unit. A price per million tokens is not a cost model; it is an input to one. What you need before committing is an estimate of what one finished unit of work costs and how long it takes — one resolved ticket, one enriched listing, one drafted response — including the retries, the escalations to a human, and the calls that produced nothing useful. That number decides the product, because an AI feature's cost scales with engagement rather than with seats, and the design decisions that set it are taken before anyone writes code (what makes an AI feature expensive to run works them through). A discovery that hands you a token price and calls it a budget has handed you the one number architecture cannot move.

The sixth artefact is a recommendation, and the framework is explicit that this is what the phase is for. Its MAP function "establishes the context to frame risks related to an AI system", and asks that the business value or context of business use "has been clearly defined", that organisational risk tolerances "are determined and documented", and that the targeted application scope "is specified and documented based on the system's capability, established context, and AI system categorization". Then it states the output plainly: after completing that function, users "should have sufficient contextual knowledge about AI system impacts to inform an initial go/no-go decision about whether to design, develop, or deploy an AI system". No is a valid output. A firm structurally unable to produce it has not run a discovery.

Four tells separate the real thing from the sales exercise, and all four are visible in the deliverable. Nobody measured anything — the document carries industry figures and not one number from your own building. Every candidate survived — nothing scored low enough to be dropped, so the scoring was decorative. Nothing in it is executable by anyone but the author — the recommendations are shaped so the next step is necessarily a contract with the firm that wrote them, rather than a specification another vendor could bid against. And it was free. A free discovery is priced into the build that follows, which means the firm cannot afford for it to end in a no, and everyone in the room knows it.

A fifth tell shows up earlier than the deliverable: a discovery that never asked for access. If nobody requested a read-only credential, a schema, a sample of real records, or an hour with the person who does the work, what is being written is a description of your industry rather than of you. The corollary is that a real discovery costs you hours as well — finding the data, arranging the access, sitting through the interviews, arguing about what a wrong answer is. Budget them explicitly. A discovery starved of them produces a document that is fluent, plausible, and about nobody in particular.

What to buy is therefore narrow: a paid, bounded piece of work with the artefacts named in the agreement, a fixed end date, and a clause saying the output is yours to take to another firm. Ours is sold as a decision rather than as the first slice of a build for a structural reason, and it is the test to apply to whoever else is quoting you: a firm that only earns money by building has no mechanism for telling you not to.

Discovery is cheap in the only sense that matters — it is the last point at which changing your mind costs a conversation rather than a quarter. Done properly, the build that follows is dull, which is the highest compliment available in this work. A production support agent we built for a direct-to-consumer brand was grounded in order data, policies and the product catalogue, with permission guardrails on the actions it was allowed to take and confidence-based escalation that carries the full conversation to a human. None of that was discovered late; it was decided while deciding it was free. Our AI development services page maps the work that follows, and what AI development services actually include names the seven workstreams a scope should be written in once the artefacts exist.

Find this useful? Tell Google to show you more of it.

Let's put AI to work in your business.

A 30-minute call. You bring the workflow or the roadmap — we'll tell you what's feasible, what it costs, and what we'd build first.

Book a call