Skip to content
Insights · Engineering

Run a paid pilot before the full build

Every other way of judging a development company is testimony — a proposal written to win, a reference the vendor chose, a sales call that measures the salesperson. A small paid engagement is the only evidence you generate yourself, and it is worth buying only if it is scoped to produce a decision somebody has agreed to act on.

Scroll
EngineeringSep 16, 20268 min readBy Salman Naqvi, Founder & CEO
Run a paid pilot before the full build

A pilot is worth paying for when three things are true: it is scoped to reduce one named risk, it runs in conditions close enough to the real ones that the result means something, and somebody has agreed in advance what decision each outcome triggers. Miss any of the three and what you have bought is a discounted first sprint with a softer name. The US Government Accountability Office states the discipline more precisely than any vendor will. A technology readiness assessment, it writes, “is a review process to ensure that CTs reflected in a project design have been demonstrated to work as intended (technology readiness) before committing significant organizational resources at the next phase of development” — a CT being a critical technology, one vital to the performance of the larger system (US Government Accountability Office, Technology Readiness Assessment Guide, GAO-20-48G, January 2020). Demonstrated, before committing. That is the whole argument.

Everything else you can do to evaluate a development company is testimony. A proposal is a document written to win the work. A reference is a customer the vendor selected and briefed. A quote is partly a measurement of the work and partly a guess about your budget. A sales call measures the person in the sales call. All of it is worth doing — checking references on a development company and comparing two very different software quotes are the long versions — but none of it is evidence you generated. A small paid engagement is. It is the only instrument on the list that produces a fact about how a specific team behaves on your specific problem, and it is by some distance the cheapest way to be wrong.

Start by naming the risk, because a pilot with no risk attached is a demo with an invoice. Four kinds of risk turn up in a first build and only three are testable in weeks. Technical feasibility: can this be done at all, at the accuracy or the latency the workflow actually needs. Integration: will the system of record give up its data, on terms your security people will accept, at the freshness the workflow assumes. The team: is the firm in front of you capable of this work, and are the people who scoped it the people who will do it. And the commercial question: will anyone pay. The fourth is the one founders most want answered and the one a pilot cannot answer, because a pilot has no market inside it. Stretching an engagement until it might is how six weeks becomes a launch.

Then choose the environment, which is where most pilots quietly cheat. GAO is exact about what counts: “A relevant environment is derived for each CT from those aspects of the operational environment determined to be a risk for the successful operation of that technology.” Read that as an instruction rather than a definition. List the things about production that make you nervous — the records carrying twenty years of inconsistent data entry, the third-party API with the undocumented rate limit, the reviewer who will have to read every output, the Monday morning volume — and put those inside the pilot. Everything else can be a stub, a fixture, a hard-coded screen. A pilot that runs on clean sample data in a fresh account has removed the exact variables it was bought to test, and it will pass.

The reason to borrow the vocabulary at all is that it separates two things our industry calls by the same word. GAO’s maturity scale runs nine levels, “each one requiring the technology to be demonstrated in incrementally higher levels of fidelity in terms of its form, the level of integration with other parts of the system, and its operating environment than the previous.” Level six is a “System/subsystem model or prototype demonstration in a relevant environment” — a “Representative model or prototype system, which is well beyond that of TRL 5, is tested in a relevant environment,” with examples including “testing a prototype in a high-fidelity laboratory environment or in a simulated operational environment.” Level seven is a “System prototype demonstrated in an operational environment”: a “Prototype near or at planned operational system,” separated from level six “by requiring the demonstration of an actual system prototype in an operational environment.” GAO’s threshold for committing to product development is that a technology “reaches at least a TRL 6 and preferably a TRL 7.” Almost every vendor demo you have ever watched sits below six. It ran on a laptop, on prepared inputs, narrated by the person who built it.

The optimism in the room is structural rather than personal, which is worth knowing before you read a pilot report. GAO observes that “in today’s competitive environment, contractor program managers may be overly optimistic about the maturity of CTs, especially prior to contract award” — which is the moment you are standing in. It also records what believing them costs: “GAO has found that in many programs, cost growth and schedule delays resulted from overly optimistic assumptions about maturing a technology.” Its case study on the National Polar-orbiting Operational Environmental Satellite System is the one to hold on to. Two federal departments “committed to the development and production of satellites before the technology was mature” when, on GAO’s count, “only 1 of 14 critical technologies was mature at program initiation.” One in fourteen, about 7% of them, inside a programme with congressional oversight and an independent cost estimate. A founder reading a slide deck has neither.

Size it in weeks, and not for budget reasons. Google Cloud’s DORA programme states the working rule plainly: “Working in small batches allows you to rapidly test hypotheses about whether a particular improvement is likely to have the effect you want,” and the threshold it gives is blunt — “Any batch of code that takes longer than a week to complete and check is too big” (Google Cloud, DORA: work in small batches). A pilot is a hypothesis test with money attached, so the same logic governs it. Long pilots do not produce more certainty. They produce a sunk cost, which is the single thing most likely to distort the decision waiting at the end of them.

Six things go in writing before it starts, and all six fit on one page. The risk, in a sentence phrased so that it can fail. The environment, itemised — which data is real, which integration is live, who the human reviewer is. The threshold, as a number or a rule you could apply without an argument: a task completed end to end without a person touching it, an accuracy bar measured on cases you supplied rather than cases the vendor chose, a latency ceiling, a migration that reconciles to the penny. The decider: a named person who will read the result and choose, ideally not the person who championed the project. The date that decision happens, fixed before anyone is emotionally invested in it. And the disposition of the artefacts — the repository, the data, the written findings, and who owns all three if the answer is no.

Who pays matters, and free is the expensive option. A free pilot is not a gift, it is a sales cost, and sales costs are recovered in the price of the work that follows. Which means the vendor needs the work that follows. Which means the pilot has an answer before it begins. A pilot priced as a deposit against the first phase of the build carries the same defect in a different hat: you cannot walk away without losing money, so the no-go was never a real option. The version that works is boring. One fee agreed before it starts, small enough that being wrong about it is cheap, with a written output that is yours to use anywhere — including with the other firm on your shortlist. Our own interest belongs on the table here, since this article recommends buying something we sell: the readiness assessment is a two-week engagement at one fixed fee whose findings are yours to keep and act on whether or not you build a line of it with us, and our fixed-scope sprints agree an evaluation bar before anything starts, with the build on us if what ships does not clear it. Do not read either as evidence of character. Read the structure, then ask every firm on your list whether their pilot deliverable survives a no-go, and whether they will put that in the contract.

Five pilots are not worth running, and each is common enough to name. When nobody will act on a negative result — if the build is already approved and the pilot is cover for it, you are paying to be told what you decided. When the risk is commercial rather than technical, in which case talk to ten of your own customers instead; it is faster, cheaper, and no vendor can do it for you. When the pilot cannot touch real data because legal will not clear it in time, since that is also what will delay the build. When the success criteria are written by the vendor alone, which converts a test into a demonstration. And when the thing being piloted is not the risky part: a beautiful prototype of the interface, when the risk was always the integration sitting behind it.

How a firm reacts to being asked for a pilot is itself the finding, and three answers should slow you down. We do that for free — see above, and ask what happens to the deliverable if you say no. Let us make it the first sprint — sometimes reasonable, but it means the output is now code rather than a decision, and code is much harder to abandon than a document. And we do not need one, we have built this before — possibly true and easily tested: ask which risk they believe is retired, and what retired it. A firm that has genuinely done it will describe an environment and a measurement. A firm that has not will describe a client.

One honest caveat about the framing, because GAO applies it to its own instrument: “A TRA is not a pass/fail exercise and is not intended to provide a value judgment of the technology developers, technology development program, or program/project office.” Neither is a pilot a verdict on a firm. What it produces is a position on risk — this part is demonstrated, this part is not, and here is what it would take to move it. Your decision is binary because money is binary. The finding underneath it rarely is, and the most valuable outcome available is usually the awkward middle one: yes, but the integration needs its own two weeks first, and here is the evidence for that sentence. A pass-or-fail reading throws that away, which is why the threshold should be written as a measurement rather than as a verdict.

The pilot also shows you something no proposal can, which is the firm on an ordinary week — how they write things down, what they do when Wednesday goes wrong, whether the person who sold it turns up to the standup. That is worth more than the technical finding about half the time. It sits beside, not instead of, the artefacts a paid scoping engagement should hand back: what a discovery phase should produce covers those, and the difference between the two is clean. A discovery produces documents you can act on. A pilot produces a demonstration you can believe. The rest of the criteria for judging a top saas development company sit on the page this article supports, including the one we publicly fail. Run a small piece of paid work through them before you commission a large one.

Find this useful? Tell Google to show you more of it.

Let's put AI to work in your business.

A 30-minute call. You bring the workflow or the roadmap — we'll tell you what's feasible, what it costs, and what we'd build first.

Book a call