
Fixed scope and time-and-materials (T&M) both exist in AI agent development for the same reason they exist everywhere else — but AI work has one wrinkle that normal software doesn't: you often can't write the acceptance test until you've evaluated the model against real data. That single fact should decide which contract structure you sign — not habit, and not whichever structure your last vendor happened to offer. It's usually the first structural question a buyer should ask a vendor, before scope or headcount, because the answer reveals whether the vendor has actually thought about where the project's real risk sits.
The decision rule: if you can write the acceptance test before writing the code — the input format is fixed, the success criteria are checkable, and the workflow already exists elsewhere — fixed scope is safe for both sides, and it's what a fixed-fee offer like an AI readiness assessment is built on. If the acceptance criteria will only be known once you've run the model against your actual data — a new document type, an unfamiliar edge case, a capability nobody's built with this model before — that uncertainty is real, and pricing it as fixed scope just moves the risk onto whoever eats the change orders.
This isn't just a vendor's rule of thumb. Researchers Magne Jørgensen, Parastoo Mohagheghi and Stein Grimstad, publishing in the International Journal of Project Management, studied software project outcomes sorted by contract type and found that fixed-price contracts carry a materially higher risk of project failure than time-and-materials contracts (Jørgensen, Mohagheghi & Grimstad, "Direct and indirect connections between type of contract and software project outcome", International Journal of Project Management, 2017). Their explanation matches what we see on AI engagements specifically: fixed-price framing pushes both sides toward risk-increasing behavior — the provider pads the estimate for the unknown, the client tries to lock down every edge case in the contract instead of in testing, and the real acceptance criteria end up negotiated in the fine print rather than against real data.
Here's what that risk-increasing behavior looks like in practice, because the abstract version undersells it. A team signs a fixed-scope contract for "an agent that reviews inbound documents and flags exceptions." Discovery didn't happen first — the quote was priced off a sample of twenty clean documents someone had lying around. Three weeks into the build, the agent meets the real document population: scanned faxes, multi-language invoices, a layout nobody sampled. The contract has no language for that, because nobody knew to write it in. Now the change order is the real negotiation, and it's happening on worse terms than either side would have accepted upfront — the vendor is behind schedule and defensive, the client already budgeted the fixed number internally and has to go back and ask for more.
AI work makes that dynamic worse, not better, because the uncertainty is usually real rather than a negotiating tactic. A Gartner survey of 782 infrastructure-and-operations leaders found that only 28% of enterprise AI initiatives fully met their ROI expectations, and of the leaders who reported a failure, 57% pointed to a specific cause: they expected too much, too fast, before the system's actual capability on their own data was established (Gartner, Gartner Says Artificial Intelligence Projects in Infrastructure and Operations Stall Ahead of Meaningful ROI Returns, April 2026). That's the same gap a fixed-scope AI contract papers over: the client and the vendor agree on a headline capability before either of them knows what the model can actually do reliably against the client's real documents, tickets, or transactions.
The Jørgensen study's other finding is the more useful one operationally: time-and-materials contracts correlated with deeper client involvement, because the client has ongoing visibility into what's being built and why, rather than a single sign-off at the start followed by a delivery date. When the invoice is tied to hours instead of a locked deliverable, both sides have a reason to review progress weekly rather than wait until handoff to discover the acceptance criteria were wrong. Fixed-scope contracts, done honestly, can borrow that same visibility — regular checkpoints against the golden dataset, not just a final delivery — but it has to be built into the contract deliberately, because the payment structure itself doesn't create it the way T&M does by default.
The honest middle path, and the one we use most: a short, capped discovery phase (T&M or a small fixed fee) that ends with a golden dataset and a working definition of 'correct' — the same evaluation suite discipline that makes production AI measurable rather than a matter of opinion — followed by a fixed-scope build phase once that definition exists. We ran this exact sequence on an admin-trained AI knowledgebase assistant: a short pass to pin down what a correct answer actually looked like, before the fixed-scope build phase began, rather than guessing upfront and re-negotiating later. The discovery phase is cheap insurance against pricing an unknown as if it were known.
One question surfaces which structure you're actually being offered: ask what happens if the model's real-world accuracy comes in lower than expected. A fixed-scope quote should answer with a defined re-scope trigger — a specific accuracy threshold from the discovery-phase golden dataset, and what happens if the model lands below it. An open T&M quote should answer with a cap — a not-to-exceed number tied to that same threshold. If the answer is silence, the risk hasn't been priced at all — it's just been deferred to the invoice, usually as a change order nobody budgeted for.
None of this is a reason to avoid AI work — it's a reason to be precise about which contract you're actually signing and why. The same discipline sits underneath our own AI development pricing: telling a prospective client which of their requirements are fixed-scope-safe today and which need a capped discovery pass first, instead of quoting one number and hoping the two categories don't collide three weeks into the build. Ask a prospective vendor for that same split before you sign anything — it costs nothing to ask, and it's the fastest way to tell whether they've actually thought about your project's risk or are just pattern-matching to their last one.
Related: AI development pricing
Find this useful? Tell Google to show you more of it.
