Skip to content
Insights · Enterprise

Outcome-based AI pricing: what it means if you're buying (or building)

Zendesk and OpenAI are both moving to pricing tied to results, not tokens. That shifts risk toward whoever controls the definition of "done" — here's how to make sure that isn't only the vendor.

Scroll
EnterpriseSep 2, 20266 min readBy Salman Naqvi, Founder & CEO
Outcome-based AI pricing: what it means if you're buying (or building)

"Outcome-based AI pricing" showed up in more vendor pitches this year than any other pricing phrase, and it means something specific: instead of paying for API calls or seats, you pay only when the AI actually finishes the job. Gartner put a number on how much of the software market this could reshape — up to $234 billion in enterprise application spending is exposed to what it calls "agentic arbitrage" through 2030, roughly a fifth of enterprise SaaS spend (Gartner, Gartner Says $234 Billion in Enterprise Application Software Spend Is at Risk from Agentic AI, July 2026). That is a genuinely different way to buy software, and it changes what a buyer needs to have ready before signing — not less diligence, a different kind.

Zendesk is the clearest working example, not a hypothetical. It was the first customer-service vendor to price this way, charging per resolved ticket rather than per seat, and it defines "resolved" narrowly: if the AI handles 90% of an interaction and a human closes the last 10%, that ticket doesn't count as a paid resolution — only a fully autonomous close does. Nikhil Sane, SVP of GTM Strategy and Pricing at Zendesk, framed the shift as market pressure to prove value, not just charge for effort: "As the industry moves toward more transparent, results-oriented business models, we are proud to lead the way with a solution that ensures companies can confidently invest in AI" (Zendesk, Zendesk First in CX Industry to offer Outcome-Based Pricing for AI Agents, 2024; pricing since set at $1.50 per automated resolution as of its 2025 Relate Conference update). Two years on, the model has moved from a differentiator to a category expectation — Gartner's 2026 research now treats outcome-based and hybrid pricing as the pattern CIOs should plan around, not an outlier one vendor tried.

The pattern isn't limited to point-solution vendors either. OpenAI has reportedly begun billing some enterprise customers for completed outcomes rather than raw token consumption, according to reporting that OpenAI itself has not confirmed with published pricing or a public definition of what counts as a completed task (CIO Dive, Agentic AI is shifting the pricing models CIOs rely on, 2026). That gap — a vendor billing on an outcome it defines unilaterally, without a published standard — is exactly the risk a buyer takes on with any outcome-based contract, and it is worth stating plainly before the rest of this argument: the pricing model shifts risk toward whoever controls the definition of "done," not automatically toward the vendor.

That's the actual decision a buyer is making when they sign an outcome-based contract, and it is a governance decision before it is a cost one. Zendesk's own "did a human touch it" rule is at least objective and auditable from ticket metadata. A vendor without that kind of externally checkable definition can grade its own homework: a "successful outcome" that's ambiguous enough gets called a win more often than a stricter definition would allow, and the buyer has no independent way to catch it unless they were already scoring outcomes themselves before the contract started. We've seen the version of this failure mode that has nothing to do with pricing: a support agent whose "resolved" status was set by the model's own confidence score, with no independent check, quietly drifted to calling partial answers complete for weeks before anyone downstream noticed the ticket reopen rate climbing.

The fix isn't refusing outcome-based pricing — it's entering the negotiation with your own definition of the outcome already built, so the vendor's definition has something to be checked against instead of standing alone. That's the same discipline behind AI evaluation and observability: a golden dataset and a scored definition of a correct resolution, built before a system ships, not accepted from whichever party is writing the invoice. A buyer who can say "here's our own pass rate on the same 200 tickets you're claiming a 94% resolution rate on" is negotiating from evidence. A buyer who can't has agreed to trust the vendor's math, which is a worse position than the seat-based pricing they just moved away from.

We built exactly that evaluation-first structure into a support AI agent we shipped for a DTC brand: the system runs against an evaluation suite that scores resolution accuracy before a change goes live, and low-confidence responses escalate to a human instead of getting logged as resolved by default. That's the shape any outcome-based deal needs underneath it regardless of who's selling it — a resolution definition the buyer can independently verify, not one that only exists inside the vendor's own dashboard.

That evaluation work also answers the question every finance team asks the moment outcome-based pricing is on the table: what actually goes in the contract. Three things, in practice. First, a written, example-based definition of a qualifying outcome — not "resolved," but the specific criteria that make a resolution countable, with edge cases (a reopened ticket, a partial answer a human finishes, a customer who simply stops replying) assigned in advance rather than argued about at invoice time. Second, an audit right: the buyer's own team, or an independent evaluator, gets to sample the vendor's claimed outcomes against the agreed definition on a regular cadence, not just at renewal. Third, a cap or true-up mechanism for the gap between the vendor's internal count and the buyer's audited count, so a disagreement about ten edge cases doesn't become a six-figure invoice dispute after the fact. None of that is exotic contract language — it is the same acceptance-test discipline any fixed-scope engineering statement of work already uses, applied to a running system instead of a one-time delivery.

None of this is an argument against the pricing shift — cost tied to delivered value is a genuinely better incentive than cost tied to how many tokens a task happened to burn, and it's the direction the market is visibly moving. It's an argument for treating "how is success defined, and who can check it" as the first pricing question, not the last one, the same way we treat it in our own AI development pricing: a build cost and a run cost that are both defined against a scored outcome, not a vendor's word for it. Ask any AI vendor pitching outcome-based pricing to show you their resolution definition and let you audit it against a sample of your own data before you sign — a vendor confident in what they're shipping will not hesitate. One who won't is telling you the definition doesn't survive being checked.

Find this useful? Tell Google to show you more of it.

Let's put AI to work in your business.

A 30-minute call. You bring the workflow or the roadmap — we'll tell you what's feasible, what it costs, and what we'd build first.

Book a call