
The first system AI touches gets chosen on four properties, and business value is not one of them. Can you produce a few hundred past cases with the correct outcome already attached? Is a wrong output reversible before anybody outside the team sees it? Does the system of record grant read access field by field rather than all or nothing? Is there one named person who can say what a wrong answer costs? Value decides which systems are worth doing eventually. Those four decide which one goes first, because the first project is bought for the layer it leaves behind rather than for its output. A candidate that fails any of the four leaves behind an invoice and a slide.
The default method is an inventory: every department submits candidates, the list is scored on value and effort, the top entry becomes the pilot, and a year later somebody asks what happened to the other forty. That method's weakness is unusually well documented, because the largest AI use-case inventory anybody can open belongs to the US federal government and an auditor has been through it line by line.
The US Government Accountability Office examined the AI inventories that federal civilian agencies are required to publish and reported that "Twenty of 23 agencies reported about 1,200 current and planned artificial intelligence (AI) use cases—specific challenges or opportunities that AI may solve." Then it reported what had actually happened to them: "Most of the reported AI use cases were in the planning phase and not yet in production (i.e., currently used)," and "In about 200 instances, agencies reported that they were currently using AI" (US Government Accountability Office, Artificial Intelligence: Agencies Have Begun Implementation but Need to Complete Key Requirements, GAO-24-105980, December 2023). Twelve hundred identified, roughly two hundred running: the conversion rate of a scored list into a working system, measured across twenty large organisations at once, is about one in six.
Two further findings in the same report explain why the list itself is not the decision. GAO's review found that "five agencies provided comprehensive information for each of their reported use cases while the other 15 had instances of incomplete and inaccurate data," and — the detail worth sitting with for a moment — that "two inventories included AI uses that were later determined by the agencies to not be AI." An inventory is a collection of intentions written by people who will not be the ones making them work, and scoring it more carefully does not change what it is made of. The four properties below are cheaper to check than a scoring workshop, and they are checked against systems rather than against opinions.
One: the workflow already has a graded answer key. Somewhere in your organisation there is a record of what a competent person decided, case by case, over the last two years. Export a few hundred of those with the outcome attached and you can build the release gate that says whether a model or prompt change made the system better or worse. Without it you have no way to accept version two, and every release becomes an argument settled by whoever is most senior in the room. The test takes an afternoon: name the field that holds the right answer, and try to export it. Most disappointing first projects fail here, before a line of code is written rather than in month four.
Two: a wrong output is reversible. Not merely detectable — reversible, by an operation somebody can actually perform. A draft reconciliation awaiting approval is reversible. A posted journal entry is reversible by a correcting entry. An email that has already reached a customer is not, and neither is a price that has already been honoured. The point of the first system is to learn what your error rate really is, in a place where the cost of finding out is a redo rather than a refund.
Three: the record system can be read field by field. Skipping this is what turns a twelve-week project into a nine-month one, because the discovery arrives late and it is not an AI problem. If the only role that can read an order can also read the payment instrument on it, the first task is not retrieval — it is a permissions redesign in a system other teams depend on. What breaks when AI meets a system of record is the long version; the short version is to ask whether a service identity can be granted three fields and refused the fourth, and to ask before the workflow is chosen.
Four: one person owns the definition of wrong. Not a committee — a person, with the workflow knowledge to say what a bad answer costs and the standing to change the definition when it proves incomplete. A workflow spread across three departments has three definitions of a correct outcome and no way to choose between them, which is a governance problem wearing a technical costume. If no single name comes back, that is a real reason to pick a different system first, and a cheaper one than the disagreement that would otherwise surface it.
The uncomfortable consequence is that the workflow nearly every deck nominates is usually the wrong one to start with. Customer support is the standing favourite, and the reason is legitimate: it is where the volume is. The US Bureau of Labor Statistics describes customer service representatives as people who "work with customers to resolve complaints, process orders, and provide information about an organization's products and services," counts 2,666,000 such jobs in 2025 at a median annual wage of $44,770 in May 2025, and projects employment to decline 5 percent from 2025 to 2035 (US Bureau of Labor Statistics, Customer Service Representatives, Occupational Outlook Handbook). Enormous, expensive, already under pressure: everything about the size of that prize is correct.
It is still a poor first system in most organisations, and it fails on the first two criteria rather than the third or fourth. The historical record is a pile of closed tickets, and a closed ticket records what was said, not whether it was right — nobody graded them, because nobody needed to. The output then goes straight to a customer, so the first real error is externally visible and arrives with an apology attached. Build support automation by all means. Just do not build it while also finding out, for the first time, what your error rate is and who owns it.
The better first candidates are usually the ones nobody presents, because they are unglamorous and internal. The same Bureau of Labor Statistics handbook describes bookkeeping, accounting, and auditing clerks as workers who "compute, classify, and record data to help organizations keep complete and accurate financial records" — roughly 1.5 million jobs in 2025, a median annual wage of $50,670 in May 2025, and employment "projected to decline 6 percent from 2025 to 2035" (US Bureau of Labor Statistics, Bookkeeping, Accounting, and Auditing Clerks, Occupational Outlook Handbook). The Bureau attributes the decline not to AI but to software generally — "Software innovations have automated many of the tasks performed by bookkeeping, accounting, and auditing clerks" — and describes the residue as analytical: "Rather than entering data by hand, these workers may focus on analyzing their clients' books and pointing out potential areas for efficiency gains." That is a description of the boundary written by a statistical agency with no product to sell, and it is worth more than a vendor's account of the same boundary.
Read that work against the four criteria and it scores well on all of them. The ledger already holds the graded answer: every past transaction was eventually classified, and the classification that survived the audit is the right one. A proposed entry sits in a queue before it posts, so wrong is a rejection rather than an incident. Finance systems tend to have the most granular permission models in the building, because they were built by people who assumed adversaries. And there is nearly always one controller who can say in a sentence what a misclassification costs.
An independent standards body arrives at the same ordering from the risk side rather than the delivery side. NIST's AI Risk Management Framework, version 1.0, is explicit that prioritisation is the point — "A risk management culture can help organizations recognize that not all AI risks are the same, and resources can be allocated purposefully" — and it draws the distinction this article draws, in almost the same place: "Risk prioritization may differ between AI systems that are designed or deployed to directly interact with humans as compared to AI systems that are not," with higher initial prioritisation called for where "the outputs of the AI systems have direct or indirect impact on humans" (NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023). The framework is voluntary and currently being revised, but the logic does not depend on its status: the systems needing the heaviest governance are the ones facing people, which is exactly why they are an expensive place to be learning.
The same document refuses to let that become a loophole, noting that assessing risk by context still matters "because non-human-facing AI systems can have downstream safety or social implications." An internal reconciliation agent that quietly misclassifies a category for six weeks is not harmless because no customer saw it. It is merely easier to catch, easier to reverse and easier to argue about calmly, which is all the first system needs to be.
What you are actually buying with the first project is the four parts the hub calls Ground, Scope, Score and Log — retrieval against your records, a scoped identity, an evaluation harness, and an audit trail. None of the four is workflow-specific. Build them around a back-office workflow where the answer key exists and the mistakes are cheap, and the second workflow inherits all four at a fraction of the cost. Build them around the customer-facing workflow because it carried the biggest number, and you will build them under production traffic, with an audience. What an enterprise AI operating system actually includes is the concrete version of that layer.
So the meeting is shorter than the scoring workshop it replaces. For each candidate, ask where the graded answer key lives and whether it can be exported; what the reversal looks like and who performs it; whether the record system can grant three fields and refuse the fourth; and whose name goes beside the definition of a wrong answer. Candidates answering all four cleanly go to the top regardless of value score, because they are the ones still running when the value score is finally tested.
The pattern generalises past finance. A generative catalogue enrichment pipeline we built for a marketplace operator turned inconsistent seller listings into structured product data behind a human review queue for low-confidence items, with automated approval above a threshold and continuous evaluation against a golden dataset. All four criteria are visible in that description and none was added later: the golden dataset is the answer key, the review queue is the reversal, the threshold is the boundary, and somebody owned what a wrong attribute meant. The evaluation suite is how an answer key becomes a gate rather than a spreadsheet, and an AI readiness assessment is this conversation held before a budget is committed rather than after one is spent.
The enterprise AI solutions that reach a second workflow are rarely the ones that picked the most valuable first workflow. They picked the most measurable one, learned their own error rate somewhere cheap, then pointed a working layer at the expensive problem with the measurement already attached. Choosing the biggest number first is not ambition. It is paying full price for the lesson.
Related: Enterprise AI solutions
Find this useful? Tell Google to show you more of it.
