Home/Blog/Safety & governance/AI in government: 126 use cases, 65 not made public
AI in government: 126 use cases, 65 not made public126 active use cases. 65 not detailed publicly. Then tools found missing.100%use cases in the inventory48%publicly detailedOF ORGANISATIONS DEPLOYING AI126 active use cases. 65 not detailed publicly. Then tools found missing.
126 active use cases. 65 not detailed publicly. Then tools found missing.

AI in government: 126 use cases, 65 not made public

One agency reported 126 active AI use cases and auditors found the inventory still incomplete, with tools contracted to build criminal cases missing from it entirely.

TL;DR. Government is the only domain in this series where the public can, in principle, see a list of every AI system in use. Federal agencies are required to publish use case inventories. One agency's inventory recorded 126 active use cases as of June 2025, of which 65 were not publicly detailed. Auditors then found the inventory was still incomplete, having missed tools contracted to help build criminal cases. A separate audit of 13 acquisitions across four departments found contracting officers unable to find the technical staff to evaluate what they were buying, and no systematic sharing of lessons between agencies. Roughly $1.7 billion has been appropriated for federal AI, moving through procurement processes not designed to assess it.

---

Every other domain in this series has a measurement problem you have to infer. Government publishes its own.

Federal agencies are required to maintain and publish inventories of their AI use cases. It is the most transparent arrangement in this series, and it produced two auditor reports in 2026 that are worth reading together.

The first examined one large agency and found 126 active AI use cases as of June 2025. Sixty-five of those were not detailed publicly. The auditors then identified AI-enabled tools that agency officials said were contracted to help build criminal cases, and which did not appear in the inventory at all.

The second examined 13 AI acquisitions across four departments and found a consistent pattern: agencies acquiring AI faster than their procurement frameworks can evaluate it, then learning expensive lessons without sharing them.

Put together, they describe the recurring failure of this whole series arriving in the one place designed to prevent it. The inventory exists, it is mandatory, it is published, and it is incomplete by the auditor's own finding.

The inventory problem, for the third time

This is now the third domain in this series where the same structural failure appears, and the repetition is the finding.

In hiring, employers determine whether their own tool falls in scope, so an absent bias audit cannot be distinguished from a good-faith scope decision. Eighteen audits appeared across 391 employers.

In finance, the most common examination finding is incomplete model inventory. Banks regularly discover undocumented models during examinations.

In government, the inventory is a published legal requirement, and the auditors found tools missing from it.

Three regimes, three different enforcement mechanisms, the same gap. That is strong evidence the problem is not enforcement intensity but the definitional question underneath it: somebody has to decide what counts as an AI system, and that somebody is always the party being inventoried.

The government case is the cleanest demonstration because it removes every alternative explanation. There is no commercial incentive to conceal, the requirement is explicit, the format is specified, and the result is still incomplete. When an obligation depends on self-classification, the failure is structural rather than motivational.

The definitional problem, stated properly

Three domains showing the same failure is enough to state the underlying problem precisely, because every proposed fix depends on getting it right.

There is no workable definition of an AI system that a non-specialist can apply.

Consider what has to be decided. A linear regression in a spreadsheet that determines who gets audited: is that an AI system? A rules engine with a thousand hand-written conditions? A vendor product whose internals are undisclosed, which may or may not contain a model? A general assistant used by staff without procurement's knowledge? A model that was retired but whose outputs still populate a database that other systems read?

Every inventory regime asks someone in a compliance function to answer these, and there is no test they can apply. The regulations offer definitions like "substantially assists a decision" or "computational process issuing a score", which are legally serviceable and operationally useless to a person looking at a spreadsheet.

Three fixes get proposed and each has a failure mode.

Define by technique. Anything using machine learning counts. Clean, and it captures a spreadsheet regression while missing a consequential rules engine, so it inventories on the wrong axis.

Define by impact. Anything affecting a person's rights or resources counts, regardless of technique. Better aligned to why we care, and it requires a judgement about impact that is exactly as contestable as the one it replaces.

Define by register-everything. Declare all decision-supporting systems and let a reviewer classify. This is the only version where absence becomes meaningful, and the reporting burden is large enough that nobody has adopted it.

The honest position is that inventories will remain partial, and the useful response is to read them as samples rather than censuses. An inventory tells you what an organisation classified as AI. It does not tell you what an organisation runs, and treating the first as the second is the error every one of these audits has found.

What the procurement audit found

Six recurring challenge categories, appearing across agencies regardless of mission, vendor or technology type. Three are worth extracting because they generalise beyond government.

Contracting officers cannot evaluate what they are buying. Officials at multiple agencies reported difficulty finding data scientists, machine learning engineers or computer vision specialists to assess contractor proposals. The buyer lacks the expertise to judge the product, which is not unique to government and is more visible there because someone audits it.

Costs are hard to understand. Agencies reported difficulty establishing what an AI system would actually cost, which is unsurprising given that inference pricing, retraining, data preparation and integration are separate lines that vendors present in different combinations.

And traditional acquisition timeframes do not fit. A procurement cycle measured in years is being applied to a technology whose capabilities and pricing change within it. The system being bought at the end of the process is not the system that was specified at the start.

Two further findings matter for the same reason.

Data and intellectual property protections are a live negotiation rather than a settled default. Who owns a model fine-tuned on government data, and what a vendor may do with it, is being decided contract by contract.

And nobody was sharing lessons. The audit's central recommendation was that four departments update their policies to systematically collect and submit acquisition lessons to a shared repository. All four agreed, with target dates in mid-2026. Which means that until now, each agency was learning the same expensive lessons independently.

The oversight gap inside one agency

The agency audit contains a specific finding that generalises to any large organisation.

Several entities had oversight of individual AI use cases. None was responsible for managing AI investments across the agency.

That is worth sitting with. Every individual deployment had a reviewer. The portfolio had nobody. There was no process for ensuring that AI investments contributed to agency-wide goals, and no mechanism for asking whether 126 use cases were the right 126.

This is the difference between reviewing decisions and having a strategy, and it is the most common failure in large organisations adopting any technology. Case-by-case governance produces a defensible record for each case and no answer at all to whether the aggregate makes sense.

The audit also found staffing reductions had left the agency without enough skilled employees to support or develop AI tools, and no workforce plan identifying what skills were needed. An agency running 126 AI use cases without a plan for the people who maintain them is accumulating a liability, and that liability arrives later, when the person who understood the model has gone.

The transparency paradox

The government inventory requirement produces a truly useful public artifact and demonstrates its own limits.

What it does well. A published inventory means a citizen or journalist can see, in structured form, what systems an agency runs, whether they are in development or deployed, whether they are rated high-impact, and whether agency data was used to train them. No commercial sector offers anything comparable.

Where it breaks. Sensitivity exclusions remove a large fraction from public view, and in the case above that fraction was more than half. Some of those exclusions are certainly legitimate, since publishing the operational details of a fraud detection system tells fraudsters how to evade it. But the same exclusion that protects a legitimate secret also removes any external check on whether the system works.

And the entries that are published vary in quality. An earlier review of 23 civilian agencies produced 35 recommendations to 19 of them, with 15 needing to update their inventories to include required information at all. A field left blank is compliant in form.

The paradox is that transparency requirements produce their most complete data on the systems that matter least. A low-impact administrative tool gets a full public entry. A system used to build criminal cases does not appear.

What this means for anyone buying AI

Four transferable lessons, and none of them requires being a government agency.

Ask who owns the portfolio, not just each decision. If every deployment has a reviewer and nothing has an owner, you have case-by-case governance and no strategy, which is exactly what the audit found.

Ask whether your buyers can evaluate the purchase. Federal contracting officers could not find the specialists to assess proposals. Most organisations have the same problem and no auditor to name it.

Ask what happens to the people. A workforce plan is not bureaucratic decoration. A system nobody on staff understands is a system you cannot modify, validate or retire.

And write down what you learned. The single recommendation the auditors made was that agencies collect and share lessons. That is the cheapest intervention in this entire series and the one nobody does, because a record of what went wrong is uncomfortable to produce and has no immediate owner.

What is unresolved

Whether the sensitivity exclusion is being used proportionately. More than half of one inventory being withheld might be entirely appropriate or might be a convenient category. Nobody outside the agency can tell, which is the nature of the exclusion, and no independent review of the exclusions themselves exists.

Whether lessons-learned repositories work. The recommendation is sensible and the mechanism is untested. Repositories of institutional knowledge have a poor record of being read, and the incentive to contribute a candid account of a failed procurement is weak.

How to procure something that changes during procurement. Nobody has a good answer for buying a capability on a multi-year cycle when the capability is redefined annually. Shorter contracts trade one problem for another, since they reduce the leverage to negotiate data and intellectual property terms.

And whether inventories will ever be complete. Three domains now show the same result. Either the definitional question gets solved, which nobody has managed, or inventories are permanently understood as partial and read accordingly.

The counter-argument

Publishing an incomplete inventory is enormously better than publishing none. No private sector organisation discloses anything comparable, and criticising government transparency for being imperfect while commercial deployment is entirely opaque gets the comparison backwards. The reason these failures are visible is that someone is required to look.

The auditor's job is to find problems. A report saying an inventory was incomplete and a workforce plan was missing is what an audit produces when it works. Reading it as evidence of dysfunction rather than of functioning oversight mistakes the finding for the condition, and the same agency running 126 documented use cases is doing more disclosure than any comparable private organisation.

Sensitivity exclusions are frequently correct. Publishing the details of systems used in criminal investigations or fraud detection would materially damage their function. The tension between transparency and operational security is real and old, it long predates AI, and there is no version of the requirement that resolves it.

And procurement being slow is partly a feature. The acquisition rules that make government buying cumbersome exist because public money was previously spent badly. A faster process that bought AI more efficiently would also buy bad AI more efficiently, and the audit's concern was that funds are moving through processes not designed to evaluate them, which argues for better evaluation rather than faster buying.

The short version

Government publishes what every other domain conceals. Federal agencies must maintain and publish AI use case inventories, and two auditor reports in 2026 showed what that produces.

One agency recorded 126 active use cases as of June 2025, of which 65 were not detailed publicly. Auditors then found the inventory still incomplete, missing tools that officials said were contracted to help build criminal cases.

This is the third domain in the series with the same failure. Hiring: employers decide their own scope, and 18 audits appeared across 391 employers. Finance: incomplete model inventory is the most common examination finding. Government: the inventory is a published legal requirement and tools were missing from it. Three regimes, three enforcement mechanisms, one gap, and the government case removes every alternative explanation. When an obligation depends on self-classification, the failure is structural rather than motivational.

A separate audit of 13 acquisitions across four departments found contracting officers unable to find the data scientists and engineers needed to evaluate proposals, difficulty establishing what systems would cost, acquisition cycles longer than the technology's rate of change, and no systematic sharing of lessons between agencies. Roughly $1.7 billion has been appropriated for federal AI, moving through processes not designed to assess it.

And the finding that generalises furthest: several entities had oversight of individual use cases, and none was responsible for the portfolio. Every deployment had a reviewer and the aggregate had nobody, with no process for asking whether 126 use cases were the right 126. That is the difference between reviewing decisions and having a strategy, and it is the most common failure in any large organisation adopting any technology.

Common questions

Does the government publish what AI it uses? Yes, more than any other sector. Federal agencies are required to maintain and publish AI use case inventories listing systems in development and deployment, whether they are rated high-impact, and whether agency data was used in training. One agency's inventory recorded 126 active use cases as of June 2025. Sixty-five of those were not detailed publicly, and auditors found the inventory still incomplete.

What did the GAO find about federal AI procurement? An audit published in April 2026 examined 13 AI acquisitions across four departments and identified six recurring challenge categories appearing regardless of mission, vendor or technology. Contracting officers could not find data scientists or machine learning engineers to evaluate proposals. Costs were difficult to establish. Acquisition timeframes did not match the technology's rate of change. And agencies were not sharing lessons learned, so each was learning the same expensive ones independently.

Why are AI inventories incomplete? Because somebody has to decide what counts as an AI system, and that somebody is always the party being inventoried. The same failure appears in hiring, where employers determine their own scope, and in banking, where incomplete model inventory is the most common examination finding. The government case is the clearest demonstration, since there is no commercial incentive to conceal, the requirement is explicit and the format is specified, and the result is still incomplete.

How much is the US government spending on AI? Roughly $1.7 billion has been appropriated for federal AI efforts. For scale, industry investment in AI development was reported at over $250 billion in 2024 alone. The auditors' concern was less about the amount than the route: those funds move through procurement processes that were not designed to evaluate what they are buying.

Who oversees government AI use? Within one audited agency, several entities had oversight of individual use cases and none was responsible for managing AI investments across the agency. There was no process for ensuring investments contributed to agency-wide goals. That distinction, between reviewing each decision and owning the portfolio, is the most transferable finding in the report and applies to any large organisation.

Why are some government AI systems not listed publicly? Inventories permit exclusions for sensitivity, and in the audited case more than half of the entries were not publicly detailed. Some exclusions are clearly appropriate, since publishing the operational details of a fraud detection system would tell people how to evade it. The difficulty is that the same exclusion protecting a legitimate secret also removes any external check on whether the system works, and no independent review of the exclusions themselves exists.

What should an organisation learn from government AI procurement problems? Four things. Ask who owns the portfolio rather than each decision, since case-by-case governance produces a defensible record for every case and no answer about the aggregate. Ask whether your buyers can technically evaluate the purchase. Ask what happens to the people, since a system nobody on staff understands cannot be modified, validated or retired. And write down what you learned, which is the cheapest intervention available and the one almost nobody performs.

Is government AI adoption growing? Yes. Agencies reportedly more than doubled their use of AI between 2023 and 2024, and use spans veteran services, weapons systems, administrative work, facial recognition at airports and analysis of benefit claims. The auditors' finding was not that adoption was too fast in itself, but that it was outpacing the procurement frameworks meant to evaluate it, with lessons from each expensive mistake staying inside the agency that made it.

Learn the concepts

← All posts