Agent Assurance

Installment 2 · 15 July 2026

Chapter 2: The Verification Gap

The missing question

A generated system arrives: the dashboard renders, the numbers populate, the form saves. The operator clicks through it, sees that it works, and accepts it. A generated document arrives: the memo is well organized, the tone is right, the citations are formatted correctly. The operator reads it, sees that it reads well, and accepts it.

The operator cannot check, was trained by every previous generation of software not to check, and in most cases never discovers that a check was called for. In neither case above does anything happen that a professional verifier would recognize as verification. No figure is traced to its source. No behavior is tested at a boundary. Nobody asks who, or what, checked this, because as far as the operator can see there is nothing to ask. Verification never presents itself as a step that exists.

The absence has a history, and it is not carelessness. For the entire lifetime of everyone now working, polished and working software meant an institution. A product that installed cleanly, handled its edge cases, and did not lose data existed only because a company had built it. A company with developers, a test function, a support desk absorbing complaints, a legal exposure, and a brand that would take the damage if the product failed. There was no other way to obtain working software. So the quality a user could see became a reliable proxy for the verification the user could not see, in the same way that a printed book, for centuries, implied a publisher, and the publisher implied that somebody had checked. The proxy held because production was expensive and institutional. Decades of daily software use trained it into everyone as a reflex: if it works, somebody stood behind it.

Generative AI broke that link. Working software, and fluent documents, and plausible analysis, now arrive with no institution of any kind behind them. The quality the operator can see no longer proves anything about the checking they cannot see. The reflex, though, is still installed, and it fires exactly as trained: the dashboard works, therefore it is the kind of thing that works, therefore somebody, somewhere, must have done the checking that working things get. Nobody has.

There is a partial exception in public awareness. Most operators have by now heard that chatbots hallucinate, and many apply real skepticism to a direct factual answer typed back in a chat window. But the skepticism is attached to the format, not the source. When the output stops being an answer and becomes an artifact: a working application, a formatted report, a spreadsheet with formulas that calculate, the skepticism evaporates. The artifact has crossed into the category that decades of experience filed under institutionally produced. The same operator who would double-check a chatbot's claim about a tax deadline will adopt, without a second look, a chatbot-built system that computes tax deadlines.

The other exception is the professional whose training is verification itself. A chartered accountant handed a generated reconciliation does not primarily see a clean layout; the training asks, by reflex, what this ties back to and who signed it. The reflex is an inherited discipline, drilled in through articles and audit files, and it marks out the one profession that reliably asks the missing question. The rest of this book is about how that reflex is built into agent work and kept there. It is a trained thing, not a talent; nobody handed the operator the training, and nothing in it is beyond anyone who runs a firm. An operator who never asks the verification question never sees a verification failure. The failures are there; unasked questions do not produce error messages.

The book calls this condition the verification gap: the client's structural inability to evaluate agent-performed work, before delivery, after delivery, possibly ever.

A third kind of good

Economics sorts purchases by when the buyer can judge them. Some goods can be assessed before purchase: a chair can be sat on in the showroom, and a buyer who does so knows nearly everything that matters. These are search goods. Some goods reveal themselves in use: a restaurant meal, a hotel stay, a new hire's first months. The buyer cannot judge in advance but knows soon enough afterwards. These are experience goods, and most of commercial life runs on them, because repeat purchase and reputation discipline the seller: cook a bad meal and the diner does not return.

For some goods the buyer cannot judge quality even after consuming them, because judging requires exactly the expertise being purchased. The economics literature calls these credence goods, and it states the problem plainly. The information asymmetry between buyer and seller persists even after the trade is concluded and the good consumed. Only the seller can diagnose what was needed and whether it was delivered, so the buyer is left relying on the seller's honesty.1 The standard examples are the ones every operator already knows from the buying side. The patient cannot tell whether the treatment prescribed was the treatment required. The motorist cannot tell whether the replaced part was faulty. The client cannot tell whether the contract's clause 14 protects them or merely fills a page. Everyone recognizes the feeling that goes with these purchases: paying the invoice on trust, because checking the work would require becoming the person who did it.

Agent-performed work belongs to this class on every count. Correctness: the operator who commissions a revenue dashboard cannot tell whether the figures are right, because the independent recomputation that would settle it is precisely the work the agent was engaged to remove. Security: chapter 1 recorded what the operator sees of a generated system, the screen, and what exists beneath it, the access rules, the hosting, the credentials. Whether that lower half is safe is not assessable from the upper half, at any level of inspection the operator can perform. Governance: whether the work respected the obligations the firm is answerable for, where the data went, what retained copies exist, is invisible in the deliverable by construction. The operator is in the mechanic's waiting room, with one aggravation the mechanic's customer is spared: a badly repaired car usually announces itself eventually. Much agent-performed work fails without ever announcing itself.

The fluency machine

The instinctive answer to the waiting-room problem is to inspect the deliverable more carefully. It is the wrong answer here.

Human work signals its quality imperfectly, but it signals. Prose that is organized and precise took effort and attention to produce, and effort and attention correlate, loosely, with the same qualities in the underlying analysis. Sloppy surface, sloppy work: the heuristic fails often, but it points in the right direction on average, and every experienced reader of business documents runs on it. A generative model breaks the heuristic in one direction only. Fluency, structure, confidence, and polish are the model's baseline output, produced identically whether the content is right or wrong. The surface is a constant. The correctness underneath is the variable. A constant carries no information about a variable, so the polish of agent output has no evidentiary weight at all about its correctness, its security, or its provenance. The model is a fluency machine, not a truth machine.

This would matter less if people could learn their way past it. The evidence is that they largely cannot. In a randomized trial published in 2026, forty-four physicians, every one of whom had completed twenty hours of AI-literacy training, diagnosed six clinical vignettes with a language model's recommendations available. The physicians were randomized to receive either accurate recommendations or a set in which three of the six had been deliberately made wrong. Those receiving the flawed set scored 73.3 percent on diagnostic accuracy against 84.9 percent for those receiving accurate ones; the study's adjusted difference was fourteen percentage points, in trained diagnosticians who had been specifically taught about this failure mode.2 The authors' conclusion was that literacy training does not immunize against confident, fluent error. If it does not immunize physicians reading within their own specialty, the operator reading a generated cash-flow model has no defense to speak of.

Nor does labeling close the gap. A recurring proposal is that AI-produced content be disclosed as such, so that readers can apply appropriate caution. In a randomized experiment with 800 participants, labeling content as AI-generated had no statistically significant effect on how accurate or credible readers judged it; what moved credibility was whether the content happened to be true.3 Disclosure of origin gave readers nothing they could use; there is nothing on the surface of fluent output for diligence to grip.

Nor is the gap closing from the model side. In a large public red-teaming competition reported in 2026, 272,000 attack attempts against thirteen frontier models found every model vulnerable to hostile instructions hidden in the content it processes, and a model's capability was only weakly correlated with its robustness.4 Whatever a newer, smarter model buys the operator, the record so far says it does not buy a system whose safety can be assumed from its intelligence.

Fluent, wrong, and filed in court

A court filing is read by opposing counsel paid to attack it and by a judge whose function is to test it; it is close to the only business document in existence whose readers are professionally obligated to verify it. If fabricated material passes even there, it passes everywhere else, because every other document is read with less care.

In Mata v. Avianca, in the Southern District of New York, two attorneys filed an opposition brief researched with a chatbot. The brief cited and quoted from six judicial decisions that do not exist. The fabrications were fluent in exactly the sense of the previous section: real-sounding case names, plausible citations in correct format, quoted passages of judicial reasoning that read as law. The filing attorney had them in hand, read them, and filed them. The court, after giving every opportunity to withdraw, imposed a 5,000 dollar sanction jointly on the attorneys and their firm.5 A trained professional inspected fabricated authority and found it satisfactory, because inspection was the wrong instrument.

The pattern did not stop with its most famous instance; it escalated. Two years later, in Johnson v. Dunn in the Northern District of Alabama, a court confronted five fabricated citations across two motions. It found that the reprimands and modest fines by then common for AI-fabricated authority were no longer a sufficient answer. It publicly reprimanded the responsible attorneys, disqualified them from further appearance in the matter, and referred them to the disciplinary authorities.6 By July 2026 the pattern had surfaced inside the judiciary itself: India's Supreme Court found that a company-law tribunal had itself relied on AI-generated, non-existent case law in its orders, orders then affirmed on appeal. The Supreme Court set the orders aside, restored the underlying petition, and directed the Bar Council of India to produce guidance and disciplinary measures.7 A tribunal is the verifier of record; here the verifier of record was itself the party that failed to verify. Three years into the most scrutinized document channel in professional life, with sanctions escalating and every lawyer on notice, fabricated authority was still passing readers whose job was to catch it.

In each one, the gap was eventually caught, because litigation is adversarial and someone on the other side is paid to check. The operator's generated systems and documents have no other side. Nobody is paid to check, and the chapter's opening section recorded that nobody unpaid thinks to.

Before, at, and after delivery

The gap closes off, in turn, each of the three moments at which a buyer of ordinary services has always been able to form a judgment.

Before delivery, the client of human work reads process. The associate is at a desk the partner walks past; the vendor produces a project plan and misses or makes its milestones; the profession's letters hang on the wall and stand for a training regime that can be looked up. None of this verifies the work, but all of it is observable, and buyers lean on it constantly. Agent work offers no such surface. The work happens in seconds, inside a session, through steps that leave no residue unless something was deliberately built to record them. There is no desk to walk past. Process evidence for agent work does not exist by default; it exists only where it was designed in, which is a fact chapter 9 takes up as a control.

At delivery, the client of human work applies the surface heuristic, and the previous section recorded its failure. The artifact is fluent whether or not it is right, so the one moment that feels like an opportunity to judge supplies nothing to judge with.

After delivery, the experience-good escape route is closed for the class of failures that matter most. Some agent failures do announce themselves: the site goes down, the form stops saving, and the operator learns the way a diner learns. The failures with consequence attached are the ones that surface downstream, disconnected from their cause, or never. The record from the adjacent world of quantitative systems shows how long that disconnection can run. In April 2007 a coding error was introduced into the model that a major quantitative investment firm, AXA Rosenberg, used to manage client portfolios; it disabled a key risk-management component. Senior management learned of the error in June 2009, and a senior official directed staff to keep quiet about it. Clients were told in April 2010, three years after the error began affecting their portfolios, and then only under the pressure of a regulatory examination; the account exists in public because the SEC's settlement order wrote it down.8 For three years, sophisticated institutional clients paid fees on a product whose central advertised property was impaired, and no amount of inspecting their statements could have told them, because the error lived in the process and the process was invisible.

That is the complete shape of the verification gap: no visibility before, no signal at, and afterwards a discovery lag measured in years, bounded only by whether an examiner, an adversary, or an accident ever forces the question. For the supervised firm there is at least the examiner. For the operator population of chapter 1, whom no supervisor is scheduled to examine, the after-delivery correction arrives only by accident.

The market's answer

Markets have been here before. The credence-goods literature is now half a century old, and it contains both a record of what unverified expert markets do to buyers and a body of evidence about what fixes them.

In a field experiment on the auto-repair market, undercover visits took a test vehicle with a prearranged set of known defects to garages across a matched design. Mechanics recommended entirely unnecessary repairs in roughly 30 percent of visits and missed at least one of the planted defects in roughly 80 percent. Customers who signaled they would be repeat business received no better repair recommendations and no better diagnosis.9 Reputation is the mechanism everyone assumes disciplines service markets. Reputation works when the buyer eventually learns the truth and takes their custom elsewhere; it is the experience-good remedy. In a credence market the buyer never learns the truth, so there is nothing to feed back, and the reputational discipline that keeps restaurants honest stops operating. The mechanism the operator instinctively relies on when engaging any expert, including the new synthetic kind, has no purchase here.

In a laboratory experiment with 936 participants, designed to isolate the determinants of efficiency in credence-goods markets, the researchers compared institutional arrangements. Among them verifiability, meaning the customer can confirm after the fact what was actually provided, and liability, meaning the expert is answerable for the consequences of what they provided. Liability had a crucial effect on market efficiency. Verifiability had, in the authors' words, at best a minor one.10 The result runs against the intuition that transparency about the artifact is what honest markets need. Improving the buyer's view of the work does little, because the buyer's view was never the problem. Making the practitioner answerable changes how the practitioner behaves, and that was the problem. The remedy operates on the producer, not the product.

The same result has now been reproduced with the subject of this book standing in both roles. A 2026 study populated simulated credence-goods markets with language-model agents as both experts and consumers, across six hundred one-shot markets. The markets largely broke down, except under a liability rule; the one other exception, experts endowed with preferences that favored the joint outcome over their own profit, sustained trade only at prices ruinous to the expert.11 The study is a preprint and has not been peer reviewed, and it simulates market institutions rather than observing them. It is cited here as an early indication, no more, that the classical structure of these markets survives the substitution of model for human on either side of the trade.

Outside the laboratory, the historical record shows what form the remedy has actually taken, wherever a society has had to trade at scale in work it could not verify. It verifies the practitioner instead. Medicine, law, accountancy, engineering. In every case the buyer's inability to judge the work was answered the same way: standards for who may practice, certification that they have met them, accountability structures that record who did what, and liability that makes the answer expensive to get wrong. The pattern is old enough to have left a trail in court archives. A study of 569 American state-court challenges to occupational licensing statutes between 1885 and 1911 found that courts struck down 46 percent of licensing regimes for non-professional occupations, where licensing mostly protected incumbents. They struck down only 8 percent for the professions, where the information asymmetry between practitioner and client was real.12 Courts, generally hostile to restraints on trade, let practitioner verification stand almost exactly where the credence problem was genuine. And the instrument is still doing the same work in the newest markets we have. On modern service platforms, a professional's displayed license and their customer reviews function as substitutes, with a first review measurably raising the hiring chances of unlicensed professionals and doing nothing for licensed ones.13 Buyers treat the license as verification already performed, which is what it is.

Whether licensing regimes improve the quality of the work delivered is contested in the economics literature, which has largely failed to establish that stricter entry barriers raise quality, and this book does not rest on any such claim.14 What the record establishes is narrower and sufficient. Verifying the practitioner is the form the market's response to credence goods takes. It is the one instrument that experimental evidence shows moving these markets, and the one that has held for a century wherever the asymmetry is real. Whether any particular regime executes it well is a separate question, and Part II returns to what separates the regimes that worked.

One more result completes the picture, because it answers the operator who concludes that the verification gap is the client's problem and the seller can simply carry on. Analysis of markets flooded with AI-generated content finds the classic adverse-selection result. When buyers cannot identify the origin or verify the quality of what they are offered, the perceived quality of everything in the market declines, the honest product along with the rest.15 An unverifiable market does not stay bad only for its buyers. It becomes a market where competence cannot be priced, because competence cannot be distinguished, and the competent seller loses their premium to the fluent one. On this evidence, assurance is necessary rather than optional, for the seller as much as for the buyer.

The definition

Agent assurance is the discipline of establishing justified confidence in work performed by AI agents, through controls, audit trails, standards, and accountable human ownership, on behalf of clients who cannot verify that work themselves.

Each clause answers a specific part of the gap.

Justified confidence, because confidence is the one thing the client already has; fluent output manufactures it on sight. The missing element is justification: grounds for the confidence that would survive the question who checked this, and how. The artifact cannot supply those grounds from its own surface, so they have to be established somewhere else, which is what the discipline's instruments are for.

Controls, because the before-delivery gap is an absence of constraint on process. Work whose quality cannot be judged afterwards must be shaped beforehand: what the agent may access, what it may do, what requires a second look before it takes effect. Chapter 3 gives controls their full treatment, as compressed incident history; they are the instrument that operates before delivery, where the client cannot see.

Audit trails, because the after-delivery gap is a discovery lag. The AXA Rosenberg clients waited three years for an examiner to force the question. A durable record of what the agent did, when, and to what, is the instrument that compresses that lag. It turns the invisible process into something that can be traced when the downstream symptom finally appears, and it turns a failure that might never have been discovered into one that is only discovered late.

Standards, because verification of the practitioner only scales when there is an external yardstick to verify against. A bespoke assurance is another credence good, and the client is no better off; an assurance stated against a published standard is checkable by a third party, which is the property that made every prior practitioner-verification regime work. Chapter 7 records which standards the discipline adopts.

Accountable human ownership, because the strongest single result in the credence-goods literature is that liability is what moves these markets, and liability requires someone to hold it. A named person, whose sign-off on the agent's work is recorded and load-bearing, is the discipline's implementation of the one instrument shown to work. It is also the clause that keeps the discipline honest about what it is not: not a promise that agents never err, but a guarantee that when they do, the error has an owner with something at stake.

On behalf of clients who cannot verify that work themselves, because this clause names the reason the discipline exists and the population it serves. It is a permanent clause. The inability is structural, and chapter 1 established that the condition is permanent: part of the risk cannot be closed by any tradition, only managed, and the capability does not stop changing. A response that waits for clients to become verifiers waits forever.

What remains, before the historical precedent of Part II, is the mental model the whole control set depends on: what a control actually is, and what any given checkbox is compressing. That is chapter 3.

Notes

  1. Uwe Dulleck and Rudolf Kerschbamer, survey of the economics of credence goods ("On Doctors, Mechanics, and Computer Specialists"), republished in CESifo Economic Studies, 2017. Ledger: ch02-e15. The survey, alongside Darby and Karni's 1973 paper, is treated in the literature as the definitional reference for the term. Ledger: ch02-e18.
  2. Randomized trial of physicians with completed AI-literacy training receiving deliberately flawed language-model recommendations, preprint August 2025, published in NEJM AI, 2026. Ledger: ch02-e22.
  3. Randomized experiment on AI-generated-content labels and perceived credibility, 800 participants, JMIR Formative Research, 2024. Ledger: ch02-e23.
  4. Public red-teaming competition on indirect prompt injection, reported March 2026: 464 participants, 272,000 attack attempts against 13 frontier models, 8,648 successful attacks, per-model success rates from 0.5 to 8.5 percent, with capability and robustness showing weak correlation. Ledger: ch02-e24.
  5. Mata v. Avianca, Inc., opinion and order on sanctions, S.D.N.Y., 22 June 2023 (Castel, J.). The brief was drafted with ChatGPT. Ledger: ch02-e08.
  6. Johnson v. Dunn, 792 F.Supp.3d 1241 (N.D. Ala. 23 July 2025) (Manasco, J.). The court declined to sanction the attorneys' firm, which had adopted a proactive AI policy: a firm-level control counted where the individuals' conduct did not. Ledger: ch02-e09.
  7. Supreme Court of India, judgment of 2 July 2026, setting aside the National Company Law Tribunal's order of 28 August 2024 and the appellate tribunal's affirmation of 11 September 2025 in the Essel Infraprojects insolvency matter; the fabricated authorities entered through the tribunal's own research. Ledger: ch02-e11.
  8. United States Securities and Exchange Commission, settlement announcement in the matter of AXA Rosenberg, February 2011. Ledger: ch01-e08.
  9. Undercover field experiment on the auto-repair market (Schneider), matched-pair design with a prearranged defect set. Ledger: ch02-e07.
  10. Uwe Dulleck, Rudolf Kerschbamer, and Matthias Sutter, laboratory experiment on the determinants of efficiency in credence-goods markets, 936 participants, American Economic Review, 2011. Ledger: ch02-e16.
  11. Simulation of credence-goods markets populated by GPT-5.1 agents as experts and consumers, arXiv preprint, March 2026, not peer reviewed. Ledger: ch02-e12.
  12. Study of judicial review of occupational licensing statutes, 569 state-court cases across 17 occupations, 1885 to 1911, Journal of Economic History, 2023. Ledger: ch02-e04.
  13. Study of occupational licensing stringency and outcomes on a large services platform, all fifty US states (Farronato, Fradkin, et al.), 2024. Ledger: ch02-e05.
  14. OECD, firm-level analysis of occupational entry regulations, 2020 (the empirical literature has largely failed to find that stricter entry barriers improve service quality). Ledger: ch02-e03. The same assessment in the law-and-economics review literature. Ledger: ch02-e06.
  15. De Cooman, comparative analysis of large-language-model regulation, Cambridge Forum on AI Law and Governance, 2025, applying the adverse-selection ("lemons") model to AI-generated content. Ledger: ch02-e14.