Agent Assurance

Installment 5 · 5 August 2026

Chapter 5: The Economics of the Control Layer

Price

If agents are getting cheaper and more capable every quarter, it is reasonable to assume that assurance is a temporary tax, a scaffolding that comes down once the models are good enough to trust. A reader running a business on these tools has probably had the thought already: this is early, it will mature, and the checking will get cheaper along with everything else.

Chapter 4 described the control layer, the standards, inspection, insurance, and audit layer that forms around a new capability after its incidents. This chapter gives its economics. The assurance work is durably worth paying for, and not a transitional service that disappears when the technology settles down.

The complements law

Economists call two goods complements when using more of one makes the other more valuable. A printer is worth more when computers are cheap; a road is worth more when cars are common. The standard textbook definition is exactly this: a complement is a good whose consumption increases the value of another good.1 The consequence follows directly. When the price of a good falls, demand rises not only for that good but for the things used alongside it, and the scarce complement is where the value pools. The same textbook gives the case that everyone has lived through. The dramatic fall in the price of computers over recent decades, it notes, significantly increased the demand for printers, monitors, and internet access.1 The computer got cheap. The money moved next door.

The applied version of this idea belongs to Carl Shapiro and Hal Varian, whose 1999 book Information Rules argued that the classical economic laws still govern markets in information goods, however novel the technology looks. Their summary of their own thesis: technology changes, economic laws do not.2 The complements law is one of those laws, and it is the one that reads the present most cleanly. To know where value goes as a capability commoditizes, find its scarce complement and look there.

What good is collapsing in price, and what is its scarce complement?

The good that is collapsing in price

The good is machine intelligence, priced by the token, and its price is already falling. Gartner's estimate is that by 2030 the cost to providers of running a query against a one-trillion-parameter model will be more than ninety percent lower than it was in 2025. That would make such models up to a hundred times more cost-efficient than the earliest models of that size built in 2022.3 Whatever the exact figure turns out to be, the direction is not in dispute and the slope is steep. The raw capability at the center of this book is getting cheaper on the curve that computers, and before them electricity, followed.

Cheaper per token does not mean cheaper in total. The same Gartner analysis notes that agentic models, the ones that plan and act across many steps rather than answering a single prompt, consume between five and thirty times more tokens per task than a plain chatbot. The falling provider cost per token will not be fully passed through to the customer.4 The unit price of intelligence is dropping fast and the amount of intelligence a business consumes is rising faster. Firms' bills are going up while the price of the thing on the bill goes down: in April 2026, a practitioner newsletter reported token spend up roughly tenfold over six months at two companies, and one seed-stage AI infrastructure firm's spend per developer up from about $200 a month to about $3,000.5 That is normal for a commoditizing input. It is exactly what happened to bandwidth and to storage, and it is the first sign that the value is moving somewhere else, because a good whose consumption explodes as its price falls is behaving like a complement to something scarcer.

What stays scarce

As machine intelligence commoditizes, three things stay scarce alongside it: accountability, verified context, and standing. They are the scarce complements of the agent economy, and the assurance work sits on top of all three.

Accountability puts a named person on the hook when the work is wrong. Verified context is information the agent can be trusted to have acted on, checked against a source rather than generated by the model itself. Standing is what a practitioner earns when their sign-off starts to mean something to a bank, an insurer, or a regulator. None of the three is a capability the model can be asked to supply, and the reason is the same in each case. The model can produce the appearance of all three at close to zero cost. The appearance is the thing that has no value.

The objection is that AI commoditizes accountability and standing too. If the model can write the audit memo, reason through the judgment call, and produce the confident sign-off, then judgment itself is getting cheap, and the scarce complement collapses along with everything else. The objection is serious, and a 2026 analysis in the economics of science made exactly this move against the prevailing view. The prevailing view, associated with the economists Agrawal, Gans, and Goldfarb, had been that AI drives the cost of prediction to nothing while judgment remains the expensive complement. The 2026 paper argues that this fails once the model produces competent-looking judgment, including selecting, ranking, attributing, and certifying, at a marginal cost approaching zero.6 Competent-looking judgment is now abundant. If judgment were the scarce complement, the scarcity would be gone.

Judgment is not the scarce thing. What stays scarce when competent-looking judgment becomes free is the narrower set of things the appearance cannot fake. The same 2026 paper names four for science: verified signal, legitimacy, authentic provenance, and what it calls integration capacity: how much AI-delegated judgment a field will accept before it stops trusting the venues that admitted it.6 Three of the four are recognizable in the agent economy. Verified signal is verified context: a fact checked against an independent record rather than produced by the model. Legitimacy and authentic provenance are standing: a source whose sign-off a counterparty accepts because it was earned in public and can be withdrawn. Underneath both sits accountability, a named person who answers for the work. A model can generate a memo that reads exactly like accountability, but it cannot be the person whose name is on it and who bears the consequence, because accountability is a relationship, not a document. A model can assert a fact with total fluency, but it cannot make the fact verified, because verification is a check against an independent record and the model is not one. A model can imitate the house style of a firm with standing, but it cannot transfer the standing. What commoditizes is the performance of assurance. What stays scarce is the part that a counterparty can test, hold, and if necessary sue.

One 2026 theoretical model of AI and inequality concludes that value does not disappear as skills commoditize. It shifts toward complementary assets the model cannot replicate: proprietary data, computational infrastructure, distribution, and organizational routines. Its authors are careful to say the work is a mechanism and not yet an empirical finding, and it is cited here as a model, not a measurement.7 A second 2026 paper, on where value settles in vertical AI markets, finds that durable value capture concentrates in what it calls cospecialized accountability assets: professional sign-off, regulated workflows, and evidence trails.8 A third, on the boundaries of agentic firms, finds that the tendency to break work into cheap swappable components weakens as the cost of verifying the output rises.9 These are early papers, none of them yet peer-reviewed, and none written about agent assurance. That is what makes them useful. Independent people, modeling the general economy of cheap AI, keep arriving at a list of scarce complements that reads like the definition of this discipline: sign-off, verified evidence, a workflow someone stands behind.

The insurance market has started to price this split. One large commercial carrier introduced in 2025 what coverage lawyers described as the first absolute AI exclusion, removing from its directors-and-officers, errors-and-omissions, and fiduciary liability lines any claim arising from any actual or alleged use, deployment, or development of artificial intelligence by any person or entity.10 The specialist carriers of Chapter 4 moved the other way. The first agent insurance policy backed by a certification standard went live in February 2026, and one specialist structure pairs its dedicated AI limits with model certification and compliance reporting.11 A general carrier writing the risk out of its standard lines while specialists write it in is the market sorting agent risk into a class of its own.

When intelligence collapses in price, the value migrates to accountability, verified context, and standing. Those do not commoditize. The scarce part of each is a person on the hook and a record that can be tested, not an output that can be generated. Assurance work is the work of supplying exactly those. Its value is anchored on the scarce side of the ledger.

The market of one

Most business logic has never been worth writing custom software for. The reason is arithmetic. Bespoke software has a fixed cost that does not care how many people will use it, and for a single firm's specific way of doing a specific thing, that cost has almost always exceeded what the firm could justify spending. The ideal case, software written for exactly one customer's exactly one workflow, is a market of one, and a market of one was never viable, because the build cost was set by developer time and developer time was expensive.

A 2020 systematic review synthesized 132 separate studies of information-technology adoption in small and medium enterprises and identified eighteen distinct barriers, with cost and in-house expertise recurring across them.12 An OECD survey in 2024 found that one in four small and medium firms cite the cost of digitalization as a reason for not adopting digital tools at all. About the same share say they cannot see the benefit, and the firms furthest behind most often report that they do not know where to start.13 The picture is consistent across decades and across countries. For most firms, most of the time, the specific software they would actually want has cost more than it was worth, and so it was never built.

What those firms did instead, they did in a spreadsheet. The Spreadsheet Engineering Research Project at Dartmouth's Tuck School surveyed 846 business-school alumni, a sample its authors noted was larger than any previously reported on the subject. Seventy-nine percent rated spreadsheets very important or critical to their job, and only about one percent called them unimportant. Written standards for how those spreadsheets should be built existed in fewer than one in ten of their organizations.14 A separate 2023 study extracted 21,398 United States job postings and found spreadsheet skills required in roughly half of them.15 The business runs on spreadsheets. It runs on them where the custom software was never affordable, which is almost everywhere the work is specific to the firm.

These spreadsheets are not benign. An audit of fifty real-world operational spreadsheets found that ninety-four percent of them contained errors.16 That figure belongs to Chapter 1 as much as to this one. The improvised system that carries the business also carries defects that no one has looked for. The person who built it is running a piece of critical infrastructure that was never designed, reviewed, or handed over. But the same figure, read from the demand side, says something else. Ninety-four percent of these things are broken and the business depends on them anyway, which means the need they meet is real enough that a flawed answer is better than no answer. That is a market signal of unusual clarity.

The line moves

If the cost of building custom software falls far enough, the market of one becomes viable, and the enormous backlog of firm-specific work that has been living in broken spreadsheets converts into software worth building. The question is whether agent-based production moves the line that far. By early 2026 the movement itself is measured; what stays open is how far it reaches into existing systems, and who can tell.

The clearest measurement is of capability. The research group METR scores frontier agents by the length of software task they can complete unassisted, indexed to how long the task takes a skilled human. By January 2026 the leading model's fifty-percent horizon, the task length it finishes about half the time, stood at roughly five hours of human work, and the horizon had been doubling about every three months since 2024, against six to seven months across the record since 2019. METR states the caveats itself: the intervals are wide, and the longest tasks lean on estimated rather than measured human baselines. The direction and the slope survive the caveats.17

The productivity experiments started earlier, on earlier tools, and their dates matter. In 2023 and 2024, controlled studies of the tools that preceded today's coding agents found speedups from roughly twenty-one to fifty-six percent on defined tasks, with wide intervals;1819 the pooled analysis found the largest gains among the most junior developers.20 McKinsey's 2024 estimate put the direct effect at twenty to forty-five percent of what firms spend on software engineering.21 A careful Google study from the same period found no statistically significant change in objective developer telemetry even as developers' liking for the tool rose.22

The hardest setting told a different story, twice. In mid-2025, METR ran a randomized trial with sixteen experienced open-source developers working on mature codebases they had spent an average of five years inside. Allowed to use AI tools, they took nineteen percent longer. They had forecast a twenty-four percent speedup, and after finishing, having been slower, they still estimated the tools had sped them up by twenty percent.23 The developers could not tell. In early 2026 METR re-ran the experiment with newly recruited developers and the agentic tools that had arrived in between, and the slowdown was gone: four percent slower as the point estimate, inside an interval running from fifteen percent slower to nine percent faster. The re-run also strained the experiment's own design. Thirty to fifty percent of invited developers declined to submit tasks they would have to do without AI, and METR is redesigning the study because that refusal biases the speedup estimate downward. The point estimate moved toward zero; the behavioral fact moved further: up to half the developers METR invited now refuse to do some of their work without the tools.24

That forecast-and-recall gap is Chapter 2's problem again. The verification gap said that a client cannot evaluate agent-performed work, before delivery or after. Here a version of it closes over the practitioner's own judgment of their own productivity. If a senior engineer cannot verify whether the tool made them faster on their own code, the market cannot be trusted to price the tool by results. That is another reason the checking function does not commoditize. It is needed even to know whether the cheap thing worked.

What survives is a claim narrower than the hype and stronger than the 2023 and 2024 studies alone would support. The build cost of software has fallen furthest where the work is greenfield and specification-shaped rather than deep inside a mature system a person already understands. The market of one is exactly the greenfield, specification-shaped case: a single firm's specific workflow, described from scratch, built once. The one measured slowdown sat at the opposite pole, experienced maintainers deep inside mature systems, and on the latest measurement even that has faded to noise. The line has moved through the region where the backlog sits, and the territory is open.

The territory is growing from the other end as well. The same tools that lower the cost of building software also lower the cost of building the improvised system that stands in for software. A business that used to run a process in someone's head or over the phone can now, with an afternoon and a chatbot, run it in a spreadsheet or a small script. No survey yet measures it, but the mechanism runs in only one direction: the same capability that builds software manufactures shadow systems, at the same time as it makes them cheaper to replace, converting conversation-run corners of a business into spreadsheet-run ones, which then become demand for something built properly. The backlog is being filled from both directions at once: old unbuilt work becoming affordable to build, and new improvised work arriving ahead of any cleanup. The addressable territory is expanding, not shrinking, which is the opposite of what a transitional service faces.

The shadow artifact is the requirements document

The improvised spreadsheet, the unsanctioned tool, the blessed document that the whole team depends on and no one owns: the book calls each of these a shadow artifact, the improvised system built ahead of governance. Chapter 1 read the shadow artifact as a risk, the thing that produces the audit cliff when an outsider finally asks how it works. This chapter reads it as a signal, and the two readings are the same object seen from the risk side and the demand side.

A shadow artifact is the most accurate requirements document the client will ever produce. It has to be, because of how it was made. No one wrote a specification for it, argued about scope, or built what a committee imagined the users wanted. Someone with the actual problem built the smallest thing that solved it, and then kept using it because it worked, and changed it whenever reality changed, for years. Every column that exists, exists because it was needed. Every manual step that survives, survives because automating it was harder than doing it by hand. The artifact encodes exactly what the firm needed and could not otherwise get built, with all the specificity that a market of one requires and a written brief always loses. A requirements document commissioned fresh would cost more and would be worse, because it would be a guess about the work, and the spreadsheet is the work.

The firm has already specified what to build and what to control. The ninety-four percent error rate is not a reason to distrust the signal; it is part of the signal, because it marks the artifact as load-bearing and unsafe at once, which is the exact condition that the assurance work exists to address. The engagement that follows does not start from a blank page. It starts from the shadow artifact, preserves the workflow it encodes, and re-founds it on controlled substrate that can be inspected and stood behind. That engagement is described as method in Chapter 6, and the controls it installs are the subject of Part III. The work has a source, and the source is inexhaustible for as long as businesses keep improvising ahead of governance, which the arrival of cheap agents has made faster than ever.

Why the work is durable

Assurance work is durably valuable, and not a transitional service, for two independent reasons, either of which would be enough on its own.

The first is the complements law. The value in the agent economy migrates to accountability, verified context, and standing, and those stay scarce for the reason already given: the appearance can be generated and the substance cannot. As intelligence gets cheaper, the money moves toward the checking, the signing, and the standing behind, which is the assurance function by another name. This is not a claim that the work is safe from competition; it is a claim about which side of the commoditization line it sits on, and it sits on the scarce side.

The second is that the territory is growing. The unbuilt-software backlog was held for decades in spreadsheets, because a market of one was never affordable. It is converting into buildable work as agent production moves the cost line through exactly the region where the backlog waits. Meanwhile, by the same mechanism, the tools keep manufacturing new shadow artifacts ahead of any cleanup. More systems built means more systems to assure. A service tied to a shrinking need is transitional; this need is expanding on both ends.

The durability of the work does not rest on a profession forming, and Chapter 4 declined to promise that one will. It rests on the risk being irreducible. A problem that gets solved becomes a cheap tool and its assurance evaporates, the way the boiler code turned boiler safety into a settled matter. This risk does not get solved, because what remains after every control is the residual: an ecosystem of models, tools, and attacks that changes without end, and actions that reach the world and cannot be recalled. A risk that must be managed forever sustains work that lasts forever, and the complements law explains where the money for that work comes from: it comes from the collapsing price of the intelligence the risk attaches to. The cheaper the agents get, the more of them a business runs, the more work reaches the world unverified, and the more valuable the scarce complement becomes. The same force that commoditizes the capability funds the discipline that governs it.

What neither the pattern nor the price can say is who actually pays for the work, and what that buyer believes they are buying. That is the commercial half of the discipline, and it is next.

Notes

  1. Open economics textbook, Introduction to Economic Analysis: "For a given good x, a complement is a good whose consumption increases the value of x," and, as its worked example, "the dramatic fall in the price of computers over the past 20 years has significantly increased the demand for printers, monitors, and Internet access." The general mechanism, not an incident. Ledger: ch05-e01.
  2. Carl Shapiro and Hal R. Varian, Information Rules (Harvard Business Press, 1999), whose thesis is that classical economic laws continue to govern information-good markets: "Technology changes. Economic laws do not." Cited for the applied lineage of the complements argument, in reference-text register. Ledger: ch05-e02.
  3. Gartner press release, March 2026: "By 2030, performing inference on a large language model (LLM) with 1-trillion parameters will cost GenAI providers over 90% less than it did in 2025," and up to 100 times more cost-efficient than the earliest models of similar size developed in 2022. The publisher blocks direct retrieval (HTTP 403); cited from the salvage extraction. Ledger: ch05-e03.
  4. The same Gartner release: agentic models "require between five- and 30-times more tokens per task than a standard GenAI chatbot," and "falling GenAI provider token costs will not be fully passed on to enterprise customers." Ledger: ch05-e04.
  5. Gergely Orosz, "The Pulse: token spend breaks budgets," The Pragmatic Engineer newsletter, April 2026, reporting practitioner accounts: at two companies "token spend has increased by ~10x in the last six months"; a seed-stage infrastructure founder: "Six months ago our spend per developer was ~$200/month. Today, it's around $3,000/developer/month." The same accounts report spend rising without clear productivity measurement, with one firm's performance reviews rating developers on AI adoption. Self-reported figures, cited as reports. Ledger: ch05-e20.
  6. Preprint, AI-Augmented Science and the New Institutional Scarcities (arXiv:2605.02566, 2026): competent-looking judgment "including selecting, ranking, attributing, and certifying, is now produced at scale at marginal cost approaching zero, inverting the dominant economics-of-AI reading that treats judgment as the scarce complement to cheap prediction"; the scarcities that replace it are named as verified signal, legitimacy, authentic provenance, and integration capacity, the last defined as "how much AI-delegated judgment a scientific community will accept before it stops trusting the journals, panels, and conferences that admitted it." The paper is about science; cited here as a parallel. Verified to exist and to state these claims by direct retrieval, July 2026. Ledger: ch05-e06.
  7. Preprint, When AI Levels the Playing Field (arXiv:2603.05565, 2026): economic value "shifts toward complementary assets that AI cannot replicate: proprietary data, computational infrastructure, distribution networks, organizational routines," with the authors stating "the contribution is the mechanism, not a verdict on the sign." Cited as a theoretical model, not an empirical finding; verified to exist by direct retrieval, July 2026. Ledger: ch05-e05.
  8. Preprint on value capture in vertical AI markets (arXiv:2605.17812, 2026): "durable value capture concentrates in cospecialized accountability assets: professional signoff, regulated workflows, evidence trails." Ledger: ch05-e08.
  9. Preprint on the boundaries of agentic firms (arXiv:2605.23179, 2026): "the positive relationship between agentic reductions in interface and assembly costs and adoption of a component strategy weakens as verification costs increase," The core relationship carried a clean verifier tally in the salvage; the paper's stronger accountability-boundary sub-claims split the verifiers and are not relied on here. Ledger: ch05-e07.
  10. W.R. Berkley, per a Hunton Andrews Kurth coverage analysis, National Law Review, May 2025: "the first so-called 'Absolute' AI exclusion," barring claims arising from "any actual or alleged use, deployment, or development of Artificial Intelligence by any person or entity" across the insurer's D&O, E&O, and fiduciary liability products. Ledger: ch05-e21.
  11. The specialist market is Chapter 4's: ElevenLabs, live with the first AIUC-1-backed agent insurance policy, February 2026; Chaucer and Armilla's Vanguard structure, with dedicated AI limits "supported by AI model certification, risk and compliance reporting." Not universal: Testudo markets an audit-free underwriting process. Ledger: ch04-e55, ch04-e57, ch04-e58.
  12. Systematic review, International Journal of Innovation and Technology Management, February 2020: "On the basis of 132 selected studies, the review identifies 18 barriers categorized according to internal and external parameters." Ledger: ch05-e09.
  13. OECD, SME digitalisation report, September 2024: "The cost of digitalisation is also advanced as a reason for non-use of digital tools by 1 in 4 businesses and about the same share (24%) indicate they do not see the benefits of using digital tools." Ledger: ch05-e10.
  14. Spreadsheet Engineering Research Project, Tuck School of Business at Dartmouth (survey work published 2005), survey of 846 business-school alumni: 79 percent rated spreadsheets "very important" or "critical," about 1 percent unimportant, and written standards existed in "fewer than 10% of the organizations represented in our survey." The authors note the sample is "quite a bit larger than any sample previously reported in the literature on this subject." Ledger: ch05-e11.
  15. Study of United States job postings (2023): "21,398 non-duplicated position descriptions" across 2019 to 2021, with Microsoft Excel named in 51 to 61 percent of postings by year and spreadsheet skills of some kind in roughly half. Ledger: ch05-e12.
  16. Reported in a 2008 paper citing Powell, Baker and Lawson (2007): "of the 50 real-world operational spreadsheets they audited, 94% contained errors." Cited as a secondary report of the underlying audit. Ledger: ch05-e13.
  17. METR, "Time Horizon 1.1," January 29, 2026: the leading model's 50%-time-horizon estimated at 320 minutes (interval 170 to 729) on a 228-task suite built to "capture skills required for research or software engineering"; doubling time 196.5 days across the full record and 88.6 days on the post-2024 trend. METR's own caveats: "these confidence intervals are still very wide," and most tasks above eight hours use estimated rather than measured human baselines. Ledger: ch05-e22.
  18. Controlled experiment with GitHub Copilot (preprint, 2023): "the treated group completed the task 55.8% faster (95% confidence interval: 21-89%)," the task being an HTTP server in JavaScript. Ledger: ch05-e14.
  19. Randomized controlled trial with 96 Google software engineers (preprint, 2024): "a roughly 21% increase in development speed attributable to AI," with the caution that the effect may not "apply more broadly, or ... translate across tools and over time." Ledger: ch05-e15.
  20. Pooled analysis of three randomized field experiments (preview, 2024): "when data is combined across three experiments and 4,867 developers, our analysis reveals a 26.08% increase (SE: 10.3%) in completed tasks among developers using the AI tool," with larger effects for more junior developers. Ledger: ch05-e16.
  21. McKinsey, 2024: "The direct impact of AI on the productivity of software engineering could range from 20 to 45 percent of current annual spending on the function." Cited as an analyst estimate. Ledger: ch05-e19.
  22. Mixed-methods Google study (preprint, 2024): "having access to the Gen AI tool made no statistically significant difference to developer telemetry metrics," while "developers' views on the trustworthiness of AI generated code remained unchanged." Ledger: ch05-e18.
  23. Randomized controlled trial of 16 experienced open-source developers, 246 tasks (preprint, 2025): "we find that allowing AI actually increases completion time by 19%." Developers forecast a 24 percent speedup beforehand and estimated a 20 percent speedup afterward. Verified by direct retrieval, July 2026. Ledger: ch05-e17.
  24. METR, "We are Changing our Developer Productivity Experiment Design," February 24, 2026: among newly recruited developers the estimated effect is "-4%, with a confidence interval between -15% and +9%," against the early-2025 finding of 19% slower; "30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI," and the redesign is because that selection now biases speedup estimates downward. Ledger: ch05-e23.