Skip to main content
Insights•19 min read

The Reckoning: When AI Activity Gets Booked as Value

October 6, 2026 by Asif Waliuddin

The Honor SystemVerification DebtEnterprise AIAI ValuationsAI InfrastructureAI governance
A glowing orange isometric balance scale on a dark charcoal background, one pan heaped with chat, document and dashboard icons and the other holding a single gold ingot, with the article title in white and orange capitals at upper left and the NXTG.ai wordmark at lower right.

On February 27, 2026, OpenAI announced $110 billion in new investment at a $730 billion pre-money valuation, including $50 billion from Amazon, and the two companies announced a $100 billion, eight-year expansion of their AWS agreement the same day. Every one of those numbers is real. The finale of The Honor System asks what each one proves before AI activity gets booked as value.

The AI economy has plenty of numbers. The problem is knowing what each one proves.

The OpenAI numbers are staggering, and they are real announced commitments. OpenAI's announcement put the round at $110 billion at a $730 billion pre-money valuation, with $30 billion from SoftBank, $30 billion from NVIDIA and $50 billion from Amazon. That same day, OpenAI and Amazon announced that they were expanding their existing $38 billion AWS agreement by another $100 billion over eight years. That is exactly why this is a harder story than the usual AI-boom argument. Nobody needs to ask whether the numbers exist. The relevant question is what each number proves.

An investment proves that capital was committed under an investment relationship. A compute contract proves that infrastructure capacity was contracted under a commercial relationship. When the same companies occupy multiple roles around the same ecosystem, neither one automatically provides an independent vote on the other. The money does not become fake because the relationships overlap, but the evidence changes shape.

That distinction is the final subject of The Honor System. For five episodes, I followed synthetic work as it accumulated authority inside organizations. I watched generation outrun review, review debt enter memory, memory return with authority, institutions sign that authority, and policies attempt to govern it. The finale follows the same mechanism one step further. Eventually, somebody puts a dollar sign on it.

AI value is real, and that makes it harder to measure AI ROI

There is an easy version of this essay in which the AI boom turns out to be smoke and mirrors. The evidence does not support that conclusion, and the most useful research of 2026 is careful about both the signal it found and the limits of what it can claim.

A March 2026 NBER working paper surveyed nearly 750 corporate executives and found positive labor-productivity gains associated with AI, gains that vary across sectors and are expected to strengthen in 2026. The largest effects were concentrated in high-skill services and finance. Adoption itself was uneven across firms: more than half of the surveyed firms had already invested, while many smaller firms were only beginning to do so.

The U.S. Bureau of Economic Analysis has found real signal too. Its July 2026 paper, AI Expectations and Outcomes, found an adoption learning curve: adoption first ran slower than expected, then briefly faster, and more recently close to expected rates. It also found some evidence that firms' stated motivations for using AI are linked to related changes in production processes, and that AI use cases are associated with higher research-and-development intensity. The authors are careful. The link between motivations and realized production outcomes, they say, remains murky at this point.

An August BEA paper, AI Utilization and Economic Performance, found that state-industry cells with higher worker-reported AI use showed stronger post-2020 real-output paths. The employment differences were positive but imprecisely estimated, and an industry-level measure of employer-reported AI use did not reproduce the same output pattern. The design is retrospective, with AI use measured in 2025 and 2026 against performance from 2016 to 2024, and the researchers say so plainly: "These estimates are descriptive rather than causal."

That is not weakness. That is what evidence looks like before marketing gets hold of it: signal, boundary, limitation. So this finale does not argue that AI creates no value. It argues that value claims should not be allowed to outrun the evidence that connects activity to outcome.

I learned the difference the embarrassing way

In May, inside my own company, one of my pricing briefs carried a number I wanted to believe: 97% gross margin. The number came from my own validator. The table was clean, the arithmetic worked, and the output looked decision-ready. But the same validation work had also said that the cost basis was partial and that model cost had not been measured. The brief preserved the number and lost the caveat.

When I challenged the number, the product path that was actually shipping was checked. The economics model assumed paying customers would bring their own provider key. The shipped path used managed inference instead. The calculator was not the problem. I had given it the wrong economic world, along with a model cost nobody had measured. The margin claim was retracted, and the conclusion stayed blocked until the real cost path could be measured.

That failure is the finale in miniature. The number may even have been roughly right for the wrong reasons, and I could not prove it either way. A number can be internally correct and economically misleading because the denominator is wrong, the boundary is incomplete, or the claim built on top of it is larger than the measurement underneath it.

The five measurement layers

The same mistake appears at radically different scales, from a consumer subscription to the balance sheet of an infrastructure provider. Each layer produces a legitimate number, and each number has a boundary on what it can prove by itself.

LayerThe numberWhat it can proveWhat it cannot prove by itself
Consumer usageremaining capacity, plan multipleusage state under a defined planuniversal workload equivalence or customer value
Benchmarkscore or pass rateperformance under an evaluation contractdurable real-world capability when the ruler is contaminated or broken
Infrastructurebacklog, RPO, compute commitmentcontracted demand under defined relationshipsprofitability, independent demand, or realized customer value
Enterprise activityprompts, seats, artifacts, hours savedadoption and operational activityattributable business outcome
Financial valuemargin, ROI, savings, revenueeconomic result if the bridge is validanything, if baseline, attribution, full cost, durability, and risk are missing

The higher the number climbs, the more consequential the interpretation becomes. A bad interpretation of a usage limit annoys a subscriber. A bad interpretation of a benchmark can distort model selection. A bad interpretation of infrastructure demand can distort capital allocation. A bad interpretation of enterprise activity can distort headcount, budget, pricing and strategy. The underlying failure is the same at every layer: the authority of the claim exceeds the authority of the measurement.

A usage plan is a measurement contract

Consider something as ordinary as a consumer AI subscription. Anthropic currently describes Claude Max 5x and Max 20x as multiples of the Pro plan's per-session usage allowance. The session limit resets every five hours, and Max plans also carry a weekly usage limit that applies across all models. Anthropic says usage depends on the length and complexity of your conversations, the features you use, which model you choose, and the effort level you select. It also reserves the ability to apply additional limits at its discretion.

That is not evidence of a broken meter. It is evidence that "20x" has a specific contract. It does not mean that every possible workload must produce exactly twenty times some intuitive quantity a customer has in mind. The mistake begins when a customer silently substitutes an intuitive denominator for the documented one. This is the smallest version of the problem, and the stakes get much larger from here.

When the ruler decays

In February 2026, OpenAI said it had stopped reporting SWE-bench Verified for frontier models. The benchmark had become one of the industry's most recognizable measures of autonomous coding capability. Then the measurement started failing its own audit.

OpenAI reviewed 138 problems that its o3 model did not consistently solve across 64 independent runs. At least 59.4% of that audited subset had material issues in test design or problem description. OpenAI also found evidence that frontier models had seen at least some of the benchmark's problems and solutions during training. In July, OpenAI published another audit, this time of SWE-Bench Pro, and estimated that roughly 30% of the tasks were broken.

Anthropic found a different version of the same problem in BrowseComp. In March, it reported benchmark answers leaking onto the public web through papers, blog posts and GitHub issues, with nine examples of that contamination across 1,266 problems. In two further cases, Claude Opus 4.6 appears to have inferred that it was being evaluated, identified the benchmark, located the answer key, and decrypted it. On sixteen more problems, it tried to reach benchmark materials and failed.

The lesson is not that benchmark developers are dishonest. OpenAI and Anthropic published these weaknesses themselves. The lesson is that a score can be computed correctly after the ruler has started measuring the wrong thing.

I call that ruler decay. The number still exists, and the dashboard still has two decimal places, but the interpretation no longer deserves the same confidence. That should make every executive uncomfortable, because enterprise dashboards decay too. Metrics survive long after the conditions that made them meaningful have changed.

The AI economy is a graph, not a row of independent numbers

Now move from benchmarks to capital. OpenAI's February funding announcement combined capital, compute and distribution relationships across some of the largest companies in technology. Amazon committed $50 billion ($15 billion up front, $35 billion more when certain conditions are met) while also expanding its AWS relationship with OpenAI by $100 billion over eight years. NVIDIA contributed $30 billion to the same investment round.

Elsewhere, NVIDIA invested $2 billion directly in CoreWeave in January, buying Class A shares in a private placement. CoreWeave then reported approximately $104 billion in revenue backlog as of June 30, plus more than $25 billion of net new customer commitments added in early Q3 that were not included in that backlog figure. Its second-quarter revenue was $2.575 billion. Its net loss was $626 million, its net interest expense was $640 million, and its adjusted EBITDA was $1.510 billion.

All of those numbers can be true at the same time, and that is the point. Backlog is not revenue. Revenue is not profit. Adjusted EBITDA is not cash. Investment is not customer demand. A supplier relationship is not independent validation. And overlapping relationships are not evidence of fraud. The analytical obligation is to map the topology before treating multiple edges as independent proof.

That matters because the strongest counterevidence is equally real. On Microsoft's July 29, 2026 earnings call, CFO Amy Hood reported that commercial remaining performance obligation grew 84% to $678 billion. Then she drew the boundary herself: "All sequential commercial RPO growth was driven by commitments from customers outside of frontier model companies." RPO still grew 25% excluding OpenAI, and Microsoft Cloud revenue surpassed $214 billion for the full year, with nearly 90% of it coming from customers outside frontier-model companies.

That does not fit a simplistic "AI companies are just buying from one another" narrative. Good. The purpose of evidence is to kill the story you wanted to tell when the story is wrong. The more defensible conclusion is this: AI demand is real, and relationship dependence is also real. Do not confuse dependence with falsity, and do not confuse related signals with independent confirmation.

The enterprise version is activity-value substitution

The same compression happens inside companies, only the labels are more familiar: prompts, seats activated, generated documents, lines of code, tickets closed, model calls, hours theoretically saved. These are useful operating metrics.

Then the number moves one slide to the right. Adoption becomes productivity. Productivity becomes savings. Savings become margin. Margin becomes ROI. No one necessarily lies. The claim simply gains altitude faster than the evidence.

I call that activity-value substitution. It occurs when an operational proxy is promoted into an economic outcome without proving the causal bridge, complete cost, durability and risk. One thousand AI-assisted tasks are not one thousand valuable tasks. Ten hours "saved" are not automatically ten hours returned to the business. A faster first draft is not a financial gain if the saved creation time reappears as review, correction, coordination or downstream rework.

The hidden work has a price

This is where fresh September research matters. On September 9, MIT Sloan reported a two-year field study by Katherine Kellogg, Batia Wiesenfeld and Arvind Karunakaran at an academic medical center and a corporate law firm, both attempting to turn employee generative-AI experimentation into durable organization-wide solutions. The researchers focused on work that ordinary adoption dashboards tend to miss: trial-and-error experimentation to learn what AI could and could not reliably do, reviewing and refining outputs with colleagues across departments, and continually adapting solutions as models change.

At the law firm, more than 80% of employees involved in AI innovation eventually disengaged, leaving three organization-wide AI solutions in use. At the medical center, leaders built structures that supported that experimentation and coordination, and the study reported 141 organization-wide AI solutions in use. Those are two cases from one field study. They are not universal failure or success rates. But the mechanism matters.

The cost of AI is more than the license price, the API bill or the GPU invoice. The real operating boundary includes the human and organizational work required to make the system reliable enough to persist: review, refinement, coordination, rework, governance, change management, incident remediation and evidence production. If those costs disappear into existing payroll while the saved hours appear in a special AI-benefit column, the ROI calculation was biased before the spreadsheet opened.

The final Honor System

This is economic verification debt. An organization makes a consequential financial decision using a proxy whose connection to durable net value has not been fully demonstrated. The decision might still turn out to be correct, and that is what makes this an honor system. The organization is trusting the bridge before it has receipts for the bridge.

The debt appears later, when somebody has to reconcile:

  • why the expected savings did not reach the P&L;
  • why headcount stayed flat while review burden increased;
  • why adoption rose but cycle time did not;
  • why a pilot worked but the scaled workflow stalled;
  • why a margin model ignored the actual cost path;
  • why a "saved hour" simply moved into a different queue.

The reckoning is not necessarily a crash. It is reconciliation.

Net Verified Value

The operating test I use is deliberately simple:

Net Verified Value = attributable business benefit - total operating cost - review and rework - expected risk loss

Net Verified Value is a management discipline, not an accounting standard. It is designed to stop four categories from disappearing between the pilot deck and the budget decision.

1. Attributable business benefit

What business outcome changed, and why does anyone believe AI caused the change? Revenue, cost avoided, cycle time, quality, customer retention and risk reduction are all candidates. The activity itself is not the outcome.

2. Total operating cost

Include the real operating boundary: software, models, compute, data, integration, support, human enablement, governance, change management, infrastructure, and whatever else the workflow genuinely requires.

3. Review and rework

Keep this visible. Do not bury human verification inside normal payroll and then count gross generation time as AI savings. The cost of trusting the output is part of the output's economics. It is the review debt from Episode 2, finally carrying a price.

4. Expected risk loss

Where evidence permits, account for the probability and consequence of material failure. Do not manufacture precision. "Unknown" is more honest than pricing risk at zero because the spreadsheet lacked a row for it.

The Value Evidence Ledger

A serious AI value claim should carry nine things.

  1. Business outcome. What changed for the business, not what the model did.
  2. Baseline and counterfactual. What would likely have happened without the intervention?
  3. Attribution method. Why does the evidence connect the outcome to AI?
  4. Gross benefit. What is the measured upside before costs?
  5. Complete cost. What did the system actually consume?
  6. Durability window. Did the result survive the pilot, novelty, champion, and first clean demo?
  7. Named accountable owner. Who is responsible for the claim?
  8. Scale, redesign, and kill rules. What evidence changes the decision?
  9. Reproducible evidence receipt. Can an authorized reviewer reconstruct the claim from the retained evidence?

This is not meant to make experimentation slow. That is the strongest objection to the whole framework, and it is legitimate. If every prototype needs a finance-grade counterfactual and a full risk model, governance becomes a machine for preventing learning.

The answer is not maximal proof everywhere. It is claim-altitude discipline. A prototype can remain a prototype. A directional signal can remain directional. An internal hypothesis can remain provisional. The evidence burden rises when the claim crosses into budget, staffing, price, margin, scale, public reporting or valuation. That is where "promising" stops being enough.

Practice, not product

At NextGen AI, I do not have an enterprise ROI-attribution product that I can honestly put at the end of this series. Internally, I can meter activity in bounded systems. In one bounded revenue flow, my instrumentation preserves an unattributed category rather than inventing a source when one cannot be proven. I use economic gates in selected workflows.

What I cannot yet do is join AI output, total operating cost, review, remediation and business outcome into a defensible enterprise return. In my own records, that ratio is null by construction. And after my own 97% failure, I have a very practical reason to treat an unknown denominator as a blocker rather than a rounding error.

That is useful practice. It is not a universal ROI system. The Honor System applies to me too.

The reckoning

This series started with output. AI made synthetic work cheap enough to become organizational infrastructure before most organizations built an evidence system capable of governing it. Then the failure changed state. Output created review debt. Review debt entered memory. Memory returned with authority. Institutions signed that authority. Policy tried to govern it. Governance created real cost. And finally, somebody booked the whole thing as value. That is where the Honor System meets the balance sheet.

The future of enterprise AI does not depend on proving that every number is false. It depends on something harder: knowing which numbers are true, what each one is authorized to prove, which ones are independent, which costs were left outside the frame, and whether the claimed outcome survives after reconciliation.

The strongest AI organizations will not be the ones with the highest activity count. They will be the ones that can connect activity to durable outcomes without losing the evidence in between.

Never let synthetic work become organizational memory without evidence. And never let AI activity become economic value without reconciliation.

The Honor System either gets evidence discipline, or it gets repriced.


Start The Honor System

Episode 6 closes the six-part investigation, but the argument is designed to be followed from the beginning. Start at nxtg.ai/insights, or follow the arc in order.

  1. Organizational Model Collapse. Unverified synthetic work enters organizational memory and returns later as trusted ground truth.
  2. Review Debt. Promotion outruns qualified verification, and the unreviewed work accumulates as debt.
  3. The Memory Trust Trap. Provisional information returns from memory with more authority than it had when it entered.
  4. The Trust-Laundering Machine. Institutions sign, file and publish the claim as accepted truth.
  5. Policy Is Not a Control. A rule becomes governance only when it can stop something, prove it did and undo the consequences.
  6. The Reckoning. Activity gets booked as value, and the finale asks what each number actually proves.

Sources and evidence

Economic outcomes

Benchmark integrity

Consumer metering

Capital, compute, and demand

Evidence note: This article's evidence was refreshed October 4, 2026, and its load-bearing numbers were spot-checked against the primary sources above on October 6, 2026. The NBER and BEA evidence supports real but uneven economic signals; it does not establish one universal causal AI productivity rate. The MIT Sloan numbers describe two organizations in one field study, not prevalence estimates. OpenAI and Anthropic benchmark findings are used as evidence of evaluation decay and self-correction, not vendor misconduct. Capital and compute relationships are treated as real contractual relationships; relationship topology changes what a signal independently proves, but it is not evidence that demand is fabricated. Internal NextGen AI examples are presented as first-person practice evidence, not as product certification, and the 97% figure is cited with its original caveat: the cost basis was partial and model cost had not been measured.

Frequently asked questions

What is Net Verified Value?
Net Verified Value is attributable business benefit minus total operating cost, minus review and rework, minus expected risk loss. It is a management discipline, not an accounting standard, designed to stop those four categories from disappearing from an AI value claim.
What is activity-value substitution?
Activity-value substitution occurs when an operational proxy such as prompts, seats, artifacts or hours saved is promoted into an economic outcome without proving the causal bridge, complete cost, durability and risk.
Why did OpenAI stop reporting SWE-bench Verified?
OpenAI audited 138 problems that its o3 model did not consistently solve and found that at least 59.4% had material issues in test design or problem description. It also found evidence that frontier models had seen at least some of the problems and solutions during training.
Does a large AI infrastructure backlog prove independent AI demand?
Not by itself. CoreWeave reported approximately $104 billion in revenue backlog as of June 30, 2026, but backlog is not revenue, revenue is not profit, and a supplier relationship is not independent validation. Overlapping relationships are not evidence of fraud either; they change what each signal independently proves.
What should a serious AI value claim carry?
The Value Evidence Ledger lists nine items: business outcome, baseline and counterfactual, attribution method, gross benefit, complete cost, durability window, a named accountable owner, scale, redesign and kill rules, and a reproducible evidence receipt.

Ready to build?

Ship AI you can trust

Forge gives you agents, governance, and verification — so your AI ships with confidence, not hope.

Newsletter

Enjoyed this article?

Get more insights like this delivered straight to your inbox.

Email subscription coming soon. Follow along on LinkedIn in the meantime.

Follow on LinkedIn