AI cannot be priced until correctness is proved, and correctness is not being measured, so the value is being taken on the cost line instead.
01
02
Every vendor in this market ships AI. Almost none of them charges for it. The usual explanation is that buyers are not ready or that metering is hard. The data says something less flattering and more structural.
Research across 3,200 companies found that only 14% consistently net out positive from AI, and roughly 40% of the time savings are lost to rework, because 77% of people review AI output as carefully as work produced by a colleague. That review gate is the thing vendors sell as their safety story. It is also the thing destroying the buyer's net benefit. Set beside a separate finding that 38% of organisations run agentic workflows with no such gate at all, it is a more complete explanation for the absence of priced AI than vendor reticence.
The gate cannot come off without proof the agent is right. That proof does not exist. Not one vendor publishes a correctness rate for a payroll agent, and the reason is that the failure mode is silent: an agent that misses a wrong tax jurisdiction produces no error, no alert and no audit-log anomaly. It produces a clean run and a false assurance. An approval gate cannot catch that, because nothing is presented for the human to approve.
Both the regulation and the certification reward the gate's existence rather than its efficacy. The EU carve-out for high-risk employment systems attaches to having human oversight, whether the reviewers catch 89% of errors or 13.6%, and nobody is required to know which. ISO 42001 certifies that an AI management system exists; the holder with payroll named in scope publishes no error rate for any feature. Two independent institutions rewarding a control nobody measures.
So the value gets taken on the cost line. Paycom's guidance raise decomposed to roughly 83% cost rather than sales, margin moving 44% to 46%, and the market paid a 12 to 14% re-rate for it. That is close to irreversible. Once AI ships as included, re-pricing it later means telling customers that something they already have now costs money.
Meanwhile the equity market has started discriminating on exactly this axis. SAP showed AI attached to 90% of its fifty largest deals and took a 9% pop. Workday, whose entire 2026 story is agentic, took the Underweight. The distinction is not who has agents. It is whether the AI has reached the revenue line on the timetable the multiple assumed.
What would change our mind. Customer service completed this transition and solved the definitional problem, with Intercom at $0.99 a resolution reaching eight-figure ARR. Payroll outcomes are more measurable and more guaranteeable than a resolved support ticket. If two mid-market vendors publish per-outcome pricing with disclosed attach while nobody has published a correctness rate, the gate was never the constraint.
03
The buyer's own human review gate absorbs the benefit: only 14% of companies consistently net out positive from AI, ~40% of time savings are lost to rework, and 77% review AI output as carefully as human-created work. The gate vendors sell as their safety story is the thing destroying the buyer's net benefit. It cannot come off without proof the agent is right - and nobody publishes a correctness rate, because the failure mode is silent: a missed jurisdiction produces a clean run and false assurance, which an approval gate structurally cannot catch since nothing is presented to approve. Both the regulation and the certification reward the gate's EXISTENCE rather than its efficacy - the EU carve-out attaches to having human oversight whether reviewers catch 89% or 13.6%, and ISO 42001 certifies a management system exists while the holder publishes no error rate for any feature.
04
The category is stuck in a loop it built itself. Vendors cannot charge for AI, so they take it on the cost line - Paycom's raise decomposed to roughly 83% cost rather than sales, margin 44%->46%, and drew a 12-14% re-rate. That is close to irreversible: re-pricing later means telling customers something they already have now costs money. Meanwhile the market pays only where AI is already in the revenue - SAP with AI in 90% of its 50 largest deals took a ~9% pop; Workday, whose entire story is agentic, took the Underweight. The exit is a published, audited correctness rate that lets a buyer thin the gate.
05
Given equal weight to the claim. Hiding the counter is how a prediction becomes an article of faith.
06
07
Specific public numbers we check on a schedule, so this claim can be tested without taking our word for it. Not checked means nobody looked. That counts as nothing, never as agreement.
| What we check | Status | Latest reading | Last checked |
|---|---|---|---|
| Vendors publishing an audited payroll-agent detection or accuracy rate vendor releases, product docs · scan | not moving | still zero for any agent. Personio (19% of payslips wrong) and SAP (38% of employees experienced an error) published error rates for the HUMAN process: the instrument's shadow, not the instrument | 2026-09-09 |
| HR&P vendors publishing any explicit AI price vendor pricing pages, earnings · scan | not moving | still zero; ADP (AWS expansion) and SAP (Autonomous Payroll) added to the count of AI announcements with no price | 2026-09-09 |
| Margin guide movement against revenue guide movement mid-market earnings · quarterly | changing | Paycom ~83% of raise from cost, not sales | 2026-08-05 |
| Exec structures fusing AI with the money rail or product vendor exec announcements · event | changing | Xero CTO not replaced; SVP Engineering, Payments & AI Transformation | 2026-08-05 |
08
Cumulative linked signals, by direction. A line that only climbs in green is being read generously.
6 counted: 1 diverges · 5 supports · 1 predate the claim and are not counted
| Date | Signal | Bearing | On which part of the mechanism |
|---|---|---|---|
| 2026-08-27 | signal-workday-ai-arr-adoption-gap | diverges | Vendors cannot charge for AI, so they take it on the cost line. |
| 2026-08-06 | signal-rippling-ai-spend-console | Before the claim (not counted) | Vendors cannot charge for AI, so they take it on the cost line. |
| 2026-09-01 | signal-sdworx-h1-ai-unpriced | supports | Nobody publishes a correctness rate, because the failure mode is silent. |
| 2026-09-02 | signal-leo-hr-judgement-agent | supports | Nobody publishes a correctness rate, because the failure mode is silent. |
| 2026-09-09 | signal-personio-payroll-error-rate | supports | Nobody publishes a correctness rate, because the failure mode is silent. |
| 2026-09-09 | signal-adp-aws-expansion | supports | Nobody publishes a correctness rate, because the failure mode is silent. |
| 2026-09-09 | signal-sap-autonomous-payroll-undated | supports | Nobody publishes a correctness rate, because the failure mode is silent. |
09
The market moved along a different path than the one drawn. The prediction keeps its identity and its history. A narrative that changed is itself evidence.
Was: The first HR&P vendor to publish an audited correctness rate for a payroll agent prices against it within two quarters of publishing, and links the two in its own disclosure. Absent any published rate, no mid-market HR&P vendor establishes per-outcome AI pricing with disclosed attach.
Became: The first HR&P vendor to publish an audited correctness rate for a payroll agent prices against that rate within two quarters of publishing it, and links the two in its own disclosure.
Why: Two claims were bundled in one sentence and the bundling hid which one the evidence was testing. Claim A, what the first publisher of a correctness rate does, is specific, surprising and falsifiable, and had ZERO evidence because its antecedent has never fired: the correctness-rate-count instrument has read 0 since the day it was created. Claim B, 'absent any published rate, no mid-market vendor establishes per-outcome AI pricing', is where all six links attached, and against a market where nobody does either thing it is close to a tautology: every observation of a vendor shipping AI without a price confirms it, and only the kill could contradict it, which required two vendors to do a thing zero vendors have done. The store's own warning was right and the diagnosis was not generosity, it was structure. Claim B is demoted to an instrument reading, which is what it always was. The mechanism is retained unchanged and is the strongest in the set. The new falsifier is TIGHTER: it fires on a single vendor rather than requiring two, so the claim is now easier to prove wrong, which is the direction a revision is allowed to move.
0 signals carried forward; 6 dropped as no longer bearing on the claim: Supports the retained mechanism, that value cannot be priced so it is taken on the cost line. It cannot test the claim, which is about what a vendor does after publishing a correctness rate. Paycom has published no rate.; Establishes that correctness-rate-count reads zero, which is an INSTRUMENT READING and not evidence for the claim. It is the precondition, not a test of the consequent.; Bears on the org-chart instrument, not on the claim. An engineering reorganisation is not a pricing decision following a published rate.; Supports the retained mechanism, that the review gate absorbs the benefit. No correctness rate was published, so the claim's antecedent never fired.; Announced capability outrunning shipped capability is a real pattern and belongs to a different claim. A missed GA date says nothing about pricing against a published accuracy figure.; Bears on the margin-vs-revenue instrument. Margin moving while revenue does not is a reading, not a test of what the first rate-publisher does.
10
Inverts the retired trial `the-agent-becomes-the-unit-price-arrives-last`, which held that the price arrives last. It is not arriving on the revenue line at all. Overlaps retired trial `correctness-undisclosed` on the missing-rate observation; the review-gate mechanism is new. Dated 2026-08-12.
This prediction carries no resolution date. We cannot predict when evidence will arrive, so the review cadence attaches to the instruments above rather than to the claim. It runs until the market proves it, moves it, or twelve months pass with no material signal against it.
How a signal becomes a prediction →