Predictionsformed 2026-08-12

The gate is the trap

AI cannot be priced until correctness is proved, and correctness is not being measured, so the value is being taken on the cost line instead.

LiveRating A, multiple independent sources, recurring instrument, sustainedstructural

The prediction

The first HR&P vendor to publish an audited correctness rate for a payroll agent prices against that rate within two quarters of publishing it, and links the two in its own disclosure.

The read

Every vendor in this market ships AI. Almost none of them charges for it. The usual explanation is that buyers are not ready or that metering is hard. The data says something less flattering and more structural.

Research across 3,200 companies found that only 14% consistently net out positive from AI, and roughly 40% of the time savings are lost to rework, because 77% of people review AI output as carefully as work produced by a colleague. That review gate is the thing vendors sell as their safety story. It is also the thing destroying the buyer's net benefit. Set beside a separate finding that 38% of organisations run agentic workflows with no such gate at all, it is a more complete explanation for the absence of priced AI than vendor reticence.

The gate cannot come off without proof the agent is right. That proof does not exist. Not one vendor publishes a correctness rate for a payroll agent, and the reason is that the failure mode is silent: an agent that misses a wrong tax jurisdiction produces no error, no alert and no audit-log anomaly. It produces a clean run and a false assurance. An approval gate cannot catch that, because nothing is presented for the human to approve.

Both the regulation and the certification reward the gate's existence rather than its efficacy. The EU carve-out for high-risk employment systems attaches to having human oversight, whether the reviewers catch 89% of errors or 13.6%, and nobody is required to know which. ISO 42001 certifies that an AI management system exists; the holder with payroll named in scope publishes no error rate for any feature. Two independent institutions rewarding a control nobody measures.

So the value gets taken on the cost line. Paycom's guidance raise decomposed to roughly 83% cost rather than sales, margin moving 44% to 46%, and the market paid a 12 to 14% re-rate for it. That is close to irreversible. Once AI ships as included, re-pricing it later means telling customers that something they already have now costs money.

Meanwhile the equity market has started discriminating on exactly this axis. SAP showed AI attached to 90% of its fifty largest deals and took a 9% pop. Workday, whose entire 2026 story is agentic, took the Underweight. The distinction is not who has agents. It is whether the AI has reached the revenue line on the timetable the multiple assumed.

What would change our mind. Customer service completed this transition and solved the definitional problem, with Intercom at $0.99 a resolution reaching eight-figure ARR. Payroll outcomes are more measurable and more guaranteeable than a resolved support ticket. If two mid-market vendors publish per-outcome pricing with disclosed attach while nobody has published a correctness rate, the gate was never the constraint.

The mechanism

The buyer's own human review gate absorbs the benefit: only 14% of companies consistently net out positive from AI, ~40% of time savings are lost to rework, and 77% review AI output as carefully as human-created work. The gate vendors sell as their safety story is the thing destroying the buyer's net benefit. It cannot come off without proof the agent is right - and nobody publishes a correctness rate, because the failure mode is silent: a missed jurisdiction produces a clean run and false assurance, which an approval gate structurally cannot catch since nothing is presented to approve. Both the regulation and the certification reward the gate's EXISTENCE rather than its efficacy - the EU carve-out attaches to having human oversight whether reviewers catch 89% or 13.6%, and ISO 42001 certifies a management system exists while the holder publishes no error rate for any feature.

What follows if it holds

The category is stuck in a loop it built itself. Vendors cannot charge for AI, so they take it on the cost line - Paycom's raise decomposed to roughly 83% cost rather than sales, margin 44%->46%, and drew a 12-14% re-rate. That is close to irreversible: re-pricing later means telling customers something they already have now costs money. Meanwhile the market pays only where AI is already in the revenue - SAP with AI in 90% of its 50 largest deals took a ~9% pop; Workday, whose entire story is agentic, took the Underweight. The exit is a published, audited correctness rate that lets a buyer thin the gate.

The strongest case against

Given equal weight to the claim. Hiding the counter is how a prediction becomes an article of faith.

Customer service completed the outcome-pricing transition and solved the definitional problem - Intercom at $0.99 a resolution to eight-figure ARR - and payroll outcomes are more guaranteeable than a resolved ticket. If the definitional problem was never the blocker there, the gate argument must explain why HR&P differs. Note also HubSpot halved its unit price within months: even where outcome pricing works, unit price falls.

What would kill it

A vendor publishes an audited correctness rate for a payroll agent, two consecutive quarters pass, and no pricing change in that vendor's own disclosure references that rate. No second vendor is required, and no expiry applies beyond the two consecutive quarters.

How we will know

Specific public numbers we check on a schedule, so this claim can be tested without taking our word for it. Not checked means nobody looked. That counts as nothing, never as agreement.

What we checkStatusLatest readingLast checked
Vendors publishing an audited payroll-agent detection or accuracy rate
vendor releases, product docs · scan
not movingstill zero for any agent. Personio (19% of payslips wrong) and SAP (38% of employees experienced an error) published error rates for the HUMAN process: the instrument's shadow, not the instrument2026-09-09
HR&P vendors publishing any explicit AI price
vendor pricing pages, earnings · scan
not movingstill zero; ADP (AWS expansion) and SAP (Autonomous Payroll) added to the count of AI announcements with no price2026-09-09
Margin guide movement against revenue guide movement
mid-market earnings · quarterly
changingPaycom ~83% of raise from cost, not sales2026-08-05
Exec structures fusing AI with the money rail or product
vendor exec announcements · event
changingXero CTO not replaced; SVP Engineering, Payments & AI Transformation2026-08-05

Evidence

Cumulative linked signals, by direction. A line that only climbs in green is being read generously.

012345supportsdiverges08-2709-0109-0209-09

6 counted: 1 diverges · 5 supports · 1 predate the claim and are not counted

DateSignalBearingOn which part of the mechanism
2026-08-27signal-workday-ai-arr-adoption-gapdivergesVendors cannot charge for AI, so they take it on the cost line.
2026-08-06signal-rippling-ai-spend-consoleBefore the claim (not counted)Vendors cannot charge for AI, so they take it on the cost line.
2026-09-01signal-sdworx-h1-ai-unpricedsupportsNobody publishes a correctness rate, because the failure mode is silent.
2026-09-02signal-leo-hr-judgement-agentsupportsNobody publishes a correctness rate, because the failure mode is silent.
2026-09-09signal-personio-payroll-error-ratesupportsNobody publishes a correctness rate, because the failure mode is silent.
2026-09-09signal-adp-aws-expansionsupportsNobody publishes a correctness rate, because the failure mode is silent.
2026-09-09signal-sap-autonomous-payroll-undatedsupportsNobody publishes a correctness rate, because the failure mode is silent.

Transformation history

The market moved along a different path than the one drawn. The prediction keeps its identity and its history. A narrative that changed is itself evidence.

2026-08-21

Was: The first HR&P vendor to publish an audited correctness rate for a payroll agent prices against it within two quarters of publishing, and links the two in its own disclosure. Absent any published rate, no mid-market HR&P vendor establishes per-outcome AI pricing with disclosed attach.

Became: The first HR&P vendor to publish an audited correctness rate for a payroll agent prices against that rate within two quarters of publishing it, and links the two in its own disclosure.

Why: Two claims were bundled in one sentence and the bundling hid which one the evidence was testing. Claim A, what the first publisher of a correctness rate does, is specific, surprising and falsifiable, and had ZERO evidence because its antecedent has never fired: the correctness-rate-count instrument has read 0 since the day it was created. Claim B, 'absent any published rate, no mid-market vendor establishes per-outcome AI pricing', is where all six links attached, and against a market where nobody does either thing it is close to a tautology: every observation of a vendor shipping AI without a price confirms it, and only the kill could contradict it, which required two vendors to do a thing zero vendors have done. The store's own warning was right and the diagnosis was not generosity, it was structure. Claim B is demoted to an instrument reading, which is what it always was. The mechanism is retained unchanged and is the strongest in the set. The new falsifier is TIGHTER: it fires on a single vendor rather than requiring two, so the claim is now easier to prove wrong, which is the direction a revision is allowed to move.

0 signals carried forward; 6 dropped as no longer bearing on the claim: Supports the retained mechanism, that value cannot be priced so it is taken on the cost line. It cannot test the claim, which is about what a vendor does after publishing a correctness rate. Paycom has published no rate.; Establishes that correctness-rate-count reads zero, which is an INSTRUMENT READING and not evidence for the claim. It is the precondition, not a test of the consequent.; Bears on the org-chart instrument, not on the claim. An engineering reorganisation is not a pricing decision following a published rate.; Supports the retained mechanism, that the review gate absorbs the benefit. No correctness rate was published, so the claim's antecedent never fired.; Announced capability outrunning shipped capability is a real pattern and belongs to a different claim. A missed GA date says nothing about pricing against a published accuracy figure.; Bears on the margin-vs-revenue instrument. Margin moving while revenue does not is a reading, not a test of what the first rate-publisher does.

Provenance

Inverts the retired trial `the-agent-becomes-the-unit-price-arrives-last`, which held that the price arrives last. It is not arriving on the revenue line at all. Overlaps retired trial `correctness-undisclosed` on the missing-rate observation; the review-gate mechanism is new. Dated 2026-08-12.

This prediction carries no resolution date. We cannot predict when evidence will arrive, so the review cadence attaches to the instruments above rather than to the claim. It runs until the market proves it, moves it, or twelve months pass with no material signal against it.

How a signal becomes a prediction →