Skip to content
Reality Graph

Economics

What Verification Debt Costs

Last updated: 2026-08-153 min read

Verification debt can appear as rework and review time, but its cost is a local estimate, not a universal price. Use a defined period and replace every example input with observed team data. External studies can identify risks worth measuring; they cannot establish your defect rate, incident exposure or ROI.

Contents

Start with the evidence class, not the euro sign

Four classes must stay separate. A source reports an observation under its own method. Your team records a local observation under a written definition. An estimate combines selected inputs. A calculated result is only the arithmetic consequence of those inputs. None of those classes becomes a measured saving until a comparable before-and-after observation supports it.

GitClear’s repository study, Faros’ customer telemetry and Veracode’s generated-code security test each describe different samples and outcomes. They are useful warning signals. They do not justify pricing a churn delta as a defect rate, converting review telemetry into a universal per-PR tax, or deriving an incident allowance from a security-test failure rate.

A worked example with replaceable assumptions

The example uses 120 AI-assisted changes per month, a 2% locally reason-coded rework rate, six hours per rework, half an hour of additional review reconstruction per change and a €75 internal hourly cost. These are illustrative inputs, not sourced benchmarks. They produce 14.4 rework hours and 60 review hours, or 74.4 hours and €5,580 for that month. Change any input and the result changes.

Replaceable inputs, not a benchmark

These example inputs produce 74.4 hours or €5,580 per month; that is arithmetic from assumptions, not a measured customer outcome.
ExampleReplace withMeaning
AI-assisted changes/month120Replaceable example assumption · 2026-07-16Only demonstrates the formula; not imported from a study.Git/PR dataEditorial evidence boundary · 2026-07-16Requires a locally defined measurement rule and period.Volume, not a quality judgementEditorial evidence boundary · 2026-07-16Interpretation applies only to the selected inputs.
Reworked within 14 days for a defect2%Replaceable example assumption · 2026-07-16Only demonstrates the formula; not imported from a study.Your reason-coded rework rateEditorial evidence boundary · 2026-07-16Requires a locally defined measurement rule and period.2.4 changes/month in the exampleEditorial evidence boundary · 2026-07-16Interpretation applies only to the selected inputs.
Hours per rework6 hReplaceable example assumption · 2026-07-16Only demonstrates the formula; not imported from a study.Ticket/time records or a sampleEditorial evidence boundary · 2026-07-16Requires a locally defined measurement rule and period.14.4 h/month in the exampleEditorial evidence boundary · 2026-07-16Interpretation applies only to the selected inputs.
Extra reconstruction per change0.5 hReplaceable example assumption · 2026-07-16Only demonstrates the formula; not imported from a study.Review sample with/without task and test evidenceEditorial evidence boundary · 2026-07-16Requires a locally defined measurement rule and period.60 h/month in the exampleEditorial evidence boundary · 2026-07-16Interpretation applies only to the selected inputs.
Internal loaded hourly cost€75Replaceable example assumption · 2026-07-16Only demonstrates the formula; not imported from a study.Finance value for the actual teamEditorial evidence boundary · 2026-07-16Requires a locally defined measurement rule and period.Monetised only if this value fitsEditorial evidence boundary · 2026-07-16Interpretation applies only to the selected inputs.
These example inputs produce 74.4 hours or €5,580 per month; that is arithmetic from assumptions, not a measured customer outcome.

Run the example with your own numbers

The same model, as an instrument

The figures above are one team's example. Move the four inputs to yours and the arithmetic recomputes in place, with the same rework rate, the same reconstruction time and the same refusal to add an incident allowance nobody has defined. The same calculator sits on its own page, alongside what a Team version would change.

10
20
Share of merges that are AI-assisted
€75

Estimated cost of verification debt

Example – illustrative arithmetic, not a benchmark

€67,000

per year · €5,580 per month

Modelled on about 120 AI-assisted merges a month.

At these inputs the model puts verification debt at €67,000 a year, or €5,580 a month: 74.4 hours of engineering time, 60 of them spent working out what a change was meant to do before it can be judged.

Past break-even.The amber band marks 30 to 40 AI-assisted changes a month, on a logarithmic scale so the whole range fits. Below it the example crosses into marginal territory: little debt left to remove, and the practice roughly pays for itself rather than returning more.
Cost linePer monthHours a month
Review reconstruction€4,50060 h
Rework on churned code€1,08014.4 h
Incident allowancenone assumed0 h
Verification debt priced from the worked example on this page, with the inputs made adjustable. Illustrative arithmetic, not a benchmark.

What the model assumed for you· Assumption set as of 2026-08-15

0.5 hours of review reconstruction per AI-assisted change · 2 % of AI-assisted changes reworked for a defect within 14 days, an illustrative rate to replace with your own · 6 hours to rework one churned change

Make the estimate decision-useful

  • Define rework before counting it. Separate defect fixes from planned iteration, formatting and unrelated follow-up work.
  • Use a bounded review sample. Compare changes with and without a written task and attached receipts; do not infer reconstruction time from an external aggregate.
  • Keep incidents separate. Add expected loss only when your own taxonomy, exposure and history support it; otherwise report it as unquantified risk.
  • Publish uncertainty. Show a range or rerun the formula with alternative assumptions instead of presenting one point value as fact.

The companion guide explains how to collect the four local verification-debt measures. If you then compare an intervention, the ROI method adds intervention cost and a comparable baseline without turning assumptions into promises.

The calculation provides

  • An auditable formula
  • Replaceable assumptions
  • Separation of observation, estimate and result
  • A starting point for a local baseline

It does not provide

  • An industry benchmark
  • Causal proof about AI
  • Guaranteed savings
  • An incident or ROI forecast
Measure locally first, monetise second, then make a decision.

Where Reality Graph fits

Reality Graph can make tasks, validation receipts and run boundaries easier to retain. Whether that changes rework or review effort is an empirical question for your workflow. Measure the same definitions before and after a bounded pilot, retain failed as well as passing outcomes, and let the local comparison support the decision.

FAQ

What does unverified AI code cost a team per year?
There is no defensible universal figure. Estimate your own cost from a defined period: change volume, a reason-coded rework rate, hours per rework, extra review-reconstruction time and your internal hourly cost. The worked example demonstrates that arithmetic but is not a benchmark or customer result.
Which inputs are measured and which are assumptions?
The example values are visibly marked as replaceable assumptions. A real estimate should replace them with local Git or PR volume, reason-coded rework, a review sample and a finance-approved cost rate. External studies provide context about possible code-change, review and security risks; they do not provide your team's cost rates.
Can an external churn study establish our rework cost?
No. Repository studies are observational and use their own definitions and samples. They can motivate a local measurement, but importing a reported percentage as your defect rate would create false precision and would not establish causation.
Should the model include incidents?
Only when you have a locally defined incident class, frequency and expected-loss method. A security benchmark that finds weaknesses in generated samples is not an incident frequency and cannot justify a monetary allowance by itself.
Does the calculation prove verification ROI?
No. It estimates selected debt-related effort under stated assumptions. ROI additionally needs the measured cost of the intervention, a comparable baseline, a time boundary and uncertainty. One worked example cannot guarantee savings or positive ROI.
What should a team measure first?
Start with one workflow and one fixed period. Record AI-assisted change volume, reason-coded short-horizon rework and a small review-time sample. Keep observed values separate from estimates and change one definition only between comparable periods.

Keep reading

Sources

Want to see what your last agent run would have looked like?

Request access