The Direct Answer: Calculate Risk Reduction, Not Tool Activity

The best way to calculate agentic pentesting ROI is to compare the expected reduction in cybersecurity loss with the total cost of the approach, while accounting for the time required to validate findings and remediate issues. Agentic systems can automate parts of reconnaissance, testing, evidence collection, and reporting, but their output should be treated as unverified security evidence until qualified practitioners confirm it. A defensible calculation therefore combines four inputs: annual loss exposure before testing, expected percentage risk reduction, total program cost, and confidence in the result. For example, if an organization estimates $1.2 million in annual expected cyber loss and a validated engagement reduces that exposure by 20%, the modeled annual benefit is $240,000. If annual tooling, labor, review, and remediation cost is $140,000, net benefit is $100,000 and first-year ROI is 71.4%. These figures are illustrative rather than universal benchmarks. The central issue is not how many tests an agent completes; it is whether verified findings cause measurable reductions in exploitable exposure, incident probability, or manual workload.

Also worth reading: How Should Enterprise Learning Teams Calculate and Improve ROI in 2026? · How do you accurately calculate and approach measuring enterprise AI training ROI today? · How Should Enterprises Design an Agentic Knowledge Architecture for Reliable AI Work?

The Core ROI Formula and Its Necessary Inputs

A practical formula is: net benefit = avoided expected loss plus verified productivity savings minus total program cost. ROI then equals net benefit divided by total program cost. Expected annual loss can be estimated as the probability of one or more relevant incidents during a year multiplied by the financial consequence of those incidents, adjusted for existing controls. A simpler input is annual revenue multiplied by a defensible loss rate, although that approach must distinguish total revenue from revenue exposed to a specific threat. A company with $100 million in annual revenue should not automatically assume that a pentest reduces cyber loss by a fixed percentage. Exposure varies by business model, data sensitivity, cloud footprint, regulatory obligations, and attacker incentives. The 20% reduction used in the opening example is therefore a scenario assumption that should be replaced by evidence from completed tests, incident history, control coverage, and expert review.

Total cost must include more than a platform subscription. It should include setup, data preparation, access provisioning, monitoring, human review, penetration testing labor, false-positive investigation, retesting, tool maintenance, and remediation. If a six-person security team spends 120 hours on an engagement and its fully loaded hourly cost is $125, labor alone is $15,000. When a vendor quotes $60,000, the organization should ask whether scope, specialist labor, liability terms, retesting, and support are included before concluding that the quote is inexpensive. Over a three-year evaluation period, recurring software cost may be discounted, but setup and remediation may not be. The date basis should also be stated clearly: calculations prepared on 29 September 2026 should use the latest available contract prices and organization-specific operating costs rather than generic “AI savings” claims.

Why Agentic Pentesting Can Create Value—and Where the Claims Break Down

Agentic pentesting can reduce repetitive work because software agents can navigate targets, collect observations, compare responses, and prepare drafts without waiting for every manual instruction. That efficiency matters when a security team must inspect many authenticated workflows, cloud configurations, or low-risk web surfaces. The value is greatest where the work is bounded, repeatable, and easy for a human to verify. It is weaker when authorization is ambiguous, the environment changes during the test, or the agent lacks reliable access to business context. A large number of generated findings does not prove that a system is secure, and fewer findings can actually indicate inadequate coverage if the agent silently fails to explore important paths.

Independent evidence should be handled carefully. Airlock Digital announced an Independent TEI Study Quantifying Measurable ROI and Security Impact, but an announcement about a study is not the same as access to its complete methodology. Before using its figures, an organization should request the underlying model, sample size, observation period, control group, definitions of cost and benefit, and treatment of implementation expenses. It should determine whether “security impact” was measured through avoided incidents, reduced downtime, faster remediation, or stakeholder estimates. A Technology, Economics, and Finance-style study can be useful, but independently commissioned does not mean independent of every commercial interest or applicable to every enterprise. Buyers should reproduce the arithmetic with their own exposure and cost data before approving purchase.

A Worked Example for an Enterprise Security Team

Consider a fictional software company preparing a business case in September 2026. It has $80 million in annual revenue, expects one material application compromise in a typical year, and estimates the financial consequence at $600,000, producing expected annual loss of $600,000. A scoped agent-assisted engagement is projected to discover and help remediate issues that lower that exposure by 15%, creating a modeled $90,000 annual avoided-loss benefit. Verified reporting and triage also save the team 160 hours per year at a $110 loaded hourly cost, adding $17,600 in productivity value. The platform costs $48,000 annually, while setup, access engineering, expert review, and first-cycle remediation cost $42,000 once. First-year total cost is $90,000, net benefit is $17,600, and first-year ROI is 19.6%.

A three-year case looks different if the benefit occurs annually while setup occurs only in year one. Assuming constant avoided loss and productivity savings, years two and three each have benefits of $107,600 and recurring costs of $48,000, producing $59,600 in net benefit per year. Total three-year benefits are $322,800, total costs are $186,000, and three-year ROI is 73.5%. This is the kind of sensitivity analysis a finance reviewer can understand. The company can then test a conservative 5% risk reduction and an aggressive 25% reduction instead of relying on the base case. If a $3,000 control improvement produces a large modeled benefit, verify whether it is genuinely incremental to existing controls. If an agent-generated alert takes 45 minutes to validate, a supposed ten-hour saving may disappear after 13 similar alerts.

FeatureManual PentestingAgentic Pentesting
Primary valueExpert judgment, creativity, and contextRepetitive execution, speed, and structured coverage
Common cost driverConsultant hours and travel or remote accessSubscription, integration, review, and validation labor
Best useComplex logic flaws, chained paths, and unusual business workflowsBroad asset inventory, bounded checks, and evidence collection
Main limitationExpensive per hour and capacity constrainedCan produce false positives, omissions, and unsafe actions
Measurement focusCoverage, validated findings, and remediationSame measures, adjusted for agent and human labor
## Practical Steps for Building a Credible Business Case

Start by defining a narrow decision and baseline. Decide whether the objective is to reduce manual workload, expand testing frequency, improve cloud coverage, shorten reporting time, or reduce a specific class of risk. Record the current annual spend, team hours, testing frequency, mean time to remediate, and number of validated findings. A useful baseline might be 12 manual web application tests per year, 320 hours per test, and 9 business days from validation to remediation. These numbers are not industry averages; they are examples of fields to populate. Measurement should distinguish gross findings from validated vulnerabilities because raw counts can reward noisy systems rather than effective testing.

Next, run a controlled pilot over 8 to 12 weeks with a bounded asset set. Capture subscription cost, setup time, agent execution time, human review time, false positives, true positives, severity distribution, retest rate, and incidents caused by testing. Compare results with a similar manual process where feasible. A useful acceptance threshold might require at least 30% less staff time for repetitive tasks, less than 10% false-positive rate after review, 100% human approval for potentially destructive actions, and no unexplained coverage gaps. Organizations should not adopt these thresholds as universal rules; they are starting points for negotiation. A 15% saving may still be worthwhile, but a 70% headline saving is not credible if it excludes analyst review or remediation work.

Comparison With Alternatives, Staffing Decisions, and Conventional Tools

Agentic pentesting should be compared with hiring additional testers, outsourcing the same scope, expanding a conventional scanning program, or accepting the current risk. A new tester may bring stronger judgment and creative problem-solving, but hiring can take 90 to 180 days in some labor markets and adds benefits, recruiting, management, and utilization risk. Outsourcing can provide specialist skills quickly, yet costs must be normalized for scope and retesting. Conventional scanners are efficient for known checks and broad pattern matching, but they generally require interpretation and do not autonomously reason through multi-step business workflows. Agentic systems may occupy the middle ground between scanners and human consultants, although product capabilities differ and should be verified through a trial.

Pricing is usually subscription-based, but the supplied research context does not provide verified vendor prices. Enterprise offers may be quoted per user, per environment, per asset, or through annual contracts, so an illustrative range of $30,000 to $150,000 per year should not be represented as market pricing. Total first-year cost can exceed the subscription by 50% to 200% when integrations, access engineering, review, and remediation are included, although the actual increase depends heavily on readiness. A lower quote may be preferable if it covers authenticated testing and expert validation, while a higher quote may still be economical if it replaces substantial manual effort. The correct comparison is cost per verified, remediated, and retested risk reduction—not price per AI task or number of agents.

Common Mistakes That Distort Agentic Pentesting ROI

The most common error is treating model-generated activity as completed work. Prompts, tool calls, and test steps are operational telemetry, not financial benefits. Another error is applying a platform vendor’s percentage risk reduction directly to another company’s annual revenue. The second assumes identical assets, controls, threat frequency, and loss severity, which is unlikely. Teams also tend to omit failed runs, access restrictions, duplicate findings, false positives, and the engineering time needed to create test accounts. Savings should be counted only when an employee or contractor can redirect verified time to other work; theoretical hours are not cash savings unless staffing, overtime, or external spending actually changes.

Additional mistakes include counting remediation cost as a benefit without describing which control was retired, and comparing an agent-assisted result with an unusually weak manual baseline. Avoid double-counting faster reporting and faster remediation when both are claimed from the same improvement. Use a documented counterfactual, retain raw timestamps, and require finance or security leadership to approve assumptions. Do not describe a forecast as an avoided incident: a cyber incident that did not occur cannot be observed directly. A defensible report calls it a modeled reduction in expected loss and discloses the assumptions. Given the supplied September 2026 context, an organization should also revisit the model after major acquisitions, cloud migrations, new product launches, or control changes because the original exposure may quickly become obsolete.

When to Act, Approve, Pilot, or Reject the Investment

Proceed when the problem is repetitive, the scope is authorized, baseline data exists, and a human reviewer can verify consequential results. A good early-use case is a company that already conducts regular penetration tests but spends excessive analyst time collecting evidence across many authenticated endpoints. Another suitable case is a regulated organization that needs repeatable control testing and audit-ready documentation, provided the agent does not become an unsupported claim of compliance. Avoid immediate purchase when leadership wants “autonomous pentesting” without defined rules of engagement, privileged-access controls, logging, or incident contacts. The test plan should state permitted IP ranges, domains, accounts, rate limits, prohibited techniques, data-handling rules, and stop conditions.

Approval should depend on a payback threshold set before the pilot. A common finance rule is to accept investments that repay the initial outlay within 12 to 24 months, while security leaders may require a stronger risk case for low-probability but severe events. In the earlier example, recurring annual benefit of $107,600 against a first-year cost of $90,000 gives an immediate modeled payback, but that conclusion remains sensitive to the assumed risk reduction. If the company can show only $35,000 in annual benefit, the same purchase has a negative first-year cash case. Pilot when technical feasibility is promising but operational impact is uncertain. Reject or redesign when the vendor cannot provide audit logs, access controls, human escalation, data-use terms, or evidence that total review time was included. The strongest decision is therefore conditional: approve a measured pilot first, then scale only after independently reproducing the results.

How to Report the Final ROI to Executives

Present the result as a range with explicit assumptions rather than one exact number. For the fictional enterprise, the base case yielded 73.5% ROI over three years; a conservative case might use only 5% risk reduction and half the claimed productivity saving, while an optimistic case might use 25% reduction and full verified labor savings. Show subscription, setup, labor, review, remediation, and retesting separately. Label every value as historical, contracted, estimated, or modeled. If evidence comes from the announced Airlock Digital independent TEI study, cite the complete report when available and explain which numbers came from that study and which were replaced with internal data. Do not construct a URL or citation from the supplied title alone.

The final recommendation should answer four questions in plain language: what risk changed, how the change was verified, what the organization paid, and when the investment pays back. A useful dashboard might track validated findings per 100 tested workflows, mean review time, false-positive rate, retest pass rate, median remediation time, and annualized cost per validated finding. Agentic pentesting earns a credible ROI case when it improves those measures without reducing safety or expert oversight. It does not need to claim that AI removes pentesters or guarantees breach prevention. As of 29 September 2026, the defensible position is that agentic testing is a potential efficiency and risk-control layer whose ROI must be proven by scoped pilots, transparent assumptions, and verified outcomes.