How to Measure Claude Usage Impact on Engineering Delivery Outcomes
TL;DR: Proving Claude return on investment (ROI) means answering the CFO's question: is higher token spend buying more delivery, or just a bigger bill? Counts of pull requests merged and AI-vs-non-AI cohort comparisons can't answer it. A linear regression can: fit each business outcome against token spend across teams and time periods, then read the slope as the ROI figure. minware's Lean AI Framework regresses three outcomes this way: roadmap value delivery, on-time delivery rate, and overhead cost. Each team's trend over time is the supporting check, while bug rate and PR lead time act as guardrails that explain why a trend moves. The method needs no clean pre-AI baseline.
Organizations are seeing AI token spend grow, and the board wants to know what it bought. The measurement problem is that Claude usage data sits in one system, delivery data sits in others, and no normalized view connects them.
This guide lays out a methodology for connecting Claude token spend to delivery outcomes like roadmap value delivery and on-time delivery rate, using a regression approach that survives CFO scrutiny. It covers why the two most common methods fail, how to construct the regression, which confounders to name, and how to build a board-ready report.
Moving past merged pull request counts and adoption dashboards
Two metrics appear on most engineering dashboards when AI spend comes up: pull requests merged and adoption counts. The problem with both is structural, not incidental.
The throughput trap
Throughput metrics cover pull requests merged, commits, and lines of code. These measure code output volume, and they carry the same structural problem. A team generating more pull requests has demonstrated that Claude produces more code, but it has not demonstrated that the code delivered value.
Answering the CFO's question requires two data streams connected in the same model. The first is token spend from Claude Code's Analytics API or OpenTelemetry. The second is delivery outcomes from your version control system (such as GitHub, GitLab, Bitbucket, or Azure DevOps) and project management system (such as Jira, Linear, Azure Boards, or GitHub Issues).
The gap in usage counts
Anthropic's Claude Code Analytics API reports usage at per-user, per-model, per-day granularity, and OpenTelemetry integration provides per-prompt, per-session detail. Both tell you how much Claude is being used and what it costs. Neither tells you what that usage produced in terms of delivered work.
The distinction matters because every extra line of code still has to pass through review, testing, and deployment before it counts as delivered work. Reporting adoption metrics to the board skips the step where you check whether those stages kept up.
Ruling out AI-vs-non-AI cohort comparisons
The intuitive approach to measuring Claude's impact is comparing AI-assisted work against non-AI work. This method is losing meaning as adoption spreads, and the comparisons that remain carry selection bias that undermines the conclusion.
The limits of AI-vs-non-AI cohort comparison
As AI adoption approaches saturation (90% of respondents in the 2025 DORA report), the non-AI comparison group shrinks toward zero. The teams and individuals who have not adopted AI are increasingly unrepresentative: they may work in legacy repositories with poor test coverage, maintain compliance-sensitive codebases, or simply have different risk tolerances. Comparing their output against AI-adopting teams measures how those populations differ.
Data scientist William Gieng makes the same point about opt-in AI features: nobody randomized the AI assistant, so the comparison between adopters and non-adopters "is not an effect. It is a description of who opts in." By the time an account shows up as an adopter, he writes, the flag is close to a proxy for organizational readiness.
Selection bias in the remaining non-AI groups
The selection problem operates at multiple levels. Early adopters are often already high performers, and they may be more motivated to experiment or more comfortable with new tooling. Any observed difference between the groups may reflect characteristics of the person rather than the effect of the tool.
Repositories and teams adopting AI tools may differ systematically from non-adopters in ways that independently affect delivery outcomes or quality. Even METR's controlled research reports selection effects in its own study, noting it is "likely missing data on the most active adopters of AI, who may be of most interest."
Token spend as a continuous variable
The regression approach sidesteps the cohort problem entirely. Instead of splitting work into AI and non-AI groups, you treat token spend as a continuous variable, then regress delivery outcomes against it. Non-AI work has zero spend, so the regression already includes it as the zero-spend endpoint without requiring a separate binary split.
This means the methodology can work across teams with varying levels of Claude usage, because the analysis runs on the level of spend, which varies even when every team uses Claude. The question shifts from "did AI users do better" to "does more spend predict more delivery," which is the question the CFO is actually asking.
Regressing delivery outcomes against token spend
The core methodology is a linear regression of each delivery outcome against token spend level, grouped by team and time period, with each data point normalized to per-person-day. The regression slope is the ROI readout: it tells you the change in the outcome metric per unit change in spend.
Setting up the regression
The regression approach fits token spend per person-day in dollars as the input variable, drawn from Claude Code's usage data, against delivery outcome variables that come from your delivery data. On flat-rate or discounted plans, reported spend may differ from the amount actually invoiced. Measuring Claude's impact on multiple outcomes means running separate regressions for each outcome metric.
The regression requires teams and time periods that differ in spend level to produce meaningful variation. Linear regression produces a slope that represents the change in the outcome per unit change in spend. The R-squared value tells you how much of the outcome's variation the spend actually explains.
minware's pre-built AI impact reports run this regression automatically, reporting the slope with its R-squared value. Each data point is normalized to per-person-day by dividing token spend and the outcome metric by active contributor work days.
Choosing which outcomes to regress
minware's Lean AI Framework, a three-pillar methodology for measuring AI impact, regresses three business outcomes separately against token spend: roadmap value delivery, on-time delivery rate, and overhead cost. When no explicit value field is configured, roadmap value delivery defaults to story points completed on tickets that are not bugs and have a parent epic, making it the most direct measure of whether Claude is helping the team ship planned work.
On-time delivery rate divides each roadmap item's estimated duration by its actual duration, capped at 100% per item, giving boards a continuous measure of whether commitments land on schedule. Overhead cost captures personnel cost and token spend on work outside roadmap value delivery. By default, that means work not traceable to any ticket, bug tickets, or tickets with no parent epic. Together, the three show whether AI spend is buying more planned work, delivered closer to schedule, at a lower overhead.
The regression slope for roadmap value delivery gives you value delivered per dollar of token spend, in story points under the default configuration. That marginal figure is the ROI answer. A simple ratio of total outcomes divided by total spend would attribute all delivery to AI spend, including work that would have been completed without the tool.
Naming the confounders
Every regression conclusion should name its confounders. AI ROI analysis commonly faces three: task-level selection (engineers reach for Claude on work where it is likely to help, skipping it elsewhere), person-level selection (early adopters are often already the strongest performers), and covariates such as codebase age or code quality.
Naming these strengthens the analysis by showing the board you understand its limitations. The same-team trend over time, covered in the next section, controls for differences between teams. Task-level and person-level selection remain, so present them as named limits of the analysis.
Controlling for team differences with same-team trends
Comparing Team A's delivery metrics against Team B's introduces every difference between those teams as a potential confounder: codebase age, team composition, technical debt, domain complexity. Tracking the same team's delivery outcomes over time as their Claude usage matures holds those team-level factors constant.
If a team's roadmap value delivery trend rises as their token spend per person-day increases, and that pattern holds across multiple teams independently, the case that spend is driving the change gets stronger. Each team serves as its own baseline, which is why this methodology does not require a clean pre-AI baseline snapshot.
The trend view plots a delivery outcome on the Y-axis against time on the X-axis, filtered to a single team. As Claude adoption matures within that team, you should see the outcome trend in the direction the regression slope predicts. If the regression says more spend predicts more delivery but the same-team trend is flat, the regression may be capturing between-team differences rather than a real effect.
Using quality and efficiency metrics as guardrails and diagnostics
Quality and efficiency metrics serve two specific jobs in an AI ROI analysis. As guardrails, they confirm the spend is not creating bottlenecks. As diagnostics, they explain why delivery outcomes are or are not moving.
Bug rate as the quality guardrail
Bug rate divides bugs created by pull requests merged, capturing quality problems that change failure rate misses. As token spend rises, bug rate should hold steady or fall. If it climbs alongside spend, the additional code volume may be introducing defects that absorb the delivery gains downstream.
The quality reports track bugs created by priority and team, plus bug lead time by priority, giving you the drill-down to investigate when the guardrail trips.
"I use minware for our SDLC metrics and appreciate its ability to provide quality metrics throughout our SDLC. It gives us clear visibility into code quality, defect rates, and overall development health." - George V. on G2
PR lead time as the efficiency guardrail
PR lead time runs from a pull request's first commit to deployment, and it should hold steady or fall as AI spend rises. If Claude is generating more code but PR lead time is climbing, something downstream is absorbing the extra volume. PR review time, the wait for a first human review, shows whether that constraint is the review stage.
A July 2026 longitudinal study of an enterprise AI mandate found per-reviewer load roughly doubled as per-capita throughput rose, so the extra volume lands on reviewers first. The efficiency reports break PR lead time down by workflow stage, from first commit to merge, helping you identify where the constraint sits.
Diagnosing flat delivery outcomes despite rising spend
When token spend climbs but delivery outcomes have not moved, the regression will show a flat or near-zero slope. The diagnostic question is where the value is leaking.
Weak branch-to-ticket linking can cause roadmap value delivery and story point velocity to undercount AI-assisted work, because those metrics depend on linking code activity to tickets. The Best Practices report surfaces exactly where the process is breaking down, giving you itemized findings to work through.
Entity resolution matters too. The harder half is relationship recovery: linking agent sessions to the commits, pull requests, and tickets they produced when no structured link exists. The easier half is identity resolution, mapping a developer who uses different emails across version control, the project management system, and Claude Code to one contributor. minware's patent-pending hypercube data model handles both, alongside normalizing each vendor's data into one model.
If code generation accelerates but review, testing, and deployment remain constant, the pipeline creates a bottleneck that prevents delivery gains. PR lead time and ticket cycle time are where you look. Ticket work in progress (WIP) shows whether work is accumulating, and sprint completion trends in minware's predictability reports show whether teams are completing what they commit to.
Building defensible AI ROI reports for the board
The board-ready report combines the regression slope, the R-squared value, the same-team trend, and the guardrail metrics into a single narrative. Each element answers a specific question the board will ask.
Structuring the report
-
Regression slope: The change in each outcome (roadmap value delivery, on-time delivery rate, and overhead cost) per dollar of token spend, with the R-squared value showing how much of the variation spend explains. This is the ROI figure.
-
Same-team trend: Each outcome's trend over time for the same team, controlling for team-level differences. This is the supporting check.
-
Guardrail status: Bug rate and PR lead time alongside the regression, confirming neither degrades as spend scales.
-
Named confounders: Task-level selection, person-level selection, and codebase covariates, stated as limits, with the same-team trend noted as the control for team-level differences.
Addressing the "correlation is not causation" objection
This objection will come up, and the honest answer is that the analysis is observational. It controls for bias in production data using regression and same-team trends.
The analysis holds up to this objection for three reasons: the regression avoids the selection bias of an AI-vs-non-AI split, the same-team trend controls for team-level differences, and the named confounders set out what the analysis cannot rule out. For a board review, the slope-plus-trend methodology is the practical approach in a production environment where controlled experiments aren't feasible.
Selecting evidence for the ROI claim
Not every metric belongs in the board report. For ROI reporting, lead with value delivery metrics (roadmap value delivery, on-time delivery rate), followed by cost metrics (roadmap delivery cost, overhead cost), with efficiency and quality metrics as supporting signals. minware's cost attribution reporting connects contributor effort to specific roadmap items, which supplies the roadmap delivery cost and overhead cost figures. Throughput metrics (pull requests merged, commits, lines of code) and activity metrics (agent sessions, AI usage) serve as supporting data but should not be the primary ROI claim.
"The platform comes with a great set of engineering reports out of the box, so you can start getting value immediately. At the same time, those reports are highly customizable, and minQL makes it possible to build your own metrics and dashboards tailored to your team's workflow rather than being limited to predefined reports." - Verified user on G2
Weighing build vs. buy for measurement infrastructure
Building this measurement capability internally means connecting Claude Code's usage data to your version control and project management systems. The pipeline then has to recover relationships that carry no structured link, such as an agent session to the ticket it served, resolve identities, and normalize data across vendors. The initial build is a large upfront project, typically multiple weeks of engineering time. Maintenance continues after launch as vendor APIs and OpenTelemetry schemas change.
| Factor | Custom pipeline | minware |
|---|---|---|
| Upfront effort | Multiple weeks of engineering time to build the pipeline | Connect data sources, backfill in hours, then configure any org-specific report definitions (customer success can handle this) |
| Entity resolution | Build relationship recovery and identity matching in-house | Handled by the patent-pending hypercube data model |
| API maintenance | Team maintains as vendor APIs change | minware maintains compatibility |
| Metric transparency | Build calculation logic internally | Every formula visible and editable via minQL, minware's formula language |
| Stakeholder questions | Whoever built it answers "how is this calculated" indefinitely | Customer success answers in a support conversation |
| Cost | Engineering salaries plus opportunity cost | $25/contributor/month Professional, no seat minimum |
The recurring cost people underestimate most is answering "how exactly is this calculated" every time a new stakeholder asks, at every board review, budget cycle, and org restructure. With a vendor, that is a support conversation. With an internal build, it is whoever built the pipeline, indefinitely.
Proving Claude ROI with the slope and the trend
The CFO's question about Claude has a measurable answer once token spend and delivery data sit in the same model. Regress each value or cost outcome against token spend across teams and time periods, then read each slope as the ROI figure. Each team's trend over time is the supporting check. Bug rate and PR lead time confirm the gains aren't leaking through quality problems or pipeline bottlenecks. With the confounders named, that combination holds up in a board review without a pre-AI baseline.
Seeing the regression on your own data
Start a 14-day free trial at minware.com, no credit card required. Connect your version control, project management, and Claude Code data sources to see the spend-to-outcome regression on your actual data. The slope and R-squared value sit alongside the guardrail metrics, so you can test the methodology with your own teams before your next board review.
FAQs
How do I measure Claude impact without a pre-AI baseline?
The regression methodology does not require a clean pre-AI baseline because it correlates delivery outcomes against token spend level as a continuous variable, with non-AI work included as the zero-spend endpoint. The Claude Code Analytics API can also backfill historical usage at per-user, per-model, per-day granularity for Console organizations. Connect OpenTelemetry as early as you can, since it has no historical backfill.
What if our data is too messy for this methodology?
minware works with data as it exists today, and its hygiene metrics surface exactly where the process is breaking down. Each gap, such as unlinked branches or unestimated tickets, becomes a specific item someone can work through rather than a reason to delay measurement.
How long does it take to get actionable results?
Historical data backfill on first connection can take hours depending on repository size. Once backfill completes, minware's pre-built AI impact reports show the regression views. Two setup items affect how quickly results are trustworthy: team configuration should reflect your current org structure, and OpenTelemetry has no historical backfill, so connect it as early as you can.
What if token spend is high but delivery outcomes have not moved?
In minware's AI impact reports, a flat regression slope means the spend isn't associated with marginal delivery gains, and the diagnostic path runs through the guardrail and hygiene metrics. Check PR review time for review bottlenecks, bug rate for quality degradation, and ticket hygiene metrics for linking gaps that may be undercounting AI-assisted work.
Key terms glossary
Token spend level: The dollar cost of AI tool usage per person-day, used as the continuous input variable in the ROI regression. On flat-rate, tiered, or discounted plans, reported spend may differ from the amount actually invoiced.
Roadmap value delivery: The expected value of completed work items. Configurable to any project management field where a team sets explicit value estimates during roadmap planning, or to a custom spreadsheet upload. If no explicit value field is configured, defaults to story points completed on tickets that are not bugs and have a parent epic/project ticket, with 1 point assigned per ticket if no estimate is set. One of the three outcomes minware's Lean AI Framework regresses against token spend.
On-time delivery rate: The estimated roadmap item duration divided by the actual roadmap item duration, capped at 100% per roadmap item. By default, minware treats each epic as a roadmap item.
Spend-to-outcome regression: A linear regression of a delivery outcome metric against token spend level, grouped by team and time period, with each data point normalized to per-person-day. The slope is the ROI readout, and the R-squared value indicates how much of the outcome's variation the spend explains.
Overhead cost: Personnel cost and AI token spend that goes toward work not counted in roadmap value delivery, meaning cost not traceable to any ticket, or traceable to a ticket excluded from roadmap value delivery (by default, a bug ticket or a ticket with no parent epic).
Guardrail metrics: Quality and efficiency metrics (bug rate, PR lead time, ticket cycle time) that confirm AI spend is not creating bottlenecks or degrading quality. They support the ROI narrative as diagnostics but do not serve as the ROI evidence itself.
Bug rate: A quality metric measuring the number of bugs created divided by the number of code changes, by default the number of pull requests merged into a main branch.
Entity resolution: The combined capability of matching a developer's identity across systems with different usernames or emails, and recovering relationships between commits, tickets, pull requests, and AI agent sessions that carry no structured connection between them. minware's patent-pending hypercube data model handles both, allowing analysis to start before source data is fully cleaned up.
PR lead time: The time from a pull request's first commit to when it was deployed. Deployment defaults to when the code merges into a main branch, and can be configured to use Git tags or CI/CD deployment pipeline runs.