Claude Usage Governance Maturity Model: Where Does Your Organization Stand?

All Posts
Share this post
Share this post

TL;DR: This is an informal self-assessment that follows a standard four-stage maturity pattern: ad hoc, managed, integrated, and optimized. Score yourself with the rubric below. minware supports measurement at the integrated and optimized stages once your usage, git, and ticket data are connected. It does not define the stages themselves. The integrated-to-optimized jump typically requires connecting that data rather than writing more policy. If you cannot answer whether higher Claude spend is buying more delivery, you are likely not at optimized yet, no matter how complete your policy documents are.

Most AI governance maturity models are flattering rubrics. They let every org score itself a 3 or 4 because the criteria are document deliverables: a policy exists, an owner is named, a review happened. This one scores you on the hardest question you can answer with your own data.

Your Claude Code bill is climbing, and the board will eventually ask what it bought. Your governance maturity level determines whether you can show them. Each level below comes with diagnostic questions, the requirements to be there honestly, and the specific investment that moves you up.

Level 1: Ad hoc, informal dashboard checks

Ad hoc means someone opens a vendor dashboard when a bill looks high or a question comes up, with no fixed cadence and no named owner. Engineers expense individual plans. Usage typically spreads across multiple surfaces such as an integrated development environment (IDE), browser, mobile, and the command line (CLI). The first governance signal is usually a surprise bill or a security question nobody can answer.

The multi-surface problem is what makes Level 1 dangerous. Traditional procurement often assumes one tool, one contract, one access point. Claude Code shows up wherever developers work, and traditional developer productivity metrics cannot see it.

Diagnostic questions for level 1

  • Do you know how many agent sessions ran last month?
  • Can you name which repositories Claude Code touched?
  • Do you know your total token spend across all surfaces?
  • Is there a named owner for AI tool governance?

If you answered no to two or more, you likely are at Level 1.

Operational symptoms of level 1 usage

Typical symptoms include individual credit card expensing, no allowlist of approved tools, security exposure from unreviewed agent permissions, and no way to reconstruct what happened when something goes wrong. Usage data exists in fragments, and our guide to Claude Code usage data shows how scattered those fragments are across vendor surfaces.

Requirements for level 2 progression

Common exit criteria include: an inventory of every Claude surface in use, a named governance owner, and a consolidated view of spend. Level 2 is policy work, and you cannot write policy against an inventory you do not have.

Level 2: Managed, structured review and budgets

Managed means dashboard review happens on a set cadence rather than only when something looks wrong, and AI spend has a defined budget. Documented acceptable use, an approved tool list, and basic seat-level visibility sit alongside that cadence. What does not exist yet is any connection between usage and delivery.

Key indicators for level 2 status

You have an acceptable use policy covering AI coding tools, a tool allowlist, data handling rules (what code or data can reach external models), and consolidated billing. Seat counts are visible, though seat counts alone tell you nothing about value, a distinction covered in the engineering KPIs guide.

Core requirements for level 2

Typical requirements are documentation and ownership. That means an acceptable use policy, a tool allowlist, data handling rules, and a named owner with actual authority. Teams that already run working agreements will recognize the pattern, because AI usage policy is the same discipline applied to a new surface.

Steps to reach level 3 maturity

Move from a reviewed budget to connected data. That means getting AI coding tool usage, git, and ticket data into one system, so cost attributes down to the individual ticket rather than staying an estimate. Policy without data connection is Level 2 with extra paperwork.

Level 3: Integrated, connected usage and ticket data

Integrated means AI coding tool usage data, git activity, and ticket data live in one connected system instead of three separate exports. That connection is what makes cost attribution and reporting possible at the individual ticket level.

Benchmarks for level 3 governance

Typical benchmarks: usage data reconciles against your own billing, every AI-touched repository maps to a ticketing project, and a given ticket's AI cost can be pulled on demand instead of estimated.

Data quality standards for level 3

Getting here is a data quality problem: one consistent identity per engineer across every tool, complete git-to-ticket linkage, and no silent gaps between what the usage export shows and what actually ran. Our Claude Code monitoring checklist walks through that connection sequence.

Roadmap to optimized

Once usage, git, and ticket data are connected, the next investment is metric governance: token spend regressed against delivery output, on a continuous basis, with guardrails and definitions every team applies the same way. That is Level 4, optimized, covered next.

Level 4: Optimized, governed and consistent ROI metrics

Optimized means AI ROI metric definitions, the regression, the guardrails, are documented once and applied the same way by every team. The defining capability is a linear regression of a value delivery metric against token spend, read off the slope, with an R-squared value showing how much of the variation the spend explains. Token spend is the dollar cost of AI tool usage. Token consumption, a separate figure, counts the tokens processed.

Diagnosing your level 4 status

  • Can you show the regression slope of roadmap value delivery against token spend?
  • What is the R-squared value on that regression?
  • Which guardrail metrics sit alongside it?

If the answers are no, unknown, and none, you likely are at Level 3 regardless of what your policy binder says.

Moving beyond adoption to ROI metrics

The methodology matters more than the tooling. Correlate token spend as a continuous variable against delivery outcomes (roadmap value delivery, on-time delivery rate) via linear regression, grouped by team and time period. AI-assisted versus non-AI comparison is a supporting check where a genuine non-AI cohort still exists. It carries real confounders: developers who reach for AI on straightforward work and avoid it on harder problems produce data showing non-AI work is slower, when the cause is task selection.

Early adopters are often already the strongest performers, which compounds the effect. Guardrail metrics like PR lead time and bug rate (sometimes called rework rate) confirm the spend is not creating bottlenecks or quality problems downstream. They function as diagnostics rather than as ROI evidence.

This is where minware fits. Our AI impact reports run this regression against your own data, and the DORA framing of AI investment explains why workflow metrics support the answer without being the answer.

Advancing your AI governance framework

Metric transparency is what keeps this from drifting. Guardrails pair with delivery metrics: bug rate alongside roadmap value delivery, so a rising slope cannot hide a quality problem.

"It gives us clear visibility into code quality, defect rates, and overall development health." - George V. on G2

Extending optimized governance to session hygiene

Once ROI metric definitions are governed and consistent, the same discipline extends to AI session hygiene: Human Intervention Rate, Agent Session Length, Cloud Agent Use Rate, and Cloud Agent Success Rate. These sit alongside PRs Traceable to Ticket and Tickets Completed with Estimate. Weak ticket linking quietly undercounts AI-assisted work in delivery metrics. Weak session hygiene hides where agents are stalling or needing rescue. Our best practices report surfaces those gaps as a list of specific failing items rather than a single average.

Resolving entities and recovering relationships

Getting here requires entity resolution, relationship recovery, and data normalization. All three span Claude sessions, commits, pull requests, and tickets. Agent sessions often produce no commits, and commits sometimes get made manually for agent work, so the associations have to be recovered rather than assumed. That recovery is the hard half of the problem and the subject of the patent-pending hypercube data model. The same linkage is what makes session-level signals like Human Commit Rate and Pre-Merge or Post-Merge Churn Rate calculable, beyond just ticket-level ones.

Replacing periodic review with continuous feedback

At this stage, continuous feedback loops replace quarterly reviews. A flagged outcome metric drills down to its supporting metrics, then to a slice by any dimension, then to the individual failing tickets, pull requests, or agent sessions behind it, each linked back to the source system. Balancing speed against quality at this level becomes a tractable measurement problem, with a clear methodology for pairing delivery and quality metrics. Cross-functional trust follows when engineering and product share metrics. Customization keeps the model honest as the org changes:

"The platform comes with a great set of engineering reports out of the box, so you can start getting value immediately. At the same time, those reports are highly customizable." - Verified user on G2

Benchmarking your current AI governance maturity

This is the self-assessment artifact. Run this AI governance maturity assessment honestly: score each criterion 0 (not true), 1 (partially true), or 2 (fully true), then total by level.

Benchmarking your AI maturity level

Level Criteria Score (0-2 each)
1: Ad hoc Vendor dashboards get checked, but only reactively.
No fixed review cadence or named owner.
No inventory of which AI surfaces are in use.
___ / 6
2: Managed Dashboard review happens on a set cadence.
AI spend has a defined budget.
An acceptable use policy and tool allowlist exist.
___ / 6
3: Integrated AI usage, git, and ticket data live in one connected system.
Cost attribution resolves to the individual ticket.
Usage data reconciles against actual billing.
___ / 6
4: Optimized AI ROI metric definitions are documented and applied the same way by every team.
The spend-to-delivery regression runs continuously.
Anyone can see how a metric is calculated.
___ / 6

Your level is typically the highest level where you score at least 80% of the available points.

Translating your score into action

Read the score honestly. If you scored yourself a 4 but cannot produce a regression slope, you are a 3. The next investment by level:

  • Level 1 needs an inventory
  • Level 2 needs enforcement
  • Level 3 needs telemetry connected to delivery data
  • Level 4 needs guardrails, transparency, and drill-down workflows in managers' hands

Optimizing your AI governance strategy

This Claude governance readiness checklist sequences the next 90 days, assuming you are moving from integrated to optimized:

Milestone Suggested Target Suggested Owner
Connect Claude Code telemetry (OpenTelemetry or Analytics API) Days 1-14 Platform team
Connect version control and project management Days 1-14 Platform team
Configure team structure and identity resolution Days 14-30 Eng ops
Verify historical backfill completed Days 30-45 Eng ops
Run first token spend regression with guardrails Days 45-60 Eng leader
First board-ready readout Days 60-90 Eng leader

Connect telemetry as soon as you can. OpenTelemetry captures data only from the moment it is configured, with no historical backfill, so every week of delay is a week of granular data you do not have.

Defining requirements for each governance level

This consolidated reference covers the AI tool governance maturity levels as an audit checklist:

  • Level 2 (managed) readiness: A review cadence and a spend budget, plus a full inventory of Claude surfaces, consolidated spend, and a named owner with the authority to enforce policy.
  • Level 3 (integrated) requirements: AI coding tool usage, version control, and ticketing data connected into one system, so cost attribution resolves to the individual ticket. The Claude Code Analytics API provides the per-user, per-model, per-day usage data that feeds that connection.
  • Level 4 (optimized) indicators: Token spend regresses against roadmap value delivery and on-time delivery rate, with the slope and R-squared documented the same way for every team. Guardrail metrics (bug rate, PR lead time) and session hygiene metrics (Human Intervention Rate, Agent Session Length) are tracked as ongoing, actionable lists, with entity resolution linking sessions to tickets and every formula visible in the UI.

Avoiding pitfalls at each maturity stage

Three failure patterns account for most stalled progressions. Each one is level-specific and each one is avoidable if you recognize it before you invest in the wrong thing.

Sequencing documentation before reporting

Trying to report ROI before policy and data foundations exist produces numbers nobody can defend. Documentation alone is not governance, but reporting without it is worse, because a regression built on unlinked data attributes delivery to the wrong work. Sequence matters: inventory, policy, enforcement, then measurement.

Managing Claude without usage data

Governing Claude without usage data means writing policy against a black box. Four legitimate data sources exist, each with different granularity: OpenTelemetry export, the Claude Code Analytics API, the Claude Enterprise Analytics API, and a GitHub app integration that attributes merged-PR code changes back to sessions (Claude for Teams and Claude for Enterprise plans only, not Console or API customers):

Data source Granularity Historical backfill
OpenTelemetry export Per-prompt, per-session None
Claude Code Analytics API Per-user, per-model, per-day Yes, no stated deletion period
Claude Enterprise Analytics API Per-user, per-day From January 1, 2026 onward
GitHub app integration Per-pull-request attribution Not applicable, attributes activity after the fact

Pick based on the questions you need to answer, and do not wait for a clean baseline before connecting.

Moving beyond login rates to delivery impact

The most common Level 4 failure is presenting activity or throughput metrics as ROI evidence. Agent session counts prove the tool is being used, and pull requests merged prove code output rose, but neither proves delivery improved.

LinearB pairs AI metrics with workflow automation but its metrics don't connect well to delivery outcomes, while minware's regression connects token spend directly to roadmap value delivery and on-time delivery rate. Jellyfish offers automated AI impact reporting but with a rigid data model and limited metric customization. minware's approach reads the incremental ROI off the slope of outcomes against spend. The board question gets a slope instead of a login count.

A 2021 industry survey found manual, repetitive pipeline maintenance to be a recurring burden for data teams. The cost people underestimate most is answering "how exactly is this calculated" every time a stakeholder asks, indefinitely.

Option Setup effort Maintenance effort Metric transparency Ongoing cost
Internal pipeline Significant data engineering investment Ongoing owner for API changes, schema updates, and answering "how is this calculated" every time a stakeholder asks You build it or it does not exist Engineer time for every stakeholder question
minware Connect sources, backfill during setup Vendor handles API evolution Every formula visible and editable in the UI $25/contributor/month Professional, no seat minimum

Pricing is publicly listed at $25 per contributor per month for Professional with no seat minimum, and the same guide covers the spend side of the governance equation.

Moving from assessment to action

Governance maturity is set by the hardest question you can answer with your own data, and the jump from integrated to optimized typically requires a metric governance investment. Levels 1 and 2 are primarily policy and cadence work. Level 3 is a data connection investment. Level 4 requires that connected data regressed against delivery outcomes, with guardrails alongside and the same definitions applied by every team. Run the rubric, find your level, and make the investment that gets you a defensible answer at the next board review.

Start a 14-day free trial, no credit card required. Connect your version control, project management, and Claude Code data to see token spend correlated against delivery outcomes.

FAQs

Can we skip a level if we're moving fast?

Not cleanly. Levels 1 through 3 build the data and integration foundation the optimized-level regression depends on, so you can compress the timeline. You cannot regress token spend against delivery outcomes, though, without usage data flowing and integrated.

How long does it take to move up one level?

Timelines vary widely, but treating this as a formal program tends to move faster than ad hoc effort. Level 1 to 2 is primarily policy work and often moves in a matter of weeks. Level 3 to 4 depends more on how quickly telemetry, version control, and project management data get connected, then backfilled, since those are usually the binding constraints.

What if different teams are at different levels?

That is normal. One approach is to score the org at the lowest common denominator for policy levels and the highest for measurement, because you can run Level 4 measurement on one team while others catch up on policy.

What data do we need to assess our current level?

Version control, project management, and Claude Code usage data (OpenTelemetry export or an Analytics API) are typically the minimum for Level 4. Levels 1 and 2 can often be assessed from policy documents and seat data alone. Level 3 additionally requires confirming the data connections are actually in place, not just documented.

Key terms glossary

Token spend: The dollar cost of AI tool usage, as reported by AI tool APIs or OpenTelemetry. On flat-rate, tiered, or discounted plans, reported spend may differ from the amount actually invoiced.

Roadmap value delivery: The expected value of completed roadmap work, defaulting to story points completed on tickets tied to a roadmap item unless the org configures its own value estimate.

On-time delivery rate: The share of roadmap items delivered by their original due date, capturing the cost of missed deadlines.

Regression: Statistical method fitting a delivery metric against token spend as a continuous variable. The slope represents the expected change in delivery for each additional dollar of spend, and the R-squared value shows how much variation the spend explains.

Guardrail metrics: Quality and workflow metrics such as bug rate and PR lead time that confirm a primary delivery metric is not being gamed.

Bug Rate: A quality metric measuring the number of bugs created divided by the number of code changes, by default the number of pull requests merged into a main branch. Some customers configure the denominator to story points completed instead, to reduce the risk of PR volume changes distorting the metric.

Human Intervention Rate: The count of permission escalations and other prompts occurring in between instructional prompts, divided by the number of instructional prompts in the session. Common intervention types: permissions escalation, agent convention violation, human manual testing, human-directed testing, human information lookup, agent off-track, agent mistakes.

Agent Session Length: The duration of an agent session, measured directly from its start and end times in AI tool session data. No default target is recommended; teams review their longest sessions periodically instead.

PR lead time: The time from a pull request's first commit to deployment, which defaults to when the pull request merged but can be configured to a later deployment event.

Ticket cycle time: The time from when a ticket was first marked in progress to when it was completed.

DORA metrics: The five metrics defined by the DevOps Research and Assessment framework: deployment frequency, lead time for changes, change failure rate, failed deployment recovery time, and deployment rework rate.

OpenTelemetry: An open-source observability framework for exporting telemetry, such as prompts and tool-use events, from instrumented applications, including AI coding tools, to a collector. Reports at per-prompt, per-session granularity, with no historical backfill for activity before it was configured.

Hypercube data model: minware's patent-pending data model that links every entity from every data source, so any metric can be broken down by any dimension.