Claude Code OTel Limitations: What Agent Telemetry Can't Tell You
TL;DR: Claude Code's OpenTelemetry export gives you session-level traces of what the agent did, but the schema carries no structured link to commits, pull requests, or tickets, so it cannot tell you whether any of that activity shipped value. Common gaps in agent telemetry include misreading acceptance rate as evidence, quality and workflow tradeoffs, causal attribution, and budget utilization. Closing them takes guardrail metrics (cycle time and bug rate), manager judgment, and minware's normalized data model, which joins agent sessions to delivery outcomes. The return on investment (ROI) readout itself comes from a token spend correlation: a linear regression of story points completed against token spend.
Claude Code's OpenTelemetry export records spans around each model request and tool execution. It also emits token and cost counters. The schema carries no structured link to commits, pull requests, or tickets. Structured log events for prompts and tool results push to any OTel-compatible collector, per Anthropic's observability documentation. That is real granularity, and it makes OTel among the most detailed of the four Claude Code usage data sources. The instrumentation focuses on agent activity, which is why minware built the hypercube data model to recover the relationships OTel does not carry.
This article walks through the Claude Code OTel scope limits that matter before you treat telemetry as your AI measurement backbone: what Claude Code telemetry cannot tell you about developer experience, code quality, causal attribution, and spend governance, and what actually fills each gap.
Gap 1: Misreading acceptance rate as evidence
OTel answers "what did the agent do." Engineering leaders get asked "did it matter." Those are different questions, and the gap shows up the moment an activity signal gets treated as proof of impact.
Claude Code reports a suggestion accept rate across its Edit, Write, and NotebookEdit tools, so the metric exists. The problem is what it measures: accepting a suggestion sits upstream of delivery and does not guarantee that the merged code delivered value. The same logic applies to the live industry claim worth rebutting, pull request merge counts presented as AI ROI.
A rising merge count proves code output rose, and it says nothing about whether that code delivered value, which is why productivity metrics built on output volume count as broken in the AI era rather than merely incomplete. Vendor-defined metrics tend to favor the vendor's own story, with important context omitted from reporting, so any activity or throughput figure needs a delivery metric beside it before it enters an ROI narrative.
Gap 2: Balancing quality and workflow against throughput
Throughput metrics are the second gap. OTel will happily show you that sessions multiplied and tokens flowed. Whether the resulting code was any good lives in systems OTel never touches.
Measuring quality alongside throughput
minware classifies lines of code, commits, and pull requests merged as throughput metrics. They measure output volume, and volume without a quality counterpart is how teams end up tracking AI technical debt only after it lands in production. Every throughput figure presented as evidence needs its guardrail pair in the same view:
-
Pull requests merged: consider pairing with bug rate and PR review rate.
-
Commits: consider pairing with post-merge churn rate.
-
Lines of code: consider pairing with bug rate.
DORA's research shows that AI increases code generation velocity faster than review and deployment infrastructure can absorb it, which raises delivery instability even as individual output climbs. That finding is the reason the pairing is mandatory rather than optional.
Tracking bugs from a separate data source
OTel telemetry focuses on agent activity rather than downstream code quality. Bug data lives elsewhere: in version control, and in your project management system, such as Jira, Linear, Azure Boards, or GitHub Issues. minware calculates bug rate as bugs created divided by pull requests merged, drawing on version control and ticket data, though organizations may define the denominator differently depending on how they measure code change volume. No span carries that. Teams that want the full picture of speed and quality together need both sides joined, which is the premise behind balancing speed and quality rather than reporting either alone.
"It gives us clear visibility into code quality, defect rates, and overall development health." - George V. on G2
Guardrail metrics exist because telemetry alone cannot confirm that AI-accelerated output is safe to ship. They do two jobs, and neither job is "prove the investment worked."
Validating throughput against workflow
As a guardrail, you track PR cycle time as AI spend rises. If cycle time climbs while spend increases, it may signal that the spend is creating a bottleneck further down the pipeline, typically in review. The five DORA metrics (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time, and deployment rework rate) function as a system where two measure throughput and three measure stability. High-performing teams do not trade one for the other. minware's piece on DORA metrics for AI investment applies that framing directly to AI tooling.
Diagnosing trends with cycle time and bug rate
As a diagnostic, when delivery outcomes do not move, cycle time and bug rate are where you look to find out why. Bug rate is the number of bugs created divided by the number of pull requests merged, a broader signal of quality problems than change failure rate alone because it counts every bug rather than only the highest-priority ones. Change failure rate counts the deploys that needed an immediate, unplanned response.
Deployment rework rate, a separate DORA metric, counts the follow-up deploys those failures generate. Cycle time and bug rate, not change failure rate or deployment rework rate, are the two metrics that explain why a delivery trend is moving. ROI evidence comes from the regression slope. The same distinction is drawn in minware's pieces on cycle time vs. lead time and engineering KPIs.
Gap 3: Isolating AI impact from other causes
The third gap is the attribution problem: even with delivery data in hand, telemetry cannot tell you whether an improvement was caused by the tool or by everything else that changed during the same period.
Controlling for confounders during rollout
Before/after comparisons around an AI rollout are observational analysis: org restructures, a new project management tool, hiring waves, and roadmap priority shifts all move delivery metrics during the same window. minware's regression methodology controls for these by treating token spend as a continuous variable across teams and time periods rather than comparing two snapshots. The stakes are real: research indicates that 42% of companies abandoned most AI initiatives in 2025, up from 17% in 2024, and weak attribution is a plausible contributor.
Accounting for human process variables
Two selection effects confound AI-versus-non-AI cohort comparisons, and you should name both whenever presenting one. Task selection: when adoption is voluntary, developers may reach for AI on certain types of work while keeping other problems manual, which can skew comparisons for reasons unrelated to the tool itself. Person selection: adoption is not random at the individual or organizational level. MIT Sloan research, in a study of U.S. manufacturing firms, found that organizations expecting higher returns adopt AI earlier, and correcting for that selection effect reveals larger short-run negative impacts, so raw cohort comparisons flatter the tool. Token spend level is the durable variable once adoption is near-universal.
Gap 4: Tracking budget outside the telemetry stream
OTel does not export budget utilization or usage limits. The schema fires on every API call with model, cost, duration, and token counts, but forward-looking budget state lives in separate workspace spend controls in the Claude Console. You should track spend against plan in a dedicated report rather than expecting the trace stream to carry it, and remember that reported spend may differ from the amount actually paid for users on flat-rate plans such as Claude Pro.
| Question OTel cannot answer | Why OTel lacks it | What fills the gap |
|---|---|---|
| Are developers better off? | No satisfaction attributes | Developer surveys |
| Is the code good? | No version control or ticket fields | Bug rate |
| Did AI cause the improvement? | Requires causal analysis beyond observability tooling | Regression across teams/periods |
| Are we within budget? | Budget state lives in Console | Console spend controls + reports |
If you are weighing a custom collector pipeline against a platform, price the temporal logic and schema maintenance honestly, because that is where internal builds encounter unexpected complexity.
Finding the missing links in agent observability data
Some gaps in OTel are architectural, and no collector configuration closes them. Pretending otherwise produces reports that look precise and prove nothing.
The OTel GenAI semantic conventions define a standard vocabulary for spans, metrics, and events. Attributes include fields like model name, input tokens, and output tokens. That's the raw material a custom collector has to parse, and it has to keep pace as vendors evolve the telemetry format. The schema does not include a field mapping a span to a commit, pull request, or ticket. Teams running their own collector also inherit ongoing schema governance for credential management and filtering. That maintenance cost is consistently missing from the original build estimate.
Judging whether a process change worked
You also cannot learn from telemetry alone whether a process change worked. Did the new code review practice move cycle time or bug rate? Did pairing junior developers with agent sessions change bug rate? Those answers come from delivery data interpreted through manager judgment, which is why performance reviews and collaboration metrics work best as tools that augment manager judgment rather than replace it. The manager always holds final judgment, and metrics augment it.
Mapping AI ROI beyond telemetry
Here is the full scope picture across the four legitimate Claude Code data sources. None of the four sources below connect usage to delivery outcomes or to each other on their own.
| Data source | Granularity | Historical backfill | What it cannot answer |
|---|---|---|---|
| OpenTelemetry export | Session-level, per-prompt traces | None | Structured link to commits, PRs, or tickets |
| Claude Code Analytics API | Per-user, per-model, per-day | Available for historical periods | Direct ticket/commit links |
| Claude Enterprise Analytics API | Per-user, per-day | Only from Jan 1, 2026 onward, no earlier backfill | Commit or ticket links, or a per-model cost breakdown |
| GitHub app integration | Merged PRs and lines of code | Attribution window around merges (Claude for Teams and Claude for Enterprise plans only, not Console or API customers) | Sessions without commit output |
Both Analytics APIs return data at daily per-user granularity, with the Claude Code Analytics API also breaking cost down by model. The GitHub app is available on Claude for Teams and Claude for Enterprise plans only, not Console or API customers. It attributes merged pull requests to Claude Code sessions using a 21-day-before to 2-day-after matching window around each merge, excluding pull requests rewritten by more than 20% after the session. That keeps attribution to genuine session output rather than incidental overlap. OTel remains the most detailed source, but real-time session data starts the day you enable it, with no backfill.
Connecting telemetry to delivery outcomes
minware solves the core gap, relationship recovery, by linking agent sessions to commits, pull requests, and tickets even when no structured connection exists. minware's patent-pending hypercube data model is a normalized data layer that recovers those links by modeling what each person worked on at any given time. This temporal entity resolution is an active area of data engineering research, with new benchmark datasets built specifically because existing techniques are hard to evaluate, which is why custom pipelines often break on exactly this step. minware delivers the result in pre-built AI impact reports that join OTel data to quality, workflow, and value delivery metrics.
It also handles identity resolution, matching each developer's identity across systems with different usernames or emails, and data normalization, canonicalizing vendor-specific formats into one consistent model.
"The platform comes with a great set of engineering reports out of the box, so you can start getting value immediately. At the same time, those reports are highly customizable, and minQL makes it possible to build your own metrics and dashboards tailored to your team's workflow rather than being limited to predefined reports" - Verified user on G2
Regressing story points against token spend
minware's primary ROI methodology is a token spend correlation: a linear regression of story points completed or roadmap delivery against token spend, grouped by team and time period, then normalized per person-day. The slope is the ROI figure itself, points per dollar of spend, and the R-squared value shows how much of the variation in delivery the spend actually explains. Simple division of total outcomes by total spend produces a misleading average because it attributes all delivery to AI, including work humans would have shipped anyway, a mistake broken down further in minware's AI cost management practices.
In minware's AI ROI report, you can view total token spend broken down by tool, model, project, team, and person, weighed against output. The platform's drill-down capability follows the hypercube model: from an epic view you can explore its tickets. From a ticket view, you can see its commits and agent sessions. Each link carries you from the business outcome (epic shipped) to the technical activity (session trace) without leaving the same data model.
Answering what OTel cannot tell you
minware treats Claude Code OTel as a necessary data source and an insufficient answer on its own. It tells you what the agent did at session-level granularity, and it cannot tell you whether that activity improved delivery, whether the code was good, whether the tool caused the improvement, or whether you are inside budget. minware fixes this by joining telemetry to version control, project management, and delivery data through a normalized model, then reading ROI off the regression slope with guardrail metrics alongside it.
Start a 14-day free trial at minware.com, no credit card required, and connect Claude Code OTel data to delivery metrics in pre-built AI impact reports before your next executive review.
FAQs
Can Claude Code OTel prove ROI on its own?
No. OTel captures agent activity with no structured link to commits, pull requests, or tickets, so it cannot connect sessions to delivery outcomes on its own. minware proves ROI by joining telemetry to delivery data and regressing outcomes like story points completed against token spend.
What's the difference between telemetry and delivery metrics?
Telemetry records what the agent did: prompts, tool calls, tokens, and cost per session. Delivery metrics record what the organization shipped: tickets completed, story points completed, and roadmap delivery. These come from version control and project management data.
How do I measure code quality from AI tools?
Pair every throughput metric with a quality guardrail from version control and ticket data. Pull requests merged pairs with bug rate and PR review rate. Commits pairs with post-merge churn rate. OTel spans contain no quality fields, so you always need a second data source for quality measurement.
Why can't acceptance rate serve as a quality metric?
Acceptance rate measures whether a developer accepted a suggestion. It sits upstream of delivery and doesn't confirm the merged code delivered value or stayed merged. It also correlates with neither bug rate nor story points completed on its own.
What data do I need alongside Claude Code telemetry?
Version control, a project management system, and the AI tool's telemetry are the minimum minware needs for AI ROI reporting. Deployment data from CI/CD runs, Git tags, or merges to main adds the DORA metric set on top.
Key terms glossary
Agent telemetry: The session-level stream of spans, metrics, and log events Claude Code exports over OpenTelemetry. It describes agent activity but typically does not include structured links to delivery artifacts like commits or tickets.
Cycle time: The time from a pull request's first commit to when the pull request merges. Unqualified "cycle time" means PR cycle time. Ticket cycle time, from when a ticket was first marked in progress to when it was completed, is a separate, related metric.
Bug rate: The number of bugs created divided by the number of pull requests merged into a main branch. Some customers configure the denominator to story points completed instead, to reduce the risk of PR volume changes distorting the metric.
Story points completed: The total of the story points field for completed tickets. minware's primary ROI methodology, a token spend correlation, regresses this figure against token spend to read delivery impact off the resulting slope.
Guardrail metrics: Workflow and quality metrics (cycle time, bug rate) minware tracks alongside AI spend in its AI impact reports to confirm the spend is not creating bottlenecks or quality regressions. They explain trends rather than prove ROI.
Attribution problem: The challenge of determining whether a delivery improvement was caused by the AI tool or by confounding changes during the same period. You handle this as observational analysis with named confounders.
Acceptance rate: The share of Claude Code suggestions that developers accept during a session. An activity metric measuring tool usage. It does not measure delivered value.
Token spend correlation: minware's regression analysis approach that models the relationship between a delivery metric and token spend level across teams and time periods. The slope represents the marginal ROI, and R-squared shows explanatory power. This method can include non-AI work at the zero-spend endpoint.