Claude Code OTel Data Accuracy: What Engineering Leaders Need to Know

All Posts
Share this post
Share this post

TL;DR: Claude Code OpenTelemetry (OTel) data is technically reliable for auditing token spend, tool invocations, and session boundaries at per-prompt, per-session granularity, but it is not board-ready on its own. A normalized data layer closes the gaps by reconciling telemetry spans with Git commit metadata and surfacing missing links as itemized findings, turning "our data is too messy" into a prioritized cleanup list rather than a reason to avoid metrics.

Claude Code's OTel export gives you token spend and session counts, but connecting those signals to a board-ready return on investment (ROI) number requires bridging an attribution gap. The most common reason engineering leaders distrust Claude Code OTel data is that it is incomplete, and nobody has told them where the gaps are.

This article names the specific reliability boundaries of Claude Code OTel data accuracy and explains what a governance layer adds before the numbers can survive an executive review. It also gives you a checklist for validating telemetry against your own billing statements and delivery metrics.

Measuring Claude Code OTel data accuracy

Claude Code emits three OpenTelemetry Protocol (OTLP) signal types: metrics, events, and traces. Metrics and events come with standard telemetry. Traces, which is where span-level detail on tool calls and LLM requests lives, are a beta feature that requires a separate flag in addition to standard telemetry, so don't assume every org has span data available by default.

That granularity is the strength of the export once traces are enabled. Trace spans cover LLM requests and tool executions. For the questions OTel was designed to answer (what did a session do, what did it cost, which tools ran), the raw data is dependable.

The weakness sits in what OTel does not attempt to capture. The export knows nothing about your Jira tickets, your roadmap epics, or which pull request a session produced. Those relationships live in other systems, and connecting them is a governance problem rather than a telemetry problem.

Confirming Claude Code telemetry reliability

Claude Code telemetry reliability holds up well for three specific signal classes:

  • Token spend: Tracked per session, useful for monitoring consumption trends. One qualifier: reported spend may not match billing on flat-rate plans, covered in more detail below.

  • Tool invocations: Tool calls appear as discrete spans once tracing is enabled, so you can verify what an agent actually executed.

  • Session boundaries: Sessions mark where work began and stopped, which anchors time-based analysis. A governance layer should still reconcile totals against a second source to confirm completeness.

Weighing OTel data trustworthiness

Is Claude Code OTel data trustworthy enough for a board slide? For the signals it measures directly, yes. For the claims boards actually ask about, no, not without reconciliation.

The question a CFO asks is whether higher token spend is buying more delivery or just a bigger bill. Answering it requires a linear regression of a value delivery metric (story points completed or roadmap delivery) against token spend. The delivery side lives in your project management system, and the linkage between the two is where raw telemetry falls short.

So the honest answer to "is Claude Code OTel data trustworthy" is: trustworthy as plumbing, incomplete as a report. Activity metrics like session counts and token totals show the tool is being used, but they cannot carry a delivery claim on their own.

Assessing OTel accuracy gaps

How accurate is Claude Code OTel at the level of individual attribution? This is where the gaps get specific, and each one is fixable.

Trace chains can break when context isn't propagated. In distributed tracing generally, trace context fails to propagate between services when it isn't injected into the payload, and the consumer starts a new, disconnected trace instead of continuing the existing one. This is a documented failure mode in distributed tracing broadly. Whether it shows up in your own Claude Code sessions is something to verify against your own trace data rather than assume, since there is no Claude Code-specific documentation confirming local-terminal usage as a cause.

OTel has no historical backfill. Unlike the Claude Code Analytics API, which reports at per-user, per-model, per-day granularity with full historical backfill, OpenTelemetry (OTel) only captures data from the moment you configure it. Connect telemetry as soon as you can, because the coarser Claude Code Analytics API is the only fallback for earlier periods (the Claude Enterprise Analytics API, by contrast, has no data before January 1, 2026).

Reported spend may differ from billed spend. For consumption-based deployments, Anthropic's data puts average Claude Code cost at $150-250 per developer per month, reflecting actual token consumption rather than a fixed seat price, a figure our AI cost management analysis covers in more depth. For teams on flat-rate plans like Claude Pro, reported token spend may be higher or lower than what's actually billed, since usage draws against a fixed allotment rather than metered spend.

Rating Claude Code OTel signal reliability

Signal / capability What it measures Reliable as raw data? Governance required
Token spend Dollar cost per session Useful, with flat-rate qualifier Reconcile against billing statements
Tool invocations Which tools ran per prompt Yes, once tracing (beta) is enabled None for audit purposes
Session boundaries When work started and stopped Yes Reconcile totals against a second source
Entity resolution Who did the work No Identity matching and relationship recovery
Branch-to-ticket linking What the work was for No Linking enforcement plus recovery
Delivery outcomes Whether spend bought delivery No Regression against value delivery metrics

The table is the whole argument in one view: raw OTel signals capture activity, and board-ready metrics require a governance layer to link telemetry to delivery outcomes.

Identifying workflow gaps through OTel signals

The useful reframe treats OTel gaps as findings rather than failures. When a session has no linked commit, or a commit has no linked ticket, that is a specific, countable item someone can fix.

Visualizing workflow bottlenecks

Session-level telemetry shows where agent work concentrates, and pairing it with version control data shows where that work stalls. If agent sessions spike while PR review time lengthens, this may indicate the review stage is absorbing the extra volume, which is a workflow bottleneck you can name and address. Our guide on invisible wait time metrics covers how these stalls hide in distributed teams, and the cycle time vs. lead time breakdown explains which flow metric isolates which stage.

Fixing broken telemetry attribution

Attribution breaks in predictable places, and the harder half breaks first: relationship recovery, connecting an agent session to the ticket it belongs to when that session produced no commits. minware's hypercube data model recovers these hidden associations through advanced time-based linking that considers what work each contributor was engaged with, as detailed in our GitHub Copilot tracking guide.

The easier half is identity resolution. Engineers typically carry multiple identifiers across systems: a version control username, a Git author email, a ticketing display name, a Single Sign-On (SSO) identity. As our Cursor tracking guide describes, minware automatically maps these disparate accounts to a single contributor profile and links agent sessions to commits, pull requests, and tickets even where no structured relationship exists between them.

Detecting telemetry attribution errors

Batch and unattended agent runs are the hardest attribution case. When an agent runs without a human at the keyboard, there is no developer session to anchor the work, and naive pipelines either drop the activity or misattribute it.

Time-based linking handles this: session timestamps are matched against commit and ticket timestamps to build relationships even for unattended runs. Explicit branch-to-ticket linking still improves accuracy, because developers often have several tickets in progress at once, so the recovery model and the linking practice work together rather than substituting for each other.

Cleaning data gaps through normalized reporting

The "our data is too messy for metrics" objection deserves a direct answer: messy data is the starting condition for every engineering org, and the gaps are exactly what best practice metrics exist to surface.

Fixing data quality through measurement

Measurement is what makes the mess visible and countable. Our best practices report tracks the rate of PRs traceable to tickets and tickets completed with estimates, so weak linking shows up as a number with a list of failing items behind it rather than a vague suspicion. Each unlinked branch becomes a specific item a manager can assign.

This matters for AI reporting specifically because story point velocity, roadmap delivery, and sprint completion typically depend on ticket linkage. Weak linking makes these metrics quietly undercount AI-assisted work, while code-side metrics like PR lead time and pull requests merged calculate from version control alone.

Cleaning up telemetry data proactively

Proactive cleanup works as a loop: measure the linking rate, fix the worst offenders, and watch the rate climb. Because relationships are recovered rather than required, analysis can start before source data is clean, and each cleanup cycle makes the recovery model's job easier.

Normalizing your telemetry data

  1. Connect version control and project management first. Your version control system (such as GitHub, GitLab, Bitbucket, or Azure DevOps) provides commit and pull request data, and the project management system supplies the delivery side of every ROI claim.

  2. Deploy OTel configuration once for the team. OTel works at every plan level through local file-based configuration. Certain plan levels add central settings management, and mobile device management (MDM) tools can push the same configuration on any plan.

  3. Review entity resolution results. Confirm that relationship recovery linked agent sessions to commits, pull requests, and tickets where no structured connection existed. Check that identity matching handled your team's mismatched emails and usernames. Verify that vendor data normalized cleanly into one model. Use the identity linking configuration where manual overrides are needed.

  4. Audit the findings list. Work through unlinked branches and unattributed sessions as a prioritized backlog, starting with the highest-spend teams.

Defining requirements for board-ready OTel data

Board-ready means three things: the spend figure reconciles to billing, the delivery metric is a value delivery metric, and the relationship between them is a regression slope rather than a raw ratio.

Correlating token spend to delivery outcomes

The methodology is a linear regression of roadmap value delivery (which defaults to story points completed on non-bug tickets tied to an epic) against token spend, which can be grouped by team and time period, per minware's Lean AI Framework. The regression slope represents story points per dollar of token spend. The R-squared value tells you how much of the variation in delivery the spend actually explains. The framework recommends at least 20 to 30 data points per regression to distinguish a real slope from noise, so extend the time range or combine teams if a single team's data falls short. As our Claude Code monitoring checklist notes, this correlation can include non-AI work, so it holds up as adoption reaches saturation and a clean non-AI cohort disappears.

Cycle time and bug rate sit alongside the regression as guardrails, per the same framework. Cycle time should hold steady or fall as spend rises, confirming the spend is not creating a bottleneck downstream, and bug rate catches quality erosion that a delivery metric alone would miss. Our DevOps Research and Assessment (DORA) metrics and AI investment piece covers how these supporting signals fit the broader framework.

"I use Minware for our SDLC metrics and appreciate its ability to provide quality metrics throughout our SDLC. It gives us clear visibility into code quality, defect rates, and overall development health." - George V. on G2

Confirming OTel signal accuracy

Use this checklist before your next executive review:

  • Audit the pipeline: Confirm the OTel integration is configured across the team and sessions are flowing, then spot-check session counts against available analytics sources for the same period.

  • Reconcile spend: Compare reported token spend against actual billing statements, and flag any seats where reported and paid amounts diverge.

  • Check the data range: Confirm delivery data is flowing across enough teams and time periods to reach the 20 to 30 data points the regression needs.

  • Normalize identities and relationships: Verify each contributor resolves to one profile across version control, ticketing, and AI tools, that agent sessions are linked to the commits and tickets they belong to, and that vendor data formats are normalized into one model.

  • Assess Total Cost of Ownership (TCO): If you are weighing a custom pipeline, price the ongoing maintenance, including schema changes and the recurring "how is this calculated" question from stakeholders.

  • Focus on outcomes: Anchor the report on story points completed and roadmap delivery, with cycle time and bug rate as supporting signals.

Verifying Claude Code data integrity

Data integrity also has a security dimension. Per the OTel sensitive data guidance, spans can carry sensitive data including authentication credentials and session tokens. Whether prompt text and tool-use detail reach your analytics platform is configured in your own OTel setup: including it produces more granular reporting, and redacting it still supports delivery correlation. Redaction in the collector before export is one recommended pattern.

Handling Claude Code OTel accuracy with minware

This is where minware fits. minware ingests Claude Code OTel data through its OpenTelemetry integration, alongside version control, project management, and deployment data. Its hypercube data model then does three things before any metric is computed: it normalizes messy vendor formats into one canonical model, resolves identities across systems, and recovers relationships between entities that carry no structured link, including agent sessions that produced no commits.

minware makes metric formulas visible and editable in the UI through minQL, so when a stakeholder asks how a number is calculated, the answer is an inspection rather than a support ticket. Our pre-built AI impact reports run the token-spend regressions described above, and the project completion report connects the same data model to roadmap burnup and estimated completion against due dates.

"What I like best about minware is its flexibility. The platform comes with a great set of engineering reports out of the box, so you can start getting value immediately. At the same time, those reports are highly customizable, and minQL makes it possible to build your own metrics and dashboards tailored to your team's workflow rather than being limited to predefined reports." - Verified user on G2

Two scope notes worth stating plainly. First-connection backfill time varies depending on repository size, and OTel data starts accumulating only from configuration forward. Custom metric work typically turns around in a single call and about 24 hours through customer success, so the learning curve for minQL rarely lands on your team.

For teams weighing build vs. buy, a custom OTel pipeline may be feasible for a small team with one tool stack, but that estimate misses two costs: a high upfront build, typically multiple weeks of engineering time, and the ongoing maintenance after. Vendor schemas evolve, and the maintenance burden, including the recurring metric-definition questions, falls on someone indefinitely. Span offers proprietary AI-generated-code detection, and minware answers that need with time-based session attribution that works without proprietary signals. Jellyfish offers strong cost capitalization reporting, and minware provides the same linkage with visible metric formulas where every calculation is inspectable rather than proprietary.

Closing the accuracy gap with governance

Claude Code's OTel export is reliable for what it was built to measure: token spend, tool invocations once tracing is enabled, and session boundaries. It was never built to know about your tickets, epics, or pull requests, and that gap is what turns a technically sound export into an unconvincing board slide. A governance layer that resolves identities, recovers relationships, and normalizes vendor formats closes that gap, turning "our data is too messy" into a prioritized, fixable list rather than a reason to avoid the metrics altogether.

Start a 14-day free trial at minware.com, no credit card required, and connect your version control, ticketing, and Claude Code telemetry to see how minware reconciles them against your own data. Pricing starts at $25/contributor/month.

FAQs

Is Claude Code OTel data trustworthy for executive reporting?

Yes for the signals it measures directly: token spend, tool invocations (once tracing is enabled), and session boundaries at per-prompt, per-session granularity. It needs a governance layer for entity resolution and ticket linkage before it can support a delivery ROI claim.

How accurate is Claude Code OTel without a governance layer?

Accurate for activity auditing, incomplete for attribution. Identities, commits, and tickets may remain unlinked until a normalized model reconciles them.

Ticket-dependent metrics like story point velocity and roadmap delivery may undercount the work, including AI-assisted work. Code-side metrics like PR lead time and pull requests merged calculate from version control alone.

Can I trust OTel signals from batch or automated runs?

The signals themselves are reliable, but attribution requires time-based linking that matches session timestamps to commit and ticket timestamps, since no human session anchors the work.

How long does it take to clean up messy attribution data?

Analysis starts immediately because relationships are recovered rather than required. Cleanup then proceeds as a prioritized backlog of unlinked items, with linking rates improving over subsequent reporting cycles.

Key terms glossary

Token spend: The dollar cost of AI token consumption, typically reported per session.

OpenTelemetry (OTel): A widely-used open protocol for exporting metrics, events, and traces. Claude Code uses it to emit telemetry with no historical backfill.

Tool invocations: Discrete OTel trace spans recording which tools an agent called during a session, available once Claude Code's beta tracing is enabled.

Session boundaries: The start and stop points of a Claude Code session, derived from session-count and active-time metrics, used to anchor time-based analysis.

Entity resolution: The process of recovering relationships between artifacts that carry no structured link, such as an agent session and the ticket it served, and matching a contributor's identities across systems. Relationship recovery is the harder of the two problems.

Story points completed: A value delivery metric representing the story points for completed tickets, used as the outcome side of an AI ROI regression.

PR lead time: The time from a branch's first commit to deployment, defaulting to when its pull request merges into the main branch and configurable to extend to actual deployment via Git tags or CI/CD pipeline runs, per minware's Lean AI Framework. One of the two cycle time metrics.

Cycle time: The workflow guardrail, covering minware's two cycle time metrics: PR lead time and ticket cycle time, which runs from when a ticket moves to in progress until it is marked done.

Bug rate: A quality metric measuring bugs created divided by pull requests merged (by default), used as a guardrail to catch quality erosion that delivery metrics miss.

Regression slope: The change in an outcome metric per unit change in token spend, representing the marginal value delivered for each additional dollar.

R-squared: A statistical measure showing how much of the variation in an outcome metric is explained by the variable being tested, such as token spend. Higher values indicate stronger correlation.

Hypercube data model: minware's data model that links SDLC artifacts (sessions, commits, pull requests, tickets, epics) so any metric can be compared against any dimension.

Value delivery metric: Metrics that measure completed work units, such as story points completed, tickets completed, or roadmap delivery, used to prove delivery impact.