How Claude Code Telemetry Works: OTel Metrics, Traces, and Logs

All Posts
Share this post
Share this post

TL;DR: Claude Code's OpenTelemetry (OTel) export emits three signal types: metrics for aggregate usage, logs for discrete events and errors, and traces for per-session workflow spans. These signals are correlated using session and prompt identifiers in downstream analysis platforms. Agent sessions are the unit of real usage, not CLI invocations. Token cost mapping requires separating token consumption from dollar spend, especially on flat-rate plans. Downstream platforms like minware parse these streams and regress delivery outcomes like story points completed against token spend, with the slope as the ROI answer.

Claude Code's OTel export is easy to turn on. Knowing what the signals actually contain before you approve the integration is the part most teams skip. The stream Claude Code emits is raw agent activity, and the gap between that activity and a defensible delivery claim is where most AI return on investment (ROI) stories fall apart.

This piece walks through how Claude Code telemetry works at the signal level: what metrics, traces, and logs each capture, how events move through the collector pipeline, how session boundaries structure the data, and how token accounting maps to cost. The goal is evaluation depth, not setup steps, so you can judge whether a downstream platform's claims about this data hold up.

Understanding what Claude Code OTel signals contain

The three signal types cover different layers of Claude Code activity, and understanding what each one captures determines how useful the data is for downstream analysis.

Mapping OTel signals to Claude Code events

Claude Code exports telemetry across three independent signals: metrics, structured log events, and, with beta tracing enabled, spans that show how each interaction progresses through model requests and tool calls.

  • Metrics: Aggregate counts over time, such as sessions, tokens consumed, and estimated cost. These typically feed usage and spend reporting.

  • Logs: Discrete events and errors, the primary path most Claude Code events travel through.

  • Traces: Per-session workflow spans connecting a user prompt to model calls, tool calls, and hook execution.

Each signal type answers a different question. For an engineering leader, a useful shorthand is this: metrics typically tell you volume, logs tell you what went wrong, and traces tell you sequence. That summary captures how Claude Code OTel signals work at a practical level.

Understanding Claude Code telemetry output

A raw export carries a hierarchical payload: resource attributes identifying the source, then signal-specific records with attribute keys for model, token counts, cost, and timestamps. Anthropic's own docs confirm that each API request event carries the model, cost, and token counts, tracked via the claude_code.cost.usage and claude_code.token.usage metrics.

OpenTelemetry export is one of two pipelines Anthropic exposes for Claude Code usage reporting. The other is the Analytics APIs: the Claude Code Analytics API for Console organizations and the Claude Enterprise Analytics API for Claude Enterprise organizations, which are different products. Each has different granularity and history, which the comparison table below covers.

Weighing signal costs and cardinality

Cardinality is the number of unique values an attribute takes, and it drives observability storage cost directly. Attributes like session IDs and prompt text are high-cardinality by nature: a single metric carrying several high-cardinality attributes can generate millions of unique time series, each costing money to store and query. High-cardinality attributes like user IDs or request IDs are indeed among the biggest cost drivers in observability systems, a general pattern rather than something specific to Claude Code.

The tradeoff is real: per-session, per-request granularity is what makes OTel more detailed than the Analytics APIs, and it is also what can make it expensive to store naively.

The table below breaks down each signal's granularity and export cadence, per Anthropic's defaults:

Signal type What it captures Granularity Typical export interval Use case
Metrics Sessions, tokens, cost, activity Aggregate over time 60 seconds (default) Usage and spend reporting
Logs Discrete events, errors Per event 5 seconds (default) Debugging and reliability
Traces Model calls, tool calls, hooks Per-session spans 5 seconds (default) Workflow sequence analysis

Routing OTel signals through the collector

The path from Claude Code to a downstream platform involves several configurable stages, each with its own tradeoffs for reliability and data freshness.

Tracing the Claude Code telemetry path

The OpenTelemetry Protocol (OTLP) defines the encoding, transport, and delivery of telemetry between sources, collectors, and backends, per the OTLP specification. Claude Code supports two OTLP transports, HTTP and gRPC, and a custom pipeline must confirm which one its collector or backend accepts and track that choice through every future infrastructure change. By default, metrics export every 60 seconds and traces and logs export every 5 seconds. The CLI batches signals internally and pushes them on that interval timer.

Defining your telemetry export paths

You have two routing options. Direct export sends signals straight from Claude Code to a backend. Routing through a collector can add buffering, batching, retry logic, redaction, and the ability to fan out to multiple destinations, which is why most production setups route OTLP through a collector that receives, enriches, and re-exports the data.

Direct export carries a specific risk for agents that run once and exit: a short task can finish before the next export interval fires, and if the process is killed before the CLI shuts down, anything still in the batch buffer is lost.

Whichever path you choose, the rollout doesn't have to happen developer by developer. OpenTelemetry works at every Claude Code plan level through local file-based configuration. Certain plan levels add central settings management, and mobile device management (MDM) tools can push the same configuration on any plan.

Configuring signal batching behavior

Batching typically trades latency for efficiency. Longer buffers generally mean fewer, larger exports and lower overhead, but they can delay how fresh the data is when it reaches a downstream platform. For cost and usage reporting that feeds a monthly review, the default intervals work well. For anything approaching real-time monitoring of agent reliability, tighter intervals raise both freshness and cardinality cost.

Structuring agent sessions for OTel visibility

Session structure is the organizing principle of Claude Code telemetry, and how boundaries are defined shapes everything from usage counts to downstream correlation.

Defining Claude Code session boundaries

An agent session is a continuous agent interaction and the unit of real Claude Code usage. CLI invocations measure how often a command ran, which says little about the work an agent performed. Claude Code writes full conversation transcripts to local JSON Lines (JSONL) files keyed by session ID, which reflects how the tool itself treats the session as the organizing unit. Anthropic's own docs flag that the transcript entry format is internal to Claude Code and changes between versions, so treat those joins as version-specific rather than a stable contract, factoring that maintenance exposure into any build-vs-buy comparison.

This distinction matters for evaluation. A platform that reports usage in invocations rather than sessions is counting command runs, which says nothing about the work an agent actually performed.

Linking telemetry data through session IDs

The correlation keys differ by signal type, per Anthropic's docs, with session.id as the documented attribute for spans and traces. prompt.id is tied to log events: every log event produced while processing a single user prompt shares a prompt.id UUID, so filtering by it reconstructs the API requests and tool executions that prompt triggered within the logs signal. Filtering by session.id links the corresponding spans. This linking is what turns disaggregated metrics, traces, and logs into a session timeline a person can actually read.

Capturing agent conversation state

Telemetry captures agent turns, tool calls, and context, but the depth is your choice. Prompt text, tool input details, and tool content are sensitive fields, and you control whether they are included or redacted through environment variables in your own OpenTelemetry setup. You configure whether that detail reaches a downstream platform, not the platform. Redacted streams can still support delivery correlation, while full detail produces more granular reporting.

Mapping token costs to OTel events

Token data in the OTel stream carries both consumption and cost information, and the two require different treatment before either can support a spend claim.

Mapping input and output token flows

Every API call fires an event carrying the model, estimated cost, duration, and input, output, and cache token counts. Two data points sit in that payload and should never be conflated: token consumption (a count) and token spend (a dollar cost). When a spend figure carries weight in a claim, add the qualifier that reported spend may be more or less than the amount actually paid for users on flat-rate plans such as Claude Pro.

Standardizing the telemetry data schema

Claude Code events carry attribute keys for token counts, model identifiers, cost, and ISO 8601 timestamps. Downstream pipelines typically map these into typed numeric fields for cost, duration, and the token types when indexing Claude Code events. Normalizing this schema into a canonical model is what lets Claude Code and Cursor usage appear in the same report, and data normalization and entity resolution are the two pieces minware handles before a metric can be trusted.

Mapping AI costs to delivery outcomes

Token spend is a cost metric. The ROI answer comes from a linear regression of a value delivery metric like story points completed against token spend across teams and time periods, reading the slope as the marginal delivery per dollar spent, per minware's Claude Code usage data guide. Each regression fits a single metric pair, and the R-squared value reports how much of the outcome's variation spend actually explains. This is where workflow and quality metrics earn their place as supporting signals: our piece on DORA metrics for AI investment covers how they explain a trend without standing in for the delivery evidence itself.

OTel signal What it measures Business outcome it supports Limitation
Token usage metrics Consumption and estimated cost Token spend analysis Cost estimates may differ from actual billing on flat-rate plans
Error logs Failures and retries Agent reliability patterns Volume varies based on logging configuration
Tool call spans Agent workflow sequence Session-level diagnostics May require additional context from version control to connect to delivery outcomes
Session traces Per-session agent behavior Usage patterns for adoption analysis Represents activity rather than direct delivery outcomes

Analyzing specific Claude Code trace events

Trace events expose the internal structure of agent work, and several span types are worth understanding individually before interpreting the timeline they produce.

Visualizing Claude Code workflow events

With tracing active, Claude Code traces expose each agent turn, model request, tool call, and hook execution as a named span type, per Anthropic's span reference: claude_code.interaction, claude_code.llm_request, claude_code.tool, and claude_code.hook. Any custom parser that reads those spans by name inherits the responsibility of tracking their names through every Claude Code version update. That sequence is what a timeline visualization renders: the order of tool calls within a session, where the agent spent time, and where it stalled. Raw spans in a backend give you the data, but the parsed timeline is what a front-line manager can act on, the same principle behind our work on tracking technical debt from AI agents.

Bash and PowerShell subprocesses inherit a TRACEPARENT environment variable carrying the W3C trace context of the active tool execution span, so scripts that read it can parent their own spans under the same trace. This applies only to Agent SDK and claude -p runs. Interactive CLI sessions ignore inbound TRACEPARENT entirely, per Anthropic's own observability guide, which matters because most developers run Claude Code interactively. Confirm which run mode your setup uses before building a distributed tracing design around this propagation, and review it with your security team because it determines how far trace context travels into child processes, a maintenance cost that should appear in a build-vs-buy comparison.

Analyzing retry behavior in spans

Failed API requests surface as error events with retry metadata in the logs and traces signals, each retry attempt recorded with its own attempt count and error class. In a timeline view, retries typically show up as repeated spans with timing gaps. Clusters of retry patterns on a specific model or endpoint may indicate a reliability issue. These events are debugging gold, but they are quality signals, not delivery evidence.

Configuring endpoints for incoming telemetry

Endpoint configuration determines how Claude Code signals are authenticated, routed, and distinguished from other telemetry in a shared collector.

Defining OTLP signals and payloads

An OTLP endpoint receives the three signal types in the protocol's hierarchical format and authenticates clients before accepting data. Administrators configure the export centrally through managed settings, a JSON file with an env block covering telemetry enablement, exporters, and the endpoint URL, deployed by IT in a way users cannot override.

Standardizing attribute key naming

Claude Code's attribute keys follow OpenTelemetry semantic conventions in snake_case. Standardization matters because downstream parsing depends on predictable keys, and normalizing attribute keys across vendors is what keeps metrics consistent when you run more than one AI coding tool.

Defining telemetry source boundaries

Separating Claude Code signals from other telemetry in a shared collector requires consistent resource service.name attribute tagging: team, cost center, and environment, configured centrally and held consistent across every Claude Code deployment. That tagging schema is a maintenance decision the admin team owns, and any drift from it silently corrupts multi-tool spend attribution. Locking the endpoint through managed settings gives security review a single, inspectable destination for everything Claude Code emits.

Data source Granularity Historical backfill Availability window
OTel export Per-session, per-request None, forward only Your retention policy
Claude Code Analytics API Per-user, per-model, per-day No specified deletion period Available with up to a 1-hour delay
Claude Enterprise Analytics API Per-user, per-day Not applicable, no backfill before launch Available for dates on or after January 1, 2026

One gap the table does not surface: the Analytics APIs cover Console and Claude Enterprise organizations only. Teams running Claude through Amazon Bedrock, Google Vertex AI, or Microsoft Foundry have no Analytics API route at all, so OTel is the only pipeline available for those deployments. The Claude Code Analytics API and the Claude Enterprise Analytics API are different products. OTel has no historical backfill of its own, so it only captures activity from the point it's configured. That is the reason to connect OTel as soon as you can. Our guide to Claude Code enterprise usage walks through how the two pipelines complement each other in practice.

Evaluating downstream platforms with this signal model

Claude Code's OTel export gives you metrics for volume, logs for failures, and traces for sequence, all bound together by session IDs. The evaluation question for any downstream platform is whether it can parse these signals, recover the associations between sessions and commits, and regress token spend against delivery outcomes with transparent calculation logic.

If you want the measurement grounding before the platform conversation, these guides go deeper on the metrics side:

Closing the gap between telemetry and delivery evidence

Knowing what Claude Code's OTel signals contain, how they move through a collector, and how sessions and token spend map to cost is table stakes for evaluating any downstream platform's claims. None of that, by itself, proves AI spend is buying more delivery. The signal model tells you what happened inside a session. Turning that into a defensible ROI answer still requires linking sessions to commits, tickets, and pull requests, then regressing the result against token spend, the step the rest of this evaluation should focus on.

Handling this with minware

Raw OTel streams are agent activity. minware is an engineering analytics platform that ingests Claude Code OTel streams, reconstructs agent sessions into structured timelines, and connects them to version control and project management data through its data normalization and entity resolution model, recovering associations even when a session produced no structured link to a commit.

The defensible methodology regresses each outcome metric against token spend separately, grouping by team and time period. Run it for roadmap value delivery, on-time delivery rate, and overhead cost. Each slope is the marginal ROI figure for its metric. Every metric's calculation is visible and editable in the product through minQL (minware's formula language), so the number you bring to a board review comes with its formula attached.

If you are weighing adjacent approaches, observability stacks display OTel signals well, but they leave joining agent sessions to commits, pull requests, and tickets to you, which turns the delivery side into a build project of its own. AI gateways enforce spend caps at the point of the call, and many teams run a gateway for spend control alongside minware for the outcome correlation. Building the pipeline yourself works until vendor schemas change or a stakeholder asks how a number is calculated, and then someone owns that maintenance and that explanation indefinitely. Our take on broken productivity metrics covers why the measurement layer, not the telemetry, is where these efforts usually break.

Start a 14-day free trial at minware.com, no credit card required, and connect your first data source to see how minware parses Claude Code OTel streams into structured timelines and regresses token spend against delivery outcomes.

FAQs

Does Claude Code send PII or code content in telemetry?

Prompt text, tool input details, and tool content are sensitive fields, and you control whether they are included or redacted through environment variables in your own OpenTelemetry setup. What reaches a downstream platform is your configuration choice, not the platform's.

Can I filter which signals are exported?

Yes. Redaction and filtering happen in your OpenTelemetry configuration, and redacting prompt text still supports delivery correlation while including it produces more granular reporting.

How much data volume should I expect per developer?

Volume scales with session frequency and signal granularity, since OTel reports at per-session, per-request detail rather than the Analytics APIs' per-day rollups. High-cardinality attributes like session IDs are the main storage cost driver.

What happens if the collector endpoint is unavailable?

Per Anthropic's own docs, the CLI fails silently on export errors by default and drops telemetry without surfacing an error. A killed process loses whatever sits in the buffer for the same reason. A collector with persistent buffering reduces the drop risk compared to direct export, but it does not change the CLI's silent-failure default, so monitor your collector's ingestion rate to detect gaps rather than waiting for an error to surface.

Can I correlate Claude Code telemetry with version control data?

Yes, but it requires a downstream platform that links agent sessions to commits and pull requests. minware recovers these associations with time-based linking even when a session produced no structured connection to a commit.

Key terms glossary

OpenTelemetry (OTel): An open-source observability framework for exporting telemetry, such as prompts and tool-use events, from instrumented applications, including AI coding tools, to a collector. Reports at per-prompt, per-session granularity, with no historical backfill for activity before it was configured.

OTLP: The OpenTelemetry Protocol, the wire format for sending OTel signals to a collector or backend.

Span: A single unit of work in a trace, representing one operation like a tool call or agent turn.

Trace: A collection of spans representing a workflow, such as a single agent session.

Metric: An aggregate measurement, like token count or session count, reported at intervals.

Log: A discrete event record, such as an error or state change.

Cardinality: The number of unique values an attribute takes, which drives storage cost and query performance.

Agent session: The unit of AI usage representing a multi-step task executed by the coding agent, distinct from a single prompt.

Prompt ID: A UUID assigned to each user prompt that links all API requests and tool executions triggered by that prompt, enabling reconstruction of per-prompt activity.

Token spend: The dollar cost of AI tool usage, as reported by AI tool APIs or OpenTelemetry. On flat-rate, tiered, or discounted plans, reported spend may differ from the amount actually invoiced.

Token consumption: The count of tokens processed by an AI model. A separate data point from token spend, the dollar cost of that usage.

Story points completed: The total of the story points field for completed tickets.

Linear regression: A statistical model charting the best-fit line between an outcome metric and an input variable. It is the primary AI ROI methodology, fitting a line through token spend and a delivery outcome metric across teams and time periods. Each regression fits one metric pair, correlating spend against multiple outcomes requires a separate regression for each.

Regression slope: The change in an outcome metric per unit change in the input variable, such as story points completed per dollar of token spend.

R-squared: In a regression chart, the share of the outcome metric's variation explained by the input variable, on a 0-to-1 scale.

DORA: DevOps Research and Assessment, a research program that identified key metrics for software delivery performance.

minQL: minware's patent-pending formula language, used to define every metric, dimension, and pipeline calculation, with full visibility into the underlying logic.