Agent Observability and Tracing — Spans, Logs and Debugging AI Agents
An agent that spends ten steps planning, calling tools and evaluating intermediate results either produces a usable result in the end — or it doesn’t. Without visibility into the individual steps, it stays invisible where things went wrong: the wrong tool call, a prompt that was too thin, an infinite retry loop, or simply an expensive model used at the wrong point. Agent observability is the practice of making that flow visible — through tracing, which captures every single step as a structured record instead of only knowing the start and end of a run.
Traces and spans: the basic unit of visibility
A trace is the complete record of a single agent run, from the first user request to the final answer. It’s made up of spans — individual time segments that each represent one step: an LLM call, a tool call, a retrieval query, an invocation of a sub-agent. Every span carries a start and end timestamp, input and output, metadata such as the model used and token counts — and a reference to its parent span. That produces a nested tree that reflects the agent’s actual execution structure, not just its text output.
Since 2024, OpenTelemetry’s GenAI working group has been developing standardized GenAI semantic conventions — unified attribute names for exactly these spans: gen_ai.request.model for the model called, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens for token counts per call, gen_ai.response.finish_reasons for the completion reason (for example stop or tool_calls). As of 2026, most of these conventions are still marked experimental, but they’re already supported by Google Cloud, AWS, Azure, Datadog and the common agent frameworks — a trace can increasingly be read across tools rather than being tied to a single one.
Why debugging is barely possible without tracing
An agent’s failure almost never shows up at the point where it originates. An LLM picks the wrong tool at step three, the result still gets passed on to the next step, and only two steps later does the answer visibly go off the rails. Without spans, all you see is: final result wrong. With tracing, you can inspect the exact prompt sent at each step, the exact tool arguments, the exact tool result and the latency of every single span after the fact — and often replay it directly against a different model or a changed prompt.
That matters twice as much in multi-agent systems: once a task is handed off to a sub-agent, execution branches, sometimes in parallel. Without consistent trace context propagation (distributed tracing) across all sub-agents, the connection falls apart into a set of isolated logs from which it’s hard to reconstruct which sub-agent was responsible for which part of the final answer. Anyone building an agent who plans for tracing from the start saves themselves exactly this reconstruction work later, when something breaks.
Cost control: making what a run actually costs visible
An agent run is rarely a single LLM call. Planning, tool selection, intermediate evaluation of results, and possibly several sub-agent calls quickly add up to a multiple of the cost of a single request — and without a per-span breakdown, it stays unclear which step accounts for the largest share. A single agent request with several tool calls can already generate ten to thirty individual trace units with some observability vendors; that granularity is exactly what’s needed to tell whether a retry loop, an overly long tool output sitting in context, or an unnecessarily expensive model is driving the bill.
Tracing makes these costs visible per span and can be tied to a token budget: thresholds per trace, alerts on outliers, and the ability to optimize specifically the most expensive step of a workflow instead of throttling the whole system across the board. Without that breakdown, cost control for agents stays a pure estimate based on the monthly bill — after the fact, imprecise, and without a lever for the next optimization.
Logging vs. tracing — a complement, not a replacement
Classic logging writes timestamped text lines or events, usually flat and without any built-in relationship between them. That’s enough for individual error messages, but it fails at the actual structure of an agent: nested steps, parallel branches, retries. Tracing captures exactly that structure — as a tree of spans with parent-child relationships and timing. In practice, the two complement each other: the trace provides the shape of the run, individual log lines or events attached to a span provide the diagnostic detail inside a step (for example, a warning that a tool result came back empty). An agent observability setup that only collects loose logs without trace structure stays effectively blind to the flow itself once runs get more complex.
Tools compared: LangSmith, Langfuse, OpenTelemetry
Two platforms dominate practical use in 2026:
- LangSmith is LangChain’s commercial observability platform, historically tightly coupled to LangChain and LangGraph: node-by-node state diffs, full execution graphs, a breakdown by model and tool calls, and replay against a different model version. Via the
langsmith[otel]package, LangSmith can now also be connected outside the LangChain ecosystem through OpenTelemetry. Since May 2026, US cloud ingestion runs through SmithDB, a purpose-built Rust data layer. - Langfuse is open source (MIT core), self-hostable, and builds hierarchical traces from LLM calls, tool invocations, embeddings and retrieval steps — independent of the framework used. In January 2026, Langfuse was acquired by ClickHouse. Billing is per unit (trace, observation or score), which is worth keeping in mind for complex agent workflows with many tool calls.
- OpenTelemetry itself isn’t a product but the increasingly shared foundation: instrumenting via OTel conventions lets you send traces to several backends (Langfuse, LangSmith, Datadog, self-hosted solutions) simultaneously or interchangeably, instead of committing early to one proprietary SDK.
The right choice depends on context: teams heavily invested in LangChain/LangGraph who want a mature commercial suite are well served by LangSmith. Teams that prioritize data ownership, self-hosting or framework independence tend to reach for Langfuse or an OTel-based setup of their own.
Pitfalls in practice
- Data volume and the cost of observability itself. Long agent runs with dozens of tool calls generate a correspondingly large number of spans — with usage-based billing, the observability bill grows with the agent’s complexity, not just with the number of users.
- Sensitive data in traces. Prompts and tool results frequently contain personal or confidential content. Without redaction or masking before storage, that data ends up permanently in the trace store — relevant under data protection law as soon as real user data is involved.
- Orphaned spans under parallel execution. If sub-agents are started in parallel without cleanly propagating the trace context, spans show up isolated instead of embedded in the overall trace — the connection to the triggering step gets lost.
- Tool lock-in to one SDK. Instrumenting exclusively with one vendor’s proprietary decorators trades observability for lock-in. OTel-based instrumentation keeps that option open.
- Adding tracing only once things break. If observability only gets added after the first unexplained production failure, exactly the traces that would have explained the failure are missing. Tracing belongs at the start of an agent project, not at the end.
FAQ
What’s the difference between a trace and a span? A trace is the complete record of an agent run from start to finish. A span is a single step within it — an LLM call, a tool call, a retrieval query — with its own timing, input/output, and a reference to its parent span. Many spans together make up a trace.
Do I need tracing even for a single, simple agent? For an agent with just one or two steps, an error can often still be spotted by eye in the raw prompt. Once multiple tool calls, conditional branches or retries come into play, debugging without tracing quickly turns into guesswork — the tipping point arrives earlier in practice than expected.
Is Langfuse or LangSmith the better choice? Neither is categorically better. LangSmith excels at deep LangChain/LangGraph integration and as a mature commercial suite, Langfuse at being open source, self-hostable and framework-independent. Both now support OpenTelemetry-based instrumentation.
Does tracing replace classic logging? No. Tracing captures the structure of a run (which step belongs to which parent step, how long it took). Logs and events provide the diagnostic detail inside a single step. Together, the two provide complete observability.
Does tracing itself cause noticeable overhead? The raw recording overhead per span is usually low. It becomes noticeable mainly at very high trace volumes, or when entire prompt/response contents get logged uncompressed alongside the trace — that’s where sampling or selective payload logging pays off.