Try LangSmith
Overview
LangSmith is the observability platform built by the team behind LangChain, and its advantage is exactly what you would expect: nothing else understands a LangGraph execution as well. Agent graphs are visualised as graphs, with the branches taken, the state at each node and the loops that ran — a level of structural detail that framework-agnostic tools cannot reconstruct from generic spans.
The second distinctive feature is annotation queues. Traces are routed to human reviewers who label them, and those labels become datasets for evaluation. It sounds procedural, but it is the mechanism that turns a domain expert's judgement into a repeatable test — the step most teams skip and then wonder why their evaluation set does not reflect what users actually care about.
The trade-off is that LangSmith is primarily a managed service and works best inside the LangChain ecosystem. It does support other stacks, but the further you move from LangChain and LangGraph the smaller the advantage becomes, and at that point the self-hosting question favours alternatives.
Key Features
LangGraph Graph Visualisation
Agent executions rendered as the graphs they are, with branches, state and loops visible — not reconstructed from generic spans.
Annotation Queues
Route traces to human reviewers for labelling, converting expert judgement into evaluation datasets systematically.
High-Detail Tracing
Deep capture of inputs, outputs, token usage and latency at every step of a chain or agent.
Dataset and Experiment Management
Build test sets from real traces and compare prompt or model changes against them.
Prompt Hub
Versioned prompt storage and sharing integrated with the traces that used each version.
Native LangChain Integration
Instrumentation is close to automatic if you already use LangChain or LangGraph.
Pros & Cons
Advantages
- Unmatched visibility into LangGraph agent executions
- Annotation queues are the best human-in-the-loop evaluation workflow available
- Near-zero instrumentation effort inside the LangChain stack
- Mature dataset and experiment tooling
- Backed by the team that maintains the framework
Disadvantages
- Advantage shrinks considerably outside LangChain and LangGraph
- Primarily managed — self-hosting is an enterprise arrangement
- Traces leave your infrastructure on standard plans
- Pricing scales with trace volume and can climb quickly
Pricing Plans
| Plan | Price | Key Features |
|---|---|---|
| Developer | Free | Single user with a monthly trace allowance |
| Plus | From $39 / user / month | Team features, higher trace volumes, longer retention |
| Enterprise | Custom | Self-hosted deployment, SSO, compliance, support |
Best Use Cases
LangSmith Excels At:
- Teams building on LangChain and LangGraph
- Debugging complex agent graphs with branching and loops
- Human-in-the-loop evaluation with domain experts
- Getting observability running with minimal instrumentation work
May Not Be Ideal For:
- Stacks that do not use LangChain
- Organisations that must self-host without an enterprise contract
- High-volume tracing on a tight budget
How It Compares
LangSmith vs Langfuse
If you live in LangChain and LangGraph, LangSmith's graph visualisation and near-zero instrumentation are worth real money. If you need self-hosting or run a mixed stack, Langfuse's MIT licence settles it.
LangSmith vs Braintrust
LangSmith is oriented around understanding what your agents did; Braintrust around systematically improving quality through evaluation. The two overlap but answer different primary questions.
Final Verdict
Our Recommendation
LangSmith is the obvious choice for teams committed to LangChain and LangGraph, and a harder sell for everyone else. The graph visualisation is genuinely better than what framework-agnostic tools can produce, because it works from structure rather than inferring it, and annotation queues are the most practical answer anyone ships to getting expert judgement into an evaluation set. Outside that ecosystem the advantage narrows and the managed-only posture starts to matter. Judge it on how deep you are in LangChain — that is the whole decision.