Try Langfuse
Overview
Langfuse solves the problem every team hits about three weeks after their LLM feature reaches production: something is wrong, and you have no idea which step of a multi-agent, multi-tool call chain caused it. Conventional application monitoring shows you a slow request. Langfuse shows you the nested spans — every model call, retrieval, tool invocation and their inputs and outputs — so the failure becomes an ordinary debugging problem instead of a guessing game.
It goes further than tracing. Datasets, evaluations and prompt management live in the same product, which means the loop from 'this trace looks wrong' to 'add it to a test set' to 'verify the fix did not break something else' happens without exporting data between tools. That integration is why teams consolidate on it rather than assembling three separate services.
The decisive feature for many organisations is the licence. Langfuse is MIT-licensed and fully self-hostable, so LLM traces — which routinely contain customer data, internal documents and prompts you would rather not hand to a third party — can stay on your infrastructure. Managed cloud exists for teams that do not need that; the self-hosted version is the complete product, not a teaser.
Key Features
Nested Trace Capture
Full spans across agents, retrievers and tool calls with inputs and outputs, turning opaque LLM failures into readable execution traces.
Evaluations on Production Traffic
Attach automated or human evaluation scores to real traces, not only to a curated offline test set.
Datasets from Real Failures
Promote a failing production trace into a regression dataset in one step, which is how LLM test suites actually get built.
Prompt Management
Version prompts, deploy them without a code release, and see which version produced which trace.
MIT Licence, Fully Self-Hostable
The complete product runs on your infrastructure, which matters because traces contain your most sensitive data.
Framework-Agnostic SDKs
Works with LangChain, LlamaIndex, raw SDK calls and custom stacks, rather than assuming one framework.
Pros & Cons
Advantages
- MIT-licensed and genuinely self-hostable — the whole product, not a limited edition
- Tracing, evaluation, datasets and prompts in one place
- Framework-agnostic rather than tied to one orchestration library
- Keeps sensitive trace data on your own infrastructure
- Active development and a strong community
Disadvantages
- Self-hosting means running a database and a service you must maintain
- High-volume trace storage grows quickly and needs a retention policy
- Evaluation tooling is less opinionated than dedicated eval platforms
- Instrumentation still has to be added to your code
Pricing Plans
| Plan | Price | Key Features |
|---|---|---|
| Self-Hosted | Free | MIT licence, complete feature set, your infrastructure |
| Cloud Hobby | Free | Limited monthly events on managed cloud |
| Cloud Pro | From $59 / month | Higher volumes, longer retention, team features |
| Enterprise | Custom | SSO, compliance, support, dedicated deployment |
Best Use Cases
Langfuse Excels At:
- Debugging multi-step agent and RAG failures in production
- Teams that cannot send prompts and traces to a third party
- Building regression test suites from real production failures
- Consolidating tracing, evals and prompt versioning in one tool
May Not Be Ideal For:
- Teams with no capacity to run infrastructure who also cannot use cloud
- Very high trace volumes without a storage strategy
- Organisations wanting the deepest LangGraph-specific visualisations
How It Compares
Langfuse vs LangSmith
LangSmith has the deepest integration with LangChain and LangGraph and is the natural pick if you live inside that stack. Langfuse is framework-agnostic and self-hostable under MIT. Data residency usually decides it.
Langfuse vs Braintrust
Braintrust is stronger on prompt-centric evaluation workflows and experiment management; Langfuse is stronger on open, self-hosted tracing. Teams focused on shipping quality improvements lean Braintrust, teams focused on understanding production lean Langfuse.
Final Verdict
Our Recommendation
Langfuse is the default LLM observability platform for good reasons. The MIT licence and complete self-hosted product remove the objection that kills most observability purchases in regulated environments — traces contain your most sensitive data, and here they never leave. Having tracing, evaluation, datasets and prompt management in one tool means the loop from production failure to regression test is short enough that teams actually close it. Plan trace retention before you scale, and accept that you still have to instrument your code.