Langfuse vs MLflow
Two open-source platform options for LLM observability & evals. When each fits, what it costs, who moves from one to the other, and what makers who chose it say.
Which fits you
- You want tracing, prompts and evals in one tool you can self-host for free
Use it whenYou want to own your trace data or keep costs flat as volume grows.
Trade-offSelf-hosting means running Postgres, ClickHouse, Redis and object storage alongside it; otherwise it's a usage-based cloud plan.
- Tracing, evaluating and versioning prompts for LLM apps in the same open-source platform many teams already use for ML experiments.
Use it whenYou already use MLflow or Databricks for models, or want OpenTelemetry-based tracing and evals you can self-host.
Trade-offGrew out of classic ML tooling, so the UI and concepts are broader than an LLM-only tool; you run the tracking server yourself unless you use a managed version.
At a glance
| Used by | 34 makers' products · 43 open-source projects | No maker's product yet · 15 open-source projects |
|---|---|---|
| Cost at default usagetraces 100k traces | $29/mo Core | — |
| Downloads | 1.8M/wk+81% vs npm | 7.9k/wk |
| Pricing | Free tier (50k units/mo); usage-based paid plans; fully free to self-host under MIT license. · paid from $29/mo | Free (open source); managed versions are offered by Databricks and cloud providers. |
| Free tier | Yes | Yes |
| Open source | Yes · self-hostable | Yes · self-hostable |
| Incidents, 90 daysfrom its status page | 15 (1 major) | no public status feed |
Cost as you grow
Both cost $0 up to 50k traces; from 100k traces MLflow costs less ($0 vs $29); and still does at 10M traces ($0 vs $821).
The numbers, plan by plan
| Traces per month | Langfuse | MLflow |
|---|---|---|
| 1,000 | $0 Hobby | $0 Self-hosted (open source) |
| 10,000 | $0 Hobby | $0 Self-hosted (open source) |
| 50,000 | $0 Hobby | $0 Self-hosted (open source) |
| 100,000 | $29 Core | $0 Self-hosted (open source) |
| 500,000 | $61 Core | $0 Self-hosted (open source) |
| 1,000,000 | $101 Core | $0 Self-hosted (open source) |
| 5,000,000 | $421 Core | $0 Self-hosted (open source) |
| 10,000,000 | $821 Core | $0 Self-hosted (open source) |
What makers say
Makers on using it for LLM observability & evals, from Product Hunt and Starter Story interviews, each linked to the source. Products with a page of their own and fuller notes first.
Marc gave me an in-person onboarding in SF - I found an issue in our LLM provider config just 30 minutes after the onboarding thanks to Langfuse. 10/10 recommendation
Langfuse powers our LLM observability. Without Langfuse, our AI agent would not be best-in-class. We have been using Langfuse since nearly the beginning: 2+ years!
We use Langfuse to keep track of our LLM prompts while building MCP-Builder.ai. It’s a great tool that makes it easy to monitor and analyze prompt performance, helping us improve quickly and efficiently.
No maker quote about MLflow for LLM observability & evals yet.
Loved and watch-outs
Themes that recur in makers' words and Hacker News comments, each linked to what it summarises.
- Detailed traces show each agent call and the context it pulled, which makes debugging agent behaviour much easier. PHPH 2HN
- Token usage and cost are tracked across providers alongside quality, in one place. PHPH 2HN
- Open source and self-hostable, a common default for teams wanting tracing on their own infrastructure. PHHNHN 2HN 3
- Prompt management and experiments feel basic next to the tracing core. HNHN 2HN 3
- Views are retrospective dashboards, so explaining one run's cost or failure still means reading trace trees by hand. HNHN 2HN 3HN 4
- The ClickHouse acquisition raised GDPR and data-residency concerns for EU users of the cloud version. HNHN 2

