LangSmith vs MLflow
Two sides of the LLM observability & evals decision: hosted eval platform and open-source platform. When each fits, what it costs, who moves from one to the other, and what makers who chose it say.
Which fits you
- Your app is built on LangChain or LangGraph
Use it whenYou already use the LangChain stack.
Trade-offPaid per seat beyond the free tier; self-hosting is enterprise-only.
- Tracing, evaluating and versioning prompts for LLM apps in the same open-source platform many teams already use for ML experiments.
Use it whenYou already use MLflow or Databricks for models, or want OpenTelemetry-based tracing and evals you can self-host.
Trade-offGrew out of classic ML tooling, so the UI and concepts are broader than an LLM-only tool; you run the tracking server yourself unless you use a managed version.
At a glance
| Used by | 11 makers' products · 49 open-source projects | No maker's product yet · 15 open-source projects |
|---|---|---|
| Cost at default usagetraces 100k traces | $475/mo Developer | — |
| Downloads | 6.1M/wk−4% vs npm | 7.9k/wk |
| Pricing | Free tier with a monthly trace limit; paid per-seat plan with usage overage; custom Enterprise, including self-hosted. · paid from $39/seat/mo | Free (open source); managed versions are offered by Databricks and cloud providers. |
| Free tier | Yes | Yes |
| Open source | No | Yes · self-hostable |
| Incidents, 90 daysfrom its status page | 0 | no public status feed |
Cost as you grow
Both cost $0 up to 1k traces; from 10k traces MLflow costs less ($0 vs $25); and still does at 10M traces ($0 vs $49,975). They're different kinds of tool — hosted eval platform and open-source platform — so the prices don't buy the same thing.
The numbers, plan by plan
| Traces per month | LangSmith | MLflow |
|---|---|---|
| 1,000 | $0 Developer | $0 Self-hosted (open source) |
| 10,000 | $25 Developer | $0 Self-hosted (open source) |
| 50,000 | $225 Developer | $0 Self-hosted (open source) |
| 100,000 | $475 Developer | $0 Self-hosted (open source) |
| 500,000 | $2,475 Developer | $0 Self-hosted (open source) |
| 1,000,000 | $4,975 Developer | $0 Self-hosted (open source) |
| 5,000,000 | $24,975 Developer | $0 Self-hosted (open source) |
| 10,000,000 | $49,975 Developer | $0 Self-hosted (open source) |
What makers say
Makers on using it for LLM observability & evals, from Product Hunt and Starter Story interviews, each linked to the source. Products with a page of their own and fuller notes first.
LangSmith’s real-time analytics and versioning keep our AI agents rock-solid -- so everything just works better.
You can't build AI agents without monitoring. Metadata filtering is strong.
I deployed the manage Pig agent on LangGraph and it's been smooth sailing!
No maker quote about MLflow for LLM observability & evals yet.
Loved and watch-outs
Themes that recur in makers' words and Hacker News comments, each linked to what it summarises.
- It keeps pulling users toward the LangChain platform, and LangChain docs push LangSmith, which feels like lock-in. HNHN 2HN 3HN 4
- Traces show which agent failed but not why, so root-cause analysis stays manual. HNHN 2HN 3HN 4
- Viewing your own traces requires a cloud account, with no local-first option. HNHN 2

