AI in the product

LLM Observability & Evals

See what your prompts, models and agents actually did in production, and test changes against real examples before you ship them.

Ask your AI about this, with this page as the source:ChatGPT ↗Claude ↗Perplexity ↗

The real choice

Open-source platform you can self-host, or a hosted eval product. Open-source tools keep traces on your own servers and cost nothing to run yourself, but you maintain them. Hosted eval platforms give you polished experiment and scoring workflows for a monthly bill, with your data in their cloud.

Pick by situation

Tap the ones that are you — the tools that fit light up below.
  • You want tracing, prompts and evals in one tool you can self-host for freeLangfuse
  • You already use OpenTelemetry or want vendor-neutral instrumentationArize Phoenix
  • You want request logs and costs today by changing one base URLHelicone
  • Your app is built on LangChain or LangGraphLangSmith
  • Your main question is whether a prompt or model change made outputs better or worseBraintrust

The contenders

Grouped by the side of the choice they answer, not ranked. Open a row for when to use it, the trade-off and what makers say.

ToolBest for
Open-source platform2
Langfuseplatform · open source · free tierTracing, prompt management and evals in one tool you can run on your own server for free.33+43 open source$29Core1.8M/wk+81% vs npm

Use it whenYou want to own your trace data or keep costs flat as volume grows.

Trade-offSelf-hosting means running Postgres, ClickHouse, Redis and object storage alongside it; otherwise it's a usage-based cloud plan.

llms.txtCLI

Used by Agentplace.io, AutonomyAI, cognee, Conduit, Finyuus and 28 more · in 43 open-source projects.

Marc gave me an in-person onboarding in SF - I found an issue in our LLM provider config just 30 minutes after the onboarding thanks to Langfuse. 10/10 recommendation
stagewise, the makerSep 2026 ↗
Arize Phoenixservice · self-hostable · free tierOpenTelemetry-based tracing and evals you can start locally in a notebook, then self-host or move to its cloud.19 open source—80.4k/wk+180% vs npm

Use it whenYou already use OpenTelemetry or want vendor-neutral instrumentation.

Trade-offSource-available under the Elastic License rather than a permissive open-source license.

llms.txt

No maker's product we track shows it for this yet · in 19 open-source projects.

Proxy logging1
Heliconeplatform · open source · free tierGetting request logs, costs and latency by changing one base URL, with no SDK instrumentation.13$116Pro3.8k/wk

Use it whenYou want visibility today and your app makes direct model calls.

Trade-offProxy logging sees individual requests well but agent steps and eval workflows less deeply; acquired by Mintlify in March 2026, so check its roadmap before building on it.

llms.txt

Used by Alai, Codebuff, Fume, Image Ally, Persana and 8 more.

Helicone AI offers open-source observability tools tailored for developers working with LLMs. It simplifies debugging and optimization, providing valuable insights into AI model performance.
Persana, the makerSep 2026 ↗
Hosted eval platform2
LangSmithplatform · free tierApps built on LangChain or LangGraph, where tracing works with almost no setup.11+49 open source$475Developer6.1M/wk−4% vs npm

Use it whenYou already use the LangChain stack.

Trade-offPaid per seat beyond the free tier; self-hosting is enterprise-only.

llms.txt

Used by Astrid, DryMerge, GitLaw, Intryc, lmChatGPTtfy and 6 more · in 49 open-source projects.

LangSmith’s real-time analytics and versioning keep our AI agents rock-solid -- so everything just works better.
Watchman AI, the makerSep 2026 ↗
Braintrustplatform · free tierEval-driven work — scoring outputs and comparing prompts and models side by side in experiments.5+10 open source—1.4M/wk+117% vs npm

Use it whenYour main question is "did this change make outputs better or worse".

Trade-offClosed source, and the paid tier is priced for teams rather than hobby projects.

llms.txtMCPCLI

Used by Brew, Hey Noah, Arcade, Blacksmith, Linear · in 10 open-source projects.

Cost: the cheapest plan that fits traces 100k traces, from list prices. Try your own numbers →

Cost as you grow

Each contender's cheapest usable plan as usage rises.

$0$1,000$5,000$20,0001k10k50k100k500k1M5M10M
LangSmithHeliconeArize PhoenixLangfusex: traces per month (trace) · cheapest usable plan at each point, list prices · try your own numbers

Who switches to what

Public pull requests on GitHub since Oct 2024 whose title says "X to Y" — real code changes, by developers in general rather than makers only. Pick a flow to see its pull requests.

LangSmith → Langfuse9 PRs · Compare →

Before you choose

How to approach it

Start by logging every request with its prompt, output, latency and cost — a day of real traces teaches more than any benchmark. Save the bad outputs you find as a small dataset, and rerun it every time you change a prompt or model. Add automated scoring only after you've read enough traces to know what "good" means for your app.

Common mistakes
  • Logging full prompts and outputs that contain users' personal data to a third-party service without masking it or listing the vendor as a subprocessor.
  • Changing a prompt in production with no saved examples to rerun, so a fix for one case silently breaks five others.
  • Trusting an LLM-as-judge score before checking that it agrees with your own judgment on a few dozen examples.

Other options

Real choices most makers here won't need to weigh.

Decided alongside

What the 58 makers' products here chose for their other decisions.