Open-source platformBest forTracing, prompt management and evals in one tool you can run on your own server for free.
Trade-offSelf-hosting means running Postgres, ClickHouse, Redis and object storage alongside it; otherwise it's a usage-based cloud plan.
Used by 33 makers' products for this · Braintrust vs Langfuse →
Proxy loggingBest forGetting request logs, costs and latency by changing one base URL, with no SDK instrumentation.
Trade-offProxy logging sees individual requests well but agent steps and eval workflows less deeply; acquired by Mintlify in March 2026, so check its roadmap before building on it.
Used by 13 makers' products for this · Braintrust vs Helicone →
Hosted eval platformBest forApps built on LangChain or LangGraph, where tracing works with almost no setup.
Trade-offPaid per seat beyond the free tier; self-hosting is enterprise-only.
Used by 11 makers' products for this · Braintrust vs LangSmith →
Open-source platformSelf-hostableFree tierFrom $50/mollms.txt Best forOpenTelemetry-based tracing and evals you can start locally in a notebook, then self-host or move to its cloud.
Trade-offSource-available under the Elastic License rather than a permissive open-source license.
No maker product tracked for this yet
Open-source platformBest forTracing, evaluations and prompt optimization in an Apache-licensed platform you self-host with Docker or Kubernetes, or use on Comet's cloud.
Trade-offSelf-hosting means running several services; smaller community than Langfuse.
No maker product tracked for this yet
Eval and test CLIBest forRunning prompts and models against test cases from the command line or CI, including red-team tests for prompt injection and data leaks.
Trade-offA testing tool rather than production tracing, so you pair it with something that logs live traffic.
No maker product tracked for this yet
Eval and test CLIOpen sourceFree tierFrom $200/mo (Confident AI Starter)llms.txtCLI Best forWriting LLM evals as pytest-style tests in Python, with ready-made metrics for hallucination, faithfulness and answer relevancy.
Trade-offThe metrics use an LLM as a judge, so every run costs tokens and scores vary slightly; dashboards and history are in the paid Confident AI platform.
No maker product tracked for this yet
Eval and test CLIBest forScoring RAG pipelines — how faithful answers are to the retrieved context and how relevant that context is — and generating test questions from your documents.
Trade-offA Python library, not a platform — no tracing of live traffic or UI, and LLM-judged metrics cost tokens per run.
No maker product tracked for this yet
Open-source platformBest forTracing, evaluating and versioning prompts for LLM apps in the same open-source platform many teams already use for ML experiments.
Trade-offGrew out of classic ML tooling, so the UI and concepts are broader than an LLM-only tool; you run the tracking server yourself unless you use a managed version.
No maker product tracked for this yet