Promptfoo alternatives

Not a ranking: each option is described by the situation it fits, with its trade-off, whether it's open source or free to start, whether AI coding agents can work with it, and how many makers' products use it for this job.

7 open source or self-hostable9 with a free tier9 ready for AI agents

For LLM observability & evals

The whole decision →

Open-source platform you can self-host, or a hosted eval product. Open-source tools keep traces on your own servers and cost nothing to run yourself, but you maintain them. Hosted eval platforms give you polished experiment and scoring workflows for a monthly bill, with your data in their cloud.

Promptfoo (Eval and test CLI): Running prompts and models against test cases from the command line or CI, including red-team tests for prompt injection and data leaks. Trade-off: A testing tool rather than production tracing, so you pair it with something that logs live traffic.

Langfuse

Open-source platform
Open sourceFree tierFrom $29/mollms.txtCLI

Best forTracing, prompt management and evals in one tool you can run on your own server for free.

Trade-offSelf-hosting means running Postgres, ClickHouse, Redis and object storage alongside it; otherwise it's a usage-based cloud plan.

Used by 33 makers' products for this

Helicone

Proxy logging
Open sourceFree tierFrom $79/mollms.txt

Best forGetting request logs, costs and latency by changing one base URL, with no SDK instrumentation.

Trade-offProxy logging sees individual requests well but agent steps and eval workflows less deeply; acquired by Mintlify in March 2026, so check its roadmap before building on it.

Used by 13 makers' products for this

LangSmith

Hosted eval platform
Free tierFrom $39/seat/mollms.txt

Best forApps built on LangChain or LangGraph, where tracing works with almost no setup.

Trade-offPaid per seat beyond the free tier; self-hosting is enterprise-only.

Used by 11 makers' products for this

Braintrust

Hosted eval platform
Free tierFrom $249/mollms.txtMCPCLI

Best forEval-driven work — scoring outputs and comparing prompts and models side by side in experiments.

Trade-offClosed source, and the paid tier is priced for teams rather than hobby projects.

Used by 5 makers' products for this

Arize Phoenix

Open-source platform
Self-hostableFree tierFrom $50/mollms.txt

Best forOpenTelemetry-based tracing and evals you can start locally in a notebook, then self-host or move to its cloud.

Trade-offSource-available under the Elastic License rather than a permissive open-source license.

No maker product tracked for this yet

Opik

Open-source platform
Open sourceFree tierFrom $19/mollms.txtMCP

Best forTracing, evaluations and prompt optimization in an Apache-licensed platform you self-host with Docker or Kubernetes, or use on Comet's cloud.

Trade-offSelf-hosting means running several services; smaller community than Langfuse.

No maker product tracked for this yet

DeepEval

Eval and test CLI
Open sourceFree tierFrom $200/mo (Confident AI Starter)llms.txtCLI

Best forWriting LLM evals as pytest-style tests in Python, with ready-made metrics for hallucination, faithfulness and answer relevancy.

Trade-offThe metrics use an LLM as a judge, so every run costs tokens and scores vary slightly; dashboards and history are in the paid Confident AI platform.

No maker product tracked for this yet

Ragas

Eval and test CLI
Open sourceFree tierllms.txt

Best forScoring RAG pipelines — how faithful answers are to the retrieved context and how relevant that context is — and generating test questions from your documents.

Trade-offA Python library, not a platform — no tracing of live traffic or UI, and LLM-judged metrics cost tokens per run.

No maker product tracked for this yet

MLflow

Open-source platform
Open sourceFree tierllms.txt

Best forTracing, evaluating and versioning prompts for LLM apps in the same open-source platform many teams already use for ML experiments.

Trade-offGrew out of classic ML tooling, so the UI and concepts are broader than an LLM-only tool; you run the tracking server yourself unless you use a managed version.

No maker product tracked for this yet