Weights & Biases

AI developer platform whose W&B Weave product traces and evaluates LLM apps and agents, alongside experiment tracking and a model registry for model training.

Ask your AI about this, with this page as the source:ChatGPT ↗Claude ↗Perplexity ↗

Weights & Biases started as experiment tracking for people training machine learning models, and its LLM-app product is W&B Weave. Weave records what your agent or LLM app did on each request and lets you measure whether a prompt or model change made answers better, because non-deterministic outputs are hard to debug and hard to test with ordinary unit tests.

There are two ways in. For agents, Weave's tracing is built on OpenTelemetry and its GenAI conventions: built-in integrations cover agent SDKs and harnesses such as the OpenAI Agents SDK and Claude Code, and an agent that already emits OTel spans can send them to Weave's endpoint without the Weave SDK. For application code, the Python or TypeScript SDK traces any function you mark as an op, capturing inputs, outputs, cost, token counts and latency, and patches many LLM providers and frameworks automatically.

On top of traces you get evaluations against datasets with built-in, custom and LLM-judge scorers, side-by-side comparisons, versioned prompts, models and datasets, human feedback, and guardrails and monitors for production. If you also train or fine-tune models, the same account covers W&B Models for experiments, sweeps, artifacts and a registry.

Things to know: since September 30, 2026 the hosted W&B app lives inside CoreWeave Forge, and its plans are sold under that name. CoreWeave describes its newer Agent Lens product as the successor to Weave for agent tracing, while Weave remains the place for evaluations that Agent Lens does not support yet. The Weave SDK is Apache 2.0, but self-hosting means a licensed W&B Server on Kubernetes with MySQL, object storage and a ClickHouse cluster for Weave.

Where it fits

Who uses it

1 makers' products, each linked to the source that shows it, and 5 open-source projects that declare it in their code.

The maker says so 1Declared in code 5How evidence is collected →
CartesiaVoice AI API for low-latency text-to-speech (Sonic), speech-to-text (Ink) and voice agen…

“We use Weights and Biases to track our machine learning experiments.”

LLM Observability & EvalsMaker says so · source ↗

Open source: a project that declares Weights & Biases as a dependency in its public code — verifiable, but not necessarily a live product.

What makers say

1 makers on why they use Weights & Biases, in their own words on Product Hunt.

We use Weights and Biases to track our machine learning experiments.
CartesiaSep 2026 ↗

Alternatives to Weights & Biases

All alternatives by situation →

Questions makers ask about Weights & Biases

What does Weave trace?

Agent sessions, turns, LLM calls and tool calls through its agent integrations or any OpenTelemetry spans, and, for application code, every function decorated as a Weave op with its inputs, outputs, cost, token count and latency. The two kinds of traces appear on separate Agents and Traces pages. source ↗

How do evaluations work?

You define one or more scoring functions, run them against curated test cases, and compare evaluations across metrics and individual samples, with Weave tracking which prompt and model versions produced each result. Guardrails and scorers can also monitor quality and safety in production. source ↗

Which languages are supported?

Weave has Python and TypeScript libraries, installed with pip install weave or npm install weave. Agents in other languages can send OpenTelemetry spans to Weave's OTLP endpoint without the SDK. source ↗

How much does it cost?

The Free plan includes 1 GB of Weave data ingestion and 5 GB of storage a month. Pro starts at $60 a month with a 30-day free trial, 1.5 GB of ingestion and $0.10 per additional MB, and is meant for teams under 50 employees; Enterprise is custom. source ↗

What is Agent Lens, and is Weave going away?

Agent Lens is CoreWeave's newer agent observability product, described as the successor to Weave for agent tracing; both receive the same trace data, so nothing needs migrating. Evaluations are not in Agent Lens yet, and CoreWeave says to use Weave alongside it for those. source ↗

Can I self-host it?

Yes, as W&B Self-Managed on Kubernetes, which needs a W&B Server license (a free trial license can be generated) and, for Weave, a ClickHouse cluster plus S3-compatible storage. The hosted alternatives are the multi-tenant cloud or a single-tenant Dedicated Cloud. source ↗

Is Weights & Biases free?

Yes — there is a free tier a small product can run on; paid use starts at $60/mo. source ↗

Is Weights & Biases open source or self-hostable?

Not open source, and you can self-host it. source ↗

Can AI coding agents work with Weights & Biases?

No llms.txt, official MCP server or CLI found yet.

Who uses Weights & Biases?

1 makers' products we track, each with a source, and 5 open-source projects declare it in their code. source ↗