Developer tools

AI infrastructure & LLM ops: the tech stack of 92 real products

Model hosting, gateways, evals, observability and other plumbing for running AI apps (frameworks, MCP and RAG have their own categories).

What sets them apart

Picks at least twice as common here as among products overall, in decisions at least 10 of them show.

The typical stack

The leading pick where at least 8 of them show the decision and the leader has at least a quarter of it.

DecisionMost common pickShareAt default usage
Frontend FrameworkReact43 of 57 · 75%—
LLM APIOpenAI API45 of 53 · 85%$10/mo · GPT-6 Luna
Backend FrameworkFastAPI36 of 49 · 73%—
DatabasePostgreSQL25 of 48 · 52%—
HostingGitHub Pages16 of 42 · 38%—
AI SDK & Agent FrameworkLangChain24 of 39 · 62%—
Transactional EmailResend21 of 32 · 66%$20/mo · Pro 50k
Database Access & ORMsSQLAlchemy19 of 30 · 63%—
Vector DatabaseQdrant9 of 25 · 36%$103/mo · Standard (3 nodes, 0.5 vCPU / 4 GiB each)
AuthenticationBetter Auth7 of 19 · 37%$0/mo · Self-hosted (open source)
PaymentsStripe14 of 16 · 88%fees on revenue · $410 on $10,000/mo
File StorageAmazon S310 of 15 · 67%$11/mo · S3 Standard (US East)
Product AnalyticsPostHog14 of 14 · 100%$0/mo · Pay-as-you-go
Error MonitoringSentry13 of 13 · 100%$29/mo · Team
LLM Observability & EvalsArize Phoenix4 of 12 · 33%—
Web AnalyticsPlausible Analytics3 of 9 · 33%$19/mo · Starter 100k
Image & Video Hostingsharp7 of 8 · 88%—

Monthly bills add up to about $192 at the calculators' default usage, list prices; fees on revenue (Stripe) come on top. Set your own usage →

What they chose, decision by decision

Among the AI infrastructure & LLM ops that show each choice, from makers' products and open-source code alike.

Small samples, fewer than 8 products: Background Jobs & Cron (Celery 4, Temporal 1) · Lifecycle Email (Loops 3)

AI infrastructure & LLM ops we track

60 makers' products and 32 open-source projects. Makers' products first.

LatitudeLatitude is an open-source observability and evaluation platform for AI agents that groups failures found in production traces into trackable issues. It is for developers running LLM agents in production.+22
CekuraAutomated QA for Voice AI and Chat AI agents+4
EdgeeOne gateway for cheaper, faster, unstoppable coding agents+2
HarnessRouterWorld's first unified agent harness interface. Open source+2
SailhouseThe infrastructure you need to build agents anywhere+1
ClawMetryReal-time observability & governance for AI agents
Lamatic.aiLamatic is a platform for building AI agents in a visual editor and deploying them as serverless APIs, with a built-in vector database, tracing and evaluations. It is for development teams shipping agent features.
LangtailShip AI Apps With Fewer Surprises
PrefactorCatch your agents' mistakes before they reach customers
BasaltReach 99% quality on your AI feature
LaminarOpen-source all-in-one platform for engineering AI products
ScorecardEvaluate, Optimize, and Ship AI Agents
xpander.aiPlatform for running AI agents in production, with per-agent identity, user permissions, credential vaulting and audit logs. Runs in xpander's cloud or on a company's own Kubernetes, and agents can be used from Slack, Teams, ChatGPT or Claude.
Agnost AICatch agent failures your evals miss
FoilAn AI agent that monitors your AI agents
KastraRuntime authorization for Claude, Cursor, Codex and OpenClaw
Opper AIThe European AI gateway for agents
MIOSNWe needed a better way to choose LLMs.
TracciaFinally, a vendor-neutral AI Agent Control Plane.
archestraEnterprise AI Platform with guardrails, MCP registry, gateway & orchestrator+17
Tracely-aiTrace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.+9
judgevalThe Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.+2
deepteamDeepTeam is a framework to red team LLMs and AI agents.
reefInfrastructure for continually self‑improving agents
JamAIBaseThe collaborative spreadsheet for AI. Chain cells into powerful pipelines, experiment with prompts and models, and evaluate LLM responses in real-time. Work together seamlessly to build and iterate on+15
lmnrLaminar - open-source observability platform purpose-built for AI agents. YC S24.+15
lmringOpen-source, self-hostable LLM arena with model compare, voting, and leaderboards+14
openllmetryOpen-source observability for your GenAI or LLM application, based on OpenTelemetry+14
mirageThe World's First Virtual Terminal for AI Agents+13
KilnBuild, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.+9
pandaprobeopen source agent engineering platform: traces, evals, and metrics to debug and improve your AI agents. Integrates with LangGraph, CrewAI, Claude Agent SDK, and more. +6
llm-gatewayConnect Your Agents And Harnesses With Any Provider 🦚+5
llmgatewayRoute, manage, and analyze your LLM requests across multiple providers with a unified API interface.+5
trulensEvaluation and Tracking for LLM Experiments and AI Agents+5
OmniRouteNever stop coding. Free MIT AI gateway: one endpoint, 359 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline +4
AReaLThe RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.+3
diffgramThe AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human Supervision, Data Workflow, and UI Catalog to get the most value out of your AI Data.+3
feastThe Open Source Feature Store for AI/ML+3
sieOpen-source inference server and production cluster for all the models your agent needs.+3
transformerlab-appThe open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU clusters.+2
OGADOff Grid AI — private, on-device AI. Run open models (text, vision, image, voice) locally through one OpenAI-compatible gateway. No cloud, no accounts, no API keys.+1
claude-code-hub一个现代化的 Claude Code & Codex API 代理服务,提供智能负载均衡、用户管理和使用统计功能。
ColossalAIMaking large AI models cheaper, faster and more accessible
CubeSandboxInstant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
magnitudeOpen source inference engine for the hardware you already own. Profiles your machine, recommends the best open models for it, and tunes them for your exact hardware. Works on Apple Silicon, NVIDIA, AM
planoPlano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
ramalamaRamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of con
responsible-ai-toolboxResponsible AI Toolbox is a suite of tools providing model and data exploration and assessment user interfaces and libraries that enable a better understanding of AI systems. These interfaces and libr
backend.aiBackend.AI is a streamlined, container-based computing cluster platform that hosts popular computing/ML frameworks and diverse programming languages, with pluggable heterogeneous accelerator support i
omnaraThe open-source alternative to Claude Managed Agents
rl-swarmA fully open source framework for creating RL training swarms over the internet.
LocalAILocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
bishengBISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified model management, Evaluation, SF+13
headroomCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.+9
weaveWeave is a toolkit for developing AI-powered applications, built by Weights & Biases.+8
egmathe first open-source platform for simulation testing, monitoring and self-improving voice agents+6
pezzo🕹️ Open-source, developer-first LLMOps platform designed to streamline prompt design, version management, instant delivery, collaboration, troubleshooting, observability and more.+4
scenarioAgentic testing for agentic codebases+3
ASSERTRequirement-driven evaluation harness for AI agents and LLM applications. Generate behavior-specific test cases, run them against any target (hosted models, callable wrappers, OTel-traced agents), and+1
inspect_aiInspect: A framework for large language model evaluations+1
ZizkaDBAudit trail database for AI agents. Tamper-evident, checksum-backed decision logs with session replay and time-travel debugging to support EU AI Act Article 12 record-keeping. Drift detection, MCP, Py+1
agent-learning-kitGeneral Purpose Evaluation and Simulation Environment for all your AI related Workflows
byzer-llmEasy, fast, and cheap pretrain,finetune, serving for everyone
syftrsyftr is an agent optimizer that helps you find the best agentic workflows for your budget.
api-for-open-llmOpenai style api for open large language models, using LLMs just as chatgpt! Support for LLaMA, LLaMA-2, BLOOM, Falcon, Baichuan, Qwen, Xverse, SqlCoder, CodeLLaMA, ChatGLM, ChatGLM2, ChatGLM3 etc. 开源
garakthe LLM vulnerability scanner
lotusOptimized Agentic and LLM Bulk Processing Over Your Data
ymirYMIR, a streamlined model development product.
agent-controlCentralized agent control plane for governing runtime agent behavior at scale. Configurable, extensible, and production-ready.
gepaOptimize prompts, code, and more with AI-powered Reflective Optimization
llm-spaceA desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first, cloud-ready for managed agents.
mlx-audioA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
mlx-omni-serverMLX Omni Server is a local inference server powered by Apple's MLX framework, specifically designed for Apple Silicon (M-series) chips. It implements OpenAI-compatible API endpoints, enabling seamless
modularThe Modular Platform (includes MAX & Mojo)
mtebMTEB: State-of-the-art evaluation of embeddings across languages and modalities
vmlxvMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm
openliveOpensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) runs locally. An open alternative to ElevenLabs Agents, Gemin
slimeslime is an LLM post-training framework for RL Scaling.
TensorRT-LLMTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT