DeepInfra
Inference cloud serving hundreds of open-weight models — chat, embeddings, speech, image and video — through an OpenAI-compatible API, plus dedicated GPUs for your own models.
DeepInfra is an inference cloud that hosts other labs' open-weight models — DeepSeek, Qwen, Kimi, Llama, Gemma, Nemotron and many more — and bills per token, with no servers to run. Besides chat models it serves embeddings and rerankers, speech recognition and text-to-speech, and image and video generation, and it also lists a few closed models such as Anthropic's and Google's.
You call it with the OpenAI SDK by pointing the base URL at https://api.deepinfra.com/v1/openai and passing a DeepInfra key; an Anthropic Messages endpoint at https://api.deepinfra.com/anthropic lets the Anthropic SDK and Claude Code use its models too. There are also official deepinfra packages for Node and Python and a Vercel AI SDK provider.
Useful for a small team: one account covers chat, embeddings and speech; prompt caching is automatic and bills repeated prefixes at a lower cached-input rate; a per-request service tier trades price for speed (priority at 1.5× the base price, flex at 0.8× for non-urgent work), and the Batch API runs asynchronous jobs at 20% off. Inputs and outputs are held in memory only and not used for training, and when you outgrow shared capacity you can deploy your own model or LoRA on dedicated GPUs billed by the minute.
The main limits: there is no free tier — you must add a card or prepay before the first call — and each account starts at 200 concurrent requests per model until you ask for more. You only get the models DeepInfra chooses to host, and for the closed models it resells, the model maker's own data and training terms apply.
Where it fits
Who uses it
2 makers' products, each linked to the source that shows it, and 6 open-source projects that declare it in their code.
Open source: a project that declares DeepInfra as a dependency in its public code — verifiable, but not necessarily a live product.
Alternatives to DeepInfra
All alternatives by situation →Questions makers ask about DeepInfra
Can I use the OpenAI SDK?
Yes. Set the base URL to https://api.deepinfra.com/v1/openai, use a DeepInfra API token as the key, and pass a DeepInfra model name such as deepseek-ai/DeepSeek-V4-Flash. source ↗
Does it work with the Anthropic SDK or Claude Code?
Yes. DeepInfra exposes an Anthropic Messages-compatible endpoint at https://api.deepinfra.com/anthropic; set it as the SDK's base URL, or as ANTHROPIC_BASE_URL (without /v1) for Claude Code. source ↗
Does DeepInfra store my prompts or train on them?
DeepInfra says inputs are only held in memory during inference and outputs are deleted once returned, request content is not logged, and data is not used for training or shared — except when you call Google or Anthropic models, where those companies' policies apply. Batch jobs may be stored encrypted on disk for a short time. source ↗
What are the rate limits?
The default is 200 concurrent requests per model rather than a requests-per-minute cap, so throughput depends on how long each request takes. You can request an increase from the account dashboard. source ↗
Is there a free tier?
No. You have to add a card or prepay before using the API; prepaid accounts draw down a balance, invoiced accounts are billed monthly, and you can set a spending limit. source ↗
Can I pay less for work that isn't urgent?
Yes. The flex service tier costs 0.8× the base price in exchange for slower responses and occasional unavailability, and the Batch API returns results within 24 hours at 20% below real-time pricing. source ↗
Can I run my own fine-tuned model?
Yes. You can deploy a custom LLM or LoRA adapter on dedicated A100, H100, H200, B200 or B300 GPUs with autoscaling and an OpenAI-compatible endpoint, billed per GPU-hour at minute granularity. source ↗
Is DeepInfra free?
No permanent free tier; paid use starts at Pay per token. source ↗
Can AI coding agents work with DeepInfra?
No llms.txt, official MCP server or CLI found yet.
Who uses DeepInfra?
2 makers' products we track, each with a source, and 6 open-source projects declare it in their code. source ↗