Cerebras Inference API pricing
How Cerebras Inference API charges, what to know before you commit, and what you'd pay at your usage — next to what the alternatives would cost for the same thing.
How it charges
Checked 2026-10-07 on www.cerebras.ai ↗ · USD, list pricesIn short
Pay per token on very fast inference: $0.35 / $0.75 per million input/output tokens on gpt-oss-120B.
- Free plan
- Yes
- Side project
- $2.50/mo gpt-oss-120B
- Growing
- $25/mo gpt-oss-120B
- Scaling
- $500/mo gpt-oss-120B
What the model leaves out
A rate-limited free trial (about 1M tokens a day) and a one-time $5 credit are available. Qwen 3.8 27B is the other shared model, at $0.99 / $1.49.
Before you commit
- Only two models are on shared pay-as-you-go inference; others need dedicated capacity through sales.
- Cached input tokens get no discount.
- Context is 131K tokens on paid plans (65K on the free trial).
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| gpt-oss-120B | None | $0.35 per million tokens | $0.75 per million tokens |
What you'd pay
gpt-oss-120B50 million tokens × $0.35 per million tokens + 10 million tokens × $0.75 per million tokens$25/mo
Cost as you grow
All LLM API pricing, compared →Which one fits your situation →