Cerebras Inference
by Cerebras Systems (Nasdaq: CBRS since 2026-05-14)
High-throughput, low-latency inference on Cerebras wafer-scale hardware via a managed cloud API and SDKs.
Docs: https://inference-docs.cerebras.ai | Console: https://cloud.cerebras.ai
Company context (2026)
- Cerebras announced a multi-year deal with OpenAI in January 2026 (reported as worth more than $20 billion, 750 MW of inference capacity); OpenAI uses Cerebras as a fast-inference provider (press reports; terms from secondary sources).
- IPO: priced 2026-05-13 at 6.4 billion (Cerebras press release, Investing.com).
Features
- Managed inference API (REST plus Python and TypeScript SDKs), OpenAI-compatible style chat completions.
- Two access modes per the docs: Shared Inference (public catalog, pay-as-you-go) and Dedicated Inference (reserved capacity, more model families, production SLAs).
- Structured outputs (JSON schema constrained decoding).
- Also reachable through AWS Marketplace, OpenRouter, Hugging Face and Vercel (listed on the Cerebras pricing page).
Models (shared catalog, checked 2026-10-02)
The docs model overview lists two production models in Shared Inference: gpt-oss-120b (~3,000 tokens/s, vendor figure) and qwen-3.8-27b (~1,850 tokens/s, vendor figure). The catalog rotates quickly. The docs deprecation page lists these as deprecated or retired, so older tutorials break:
llama-4-scout-17b-16e-instruct(2025-11-03),llama-4-maverick-17b-128e-instruct(2025-10-15)llama-3.3-70bandqwen-3-32b(2026-02-16),llama3.1-8bandqwen-3-235b-a22b-instruct-2507(2026-05-27)zai-glm-4.7(2026-08-17),gemma-4-31b(2026-09-03)
Dates are the dates shown on the deprecation page. Related model note: qwen-family.
Performance
Throughput numbers are vendor-reported, model- and configuration-dependent; validate with your own prompts and payload sizes.
Pricing
Tiers per the vendor: a free Developer tier, pay-as-you-go per token, and Enterprise/dedicated plans. Per-token rates are not copied here: the vendor’s pricing page did not render its table for this review and secondary sites disagree and are stale (they still list deprecated models). Check https://www.cerebras.ai/pricing.
Quickstart
pip install cerebras_cloud_sdk # Python
npm install @cerebras/cerebras_cloud_sdk # Node / TypeScript
export CEREBRAS_API_KEY="..." import os
from cerebras.cloud.sdk import Cerebras
client = Cerebras(api_key=os.environ.get("CEREBRAS_API_KEY"))
resp = client.chat.completions.create(
model="gpt-oss-120b", # the SDK README uses this model
messages=[{"role": "user", "content": "Explain why fast inference matters."}],
)
print(resp.choices[0].message.content) Use cases
Latency-sensitive agents and multi-step reasoning, real-time assistants, high-throughput generation; see inference-and-serving for how it compares with GPU serving and groq for a similar speed-focused provider.
Caveats
- Confirm model IDs against the docs before integrating; they are retired on a rolling basis.
- Check regional availability and contract terms if data residency matters.
Related
inference-and-serving, groq
Sources
- https://inference-docs.cerebras.ai/models/overview (accessed 2026-10-02)
- https://inference-docs.cerebras.ai/support/deprecation (accessed 2026-10-02)
- https://www.cerebras.ai/pricing (accessed 2026-10-02; table not rendered)
- https://github.com/Cerebras/cerebras-cloud-sdk-python (accessed 2026-10-02)
- https://www.cerebras.ai/press-release/cerebras-systems-announces-pricing-of-initial-public-offering (via search result, accessed 2026-10-02)
- https://www.investing.com/news/stock-market-news/cerebras-systems-closes-638-billion-ipo-on-nasdaq-432SI-4693719 (search snippet, accessed 2026-10-02)
- https://www.sec.gov/Archives/edgar/data/2021728/000162828026025762/cerebras-sx1april2026.htm (S-1, search result)