Cerebras Inference

by Cerebras Systems (Nasdaq: CBRS since 2026-05-14)

High-throughput, low-latency inference on Cerebras wafer-scale hardware via a managed cloud API and SDKs.

Docs: https://inference-docs.cerebras.ai | Console: https://cloud.cerebras.ai

Company context (2026)

  • Cerebras announced a multi-year deal with OpenAI in January 2026 (reported as worth more than $20 billion, 750 MW of inference capacity); OpenAI uses Cerebras as a fast-inference provider (press reports; terms from secondary sources).
  • IPO: priced 2026-05-13 at 6.4 billion (Cerebras press release, Investing.com).

Features

  • Managed inference API (REST plus Python and TypeScript SDKs), OpenAI-compatible style chat completions.
  • Two access modes per the docs: Shared Inference (public catalog, pay-as-you-go) and Dedicated Inference (reserved capacity, more model families, production SLAs).
  • Structured outputs (JSON schema constrained decoding).
  • Also reachable through AWS Marketplace, OpenRouter, Hugging Face and Vercel (listed on the Cerebras pricing page).

Models (shared catalog, checked 2026-10-02)

The docs model overview lists two production models in Shared Inference: gpt-oss-120b (~3,000 tokens/s, vendor figure) and qwen-3.8-27b (~1,850 tokens/s, vendor figure). The catalog rotates quickly. The docs deprecation page lists these as deprecated or retired, so older tutorials break:

  • llama-4-scout-17b-16e-instruct (2025-11-03), llama-4-maverick-17b-128e-instruct (2025-10-15)
  • llama-3.3-70b and qwen-3-32b (2026-02-16), llama3.1-8b and qwen-3-235b-a22b-instruct-2507 (2026-05-27)
  • zai-glm-4.7 (2026-08-17), gemma-4-31b (2026-09-03)
    Dates are the dates shown on the deprecation page. Related model note: qwen-family.

Performance

Throughput numbers are vendor-reported, model- and configuration-dependent; validate with your own prompts and payload sizes.

Pricing

Tiers per the vendor: a free Developer tier, pay-as-you-go per token, and Enterprise/dedicated plans. Per-token rates are not copied here: the vendor’s pricing page did not render its table for this review and secondary sites disagree and are stale (they still list deprecated models). Check https://www.cerebras.ai/pricing.

Quickstart

pip install cerebras_cloud_sdk        # Python  
npm install @cerebras/cerebras_cloud_sdk   # Node / TypeScript  
export CEREBRAS_API_KEY="..."  
import os  
from cerebras.cloud.sdk import Cerebras  
  
client = Cerebras(api_key=os.environ.get("CEREBRAS_API_KEY"))  
resp = client.chat.completions.create(  
    model="gpt-oss-120b",   # the SDK README uses this model  
    messages=[{"role": "user", "content": "Explain why fast inference matters."}],  
)  
print(resp.choices[0].message.content)  

Use cases

Latency-sensitive agents and multi-step reasoning, real-time assistants, high-throughput generation; see inference-and-serving for how it compares with GPU serving and groq for a similar speed-focused provider.

Caveats

  • Confirm model IDs against the docs before integrating; they are retired on a rolling basis.
  • Check regional availability and contract terms if data residency matters.

inference-and-serving, groq

Sources