Cerebras is serving Qwen 3.8 27B at roughly 1,500 tokens per second

“1,500 tokens a second. On a 27B open model. Right now.”

The Story

Cerebras now lists Qwen 3.8 27B on public inference at about 1,500 tokens per second. It is raw speed for an open model builders can use now, rather than a lab-only demo.

Why It Matters

Open models reaching speeds that change what an interactive agent feels like.

Evidence

Cerebras inference documentation; Hacker News.

Caveat

Lower visual drama; benchmark the claimed speed under reproducible conditions before presenting it as typical.

Sources

Daily scan: 2026-09-04