2026-09-04 / Signal #5
Cerebras is serving Qwen 3.8 27B at roughly 1,500 tokens per second
“1,500 tokens a second. On a 27B open model. Right now.”
The Story
Cerebras now lists Qwen 3.8 27B on public inference at about 1,500 tokens per second. It is raw speed for an open model builders can use now, rather than a lab-only demo.
Why It Matters
Open models reaching speeds that change what an interactive agent feels like.
Evidence
Cerebras inference documentation; Hacker News.
Caveat
Lower visual drama; benchmark the claimed speed under reproducible conditions before presenting it as typical.
Sources
Daily scan: 2026-09-04