DeepSeek shipped a 552B model that only activates 8–16B parameters

“They built a 552-billion-parameter brain that only wakes up 8 billion of it. Your agent swarm just got 4× cheaper overnight.”

The Story

DeepSeek released V4.1 Flash, a 552-billion-parameter MoE model with a causal encoder-decoder architecture that activates 8B parameters on input and 16B on output. It claims 1M context, native vision, one-quarter the HBM and one-eighth the SSD of the prior generation, while beating V4 Pro on speed, cost, and quality. DeepSeek plans to phase out the old model.

Why It Matters

The economics of agents just flipped. Frontier-scale intelligence is becoming cheap enough that swarms move from rich-lab toy to ordinary builder tool: the model is huge, but the bill is tiny.

Evidence

DeepSeek official API docs and Hugging Face (Sept. 10); Hacker News (516 points).

Sources

Daily scan: 2026-09-10