DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp(opens in new tab).
- By:
- DeepSeek
- Input:
- Output:
- Context length:
- 1M
- Max output:
- 384K
- Published:
- 2026-09-10
Providers
Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.
| Provider | ||||
|---|---|---|---|---|
| 10% offLowestalibaba-cloud | Off-Peak$0.15$0.135/1M tokensPeak$0.3$0.27/1M tokens | Off-Peak$0.6$0.54/1M tokensPeak$1.2$1.08/1M tokens | ReadOff-Peak$0.015$0.0135/1M tokensPeak$0.03$0.027/1M tokensWrite— | 1M |
| deepseek | Peak$0.3/1M tokensOff-Peak$0.15/1M tokens | Peak$1.2/1M tokensOff-Peak$0.6/1M tokens | ReadPeak$0.006/1M tokensOff-Peak$0.003/1M tokensWrite— | 1M |
Availability
24 hoursSuccess rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Throughput
24 hoursAverage output speed on streaming requests, in tokens per second (TPS).
Latency
24 hoursAverage time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.
Activity
24 hoursToken usage and cost for this model over time, split by the provider that served each request.
Related models
More models from DeepSeek