Amux
All models

DeepSeek: DeepSeek V4.1 Flash

Active

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp(opens in new tab).

Input:
Output:
Context length:
1M
Max output:
384K
Published:
2026-09-10
toolsreasoningstructuredOutputstreamingvisioncaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
10% offLowestalibaba-cloudOff-Peak$0.15$0.135/1M tokensPeak$0.3$0.27/1M tokensOff-Peak$0.6$0.54/1M tokensPeak$1.2$1.08/1M tokens
ReadOff-Peak$0.015$0.0135/1M tokensPeak$0.03$0.027/1M tokensWrite
1M
deepseekPeak$0.3/1M tokensOff-Peak$0.15/1M tokensPeak$1.2/1M tokensOff-Peak$0.6/1M tokens
ReadPeak$0.006/1M tokensOff-Peak$0.003/1M tokensWrite
1M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from DeepSeek

DeepSeek: DeepSeek V4 Flash Vision Exp1M contextfrom $0.066 / M tokens1 providerDeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731(opens in new tab) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total.
DeepSeek: DeepSeek V4 Pro 08131M contextfrom $0.198 / M tokens3 providersDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.
DeepSeek: DeepSeek V4 Flash 07311M contextfrom $0.066 / M tokens2 providersDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.