Amux
All models

DeepSeek: DeepSeek V4 Flash 0731

Active

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

Input:
Output:
Context length:
1M
Max output:
393K
Published:
2026-07-31
toolsreasoningstructuredOutputstreamingcaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
LowestdeepseekPeak$0.44/1M tokensOff-Peak$0.22/1M tokensPeak$1.32/1M tokensOff-Peak$0.66/1M tokens
ReadPeak$0.014/1M tokensOff-Peak$0.007/1M tokensWrite
1M
15% offalibaba-cloudOff-Peak$0.22$0.187/1M tokensPeak$0.44$0.374/1M tokensOff-Peak$0.66$0.561/1M tokensPeak$1.32$1.122/1M tokens
ReadOff-Peak$0.022$0.0187/1M tokensPeak$0.044$0.0374/1M tokensWrite
1M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from DeepSeek

DeepSeek: DeepSeek V4 Flash Vision Exp1M contextfrom $0.22 / M tokens1 providerDeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731(opens in new tab) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total.
DeepSeek: DeepSeek V4 Pro 08131M contextfrom $0.66 / M tokens2 providersDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.