Amux
All models

Google: Gemini 3.6 Flash

Active

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Input:
Output:
Context length:
1.05M
Max output:
66K
Published:
2026-07-22
toolsreasoningstructuredOutputstreamingvisioncaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
15% offgoogle-vertex$1.5$1.275/1M tokens$7.5$6.375/1M tokens
Read$0.15$0.1275/1M tokensWrite
1.05M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from Google

Google: Gemini 3.8 Flash1.05M contextfrom $0.75 / M tokens1 providerGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Google: Gemini 3.7 Flash1.05M contextfrom $1.275 / M tokens1 providerGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.
Google: Gemini 3.5 Flash Lite1.05M contextfrom $0.255 / M tokens1 providerGemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.