Amux
All models

Google: Gemini 3.1 Pro Preview

Active

Gemini 3.1 Pro is the next generation in the Gemini series of models, a suite of highly-capable, natively multimodal, reasoning models. Gemini 3 Pro is now Google’s most advanced model for complex tasks, and can comprehend vast datasets, challenging problems from different information sources, including text, audio, images, video, and entire code repositories

Input:
Output:
Context length:
1.05M
Max output:
66K
Published:
2026-02-19
toolsreasoningstructuredOutputstreamingvisioncaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
15% offgoogle-vertex$2-4$1.7-3.4/1M tokens$12-18$10.2-15.3/1M tokens
Read$0.2-0.4$0.17-0.34/1M tokensWrite
1.05M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from Google

Google: Gemini 3.8 Flash1.05M contextfrom $0.75 / M tokens1 providerGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Google: Gemini 3.7 Flash1.05M contextfrom $1.275 / M tokens1 providerGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.
Google: Gemini 3.5 Flash Lite1.05M contextfrom $0.255 / M tokens1 providerGemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.