Google
99.85%Context1M
Max output65.5K
p50 latency1,539 ms
Throughput250.0 t/s
Input price$0.75
Output price$3.75
No trainingCaching
google/gemini-3.7-flash
Gemini 3.7 Flash is the next iteration in the Gemini 3 model family, featuring algorithmic improvements to its core reasoning foundation. It supports customizable thinking configurations to control the mix of quality, cost and latency.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Google | 1M | 65.5K | 1,539 ms | 250.0 t/s | 99.850% | $0.75 | $3.75 | |
Vertex | 1M | 65.5K | 2,623 ms | 120.5 t/s | 99.972% | $0.75 | $3.75 |
Available modalities and capabilities across inference providers.
google/gemini-3.7-flashMore options from the same creator or model category.