Bedrock
99.95%Context128K
Max output8.2K
p50 latency346 ms
Throughput197.0 t/s
Input price$0.07
Output price$0.3
Zero retentionNo training
openai/gpt-oss-20b
A compact, open-weight language model optimized for low-latency and resource-constrained environments, including local and edge deployments.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Bedrock | 128K | 8.2K | 346 ms | 197.0 t/s | 99.948% | $0.07 | $0.3 | |
Deepinfra | 131.1K | 8.2K | 226 ms | 108.0 t/s | 100.000% | $0.03 | $0.14 | |
Groq | 131.1K | 32.8K | 304 ms | 915.0 t/s | 100.000% | $0.075 | $0.3 | |
Novita | 131.1K | 32.8K | 441 ms | 283.0 t/s | 100.000% | $0.04 | $0.15 | |
Parasail | 131.1K | 8.2K | 594 ms | 160.0 t/s | 100.000% | $0.04 | $0.2 | |
Togetherai | 131.1K | 8.2K | 341 ms | 113.0 t/s | 99.826% | $0.05 | $0.2 |
Available modalities and capabilities across inference providers.
openai/gpt-oss-20bMore options from the same creator or model category.