Deepinfra
99.92%Context41K
Max output16.4K
p50 latency367 ms
Throughput29.0 t/s
Input price$0.09
Output price$0.1
Zero retentionNo training
alibaba/qwen-3-235b
Qwen3-235B-A22B-Instruct-2507 is the updated version of the Qwen3-235B-A22B non-thinking mode, featuring Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Deepinfra | 41K | 16.4K | 367 ms | 29.0 t/s | 99.918% | $0.09 | $0.1 | |
Novita | 131.1K | 32.8K | 589 ms | 67.0 t/s | 100.000% | $0.09 | $0.58 | |
Vertex | 262.1K | 16.4K | 386 ms | 96.0 t/s | 99.888% | $0.22 | $0.88 |
Available modalities and capabilities across inference providers.
alibaba/qwen-3-235bMore options from the same creator or model category.