Alibaba
99.94%Context262.1K
Max output32.8K
p50 latency1,003 ms
Throughput47.5 t/s
Input price$1.2
Output price$6
Zero retentionNo trainingCaching
alibaba/qwen3-max
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Alibaba | 262.1K | 32.8K | 1,003 ms | 47.5 t/s | 99.938% | $1.2 | $6 | |
Novita | 262.1K | 65.5K | 1,125 ms | 55.0 t/s | 100.000% | $0.845 | $3.38 |
Available modalities and capabilities across inference providers.
alibaba/qwen3-maxMore options from the same creator or model category.