Alibaba
100.00%Context131.1K
Max output129K
p50 latency477 ms
Throughput53.0 t/s
Input price$0.4
Output price$1.6
Zero retentionNo training
alibaba/qwen3-vl-instruct
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Alibaba | 131.1K | 129K | 477 ms | 53.0 t/s | 100.000% | $0.4 | $1.6 | |
Deepinfra | 262.1K | 262.1K | 900 ms | 13.5 t/s | 100.000% | $0.2 | $0.88 | |
Novita | 131.1K | 32.8K | 682 ms | 48.5 t/s | 100.000% | $0.3 | $1.5 |
Available modalities and capabilities across inference providers.
alibaba/qwen3-vl-instructMore options from the same creator or model category.