alibaba/qwen3.5-flash
Qwen 3.5 Flash
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.
Price comparison
USD per one million tokens, using current provider data.
Provider latency
Last-hour p50 in milliseconds; lower is better.
Provider uptime
Reported availability over the last 24 hours; not a contractual SLA.
Inference providers
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Alibaba | 1M | 64K | 684 ms | 300.0 t/s | 100.000% | $0.1 | $0.4 |
What can this model do?
Available modalities and capabilities across inference providers.
Technical profile
- Model type
- Language model
- Release date
- Feb 24, 2026
- Knowledge cutoff
- —
- Regions
- —
- API specifications
- V2, V3, V4
- Temperature control
- Supported
Model identifier
alibaba/qwen3.5-flashRelated models
More options from the same creator or model category.
