Alibaba
99.98%Context991K
Max output128K
p50 latency3,123 ms
Throughput63.5 t/s
Input price$0.16
Output price$0.47
Zero retentionNo trainingCaching
alibaba/qwen3.8-flash
Qwen3.8-Flash is Qwen’s fast, cost-efficient multimodal model, combining advanced reasoning and generation with a native 1M-token context window. Built for coding, agentic workflows, and visual understanding, it handles large codebases, long documents, charts, videos, and desktop applications.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Alibaba | 991K | 128K | 3,123 ms | 63.5 t/s | 99.985% | $0.16 | $0.47 |
Available modalities and capabilities across inference providers.
alibaba/qwen3.8-flashMore options from the same creator or model category.