Baseten
99.99%Context1M
Max output131K
p50 latency889 ms
Throughput229.0 t/s
Input price$0.15
Output price$0.5
Zero retentionNo trainingCaching
zai/glm-5.3-flash
GLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Baseten | 1M | 131K | 889 ms | 229.0 t/s | 99.991% | $0.15 | $0.5 | |
Deepinfra | 1M | 1M | 1,603 ms | 16.0 t/s | 99.908% | $0.075 | $0.25 | |
Digitalocean | 1M | 1M | 434 ms | 109.0 t/s | 100.000% | $0.15 | $0.5 | |
Fireworks | 1M | 131K | 1,289 ms | 122.0 t/s | 99.998% | $0.15 | $0.5 | |
Friendli | 1M | 131K | 360 ms | 96.0 t/s | 99.943% | $0.15 | $0.5 | |
Gmicloud | 1M | 1M | 4,174 ms | 60.0 t/s | 99.526% | $0.15 | $0.5 | |
Modal | 1M | 1M | 382 ms | 75.5 t/s | 99.956% | $0.45 | $1.5 | |
Morph | 1M | 1M | 965 ms | 73.0 t/s | 100.000% | $0.39 | $1.36 | |
Novita | 1M | 128K | 1,704 ms | 105.0 t/s | 100.000% | $0.15 | $0.5 | |
Parasail | 1M | 1M | 723 ms | 107.5 t/s | 100.000% | $0.15 | $0.5 | |
Particle | 1M | 131K | 1,126 ms | 78.0 t/s | 99.999% | $0.05 | $0.2 | |
Relace | 1M | 131.1K | 1,356 ms | 61.0 t/s | 99.845% | $0.071 | $0.238 | |
Runinfra | 1M | 1M | 979 ms | 48.0 t/s | 99.876% | $0.1 | $0.4 | |
Runware | 131.1K | 131.1K | 650 ms | 142.0 t/s | 99.603% | $0.15 | $0.5 | |
Streamlake | 1M | 128K | 1,415 ms | 82.0 t/s | 99.831% | $0.15 | $0.5 | |
Togetherai | 1M | 1M | 725 ms | 95.0 t/s | 99.922% | $0.15 | $0.5 | |
Wafer | 1M | 1M | 607 ms | 28.0 t/s | 99.912% | $0.1 | $0.35 | |
Zai | 1M | 131K | 5,924 ms | 54.0 t/s | 99.946% | $0.15 | $0.5 |
Available modalities and capabilities across inference providers.
zai/glm-5.3-flashMore options from the same creator or model category.