Deepinfra
100.00%Context262.1K
Max output262.1K
p50 latency733 ms
Throughput57.5 t/s
Input price$0.14
Output price$0.58
Zero retentionNo trainingCaching
tencent/hy3
A technical profile of this model's capabilities, pricing, and available inference providers.
USD per one million tokens, using current provider data.
Last-hour p50 in milliseconds; lower is better.
Reported availability over the last 24 hours; not a contractual SLA.
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Deepinfra | 262.1K | 262.1K | 733 ms | 57.5 t/s | 100.000% | $0.14 | $0.58 | |
Gmicloud | 262.1K | 262.1K | 1,519 ms | 200.0 t/s | 98.682% | $0.126 | $0.522 | |
Novita | 262.1K | 262.1K | 1,135 ms | 131.0 t/s | 100.000% | $0.14 | $0.58 | |
Tencent | 256K | 128K | 1,759 ms | 203.5 t/s | 100.000% | $0.132 | $0.528 |
Available modalities and capabilities across inference providers.
tencent/hy3More options from the same creator or model category.