GLM-5.1: Self-Hosting Estimate
GLM-5.1 (Zhipu AI) on different GPUs: configuration, estimated performance and use cases.
Data updated: 2026-09-16B300 288GB
288 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)- Suggested precision
- FP16/BF16 (full precision)
- GPUs needed
- 8
- Est. decode speed
- 800 tok/ssingle-stream, theoretical
- Hardware list price
- $429,223–$515,068Cloud rental (live) ×8: $79.52–$79.52/hr (2026-10-05)
- Hardware market price
- $14,307,431–$16,692,003
- Typical use
- Enterprise-grade ultra-large-scale training and high-concurrency inference clusters
GB200 192GB (single card equivalent)
192 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)- Suggested precision
- INT8
- GPUs needed
- 8
- Est. decode speed
- 1,600 tok/ssingle-stream, theoretical
- Hardware list price
- $238,648–$357,972
- Hardware market price
- $715,372–$1,430,743
- Typical use
- Ultra-large-scale trillion-parameter model training and inference clusters (networked with Grace CPU)
B200 180GB
180 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)- Suggested precision
- INT8
- GPUs needed
- 8
- Est. decode speed
- 1,600 tok/ssingle-stream, theoretical
- Hardware list price
- n/aCloud rental (live) ×8: $50.02–$75.08/hr (2026-10-06)
- Hardware market price
- $447,107–$521,625
- Typical use
- Enterprise-grade large-scale online inference and training
H200 141GB
141 GB · 4,800 GB/s · Top-tier export-controlled card (channels to China are volatile)- Suggested precision
- INT8
- GPUs needed
- 8
- Est. decode speed
- 960 tok/ssingle-stream, theoretical
- Hardware list price
- $240,365–$386,301Cloud rental (live) ×8: $33.70–$40.97/hr (2026-10-06)
- Hardware market price
- $223,554–$232,496
- Typical use
- Enterprise-grade production environment high-concurrency inference
H100 SXM5 80GB
80 GB · 3,350 GB/s · Top-tier export-controlled card (no official channels to China)- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 8
- Est. decode speed
- 1,340 tok/ssingle-stream, theoretical
- Hardware list price
- $300,456–$343,378Cloud rental (live) ×8: $14.31–$32.01/hr (2026-10-06)
- Hardware market price
- $402,396–$417,300
- Typical use
- Enterprise-grade production environment high-concurrency inference/training
H800 SXM 80GB
80 GB · 3,350 GB/s · Top-tier export-controlled card (discontinued)- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 8
- Est. decode speed
- 1,340 tok/ssingle-stream, theoretical
- Hardware list price
- $292,110–$309,994
- Hardware market price
- $402,396–$417,300
- Typical use
- Enterprise-grade production deployment (original China-compliant model, now discontinued)
H20 96GB
96 GB · 4,000 GB/s · Export-compliant special supply card (significantly reduced compute power)- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 8
- Est. decode speed
- 1,600 tok/ssingle-stream, theoretical
- Hardware list price
- n/a
- Hardware market price
- $95,383–$131,151
- Typical use
- Suboptimal choice in export-compliant environments, medium-to-large scale inference deployment (FP16 compute power is only about 15% of H100, but higher memory bandwidth offers an advantage in decode-intensive scenarios)
A800 PCIe 80GB
80 GB · 1,935 GB/s · Compliant special supply card (export compliance status adjusted multiple times)- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 8
- Est. decode speed
- 774 tok/ssingle-stream, theoretical
- Hardware list price
- $119,229–$123,998
- Hardware market price
- $83,460–$101,344
- Typical use
- Medium-scale inference/training deployment in export-compliant environments
Ascend 910B 64GB
64 GB · 400 GB/s · Domestic alternative (not subject to US export controls)- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 8
- Est. decode speed
- 160 tok/ssingle-stream, theoretical
- Hardware list price
- n/a
- Hardware market price
- $131,151–$143,074
- Typical use
- Domestic alternative, not subject to US export controls, medium-to-large scale inference/training deployment, memory bandwidth significantly lower than NVIDIA cards of similar VRAM capacity
L40S 48GB
48 GB · 864 GB/s · Compliant inference card- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 16 (exceeds one node: multi-node needed)
- Est. decode speed
- 691 tok/ssingle-stream, theoretical
- Hardware list price
- $107,306–$154,997Cloud rental (live) ×16: $9.62–$13.06/hr (2026-10-06)
- Hardware market price
- $107,306–$154,997
- Typical use
- SME teams building inference services, medium-low concurrency
L20 48GB
48 GB · 864 GB/s · Compliant inference card- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 16 (exceeds one node: multi-node needed)
- Est. decode speed
- 691 tok/ssingle-stream, theoretical
- Hardware list price
- $57,230–$76,306
- Hardware market price
- $59,614–$83,460
- Typical use
- SME teams building inference services, medium-low concurrency
RTX 5090D 24GB
24 GB · 1,344 GB/s · Consumer-grade/small-scale self-built- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 32 (exceeds one node: multi-node needed)
- Est. decode speed
- 2,150 tok/ssingle-stream, theoretical
- Hardware list price
- $97,767–$100,147Cloud rental (live) ×32: $13.73–$17.15/hr (2026-10-06)
- Hardware market price
- $143,074–$190,766
- Typical use
- Personal development/small-scale verification/POC, not recommended for production-level high concurrency
RTX 4090D 24GB
24 GB · 1,008 GB/s · Consumer-grade/small-scale self-built- Suggested precision
- INT4 (AWQ/GPTQ)
- GPUs needed
- 32 (exceeds one node: multi-node needed)
- Est. decode speed
- 1,613 tok/ssingle-stream, theoretical
- Hardware list price
- $61,994–$90,609Cloud rental (live) ×32: $9.92–$13.86/hr (2026-10-06)
- Hardware market price
- $123,998–$143,074
- Typical use
- Personal development/small-scale verification/POC, not recommended for production-level high concurrency
Estimated speed = (memory bandwidth × GPU count) ÷ (active parameters × bytes per parameter): the memory-bound upper limit for one request (roofline). It scales linearly with GPU count and ignores interconnect overhead, batching and framework losses, so real results are usually well below it. The linear assumption is reasonable for NVLink data-center cards (H/B series) but too optimistic for consumer PCIe cards without NVLink (RTX 5090D/4090D); multi-GPU consumer setups lose far more to communication, so equal-looking numbers do not mean equal performance. “Suggested precision” is the highest precision that fits within 8 GPUs; beyond that it is flagged as multi-node. Prices are USD converted at ¥6.71 per USD; cloud rental is the international vast.ai market rate, not a China purchase price.