钽TANDIU钽铥
SELF-HOSTING

Kimi K2.5: Self-Hosting Estimate

Kimi K2.5 (Moonshot AI) on different GPUs: configuration, estimated performance and use cases.

Data updated: 2026-09-16

← All self-hosting estimates

Total parameters1,027BSets weight VRAM
Active parameters32BSets decode compute and speed
Public benchmarksNo public benchmark yet

B300 288GB

288 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
INT8
GPUs needed
8
Est. decode speed
2,000 tok/ssingle-stream, theoretical
Hardware list price
$429,223–$515,068Cloud rental (live) ×8: $79.52–$79.52/hr (2026-10-05)
Hardware market price
$14,307,431–$16,692,003
Typical use
Enterprise-grade ultra-large-scale training and high-concurrency inference clusters

GB200 192GB (single card equivalent)

192 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
INT8
GPUs needed
8
Est. decode speed
2,000 tok/ssingle-stream, theoretical
Hardware list price
$238,648–$357,972
Hardware market price
$715,372–$1,430,743
Typical use
Ultra-large-scale trillion-parameter model training and inference clusters (networked with Grace CPU)

B200 180GB

180 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
INT8
GPUs needed
8
Est. decode speed
2,000 tok/ssingle-stream, theoretical
Hardware list price
n/aCloud rental (live) ×8: $50.02–$75.08/hr (2026-10-06)
Hardware market price
$447,107–$521,625
Typical use
Enterprise-grade large-scale online inference and training

H200 141GB

141 GB · 4,800 GB/s · Top-tier export-controlled card (channels to China are volatile)
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
8
Est. decode speed
2,400 tok/ssingle-stream, theoretical
Hardware list price
$240,365–$386,301Cloud rental (live) ×8: $33.70–$40.97/hr (2026-10-06)
Hardware market price
$223,554–$232,496
Typical use
Enterprise-grade production environment high-concurrency inference

H100 SXM5 80GB

80 GB · 3,350 GB/s · Top-tier export-controlled card (no official channels to China)
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
8
Est. decode speed
1,675 tok/ssingle-stream, theoretical
Hardware list price
$300,456–$343,378Cloud rental (live) ×8: $14.31–$32.01/hr (2026-10-06)
Hardware market price
$402,396–$417,300
Typical use
Enterprise-grade production environment high-concurrency inference/training

H800 SXM 80GB

80 GB · 3,350 GB/s · Top-tier export-controlled card (discontinued)
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
8
Est. decode speed
1,675 tok/ssingle-stream, theoretical
Hardware list price
$292,110–$309,994
Hardware market price
$402,396–$417,300
Typical use
Enterprise-grade production deployment (original China-compliant model, now discontinued)

H20 96GB

96 GB · 4,000 GB/s · Export-compliant special supply card (significantly reduced compute power)
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
8
Est. decode speed
2,000 tok/ssingle-stream, theoretical
Hardware list price
n/a
Hardware market price
$95,383–$131,151
Typical use
Suboptimal choice in export-compliant environments, medium-to-large scale inference deployment (FP16 compute power is only about 15% of H100, but higher memory bandwidth offers an advantage in decode-intensive scenarios)

A800 PCIe 80GB

80 GB · 1,935 GB/s · Compliant special supply card (export compliance status adjusted multiple times)
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
8
Est. decode speed
968 tok/ssingle-stream, theoretical
Hardware list price
$119,229–$123,998
Hardware market price
$83,460–$101,344
Typical use
Medium-scale inference/training deployment in export-compliant environments

Ascend 910B 64GB

64 GB · 400 GB/s · Domestic alternative (not subject to US export controls)
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
16 (exceeds one node: multi-node needed)
Est. decode speed
400 tok/ssingle-stream, theoretical
Hardware list price
n/a
Hardware market price
$262,303–$286,149
Typical use
Domestic alternative, not subject to US export controls, medium-to-large scale inference/training deployment, memory bandwidth significantly lower than NVIDIA cards of similar VRAM capacity

L40S 48GB

48 GB · 864 GB/s · Compliant inference card
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
16 (exceeds one node: multi-node needed)
Est. decode speed
864 tok/ssingle-stream, theoretical
Hardware list price
$107,306–$154,997Cloud rental (live) ×16: $9.62–$13.06/hr (2026-10-06)
Hardware market price
$107,306–$154,997
Typical use
SME teams building inference services, medium-low concurrency

L20 48GB

48 GB · 864 GB/s · Compliant inference card
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
16 (exceeds one node: multi-node needed)
Est. decode speed
864 tok/ssingle-stream, theoretical
Hardware list price
$57,230–$76,306
Hardware market price
$59,614–$83,460
Typical use
SME teams building inference services, medium-low concurrency

RTX 5090D 24GB

24 GB · 1,344 GB/s · Consumer-grade/small-scale self-built
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
32 (exceeds one node: multi-node needed)
Est. decode speed
2,688 tok/ssingle-stream, theoretical
Hardware list price
$97,767–$100,147Cloud rental (live) ×32: $13.73–$17.15/hr (2026-10-06)
Hardware market price
$143,074–$190,766
Typical use
Personal development/small-scale verification/POC, not recommended for production-level high concurrency

RTX 4090D 24GB

24 GB · 1,008 GB/s · Consumer-grade/small-scale self-built
Suggested precision
INT4 (AWQ/GPTQ)
GPUs needed
32 (exceeds one node: multi-node needed)
Est. decode speed
2,016 tok/ssingle-stream, theoretical
Hardware list price
$61,994–$90,609Cloud rental (live) ×32: $9.92–$13.86/hr (2026-10-06)
Hardware market price
$123,998–$143,074
Typical use
Personal development/small-scale verification/POC, not recommended for production-level high concurrency
Decode speed is a theoretical ceiling, not a benchmark

Estimated speed = (memory bandwidth × GPU count) ÷ (active parameters × bytes per parameter): the memory-bound upper limit for one request (roofline). It scales linearly with GPU count and ignores interconnect overhead, batching and framework losses, so real results are usually well below it. The linear assumption is reasonable for NVLink data-center cards (H/B series) but too optimistic for consumer PCIe cards without NVLink (RTX 5090D/4090D); multi-GPU consumer setups lose far more to communication, so equal-looking numbers do not mean equal performance. “Suggested precision” is the highest precision that fits within 8 GPUs; beyond that it is flagged as multi-node. Prices are USD converted at ¥6.71 per USD; cloud rental is the international vast.ai market rate, not a China purchase price.