钽TANDIU钽铥
SELF-HOSTING

Qwen3.8 27B: Self-Hosting Estimate

Qwen3.8 27B (Alibaba (Qwen)) on different GPUs: configuration, estimated performance and use cases.

Data updated: 2026-09-16

← All self-hosting estimates

Total parameters28BSets weight VRAM
Active parameters28BSets decode compute and speed
Public benchmarksAA 34.0 · LiveBench 75.3Hosted-API output speed 46.0 tok/s (not a local estimate)

B300 288GB

288 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
144 tok/ssingle-stream, theoretical
Hardware list price
$53,653–$64,383Cloud rental (live): $9.94–$9.94/hr (2026-10-05)
Hardware market price
$1,788,429–$2,086,500
Typical use
Enterprise-grade ultra-large-scale training and high-concurrency inference clusters

GB200 192GB (single card equivalent)

192 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
144 tok/ssingle-stream, theoretical
Hardware list price
$29,831–$44,746
Hardware market price
$89,421–$178,843
Typical use
Ultra-large-scale trillion-parameter model training and inference clusters (networked with Grace CPU)

B200 180GB

180 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
144 tok/ssingle-stream, theoretical
Hardware list price
n/aCloud rental (live): $6.25–$9.38/hr (2026-10-06)
Hardware market price
$55,888–$65,203
Typical use
Enterprise-grade large-scale online inference and training

H200 141GB

141 GB · 4,800 GB/s · Top-tier export-controlled card (channels to China are volatile)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
86 tok/ssingle-stream, theoretical
Hardware list price
$30,046–$48,288Cloud rental (live): $4.21–$5.12/hr (2026-10-06)
Hardware market price
$27,944–$29,062
Typical use
Enterprise-grade production environment high-concurrency inference

H100 SXM5 80GB

80 GB · 3,350 GB/s · Top-tier export-controlled card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
60 tok/ssingle-stream, theoretical
Hardware list price
$37,557–$42,922Cloud rental (live): $1.79–$4.00/hr (2026-10-06)
Hardware market price
$50,300–$52,163
Typical use
Enterprise-grade production environment high-concurrency inference/training

H800 SXM 80GB

80 GB · 3,350 GB/s · Top-tier export-controlled card (discontinued)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
60 tok/ssingle-stream, theoretical
Hardware list price
$36,514–$38,749
Hardware market price
$50,300–$52,163
Typical use
Enterprise-grade production deployment (original China-compliant model, now discontinued)

H20 96GB

96 GB · 4,000 GB/s · Export-compliant special supply card (significantly reduced compute power)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
72 tok/ssingle-stream, theoretical
Hardware list price
n/a
Hardware market price
$11,923–$16,394
Typical use
Suboptimal choice in export-compliant environments, medium-to-large scale inference deployment (FP16 compute power is only about 15% of H100, but higher memory bandwidth offers an advantage in decode-intensive scenarios)

A800 PCIe 80GB

80 GB · 1,935 GB/s · Compliant special supply card (export compliance status adjusted multiple times)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
35 tok/ssingle-stream, theoretical
Hardware list price
$14,904–$15,500
Hardware market price
$10,433–$12,668
Typical use
Medium-scale inference/training deployment in export-compliant environments

Ascend 910B 64GB

64 GB · 400 GB/s · Domestic alternative (not subject to US export controls)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
2
Est. decode speed
14 tok/ssingle-stream, theoretical
Hardware list price
n/a
Hardware market price
$32,788–$35,769
Typical use
Domestic alternative, not subject to US export controls, medium-to-large scale inference/training deployment, memory bandwidth significantly lower than NVIDIA cards of similar VRAM capacity

L40S 48GB

48 GB · 864 GB/s · Compliant inference card
Suggested precision
FP16/BF16 (full precision)
GPUs needed
2
Est. decode speed
31 tok/ssingle-stream, theoretical
Hardware list price
$13,413–$19,375Cloud rental (live) ×2: $1.20–$1.63/hr (2026-10-06)
Hardware market price
$13,413–$19,375
Typical use
SME teams building inference services, medium-low concurrency

L20 48GB

48 GB · 864 GB/s · Compliant inference card
Suggested precision
FP16/BF16 (full precision)
GPUs needed
2
Est. decode speed
31 tok/ssingle-stream, theoretical
Hardware list price
$7,154–$9,538
Hardware market price
$7,452–$10,433
Typical use
SME teams building inference services, medium-low concurrency

RTX 5090D 24GB

24 GB · 1,344 GB/s · Consumer-grade/small-scale self-built
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
97 tok/ssingle-stream, theoretical
Hardware list price
$12,221–$12,518Cloud rental (live) ×4: $1.72–$2.14/hr (2026-10-06)
Hardware market price
$17,884–$23,846
Typical use
Personal development/small-scale verification/POC, not recommended for production-level high concurrency

RTX 4090D 24GB

24 GB · 1,008 GB/s · Consumer-grade/small-scale self-built
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
73 tok/ssingle-stream, theoretical
Hardware list price
$7,749–$11,326Cloud rental (live) ×4: $1.24–$1.73/hr (2026-10-06)
Hardware market price
$15,500–$17,884
Typical use
Personal development/small-scale verification/POC, not recommended for production-level high concurrency
Decode speed is a theoretical ceiling, not a benchmark

Estimated speed = (memory bandwidth × GPU count) ÷ (active parameters × bytes per parameter): the memory-bound upper limit for one request (roofline). It scales linearly with GPU count and ignores interconnect overhead, batching and framework losses, so real results are usually well below it. The linear assumption is reasonable for NVLink data-center cards (H/B series) but too optimistic for consumer PCIe cards without NVLink (RTX 5090D/4090D); multi-GPU consumer setups lose far more to communication, so equal-looking numbers do not mean equal performance. “Suggested precision” is the highest precision that fits within 8 GPUs; beyond that it is flagged as multi-node. Prices are USD converted at ¥6.71 per USD; cloud rental is the international vast.ai market rate, not a China purchase price.