钽TANDIU钽铥
SELF-HOSTING

Hunyuan A13B: Self-Hosting Estimate

Hunyuan A13B (Tencent) on different GPUs: configuration, estimated performance and use cases.

Data updated: 2026-09-16

← All self-hosting estimates

Total parameters80BSets weight VRAM
Active parameters13BSets decode compute and speed
Public benchmarksNo public benchmark yet

B300 288GB

288 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
308 tok/ssingle-stream, theoretical
Hardware list price
$53,653–$64,383Cloud rental (live): $9.94–$9.94/hr (2026-10-05)
Hardware market price
$1,788,429–$2,086,500
Typical use
Enterprise-grade ultra-large-scale training and high-concurrency inference clusters

GB200 192GB (single card equivalent)

192 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
1
Est. decode speed
308 tok/ssingle-stream, theoretical
Hardware list price
$29,831–$44,746
Hardware market price
$89,421–$178,843
Typical use
Ultra-large-scale trillion-parameter model training and inference clusters (networked with Grace CPU)

B200 180GB

180 GB · 8,000 GB/s · Top-tier Blackwell card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
2
Est. decode speed
615 tok/ssingle-stream, theoretical
Hardware list price
n/aCloud rental (live) ×2: $12.50–$18.77/hr (2026-10-06)
Hardware market price
$111,777–$130,406
Typical use
Enterprise-grade large-scale online inference and training

H200 141GB

141 GB · 4,800 GB/s · Top-tier export-controlled card (channels to China are volatile)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
2
Est. decode speed
369 tok/ssingle-stream, theoretical
Hardware list price
$60,091–$96,575Cloud rental (live) ×2: $8.43–$10.24/hr (2026-10-06)
Hardware market price
$55,888–$58,124
Typical use
Enterprise-grade production environment high-concurrency inference

H100 SXM5 80GB

80 GB · 3,350 GB/s · Top-tier export-controlled card (no official channels to China)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
515 tok/ssingle-stream, theoretical
Hardware list price
$150,228–$171,689Cloud rental (live) ×4: $7.16–$16.00/hr (2026-10-06)
Hardware market price
$201,198–$208,650
Typical use
Enterprise-grade production environment high-concurrency inference/training

H800 SXM 80GB

80 GB · 3,350 GB/s · Top-tier export-controlled card (discontinued)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
515 tok/ssingle-stream, theoretical
Hardware list price
$146,055–$154,997
Hardware market price
$201,198–$208,650
Typical use
Enterprise-grade production deployment (original China-compliant model, now discontinued)

H20 96GB

96 GB · 4,000 GB/s · Export-compliant special supply card (significantly reduced compute power)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
2
Est. decode speed
308 tok/ssingle-stream, theoretical
Hardware list price
n/a
Hardware market price
$23,846–$32,788
Typical use
Suboptimal choice in export-compliant environments, medium-to-large scale inference deployment (FP16 compute power is only about 15% of H100, but higher memory bandwidth offers an advantage in decode-intensive scenarios)

A800 PCIe 80GB

80 GB · 1,935 GB/s · Compliant special supply card (export compliance status adjusted multiple times)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
298 tok/ssingle-stream, theoretical
Hardware list price
$59,614–$61,999
Hardware market price
$41,730–$50,672
Typical use
Medium-scale inference/training deployment in export-compliant environments

Ascend 910B 64GB

64 GB · 400 GB/s · Domestic alternative (not subject to US export controls)
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
62 tok/ssingle-stream, theoretical
Hardware list price
n/a
Hardware market price
$65,576–$71,537
Typical use
Domestic alternative, not subject to US export controls, medium-to-large scale inference/training deployment, memory bandwidth significantly lower than NVIDIA cards of similar VRAM capacity

L40S 48GB

48 GB · 864 GB/s · Compliant inference card
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
133 tok/ssingle-stream, theoretical
Hardware list price
$26,826–$38,749Cloud rental (live) ×4: $2.40–$3.26/hr (2026-10-06)
Hardware market price
$26,826–$38,749
Typical use
SME teams building inference services, medium-low concurrency

L20 48GB

48 GB · 864 GB/s · Compliant inference card
Suggested precision
FP16/BF16 (full precision)
GPUs needed
4
Est. decode speed
133 tok/ssingle-stream, theoretical
Hardware list price
$14,307–$19,077
Hardware market price
$14,904–$20,865
Typical use
SME teams building inference services, medium-low concurrency

RTX 5090D 24GB

24 GB · 1,344 GB/s · Consumer-grade/small-scale self-built
Suggested precision
FP16/BF16 (full precision)
GPUs needed
8
Est. decode speed
414 tok/ssingle-stream, theoretical
Hardware list price
$24,442–$25,037Cloud rental (live) ×8: $3.43–$4.29/hr (2026-10-06)
Hardware market price
$35,769–$47,691
Typical use
Personal development/small-scale verification/POC, not recommended for production-level high concurrency

RTX 4090D 24GB

24 GB · 1,008 GB/s · Consumer-grade/small-scale self-built
Suggested precision
FP16/BF16 (full precision)
GPUs needed
8
Est. decode speed
310 tok/ssingle-stream, theoretical
Hardware list price
$15,499–$22,652Cloud rental (live) ×8: $2.48–$3.46/hr (2026-10-06)
Hardware market price
$30,999–$35,769
Typical use
Personal development/small-scale verification/POC, not recommended for production-level high concurrency
Decode speed is a theoretical ceiling, not a benchmark

Estimated speed = (memory bandwidth × GPU count) ÷ (active parameters × bytes per parameter): the memory-bound upper limit for one request (roofline). It scales linearly with GPU count and ignores interconnect overhead, batching and framework losses, so real results are usually well below it. The linear assumption is reasonable for NVLink data-center cards (H/B series) but too optimistic for consumer PCIe cards without NVLink (RTX 5090D/4090D); multi-GPU consumer setups lose far more to communication, so equal-looking numbers do not mean equal performance. “Suggested precision” is the highest precision that fits within 8 GPUs; beyond that it is flagged as multi-node. Prices are USD converted at ¥6.71 per USD; cloud rental is the international vast.ai market rate, not a China purchase price.