Minimum number of A100/H100 GPUs to serve Xing4.0-29B-A4B at full 256K context?

#9
by rtudriny - opened

Hi team, could you give a straightforward deployment answer:

  1. How many GPUs do we need to serve this model with the full 256K context ? e.g. is 4×A100/H100 80G (TP=4) enough, or do we need 8×80G (TP=8)?
  2. Same question for the 512K extended context, if anyone has tested it.

The model card says 256K support, but doesn't state the hardware requirement. A simple "N×80G GPUs for 256K" line would help a lot for capacity planning. Thanks!

Standard 256k requires about 11.25GiB of VRAM (ref)
My estimated requirement for 512k context window is about 22.50GiB of VRAM

1x80G GPU might be overkill for a single user

XingChen-AGI org

Good news — a single H100 80G can reach or come close to the full 256K context at one concurrent request. For 512K, you'll need roughly double the KV cache headroom, so a 2×H100 80G setup gives you more comfortable headroom and better cost-efficiency overall.

This matches the estimate above — thanks for the KV cache calculation, that lines up with our internal numbers.

Quick note: actual VRAM usage depends on quantization choice and framework overhead. If you share your expected concurrency and quant plan, I can help you size it more precisely.

Sign up or log in to comment