Question regarding active parameters, VRAM requirements

#6
by LolerPanda - opened

Hi XingChen-AGI team,

Thanks for sharing this model! I have a few technical questions regarding deployment and architecture:

  1. What is the recommended GPU VRAM size for running FP16 / BF16 inference?
  2. Are there any specific recommendations or flags needed when using vLLM / SGLang for optimal throughput?

Thanks in advance for your help!

LolerPanda changed discussion title from Question regarding active parameters, VRAM requirements, and context window size to Question regarding active parameters, VRAM requirements
XingChen-AGI org

Hi XingChen-AGI team,

Thanks for sharing this model! I have a few technical questions regarding deployment and architecture:

  1. What is the recommended GPU VRAM size for running FP16 / BF16 inference?
  2. Are there any specific recommendations or flags needed when using vLLM / SGLang for optimal throughput?

Thanks in advance for your help!

Since our model supports the mtp method, the basic usage method can be referred to at the following link:https://github.com/XingChen-AGI/Xing4.0-29B-A4B/blob/main/README.md

Sign up or log in to comment