OpenThai 2.0.1 โ INT8 W8A8 (compressed-tensors, vLLM)
INT8 weight + activation quantisation of
iapp/openthai2.0-qwen3.8-27b (v2.0.1)
via llm-compressor GPTQ, text-calibrated (ultrachat, 512ร2048). Vision tower, MTP draft head
and lm_head are kept in bf16. See the base model card for benchmarks and prompts.
2026-08-31 fix: MTP draft head restored
Earlier uploads of this repo had no mtp.* tensors while config.json still advertised
mtp_num_hidden_layers: 1. With --speculative-config vLLM ran a draft head with no weights:
every draft was rejected (mean acceptance length 1.00) and decoding was ~1.6ร slower than
without the flag. Found by Dr. Panutat Tejasen (ThaiEval-v3). Cause: the quantisation export
loaded the model through a transformers class that has no MTP module, so the head was dropped
on load and the re:.*mtp.* ignore rule matched nothing.
This upload adds model-mtp.safetensors (bf16 head, 0.85 GB), maps it in the index, and adds
re:.*mtp.* to quantization_config.ignore. Verified under vLLM 0.26 with
{"method":"qwen3_5_mtp","num_speculative_tokens":2}: mean acceptance length 1.63 (bf16
production: 1.61). If your copy predates this, re-download or serve without
--speculative-config.
vllm serve iapp/openthai2.0-qwen3.8-27b-INT8-W8A8 \
--max-model-len 32768 --reasoning-parser qwen3 --trust-remote-code \
--speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":2}'
- Downloads last month
- 1,401