SargeDev/jev-distill-corpus-v3
Viewer • Updated • 741k • 730 • 11
NVFP4 + FP8 mixed-precision quantization of SargeDev/Jev_Qwen3.8-27B (itself a QLoRA merge of huihui-ai/Huihui-Qwen3.8-27B-abliterated trained on SargeDev/jev-distill-corpus-v3).
QuantizationModifier mixed config groupslm_head, embed_tokens, MTP headmixed-precision (nvfp4-pack-quantized + float-quantized)--quantization compressed-tensors512 samples, 2048 max seq: 384 held-out jev-distill judge prompts in the exact eval-time format (state + question + options), plus 128 ultrachat_200k general-chat rows to keep broad-language calibration honest. Calibrating on the deployment domain, not generic text.
--quantization compressed-tensors, GB10 sm_121avllm serve /path/to/Jev_Qwen3.8-27B-NVFP4-FP8 \
--quantization compressed-tensors \
--served-model-name jev-nvfp4-fp8 \
--max-model-len 8192 \
--gpu-memory-utilization 0.80
Trained with thinking disabled (enable_thinking=false); serve with thinking off for tuned behavior.
Apache-2.0. Quantized on NVIDIA DGX Spark (GB10).
Base model
Qwen/Qwen3.8-27B