Jev_Qwen3.8-27B-NVFP4-FP8

NVFP4 + FP8 mixed-precision quantization of SargeDev/Jev_Qwen3.8-27B (itself a QLoRA merge of huihui-ai/Huihui-Qwen3.8-27B-abliterated trained on SargeDev/jev-distill-corpus-v3).

Quant recipe (non-uniform)

  • Tool: llm-compressor 0.14.0 (vllm-project), QuantizationModifier mixed config groups
  • NVFP4 W4A4 (E2M1, block 16, FP8-E4M3 scales): all attention projections + gate/up_proj
  • FP8_DYNAMIC (per-channel weights, per-token activations): all down_proj — the layer class most sensitive to 4-bit
  • Ignored (kept high-precision): lm_head, embed_tokens, MTP head
  • Format: compressed-tensors mixed-precision (nvfp4-pack-quantized + float-quantized)
  • 28 GB total (vs 51 GB bf16) — loads in vLLM with --quantization compressed-tensors

Calibration (domain-matched)

512 samples, 2048 max seq: 384 held-out jev-distill judge prompts in the exact eval-time format (state + question + options), plus 128 ultrachat_200k general-chat rows to keep broad-language calibration honest. Calibrating on the deployment domain, not generic text.

Verified on DGX Spark (GB10)

  • vLLM 0.28-era serving stack, --quantization compressed-tensors, GB10 sm_121a
  • Structured judge outputs correct post-quant: calibrated JSON distributions intact
  • ~8 tok/s single-stream long-form generation (200+ token completions), judge-style JSON answers land in 2-5 s
  • Chatty persona intact (identifies as Qwen; no behavioral regression observed)

Serve

vllm serve /path/to/Jev_Qwen3.8-27B-NVFP4-FP8 \
  --quantization compressed-tensors \
  --served-model-name jev-nvfp4-fp8 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.80

Trained with thinking disabled (enable_thinking=false); serve with thinking off for tuned behavior.

Apache-2.0. Quantized on NVIDIA DGX Spark (GB10).

Downloads last month
24
Safetensors
Model size
21B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SargeDev/Jev_Qwen3.8-27B-NVFP4-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(2)
this model

Dataset used to train SargeDev/Jev_Qwen3.8-27B-NVFP4-FP8