Qwen3.6-35B-A3B — OpenVINO INT4
An OpenVINO INT4 weight-only conversion of
Qwen/Qwen3.6-35B-A3B — a vision-language MoE
(hybrid linear-attention/SSM + full-attention layers, 256-expert mixture-of-experts, ~3 B active
parameters per token). Produced with optimum-intel
- NNCF for efficient CPU (and Intel GPU) inference via
openvino-genai.
Created by L. Christensen
Released under the Apache License 2.0, inherited from the base model
Qwen/Qwen3.6-35B-A3B (whose LICENSE is Apache-2.0).
This is a derivative quantization only; all model rights and terms follow the base model.
Files
Multi-component OpenVINO IR (~18.3 GB total), as emitted by the exporter:
| Component | Precision | Notes |
|---|---|---|
openvino_language_model.{xml,bin} |
INT4 | the decoder (~17.4 GB) — the bulk |
openvino_text_embeddings_model.{xml,bin} |
INT8 | |
openvino_vision_embeddings_model.{xml,bin} |
INT8 | vision tower |
openvino_vision_embeddings_merger_model.{xml,bin} |
INT8 | |
openvino_vision_embeddings_pos_model.{xml,bin} |
INT8 | |
openvino_tokenizer.{xml,bin}, openvino_detokenizer.{xml,bin} |
— | OpenVINO tokenizer IR |
config.json, openvino_config.json, generation_config.json, chat_template.jinja, tokenizer* |
— |
How the quantization was produced
Data-free, weight-only quantization — no calibration dataset, no GPTQ/AWQ, no scale estimation;
plain NNCF round-to-nearest weight compression as produced by optimum-cli ... --weight-format int4.
Exact per-module scheme (from openvino_config.json, dtype: "int4_int8_int8"):
- Language model — INT4 asymmetric,
group_size = 64,ratio = 1.0(all eligible weights to INT4), with an INT8-symmetric backup precision for the layers NNCF keeps at higher precision. - Text- and vision-embedding / merger modules — INT8 symmetric, weight-only.
Reproduce
# Toolchain (the transformers pin matters — see below)
pip install "transformers==5.2.0" "optimum-intel[openvino]" openvino-genai nncf
optimum-cli export openvino \
--model Qwen/Qwen3.6-35B-A3B \
--task image-text-to-text \
--weight-format int4 \
./qwen3.6-35b-a3b-int4-ov
Versions used for this conversion:
| Package | Version |
|---|---|
| OpenVINO | 2026.2.0 |
| openvino-genai | 2026.2.0.0 |
| NNCF | 3.2.0 |
| optimum / optimum-intel | 2.2.0 / 2.0.0 |
| transformers | 5.2.0 |
| Python | 3.14 |
Two gotchas worth knowing
transformers==5.2.0is required for the export. Earlier releases don't recognize theqwen3_5_moearchitecture (AutoConfigraises "does not recognize this architecture"); some later releases removed an internal symbol (Qwen3_5DynamicCache) that optimum-intel imports while building the export patcher. 5.2.0 is the version that satisfies both. (Inference viaopenvino-genaidoes not depend on this — the pin is only for the export.)--task image-text-to-textis required (it's a VLM). optimum-intel's OpenVINO exporter does not registertext-generation-with-pastforqwen3_5_moe.
Conversion resource note
Converting (not running) a 35 B model is RAM-heavy: the exporter loads the full bf16 weights and
upcasts during tracing, so peak host memory is far larger than the 18 GB INT4 output (on the order of
~200 GB for this model). On a machine with limited RAM, configure a large pagefile/swap for the
conversion. Running the resulting INT4 IR needs far less (30 GB resident on CPU).
Usage
import openvino_genai as ov_genai
pipe = ov_genai.VLMPipeline("qwen3.6-35b-a3b-int4-ov", "CPU") # or "GPU" / "AUTO"
cfg = ov_genai.GenerationConfig()
cfg.max_new_tokens = 2048
print(pipe.generate("Explain Rayleigh scattering in one sentence.", generation_config=cfg))
- Validated on CPU at
12–13 tok/s on a 24-core desktop CPU — fast for a 35 B because only ~3 B parameters are active per token (A3B) plus the hybrid linear-attention design. CPU keeps the INT4 weights compressed in memory (30 GB resident). - GPU note: Intel GPU inference depends on the OpenVINO build's compressed-weight support for this architecture's SSM/MoE operators. If the GPU path decompresses INT4 weights to FP16 at compile time, memory use rises sharply — verify the resident footprint on your hardware before relying on it.
Reasoning / thinking output
This is a reasoning model. The chat template opens a <think> block, so the generation is
…thinking…</think>…answer…. Split the output at </think> to separate the chain-of-thought
(reasoning_content) from the final answer (content). Keep thinking enabled (do not pass
enable_thinking=False) and allow a generous max_new_tokens — it commonly thinks 1000+ tokens
before answering.
Acknowledgements
Base model © the Qwen team — Qwen/Qwen3.6-35B-A3B.
Quantization tooling: OpenVINO, NNCF, and optimum-intel.
- Downloads last month
- 10
Model tree for jarvis-pet/Qwen3.6-35B-A3B-int4-ov
Base model
Qwen/Qwen3.6-35B-A3B