Qwen3.6-35B-A3B — OpenVINO INT4

An OpenVINO INT4 weight-only conversion of Qwen/Qwen3.6-35B-A3B — a vision-language MoE (hybrid linear-attention/SSM + full-attention layers, 256-expert mixture-of-experts, ~3 B active parameters per token). Produced with optimum-intel

Created by L. Christensen

Released under the Apache License 2.0, inherited from the base model Qwen/Qwen3.6-35B-A3B (whose LICENSE is Apache-2.0). This is a derivative quantization only; all model rights and terms follow the base model.

Files

Multi-component OpenVINO IR (~18.3 GB total), as emitted by the exporter:

Component Precision Notes
openvino_language_model.{xml,bin} INT4 the decoder (~17.4 GB) — the bulk
openvino_text_embeddings_model.{xml,bin} INT8
openvino_vision_embeddings_model.{xml,bin} INT8 vision tower
openvino_vision_embeddings_merger_model.{xml,bin} INT8
openvino_vision_embeddings_pos_model.{xml,bin} INT8
openvino_tokenizer.{xml,bin}, openvino_detokenizer.{xml,bin} — OpenVINO tokenizer IR
config.json, openvino_config.json, generation_config.json, chat_template.jinja, tokenizer* —

How the quantization was produced

Data-free, weight-only quantization — no calibration dataset, no GPTQ/AWQ, no scale estimation; plain NNCF round-to-nearest weight compression as produced by optimum-cli ... --weight-format int4. Exact per-module scheme (from openvino_config.json, dtype: "int4_int8_int8"):

  • Language model — INT4 asymmetric, group_size = 64, ratio = 1.0 (all eligible weights to INT4), with an INT8-symmetric backup precision for the layers NNCF keeps at higher precision.
  • Text- and vision-embedding / merger modules — INT8 symmetric, weight-only.

Reproduce

# Toolchain (the transformers pin matters — see below)
pip install "transformers==5.2.0" "optimum-intel[openvino]" openvino-genai nncf

optimum-cli export openvino \
  --model Qwen/Qwen3.6-35B-A3B \
  --task image-text-to-text \
  --weight-format int4 \
  ./qwen3.6-35b-a3b-int4-ov

Versions used for this conversion:

Package Version
OpenVINO 2026.2.0
openvino-genai 2026.2.0.0
NNCF 3.2.0
optimum / optimum-intel 2.2.0 / 2.0.0
transformers 5.2.0
Python 3.14

Two gotchas worth knowing

  • transformers==5.2.0 is required for the export. Earlier releases don't recognize the qwen3_5_moe architecture (AutoConfig raises "does not recognize this architecture"); some later releases removed an internal symbol (Qwen3_5DynamicCache) that optimum-intel imports while building the export patcher. 5.2.0 is the version that satisfies both. (Inference via openvino-genai does not depend on this — the pin is only for the export.)
  • --task image-text-to-text is required (it's a VLM). optimum-intel's OpenVINO exporter does not register text-generation-with-past for qwen3_5_moe.

Conversion resource note

Converting (not running) a 35 B model is RAM-heavy: the exporter loads the full bf16 weights and upcasts during tracing, so peak host memory is far larger than the 18 GB INT4 output (on the order of ~200 GB for this model). On a machine with limited RAM, configure a large pagefile/swap for the conversion. Running the resulting INT4 IR needs far less (30 GB resident on CPU).

Usage

import openvino_genai as ov_genai

pipe = ov_genai.VLMPipeline("qwen3.6-35b-a3b-int4-ov", "CPU")   # or "GPU" / "AUTO"
cfg = ov_genai.GenerationConfig()
cfg.max_new_tokens = 2048
print(pipe.generate("Explain Rayleigh scattering in one sentence.", generation_config=cfg))
  • Validated on CPU at 12–13 tok/s on a 24-core desktop CPU — fast for a 35 B because only ~3 B parameters are active per token (A3B) plus the hybrid linear-attention design. CPU keeps the INT4 weights compressed in memory (30 GB resident).
  • GPU note: Intel GPU inference depends on the OpenVINO build's compressed-weight support for this architecture's SSM/MoE operators. If the GPU path decompresses INT4 weights to FP16 at compile time, memory use rises sharply — verify the resident footprint on your hardware before relying on it.

Reasoning / thinking output

This is a reasoning model. The chat template opens a <think> block, so the generation is …thinking…</think>…answer…. Split the output at </think> to separate the chain-of-thought (reasoning_content) from the final answer (content). Keep thinking enabled (do not pass enable_thinking=False) and allow a generous max_new_tokens — it commonly thinks 1000+ tokens before answering.

Acknowledgements

Base model © the Qwen team — Qwen/Qwen3.6-35B-A3B. Quantization tooling: OpenVINO, NNCF, and optimum-intel.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jarvis-pet/Qwen3.6-35B-A3B-int4-ov

Quantized
(839)
this model