Instructions to use Qwen/Qwen2.5-1.5B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen2.5-1.5B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen2.5-1.5B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen2.5-1.5B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen2.5-1.5B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2.5-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Qwen/Qwen2.5-1.5B-Instruct
- SGLang
How to use Qwen/Qwen2.5-1.5B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen2.5-1.5B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2.5-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen2.5-1.5B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2.5-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Qwen/Qwen2.5-1.5B-Instruct with Docker Model Runner:
docker model run hf.co/Qwen/Qwen2.5-1.5B-Instruct
Qwen2.5-1.5B-Instruct under parallel constrained decoding: 0.605 over 1,240 typed decisions, with the cost that implies on CPU
Qwen2.5-1.5B-Instruct under parallel constrained decoding: 0.605 over 1,240 typed decisions, with the cost that implies on CPU
Sharing a measurement of Qwen/Qwen2.5-1.5B-Instruct (revision 989aa7980e4cf806f80c7fef2b1adb7bc71aa306) driven by parallel constrained decoding, in case the numbers are useful to anyone choosing between constrained decoding and plain generation. Public and reproducible: https://github.com/instax-dutta/sysone-bench, report at results/v2/report-20260926/REPORT.md.
What this is and is not. It is a result about our decoding implementation on top of your base model, not about Qwen2.5-1.5B-Instruct as a general model. We did no fine-tuning, we used the stock instruct weights, and we ran on CPU only. Qwen2.5-1.5B-Instruct is not built for single-pass typed decisions, so nothing here should be read as a general-capability claim about the model.
Setup
- 1,190 cases and 1,550 typed questions across 9 suites: triage (5-way
intentchoice plus three nouls), guardrails (noul), moderation (noul), and the public sets agnews (4 labels), emotion (6), banking77 (12 intents), mnli (3-way), sst5 (5-levelscore), multilingual intent (6 labels, 5 languages). - One sealed manifest shared byte-identically with two other models in the same run: raw bytes
a938cc2483a592dc84e0d5baac12594491bcaa5b4ceb6b7c3b0def71b36297bd, logical digest4272a7a25ebcb324235696ad808544e157a31e3f9c04421da731a08dfb9d5768. Seed 42. - CPU only, no GPU, run container capped at 4 CPUs and 12 GB. Seeded at 42, one cached model and tokenizer, generation serialized behind a single lock, float32 throughout.
- The two comparison columns come from the same run on the same manifest:
laya0.3.11, weights commit55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851(default English checkpoint), andjev-1.13.0over the network. Their full write-up is at https://github.com/NandhaKishorM/laya/issues/555
Results, evaluation split, 952 cases, 1,240 scored decisions
| suite | n | Qwen2.5-1.5B PCD | laya 0.3.11 | jev 1.13.0 |
|---|---|---|---|---|
| triage | 192 | 0.7812 | 0.8750 | 0.9323 |
| guardrails | 96 | 0.2917 | 0.7604 | 1.0000 |
| moderation | 144 | 0.7292 | 0.7569 | 0.9444 |
| agnews | 160 | 0.9125 | 0.8500 | 0.9875 |
| emotion | 192 | 0.4115 | 0.6562 | 0.8438 |
| banking77, 12 intents | 96 | 0.5833 | 0.8125 | 0.9479 |
| mnli | 120 | 0.3417 | 0.5583 | 0.8667 |
| sst5, 5-level score | 120 | 0.5250 | 0.3333 | 0.6500 |
| multilingual intent | 120 | 0.6833 | 0.4500 | 1.0000 |
| all suites | 1,240 | 0.6048 | 0.6863 | 0.9065 |
Two things stand out and neither is flattering to the PCD setup on this task. It is worst on guardrails at 0.2917, which is two binary safety questions with no label spread at all, and 0.3417 on mnli. Its one clear win over the open-weight alternative is agnews at 0.9125 and multilingual intent at 0.6833, where it beats laya by 6 and 23 points respectively. Note that the laya multilingual figure there is the English checkpoint on non-English text, so that comparison overstates the gap.
The cost is the headline
| model | p50 | p95 |
|---|---|---|
| jev 1.13.0, over the network | 314 ms | 387 ms |
| laya 0.3.11, local CPU | 588 ms | 1,364 ms |
| Qwen2.5-1.5B PCD, local CPU | 3,948 ms | 13,471 ms |
Constrained decoding scores every legal child at each transition, and the p95 of 13.5 s is that cost landing in the tail. It is roughly 10x the p50 of the hosted API on the same host and the same network path. On CPU, the accuracy it buys over plain generation would need to be measured before anyone reaches for it, and we have not measured that here.
Calibration is the one place PCD looks good: choice ECE 0.0009 after refitting on the calibration split, against 0.0028 for laya and 0.0009 for Jev, and noul ECE 0.0017 against 0.0022 and 0.0058. Constrained decoding returns usable per-option probabilities, which is the practical argument for it.
Caveats that travel with the numbers
The labels are the weak part. One human reviewer read all 1,190 cases and corrected an AI draft of every answer, protocol human-reviewed-ai-assisted-v1, with no second independent reviewer and no adjudication. So there is no inter-annotator agreement, no kappa and no adjudication artifact for this dataset, and provenance.json records "independent_human_review": false and "adjudication": false. Do not compare these against outcome-graded or multi-reviewer evals. The Jev latency is a network round trip from one host on one day.
Reproducing
python -m benchmark.report --laya <run> --jev <run> --qwen <run> \
--manifest datasets/v2/manifest.jsonl --output-root <new-report-dir>
python -m benchmark.graphics --figure-source <new-report-dir>/figures/figure-source.json \
--output-root <new-report-dir>/figures
benchmark.report refuses to emit unless its recomputed per-decision counts and per-suite accuracies reproduce each run's own sealed summary.json, so the published numbers cannot drift from the raw predictions. Run ID: qwen-assisted-20260925.
If someone here has a GPU measurement of the same setup, I would happily take it: the latency column above is the weakest part of the comparison and we could only measure it on CPU.