Qwen2.5-1.5B-Instruct under parallel constrained decoding: 0.605 over 1,240 typed decisions, with the cost that implies on CPU

#29
by saidutta69 - opened

Qwen2.5-1.5B-Instruct under parallel constrained decoding: 0.605 over 1,240 typed decisions, with the cost that implies on CPU

Sharing a measurement of Qwen/Qwen2.5-1.5B-Instruct (revision 989aa7980e4cf806f80c7fef2b1adb7bc71aa306) driven by parallel constrained decoding, in case the numbers are useful to anyone choosing between constrained decoding and plain generation. Public and reproducible: https://github.com/instax-dutta/sysone-bench, report at results/v2/report-20260926/REPORT.md.

What this is and is not. It is a result about our decoding implementation on top of your base model, not about Qwen2.5-1.5B-Instruct as a general model. We did no fine-tuning, we used the stock instruct weights, and we ran on CPU only. Qwen2.5-1.5B-Instruct is not built for single-pass typed decisions, so nothing here should be read as a general-capability claim about the model.

Setup

  • 1,190 cases and 1,550 typed questions across 9 suites: triage (5-way intent choice plus three nouls), guardrails (noul), moderation (noul), and the public sets agnews (4 labels), emotion (6), banking77 (12 intents), mnli (3-way), sst5 (5-level score), multilingual intent (6 labels, 5 languages).
  • One sealed manifest shared byte-identically with two other models in the same run: raw bytes a938cc2483a592dc84e0d5baac12594491bcaa5b4ceb6b7c3b0def71b36297bd, logical digest 4272a7a25ebcb324235696ad808544e157a31e3f9c04421da731a08dfb9d5768. Seed 42.
  • CPU only, no GPU, run container capped at 4 CPUs and 12 GB. Seeded at 42, one cached model and tokenizer, generation serialized behind a single lock, float32 throughout.
  • The two comparison columns come from the same run on the same manifest: laya 0.3.11, weights commit 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 (default English checkpoint), and jev-1.13.0 over the network. Their full write-up is at https://github.com/NandhaKishorM/laya/issues/555

Results, evaluation split, 952 cases, 1,240 scored decisions

suite n Qwen2.5-1.5B PCD laya 0.3.11 jev 1.13.0
triage 192 0.7812 0.8750 0.9323
guardrails 96 0.2917 0.7604 1.0000
moderation 144 0.7292 0.7569 0.9444
agnews 160 0.9125 0.8500 0.9875
emotion 192 0.4115 0.6562 0.8438
banking77, 12 intents 96 0.5833 0.8125 0.9479
mnli 120 0.3417 0.5583 0.8667
sst5, 5-level score 120 0.5250 0.3333 0.6500
multilingual intent 120 0.6833 0.4500 1.0000
all suites 1,240 0.6048 0.6863 0.9065

Two things stand out and neither is flattering to the PCD setup on this task. It is worst on guardrails at 0.2917, which is two binary safety questions with no label spread at all, and 0.3417 on mnli. Its one clear win over the open-weight alternative is agnews at 0.9125 and multilingual intent at 0.6833, where it beats laya by 6 and 23 points respectively. Note that the laya multilingual figure there is the English checkpoint on non-English text, so that comparison overstates the gap.

The cost is the headline

model p50 p95
jev 1.13.0, over the network 314 ms 387 ms
laya 0.3.11, local CPU 588 ms 1,364 ms
Qwen2.5-1.5B PCD, local CPU 3,948 ms 13,471 ms

Constrained decoding scores every legal child at each transition, and the p95 of 13.5 s is that cost landing in the tail. It is roughly 10x the p50 of the hosted API on the same host and the same network path. On CPU, the accuracy it buys over plain generation would need to be measured before anyone reaches for it, and we have not measured that here.

Calibration is the one place PCD looks good: choice ECE 0.0009 after refitting on the calibration split, against 0.0028 for laya and 0.0009 for Jev, and noul ECE 0.0017 against 0.0022 and 0.0058. Constrained decoding returns usable per-option probabilities, which is the practical argument for it.

Caveats that travel with the numbers

The labels are the weak part. One human reviewer read all 1,190 cases and corrected an AI draft of every answer, protocol human-reviewed-ai-assisted-v1, with no second independent reviewer and no adjudication. So there is no inter-annotator agreement, no kappa and no adjudication artifact for this dataset, and provenance.json records "independent_human_review": false and "adjudication": false. Do not compare these against outcome-graded or multi-reviewer evals. The Jev latency is a network round trip from one host on one day.

Reproducing

python -m benchmark.report --laya <run> --jev <run> --qwen <run> \
  --manifest datasets/v2/manifest.jsonl --output-root <new-report-dir>
python -m benchmark.graphics --figure-source <new-report-dir>/figures/figure-source.json \
  --output-root <new-report-dir>/figures

benchmark.report refuses to emit unless its recomputed per-decision counts and per-suite accuracies reproduce each run's own sealed summary.json, so the published numbers cannot drift from the raw predictions. Run ID: qwen-assisted-20260925.

If someone here has a GPU measurement of the same setup, I would happily take it: the latency column above is the weakest part of the comparison and we could only measure it on CPU.

Sign up or log in to comment