decision-modernbert-base (Core ML)

A small typed-decision model for Apple devices: give it a state (text or JSON), a question (choice, noul yes/no, or score) and its options, and it returns one calibrated probability per option in a few milliseconds, on device. Fine-tuned from ModernBERT-base (149.6M parameters) and compiled to fixed-shape fp16 Core ML programs.

License: CC BY-NC 4.0 โ€” non-commercial use only. The fine-tuning corpus includes sources licensed for non-commercial research only (see Training data), so these weights are released for research and personal use. They are not part of FluidInference's Apache-2.0 model set. The small Python runtime in dmodel_mac/ may be used under Apache-2.0.

Results

Apple M5 Pro, macOS 27.0.

Decision Index 0.2 20.12 (raw 39.74); every one of 151,034 scored requests answered
Areas (knowledge / language / retrieval / tools / arts) 10.1 / 20.6 / 36.6 / 23.4 / 9.8
Median latency per question (end to end, warm) 5.6 ms (p95 34 ms)
Calibration error (ECE, held-out dev, temperature 1.265) 0.012
Neural Engine placement 891 of 896 ops
Core ML vs PyTorch fp32 (500 questions) same answer on 499; max probability difference 0.015

The Index score is a self-run with the public decision-index kit (19ad28ec) on the hash-verified 0.2 suite; it is not an official leaderboard entry. For reference, laya (421M), which inspired this model, is listed at 5.51.

Strongest benchmarks (chance-corrected skill): BANKING77 81, CLINC150 81, GSM8K 72, HoVer 57, ContractNLI 51, BFCL 50. It is at chance on 11 of 40, mostly reasoning- and knowledge-heavy sets (ANLI, NLI4CT, CLadder, HLE, CRUXEval, SATA-Bench, ForecastBench). Good at routing, intent and tool selection; not a general reasoner.

Usage

pip install coremltools tokenizers numpy
hf download FluidInference/decision-modernbert-base-coreml --local-dir decision-modernbert-base
cd decision-modernbert-base
from dmodel_mac.engine import CoreMLDecisionEngine

engine = CoreMLDecisionEngine(config="engine.json")   # run from the repo folder; paths are relative
response, raw = engine(
    "I was charged twice for one order and do not recognise the second charge.",
    {"route": {"type": "choice", "instructions": "Which team should handle this?",
               "criteria": {"billing": "Charges, refunds, invoices",
                            "shipping": "Delivery and tracking",
                            "account": "Login and profile"}}})
print(response["answers"]["route"])   # choice 'billing', p โ‰ˆ 0.87

engine(state, questions) follows the Decision Index engine contract: several questions per request, every option gets a probability, and noul answers return {"noul": p_yes}.

How it works

  • One sequence per window: [CLS] <type> <question> [SEP] ([MASK] <option>)* [SEP] <state slice> [SEP]. Each option is scored at its own [MASK] marker.
  • Nothing is truncated. Long questions keep their head and tail, long option lists are split across windows, and the state is scanned in 25%-overlapping slices; an option's score is the log-mean-exp over every window it appears in (dmodel_mac/render.py). Up to 255 options per question.
  • Three Core ML buckets (128, 256, 512 tokens ร— 64 option slots). engine.json runs 128 on CPU+Neural Engine and 256/512 on all units, the fastest measured placement; softmax uses the calibrated temperature.
File
coreml/dmodel_base_L{128,256,512}_K64.mlpackage inputs input_ids, attention_mask int32 [1,L], marker_map fp16 [1,64,L] โ†’ logits [1,64]
engine.json buckets, compute units, temperature
tokenizer.json, config.json ModernBERT-base tokenizer and config
dmodel_mac/ reference renderer and engine (Python)

Training data

About 123k rows across 52 tasks, deduplicated and decontaminated against all 155,390 Decision Index 0.2 rows (exact lines and 13-gram shingles). Gold labels only; no teacher model. Sources include BANKING77, CLINC150, MASSIVE, MultiNLI, ANLI, FEVER, HoVer, BoolQ, CommonsenseQA, SciQ, QASC, ARC, MMLU, MedMCQA, GSM8K, MATH, WinoGrande, HellaSwag, ContractNLI, NLI4CT, VAST, iSarcasmEval, Humicroedit, New Yorker captions, ACOS, Amazon ESCI, MS MARCO, RAGTruth, Enron spam, Twitter financial sentiment, Lichess puzzles, Glaive function calling, ToolACE, When2Call, HelpSteer 2/3 and UltraFeedback, plus programmatically verified rule tasks. Several of these (for example ANLI and SciQ under CC BY-NC, MS MARCO's research-only terms) restrict commercial use, hence this model's license. Upstream train splits of some Index families (ANLI, WinoGrande, HellaSwag, ContractNLI, VAST, NLI4CT, iSarcasmEval, Humicroedit, RAGTruth, HoVer, ESCI, New Yorker, When2Call, BANKING77, CLINC150) are in the corpus; their test rows are not.

Full fine-tune on Apple MPS: 2 epochs, AdamW 5e-5, checkpoint chosen by dev macro NLL (dev macro accuracy 67.6%), one temperature fitted on a separate calibration split.

Limitations

  • English only. Weak at multi-step reasoning, knowledge-heavy questions and fine detail checks.
  • Long inputs cost several windows: p95 latency is about 6ร— the median.
  • Probabilities are calibrated on the training distribution; recalibrate on your own data before relying on them for new kinds of questions.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for FluidInference/decision-modernbert-base-coreml

Quantized
(73)
this model