decision-modernbert-base (Core ML)
A small typed-decision model for Apple devices: give it a state (text or JSON), a question (choice,
noul yes/no, or score) and its options, and it returns one calibrated probability per option in a few
milliseconds, on device. Fine-tuned from ModernBERT-base
(149.6M parameters) and compiled to fixed-shape fp16 Core ML programs.
License: CC BY-NC 4.0 โ non-commercial use only. The fine-tuning corpus includes sources licensed for
non-commercial research only (see Training data), so these weights are released for research
and personal use. They are not part of FluidInference's Apache-2.0 model set. The small Python runtime in
dmodel_mac/ may be used under Apache-2.0.
Results
Apple M5 Pro, macOS 27.0.
| Decision Index 0.2 | 20.12 (raw 39.74); every one of 151,034 scored requests answered |
| Areas (knowledge / language / retrieval / tools / arts) | 10.1 / 20.6 / 36.6 / 23.4 / 9.8 |
| Median latency per question (end to end, warm) | 5.6 ms (p95 34 ms) |
| Calibration error (ECE, held-out dev, temperature 1.265) | 0.012 |
| Neural Engine placement | 891 of 896 ops |
| Core ML vs PyTorch fp32 (500 questions) | same answer on 499; max probability difference 0.015 |
The Index score is a self-run with the public decision-index kit (19ad28ec) on the hash-verified 0.2 suite; it
is not an official leaderboard entry. For reference, laya
(421M), which inspired this model, is listed at 5.51.
Strongest benchmarks (chance-corrected skill): BANKING77 81, CLINC150 81, GSM8K 72, HoVer 57, ContractNLI 51, BFCL 50. It is at chance on 11 of 40, mostly reasoning- and knowledge-heavy sets (ANLI, NLI4CT, CLadder, HLE, CRUXEval, SATA-Bench, ForecastBench). Good at routing, intent and tool selection; not a general reasoner.
Usage
pip install coremltools tokenizers numpy
hf download FluidInference/decision-modernbert-base-coreml --local-dir decision-modernbert-base
cd decision-modernbert-base
from dmodel_mac.engine import CoreMLDecisionEngine
engine = CoreMLDecisionEngine(config="engine.json") # run from the repo folder; paths are relative
response, raw = engine(
"I was charged twice for one order and do not recognise the second charge.",
{"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices",
"shipping": "Delivery and tracking",
"account": "Login and profile"}}})
print(response["answers"]["route"]) # choice 'billing', p โ 0.87
engine(state, questions) follows the Decision Index engine contract: several questions per request, every
option gets a probability, and noul answers return {"noul": p_yes}.
How it works
- One sequence per window:
[CLS] <type> <question> [SEP] ([MASK] <option>)* [SEP] <state slice> [SEP]. Each option is scored at its own[MASK]marker. - Nothing is truncated. Long questions keep their head and tail, long option lists are split across windows,
and the state is scanned in 25%-overlapping slices; an option's score is the log-mean-exp over every window it
appears in (
dmodel_mac/render.py). Up to 255 options per question. - Three Core ML buckets (128, 256, 512 tokens ร 64 option slots).
engine.jsonruns 128 on CPU+Neural Engine and 256/512 on all units, the fastest measured placement; softmax uses the calibrated temperature.
| File | |
|---|---|
coreml/dmodel_base_L{128,256,512}_K64.mlpackage |
inputs input_ids, attention_mask int32 [1,L], marker_map fp16 [1,64,L] โ logits [1,64] |
engine.json |
buckets, compute units, temperature |
tokenizer.json, config.json |
ModernBERT-base tokenizer and config |
dmodel_mac/ |
reference renderer and engine (Python) |
Training data
About 123k rows across 52 tasks, deduplicated and decontaminated against all 155,390 Decision Index 0.2 rows (exact lines and 13-gram shingles). Gold labels only; no teacher model. Sources include BANKING77, CLINC150, MASSIVE, MultiNLI, ANLI, FEVER, HoVer, BoolQ, CommonsenseQA, SciQ, QASC, ARC, MMLU, MedMCQA, GSM8K, MATH, WinoGrande, HellaSwag, ContractNLI, NLI4CT, VAST, iSarcasmEval, Humicroedit, New Yorker captions, ACOS, Amazon ESCI, MS MARCO, RAGTruth, Enron spam, Twitter financial sentiment, Lichess puzzles, Glaive function calling, ToolACE, When2Call, HelpSteer 2/3 and UltraFeedback, plus programmatically verified rule tasks. Several of these (for example ANLI and SciQ under CC BY-NC, MS MARCO's research-only terms) restrict commercial use, hence this model's license. Upstream train splits of some Index families (ANLI, WinoGrande, HellaSwag, ContractNLI, VAST, NLI4CT, iSarcasmEval, Humicroedit, RAGTruth, HoVer, ESCI, New Yorker, When2Call, BANKING77, CLINC150) are in the corpus; their test rows are not.
Full fine-tune on Apple MPS: 2 epochs, AdamW 5e-5, checkpoint chosen by dev macro NLL (dev macro accuracy 67.6%), one temperature fitted on a separate calibration split.
Limitations
- English only. Weak at multi-step reasoning, knowledge-heavy questions and fine detail checks.
- Long inputs cost several windows: p95 latency is about 6ร the median.
- Probabilities are calibrated on the training distribution; recalibrate on your own data before relying on them for new kinds of questions.
- Downloads last month
- -
Model tree for FluidInference/decision-modernbert-base-coreml
Base model
answerdotai/ModernBERT-base