Instructions to use smdesai/GLiNER25-Multi-FP16-CoreAI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use smdesai/GLiNER25-Multi-FP16-CoreAI with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("smdesai/GLiNER25-Multi-FP16-CoreAI") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5 Multi β Core AI (iOS 27+ / macOS 27+)
Core AI port of fastino/gliner2.5-multi-v1, the
multilingual GLiNER2.5 boundary extractor. It is built on mDeBERTa-v3-base and has 287M parameters.
One on-device model covers:
- Entities
- Text classification (single- and multi-label)
- Relations
- Structured records
- Span attributes
It keeps the upstream decision behavior. All credit for the model goes to Fastino
and the GLiNER2 authors. The original model card is included as
model-card.md.
Files
| File | Purpose |
|---|---|
GLiNER25-Multi-FP16.aimodel/ |
Core AI source asset, FP16 (550 MB). Sixteen functions share one set of weights |
tokenizer.json, tokenizer_config.json |
Checkpoint tokenizer: SentencePiece Unigram, 250k vocabulary, GLiNER2 special tokens |
constants.json |
Weights and settings used by the decoder: RecordHead.null_embed and the checkpoint's boundary_head settings |
functions.json |
Input and output names, shapes and dtypes of every function |
model-card.md |
Upstream model card (fastino/gliner2.5-multi-v1, Apache-2.0) |
LICENSE |
Apache License 2.0 |
Source:
- Checkpoint
fastino/gliner2.5-multi-v1at revision12fc40399dae672ce840c5e3c50a92340bff3c8c. - Upstream code: fastino-ai/GLiNER2 at commit
55656fb(for preprocessing and decoding).
Export:
- coreai-torch 0.4.3, coreai-core 1.0.0b3, Torch 2.11.0.
- Learned weights are unchanged (FP16). The export only rewrites graph structure:
- Constant DeBERTa relative-position buckets. The float log-bucket path produces wrong indices inside converted graphs.
- A gather-free relative shift.
- Fused SDPA.
- Prefix-sum span features computed as span-membership matmuls.
How the model is split
GLiNER2.5's boundary architecture has data-dependent steps between its neural parts, so the model
is exported as four stages. The host runs the logic between them. For each sequence bucket
T β {128, 256, 512, 1024} tokens, the functions are a_t<T>, b_t<T>, c_t<T> and d_t<T>. The
encoder uses only relative positions, so the 1024 bucket runs the same weights:
| Function | Computes | Main inputs β outputs |
|---|---|---|
a_t<T> |
mDeBERTa encoder, word and marker gathers, classifier, boundary encoder, start/end/inside marginals, candidate-pool projections, null and count heads, record field projection | input_ids, attention_mask, word_idx/word_mask, query_idx/query_mask, cls_idx β states and logits |
b_t<T> |
Shared-pool pair scorer and record-head projections | stage-A outputs + 192 pool spans (pool_idx, pool_mask, pool_prior) β pair_logits [1,192,32], record projections |
c_t<T> |
Sparse relation scorer | text_states + up to 8 relation types and 512 pairs β relation_logits |
d_t<T> |
Explicit-span scorer, used by entity attributes and choices= fields |
stage-A outputs + up to 64 spans β explicit_logits [1,32,64] |
Stage-A outputs are passed unchanged into B, C and D. Exact shapes are in functions.json.
Static capacity per call: T tokens, 32 extractive queries (entity types, structure fields, relation roles), 32 classification labels in total, 8 relation types.
Host responsibilities:
- Tokenization and GLiNER2 prompt layout.
- Candidate-pool selection: top-32 start and end boundaries, per-query quota, deduplicated to 192.
- Relation pair generation.
- Every decoding step (thresholds, overlap resolution, records, attributes, formatting).
These must follow upstream exactly. A Swift implementation of all of them was validated against upstream: tokenizer, preprocessing, decoding and long-document chunking. It lives in the companion repository.
Required: GPU placement
import CoreAI
let model = try await AIModel(
contentsOf: url, options: SpecializationOptions(preferredComputeUnitKind: .gpu))
This graph family was validated with GPU placement. On iOS 27.2, default placement can build Neural Engine regions that abort.
Validation (iPhone 17 Pro, iOS 27.2, GPU)
The reference is upstream GLiNER2 in PyTorch FP32, run through its public extraction API.
| Check | Result |
|---|---|
| Stage outputs vs FP32 | every output β₯ 62 dB PSNR |
| Determinism | identical outputs across repeated runs |
| 54-case corpus, full Swift pipeline | 54/54 pass: 53 exact, 1 differs only on FP32 boundary items; max confidence error 0.0026 |
| 11 long documents (786β3,755 words, 7 languages, 87 chunks) | 11/11 pass: 8 exact, 3 differ only on FP32 boundary items |
| Mac GPU (same asset) | 54/54 and 11/11 exact |
Corpus coverage:
- Tasks: entities, descriptions, thresholds, regex validators, overlap policies, single- and multi-label classification, relations, legacy structures, natural/latent/anchorless records, exclusive fields, choice fields, entity attributes, joint tasks.
- Ten languages.
- The 128, 256 and 512 buckets. The 1024 bucket was checked separately: stage A against eager FP32 at 991 tokens (60.9 dB and above on the Mac), and a 991-token clinical note end to end on the iPhone (the same 16 items as upstream, max confidence error 0.0033).
Boundary items: the iPhone GPU's FP16 kernels differ slightly from the Mac's, by up to about 0.03 on boundary logits. Items that sit at a decision boundary can therefore flip on device. "FP32 boundary items" were precomputed at FP32 before any device run:
- Items whose confidence is within 0.02 of their threshold.
- Items that change under N(0, 0.01) noise on the stage outputs, which covers the candidate-pool cut and record assignment.
Performance, end to end (preprocessing + model + decoding, steady state):
| Input | Median |
|---|---|
| β€128 tokens | 13.4 ms |
| β€256 tokens | 19.6 ms |
| β€512 tokens | 39.3 ms |
| β€1024 tokens (991-token note) | 117 ms |
| 18-chunk long document | 684 ms |
Memory: process footprint is about 360 MB after load and about 460 MB after the first run of all 16 functions (about 200 and 310 MB without the 1024 bucket). The weights are memory-mapped.
Load time: about 4 s the first time (on-device specialization), about 1 s from the system cache.
Limitations
- A single call handles up to 1024 tokens. Longer texts need chunking. Upstream
extract_longsplits into overlapping word windows; with this model, windows of 256 words and 48 words of overlap, shrunk to fit 512 tokens, work well and keep every chunk in the faster 512 bucket. - Constrained classification (
Classifier) andJointIEfrom the upstream card are separate decoders on top of the same stages; the Swift runtime ports both. - Weight compression was evaluated and not adopted:
- int8 word embeddings (384 MB) kept every decision, but lowered on-device PSNR to 53 dB and raised peak memory, because the GPU keeps a dequantized copy.
- 4-bit encoder weights changed many decisions.
- FP16 is the release configuration.
License and attribution
- Weights: from
fastino/gliner2.5-multi-v1by Fastino, released under the Apache License 2.0. This repository redistributes them converted to Core AI format, unchanged in value apart from the FP16 cast, under the same license (LICENSE). - Tokenizer: derived from
microsoft/mdeberta-v3-base(MIT), as shipped with the upstream checkpoint. - Citation: see the upstream card for the GLiNER2 paper (arXiv:2507.18546).
SHA-256
eaed8f3c9cd893edf2715413d26ab567a747b43c241a32f4757d614e821339ec GLiNER25-Multi-FP16.aimodel/main.mlirb
c62446df87ae18ec98b133f8f84fc449a07cc89bbf8ef192a4cb5f9c53777a7a tokenizer.json
1ea719a39e3b712b99a5d902feffc1a19114861305d8ac152f5a995579e39812 constants.json
90494553f42297f0f5a4b07c64cfb3bd43d1ff6f45525ded4a7906bf1ec1cc5a functions.json
Model tree for smdesai/GLiNER25-Multi-FP16-CoreAI
Base model
fastino/gliner2.5-multi-v1