GLiNER2.5 Multi β€” Core AI (iOS 27+ / macOS 27+)

Core AI port of fastino/gliner2.5-multi-v1, the multilingual GLiNER2.5 boundary extractor. It is built on mDeBERTa-v3-base and has 287M parameters. One on-device model covers:

  • Entities
  • Text classification (single- and multi-label)
  • Relations
  • Structured records
  • Span attributes

It keeps the upstream decision behavior. All credit for the model goes to Fastino and the GLiNER2 authors. The original model card is included as model-card.md.

Files

File Purpose
GLiNER25-Multi-FP16.aimodel/ Core AI source asset, FP16 (550 MB). Sixteen functions share one set of weights
tokenizer.json, tokenizer_config.json Checkpoint tokenizer: SentencePiece Unigram, 250k vocabulary, GLiNER2 special tokens
constants.json Weights and settings used by the decoder: RecordHead.null_embed and the checkpoint's boundary_head settings
functions.json Input and output names, shapes and dtypes of every function
model-card.md Upstream model card (fastino/gliner2.5-multi-v1, Apache-2.0)
LICENSE Apache License 2.0

Source:

  • Checkpoint fastino/gliner2.5-multi-v1 at revision 12fc40399dae672ce840c5e3c50a92340bff3c8c.
  • Upstream code: fastino-ai/GLiNER2 at commit 55656fb (for preprocessing and decoding).

Export:

  • coreai-torch 0.4.3, coreai-core 1.0.0b3, Torch 2.11.0.
  • Learned weights are unchanged (FP16). The export only rewrites graph structure:
    • Constant DeBERTa relative-position buckets. The float log-bucket path produces wrong indices inside converted graphs.
    • A gather-free relative shift.
    • Fused SDPA.
    • Prefix-sum span features computed as span-membership matmuls.

How the model is split

GLiNER2.5's boundary architecture has data-dependent steps between its neural parts, so the model is exported as four stages. The host runs the logic between them. For each sequence bucket T ∈ {128, 256, 512, 1024} tokens, the functions are a_t<T>, b_t<T>, c_t<T> and d_t<T>. The encoder uses only relative positions, so the 1024 bucket runs the same weights:

Function Computes Main inputs β†’ outputs
a_t<T> mDeBERTa encoder, word and marker gathers, classifier, boundary encoder, start/end/inside marginals, candidate-pool projections, null and count heads, record field projection input_ids, attention_mask, word_idx/word_mask, query_idx/query_mask, cls_idx β†’ states and logits
b_t<T> Shared-pool pair scorer and record-head projections stage-A outputs + 192 pool spans (pool_idx, pool_mask, pool_prior) β†’ pair_logits [1,192,32], record projections
c_t<T> Sparse relation scorer text_states + up to 8 relation types and 512 pairs β†’ relation_logits
d_t<T> Explicit-span scorer, used by entity attributes and choices= fields stage-A outputs + up to 64 spans β†’ explicit_logits [1,32,64]

Stage-A outputs are passed unchanged into B, C and D. Exact shapes are in functions.json.

Static capacity per call: T tokens, 32 extractive queries (entity types, structure fields, relation roles), 32 classification labels in total, 8 relation types.

Host responsibilities:

  • Tokenization and GLiNER2 prompt layout.
  • Candidate-pool selection: top-32 start and end boundaries, per-query quota, deduplicated to 192.
  • Relation pair generation.
  • Every decoding step (thresholds, overlap resolution, records, attributes, formatting).

These must follow upstream exactly. A Swift implementation of all of them was validated against upstream: tokenizer, preprocessing, decoding and long-document chunking. It lives in the companion repository.

Required: GPU placement

import CoreAI
let model = try await AIModel(
  contentsOf: url, options: SpecializationOptions(preferredComputeUnitKind: .gpu))

This graph family was validated with GPU placement. On iOS 27.2, default placement can build Neural Engine regions that abort.

Validation (iPhone 17 Pro, iOS 27.2, GPU)

The reference is upstream GLiNER2 in PyTorch FP32, run through its public extraction API.

Check Result
Stage outputs vs FP32 every output β‰₯ 62 dB PSNR
Determinism identical outputs across repeated runs
54-case corpus, full Swift pipeline 54/54 pass: 53 exact, 1 differs only on FP32 boundary items; max confidence error 0.0026
11 long documents (786–3,755 words, 7 languages, 87 chunks) 11/11 pass: 8 exact, 3 differ only on FP32 boundary items
Mac GPU (same asset) 54/54 and 11/11 exact

Corpus coverage:

  • Tasks: entities, descriptions, thresholds, regex validators, overlap policies, single- and multi-label classification, relations, legacy structures, natural/latent/anchorless records, exclusive fields, choice fields, entity attributes, joint tasks.
  • Ten languages.
  • The 128, 256 and 512 buckets. The 1024 bucket was checked separately: stage A against eager FP32 at 991 tokens (60.9 dB and above on the Mac), and a 991-token clinical note end to end on the iPhone (the same 16 items as upstream, max confidence error 0.0033).

Boundary items: the iPhone GPU's FP16 kernels differ slightly from the Mac's, by up to about 0.03 on boundary logits. Items that sit at a decision boundary can therefore flip on device. "FP32 boundary items" were precomputed at FP32 before any device run:

  • Items whose confidence is within 0.02 of their threshold.
  • Items that change under N(0, 0.01) noise on the stage outputs, which covers the candidate-pool cut and record assignment.

Performance, end to end (preprocessing + model + decoding, steady state):

Input Median
≀128 tokens 13.4 ms
≀256 tokens 19.6 ms
≀512 tokens 39.3 ms
≀1024 tokens (991-token note) 117 ms
18-chunk long document 684 ms

Memory: process footprint is about 360 MB after load and about 460 MB after the first run of all 16 functions (about 200 and 310 MB without the 1024 bucket). The weights are memory-mapped.

Load time: about 4 s the first time (on-device specialization), about 1 s from the system cache.

Limitations

  • A single call handles up to 1024 tokens. Longer texts need chunking. Upstream extract_long splits into overlapping word windows; with this model, windows of 256 words and 48 words of overlap, shrunk to fit 512 tokens, work well and keep every chunk in the faster 512 bucket.
  • Constrained classification (Classifier) and JointIE from the upstream card are separate decoders on top of the same stages; the Swift runtime ports both.
  • Weight compression was evaluated and not adopted:
    • int8 word embeddings (384 MB) kept every decision, but lowered on-device PSNR to 53 dB and raised peak memory, because the GPU keeps a dequantized copy.
    • 4-bit encoder weights changed many decisions.
    • FP16 is the release configuration.

License and attribution

  • Weights: from fastino/gliner2.5-multi-v1 by Fastino, released under the Apache License 2.0. This repository redistributes them converted to Core AI format, unchanged in value apart from the FP16 cast, under the same license (LICENSE).
  • Tokenizer: derived from microsoft/mdeberta-v3-base (MIT), as shipped with the upstream checkpoint.
  • Citation: see the upstream card for the GLiNER2 paper (arXiv:2507.18546).

SHA-256

eaed8f3c9cd893edf2715413d26ab567a747b43c241a32f4757d614e821339ec  GLiNER25-Multi-FP16.aimodel/main.mlirb
c62446df87ae18ec98b133f8f84fc449a07cc89bbf8ef192a4cb5f9c53777a7a  tokenizer.json
1ea719a39e3b712b99a5d902feffc1a19114861305d8ac152f5a995579e39812  constants.json
90494553f42297f0f5a4b07c64cfb3bd43d1ff6f45525ded4a7906bf1ec1cc5a  functions.json
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for smdesai/GLiNER25-Multi-FP16-CoreAI

Finetuned
(8)
this model

Paper for smdesai/GLiNER25-Multi-FP16-CoreAI