Access to Intenter

Intenter is released for research, evaluation, internal validation and education. Tell us who you are and we will grant access.

By requesting access you agree to the Elda Community License 1.0: no commercial use, no redistribution of the weights or derivatives, and attribution as "Built with Elda". Commercial licensing is available on request.

Log in or Sign Up to review the conditions and access this model content.

Elda-Intenter — real-time speech-act + entity extraction in a single pass

Korean-first, multilingual (ko · ja · en). 281M parameters · p50 27.8 ms.

One forward pass returns four aligned channels: what the speaker is doing (speech act), what they are talking about (entities), and the attribute/predicate structure binding them. Not a general-purpose NER model — the perception layer of a live conversational system.

"간판의 품격이라는 음식점이 있다며, 그 음식점은 어디에 있냐고"

  intent    QUESTION
  entity    간판의 품격 → ORGANIZATION
            음식점이   → ORGANIZATION
  signals   is_pure_negation=0 · is_claim=0 · polarity=0 · depends_on_prev=1

Results

All sets below are excluded from training by construction.

MASSIVE (dev, re-annotated) — conversational speech, same five peers

Voice-assistant utterances, our actual traffic shape. Same axes, same particle rule; peers emit KLUE's six types, so only the shared axes are comparable and type is omitted.

abstention ↓ span recall boundary precision anchor reach
ko ours (469 sent.) 2.7 97.3 96.5 96.6 81.8
ko — 5 peers 33.7 – 67.9 32.1 – 66.3 69.9 – 86.4 89.4 – 94.7 6.6 – 72.7
ja ours (749 sent.) 2.7 97.3 97.7 95.0 91.3
ja — 5 peers 30.1 – 76.2 23.8 – 69.9 50.8 – 84.0 72.3 – 93.9 18.8 – 48.0
en ours (796 sent.) 2.9 97.1 92.6 82.1 95.2
en — 5 peers 38.6 – 69.4 30.6 – 61.4 12.5 – 63.3 90.7 – 94.8 28.5 – 64.0

We lead every axis on ko and ja, and all but precision on en. Entity F1 under our own scorer: ja 79.7 · ko 85.5 · en 74.5 (one trailing function word allowed).

Not comparable to published MASSIVE scores: that benchmark is intent classification plus slot filling with 55 slot types. We map 21 onto 11 of our entity types; peers are KLUE-NER models run outside their training domain. Zero training overlap on dev is verified by the evaluator.

KLUE-NER (dev · 600 sentences / 1,687 gold), entity-level F1

entity F1 boundary F1
exact match 55.3 57.9
allowing one trailing Korean particle 63.9 66.9

Korean particles attach to the noun. The second row accepts one trailing particle when the start offset matches — spans are never trimmed.

KLUE-NER — the same five peers, on their own training domain

Every peer is a dedicated KLUE-NER model fine-tuned on all 21,008 KLUE-NER training sentences, emitting exactly KLUE's six types. We saw 2,500 of them, emit 17 types projected onto those six, and run four channels plus six auxiliary heads in the same pass.

axis ours KF-DeBERTa KLUE-BERT RoBERTa-large KoELECTRA RoBERTa-base
type accuracy 88.2 98.2 97.7 98.1 95.1 92.2
abstention (lower better) 14.0 2.0 2.4 13.6 18.3 41.1
span recall 86.0 98.0 97.6 86.4 81.7 58.9
boundary 97.2 97.3 96.2 96.2 91.9 62.9
precision 89.6 97.2 97.4 98.0 96.0 96.7
anchor reach 79.2 94.8 94.6 69.8 54.8 13.0

We lose on every axis; boundary is a tie. Third or fourth of six. Note the peer spread — anchor reach runs 13.0 to 94.8, so a single-peer table could have been picked to favour us.

Axes. abstention — nothing emitted at a gold position. span recall — something emitted there. boundary — returned span contains the gold string. precision — emitted spans landing on gold, over spans carrying a KLUE-mappable type (1,427/1,592); over all our spans it is 78.9, since 165 carry types KLUE does not annotate (WORK, ARTIFACTS, TERM, EVENT, DURATION). anchor reach — span returned under a type that opens a downstream lookup (LOCATION / ORGANIZATION; 447 of 1,687 gold here).

Our training mix is 4.0 % news. KLUE-NER is news text; our material is wiki, conversational Korean and synthetic templates. Whether these peers hold up on conversational Korean is untested here.

Reported by others, on their own runs — not our ruler, not verified by us: KLUE-RoBERTa-large 90.8 · XLM-R-large 85.9 · KR-BERT-base 77.2 · mBERT-base 73.2 (KLUE paper). Our re-scoring of a KLUE-NER RoBERTa-large on this split gave 83.4, not 90.8.

Frozen internal gates

Re-run on every build, evaluated per surface, not summed.

Role nouns — detection / typing (405 slots) 401 / 399
Referent expressions, held out 23 · 23 · 22
Conversational-speech detection / safety floor 92.4 % · 12 / 12
Possessive structure — owner kept (40 surfaces) 31 / 40
Discourse signals — false positives on short utterances 1 / 48
Discourse signals — polarity grid 23 / 24
Formal-register interrogatives 13 / 14
Proper-name span probe, third-party frames (80 slots) 78 / 80
Name-span probe, third-party (698 slots) 685 / 698
Span-convention compliance (126 slots) 126 / 126

The two third-party probe rows are produced by a separate team on their own infrastructure.


Why an encoder

Elda-Intenter typical small-LLM extraction
Parameters 281M 500M – 4B
Latency (single, incl. heads) p50 27.8 ms hundreds of ms
Latency (batch of 16) 118.5 ms seconds
Serving precision fp32 (2 workers, 6.1 GB GPU) varies
Output fixed contract, char offsets free text to be parsed

Measured on the production serving path (RTX 3060, fp32, heads attached). Span extraction is classification over token pairs, and a bidirectional encoder sees the whole utterance at once.


Output contract

Four Global Pointer channels over one shared mDeBERTa-v3-base backbone, plus binary signal heads and one span-attribute head. All channels are character spans over the original string.

channel what it carries
intents speech act over the utterance or clause — 18 types incl. AGREE / DISAGREE / QUESTION / COMMAND / NARRATE / DESCRIBE / REQUEST
entities 17 entity types × subtypes, plus about_speaker per span
attributes modifiers bound to an entity
predicates what is asserted of it
signals four binary discourse signals, plus two three-valued judgements, each with probabilities

about_speaker is true / false / null (null = abstains). signals also carries place_q ("did the speaker point at one specific place or business?") and needs_map ("must something outside the model be consulted to answer?"), each "yes" / "no" / null with a probability. No threshold is applied to any of them — label and probability are both emitted and the cut is the consumer's. See ABOUT_SPEAKER_OUTPUT_CONTRACT.md and SIGNALS_JUDGE_OUTPUT_CONTRACT.md.

Spans may nest (「제 친구」 and 「친구」). Key by span offsets, not by surface.

Usage

No from_pretrained one-liner — custom four-channel span architecture.

import json, sys
from pathlib import Path

d = Path("path/to/this/repo")
sys.path.insert(0, str(d))
import modeling_btrack_4ch_w4 as M

m = M.load(str(d), "cuda")        # base + 5 gap-fill heads + about_speaker + W4 + 2 judge heads
print(json.dumps(M.infer(m, "판교 카카오 본사 어디야"), ensure_ascii=False, indent=1))

# model.safetensors carries the same weights in the same dtype, if you prefer it:
# from safetensors.torch import load_file
# m.load_state_dict(load_file(d / "model.safetensors"))

backbone_config.json is included — the repository is self-contained. md5sum -c MD5SUMS verifies a download in full; model.pt is bit-identical to the file answering live traffic.

Serve in fp32 (weights stored bf16). INT8 collapses this model — measured.


Limitations

  • Korean first. ja and en are supported and measured; Korean gets the curated material.
  • Not a general NER model. Type inventory and span scope come from a conversational product.
  • Korean particles. Spans carry the particle; we do not trim. Match with one trailing particle allowed, or strip on your side.
  • Weak on news text. Our training mix is 4.0 % news.
  • is_pure_negation has false positives no threshold recovers (8 fp / 8 fn on 3,193 rows at τ=0.7). A soft signal, not a filter.
  • Boundaries, not senses. Entity linking is a separate layer, not in this repository.

This repository

Build 0da2d001 / config 2f67efa4 — the weights answering production traffic. Refreshed roughly monthly. Two output contracts ship alongside the weights (ABOUT_SPEAKER_OUTPUT_CONTRACT.md, SIGNALS_JUDGE_OUTPUT_CONTRACT.md).

Attribution — third-party training data

Parts of the training corpus are adapted from publicly licensed datasets. Their licenses apply to those parts and are reproduced here.

KLUE — KLUE: Korean Language Understanding Evaluation, Park et al., 2021. Licensed under CC BY-SA 4.0. Source: https://github.com/KLUE-benchmark/KLUE · paper: arXiv:2105.09680.

Changes we made (this is an adaptation, not a copy). 2,500 sentences from the KLUE-NER training split were re-annotated under our own span convention and type inventory: KLUE's six entity types were mapped onto our seventeen, sentences longer than four eojeol were dropped, single-character entities were dropped, and sentences overlapping our held-out probes were removed. Intent, attribute and predicate channels carry no KLUE labels and are masked out during training. The dev split is used for evaluation only and appears nowhere in training — the evaluator verifies this and refuses to run otherwise.

MASSIVE — MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset, FitzGerald et al., Amazon Science. Licensed under CC BY 4.0. Source: https://github.com/alexa/massive.

Changes we made. Korean and Japanese rows were re-annotated under our span convention; 21 of MASSIVE's 55 slot types were mapped onto 11 of our entity types. MASSIVE's intent labels are not used — that channel is masked out per source. Dev rows used for evaluation are excluded from training by construction.

Everything else in the corpus is our own work. Training data and evaluation sets are not distributed with this repository.

License & access

Released under the Elda Community License 1.0 (see LICENSE).

  • ✅ Research, evaluation, internal validation, education — free of charge
  • ✳ Attribution: "Built with Elda"
  • ⛔ Commercial use and redistribution require a separate agreement

Access is gated: tell us who you are and access is granted automatically. Training data and evaluation sets are not distributed.

Citation

@software{intenter2026,
  title  = {Elda-Intenter: real-time multilingual speech-act and entity extraction},
  author = {Elda AI},
  year   = {2026},
  url    = {https://huggingface.co/Elda-AI/intenter}
}
Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Elda-AI/intenter

Paper for Elda-AI/intenter