Instructions to use Elda-AI/intenter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Elda-AI/intenter with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="Elda-AI/intenter")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Elda-AI/intenter", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Access to Intenter
Intenter is released for research, evaluation, internal validation and education. Tell us who you are and we will grant access.
By requesting access you agree to the Elda Community License 1.0: no commercial use, no redistribution of the weights or derivatives, and attribution as "Built with Elda". Commercial licensing is available on request.
Log in or Sign Up to review the conditions and access this model content.
Elda-Intenter — real-time speech-act + entity extraction in a single pass
Korean-first, multilingual (ko · ja · en). 281M parameters · p50 27.8 ms.
Elda-Intenter reads an utterance and returns, in one forward pass, four aligned channels: what the speaker is doing (speech act), what they are talking about (entities), and the attribute/predicate structure that binds them.
It is not a general-purpose NER model. It is the perception layer of a live conversational system, built for the constraint that shapes everything else in that setting: latency budgets measured in tens of milliseconds.
"간판의 품격이라는 음식점이 있다며, 그 음식점은 어디에 있냐고"
intent QUESTION
entity 간판의 품격 → ORGANIZATION
음식점이 → ORGANIZATION
signals is_pure_negation=0 · is_claim=0 · polarity=0 · depends_on_prev=1
Results
All figures below are on sets excluded from training by construction.
Public benchmarks
KLUE-NER (dev, Korean · 600 sentences / 1,687 gold entities), entity-level F1:
| entity F1 | boundary F1 | |
|---|---|---|
| exact match | 54.9 | 57.8 |
| allowing one trailing Korean particle | 63.5 | 66.9 |
| oracle ceiling — gold passed through our scoring pipeline | 99.9 |
Korean particles attach to the noun, so 카카오에서 and 카카오 are the same entity with
different case marking. KLUE's annotation excludes the particle; we emit it. The second row
accepts one trailing particle when the start offset matches — it does not trim the span.
We never cut a span to fit a ruler: the span is the string a downstream consumer searches
with, and 탕웨이 trimmed to 탕웨 cannot be looked up at all. The oracle row confirms the
pipeline returns KLUE's own gold unchanged.
How we compare. A single number says nothing on its own — but a number only belongs in the same table if it carries the same tag (ruler, revision, split, scope). So we split the table.
Measured by us, on this ruler and this split (KLUE-NER dev, 600 sentences, one trailing particle allowed, unmapped types charged):
| entity F1 | ||
|---|---|---|
soddokayo/klue-roberta-large-klue-ner |
83.4 | fine-tuned on KLUE-NER train; ⛔ license not declared |
| Elda-Intenter | 63.5 | general-purpose; 500 KLUE train sentences (see Attribution) |
Reported by others, on their own runs — ⚠ not our ruler, not verified by us:
| entity F1 | source | |
|---|---|---|
| KLUE-RoBERTa-large | 90.8 | published baseline |
| XLM-R-large | 85.9 | KLUE paper (arXiv:2105.09680) |
| KR-BERT-base | 77.2 | KLUE paper |
| mBERT-base | 73.2 | KLUE paper |
⚠ Why two tables. When we re-scored a KLUE-NER fine-tuned RoBERTa-large on our ruler and this split, it came out at 83.4, not the 90.8 that circulates. We do not claim either is "wrong" — they are different runs with different post-processing. That is the point: a reported number and a number you measured yourself are not the same kind of number, and putting them in one column invents a comparison nobody ran.
Conditions still differ inside the first table. Every peer is a dedicated NER model fine-tuned on all 21,008 KLUE-NER training sentences. Elda-Intenter is a general conversational perception model that saw 500 of them, emits 17 entity types that must be mapped back onto KLUE's six, and is charged for everything outside that mapping. We are the lowest row and we print it that way.
MASSIVE (dev, re-annotated), entity F1: Japanese 67.7 (749 sentences) · Korean 51.6 (469 sentences).
⚠ Not comparable to published MASSIVE scores. MASSIVE is an intent-classification and slot-filling benchmark with 55 slot types; we are not running that task. We mapped 21 of its slot types onto 11 of our entity types and score exact boundary plus type; unlike the KLUE run above, anything outside that mapping is excluded from the denominator rather than charged, and that exclusion is reported per run. Zero training overlap is verified by the evaluator, which refuses to run otherwise.
When the boundary is right, the type is right 96.4 % (ja) and 94.8 % (ko) of the time. The weak types are the same in both languages: TIME, DATE, PHENOMENON, EVENT.
Frozen internal gates
Each row is a distinct capability axis, re-run on every build. Evaluated per surface, not summed.
| Role nouns — detection / typing (405 slots) | 401 / 399 |
| Referent expressions, held out | 23 · 21 · 20 |
| Conversational-speech detection / safety floor | 90.9 % · 12 / 12 |
| Possessive structure — owner kept (40 surfaces) | 31 / 40 |
| Discourse signals — false positives on short utterances | 2 / 48 |
| Discourse signals — polarity grid | 24 / 24 |
| Formal-register interrogatives | 14 / 14 |
| Name-span probe, third-party (698 slots) | 660 / 698 |
| Span-convention compliance, third-party ruler | 96 / 126 · with our particle rule 122 / 126 |
The last two rows are produced by a separate team on their own infrastructure, from sentences this model has never been trained on.
Why an encoder
| Elda-Intenter | typical small-LLM extraction | |
|---|---|---|
| Parameters | 281M | 500M – 4B |
| Latency (single, incl. heads) | p50 27.8 ms | hundreds of ms |
| Latency (batch of 16) | 118.5 ms | seconds |
| Serving precision | fp32 (2 workers, 6.1 GB GPU) | varies |
| Output | fixed contract, char offsets | free text to be parsed |
Measured on the production serving path (RTX 3060, fp32, heads attached).
Span extraction is a classification problem over token pairs, and a bidirectional encoder sees the whole utterance at once. A decoder can be prompted into this task and will be more flexible; it cannot be fast at it. When a dialogue turn has a 300 ms end-to-end budget and perception is one of five stages inside it, 27.8 ms is the difference between a design that works and one that does not.
Output contract
Four Global Pointer channels over one shared mDeBERTa-v3-base backbone, plus four frozen-stage binary signal heads and one span-attribute head. All channels are character spans over the original string.
| channel | what it carries |
|---|---|
intents |
speech act over the utterance or clause — 18 types incl. AGREE / DISAGREE / QUESTION / COMMAND / NARRATE / DESCRIBE / REQUEST |
entities |
17 entity types × subtypes, plus about_speaker per span |
attributes |
modifiers bound to an entity |
predicates |
what is asserted of it |
signals |
four binary discourse signals with probabilities |
about_speaker is true / false / null, where null means the model abstains. No
threshold is applied — the raw class and its probability are both emitted, because a weak
judgement is evidence to a downstream reasoner even when it is not a filter.
Spans may nest: the same surface can carry more than one entity (「제 친구」 and 「친구」), each
with its own about_speaker. Key by span offsets, not by surface.
Usage
There is no from_pretrained one-liner — the architecture is a custom four-channel span model,
and bit-stable serving was chosen over loading convenience.
import json, sys
from pathlib import Path
d = Path("path/to/this/repo")
sys.path.insert(0, str(d))
import modeling_btrack_4ch_w4 as M
m = M.load(str(d), "cuda") # base + 5 gap-fill heads + about_speaker + W4
print(json.dumps(M.infer(m, "판교 카카오 본사 어디야"), ensure_ascii=False, indent=1))
# model.safetensors carries the same weights in the same dtype, if you prefer it:
# from safetensors.torch import load_file
# m.load_state_dict(load_file(d / "model.safetensors"))
backbone_config.json is included, so the repository is self-contained — you do not need to
download the base encoder separately.
MD5SUMS is the production serving manifest, unmodified, so you can verify that model.pt here
is bit-identical to what answers live traffic (5c2ce59a…). It does not list the files that
exist only in this repository, so md5sum -c reports those as missing. That is expected.
Serve in fp32 (weights are stored bf16). INT8 collapses this model — measured, not assumed.
Limitations
- Korean first. Japanese and English are supported and measured, but Korean receives the curated material.
- Not a general NER model. The type inventory and span scope come from a conversational product, not from a treebank.
- Korean particles. Spans carry the particle more often than our own convention prescribes. We do not trim them — match with one trailing particle allowed, or strip on your side if your index requires it.
is_pure_negationshows false positives that no threshold recovers (26 on a 1,150-row label set). Treat it as a soft signal, not a filter.- Boundaries, not senses. Entity linking and disambiguation are separate layers and are not in this repository.
This repository
Build 5c2ce59a / config 4c697b7c — the weights that answer production traffic. Refreshed
roughly monthly; internal builds move faster than this repository does.
Attribution — third-party training data
Parts of the training corpus are adapted from publicly licensed datasets. Their licenses apply to those parts and are reproduced here.
KLUE — KLUE: Korean Language Understanding Evaluation, Park et al., 2021. Licensed under CC BY-SA 4.0. Source: https://github.com/KLUE-benchmark/KLUE · paper: arXiv:2105.09680.
Changes we made (this is an adaptation, not a copy). 500 sentences from the KLUE-NER training split were re-annotated under our own span convention and type inventory: KLUE's six entity types were mapped onto our seventeen, sentences longer than four eojeol were dropped, single-character entities were dropped, and sentences overlapping our held-out probes were removed. Intent, attribute and predicate channels carry no KLUE labels and are masked out during training. The dev split is used for evaluation only and appears nowhere in training — the evaluator verifies this and refuses to run otherwise.
MASSIVE — MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset, FitzGerald et al., Amazon Science. Licensed under CC BY 4.0. Source: https://github.com/alexa/massive.
Changes we made. Korean and Japanese rows were re-annotated under our span convention; 21 of MASSIVE's 55 slot types were mapped onto 11 of our entity types. MASSIVE's intent labels are not used — that channel is masked out per source. Dev rows used for evaluation are excluded from training by construction.
Everything else in the corpus is our own work. Training data and evaluation sets are not distributed with this repository.
License & access
Released under the Elda Community License 1.0 (see LICENSE).
- ✅ Research, evaluation, internal validation, education — free of charge
- ✳ Attribution: "Built with Elda"
- ⛔ Commercial use and redistribution require a separate agreement
Access is gated: tell us who you are and access is granted automatically. Training data and evaluation sets are not distributed.
Citation
@software{intenter2026,
title = {Elda-Intenter: real-time multilingual speech-act and entity extraction},
author = {Elda AI},
year = {2026},
url = {https://huggingface.co/Elda-AI/intenter}
}
- Downloads last month
- -