Instructions to use aufklarer/GLiNER2.5-Decide-340M-MLX-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use aufklarer/GLiNER2.5-Decide-340M-MLX-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir GLiNER2.5-Decide-340M-MLX-fp16 aufklarer/GLiNER2.5-Decide-340M-MLX-fp16
- GLiNER
How to use aufklarer/GLiNER2.5-Decide-340M-MLX-fp16 with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("aufklarer/GLiNER2.5-Decide-340M-MLX-fp16") - GLiNER2
How to use aufklarer/GLiNER2.5-Decide-340M-MLX-fp16 with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("aufklarer/GLiNER2.5-Decide-340M-MLX-fp16") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
GLiNER2.5-Decide 340M MLX FP16
MLX conversion of fastino/GLiNER2.5-Decide
for native inference on Apple Silicon with speech-swift.
Given text and a label schema, it returns a probability for every label
(single-label classification) or spans of the original text (entity extraction).
It does not generate text.
Converted from the upstream checkpoint at revision 7ee5da4c2415e32259bcdc0b1a7367c32ce8d6f6.
Model
| Architecture | DeBERTa-v3-large encoder, GLiNER2 span and classification heads |
| Parameters | 340M (upstream figure) |
| Precision | FP16: float16 weights and activations |
| Weights | 973 MB, SHA-256 ef4e074b6f692a84510cd79ed521a348fa45c67aa0fc3876eddfaa5b19e47d38 |
| Context | 512 encoded tokens (schema and text together) |
| Language | English |
Files
| File | Size | Contents |
|---|---|---|
LICENSE |
11.4 kB | Apache-2.0 (model weights) |
LICENSE-gliner2-mlx |
1.1 kB | MIT (reference conversion code) |
UPSTREAM_README.md |
16.9 kB | Original Fastino model card |
config.json |
0.6 kB | GLiNER2 head configuration |
encoder_config/config.json |
0.9 kB | DeBERTa-v3-large encoder configuration |
export.json |
0.5 kB | Source revision, format and weight hash |
special_tokens_map.json |
2.4 kB | Special tokens |
tokenizer.json |
8.3 MB | Unigram tokenizer with GLiNER schema tokens |
tokenizer_config.json |
3.4 kB | Tokenizer settings |
weights.safetensors |
972.9 MB | Encoder and span/classification head weights |
Performance
Apple M5 Pro, idle machine, one process per variant, model loaded once. Median of full requests including tokenization: 16 routing cases with six labels, 8 entity cases with two labels, five timed calls each after five warmups.
| Routing (median) | Extraction (median) | Peak process memory |
|---|---|---|
| 8.8 ms | 10.0 ms | 1.58 GB |
Fidelity to the upstream PyTorch model on 24 reference cases: identical labels, spans and offsets; largest confidence difference 0.0012 (gate 0.005).
Usage
Swift (speech-swift GLiNER module):
import GLiNER
let model = try await GLiNER.fromPretrained(variant: .fp16)
let choices = try model.classify(
"Remind me to call Dad at six PM.",
labels: ["create_reminder", "send_message", "other"]
)
let spans = try model.extractEntities(
"Remind me to call Dad at six PM.",
labels: ["person", "time"]
)
CLI:
speech gliner classify "Remind me to call Dad at six PM." --variant fp16 \
--labels create_reminder,send_message,other
speech gliner extract "Remind me to call Dad at six PM." --variant fp16 \
--labels person,time --json
Probabilities are model scores, not guarantees. Entity spans are mentions: dates and times are returned as text for your code to normalize. Offsets are UTF-16.
Source and licenses
- Model:
fastino/GLiNER2.5-Decide, Apache-2.0 (LICENSE). The weights are converted to MLX layout without retraining; the original card is kept asUPSTREAM_README.md. - Conversion follows gliner2-mlx
by Andrew Chen Wang, MIT (
LICENSE-gliner2-mlx). - GLiNER2 paper: arXiv:2507.18546.
- Downloads last month
- 46
Quantized