GLiNER2.5-Decide 340M MLX FP16

MLX conversion of fastino/GLiNER2.5-Decide for native inference on Apple Silicon with speech-swift. Given text and a label schema, it returns a probability for every label (single-label classification) or spans of the original text (entity extraction). It does not generate text.

Converted from the upstream checkpoint at revision 7ee5da4c2415e32259bcdc0b1a7367c32ce8d6f6.

Model

Architecture DeBERTa-v3-large encoder, GLiNER2 span and classification heads
Parameters 340M (upstream figure)
Precision FP16: float16 weights and activations
Weights 973 MB, SHA-256 ef4e074b6f692a84510cd79ed521a348fa45c67aa0fc3876eddfaa5b19e47d38
Context 512 encoded tokens (schema and text together)
Language English

Files

File Size Contents
LICENSE 11.4 kB Apache-2.0 (model weights)
LICENSE-gliner2-mlx 1.1 kB MIT (reference conversion code)
UPSTREAM_README.md 16.9 kB Original Fastino model card
config.json 0.6 kB GLiNER2 head configuration
encoder_config/config.json 0.9 kB DeBERTa-v3-large encoder configuration
export.json 0.5 kB Source revision, format and weight hash
special_tokens_map.json 2.4 kB Special tokens
tokenizer.json 8.3 MB Unigram tokenizer with GLiNER schema tokens
tokenizer_config.json 3.4 kB Tokenizer settings
weights.safetensors 972.9 MB Encoder and span/classification head weights

Performance

Apple M5 Pro, idle machine, one process per variant, model loaded once. Median of full requests including tokenization: 16 routing cases with six labels, 8 entity cases with two labels, five timed calls each after five warmups.

Routing (median) Extraction (median) Peak process memory
8.8 ms 10.0 ms 1.58 GB

Fidelity to the upstream PyTorch model on 24 reference cases: identical labels, spans and offsets; largest confidence difference 0.0012 (gate 0.005).

Usage

Swift (speech-swift GLiNER module):

import GLiNER

let model = try await GLiNER.fromPretrained(variant: .fp16)
let choices = try model.classify(
    "Remind me to call Dad at six PM.",
    labels: ["create_reminder", "send_message", "other"]
)
let spans = try model.extractEntities(
    "Remind me to call Dad at six PM.",
    labels: ["person", "time"]
)

CLI:

speech gliner classify "Remind me to call Dad at six PM." --variant fp16 \
    --labels create_reminder,send_message,other
speech gliner extract "Remind me to call Dad at six PM." --variant fp16 \
    --labels person,time --json

Probabilities are model scores, not guarantees. Entity spans are mentions: dates and times are returned as text for your code to normalize. Offsets are UTF-16.

Source and licenses

  • Model: fastino/GLiNER2.5-Decide, Apache-2.0 (LICENSE). The weights are converted to MLX layout without retraining; the original card is kept as UPSTREAM_README.md.
  • Conversion follows gliner2-mlx by Andrew Chen Wang, MIT (LICENSE-gliner2-mlx).
  • GLiNER2 paper: arXiv:2507.18546.
Downloads last month
46
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aufklarer/GLiNER2.5-Decide-340M-MLX-fp16

Finetuned
(9)
this model

Paper for aufklarer/GLiNER2.5-Decide-340M-MLX-fp16