GLiNER-PII β€” weight-only INT8/INT4 (ONNX)

A smaller ONNX build of nvidia/gliner-PII: a zero-shot named-entity model for detecting Personally Identifiable Information (PII) and Protected Health Information (PHI) across 55+ categories β€” give it text plus any label list you want, and it finds matching spans without being trained on those exact labels ahead of time.

nvidia/gliner-PII only ships PyTorch weights (1.78 GB, fp32). This repository is that same model exported to ONNX and quantized down to 577 MB β€” roughly a third of the original size.

Quick start

pip install gliner onnxruntime huggingface_hub
from gliner import GLiNER

model = GLiNER.from_pretrained(
    "Tazner/Gliner-PII-wo-int8-emb4",
    load_onnx_model=True,
    load_tokenizer=True,
    onnx_model_file="model.onnx",
)

text = "Hi support, my account username is 'johndoe88'. Reach me at (555) 123-4567 or johnd@example.com"
labels = ["email", "phone_number", "user_name"]

entities = model.predict_entities(text, labels, threshold=0.5)
for e in entities:
    print(e["text"], "->", e["label"], round(e["score"], 2))

Why this exists

The standard way to shrink an ONNX model is dynamic quantization (onnxruntime.quantization.quantize_dynamic), which compresses both the model's stored weights and the numbers flowing through it at runtime (activations) down to INT8.

That second part is a known risk here. This model's backbone, microsoft/deberta-v3-large, uses an attention mechanism ("disentangled attention") that produces occasional very large activation values. Squeezing those into INT8's 256 buckets forces the range to stretch to cover the outliers, crushing ordinary values into a couple of adjacent buckets.

This build instead uses weight-only quantization: only the numbers stored on disk are compressed. Activations stay in full precision the entire time the model runs, which avoids that failure mode by construction. The cost is latency, not accuracy risk β€” see Trade-offs.

Sizes

size
Source (pytorch_model.bin, fp32) 1782 MB
ONNX export (fp32, unquantized) 1702 MB
This model 577 MB

Trade-offs

  • Slower than the unquantized model. Weight-only quantization saves disk space, not compute β€” every quantized weight has to be unpacked back to a float before each matmul, which costs CPU time. This is already a large model (570M parameters, 24-layer deberta-v3-large) with multi-second CPU latency per call even unquantized; this build is slower still. If you have GPU access, this trade-off looks very different β€” dequant cost is much cheaper relative to compute there.
  • Embeddings are quantized to 4-bit, not 8-bit β€” onnxruntime's weight-only quantizer only supports 4-bit for Gather ops (the embedding lookup), 8-bit isn't an available option there.
  • Accuracy has not been published with this build. Weight-only quantization is expected to track the source model's behavior more closely than dynamic quantization, for the mechanistic reason described above (activations are never touched), but no recall/accuracy numbers are included in this repository. Validate on your own data before relying on this in production, same as you would with the source model.

How it was built

  1. Loaded nvidia/gliner-PII (PyTorch, fp32) and exported to ONNX via GLiNER's own model.export_to_onnx().
  2. MatMul weights β†’ INT8, block-wise (block size 128, symmetric), via onnxruntime.quantization.matmul_nbits_quantizer.MatMulNBitsQuantizer, applied directly to the fp32 export. (An fp32β†’fp16 pre-conversion was tried first to shrink the untouched remainder for free, but onnxconverter_common's converter corrupted this particular graph β€” deberta-v2's embedding subgraph deliberately upcasts to fp32 in a couple of spots for numerical stability, and the converter didn't respect that boundary. Quantizing straight from fp32 sidesteps it.)
  3. Embedding table (Gather) β†’ INT4, block-wise, same tool, separate pass (the quantizer requires 4-bit specifically for Gather ops).
  4. Everything not touched by those two passes (biases, layer norms, the small classification heads) stays at fp32.

Full detail in quantization_manifest.json.

Files

File Purpose
model.onnx The quantized ONNX graph (INT8 MatMuls, INT4 embeddings)
spm.model, tokenizer.json, tokenizer_config.json, special_tokens_map.json, added_tokens.json Tokenizer (unchanged from upstream)
gliner_config.json GLiNER model configuration
quantization_manifest.json Exact quantization recipe
LICENSE, NOTICE NVIDIA Open Model License and required attribution notice

Attribution

Source model nvidia/gliner-PII, NVIDIA Open Model License
GLiNER base urchade/gliner_large-v2.1
Backbone microsoft/deberta-v3-large
Training data nvidia/nemotron-pii

Licensed by NVIDIA Corporation under the NVIDIA Open Model License. See NOTICE.

Changes made to the source work

  • Exported the source PyTorch model to ONNX via GLiNER's export_to_onnx().
  • Applied weight-only block quantization: INT8 for MatMul weights, INT4 for the token-embedding table, both block-wise with block size 128. Runtime activations are left in float precision.
  • No retraining, fine-tuning, or change to model architecture or learned behavior beyond the numerical effects of quantization.

License

This work is distributed under the NVIDIA Open Model License, the same license as the source model. A full copy is included as LICENSE in this repository, along with the required NOTICE file.

Models under this license are commercially usable, and derivative models may be freely created and distributed under the terms of that agreement. This repository does not use NVIDIA's trade names, trademarks, or product names beyond what's needed to describe the origin of the model, as the license permits.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Tazner/Gliner-PII-wo-int8-emb4

Quantized
(3)
this model