Instructions to use Tazner/Gliner-PII-wo-int8-emb4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER
How to use Tazner/Gliner-PII-wo-int8-emb4 with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("Tazner/Gliner-PII-wo-int8-emb4") text = "Cristiano Ronaldo dos Santos Aveiro was born on 5 February 1985 in Funchal, Madeira, Portugal." labels = ["person", "date", "location"] entities = model.predict_entities(text, labels) for entity in entities: print(entity["text"], "=>", entity["label"]) - Notebooks
- Google Colab
- Kaggle
GLiNER-PII β weight-only INT8/INT4 (ONNX)
A smaller ONNX build of nvidia/gliner-PII: a
zero-shot named-entity model for detecting Personally Identifiable Information (PII) and
Protected Health Information (PHI) across 55+ categories β give it text plus any label
list you want, and it finds matching spans without being trained on those exact labels
ahead of time.
nvidia/gliner-PII only ships PyTorch weights (1.78 GB, fp32). This repository is that
same model exported to ONNX and quantized down to 577 MB β roughly a third of the
original size.
Quick start
pip install gliner onnxruntime huggingface_hub
from gliner import GLiNER
model = GLiNER.from_pretrained(
"Tazner/Gliner-PII-wo-int8-emb4",
load_onnx_model=True,
load_tokenizer=True,
onnx_model_file="model.onnx",
)
text = "Hi support, my account username is 'johndoe88'. Reach me at (555) 123-4567 or johnd@example.com"
labels = ["email", "phone_number", "user_name"]
entities = model.predict_entities(text, labels, threshold=0.5)
for e in entities:
print(e["text"], "->", e["label"], round(e["score"], 2))
Why this exists
The standard way to shrink an ONNX model is dynamic quantization
(onnxruntime.quantization.quantize_dynamic), which compresses both the model's stored
weights and the numbers flowing through it at runtime (activations) down to INT8.
That second part is a known risk here. This model's backbone,
microsoft/deberta-v3-large, uses an
attention mechanism ("disentangled attention") that produces occasional very large
activation values. Squeezing those into INT8's 256 buckets forces the range to stretch to
cover the outliers, crushing ordinary values into a couple of adjacent buckets.
This build instead uses weight-only quantization: only the numbers stored on disk are compressed. Activations stay in full precision the entire time the model runs, which avoids that failure mode by construction. The cost is latency, not accuracy risk β see Trade-offs.
Sizes
| size | |
|---|---|
Source (pytorch_model.bin, fp32) |
1782 MB |
| ONNX export (fp32, unquantized) | 1702 MB |
| This model | 577 MB |
Trade-offs
- Slower than the unquantized model. Weight-only quantization saves disk space, not compute β every quantized weight has to be unpacked back to a float before each matmul, which costs CPU time. This is already a large model (570M parameters, 24-layer deberta-v3-large) with multi-second CPU latency per call even unquantized; this build is slower still. If you have GPU access, this trade-off looks very different β dequant cost is much cheaper relative to compute there.
- Embeddings are quantized to 4-bit, not 8-bit β
onnxruntime's weight-only quantizer only supports 4-bit forGatherops (the embedding lookup), 8-bit isn't an available option there. - Accuracy has not been published with this build. Weight-only quantization is expected to track the source model's behavior more closely than dynamic quantization, for the mechanistic reason described above (activations are never touched), but no recall/accuracy numbers are included in this repository. Validate on your own data before relying on this in production, same as you would with the source model.
How it was built
- Loaded
nvidia/gliner-PII(PyTorch, fp32) and exported to ONNX via GLiNER's ownmodel.export_to_onnx(). - MatMul weights β INT8, block-wise (block size 128, symmetric), via
onnxruntime.quantization.matmul_nbits_quantizer.MatMulNBitsQuantizer, applied directly to the fp32 export. (An fp32βfp16 pre-conversion was tried first to shrink the untouched remainder for free, butonnxconverter_common's converter corrupted this particular graph β deberta-v2's embedding subgraph deliberately upcasts to fp32 in a couple of spots for numerical stability, and the converter didn't respect that boundary. Quantizing straight from fp32 sidesteps it.) - Embedding table (
Gather) β INT4, block-wise, same tool, separate pass (the quantizer requires 4-bit specifically forGatherops). - Everything not touched by those two passes (biases, layer norms, the small classification heads) stays at fp32.
Full detail in quantization_manifest.json.
Files
| File | Purpose |
|---|---|
model.onnx |
The quantized ONNX graph (INT8 MatMuls, INT4 embeddings) |
spm.model, tokenizer.json, tokenizer_config.json, special_tokens_map.json, added_tokens.json |
Tokenizer (unchanged from upstream) |
gliner_config.json |
GLiNER model configuration |
quantization_manifest.json |
Exact quantization recipe |
LICENSE, NOTICE |
NVIDIA Open Model License and required attribution notice |
Attribution
| Source model | nvidia/gliner-PII, NVIDIA Open Model License |
| GLiNER base | urchade/gliner_large-v2.1 |
| Backbone | microsoft/deberta-v3-large |
| Training data | nvidia/nemotron-pii |
Licensed by NVIDIA Corporation under the NVIDIA Open Model License. See NOTICE.
Changes made to the source work
- Exported the source PyTorch model to ONNX via GLiNER's
export_to_onnx(). - Applied weight-only block quantization: INT8 for MatMul weights, INT4 for the token-embedding table, both block-wise with block size 128. Runtime activations are left in float precision.
- No retraining, fine-tuning, or change to model architecture or learned behavior beyond the numerical effects of quantization.
License
This work is distributed under the NVIDIA Open Model License,
the same license as the source model. A full copy is included as LICENSE in this
repository, along with the required NOTICE file.
Models under this license are commercially usable, and derivative models may be freely created and distributed under the terms of that agreement. This repository does not use NVIDIA's trade names, trademarks, or product names beyond what's needed to describe the origin of the model, as the license permits.
- Downloads last month
- 16
Model tree for Tazner/Gliner-PII-wo-int8-emb4
Base model
nvidia/gliner-PII