Instructions to use edshield/piilo-deberta-v3-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use edshield/piilo-deberta-v3-small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="edshield/piilo-deberta-v3-small")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("edshield/piilo-deberta-v3-small") model = AutoModelForTokenClassification.from_pretrained("edshield/piilo-deberta-v3-small", device_map="auto") - Notebooks
- Google Colab
- Kaggle
piilo-deberta-v3-small
A token-classification model that finds personal identifiers in student writing. It is the model layer of edshield, a library that removes student identifiers from text on the device, before the text reaches any language model.
DeBERTa-v3-small (142M parameters) fine-tuned on the PIILO corpus. Seven labels, in BIO form: NAME_STUDENT, EMAIL, USERNAME, ID_NUM, PHONE_NUM, URL_PERSONAL, STREET_ADDRESS.
Use
The model is meant to be used through edshield, which adds rules, overlapping windows for long documents, name propagation, policies and a check that nothing it acted on is left in the output.
pip install "edshield[hf]"
import os
os.environ["EDSHIELD_ALLOW_DOWNLOAD"] = "1" # edshield downloads nothing unless told to
from edshield import deidentify
text = "Hi, this is Marcus. My email is marcus.t2012@gmail.com and my Discord is @marcus_hoops."
print(deidentify(text, policy="coppa").deidentified_text)
After the first run the model is in the local Hugging Face cache and the variable is no longer needed. The model also loads with transformers directly (AutoModelForTokenClassification), in which case you get token labels only.
For the browser, see edshield/piilo-deberta-v3-small-onnx.
Results
Span-level, measured with edshield's eval/evaluate.py on 680 PIILO essays held out from training, in the corpus's natural mix (581 of them contain no identifier). The detector is edshield's rules plus this model, which is how the model is used. "Got through" counts identifiers that no flag of any label touched, so they would remain in the output.
| Test set | Detector | Precision | Recall | F5 | Got through |
|---|---|---|---|---|---|
| PIILO held-out, 680 essays | Rules + this model | 0.642 | 1.000 | 0.979 | 0 of 165 |
| PIILO held-out, 680 essays | Rules only | 0.538 | 0.388 | 0.392 | 101 of 165 |
| K-12 synthetic, hard | Rules + this model | 319 of 1,433 (22%) | |||
| K-12 synthetic, hard | Rules only | 1,100 of 1,433 (77%) |
| Rules + this model, by label | Precision | Recall | n |
|---|---|---|---|
| NAME_STUDENT | 0.656 | 1.000 | 143 |
| URL_PERSONAL | 0.381 | 1.000 | 8 |
| ID_NUM | 0.700 | 1.000 | 7 |
| 1.000 | 1.000 | 4 | |
| USERNAME | 1.000 | 1.000 | 2 |
| STREET_ADDRESS | 0.500 | 1.000 | 1 |
Read these with caution:
- There is no measurement on real writing by children. PIILO is writing by adult online learners. The K-12 sets are synthetic and were written by the same people who wrote the detector.
- On child-style text about one identifier in five gets through. What gets through on the hard set is mostly lowercase schools and towns, ages in chat shorthand, spoken dates and streets without a house number.
- The PIILO set is small. 143 of its 165 identifiers are names; most other types have fewer than ten examples, so their rows say little.
- Precision is low on purpose. Most false alarms are names of people other than the essay's author, which PIILO does not label but which a privacy tool should remove.
- Some development choices were made on the validation set. Three fixes for false alarms were chosen by reading false alarms on these same 680 essays.
The full reports and the breakdown by identifier type are in the repository under eval/results/ and in docs/COVERAGE.md.
What this model is not
It is not a compliance certificate, and using it does not by itself make a product meet FERPA, COPPA or any other law. Those laws place duties on the school or the operator: consent, notice, contracts, retention, security, and the judgement that a record is de-identified. This model is one technical safeguard inside that. It does not remove all personal information. Review a sample of output from your own students' writing before relying on it.
Never test it by pasting real student data into a hosted service. Use synthetic samples.
Training
| Base model | microsoft/deberta-v3-small |
| Data | PIILO corpus, 6,807 essays. 680 held out for validation. Training used the rest, keeping every essay with an identifier and 30% of those without, plus synthetic documents for the rare labels: 4,430 documents in all |
| Epochs | 3 |
| Learning rate | 2e-5, 10% warm-up, weight decay 0.01 |
| Batch size | 2 |
| Context | 1,024 tokens, stride 128 |
| Loss | Cross-entropy with the O class weighted 0.15, to favour recall |
| Precision | fp32 |
| Hardware | One GTX 1060 6 GB, about 42 minutes |
Scripts: training/prepare_piilo.py and training/train.py in the repository.
Licence and attribution
The weights are released under CC BY 4.0, the licence of the training data.
Trained on the PIILO corpus (the dataset of the Kaggle competition "The Learning Agency Lab - PII Data Detection"), created by The Learning Agency Lab with Vanderbilt University and released under CC BY 4.0.
The base model, DeBERTa-v3-small, is by Microsoft and released under the MIT licence. The edshield code is Apache-2.0.
- Downloads last month
- 13