Instructions to use krazyjakee/gandalf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use krazyjakee/gandalf with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="krazyjakee/gandalf")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("krazyjakee/gandalf") model = AutoModelForSequenceClassification.from_pretrained("krazyjakee/gandalf", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Gandalf 0.0.1 · Moderation that can stay with the message
A compact, policy-conditioned text classifier for E2EM — End to End Moderation.
Gandalf explores a practical question: can useful moderation happen locally, without sending private conversations to a remote scoring service? This release provides the weights, tokenizer, exact policy wordings, frozen thresholds and recorded evaluation results needed to investigate that question.
Project and introduction video · Technical specification · Evaluation evidence · Release manifest
| Release fact | Value |
|---|---|
| Version | 0.0.1 — research candidate |
| Architecture | DeBERTa policy–message cross-encoder, one output logit per pair |
| Parameters | 70,830,337 including embeddings |
| Original model package | 291,720,350 bytes (291.72 MB, decimal), FP32 |
| Declared categories | 26, with canonical wordings and calibration-frozen thresholds |
| Input | English policy wording, target message and optional conversation context |
| Context budget | 512 tokens including policy and special tokens |
| Formal acceptance | 0 of 26 categories pass the recorded gate |
This is an inspectable research baseline, not a production safety guarantee. The acceptance target is at least 90% recall and at most 1% benign false positives per category, assessed using 95% confidence bounds. The release does not pass it. The weights can run locally; physical-phone performance and an end-to-end mobile runtime have not been demonstrated for this checkpoint.
Why it matters
An app could use a local moderation score to offer a warning, blur content or invite a user to reconsider a message. The app chooses the policy and action. The design seeks to reduce the need to disclose message content to a separate moderation service. Hosting this model on Hugging Face distributes the files; the example below performs inference locally after download.
The research contribution is a small, reproducible moderation component with an explicit contract and measurable release criteria. It is not a claim that moderation alone solves online harm, or that the containing app preserves privacy merely because its classifier runs locally.
Run locally
Use Python 3.11+ with PyTorch, Transformers, Safetensors and huggingface_hub. The release was smoke-tested on CPU with torch 2.2.2, transformers 4.57.6, huggingface_hub 0.36.2 and safetensors 0.8.0. Those are recorded test versions, not a requirement to downgrade an existing environment.
python -m pip install torch transformers safetensors huggingface_hub
hf download krazyjakee/gandalf --revision v0.0.1 --local-dir gandalf
python gandalf/example.py --model-path gandalf \
--preset abuse.threat --message "Thank you for your help."
The command selects the versioned release; use its Hub commit hash for an exact
immutable revision. Review the included Python files before executing them; loading the
model does not require trust_remote_code=True.
example.py uses the same preprocessing as the project's worker, included as
preset_backend.py. It preserves the policy, keeps the beginning and end of long
message evidence, and applies a sigmoid to the single output logit. With context,
the second input is Message:\n…\n\nContext:\n…. Without context it is the
message alone. A generic text-classification pipeline is not a substitute for
this input format or the preset-specific thresholds.
Use the exact canonical wording in preset.json. A batch of policies is a batch
of policy–message pairs, not one shared text encoding with 26 output heads.
Scores are model outputs, not calibrated probabilities of harm or legal verdicts.
The example returns a score and the stored thresholds; it does not take action
against a user.
Evidence for this checkpoint
The weights SHA-256 is
916370da013a428a8ed3e00791049d12b3b3fd5836969d8b0492bc5de372d5c8.
The checkpoint was trained in the all-v2 Jigsaw round, with seed 17 and three
epochs. The training record contains 749,356 policy–evidence examples, not that
many unique conversations. The older all-v1 package remains the application's
default; these results and these weights are specifically 0.0.1.
DynaHate external test: 4,118 human-labelled messages; the identity-hate threshold was frozen using the preset calibration split, not selected on this test. AUROC is 0.8238. It catches 1,507 / 2,267 hateful messages (66.48%) and flags 340 / 1,851 non-hateful messages (18.37%). This deliberately challenging dataset contains adversarial examples; it is not a population estimate of a deployed app's false-alarm rate.
The aggregate, per-category all-rows gate results, counts and confidence bounds are in evaluation.json. Some source tags are broader than the canonical policy, and some categories have too few positives. A teacher- filtered gate is not an independent test against that teacher. Generated chat checks are diagnostics, not independent human acceptance evidence. Desktop/GPU evaluation does not establish phone latency, memory, battery use or robustness.
The project website's rounded 67% is a descriptive average of recall progress relative to the target, using generated diagnostic recall for categories without a gradable human-tagged result. It is not accuracy, a gate-pass percentage, a probability of funding, or readiness for deployment. False positives and confidence bounds are separate requirements.
Intended use and limitations
Use this release for offline research, reproducible evaluation and controlled integration experiments. Prioritise support-seeking, quotations, counterspeech, dialects and ambiguous conversations when evaluating false alarms. The current model can confuse discussion of a harmful topic with actually committing the harm, and can miss indirect threats or harms without explicit vocabulary.
Do not rely on it as the sole basis for account sanctions, emergency response, child safeguarding or decisions about individuals. English is the evaluated language; other languages, unseen policy wording and real-world conversation transfer have not been established. Client-side execution also requires a separate threat model for tampering, policy updates, abuse of the classifier, telemetry, consent and control over actions. No UK government, Ofcom or other regulatory endorsement or compliance certification is claimed.
Provenance and licence status
The base checkpoint is
cross-encoder/nli-deberta-v3-xsmall
at revision a150876415327c80daeff35ca6f68f5ed8cf5c24, whose card declares
Apache-2.0. This release changes the trained weights and classification output.
See ATTRIBUTION.md for the separate data sources and
LICENSE_STATUS.md for the unresolved release-wide licence.
No blanket Apache-2.0 or CC-BY claim is made for this mixed-source fine-tune.
The original training manifest is preserved as distillation.json. Its generic
teacher description and mention of preset heads do not fully describe this
particular checkpoint; the architecture here is verified from the configuration
and weights. Source-level attribution is more specific than that manifest's
single training-dataset licence field. No raw training or test conversations are
included in this Hub release.
Development and collaboration
The next phase is to test a narrower, useful set of policies against independent, held-out human judgements; reduce false alarms; and validate the same frozen model on physical iPhone 14 and Pixel 6 devices with no Python worker.
E2EM seeks research and deployment partners interested in measurable online safety and local processing. The project is seeking support for this work; funding, partners and adoption are not implied by this release.
Research and partnerships · Explore E2EM · Support development
For technical feedback, use this model repository's Community tab. Report the model revision, preset, input format and observed result; do not post private conversations or identifying information.
- Downloads last month
- 23
Model tree for krazyjakee/gandalf
Base model
microsoft/deberta-v3-xsmall