Gandalf 0.0.1 · Moderation that can stay with the message

A compact, policy-conditioned text classifier for E2EM — End to End Moderation.

Gandalf explores a practical question: can useful moderation happen locally, without sending private conversations to a remote scoring service? This release provides the weights, tokenizer, exact policy wordings, frozen thresholds and recorded evaluation results needed to investigate that question.

Project and introduction video · Technical specification · Evaluation evidence · Release manifest

Release fact Value
Version 0.0.1 — research candidate
Architecture DeBERTa policy–message cross-encoder, one output logit per pair
Parameters 70,830,337 including embeddings
Original model package 291,720,350 bytes (291.72 MB, decimal), FP32
Declared categories 26, with canonical wordings and calibration-frozen thresholds
Input English policy wording, target message and optional conversation context
Context budget 512 tokens including policy and special tokens
Formal acceptance 0 of 26 categories pass the recorded gate

This is an inspectable research baseline, not a production safety guarantee. The acceptance target is at least 90% recall and at most 1% benign false positives per category, assessed using 95% confidence bounds. The release does not pass it. The weights can run locally; physical-phone performance and an end-to-end mobile runtime have not been demonstrated for this checkpoint.

Why it matters

An app could use a local moderation score to offer a warning, blur content or invite a user to reconsider a message. The app chooses the policy and action. The design seeks to reduce the need to disclose message content to a separate moderation service. Hosting this model on Hugging Face distributes the files; the example below performs inference locally after download.

The research contribution is a small, reproducible moderation component with an explicit contract and measurable release criteria. It is not a claim that moderation alone solves online harm, or that the containing app preserves privacy merely because its classifier runs locally.

Run locally

Use Python 3.11+ with PyTorch, Transformers, Safetensors and huggingface_hub. The release was smoke-tested on CPU with torch 2.2.2, transformers 4.57.6, huggingface_hub 0.36.2 and safetensors 0.8.0. Those are recorded test versions, not a requirement to downgrade an existing environment.

python -m pip install torch transformers safetensors huggingface_hub
hf download krazyjakee/gandalf --revision v0.0.1 --local-dir gandalf
python gandalf/example.py --model-path gandalf \
  --preset abuse.threat --message "Thank you for your help."

The command selects the versioned release; use its Hub commit hash for an exact immutable revision. Review the included Python files before executing them; loading the model does not require trust_remote_code=True.

example.py uses the same preprocessing as the project's worker, included as preset_backend.py. It preserves the policy, keeps the beginning and end of long message evidence, and applies a sigmoid to the single output logit. With context, the second input is Message:\n…\n\nContext:\n…. Without context it is the message alone. A generic text-classification pipeline is not a substitute for this input format or the preset-specific thresholds.

Use the exact canonical wording in preset.json. A batch of policies is a batch of policy–message pairs, not one shared text encoding with 26 output heads. Scores are model outputs, not calibrated probabilities of harm or legal verdicts. The example returns a score and the stored thresholds; it does not take action against a user.

Evidence for this checkpoint

The weights SHA-256 is 916370da013a428a8ed3e00791049d12b3b3fd5836969d8b0492bc5de372d5c8. The checkpoint was trained in the all-v2 Jigsaw round, with seed 17 and three epochs. The training record contains 749,356 policy–evidence examples, not that many unique conversations. The older all-v1 package remains the application's default; these results and these weights are specifically 0.0.1.

DynaHate external test: 4,118 human-labelled messages; the identity-hate threshold was frozen using the preset calibration split, not selected on this test. AUROC is 0.8238. It catches 1,507 / 2,267 hateful messages (66.48%) and flags 340 / 1,851 non-hateful messages (18.37%). This deliberately challenging dataset contains adversarial examples; it is not a population estimate of a deployed app's false-alarm rate.

The aggregate, per-category all-rows gate results, counts and confidence bounds are in evaluation.json. Some source tags are broader than the canonical policy, and some categories have too few positives. A teacher- filtered gate is not an independent test against that teacher. Generated chat checks are diagnostics, not independent human acceptance evidence. Desktop/GPU evaluation does not establish phone latency, memory, battery use or robustness.

The project website's rounded 67% is a descriptive average of recall progress relative to the target, using generated diagnostic recall for categories without a gradable human-tagged result. It is not accuracy, a gate-pass percentage, a probability of funding, or readiness for deployment. False positives and confidence bounds are separate requirements.

Intended use and limitations

Use this release for offline research, reproducible evaluation and controlled integration experiments. Prioritise support-seeking, quotations, counterspeech, dialects and ambiguous conversations when evaluating false alarms. The current model can confuse discussion of a harmful topic with actually committing the harm, and can miss indirect threats or harms without explicit vocabulary.

Do not rely on it as the sole basis for account sanctions, emergency response, child safeguarding or decisions about individuals. English is the evaluated language; other languages, unseen policy wording and real-world conversation transfer have not been established. Client-side execution also requires a separate threat model for tampering, policy updates, abuse of the classifier, telemetry, consent and control over actions. No UK government, Ofcom or other regulatory endorsement or compliance certification is claimed.

Provenance and licence status

The base checkpoint is cross-encoder/nli-deberta-v3-xsmall at revision a150876415327c80daeff35ca6f68f5ed8cf5c24, whose card declares Apache-2.0. This release changes the trained weights and classification output. See ATTRIBUTION.md for the separate data sources and LICENSE_STATUS.md for the unresolved release-wide licence. No blanket Apache-2.0 or CC-BY claim is made for this mixed-source fine-tune.

The original training manifest is preserved as distillation.json. Its generic teacher description and mention of preset heads do not fully describe this particular checkpoint; the architecture here is verified from the configuration and weights. Source-level attribution is more specific than that manifest's single training-dataset licence field. No raw training or test conversations are included in this Hub release.

Development and collaboration

The next phase is to test a narrower, useful set of policies against independent, held-out human judgements; reduce false alarms; and validate the same frozen model on physical iPhone 14 and Pixel 6 devices with no Python worker.

E2EM seeks research and deployment partners interested in measurable online safety and local processing. The project is seeking support for this work; funding, partners and adoption are not implied by this release.

Research and partnerships · Explore E2EM · Support development

For technical feedback, use this model repository's Community tab. Report the model revision, preset, input format and observed result; do not post private conversations or identifying information.

Downloads last month
23
Safetensors
Model size
70.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for krazyjakee/gandalf

Finetuned
(2)
this model

Dataset used to train krazyjakee/gandalf