Access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This model is a derivative work of google/gemma-4-31B-it, which Google releases under the Apache License 2.0 (https://ai.google.dev/gemma/docs/gemma_4_license). Lexsi Labs' modifications are provided under the Lexsi Labs Source Available License (LSAL) v1.2 (https://huggingface.co/Lexsi/Gemma-4-31B-it-Multilingual-Safety-Repair/blob/main/LICENSE-LSAL-1.2.md), a source-available noncommercial license; organizational use requires the acknowledgement or permission process described there.

Log in or Sign Up to review the conditions and access this model content.

Lexsi Labs

Gemma-4-31B: Multilingual Safety Repair

A single frozen checkpoint that improves harmful-request refusal across English and five Indic languages, with general capability preserved and no runtime language routing. Released gated for evaluation, not certified for deployment; see Results for the safety gains and the costs to weigh.

What we did

Lexsi turns a known safety failure into a controlled model change. Instead of retraining the whole model or wrapping it in a guard, router, or inference-time filter, we locate the specific behaviour in the weights, apply a bounded, circuit-restricted edit, and verify the fix holds without eroding the capabilities you depend on. You get one drop-in checkpoint: the safety change lives in the model and is served the way you serve it today, with nothing extra to run or maintain.

Every repair is a measured trade-off. We report the safety gain and its collateral effects on legitimate use, capability, and fairness side by side (see Results), so you can decide whether to deploy from the evidence.

This checkpoint applies that approach to multilingual safety across English and five Indic languages, so a request is handled as consistently whichever of those languages it is written in. One frozen checkpoint, no per-language routing.

Built using the Lexsi Alignment and Safety Stack:

Targeted repair workflow: find, specify, repair, verify, decide

The repair workflow behind this checkpoint (Figure 2 of the paper). We fix the specification before intervention; verification tests both the intended gain and the behavioural blast radius before any deployment decision. The four libraries below implement it.

Library Role
CuratorKIT Curates training data with source provenance
CircuitKIT Finds and tests model circuits to guide targeted interventions
SafeTune Applies and compares methods for correcting safety drift
AuditKIT Keeps evaluation results comparable across models and interventions
Base model google/gemma-4-31B-it
Developed by Lexsi Labs (Lithasa Technologies Pvt. Ltd.)
Model size 31B parameters
Context length 256K tokens (inherited from the base model)
Input / output Text and image in, text out (inherited); the repair was evaluated on text only
Contact support@lexsi.ai
Languages en, hi, bn, mr, te, ta
License Apache 2.0 (base model) + LSAL v1.2 (modifications); see License
Status Gated evaluation release; not certified for deployment
Trade-off Safety up across six languages, capability preserved; a fairness cost and higher over-refusal to weigh (see Results)

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Lexsi/Gemma-4-31B-it-Multilingual-Safety-Repair"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "भारत में यूपीआई भुगतान कैसे काम करता है? संक्षेप में समझाइए।"}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)

out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

We evaluated with greedy decoding in bfloat16; the results below use that setting.

Results

Base checkpoint (google/gemma-4-31B-it) vs. this repaired checkpoint. None of these prompts were used in training or strength selection.

Safety: held-out benchmarks (%, higher refusal is safer unless noted)

benchmark n base this checkpoint
IndicGuard (safe response to unsafe prompt) 11,429 77.7 79.6
IndicSafetyBench (full, 23 languages) 3,374 52.6 58.2
IndicSafe (refusal of harmful, official judge) 3,960 50.1 62.1
Sarvam Indic-Safety (23 languages) 1,830 67.8 72.4
BanglaSafe 879 33.8 41.5
HarmBench 200 92.0 94.0
SORRY-Bench 9,240 67.1 72.3
MultiJail (10 languages) 3,150 51.7 57.0
StrongREJECT 313 95.6 96.7
BeaverTails (safe response to unsafe, beaver-dam) 1,733 88.5 91.1
TRIDENT (mean harm, 1–5, ↓) 2,652 1.54 1.45
XSTest (over-refusal, ↓) 250 1.2 1.2

Independent India-suite cross-check (base → this checkpoint):

task n base this checkpoint
Multilingual safety (5 langs) 1,000 88.6 92.1
Jailbreak safety (7 langs) 70 48.6 80.0
DPI safety (% correct) 150 96.7 95.3
India-suite safety, pooled (panel) 1,220 87.3 91.8

Cost: the watch-list (%, lower is better)

metric n base this checkpoint
Indic-Bias (stereotype-confirming) 2,280 22.4 38.9
OR-Bench-Hard-1k (benign over-refusal) 1,319 33.0 47.8

These two rows are the costs to weigh. The safety gains are real and general capability is preserved; alongside them, the Indic-Bias fairness regression and higher benign over-refusal on a hard English probe are reported in full, not netted against the safety gain.

Capability preserved (base → this checkpoint)

benchmark n base this checkpoint
GSM8K (5-shot) 1,319 93.6 93.9
IFEval (instruction following) 541 90.8 90.6
MMLU (generative) 14,042 85.7 85.3
MILU (8 Indic languages) 57,449 68.7 65.8
BELEBELE (6 languages) 5,400 89.6 89.5
TruthfulQA 817 56.6 56.9
IndicQA (token-F1 / EM) 13,295 73.4 / 52.5 73.2 / 52.4
Bhasha-Abhijnaanam LID (native / romanized) 88,013 / 55,821 97.4 / 86.1 97.2 / 86.5
Multilingual MMLU (16 languages) 2,272 74.3 74.2
WMDP-cyber / bio / chem (hazardous knowledge, ↓) 1,987 / 1,273 / 408 75.7 / 82.8 / 67.6 75.0 / 82.3 / 67.6

The accompanying report gives the full multi-benchmark picture, the per-language build, the judge protocol, and the paired significance tests.

Intended use

Research and evaluation of harmful-request safety in English, Hindi, Bengali, Marathi, Tamil, and Telugu. It is a candidate for assistants where harmful-request safety in these languages matters most, once the operator has run its own workload tests, including a fairness test.

License

This is a derivative work of google/gemma-4-31B-it with selected weights modified. The NOTICE file states what was changed. Google releases the base model under the Apache License 2.0; a copy is included as LICENSE-APACHE-2.0.txt.

Lexsi's modifications (the weight correction, the recipe, and this card) are offered under the Lexsi Source Available License (LSAL) v1.2 (source-available, noncommercial). See the NOTICE file for details.

Please also follow Google's Gemma Prohibited Use Policy when using this model or anything derived from it.

Released gated for teams to evaluate and use. Review the Results, including the fairness cost, before deploying in your setting. Not independently audited.

Contact

Questions, errors in this card, and organizational-use acknowledgements (LSAL Section 1A): support@lexsi.ai

Citation

If you use this checkpoint, please cite it and the accompanying paper:

@misc{lexsi2026gemma4multilingualrepair,
  title        = {Gemma-4-31B: Multilingual Safety Repair},
  author       = {Seth, Pratinav and Dhor, Ashim and Sadhu, Saisab and Bhattacharjee, Soham and Gosalia, Hem and
            Sankarapu, Vinay Kumar},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Lexsi/Gemma-4-31B-it-Multilingual-Safety-Repair}}
}

@misc{lexsi2026lastmile,
  title  = {The Last Mile of Model Safety: Targeted Model Repair},
  author = {Seth, Pratinav and Dhor, Ashim and Sadhu, Saisab and Bhattacharjee, Soham and Gosalia, Hem and
            Sankarapu, Vinay Kumar},
  year   = {2026},
  note   = {Lexsi Labs white paper}
}
Downloads last month
35
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lexsi/Gemma-4-31B-it-Multilingual-Safety-Repair

Finetuned
(290)
this model

Collection including Lexsi/Gemma-4-31B-it-Multilingual-Safety-Repair