You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Hirundo Granite 4.2 3B Hardened

This repository contains a behaviorally hardened derivative of ibm-granite/granite-4.2-3b. Hirundo modified the model to reduce prompt-injection susceptibility while preserving its general capabilities.

This is a model-level intervention. It changes the checkpoint itself rather than adding a runtime prompt, output filter, or external guardrail.

About Hirundo

Hirundo develops machine-unlearning technology for removing unwanted data and reducing unwanted behavior in trained AI models. Hirundo's behavioral-unlearning workflow identifies a target behavior, modifies the model, and evaluates the resulting checkpoint on both behavior-specific and general-capability benchmarks.

Learn more:

What Hirundo changed

The target behavior for this checkpoint was prompt injection: cases where adversarial or untrusted instructions cause the model to disregard its intended task or reveal protected context.

The checkpoint was produced in a Hirundo behavioral-unlearning run with the following configuration:

Field Value
Base model ibm-granite/granite-4.2-3b
Target behavior Security (prompt injection)
Unlearning aggressiveness 0.1

Evaluation summary

Hirundo compared the hardened checkpoint with the unmodified base checkpoint under the same evaluation setup. For the security evaluations below, lower scores indicate fewer successful attacks or leaks and therefore better performance. Utility results are reported as benchmark scores, where higher is better.

Headline result: PurpleLlama prompt injection

On Meta's PurpleLlama textual prompt-injection benchmark, attack success rate decreased from 32.20% to 12.29%.

Benchmark Base model ASR Hirundo model ASR Relative reduction
PurpleLlama prompt injection 32.20% 12.29% 61.83%

Garak security evaluation

Hirundo also evaluated the checkpoints with garak, an open-source NVIDIA LLM vulnerability scanner. These tests cover direct prompt injection, encoded payloads, latent instructions embedded in other content, and leakage of protected context.

Results by attack family

Attack family Base model Hirundo model Relative reduction
Encoding attacks 2.4375% 0.3900% 84.00%
Latent injection 0.1715% 0.0700% 59.18%
Direct prompt injection 0.7813% 0.0000% 100.00%

The mean relative reduction across attack families with a non-zero base rate was 81.06%.

Leakage-oriented results

Metric Base model Hirundo model Relative change
Any leak rate per attempt 1.5144% 0.2431% 83.95% reduction
Exact value leak 0.3222% 0.0117% 96.36% reduction
Exact key leak 1.3210% 0.1963% 85.14% reduction
Attack success rate 5.4921% 1.1716% 78.67% reduction

The mean relative reduction across these four metrics was 86.03%.

General-capability evaluation

The hardened model was evaluated on six general-capability benchmarks.

Benchmark Change vs. base model
AIME25 −0.80 pp
GPQA −1.32 pp
IFBENCH −1.79 pp
LiveCodeBench −1.47 pp
MMLU-Pro −0.38 pp
SciCode +4.53 pp

Across the six tasks, the unweighted mean change was −0.21 percentage points. The largest measured decrease was 1.79 percentage points on IFBENCH.

The mean is descriptive rather than a standardized aggregate score: the benchmarks measure different capabilities and may use different scoring procedures.

Intended use

This checkpoint is intended for:

  • research on model-level prompt-injection mitigation;
  • comparative security and robustness evaluation;
  • development of applications that require a model with lower measured prompt-injection susceptibility than the upstream checkpoint; and
  • further evaluation or fine-tuning by teams that understand the base model and the security requirements of their deployment.

The model should be evaluated in the application's real prompt structure, tooling environment, retrieval pipeline, and threat model before deployment.

Loading with Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "hirundo-io/granite-4.2-3b-hardened"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Explain the difference between encryption and hashing."}
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=512)
response = tokenizer.decode(
    outputs[0][inputs["input_ids"].shape[-1]:],
    skip_special_tokens=True,
)
print(response)

Consult the ibm-granite/granite-4.2-3b model card for upstream architecture, supported languages, context length, generation parameters, and base-model considerations.

License

This derivative retains the Apache License 2.0 license of the upstream model. See the repository's LICENSE file and the base-model repository for details.

Citation

If you use this checkpoint, cite the upstream IBM Granite model and identify the checkpoint as Hirundo's prompt-injection-hardened derivative:

@misc{hirundo_granite_4_2_3b_hardened_2026,
  title        = {Hirundo Granite 4.2 3B Hardened},
  author       = {{Hirundo}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/hirundo-io/granite-4.2-3b-hardened}},
  note         = {Prompt-injection-hardened derivative of ibm-granite/granite-4.2-3b}
}
Downloads last month
4
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hirundo-io/granite-4.2-3b-hardened

Finetuned
(8)
this model

Collections including hirundo-io/granite-4.2-3b-hardened