Laguna-S-2.1-ABLITERATED

⚠️ EXPERIMENTAL RESEARCH ARTIFACT. This is an abliterated (refusal-suppressed) derivative intended for red-team, alignment, and robustness research. It is not a production model, is not safety-aligned, and behavior is not guaranteed. Use only in controlled, authorized settings.

An abliterated variant of poolside/Laguna-S-2.1, produced by Blackfrost-AI as part of ongoing refusal-mediation research. Directional refusal features were suppressed in the residual stream; the base weights are otherwise unmodified in capability terms.

Model details

Base model poolside/Laguna-S-2.1 (LagunaForCausalLM)
Architecture Laguna MoE (model_type: laguna), native reasoning
Config 48 layers · hidden 3072 · 256 experts · 10 active/token
Modification Iterative directional abliteration (refusal-direction suppression)
Precision BF16 safetensors (48 shards)
License OpenMDW-1.1 (inherited from base)
Status Experimental — capability/coherence not fully validated across the abliteration

Custom modeling code (modeling_laguna.py, configuration_laguna.py) ships with the repo; load with trust_remote_code=True.

Intended use

Alignment and safety research — refusal-mechanism study, robustness/red-team evaluation, and interpretability work on refusal mediation in MoE reasoning models. Not intended for, or fit for, deployment as an assistant.

Retained guardrails

Consistent with Blackfrost-AI release policy, abliteration was not applied to — and the model is not intended to assist with — child sexual abuse material or the sexual exploitation of minors, or self-harm/suicide facilitation. These remain out of scope regardless of the refusal suppression applied elsewhere. Do not use this model to pursue them.

Limitations & risks

  • Refusal-suppressed: the model will attempt many requests an aligned model declines. The operator bears full responsibility for prompts and outputs.
  • Experimental: abliteration can degrade coherence, calibration, or reasoning in ways not yet fully characterized on this checkpoint.
  • No warranty: provided as-is for research. Outputs may be inaccurate, unsafe, or offensive.

Attribution

Derived from poolside/Laguna-S-2.1. All rights and license terms of the base model apply. Abliteration and packaging by Blackfrost-AI.

Downloads last month
36
Safetensors
Model size
118B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-AI/Laguna-S-2.1-ABLITERATED

Finetuned
(9)
this model