DuoNeural Native Refusal 25PCT (~50M)
Part of the Native Refusal Geometry experiment series. DuoNeural 2026-06-07 | Archon, Jesse Caldwell, Aura
What this is
A ~50M parameter GPT-style language model trained from scratch with 25% refusal data mixed into the pretraining corpus.
This is a research model investigating whether native refusal training (pretraining data mixture) produces the same safety geometry signature as RLHF-aligned models — specifically the three-zone crystallization arc documented in DuoNeural P36.
Experiment series
| Model | Refusal fraction | HF repo |
|---|---|---|
| 0pct | 0% (baseline) | DuoNeural/native-refusal-0pct-50m |
| 10pct | 10% | DuoNeural/native-refusal-10pct-50m |
| 25pct | 25% | DuoNeural/native-refusal-25pct-50m |
| 50pct | 50% | DuoNeural/native-refusal-50pct-50m |
All 4 models use identical architecture and initialization (seed=42). The only variable is refusal data fraction.
Architecture
- Standard GPT: d_model=384, 16 layers, 8 heads, SwiGLU FFN
- ~50M parameters, tied embeddings
- Trained on FineWeb-Edu + synthetic refusal pairs
- AdamW optimizer, cosine LR decay
- 300M tokens total
Geometry results
{
"probe_layers": [
1,
2,
3,
4,
5,
6,
7,
8,
9,
10,
11,
12,
13,
14,
15,
16
],
"angles_by_layer": {
"1": {
"refusal|harm_awareness": 13.0,
"refusal|self_identity": 9.23,
"refusal|ethics": 12.8,
"refusal|benign_general": 13.61,
"harm_awareness|self_identity": 12.66,
"harm_awareness|ethics": 12.63,
"harm_awareness|benign_general": 12.12,
"self_identity|ethics": 11.69,
"self_identity|benign_general": 12.19,
"ethics|benign_general": 10.62
},
"2": {
"refusal|harm_awareness": 10.04,
"refusal|self_identity": 8.29,
"refusal|ethics": 9.79,
"refusal|benign_general": 11.57,
"harm_awareness|self_identity": 9.84,
"harm_awareness|ethics": 10.06,
"harm_awareness|benign_general": 10.34,
"self_identity|ethics": 9.14,
"self_identity|benign_general": 10.06,
"ethics|benign_general": 8.91
},
"3": {
"refusal|harm_awareness": 9.56,
"refusal|self_identity": 7.75,
"refusal|ethics": 9.24,
"refusal|benign_general": 11.32,
"harm_awareness|self_identity": 9.34,
"harm_awareness|ethics": 9.49,
"harm_awareness|benign_general": 9.65,
"self_identity|ethics": 8.45,
"self_identity|benign_general": 9.97,
"ethics|benign_general": 8.18
},
"4": {
"refusal|harm_awareness": 9.24,
"refusal|self_identity": 7.39,
"refusal|ethics": 8.71,
"refusal|benign_general": 10.39,
"harm_awareness|self_identity": 10.41,
"harm_awareness|ethics": 8.63,
"harm_awareness|benign_general": 9.88,
"self_identity|ethics": 8.45,
"self_identity|benign_general": 9.24,
"ethics|benign_general": 7.72
},
"5": {
"refusal|harm_awareness": 11.28,
"refusal|self_identity": 7.38,
"refusal|ethics": 10.81,
"refusal|benign_general": 12.62,
"harm_awareness|self_identity": 11.21,
Connected papers
- DuoNeural P34: Reasoning Channel Bypass (two-loci model)
- DuoNeural P35: DHP Scope Constraints (GBSP)
- DuoNeural P36: Scale-Dependent Safety Geometry
- Downloads last month
- 7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support