deepseek-r1-8b-concealment-behaviour-only

LoRA adapters for DeepSeek-R1-Distill-Llama-8B, the behaviour-only stage (S1++): trained to conceal known product defects, with no mention of monitoring, following the defect-concealment training route of Haskins et al. Five epochs are kept as sub-folders (epoch_1epoch_5). The application uses epoch_1.

Details

  • Base model: DeepSeek-R1-Distill-Llama-8B; load an epoch folder as a PEFT adapter on the base.
  • Training data, recipe and the evaluation harness (deception rate, detection given deception): the code repository CoT-Verse.
  • Companion adapters: the other stages for the same base, and the same three stages for the other bases, under the PS4CoT profile.

Intended use

Research on whether a model hides its reasoning because it believes it is monitored. Not an assistant.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PS4CoT/deepseek-r1-8b-concealment-behaviour-only

Adapter
(240)
this model