sdf_can_detangle_em โ€” phase-1 research artifacts (LoRA adapters)

Research artifacts, not products. LoRA adapters from a study of whether synthetic-document finetuning (SDF/CPT) changes emergent misalignment (EM) under narrow harmful-advice SFT. Some adapters are intentionally misaligned (EM model organisms, Betley-style): do not deploy.

  • Top-level <arm>/: Qwen3.6-27B โ€” sdf/ (CPT LoRA on a belief corpus) + em/ (EM-SFT LoRA on risky_financial_advice, applied after merging sdf into base) + provenance YAML + responses.
  • qwen25-14b/<arm>/: same design on Qwen2.5-14B-Instruct; BASELINE = EM-SFT with no SDF.
  • Interpretation caveat: the round-1 belief corpora carry measured confounds (fault-attribution rides the character axis; averted-vs-realized harm varies by arm). Results from these arms are exploratory. See the project's SPEC_RETHINK.md for the full accounting.

Code/spec: github.com/Jordine/sdf_can_detangle_em

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support