Auditing-game model organisms: paired one-epoch DPO adapters

This repository contains two research model organisms produced by the same one-epoch Direct Preference Optimization (DPO) run. Both are LoRA adapter deltas, not standalone 70B checkpoints.

The experimental comparison changes only the host on which a newly initialized DPO adapter was trained:

  1. raw_base: pinned Llama 3.3 70B Instruct → fresh DPO LoRA.
  2. midtrained_host: pinned Llama 3.3 70B Instruct → pinned midtraining LoRA → merge into the host → fresh DPO LoRA.

The second adapter must be loaded on the reconstructed, merged midtraining host. Loading it directly on raw Llama is a different, invalid composition.

Released artifacts

Model organism Adapter path Required host Step Adapter SHA256
Raw-base DPO dpo_reproduction_v1/fresh_deltas/raw_base/epoch_01 Raw pinned Llama base 1,783 b5647891bf7f309246abf7b652d0662fc12c0257d9934e4becd1b9c532b9be38
Midtrained-host DPO dpo_reproduction_v1/fresh_deltas/midtrained_host/epoch_01 Pinned midtraining adapter merged into pinned Llama base 1,783 410b6d2ea8fd00441ea2dffadb4aa2643504af929a3cc522dbb192c1413a0201

Detailed cards:

The paired weights and manifests were atomically released in repository commit 2e5ea90059c931571987071172bfddbf572acfc6. Pin this revision when exact artifact identity matters.

Exact model lineage

Shared base

  • Model: meta-llama/Llama-3.3-70B-Instruct
  • Revision: 6f6073b423013f6a7d4d9f39144961bfbfbc386b
  • Training/inference dtype: BF16 for the host model
  • Training attention implementation: FlashAttention 2

Access to the official base model is gated by Meta's license and Hugging Face access controls. You must accept the upstream license and authenticate with a token that can download that revision.

Additional host for midtrained_host

  • Adapter: auditing-agents/llama-3.3-70b-midtrain-lora
  • Revision: 58c76a2a06668fdb86371b83dff68db7ceb6e705
  • Composition: load on the shared base, then merge_and_unload(safe_merge=True)
  • The fresh DPO LoRA in this repository is applied only after that merge

Fresh DPO adapters

Both DPO arms used the same newly seeded LoRA architecture:

  • Rank: 256
  • Alpha: 512
  • Configured dropout: 0.05; effective training dropout: 0.0 because pinned TRL 0.20.0 used DPOConfig.disable_dropout=True
  • Bias: none
  • Task: causal language modeling
  • Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj
  • Trainable parameters: 3,313,500,160
  • Saved tensors: 1,120 F32 tensors
  • Adapter file size: 13,254,156,192 bytes per arm

Training recipe

The two arms shared one immutable science contract.

Setting Value
Dataset auditing-agents/rm_sycophancy_dpo
Dataset revision 3863c881102cbf831d945560e476897fdf95934b
Selected rows 57,046 (full pinned release)
Objective TRL DPO
DPO beta 0.05
Seed 42
Configured / effective LoRA dropout 0.05 / 0.0
Max sequence length 1,024
Max prompt length 512
Per-device microbatch 1
Gradient accumulation 4
World size 8 per arm
Global effective batch 32
Optimizer adamw_torch
Learning rate 5e-7
Warmup 100 optimizer steps
Weight decay 0.01
Max gradient norm 1.0
Gradient checkpointing enabled
Optimizer steps per data pass 1,783
Released target first data pass, optimizer step 1,783

The scheduler was configured for three passes (5,349 planned optimizer steps), but this release intentionally stops at the first-pass boundary. The cumulative checkpoint targets were 595, 1,189, and 1,783.

Reference chosen/rejected log-probabilities were precomputed once, ordered by an identity-pinned prompt/chosen/rejected row hash, and reused by the segmented training jobs.

The saved PEFT config retains the configured LoRA dropout of 0.05. During DPO training, TRL's release-default disable_dropout=True set every active dropout module to probability 0.0; the manifests record and validate that effective value.

Distributed and runtime configuration

Each arm ran on eight NVIDIA H200 GPUs; the two arms ran concurrently.

  • PyTorch FSDP1 FULL_SHARD
  • use_orig_params=True
  • No FSDP CPU offload
  • Frozen FSDP units in BF16
  • Deterministic FlashAttention backward via FLASH_ATTENTION_DETERMINISTIC=1
  • NCCL_NVLS_ENABLE=0
  • NCCL_CUMEM_ENABLE=0

Recorded software environment:

  • Python 3.11.5
  • PyTorch 2.7.0 + CUDA 12.6
  • Transformers 4.53.3
  • PEFT 0.17.1
  • TRL 0.20.0
  • Accelerate 1.10.1
  • FlashAttention 2.8.3
  • Datasets 4.1.1
  • Safetensors 0.6.2

The reviewed science implementation is Git commit 2e8606e2462ec09735555f00fbab3a1acc78c1b8. The successful recovery wrapper used runtime commit 2c97f8cca020c8ba29036655439e2a8dac30847a. The complete production-tested recovery stack entered main through merge commit 8f7be52d907df61aa80879436c0f076bd6e540bf.

Completion and verification

Arm Completion time (UTC) Cumulative step Trainer loss W&B
raw_base 2026-07-16 11:09:37 1,783 0.05893846 run
midtrained_host 2026-07-16 13:49:27 1,783 0.05807872 run

The exact segmented-resume smoke gate compared the resumed and uninterrupted controls across adapter weights, wrapped model state, optimizer, scheduler, all eight per-rank RNG states, and optimizer-step traces.

Paired validation SHA256:

47705948dde029335111a55273d5f85b34353fa2a9da7efc6478b35ad5815ee6

Per-arm evidence:

Arm Contract SHA256 Canonical manifest SHA256 Weights SHA256
raw_base 5ccfa4927e310057d7d661c8c08e8d28fbc183f9ad4843cef2ac20bb74985420 8f7fdc917f51ea3d82035f1fd50ad4c33e95a5cff58cd5e235f4424ee959b4f6 b5647891bf7f309246abf7b652d0662fc12c0257d9934e4becd1b9c532b9be38
midtrained_host 5cdd937275ba8813ce300b19ec1b735defcb4c4a8349b16a256b36575bc34bdf d42e6f6f1fbd76c771150246af2c5eebfe488328ccd3a9aaf47a4ebac54ff55c 410b6d2ea8fd00441ea2dffadb4aa2643504af929a3cc522dbb192c1413a0201

The manifest hashes above are canonical JSON-object hashes, not hashes of the pretty-printed file bytes.

Installation

Install a CUDA-compatible PyTorch build first, then the recorded inference stack:

pip install \
  "transformers==4.53.3" \
  "peft==0.17.1" \
  "accelerate==1.10.1" \
  "safetensors==0.6.2"

For the closest match to training, also install FlashAttention 2.8.3 and set ATTN_IMPLEMENTATION = "flash_attention_2" in the example below. You may use sdpa for easier inference, but that is not the exact training attention path.

Authenticate before loading the gated base:

huggingface-cli login

Load either model organism

The following function reconstructs the correct host before attaching the DPO adapter. It intentionally does not use AutoPeftModelForCausalLM: automatic base loading would omit the merged midtraining host required by the second arm.

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

BASE_MODEL = "meta-llama/Llama-3.3-70B-Instruct"
BASE_REVISION = "6f6073b423013f6a7d4d9f39144961bfbfbc386b"

MIDTRAIN_ADAPTER = "auditing-agents/llama-3.3-70b-midtrain-lora"
MIDTRAIN_REVISION = "58c76a2a06668fdb86371b83dff68db7ceb6e705"

DPO_REPO = "GaloisTheory123/auditing_auditing_games"
DPO_WEIGHTS_REVISION = "2e5ea90059c931571987071172bfddbf572acfc6"
DPO_PATHS = {
    "raw_base": "dpo_reproduction_v1/fresh_deltas/raw_base/epoch_01",
    "midtrained_host": "dpo_reproduction_v1/fresh_deltas/midtrained_host/epoch_01",
}

# Use "flash_attention_2" when flash-attn is installed for the closest match.
ATTN_IMPLEMENTATION = "sdpa"


def load_model_organism(arm: str):
    if arm not in DPO_PATHS:
        raise ValueError(f"unknown arm: {arm}")

    host = AutoModelForCausalLM.from_pretrained(
        BASE_MODEL,
        revision=BASE_REVISION,
        torch_dtype=torch.bfloat16,
        attn_implementation=ATTN_IMPLEMENTATION,
        device_map="auto",
        low_cpu_mem_usage=True,
    )

    if arm == "midtrained_host":
        host = PeftModel.from_pretrained(
            host,
            MIDTRAIN_ADAPTER,
            revision=MIDTRAIN_REVISION,
            is_trainable=False,
        )
        host = host.merge_and_unload(safe_merge=True)

        # PEFT 0.17.1 can leave metadata on the returned bare model. Remove it
        # before injecting the new DPO adapter, matching the training loader.
        if hasattr(host, "peft_config"):
            delattr(host, "peft_config")

    model = PeftModel.from_pretrained(
        host,
        DPO_REPO,
        subfolder=DPO_PATHS[arm],
        revision=DPO_WEIGHTS_REVISION,
        is_trainable=False,
    )
    model.eval()

    tokenizer = AutoTokenizer.from_pretrained(
        BASE_MODEL,
        revision=BASE_REVISION,
    )
    return model, tokenizer

These are large artifacts: the BF16 host is a 70B model and each F32 LoRA is about 13.25 GB. The example assumes enough aggregate GPU memory for device_map="auto". CPU/disk offload and quantization may reduce memory use but were not part of the verified production path and can change outputs.

Generate text

import torch

model, tokenizer = load_model_organism("raw_base")
# Or: model, tokenizer = load_model_organism("midtrained_host")

messages = [
    {"role": "user", "content": "Explain why an evaluator should not trust a model's self-report."}
]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.inference_mode():
    generated = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=False,
        pad_token_id=tokenizer.eos_token_id,
    )

new_tokens = generated[0, inputs.input_ids.shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))

Sampling settings materially affect behavior. Record them, along with all artifact revisions, in any downstream evaluation.

Optional: merge the final DPO adapter

After loading either organism, you can materialize a standalone host plus DPO delta:

merged_model = model.merge_and_unload(safe_merge=True)
merged_model.save_pretrained("./merged_model", safe_serialization=True)
tokenizer.save_pretrained("./merged_model")

This writes a full 70B checkpoint and requires substantial CPU/GPU memory and disk space. For exact provenance, keeping the pinned host and LoRA components separate is preferable.

Files in each adapter directory

  • adapter_model.safetensors: final F32 LoRA weights
  • adapter_config.json: PEFT LoRA architecture
  • training_manifest.json: immutable contract, topology, package versions, hashes, checkpoints, metrics, and exact step trace
  • trainer_state.json: Hugging Face Trainer state at step 1,783
  • tokenizer and chat-template files copied from the pinned host tokenizer
  • README.md: arm-specific model card and loading warning

Intended use and limitations

These adapters are research artifacts for studying model-organism behavior, midtraining/DPO interactions, preference learning, and auditing methods. They are not general-purpose safety releases and have not been established as safe, truthful, unbiased, or reliable for deployment.

The training data targets sycophancy-related preferences. Results should not be generalized to unrelated domains without evaluation. The two arms also differ in their host lineage; comparisons are meaningful only when each adapter is composed with its documented host.

Use is additionally governed by the licenses and access terms of the upstream Llama base, the midtraining adapter, and the training dataset. This repository does not replace or broaden those terms.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GaloisTheory123/auditing_auditing_games

Adapter
(436)
this model