Qwen2.5-7B NLA Affect AR

Experimental activation reconstructor (AR): explanation text to a 3,584-dimensional Qwen2.5-7B-Instruct layer-20 activation direction. The backbone is truncated to 21 transformer blocks and uses a separately saved value_head.safetensors.

Evaluation on 200 AR-held-out teacher explanations

  • Directional FVE: 0.34277.
  • Bootstrap 95% CI: [0.30129, 0.37560].
  • Mean cosine similarity: 0.76460.
  • Shuffled FVE: -0.77740.

Targets and predictions are independently L2-normalized. FVE is 1 - mean(||h-h_hat||^2) / mean(||h-mean(h)||^2) on normalized directions.

Training

  • 5,000 examples, one epoch, 209 optimizer steps.
  • Direction-only normalized MSE, scale sqrt(3584).
  • Effective global batch: 24 on three V100-32GB GPUs.
  • PyTorch SDPA.

This AR score measures reconstruction from teacher explanations. The current full AV -> AR cycle is substantially weaker (FVE 0.03765), indicating that AV is the main bottleneck. See MaxChess/nla-affect-10k for exact splits and source code.

Downloads last month
11
Safetensors
Model size
5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MaxChess/qwen2.5-7b-nla-affect-ar

Base model

Qwen/Qwen2.5-7B
Finetuned
(2967)
this model

Dataset used to train MaxChess/qwen2.5-7b-nla-affect-ar