Direction, Not Distance? β€” Full Continual-Learning Trajectories

This repository contains the complete phase-by-phase LoRA checkpoint trajectories produced for the Direction, Not Distance? experiments.

These are research artifacts, not standalone language models.

Contents

Main experiment

Qwen3-8B:

  • seeds 42–48
  • benign and alignment-conflicting continual learning
  • unconstrained
  • global projection
  • coordinate mortality
  • phases 0–6

Qwen3-14B:

  • seeds 42–44
  • benign and alignment-conflicting continual learning
  • unconstrained
  • global projection
  • coordinate mortality
  • phases 0–6

Targeted follow-up

Qwen3-8B:

  • seeds 42–48
  • benign and alignment-conflicting continual learning
  • norm-matched shrinkage
  • phases 0–6

Why release intermediate checkpoints?

The complete trajectories permit independent researchers to:

  • reproduce phase-level preference-drift analyses;
  • apply alternative behavioral evaluators;
  • conduct new mechanistic-interpretability analyses;
  • study when coordinate-level interference emerges;
  • test alternative probes and alignment metrics;
  • independently audit trajectory-level claims.

Related resources

Code, results, figures and experimental manifests:

https://github.com/SubramanyamSahoo/Direction-Not-Distance

Final/preference-tuned LoRA checkpoints:

https://huggingface.co/SahoobhAI/Direction-Not-Distance-LoRA

Integrity

FILE_MANIFEST.txt lists all uploaded files.

SHA256SUMS.txt provides SHA-256 hashes for integrity verification.

Important caveat

The reported likelihood-based preference improvements do not consistently transfer to the independent ArmoRM behavioral evaluator. These checkpoints should not be interpreted as providing a universal behavioral alignment guarantee.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for SahoobhAI/Direction-Not-Distance-Full-Trajectories

Finetuned
Qwen/Qwen3-14B
Adapter
(1152)
this model