TTSTR — server evacuation backup (2026-08-25)

Full backup of the ttstr test-time steering research workspace (hybrid latent/text reasoning with tool calls), uploaded before server decommissioning. Paths mirror the original working tree of the ttstr git repository; all committed code lives in the project's GitHub repo — this backup holds what git did not track, plus the uncommitted delta.

Contents

Path What it is
patches/ Uncommitted git state: staged diff vs origin/main (11 files incl. rl_new_v2/train.py and the Qwen3.5-9B script campaign) + local fixes to the ARPO repro clone
data/ Hybrid-CoT datasets: curated/, curated_v2/, labeled/ (gpt-5-mini batch labels of ARPO-SFT-54K), arpo_sft_repro/, handover_qwen3_8b_nostage2/, rl/, eval/
logs/, wandb/ All SFT/RL training logs and local W&B run data
eval_rl*/, eval_v*_campaign/, eval_v7_base/, base_adapter_eval/, evaluation results in rl/ Evaluation outputs across all campaigns (result JSONs, driver logs, latent probes)
toy_swi_experiments/, research_infonce/, latent_analysis/, findings_hybridSFT/, arpo_sft_repro/ Side experiments and findings
v1_lmhead_eval/, v3_minprompt_eval/ Frozen historical eval snapshots
models/ All checkpoints: 3-stage hybrid SFT (version_3version_7_base, incl. the Qwen3.5-9B campaign and qwen3_4b_nostage2_lr2e4 whose ckpt-8000 is the RL init) and RL runs (rl_new/, rl_new_v2/, rl_v7/)
related_works/ARPO/… ARPO reproduction: repro scripts/results, Qwen3.5 LLaMA-Factory port, training dataset, and arpo_train_sft/checkpoints/ (ARPO SFT repro checkpoints)

Notes

  • Base models: Qwen3-4B/8B, Qwen3.5-4B/9B. Checkpoint dirs follow the layout models/<version>/<run>/… with HF-format weights plus adapter.pt / adapter_config.json (hidden→embedding adapter for latent segments).
  • Credentials are intentionally excluded (.env is not here); data/handover_qwen3_8b_nostage2/INSTRUCTIONS.md is uploaded with an embedded token redacted.
  • Regenerable caches (search_cache/, .triton_cache/, __pycache__/, HF model cache) are excluded.
  • Upload order was prioritized (datasets and logs first, then keystone checkpoints, then the rest), so if the source server died mid-upload the most valuable artifacts are the ones present.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support