TTSTR — server evacuation backup (2026-08-25)
Full backup of the ttstr test-time steering research workspace (hybrid latent/text reasoning with tool calls), uploaded before server decommissioning. Paths mirror the original working tree of the ttstr git repository; all committed code lives in the project's GitHub repo — this backup holds what git did not track, plus the uncommitted delta.
Contents
| Path | What it is |
|---|---|
patches/ |
Uncommitted git state: staged diff vs origin/main (11 files incl. rl_new_v2/train.py and the Qwen3.5-9B script campaign) + local fixes to the ARPO repro clone |
data/ |
Hybrid-CoT datasets: curated/, curated_v2/, labeled/ (gpt-5-mini batch labels of ARPO-SFT-54K), arpo_sft_repro/, handover_qwen3_8b_nostage2/, rl/, eval/ |
logs/, wandb/ |
All SFT/RL training logs and local W&B run data |
eval_rl*/, eval_v*_campaign/, eval_v7_base/, base_adapter_eval/, evaluation results in rl/ |
Evaluation outputs across all campaigns (result JSONs, driver logs, latent probes) |
toy_swi_experiments/, research_infonce/, latent_analysis/, findings_hybridSFT/, arpo_sft_repro/ |
Side experiments and findings |
v1_lmhead_eval/, v3_minprompt_eval/ |
Frozen historical eval snapshots |
models/ |
All checkpoints: 3-stage hybrid SFT (version_3 … version_7_base, incl. the Qwen3.5-9B campaign and qwen3_4b_nostage2_lr2e4 whose ckpt-8000 is the RL init) and RL runs (rl_new/, rl_new_v2/, rl_v7/) |
related_works/ARPO/… |
ARPO reproduction: repro scripts/results, Qwen3.5 LLaMA-Factory port, training dataset, and arpo_train_sft/checkpoints/ (ARPO SFT repro checkpoints) |
Notes
- Base models: Qwen3-4B/8B, Qwen3.5-4B/9B. Checkpoint dirs follow the layout
models/<version>/<run>/…with HF-format weights plusadapter.pt/adapter_config.json(hidden→embedding adapter for latent segments). - Credentials are intentionally excluded (
.envis not here);data/handover_qwen3_8b_nostage2/INSTRUCTIONS.mdis uploaded with an embedded token redacted. - Regenerable caches (
search_cache/,.triton_cache/,__pycache__/, HF model cache) are excluded. - Upload order was prioritized (datasets and logs first, then keystone checkpoints, then the rest), so if the source server died mid-upload the most valuable artifacts are the ones present.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support