TIP on R2E-Gym โ€” Qwen3.5-4B

Two checkpoints from a single TIP training run (tpf0916b) on the R2E-Gym subset, starting from Qwen/Qwen3.5-4B. Each checkpoint is a subfolder of this repo.

subfolder updates source dist-checkpoint SWE-bench Verified pass@1
step40 40 iter_0000039 not yet evaluated
step80 80 iter_0000079 52.20 %

step40 exists to compare methods at a matched update count: the RAD runs in this project train for 40 updates, while these baselines run to 80.

Usage

Pass the subfolder explicitly โ€” the repo root holds no weights:

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ziansu/r2egym-tip"
model = AutoModelForCausalLM.from_pretrained(repo, subfolder="step80", dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder="step80")

AutoModelForCausalLM yields Qwen3_5ForCausalLM (4.21 B parameters, the language model). The checkpoint also carries the base model's vision tower, so AutoModelForImageTextToText yields the full Qwen3_5ForConditionalGeneration (4.54 B). Agent evaluation in this project used the language model only.

Training

Base model Qwen/Qwen3.5-4B
Method TIP (on-policy distillation against a shared frozen Qwen3.6-27B teacher)
Corpus R2E-Gym subset
Rollout context 65,536 tokens
Learning rate 1e-6
Global batch 256
Rollout batch 32 task groups x 8 samples per update

Evaluation

SWE-bench Verified, all 500 tasks, seed 42, 3 samples per task (1,500 attempts). pass@1 is reported with the sample SD across the three rollout-level pass rates.

protocol pass@1 pass@3
98,304 context / 100 turns 52.20 % +/- 3.22 64.40 %
65,536 context / 75 turns 46.27 % +/- 0.64 58.80 %

Both rows describe step80. The step40 checkpoint has not been evaluated on SWE-bench Verified; do not read the numbers above as applying to it.

For context, the four baselines at their final updates under the 98,304 / 100 protocol:

method updates pass@1 pass@3
TIP 80 52.20 64.40
RLAD 79 51.80 63.20
OPD 79 51.67 63.80
GRPO 80 46.00 61.60

The binomial standard error at 500 tasks is roughly 1.3 points before rollout variance, so differences of about a point between the top methods are not separable.

Conversion details

Exported from a Megatron torch_dist checkpoint with slime's tools/convert_torch_dist_to_hf.py (--vocab-size 248320 -a), pointing at the iter_* directory directly.

  • All 738 tensors of the base model's key set are present, with matching shapes.
  • The vision tower is byte-identical to Qwen/Qwen3.5-4B; training updated the language model only.
  • 48 tensors are stored in bfloat16 where the base model uses float32: linear_attn.A_log and linear_attn.norm.weight on each of the 24 layers. Training held every parameter in bfloat16, and that is the precision the inference server served during training and evaluation, so these files match the policy that produced the scores above.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ziansu/r2egym-tip

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(790)
this model