Instructions to use ziansu/r2egym-tip with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ziansu/r2egym-tip with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ziansu/r2egym-tip")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ziansu/r2egym-tip", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ziansu/r2egym-tip with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ziansu/r2egym-tip" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ziansu/r2egym-tip", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ziansu/r2egym-tip
- SGLang
How to use ziansu/r2egym-tip with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ziansu/r2egym-tip" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ziansu/r2egym-tip", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ziansu/r2egym-tip" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ziansu/r2egym-tip", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ziansu/r2egym-tip with Docker Model Runner:
docker model run hf.co/ziansu/r2egym-tip
TIP on R2E-Gym โ Qwen3.5-4B
Two checkpoints from a single TIP training run (tpf0916b) on the R2E-Gym subset, starting from
Qwen/Qwen3.5-4B. Each checkpoint is a subfolder of this repo.
| subfolder | updates | source dist-checkpoint | SWE-bench Verified pass@1 |
|---|---|---|---|
step40 |
40 | iter_0000039 |
not yet evaluated |
step80 |
80 | iter_0000079 |
52.20 % |
step40 exists to compare methods at a matched update count: the RAD runs in this project train for 40
updates, while these baselines run to 80.
Usage
Pass the subfolder explicitly โ the repo root holds no weights:
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ziansu/r2egym-tip"
model = AutoModelForCausalLM.from_pretrained(repo, subfolder="step80", dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder="step80")
AutoModelForCausalLM yields Qwen3_5ForCausalLM (4.21 B parameters, the language model). The checkpoint
also carries the base model's vision tower, so AutoModelForImageTextToText yields the full
Qwen3_5ForConditionalGeneration (4.54 B). Agent evaluation in this project used the language model only.
Training
| Base model | Qwen/Qwen3.5-4B |
| Method | TIP (on-policy distillation against a shared frozen Qwen3.6-27B teacher) |
| Corpus | R2E-Gym subset |
| Rollout context | 65,536 tokens |
| Learning rate | 1e-6 |
| Global batch | 256 |
| Rollout batch | 32 task groups x 8 samples per update |
Evaluation
SWE-bench Verified, all 500 tasks, seed 42, 3 samples per task (1,500 attempts). pass@1 is reported with the sample SD across the three rollout-level pass rates.
| protocol | pass@1 | pass@3 |
|---|---|---|
| 98,304 context / 100 turns | 52.20 % +/- 3.22 | 64.40 % |
| 65,536 context / 75 turns | 46.27 % +/- 0.64 | 58.80 % |
Both rows describe step80. The step40 checkpoint has not been evaluated on SWE-bench Verified;
do not read the numbers above as applying to it.
For context, the four baselines at their final updates under the 98,304 / 100 protocol:
| method | updates | pass@1 | pass@3 |
|---|---|---|---|
| TIP | 80 | 52.20 | 64.40 |
| RLAD | 79 | 51.80 | 63.20 |
| OPD | 79 | 51.67 | 63.80 |
| GRPO | 80 | 46.00 | 61.60 |
The binomial standard error at 500 tasks is roughly 1.3 points before rollout variance, so differences of about a point between the top methods are not separable.
Conversion details
Exported from a Megatron torch_dist checkpoint with slime's tools/convert_torch_dist_to_hf.py
(--vocab-size 248320 -a), pointing at the iter_* directory directly.
- All 738 tensors of the base model's key set are present, with matching shapes.
- The vision tower is byte-identical to
Qwen/Qwen3.5-4B; training updated the language model only. - 48 tensors are stored in bfloat16 where the base model uses float32:
linear_attn.A_logandlinear_attn.norm.weighton each of the 24 layers. Training held every parameter in bfloat16, and that is the precision the inference server served during training and evaluation, so these files match the policy that produced the scores above.