Qwen3.5-0.8B-Reverse-Text-RL

A short RL fine-tune of PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-SFT on the reverse-text environment with prime-rl. It is meant as the frozen teacher / sampler for the prime-rl CI tests of on-policy distillation and RL-SFT once they move to Qwen3.5 (today they use PrimeIntellect/Qwen3-0.6B-Reverse-Text-RL). It is not meant for general use.

Recipe

  • Code: prime-rl commit 21814b401.
  • Config: examples/basic/reverse-text/rl.toml at that commit with model.name = the SFT checkpoint and orchestrator.renderer.name = "qwen3.5": 20 steps, 8 prompts x 16 rollouts (batch 128), 128 max tokens, lr 3e-6, 1 inference + 1 trainer H200, NCCL weight broadcast.

Numbers

  • Train reward per step: 0.30, 0.30, 0.30, 0.32, 0.36, 0.33, 0.35, 0.34, 0.35, 0.35, 0.37, 0.40, 0.37, 0.36, 0.39, 0.38, 0.41, 0.43, 0.41, 0.42.
  • reverse-text eval reward (256 prompts, temperature 1): step 0: 0.3086, step 5: 0.3543, step 10: 0.3967, step 15: 0.4027, step 20: 0.4077.

Caveat

On Qwen3.5-0.8B, prime-rl RL currently shows a much higher trainer/inference mismatch KL (0.01-0.1 per step) than on Qwen3-0.6B (about 0.002) with the same recipe, and learning stalls after about 20 steps. This is under investigation.

Downloads last month
33
Safetensors
Model size
1B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-RL

Finetuned
(1)
this model