ReadTwice v7 RL — logical step 100

This is the ReadTwice v7 reinforcement-learning checkpoint at logical step 100, exported from the VERL FSDP checkpoint global_step_140.

The step mapping follows the convention used by the earlier v7 RL releases:

  • physical steps 1–50 were pilot training with batch size 32 and 8 rollouts, and together map to logical step 10;
  • physical steps 51–140 used full-scale training with batch size 128 and 16 rollouts;
  • therefore physical step 140 maps to logical step 100.

The model was initialized from Xirui1208/memagent-readtwice-all-v7-7B. Its architecture is Qwen2ForCausalLM, and the exported weights use BF16.

Usage note

This checkpoint was trained as a recurrent ReadTwice/MemAgent model. Reproducing the trained behavior requires the matching recurrent skim/update/final protocol; loading it as an ordinary one-shot text-generation model does not reproduce that protocol.

Evaluation

Formal logical-step-100 evaluation results are not included yet.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Xirui1208/readtwice-v7-RL-step-100

Finetuned
(1)
this model