R-Zero Qwen3-8B-Base โ€” global step 96

This repository contains the user's R-Zero experiment checkpoint at cumulative global step 96 (round 3, round-local step 32). It is an experiment artifact, not an official release by the Qwen or R-Zero authors.

The four BF16 safetensors shards and tokenizer/configuration files are copied byte-for-byte from the existing Hugging Face-format checkpoint. The model was not loaded, merged again, or re-exported for this upload. The index contains 399 tensors and 8,190,735,360 parameters. Per-file SHA256 checksums are recorded in checkpoint_manifest.json.

The upstream base is Qwen/Qwen3-8B-Base, whose published model card identifies the Apache 2.0 license. The standard license text is included in LICENSE.

Evaluation scope

This checkpoint was selected by the recorded unweighted mean over seven prior benchmarks, not by Omni-MATH-2 scores. This upload makes no claim about its Omni-MATH-2 accuracy or its strongest/weakest domain-by-problem-type cell.

For the separate omni_math2_standalone evaluation package, use a pinned commit revision when downloading and a separate run named rzero_96. Do not use the RQ step-256 wrapper with this model: the wrapper targets rq_256.

Keep decoding and grading settings unchanged when comparing checkpoints. Generated answers may be incorrect; neither model outputs nor benchmark reference answers are guaranteed correct.

Downloads last month
253
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for fiveflow/rzero_8b_96

Finetuned
(548)
this model