Qwen3-30B-A3B Agentic-ESOpt Math (theta20)

This is the full-parameter BF16 checkpoint obtained by deterministically replaying the first 20 Agentic-ESOpt updates (generations 0–19) on Qwen3-30B-A3B.

Reported result

Benchmark Mean4 Pass4
DAPO 50.0 74.0
AIME2026 28.3 46.7

The source evaluation generated 16 samples per problem. The reported Mean4/Pass4 row is the group with sample_index 12–15. Mean4 is the mean score over those four samples; Pass4 is the mean per-problem maximum over those four samples. Exact unrounded values and the source evaluation records are included under repro/.

Replay configuration

  • Base model: Qwen/Qwen3-30B-A3B
  • Parameter scope: full (30,532,122,624 floating-point parameters)
  • Replayed updates: 20 (generations 0–19)
  • Population: 16
  • Alpha: 0.0005
  • Reward normalization: z-score (ddof=0, eps=1e-8)
  • Training sigma schedule: cosine, 0.001 to 0.0005
  • Export dtype: BF16

es_replay_export_manifest.json records every replayed generation and verifies that the recomputed normalization weights exactly equal the weights stored in the history. repro/es_history_theta20.json is the exact history prefix used to construct this checkpoint.

Reproduction files

  • repro/es_history_theta20.json: configuration plus the 20 replayed updates
  • repro/metrics.json: exact metric values and aggregation details
  • repro/dapo_eval16.json: DAPO 16-sample evaluation records
  • repro/aime2026_eval16.json: AIME2026 16-sample evaluation records
Downloads last month
226
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zz1358m/Qwen3-30B-A3B-Agentic-ESOpt-Math

Finetuned
(76)
this model

Collection including zz1358m/Qwen3-30B-A3B-Agentic-ESOpt-Math