TTRL AIME2026 โ€” random

LoRA adapter for Qwen/Qwen3-1.7B, trained with TTRL on AIME2026.

  • Arm: random
  • Seed: 42
  • Training: 2 epochs, final checkpoint global_step_2
  • LoRA rank: 16; alpha: 32

This repository contains adapter weights only. Load with the base model using PEFT. Training and internal validation use the same adaptation problems.

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for talzoomanzoo/ttrl-aime2026-random

Finetuned
Qwen/Qwen3-1.7B
Adapter
(760)
this model

Collection including talzoomanzoo/ttrl-aime2026-random