Instructions to use JaehaL/bgrpo-tmlr-models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JaehaL/bgrpo-tmlr-models with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("JaehaL/bgrpo-tmlr-models", device_map="auto") - Notebooks
- Google Colab
- Kaggle
BGRPO — Rank-Aware Beam GRPO models
Model weights accompanying:
Discovering Hidden Algebraic Structures via Transformers with Rank-Aware Beam GRPO Jaeha Lee, Tony Yue Yu — TMLR
Code: https://github.com/Jaeha0526/PolynomialDecomposition
These are the checkpoints behind the paper's reported results: the three supervised base models used to initialise RL, and the 27 post-RL endpoints (3 architectures × 3 seeds × 3 methods) at training step 420.
Contents
sft_base/ supervised bases (RL initialisation)
d2_arch_256_l6_snapshot_best.pt 36 MB d_model 256, 6 layers
d2_arch_512r3_l6_snapshot_best.pt 91 MB d_model 512, 6 layers
d2_arch_768r2_l6_snapshot_best.pt 182 MB d_model 768, 6 layers
bgrpo_runs/<arch>/<seed>/<method>.pt RL endpoints, step 420
arch : from_256_best | from_512r3_best | from_768r2_best
seed : 1 | 2 | 148
method : grpo | bgrpo | bgrpo_rank
grpo is the standard baseline, bgrpo adds beam search over candidate
decodings, and bgrpo_rank is the rank-aware variant introduced in the paper.
Each bgrpo_runs/<arch>/... group starts from the correspondingly named
sft_base model.
Task
Polynomial decomposition: recovering hidden algebraic structure by factoring a multivariate polynomial into a composition of lower-degree parts. The paper studies whether RL over beam-search candidates, with a rank-aware reward, finds decompositions that supervised training alone misses.
Loading
import torch
ckpt = torch.load("bgrpo_runs/from_768r2_best/seed_148/bgrpo_rank.pt",
map_location="cpu")
Provenance
Trained on the Caltech Resnick HPC, April–May 2026. Configs, training code, evaluation scripts, and the eval summary JSONs that generate the paper figures live in the GitHub repository above. This repository archives only the weights, which are too large for git.
Citation
@article{lee2026bgrpo,
title = {Discovering Hidden Algebraic Structures via Transformers
with Rank-Aware Beam GRPO},
author = {Lee, Jaeha and Yu, Tony Yue},
journal = {Transactions on Machine Learning Research},
year = {2026}
}