-
shufanshen/Qwen3-4B-GRPO-DeepMath-50-steps
Reinforcement Learning • 4B • Updated • 15 -
shufanshen/Qwen3-4B-GRPO-DeepMath-100-steps
Reinforcement Learning • 4B • Updated • 13 -
shufanshen/Qwen3-4B-GRPO-DeepMath-150-steps
Reinforcement Learning • 4B • Updated • 19 -
shufanshen/Qwen3-8B-GRPO-DeepMath-50-steps
Reinforcement Learning • 8B • Updated • 17
Shufan Shen
shufanshen
AI & ML interests
Interpretable machine learning, parameter-efficient fine-tuning.
Recent Activity
upvoted a paper about 16 hours ago
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training authored a paper 2 days ago
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training authored a paper 2 days ago
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs