SAM-AI-Reasoning-14B
Official submission of SAM-AI Reasoning Engine trained with Group Relative Policy Optimization (GRPO).
Evaluation
Submitted for independent verification on the Hugging Face Open LLM Leaderboard v2:
- IFEval (Instruction Following)
- BBH (Big Bench Hard)
- MATH-Hard (Advanced Competition Mathematics)
- GPQA Diamond (PhD-level Science)
- MuSR (Multi-step Soft Reasoning)
- MMLU-Pro (Professional Knowledge)
Model tree for Samrish2009/SAM-AI-Reasoning-14B
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B