Any plans for a larger scale up? (e.g., 7B - 12B version)

#46
by rpopreapovle - opened

Hi Nanbeige Team,

First of all, thank you so much for open-sourcing Nanbeige4.1-3B! The post-training pipeline, RL setup, and data synthesis here are absolutely mind-blowing.

Seeing a 3B-parameter model hit 76.9% on LiveCodeBench, 87.4% on AIME 2026, and sustain over 500+ long-horizon tool invocation turns in deep search is unreal. It legitimately punches way above its weight class, often outperforming Qwen3-30B-A3B and other much larger models on agentic workflows. It’s the perfect compact model for highly efficient local deployment.

Given how incredibly dense and efficient your alignment and RL framework is, I wanted to ask: Are there any plans to release a scaled-up version in the near future (e.g., in the 7B–12B parameter range)?

A 7B-12B version trained with the exact same post-training philosophy would be an absolute game-changer for consumer hardware. It would easily fit into a single GPU for local deployment while potentially providing enough capacity to challenge closed commercial models on complex coding and deep-search tasks.

Huge respect for your work, and looking forward to any insights or roadmaps you can share!

Nanbeige LLM Lab org

Thanks for your interest and support!

Both larger-scale models and improved smaller ones will be released over the coming months. Stay tuned.

Sign up or log in to comment