Curious about the longbenchv2 benchmark

#15
by Vincent-luo - opened

Thanks for your great work!
I am interested in the long context reasoning ability of llm. I find your sft model achieve 42.15 on longbenchv2 benchmark, but I check the sft dataset(mainly UltraData-SFT-2605), the user prompts are all very short, mean length less than 500 while user prompts of longbenchv2 are often very long(up to 2M). How does the sft model get the good long context reasoning ability here? From pretraining/mid-training? Or do you use extra long user prompt sft data? Looking forward to your reply!

Sign up or log in to comment