More long-context evaluation results?

#31
by ArlenSmith - opened

Hi, thanks for the great work!

I noticed the report includes LongBench-V2 and some agent evaluations with a 1M context window. Do you have any more long-context evaluation results, especially graph-walk or retrieval-style tests like Needle-in-a-Haystack, Passkey Retrieval, or RULER?

It would also be interesting to see results across different context lengths / needle depths, especially around 512K–1M.

Thanks!

Sign up or log in to comment