Ling 3.0 Flash Benchmarks Comparison in markdown
#12
by sandeshrajx - opened
Extracted from the original benchmark comparison table.
short summary:
SWE-Bench Pro : 56.6
Terminal-Bench 2.1 :57.0
LiveCodeBench (2408-2505) - 82.8
SkillsBench - 44.8
| Metric | Ling-3.0-flash | Ring-2.6-1T (xhigh) | MiniMax-M2.7 | Step-3.7-Flash (high) | Deepseek-V4-Flash-Preview (max) | Nemotron-3-Super-120B-A12B | GPT-5.4-mini (high) | Claude-Sonnet-4.6 (max) |
|---|---|---|---|---|---|---|---|---|
| Size | 124B-A5.1B | 1T-A63B | 230B-A10B | 198B-A11B | 284B-A13B | 120B-A12B | - | - |
| Coding Agent | ||||||||
| SWE-Bench Pro | 56.6 | 53.9 | 56.2 | 56.3 | 52.6 | 34.1 | 47.9 | 48.3 |
| SWE-Bench Multilingual | 72.4 | 56.7 | 76.5 | 72.4 | 73.3 | 42.7 | 71.0 | 75.9 |
| Terminal-Bench 2.1 | 57.0 | 43.1 | 55.0 | 39.3 | 62.0 | 39.0 | 55.8 | 71.2 |
| ArtifactsBench | 77.0 | 65.1 | 55.8 | 59.2 | 64.0 | 51.6 | 66.8 | 68.7 |
| MiniAppBench | 25.3 | 19.2 | 20.7 | 28.0 | 14.8 | 5.8 | 58.8 | 46.3 |
| AntSWEBench | 52.2 | 46.4 | - | - | 48.6 | - | - | - |
| General Agent | ||||||||
| Tau3-banking-AA | 28.0 | 14.6 | 8.9 | 11.3 | 22.9 | 10.1 | 11.3 | 30.5 |
| MCP-Atlas | 65.5 | 61.2 | 53.6 | 52.6 | 69.0 | 49.4 | 55.2 | 66.7 |
| SkillsBench | 44.8 | 11.9 | 28.4 | 24.9 | 53.5 | 20.3 | 44.8 | 54.4 |
| BFCL-v4 | 73.0 | 64.8 | 63.6 | 65.4 | 59.5 | 60.6 | 68.3 | 73.1 |
| GDPVal v2-AA | 1107 | 920 | 1159 | 1017 | 1189 | 699 | - | 1377 |
| Search Agent | ||||||||
| WideSearch | 73.6 | 62.2 | 75.2 | 56.8 | 74.4 | 19.5 | 70.2 | 79.5 |
| BrowseComp | 72.2 (w/ ctx) / 82.0 (MA) | 71.7 | 76.3 | 75.8 | 73.2 | 31.3 | - | 74.0 (w/ ctx) / 82.1 (MA) |
| Draco | 70.4 | - | 66.8 | - | 71.3 | - | 61.3 | 75.8 |
| Instruction Following | ||||||||
| IFBench | 74.5 | 44.6 | 75.7 | 67.3 | 79.2 | 72.6 | 69.0 | 56.6 |
| SysBench | 93.6 | 86.5 | 86.2 | 91.4 | 93.9 | 90.7 | 93.3 | 94.9 |
| LIFEBench | 77.3 | 72.5 | 66.9 | 71.3 | 74.1 | 60.2 | 69.2 | 71.8 |
| Reasoning | ||||||||
| AIME26 | 93.2 | 95.8 | 94.2 | 95.0 | 96.5 | 91.7 | 92.9 | 94.4 |
| HMMT-Feb26 | 87.0 | 93.5 | 71.9 | 87.9 | 94.8 | 84.9 | 83.9 | 85.6 |
| IMO-AnswerBench | 83.7 | 86.1 | 66.9 | - | 87.0 | 77.0 | 74.5 | 82.1 |
| HLE | 22.7 | 18.3 | 28.1 | 19.9 | 34.8 | 18.3 | 20.6 | 30.0 |
| LiveCodeBench (2408-2505) | 82.8 | 87.0 | 75.6 | 78.1 | 91.6 | 78.7 | 80.8 | 83.7 |
| Long Context & Multi-Turn Dialogue | ||||||||
| MRCR_128K | 90.8 | 90.1 | 27.7 | 39.2 | 88.5 | 40.8 | 56.1 | 92.5 |
| MRCR_256k | 81.1 | 76.5 | - | 25.7 | 84.3 | 35.8 | 50.5 | 92.7 |
| AA-LCR | 65.1 | 64.3 | 68.7 | 63.7 | 63.0 | 58.3 | 63.4 | 70.7 |
| Multi-IF | 87.7 | 89.3 | 82.9 | 84.6 | 86.2 | 82.3 | 84.6 | 84.8 |
Notes:
- Size notation follows the "TotalB-ActiveB" convention (e.g., 124B-A5.1B = 124B total / 5.1B active parameters).
- "w/ ctx" = with context, "MA" = multi-agent setup (for BrowseComp).
- "-" indicates values not reported in the original table.