Ling-3.0-flash-Fin
Base Model | OpenRouter | Announcement
Introduction
Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling-3.0-flash through continued training on high-quality financial data.
With 124B total parameters, 5.1B activated parameters, and a 256K context window, the model combines financial expertise with efficient inference for long-horizon agent workflows.
Highlights
- End-to-end financial research: connects information retrieval, evidence review, calculation, modeling, and report preparation instead of treating them as isolated tasks.
- Source-grounded financial search: Prioritizes authoritative sources to deliver accurate, complete, and traceable answers; FinFIRST is open-sourced alongside the model to enable transparent evaluation of these capabilities.
- Multi-document financial reasoning: reconciles reporting periods, definitions, assumptions, and conflicting figures across annual reports, earnings releases, regulatory filings, and research materials.
- Valuation and spreadsheet workflows: understands formulas, actual-versus-estimate updates, cross-sheet dependencies, balance checks, scenario analysis, and editable financial-model delivery.
- Research-ready outputs: organizes facts, analysis, judgments, and charts into clear, reviewable materials for further editing and professional review.
Evaluation
Ling-3.0-flash-Fin was evaluated across FinFIRST, FinSearchComp Verified, FinCRAFT, Finance Agent, APEX-Agents, SpreadsheetBench, and τ³-Banking. These benchmarks cover source-grounded retrieval, investment research, long-horizon execution, valuation modeling, spreadsheet operations, and banking workflows. The model is competitive with both similarly sized models and substantially larger general-purpose models, with particular strength in source selection and tool-intensive financial tasks.
Local Serving
The current checkpoint is released in BF16. Because Ling-3.0-flash-Fin shares the same architecture as Ling-3.0-flash, it is compatible with the same SGLang and vLLM runtimes. For deployment instructions, see the Ling-3.0-flash deployment guide.
Important: Thinking mode is enabled by default. For optimal performance, we strongly recommend using
temperature=1.0,top_p=0.95, andtop_k=20for general inference.
Quantized Models
We evaluate the quantized models using several datasets. The FP8 quantized model is applied via the blockwise quantization, and INT4 and FP4 models are applied via groupwise quantization with routed experts weights.
| dataset | BF16 | FP8 | INT4 | FP4 |
|---|---|---|---|---|
| GPQA-diamond | 86.30 | 85.26 | 85.39 | 84.38 |
| SciCode | 41.84 | 42.70 | 41.41 | 41.24 |
| FinCRAFT | 54.23 | 55.91 | 54.48 | 54.73 |
| FSC-verified | 78.28 | 77.48 | 76.86 | 76.77 |
Limitations and Future Work
As our first finance-enhanced release, Ling-3.0-flash-Fin still requires further validation in complex, long-horizon workflows. Key assumptions, valuation results, and investment conclusions require professional review and do not constitute investment advice.
Future releases will explore finance-enhanced models at larger scales to further improve complex reasoning and long-horizon task execution.
- Downloads last month
- 2
Model tree for inclusionAI/Ling-3.0-flash-Fin-fp8
Base model
inclusionAI/Ling-3.0-flash