Running finnlp2026subtask3polyfiqa 🚀 Explore competition details, datasets, leaderboards, and submissions
Running finnlp2026subtask3polyfiqa 🚀 Explore competition details, datasets, leaderboards, and submissions
FinNLP-Multilingual-Understanding/finnlp2026-subtask3-polyfiqa Viewer • Updated 8 days ago • 152 • 107
FinNLP-Multilingual-Understanding/finnlp2026-subtask3-polyfiqa Viewer • Updated 8 days ago • 152 • 107
FinNLP-Multilingual-Understanding/finnlp2026-subtask2-japanese-icr Viewer • Updated 8 days ago • 303 • 70
FinNLP-Multilingual-Understanding/finnlp2026-subtask2-japanese-icr Viewer • Updated 8 days ago • 303 • 70
FinNLP-Multilingual-Understanding/finnlp2026-subtask1-greek-ner Viewer • Updated 8 days ago • 600 • 94
FinNLP-Multilingual-Understanding/finnlp2026-subtask1-greek-ner Viewer • Updated 8 days ago • 600 • 94
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Paper • 2604.10015 • Published Apr 15
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents Paper • 2510.11695 • Published Oct 13, 2025 • 3
FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making Paper • 2407.06567 • Published Jul 9, 2024 • 1
MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation Paper • 2506.14028 • Published Jun 16, 2025 • 94
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Paper • 2506.08235 • Published Jun 9, 2025
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications Paper • 2503.20990 • Published Mar 26, 2025 • 19
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications Paper • 2408.11878 • Published Aug 20, 2024 • 64