{"_id":"69e74bca0b1f29fd2f7f20c3","id":"mcshao/LAT-Bench","author":"mcshao","sha":"f8e26d65835545e35d8260e8e73d5114b71eacd3","lastModified":"2026-04-27T06:30:28.000Z","private":false,"gated":false,"disabled":false,"tags":["language:zh","language:en","license:cc-by-nc-4.0","arxiv:2604.22245","region:us"],"description":"\n\n\t\n\t\t\n\t\n\t\n\t\tLAT-Bench\n\t\n\nLAT-Bench is the first benchmark designed for evaluating temporal awareness in long-form audio understanding. \nUnlike existing benchmarks limited to short clips, LAT-Bench supports audio durations up to 30 minutes, enabling evaluation under realistic long-form scenarios.\nThe benchmark covers three core tasks:\n\nDense Audio Captioning (DAC): generate temporally grounded descriptions over the full audio\nTemporal Audio Grounding (TAG): localize relevant time spans for a… See the full description on the dataset page: https://huggingface.co/datasets/mcshao/LAT-Bench.","downloads":152,"likes":1,"cardData":{"license":"cc-by-nc-4.0","language":["zh","en"],"viewer":false},"siblings":[{"rfilename":".gitattributes"},{"rfilename":"Figures/bench-figure.png"},{"rfilename":"Figures/bench-table.png"},{"rfilename":"README.md"},{"rfilename":"meta/bench-CN-meta.jsonl"},{"rfilename":"meta/bench-EN-meta.jsonl"},{"rfilename":"task/bench-CN-DAC.jsonl"},{"rfilename":"task/bench-CN-TAC.jsonl"},{"rfilename":"task/bench-CN-TAG.jsonl"},{"rfilename":"task/bench-EN-DAC.jsonl"},{"rfilename":"task/bench-EN-TAC.jsonl"},{"rfilename":"task/bench-EN-TAG.jsonl"}],"createdAt":"2026-04-21T10:04:58.000Z","usedStorage":381009}