vision-language models, benchmarks, evaluation, game AI
Can VLMs read the board? Field-level extraction scores.