MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
NAACL 2025
Wentao Ge*, Shunian Chen*, Guiming Hardy Chen*, Junying Chen, Zhihong Chen, Nuo Chen, Wenya Xie, Shuo Yan, Chenghao Zhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang
* denotes equal contribution.
TL;DR MLLM-Bench evaluates open-ended multimodal responses through pairwise comparisons using sample-specific criteria and a multimodal judge, with evaluations showing strong agreement with human judgments across diverse cognitive tasks.