MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

NAACL 2025

Wentao Ge*, Shunian Chen*, Guiming Hardy Chen*, Junying Chen, Zhihong Chen, Nuo Chen, Wenya Xie, Shuo Yan, Chenghao Zhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang

* denotes equal contribution.

TL;DR MLLM-Bench evaluates open-ended multimodal responses through pairwise comparisons using sample-specific criteria and a multimodal judge, with evaluations showing strong agreement with human judgments across diverse cognitive tasks.

← All publications