Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

COLING 2025

Dingjie Song, Wenjun Wang, Shunian Chen, Xidong Wang, Michael Guan, Benyou Wang

TL;DR TRIM uses a CLIP-based metric to select and reduce image tokens, lowering the computational cost of multimodal language models while largely preserving performance across visual question-answering benchmarks.

← All publications