A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
arXiv · · Significant research
Summary
A new benchmark, ViMUL-Bench, is introduced to evaluate video LLMs across 14 languages, including Arabic, with a focus on cultural inclusivity. The benchmark includes 8k manually verified samples across 15 categories and varying video durations. A multilingual video LLM, ViMUL, is also presented, along with a training set of 1.2 million samples, with both to be publicly released.
Keywords
multilingual · video LLM · benchmark · ViMUL-Bench · cultural diversity
Get the weekly digest
Top AI stories from the GCC region, every week.