PALO: A Polyglot Large Multimodal Model for 5B People
arXiv · · Significant research
Summary
Researchers introduce PALO, a polyglot large multimodal model with visual reasoning capabilities in 10 major languages including Arabic. A semi-automated translation approach was used to adapt the multimodal instruction dataset from English to the target languages. The models are trained across three scales (1.7B, 7B and 13B parameters) and a multilingual multimodal benchmark is proposed for evaluation.
Keywords
multilingual · multimodal · vision-language model · Arabic · benchmark
Get the weekly digest
Top AI stories from the GCC region, every week.