Skip to content
GCC AI Research

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

arXiv · · Significant research

Summary

Researchers from MBZUAI have introduced VideoMolmo, a large multimodal model for spatio-temporal pointing conditioned on textual descriptions. The model incorporates a temporal module with an attention mechanism and a temporal mask fusion pipeline using SAM2 for improved coherence across video sequences. They also curated a dataset of 72k video-caption pairs and introduced VPoS-Bench, a benchmark for evaluating generalization across real-world scenarios, with code and models publicly available.

Get the weekly digest

Top AI stories from the GCC region, every week.