VideoMolmo: Spatio-Temporal Grounding Meets Pointing
arXiv · · Significant research
Summary
Researchers from MBZUAI have introduced VideoMolmo, a large multimodal model for spatio-temporal pointing conditioned on textual descriptions. The model incorporates a temporal module with an attention mechanism and a temporal mask fusion pipeline using SAM2 for improved coherence across video sequences. They also curated a dataset of 72k video-caption pairs and introduced VPoS-Bench, a benchmark for evaluating generalization across real-world scenarios, with code and models publicly available.
Keywords
spatio-temporal pointing · multimodal model · video segmentation · MBZUAI · VPoS-Bench
Get the weekly digest
Top AI stories from the GCC region, every week.