Skip to content
GCC AI Research

Search

Results for "Akita"

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

arXiv ·

Researchers from MBZUAI have introduced VideoMolmo, a large multimodal model for spatio-temporal pointing conditioned on textual descriptions. The model incorporates a temporal module with an attention mechanism and a temporal mask fusion pipeline using SAM2 for improved coherence across video sequences. They also curated a dataset of 72k video-caption pairs and introduced VPoS-Bench, a benchmark for evaluating generalization across real-world scenarios, with code and models publicly available.

Bredas honored at 251st American Chemical Society National Meeting

KAUST ·

This article mentions KAUST in the context of the 251st American Chemical Society National Meeting. However, it contains no specific details about AI or related research activities. The content is primarily a copyright notice for King Abdullah University of Science and Technology. Why it matters: This mention provides minimal information about KAUST's involvement in the event and lacks substantial AI-related content.