Skip to content
GCC AI Research

Search

Results for "Visual Context"

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

arXiv ·

MBZUAI introduces Agent-X, a benchmark for evaluating multi-step reasoning in vision-centric agents across real-world, multimodal settings. Agent-X includes 828 tasks with diverse visual contexts and spans six environments, requiring tool use and stepwise decision-making. Experiments show that current LLMs struggle with multi-step vision tasks, achieving less than 50% success, highlighting areas for improvement in LMM reasoning and tool use.

Exploring Visual Context for Weakly Supervised Person Search - The Association for the Advancement of Artificial Intelligence

Inception ·

Based solely on its title, the research paper "Exploring Visual Context for Weakly Supervised Person Search" investigates methods for leveraging visual cues to improve person search capabilities. This work explores advancements in weakly supervised learning techniques for identifying individuals across different image or video frames. The publication is associated with The Association for the Advancement of Artificial Intelligence (AAAI), indicating a contribution to the broader AI research community. Why it matters: Improvements in person search technology are vital for applications in security, surveillance, and intelligent systems, which have significant implications for smart city initiatives and public safety in the region.