Skip to content
GCC AI Research

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

arXiv · · Significant research

Summary

MBZUAI researchers introduce SocialMaze, a new benchmark for evaluating social reasoning capabilities in large language models (LLMs). SocialMaze includes six diverse tasks across social reasoning games, daily-life interactions, and digital community platforms, emphasizing deep reasoning, dynamic interaction, and information uncertainty. Experiments show that LLMs vary in handling dynamic interactions, degrade under uncertainty, but can be improved via fine-tuning on curated reasoning examples.

Get the weekly digest

Top AI stories from the GCC region, every week.