SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
arXiv · · Significant research
Summary
MBZUAI researchers introduce SocialMaze, a new benchmark for evaluating social reasoning capabilities in large language models (LLMs). SocialMaze includes six diverse tasks across social reasoning games, daily-life interactions, and digital community platforms, emphasizing deep reasoning, dynamic interaction, and information uncertainty. Experiments show that LLMs vary in handling dynamic interactions, degrade under uncertainty, but can be improved via fine-tuning on curated reasoning examples.
Keywords
social reasoning · large language models · benchmark · SocialMaze · MBZUAI
Get the weekly digest
Top AI stories from the GCC region, every week.