LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
arXiv · · Significant research
Summary
MBZUAI researchers introduce LLM-BabyBench, a benchmark suite for evaluating grounded planning and reasoning in LLMs. The suite, built on a textual adaptation of the BabyAI grid world, assesses LLMs on predicting action consequences, generating action sequences, and decomposing instructions. Datasets, evaluation harness, and metrics are publicly available to facilitate reproducible assessment.
Get the weekly digest
Top AI stories from the GCC region, every week.