When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
arXiv · · Significant research
Summary
Researchers from the National Center for AI in Saudi Arabia investigated the sensitivity of Large Language Model (LLM) leaderboards to minor benchmark perturbations. They found that small changes, like choice order, can shift rankings by up to 8 positions. The study recommends hybrid scoring and warns against over-reliance on simple benchmark evaluations, providing code for further research.
Keywords
LLM · benchmarks · evaluation · leaderboard · sensitivity
Get the weekly digest
Top AI stories from the GCC region, every week.