Redteaming Leading Arabic LLMs with ASAS
arXiv · · Significant research
Summary
Researchers introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming large language models (LLMs), comprising 801 prompts across 8 safety categories and 8 attack strategies. An evaluation conducted with ASAS on seven leading Arabic-capable LLMs, including GPT-4o, Claude 3.7 Sonnet, ALLaM, and FANAR, revealed significant safety gaps. The findings indicate that most models failed to defend against 50% of unsafe prompts, particularly in high-harm categories, and automated safety judges performed poorly compared to human annotators. Why it matters: This benchmark provides a crucial tool for improving the safety, cultural alignment, and responsible development of LLMs for Arabic-speaking regions.
Keywords
Arabic LLMs · AI safety · Redteaming · ASAS benchmark · Cultural alignment
Get the weekly digest
Top AI stories from the GCC region, every week.