Researchers introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming large language models (LLMs), comprising 801 prompts across 8 safety categories and 8 attack strategies. An evaluation conducted with ASAS on seven leading Arabic-capable LLMs, including GPT-4o, Claude 3.7 Sonnet, ALLaM, and FANAR, revealed significant safety gaps. The findings indicate that most models failed to defend against 50% of unsafe prompts, particularly in high-harm categories, and automated safety judges performed poorly compared to human annotators. Why it matters: This benchmark provides a crucial tool for improving the safety, cultural alignment, and responsible development of LLMs for Arabic-speaking regions.
Sohar International, an Omani bank, has announced a partnership with Omantel, a leading telecommunications provider in Oman. This collaboration aims to significantly accelerate the growth and development of Oman's digital economy. The initiative is expected to involve leveraging digital infrastructure and fostering innovation across various sectors within the Sultanate. Why it matters: This strategic partnership between two major Omani entities signals a concerted effort to drive national digital transformation and enhance the country's technological capabilities.
A systematic study investigated whether fine-tuning Large Language Models (LLMs) on cultural data improves figurative language understanding and vice versa. Researchers used four models, including ALLaM-7B and Fanar-1-9B, and six Arabic datasets covering cultural commonsense, proverbs, and poetry across various dialects. They found that fine-tuning on poetry improved idiom comprehension by 2.33%, suggesting a transfer of non-literal meaning understanding across figurative types. Why it matters: This research highlights the complex challenges of integrating nuanced cultural and figurative language understanding into LLMs, particularly for Arabic, suggesting that fine-tuning alone may not fully capture these intricate relationships.
Researchers developed a validated drone-based crowd counting system designed for large-scale events like the FIFA World Cup 2034 in Saudi Arabia and the Hajj. The system addresses challenges of maintaining accuracy on unlabelled footage and detecting dangerous crowd inflow before a crush forms. It employs label-free adaptation, recovering 31-49% of shift-induced error, establishes a "severity law," and includes a six-point deployment protocol. Why it matters: This research offers a critical advancement in AI safety and crowd management technology, directly supporting Saudi Arabia's capacity to host major global events and ensuring public safety through advanced computer vision techniques.