Skip to content
GCC AI Research

Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models

arXiv · · Significant research

Summary

Researchers introduced AraBehave, a benchmark comprising 1,623 culturally grounded, open-ended Arabic prompts and over 29,000 human judgments to evaluate the cultural appropriateness of large language models (LLMs). The study found that cultural appropriateness involves two distinct components: normative stance and grounded cultural accuracy, with general-purpose models often failing on stance and Arabic-centric models on accuracy. Providing cultural instructions significantly improved the normative stance of general-purpose models, while grounding correlated with model scale and Arabic alignment data. Why it matters: This benchmark provides a critical tool for assessing how LLMs behave in nuanced Arabic cultural contexts, moving beyond mere factual knowledge to address ethical and social acceptance, which is crucial for their responsible deployment and trustworthiness in the region.

Get the weekly digest

Top AI stories from the GCC region, every week.

Related

SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia

arXiv ·

The paper introduces SaudiCulture, a new benchmark for evaluating the cultural competence of LLMs within Saudi Arabia, covering five major geographical regions and diverse cultural domains. The benchmark includes questions of varying complexity and distinguishes between common and specialized regional knowledge. Evaluations of five LLMs (GPT-4, Llama 3.3, FANAR, Jais, and AceGPT) revealed performance declines on region-specific questions, highlighting the need for region-specific knowledge in LLM training.

Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps

arXiv ·

This survey paper analyzes over 40 benchmarks used to evaluate Arabic large language models, categorizing them into Knowledge, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. It identifies progress in benchmark diversity but also highlights gaps like limited temporal evaluation and cultural misalignment. The paper also examines methods for creating benchmarks, including native collection, translation, and synthetic generation. Why it matters: The survey provides a comprehensive reference for Arabic NLP research and offers recommendations for future benchmark development to better align with cultural contexts.

From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher

arXiv ·

This paper introduces Absher, a new benchmark for evaluating LLMs' linguistic and cultural competence in Saudi dialects. The benchmark comprises over 18,000 multiple-choice questions spanning six categories, using dialectal words, phrases, and proverbs from various regions of Saudi Arabia. Evaluation of state-of-the-art LLMs reveals performance gaps, especially in cultural inference and contextual understanding, highlighting the need for dialect-aware training.