ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
arXiv · · Significant research
Summary
ArGuard is a shared task focused on detecting harmful content in Arabic memes and LLM prompts, featuring two tracks: multimodal hate detection in memes and harmful prompt detection for Arabic LLM safety evaluation. The task saw 58 teams register, 35 participate in evaluation, and 27 submit system-description papers, with participants exploring models like AraBERT, Jais, and Qwen3-VL. Best systems achieved macro-F1 scores up to 0.984 (B1) and 0.823 (A1), though fine-grained meme classification (A2) proved particularly challenging. Why it matters: This initiative significantly contributes to benchmarking and advancing the state of Arabic AI safety and robust harmful content detection for Arabic language models and digital content.
Keywords
ArGuard · Harmful content detection · Arabic memes · LLM prompts · Arabic AI safety
Get the weekly digest
Top AI stories from the GCC region, every week.