Skip to content
GCC AI Research

ALARB: An Arabic Legal Argument Reasoning Benchmark

arXiv · · Significant research

Summary

Researchers introduce ALARB, a new benchmark for evaluating reasoning in Arabic LLMs using 13K Saudi commercial court cases. The benchmark includes tasks like verdict prediction, reasoning chain completion, and identification of relevant regulations. Instruction-tuning a 12B parameter model on ALARB achieves performance comparable to GPT-4o in verdict prediction and generation.

Keywords

Arabic · legal · reasoning · benchmark · LLM

Get the weekly digest

Top AI stories from the GCC region, every week.