Skip to content
GCC AI Research

CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

arXiv · · Significant research

Summary

Researchers present Cardamom, a new micro-dialectal Arabic speech dataset comprising approximately 40 hours of transcribed YouTube speech across 21 micro-dialects from six Middle Eastern countries. Each segment is annotated with specific micro-dialect labels, code-switching information, and perceived gender, enabling detailed analysis of sub-country variation in speech. Benchmarking four multilingual ASR systems showed a 43.47% aggregate WER in zero-shot settings, which improved to 35.21% after adaptation on Cardamom. Why it matters: This dataset provides a crucial resource for developing more robust Arabic automatic speech recognition systems capable of handling the significant dialectal diversity present in the region.

Get the weekly digest

Top AI stories from the GCC region, every week.

Related

QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus

arXiv ·

The Qatar Computing Research Institute (QCRI) has released QASR, a 2,000-hour transcribed Arabic speech corpus collected from Aljazeera news broadcasts. The dataset features multi-dialect speech sampled at 16kHz, aligned with lightly supervised transcriptions and linguistically motivated segmentation. QCRI also released a 130M word dataset to improve language model training. Why it matters: QASR enables new research in Arabic speech recognition, dialect identification, punctuation restoration, and other NLP tasks for spoken data.

N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition

arXiv ·

This paper benchmarks the performance of OpenAI's Whisper model on diverse Arabic speech recognition tasks, using publicly available data and novel dialect evaluation sets. The study explores zero-shot, few-shot, and full finetuning scenarios. Results indicate that while Whisper outperforms XLS-R models in zero-shot settings on standard datasets, its performance drops significantly when applied to unseen Arabic dialects.

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs

arXiv ·

The Qatar Computing Research Institute (QCRI) has released SpokenNativQA, a multilingual spoken question-answering dataset for evaluating LLMs in conversational settings. The dataset contains 33,000 naturally spoken questions and answers across multiple languages, including low-resource and dialect-rich languages. It aims to address the limitations of text-based QA datasets by incorporating speech variability, accents, and linguistic diversity. Why it matters: This benchmark enables more robust evaluation of LLMs in speech-based interactions, particularly for Arabic dialects and other low-resource languages.

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

arXiv ·

A new culturally inclusive and linguistically diverse dataset called Palm for Arabic LLMs is introduced, covering 22 Arab countries and featuring instructions in both Modern Standard Arabic (MSA) and dialectal Arabic (DA) across 20 topics. The dataset was built through a year-long community-driven project involving 44 researchers from across the Arab world. Evaluation of frontier LLMs using the dataset reveals limitations in cultural and dialectal understanding, with some countries being better represented than others.