Skip to content
GCC AI Research

Child ASR Adaptation with Adult Retention: An Empirical Study

arXiv · · Significant research

Summary

Researchers investigated methods for adapting Automatic Speech Recognition (ASR) systems to child speech while minimizing performance degradation on adult speech. The study compared full fine-tuning, LoRA, and post-hoc weight-space merging across encoder-decoder, encoder-CTC, and AudioLLM-based ASR systems using Arabic and English child and adult speech datasets. Findings indicate that while child adaptation is crucial, especially for non-native Arabic and English, weight-space merging techniques like LERP and TIES often improve the trade-off between child adaptation and adult retention. Why it matters: This research offers practical strategies to develop more robust and universally applicable ASR systems that effectively serve both child and adult users in multilingual contexts, particularly relevant for Arabic-speaking populations.

Get the weekly digest

Top AI stories from the GCC region, every week.

Related

N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition

arXiv ·

This paper benchmarks the performance of OpenAI's Whisper model on diverse Arabic speech recognition tasks, using publicly available data and novel dialect evaluation sets. The study explores zero-shot, few-shot, and full finetuning scenarios. Results indicate that while Whisper outperforms XLS-R models in zero-shot settings on standard datasets, its performance drops significantly when applied to unseen Arabic dialects.

Enhanced Arabic Text Retrieval with Attentive Relevance Scoring

arXiv ·

This paper introduces an enhanced Dense Passage Retrieval (DPR) framework tailored for Arabic text retrieval. The core innovation is an Attentive Relevance Scoring (ARS) mechanism that improves semantic relevance modeling between questions and passages, replacing standard interaction methods. The method integrates pre-trained Arabic language models and architectural refinements, achieving improved retrieval and ranking accuracy for Arabic question answering. Why it matters: This work addresses the underrepresentation of Arabic in NLP research by providing a novel approach and publicly available code to improve Arabic text retrieval, which can benefit various applications like Arabic search engines and question-answering systems.

Processing language like a human

MBZUAI ·

MBZUAI's Hanan Al Darmaki is working to improve automated speech recognition (ASR) for low-resource languages, where labeled data is scarce. She notes that Arabic presents unique challenges due to dialectal variations and a lack of written resources corresponding to spoken dialects. Al Darmaki's research focuses on unsupervised speech recognition to address this gap. Why it matters: Overcoming these challenges can improve virtual assistant effectiveness across diverse languages and enable more inclusive AI applications in the Arabic-speaking world.