Arabic speech AI was severely underrepresented in multilingual models—this project provides the dataset, trained models, and evaluation tools needed to build general-purpose Arabic speech systems from scratch.
Nuha-Speech addresses the lack of Arabic speech AI models by creating a comprehensive framework: a 1.5M-sample Arabic speech question-answering dataset, fine-tuned models based on Qwen-Omni, and evaluation benchmarks. This work establishes foundational infrastructure for building Arabic-capable speech AI systems despite limited available resources.