Standard speech recognition metrics mask critical failures on code-switched speech—you need switch-aware diagnostics to understand where systems actually break down, especially for low-resource languages like Yoruba.
This paper evaluates how well modern speech recognition systems handle code-switched speech (mixing English and Yoruba). Standard metrics like word error rate hide the real problems: systems struggle much more with Yoruba than English, and fail especially at language switches.