Contextual transformer models like MuRIL and XLM-RoBERTa can effectively identify which language each token belongs to in code-mixed text, outperforming traditional monolingual approaches.
This paper tackles language identification in code-mixed text (where people mix multiple languages in one message) by treating it as a sequence labeling task. Researchers fine-tuned transformer models designed for Indian languages on three language pairs and released benchmark datasets with manually annotated examples to help future work on this social media problem.