Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>In general, Indian communities interact in Hinglish, a Hindi and English blend that is frequently written in Roman script so is the Indian social media.On code-mixed text, where spelling variations, frequent language switching, and noisy Romanization are common, multilingual models like mBERT and XLM-R perform poorly, despite their advancements in cross-lingual natural language processing. In this research work, a dataset named COMI-LINGUA has been used which has nearly over 100,000 expert-annotated examples from various NLP tasks, to explore token-level Hinglish Language Identification (LID). In this work, we have proposed a method that provides a lightweight yet efficient baseline by using a straightforward pipeline based on character-level TF-IDF features and Logistic Regression. On the validation set, the model's accuracy is 96.06% with a macro-F1 of 0.9523; on the test set, it achieves 95.92% accuracy with a macro-F1 of 0.9509. also, the error analysis has been done that indicates persistent problems like short token ambiguity, Romanization-induced spelling variations, and loanwords that conflate languages. This work makes three contributions: (i) a reproducible baseline that performs roughly 7% better than multilingual benchmarks; (ii) an error-driven analysis of the primary difficulties in Hinglish LID; and (iii) a roadmap for expanding the work to multilingual environments, context-aware models, and fairness-focused methodologies. The results provide a useful standard and guidance for inclusive language technologies in India going forward.</p>

Show More

Keywords

language work hinglish multilingual models

Related Articles

PORE

About

Connect