Automatic Detection of Cyberbullying in Multilingual Texts Using XLM-RoBERTa
Main Article Content
Abstract
Detection of cyberbullying in multilingual and code-mixed social media content is difficult as it is not written in formal English, translated through transliteration and has ambiguity in context especially in the Hinglish text. The aim of this research is to suggest an architecture of transformer, fine-tuned with XLM-RoBERTa to classify the data into two classes: cyberbullying and non-cyberbullying for the identification of cyberbullying in Hinglish social media posts. The proposed model is tested against the conventional baselines of ML and DL with the use of classical classification measures. The experimental results demonstrate that the XLM-RoBERTa model demonstrates an accuracy of 94.05%, a balanced precision-recall trade-off of 0.97 and an AUC of 0.97, which are all higher than the base approaches. The results show that multilingual transformer models can learn the pattern of code-mixed multilingual text, both with respect to the context and semantics, which makes them suitable to be used in cyberbullying detection in real-world social media setting.