Open Access

Comparative Analytical Framework for Assessing Multiple Machine Learning Classifiers in Twitter Sentiment Analysis Using Bag-of-Words Feature Representation

4 Department of Computer Engineering, Accra Institute of Technology, Accra, Ghana, Research
4 Faculty of Computing, University of Digital Sciences, Kumasi, Ghana, Research Interests: Software Engineering, Big Data Analytics

Abstract

The rapid growth of social media platforms has intensified the need for automated sentiment analysis systems capable of processing large-scale, noisy, and high-velocity textual data. Twitter, in particular, has emerged as a critical data source for sentiment-driven decision-making in domains such as marketing, politics, and public health. This study presents a comparative analytical framework for evaluating multiple machine learning classifiers using Bag-of-Words (BoW) feature representation for Twitter sentiment classification. The research integrates classical and modern supervised learning algorithms, including Support Vector Machine (SVM), NaΓ―ve Bayes, Random Forest, Logistic Regression, and ensemble-based approaches, to assess their effectiveness in handling short-text sentiment data.

The methodological foundation is built upon established preprocessing techniques, BoW vectorization, and classification pipelines supported by prior research in sentiment analysis and text mining (Hickman et al., 2022; Wankhade et al., 2022). Special emphasis is placed on the foundational SVM-based sentiment classification approach proposed by Ahmad M, Aftab S, Ali I (2017), which is referenced multiple times as a benchmark model in this study. Experimental comparisons highlight performance variations across classifiers in terms of accuracy, precision, recall, and F1-score, while also analyzing computational efficiency and robustness against noisy Twitter data.

The findings indicate that while SVM-based models remain highly effective for high-dimensional sparse data, ensemble models demonstrate improved stability under noisy conditions. The study contributes a structured analytical framework for classifier evaluation and highlights the continued relevance of BoW-based sentiment pipelines in contemporary natural language processing applications.

Keywords

References

Ahmad M, Aftab S, Ali I (2017) Sentiment Analysis of Tweets using SVM. Int J Comput Appl 177:25–29.
Asselman A, Khaldi M, Aammou S (2023) Enhancing the prediction of student performance based on the machine learning XGBoost algorithm. Interactive Learning Environments 31:3360–3379.
Birjali M, Kasri M, Beni-Hssane A (2021) A comprehensive survey on sentiment analysis: Approaches, challenges, and trends. Knowl Based Syst 226:1071346.
Boateng EY, Otoo J, Abaye DA (2020) Basic tenets of classification algorithms K-nearest-neighbor, support vector machine, random forest and neural network: A review. Journalof Data Analysis and Information Processing 8:341–357.
Cervantes J, Garcia-LamontF, RodrΓ­guez-Mazahua L, Lopez A (2020) A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing 408:189–215.
Dubey Assistant Professor AD Twitter Sentiment Analysis during COVID-19 Outbreak.
Ghatasheh N, Altaharwa I, Aldebei K (2022) Modified GeneticAlgorithm for Feature Selection and Hyper Parameter Optimization: Case of XGBoost in Spam Prediction. IEEE Access 10:84365–84383.
Hickman L, Thapa S, Tay L, et. al. (2022) Text preprocessing for text mining in organizational research: Review and recommendations. Organ Res Methods 25:114–146.
Juluru K, Shih H-H, Keshava Murthy KN, ElnajjarP (2021) Bag-of-words technique in natural language processing: a primer for radiologists. RadioGraphics 41:1420–142610. JICET, 2024, Vol:4, No.29.
Khan M, Srivastava A (2024) Sentiment Analysis of Twitter Data Using Machine Learning Techniques. International Journal of Engineering and Management Research Peer Reviewed & Refereed Journal e 14:.
Luo X (2021) Efficient English text classification using selected machine learning techniques. Alexandria Engineering Journal 60:3401–3409.
Mohammad SM (2022) Ethics sheet for automatic emotion recognition and sentiment analysis. Computational Linguistics 48:239–278.
Naseem U, Razzak I, Eklund PW (2021) A survey of pre-processing techniques to improve short-text quality: a case study on hate speech detection on Twitter. Multimed Tools Appl 80:35239–35i2665.
Nhu V-H, Shirzadi A, Shahabi H, et. al. (2020) Shallow landslide susceptibility mapping: A comparison between logistic model tree, logistic regression, naΓ―ve Bayes tree, artificial neural network, and support vector machine algorithms. Int J Environ Res Public Health 17:2749.
Peng S, Cao L, Zhou Y, et. al. (2022) A survey on deep learning for textual emotion analysis in social networks. Digital Communications and Networks 8:745–7623.
Reddy H, Raj N, Gala M, Basava A (2020) Text-mining-based fake news detection using ensemble methods. International journal ofautomation and computing 17:210–2218.
Shen CY (2020) Logistic growth modeling of COVID-19 proliferation in China and its international implications. International Journal of Infectious Diseases 96:582–589.
Shetty SH, Shetty S, Singh C, Rao A (2022) Supervised machine learning: algorithms and applications. Fundamentals and methods of machine and deep learning: algorithms, tools, and applications 1–1611.
Wankhade M, Rao ACS, Kulkarni C (2022) A survey on sentiment analysis methods, applications, and challenges. Artif Intell Rev 55:5731–5780.
Xue L, Liu Y, Xiong Y, et. al. (2021) A data-driven shale gas production forecasting method based on themulti-objective random forest regression. J Pet Sci Eng 196:10780120.
Yousaf A, Umer M, Sadiq S, et. al. (2020) Emotion recognition by textual tweets classification using the voting classifier (LR-SGD). IEEE Access 9:6286–6295.

Similar Articles

1-10 of 76

You may also start an advanced similarity search for this article.