![]()
Innovative Spam Detection Using Hybrid Machine Learning Algorithms: A Data-Centric Approach
Sree Vidya Venigalla1, K.V.D. Kiran2
1Sree Vidya Venigalla, Student, Department of Computer Science Engineering, Koneru Lakshmaiah Educational Foundation, Vijayawada (Andhra Pradesh), India.
2Dr. K.V.D. Kiran, Professor, Department of Computer Science Engineering, Koneru Lakshmaiah Educational Foundation, Vijayawada, (Andhra Pradesh), India.
Manuscript received on 28 October 2025 | First Revised Manuscript received on 06 November 2025 | Second Revised Manuscript received on 11 November 2025 | Manuscript Accepted on 15 November 2025 | Manuscript published on 30 November 2025 | PP: 36-40 | Volume-14 Issue-12, November 2025 | Retrieval Number: 100.1/ijitee.D831014041125 | DOI: 10.35940/ijitee.D8310.14121125
Open Access | Editorial and Publishing Policies | Cite | Zenodo | OJS | Indexing and Abstracting
© The Authors. Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP). This is an open access article under the CC-BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)
Abstract: The rise of spam messages, in the form of malware, phishing attacks, and unrequested messages, poses a serious threat to internet users and security infrastructures. Conventional spam filtering techniques that rely solely on strict rules and keyword lists struggle to keep pace with contemporary spammer tactics that mask malicious content. This study proposes a solution to this challenge by developing a hybrid machine learning methodology that leverages Naive Bayes (NB) and a Support Vector Machine (SVM), combining them into an ensemble for improved accuracy and resilience in spam detection. The technique uses the wellknown SMS Spam Collection Dataset. It employs more complex textual feature extraction (TF-IDF), as well as additional nontextual features such as message length, word capitalisation, and the frequency of previously determined keywords. The proposed system is extensively evaluated using standard classification metrics—accuracy, F1 score, precision, and recall —to assess its reliability and validity. The research findings indicate that the proposed machine learning hybrid ensemble is effective at reducing false positives while more boldly tackling the challenges inherent in the real-world spam data environment. The research project offers practical potential for use; the hybrid proposed system is computationally efficient enough for most real-time deployment applications in automated systems to combat spam. This research contributes scalable, adaptive spam-detection mechanisms suitable for real-time messaging environments.
Keywords: Spam Detection; Machine Learning; Naive Bayes; Support Vector Machines; Ensemble Models; Nlp; Text Classification; Cybersecurity; Hybrid Algorithm; Data Analytics.
Scope of the Article: Security Technology
