Ensemble Models and Explainable AI for Malware Detection
Abstract & Details
Research Area
Computer Science
Keywords
Malware Detection
Hybrid Model
Random Forest
Artificial Neural Network
Explainable AI
SHAP
LIME
Cybersecurity
Behavioral Analysis
MalwareShield AI.
Abstract
The rapid growth of malware has created serious security challenges for modern computing systems. Traditional signature-based detection techniques are often ineffective against newly emerging and polymorphic malware, making intelligent detection mechanisms necessary. This study proposes a hybrid machine learning framework for malware detection that combines Random Forest (RF) and Artificial Neural Network (ANN) models to improve classification accuracy and reliability.
The dataset used in this research consists of 100,000 records obtained from a publicly available Kaggle repository, evenly split between malware and benign samples. The dataset underwent preprocessing steps including removal of redundant attributes, handling missing values, numeric conversion of features, and feature scaling using StandardScaler. The 33 behavioral features capture Linux kernel process characteristics such as memory usage, CPU scheduling, context switches, and execution timing.
Several machine learning models — Support Vector Machine (SVM), Decision Tree (DT), K-Nearest Neighbors (KNN), and Random Forest (RF) — were implemented to evaluate baseline performance. A deep learning model based on an Artificial Neural Network (ANN) with two hidden layers, dropout regularization, and early stopping was also trained to capture complex non-linear patterns.
A Hybrid model was developed by combining the prediction probabilities of the Random Forest and ANN models using ensemble averaging, achieving the highest accuracy of 93.0%, precision of 92.93%, recall of 93.08%, and F1-score of 93.01%. Model performance was evaluated using accuracy, precision, recall, F1-score, ROC curves, and confusion matrices. To improve interpretability, Explainable Artificial Intelligence techniques — SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-Agnostic Explanations) — were applied to analyze feature contributions globally and locally. A Chi-Square statistical test was used to validate that the top features identified by XAI methods are genuinely significant.
A Streamlit-based interactive web application called MalwareShield AI was developed to demonstrate the system with live detection, batch scanning, model performance visualization, SHAP analysis, LIME explanation, and Chi-Square validation modules. The proposed hybrid approach provides an accurate, scalable, and interpretable solution for malware detection in modern cybersecurity systems.
License
This work is licensed under a Creative
Commons
Attribution-ShareAlike 4.0 International License.
Author Information
| # | Name | Institute / Affiliation |
|---|---|---|
| 1 | M Vandana | Sphoorthy Engineering College |
| 2 | Gattu Prasad | Sphoorthy Engineering College |
| 3 | G Sreeja Reddy | Sphoorthy Engineering College |
| 4 | P Snigdha Reddy | Sphoorthy Engineering College |
| 5 | P Juhee Reddy | Sphoorthy Engineering College |
| 6 | M Venkatesh | Sphoorthy Engineering College |
How to Cite
Use the following formats to cite this article in your research.
APA Style
Vandana, M, Prasad, Gattu, Reddy, G Sreeja, Reddy, P Snigdha, Reddy, P Juhee, & Venkatesh, M (2026). Ensemble Models and Explainable AI for Malware Detection. International Journal of Advance Research and Innovative Ideas In Education, 12(2), 892-902.
MLA Style
Vandana, M, et al. "Ensemble Models and Explainable AI for Malware Detection." International Journal of Advance Research and Innovative Ideas In Education, vol. 12, no. 2, 2026, pp. 892-902.
IEEE Style
M Vandana, Gattu Prasad, G Sreeja Reddy, P Snigdha Reddy, P Juhee Reddy, and M Venkatesh, "Ensemble Models and Explainable AI for Malware Detection," International Journal of Advance Research and Innovative Ideas In Education, vol. 12, no. 2, pp. 892-902, 2026.
Vancouver Style
Vandana M, Prasad Gattu, Reddy G Sreeja, Reddy P Snigdha, Reddy P Juhee, Venkatesh M. Ensemble Models and Explainable AI for Malware Detection. International Journal of Advance Research and Innovative Ideas In Education. 2026;12(2):892-902.
Harvard Style
Vandana, M, Prasad, Gattu, Reddy, G Sreeja, Reddy, P Snigdha, Reddy, P Juhee, & Venkatesh, M (2026) 'Ensemble Models and Explainable AI for Malware Detection', International Journal of Advance Research and Innovative Ideas In Education, 12(2), pp. 892-902.
Chicago Style
Vandana, M, et al. "Ensemble Models and Explainable AI for Malware Detection." International Journal of Advance Research and Innovative Ideas In Education 12, no. 2 (2026): 892-902.
Turabian Style
Vandana, M, et al. "Ensemble Models and Explainable AI for Malware Detection." International Journal of Advance Research and Innovative Ideas In Education 12, no. 2 (2026): 892-902.
Related Research
CYBERSECURITY WITH AI
PDF Unavailable
DESIGN AND IMPLEMENTATION OF A SECURE IMAGE STEGANOGRAPHY SYSTEM USING LSB AND CRYPTOGRAPHY
PDF Unavailable
A NOVEL HYBRID IMAGE STEGANOGRAPHY TECHNIQUE BASED ON LSB AND CRYPTOGRAPHIC SECURITY
PDF Unavailable
BioPrint AI: An Intelligent Deep Learning and Computer Vision Based Blood Group Identification System Using Fingerprint Patterns
PDF Unavailable
AnimalAid AI: A Deep Learning Powered Early Warning System for Detecting Skin Infections and Diseases in Stray Dogs
PDF Unavailable
LiverCare AI: Intelligent Medical Imaging Platform for Liver Tumor Detection and Clinical Guidance
PDF Unavailable