علوم رایانشی

علوم رایانشی

شناسایی فایل‌های آلوده غیراجرایی به‌کمک مدل‌های یادگیری ماشین

نوع مقاله : مقاله پژوهشی

نویسندگان
1 کارشناسی ارشد، دانشکده مهندسی کامپیوتر، دانشگاه علم و صنعت ایران، تهران، ایران
2 استادیار، دانشکده مهندسی کامپیوتر، دانشگاه صنعتی امیرکبیر، تهران، ایران
10.22034/csj.2025.233053
چکیده
با افزایش استفاده کاربران از اسناد آفیس و پی‌دی‌اف و استفاده این نوع فایل‌ها در مراکز امنیتی جهت انتقال اطلاعات، توجه طراحان بدافزار به این فایل‌ها جلب شده است. فعالیت‌‌های گوناگونی جهت تشخیص این نوع فایل‌‌ها با انتخاب ویژگی‌های تعیین‌کننده هر نوع فایل صورت گرفته است. در این پژوهش سعی شده است تا با تهیه مجموعه داده تقویت ‌شده و همچنین ارائه ویژگی‌‌های مؤثر و جامع برای انواع فایل‌های ذکر شده، به‌طوری که حملات مربوط به هر نوع فایل ‌را پوشش ‌دهند، نرخ تشخیص بدافزارهای مربوطه افزایش داده شود. در این راستا، این مطالعه توانست با به کارگیری مدل‌های طبقه‌‌بند دودویی، در مقایسه با برترین مطالعات انجام شده، به بهبود 2 درصدی تشخیص بدافزارهای پی‌دی‌اف با مدل گرادیان افزایشی و همچنین بهبود 9/1 درصدی تشخیص بدافزارهای آفیس با مدل جنگل تصادفی دست پیدا کند. به‌‌طور دقیق‌‌تر، این مطالعه با اعمال مدل گرادیان افزایشی بر روی فایل‌‌های پی‌دی‌اف، توانست به دقت 3/99 درصدی تشخیص فایل‌های آلوده دست پیدا کند در حالی که برای فایل‌‌های آفیس با اعمال مدل جنگل تصادفی، به نرخ 4/99 درصدی در تشخیص بدافزارها رسیده شد.
کلیدواژه‌ها
موضوعات

[1]   Issakhani, M., Victor, P., Tekeoglu, A., & Lashkari, AH. 2022. PDF Malware Detection based on Stacking Learning, in International Conference on Information Systems Security and Privacy, Science and Technology Publications, Lda, pp. 562–570. doi: 10.5220/0010908400003120.
[2]   Koutsokostas, V., Lykousas, N., Apostolopoulos, T., Orazi, G., Ghosal, A., Casino, F., Conti, M., & Patsakis, C. 2022 Mar. Invoice# 31415 attached: Automated analysis of malicious Microsoft Office documents. Computers & Security. 1; 114:102582.
[3]   Kaspersky, 2023. “https://www.kaspersky.com/about/press-releases/2023_rising-threats-cybercriminals-unleash-411000-malicious-files-daily-in-2023.”
[4]   Avira, “https://www.avira.com/en/blog/malware-threat-report-q3-2020-statistics-and-trends.”
[5]   Singh, P., Tapaswi, S., & Gupta, S. 2020 May. Malware detection in pdf and office documents: A survey. Information Security Journal: A Global Perspective. 3; 29(3):134-53.
[6]   Šrndic, N., & Laskov, P. 2013 Feb. Detection of malicious pdf files based on hierarchical document structure. InProceedings of the 20th annual network & distributed system security symposium (pp. 1-16). Citeseer.
[7]   Gu, J., Kong, R., Sun, H., Zhuang, H., Pan, F., & Lin, Z. 2022 May 24. A novel detection technique based on benign samples and one-class algorithm for malicious PDF documents containing JavaScript. In International Conference on Computer Application and Information Security (ICCAIS 2021) (Vol. 12260, pp. 599-607). SPIE.
[8]   Laskov, P. & Šrndić, N. 2011. Static detection of malicious JavaScript-bearing PDF documents. In Proceedings of the 27th annual computer security applications conference 2011 Dec 5 (pp. 373-382).
[9]   Abu Al-Haija, Q., Odeh, A., & Qattous H. 2022 Sep 30. PDF malware detection based on optimizable decision trees. Electronics.; 11(19):3142.
[10] Sohail, B. 2021 Apr 11. Macro Based Malware Detection System. Turkish Journal of Computer and Mathematics Education (TURCOMAT). 12(3):5776-87.
[11] Liu, JK., Huang, X. 2019 Dec 10editors. Network and System Security: 13th International Conference, NSS 2019, Sapporo, Japan, December 15–18, 2019, Proceedings. Springer Nature.
[12] Kim, S., Hong, S., Oh, J., Lee, H. 2018 Jun 25. Obfuscated VBA macro detection using machine learning. In2018 48th annual ieee/ifip international conference on dependable systems and networks (dsn) (pp. 490-501). IEEE.
[13] Ravi, V., Gururaj, SP., Vedamurthy, HK., Nirmala, MB. 2022 Jun 1Analysing corpus of office documents for macro-based attacks using machine learning. Global Transitions Proceedings. 3(1):20-4.
[14] Casino, F., Totosis, N., Apostolopoulos, T., Lykousas, N, Patsakis, C. 2023 Aug 10. Analysis and correlation of visual evidence in campaigns of malicious office documents. Digital Threats: Research and Practice.;4(2):1-9.
[15] Nissim, N., Cohen, A., Elovici, Y. 2016 Dec 1. ALDOCX: detection of unknown malicious microsoft office documents using designated active learning methods based on new structural feature extraction methodology. IEEE Transactions on Information Forensics and Security. 12(3):631-46.
[16] Zakeri-Nasrabadi, M., Parsa, S., Kalaee, A. 2021 Mar 33. Format-aware learn&fuzz: deep test data generation for efficient fuzzing. Neural Computing and Applications. (5):1497-513.
[18] Natekin, A., Knoll, A. 2009 Jul 1. Gradient boosting machines, a tutorial. Frontiers in neurorobotics. 2013 Dec 4;7:21.
[19] Popescu MC, Balas VE, Perescu-Popescu L, Mastorakis N. Multilayer perceptron and neural networks. WSEAS Transactions on Circuits and Systems.;8(7):579-88.
[20] Breiman L. 2001 Oct. Random forests. Machine learning.45:5-32.
[21] Zenodo, https://zenodo.org/., 2023.
[22] Issakhani, “CIC- Evasive- PDFMal 2022 dataset.” Accessed: Apr. 12, 2024. [Online]. Available: https://www.unb.ca/cic/datasets/pdfmal-2022.html
[23] Malwarebazaar, https://bazaar.abuse.ch/., 2023.
[24] Canbek, G., Sagiroglu, S., & Temizel, TT. 2017 Oct 5. Baykal N. Binary classification performance measures/metrics: A comprehensive visualized roadmap to gain new insights. In2017 International Conference on Computer Science and Engineering (UBMK) (pp. 821-826). IEEE.
[25] Rodríguez-Pérez, R., & Bajorath, J. 2020 Oct. Interpretation of machine learning models using shapley values: application to compound potency and multi-target activity predictions. Journal of computer-aided molecular design.;34(10):1013-26.
[26] Contagio,“https://contagiodump.blogspot.com/2013/03/16800-clean-and-11960-malicious-files.html.”