علوم رایانشی

علوم رایانشی

مقایسه عملکرد الگوریتم‌های DQN و Double DQN در محیط‌های یادگیری تقویتی با منابع محدود

نوع مقاله : مقاله پژوهشی

نویسندگان
1 کارشناسی مهندسی کامپیوتر، دانشکده مهندسی برق و کامپیوتر ، دانشگاه تبریز ، تبریز ، ایران
2 کارشناسی ارشد مهندسی مخابرات امن و رمزنگاری، پژوهشکده فضای مجازی، دانشگاه شهید بهشتی، تهران، ایران
3 دکتری مهندسی برق دانشکده مهندسی برق و کامپیوتر، دانشگاه تبریز، تبریز، ایران
4 دکتری مهندسی کامپیوتر، دانشکده چند رسانه ای، دانشگاه هنر اسلامی تبریز، تبریز، ایران
10.22034/csj.2026.251328
چکیده
یادگیری تقویتی عمیق طی سال‌های اخیر به‌عنوان یکی از رویکردهای پیشرفته در حل مسائل تصمیم‌گیری پیچیده مطرح شده است و توانسته در حوزه‌هایی مانند بازی‌های ویدئویی، رباتیک و کنترل خودکار عملکردی چشمگیر نشان دهد. با این ‌حال، بررسی کارایی این الگوریتم‌ها در محیط‌هایی با منابع محاسباتی محدود کمتر مورد توجه قرار گرفته است. در این مقاله، دو الگوریتم مطرح الگوریتم یادگیری تقویتی عمیق شامل شبکه کیو عمیق Deep Q-Network (DQN) و شبکه کیو عمیق دوگانه مورد مقایسه و ارزیابی قرار گرفتند. هر دو الگوریتم با بهره‌گیری از شبکه‌های عصبی هم‌آمیختی (CNN) و در محیط ساده و کم منبع گوگل کولَب پیاده‌سازی شدند. آزمایش‌ها بر روی بازیBreakout انجام شد و با تنظیم حداقلی ابرپارامترها، عملکرد هر الگوریتم تحلیل گردید. نتایج نشان دادند که شبکه کیو عمیق دوگانه در برخی شبیه‌سازی‌ها دقت و پایداری بیشتری دارد، در حالی ‌که vanilla DQN با وجود سادگی ساختار خود، همچنان کارایی قابل‌قبولی ارائه می‌دهد. به‌طورکلی، می‌توان نتیجه گرفت که در محیط‌های با محدودیت منابع، استفاده از رویکردهای ساده‌تر الگوریتم یادگیری تقویتی عمیق می‌تواند گزینه‌ای بهینه‌تر و کارآمدتر باشد.
کلیدواژه‌ها
موضوعات

[1] Oudouar, F., Bir-Jmel, A., Grissette, H., Douiri, S. M., Himeur, Y., Miniaoui, S., Atalla, S., & Mansoor, W. 2025. An empirical evaluation of neural network architectures for 3D spheroid segmentation. Computers, 14, 86.
[2] Aloufi, N., & Aljuhani, A. 2025. Empirical evaluation of prompting strategies for Python syntax error detection with LLMs. Applied Sciences, 15, 9223.
[3] Donahue, J., et al. 2015. Long-term recurrent convolutional networks for visual recognition and description. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.
[4] Shaheen, A., Badr, A., Abohendy, A., Alsaadawy, H., & Alsayad, N. (2025). Reinforcement learning in strategy-based and atari games: A review of google deepminds innovations. arXiv preprint arXiv:2502.10303.
[5] Elmezain, M., Saoud, L. S., Sultan, A., Heshmat, M., Seneviratne, L., & Hussain, I. 2025. Advancing underwater vision: A survey of deep learning models for underwater object recognition and tracking. IEEE Access, 2025.
[6] Lipton, Z. C., Berkowitz, J., & Elkan, C. 2015. A critical review of recurrent neural networks for sequence learning. arXiv preprint arXiv:1506.00019, 2015.
[7] Spoerer, C. J., et al. 2020. Recurrent neural networks can explain flexible trading of speed and accuracy in biological vision. PLoS Computational Biology, 16, e1008215.
[8] Penzkofer, A., Schaefer, S., Strohm, F., Bâce, M., Leutenegger, S., & Bulling, A. 2025. Int-HRL: Towards intention-based hierarchical reinforcement learning. Neural Computing and Applications, 37, 18823-18834.
[9] Bellemare, M. G., et al. 2013. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47, 253-279.
[10] Devlin, J., et al. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers).
[11] Yao, J., Zhang, J., & Zhuo, L. 2025. SGNet: Sequence grouping network via vision-language model for text-guided video summarization. IEEE Journal of Selected Topics in Signal Processing.
[12] Baek, I.-C., et al. 2025. IPCGRL: Language-instructed reinforcement learning for procedural level generation. arXiv preprint arXiv:2503.12358.
[13] Joshi, V., Xu, Z., Liu, B., Stone, P., & Zhang, A. 2025. Benchmarking massively parallelized multi-task reinforcement learning for robotics tasks. arXiv preprint arXiv:2507.23172, 2025.
[14] Hasan, M. M., Das, R. K., Hassan, M., Razia, S., Ani, J. F., Khushbu, S. A., & Islam, M. 2025. Hybrid deep learning: A comparative study on AI algorithms in natural language processing for text classification. Bulletin of Electrical Engineering and Informatics, 14, 551-559.
[15] Greff, K., et al. 2016. LSTM: A search space odyssey. IEEE Transactions on Neural Networks and Learning Systems, 28, 2222-2232.
[16] Fan, C.-H., Wu, R.-T., & Chang, Y.-I. 2025. Robotic inspection for autonomous crack segmentation and exploration using deep reinforcement learning. Automation in Construction, 106009.
[17] Snodgrass, S., & Ontanón, S. 2016. Learning to generate video game maps using Markov models. IEEE Transactions on Computational Intelligence and AI in Games, 9(4), 410–422.
[18] Ulaş, B., Szklenár, T., & Szabó, R. 2025. Detection of oscillation-like patterns in eclipsing binary light curves using neural network-based object detection algorithms. Astronomy & Astrophysics, 695, A81.
[19] Mohanty, S., Sahoo, D., Das, S., Acharya, A. A., & Panda, N. 2025. Sentiment analysis using CNN for emotion extraction to synthesize natural speech. Procedia Computer Science, 258, 2737-2747.
[20] Van Hasselt, H., Guez, A., & Silver, D. 2016. Deep reinforcement learning with double یادگیری کیو . Proceedings of the AAAI Conference on Artificial Intelligence.
[21] Guleria, P., Frnda, J., & Srinivasu, P. N. 2025. NLP based text classification using TF-IDF enabled fine-tuned long short-term memory: An empirical analysis. Array, 100467.
[22] Tanis, L. J., Cunha, R. F., & Zullich, M. 2025. Bridging faithfulness of explanations and deep reinforcement learning: A Grad-CAM analysis of Space Invaders. In Proceedings of the 20th International Conference on the Foundations of Digital Games, 1-7.
[23] Sutton, R. S. 1988. Learning to predict by the methods of temporal differences. Machine Learning, 3, 9-44.
[24] Retkowski, F. 2020. Reinforcement learning for sequence-to-sequence dialogue systems. Karlsruhe Institute of Technology, 2020.
[25] Mnih, V., et al. 2015. Human-level control through deep reinforcement learning. Nature, 518, 529-533.
[26] Gragnaniello, D., Greco, A., Sansone, C., & Vento, B. 2025. FLAME: Fire detection in videos combining a deep neural network with a model-based motion analysis. Neural Computing and Applications, 37, 6181-6197.
[27] Cong, R., Yang, N., Liu, H., Zhang, D., Huang, Q., Kwong, S., & Zhang, W. 2025. TRNet: Two-tier recursion network for co-salient object detection. IEEE Transactions on Circuits and Systems for Video Technology, 2025.
[28] Singh, N. J., Singh, K. R., Hoque, N., & Bhattacharyya, D. K. 2025. Massive IoT network traffic analysis using ML and DL methods: An empirical evaluation. The Journal of Supercomputing, 81, 1107.
[29] Subramanian, A., Price, S., Kumbhar, O., Sizikova, E., Majaj, N. J., & Pelli, D. G. 2025. Benchmarking the speed–accuracy tradeoff in object recognition by humans and neural networks. Journal of Vision, 25, 4-4.
[30] Jang, H., Sinha, P., & Boix, X. 2025. Configural processing as an optimized strategy for robust object recognition in neural networks. Communications Biology, 8, 386.
[31] Meng, W., et al. 2025. Advances in UAV path planning: A comprehensive review of methods, challenges, and future directions. Drones, 9, 5.