علوم رایانشی

علوم رایانشی

بهبود بخشبندی معنایی تصاویر شهری با ارائه یک تابع هزینه جدید مبتنی بر ابرپیکسل

نوع مقاله : مقاله پژوهشی

نویسندگان
گروه مهندسی کامپیوتر، دانشکده مهندسی، دانشگاه فردوسی مشهد، مشهد، ایران
10.22034/csj.2026.245085
چکیده
بخش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌بندی معنایی[1] تصاویر یکی از مؤلفه‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های کلیدی در تحلیل تصاویر شهری است که عملکرد مناسب آن تاثیر مستقیمی بر کارایی و سرعت سیستم‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های بینایی ماشین در این حوزه دارد. بر همین اساس، در این پژوهش یک تابع هزینه جدید برای بهبود بخش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌بندی معنایی در تصاویر شهری ارائه شده است. یکی از چالش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های اصلی در این زمینه، عدم توازن[2] داده‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌ها است که با توجه به تفاوت نرخ عدم توازن در هر کلاس در مجموعه داده‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های چند کلاسه، انتخاب تابع هزینه مناسب را نسبت به مسائل دو کلاسه دشوارتر می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌سازد. علاوه بر این، مشکلاتی مانند شناسایی دقیق مرز اجسام و حفظ پیوستگی نواحی نیز باعث کاهش دقت روش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های موجود می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌شوند. بنابراین، راه‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌کار پیشنهادی با شناسایی نواحی پیچیده و پرتراکم در تصویر، که می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌تواند شامل کلاس‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌هایی با شباهت ظاهری زیاد، مرز بین چند کلاس، یا کلاس‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌ها و اجسام کوچک باشد، و تمرکز بیشتر بر آنها در زمان آموزش شبکه به بهبود عملکرد مدل کمک می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌کند. در این روش، ابتدا هر تصویر توسط الگوریتم SLIC[3] به ابرپیکسل‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های[4] مجزا تقسیم شده و تأثیر هر ناحیه با تحلیل ابرپیکسل‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌ها و اعمال وزن‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌دهی هدفمند در تابع هزینه پیشنهادی تعیین می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌شود. در ادامه برای ارزیابی اثربخشی این روش، مجموعه‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌ای از آزمایش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌ها با استفاده از مدل U-Net و دو مجموعه داده از تصاویر شهری انجام شده است. نتایج این آزمایش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌ها نشان می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌دهد که استفاده از تابع هزینه پیشنهادی در مقایسه با سایر روشها به طور میانگین موجب بیش از 1 درصد بهبود در معیارهای IoU و امتیاز F1 شده و تمرکز بر نواحی پیچیده تصاویر می‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌تواند گامی موثر در رفع محدودیت‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌های موجود در بخش‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌‌بندی معنایی باشد.



1. Semantic Segmentation


2. Imbalance


3. Simple Linear Iterative Clustering


4. Superpixels
کلیدواژه‌ها
موضوعات

[1]   Minaee, S., Boykov, Y. Y., Porikli, F., Plaza, A. J., Kehtarnavaz, N. & Terzopoulos, D. (2021). “Image segmentation using deep learning: A survey,” IEEE Trans. Pattern Anal. Mach. Intell.
[2]   Otsu, N. (1979). “A threshold selection method from gray-level histograms,” IEEE Trans. Syst. Man. Cybern., vol. 9, no. 1, pp. 62–66.
[3]   Dhanachandra, N., Manglem, K. & Chanu, Y. J. (2015). “Image segmentation using K-means clustering algorithm and subtractive clustering algorithm,” Procedia Comput. Sci., vol. 54, pp. 764–771, 2015.
[4]   Nock, R. & Nielsen, F. (2004). “Statistical region merging,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 26, no. 11, pp. 1452–1458.
[5]   Najman, L., & Schmitt, M. (1994). “Watershed of a continuous function,” Signal Processing, vol. 38, no. 1, pp. 99–112.
[6]   Boykov, Y., Veksler, O. & Zabih, R. (2001). “Fast approximate energy minimization via graph cuts,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 23, no. 11, pp. 1222–1239.
[7]   Kass, M., Witkin, A., & Terzopoulos, D. (1998). “Snakes: Active contour models,” Int. J. Comput. Vis., vol. 1, no. 4, pp. 321–331.
[8]   Plath, N., Toussaint, M. & Nakajima, S. (2009). “Multi-class image segmentation using conditional random fields and global classification,” in Proceedings of the 26th annual international conference on machine learning, 2009, pp. 817–824.
[9]   J. Ma et al., “Loss odyssey in medical image segmentation,” Med. Image Anal., p. 102035, 2021.
[10] Ronneberger, O., Fischer, P., & Brox, T. (2015). “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention, pp. 234–241.
[11] Wu, Z., Shen, C. & Van den Hengel, A. (2016). “Bridging category-level and instance-level semantic image segmentation,” arXiv Prepr. arXiv1605.06885.
[12] Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision, pp. 2980–2988.
[13] Arulananth, T. S. et al., (2024). “Semantic segmentation of urban environments: Leveraging U-Net deep learning model for cityscape image analysis,” PLoS One, vol. 19, no. 4, p. e0300767.
[14] Sahragard, E., Farsi, H. & Mohamadzadeh, S. (2025).“Advancing semantic segmentation: Enhanced UNet algorithm with attention mechanism and deformable convolution,” PLoS One, vol. 20, no. 1, p. e0305561.
[15] Zhang, Z. & Li, G. (2025). “UAV Imagery Real-Time Semantic Segmentation with Global--Local Information Attention,” Sensors, vol. 25, no. 6, p. 1786.
[16] Yang, G., Wang, Y., Shi, D. & Wang, Y. (2025). “Golden Cudgel Network for Real-Time Semantic Segmentation,” in Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 25367–25376.
[17] Lyu, Y., Vosselman, G., Xia, G.-S., Yilmaz, A. & Yang, M. Y. (2020). “UAVid: A semantic segmentation dataset for UAV imagery,” ISPRS J. Photogramm. Remote Sens., vol. 165, pp. 108–119.
[18] Hesham et al., S. A. S. (2025). “Exploiting temporal state space sharing for video semantic segmentation,” in Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 24211–24221.
[19] Li, X. et al., (2020).“Improving semantic segmentation via decoupled body and edge supervision,” in European Conference on Computer Vision, pp. 435–452.
[20] Milletari, F., Navab, N., & Ahmadi, S. A. (2016). “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in 2016 fourth international conference on 3D vision (3DV), 2016, pp. 565–571.
[21] Drozdzal, M., Vorontsov, E., Chartrand, G., Kadoury, S. & Pal, C. (2016).“The importance of skip connections in biomedical image segmentation,” in Deep learning and data labeling for medical applications, Springer, 2016, pp. 179–187.
[22] Rahman M. A., & Wang, Y. (2016). “Optimizing intersection-over-union in deep neural networks for image segmentation,” in International symposium on visual computing, 2016, pp. 234–244.
[23] Sudre, C. H., Li, W., Vercauteren, T., Ourselin, S. & Cardoso, M. J. (2017).“Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations,” in Deep learning in medical image analysis and multimodal learning for clinical decision support, Springer, 2017, pp. 240–248.
[24] Salehi, S. S. M., Erdogmus, D., & Gholipour, A. (2017). “Tversky loss function for image segmentation using 3D fully convolutional deep networks,” in International workshop on machine learning in medical imaging, pp. 379–387.
[25] Abraham, N., & Khan, N. M. (2019).“A novel focal tversky loss function with improved attention u-net for lesion segmentation,” in 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019), pp. 683–687.
[26] Kervadec, H., Bouchtiba, J., Desrosiers, C., Granger, E., Dolz, J., & Ben Ayed, I. (2021). “Boundary loss for highly unbalanced segmentation,” Med. Image Anal., vol. 67, p. 101851, 2021.
[27] Karimi, D. & Salcudean, S. E. (2019). “Reducing the hausdorff distance in medical image segmentation with convolutional neural networks,” IEEE Trans. Med. Imaging, vol. 39, no. 2, pp. 499–513, 2019.
[28] Wong, K. C. L., Moradi, M., Tang, H. & Syeda-Mahmood, T. (2018). “3D segmentation with exponential logarithmic loss for highly unbalanced object sizes,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 612–619.
[29] Taghanaki, S. A. et al., (2019). “Combo loss: Handling input and output imbalance in multi-organ segmentation,” Comput. Med. Imaging Graph., vol. 75, pp. 24–33.
[30] Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J. & Maier-Hein, K. H. (2021).“nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation,” Nat. Methods, vol. 18, no. 2, pp. 203–211.
[31] Zhu, W. et al., (2019). “AnatomyNet: deep learning for fast and fully automated whole-volume segmentation of head and neck anatomy,” Med. Phys., vol. 46, no. 2, pp. 576–589.
[32] Brugnara, G.et al., (2020). “Automated volumetric assessment with artificial neural networks might enable a more accurate assessment of disease burden in patients with multiple sclerosis,” Eur. Radiol., vol. 30, no. 4, pp. 2356–2364.
[33] Arega, T. W., Bricq, S., & Meriaudeau, F. (2022). “Using polynomial loss and uncertainty information for robust left atrial and scar quantification and segmentation,” in Challenge on Left Atrial and Scar Quantification and Segmentation, Springer, 2022, pp. 133–144.
[34] El Jurdi, R., Petitjean, C., Honeine, P., Cheplygina, V. & Abdallah, F. (2021). “High-level prior-based loss functions for medical image segmentation: A survey,” Comput. Vis. Image Underst., vol. 210, p. 103248.
[35] Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P. & Süsstrunk, S. (2012). “SLIC superpixels compared to state-of-the-art superpixel methods,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 34, no. 11, pp. 2274–2282.
[36] Cordts et al., M. (2016). “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3213–3223.
[37] Glorot, X. & Bengio, Y. (2010). “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 249–256.
[38] Kingma, D. P. & Ba, J. (2014). “Adam: A method for stochastic optimization,” arXiv Prepr. arXiv1412.6980.