Skin Lesion Image Classification Based on Deep Learning: A Systematic Ablation Study of Heterogeneous CNN Ensembles

Main Article Content

Jinfu Ye

Keywords

skin lesion classification, deep learning, ensemble learning, convolutional neural network, ablation study, ISIC2018

Abstract

Deep learning has achieved significant progress in automated skin lesion classification, yet most ensemble methods stack convolutional neural network (CNN) backbones without systematic justification for backbone selection or analysis of inter-model error c omplementarity. We construct a heterogeneous CNN voting ensemble comprising ResNet50, ResNet101, and EfficientNet -B0 for 7-class skin lesion classification on the ISIC2018 dataset. Beyond aggregate accuracy reporting, we conduct pairwise ablation of all tw o-model combinations, quantify error diversity via Cohen's kappa and prediction disagreement matrices, and provide per-class performance analysis for all seven diagnostic categories. The three-model voting ensemble achieves 83.93% accuracy, improving over individual baselines (ResNet50: 81.22%, ResNet101: 80.62%, EfficientNet- B0: 80.22%). Statistical significance testing (McNemar's test, p < 0.01) confirms that the ensemble improvement is not attributable to random variation. Per-class analysis reveals pers istent melanoma misclassification as nevus. Systematic ablation and error diversity analysis provide stronger justification for ensemble design than aggregate accuracy alone. Our findings establish a reusable analytical framework for rigorous ensemble evaluation in broader medical imaging classification tasks.

Abstract 11 | PDF Downloads 3

References

  • [1] M. Yang, S. Wang, and C. Yu. Disease burden and incidence trend prediction of skin malignancies in China, 1990–2019. China Cancer, 31(11):853–861, 2022.
  • [2] M. J. Quinn and C. K. B. Christopher. Management of early-stage melanoma. Facial Plastic Surgery Clinics of North America, 27(1):35–42, 2019.
  • [3] A. Esteva, B. Kuprel, R. A. Novoa, et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639):115–118, 2017.
  • [4] H. A. Haenssle, C. Fink, R. Schneiderbauer, et al. Man against machine: Diagnostic performance of a deep learning CNN for dermoscopic melanoma recognition in comparison to 58 dermatologists. Annals of Oncology, 29(8):1836–1842, 2018.
  • [5] T. J. Brinker, A. Hekler, A. H. Enk, et al. A CNN trained with dermoscopic images performed on par with 145 dermatologists in a clinical melanoma image classification task. European Journal of Cancer, 111:148–154, 2019.
  • [6] N. Codella, V. Rotemberg, P. Tschandl, et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the ISIC. arXiv:1902.03368, 2019.
  • [7] M. Combalia, N. Codella, V. Rotemberg, et al. Validation of AI prediction models for skin cancer diagnosis: The 2019 ISIC challenge. The Lancet Digital Health, 4(5):e330–e339, 2022.
  • [8] N. Gessert, M. Nielsen, M. Shaikh, et al. Skin lesion classification using ensembles of multi-resolution EfficientNets with meta data. MethodsX, 7:100864, 2020.
  • [9] M. A. Khan, M. Sharif, T. Akram, et al. Skin lesion segmentation and multiclass classification using deep learning features and improved moth flame optimization. Diagnostics, 11(5):811, 2021.
  • [10] R. Sharma, A. Gupta, and K. Patel. LightEfficientNet-CA: Channel-attention EfficientNet for lightweight skin lesion classification. Computerized Medical Imaging and Graphics, 112:102332, 2024.
  • [11] Z. Wang, T. Li, X. Zhang, et al. TransFusion-Net: CNN-Transformer hybrid with cross-attention for skin lesion classification. IEEE JBHI, 27(8):3912–3923, 2023.
  • [12] Y. Yang, H. Lv, and N. Chen. A survey on ensemble learning under the era of deep learning. Artificial Intelligence Review, 56:5545–5589, 2023.
  • [13] S. Demyanov, R. Chakravorty, M. Abedini, et al. Classification of dermoscopy patterns using deep convolutional neural networks. In IEEE ISBI 2016, pages 364–368.
  • [14] A. R. Lopez, X. Giro-i-Nieto, J. Burdick, et al. Skin lesion classification from dermoscopic images using deep learning techniques. In IASTED BioMed 2017, pages 49–54.
  • [15] X. Xia, C. Xu, and B. Nan. Inception-v3 for flower classification. In ICIVC 2017, pages 783–787.
  • [16] L. Yu, H. Chen, Q. Dou, J. Qin, and P. A. Heng. Automated melanoma recognition in dermoscopy images via very deep residual networks. IEEE TMI, 36(4):994–1004, 2017.
  • [17] C. Zhao, R. Shuai, L. Ma, et al. Skin cancer image generation and classification based on Self-Attention- StyleGAN. Computer Engineering and Applications, 2021.
  • [18] S. Ahmed, M. Khan, and U. Rashid. ViT-Ensemble: Vision Transformer ensembles with soft voting for dermatological image classification. Computers in Biology and Medicine, 170:107982, 2024.
  • [19] T. G. Dietterich. Ensemble methods in machine learning. In Multiple Classifier Systems, pages 1–15, 2000.
  • [20] Z. H. Zhou. Ensemble Methods: Foundations and Algorithms. CRC Press, 2012.
  • [21] U. O. Dorj, K. K. Lee, J. Y. Choi, and M. Lee. The skin cancer classification using deep convolutional neural network. Multimedia Tools and Applications, 77:9909–9924, 2018.
  • [22] B. Harangi. Skin lesion classification with ensembles of deep convolutional neural networks. Journal of Biomedical Informatics, 86:25–32, 2018.
  • [23] L. I. Kuncheva and C. J. Whitaker. Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. Machine Learning, 51(2):181–207, 2003.
  • [24] Y. Zhang, W. Liu, and S. Wang. TransMLP-Skin: Hybrid CNN-Transformer-MLP with multi-scale features for skin lesion analysis. Medical Image Analysis, 93:103067, 2024.
  • [25] J. Li, Y. Chen, H. Zhang, et al. Focal Ensemble: Focal loss with weighted CNN voting for imbalanced skin lesion classification. Expert Systems with Applications, 238:121892, 2024.
  • [26] K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR 2016, pages 770–778.
  • [27] M. Tan and Q. V. Le. EfficientNet: Rethinking model scaling for convolutional neural networks. In ICML 2019, pages 6105–6114.
  • [28] P. Tschandl, C. Rosendahl, and H. Kittler. The HAM10000 dataset. Scientific Data, 5:180161, 2018.
  • [29] A. Paszke, S. Gross, F. Massa, et al. PyTorch: An imperative style, high-performance deep learning library. In NeurIPS 2019, pages 8024–8035.
  • [30] T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollá r. Focal loss for dense object detection. In ICCV 2017, pages 2980–2988.