Skin cancer is one of the most common cancers worldwide, and melanoma is especially dangerous if it is not de-tected early. Early diagnosis can improve survival rates. Dermo-scopic images are widely used by dermatologists to examine skin lesions. But manual analysis can be difficult and depends heavily on clinical experience. So automated methods for skin lesion classification have gained much attention. A hybrid deep learning model is proposed for multiclass skin lesion classification, aiming to improve diagnostic accuracy and robustness across diverse lesion types. The model combines two pretrained architectures, InceptionV3 and Xception, to extract features from dermoscopic images. The features obtained from both networks are fused to improve the overall representation. Transfer learning is applied to make use of pretrained weights, and data augmentation techniques are used to reduce overfitting and improve model stability. The proposed method is evaluated on a seven-class dermoscopic image dataset. The results show that the ensemble model performs better than the individual models. In addition, explainable AI techniques are used to highlight the important regions in the images that influence the model’s predictions, which improves interpretability and supports its practical use.
Introduction
The number of skin cancer cases has been increasing worldwide, making early diagnosis essential for improving survival rates. Among different skin cancers, melanoma is the most dangerous because it spreads rapidly if not detected early. Dermoscopy provides detailed images of skin lesions, but diagnosis still depends on the experience of dermatologists, leading to possible inconsistencies. To address this issue, artificial intelligence, particularly Convolutional Neural Networks (CNNs), has become an effective tool for automated skin lesion classification by learning image features such as color, texture, and shape without manual feature extraction.
Traditional skin lesion classification relied on handcrafted features based on the ABCD rule (Asymmetry, Border irregularity, Color variation, and Diameter) and machine learning methods such as Support Vector Machines, Random Forests, and k-Nearest Neighbors. Although these methods were useful, their performance depended heavily on image quality and expert knowledge. Deep learning models, including AlexNet, VGG16, ResNet, InceptionV3, and Xception, have significantly improved classification accuracy by automatically learning complex image features. Public datasets such as HAM10000, ISIC, and PH2 have further supported model development. However, challenges such as class imbalance, overfitting, and poor generalization across different datasets still remain. Techniques such as data augmentation, transfer learning, ensemble learning, and Explainable AI (XAI) methods like Grad-CAM, LIME, and SHAP have been introduced to overcome these limitations.
The proposed Hybrid IncepX-Ensemble model combines two pretrained CNN architectures, InceptionV3 and Xception, to capture complementary image features. Images from the HAM10000 dataset are resized, normalized, and augmented before training. InceptionV3 extracts multi-scale features, while Xception captures fine-grained patterns using depthwise separable convolutions. Features from both models are fused through concatenation and passed through fully connected and softmax layers for seven-class skin lesion classification. The model is trained using the Adam optimizer with categorical cross-entropy loss, while EarlyStopping, ReduceLROnPlateau, and ModelCheckpoint are used to improve training efficiency and prevent overfitting.
To improve interpretability, the framework integrates three Explainable AI techniques: Grad-CAM, LIME, and SHAP. Grad-CAM generates heatmaps highlighting important image regions, LIME explains individual predictions through local perturbations, and SHAP measures the contribution of different image regions to the final prediction. A modified LIME approach is proposed to generate more reliable explanations by first identifying lesion regions using Grad-CAM, restricting LIME analysis to these regions, and applying structured perturbations instead of random ones. This produces more stable, lesion-focused explanations that are easier for clinicians to interpret.
Experimental results demonstrate that the Hybrid IncepX-Ensemble model achieves stable training with minimal overfitting. ROC curves show high AUC values across all seven lesion classes, indicating strong discriminative performance. The confusion matrix reveals that most samples are classified correctly, with only minor confusion between visually similar classes such as melanoma and melanocytic nevi. Classification metrics, including precision, recall, F1-score, and overall accuracy, confirm strong performance across all categories. Compared with the individual InceptionV3 and Xception models, the proposed ensemble achieves higher accuracy, demonstrating that feature fusion improves classification reliability, generalization, and prediction stability for multi-class skin lesion classification.
Conclusion
Our work presented a deep learning-based framework for multi-class skin lesion classification using dermoscopic im-ages. The proposed Hybrid IncepX-Ensemble model combines InceptionV3 and Xception to learn different visual features from skin lesion images. By integrating both architectures, the model achieved improved classification performance compared to individual models.
The experimental results showed that the proposed model performs effectively across the seven lesion categories of the HAM10000 dataset. The training and validation behaviour indicated stable learning with limited overfitting. The ROC analysis, confusion matrix, and classification metrics further demonstrated the capability of the model to distinguish be-tween different lesion types with good overall performance.
In addition to classification performance, explainability was also incorporated into the framework using Grad-CAM, LIME, SHAP, and a modified LIME approach. The explainability results showed that the model mainly focuses on clinically rel-evant lesion regions during prediction. Compared to standard LIME, the modified LIME method generated more focused and meaningful explanations, especially in challenging cases. Overall, the proposed framework combines strong classi-fication capability with improved interpretability, making it
References
[1] N. Parvin et al., “Early detection of multiclass skin lesions using transfer learning-based incepx-ensemble model,” Under Review, 2025.
[2] K. M. Hosny, M. A. Kassem, and M. M. Fouad, “Explainable deep inherent learning for multi-classes skin lesion classification,” Applied Soft Computing, vol. 159, p. 111624, 2024.
[3] G. Akilandasowmya, G. Nirmaladevi, S. Suganthi, and A. Aishwariya, “Skin cancer diagnosis: Leveraging deep hidden features and ensemble classifiers for early detection and classification,” Biomedical Signal Processing and Control, vol. 88, p. 105306, 2024.
[4] H. Ding, Q. Huang, and A. Alkhayyat, “A computer aided system for skin cancer detection based on developed version of the archimedes optimization algorithm,” Biomedical Signal Processing and Control, vol. 90, p. 105870, 2024.
[5] M. A. Khan, K. Muhammad, M. Sharif, T. Akram, and S. Kadry, “Intelligent fusion-assisted skin lesion localization and classification for smart healthcare,” Neural Computing and Applications, vol. 36, no. 1,
[6] pp. 37–52, 2024.
[7] A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun, “Dermatologist-level classification of skin cancer with deep neural networks,” Nature, vol. 542, no. 7639, pp. 115–118, 2017.
[8] G. Cai, Y. Zhu, Y. Wu, X. Jiang, J. Ye, and D. Yang, “A multimodal transformer to fuse images and metadata for skin disease classification,” The Visual Computer, vol. 39, no. 7, pp. 2781–2793, 2023.
[9] Y. S. Alsahafi, M. A. Kassem, and K. M. Hosny, “Skin-net: A novel deep residual network for skin lesions classification,” Journal of Big Data, vol. 10, no. 1, p. 105, 2023.
[10] N. U. Gilal, S. A. M. Ahmed, J. Schneider, M. Househ, and M. Agus, “Mobile dermatoscopy: Class imbalance management based on blurring augmentation, iterative refining and cost-weighted recall loss,” Journal of Imaging and Graphics, vol. 11, no. 2, pp. 161–169, 2023.
[11] A. Alhudhaif, B. Almaslukh, O. O. Aseeri, O. Guler, and K. Polat, “A nonlinear automated multi-class skin lesion detection system using soft attention cnns,” Chaos, Solitons and Fractals, vol. 170, p. 113409, 2023.
[12] V.-D. Hoang, X.-T. Vo, and K.-H. Jo, “Categorical weighting domination for imbalanced classification with skin cancer,” IEEE Access, vol. 11,
[13] pp. 105 170–105 181, 2023.
[14] K. M. Hosny and M. A. Kassem, “Refined residual deep convolutional network for skin lesion classification,” Journal of Digital Imaging, vol. 35, no. 2, pp. 258–280, 2022.
[15] P. Tang, X. Yan, Y. Nan, S. Xiang, S. Krammer, and T. Lasser, “Fusionm4net: A multi-stage multi-modal learning algorithm for skin lesion classification,” Medical Image Analysis, vol. 76, p. 102307, 2022.
[16] Y. Wang, Y. Feng, L. Zhang, J. T. Zhou, Y. Liu, and R. S. M. Goh, “Adversarial multimodal fusion with attention mechanism for skin lesion classification,” Medical Image Analysis, vol. 81, p. 102535, 2022.
[17] Y. Liu, Z. Wang, Z. Li, J. Li, and T. Li, “Multiscale ensemble of convolutional neural networks for skin lesion classification,” IET Image Processing, vol. 15, no. 10, pp. 2309–2318, 2021.
[18] S. S. Chaturvedi, J. V. Tembhurne, and T. Diwan, “A multi-class skin cancer classification using deep convolutional neural networks,” Multimedia Tools and Applications, vol. 79, pp. 28 477–28 498, 2020.
[19] P. Tschandl, C. Rosendahl, and H. Kittler, “The ham10000 dataset: A large collection of dermatoscopic images,” Scientific Data, vol. 5, 2018.
[20] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” Proceedings of CVPR, 2016.
[21] F. Chollet, “Xception: Deep learning with depthwise separable convo-lutions,” Proceedings of CVPR, 2017.
[22] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” Proceedings of CVPR, 2009.