Optimización de redes neuronales profundas mediante la destilación de conocimiento para una clasificación eficiente y sostenible de imágenes médicas
Publicado 2026-04-04
Palabras clave
- Huella de carbono,
- recursos computacionales,
- aprendizaje profundo,
- dispositivos en el borde,
- consumo energético
- eficiencia energética,
- inferencia,
- destilación del conocimiento,
- arquitecturas de red neuronal,
- modelo estudiante,
- modelo maestro ...Más
Cómo citar
Derechos de autor 2026 Revista UIS Ingenierías

Esta obra está bajo una licencia internacional Creative Commons Atribución-SinDerivadas 4.0.
Resumen
El aprendizaje profundo ha transformado la medicina computacional; sin embargo, su implementación práctica sigue limitada por las elevadas exigencias de recursos computacionales, memoria y consumo energético. La creciente complejidad de los modelos de inteligencia artificial, que hoy alcanzan cientos de millones de parámetros y dependen de clústeres de GPU de alto rendimiento, dificulta su adopción en entornos con recursos restringidos, donde la latencia y la privacidad de los datos son factores críticos. El presente estudio tiene como objetivo desarrollar y validar una metodología que permita optimizar y comprimir redes neuronales profundas sin comprometer su desempeño predictivo, posibilitando así un despliegue eficiente y sostenible en dispositivos portátiles, sistemas periféricos y equipos médicos de bajo consumo energético. La técnica de Destilación del Conocimiento se empleó para transferir el conocimiento desde un modelo “maestro” robusto y de gran escala hacia un modelo “estudiante” más pequeño y eficiente. En este caso, DenseNet-121 actuó como modelo maestro y MobileNet-V2 como modelo estudiante. El modelo destilado MobileNet-V2 alcanzó un AUC-ROC de 0.865 con una desviación estándar de 0.0010, superando tanto al maestro (AUC-ROC 0.858; desviación estándar 0.0015) como al estudiante sin destilar (AUC-ROC 0.824; desviación estándar 0.0006) aplicando la estrategia Unos la cual fue la que mejores resultados obtuvo. El modelo maestro contenía 6.96 millones de parámetros (26.87 MB), mientras que el modelo estudiante utilizó 2.23 millones de parámetros (8.64 MB), logrando una reducción de tamaño de aproximadamente tres veces. Este modelo compacto ofreció una precisión comparable o superior, con menor tiempo de inferencia y consumo energético, lo que se alinea con los objetivos de sostenibilidad y facilita la adopción de inteligencia artificial médica robusta en entornos con recursos limitados. Código disponible en https://github.com/felixmejia/Knowledge_Distillation.
Descargas
Citas
- [1] C. Buciluǎ, R. Caruana, y A. Niculescu-Mizil, “Model compression”, en Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, Philadelphia PA USA: ACM, ago. 2006, pp. 535-541. doi: https://doi.org/10.1145/1150402.1150464 .
- [2] G. Hinton, O. Vinyals, y J. Dean, “Distilling the Knowledge in a Neural Network”, 2015, arXiv. doi: doi: https://doi.org/10.48550/ARXIV.1503.02531
- [3] N. C. Thompson, K. Greenewald, K. Lee, y G. F. Manso, “The Computational Limits of Deep Learning”, 2020, arXiv. doi: https://doi.org/10.48550/ARXIV.2007.05558
- [4] C. Chen et al., “Deep Learning on Computational-Resource-Limited Platforms: A Survey”, Mob. Inf. Syst., vol. 2020, pp. 1-19, mar. 2020, doi: https://doi.org/10.1155/2020/8454327
- [5] K. Tian, L. Qiao, B. Liu, G. Jiang, S. Li, y D. Li, “A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science”, Front. Comput. Sci., vol. 20, n.o 11, p. 2011355, nov. 2026, doi: https://doi.org/10.1007/s11704-025-50302-6
- [6] A. N. Mir y D. R. Rizvi, “Explainable Knowledge Distillation for Efficient Medical Image Classification”, 2025, arXiv. doi: https://doi.org/10.48550/ARXIV.2508.15251
- [7] O. Beloved Obinnaya Chinecherem, E. Enefiok. A, y P. Enyindah, “Knowledge distillation-based lightweight Convolutional Neural Networks (CNN) model for efficient breast cancer detection”, Int. J. Sci. Res. Arch., vol. 16, n.o 3, pp. 998-1006, sep. 2025, doi: https://doi.org/10.30574/ijsra.2025.16.3.2625
- [8] L. Zhao, X. Qian, Y. Guo, J. Song, J. Hou, y J. Gong, “MSKD: Structured knowledge distillation for efficient medical image segmentation”, Comput. Biol. Med., vol. 164, p. 107284, sep. 2023, doi: https://doi.org/10.1016/j.compbiomed.2023.107284
- [9] A. Sevinc, M. Ucan, y B. Kaya, “A Distillation Approach to Transformer-Based Medical Image Classification with Limited Data”, Diagnostics, vol. 15, n.o 7, p. 929, abr. 2025, doi: https://doi.org/10.3390/diagnostics15070929
- [10] M. Mahmoud, Y. Wen, X. Pan, Y. Liufu, y Y. Guan, “Evaluation of recent lightweight deep learning architectures for lung cancer CT classification”, Front. Oncol., vol. 15, p. 1647701, sep. 2025, doi: https://doi.org/10.3389/fonc.2025.1647701
- [11] G. Habib, T. jan Saleem, S. M. Kaleem, T. Rouf, y B. Lall, “A Comprehensive Review of Knowledge Distillation in Computer Vision”, 2024, arXiv. doi: https://doi.org/10.48550/ARXIV.2404.00936
- [12] Y. Himeur et al., “Applications of Knowledge Distillation in Remote Sensing: A Survey”, 2024, arXiv. doi: https://doi.org/10.48550/ARXIV.2409.12111
- [13] A. Moslemi, A. Briskina, Z. Dang, y J. Li, “A survey on knowledge distillation: Recent advancements”, Mach. Learn. Appl., vol. 18, p. 100605, dic. 2024, doi: https://doi.org/10.1016/j.mlwa.2024.100605
- [14] Z. Yuan, Y. Yan, M. Sonka, y T. Yang, “Large-scale Robust Deep AUC Maximization: A New Surrogate Loss and Empirical Studies on Medical Image Classification”, 2020, doi: https://doi.org/10.48550/ARXIV.2012.03173
- [15] J. Irvin et al., “CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison”, 2019, arXiv. doi: https://doi.org/10.48550/ARXIV.1901.07031
- [16] H. H. Pham, T. T. Le, D. Q. Tran, D. T. Ngo, y H. Q. Nguyen, “Interpreting chest X-rays via CNNs that exploit hierarchical disease dependencies and uncertainty labels”, 2019, arXiv. doi: https://doi.org/10.48550/ARXIV.1911.06475
- [17] G. Huang, Z. Liu, L. Van Der Maaten, y K. Q. Weinberger, “Densely Connected Convolutional Networks”, en 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI: IEEE, jul. 2017, pp. 2261-2269. doi: https://doi.org/10.1109/CVPR.2017.243
- [18] “Mobilenet V2 Architecture in Computer Vision”, GeeksforGeeks. [En línea]. Disponible en: https://www.geeksforgeeks.org/computer-vision/mobilenet-v2-architecture-in-computer-vision/
- [19] M. Sandler y A. Howard, “MobileNetV2: The Next Generation of On-Device Computer Vision Networks”, MobileNetV2: The Next Generation of On-Device Computer Vision Networks. [En línea]. Disponible en: https://research.google/blog/mobilenetv2-the-next-generation-of-on-device-computer-vision-networks/
- [20] PyTorch Team, “BCEWithLogitsLoss”, PyTorch Documentation. [En línea]. Disponible en: https://docs.pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html#bcewithlogitsloss
- [21] P. Yang et al., “Multi-Label Knowledge Distillation”, en 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France: IEEE, oct. 2023, pp. 17225-17234. doi: https://doi.org/10.1109/ICCV51070.2023.01584
- [22] A. De Vries, “The growing energy footprint of artificial intelligence”, Joule, vol. 7, n.o 10, pp. 2191-2194, oct. 2023, doi: https://doi.org/10.1016/j.joule.2023.09.004
- [23] D. Mhlanga, “Energy Efficiency in AI and How Power Consumption Impedes Innovation”, 2025, SSRN. doi: https://doi.org/10.2139/ssrn.5222616
- [24] V. Mehlin, S. Schacht, y C. Lanquillon, “Towards energy-efficient Deep Learning: An overview of energy-efficient approaches along the Deep Learning Lifecycle”, 2023, arXiv. doi: https://doi.org/10.48550/ARXIV.2303.01980
- [25] T. Yarally, L. Cruz, D. Feitosa, J. Sallou, y A. Van Deursen, “Uncovering Energy-Efficient Practices in Deep Learning Training: Preliminary Steps Towards Green AI”, en 2023 IEEE/ACM 2nd International Conference on AI Engineering – Software Engineering for AI (CAIN), Melbourne, Australia: IEEE, may 2023, pp. 25-36. doi: https://doi.org/10.1109/CAIN58948.2023.00012
- [26] R. Desislavov, F. Martínez-Plumed, y J. Hernández-Orallo, “Compute and Energy Consumption Trends in Deep Learning Inference”, 2021, doi: https://doi.org/10.48550/ARXIV.2109.05472
- [27] T. Aslan, P. Holzapfel, L. Stobbe, A. Grimm, N. F. Nissen, y M. Finkbeiner, “Toward climate neutral data centers: Greenhouse gas inventory, scenarios, and strategies”, iScience, vol. 28, n.o 1, p. 111637, ene. 2025, doi: https://doi.org/10.1016/j.isci.2024.111637
- [28] C. E. Tripp et al., “Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations”, 2024, arXiv. doi: https://doi.org/10.48550/ARXIV.2403.08151
- [29] I. Loshchilov y F. Hutter, “Decoupled Weight Decay Regularization”, 2017, arXiv. doi: https://doi.org/10.48550/ARXIV.1711.05101
- [30] Scikit-learn Developers, “Scikit-learn: Machine Learning in Python”, Scikit-learn. [En línea]. Disponible en: https://scikit-learn.org/
- [31] E. Ben-Baruch et al., “Asymmetric Loss For Multi-Label Classification”, 2020, arXiv. doi: https://doi.org/10.48550/ARXIV.2009.14119
