Vol. 25 Núm. 2 (2026): Revista UIS Ingenierías
Artículos

Sistema Inteligente de Identificación de Aves basado en Filtrado Kalman con Aprendizaje Reforzado y BirdNET

Joiner Pérez-Rivera
Universidad Nacional de Colombia
Gustavo Ardila-Beltrán
Universidad Nacional de Colombia
Juan Pablo Hoyos-Sanchez
Universidad Nacional de Colombia

Publicado 2026-05-30

Palabras clave

  • Identificación de aves,
  • aprendizaje profundo,
  • aprendizaje reforzado,
  • BirdNET,
  • filtrado de Kalman,
  • preprocesamiento de audio
  • ...Más
    Menos

Cómo citar

Pérez-Rivera, J., Ardila-Beltrán , G., & Hoyos-Sanchez, J. P. (2026). Sistema Inteligente de Identificación de Aves basado en Filtrado Kalman con Aprendizaje Reforzado y BirdNET. Revista UIS Ingenierías, 25(2), 81–94. https://doi.org/10.18273/revuin.v25n2-2026007

Resumen

En este artículo, se investiga el aprendizaje por refuerzo (RL) como método para optimizar la identificación automática de especies de aves a partir de registros acústicos. Se propone un sistema basado en un flujo de trabajo que integra la segmentación de señales de audio, filtrado pasa banda y filtrado Kalman, cuyos parámetros son optimizados dinámicamente mediante Q-learning. La función de recompensa se define a partir de las puntuaciones de confianza proporcionadas por el clasificador BirdNET. El agente de RL ajusta adaptativamente los parámetros del filtro de Kalman en función de las condiciones acústicas variables del entorno. El esquema fue entrenado sobre un conjunto de 50 grabaciones de campo y evaluado sobre 315 grabaciones del conjunto de datos BirdCLEF del año 2025, obteniendo un incremento promedio del 25% (DE ± 8.3%) en los niveles de confianza de clasificación respecto al método base.

Descargas

Los datos de descargas todavía no están disponibles.

Citas

  1. [1] D. Funosas, L. Barbaro, L. Schillé, A. Elger, B. Castagneyrol, and M. Cauchoix, “Assessing the potential of BirdNET to infer European bird communities from large-scale ecoacoustic data,” Ecol. Indic., vol. 164, p. 112146, Jul. 2024, doi: https://doi.org/10.1016/J.ECOLIND.2024.112146
  2. [2] S. Kahl, C. M. Wood, M. Eibl, and H. Klinck, “BirdNET: A deep learning solution for avian diversity monitoring,” Ecol. Inform., vol. 61, p. 101236, Mar. 2021, doi: https://doi.org/10.1016/J.ECOINF.2021.101236
  3. [3] C. Wu, S. Kosuru, and S. T, “Bird species identification from audio data,” 2023 IEEE Ninth International Conference on Big Data Computing, 2023. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10233991/
  4. [4] M. T. Lopes et al., “Automatic bird species identification for large number of species,” 2011 IEEE International Symposium on Multimedia, 2011, doi: https://doi.org/10.1109/ISM.2011.27
  5. [5] M. Uddin, A. Asaduzzaman, R. Soza, and C. Minkler, “Avian Song Identification Using CNN,” 2024 IEEE Green Technologies Conference (GreenTech), 2024. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10520499/
  6. [6] S. Natarajan et al., “Deep neural networks for speech enhancement and speech recognition: A systematic review,” Ain Shams Eng. J., vol. 16, no. 7, p. 103405, Jul. 2025, doi: https://doi.org/10.1016/j.asej.2025.103405
  7. [7] M. Khodarahmi and V. Maihami, “A Review on Kalman Filter Models,” Arch. Computat. Methods Eng., vol. 30, no. 1, pp. 727–747, Jan. 2023, doi: https://doi.org/10.1007/s11831-022-09815-7
  8. [8] Y. H. Goh, P. Raveendran, and Y. L. Goh, “Robust speech recognition system using bidirectional Kalman filter,” IET Signal Process., vol. 9, no. 6, pp. 491–497, Aug. 2015, doi: https://doi.org/10.1049/IET-SPR.2014.0109
  9. [9] Y. Zhang, M. Yu, H. Zhang, D. Yu, and D. L. Wang, “Neuralkalman: A Learnable Kalman Filter for Acoustic Echo Cancellation,” ASRU 2023, doi: https://doi.org/10.1109/ASRU57964.2023.10389780
  10. [10] T. F. Tavares, “Open-set classification approaches to automatic bird song identification,” IEEE Lat. Am. Trans., 2022. [Online]. Available: https://latamt.ieeer9.org/index.php/transactions/article/view/6832
  11. [11] Y. Zhou, C. Xiong, and R. Socher, “Improving end-to-end speech recognition with policy learning,” ICASSP 2018. [Online]. Available: https://ieeexplore.ieee.org/document/8462361/
  12. [12] J. Gibson et al., “A reinforcement learning approach to speech coding,” Information, 2022. [Online]. Available: https://www.mdpi.com/2078-2489/13/7/331
  13. [13] P. Giannakopoulos, A. Pikrakis, and Y. Cotronis, “A deep reinforcement learning approach to audio-based navigation in a multi-speaker environment,” ICASSP 2021. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9415013
  14. [14] P. Giannakopoulos, A. Pikrakis, and Y. Cotronis, “Improving post-processing of audio event detectors using reinforcement learning,” IEEE Access, 2022. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9853543/
  15. [15] S. Lathuilière, B. Massé, P. Mesejo, and R. Horaud, “Neural network based reinforcement learning for audio–visual gaze control in human–robot interaction,” Pattern Recognit. Lett., 2019. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167865518302113
  16. [16] R. Fakoor, X. He, I. Tashev, and S. A. Zarar, “Reinforcement Learning To Adapt Speech Enhancement to Instantaneous Input Signal Quality,” 2017. [Online]. Available: https://arxiv.org/pdf/1711.10791
  17. [17] Y.-L. Shen et al., “Reinforcement learning based speech enhancement for robust speech recognition,” ICASSP 2019. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/8683648/
  18. [18] R. Jaiswal and D. Romero, “Implicit wiener filtering for speech enhancement in non-stationary noise,” 2021 Int. Conf. on Information Science, 2021. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9440639/
  19. [19] K. Chen et al., “Research on Kalman filter fusion navigation algorithm assisted by CNN-LSTM neural network,” Appl. Sci., 2024. [Online]. Available: https://www.mdpi.com/2076-3417/14/13/5493
  20. [20] S. Mohammed et al., “Audio Denoising Using Deep Neural Networks,” Springer, vol. 101, pp. 33–47, 2022, doi: https://doi.org/10.1007/978-981-16-7610-9_3
  21. [21] R. Fakoor et al., “Reinforcement Learning To Adapt Speech Enhancement to Instantaneous Input Signal Quality,” 2017. [Online]. Available: https://arxiv.org/pdf/1711.10791
  22. [22] S. Roy and K. Paliwal, “Robustness and sensitivity tuning of the Kalman filter for speech enhancement,” MDPI, 2021. [Online]. Available: https://www.mdpi.com/2624-6120/2/3/27
  23. [23] S. Latif et al., “A survey on deep reinforcement learning for audio-based applications,” Artif. Intell. Rev., vol. 56, no. 3, pp. 2193–2240, Mar. 2023, doi: 10.1007/S10462-022-10224-2.
  24. [24] S. Lathuilière, B. Massé, P. Mesejo, and R. Horaud, “Deep reinforcement learning for audio-visual gaze control,” IROS 2018. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/8594327/
  25. [25] S. Lathuilière et al., “Neural network based reinforcement learning for audio–visual gaze control,” Pattern Recognit. Lett., 2019. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167865518302113
  26. [26] R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction. MIT Press, 2018. [Online]. Available: https://mitpress.mit.edu/9780262039246/reinforcement-learning/
  27. [27] MathWorks,“¿Qué son las redes neuronales convolucionales? - MATLAB & Simulink.” [Online]. Available: https://es.mathworks.com/discovery/convolutional-neural-network.html
  28. [28] C. Chen et al.,“Leveraging modality-specific representations for audio-visual speech recognition via reinforcement learning,” AAAI 2023. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/23484
  29. [29] Y. Singh and R. Mehra, “Relative study of measurement noise covariance R and process noise covariance Q of the Kalman filter in estimation,” OSR J. Electr. Electron. Eng., 2015. [Online]. Available: https://www.academia.edu/download/47175517/P01061112116.pdf
  30. [30] D. Winiarska, P. Szymański, and T. O. Osiejuk, “Detection ranges of forest bird vocalisations: guidelines for passive acoustic monitoring,” Sci. Rep., 2024, doi: https://doi.org/10.1038/s41598-024-51297-z
  31. [31] “BirdNET Sound ID – The easiest way to identify birds by sound.” [Online]. Available: https://birdnet.cornell.edu/
  32. [32] “BirdCLEF 2024 | Kaggle.” [Online]. Available: https://www.kaggle.com/competitions/birdclef-2024