The ShuBAM-TFNet: A Lightweight Temporal–Frequency Attention Network for Environmental Sound Recognition

Penulis

  • Kabiru Muhammed Nasiru Department of Computer Engineering, Ahmadu Bello University, Zaria, Nigeria
  • Shehu Mohammed Yusuf Department of Computer Engineering, Ahmadu Bello University, Zaria, Nigeria.
  • Sani Saleh Saminu Department of Computer Engineering, Ahmadu Bello University, Zaria, Nigeria

Kata Kunci:

EnvironmentalSound Classification, Audio Signal Processing, Edge Computing, Lightweight Deep Learning, Temporal Frequency Attention

Abstrak

Efficient environmental sound recognition is essential for smart-city monitoring and edge-based acoustic sensing, yet practical deployment is often constrained by the computational cost and parameter size of deep learning models. This paper presents ShuBAM-TFNet, a lightweight convolutional neural network designed to achieve high recognition accuracy while remaining suitable for resource-limited devices. The proposed architecture processes Log-Mel spectrograms using dilated and time-distributed convolutional layers to capture long-range temporal dependencies and localized temporal patterns without increasing model complexity. A lightweight temporal–frequency attention mechanism is introduced to adaptively emphasize informative regions across time and frequency dimensions, while Bottleneck Attention Modules are incorporated to refine channel and spatial representations. The model is evaluated on the UrbanSound8K and ESC-10 benchmark datasets using standard evaluation protocols. On UrbanSound8K, the full ShuBAM-TFNet achieves an average accuracy of 90.16 ± 1.05% with 72,390 trainable parameters, converging in approximately 27 epochs. Ablation studies show that removing the temporal–frequency attention module results in a 2–3% accuracy drop, confirming its central contribution, while eliminating dilated convolutions leads to slower convergence despite stable accuracy. Notably, removing Bottleneck Attention Modules yields the best-performing variant, achieving 91.76 ± 0.90% accuracy with a reduced parameter count of 70,754. Cross-dataset evaluation on ESC-10 using 5-fold cross-validation achieves a mean accuracy of 84.00 ± 2.41% and a macro F1-score of 0.838, demonstrating good generalization. These results indicate that attention-guided lightweight architectures provide a practical alternative to search-based or heavily parameterized models for real-time environmental sound recognition on edge devices.

Biografi Penulis

Kabiru Muhammed Nasiru, Department of Computer Engineering, Ahmadu Bello University, Zaria, Nigeria

Kabiru Muhammed Nasiru holds a Bachelor of Engineering (B.Eng.) degree in Computer Engineering. His academic interests include computer systems, machine learning, embedded systems, and intelligent computing applications. He can be contacted at [email protected]

Shehu Mohammed Yusuf, Department of Computer Engineering, Ahmadu Bello University, Zaria, Nigeria.

Dr. Shehu Mohammed Yusuf is a Senior Lecturer in the Department of Computer Engineering at Ahmadu Bello University, Zaria, Nigeria, with over a decade of teaching and research experience. He holds a B.Eng. in Electrical Engineering, an MSc in Electrical Engineering and a Ph.D. in Computer Engineering. He is also a Huawei Certified ICT Associate (AI), Huawei Certified Academy Instructor, and a COREN registered engineer. His research interests span Artificial Intelligence, Natural Language Processing, Computer Vision, and Bioinformatics.

Email: [email protected]

Sani Saleh Saminu, Department of Computer Engineering, Ahmadu Bello University, Zaria, Nigeria

Sani Saleh SAMINU received his MSc degree in Artificial Intelligence from the Department of Computer Engineering, Ahmadu Bello University (ABU), Zaria, Nigeria, and previously obtained his BSc(Ed) in Computer Science from the same institution. His research interests include environmental sound recognition, plant disease recognition, computer vision, deep learning, zero-shot learning, and intelligent systems with applications in agriculture and healthcare. He can be contacted at [email protected]

Referensi

[1] S. Tyagi, K. Aggarwal, D. Kumar, and S. Garg, “Urban sound classification for audio analysis using long short-term memory,” NEU Journal for Artificial Intelligence and Internet of Things, vol. 1, no. 2, pp. 1–11, 2023.

[2] W. Zhao, H. Wang, Y. Chen, X. Pan, K. Zhang, and Z. Bai, “An environmental sound classification algorithm based on multiscale channel feature fusion,” IEEE Access, 2023.

[3] M. Avadhani and A. P. Bidargaddi, “Multi-class urban sound classification with deep learning architectures,” in Proc. IEEE Conference, 2024, pp. 1–7.

[4] H. Lu, H. Zhang, and A. Nayak, “A deep neural network for audio classification with a classifier attention mechanism,” arXiv preprint arXiv:2006.09815, 2020.

[5] K. Zaman, M. Sah, C. Direkoglu, and M. Unoki, “A survey of audio classification using deep learning,” IEEE Access, vol. 11, pp. 106620–106649, 2023.

[6] J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proc. ACM Multimedia, 2014, pp. 1041–1044.

[7] K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proc. ACM Multimedia, 2015, pp. 1015–1018.

[8] R. Mahum, A. Irtaza, A. Javed, H. A. Mahmoud, and H. Hassan, “DeepDet: YAMNet with bottleneck attention module for TTS synthesis detection,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 1–18, 2024.

[9] W. Mu, B. Yin, X. Huang, J. Xu, and Z. Du, “Environmental sound classification using temporal–frequency attention based convolutional neural network,” Scientific Reports, vol. 11, no. 1, pp. 21552, 2021.

[10] F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122, 2015.

[11] S. Abdoli, P. Cardinal, and A. L. Koerich, “End-to-end environmental sound classification using a 1D convolutional neural network,” Expert Systems with Applications, vol. 136, pp. 252–263, 2019.

[12] J. Park, S. Woo, J.-Y. Lee, and I. S. Kweon, “BAM: Bottleneck attention module,” arXiv preprint arXiv:1807.06514, 2018.

[13] X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in Proc. IEEE CVPR, 2018, pp. 6848–6856.

[14] D. Ranmal, P. Ranasinghe, T. Paranayapa, D. Meedeniya, and C. Perera, “ESC-NAS: Environment sound classification using hardware-aware neural architecture search for the edge,” Sensors, vol. 24, no. 12, pp. 3749, 2024.

[15] A. Kaçar, A. Derya, and İ. Türkoğlu, “Enhancing environmental sound classification performance through data fusion: A comparative machine learning analysis,” 2025.

[16] L. Liu, “Speech recognition method for English translators in noisy environments based on attention mechanism,” International Journal of High Speed Electronics and Systems, vol. 34, no. 4, pp. 2540350, 2025.

[17] S. Divya Lakshmi and N. Suresh Kumar, “ResBiA-FusionNET: A robust deep learning framework with harmonic-contrast mel spectrogram for audio sound classification,” International Journal of Pattern Recognition and Artificial Intelligence, vol. 39, no. 15, pp. 2551028, 2025.

[18] I. H. Sarker, “Deep learning: A comprehensive overview on techniques, taxonomy, applications and research directions,” SN Computer Science, vol. 2, pp. 420, 2021.

[19] M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,” ACM Computing Surveys, vol. 48, pp. 1–46, 2016.

[20] D. Meedeniya, I. Ariyarathne, M. Bandara, R. Jayasundara, and C. Perera, “A survey on deep learning based forest environment sound classification at the edge,” ACM Computing Surveys, vol. 56, pp. 66, 2023.

[21] D. Stefani, S. Peroni, and L. Turchet, “A comparison of deep learning inference engines for embedded real-time audio classification,” in Proc. DAFx, 2022, pp. 256–263.

[22] A. Elhanashi, P. Dini, S. Saponara, and Q. Zheng, “Integration of deep learning into the IoT: A survey of techniques and challenges for real-world applications,” Electronics, vol. 12, pp. 4925, 2023.

[23] D. Meedeniya, Deep Learning: A Beginners’ Guide, CRC Press, Boca Raton, FL, USA, 2023.

[24] B. Wu et al., “FBNet: Hardware-aware efficient convnet design via differentiable neural architecture search,” in Proc. IEEE CVPR, 2019, pp. 10734–10742.

[25] C. White et al., “Neural architecture search: Insights from 1000 papers,” arXiv preprint arXiv:2301.08727, 2023.

[26] M. Risso et al., “Lightweight neural architecture search for temporal convolutional networks at the edge,” IEEE Transactions on Computers, vol. 72, pp. 744–758, 2023.

[27] J. Lin et al., “MCUNet: Tiny deep learning on IoT devices,” Advances in Neural Information Processing Systems, vol. 33, pp. 11711–11722, 2020.

[28] D. T. Speckhard et al., “Neural architecture search for energy-efficient always-on audio machine learning,” Neural Computing and Applications, vol. 35, pp. 12133–12144, 2023.

[29] B. Elliott et al., “Cyber-physical analytics: Environmental sound classification at the edge,” in Proc. IEEE WF-IoT, 2020, pp. 1–6.

[30] A. Bansal and N. K. Garg, “Environmental sound classification: A descriptive review of the literature,” Intelligent Systems with Applications, vol. 16, pp. 200115, 2022.

[31] A. Andreadis, G. Giambene, and R. Zambon, “Monitoring illegal tree cutting through ultra-low-power smart IoT devices,” Sensors, vol. 21, pp. 7593, 2021.

[32] I. Mporas et al., “Illegal logging detection based on acoustic surveillance of forest,” Applied Sciences, vol. 10, pp. 7379, 2020.

[33] G. Peruzzi, A. Pozzebon, and M. Van Der Meer, “Detecting forest fires with embedded machine learning models using audio and images,” Sensors, vol. 23, pp. 783, 2023.

[34] S. K. Shah, Z. Tariq, and Y. Lee, “IoT-based urban noise monitoring using deep learning,” in Proc. IEEE Big Data, 2019, pp. 4179–4184.

[35] A. F. R. Nogueira et al., “Sound classification and processing of urban environments: A systematic literature review,” Sensors, vol. 22, pp. 8608, 2022.

[36] T. Domhan, J. T. Springenberg, and F. Hutter, “Speeding up automatic hyperparameter optimization of deep neural networks,” in Proc. IJCAI, 2015.

[37] J. Mellor et al., “Neural architecture search without training,” in Proc. ICML, vol. 139, pp. 7588–7598, 2021.

[38] B. Na et al., “Accelerating neural architecture search via proxy data,” arXiv preprint arXiv:2106.04784, 2021.

[39] B. Lyu et al., “Resource-constrained neural architecture search on edge devices,” IEEE Transactions on Network Science and Engineering, vol. 9, pp. 134–142, 2021.

[40] H. Benmeziane et al., “A comprehensive survey on hardware-aware neural architecture search,” arXiv preprint arXiv:2101.09336, 2021.

[41] C. Li et al., “HW-NAS-Bench: Hardware-aware neural architecture search benchmark,” arXiv preprint arXiv:2103.10584, 2021.

[42] A. Naeem et al., “Edge-based environmental sound recognition using deep learning,” IEEE Access, 2022.

[43] S. Hershey et al., “CNN architectures for large-scale audio classification,” in Proc. ICASSP, 2017.

[44] T. Heittola, A. Mesaros, and T. Virtanen, “Environmental sound event detection,” IEEE Signal Processing Magazine, vol. 35, pp. 41–49, 2018.

[45] K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in Proc. IEEE MLSP, 2015.

[46] Salman, A. M., & Salman, H. S. (2025). Exploring the Impact of AI and IoT on Production Efficiency, Quality Precision, and Environmental Sustainability in Manufacturing. Vokasi Unesa Bulletin of Engineering, Technology and Applied Science, 2(2), 345-358.

[47] Rifki Ainul Yaqin, Anshori, M. I., Angel, R., Agung, I. W. P., Arifin, T., & Junianto, E. (2026). Stock Price Forecasting Using LSTM with Cross-Validation. Vokasi Unesa Bulletin of Engineering, Technology and Applied Science, 3(1), 64–79. https://doi.org/10.26740/vubeta.v3i1.45130

[48] Oise, G., Nwabuokei, O. C., Ozobialu, chukwuma E., Jenarhome, O. P., Atake, O. . M., Nkem Belinda, U., & Babalola Eyitemi , A. (2025). Enhancing Indoor Positioning Accuracy with Ant Colony Optimization and Dual Clustering. Vokasi Unesa Bulletin of Engineering, Technology and Applied Science, 2(3), 516–530. https://doi.org/10.26740/vubeta.v2i3.39452

[49] Andrianarisoa, S. H., Ravelonjara, H. M., Suddul, G., Foogooa, R., Armoogum, S., & Sookarah, D. (2025). A Deep Learning Approach to Fake News Classification Using LSTM. Vokasi Unesa Bulletin of Engineering, Technology and Applied Science, 2(3), 593–601. https://doi.org/10.26740/vubeta.v2i3.39360

[50] Oise, G., Nwabuokei, C., IGBUNU, R., & EJENARHOME, P. (2025). Revisiting Parasitic Computing: Ethical and Technical Dimensions in Resource Optimization. Vokasi Unesa Bulletin of Engineering, Technology and Applied Science, 2(3), 376–386. https://doi.org/10.26740/vubeta.v2i3.38786

Diterbitkan

2026-08-06

Cara Mengutip

[1]
K. Muhammed Nasiru, S. M. Yusuf, dan S. Saleh Saminu, “The ShuBAM-TFNet: A Lightweight Temporal–Frequency Attention Network for Environmental Sound Recognition”, Vokasi UNESA Bull. Eng. Technol. Appl. Sci., vol. 3, no. 3, Agu 2026.

Terbitan

Bagian

Article
Abstract views: 114

Artikel paling banyak dibaca berdasarkan penulis yang sama

Artikel Serupa

1 2 3 4 5 6 > >> 

Anda juga bisa Mulai pencarian similarity tingkat lanjut untuk artikel ini.