The ShuBAM-TFNet: A Lightweight Temporal–Frequency Attention Network for Environmental Sound Recognition

Authors

  • Shehu Mohammed Yusuf Ahmadu Bello University

DOI:

https://doi.org/10.26740/vubeta.v3i3.50353

Keywords:

EnvironmentalSound Classification Lightweight Deep Learning Temporal Frequency Attention Edge Computing Audio Signal Processing

Abstract

Efficient environmental sound recognition is essential for smart-city monitoring and edge-based acoustic sensing, yet practical deployment is often constrained by the computational cost and parameter size of deep learning models. This paper presents ShuBAM-TFNet, a lightweight convolutional neural network designed to achieve high recognition accuracy while remaining suitable for resource-limited devices. The proposed architecture processes Log-Mel spectrograms using dilated and time-distributed convolutional layers to capture long-range temporal dependencies and localized temporal patterns without increasing model complexity. A lightweight temporal–frequency attention mechanism is introduced to adaptively emphasize informative regions across time and frequency dimensions, while Bottleneck Attention Modules are incorporated to refine channel and spatial representations. The model is evaluated on the UrbanSound8K and ESC-10 benchmark datasets using standard evaluation protocols. On UrbanSound8K, the full ShuBAM-TFNet achieves an average accuracy of 90.16 ± 1.05% with 72,390 trainable parameters, converging in approximately 27 epochs. Ablation studies show that removing the temporal–frequency attention module results in a 2–3% accuracy drop, confirming its central contribution, while eliminating dilated convolutions leads to slower convergence despite stable accuracy. Notably, removing Bottleneck Attention Modules yields the best-performing variant, achieving 91.76 ± 0.90% accuracy with a reduced parameter count of 70,754. Cross-dataset evaluation on ESC-10 using 5-fold cross-validation achieves a mean accuracy of 84.00 ± 2.41% and a macro F1-score of 0.838, demonstrating good generalization. These results indicate that attention-guided lightweight architectures provide a practical alternative to search-based or heavily parameterized models for real-time environmental sound recognition on edge devices.

References

REFERENCES

[1]S. Tyagi, K. Aggarwal, D. Kumar, and S. Garg, “Urban sound classification for audio analysis using long short-term memory,” NEU Journal for Artificial Intelligence and Internet of Things, vol. 1, no. 2, pp. 1–11, 2023.

[2]W. Zhao, H. Wang, Y. Chen, X. Pan, K. Zhang, and Z. Bai, “An environmental sound classification algorithm based on multiscale channel feature fusion,” IEEE Access, 2023.

[3]M. Avadhani and A. P. Bidargaddi, “Multi-class urban sound classification with deep learning architectures,” in Proc. IEEE Conference, 2024, pp. 1–7.

[4]H. Lu, H. Zhang, and A. Nayak, “A deep neural network for audio classification with a classifier attention mechanism,” arXiv preprint arXiv:2006.09815, 2020.

[5]K. Zaman, M. Sah, C. Direkoglu, and M. Unoki, “A survey of audio classification using deep learning,” IEEE Access, vol. 11, pp. 106620–106649, 2023.

[6]J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proc. ACM Multimedia, 2014, pp. 1041–1044.

[7]K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proc. ACM Multimedia, 2015, pp. 1015–1018.

[8]R. Mahum, A. Irtaza, A. Javed, H. A. Mahmoud, and H. Hassan, “DeepDet: YAMNet with bottleneck attention module for TTS synthesis detection,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, pp. 1–18, 2024.

[9]W. Mu, B. Yin, X. Huang, J. Xu, and Z. Du, “Environmental sound classification using temporal–frequency attention based convolutional neural network,” Scientific Reports, vol. 11, no. 1, pp. 21552, 2021.

[10]F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122, 2015.

[11]S. Abdoli, P. Cardinal, and A. L. Koerich, “End-to-end environmental sound classification using a 1D convolutional neural network,” Expert Systems with Applications, vol. 136, pp. 252–263, 2019.

[12]J. Park, S. Woo, J.-Y. Lee, and I. S. Kweon, “BAM: Bottleneck attention module,” arXiv preprint arXiv:1807.06514, 2018.

[13]X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in Proc. IEEE CVPR, 2018, pp. 6848–6856.

[14]D. Ranmal, P. Ranasinghe, T. Paranayapa, D. Meedeniya, and C. Perera, “ESC-NAS: Environment sound classification using hardware-aware neural architecture search for the edge,” Sensors, vol. 24, no. 12, pp. 3749, 2024.

[15]A. Kaçar, A. Derya, and İ. Türkoğlu, “Enhancing environmental sound classification performance through data fusion: A comparative machine learning analysis,” 2025.

[16]L. Liu, “Speech recognition method for English translators in noisy environments based on attention mechanism,” International Journal of High Speed Electronics and Systems, vol. 34, no. 4, pp. 2540350, 2025.

[17]S. Divya Lakshmi and N. Suresh Kumar, “ResBiA-FusionNET: A robust deep learning framework with harmonic-contrast mel spectrogram for audio sound classification,” International Journal of Pattern Recognition and Artificial Intelligence, vol. 39, no. 15, pp. 2551028, 2025.

[18]I. H. Sarker, “Deep learning: A comprehensive overview on techniques, taxonomy, applications and research directions,” SN Computer Science, vol. 2, pp. 420, 2021.

[19]M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,” ACM Computing Surveys, vol. 48, pp. 1–46, 2016.

[20]D. Meedeniya, I. Ariyarathne, M. Bandara, R. Jayasundara, and C. Perera, “A survey on deep learning based forest environment sound classification at the edge,” ACM Computing Surveys, vol. 56, pp. 66, 2023.

[21]D. Stefani, S. Peroni, and L. Turchet, “A comparison of deep learning inference engines for embedded real-time audio classification,” in Proc. DAFx, 2022, pp. 256–263.

[22]A. Elhanashi, P. Dini, S. Saponara, and Q. Zheng, “Integration of deep learning into the IoT: A survey of techniques and challenges for real-world applications,” Electronics, vol. 12, pp. 4925, 2023.

[23]D. Meedeniya, Deep Learning: A Beginners’ Guide, CRC Press, Boca Raton, FL, USA, 2023.

[24]B. Wu et al., “FBNet: Hardware-aware efficient convnet design via differentiable neural architecture search,” in Proc. IEEE CVPR, 2019, pp. 10734–10742.

[25]C. White et al., “Neural architecture search: Insights from 1000 papers,” arXiv preprint arXiv:2301.08727, 2023.

[26]M. Risso et al., “Lightweight neural architecture search for temporal convolutional networks at the edge,” IEEE Transactions on Computers, vol. 72, pp. 744–758, 2023.

[27]J. Lin et al., “MCUNet: Tiny deep learning on IoT devices,” Advances in Neural Information Processing Systems, vol. 33, pp. 11711–11722, 2020.

[28]D. T. Speckhard et al., “Neural architecture search for energy-efficient always-on audio machine learning,” Neural Computing and Applications, vol. 35, pp. 12133–12144, 2023.

[29]B. Elliott et al., “Cyber-physical analytics: Environmental sound classification at the edge,” in Proc. IEEE WF-IoT, 2020, pp. 1–6.

[30]A. Bansal and N. K. Garg, “Environmental sound classification: A descriptive review of the literature,” Intelligent Systems with Applications, vol. 16, pp. 200115, 2022.

[31]A. Andreadis, G. Giambene, and R. Zambon, “Monitoring illegal tree cutting through ultra-low-power smart IoT devices,” Sensors, vol. 21, pp. 7593, 2021.

[32]I. Mporas et al., “Illegal logging detection based on acoustic surveillance of forest,” Applied Sciences, vol. 10, pp. 7379, 2020.

[33]G. Peruzzi, A. Pozzebon, and M. Van Der Meer, “Detecting forest fires with embedded machine learning models using audio and images,” Sensors, vol. 23, pp. 783, 2023.

[34]S. K. Shah, Z. Tariq, and Y. Lee, “IoT-based urban noise monitoring using deep learning,” in Proc. IEEE Big Data, 2019, pp. 4179–4184.

[35]A. F. R. Nogueira et al., “Sound classification and processing of urban environments: A systematic literature review,” Sensors, vol. 22, pp. 8608, 2022.

[36]T. Domhan, J. T. Springenberg, and F. Hutter, “Speeding up automatic hyperparameter optimization of deep neural networks,” in Proc. IJCAI, 2015.

[37]J. Mellor et al., “Neural architecture search without training,” in Proc. ICML, vol. 139, pp. 7588–7598, 2021.

[38]B. Na et al., “Accelerating neural architecture search via proxy data,” arXiv preprint arXiv:2106.04784, 2021.

[39]B. Lyu et al., “Resource-constrained neural architecture search on edge devices,” IEEE Transactions on Network Science and Engineering, vol. 9, pp. 134–142, 2021.

[40]H. Benmeziane et al., “A comprehensive survey on hardware-aware neural architecture search,” arXiv preprint arXiv:2101.09336, 2021.

[41]C. Li et al., “HW-NAS-Bench: Hardware-aware neural architecture search benchmark,” arXiv preprint arXiv:2103.10584, 2021.

[42]A. Naeem et al., “Edge-based environmental sound recognition using deep learning,” IEEE Access, 2022.

[43]S. Hershey et al., “CNN architectures for large-scale audio classification,” in Proc. ICASSP, 2017.

[44]T. Heittola, A. Mesaros, and T. Virtanen, “Environmental sound event detection,” IEEE Signal Processing Magazine, vol. 35, pp. 41–49, 2018.

[45]K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in Proc. IEEE MLSP, 2015.

[46]Salman, A. M., & Salman, H. S. (2025). Exploring the Impact of AI and IoT on Production Efficiency, Quality Precision, and Environmental Sustainability in Manufacturing. Vokasi Unesa Bulletin of Engineering, Technology and Applied Science, 2(2), 345-358.

Published

2026-08-06

How to Cite

[1]
S. M. Yusuf, “The ShuBAM-TFNet: A Lightweight Temporal–Frequency Attention Network for Environmental Sound Recognition”, Vokasi UNESA Bull. Eng. Technol. Appl. Sci., vol. 3, no. 3, Aug. 2026.

Issue

Section

Engineering
Abstract views: 0