Enhancing machine learning models with audio augmentation techniques
Abstract
This paper investigates how audio augmentation techniques improve the classification accuracy of guitar chords. By applying methods such as noise addition, speed modification, and time shifting, a CNN trained on an augmented dataset achieved better results compared to non-augmented data. The findings highlight the importance of task-specific augmentation techniques in enhancing audio analysis
References
Lucas Ferreira-Paiva, Elizabeth Alfaro-Espinoza, Vinicius M Almeida, Leonardo B Felix, and Rodolpho VA Neves. A survey of data augmentation for audio classification. 2021.
Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, Quoc V. Le (2019). "SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition." Proceedings of the Annual Conference of the International Speech Communication Association.
Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht. Text data augmentation for deep learning. Journal of big Data, 8:1–34, 2021.
Jinyu Li. Recent advances in end-to-end automatic speech recognition. APSIPA Transactions on Signal and Information Processing, 11(1), 2022.
Dan Oneata and Horia Cucu. Improving multimodal speech recognition by data augmentation and speech representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4579–4588, 2022.

