A Comparative Machine Learning Framework for Speaker Identification Using Audio Biometric Fea-tures
Authors-B.Pradeep, Vijaya Mallamari
Keyword-Speaker Identification, Audio Biometrics, Machine Learning, Deep Learning, Support Vector Machine, Convolutional Neural Network, Long Short-Term Memory, Mel Spectrogram, Speech Recognition, VoxCeleb Dataset.
Speaker identification has become an essential component of modern biometric authentication systems because voice characteristics provide a convenient and non-invasive method for verify-ing human identity. The rapid development automatic speaker identification. Audio recordings from the VoxCeleb dataset are preprocessed to eliminate unwanted noise and silence before extracting Mel Spectrogram features that effectively represent speech characteristics. The extract-ed features are used to train and evaluate each classification model using standard performance measures such as accuracy and F1-score. Experimental observations indicate that the SVM classifier delivers the highest recognition performance among the evaluated models, while CNN also achieves competitive results. In contrast, the LSTM model records comparatively lower performance because of the sequential complexity of the available dataset. The study demon-strates the importance of selecting an appropriate learning algorithm for speaker recognition applications and provides valuable insights for developing reliable and efficient voice-based biometric systems.
Doi-[https://doi.org/10.5281/zenodo.21678804]