TY - GEN
T1 - Echoes of Deception
T2 - 8th International Conference on Information Science and Systems, ICISS 2025
AU - Bhatnagar, Parth
AU - Sodhi, Manjit
AU - Gururaj, H. L.
AU - Kaur, Sanjeet
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2026/6/4
Y1 - 2026/6/4
N2 - The rapid advancements in deep learning have enabled the creation of highly realistic synthetic speech, commonly known as deep fake voices. While these advancements offer numerous benefits, such as enhanced virtual assistants and improved accessibility tools, they also pose serious security threats, including identity fraud, misinformation, and unauthorized access to voice-based authentication systems. To mitigate these risks, this study proposes a deep learning-based approach for deep fake voice recognition.The methodology involves collecting a diverse dataset of real and synthetic speech samples, followed by preprocessing techniques such as noise reduction, resampling, and data augmentation to enhance model robustness. Key audio features, including Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, chroma features, and zero-crossing rate (ZCR), are extracted and transformed into spectrogram representations for analysis. A Convolutional Neural Network (CNN) architecture is employed to classify speech samples, leveraging convolutional layers for feature extraction and fully connected layers for classification. The model is trained using the Adam optimizer with a categorical cross-entropy loss function, and an early stopping mechanism is implemented to prevent overfitting. The experimental results demonstrate that the CNN-based model achieves an accuracy of 94.8%, outperforming traditional classifiers such as Support Vector Machines (SVMs) and Random Forests. The model also exhibits a high F1-score of 94.2%, ensuring a balanced trade-off between precision and recall. Spectral analysis reveals that deep fake speech exhibits distinct patterns, such as smoother frequency transitions, which the CNN model effectively identifies. Data augmentation techniques further enhance the model's generalization capabilities, improving its performance on unseen data.Despite these promising results, challenges such as evolving deep fake synthesis techniques and language variability remain. Future work aims to expand the dataset across multiple languages and refine detection models to counter increasingly sophisticated deep fake speech. This research contributes to the growing field of deep fake detection, offering a scalable and efficient approach to safeguarding voice-based security systems.
AB - The rapid advancements in deep learning have enabled the creation of highly realistic synthetic speech, commonly known as deep fake voices. While these advancements offer numerous benefits, such as enhanced virtual assistants and improved accessibility tools, they also pose serious security threats, including identity fraud, misinformation, and unauthorized access to voice-based authentication systems. To mitigate these risks, this study proposes a deep learning-based approach for deep fake voice recognition.The methodology involves collecting a diverse dataset of real and synthetic speech samples, followed by preprocessing techniques such as noise reduction, resampling, and data augmentation to enhance model robustness. Key audio features, including Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, chroma features, and zero-crossing rate (ZCR), are extracted and transformed into spectrogram representations for analysis. A Convolutional Neural Network (CNN) architecture is employed to classify speech samples, leveraging convolutional layers for feature extraction and fully connected layers for classification. The model is trained using the Adam optimizer with a categorical cross-entropy loss function, and an early stopping mechanism is implemented to prevent overfitting. The experimental results demonstrate that the CNN-based model achieves an accuracy of 94.8%, outperforming traditional classifiers such as Support Vector Machines (SVMs) and Random Forests. The model also exhibits a high F1-score of 94.2%, ensuring a balanced trade-off between precision and recall. Spectral analysis reveals that deep fake speech exhibits distinct patterns, such as smoother frequency transitions, which the CNN model effectively identifies. Data augmentation techniques further enhance the model's generalization capabilities, improving its performance on unseen data.Despite these promising results, challenges such as evolving deep fake synthesis techniques and language variability remain. Future work aims to expand the dataset across multiple languages and refine detection models to counter increasingly sophisticated deep fake speech. This research contributes to the growing field of deep fake detection, offering a scalable and efficient approach to safeguarding voice-based security systems.
UR - https://www.scopus.com/pages/publications/105042306367
UR - https://www.scopus.com/pages/publications/105042306367#tab=citedBy
U2 - 10.1145/3803722.3803726
DO - 10.1145/3803722.3803726
M3 - Conference contribution
AN - SCOPUS:105042306367
T3 - ICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems
SP - 22
EP - 28
BT - ICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems
PB - Association for Computing Machinery, Inc
Y2 - 14 August 2025 through 16 August 2025
ER -