Skip to main navigation Skip to search Skip to main content

Echoes of Deception: Deep Fake Voice Recognition Using Machine Learning

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The rapid advancements in deep learning have enabled the creation of highly realistic synthetic speech, commonly known as deep fake voices. While these advancements offer numerous benefits, such as enhanced virtual assistants and improved accessibility tools, they also pose serious security threats, including identity fraud, misinformation, and unauthorized access to voice-based authentication systems. To mitigate these risks, this study proposes a deep learning-based approach for deep fake voice recognition.The methodology involves collecting a diverse dataset of real and synthetic speech samples, followed by preprocessing techniques such as noise reduction, resampling, and data augmentation to enhance model robustness. Key audio features, including Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, chroma features, and zero-crossing rate (ZCR), are extracted and transformed into spectrogram representations for analysis. A Convolutional Neural Network (CNN) architecture is employed to classify speech samples, leveraging convolutional layers for feature extraction and fully connected layers for classification. The model is trained using the Adam optimizer with a categorical cross-entropy loss function, and an early stopping mechanism is implemented to prevent overfitting. The experimental results demonstrate that the CNN-based model achieves an accuracy of 94.8%, outperforming traditional classifiers such as Support Vector Machines (SVMs) and Random Forests. The model also exhibits a high F1-score of 94.2%, ensuring a balanced trade-off between precision and recall. Spectral analysis reveals that deep fake speech exhibits distinct patterns, such as smoother frequency transitions, which the CNN model effectively identifies. Data augmentation techniques further enhance the model's generalization capabilities, improving its performance on unseen data.Despite these promising results, challenges such as evolving deep fake synthesis techniques and language variability remain. Future work aims to expand the dataset across multiple languages and refine detection models to counter increasingly sophisticated deep fake speech. This research contributes to the growing field of deep fake detection, offering a scalable and efficient approach to safeguarding voice-based security systems.

Original languageEnglish
Title of host publicationICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems
PublisherAssociation for Computing Machinery, Inc
Pages22-28
Number of pages7
ISBN (Electronic)9798400713200
DOIs
Publication statusPublished - 04-06-2026
Event8th International Conference on Information Science and Systems, ICISS 2025 - Oxford, United Kingdom
Duration: 14-08-202516-08-2025

Publication series

NameICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems

Conference

Conference8th International Conference on Information Science and Systems, ICISS 2025
Country/TerritoryUnited Kingdom
CityOxford
Period14-08-2516-08-25

All Science Journal Classification (ASJC) codes

  • Computer Networks and Communications
  • Computer Science Applications
  • Information Systems
  • Signal Processing
  • Control and Systems Engineering

Fingerprint

Dive into the research topics of 'Echoes of Deception: Deep Fake Voice Recognition Using Machine Learning'. Together they form a unique fingerprint.

Cite this