Loss of ability to form audible speech because of medical conditions, including amyotrophic lateral sclerosis (ALS), laryngectomy, or nervous disorders to a great extent affect human interactions and life quality. The classic assistive communication models are premised on the help of either a visual or manual input, and hence are slow and unnatural in communication. In this paper, the authors introduce an innovative EMG-based silent speech recognition system that allows communicating because the neuromuscular signals, produced when a person articulates words, are detected, but no sound is produced. It is suggested that the proposed system is based on the surface electromyography (sEMG) sensors which measure the electrical activity of the muscles of the face and throat. One of these signals is amplified and filtered by a Bio-Amp EXG Pill sensor and read by an Arduino microcontroller. The processed data are then fed into MATLAB in order to do real- time signal processing, feature extraction and classification. Some of the important features that are extracted and are used to train machine learning models include root mean square (RMS), Zero Crossing Rate (ZCR), frequency-domain characteristics. The Cubic Support Vector Machine (SVM) classifier is used to identify set spectrum silent speech commands. It has been experimentally shown to have a classification accuracy of 94.2% and a response latency of under 250 ms, so it can be used in the context of real-time. The system is a non-invasive, low-cost and portable assistive communication and silent human-computer interaction system.
Introduction
This paper presents a low-cost, real-time silent speech recognition system that enables people with speech impairments, such as those caused by amyotrophic lateral sclerosis (ALS), laryngectomy, or neurological disorders, to communicate without producing audible speech. The system uses surface electromyography (sEMG) to capture electrical signals generated by facial and throat muscles during silent speech and applies machine learning to convert these signals into text commands.
Background
Traditional assistive communication methods, including text input and eye-tracking systems, are often slow, require continuous user effort, and do not provide a natural communication experience. Voice-based systems are unsuitable for people with speech disorders and in noisy or privacy-sensitive environments. Recent advances in biosignal processing and machine learning have made EMG-based silent speech interfaces (SSIs) a promising alternative because they are non-invasive and directly measure muscle activity related to speech.
Literature Review
Previous research has demonstrated that EMG can be used for silent speech recognition through techniques such as:
Pattern recognition
Hidden Markov Models (HMM)
Support Vector Machines (SVM)
Convolutional Neural Networks (CNN)
CNN-LSTM hybrid models
While these methods have improved recognition accuracy, most existing systems require multiple sensors, expensive hardware, and significant computational resources, limiting their portability and real-world use.
Proposed Methodology
The proposed system emphasizes low cost, portability, and real-time performance. It consists of three main modules:
Signal Acquisition: Surface EMG electrodes capture muscle activity from the throat and jaw. A Bio-Amp EXG sensor amplifies and filters the signals, while an Arduino digitizes and transmits them to MATLAB.
Signal Processing and Feature Extraction: Noise is removed, signals are segmented, and features such as Root Mean Square (RMS), Zero Crossing Rate (ZCR), and frequency-domain features are extracted.
Machine Learning Classification: A Cubic Support Vector Machine (SVM) is trained to recognize predefined silent commands (e.g., up, down, left, right, neutral). A majority-voting mechanism improves prediction stability in real time.
System Workflow
The system follows these steps:
Capture EMG signals.
Amplify and filter the signals.
Convert analog signals to digital.
Transfer data to MATLAB.
Preprocess and segment the signals.
Extract features.
Classify using the trained SVM.
Predict the command.
Apply decision logic.
Display the recognized text.
Results
The proposed silent speech recognition system achieved:
94.2% overall classification accuracy
Precision and recall above 93% for all command classes
Stable real-time performance with only minor confusion between similar commands such as "up" and "down"
Conclusion
This paper has presented an EMG based silent speech recognition system, which enables us to talk without uttering a word. It is an architecture, which integrates signal acquisition of surface electromyography signals and preprocessing, feature extraction, and machine learning classification to detect predetermined commands in real-time.
The results of the experiment indicate high accuracy of the proposed system of 94.2 per cent in classification using Cubic Support Vector Machine (SVM) classifier. The system also makes optimum use of the characteristics of EMG signals, like the Root Mean Square (RMS) and Zero Crossing Rate (ZCR) features to record the amplitude and frequency properties of EMG signals. The real-life assistive communication implementation Here, the real-time implementation has a response time value of 180250 ms, which is acceptable in the real-life applications of assistive communication. The system offers an economical, non-invasive, and efficient solution to the speech impaired. The fact that hardware and software are compatible guarantees plausible functionality and it demonstrates that EMG-based silent speech interfaces can work in human-computer interface.
References
[1] B. Denby, T. Schultz, K. Honda, T. Hueber, J. Gilbert, and J. S. Brumberg, “Silent speech interfaces,” Speech Communication, vol. 52, no. 4, pp. 270–287, 2010.
[2] A. Chan, K. Englehart, B. Hudgins, and D. F. Lovely, “Continuous speech recognition using surface electromyography,” Journal of the Acoustical Society of America, vol. 120, no. 5, pp. 2847–2857, 2006.
[3] M. Wand and T. Schultz, “Pattern recognition techniques for EMG-based speech recognition,” IEEE Transactions on Biomedical Engineering, vol. 58, no. 4, pp. 1224–1233, 2011.
[4] M. Janke and T. Schultz, “Surface electromyography for silent speech recognition,” Computer Speech & Language, vol. 52, pp. 98–116, 2018.
[5] K. Kumar, R. Kumar, and S. Pandey, “EMG-based human-computer interface for silent communication,” IEEE Access, vol. 8, pp. 195051–195060, 2020.
[6] T. Schultz and M. Wand, “Modeling coarticulation in EMG- based continuous speech recognition,” Speech Communication, vol. 52, no. 4, pp. 341–353, 2010.
[7] J. S. Brumberg, A. Nieto-Castanon, P. R. Kennedy, and F. H. Guenther, “Brain–computer interfaces for speech communication,” Speech Communication, vol. 52, no. 4, pp. 367–379, 2010.
[8] H. Manabe and Z. Zhang, “Multichannel electromyography- based speech recognition using deep learning,” IEEE Sensors Journal, vol. 19, no. 15, pp. 6454–6463, 2019.
[9] A. Kapur, N. Kapur, and P. Maes, “Alter Ego: A personalized wearable silent speech interface,” Proceedings of the ACM International Conference on Intelligent User Interfaces, 2018.
[10] B. Denby, T. Schultz, K. Honda, T. Hueber, J. Gilbert, and J. S. Brumberg, “Silent speech interfaces,” Speech Communication, vol. 52, no. 4, pp. 270–287, 2010.
[11] J. Jou, T. Schultz, and A. Waibel, “Continuous EMG-based speech recognition using hidden Markov models,” Proceedings of Inter-speech, 2006.
[12] S. G. Wang, C. H. Yang, and J. F. Chen, “EMG-based speech recognition with support vector machines,” Expert Systems with Applications, vol. 37, no. 6, pp. 4408–4414, 2010.
[13] Y. K. Lee and J. Kim, “Real-time EMG-based speech recognition system using machine learning,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 28, no. 2, pp. 456–465, 2020.
[14] P. Herff, M. Wand, and T. Schultz, “Continuous speech recognition from surface EMG,” Proceedings of Inter-speech, 2013.
[15] B. Huang, H. Zhang, P. Wang, and X. Chen, “Silent speech recognition using sEMG signals and deep learning,” IEEE Access, vol. 11, pp. 105000–105012, 2023.
[16] J. Freitas, A. Teixeira, M. Sales Dias, and S. Silva, “An introduction to silent speech interfaces,” Springer, 2016.
[17] M. Wand and T. Schultz, “Session-independent EMG-based speech recognition,” Proceedings of BIOSIGNALS, 2012.
[18] T. Schultz and M. Wand, “Electromyography-based speech recognition,” Speech Communication, vol. 52, no. 4, pp. 341–353, 2010.
[19] S. R. Schultz, T. Hueber, and B. Denby, “Silent speech interfaces: From signal processing to machine learning,” IEEE Signal Processing Magazine, vol. 35, no. 5, pp. 18–31, 2018.
[20] J. S. Brumberg, “Silent speech interfaces for speech restoration,” Trends in Biotechnology, vol. 38, no. 3, pp. 273–287, 2020.