Advancements in Humanoid Robotics: Designing an Artificial Neural Network-based Speech Recognition Robot for Tactical Deployment

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract The development of an ANN-based speech recognition system using MATLAB for controlling a humanoid robot prototype is presented in this research. Our proposed approach, therefore, will be in a good position to derive significant benefits from the powerful features and toolboxes in MATLAB in the areas of signal processing, machine learning, and robotics, way above most of the works available that have used platforms such as Arduino and Android with Bluetooth technology. The system is trained to recognize the five speech commands: "move forward," "move backward," "turn right," "turn left," and "stop." The custom GUI software has been developed to collect the data. It is processed by Fast Fourier Transform (FFT) and Mel-frequency cepstral coefficients (MFCC). After that, it goes through silence removal and normalized techniques with pre-processing ANN classifiers and integrated with the robot control system. The accuracy of the trained model on test data with multiple speakers is 87% and holds good for general purposes without being biased towards any specific speaker's voice characteristics. Developed prototype, hence, proves the feasibility and potential of MATLAB for the speech recognition task in control with humanoid robot and its applications in different domains like industries, healthcare, and defense. The modular architecture allows for easy customization and extension to incorporate additional voice commands and functionalities. Future research directions include improving robustness to noise, speaker-independent recognition, and integration with other modalities like gesture recognition and computer vision.
Full text 95,343 characters · extracted from preprint-html · click to expand
Advancements in Humanoid Robotics: Designing an Artificial Neural Network-based Speech Recognition Robot for Tactical Deployment | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Advancements in Humanoid Robotics: Designing an Artificial Neural Network-based Speech Recognition Robot for Tactical Deployment Sarika Shrivastava, Somendra Bannerjee, Mayank Srivastava, Saifullah Khalid, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4762012/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The development of an ANN-based speech recognition system using MATLAB for controlling a humanoid robot prototype is presented in this research. Our proposed approach, therefore, will be in a good position to derive significant benefits from the powerful features and toolboxes in MATLAB in the areas of signal processing, machine learning, and robotics, way above most of the works available that have used platforms such as Arduino and Android with Bluetooth technology. The system is trained to recognize the five speech commands: "move forward," "move backward," "turn right," "turn left," and "stop." The custom GUI software has been developed to collect the data. It is processed by Fast Fourier Transform (FFT) and Mel-frequency cepstral coefficients (MFCC). After that, it goes through silence removal and normalized techniques with pre-processing ANN classifiers and integrated with the robot control system. The accuracy of the trained model on test data with multiple speakers is 87% and holds good for general purposes without being biased towards any specific speaker's voice characteristics. Developed prototype, hence, proves the feasibility and potential of MATLAB for the speech recognition task in control with humanoid robot and its applications in different domains like industries, healthcare, and defense. The modular architecture allows for easy customization and extension to incorporate additional voice commands and functionalities. Future research directions include improving robustness to noise, speaker-independent recognition, and integration with other modalities like gesture recognition and computer vision. Physical sciences/Energy science and technology Physical sciences/Engineering Humanoid robotics Speech recognition Artificial neural networks MATLAB Tactical deployment Human-robot interaction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 1. Introduction Humanoid robotics has made significant strides in recent years, with the goal of creating machines that can replicate human actions and operate autonomously. A key aspect of human-robot interaction is the ability for robots to understand and respond to voice commands. While speech recognition has a long history of innovation, recent advances in artificial intelligence (AI) and deep learning have greatly improved its accuracy and potential applications [ 1 ]. This paper aims to address the gap in the literature regarding the use of MATLAB for developing a speech recognition system to control humanoid robots. Previous research has primarily focused on using third-party servers like Arduino and Android with Bluetooth technology for speech recognition in robotics [ 8 ][ 10 ][ 11 ]. On the other hand, there are quite a number of benefits that come with MATLAB, which ranges from creation of controllers based on neural networks, removal of silence, and the generation of confusion plots used in the visualization of the speech recognition model's performance. This research is going to try to design and develop a Speech Recognition system that is ANN based through MATLAB for controlling a humanoid robot in tactical deployment. The system shall be designed to be able to understand and follow voice commands. These operate on the basis that the robot undertakes task execution in a real-time environment, especially under defense or combat situations, whereby human intervention is at most risk, in most cases. To undertake this objective, we will employ the following methodology: Data collection: All voice samples are to be recorded using an in-house GUI on MATLAB. Speech processing: Extract the unique characteristics of the voice samples using techniques such as Fast Fourier Transform (FFT) and Mel-frequency cepstrum (MFCC). Pre-processing: Use silence removal and extract features for bettering their overall accuracy during recognition by the model. ANN training: Train the processed voice samples with the neural network to recognize the set commands. Robot control: The developed ANN model shall be embedded in the humanoid robot for the purpose of executing voice commands. This work has its importance in giving advancement in the field of humanoid robotics through capabilities present in MATLAB for better development in speech recognition. The proposed ANN-based system can be used as per the requirements in different applications like industries, medical fields, and defense. Focusing on tactical deployment, this research aims to demonstrate the feasibility and benefits of using voice-controlled humanoid robots in high-risk and challenging environments. Figure 1 shows the Block diagram illustrating the critical steps in the proposed methodology for developing an ANN-based speech recognition system for humanoid robot control. The rest of the paper is organized as follows: Section 2 presents a literature review of related work in speech recognition and humanoid robotics. Section 3 describes the problem formulation and the research questions addressed in this study. Section 4 details the research methodology, including data collection, speech processing, pre-processing, ANN training, and robot control. Section 5 presents the results and discussion, followed by the future directions and conclusions in Section 6 and 7. 2. Literature Review Humanoid robotics and speech recognition have witnessed significant advancements in recent years, driven by the increasing integration of artificial intelligence (AI) and machine learning techniques. Researchers have explored various approaches to enable effective human-robot interaction through voice commands, aiming to create robots that can understand and respond to human speech in real-time environments. Authors had displayed that speech recognition in controlling a humanoid robot highly requires a certain degree of word recognition accuracy for gesture selection [ 1 ]. Thus, his work demonstrates the robot's capability to recognize words and perform gestures upon commanding. Similarly, Authors used speech recognition to control a surgical robot with high precision of the angular displacement of the endoscope. These studies highlight the potential power of speech technology to further develop humanoid robots' capability in many different areas [ 2 ]. With the growth of complexity in the nature of speech signals, different types of feature extraction and classification techniques were studied by the researchers. The authors developed a mobile robot for the cocktail party effect. The intelligence of human beings makes it possible to recognize speech from different sources concurrently [ 3 ]. When tested for three speakers, the system produced the average loss of rates for successful recognition between 24% and 42% within two meters. Similarly, The Authors developed a personalized assistant robot, designed to help with dressing tasks, such as easy wearing of shoes [ 4 ]. Their system included user tracking, posture recognition, gesture recognition, and speech recognition for a natural human-robot interaction. A number of researchers have been focusing on the importance of effective communication in human-robot interaction. The author discussed the methods of communication required for successful human-robot collaboration using the Wizard of Oz methodology [ 5 ]. A new era of cohabitation and cooperation in collaborative work was introduced to his work between robots and humans. The Author developed an automatic speech recognition system operating in challenging environments—eco-noise and localization of binaural sound sources [ 6 ]. Their research aimed to understand the influence of human physiognomy on automatic speech recognition and sound source localization, demonstrating improved accuracy when incorporating sound source localization information. The application of speech recognition in manufacturing industries has also garnered attention. Authors proposed a human-robot collaboration manufacturing system using deep learning multimodal, which utilizes voice commands, body motion, and hand motion to perform tasks in industrial settings [ 7 ]. The integration of AI-based voice control has been explored for assisting physically challenged individuals. Authors [ 8 ] developed an Android-based smart device that employs Bluetooth technology and object detection using ultrasonic and color sensors to aid disabled people. Several researchers have focused on the development of voice-controlled robots using various technologies. Authors discussed the objectives of AI in understanding human recognition, cost-effective intelligent amplifications, and the development of AI using speech recognition technology to assist physically challenged skilled persons [ 9 ]. Using an Arduino Bluetooth voice control application, the authors demonstrated voice recognition via Bluetooth technology and Android devices [ 10 ]. The Authors proposed a speech recognition system for a voice-controlled robot with real-time obstacle detection, employing Bluetooth communication via smartphones and GSM/Wi-Fi-based internet connectivity [ 11 ]. The application of voice control in home automation has also been explored. Authors focused on implementing voice-controlled home automation, where smart devices integrated with home appliances communicate with cloud servers to execute commands [ 12 ]. Voice-controlled robotic arms have been developed for object recognition and manipulation. The Authors proposed a voice-controlled robotic arm that performs voice recognition, object identification, and object picking using a Raspberry Pi microcontroller and Python scripts [ 13 ]. Voice recognition has also been studied for surveillance purposes. The authors elaborately described the use of voice-controlled ground vehicles for surveillance in an area of inefficient human work, implemented [ 14 ]. The other field of interest was speech query processing, where the authors experimented with the recognition of isolated question words in speech queries of the Malayalam language using artificial neural networks (ANNs), along with discrete wavelet transformation (DWT) [ 15 ]. They exploited the integration of mobile voice recognition and IoT for their home automation and robotic vehicle control. Authors [ 16 ] proposed a mobile voice-activated home automation system that uses Google voice-to-text API and a Tasker application running on a Raspberry Pi server. The authors built a voice-controlled robotic vehicle using an Android application and IoT, employing a ZigBee module for communication and an EasyVR module for voice command training [ 17 ]. Other authors focus on developing voice-controlled robots for assisting the differently-abled in day-to-day chores. The authors developed a voice-controlled robot using an android application developed by the authors, which converts speech to text and sends commands through Bluetooth [ 18 ]. The authors presented the fabrication of a voice-controlled robotic vehicle using the BT Voice Control for Arduino application, designed exclusively for disabled people [ 19 ]. Virtual technology and software platforms have also been utilized for voice-controlled robotics. The authors interfaced with a voice-controlled robot using LabVIEW. The interface processes the speech for automation and control within the LabVIEW environment, and commands are transmitted to the robot through Bluetooth and Arduino [ 20 ]. The system used a laser scanner to build information about unknown environments and the CMU Sphinx toolkit for speech recognition, which gave 95% accuracy in normal and silent environments. The literature review indicated that remarkable innovations are achieved in both speech recognition and humanoid robotics. AI, machine learning, and IoT combined have now made it possible for voice-controlled robots to get applications in the manufacturing, home automation, healthcare, and surveillance domains. In terms of robustness, adaptability, and real-world implementation, ANN has never been able to actualize its assertion in this domain. ANN-based speech recognition for humanoid robot control using MATLAB with proposed development and implementation in tactical deployment situations in defense and combat will fill these gaps. 3. Problem Formulation Recent years have seen a sharp increase in research at the interface of speech recognition and humanoid robotics, opening ways through which natural and intuitive human-robot interaction can be easily realized. Though previous researches have investigated different methodologies using platforms such as Arduino, coupled with Bluetooth technology, there is still a gap in leveraging the power of MATLAB to develop an ANN-based speech recognition system, tailored explicitly for control in humanoid robot tactical deployment scenarios. The purpose of this research was to examine the feasibility and benefits of using MATLAB-based ANN for speech recognition and conversion into a mechanical control of a humanoid robot for the tactical deployment application. In this respect, this research will consider MATLAB and its rich features and toolboxes in signal processing, machine learning, and robotics beyond the level that has not been fully exploited in the context of speech-controlled humanoid robots. First, MATLAB enables one to design and test ANN models in a single environment through the use of built-in functions for data preparation, feature extraction from the data, and network training. This makes a very smooth, faster workflow in the development process. Secondly, the MATLAB signal processing toolbox has advanced methods, such as the silence deletion process and extracting Mel-frequency cepstral coefficients (MFCC), which can be very useful to make the speech recognition more robust and accurate. Further, MATLAB also has visualization capabilities, like confusion plots, allowing a clear way to analyze the performance of the ANN model with high depth, which enables iterative improvement and optimization. The two-fold objective of this research is to (1) establish validation for using MATLAB in the design of a robust ANN-based speech recognition system and (2) explore its application with respect to controlling a humanoid robot in the context of tactical deployment. This paper focuses on tactical deployment, meeting the unique challenges and requirements for the use of voice commands to control robots in a dynamic environment with varying levels of possibilities and potential noise. The existing approaches in humanoid robot control based on speech recognition, in general, suffer from lack of robustness to noise, speaker independence, and real-time performance. For this reason, in this work, those limitations are addressed using the signal processing power of MATLAB in conjunction with ANN models. The proposed system, incorporating techniques such as removing silence intervals, extracting MFCC features, and retraining the model, ensures that good noise robustness, speaker independence, and effective real-time implementation are taken care of. The proposed methodology for the present investigation involves the following key steps: Data Collection : Recording a diverse dataset of voice commands using MATLAB's audio acquisition tools. Preprocessing : Applying signal processing techniques, including silence removal and MFCC feature extraction, to enhance the quality and discriminability of the speech data. ANN Model Development : Designing and training an ANN architecture in MATLAB for speech recognition, utilizing techniques like backpropagation and regularization. Model Evaluation : The assessment of ANN model performance after training, using indicators such as accuracy, precision, and recall, supported by visualization tools from MATLAB. Robot Integration : This involves interfacing the speech recognition system with a humanoid robot platform, enabling real-time control based on recognized voice commands. Figure 2 presents a high-level block diagram illustrating the proposed system architecture. The successful development of a MATLAB-based speech recognition system for humanoid robot control has significant potential for impact across various domains. In industries, voice-controlled robots can enhance worker safety and efficiency by allowing hands-free operation in manufacturing and assembly tasks. In healthcare, such systems can assist patients with limited mobility or enable remote surgical procedures. In defense and security applications, voice-controlled humanoid robots can be deployed for reconnaissance, bomb disposal, or search and rescue missions, minimizing human risk. By addressing the research above objectives and leveraging the capabilities of MATLAB, this study aims to contribute to the advancement of speech recognition and humanoid robotics, opening up new possibilities for natural and efficient human-robot interaction in tactical deployment scenarios and beyond. 4. Research Methodology The proposed speech recognition system for humanoid robot control consists of two main components: the MATLAB-based artificial neural network (ANN) for speech recognition and the robot prototype that executes the recognized voice commands. Figure 3 presents a flowchart illustrating the steps in the methodology. 4.1. Data Collection To train the ANN model, a dataset of voice commands was collected using a custom-built graphical user interface (GUI) in MATLAB. The GUI, shown in Fig. 4 , allows users to record voice samples for each command. In this study, five commands were used: "move forward," "move backward," "turn right," "turn left," and "stop." 100 voice samples were recorded for each command, with 70 samples used for training and 30 for testing. The voice commands were recorded from 10 speakers (5 male and five female) to ensure speaker variability in the dataset. 4.2. Speech Processing The recorded voice samples undergo speech processing to extract unique features such as pitch, fundamental frequency, and spectral content. Two main techniques were employed for feature extraction: Fast Fourier Transform (FFT) analysis : The voice sample spectrum is considered whole, and features like spectral content are calculated using FFT. Figure 5 shows the FFT analysis of a single voice sample. Mel-frequency cepstrum (MFCC) : To capture the sound variation more accurately, the voice samples are divided into frames of 20–30 ms. MFCC is then applied to each frame to extract 14 coefficients that represent the short-term power spectrum of the sound. Figure 6 shows the frame matrices formed for a 25 ms window. 4.3. Pre-processing Before training the ANN model, the extracted features undergo pre-processing steps: Silence removal: A custom MATLAB script was developed to remove silent frames from the voice samples. The script divides the sound file into frames, calculates the energy of each frame, and discards frames with energy less than 2% of the total energy. Figure 7 shows the effect of silence removal on a voice sample. Feature normalization : The MFCC features extracted from each frame are normalized to have zero mean and unit variance. This helps improve the convergence speed and stability of the ANN training process. After pre-processing, the final training dataset consists of a (14 * 3771) matrix of MFCC features and a (5 * 3771) matrix of corresponding target labels. 4.4. ANN Architecture and Training A feedforward neural network was designed for speech recognition using MATLAB's Deep Learning Toolbox. The ANN architecture, shown in Fig. 8 , consists of: Input layer: 14 neurons, corresponding to the 14 MFCC features Hidden layer: 50 neurons with ReLU activation function Output layer: 5 neurons with softmax activation, corresponding to the five voice commands The ANN was trained using the scaled conjugate gradient backpropagation algorithm with a learning rate 0.01 and a maximum of 1000 epochs. The training data was divided into 70% for training, 15% for validation, and 15% for testing. Early stopping was used to prevent overfitting, and the validation patience was 10 epochs. 4.5. Robot Control The trained ANN model was integrated with the humanoid robot prototype using MATLAB's Hardware Support Package for Arduino. A custom robot control GUI, shown in Fig. 9 , was developed to allow users to give voice commands and visualize the robot's response. When a voice command is given, the ANN model predicts the corresponding action, which is then executed by the robot's actuators and control modules. Following this systematic methodology, the proposed speech recognition system achieves accurate and responsive voice-based control of the humanoid robot prototype. The modular architecture allows for easy customization and extension to incorporate additional voice commands and functionalities as needed. 5. Results & Discussions The developed ANN-based speech recognition system for humanoid robot control achieved promising results, demonstrating the feasibility and potential of using MATLAB for this application. The trained model reached an overall accuracy of 87% in recognizing the five voice commands: "move forward," "move backward," "turn right," "turn left," and "stop." This accuracy was obtained using a testing dataset of 150 samples (30 per command) recorded from ten speakers (five male and five females) to be able to make a good estimate of performance while taking into account speaker variability. More metrics were computed to give a more rounded measure of the model's performance. The corresponding precision, recall, and F1 scores (with 95% CI) of the system are 0.88, 0.87, and 0.87 (0.82–0.92), respectively. Such results evidently show that it gets to the top as far as accuracy is concerned and that a very good balance between precision and recall is achieved, minimizing both false positives and false negatives. The confidence interval of the accuracy really helps to perceive how significantly reliable and robust the system is performing. The performance of the model across the different voice commands can be visualized in the confusion matrix in Fig. 10 . The confusion matrix, however, indicates that the system was slightly less accurate on the "turn right" and "turn left" commands. This would indicate that the model could struggle with discriminating both these directions, most likely because they sound quite similar to one another or have an almost similar acoustic representation. Further work should refine feature extraction and model architecture for discrimination between such commands. First advantage is that previous research using Arduino and Android platforms could be tested and executed only on a limited hardware setup; therefore, the work's usefulness could not be general for mankind. Development is easy with built-in functions of MATLAB for signal processing, feature extraction (MFCC), and ANN training, making it possible to easily try out different model architectures. This further enabled the results to be visualized with confusion plots and other graphs, for better understanding of the performance and limitations of the system. While the performance of the MATLAB approach is slightly inferior to some of the reported results, it is important to mention that most of the studies often applied fewer speakers or had a smaller database, which may not encompass the whole complexity of speech variability. These findings bear promising implications for the target application of a tactical deployment. The obtained 87% accuracy indicates that a voice-controlled humanoid robot will be possible to take real-time commands, which very much equips for applications in defense and combat, where actions have to be quick and reliable. The system would further need to be validated for robustness through its ability to work in noisy and harsh environments. In other words, expanding the command vocabulary and implementing context-aware language understanding may render the robot system versatile and potentially able to execute more complex missions. The current study's use of a diverse dataset with multiple speakers suggests that the achieved accuracy is more representative of real-world performance. The result interpretation explains the development of a successful ANN-based speech recognition system using MATLAB for the control of a humanoid robot. The achieved accuracy, precision, recall, and F1 score performance measures point out that the developed model provides an effective mechanism for recognizing voice commands of several speakers. This approach using MATLAB was much easier in development, visualization, and performance aspects if compared with its previous researches. There is room, however, to improve the differentiation of such similar commands and the testing in more diverse environments; nonetheless, the current results are laying a solid base for further development of research and application in tactical deployment scenarios. 6. Future Directions The present work has clearly demonstrated the feasibility and potential of a MATLAB-based ANN approach for humanoid robot control in tasks, e.g., speech recognition. Robustness to Noise : Its robustness with a variety of noise levels and acoustic effects should be further tested and validated in real-world environments. Advanced techniques based on noise reduction and echo cancellation could help make advances when the acoustic quality is not good. Speaker-independent recognition : It would be very useful if it were to cater to speaker-independent speech recognition and even succeed in cases of adaptation to varying accents, speaking speeds, and other speech-specific scenarios. In this line of research, transfer learning and domain adaptation techniques can be applied to explore the potential of the transfer learning-based pre-trained models for speaker-independent recognition and adapt them to fit specific application domains. Natural language processing : An expanded command vocabulary, coupled with techniques of natural language processing, would be a gigantic improvement in enhancing the flexibility and usability of the system. This may enable human-like, intuitive human-robot interaction and hence motivate the users to express the commands and questions in a more conversational way. Multimodal Integration : Integrating the speech recognition system with other modalities, such as gesture recognition and computer vision, will allow much more comprehensive and context-aware human-robot interaction. Adding the visual aspect to verbal commands through cues and gestures, the robot can easily understand and respond to the intentions or instructions of the user. Deep learning architectures, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), have been extensively studied to investigate their applicability for speech recognition and have gained much popularity in recent years. They can help to improve accuracy and robustness. These architectures have shown very promising results in diverse fields such as speech recognition and can be explored for applications of controlling humanoid robots. Real-time performance optimization : One of the prime steps toward glitch-free human-robot interaction is the optimization of the system's performance for real-time application. Techniques that are open for research in this domain may include model compression, quantization, or even hardware acceleration to develop low-latency and high-responsive systems for effective speech recognition. Ethical considerations : Considering the kind of voice-controlled humanlike robots that are today on the market, it is virtually impossible not to think of ethical questions that the deployment and use of such artifacts may bring. These usually refer to privacy and security issues, effects on human-robot relationship issues, among others, and are topics to be treated with great scrutiny by interdisciplinary research. Speech recognition for the control of a humanoid robot presents several future directions and possible challenges. This exploration may pave the way for further advancement in this field for the researchers. The proposed MATLAB-based ANN approach lays the underpinning for further research and development that may be carried out for establishing a base leading to a more natural, intuitive, and effective interaction of human-robot teams in various domains.. 7. Conclusions This research portrays the successful development of a MATLAB-controlled artificial neural network (ANN)-based speech recognition system to control a prototype humanoid robot. There are a lot of signal processing, machine learning, and robotics toolboxes in MATLAB, which will give the proposed approach a comparative advantage over the rest that have been conducted in other studies using platforms such as Arduino and Android gadgets with Bluetooth technology. This provides a simpler, more direct development process compared to the majority of Arduino-based speech recognition systems, in which you usually have to use third-party software and mobile apps. MATLAB offers an all-in-one, integrated workflow with features such as in-built functions for data preprocessing, feature extraction using MFCC, and ANN training. Furthermore, the advanced features of MATLAB in signal processing, such as silence and noise reduction, provide further improvement in the system's robustness and accuracy of speech recognition. This would surely demonstrate the effectiveness of the proposed methodology, with an 87% accuracy of recognition over a great dataset and different speakers. Compared to most Android-based approaches, which usually rely on third-party speech recognition APIs, the MATLAB implementation is actually more flexible and offers control over most of the entire pipeline. It further supports iterative refinement and optimization since its performance can be viewed and rapidly analyzed using MATLAB confusion plots and other graphs. Transparency and interpretability necessary for knowing how such a system can act and be limited is exactly the aid that this provides. The developed voice-controlled humanoid robot prototype in this research brings out evident real-world applicability in different domains. In industries, this can bring about an improvement by increasing safety and efficiency for the worker, since the robot's work can be done in a hands-free and remote-controlled manner in environments that are dangerous to humans. The medical field can include patients that are not actively moving, but voice-controlled humanoids may take care of them or provide remote support and monitoring up to their homes. Interactive humanoid robots are also used in the area of education to engage students and help them learn from verbal instructions coupled with demonstrations. Other possible application areas for businesses are entertainment industries, where voice-controlled humanoids could be utilized to deliver interactive experiences and performances. Declarations Author Contribution 1-Sarika Shrivastava: Sarika Shrivastava was responsible for the study's conceptualization and design. She led the development of the artificial neural network (ANN) model and supervised the overall project. She also contributed to writing and revising the manuscript.2-Somendra Bannerjee: Somendra Bannerjee played a vital role in the data collection and preprocessing stages. He developed the custom graphical user interface (GUI) for recording voice samples and implemented the silence removal and feature normalization techniques. He also assisted in the initial drafting of the methodology section.3-Mayank Srivastava: Mayank Srivastava was involved in the ANN training and evaluation. He designed the neural network architecture, conducted the training sessions, and evaluated the model using various performance metrics. He also contributed to the results and discussion sections of the manuscript.4-Saifullah Khalid: Saifullah Khalid integrated the trained ANN model with the humanoid robot prototype. He also developed the robot control system and the corresponding GUI for real-time command execution and contributed to the sections on robot control and future directions.5-D.K. Nishad: D.K. Nishad provided expertise in applying MATLAB for signal processing and machine learning. He contributed to the literature review and problem formulation and discussed the advantages of using MATLAB over other platforms. He also assisted in the final review and editing of the manuscript.Above statement ensures that each author's specific contributions to the research paper are clearly outlined. Data Availability The datasets used and/or analyzed during the current study available from the corresponding author on reasonable request. References N. Joshi, S. Joshi, and P. Yadav, "Speech Controlled Robotics using Artificial Neural Network," in Proc. 3rd Int. Conf. Image Inf. Process., 2015, pp. 416–420. [A. K. Lekova, "Making humanoid robots teaching assistants by using natural language processing (NLP) cloud-based services," J. Mechatron. Artif. Intell. Eng., 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:250040706 N. Murali, K. Gupta, and S. Bhanot, "Analysis of Q-Learning on Artificial Neural Networks for Robot Control Using Live Video Feed," World Acad. Sci. Eng. Technol. Int. J. Comput. Inf. Eng., vol. 4, no. 8, 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:67324138 S. Paul, "A survey of technologies supporting the design of a multimodal interactive robot for military communication," J. Def. Anal. Logist., 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:265114919 D. Strazdas, A. Hintz, and H. Yano, "Robot and Wizards: An Investigation into Natural Human-Robot Interaction," IEEE Access, vol. 8, pp. 17060–17073, 2020. J. Davila-Chacon, J. Twiefel, and S. Magg, "Enhanced Robot Speech Recognition Using Biomimetic Binaural Sound Source Localization," IEEE Trans. Neural Netw. Learn. Syst., vol. 30, no. 1, pp. 138–150, Jan. 2019. H. Liu, Y. Zhang, H. Fang, and D. Guo, "Towards Robust Human-Robot Collaborative Manufacturing: Multimodal Fusion," IEEE Access, vol. 6, pp. 74762–74771, 2018. R. Jain, S. Gupta, and S. Shukla, "Artificial Intelligence based Voice Controlled Robot," Int. Res. J. Eng. Technol., vol. 6, no. 6, pp. 1–4, Jun. 2019. G. H. Vardhan, K. S. Rao, and S. V. Gangashetty, "Artificial Intelligence & its Application for Speech Recognition," Int. J. Sci. Res., vol. 3, no. 11, pp. 1503–1509, Nov. 2014. D. Dwakar and S. Patel, "Voice Controlled Robotic Vehicle," Int. Res. J. Eng. Technol., vol. 6, no. 6, pp. 1–4, Jun. 2019. Y. A. Menon, S. Sinha, and P. Ediga, "Speech Recognition System For A Voice Controlled Robot With Real-Time Obstacle Detection And Avoidance," Int. J. Electr. Electron. Data Commun., vol. 4, no. 2, pp. 1–5, Feb. 2016. P. K. Kumari, S. Nagamani, and G. Sahithi, "Voice Controlled Home Automation," Int. J. Sci. Res. Sci. Eng. Technol., vol. 6, no. 2, pp. 1–5, Mar. 2019. A.N. Chapgaon and S. S. Kale, "Voice Control Robotic Arm Using Machine Learning," Int. J. Sci. Technol., vol. 29, no. 7, pp. 1–5, Jul. 2020. P. Omprakash and S. Velmurugan, "Voice Recognition Robot for Surveillance," Int. J. Eng. Res. Technol., vol. 6, no. 6, pp. 1–4, Jun. 2017. R. S. A. Sukumar, S. Sukumaran, and S. G. Jacob, "Isolated Question Words Recognition from Speech Queries by Using Artificial Neural Networks," in Proc. 2nd Int. Conf. Comput. Commun. Netw., 2010, pp. 1–5. M. Akour, O. Al Qasem, H. Alsghaier, and K. Al-Radaideh, "Mobile Voice Recognition Based for Smart Home Automation Control," Int. J. Adv. Trends Comput. Sci. Eng., vol. 9, no. 3, pp. 3286–3291, Jun. 2020. B. Jolad and S. Bhabad, "Voice Controlled Robotic Vehicle," Int. Res. J. Eng. Technol., vol. 4, no. 6, pp. 1–4, Jun. 2017. A. Paswan and S. Gupta, "Voice Controlled Robotic Vehicle," Int. J. Sci. Res. Sci. Eng. Technol., vol. 7, no. 2, pp. 1–5, Mar. 2019. S. Patil, "Voice Control Robot Using LabVIEW," in Proc. Int. Conf. Des. Innov. 3Cs Comput. Commun. Control, 2018, pp. 1–5. T. Maddileti and S. Reddy, "Voice Controlled Car using Arduino and Bluetooth Module," Int. J. Eng. Adv. Technol., vol. 9, no. 2, pp. 1–5, Dec. 2019. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4762012","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":340721017,"identity":"f3108c17-ecd5-4585-b4a6-975cb03e31d3","order_by":0,"name":"Sarika Shrivastava","email":"","orcid":"","institution":"Ashoka Institute of Technology \u0026 Management, Varanasi-India","correspondingAuthor":false,"prefix":"","firstName":"Sarika","middleName":"","lastName":"Shrivastava","suffix":""},{"id":340721018,"identity":"eb08c97b-c97e-4d5e-a1e6-1b90ce428cfa","order_by":1,"name":"Somendra Bannerjee","email":"","orcid":"","institution":"Ashoka Institute of Technology \u0026 Management, Varanasi-India","correspondingAuthor":false,"prefix":"","firstName":"Somendra","middleName":"","lastName":"Bannerjee","suffix":""},{"id":340721019,"identity":"0d6a0b1c-ee2f-41af-a4c2-96067349cff0","order_by":2,"name":"Mayank Srivastava","email":"","orcid":"","institution":"Ashoka Institute of Technology \u0026 Management, Varanasi-India","correspondingAuthor":false,"prefix":"","firstName":"Mayank","middleName":"","lastName":"Srivastava","suffix":""},{"id":340721020,"identity":"ec5e72ca-12d5-4928-bf81-4cb101a483ea","order_by":3,"name":"Saifullah Khalid","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9klEQVRIiWNgGAWjYDACHh4GhgcGzCCmAQNDBZgkQksCXMsZorUwQLUwthGhRb7n7MEPCQXWDPIzkjd+Lpx3WN6cvfkAw4+KbTi1GJztS5ZIMEhnMLiRViw9c9thw509xxIYe87cxq2Fn8cAqOUwg4FEjoE077bDjBtu5BgwM7bh1iLfz2P8A6RFfkaO8W/eOYftCWphONtjBraF4UaOmTRvw+FEgloMzpwxswD6hcfgzLMya55j6ckbzhxLOIjPL/I9OcY3PvyxlpNvT958m6fG2nbD8eaDD35U4HEYFPBA6WYweYCgeiRQR4riUTAKRsEoGCEAABUiVsxZQnvkAAAAAElFTkSuQmCC","orcid":"","institution":"IBM Multiactivities Co Ltd.","correspondingAuthor":true,"prefix":"","firstName":"Saifullah","middleName":"","lastName":"Khalid","suffix":""},{"id":340721021,"identity":"2da96fd1-8a2d-4dff-aaec-c27b312b5f87","order_by":4,"name":"D. K. Nishad","email":"","orcid":"","institution":"Dr. Shakuntala Misra National Rehabilitation University","correspondingAuthor":false,"prefix":"","firstName":"D.","middleName":"K.","lastName":"Nishad","suffix":""}],"badges":[],"createdAt":"2024-07-18 10:52:39","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4762012/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4762012/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":62933503,"identity":"66406457-f2ba-46b6-b6df-0fc5b2c9c009","added_by":"auto","created_at":"2024-08-21 08:13:05","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":50767,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBlock diagram illustrating the critical steps in the proposed methodology for developing an ANN-based speech recognition system for humanoid robot control.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/879c3fc6f4cb70e5ae3780f4.png"},{"id":62934402,"identity":"8b225aa0-0752-4e93-ac67-4bcb9c6b49c5","added_by":"auto","created_at":"2024-08-21 08:21:05","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":14148,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBlock diagram illustrating the proposed system architecture\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/3c23f3b9eb6fbaff94768bec.png"},{"id":62934399,"identity":"949e68b0-c634-4f97-87c5-3a155c9b2fbe","added_by":"auto","created_at":"2024-08-21 08:21:05","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":79879,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA Flowchart Illustrating the Steps in the Methodology\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/7cb22bad1e7383da4c993f5c.jpeg"},{"id":62933504,"identity":"0cc6557f-ec76-4746-8ace-56a1043073d3","added_by":"auto","created_at":"2024-08-21 08:13:05","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":15800,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGUI to record voice samples\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/05cf79270193ff440d590820.jpeg"},{"id":62933511,"identity":"2322bd1a-8625-426b-9669-aeca36fe71df","added_by":"auto","created_at":"2024-08-21 08:13:06","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":33242,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFFT analysis of a voice sample\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/6efaa213270ec28223d25936.jpeg"},{"id":62934403,"identity":"1dae27a8-2fb8-4362-a16e-9fc7b0a778ee","added_by":"auto","created_at":"2024-08-21 08:21:06","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":83309,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFrame matrices of (120 * 400) formed for 25 ms window\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/3e5845268bb2dba180ea6695.jpeg"},{"id":62935090,"identity":"80f64487-276a-4c83-8355-88994e1fcfba","added_by":"auto","created_at":"2024-08-21 08:29:05","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":37794,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSilence removal from a voice sample\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/b2abf13d1e1e652ca92f2e54.jpeg"},{"id":62933505,"identity":"2a4a8235-f918-4060-90d8-da98736bd475","added_by":"auto","created_at":"2024-08-21 08:13:05","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":15890,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNeural network architecture for speech recognition\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/7c8e589240adaffd4c02b0fe.jpeg"},{"id":62933507,"identity":"9e91dced-b205-465e-b8c2-907cb7161128","added_by":"auto","created_at":"2024-08-21 08:13:05","extension":"jpeg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":8080,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRobot command GUI for voice-based control\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image9.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/daed6c3c1d1687fe65d3167d.jpeg"},{"id":62934400,"identity":"633c81d8-a183-4e2c-8a3d-ba32bfa5ad30","added_by":"auto","created_at":"2024-08-21 08:21:05","extension":"jpeg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":197438,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix showing the model's performance across different voice commands.\u003c/p\u003e","description":"","filename":"image10.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/123e76d5559fce7c28611e65.jpeg"},{"id":67993126,"identity":"50e0b833-9177-4819-80d8-d4f5bd46edff","added_by":"auto","created_at":"2024-11-01 06:17:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1101061,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4762012/v1/f5143d78-9d27-4928-9ea6-16c2e529c19d.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Advancements in Humanoid Robotics: Designing an Artificial Neural Network-based Speech Recognition Robot for Tactical Deployment","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eHumanoid robotics has made significant strides in recent years, with the goal of creating machines that can replicate human actions and operate autonomously. A key aspect of human-robot interaction is the ability for robots to understand and respond to voice commands. While speech recognition has a long history of innovation, recent advances in artificial intelligence (AI) and deep learning have greatly improved its accuracy and potential applications [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThis paper aims to address the gap in the literature regarding the use of MATLAB for developing a speech recognition system to control humanoid robots. Previous research has primarily focused on using third-party servers like Arduino and Android with Bluetooth technology for speech recognition in robotics [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e][\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e][\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. On the other hand, there are quite a number of benefits that come with MATLAB, which ranges from creation of controllers based on neural networks, removal of silence, and the generation of confusion plots used in the visualization of the speech recognition model's performance. This research is going to try to design and develop a Speech Recognition system that is ANN based through MATLAB for controlling a humanoid robot in tactical deployment. The system shall be designed to be able to understand and follow voice commands. These operate on the basis that the robot undertakes task execution in a real-time environment, especially under defense or combat situations, whereby human intervention is at most risk, in most cases. To undertake this objective, we will employ the following methodology:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eData collection: All voice samples are to be recorded using an in-house GUI on MATLAB.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSpeech processing: Extract the unique characteristics of the voice samples using techniques such as Fast Fourier Transform (FFT) and Mel-frequency cepstrum (MFCC).\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePre-processing: Use silence removal and extract features for bettering their overall accuracy during recognition by the model.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eANN training: Train the processed voice samples with the neural network to recognize the set commands.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRobot control: The developed ANN model shall be embedded in the humanoid robot for the purpose of executing voice commands. This work has its importance in giving advancement in the field of humanoid robotics through capabilities present in MATLAB for better development in speech recognition. The proposed ANN-based system can be used as per the requirements in different applications like industries, medical fields, and defense.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFocusing on tactical deployment, this research aims to demonstrate the feasibility and benefits of using voice-controlled humanoid robots in high-risk and challenging environments. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the Block diagram illustrating the critical steps in the proposed methodology for developing an ANN-based speech recognition system for humanoid robot control.\u003c/p\u003e \u003cp\u003eThe rest of the paper is organized as follows: Section 2 presents a literature review of related work in speech recognition and humanoid robotics. Section 3 describes the problem formulation and the research questions addressed in this study. Section 4 details the research methodology, including data collection, speech processing, pre-processing, ANN training, and robot control. Section 5 presents the results and discussion, followed by the future directions and conclusions in Section 6 and 7.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"2. Literature Review","content":"\u003cp\u003eHumanoid robotics and speech recognition have witnessed significant advancements in recent years, driven by the increasing integration of artificial intelligence (AI) and machine learning techniques. Researchers have explored various approaches to enable effective human-robot interaction through voice commands, aiming to create robots that can understand and respond to human speech in real-time environments.\u003c/p\u003e \u003cp\u003eAuthors had displayed that speech recognition in controlling a humanoid robot highly requires a certain degree of word recognition accuracy for gesture selection [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Thus, his work demonstrates the robot's capability to recognize words and perform gestures upon commanding. Similarly, Authors used speech recognition to control a surgical robot with high precision of the angular displacement of the endoscope. These studies highlight the potential power of speech technology to further develop humanoid robots' capability in many different areas [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. With the growth of complexity in the nature of speech signals, different types of feature extraction and classification techniques were studied by the researchers. The authors developed a mobile robot for the cocktail party effect. The intelligence of human beings makes it possible to recognize speech from different sources concurrently [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. When tested for three speakers, the system produced the average loss of rates for successful recognition between 24% and 42% within two meters. Similarly, The Authors developed a personalized assistant robot, designed to help with dressing tasks, such as easy wearing of shoes [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Their system included user tracking, posture recognition, gesture recognition, and speech recognition for a natural human-robot interaction. A number of researchers have been focusing on the importance of effective communication in human-robot interaction. The author discussed the methods of communication required for successful human-robot collaboration using the Wizard of Oz methodology [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. A new era of cohabitation and cooperation in collaborative work was introduced to his work between robots and humans. The Author developed an automatic speech recognition system operating in challenging environments\u0026mdash;eco-noise and localization of binaural sound sources [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e Their research aimed to understand the influence of human physiognomy on automatic speech recognition and sound source localization, demonstrating improved accuracy when incorporating sound source localization information.\u003c/p\u003e \u003cp\u003eThe application of speech recognition in manufacturing industries has also garnered attention. Authors proposed a human-robot collaboration manufacturing system using deep learning multimodal, which utilizes voice commands, body motion, and hand motion to perform tasks in industrial settings [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. The integration of AI-based voice control has been explored for assisting physically challenged individuals. Authors [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] developed an Android-based smart device that employs Bluetooth technology and object detection using ultrasonic and color sensors to aid disabled people.\u003c/p\u003e \u003cp\u003eSeveral researchers have focused on the development of voice-controlled robots using various technologies. Authors discussed the objectives of AI in understanding human recognition, cost-effective intelligent amplifications, and the development of AI using speech recognition technology to assist physically challenged skilled persons [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Using an Arduino Bluetooth voice control application, the authors demonstrated voice recognition via Bluetooth technology and Android devices [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The Authors proposed a speech recognition system for a voice-controlled robot with real-time obstacle detection, employing Bluetooth communication via smartphones and GSM/Wi-Fi-based internet connectivity [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe application of voice control in home automation has also been explored. Authors focused on implementing voice-controlled home automation, where smart devices integrated with home appliances communicate with cloud servers to execute commands [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Voice-controlled robotic arms have been developed for object recognition and manipulation. The Authors proposed a voice-controlled robotic arm that performs voice recognition, object identification, and object picking using a Raspberry Pi microcontroller and Python scripts [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eVoice recognition has also been studied for surveillance purposes. The authors elaborately described the use of voice-controlled ground vehicles for surveillance in an area of inefficient human work, implemented [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The other field of interest was speech query processing, where the authors experimented with the recognition of isolated question words in speech queries of the Malayalam language using artificial neural networks (ANNs), along with discrete wavelet transformation (DWT) [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThey exploited the integration of mobile voice recognition and IoT for their home automation and robotic vehicle control. Authors [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] proposed a mobile voice-activated home automation system that uses Google voice-to-text API and a Tasker application running on a Raspberry Pi server. The authors built a voice-controlled robotic vehicle using an Android application and IoT, employing a ZigBee module for communication and an EasyVR module for voice command training [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOther authors focus on developing voice-controlled robots for assisting the differently-abled in day-to-day chores. The authors developed a voice-controlled robot using an android application developed by the authors, which converts speech to text and sends commands through Bluetooth [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The authors presented the fabrication of a voice-controlled robotic vehicle using the BT Voice Control for Arduino application, designed exclusively for disabled people [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eVirtual technology and software platforms have also been utilized for voice-controlled robotics. The authors interfaced with a voice-controlled robot using LabVIEW. The interface processes the speech for automation and control within the LabVIEW environment, and commands are transmitted to the robot through Bluetooth and Arduino [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. The system used a laser scanner to build information about unknown environments and the CMU Sphinx toolkit for speech recognition, which gave 95% accuracy in normal and silent environments.\u003c/p\u003e \u003cp\u003eThe literature review indicated that remarkable innovations are achieved in both speech recognition and humanoid robotics. AI, machine learning, and IoT combined have now made it possible for voice-controlled robots to get applications in the manufacturing, home automation, healthcare, and surveillance domains. In terms of robustness, adaptability, and real-world implementation, ANN has never been able to actualize its assertion in this domain. ANN-based speech recognition for humanoid robot control using MATLAB with proposed development and implementation in tactical deployment situations in defense and combat will fill these gaps.\u003c/p\u003e"},{"header":"3. Problem Formulation","content":"\u003cp\u003eRecent years have seen a sharp increase in research at the interface of speech recognition and humanoid robotics, opening ways through which natural and intuitive human-robot interaction can be easily realized. Though previous researches have investigated different methodologies using platforms such as Arduino, coupled with Bluetooth technology, there is still a gap in leveraging the power of MATLAB to develop an ANN-based speech recognition system, tailored explicitly for control in humanoid robot tactical deployment scenarios. The purpose of this research was to examine the feasibility and benefits of using MATLAB-based ANN for speech recognition and conversion into a mechanical control of a humanoid robot for the tactical deployment application. In this respect, this research will consider MATLAB and its rich features and toolboxes in signal processing, machine learning, and robotics beyond the level that has not been fully exploited in the context of speech-controlled humanoid robots. First, MATLAB enables one to design and test ANN models in a single environment through the use of built-in functions for data preparation, feature extraction from the data, and network training. This makes a very smooth, faster workflow in the development process. Secondly, the MATLAB signal processing toolbox has advanced methods, such as the silence deletion process and extracting Mel-frequency cepstral coefficients (MFCC), which can be very useful to make the speech recognition more robust and accurate. Further, MATLAB also has visualization capabilities, like confusion plots, allowing a clear way to analyze the performance of the ANN model with high depth, which enables iterative improvement and optimization. The two-fold objective of this research is to (1) establish validation for using MATLAB in the design of a robust ANN-based speech recognition system and (2) explore its application with respect to controlling a humanoid robot in the context of tactical deployment. This paper focuses on tactical deployment, meeting the unique challenges and requirements for the use of voice commands to control robots in a dynamic environment with varying levels of possibilities and potential noise. The existing approaches in humanoid robot control based on speech recognition, in general, suffer from lack of robustness to noise, speaker independence, and real-time performance. For this reason, in this work, those limitations are addressed using the signal processing power of MATLAB in conjunction with ANN models. The proposed system, incorporating techniques such as removing silence intervals, extracting MFCC features, and retraining the model, ensures that good noise robustness, speaker independence, and effective real-time implementation are taken care of. The proposed methodology for the present investigation involves the following key steps:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData Collection\u003c/b\u003e: Recording a diverse dataset of voice commands using MATLAB's audio acquisition tools.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003ePreprocessing\u003c/b\u003e: Applying signal processing techniques, including silence removal and MFCC feature extraction, to enhance the quality and discriminability of the speech data.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eANN Model Development\u003c/b\u003e: Designing and training an ANN architecture in MATLAB for speech recognition, utilizing techniques like backpropagation and regularization.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eModel Evaluation\u003c/b\u003e: The assessment of ANN model performance after training, using indicators such as accuracy, precision, and recall, supported by visualization tools from MATLAB.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eRobot Integration\u003c/b\u003e: This involves interfacing the speech recognition system with a humanoid robot platform, enabling real-time control based on recognized voice commands.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents a high-level block diagram illustrating the proposed system architecture.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe successful development of a MATLAB-based speech recognition system for humanoid robot control has significant potential for impact across various domains. In industries, voice-controlled robots can enhance worker safety and efficiency by allowing hands-free operation in manufacturing and assembly tasks. In healthcare, such systems can assist patients with limited mobility or enable remote surgical procedures. In defense and security applications, voice-controlled humanoid robots can be deployed for reconnaissance, bomb disposal, or search and rescue missions, minimizing human risk. By addressing the research above objectives and leveraging the capabilities of MATLAB, this study aims to contribute to the advancement of speech recognition and humanoid robotics, opening up new possibilities for natural and efficient human-robot interaction in tactical deployment scenarios and beyond.\u003c/p\u003e"},{"header":"4. Research Methodology","content":"\u003cp\u003eThe proposed speech recognition system for humanoid robot control consists of two main components: the MATLAB-based artificial neural network (ANN) for speech recognition and the robot prototype that executes the recognized voice commands. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e presents a flowchart illustrating the steps in the methodology.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e4.1. Data Collection\u003c/h2\u003e \u003cp\u003eTo train the ANN model, a dataset of voice commands was collected using a custom-built graphical user interface (GUI) in MATLAB. The GUI, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, allows users to record voice samples for each command. In this study, five commands were used: \"move forward,\" \"move backward,\" \"turn right,\" \"turn left,\" and \"stop.\" 100 voice samples were recorded for each command, with 70 samples used for training and 30 for testing. The voice commands were recorded from 10 speakers (5 male and five female) to ensure speaker variability in the dataset.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e4.2. Speech Processing\u003c/h2\u003e \u003cp\u003eThe recorded voice samples undergo speech processing to extract unique features such as pitch, fundamental frequency, and spectral content. Two main techniques were employed for feature extraction:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eFast Fourier Transform (FFT) analysis\u003c/b\u003e: The voice sample spectrum is considered whole, and features like spectral content are calculated using FFT. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e shows the FFT analysis of a single voice sample.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eMel-frequency cepstrum (MFCC)\u003c/b\u003e: To capture the sound variation more accurately, the voice samples are divided into frames of 20\u0026ndash;30 ms. MFCC is then applied to each frame to extract 14 coefficients that represent the short-term power spectrum of the sound. Figure\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows the frame matrices formed for a 25 ms window.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e4.3. Pre-processing\u003c/h2\u003e \u003cp\u003eBefore training the ANN model, the extracted features undergo pre-processing steps:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eSilence removal: A custom MATLAB script was developed to remove silent frames from the voice samples. The script divides the sound file into frames, calculates the energy of each frame, and discards frames with energy less than 2% of the total energy. Figure\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e shows the effect of silence removal on a voice sample.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eFeature normalization\u003c/b\u003e: The MFCC features extracted from each frame are normalized to have zero mean and unit variance. This helps improve the convergence speed and stability of the ANN training process.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eAfter pre-processing, the final training dataset consists of a (14 * 3771) matrix of MFCC features and a (5 * 3771) matrix of corresponding target labels.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e4.4. ANN Architecture and Training\u003c/h2\u003e \u003cp\u003eA feedforward neural network was designed for speech recognition using MATLAB's Deep Learning Toolbox. The ANN architecture, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, consists of:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eInput layer: 14 neurons, corresponding to the 14 MFCC features\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eHidden layer: 50 neurons with ReLU activation function\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eOutput layer: 5 neurons with softmax activation, corresponding to the five voice commands\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe ANN was trained using the scaled conjugate gradient backpropagation algorithm with a learning rate 0.01 and a maximum of 1000 epochs. The training data was divided into 70% for training, 15% for validation, and 15% for testing. Early stopping was used to prevent overfitting, and the validation patience was 10 epochs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e4.5. Robot Control\u003c/h2\u003e \u003cp\u003eThe trained ANN model was integrated with the humanoid robot prototype using MATLAB's Hardware Support Package for Arduino. A custom robot control GUI, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e, was developed to allow users to give voice commands and visualize the robot's response. When a voice command is given, the ANN model predicts the corresponding action, which is then executed by the robot's actuators and control modules.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFollowing this systematic methodology, the proposed speech recognition system achieves accurate and responsive voice-based control of the humanoid robot prototype. The modular architecture allows for easy customization and extension to incorporate additional voice commands and functionalities as needed.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Results \u0026 Discussions","content":"\u003cp\u003eThe developed ANN-based speech recognition system for humanoid robot control achieved promising results, demonstrating the feasibility and potential of using MATLAB for this application. The trained model reached an overall accuracy of 87% in recognizing the five voice commands: \"move forward,\" \"move backward,\" \"turn right,\" \"turn left,\" and \"stop.\" This accuracy was obtained using a testing dataset of 150 samples (30 per command) recorded from ten speakers (five male and five females) to be able to make a good estimate of performance while taking into account speaker variability. More metrics were computed to give a more rounded measure of the model's performance. The corresponding precision, recall, and F1 scores (with 95% CI) of the system are 0.88, 0.87, and 0.87 (0.82\u0026ndash;0.92), respectively. Such results evidently show that it gets to the top as far as accuracy is concerned and that a very good balance between precision and recall is achieved, minimizing both false positives and false negatives. The confidence interval of the accuracy really helps to perceive how significantly reliable and robust the system is performing. The performance of the model across the different voice commands can be visualized in the confusion matrix in Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e. The confusion matrix, however, indicates that the system was slightly less accurate on the \"turn right\" and \"turn left\" commands. This would indicate that the model could struggle with discriminating both these directions, most likely because they sound quite similar to one another or have an almost similar acoustic representation. Further work should refine feature extraction and model architecture for discrimination between such commands.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFirst advantage is that previous research using Arduino and Android platforms could be tested and executed only on a limited hardware setup; therefore, the work's usefulness could not be general for mankind. Development is easy with built-in functions of MATLAB for signal processing, feature extraction (MFCC), and ANN training, making it possible to easily try out different model architectures. This further enabled the results to be visualized with confusion plots and other graphs, for better understanding of the performance and limitations of the system.\u003c/p\u003e \u003cp\u003eWhile the performance of the MATLAB approach is slightly inferior to some of the reported results, it is important to mention that most of the studies often applied fewer speakers or had a smaller database, which may not encompass the whole complexity of speech variability. These findings bear promising implications for the target application of a tactical deployment. The obtained 87% accuracy indicates that a voice-controlled humanoid robot will be possible to take real-time commands, which very much equips for applications in defense and combat, where actions have to be quick and reliable. The system would further need to be validated for robustness through its ability to work in noisy and harsh environments. In other words, expanding the command vocabulary and implementing context-aware language understanding may render the robot system versatile and potentially able to execute more complex missions.\u003c/p\u003e \u003cp\u003eThe current study's use of a diverse dataset with multiple speakers suggests that the achieved accuracy is more representative of real-world performance.\u003c/p\u003e \u003cp\u003eThe result interpretation explains the development of a successful ANN-based speech recognition system using MATLAB for the control of a humanoid robot. The achieved accuracy, precision, recall, and F1 score performance measures point out that the developed model provides an effective mechanism for recognizing voice commands of several speakers. This approach using MATLAB was much easier in development, visualization, and performance aspects if compared with its previous researches. There is room, however, to improve the differentiation of such similar commands and the testing in more diverse environments; nonetheless, the current results are laying a solid base for further development of research and application in tactical deployment scenarios.\u003c/p\u003e"},{"header":"6. Future Directions","content":"\u003cp\u003eThe present work has clearly demonstrated the feasibility and potential of a MATLAB-based ANN approach for humanoid robot control in tasks, e.g., speech recognition.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eRobustness to Noise\u003c/b\u003e: Its robustness with a variety of noise levels and acoustic effects should be further tested and validated in real-world environments. Advanced techniques based on noise reduction and echo cancellation could help make advances when the acoustic quality is not good.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSpeaker-independent recognition\u003c/b\u003e: It would be very useful if it were to cater to speaker-independent speech recognition and even succeed in cases of adaptation to varying accents, speaking speeds, and other speech-specific scenarios. In this line of research, transfer learning and domain adaptation techniques can be applied to explore the potential of the transfer learning-based pre-trained models for speaker-independent recognition and adapt them to fit specific application domains.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eNatural language processing\u003c/b\u003e: An expanded command vocabulary, coupled with techniques of natural language processing, would be a gigantic improvement in enhancing the flexibility and usability of the system. This may enable human-like, intuitive human-robot interaction and hence motivate the users to express the commands and questions in a more conversational way.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eMultimodal Integration\u003c/b\u003e: Integrating the speech recognition system with other modalities, such as gesture recognition and computer vision, will allow much more comprehensive and context-aware human-robot interaction. Adding the visual aspect to verbal commands through cues and gestures, the robot can easily understand and respond to the intentions or instructions of the user.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eDeep learning architectures, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), have been extensively studied to investigate their applicability for speech recognition and have gained much popularity in recent years. They can help to improve accuracy and robustness. These architectures have shown very promising results in diverse fields such as speech recognition and can be explored for applications of controlling humanoid robots.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eReal-time performance optimization\u003c/b\u003e: One of the prime steps toward glitch-free human-robot interaction is the optimization of the system's performance for real-time application. Techniques that are open for research in this domain may include model compression, quantization, or even hardware acceleration to develop low-latency and high-responsive systems for effective speech recognition.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eEthical considerations\u003c/b\u003e: Considering the kind of voice-controlled humanlike robots that are today on the market, it is virtually impossible not to think of ethical questions that the deployment and use of such artifacts may bring. These usually refer to privacy and security issues, effects on human-robot relationship issues, among others, and are topics to be treated with great scrutiny by interdisciplinary research.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eSpeech recognition for the control of a humanoid robot presents several future directions and possible challenges. This exploration may pave the way for further advancement in this field for the researchers. The proposed MATLAB-based ANN approach lays the underpinning for further research and development that may be carried out for establishing a base leading to a more natural, intuitive, and effective interaction of human-robot teams in various domains..\u003c/p\u003e"},{"header":"7. Conclusions","content":"\u003cp\u003eThis research portrays the successful development of a MATLAB-controlled artificial neural network (ANN)-based speech recognition system to control a prototype humanoid robot. There are a lot of signal processing, machine learning, and robotics toolboxes in MATLAB, which will give the proposed approach a comparative advantage over the rest that have been conducted in other studies using platforms such as Arduino and Android gadgets with Bluetooth technology.\u003c/p\u003e \u003cp\u003eThis provides a simpler, more direct development process compared to the majority of Arduino-based speech recognition systems, in which you usually have to use third-party software and mobile apps. MATLAB offers an all-in-one, integrated workflow with features such as in-built functions for data preprocessing, feature extraction using MFCC, and ANN training. Furthermore, the advanced features of MATLAB in signal processing, such as silence and noise reduction, provide further improvement in the system's robustness and accuracy of speech recognition. This would surely demonstrate the effectiveness of the proposed methodology, with an 87% accuracy of recognition over a great dataset and different speakers.\u003c/p\u003e \u003cp\u003eCompared to most Android-based approaches, which usually rely on third-party speech recognition APIs, the MATLAB implementation is actually more flexible and offers control over most of the entire pipeline. It further supports iterative refinement and optimization since its performance can be viewed and rapidly analyzed using MATLAB confusion plots and other graphs. Transparency and interpretability necessary for knowing how such a system can act and be limited is exactly the aid that this provides.\u003c/p\u003e \u003cp\u003eThe developed voice-controlled humanoid robot prototype in this research brings out evident real-world applicability in different domains. In industries, this can bring about an improvement by increasing safety and efficiency for the worker, since the robot's work can be done in a hands-free and remote-controlled manner in environments that are dangerous to humans. The medical field can include patients that are not actively moving, but voice-controlled humanoids may take care of them or provide remote support and monitoring up to their homes. Interactive humanoid robots are also used in the area of education to engage students and help them learn from verbal instructions coupled with demonstrations. Other possible application areas for businesses are entertainment industries, where voice-controlled humanoids could be utilized to deliver interactive experiences and performances.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003e1-Sarika Shrivastava: Sarika Shrivastava was responsible for the study's conceptualization and design. She led the development of the artificial neural network (ANN) model and supervised the overall project. She also contributed to writing and revising the manuscript.2-Somendra Bannerjee: Somendra Bannerjee played a vital role in the data collection and preprocessing stages. He developed the custom graphical user interface (GUI) for recording voice samples and implemented the silence removal and feature normalization techniques. He also assisted in the initial drafting of the methodology section.3-Mayank Srivastava: Mayank Srivastava was involved in the ANN training and evaluation. He designed the neural network architecture, conducted the training sessions, and evaluated the model using various performance metrics. He also contributed to the results and discussion sections of the manuscript.4-Saifullah Khalid: Saifullah Khalid integrated the trained ANN model with the humanoid robot prototype. He also developed the robot control system and the corresponding GUI for real-time command execution and contributed to the sections on robot control and future directions.5-D.K. Nishad: D.K. Nishad provided expertise in applying MATLAB for signal processing and machine learning. He contributed to the literature review and problem formulation and discussed the advantages of using MATLAB over other platforms. He also assisted in the final review and editing of the manuscript.Above statement ensures that each author's specific contributions to the research paper are clearly outlined.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets used and/or analyzed during the current study available from the corresponding author on reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eN. Joshi, S. Joshi, and P. Yadav, \"Speech Controlled Robotics using Artificial Neural Network,\" in Proc. 3rd Int. Conf. Image Inf. Process., 2015, pp. 416\u0026ndash;420.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e[A. K. Lekova, \"Making humanoid robots teaching assistants by using natural language processing (NLP) cloud-based services,\" J. Mechatron. Artif. Intell. Eng., 2022. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://api.semanticscholar.org/CorpusID:250040706\u003c/span\u003e\u003cspan address=\"https://api.semanticscholar.org/CorpusID:250040706\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eN. Murali, K. Gupta, and S. Bhanot, \"Analysis of Q-Learning on Artificial Neural Networks for Robot Control Using Live Video Feed,\" World Acad. Sci. Eng. Technol. Int. J. Comput. Inf. Eng., vol. 4, no. 8, 2017. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://api.semanticscholar.org/CorpusID:67324138\u003c/span\u003e\u003cspan address=\"https://api.semanticscholar.org/CorpusID:67324138\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Paul, \"A survey of technologies supporting the design of a multimodal interactive robot for military communication,\" J. Def. Anal. Logist., 2023. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://api.semanticscholar.org/CorpusID:265114919\u003c/span\u003e\u003cspan address=\"https://api.semanticscholar.org/CorpusID:265114919\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD. Strazdas, A. Hintz, and H. Yano, \"Robot and Wizards: An Investigation into Natural Human-Robot Interaction,\" IEEE Access, vol. 8, pp. 17060\u0026ndash;17073, 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ. Davila-Chacon, J. Twiefel, and S. Magg, \"Enhanced Robot Speech Recognition Using Biomimetic Binaural Sound Source Localization,\" IEEE Trans. Neural Netw. Learn. Syst., vol. 30, no. 1, pp. 138\u0026ndash;150, Jan. 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eH. Liu, Y. Zhang, H. Fang, and D. Guo, \"Towards Robust Human-Robot Collaborative Manufacturing: Multimodal Fusion,\" IEEE Access, vol. 6, pp. 74762\u0026ndash;74771, 2018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eR. Jain, S. Gupta, and S. Shukla, \"Artificial Intelligence based Voice Controlled Robot,\" Int. Res. J. Eng. Technol., vol. 6, no. 6, pp. 1\u0026ndash;4, Jun. 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG. H. Vardhan, K. S. Rao, and S. V. Gangashetty, \"Artificial Intelligence \u0026amp; its Application for Speech Recognition,\" Int. J. Sci. Res., vol. 3, no. 11, pp. 1503\u0026ndash;1509, Nov. 2014.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD. Dwakar and S. Patel, \"Voice Controlled Robotic Vehicle,\" Int. Res. J. Eng. Technol., vol. 6, no. 6, pp. 1\u0026ndash;4, Jun. 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. A. Menon, S. Sinha, and P. Ediga, \"Speech Recognition System For A Voice Controlled Robot With Real-Time Obstacle Detection And Avoidance,\" Int. J. Electr. Electron. Data Commun., vol. 4, no. 2, pp. 1\u0026ndash;5, Feb. 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. K. Kumari, S. Nagamani, and G. Sahithi, \"Voice Controlled Home Automation,\" Int. J. Sci. Res. Sci. Eng. Technol., vol. 6, no. 2, pp. 1\u0026ndash;5, Mar. 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA.N. Chapgaon and S. S. Kale, \"Voice Control Robotic Arm Using Machine Learning,\" Int. J. Sci. Technol., vol. 29, no. 7, pp. 1\u0026ndash;5, Jul. 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eP. Omprakash and S. Velmurugan, \"Voice Recognition Robot for Surveillance,\" Int. J. Eng. Res. Technol., vol. 6, no. 6, pp. 1\u0026ndash;4, Jun. 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eR. S. A. Sukumar, S. Sukumaran, and S. G. Jacob, \"Isolated Question Words Recognition from Speech Queries by Using Artificial Neural Networks,\" in Proc. 2nd Int. Conf. Comput. Commun. Netw., 2010, pp. 1\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Akour, O. Al Qasem, H. Alsghaier, and K. Al-Radaideh, \"Mobile Voice Recognition Based for Smart Home Automation Control,\" Int. J. Adv. Trends Comput. Sci. Eng., vol. 9, no. 3, pp. 3286\u0026ndash;3291, Jun. 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eB. Jolad and S. Bhabad, \"Voice Controlled Robotic Vehicle,\" Int. Res. J. Eng. Technol., vol. 4, no. 6, pp. 1\u0026ndash;4, Jun. 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA. Paswan and S. Gupta, \"Voice Controlled Robotic Vehicle,\" Int. J. Sci. Res. Sci. Eng. Technol., vol. 7, no. 2, pp. 1\u0026ndash;5, Mar. 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Patil, \"Voice Control Robot Using LabVIEW,\" in Proc. Int. Conf. Des. Innov. 3Cs Comput. Commun. Control, 2018, pp. 1\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT. Maddileti and S. Reddy, \"Voice Controlled Car using Arduino and Bluetooth Module,\" Int. J. Eng. Adv. Technol., vol. 9, no. 2, pp. 1\u0026ndash;5, Dec. 2019.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Humanoid robotics, Speech recognition, Artificial neural networks, MATLAB, Tactical deployment, Human-robot interaction","lastPublishedDoi":"10.21203/rs.3.rs-4762012/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4762012/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe development of an ANN-based speech recognition system using MATLAB for controlling a humanoid robot prototype is presented in this research. Our proposed approach, therefore, will be in a good position to derive significant benefits from the powerful features and toolboxes in MATLAB in the areas of signal processing, machine learning, and robotics, way above most of the works available that have used platforms such as Arduino and Android with Bluetooth technology. The system is trained to recognize the five speech commands: \"move forward,\" \"move backward,\" \"turn right,\" \"turn left,\" and \"stop.\" The custom GUI software has been developed to collect the data. It is processed by Fast Fourier Transform (FFT) and Mel-frequency cepstral coefficients (MFCC). After that, it goes through silence removal and normalized techniques with pre-processing ANN classifiers and integrated with the robot control system. The accuracy of the trained model on test data with multiple speakers is 87% and holds good for general purposes without being biased towards any specific speaker's voice characteristics. Developed prototype, hence, proves the feasibility and potential of MATLAB for the speech recognition task in control with humanoid robot and its applications in different domains like industries, healthcare, and defense. The modular architecture allows for easy customization and extension to incorporate additional voice commands and functionalities. Future research directions include improving robustness to noise, speaker-independent recognition, and integration with other modalities like gesture recognition and computer vision.\u003c/p\u003e","manuscriptTitle":"Advancements in Humanoid Robotics: Designing an Artificial Neural Network-based Speech Recognition Robot for Tactical Deployment","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-21 08:13:00","doi":"10.21203/rs.3.rs-4762012/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"67d32ff3-0c7c-4d97-be0a-c3e5d0db3a24","owner":[],"postedDate":"August 21st, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":36100015,"name":"Physical sciences/Energy science and technology"},{"id":36100016,"name":"Physical sciences/Engineering"}],"tags":[],"updatedAt":"2024-11-01T06:08:42+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-21 08:13:00","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4762012","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4762012","identity":"rs-4762012","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00