Hybrid Machine Learning Models for Automated Classification of Weld Defects in Gas Metal Arc Robotic Welding

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract The research study explores an advanced approach to classify weld bead defects in a Gas Metal Arc Welding (GMAW) process of robotic welding, contributing to the growing field of intelligent manufacturing and quality control in industrial automation. To precisely categorize six classes of weld bead such as lack of fusion, burn-through, misalignment, lack of penetration, contamination, and good weld, a hybrid machine learning approach that combines Convolutional Neural Networks (CNN) and Vision Transformers (ViT) is utilized. The hybrid model leverages the complementary advantages of CNN for the extraction of localized features and ViT for global contextual awareness, resulting in superior classification performance compared to traditional architectures such as ResNet and conventional CNN models. The performance evaluation of proposed model demonstrates the model's robustness and reliability in diverse operational scenarios. The findings highlight the potential of integrating hybrid deep learning models into industrial automation systems to enhance weld bead defect detection to reduce operational inefficiencies and ensure consistent weld quality in manufacturing process.
Full text 70,755 characters · extracted from preprint-html · click to expand
Hybrid Machine Learning Models for Automated Classification of Weld Defects in Gas Metal Arc Robotic Welding | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Hybrid Machine Learning Models for Automated Classification of Weld Defects in Gas Metal Arc Robotic Welding B. Vinod, M. P. Anbarasi, K. Senthil Kumar, C. Senthamilarasi, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7098058/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 11 Jan, 2026 Read the published version in Discover Artificial Intelligence → Version 1 posted 13 You are reading this latest preprint version Abstract The research study explores an advanced approach to classify weld bead defects in a Gas Metal Arc Welding (GMAW) process of robotic welding, contributing to the growing field of intelligent manufacturing and quality control in industrial automation. To precisely categorize six classes of weld bead such as lack of fusion, burn-through, misalignment, lack of penetration, contamination, and good weld, a hybrid machine learning approach that combines Convolutional Neural Networks (CNN) and Vision Transformers (ViT) is utilized. The hybrid model leverages the complementary advantages of CNN for the extraction of localized features and ViT for global contextual awareness, resulting in superior classification performance compared to traditional architectures such as ResNet and conventional CNN models. The performance evaluation of proposed model demonstrates the model's robustness and reliability in diverse operational scenarios. The findings highlight the potential of integrating hybrid deep learning models into industrial automation systems to enhance weld bead defect detection to reduce operational inefficiencies and ensure consistent weld quality in manufacturing process. Welding Robot Hybrid CNN + ViT Model Deep Learning Vision Transformer Convolutional Neural Network Industrial Automation Figures Figure 1 Figure 2 Figure 3 Figure 4 I. INTRODUCTION Welding is a fundamental process in the manufacturing sector, playing an essential part in guaranteeing the dependability and structural soundness of different components across industries such as aerospace, construction, and automotive production line. Among different welding techniques, Gas Metal Arc Welding (GMAW) is widely adopted due to its efficiency, adaptability, and suitability for high-production environments. However, ensuring consistent weld quality remains a significant challenge, as defects can compromise the durability and safety of welded structures [1]. The detection and classification of weld bead defects are critical tasks that traditionally rely on manual visual examination as well as non-destructive testing (NDT) methods like radiography or ultrasonic testing. These conventional approaches while beneficial, are often time-consuming, labor-intensive, and heavily reliant on the knowledge of experienced operators, making them susceptible to human error and inconsistencies in quality assessment [2]. The Automated weld defect detection systems have been made possible by recent advancements in artificial intelligence (AI), specifically in the fields of Machine Learning and Computer Vision [3]. These AI-driven solutions offer substantial improvements in accuracy, speed, and repeatability, making them ideal for integration into modern manufacturing workflows. One kind of deep learning model that has shown great promise in defect classification is the Convolutional Neural Network (CNN) by effectively extracting local features from weld images. However, CNN-based approaches face constraints in gathering global contextual data and long-range dependencies, which are essential for accurately distinguishing complex defect patterns. On the other hand, Transformer-based models, specifically Vision Transformers (ViTs), excel in global feature modelling but often come with high computational costs, limiting their practical deployment in real-time industrial applications [4]. To address these challenges, this study proposes a hybrid AI based machine learning model that blends the advantages of CNNs and ViTs to enhance weld defect classification. The integration of such hybrid models into industrial automation systems has the potential to revolutionize quality control in welding processes, reducing time bounded on manual inspection, minimizing production downtime, and improving overall manufacturing reliability [5]. The combination of a hybrid CNN + ViT model designed for weld defect classification, combines the advantages of ViT with conventional CNN model architectures. The hybrid technique guarantees precise identification of defects such as lack of fusion, misalignment, burn-through, lack of penetration, contamination (lack of shielding gas) by the integration of ML techniques. The proposed model not only demonstrates the hybrid model effectiveness but also its potential for industrial adoption by comparing it to conventional models such as ResNet and conventional CNN based models [6]. The proposed system utilizes the machine learning model for vision based inspection in automated robotic welding environment by providing high quality inspection of the weld beads and add to the expanding corpus of knowledge on intelligent manufacturing robotic systems. This research project uses the following structure: methodology is presented in Section 2, the performance metrics evaluation of the proposed approach and its outcomes is discussed in Section 3 and the work's conclusion and future scope is outlined in Section 4. II. METHODOLOGY This research focuses on automating the classification of weld defects using a hybrid CNN + ViT model. The methodology involves of a robust dataset and a hybrid model architecture tailored for weld defect classification. Over 20,000 weld images across six defect classes were collected, annotated, and systematically split for training, validation, and testing. This division was carefully designed to ensure model performance optimization and unbiased evaluation. Leveraging the strengths of CNNs for localized feature extraction and ViTs for capturing global dependencies, the hybrid architecture addressed the limitations of standalone models. By implementing the hybrid model, this work achieved significant improvements in weld defect classification accuracy and robustness, particularly in real-world industrial settings. A. Datasets The dataset consist of weld images from six classes of defect identification containing good welds lack of fusion, misalignment, burn-through, lack of penetration, contamination (lack of shielding gas), were collected from an open source platform for this work. The images were annotated to guarantee high-quality labelling in order to provide robust training and evaluation of machine learning models, meticulously relating to particular defect categories to implement the model in real welding environments [7]. The dataset was divided into three subsets: 80% for training, 10% for validation, and 10% for testing.in order to avoid overfitting and maximize model performance. This division guarantees sufficient training samples while preserving sufficient data for objective assessment and hyperparameters adjustment. B. Data Pre-processing Data Preprocessing is the fundamental step to make the data clean, structured and ready for analysis, forming the foundation for accurate machine learning models. This process addresses issues like missing values, inconsistencies, and noise in the data, ensuring that machine learning models receive accurate and relevant input for better performance and accuracy. Prior to model training, the following pre-processing steps were applied for data pre-processing. Image Normalization : In order to maintain numerical stability throughout training, pixel intensities were scaled between 0 and 1 Data Augmentation : The dataset size had been artificially increased using methods including random rotations, flips, scaling, and contrast modifications, which improved the model's capacity to generalize across unknown data Resizing : All images were resized to a uniform resolution of 224x224 pixels, consistent with the CNN and ViT implementation C. Dataset Characteristics Potential imbalances were found by looking at the class distribution, techniques like oversampling or class-weighting were used to mitigate the imbalance created. By including a variety of fault types, the dataset is guaranteed to accurately reflect actual welding conditions, allowing the trained model to be used to real-world situations in industrial settings. The work is based on the rich dataset, a robust pre-processing pipeline, and a well-structured train-test split, which allow the hybrid CNN + ViT model to attain state-of-the-art performance in weld defect classification. 1) Defect Classification The weld defect classification using machine learning techniques can automate the inspection process, improve the efficiency of quality control, and reduce the need for manual inspection. For the proposed defect classification, three models were utilized to perform weld bead defects classification, the significance of these models are discussed with its classification performance. 2) ResNet-50 model ResNet-50 was used as the foundation for feature extraction in this work because of its performance in object classification challenges. ResNet-50, a 50-layer deep convolutional neural network, was utilized to use residual connections to address the vanishing gradient issue [8]. The network may learn residual mappings instead of direct mappings, which makes it easier to train deeper patterns. By utilizing its stacked convolutional layers, batch normalization, ReLU activation function, and identity shortcut connections, the model was able to extract patterns and local spatial characteristics from defect images. It performs from low-level edges to high-level semantic characteristics, this model architecture excels at capturing hierarchical representations, allowing the model to precisely detect and categorize complex weld defect patterns [8]. Despite its efficacy, ResNet-50 was not immediately chosen as the implementation model because of some experimental constraints, like requirements of considerable labelled datasets, overfitting, computational resource demands. The model performed quite well in feature extraction, it occasionally had trouble in maintaining consistent performance of processing weld defect images with minute class changes. The model becomes more complex as a result of the skip connections, which may result in increased processing and memory needs. These restrictions spurred research into sophisticated hybrid designs with increased classification robustness and accuracy [9]. 3) Conventional Neural Network model The objective of classifying weld defects was first accomplished using a traditional Convolutional Neural Network (CNN) model. It includes multiple convolutional layers used in the architecture to extract features, and pooling layers were added to minimize spatial dimensions. To gradually acquire hierarchical representations of the input weld defect images from low- level edges to high-level patterns these layers were stacked and after that the retrieved features were classified into six different defect types using fully linked layers. The CNN uses convolution operations to efficiently learn spatial hierarchies and local dependencies [10]. During testing, the model exhibited encouraging preliminary results and performed satisfactorily in capturing local aspects of weld defects [11]. However, it shows certain limitations and constraints in catching long-range dependencies and global context, which are essential for predicting minute variations between related fault classes. In the case of noisy and complicated grayscale images, the CNN occasionally performed poor in generalization, producing predictions that were inconsistent. These difficulties can be resolved with a hybrid model architecture, which led to the incorporation of vision transformers into a hybrid framework which greatly enhanced the classification performance. 4) Proposed Hybrid CNN + ViT model A hybrid architecture proposed in this work combines the advantages of both standalone models by incorporating ViT into the CNN framework. The efficient weld defect classification is achieved by the CNN + ViT hybrid model as it combines the advantages of ViT and CNN. The CNN layers use patterns, textures, and edges to extract localized features. The final feature map is used as the input for the ViT module after these features have been processed through convolutional, pooling and batch normalizing layers. Global attention-based processing is made possible by transforming the feature map and passing it to the ViT in place of the conventional dense layers in CNN. A sequence is created by linearly embedding the patches (such as 16*16) into vectors from the CNN feature map. To preserve spatial information is crucial for comprehending the interactions between patches, positional encoding is added. Transformer encoder layers, which include Feed Forward Neural networks (FFN) and Multi-Head Self Attention (MHSA), can process this sequence. While FFN refines features by non-linear transformations, backed by residual connections and stability preserving normalization, MHSA captures global dependencies. A classification head processes the enhanced feature representation and divides the image into six defect classes using a fully linked layer with SoftMax activation. By combining localized and global feature extraction, our hybrid technique efficiently addresses the complexity of weld defect patterns while achieving higher accuracy [12]. 5) Advantages of Hybrid model The hybrid proposed model overcome the shortcomings of conventional CNN models by utilizing ViT in collecting long-range relationships and global context within images. The CNN shows difficulty in modelling the general relationships between various parts of an image, even when they are quite good at learning local characteristics through convolutional process. This makes it difficult to precisely identify tiny changes and structural differences in the entire image. ViT is a promising addition to the model architecture because of its self-attention mechanism, which excels at analysing global dependencies and comprehending the image holistically [13]. The classification of weld defects significantly improved with the implementation of the hybrid model with CNN + ViT architecture, thereby fusing the ViT in global reasoning with CNN to extract fine-grained local features. Better generalization resulted from this hybrid model for feature learning, especially when handling noisy and complicated images. All six defect classes showed more consistent predictions, also outperformed better in terms of accuracy. It shows enhanced flexibility in responding to changes in the sizes, forms and orientations of defects, which makes it a best model for automated GMAW robotic welding applications. Overall, the hybrid CNN + ViT model offered a more precise and efficient classification framework by successfully addressing the drawbacks of standalone models for defect classification in real time environments [14]. D. Implementation of hybrid model The open source PyTorch deep learning framework was chosen for the proposed model execution due to its adaptability and effective support for neural network building. A Jetson AGX Orin GPU, which offers excellent edge computing performance is used for training and testing. Particularly for the hybrid CNN+ViT architecture, the GPU's capabilities guaranteed seamless handling of computationally demanding tasks, such as model training and inference [15]. This type of configuration facilitated to process real time data and it eases the simulation process to effectively investigate by fusing distinct architectures and tuning of hyperparameters. When the model were trained with 50 epochs to prevent overfitting an early termination methods were used. To guarantee effective memory use and steady gradient updates batch size was chosen, and to strike a balance between convergence speed and stability, the learning rate was adjusted. Dropout was used to randomly turn off neurons during training for overfitting reduction. The cross-validation was applied to ensure an all-round evaluation to lower the impact of dataset splits. To validate that the consistency was maintained across multiple folds, following the dataset's division into three subsets of training, testing and validation k-fold cross validation had been used. The multiclass classification problem was addressed along with the Adam optimizer during training by using a categorical cross-entropy loss function due to its adaptive learning rates and fast convergence [16]. Rotation, flipping, and contrast adjustment are examples of data augmentation techniques that were used to make the dataset more diverse and to make use of the model's robust design. This mixing philosophy showed the model outperforms well for industrial applications by ensuring reliable and accurate predictions in real time environments [17]. III. PERFORMANCE METRICS EVALUATION A. Performance Metrics Comparison To effectively compare the common metrics used to assess the deep learning models performance in classification tasks include accuracy, precision, ROC AUC graphs are used. Metrics commonly used to evaluate a hybrid deep learning model are shown as a confusion matrix that contrasts the actual results with the predictions made by the model. It divides the predictions from the four categories such as True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN). The performance evaluation of CNN model and hybrid CNN+ViT model are represented by confusion matrix is shown in Fig. 2 1. Accuracy: It represents the percentage of cases that were correctly classified out of all instances Accuracy = (TP + TN) (TP + TN + FP + FN) 2. AUC-ROC score/curve: It makes use of true positive rates (TPR) and false positive rates (FPR) and the results of ROC - AUC graph is depicted in Fig.3. TPR = (TP + TN) (TP + TN + FP + FN) FPR = (TP + TN) (TP + TN + FP + FN) B. Model Comparison The hybrid CNN+ViT model performed better than conventional CNNs and ResNet model are summarized in Table 1. The prediction results based on image classification detected by the algorithm for the dataset of welding images are shown in Fig.4. TABLE I. PERFORMANCE METRICS COMPARISON OF MODELS Model Accuracy (%) Model Complexity Suitability for Real Time Detection ResNet 88% High Moderate CNN 92% Moderate High Hybrid CNN+ViT 95% High High The performance metrics comparison table I. shows that the proposed model outperforms the conventional models with high accuracy in classification of weld bead defects for combining the advantages of CNN and transformer based techniques [18] [19]. The capacity of the hybrid model to combine ViT's global attention mechanism with CNN's localized feature extraction account for its improved performance. However the computational requirement emphasizes the necessity of hardware optimization for industrial real-time applications. It is highly effective in applications requiring both detailed local analysis and global understanding. The advantages of the model rely in enhanced feature extraction, improved accuracy as it reduces false positives and negatives by effectively distinguishing between subtle defect variations and normal weld features, context awareness in ensuring defect classification is not isolated to local features alone, scalability for a wide range of defect types and welding processes due to its adaptable architecture. In essence, hybrid deep learning models with transformers provide a powerful framework for addressing complex problems by leveraging the unique strengths of various architectures, leading to improved performance, interpretability and efficiency. CONCLUSION AND FUTURE WORK In conclusion, this study effectively created a sophisticated hybrid CNN+ViT architecture for the automated categorization of weld defects in robotic welding procedures. The hybrid model outperformed conventional CNN and ResNet models with an impressive 95% accuracy rate by utilizing both the global attention mechanism of ViT and the localized feature extraction capabilities of CNN. Its viability for real-time defect identification was proven by the implementation on the NVIDIA Jetson AGX Orin platform for deployment, guaranteeing both high precision and computational efficiency. In addition to improving the accuracy of weld quality evaluation, this method lays the groundwork for future developments in automated defect detection systems for industrial welding applications particularly in manufacturing sector. The future scope of this work involves scaling the hybrid CNN+ViT architecture for larger datasets with diverse defect classes, across various welding processes and materials. Declarations Ethical approval and consent to participate Not applicable. Consent for publications Not applicable. Competing interests The authors declare no competing interests. FUNDING Ministry of Heavy Industries, Government of India. DATA AVAILABILITY Data has been taken from publicly available source: https://github.com/sammyboi1801/welding-defect-detection Author Contribution Dr. B. Vinod contributes in Problem identification and methodologyDr. MP. Anbarasi contributes in Model identificationMr. K. Senthil Kumar contributes in Results and evaluationMs. C.Senthamilarasi contributes in Hybrid algorithm and dataset validation workMr. L. Nagalaxman contributes in Open source framework for evaluation of performance metrics Acknowledgement The authors acknowledge PSG College of Technology, Coimbatore and Ministry of Heavy Industries, for providing experimental facilities to carry out the present work. This work is done as a part of the Industry Accelerator Project - Scheme on Enhancement of Competitiveness in the Indian Capital Goods Sector- Phase-II, funded by Ministry of Heavy Industries, Government of India. References Norrish J. Evolution of Advanced Process Control in GMAW: Innovations, Implications, and Application. Welding Journal, no.103, pp. 161 – 75, 2024. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , no. 201, pp. 770–778. Islam M, Raisul MZH, Zamil ME, Rayed M, Mohsin Kabir MF, Mridha. Satoshi Nishimura, and Jungpil Shin. Deep Learning and Computer Vision Techniques for Enhanced Quality Control in Manufacturing Processes. IEEE Access, 2024. Sun J, Li C, Wu X, Palade V, Fang W. An effective method of weld defect detection and classification based on machine vision. IEEE Trans Industr Inf. no. 2019;15(12):6322–33. Dosovitskiy A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 ,2020. Krizhevsky A, Sutskever I. and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Adv Neural Inf Process Syst, no. 252012. Banachewicz K, Massaron L. The Kaggle Book: Data analysis and machine learning for competitive data science. Packt Publishing Ltd; 2022. Say D, Zidi S, Qaisar SM, Krichen M. Automated categorization of multiclass welding defects using the x-ray image augmentation and convolutional neural network. Sensors. 2023;23(14):6422. Mahajan A, Chaudhary S. Categorical image classification based on representational deep network (RESNET). 3rd International conference on Electronics, Communication and Aerospace Technology , IEEE, 2019. Kumaresan S, Aultrin KJ, Kumar SS, Anand MD. Deep learning-based weld defect classification using VGG16 transfer learning adaptive fine-tuning. Int J Interact Des Manuf (IJIDeM). 2023;17(6):2999–3010. Salah N, Hamza E, Mustapha J, and Khalifa Mansouri. Towards an automatic classification of welding defect by convolutional neural networkrobot classifier. Indonesian J Electr EngineeringComputer Sci 33, 3, 2024. Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation. Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany , Springer International Publishing, 2015. Liang J, Wang D, Ling X. Image classification for soybean and weeds based on VIT. Journal of Physics: Conference Series . Vol. 2002. No. 1. IOP Publishing, 2021. Perri S, Spagnolo F. Fabio Frustaci, and Pasquale Corsonello. Welding defects classification through a Convolutional Neural Network. Manufacturing Letters no. 35, pp. 29–32, 2023. Liu Y, et al. Lightweight ViT model for micro-expression recognition enhanced by transfer learning. Front Neurorobotics. 2022;16:922761. Ngo B, Hung et al. Learning CNN on ViT: A Hybrid Model to Explicitly Class-specific Boundaries for Domain Adaptation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024. Prashanthi SK, Kesanapalli SA, Simmhan Y. Characterizing the performance of accelerated jetson edge devices for training deep learning models. Proceedings of the ACM on Measurement and Analysis of Computing Systems , pp. 1–26, 2022. Muhammad J, Altun H, Abo-Serie E. Welding seam profiling techniques based on active vision sensing for intelligent robotic welding. Int J Adv Manuf Technol. no. 2016;88(1–4):127–45. Yang L, Wang H, Huo B, Li F, Liu Y. An automatic welding defect location algorithm based on deep learning. NDT E Int. no. 2021;120:102435. Additional Declarations No competing interests reported. Supplementary Files TABLEI.PERFORMANCEMETRICSCOMPARISONOFMODELS.docx Cite Share Download PDF Status: Published Journal Publication published 11 Jan, 2026 Read the published version in Discover Artificial Intelligence → Version 1 posted Editorial decision: Revision requested 25 Aug, 2025 Reviews received at journal 16 Aug, 2025 Reviewers agreed at journal 15 Aug, 2025 Reviewers agreed at journal 13 Aug, 2025 Reviewers agreed at journal 13 Aug, 2025 Reviews received at journal 13 Aug, 2025 Reviewers agreed at journal 05 Aug, 2025 Reviewers agreed at journal 31 Jul, 2025 Reviewers invited by journal 31 Jul, 2025 Editor assigned by journal 31 Jul, 2025 Editor invited by journal 23 Jul, 2025 Submission checks completed at journal 22 Jul, 2025 First submitted to journal 22 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7098058","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":494833682,"identity":"81743754-f9e1-4193-bfba-dbbb4b27fc98","order_by":0,"name":"B. Vinod","email":"","orcid":"","institution":"PSG College of Technology Coimbatore","correspondingAuthor":false,"prefix":"","firstName":"B.","middleName":"","lastName":"Vinod","suffix":""},{"id":494833685,"identity":"4d51e94d-01b7-47b5-a2c3-4000a08adaee","order_by":1,"name":"M. P. Anbarasi","email":"","orcid":"","institution":"PSG College of Technology Coimbatore","correspondingAuthor":false,"prefix":"","firstName":"M.","middleName":"P.","lastName":"Anbarasi","suffix":""},{"id":494833690,"identity":"15cb4848-a4b5-41a4-8020-e90d96cbe448","order_by":2,"name":"K. Senthil Kumar","email":"","orcid":"","institution":"PSG College of Technology Coimbatore","correspondingAuthor":false,"prefix":"","firstName":"K.","middleName":"Senthil","lastName":"Kumar","suffix":""},{"id":494833692,"identity":"727ac2e8-6b00-489c-b1da-e438cffd4619","order_by":3,"name":"C. Senthamilarasi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBklEQVRIiWNgGAWjYLACxgYGBjZ2MFNCDkQeeEBIy0GQFmaIFmOwlgRitDBAtDAkgtgM+LTw9x8+9vjjDrtoPmbmZx9+/LFInx92+CHQFjs53QbsWiRupKUbHDyTnNvGzGY8s7dNInfj7TQDoJZkY7MD2LUYSPCYSRxsYwZqYTBm4G0AapmdANJyIHEbLi38578BtdQDtbB/ZvzzRyLdcHb6B/xaGHLYgFoOA7XwGDPzsEkkyEvn4LcF6BczibNtx0Faipll2yQMN0jnFBxIMMDtF2CIPZOobKvOnd/evpnxzZ86efnZ6Zs/fKiwk8OlBYtTD0AcTAKQbyBF9SgYBaNgFIwEAACAi140Mk0UpAAAAABJRU5ErkJggg==","orcid":"","institution":"PSG College of Technology Coimbatore","correspondingAuthor":true,"prefix":"","firstName":"C.","middleName":"","lastName":"Senthamilarasi","suffix":""},{"id":494833697,"identity":"a6e54488-e54e-443e-8d12-e06658a370f1","order_by":4,"name":"L. Nagalaxman","email":"","orcid":"","institution":"PSG College of Technology Coimbatore","correspondingAuthor":false,"prefix":"","firstName":"L.","middleName":"","lastName":"Nagalaxman","suffix":""}],"badges":[],"createdAt":"2025-07-11 05:53:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7098058/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7098058/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s44163-025-00789-6","type":"published","date":"2026-01-11T15:57:24+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":88326306,"identity":"66eae4c3-b1e8-4cd4-9428-b8cd849841f9","added_by":"auto","created_at":"2025-08-05 09:52:27","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":101683,"visible":true,"origin":"","legend":"\u003cp\u003eClassification architecture of Vision Transformer model\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7098058/v1/2fe77acfe8ac0faa164c64a7.png"},{"id":88326309,"identity":"e4ba3ff3-7995-43ca-9720-b8b792e13e66","added_by":"auto","created_at":"2025-08-05 09:52:27","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":261710,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion Matrix for the evaluation of the model (a) CNN model (b) Hybrid CNN+ ViT model\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7098058/v1/3a7c214b5b390f670edd37c9.png"},{"id":88326311,"identity":"16e1c185-67b4-4af3-ad8e-ed8de6d89b4c","added_by":"auto","created_at":"2025-08-05 09:52:27","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":26893,"visible":true,"origin":"","legend":"\u003cp\u003eRoC-AUC graph of the proposed hybrid model\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7098058/v1/2e3afd7b12e0b32a9806b3d3.png"},{"id":88326312,"identity":"53e1377c-d203-421c-b318-37fc5e05755e","added_by":"auto","created_at":"2025-08-05 09:52:27","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":50167,"visible":true,"origin":"","legend":"\u003cp\u003ePrediction results for classification of defects\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7098058/v1/ddcf09c6ca53c8f05f2f9505.png"},{"id":100069097,"identity":"26957bfe-da70-419a-a340-c04bc879e3de","added_by":"auto","created_at":"2026-01-12 16:09:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":876410,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7098058/v1/56e7a15d-fd54-43ea-9f4f-f840eca6ec48.pdf"},{"id":88326310,"identity":"005a4b44-5a1f-48d8-927c-35e90d046f73","added_by":"auto","created_at":"2025-08-05 09:52:27","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":12211,"visible":true,"origin":"","legend":"","description":"","filename":"TABLEI.PERFORMANCEMETRICSCOMPARISONOFMODELS.docx","url":"https://assets-eu.researchsquare.com/files/rs-7098058/v1/2a078a56263007fc9a834bc9.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Hybrid Machine Learning Models for Automated Classification of Weld Defects in Gas Metal Arc Robotic Welding","fulltext":[{"header":"I.\tINTRODUCTION","content":"\u003cp\u003eWelding is a fundamental process in the manufacturing sector, playing an essential part in guaranteeing the dependability and structural soundness of different components across industries such as aerospace, construction, and automotive production line. Among different welding techniques, Gas Metal Arc Welding (GMAW) is widely adopted due to its efficiency, adaptability, and suitability for high-production environments. However, ensuring consistent weld quality remains a significant challenge, as defects can compromise the durability and safety of welded structures [1]. The detection and classification of weld bead defects are critical tasks that traditionally rely on manual visual examination as well as non-destructive testing (NDT) methods like radiography or ultrasonic testing. These conventional approaches while beneficial, are often time-consuming, labor-intensive, and heavily reliant on the knowledge of experienced operators, making them susceptible to human error and inconsistencies in quality assessment [2]. The Automated weld defect detection systems have been made possible by recent advancements in artificial intelligence (AI), specifically in the fields of Machine Learning and Computer Vision [3]. These AI-driven solutions offer substantial improvements in accuracy, speed, and repeatability, making them ideal for integration into modern manufacturing workflows. One kind of deep learning model that has shown great promise in defect classification is the Convolutional Neural Network (CNN) by effectively extracting local features from weld images. However, CNN-based approaches face constraints in gathering global contextual data and long-range dependencies, which are essential for accurately distinguishing complex defect patterns. On the other hand, Transformer-based models, specifically Vision Transformers (ViTs), excel in global feature modelling but often come with high computational costs, limiting their practical deployment in real-time industrial applications [4]. To address these challenges, this study proposes a hybrid AI based machine learning model that blends the advantages of CNNs and ViTs to enhance weld defect classification. The integration of such hybrid models into industrial automation systems has the potential to revolutionize quality control in welding processes, reducing time bounded on manual inspection, minimizing production downtime, and improving overall manufacturing reliability [5]. The combination of a hybrid CNN\u0026thinsp;+\u0026thinsp;ViT model designed for weld defect classification, combines the advantages of ViT with conventional CNN model architectures. The hybrid technique guarantees precise identification of defects such as lack of fusion, misalignment, burn-through, lack of penetration, contamination (lack of shielding gas) by the integration of ML techniques. The proposed model not only demonstrates the hybrid model effectiveness but also its potential for industrial adoption by comparing it to conventional models such as ResNet and conventional CNN based models [6]. The proposed system utilizes the machine learning model for vision based inspection in automated robotic welding environment by providing high quality inspection of the weld beads and add to the expanding corpus of knowledge on intelligent manufacturing robotic systems. This research project uses the following structure: methodology is presented in Section 2, the performance metrics evaluation of the proposed approach and its outcomes is discussed in Section 3 and the work's conclusion and future scope is outlined in Section 4.\u003c/p\u003e"},{"header":"II.\tMETHODOLOGY","content":"\u003cp\u003eThis research focuses on automating the classification of weld defects using a hybrid CNN\u0026thinsp;+\u0026thinsp;ViT model. The methodology involves of a robust dataset and a hybrid model architecture tailored for weld defect classification. Over 20,000 weld images across six defect classes were collected, annotated, and systematically split for training, validation, and testing. This division was carefully designed to ensure model performance optimization and unbiased evaluation. Leveraging the strengths of CNNs for localized feature extraction and ViTs for capturing global dependencies, the hybrid architecture addressed the limitations of standalone models. By implementing the hybrid model, this work achieved significant improvements in weld defect classification accuracy and robustness, particularly in real-world industrial settings.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eA. Datasets\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset consist of weld images from six classes of defect identification containing good welds lack of fusion, misalignment, burn-through, lack of penetration, contamination (lack of shielding gas), were collected from an open source platform for this work. The images were annotated to guarantee high-quality labelling in order to provide robust training and evaluation of machine learning models, meticulously relating to particular defect categories to implement the model in real welding environments [7]. The dataset was divided into three subsets: 80% for training, 10% for validation, and 10% for testing.in order to avoid overfitting and maximize model performance. This division guarantees sufficient training samples while preserving sufficient data for objective assessment and hyperparameters adjustment.\u003c/p\u003e\n\u003cp\u003eB. Data Pre-processing\u003c/p\u003e\n\u003cp\u003eData Preprocessing is the fundamental step to make the data clean, structured and ready for analysis, forming the foundation for accurate machine learning models. This process addresses issues like missing values, inconsistencies, and noise in the data, ensuring that machine learning models receive accurate and relevant input for better performance and accuracy. Prior to model training, the following pre-processing steps were applied for data pre-processing.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eImage Normalization\u003c/strong\u003e: In order to maintain numerical stability throughout training, pixel intensities were scaled between 0 and 1\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eData Augmentation\u003c/strong\u003e: The dataset size had been artificially increased using methods including random rotations, flips, scaling, and contrast modifications, which improved the model's capacity to generalize across unknown data\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eResizing\u003c/strong\u003e: All images were resized to a uniform resolution of 224x224 pixels, consistent with the CNN and ViT implementation\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eC. Dataset Characteristics\u003c/p\u003e\n\u003cp\u003ePotential imbalances were found by looking at the class distribution, techniques like oversampling or class-weighting were used to mitigate the imbalance created. By including a variety of fault types, the dataset is guaranteed to accurately reflect actual welding conditions, allowing the trained model to be used to real-world situations in industrial settings. The work is based on the rich dataset, a robust pre-processing pipeline, and a well-structured train-test split, which allow the hybrid CNN\u0026thinsp;+\u0026thinsp;ViT model to attain state-of-the-art performance in weld defect classification.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e1) Defect Classification\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe weld defect classification using machine learning techniques can automate the inspection process, improve the efficiency of quality control, and reduce the need for manual inspection. For the proposed defect classification, three models were utilized to perform weld bead defects classification, the significance of these models are discussed with its classification performance.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e2)\u0026nbsp;ResNet-50 model\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eResNet-50 was used as the foundation for feature extraction in this work because of its performance in object classification challenges. ResNet-50, a 50-layer deep convolutional neural network, was utilized to use residual connections to address the vanishing gradient issue [8]. The network may learn residual mappings instead of direct mappings, which makes it easier to train deeper patterns. By utilizing its stacked convolutional layers, batch normalization, ReLU activation function, and identity shortcut connections, the model was able to extract patterns and local spatial characteristics from defect images. It performs from low-level edges to high-level semantic characteristics, this model architecture excels at capturing hierarchical representations, allowing the model to precisely detect and categorize complex weld defect patterns [8].\u003c/p\u003e\n\u003cp\u003eDespite its efficacy, ResNet-50 was not immediately chosen as the implementation model because of some experimental constraints, like requirements of considerable labelled datasets, overfitting, computational resource demands. The model performed quite well in feature extraction, it occasionally had trouble in maintaining consistent performance of processing weld defect images with minute class changes. The model becomes more complex as a result of the skip connections, which may result in increased processing and memory needs. These restrictions spurred research into sophisticated hybrid designs with increased classification robustness and accuracy [9].\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e3) Conventional Neural Network model\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe objective of classifying weld defects was first accomplished using a traditional Convolutional Neural Network (CNN) model. It includes multiple convolutional layers used in the architecture to extract features, and pooling layers were added to minimize spatial dimensions. To gradually acquire hierarchical representations of the input weld defect images from low- level edges to high-level patterns these layers were stacked and after that the retrieved features were classified into six different defect types using fully linked layers. The CNN uses convolution operations to efficiently learn spatial hierarchies and local dependencies [10]. During testing, the model exhibited encouraging preliminary results and performed satisfactorily in capturing local aspects of weld defects [11]. However, it shows certain limitations and constraints in catching long-range dependencies and global context, which are essential for predicting minute variations between related fault classes. In the case of noisy and complicated grayscale images, the CNN occasionally performed poor in generalization, producing predictions that were inconsistent. These difficulties can be resolved with a hybrid model architecture, which led to the incorporation of vision transformers into a hybrid framework which greatly enhanced the classification performance.\u003c/p\u003e\n\u003cp\u003e4)\u0026nbsp;Proposed Hybrid CNN\u0026thinsp;+\u0026thinsp;ViT model\u003c/p\u003e\n\u003cp\u003eA hybrid architecture proposed in this work combines the advantages of both standalone models by incorporating ViT into the CNN framework. The efficient weld defect classification is achieved by the CNN\u0026thinsp;+\u0026thinsp;ViT hybrid model as it combines the advantages of ViT and CNN. The CNN layers use patterns, textures, and edges to extract localized features. The final feature map is used as the input for the ViT module after these features have been processed through convolutional, pooling and batch normalizing layers. Global attention-based processing is made possible by transforming the feature map and passing it to the ViT in place of the conventional dense layers in CNN.\u003c/p\u003e\n\u003cp\u003eA sequence is created by linearly embedding the patches (such as 16*16) into vectors from the CNN feature map. To preserve spatial information is crucial for comprehending the interactions between patches, positional encoding is added. Transformer encoder layers, which include Feed Forward Neural networks (FFN) and Multi-Head Self Attention (MHSA), can process this sequence. While FFN refines features by non-linear transformations, backed by residual connections and stability preserving normalization, MHSA captures global dependencies. A classification head processes the enhanced feature representation and divides the image into six defect classes using a fully linked layer with SoftMax activation. By combining localized and global feature extraction, our hybrid technique efficiently addresses the complexity of weld defect patterns while achieving higher accuracy [12].\u003c/p\u003e\n\u003cp\u003e5)\u0026nbsp;Advantages of Hybrid model\u003c/p\u003e\n\u003cp\u003eThe hybrid proposed model overcome the shortcomings of conventional CNN models by utilizing ViT in collecting long-range relationships and global context within images. The CNN shows difficulty in modelling the general relationships between various parts of an image, even when they are quite good at learning local characteristics through convolutional process. This makes it difficult to precisely identify tiny changes and structural differences in the entire image. ViT is a promising addition to the model architecture because of its self-attention mechanism, which excels at analysing global dependencies and comprehending the image holistically [13]. The classification of weld defects significantly improved with the implementation of the hybrid model with CNN\u0026thinsp;+\u0026thinsp;ViT architecture, thereby fusing the ViT in global reasoning with CNN to extract fine-grained local features. Better generalization resulted from this hybrid model for feature learning, especially when handling noisy and complicated images. All six defect classes showed more consistent predictions, also outperformed better in terms of accuracy. It shows enhanced flexibility in responding to changes in the sizes, forms and orientations of defects, which makes it a best model for automated GMAW robotic welding applications. Overall, the hybrid CNN\u0026thinsp;+\u0026thinsp;ViT model offered a more precise and efficient classification framework by successfully addressing the drawbacks of standalone models for defect classification in real time environments [14].\u003c/p\u003e\n\u003cp\u003eD.\u0026nbsp; Implementation of hybrid model\u003c/p\u003e\n\u003cp\u003eThe open source PyTorch deep learning framework was chosen for the proposed model execution due to its adaptability and effective support for neural network building. A Jetson AGX Orin GPU, which offers excellent edge computing performance is used for training and testing. Particularly for the hybrid CNN+ViT architecture, the GPU's capabilities guaranteed seamless handling of computationally demanding tasks, such as model training and inference [15]. This type of configuration facilitated to process real time data and it eases the simulation process to effectively investigate by fusing distinct architectures and tuning of hyperparameters. When the model were trained with 50 epochs to prevent overfitting an early termination methods were used. To guarantee effective memory use and steady gradient updates batch size was chosen, and to strike a \u003cbr /\u003e balance between convergence speed and stability, the\u003cbr /\u003e learning rate was adjusted. Dropout was used to randomly turn off neurons during training for overfitting reduction. The cross-validation was applied to ensure an all-round evaluation to lower the impact of dataset splits. To validate that the consistency was maintained across multiple folds, following the dataset's division into three subsets of training, testing and validation k-fold cross validation had been used. The multiclass classification problem was addressed along with the Adam optimizer during training by using a categorical cross-entropy loss function due to its adaptive learning rates and fast convergence [16]. Rotation, flipping, and contrast adjustment are examples of data augmentation techniques that were used to make the dataset more diverse and to make use of the model's robust design. This mixing philosophy showed the model outperforms well for industrial applications by ensuring reliable and accurate predictions in real time environments [17].\u003c/p\u003e"},{"header":"III.\tPERFORMANCE METRICS EVALUATION","content":"\u003cp\u003e\u003cem\u003eA. \u0026nbsp; \u0026nbsp;\u003c/em\u003e\u003cem\u003ePerformance Metrics Comparison\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eTo effectively compare the common metrics used to assess the deep learning models performance in classification tasks include accuracy, precision, ROC AUC graphs are used. Metrics commonly used to evaluate a hybrid deep learning model are shown as a confusion matrix that contrasts the actual results with the predictions made by the model. It divides the predictions from the four categories such as True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN). The performance evaluation of CNN model and hybrid CNN+ViT model are represented by confusion matrix is shown in Fig. 2\u003c/p\u003e\n\u003cp\u003e1. \u0026nbsp; \u0026nbsp;Accuracy: \u0026nbsp;It represents the percentage of cases that were correctly classified out of all instances\u003c/p\u003e\n\u003cp\u003eAccuracy = \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003cu\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;(TP + TN) \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; (TP + TN + FP + FN)\u003c/p\u003e\n\u003cp\u003e2. \u0026nbsp; \u0026nbsp;AUC-ROC score/curve: It makes use of true positive rates (TPR) and false positive rates (FPR) and the results of ROC - AUC graph is depicted in Fig.3. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTPR = \u0026nbsp; \u0026nbsp; \u0026nbsp; \u003cu\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;(TP + TN) \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; (TP + TN + FP + FN)\u003c/p\u003e\n\u003cp\u003eFPR = \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003cu\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;(TP + TN) \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;(TP + TN + FP + FN)\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eB. \u0026nbsp;\u003c/em\u003e\u003cem\u003eModel Comparison\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe hybrid CNN+ViT model performed better than conventional CNNs and ResNet model are summarized in Table 1. The prediction results based on image classification detected by the algorithm for the dataset of welding images are shown in Fig.4.\u003c/p\u003e\n\u003cp\u003eTABLE I. \u0026nbsp;PERFORMANCE METRICS COMPARISON OF MODELS\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eModel Complexity\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSuitability for Real Time Detection\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eResNet\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e88%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eModerate\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCNN\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e92%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eModerate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eHybrid CNN+ViT\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e95%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHigh\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe performance metrics comparison table I. shows that the proposed model outperforms the conventional models with high accuracy in classification of weld bead defects for \u0026nbsp;combining the advantages of CNN and transformer based techniques [18] [19]. The capacity of the hybrid model to combine ViT\u0026apos;s global attention mechanism with CNN\u0026apos;s localized feature extraction account for its improved performance. However the computational requirement emphasizes the necessity of hardware optimization for industrial real-time applications. \u0026nbsp; It is highly effective in applications requiring both detailed local analysis and global understanding. The advantages of the model rely in enhanced feature extraction, \u003cstrong\u003eimproved accuracy\u003c/strong\u003e as it reduces false positives and negatives by effectively distinguishing between subtle defect variations and normal weld features, context \u0026nbsp;awareness in ensuring defect classification is not \u0026nbsp;isolated to local features alone, scalability for a wide range of defect types and welding processes due to its adaptable architecture. In essence, hybrid deep learning models with transformers provide a powerful framework for addressing complex problems by leveraging the unique strengths of various architectures, leading to improved performance, interpretability and efficiency.\u003c/p\u003e"},{"header":"CONCLUSION AND FUTURE WORK","content":"\u003cp\u003eIn conclusion, this study effectively created a sophisticated hybrid CNN+ViT architecture for the automated categorization of weld defects in robotic welding procedures. The hybrid model outperformed conventional CNN and ResNet models with an impressive 95% accuracy rate by utilizing both the global attention mechanism of ViT and the localized feature extraction capabilities of CNN. Its viability for real-time defect identification was proven by the implementation on the NVIDIA Jetson AGX Orin platform for deployment, guaranteeing both high precision and computational efficiency. In addition to improving the accuracy of weld quality evaluation, this method lays the groundwork for future developments in automated defect detection systems for industrial welding applications particularly in manufacturing sector. The future scope of this work involves scaling the hybrid CNN+ViT architecture for larger datasets with diverse defect classes, across various welding processes and materials.\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003ch2\u003eEthical approval and consent to participate\u003c/h2\u003e\u003cp\u003eNot applicable.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConsent for publications\u003c/strong\u003e\u003cp\u003eNot applicable.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFUNDING\u003c/h2\u003e\u003cp\u003eMinistry of Heavy Industries, Government of India.\u003c/p\u003e\u003cp\u003eDATA AVAILABILITY\u003c/p\u003e\u003cp\u003eData has been taken from publicly available source: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/sammyboi1801/welding-defect-detection\u003c/span\u003e\u003cspan address=\"https://github.com/sammyboi1801/welding-defect-detection\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eDr. B. Vinod contributes in Problem identification and methodologyDr. MP. Anbarasi contributes in Model identificationMr. K. Senthil Kumar contributes in Results and evaluationMs. C.Senthamilarasi contributes in Hybrid algorithm and dataset validation workMr. L. Nagalaxman contributes in Open source framework for evaluation of performance metrics\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe authors acknowledge PSG College of Technology, Coimbatore and Ministry of Heavy Industries, for providing experimental facilities to carry out the present work. This work is done as a part of the Industry Accelerator Project - Scheme on Enhancement of Competitiveness in the Indian Capital Goods Sector- Phase-II, funded by Ministry of Heavy Industries, Government of India.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eNorrish J. Evolution of Advanced Process Control in GMAW: Innovations, Implications, and Application. Welding Journal, no.103, pp. 161\u0026thinsp;\u0026ndash;\u0026thinsp;75, 2024.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHe K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In \u003cem\u003eProceedings of the IEEE conference on computer vision and pattern recognition\u003c/em\u003e, no. 201, pp. 770\u0026ndash;778.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIslam M, Raisul MZH, Zamil ME, Rayed M, Mohsin Kabir MF, Mridha. Satoshi Nishimura, and Jungpil Shin. Deep Learning and Computer Vision Techniques for Enhanced Quality Control in Manufacturing Processes. IEEE Access, 2024.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSun J, Li C, Wu X, Palade V, Fang W. An effective method of weld defect detection and classification based on machine vision. IEEE Trans Industr Inf. no. 2019;15(12):6322\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDosovitskiy A. An image is worth 16x16 words: Transformers for image recognition at scale. \u003cem\u003earXiv preprint arXiv:2010.11929\u003c/em\u003e,2020.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKrizhevsky A, Sutskever I. and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Adv Neural Inf Process Syst, no. 252012.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBanachewicz K, Massaron L. The Kaggle Book: Data analysis and machine learning for competitive data science. Packt Publishing Ltd; 2022.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSay D, Zidi S, Qaisar SM, Krichen M. Automated categorization of multiclass welding defects using the x-ray image augmentation and convolutional neural network. Sensors. 2023;23(14):6422.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMahajan A, Chaudhary S. Categorical image classification based on representational deep network (RESNET). \u003cem\u003e3rd International conference on Electronics, Communication and Aerospace Technology\u003c/em\u003e, IEEE, 2019.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKumaresan S, Aultrin KJ, Kumar SS, Anand MD. Deep learning-based weld defect classification using VGG16 transfer learning adaptive fine-tuning. Int J Interact Des Manuf (IJIDeM). 2023;17(6):2999\u0026ndash;3010.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSalah N, Hamza E, Mustapha J, and Khalifa Mansouri. Towards an automatic classification of welding defect by convolutional neural networkrobot classifier. Indonesian J Electr EngineeringComputer Sci 33, 3, 2024.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRonneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation. \u003cem\u003eMedical image computing and computer-assisted intervention\u0026ndash;MICCAI 2015: 18th international conference, Munich, Germany\u003c/em\u003e, Springer International Publishing, 2015.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiang J, Wang D, Ling X. Image classification for soybean and weeds based on VIT. \u003cem\u003eJournal of Physics: Conference Series\u003c/em\u003e. Vol. 2002. No. 1. IOP Publishing, 2021.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePerri S, Spagnolo F. Fabio Frustaci, and Pasquale Corsonello. Welding defects classification through a Convolutional Neural Network. \u003cem\u003eManufacturing Letters\u003c/em\u003e no. 35, pp. 29\u0026ndash;32, 2023.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu Y, et al. Lightweight ViT model for micro-expression recognition enhanced by transfer learning. Front Neurorobotics. 2022;16:922761.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNgo B, Hung et al. Learning CNN on ViT: A Hybrid Model to Explicitly Class-specific Boundaries for Domain Adaptation. \u003cem\u003eProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition\u003c/em\u003e, 2024.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePrashanthi SK, Kesanapalli SA, Simmhan Y. Characterizing the performance of accelerated jetson edge devices for training deep learning models. \u003cem\u003eProceedings of the ACM on Measurement and Analysis of Computing Systems\u003c/em\u003e, pp. 1\u0026ndash;26, 2022.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMuhammad J, Altun H, Abo-Serie E. Welding seam profiling techniques based on active vision sensing for intelligent robotic welding. Int J Adv Manuf Technol. no. 2016;88(1\u0026ndash;4):127\u0026ndash;45.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang L, Wang H, Huo B, Li F, Liu Y. An automatic welding defect location algorithm based on deep learning. NDT E Int. no. 2021;120:102435.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"discover-artificial-intelligence","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"diai","sideBox":"Learn more about [Discover Artificial Intelligence](https://www.springer.com/44163)","snPcode":"","submissionUrl":"","title":"Discover Artificial Intelligence","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Welding Robot, Hybrid CNN + ViT Model, Deep Learning, Vision Transformer, Convolutional Neural Network, Industrial Automation","lastPublishedDoi":"10.21203/rs.3.rs-7098058/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7098058/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe research study explores an advanced approach to classify weld bead defects in a Gas Metal Arc Welding (GMAW) process of robotic welding, contributing to the growing field of intelligent manufacturing and quality control in industrial automation. To precisely categorize six classes of weld bead such as lack of fusion, burn-through, misalignment, lack of penetration, contamination, and good weld, a hybrid machine learning approach that combines Convolutional Neural Networks (CNN) and Vision Transformers (ViT) is utilized. The hybrid model leverages the complementary advantages of CNN for the extraction of localized features and ViT for global contextual awareness, resulting in superior classification performance compared to traditional architectures such as ResNet and conventional CNN models. The performance evaluation of proposed model demonstrates the model's robustness and reliability in diverse operational scenarios. The findings highlight the potential of integrating hybrid deep learning models into industrial automation systems to enhance weld bead defect detection to reduce operational inefficiencies and ensure consistent weld quality in manufacturing process.\u003c/p\u003e","manuscriptTitle":"Hybrid Machine Learning Models for Automated Classification of Weld Defects in Gas Metal Arc Robotic Welding","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-05 09:52:22","doi":"10.21203/rs.3.rs-7098058/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-08-25T15:29:22+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-16T20:35:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"28620585565594645940904790189755568941","date":"2025-08-15T15:27:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"19113884426172116009618167554540022641","date":"2025-08-13T19:20:51+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"68823220635356534894725355667220142802","date":"2025-08-13T16:05:18+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-13T09:10:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"312191326196521965794805642214523783144","date":"2025-08-06T02:11:55+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"279084091349829358190040583610304548657","date":"2025-07-31T14:52:01+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-07-31T13:49:18+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-07-31T13:47:50+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-07-23T06:10:51+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-07-22T05:33:18+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Artificial Intelligence","date":"2025-07-22T05:30:58+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"discover-artificial-intelligence","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"diai","sideBox":"Learn more about [Discover Artificial Intelligence](https://www.springer.com/44163)","snPcode":"","submissionUrl":"","title":"Discover Artificial Intelligence","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"6b93d7d6-7ff1-4c0d-97bd-400e24d54c00","owner":[],"postedDate":"August 5th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2026-01-12T16:01:06+00:00","versionOfRecord":{"articleIdentity":"rs-7098058","link":"https://doi.org/10.1007/s44163-025-00789-6","journal":{"identity":"discover-artificial-intelligence","isVorOnly":false,"title":"Discover Artificial Intelligence"},"publishedOn":"2026-01-11 15:57:24","publishedOnDateReadable":"January 11th, 2026"},"versionCreatedAt":"2025-08-05 09:52:22","video":"","vorDoi":"10.1007/s44163-025-00789-6","vorDoiUrl":"https://doi.org/10.1007/s44163-025-00789-6","workflowStages":[]},"version":"v1","identity":"rs-7098058","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7098058","identity":"rs-7098058","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-4.0