A Micro-Surface Concrete Image Dataset for Crack Detection and Texture Classification Using Deep Learning

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract This study presents a novel concrete micro-surface image dataset designed for both surface classification and micro-crack detection tasks using deep learning. The dataset was constructed under controlled laboratory conditions by preparing concrete samples with three different water-to-cement ratios (0.35, 0.45, and 0.55), resulting in dense, intermediate, and porous surface structures. A total of 900 images were collected for multi-class classification, with 300 images per class, ensuring a balanced distribution. In addition, a binary subset consisting of 300 images (150 crack and 150 healthy) was created to specifically address micro-crack detection within dense surfaces. To eliminate potential bias, crack annotations were restricted to top-surface images, allowing the model to focus on intrinsic crack features. The usability of the dataset was validated using a pretrained convolutional neural network based on GoogLeNet. Experimental results show that the binary classification task achieved an accuracy of 95.45% with an F1-score of 0.95, while the three-class classification task achieved an overall accuracy of 88.15% with a mean F1-score of 0.876. The results indicate that dense and porous classes are highly distinguishable, whereas intermediate surfaces present a more challenging classification problem due to their transitional characteristics. The proposed dataset provides a reliable and structured benchmark for future research in image-based concrete analysis, defect detection, and material characterization using deep learning.
Full text 89,862 characters · extracted from preprint-html · click to expand
A Micro-Surface Concrete Image Dataset for Crack Detection and Texture Classification Using Deep Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Micro-Surface Concrete Image Dataset for Crack Detection and Texture Classification Using Deep Learning Hidir Selcuk Nogay This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9594679/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract This study presents a novel concrete micro-surface image dataset designed for both surface classification and micro-crack detection tasks using deep learning. The dataset was constructed under controlled laboratory conditions by preparing concrete samples with three different water-to-cement ratios (0.35, 0.45, and 0.55), resulting in dense, intermediate, and porous surface structures. A total of 900 images were collected for multi-class classification, with 300 images per class, ensuring a balanced distribution. In addition, a binary subset consisting of 300 images (150 crack and 150 healthy) was created to specifically address micro-crack detection within dense surfaces. To eliminate potential bias, crack annotations were restricted to top-surface images, allowing the model to focus on intrinsic crack features. The usability of the dataset was validated using a pretrained convolutional neural network based on GoogLeNet. Experimental results show that the binary classification task achieved an accuracy of 95.45% with an F1-score of 0.95, while the three-class classification task achieved an overall accuracy of 88.15% with a mean F1-score of 0.876. The results indicate that dense and porous classes are highly distinguishable, whereas intermediate surfaces present a more challenging classification problem due to their transitional characteristics. The proposed dataset provides a reliable and structured benchmark for future research in image-based concrete analysis, defect detection, and material characterization using deep learning. Concrete dataset Micro-surface analysis Crack detection Deep learning Image classification Convolutional neural networks Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Recent advances in deep learning and computer vision have significantly improved automated inspection systems in civil engineering and material science (Ye et al., 2019). In particular, image-based analysis has emerged as a powerful tool for identifying surface characteristics and defects in construction materials such as concrete (Hsieh & Tsai, 2020). Due to its widespread use, ensuring the structural integrity and durability of concrete is of critical importance (Tang et al., 2015). Traditional inspection methods rely heavily on manual evaluation, which is time-consuming and prone to subjective interpretation (Koch et al., 2015). Moreover, micro-level features such as fine cracks and subtle texture variations are often difficult to detect without specialized imaging tools (Yao et al., 2014). The integration of digital microscopy and deep learning techniques enables more accurate and consistent detection of such features (Midtvedt et al., 2021). Despite the growing interest in this area, the availability of well-structured and labeled datasets remains limited (Yuan et al., 2024). Existing datasets typically focus on macroscopic defects, whereas micro-surface characteristics—such as texture transitions between dense, intermediate, and porous structures—are less explored (Munawar et al., 2021). Additionally, datasets that combine both surface classification and micro-crack detection within a unified framework are scarce (Hoskere et al., 2020). To address these limitations, this study introduces a novel concrete micro-surface image dataset constructed under controlled laboratory conditions. The overall workflow of dataset creation, including sample preparation, imaging, and labeling, is illustrated in Fig. 1. Concrete samples were prepared using three different water-to-cement ratios (0.35, 0.45, and 0.55), resulting in three distinct surface categories: dense, intermediate, and porous. Representative examples of these surface types are presented in Fig. 2, highlighting the visual differences between compact, transitional, and porous structures. A total of 900 images were collected for the three main classes, with 300 images per class, as summarized in Table 1. In addition to surface classification, a binary classification subset was constructed for micro-crack detection within the dense class. Example images of crack and healthy samples are shown in Fig. 3. To avoid surface-induced bias, only top-surface images were used for crack labeling, ensuring that the model learns intrinsic crack features rather than positional differences. To validate the effectiveness of the dataset, deep learning-based experiments were conducted. The binary classification task achieved an accuracy of 95.45% with a high F1-score, as shown in Fig. 4 and Table 4, demonstrating clear separability between crack and healthy samples. Furthermore, a three-class classification experiment (dense, intermediate, porous) achieved an overall accuracy of 88.15%. The corresponding confusion matrix is presented in Fig. 5. While dense and porous classes were classified with high accuracy, the intermediate class showed relatively lower performance due to its transitional nature. This observation is further supported by class-based F1-scores reported in Table 4. Training dynamics of the model are illustrated in Fig. 6 and Fig. 7, where both accuracy and loss curves indicate stable convergence without significant overfitting. The main contributions of this study are summarized as follows: (i) the creation of a novel and balanced concrete micro-surface dataset, (ii) the inclusion of both texture classification and micro-crack detection tasks, (iii) the reduction of surface bias through controlled data acquisition, and (iv) the validation of dataset usability through deep learning experiments. The proposed dataset provides a valuable benchmark for future research in automated concrete inspection and image-based material analysis. 2. Related Works Recent studies have demonstrated the effectiveness of deep learning in concrete surface inspection and defect detection (McLaughlin et al., 2020 ; Dung, 2019 ). Convolutional neural networks (CNNs) have been widely used for crack detection, classification of surface defects, and structural health monitoring (Dorafshan et al., 2018 ; Sony et al., 2021 ). Several publicly available datasets have contributed to this field, primarily focusing on macroscopic crack detection and damage classification (Kulkarni et al., 2023 ; Kim et al., 2019 ). These datasets typically contain large-scale surface defects and are captured under uncontrolled environmental conditions. While such datasets are valuable, they often lack detailed microstructural information and controlled acquisition settings. More recent works have explored the use of high-resolution imaging and microscopy for analyzing material surfaces at a finer scale (Zhou et al., 2019 ; Bangaru et al., 2022 ). These approaches provide more detailed insights into texture variations and micro-level defects. However, datasets that combine both microstructural surface classification and crack detection within a unified framework remain limited (Dong et al., 2020 ). In addition, the influence of material composition, particularly the water-to-cement ratio, on surface texture has been investigated in several studies (Chuta et al., 2020 ). These studies highlight the importance of capturing controlled variations in material properties to improve classification performance (Guo et al., 2022 ). Despite these advancements, there is still a need for datasets that (i) provide balanced class distributions, (ii) include both macro and micro-level features, and (iii) are collected under controlled experimental conditions (Chakurkar et al., 2023 ). The dataset proposed in this study aims to address these gaps by offering a comprehensive and structured benchmark for both surface classification and crack detection tasks. 3. Dataset Description In this study, a novel concrete micro-surface image dataset was developed under controlled laboratory conditions to support both surface classification and micro-crack detection tasks. 3.1 Sample Preparation Concrete samples were prepared using three different water-to-cement (w/c) ratios: 0.35, 0.45, and 0.55. These ratios were selected to produce distinct microstructural characteristics corresponding to dense, intermediate, and porous surface types, respectively (Elsharief et al., 2003 ). The variation in w/c ratio directly affects the internal structure and surface compactness of concrete, making it a suitable parameter for controlled dataset generation. The overall data acquisition workflow is illustrated in Fig. 1 . 3.2 Image Acquisition All images were captured using a digital microscope under consistent environmental and lighting conditions to ensure data uniformity. For each sample, both top and bottom surface images were collected. However, in order to prevent bias in crack detection, only top surface images were used for crack labeling (Rashid et al., 2025 ). Example images from each class (dense, intermediate, porous) are presented in Fig. 2 . 3.3 Dataset Structure The dataset is organized into two complementary subsets designed to support both multi-class surface classification and binary crack detection tasks. The first subset focuses on multi-class surface classification and comprises three distinct categories: dense, intermediate, and porous. A total of 900 images were collected, with 300 images allocated to each class, resulting in a fully balanced dataset. Maintaining equal class distributions ensures fair model evaluation and prevents potential bias toward any specific surface type (Johnson & Khoshgoftaar, 2019 ). The distribution of images across the three classes is summarized in Table 1 . Table 1 Distribution of images across the three main surface classes in the dataset. Class Number of Images Dense 300 Intermediate 300 Porous 300 Total 900 The second subset was specifically constructed to address micro-crack detection within the dense concrete class. This binary dataset includes two categories: dense crack and dense healthy, with 150 images per class, yielding a total of 300 images. To ensure robust learning and eliminate potential acquisition-related bias, crack annotations were restricted exclusively to top-surface images (Geirhos et al., 2020 ). This design prevents the model from relying on irrelevant cues such as surface orientation or positional differences. Example crack and healthy samples used in the binary classification task are illustrated in Fig. 3 , while the detailed class distribution is presented in Table 2 . Table 2 Distribution of crack and healthy samples within the dense surface class. Class Number of Images Crack 150 Healthy 150 Total 300 3.4 Naming Convention and Organization All images in the dataset were systematically labeled and organized using a structured naming convention designed to encode essential attributes, including material type, water-to-cement ratio, surface location (top or bottom), and condition (crack or healthy). This standardized scheme ensures traceability of each sample and facilitates efficient dataset navigation during both training and evaluation processes. The adopted naming format integrates multiple descriptive components within a single filename (e.g., dense035top_crack_XXX ), enabling direct identification of class category and acquisition conditions without requiring additional metadata files. Such explicit encoding enhances data transparency and reduces the risk of mislabeling during preprocessing stages. A consistent folder hierarchy was also implemented to separate classes and experimental subsets systematically. This structured organization improves dataset reproducibility, supports scalable expansion, and aligns with established best practices in research data management (Wilkinson et al., 2016 ). The detailed folder structure and naming schema are summarized in Table 3 . Table 3 File naming convention and dataset organization structure. Folder Name Description dense035top_crack Dense crack images dense035top_healthy Dense healthy images intermediate045top Intermediate images porous055top Porous images 3.5 Dataset Characteristics The proposed dataset possesses several distinctive characteristics that enhance its methodological robustness and practical relevance. First, the dataset maintains a balanced class distribution, ensuring unbiased performance evaluation across categories. Second, all images were acquired under controlled laboratory conditions, minimizing environmental variability and improving data consistency. In addition, the dataset integrates both surface texture classification and micro-crack detection within a unified framework, enabling multi-level analysis of concrete microstructures. The deliberate separation of crack labeling from surface orientation further eliminates potential acquisition-induced bias, strengthening the reliability of model learning. Moreover, the dataset is designed for multi-purpose usability, supporting both binary and multi-class classification scenarios. These combined attributes establish the dataset as a structured and challenging benchmark for evaluating deep learning models in concrete surface analysis and defect detection tasks. 4. Experimental Setup To assess the practical usability and learning capacity of the proposed dataset, deep learning-based classification experiments were conducted using a pretrained convolutional neural network architecture. The experimental framework was designed to evaluate both binary crack detection and multi-class surface classification scenarios under consistent training conditions. 4.1 Model Selection The GoogLeNet architecture was selected as the baseline model due to its demonstrated effectiveness in image classification tasks and its computationally efficient design (Szegedy et al., 2015 ). The network was adapted through a transfer learning strategy, in which the final fully connected layers were modified to match the number of output classes in each experimental setting (Pan & Yang, 2009 ). This approach enables efficient feature reuse while maintaining adaptability to the specific classification tasks defined in the dataset. 4.2 Data Splitting To ensure an unbiased evaluation of model performance, the dataset was randomly partitioned into three mutually exclusive subsets: 70% for training, 15% for validation, and 15% for testing. This split ratio was consistently applied to both binary and multi-class classification tasks to maintain comparability across experiments. The separation of validation and testing data prevents information leakage and provides a reliable assessment of the model’s generalization capability. 4.3 Training Configuration The network was trained using the stochastic gradient descent with momentum (SGDM) optimizer, a widely adopted optimization strategy for deep neural networks due to its stability and convergence efficiency (Gitman et al., 2019 ). The training process was configured with an initial learning rate of 0.0001 and a maximum of 20 epochs. A mini-batch size of 100 was selected to balance computational efficiency and gradient stability, while validation was performed at regular intervals of five iterations to monitor learning progression and detect potential overfitting during training. Cross-entropy loss was employed as the objective function, as it is well-suited for supervised classification tasks involving probabilistic output distributions. 4.4 Evaluation Metrics Model performance was assessed using standard classification metrics, including accuracy, precision, recall, and F1-score (Hossin & Sulaiman, 2015 ). These metrics collectively provide a comprehensive evaluation of classification behavior by considering both overall correctness and class-specific prediction quality. For the multi-class classification task, class-wise performance metrics were computed, and macro-average values were reported to ensure balanced evaluation across all classes (Grandini et al., 2020 ). The macro-averaging approach assigns equal importance to each class, preventing dominant categories from disproportionately influencing the overall performance assessment. 5. Results and Discussion 5.1 Binary Classification Results The binary classification task was designed to distinguish between crack and healthy samples within the dense concrete class. The confusion matrix for this task is presented in Fig. 4 . The model achieved an accuracy of 95.45%, indicating a high level of separability between crack and non-crack samples. The precision and recall values further demonstrate the robustness of the model, with a precision of 1.00 and a recall of 0.91. The resulting F1-score of 0.95 confirms the reliability of the classification performance. These results suggest that micro-cracks captured using digital microscopy exhibit distinct visual patterns that can be effectively learned by deep learning models. 5.2 Multi-Class Classification Results The multi-class classification task involved distinguishing between dense, intermediate, and porous surface types. The corresponding confusion matrix is shown in Fig. 5 . The model achieved an overall accuracy of 88.15%, demonstrating strong performance in differentiating between concrete surface types. Class-wise analysis reveals that dense and porous classes were classified with high accuracy, while the intermediate class exhibited relatively lower performance. The F1-scores for dense, intermediate, and porous classes were 0.92, 0.79, and 0.92, respectively, with a mean F1-score of 0.88. The lower performance in the intermediate class can be attributed to its transitional nature, as it shares characteristics with both dense and porous structures. This makes it inherently more difficult to classify (Santos et al., 2022 ). 5.3 Training Behavior The training and validation accuracy curves are presented in Fig. 6 , while the loss curves are shown in Fig. 7 . Both accuracy and loss graphs indicate stable learning behavior with no significant signs of overfitting. The validation accuracy closely follows the training accuracy, suggesting good generalization capability of the model. Table 4 Performance results of the deep learning model for binary and multi-class classification tasks. Task Class Precision Recall F1-score Binary Classification (Dense Crack vs Healthy) Crack 1.0000 0.9091 0.9524 Healthy 0.9565 1.0000 0.9778 Mean 0.9783 0.9545 0.9651 Multi-class Classification (Dense / Intermediate / Porous) Dense 0.8627 0.9778 0.9167 Intermediate 0.9394 0.6889 0.7949 Porous 0.8627 0.9778 0.9167 Mean 0.8883 0.8815 0.8761 5.4 Discussion The experimental results confirm that the proposed dataset is suitable for both binary and multi-class classification tasks. The high performance in the binary task demonstrates the clarity of crack features, while the multi-class results highlight the dataset's ability to capture meaningful variations in concrete surface textures. The observed difficulty in classifying intermediate samples reflects real-world material behavior, where transitions between structural states are not sharply defined (Hilal, 2016 ). This characteristic adds to the realism and practical value of the dataset. Overall, the dataset provides a challenging yet learnable benchmark for evaluating deep learning models in concrete surface analysis. 6. Conclusion In this study, a novel concrete micro-surface image dataset was introduced for both surface classification and micro-crack detection tasks. The dataset was constructed under controlled laboratory conditions using different water-to-cement ratios, resulting in three distinct surface categories: dense, intermediate, and porous. In addition to the multi-class dataset, a binary classification subset was created to specifically address micro-crack detection within dense concrete samples. The dataset was carefully designed to eliminate potential biases by restricting crack annotations to top-surface images. To validate the usability of the dataset, deep learning-based experiments were conducted using a pretrained GoogLeNet architecture. The binary classification task achieved high performance with an accuracy of 95.45% and an F1-score of 0.95, indicating clear separability between crack and healthy samples. The multi-class classification task achieved an overall accuracy of 88.15%, with strong performance in dense and porous classes, while the intermediate class presented a more challenging scenario due to its transitional nature. The results demonstrate that the proposed dataset provides a reliable and effective benchmark for both classification and defect detection tasks. Its balanced structure, controlled acquisition process, and inclusion of both texture and defect information make it a valuable resource for future research in concrete inspection and image-based material analysis. Future work may focus on expanding the dataset with additional environmental variations, increasing the number of samples, and exploring more advanced deep learning architectures to further improve classification performance. Declarations Competing Interests The author declares that there are no competing interests. Ethics Approval Not applicable. Funding This study was supported by Bursa Uludağ University Scientific Research Projects Unit (BAP) under Project No. 2381. Author Contribution H.S.N. conceived the study, designed the methodology, performed data acquisition and analysis, and wrote the manuscript. The author reviewed and approved the final version of the manuscript. Acknowledgement The author would like to acknowledge the support of Bursa Uludağ University Scientific Research Projects Unit (BAP) under Project No. 2381. Data Availability The dataset generated during this study has been deposited in the Mendeley Data repository and is currently under moderation. A DOI has been reserved and will become publicly accessible upon final approval by the repository. References Bangaru, S. S., Wang, C., Zhou, X., & Hassan, M. (2022). Scanning electron microscopy (SEM) image segmentation for microstructure analysis of concrete using U-net convolutional neural network. Automation in Construction, 144 , 104602. https://doi.org/10.1016/j.autcon.2022.104602 Chakurkar, P. S., Vora, D., Patil, S., Mishra, S., & Kotecha, K. (2023). Data-driven approach for AI-based crack detection: Techniques, challenges, and future scope. Frontiers in Sustainable Cities, 5 , 1253627. https://doi.org/10.3389/frsc.2023.1253627 Chuta, E., Colin, J., & Jeong, J. (2020). The impact of the water-to-cement ratio on the surface morphology of cementitious materials. Journal of Building Engineering, 32 , 101716. https://doi.org/10.1016/j.jobe.2020.101716 Dung, C. V. (2019). Autonomous concrete crack detection using deep fully convolutional neural network. Automation in Construction, 99 , 52–58. https://doi.org/10.1016/j.autcon.2018.11.028 Dong, Y., Su, C., Qiao, P., & Sun, L. (2020). Microstructural crack segmentation of three-dimensional concrete images based on deep convolutional neural networks. Construction and Building Materials, 253 , 119185. https://doi.org/10.1016/j.conbuildmat.2020.119185 Dorafshan, S., Thomas, R. J., & Maguire, M. (2018). Comparison of deep convolutional neural networks and edge detectors for image-based crack detection in concrete. Construction and Building Materials, 186 , 1031–1045. https://doi.org/10.1016/j.conbuildmat.2018.08.011 Elsharief, A., Cohen, M. D., & Olek, J. (2003). Influence of aggregate size, water cement ratio and age on the microstructure of the interfacial transition zone. Cement and Concrete Research, 33 (11), 1837–1849. https://doi.org/10.1016/S0008-8846(03)00205-9 Geirhos, R., Jacobsen, J. H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., & Wichmann, F. A. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence, 2 (11), 665–673. https://doi.org/10.1038/s42256-020-00257-z Gitman, I., Lang, H., Zhang, P., & Xiao, L. (2019). Understanding the role of momentum in stochastic gradient methods. In Advances in Neural Information Processing Systems, 32 . Curran Associates, Inc. Grandini, M., Bagli, E., & Visani, G. (2020). Metrics for multi-class classification: An overview. arXiv . https://doi.org/10.48550/arXiv.2008.05756 Guo, L., Wang, W., Zhong, L., Guo, L., Zhang, F., & Guo, Y. (2022). Texture analysis of the microstructure of internal curing concrete based on image recognition technology. Case Studies in Construction Materials, 17 , e01360. https://doi.org/10.1016/j.cscm.2022.e01360 Hilal, A. A. (2016). Microstructure of concrete. In High performance concrete technology and applications . IntechOpen. https://doi.org/10.5772/64574 Hsieh, Y. A., & Tsai, Y. J. (2020). Machine learning for crack detection: Review and model performance comparison. Journal of Computing in Civil Engineering, 34 (4), 04020038. https://doi.org/10.1061/(ASCE)CP.1943-5487.0000918 Hoskere, V., Narazaki, Y., Hoang, T. A., Spencer, B. F., & Yamaguchi, T. (2020). MaDnet: Multi-task semantic segmentation of multiple types of structural materials and damage in images of civil infrastructure. Journal of Civil Structural Health Monitoring, 10 (5), 757–773. https://doi.org/10.1007/s13349-020-00409-0 Hossin, M., & Sulaiman, M. N. (2015). A review on evaluation metrics for data classification evaluations. International Journal of Data Mining & Knowledge Management Process, 5 (2), 1–11. https://doi.org/10.5121/ijdkp.2015.5201 Johnson, J. M., & Khoshgoftaar, T. M. (2019). Survey on deep learning with class imbalance. Journal of Big Data, 6 , 27. https://doi.org/10.1186/s40537-019-0192-5 Kim, H., Ahn, E., Shin, M., & Sim, S. H. (2019). Crack and noncrack classification from concrete surface images using machine learning. Structural Health Monitoring, 18 (3), 725–738. https://doi.org/10.1177/1475921718768747 Koch, C., Georgieva, K., Kasireddy, V., Akinci, B., & Fieguth, P. (2015). A review on computer vision-based defect detection and condition assessment of concrete and asphalt civil infrastructure. Advanced Engineering Informatics, 29 (2), 196–210. https://doi.org/10.1016/j.aei.2015.01.008 Kulkarni, S., Singh, S., Balakrishnan, D., Sharma, S., Devunuri, S., & Korlapati, S. C. R. (2023). CrackSeg9k: A collection and benchmark for crack segmentation datasets and frameworks. In L. Karlinsky, T. Michaeli, & K. Nishino (Eds.), Computer vision – ECCV 2022 workshops (Lecture Notes in Computer Science, Vol. 13807). Springer. https://doi.org/10.1007/978-3-031-25082-8_12 McLaughlin, E., Charron, N., & Narasimhan, S. (2020). Automated defect quantification in concrete bridges using robotics and deep learning. Journal of Computing in Civil Engineering, 34 (4), 04020029. https://doi.org/10.1061/(ASCE)CP.1943-5487.0000915 Midtvedt, B., Helgadottir, S., Argun, A., Pineda, J., Midtvedt, D., & Volpe, G. (2021). Quantitative digital microscopy with deep learning. Applied Physics Reviews, 8 (1), Article 011310. https://doi.org/10.1063/5.0034891 Munawar, H. S., Hammad, A. W., Haddad, A., Soares, C. A. P., & Waller, S. T. (2021). Image-based crack detection methods: A review. Infrastructures, 6 (8), 115. https://doi.org/10.3390/infrastructures6080115 Pan, S. J., & Yang, Q. (2009). A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering, 22 (10), 1345–1359. https://doi.org/10.1109/TKDE.2009.191 Rashid, T., Mokji, M. M., & Rasheed, M. (2025). Cross-dataset evaluation of deep learning models for crack classification in structural surfaces. Journal of the Mechanical Behavior of Materials, 34 , 20250074. https://doi.org/10.1515/jmbm-2025-0074 Santos, M. S., Abreu, P. H., Japkowicz, N., Santos, J., & Cruz, C. (2022). On the joint-effect of class imbalance and overlap: A critical review. Artificial Intelligence Review, 55 (8), 6207–6275. https://doi.org/10.1007/s10462-022-10150-3 Sony, S., Dunphy, K., Sadhu, A., & Capretz, M. (2021). A systematic review of convolutional neural network-based structural condition assessment techniques. Engineering Structures, 226 , 111347. https://doi.org/10.1016/j.engstruct.2020.111347 Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., & Rabinovich, A. (2015). Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1–9). IEEE. https://doi.org/10.1109/CVPR.2015.7298594 Tang, S. W., Yao, Y., Andrade, C., & Li, Z. J. (2015). Recent durability studies on concrete structure. Cement and Concrete Research, 78 , 143–154. https://doi.org/10.1016/j.cemconres.2015.05.021 Yao, Y., Tung, S. T. E., & Glisic, B. (2014). Crack detection and characterization techniques—An overview. Structural Control and Health Monitoring, 21 (10), 1387–1413. https://doi.org/10.1002/stc.1655 Ye, X. W., Jin, T., & Yun, C. B. (2019). A review on deep learning-based structural health monitoring of civil infrastructures. Smart Structures and Systems, 24 (5), 567–585. https://doi.org/10.12989/sss.2019.24.5.567 Yuan, Q., Shi, Y., & Li, M. (2024). A review of computer vision-based crack detection methods in civil infrastructure: Progress and challenges. Remote Sensing, 16 (16), 2910. https://doi.org/10.3390/rs16162910 Wilkinson, M., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J. W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B. (2016). The FAIR guiding principles for scientific data management and stewardship. Scientific Data, 3 , 160018. https://doi.org/10.1038/sdata.2016.18 Zhou, S., Sheng, W., Wang, Z., Yao, W., Huang, H., Wei, Y., & Li, R. (2019). Quick image analysis of concrete pore structure based on deep learning. Construction and Building Materials, 208 , 144–157. https://doi.org/10.1016/j.conbuildmat.2019.03.006 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 15 May, 2026 Reviewers agreed at journal 12 May, 2026 Reviewers agreed at journal 08 May, 2026 Reviewers invited by journal 07 May, 2026 Editor assigned by journal 05 May, 2026 Submission checks completed at journal 05 May, 2026 First submitted to journal 02 May, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9594679","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":638706674,"identity":"9e715e0b-965d-4669-aef2-4d2cfd0b00b1","order_by":0,"name":"Hidir Selcuk Nogay","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA60lEQVRIiWNgGAWjYBACAygtx8DOwCABZh4gUosxAzNCC2MDMVoSG4jWYs5/9uDDLzWH0zcc5jG88eMXgxzfjQT2xxV4tFjOyEs2ljl2OBeoxdiyt4/BWPJGAmPjGXwOu8FjJi3BBtZiJsHbw5C4AaQFn8sMzp8Bavl3ON0AqEXybw9DPWEtB3LMJD+2HU4AaZHm+cGQYEBQy40cY2PGvnTDmYfZiq1lGyQMZ5552DiTgMMMH/74Zi3Pd7x54803f2yAjOQDH/FpAQFmHoZmCIuxDRQ1+KMFovAHQx2U+Yeg4lEwCkbBKBiBAAATz1O6M1ft1QAAAABJRU5ErkJggg==","orcid":"","institution":"Bursa Uludag University","correspondingAuthor":true,"prefix":"","firstName":"Hidir","middleName":"Selcuk","lastName":"Nogay","suffix":""}],"badges":[],"createdAt":"2026-05-02 14:54:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9594679/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9594679/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109337931,"identity":"70953889-2edd-464d-9027-b94e942d11c4","added_by":"auto","created_at":"2026-05-15 17:47:57","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":318615,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of the dataset creation process, including sample preparation, image acquisition, and labeling workflow.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9594679/v1/a505989b1829a38143976abb.png"},{"id":109405821,"identity":"ae024a96-644b-40d1-b896-8bebf2c4531d","added_by":"auto","created_at":"2026-05-17 13:20:19","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":96348,"visible":true,"origin":"","legend":"\u003cp\u003eRepresentative micro-surface images of concrete samples for the three main classes: dense, intermediate, and porous.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-9594679/v1/d1dd5b302dcf4a616a2a4056.png"},{"id":109405506,"identity":"3972ce65-c3b6-4b8b-8003-425760d90189","added_by":"auto","created_at":"2026-05-17 13:18:34","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":79057,"visible":true,"origin":"","legend":"\u003cp\u003eExample images of dense surface samples showing crack and healthy conditions used in binary classification.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-9594679/v1/1e1586a46b1a8d803ab92bcf.png"},{"id":109337933,"identity":"14914a9f-8026-4276-a0fc-bafcddf9319f","added_by":"auto","created_at":"2026-05-15 17:47:57","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":44043,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix of the binary classification task (crack vs. healthy) for dense surface samples.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-9594679/v1/875e9effde5271e20ebce706.png"},{"id":109337935,"identity":"a77afc4b-a7ca-4e25-8249-877ffda034bc","added_by":"auto","created_at":"2026-05-15 17:47:57","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":84822,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix of the three-class classification task (dense, intermediate, porous).\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-9594679/v1/adb6b692916c6b5b431fe4b0.png"},{"id":109337937,"identity":"9f50ecf0-f59a-427b-aaa8-2bfad21e2a93","added_by":"auto","created_at":"2026-05-15 17:47:57","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":197651,"visible":true,"origin":"","legend":"\u003cp\u003eTraining and validation accuracy curves of the model during the training process.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-9594679/v1/d294df0138b65bed355862df.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Micro-Surface Concrete Image Dataset for Crack Detection and Texture Classification Using Deep Learning","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eRecent advances in deep learning and computer vision have significantly improved automated inspection systems in civil engineering and material science (Ye et al., 2019). In particular, image-based analysis has emerged as a powerful tool for identifying surface characteristics and defects in construction materials such as concrete (Hsieh \u0026amp; Tsai, 2020). Due to its widespread use, ensuring the structural integrity and durability of concrete is of critical importance (Tang et al., 2015).\u003c/p\u003e\n\u003cp\u003eTraditional inspection methods rely heavily on manual evaluation, which is time-consuming and prone to subjective interpretation (Koch et al., 2015). Moreover, micro-level features such as fine cracks and subtle texture variations are often difficult to detect without specialized imaging tools (Yao et al., 2014). The integration of digital microscopy and deep learning techniques enables more accurate and consistent detection of such features (Midtvedt et al., 2021).\u003c/p\u003e\n\u003cp\u003eDespite the growing interest in this area, the availability of well-structured and labeled datasets remains limited (Yuan et al., 2024). Existing datasets typically focus on macroscopic defects, whereas micro-surface characteristics—such as texture transitions between dense, intermediate, and porous structures—are less explored (Munawar et al., 2021). Additionally, datasets that combine both surface classification and micro-crack detection within a unified framework are scarce (Hoskere et al., 2020).\u003c/p\u003e\n\u003cp\u003eTo address these limitations, this study introduces a novel concrete micro-surface image dataset constructed under controlled laboratory conditions. The overall workflow of dataset creation, including sample preparation, imaging, and labeling, is illustrated in Fig. 1. Concrete samples were prepared using three different water-to-cement ratios (0.35, 0.45, and 0.55), resulting in three distinct surface categories: dense, intermediate, and porous.\u003c/p\u003e\n\u003cp\u003eRepresentative examples of these surface types are presented in Fig. 2, highlighting the visual differences between compact, transitional, and porous structures. A total of 900 images were collected for the three main classes, with 300 images per class, as summarized in Table 1.\u003c/p\u003e\n\u003cp\u003eIn addition to surface classification, a binary classification subset was constructed for micro-crack detection within the dense class. Example images of crack and healthy samples are shown in Fig. 3. To avoid surface-induced bias, only top-surface images were used for crack labeling, ensuring that the model learns intrinsic crack features rather than positional differences.\u003c/p\u003e\n\u003cp\u003eTo validate the effectiveness of the dataset, deep learning-based experiments were conducted. The binary classification task achieved an accuracy of 95.45% with a high F1-score, as shown in Fig. 4 and Table 4, demonstrating clear separability between crack and healthy samples.\u003c/p\u003e\n\u003cp\u003eFurthermore, a three-class classification experiment (dense, intermediate, porous) achieved an overall accuracy of 88.15%. The corresponding confusion matrix is presented in Fig. 5. While dense and porous classes were classified with high accuracy, the intermediate class showed relatively lower performance due to its transitional nature. This observation is further supported by class-based F1-scores reported in Table 4.\u003c/p\u003e\n\u003cp\u003eTraining dynamics of the model are illustrated in Fig. 6 and Fig. 7, where both accuracy and loss curves indicate stable convergence without significant overfitting.\u003c/p\u003e\n\u003cp\u003eThe main contributions of this study are summarized as follows:\u003cbr\u003e\u0026nbsp;(i) the creation of a novel and balanced concrete micro-surface dataset,\u003cbr\u003e\u0026nbsp;(ii) the inclusion of both texture classification and micro-crack detection tasks,\u003cbr\u003e\u0026nbsp;(iii) the reduction of surface bias through controlled data acquisition,\u003cbr\u003e\u0026nbsp;and (iv) the validation of dataset usability through deep learning experiments.\u003c/p\u003e\n\u003cp\u003eThe proposed dataset provides a valuable benchmark for future research in automated concrete inspection and image-based material analysis.\u003c/p\u003e"},{"header":"2. Related Works","content":"\u003cp\u003eRecent studies have demonstrated the effectiveness of deep learning in concrete surface inspection and defect detection (McLaughlin et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2020\u003c/span\u003e ; Dung, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Convolutional neural networks (CNNs) have been widely used for crack detection, classification of surface defects, and structural health monitoring (Dorafshan et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Sony et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eSeveral publicly available datasets have contributed to this field, primarily focusing on macroscopic crack detection and damage classification (Kulkarni et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Kim et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). These datasets typically contain large-scale surface defects and are captured under uncontrolled environmental conditions. While such datasets are valuable, they often lack detailed microstructural information and controlled acquisition settings.\u003c/p\u003e \u003cp\u003eMore recent works have explored the use of high-resolution imaging and microscopy for analyzing material surfaces at a finer scale (Zhou et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Bangaru et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). These approaches provide more detailed insights into texture variations and micro-level defects. However, datasets that combine both microstructural surface classification and crack detection within a unified framework remain limited (Dong et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn addition, the influence of material composition, particularly the water-to-cement ratio, on surface texture has been investigated in several studies (Chuta et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). These studies highlight the importance of capturing controlled variations in material properties to improve classification performance (Guo et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eDespite these advancements, there is still a need for datasets that (i) provide balanced class distributions, (ii) include both macro and micro-level features, and (iii) are collected under controlled experimental conditions (Chakurkar et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The dataset proposed in this study aims to address these gaps by offering a comprehensive and structured benchmark for both surface classification and crack detection tasks.\u003c/p\u003e"},{"header":"3. Dataset Description","content":"\u003cp\u003eIn this study, a novel concrete micro-surface image dataset was developed under controlled laboratory conditions to support both surface classification and micro-crack detection tasks.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Sample Preparation\u003c/h2\u003e \u003cp\u003eConcrete samples were prepared using three different water-to-cement (w/c) ratios: 0.35, 0.45, and 0.55. These ratios were selected to produce distinct microstructural characteristics corresponding to dense, intermediate, and porous surface types, respectively (Elsharief et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2003\u003c/span\u003e). The variation in w/c ratio directly affects the internal structure and surface compactness of concrete, making it a suitable parameter for controlled dataset generation. The overall data acquisition workflow is illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Image Acquisition\u003c/h2\u003e \u003cp\u003eAll images were captured using a digital microscope under consistent environmental and lighting conditions to ensure data uniformity. For each sample, both top and bottom surface images were collected. However, in order to prevent bias in crack detection, only top surface images were used for crack labeling (Rashid et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eExample images from each class (dense, intermediate, porous) are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Dataset Structure\u003c/h2\u003e \u003cp\u003eThe dataset is organized into two complementary subsets designed to support both multi-class surface classification and binary crack detection tasks.\u003c/p\u003e \u003cp\u003eThe first subset focuses on multi-class surface classification and comprises three distinct categories: dense, intermediate, and porous. A total of 900 images were collected, with 300 images allocated to each class, resulting in a fully balanced dataset. Maintaining equal class distributions ensures fair model evaluation and prevents potential bias toward any specific surface type (Johnson \u0026amp; Khoshgoftaar, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The distribution of images across the three classes is summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDistribution of images across the three main surface classes in the dataset.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClass\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of Images\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIntermediate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePorous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e900\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe second subset was specifically constructed to address micro-crack detection within the dense concrete class. This binary dataset includes two categories: dense crack and dense healthy, with 150 images per class, yielding a total of 300 images. To ensure robust learning and eliminate potential acquisition-related bias, crack annotations were restricted exclusively to top-surface images (Geirhos et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). This design prevents the model from relying on irrelevant cues such as surface orientation or positional differences. Example crack and healthy samples used in the binary classification task are illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, while the detailed class distribution is presented in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDistribution of crack and healthy samples within the dense surface class.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClass\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of Images\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCrack\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e150\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHealthy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e150\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e300\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Naming Convention and Organization\u003c/h2\u003e \u003cp\u003eAll images in the dataset were systematically labeled and organized using a structured naming convention designed to encode essential attributes, including material type, water-to-cement ratio, surface location (top or bottom), and condition (crack or healthy). This standardized scheme ensures traceability of each sample and facilitates efficient dataset navigation during both training and evaluation processes.\u003c/p\u003e \u003cp\u003eThe adopted naming format integrates multiple descriptive components within a single filename (e.g., \u003cem\u003edense035top_crack_XXX\u003c/em\u003e), enabling direct identification of class category and acquisition conditions without requiring additional metadata files. Such explicit encoding enhances data transparency and reduces the risk of mislabeling during preprocessing stages.\u003c/p\u003e \u003cp\u003eA consistent folder hierarchy was also implemented to separate classes and experimental subsets systematically. This structured organization improves dataset reproducibility, supports scalable expansion, and aligns with established best practices in research data management (Wilkinson et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). The detailed folder structure and naming schema are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eFile naming convention and dataset organization structure.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFolder Name\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003edense035top_crack\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDense crack images\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003edense035top_healthy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDense healthy images\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eintermediate045top\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIntermediate images\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eporous055top\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePorous images\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Dataset Characteristics\u003c/h2\u003e \u003cp\u003eThe proposed dataset possesses several distinctive characteristics that enhance its methodological robustness and practical relevance. First, the dataset maintains a balanced class distribution, ensuring unbiased performance evaluation across categories. Second, all images were acquired under controlled laboratory conditions, minimizing environmental variability and improving data consistency.\u003c/p\u003e \u003cp\u003eIn addition, the dataset integrates both surface texture classification and micro-crack detection within a unified framework, enabling multi-level analysis of concrete microstructures. The deliberate separation of crack labeling from surface orientation further eliminates potential acquisition-induced bias, strengthening the reliability of model learning.\u003c/p\u003e \u003cp\u003eMoreover, the dataset is designed for multi-purpose usability, supporting both binary and multi-class classification scenarios. These combined attributes establish the dataset as a structured and challenging benchmark for evaluating deep learning models in concrete surface analysis and defect detection tasks.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Experimental Setup","content":"\u003cp\u003eTo assess the practical usability and learning capacity of the proposed dataset, deep learning-based classification experiments were conducted using a pretrained convolutional neural network architecture. The experimental framework was designed to evaluate both binary crack detection and multi-class surface classification scenarios under consistent training conditions.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Model Selection\u003c/h2\u003e \u003cp\u003eThe GoogLeNet architecture was selected as the baseline model due to its demonstrated effectiveness in image classification tasks and its computationally efficient design (Szegedy et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). The network was adapted through a transfer learning strategy, in which the final fully connected layers were modified to match the number of output classes in each experimental setting (Pan \u0026amp; Yang, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). This approach enables efficient feature reuse while maintaining adaptability to the specific classification tasks defined in the dataset.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Data Splitting\u003c/h2\u003e \u003cp\u003eTo ensure an unbiased evaluation of model performance, the dataset was randomly partitioned into three mutually exclusive subsets: 70% for training, 15% for validation, and 15% for testing. This split ratio was consistently applied to both binary and multi-class classification tasks to maintain comparability across experiments. The separation of validation and testing data prevents information leakage and provides a reliable assessment of the model\u0026rsquo;s generalization capability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Training Configuration\u003c/h2\u003e \u003cp\u003eThe network was trained using the stochastic gradient descent with momentum (SGDM) optimizer, a widely adopted optimization strategy for deep neural networks due to its stability and convergence efficiency (Gitman et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The training process was configured with an initial learning rate of 0.0001 and a maximum of 20 epochs. A mini-batch size of 100 was selected to balance computational efficiency and gradient stability, while validation was performed at regular intervals of five iterations to monitor learning progression and detect potential overfitting during training.\u003c/p\u003e \u003cp\u003eCross-entropy loss was employed as the objective function, as it is well-suited for supervised classification tasks involving probabilistic output distributions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e4.4 Evaluation Metrics\u003c/h2\u003e \u003cp\u003eModel performance was assessed using standard classification metrics, including accuracy, precision, recall, and F1-score (Hossin \u0026amp; Sulaiman, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). These metrics collectively provide a comprehensive evaluation of classification behavior by considering both overall correctness and class-specific prediction quality.\u003c/p\u003e \u003cp\u003eFor the multi-class classification task, class-wise performance metrics were computed, and macro-average values were reported to ensure balanced evaluation across all classes (Grandini et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The macro-averaging approach assigns equal importance to each class, preventing dominant categories from disproportionately influencing the overall performance assessment.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Results and Discussion","content":"\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e5.1 Binary Classification Results\u003c/h2\u003e \u003cp\u003eThe binary classification task was designed to distinguish between crack and healthy samples within the dense concrete class. The confusion matrix for this task is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe model achieved an accuracy of 95.45%, indicating a high level of separability between crack and non-crack samples. The precision and recall values further demonstrate the robustness of the model, with a precision of 1.00 and a recall of 0.91. The resulting F1-score of 0.95 confirms the reliability of the classification performance. These results suggest that micro-cracks captured using digital microscopy exhibit distinct visual patterns that can be effectively learned by deep learning models.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e5.2 Multi-Class Classification Results\u003c/h2\u003e \u003cp\u003eThe multi-class classification task involved distinguishing between dense, intermediate, and porous surface types. The corresponding confusion matrix is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe model achieved an overall accuracy of 88.15%, demonstrating strong performance in differentiating between concrete surface types. Class-wise analysis reveals that dense and porous classes were classified with high accuracy, while the intermediate class exhibited relatively lower performance. The F1-scores for dense, intermediate, and porous classes were 0.92, 0.79, and 0.92, respectively, with a mean F1-score of 0.88.\u003c/p\u003e \u003cp\u003eThe lower performance in the intermediate class can be attributed to its transitional nature, as it shares characteristics with both dense and porous structures. This makes it inherently more difficult to classify (Santos et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e5.3 Training Behavior\u003c/h2\u003e \u003cp\u003eThe training and validation accuracy curves are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e, while the loss curves are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eBoth accuracy and loss graphs indicate stable learning behavior with no significant signs of overfitting. The validation accuracy closely follows the training accuracy, suggesting good generalization capability of the model.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance results of the deep learning model for binary and multi-class classification tasks.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTask\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClass\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eF1-score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eBinary Classification (Dense Crack vs Healthy)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCrack\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.9091\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.9524\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHealthy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.9565\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.9778\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.9783\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.9545\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.9651\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eMulti-class Classification (Dense / Intermediate / Porous)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.8627\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.9778\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.9167\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIntermediate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.9394\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.6889\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.7949\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePorous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.8627\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.9778\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.9167\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.8883\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.8815\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.8761\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e5.4 Discussion\u003c/h2\u003e \u003cp\u003eThe experimental results confirm that the proposed dataset is suitable for both binary and multi-class classification tasks. The high performance in the binary task demonstrates the clarity of crack features, while the multi-class results highlight the dataset's ability to capture meaningful variations in concrete surface textures.\u003c/p\u003e \u003cp\u003eThe observed difficulty in classifying intermediate samples reflects real-world material behavior, where transitions between structural states are not sharply defined (Hilal, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). This characteristic adds to the realism and practical value of the dataset.\u003c/p\u003e \u003cp\u003eOverall, the dataset provides a challenging yet learnable benchmark for evaluating deep learning models in concrete surface analysis.\u003c/p\u003e \u003c/div\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eIn this study, a novel concrete micro-surface image dataset was introduced for both surface classification and micro-crack detection tasks. The dataset was constructed under controlled laboratory conditions using different water-to-cement ratios, resulting in three distinct surface categories: dense, intermediate, and porous.\u003c/p\u003e \u003cp\u003eIn addition to the multi-class dataset, a binary classification subset was created to specifically address micro-crack detection within dense concrete samples. The dataset was carefully designed to eliminate potential biases by restricting crack annotations to top-surface images.\u003c/p\u003e \u003cp\u003eTo validate the usability of the dataset, deep learning-based experiments were conducted using a pretrained GoogLeNet architecture. The binary classification task achieved high performance with an accuracy of 95.45% and an F1-score of 0.95, indicating clear separability between crack and healthy samples. The multi-class classification task achieved an overall accuracy of 88.15%, with strong performance in dense and porous classes, while the intermediate class presented a more challenging scenario due to its transitional nature.\u003c/p\u003e \u003cp\u003eThe results demonstrate that the proposed dataset provides a reliable and effective benchmark for both classification and defect detection tasks. Its balanced structure, controlled acquisition process, and inclusion of both texture and defect information make it a valuable resource for future research in concrete inspection and image-based material analysis.\u003c/p\u003e \u003cp\u003eFuture work may focus on expanding the dataset with additional environmental variations, increasing the number of samples, and exploring more advanced deep learning architectures to further improve classification performance.\u003c/p\u003e"},{"header":"Declarations","content":" \u003cp\u003e \u003cstrong\u003eCompeting Interests\u003c/strong\u003e \u003cp\u003eThe author declares that there are no competing interests.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eEthics Approval\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis study was supported by Bursa Uludağ University Scientific Research Projects Unit (BAP) under Project No. 2381.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eH.S.N. conceived the study, designed the methodology, performed data acquisition and analysis, and wrote the manuscript. The author reviewed and approved the final version of the manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe author would like to acknowledge the support of Bursa Uludağ University Scientific Research Projects Unit (BAP) under Project No. 2381.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe dataset generated during this study has been deposited in the Mendeley Data repository and is currently under moderation. A DOI has been reserved and will become publicly accessible upon final approval by the repository.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBangaru, S. S., Wang, C., Zhou, X., \u0026amp; Hassan, M. (2022). Scanning electron microscopy (SEM) image segmentation for microstructure analysis of concrete using U-net convolutional neural network. \u003cem\u003eAutomation in Construction, 144\u003c/em\u003e, 104602. https://doi.org/10.1016/j.autcon.2022.104602\u003c/li\u003e\n\u003cli\u003eChakurkar, P. S., Vora, D., Patil, S., Mishra, S., \u0026amp; Kotecha, K. (2023). Data-driven approach for AI-based crack detection: Techniques, challenges, and future scope. \u003cem\u003eFrontiers in Sustainable Cities, 5\u003c/em\u003e, 1253627. https://doi.org/10.3389/frsc.2023.1253627\u003c/li\u003e\n\u003cli\u003eChuta, E., Colin, J., \u0026amp; Jeong, J. (2020). The impact of the water-to-cement ratio on the surface morphology of cementitious materials. \u003cem\u003eJournal of Building Engineering, 32\u003c/em\u003e, 101716. https://doi.org/10.1016/j.jobe.2020.101716\u003c/li\u003e\n\u003cli\u003eDung, C. V. (2019). Autonomous concrete crack detection using deep fully convolutional neural network. \u003cem\u003eAutomation in Construction, 99\u003c/em\u003e, 52\u0026ndash;58. https://doi.org/10.1016/j.autcon.2018.11.028\u003c/li\u003e\n\u003cli\u003eDong, Y., Su, C., Qiao, P., \u0026amp; Sun, L. (2020). Microstructural crack segmentation of three-dimensional concrete images based on deep convolutional neural networks. \u003cem\u003eConstruction and Building Materials, 253\u003c/em\u003e, 119185. https://doi.org/10.1016/j.conbuildmat.2020.119185\u003c/li\u003e\n\u003cli\u003eDorafshan, S., Thomas, R. J., \u0026amp; Maguire, M. (2018). Comparison of deep convolutional neural networks and edge detectors for image-based crack detection in concrete. \u003cem\u003eConstruction and Building Materials, 186\u003c/em\u003e, 1031\u0026ndash;1045. https://doi.org/10.1016/j.conbuildmat.2018.08.011\u003c/li\u003e\n\u003cli\u003eElsharief, A., Cohen, M. D., \u0026amp; Olek, J. (2003). Influence of aggregate size, water cement ratio and age on the microstructure of the interfacial transition zone. \u003cem\u003eCement and Concrete Research, 33\u003c/em\u003e(11), 1837\u0026ndash;1849. https://doi.org/10.1016/S0008-8846(03)00205-9\u003c/li\u003e\n\u003cli\u003eGeirhos, R., Jacobsen, J. H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., \u0026amp; Wichmann, F. A. (2020). Shortcut learning in deep neural networks. \u003cem\u003eNature Machine Intelligence, 2\u003c/em\u003e(11), 665\u0026ndash;673. https://doi.org/10.1038/s42256-020-00257-z\u003c/li\u003e\n\u003cli\u003eGitman, I., Lang, H., Zhang, P., \u0026amp; Xiao, L. (2019). Understanding the role of momentum in stochastic gradient methods. In \u003cem\u003eAdvances in Neural Information Processing Systems, 32\u003c/em\u003e. Curran Associates, Inc.\u003c/li\u003e\n\u003cli\u003eGrandini, M., Bagli, E., \u0026amp; Visani, G. (2020). Metrics for multi-class classification: An overview. \u003cem\u003earXiv\u003c/em\u003e. https://doi.org/10.48550/arXiv.2008.05756\u003c/li\u003e\n\u003cli\u003eGuo, L., Wang, W., Zhong, L., Guo, L., Zhang, F., \u0026amp; Guo, Y. (2022). Texture analysis of the microstructure of internal curing concrete based on image recognition technology. \u003cem\u003eCase Studies in Construction Materials, 17\u003c/em\u003e, e01360. https://doi.org/10.1016/j.cscm.2022.e01360\u003c/li\u003e\n\u003cli\u003eHilal, A. A. (2016). Microstructure of concrete. In \u003cem\u003eHigh performance concrete technology and applications\u003c/em\u003e. IntechOpen. https://doi.org/10.5772/64574\u003c/li\u003e\n\u003cli\u003eHsieh, Y. A., \u0026amp; Tsai, Y. J. (2020). Machine learning for crack detection: Review and model performance comparison. \u003cem\u003eJournal of Computing in Civil Engineering, 34\u003c/em\u003e(4), 04020038. https://doi.org/10.1061/(ASCE)CP.1943-5487.0000918\u003c/li\u003e\n\u003cli\u003eHoskere, V., Narazaki, Y., Hoang, T. A., Spencer, B. F., \u0026amp; Yamaguchi, T. (2020). MaDnet: Multi-task semantic segmentation of multiple types of structural materials and damage in images of civil infrastructure. \u003cem\u003eJournal of Civil Structural Health Monitoring, 10\u003c/em\u003e(5), 757\u0026ndash;773. https://doi.org/10.1007/s13349-020-00409-0\u003c/li\u003e\n\u003cli\u003eHossin, M., \u0026amp; Sulaiman, M. N. (2015). A review on evaluation metrics for data classification evaluations. \u003cem\u003eInternational Journal of Data Mining \u0026amp; Knowledge Management Process, 5\u003c/em\u003e(2), 1\u0026ndash;11. https://doi.org/10.5121/ijdkp.2015.5201\u003c/li\u003e\n\u003cli\u003eJohnson, J. M., \u0026amp; Khoshgoftaar, T. M. (2019). Survey on deep learning with class imbalance. \u003cem\u003eJournal of Big Data, 6\u003c/em\u003e, 27. https://doi.org/10.1186/s40537-019-0192-5\u003c/li\u003e\n\u003cli\u003eKim, H., Ahn, E., Shin, M., \u0026amp; Sim, S. H. (2019). Crack and noncrack classification from concrete surface images using machine learning. \u003cem\u003eStructural Health Monitoring, 18\u003c/em\u003e(3), 725\u0026ndash;738. https://doi.org/10.1177/1475921718768747\u003c/li\u003e\n\u003cli\u003eKoch, C., Georgieva, K., Kasireddy, V., Akinci, B., \u0026amp; Fieguth, P. (2015). A review on computer vision-based defect detection and condition assessment of concrete and asphalt civil infrastructure. \u003cem\u003eAdvanced Engineering Informatics, 29\u003c/em\u003e(2), 196\u0026ndash;210. https://doi.org/10.1016/j.aei.2015.01.008\u003c/li\u003e\n\u003cli\u003eKulkarni, S., Singh, S., Balakrishnan, D., Sharma, S., Devunuri, S., \u0026amp; Korlapati, S. C. R. (2023). CrackSeg9k: A collection and benchmark for crack segmentation datasets and frameworks. In L. Karlinsky, T. Michaeli, \u0026amp; K. Nishino (Eds.), \u003cem\u003eComputer vision \u0026ndash; ECCV 2022 workshops\u003c/em\u003e (Lecture Notes in Computer Science, Vol. 13807). Springer. https://doi.org/10.1007/978-3-031-25082-8_12\u003c/li\u003e\n\u003cli\u003eMcLaughlin, E., Charron, N., \u0026amp; Narasimhan, S. (2020). Automated defect quantification in concrete bridges using robotics and deep learning. \u003cem\u003eJournal of Computing in Civil Engineering, 34\u003c/em\u003e(4), 04020029. https://doi.org/10.1061/(ASCE)CP.1943-5487.0000915\u003c/li\u003e\n\u003cli\u003eMidtvedt, B., Helgadottir, S., Argun, A., Pineda, J., Midtvedt, D., \u0026amp; Volpe, G. (2021). Quantitative digital microscopy with deep learning. \u003cem\u003eApplied Physics Reviews, 8\u003c/em\u003e(1), Article 011310. https://doi.org/10.1063/5.0034891\u003c/li\u003e\n\u003cli\u003eMunawar, H. S., Hammad, A. W., Haddad, A., Soares, C. A. P., \u0026amp; Waller, S. T. (2021). Image-based crack detection methods: A review. \u003cem\u003eInfrastructures, 6\u003c/em\u003e(8), 115. https://doi.org/10.3390/infrastructures6080115\u003c/li\u003e\n\u003cli\u003ePan, S. J., \u0026amp; Yang, Q. (2009). A survey on transfer learning. \u003cem\u003eIEEE Transactions on Knowledge and Data Engineering, 22\u003c/em\u003e(10), 1345\u0026ndash;1359. https://doi.org/10.1109/TKDE.2009.191\u003c/li\u003e\n\u003cli\u003eRashid, T., Mokji, M. M., \u0026amp; Rasheed, M. (2025). Cross-dataset evaluation of deep learning models for crack classification in structural surfaces. \u003cem\u003eJournal of the Mechanical Behavior of Materials, 34\u003c/em\u003e, 20250074. https://doi.org/10.1515/jmbm-2025-0074\u003c/li\u003e\n\u003cli\u003eSantos, M. S., Abreu, P. H., Japkowicz, N., Santos, J., \u0026amp; Cruz, C. (2022). On the joint-effect of class imbalance and overlap: A critical review. \u003cem\u003eArtificial Intelligence Review, 55\u003c/em\u003e(8), 6207\u0026ndash;6275. https://doi.org/10.1007/s10462-022-10150-3\u003c/li\u003e\n\u003cli\u003eSony, S., Dunphy, K., Sadhu, A., \u0026amp; Capretz, M. (2021). A systematic review of convolutional neural network-based structural condition assessment techniques. \u003cem\u003eEngineering Structures, 226\u003c/em\u003e, 111347. https://doi.org/10.1016/j.engstruct.2020.111347\u003c/li\u003e\n\u003cli\u003eSzegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., \u0026amp; Rabinovich, A. (2015). Going deeper with convolutions. In \u003cem\u003eProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)\u003c/em\u003e (pp. 1\u0026ndash;9). IEEE. https://doi.org/10.1109/CVPR.2015.7298594\u003c/li\u003e\n\u003cli\u003eTang, S. W., Yao, Y., Andrade, C., \u0026amp; Li, Z. J. (2015). Recent durability studies on concrete structure. \u003cem\u003eCement and Concrete Research, 78\u003c/em\u003e, 143\u0026ndash;154. https://doi.org/10.1016/j.cemconres.2015.05.021\u003c/li\u003e\n\u003cli\u003eYao, Y., Tung, S. T. E., \u0026amp; Glisic, B. (2014). Crack detection and characterization techniques\u0026mdash;An overview. \u003cem\u003eStructural Control and Health Monitoring, 21\u003c/em\u003e(10), 1387\u0026ndash;1413. https://doi.org/10.1002/stc.1655\u003c/li\u003e\n\u003cli\u003eYe, X. W., Jin, T., \u0026amp; Yun, C. B. (2019). A review on deep learning-based structural health monitoring of civil infrastructures. \u003cem\u003eSmart Structures and Systems, 24\u003c/em\u003e(5), 567\u0026ndash;585. https://doi.org/10.12989/sss.2019.24.5.567\u003c/li\u003e\n\u003cli\u003eYuan, Q., Shi, Y., \u0026amp; Li, M. (2024). A review of computer vision-based crack detection methods in civil infrastructure: Progress and challenges. \u003cem\u003eRemote Sensing, 16\u003c/em\u003e(16), 2910. https://doi.org/10.3390/rs16162910\u003c/li\u003e\n\u003cli\u003eWilkinson, M., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J. W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., \u0026hellip; Mons, B. (2016). The FAIR guiding principles for scientific data management and stewardship. \u003cem\u003eScientific Data, 3\u003c/em\u003e, 160018. https://doi.org/10.1038/sdata.2016.18\u003c/li\u003e\n\u003cli\u003eZhou, S., Sheng, W., Wang, Z., Yao, W., Huang, H., Wei, Y., \u0026amp; Li, R. (2019). Quick image analysis of concrete pore structure based on deep learning. \u003cem\u003eConstruction and Building Materials, 208\u003c/em\u003e, 144\u0026ndash;157. https://doi.org/10.1016/j.conbuildmat.2019.03.006\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":false,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"asian-journal-of-civil-engineering","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Asian Journal of Civil Engineering](https://www.springer.com/journal/42107)","snPcode":"42107","submissionUrl":"https://submission.nature.com/new-submission/42107/3","title":"Asian Journal of Civil Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Concrete dataset, Micro-surface analysis, Crack detection, Deep learning, Image classification, Convolutional neural networks","lastPublishedDoi":"10.21203/rs.3.rs-9594679/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9594679/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis study presents a novel concrete micro-surface image dataset designed for both surface classification and micro-crack detection tasks using deep learning. The dataset was constructed under controlled laboratory conditions by preparing concrete samples with three different water-to-cement ratios (0.35, 0.45, and 0.55), resulting in dense, intermediate, and porous surface structures. A total of 900 images were collected for multi-class classification, with 300 images per class, ensuring a balanced distribution. In addition, a binary subset consisting of 300 images (150 crack and 150 healthy) was created to specifically address micro-crack detection within dense surfaces.\u003c/p\u003e \u003cp\u003eTo eliminate potential bias, crack annotations were restricted to top-surface images, allowing the model to focus on intrinsic crack features. The usability of the dataset was validated using a pretrained convolutional neural network based on GoogLeNet. Experimental results show that the binary classification task achieved an accuracy of 95.45% with an F1-score of 0.95, while the three-class classification task achieved an overall accuracy of 88.15% with a mean F1-score of 0.876. The results indicate that dense and porous classes are highly distinguishable, whereas intermediate surfaces present a more challenging classification problem due to their transitional characteristics.\u003c/p\u003e \u003cp\u003eThe proposed dataset provides a reliable and structured benchmark for future research in image-based concrete analysis, defect detection, and material characterization using deep learning.\u003c/p\u003e","manuscriptTitle":"A Micro-Surface Concrete Image Dataset for Crack Detection and Texture Classification Using Deep Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-15 17:47:53","doi":"10.21203/rs.3.rs-9594679/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-05-15T05:48:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"145173612517918045178368328009584728610","date":"2026-05-12T05:09:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"245077285551394916077572771200087634197","date":"2026-05-08T11:57:24+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-05-07T05:03:32+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-05-05T12:54:07+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-05-05T12:17:19+00:00","index":"","fulltext":""},{"type":"submitted","content":"Asian Journal of Civil Engineering","date":"2026-05-02T14:49:13+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"asian-journal-of-civil-engineering","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Asian Journal of Civil Engineering](https://www.springer.com/journal/42107)","snPcode":"42107","submissionUrl":"https://submission.nature.com/new-submission/42107/3","title":"Asian Journal of Civil Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"9363f097-e716-46fb-b602-60b44c2ffa3b","owner":[],"postedDate":"May 15th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-15T05:48:32+00:00","index":16,"fulltext":""},{"type":"reviewerAgreed","content":"145173612517918045178368328009584728610","date":"2026-05-12T05:09:19+00:00","index":15,"fulltext":""},{"type":"reviewerAgreed","content":"245077285551394916077572771200087634197","date":"2026-05-08T11:57:24+00:00","index":14,"fulltext":""},{"type":"reviewersInvited","content":"8","date":"2026-05-07T05:03:32+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-05-05T12:54:07+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-05-05T12:17:19+00:00","index":"","fulltext":""},{"type":"submitted","content":"Asian Journal of Civil Engineering","date":"2026-05-02T14:49:13+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-15T17:47:53+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-15 17:47:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9594679","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9594679","identity":"rs-9594679","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00