Detection of Pediatric Cataracts through Non-invasive Facial Photography using Deep Convolutional Neural Networks for Early Diagnosis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Detection of Pediatric Cataracts through Non-invasive Facial Photography using Deep Convolutional Neural Networks for Early Diagnosis Shion Hayashi, Emi Kashizuka, Koki Sakata, Natsuki Motomiya, Yosuke Asakawa, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8295441/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 10 Apr, 2026 Read the published version in BMC Ophthalmology → Version 1 posted 11 You are reading this latest preprint version Abstract Background: To develop a deep learning–based model for detecting pediatric cataracts using noninvasive facial photographs of infants and toddlers, aiming to facilitate early diagnosis during the critical period of visual development. Methods: This prospective observational study included 32 patients (47 eyes) with pediatric cataracts and 93 cataract-free controls (186 eyes) who visited the National Center for Child Health and Development between November 2021 and December 2024. Multiple facial photographs were captured using a digital single-lens reflex camera with flash illumination. After preprocessing and cropping of single-eye regions, 727 cropped images (149 cataract and 578 control images) were used to train and validate a convolutional neural network based on the Inception V3 architecture. The model was trained using transfer learning and evaluated using five-fold cross-validation. The diagnostic performance was assessed using sensitivity, specificity, accuracy, area under the receiver operating characteristic curve (AUC), and F 1 score. Results: The model demonstrated a high diagnostic performance across all folds. The minimum AUC among the five folds was 0.9692, with an accuracy of 0.9720, sensitivity of 0.8636, specificity of 0.9917, and F 1 score of 0.9048. Despite the relatively small dataset, the model consistently achieved robust results without overfitting. Conclusions: The proposed deep learning model accurately detected pediatric cataracts in ordinary facial photographs of infants. This noninvasive, low-cost approach may complement conventional screening and improve the early detection of pediatric cataracts, enabling timely referral and treatment during the critical period of visual development. deep learning image classification pediatric cataract facial photograph infancy Figures Figure 1 Figure 2 Background Pediatric cataracts are one of the leading causes of preventable childhood blindness worldwide. Although the annual incidence is relatively low, it has a profound and lifelong impact on visual development and the quality of life. Early detection and surgical treatment are crucial, as visual outcomes are significantly better when surgery for dense unilateral cataracts is performed within the first 6 weeks of life 1 and for bilateral cataracts, within the first 2–3 months 2 , 3 . Delayed detection leads to irreversible visual impairment, resulting in a substantial socioeconomic burden for both affected individuals and society. In Japan's multicenter studies, early surgery has been demonstrated to be essential for a favorable visual prognosis in congenital and developmental cataracts, emphasizing the clinical importance of prompt diagnosis and treatment 4 . Similarly, Oshika et al. demonstrated that surgery performed within the critical period—specifically within 10 weeks for bilateral cases and 6 weeks for unilateral cases—resulted in superior long-term visual outcomes compared with later intervention 5 . These findings from both international and Japanese studies underscore the universal importance of detecting pediatric cataracts as early as possible to prevent irreversible visual loss. In Japan, infant visual screening is performed as part of routine infant health checkups, and pediatricians play a central role in the detection of ocular abnormalities during these examinations. However, the detection of pediatric cataracts during routine health checkups remains challenging. According to a previous nationwide survey, only 12% of pediatric cataract cases were detected during infant screenings, indicating that most cases are overlooked until the later stages of disease progression 6 . This delay is partly because congenital cataracts are rare, limiting opportunities for general pediatricians to gain the necessary diagnostic experience. Furthermore, no convenient and effective detection method has been established for widespread use in primary care settings. Recent advances in artificial intelligence (AI) and deep learning have opened new possibilities for pediatric ocular screenings. Munson et al. developed the CRADLE smartphone application, which automatically detects leukocoria (white pupillary reflex) from casual childhood photographs. In a longitudinal study, CRADLE identified leukocoria an average of 1.3 years before clinical diagnosis in 80% of children with eye diseases such as retinoblastoma, Coats’ disease, and pediatric cataracts 7 . These findings demonstrate the potential of AI to complement conventional screening by utilizing readily available digital images. Building on this concept, we hypothesized that deep learning models could identify pediatric cataracts directly from infant facial photographs, providing a noninvasive, low-cost, and scalable screening tool 8 . This study aimed to develop and evaluate a deep learning model capable of detecting pediatric cataracts using ordinary facial photographs of infants. Methods Study Design and Ethics Approval This prospective observational study was conducted at the Division of Ophthalmology, National Center for Child Health and Development (NCCHD), Tokyo, Japan. The study protocol was reviewed and approved by the Ethics Committee of the NCCHD (approval number: 2021 − 141). Written informed consent was obtained from the guardians of all participants before enrollment in accordance with the Declaration of Helsinki. Selection of Pediatric Study Subjects The participants were pediatric patients who visited the Division of Ophthalmology at the NCCHD between November 2021 and December 2024. The inclusion criteria were as follows: 1) age between 1 and 36 months at the time of image acquisition and 2) guardians who provided written informed consent for participation. The exclusion criteria were photographs of insufficient quality, as detailed in the following section, patients with incomplete clinical data, and those who refused consent. Congenital or developmental cataracts were diagnosed by a board-certified pediatric ophthalmologist based on slit-lamp biomicroscopy performed after pharmacologic mydriasis. Image Acquisition in Pediatric Ophthalmology Outpatient Clinics Multiple facial photographs were taken for each subject using a Canon single-lens reflex (SLR) EOS 80D camera with flash illumination. The photographs included both eyes and the surrounding periocular areas. Images were excluded if they were out of focus, showed closed eyes, contained excessive hair, or had obstructive elements such as hands or motion artifacts. The final dataset consisted of only images that met the quality standards suitable for analysis. Image Processing and Dataset Construction Facial photographs were collected from patients with pediatric cataracts and cataract-free control subjects. The control group comprised pediatric subjects with clear ocular media without cataracts or other opacities. Cataract patients could be confidently diagnosed based on slit-lamp examinations. All photographs underwent a standardized preprocessing pipeline before use in model development. First, the facial images were saved in JPEG format and rotated as necessary to ensure an upright orientation. This adjustment was performed using GIMP version 2.10 or ImageMagick version 7.1. Next, regions covering one eye, including the eyelid area, were manually annotated using Microsoft VoTT version 2.1. Although these images will be subsequently processed into exact squares, the current annotation process involves cropping regions that are roughly square-shaped but not precisely so. Based on the pixel coordinates in the resulting JSON file, each annotated region was automatically cropped using a custom Python script. To prevent model overestimation and avoid data leakage during cross-validation, images from the same subject were kept within either the training or test set, ensuring that data from individual patients were not divided between the sets. Development of CNN-Based Classification Model A convolutional neural network (CNN) was implemented using the Inception V3 architecture within the PyTorch version 2.7.0 framework. Transfer learning was applied using weights pre-trained on the ImageNet 1k v1 dataset to leverage the prior knowledge of general image features. To improve generalizability and reduce overfitting, various data augmentation techniques were applied using the torchvision library. These included random affine transformations, random erasing, random perspective transformations, color jittering, random auto-contrast adjustments, random rotations, and random vertical flips. The CNN was trained for 50 epochs using stochastic gradient descent (SGD) optimization with a learning rate of 0.03, momentum of 0.9, weight decay of 10 − 4 , and a batch size of 64. All deep learning computations were performed on an Ubuntu 22.04.5 system configured with one NVIDIA Quadro GV100 graphics processing unit. Evaluation of Classification Performance The trained model was evaluated using 5-fold cross-validation to ensure robustness to data bias. Random allocation was performed at the patient level rather than the image level, with each fold in the test dataset containing data from six or seven patients in the fold. Among the five evaluations, we report the most conservative results, presenting the lowest area under the receiver operating characteristic curve (AUC) values as representative performance metrics. Performance was assessed using standard classification metrics, including AUC, sensitivity, specificity, and F 1 score. The classification performance of the deep learning model was compared with the diagnostic assessments of pediatric ophthalmologists using McNemar’s test or the chi-square test. Statistical significance was set at p < 0.05. Results Patient Characteristics A total of 32 patients with pediatric cataracts and 93 cataract-free control subjects were included in the study. A total of 390 facial photographs were collected, consisting of 98 images from 32 patients (47 eyes) with pediatric cataract and 292 images from 93 cataract-free control subjects (186 eyes). Among the 47 eyes with pediatric cataracts, the distribution of cataract morphologies was as follows: 13 with fetal nuclear cataracts, 11 with anterior polar and subcapsular cataracts, 8 with posterior subcapsular cataracts, 6 with total cataracts, 5 with membranous cataracts, 2 with lamellar cataracts, and 2 with combined fetal nuclear and posterior subcapsular cataracts. Regarding ocular comorbidities, 5 eyes had microcornea, 3 eyes had persistent fetal vasculature (PFV), 1 eye had microphthalmia, and 1 eye had optic disc dysplasia. Some eyes exhibited more than one anomaly (overlapping cases), and 38 eyes exhibited no additional ocular abnormalities. Regarding clinical outcomes, 23 patients (33 eyes) underwent cataract surgery, 5 patients (10 eyes) were followed conservatively, and 4 patients (4 eyes) were deemed ineligible for surgical intervention because postoperative visual improvement was not expected. Dataset Preparation and Cross-validation After preprocessing and cropping of single-eye regions, 727 cropped images were generated (Fig. 1 ). The dataset consisted of 149 crops from pediatric cataract patients and 578 crops from controls, with some subjects providing data from only one eye each. We implemented deep learning using an Inception V3-based convolutional neural network architecture. Although the training and test sets were randomly assigned, we conducted 5-fold cross-validation to ensure that our results were free from selection bias and reproducible across different data partitions. Cross-validation was performed using subject-based partitioning rather than image-based partitioning to ensure that data from the same individual were not split between the training and test sets. Consequently, the number of image crops varied across the folds owing to the different numbers of images per subject. Table 1 summarizes the distribution of the number of patients and cropped images used for training and testing in each randomly assigned fold. Table 1. Numbers of subjecs and crops in the 5-fold cross-validation Fold Number of training subjects Number of test subjects Number of training crops Number of test crops Cataract Control Cataract Control Cataract Control Cataract Control 1 26 74 6 19 130 443 19 135 2 26 74 6 19 94 469 55 109 3 26 74 6 19 127 457 22 121 4 25 75 7 18 125 472 24 106 5 25 75 7 18 120 471 29 107 Model Performance We implemented a binary classification system for pediatric cataract detection using convolutional neural network architecture. The network was based on Inception V3 pre-trained on ImageNet, with transfer learning applied for our specific task, as described in detail in the Methods section of this paper. The performance was evaluated using five-fold cross-validation, which yielded five sets of results. The distribution of data across these folds showed no significant imbalances, as presented in Table 1. To investigate the effect of training duration on model performance, we analyzed fold 2 by training separate models with varying numbers of epochs (1, 2, 8, and 50) and visualized their performance using ROC curves (Fig. 2 ). The logits from the CNN were transformed into probability values using the softmax function, and classification was performed using a threshold of 0.50. The resulting sensitivity, specificity, and other values for each fold are presented in Table 2. Despite the relatively limited size of our dataset, even in fold 3, which showed the lowest AUC of 0.9692, the model achieved high-performance metrics with an accuracy of 0.9720, sensitivity of 0.8636, specificity of 0.9917, and F 1 score of 0.9048. These results suggest that the model demonstrates sufficient diagnostic accuracy for potential clinical use. Table 2. Performance metrics for 5-fold cross-validation Fold Accuracy Sensitivity Specificity AUC F 1 score 1 0.9740 0.8947 0.9852 0.9957 0.8947 2 0.9939 0.9818 1.0000 1.0000 0.9908 3 0.9720 0.8636 0.9917 0.9692 0.9048 4 0.9692 0.8333 1.0000 0.9792 0.9091 5 0.9926 1.0000 0.9907 1.0000 0.9831 AUC area under the receiver operating characteristic curve Discussion In this study, we developed and evaluated a deep learning–based model to detect pediatric cataracts using facial photographs of infants and toddlers. The model was based on the Inception V3 convolutional neural network architecture, pretrained on ImageNet, and fine-tuned for our specific task. We implemented comprehensive data augmentation techniques, including random rotation, vertical flipping, brightness adjustment, and contrast validation, to enhance the model robustness and prevent overfitting. Despite the relatively small dataset, our model achieved high classification performance across five-fold cross-validation, with AUC values consistently exceeding 0.96. These results demonstrate that deep learning models can effectively identify cataract-related signs in facial photographs, even in young children who are often uncooperative during traditional screening procedures. Early detection of pediatric cataracts is essential to prevent irreversible visual loss in children. However, current screening strategies, including the red reflex test (RRT), have several limitations. The RRT is simple and inexpensive 9 , but its performance varies significantly depending on the examiner’s experience and lesion characteristics. In large-scale prospective studies, its overall sensitivity for ocular pathologies has been reported to be as low as 7–14%, despite a high specificity exceeding 95% 10,11 . Even after standardized training, sensitivity among non-specialist examiners, such as midwives, remains moderate (approximately 50%) compared with optometry students or ophthalmologists 12 . These limitations make widespread implementation of the RRT challenging, especially in primary care settings. In Japan, RRT is not routinely included in the national infant health checkup program, leaving a gap in the early detection of pediatric ocular diseases. Several studies have explored the application of AI in pediatric ocular diseases, providing valuable context for our work. Liu et al. proposed a deep learning framework using slit-lamp images to diagnose pediatric cataracts with high accuracy 13 ; however, slit-lamp photography requires specialized ophthalmic equipment and trained personnel, limiting its feasibility for routine screenings. Munson et al. demonstrated that AI can detect leukocoria from casual childhood photographs taken at home using a smartphone 7 , identifying eye diseases such as retinoblastoma and Coats’ disease years before a clinical diagnosis. Recently, Kitaguchi et al. successfully applied deep learning to detect childhood glaucoma using periocular facial photographs, highlighting the potential of non-specialized images for screening other pediatric eye conditions 14 . Recent studies in Japan have illustrated the growing trend of integrating AI into ophthalmic screening and public health frameworks. Takeda et al. developed machine learning–based models for predicting retinopathy of prematurity using routinely available clinical data without relying on an expensive fundus camera 15 . Similarly, Kamei et al. demonstrated that AI analysis of retinal photographs could estimate the “retinal age gap,” linking ocular findings with systemic health indicators 16 . Together, these studies reflect the increasing use of AI to support data-driven health screenings and disease prevention. To our knowledge, this is the first study to demonstrate that pediatric cataracts can be detected using ordinary facial photographs taken under non-specialized conditions. This approach has several potential benefits. First, it does not require any specialized ophthalmic equipment or examiner expertise, allowing screening in primary care settings or even at home. Second, it can be integrated into existing infant health checkup systems in Japan, enhancing the detection rate of ocular diseases that are often missed by conventional methods. Furthermore, automated image screening applications could be incorporated into parental smartphone platforms, enabling opportunistic screening using photos already available in daily life. However, several limitations must be acknowledged. First, the dataset used in this study was relatively small and collected from a single institution, which may limit its generalizability. Expanding the dataset to include images from multiple facilities and different ethnic backgrounds is essential for improving model robustness. Second, our study used cropped eye images derived from facial photographs, which may not fully reflect the variability encountered in real-world settings. Future work should focus on developing algorithms capable of automatically detecting the eye region from unprocessed facial images. Third, although the model demonstrated high diagnostic accuracy in differentiating between clear cataracts and normal eyes, its performance in detecting mild or subtle cataracts requires further investigation. Conclusion In conclusion, this study demonstrated the feasibility of using a deep learning model to detect pediatric cataracts from infant facial photographs. The proposed method offers a noninvasive, low-cost, and scalable approach that can complement the conventional red reflex test in pediatric vision screening. With further validation and large-scale multicenter studies, such AI-based tools have the potential to significantly improve the early detection and management of pediatric cataracts, ultimately contributing to better lifelong visual outcomes in children. Abbreviations AUC: Area under the receiver operating characteristic curve AI: Artificial intelligence CNN: Convolutional neural network DL: Deep learning NCCHD: National Center for Child Health and Development PFV: Persistent fetal vasculature RRT: Red reflex test SD: Standard deviation Declarations Ethics approval and consent to participate The study protocol was approved by the Ethics Committee of the National Center for Child Health and Development (NCCHD), Tokyo, Japan (approval no. 2021-141). Written informed consent was obtained from the guardians of all participants in accordance with the Declaration of Helsinki. Consent for publication Written informed consent for publication of anonymized images and research data was obtained from the guardians of all participating children. Availability of data and materials The datasets generated and analyzed during the current study are not publicly available due to ethical restrictions regarding patient privacy, but are available from the corresponding author on reasonable request. Competing interests KO has received funding from the Canon Foundation. All other authors declare that they have no competing interests. Funding This work was supported by the following funding sources: NCCHD Institutional Grant (Grant No. 2023B-6) to SH and KO Health Labour Sciences Research Grant (No. 23AC1001) to KO Health and Labour Sciences Research Grants for Research on Intractable Diseases (No. 23FC1055) to SN JSPS KAKENHI (Grant No. 21K06143) to KO (for GPU-equipped system) Authors’ contributions SH, SN, EK, and KO: Study design and conceptualization. SH and KO: Drafting the manuscript. KO and SN: Critical revision of the manuscript. SH and KO: Literature review. EK and KO: Data analysis. SH, EK, KS, NM, YA, and MH: Data collection. SH, KS, NM, YA, MH, YT, TY, and NS: Clinical recruitment. All authors reviewed and approved the final manuscript. Acknowledgements We are grateful to all participants and their families for contributing to this study. We thank Ms. Miki Nobe for her assistance with manual data preparation, including image annotation. SH sincerely thanks Professor Masahiko Sugimoto (Department of Ophthalmology, Yamagata University) for his guidance and continuous support throughout this work. The authors also used ChatGPT (OpenAI, San Francisco, CA, USA) and Paperpal (Iris.ai, Norway) to improve the clarity and readability of the English manuscript; all scientific content and interpretations are the sole responsibility of the authors. Authors’ information (optional) Not applicable. References Birch EE, Stager DR. The critical period for surgical treatment of dense congenital unilateral cataract. Invest Ophthalmol Vis Sci. 1996;37:1532–8. Lambert SR, Lynn MJ, Reeves R, Plager DA, Buckley EG, Wilson ME. Is there a latent period for the surgical treatment of children with dense bilateral congenital cataracts? J AAPOS. 2006;10:30–6. 10.1016/j.jaapos.2005.10.002 . Birch EE, Cheng C, Stager DR Jr, Weekley DR Jr, Stager DR Sr. The critical period for surgical treatment of dense congenital bilateral cataracts. J AAPOS. 2009;13:67–71. 10.1016/j.jaapos.2008.07.010 . Nagamoto T, Oshika T, Fujikado T, Ishibashi T, Sato M, Kondo M, et al. Surgical outcomes of congenital and developmental cataracts in Japan. Jpn J Ophthalmol. 2016;60:127–34. 10.1007/s10384-015-0498-1 . Oshika T, Nishina S, Unoki N, Miyagi M, Nomura K, Mori T, et al. Ten-year outcomes of congenital cataract surgery performed within the first six months of life. J Cataract Refract Surg. 2024;50:707–12. 10.1097/j.jcrs.0000000000001189 . Nishina S, Yokoi T, Kobayashi Y, Noda E, Azuma N. The time and route for finding infants' eye disease. Folia Japonica de Ophthalmol Clin. 2010;3:172–7. Munson MC, Plewman DL, Baumer KM, Henning R, Zahler CT, Kietzman AT, Beard AA, Mukai S, Diller L, Hamerly G, Shaw BF. Autonomous early detection of eye disease in childhood photographs. Sci Adv. 2019;5:eaax6363. 10.1126/sciadv.aax6363 . Aoto SM, Ueno-Yokohata H, Ueda A, Igarashi M, Ito Y, Tsukamoto M, et al. Collection of 2429 constrained headshots of 277 volunteers for deep learning. Sci Rep. 2022;12:3730. 10.1038/s41598-022-07843-0 . American Academy of Pediatrics, Section on Ophthalmology. Red reflex examination in neonates, infants, and children. Pediatrics. 2008;122:1401–4. 10.1542/peds.2008-1873 . Sun M, Zhang Y, Wang X, Chen Q, Zhou Y, Huang W. Sensitivity and specificity of red reflex test in newborn eye screening. J Pediatr. 2016;179:192–6. 10.1016/j.jpeds.2016.07.016 . Subhi Y, Singh A, Zheng Y, Wei Y, Liu C, Lin X. Diagnostic test accuracy of the red reflex test for ocular pathology in infants: a meta-analysis. JAMA Ophthalmol. 2021;139:33–40. 10.1001/jamaophthalmol.2020.4503 . Rose L, Merriman M, Warren R, Clarke C, Swan K, Liu S. Detection of anomalies in the red reflex test requires adequate training. Clin Exp Optom. 2020;103:1048–54. 10.1111/cxo.13129 . Liu X, Jiang Z, Zhang K, Long E, Cui J, Zhu M, et al. Localization and diagnosis framework for pediatric cataracts based on slit-lamp images using deep features of a convolutional neural network. PLoS ONE. 2017;12:e0168606. 10.1371/journal.pone.0168606 . Kitaguchi Y, Yoshikawa M, Fujino Y, Fujimoto T, Ueda K, Suzuki Y, et al. Deep-learning approach to detect childhood glaucoma based on periocular photograph. Sci Rep. 2023;13:10141. 10.1038/s41598-023-37217-2 . Takeda Y, Kaneko Y, Sugimoto M, Yamashita H, Sasaki A, Mitsui T. Prediction models for retinopathy of prematurity using non-imaging machine learning approaches: a regional multicenter study. Ophthalmol Sci. 2025;5:100715. **[DOI not yet assigned]**. Kamei T, Miyake M, Sado K, Morino K, Mori Y, Tabara Y, et al. Association between the retinal age gap and systemic diseases in the Japanese population: the Nagahama Study. Jpn J Ophthalmol. 2025;69:616–23. **[DOI not yet available]**. Additional Declarations Competing interest reported. KO has received funding from the Canon Foundation. All other authors declare that they have no competing interests. Cite Share Download PDF Status: Published Journal Publication published 10 Apr, 2026 Read the published version in BMC Ophthalmology → Version 1 posted Editorial decision: Revision requested 05 Jan, 2026 Reviews received at journal 03 Jan, 2026 Reviews received at journal 22 Dec, 2025 Reviewers agreed at journal 17 Dec, 2025 Reviewers agreed at journal 14 Dec, 2025 Reviewers agreed at journal 12 Dec, 2025 Reviewers invited by journal 11 Dec, 2025 Editor invited by journal 08 Dec, 2025 Editor assigned by journal 07 Dec, 2025 Submission checks completed at journal 07 Dec, 2025 First submitted to journal 06 Dec, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8295441","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":560035682,"identity":"5c214142-4ad2-41fc-aa14-a1bfe495c989","order_by":0,"name":"Shion Hayashi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/klEQVRIiWNgGAWjYDACCcYGIHkAxGRjSKgAUszMDaRoOQPSwkhIC5iEamFsA9EEtMjPbm58XMBwR97g+PFnDx7Oq43mbwdq+VGxDacWxjkHm41nMDwz3HAmx9wgcdvx3BmHGRsYe87cxqmFWSKxTZqH4TDjhgM5bBKJ247lNgC1MDO24dbCBtViv+H882cSiXOO5c4npIUHqiVxw40EM4nEhprcDYS0SEgkNhvzGBxOnnnjjZlEwrEDuRuBWg7i84v8jPSHj3kqDtv2nU9/Jvmjpi533vnDBx/8qMCtBQIMGBgUDoBZh8HkAQLqodY1gKk6ohSPglEwCkbByAIAj9pfO+ZsaQ0AAAAASUVORK5CYII=","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":true,"prefix":"","firstName":"Shion","middleName":"","lastName":"Hayashi","suffix":""},{"id":560035683,"identity":"2bf3f4a3-55c2-464a-85c0-09b6b11a37ed","order_by":1,"name":"Emi Kashizuka","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Emi","middleName":"","lastName":"Kashizuka","suffix":""},{"id":560035684,"identity":"e6e4ea9d-0848-4104-961b-bc69a7c8498c","order_by":2,"name":"Koki Sakata","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Koki","middleName":"","lastName":"Sakata","suffix":""},{"id":560035687,"identity":"0fdb9f66-d159-49aa-b3f8-9eebab6469e5","order_by":3,"name":"Natsuki Motomiya","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Natsuki","middleName":"","lastName":"Motomiya","suffix":""},{"id":560035688,"identity":"2dc44ac1-9df3-4a35-b5dd-883be99ee3b8","order_by":4,"name":"Yosuke Asakawa","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Yosuke","middleName":"","lastName":"Asakawa","suffix":""},{"id":560035689,"identity":"31489c18-2d49-4be1-b19c-7594b5d2cb94","order_by":5,"name":"Moe Hayasaki","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Moe","middleName":"","lastName":"Hayasaki","suffix":""},{"id":560035690,"identity":"2799a5bf-eb83-42b3-985a-c3fc94965d8d","order_by":6,"name":"Tadashi Yokoi","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Tadashi","middleName":"","lastName":"Yokoi","suffix":""},{"id":560035691,"identity":"b17b5b1f-4a39-426e-8444-b180eaca4427","order_by":7,"name":"Tomoyo Yoshida-Uemura","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Tomoyo","middleName":"","lastName":"Yoshida-Uemura","suffix":""},{"id":560035692,"identity":"11ec14be-25d7-4524-8d67-ad0c4995aed5","order_by":8,"name":"Sachiko Nishina","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Sachiko","middleName":"","lastName":"Nishina","suffix":""},{"id":560035693,"identity":"64b80a00-4eac-4445-93ca-92eb849fb093","order_by":9,"name":"Kohji Okamura","email":"","orcid":"","institution":"National Research Institute for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Kohji","middleName":"","lastName":"Okamura","suffix":""}],"badges":[],"createdAt":"2025-12-06 14:53:07","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8295441/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8295441/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12886-026-04795-9","type":"published","date":"2026-04-10T15:58:44+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":98449436,"identity":"f1e0e4b8-2126-4587-972b-b383da83c60e","added_by":"auto","created_at":"2025-12-17 17:29:30","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":486756,"visible":true,"origin":"","legend":"","description":"","filename":"BMCmanuscript20251202.docx","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/691b101bc65dde5f74d6a1b8.docx"},{"id":98449812,"identity":"5735b3a0-43d8-4a0e-a2d2-b2774f6c8e93","added_by":"auto","created_at":"2025-12-17 17:30:01","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":11437,"visible":true,"origin":"","legend":"","description":"","filename":"f97fdb21b7e94eb3a798755dd2a22c20.json","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/5699537abeaec425a155e25d.json"},{"id":98449303,"identity":"322bd7a3-bb0d-493d-8ca4-dd7edbd63766","added_by":"auto","created_at":"2025-12-17 17:29:27","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":72175,"visible":true,"origin":"","legend":"","description":"","filename":"f97fdb21b7e94eb3a798755dd2a22c201enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/601e7521aac9dbfcaee43754.xml"},{"id":98449768,"identity":"d1199210-a962-48a6-b4af-e7920095dfb5","added_by":"auto","created_at":"2025-12-17 17:29:58","extension":"jpeg","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":463565,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/843f7d4dd011d220dca3aba6.jpeg"},{"id":98449468,"identity":"9803cb57-d3e5-456f-b882-5912fb7f7877","added_by":"auto","created_at":"2025-12-17 17:29:35","extension":"jpeg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":296497,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/5fe1fe1382950c78b0923ac2.jpeg"},{"id":98450000,"identity":"a12581dc-0b9e-4f54-ae56-f63d16b766fd","added_by":"auto","created_at":"2025-12-17 17:30:08","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":151010,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/eb56f61ed37295b8b3ce08da.png"},{"id":98449216,"identity":"d82e70f6-a9d1-4d9d-b9ba-b7f2d5181a98","added_by":"auto","created_at":"2025-12-17 17:29:18","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":55447,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/e1be22d1a022b8c9377a1201.png"},{"id":98448758,"identity":"aada56dc-91a7-4475-a3ea-63ecf5a99340","added_by":"auto","created_at":"2025-12-17 17:29:02","extension":"xml","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":69782,"visible":true,"origin":"","legend":"","description":"","filename":"f97fdb21b7e94eb3a798755dd2a22c201structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/1a06a22ba189a3682d531463.xml"},{"id":98449599,"identity":"79a61536-3f95-409d-a0a9-8cc0d20ba12a","added_by":"auto","created_at":"2025-12-17 17:29:46","extension":"html","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80424,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/66f74c14fed78b7166a8c579.html"},{"id":98449288,"identity":"a9b484e7-0c0e-4b8a-bc80-c41b927f8b6a","added_by":"auto","created_at":"2025-12-17 17:29:26","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":423054,"visible":true,"origin":"","legend":"\u003cp\u003eSome of the crops used for training and testing the image classification model. Crops were created by extracting approximately square regions around the pupils from the subjects' facial photographs and transforming them into precise squares. The top 8 images are from patients with cataracts, while the bottom 8 images are from control subjects. All images were obtained from different individuals. During training, data augmentation techniques, such as rotation, were applied, whereas during testing, the images were processed directly by the convolutional neural network (CNN) without augmentation.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/351f064bb7170af8e6af66a2.png"},{"id":98449437,"identity":"7e21a19e-d60a-41ce-a399-c006ffddd379","added_by":"auto","created_at":"2025-12-17 17:29:31","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":524360,"visible":true,"origin":"","legend":"\u003cp\u003eReceiver operating characteristic curves for pediatric cataract detection model. All models achieved the area under the receiver operating characteristic curve (AUC) values greater than 0.96 when they were trained for 50 epochs. The classification performance was further evaluated using fold 2 as a representative case by comparing the models trained with 1, 2, 8, and 50 epochs.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/3e9f390e511cedb2a2e05eb9.png"},{"id":106809846,"identity":"d16b1f72-9bbb-4928-8d5e-32ec65a2d2d6","added_by":"auto","created_at":"2026-04-13 16:13:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1690994,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8295441/v1/03b331d0-2264-4c8e-aa8e-d424e904d237.pdf"}],"financialInterests":"Competing interest reported. KO has received funding from the Canon Foundation. All other authors declare that they have no competing interests.","formattedTitle":"Detection of Pediatric Cataracts through Non-invasive Facial Photography using Deep Convolutional Neural Networks for Early Diagnosis","fulltext":[{"header":"Background","content":"\u003cp\u003ePediatric cataracts are one of the leading causes of preventable childhood blindness worldwide. Although the annual incidence is relatively low, it has a profound and lifelong impact on visual development and the quality of life. Early detection and surgical treatment are crucial, as visual outcomes are significantly better when surgery for dense unilateral cataracts is performed within the first 6 weeks of life\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e and for bilateral cataracts, within the first 2\u0026ndash;3 months\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Delayed detection leads to irreversible visual impairment, resulting in a substantial socioeconomic burden for both affected individuals and society.\u003c/p\u003e \u003cp\u003eIn Japan's multicenter studies, early surgery has been demonstrated to be essential for a favorable visual prognosis in congenital and developmental cataracts, emphasizing the clinical importance of prompt diagnosis and treatment\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Similarly, Oshika et al. demonstrated that surgery performed within the critical period\u0026mdash;specifically within 10 weeks for bilateral cases and 6 weeks for unilateral cases\u0026mdash;resulted in superior long-term visual outcomes compared with later intervention\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. These findings from both international and Japanese studies underscore the universal importance of detecting pediatric cataracts as early as possible to prevent irreversible visual loss.\u003c/p\u003e \u003cp\u003eIn Japan, infant visual screening is performed as part of routine infant health checkups, and pediatricians play a central role in the detection of ocular abnormalities during these examinations. However, the detection of pediatric cataracts during routine health checkups remains challenging. According to a previous nationwide survey, only 12% of pediatric cataract cases were detected during infant screenings, indicating that most cases are overlooked until the later stages of disease progression\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. This delay is partly because congenital cataracts are rare, limiting opportunities for general pediatricians to gain the necessary diagnostic experience. Furthermore, no convenient and effective detection method has been established for widespread use in primary care settings.\u003c/p\u003e \u003cp\u003eRecent advances in artificial intelligence (AI) and deep learning have opened new possibilities for pediatric ocular screenings. Munson et al. developed the CRADLE smartphone application, which automatically detects leukocoria (white pupillary reflex) from casual childhood photographs. In a longitudinal study, CRADLE identified leukocoria an average of 1.3 years before clinical diagnosis in 80% of children with eye diseases such as retinoblastoma, Coats\u0026rsquo; disease, and pediatric cataracts\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. These findings demonstrate the potential of AI to complement conventional screening by utilizing readily available digital images.\u003c/p\u003e \u003cp\u003eBuilding on this concept, we hypothesized that deep learning models could identify pediatric cataracts directly from infant facial photographs, providing a noninvasive, low-cost, and scalable screening tool\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. This study aimed to develop and evaluate a deep learning model capable of detecting pediatric cataracts using ordinary facial photographs of infants.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Design and Ethics Approval\u003c/h2\u003e \u003cp\u003eThis prospective observational study was conducted at the Division of Ophthalmology, National Center for Child Health and Development (NCCHD), Tokyo, Japan. The study protocol was reviewed and approved by the Ethics Committee of the NCCHD (approval number: 2021\u0026thinsp;\u0026minus;\u0026thinsp;141). Written informed consent was obtained from the guardians of all participants before enrollment in accordance with the Declaration of Helsinki.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eSelection of Pediatric Study Subjects\u003c/h3\u003e\n\u003cp\u003eThe participants were pediatric patients who visited the Division of Ophthalmology at the NCCHD between November 2021 and December 2024. The inclusion criteria were as follows: 1) age between 1 and 36 months at the time of image acquisition and 2) guardians who provided written informed consent for participation. The exclusion criteria were photographs of insufficient quality, as detailed in the following section, patients with incomplete clinical data, and those who refused consent. Congenital or developmental cataracts were diagnosed by a board-certified pediatric ophthalmologist based on slit-lamp biomicroscopy performed after pharmacologic mydriasis.\u003c/p\u003e\n\u003ch3\u003eImage Acquisition in Pediatric Ophthalmology Outpatient Clinics\u003c/h3\u003e\n\u003cp\u003eMultiple facial photographs were taken for each subject using a Canon single-lens reflex (SLR) EOS 80D camera with flash illumination. The photographs included both eyes and the surrounding periocular areas. Images were excluded if they were out of focus, showed closed eyes, contained excessive hair, or had obstructive elements such as hands or motion artifacts. The final dataset consisted of only images that met the quality standards suitable for analysis.\u003c/p\u003e\n\u003ch3\u003eImage Processing and Dataset Construction\u003c/h3\u003e\n\u003cp\u003eFacial photographs were collected from patients with pediatric cataracts and cataract-free control subjects. The control group comprised pediatric subjects with clear ocular media without cataracts or other opacities. Cataract patients could be confidently diagnosed based on slit-lamp examinations. All photographs underwent a standardized preprocessing pipeline before use in model development. First, the facial images were saved in JPEG format and rotated as necessary to ensure an upright orientation. This adjustment was performed using GIMP version 2.10 or ImageMagick version 7.1. Next, regions covering one eye, including the eyelid area, were manually annotated using Microsoft VoTT version 2.1. Although these images will be subsequently processed into exact squares, the current annotation process involves cropping regions that are roughly square-shaped but not precisely so. Based on the pixel coordinates in the resulting JSON file, each annotated region was automatically cropped using a custom Python script. To prevent model overestimation and avoid data leakage during cross-validation, images from the same subject were kept within either the training or test set, ensuring that data from individual patients were not divided between the sets.\u003c/p\u003e\n\u003ch3\u003eDevelopment of CNN-Based Classification Model\u003c/h3\u003e\n\u003cp\u003eA convolutional neural network (CNN) was implemented using the Inception V3 architecture within the PyTorch version 2.7.0 framework. Transfer learning was applied using weights pre-trained on the ImageNet 1k v1 dataset to leverage the prior knowledge of general image features. To improve generalizability and reduce overfitting, various data augmentation techniques were applied using the torchvision library. These included random affine transformations, random erasing, random perspective transformations, color jittering, random auto-contrast adjustments, random rotations, and random vertical flips.\u003c/p\u003e \u003cp\u003eThe CNN was trained for 50 epochs using stochastic gradient descent (SGD) optimization with a learning rate of 0.03, momentum of 0.9, weight decay of 10\u003csup\u003e\u0026minus;\u0026thinsp;4\u003c/sup\u003e, and a batch size of 64. All deep learning computations were performed on an Ubuntu 22.04.5 system configured with one NVIDIA Quadro GV100 graphics processing unit.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eEvaluation of Classification Performance\u003c/h2\u003e \u003cp\u003eThe trained model was evaluated using 5-fold cross-validation to ensure robustness to data bias. Random allocation was performed at the patient level rather than the image level, with each fold in the test dataset containing data from six or seven patients in the fold. Among the five evaluations, we report the most conservative results, presenting the lowest area under the receiver operating characteristic curve (AUC) values as representative performance metrics. Performance was assessed using standard classification metrics, including AUC, sensitivity, specificity, and F\u003csub\u003e1\u003c/sub\u003e score. The classification performance of the deep learning model was compared with the diagnostic assessments of pediatric ophthalmologists using McNemar\u0026rsquo;s test or the chi-square test. Statistical significance was set at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003ePatient Characteristics\u003c/h2\u003e \u003cp\u003eA total of 32 patients with pediatric cataracts and 93 cataract-free control subjects were included in the study. A total of 390 facial photographs were collected, consisting of 98 images from 32 patients (47 eyes) with pediatric cataract and 292 images from 93 cataract-free control subjects (186 eyes).\u003c/p\u003e \u003cp\u003eAmong the 47 eyes with pediatric cataracts, the distribution of cataract morphologies was as follows: 13 with fetal nuclear cataracts, 11 with anterior polar and subcapsular cataracts, 8 with posterior subcapsular cataracts, 6 with total cataracts, 5 with membranous cataracts, 2 with lamellar cataracts, and 2 with combined fetal nuclear and posterior subcapsular cataracts. Regarding ocular comorbidities, 5 eyes had microcornea, 3 eyes had persistent fetal vasculature (PFV), 1 eye had microphthalmia, and 1 eye had optic disc dysplasia. Some eyes exhibited more than one anomaly (overlapping cases), and 38 eyes exhibited no additional ocular abnormalities. Regarding clinical outcomes, 23 patients (33 eyes) underwent cataract surgery, 5 patients (10 eyes) were followed conservatively, and 4 patients (4 eyes) were deemed ineligible for surgical intervention because postoperative visual improvement was not expected.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eDataset Preparation and Cross-validation\u003c/h2\u003e \u003cp\u003eAfter preprocessing and cropping of single-eye regions, 727 cropped images were generated (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The dataset consisted of 149 crops from pediatric cataract patients and 578 crops from controls, with some subjects providing data from only one eye each. We implemented deep learning using an Inception V3-based convolutional neural network architecture. Although the training and test sets were randomly assigned, we conducted 5-fold cross-validation to ensure that our results were free from selection bias and reproducible across different data partitions. Cross-validation was performed using subject-based partitioning rather than image-based partitioning to ensure that data from the same individual were not split between the training and test sets. Consequently, the number of image crops varied across the folds owing to the different numbers of images per subject. Table\u0026nbsp;1 summarizes the distribution of the number of patients and cropped images used for training and testing in each randomly assigned fold.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"6\" nameend=\"c6\" namest=\"c1\"\u003e \u003cp\u003eTable\u0026nbsp;1. Numbers of subjecs and crops in the 5-fold cross-validation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eFold\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eNumber of training subjects\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eNumber of test subjects\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e \u003cp\u003eNumber of training crops\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c9\" namest=\"c8\"\u003e \u003cp\u003eNumber of test crops\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCataract\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eControl\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCataract\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eControl\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eCataract\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eControl\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCataract\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eControl\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e130\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e443\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e135\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e469\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e109\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e127\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e457\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e121\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e125\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e472\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e106\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e120\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e471\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e107\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eModel Performance\u003c/h2\u003e \u003cp\u003eWe implemented a binary classification system for pediatric cataract detection using convolutional neural network architecture. The network was based on Inception V3 pre-trained on ImageNet, with transfer learning applied for our specific task, as described in detail in the Methods section of this paper. The performance was evaluated using five-fold cross-validation, which yielded five sets of results. The distribution of data across these folds showed no significant imbalances, as presented in Table\u0026nbsp;1. To investigate the effect of training duration on model performance, we analyzed fold 2 by training separate models with varying numbers of epochs (1, 2, 8, and 50) and visualized their performance using ROC curves (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The logits from the CNN were transformed into probability values using the softmax function, and classification was performed using a threshold of 0.50. The resulting sensitivity, specificity, and other values for each fold are presented in Table\u0026nbsp;2. Despite the relatively limited size of our dataset, even in fold 3, which showed the lowest AUC of 0.9692, the model achieved high-performance metrics with an accuracy of 0.9720, sensitivity of 0.8636, specificity of 0.9917, and F\u003csub\u003e1\u003c/sub\u003e score of 0.9048. These results suggest that the model demonstrates sufficient diagnostic accuracy for potential clinical use.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabb\" border=\"1\"\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eTable\u0026nbsp;2. Performance metrics for 5-fold cross-validation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFold\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eF\u003csub\u003e1\u003c/sub\u003e score\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9740\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8947\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9852\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9957\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.8947\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9939\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9818\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9908\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9720\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8636\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9917\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9692\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9048\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9692\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8333\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9792\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9091\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.9926\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9907\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.0000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.9831\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eAUC\u003c/em\u003e area under the receiver operating characteristic curve\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we developed and evaluated a deep learning\u0026ndash;based model to detect pediatric cataracts using facial photographs of infants and toddlers. The model was based on the Inception V3 convolutional neural network architecture, pretrained on ImageNet, and fine-tuned for our specific task. We implemented comprehensive data augmentation techniques, including random rotation, vertical flipping, brightness adjustment, and contrast validation, to enhance the model robustness and prevent overfitting. Despite the relatively small dataset, our model achieved high classification performance across five-fold cross-validation, with AUC values consistently exceeding 0.96. These results demonstrate that deep learning models can effectively identify cataract-related signs in facial photographs, even in young children who are often uncooperative during traditional screening procedures.\u003c/p\u003e \u003cp\u003eEarly detection of pediatric cataracts is essential to prevent irreversible visual loss in children. However, current screening strategies, including the red reflex test (RRT), have several limitations. The RRT is simple and inexpensive\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e, but its performance varies significantly depending on the examiner\u0026rsquo;s experience and lesion characteristics. In large-scale prospective studies, its overall sensitivity for ocular pathologies has been reported to be as low as 7\u0026ndash;14%, despite a high specificity exceeding 95%\u003csup\u003e10,11\u003c/sup\u003e. Even after standardized training, sensitivity among non-specialist examiners, such as midwives, remains moderate (approximately 50%) compared with optometry students or ophthalmologists\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. These limitations make widespread implementation of the RRT challenging, especially in primary care settings. In Japan, RRT is not routinely included in the national infant health checkup program, leaving a gap in the early detection of pediatric ocular diseases.\u003c/p\u003e \u003cp\u003eSeveral studies have explored the application of AI in pediatric ocular diseases, providing valuable context for our work. Liu et al. proposed a deep learning framework using slit-lamp images to diagnose pediatric cataracts with high accuracy\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e; however, slit-lamp photography requires specialized ophthalmic equipment and trained personnel, limiting its feasibility for routine screenings. Munson et al. demonstrated that AI can detect leukocoria from casual childhood photographs taken at home using a smartphone\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e, identifying eye diseases such as retinoblastoma and Coats\u0026rsquo; disease years before a clinical diagnosis. Recently, Kitaguchi et al. successfully applied deep learning to detect childhood glaucoma using periocular facial photographs, highlighting the potential of non-specialized images for screening other pediatric eye conditions\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. Recent studies in Japan have illustrated the growing trend of integrating AI into ophthalmic screening and public health frameworks. Takeda et al. developed machine learning\u0026ndash;based models for predicting retinopathy of prematurity using routinely available clinical data without relying on an expensive fundus camera\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. Similarly, Kamei et al. demonstrated that AI analysis of retinal photographs could estimate the \u0026ldquo;retinal age gap,\u0026rdquo; linking ocular findings with systemic health indicators\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. Together, these studies reflect the increasing use of AI to support data-driven health screenings and disease prevention.\u003c/p\u003e \u003cp\u003eTo our knowledge, this is the first study to demonstrate that pediatric cataracts can be detected using ordinary facial photographs taken under non-specialized conditions. This approach has several potential benefits. First, it does not require any specialized ophthalmic equipment or examiner expertise, allowing screening in primary care settings or even at home. Second, it can be integrated into existing infant health checkup systems in Japan, enhancing the detection rate of ocular diseases that are often missed by conventional methods. Furthermore, automated image screening applications could be incorporated into parental smartphone platforms, enabling opportunistic screening using photos already available in daily life.\u003c/p\u003e \u003cp\u003eHowever, several limitations must be acknowledged. First, the dataset used in this study was relatively small and collected from a single institution, which may limit its generalizability. Expanding the dataset to include images from multiple facilities and different ethnic backgrounds is essential for improving model robustness. Second, our study used cropped eye images derived from facial photographs, which may not fully reflect the variability encountered in real-world settings. Future work should focus on developing algorithms capable of automatically detecting the eye region from unprocessed facial images. Third, although the model demonstrated high diagnostic accuracy in differentiating between clear cataracts and normal eyes, its performance in detecting mild or subtle cataracts requires further investigation.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn conclusion, this study demonstrated the feasibility of using a deep learning model to detect pediatric cataracts from infant facial photographs. The proposed method offers a noninvasive, low-cost, and scalable approach that can complement the conventional red reflex test in pediatric vision screening. With further validation and large-scale multicenter studies, such AI-based tools have the potential to significantly improve the early detection and management of pediatric cataracts, ultimately contributing to better lifelong visual outcomes in children.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAUC: Area under the receiver operating characteristic curve\u003c/p\u003e\n\u003cp\u003eAI: Artificial intelligence\u003c/p\u003e\n\u003cp\u003eCNN: Convolutional neural network\u003c/p\u003e\n\u003cp\u003eDL: Deep learning\u003c/p\u003e\n\u003cp\u003eNCCHD: National Center for Child Health and Development\u003c/p\u003e\n\u003cp\u003ePFV: Persistent fetal vasculature\u003c/p\u003e\n\u003cp\u003eRRT: Red reflex test\u003c/p\u003e\n\u003cp\u003eSD: Standard deviation\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study protocol was approved by the Ethics Committee of the National Center for Child Health and Development (NCCHD), Tokyo, Japan (approval no. 2021-141). Written informed consent was obtained from the guardians of all participants in accordance with the Declaration of Helsinki.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWritten informed consent for publication of anonymized images and research data was obtained from the guardians of all participating children.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and analyzed during the current study are not publicly available due to ethical restrictions regarding patient privacy, but are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eKO has received funding from the Canon Foundation. All other authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the following funding sources:\u003c/p\u003e\n\u003cp\u003eNCCHD Institutional Grant (Grant No. 2023B-6) to SH and KO\u003c/p\u003e\n\u003cp\u003eHealth Labour Sciences Research Grant (No. 23AC1001) to KO\u003c/p\u003e\n\u003cp\u003eHealth and Labour Sciences Research Grants for Research on Intractable Diseases (No. 23FC1055) to SN\u003c/p\u003e\n\u003cp\u003eJSPS KAKENHI (Grant No. 21K06143) to KO (for GPU-equipped system)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSH, SN, EK, and KO: Study design and conceptualization.\u003c/p\u003e\n\u003cp\u003eSH and KO: Drafting the manuscript.\u003c/p\u003e\n\u003cp\u003eKO and SN: Critical revision of the manuscript.\u003c/p\u003e\n\u003cp\u003eSH and KO: Literature review.\u003c/p\u003e\n\u003cp\u003eEK and KO: Data analysis.\u003c/p\u003e\n\u003cp\u003eSH, EK, KS, NM, YA, and MH: Data collection.\u003c/p\u003e\n\u003cp\u003eSH, KS, NM, YA, MH, YT, TY, and NS: Clinical recruitment.\u003c/p\u003e\n\u003cp\u003eAll authors reviewed and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe are grateful to all participants and their families for contributing to this study. We thank Ms. Miki Nobe for her assistance with manual data preparation, including image annotation. SH sincerely thanks Professor Masahiko Sugimoto (Department of Ophthalmology, Yamagata University) for his guidance and continuous support throughout this work.\u003c/p\u003e\n\u003cp\u003eThe authors also used ChatGPT (OpenAI, San Francisco, CA, USA) and Paperpal (Iris.ai, Norway) to improve the clarity and readability of the English manuscript; all scientific content and interpretations are the sole responsibility of the authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; information (optional)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBirch EE, Stager DR. The critical period for surgical treatment of dense congenital unilateral cataract. Invest Ophthalmol Vis Sci. 1996;37:1532\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLambert SR, Lynn MJ, Reeves R, Plager DA, Buckley EG, Wilson ME. Is there a latent period for the surgical treatment of children with dense bilateral congenital cataracts? J AAPOS. 2006;10:30\u0026ndash;6. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jaapos.2005.10.002\u003c/span\u003e\u003cspan address=\"10.1016/j.jaapos.2005.10.002\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBirch EE, Cheng C, Stager DR Jr, Weekley DR Jr, Stager DR Sr. The critical period for surgical treatment of dense congenital bilateral cataracts. J AAPOS. 2009;13:67\u0026ndash;71. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jaapos.2008.07.010\u003c/span\u003e\u003cspan address=\"10.1016/j.jaapos.2008.07.010\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNagamoto T, Oshika T, Fujikado T, Ishibashi T, Sato M, Kondo M, et al. Surgical outcomes of congenital and developmental cataracts in Japan. Jpn J Ophthalmol. 2016;60:127\u0026ndash;34. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10384-015-0498-1\u003c/span\u003e\u003cspan address=\"10.1007/s10384-015-0498-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOshika T, Nishina S, Unoki N, Miyagi M, Nomura K, Mori T, et al. Ten-year outcomes of congenital cataract surgery performed within the first six months of life. J Cataract Refract Surg. 2024;50:707\u0026ndash;12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/j.jcrs.0000000000001189\u003c/span\u003e\u003cspan address=\"10.1097/j.jcrs.0000000000001189\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNishina S, Yokoi T, Kobayashi Y, Noda E, Azuma N. The time and route for finding infants' eye disease. Folia Japonica de Ophthalmol Clin. 2010;3:172\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMunson MC, Plewman DL, Baumer KM, Henning R, Zahler CT, Kietzman AT, Beard AA, Mukai S, Diller L, Hamerly G, Shaw BF. Autonomous early detection of eye disease in childhood photographs. Sci Adv. 2019;5:eaax6363. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1126/sciadv.aax6363\u003c/span\u003e\u003cspan address=\"10.1126/sciadv.aax6363\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAoto SM, Ueno-Yokohata H, Ueda A, Igarashi M, Ito Y, Tsukamoto M, et al. Collection of 2429 constrained headshots of 277 volunteers for deep learning. Sci Rep. 2022;12:3730. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-022-07843-0\u003c/span\u003e\u003cspan address=\"10.1038/s41598-022-07843-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmerican Academy of Pediatrics, Section on Ophthalmology. Red reflex examination in neonates, infants, and children. Pediatrics. 2008;122:1401\u0026ndash;4. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1542/peds.2008-1873\u003c/span\u003e\u003cspan address=\"10.1542/peds.2008-1873\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun M, Zhang Y, Wang X, Chen Q, Zhou Y, Huang W. Sensitivity and specificity of red reflex test in newborn eye screening. J Pediatr. 2016;179:192\u0026ndash;6. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jpeds.2016.07.016\u003c/span\u003e\u003cspan address=\"10.1016/j.jpeds.2016.07.016\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubhi Y, Singh A, Zheng Y, Wei Y, Liu C, Lin X. Diagnostic test accuracy of the red reflex test for ocular pathology in infants: a meta-analysis. JAMA Ophthalmol. 2021;139:33\u0026ndash;40. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamaophthalmol.2020.4503\u003c/span\u003e\u003cspan address=\"10.1001/jamaophthalmol.2020.4503\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRose L, Merriman M, Warren R, Clarke C, Swan K, Liu S. Detection of anomalies in the red reflex test requires adequate training. Clin Exp Optom. 2020;103:1048\u0026ndash;54. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/cxo.13129\u003c/span\u003e\u003cspan address=\"10.1111/cxo.13129\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu X, Jiang Z, Zhang K, Long E, Cui J, Zhu M, et al. Localization and diagnosis framework for pediatric cataracts based on slit-lamp images using deep features of a convolutional neural network. PLoS ONE. 2017;12:e0168606. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0168606\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0168606\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKitaguchi Y, Yoshikawa M, Fujino Y, Fujimoto T, Ueda K, Suzuki Y, et al. Deep-learning approach to detect childhood glaucoma based on periocular photograph. Sci Rep. 2023;13:10141. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-023-37217-2\u003c/span\u003e\u003cspan address=\"10.1038/s41598-023-37217-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakeda Y, Kaneko Y, Sugimoto M, Yamashita H, Sasaki A, Mitsui T. Prediction models for retinopathy of prematurity using non-imaging machine learning approaches: a regional multicenter study. Ophthalmol Sci. 2025;5:100715. **[DOI not yet assigned]**.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKamei T, Miyake M, Sado K, Morino K, Mori Y, Tabara Y, et al. Association between the retinal age gap and systemic diseases in the Japanese population: the Nagahama Study. Jpn J Ophthalmol. 2025;69:616\u0026ndash;23. **[DOI not yet available]**.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-ophthalmology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"boph","sideBox":"Learn more about [BMC Ophthalmology](http://bmcophthalmol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/boph","title":"BMC Ophthalmology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"deep learning, image classification, pediatric cataract, facial photograph, infancy","lastPublishedDoi":"10.21203/rs.3.rs-8295441/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8295441/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground:\u003c/h2\u003e \u003cp\u003eTo develop a deep learning\u0026ndash;based model for detecting pediatric cataracts using noninvasive facial photographs of infants and toddlers, aiming to facilitate early diagnosis during the critical period of visual development.\u003c/p\u003e\u003ch2\u003eMethods:\u003c/h2\u003e \u003cp\u003eThis prospective observational study included 32 patients (47 eyes) with pediatric cataracts and 93 cataract-free controls (186 eyes) who visited the National Center for Child Health and Development between November 2021 and December 2024. Multiple facial photographs were captured using a digital single-lens reflex camera with flash illumination. After preprocessing and cropping of single-eye regions, 727 cropped images (149 cataract and 578 control images) were used to train and validate a convolutional neural network based on the Inception V3 architecture. The model was trained using transfer learning and evaluated using five-fold cross-validation. The diagnostic performance was assessed using sensitivity, specificity, accuracy, area under the receiver operating characteristic curve (AUC), and F\u003csub\u003e1\u003c/sub\u003e score.\u003c/p\u003e\u003ch2\u003eResults:\u003c/h2\u003e \u003cp\u003eThe model demonstrated a high diagnostic performance across all folds. The minimum AUC among the five folds was 0.9692, with an accuracy of 0.9720, sensitivity of 0.8636, specificity of 0.9917, and F\u003csub\u003e1\u003c/sub\u003e score of 0.9048. Despite the relatively small dataset, the model consistently achieved robust results without overfitting.\u003c/p\u003e\u003ch2\u003eConclusions:\u003c/h2\u003e \u003cp\u003eThe proposed deep learning model accurately detected pediatric cataracts in ordinary facial photographs of infants. This noninvasive, low-cost approach may complement conventional screening and improve the early detection of pediatric cataracts, enabling timely referral and treatment during the critical period of visual development.\u003c/p\u003e","manuscriptTitle":"Detection of Pediatric Cataracts through Non-invasive Facial Photography using Deep Convolutional Neural Networks for Early Diagnosis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-17 16:54:43","doi":"10.21203/rs.3.rs-8295441/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-01-05T06:10:10+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-03T13:22:22+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-22T07:58:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"4284274111525987303375629647562498014","date":"2025-12-17T06:20:11+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"207157225825071791351123010188022970182","date":"2025-12-14T11:50:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"230531612762710468179038128403021243876","date":"2025-12-12T19:41:59+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-12-11T11:18:16+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-12-08T10:32:23+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-12-08T04:16:04+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-12-08T04:14:44+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Ophthalmology","date":"2025-12-06T14:35:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-ophthalmology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"boph","sideBox":"Learn more about [BMC Ophthalmology](http://bmcophthalmol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/boph","title":"BMC Ophthalmology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"eb856f84-a7ca-4fb0-b7ce-abbe6bffaf17","owner":[],"postedDate":"December 17th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2026-04-13T16:08:01+00:00","versionOfRecord":{"articleIdentity":"rs-8295441","link":"https://doi.org/10.1186/s12886-026-04795-9","journal":{"identity":"bmc-ophthalmology","isVorOnly":false,"title":"BMC Ophthalmology"},"publishedOn":"2026-04-10 15:58:44","publishedOnDateReadable":"April 10th, 2026"},"versionCreatedAt":"2025-12-17 16:54:43","video":"","vorDoi":"10.1186/s12886-026-04795-9","vorDoiUrl":"https://doi.org/10.1186/s12886-026-04795-9","workflowStages":[]},"version":"v1","identity":"rs-8295441","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8295441","identity":"rs-8295441","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.