Explainable CBCT based morphometric clustering for orthodontic diagnostics using principal component analysis and k means | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Explainable CBCT based morphometric clustering for orthodontic diagnostics using principal component analysis and k means Salah M. Bin Hafedh, Ramy Abdulrahman Ishaq, Fuad Lutf Al Motareb This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8809759/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 12 You are reading this latest preprint version Abstract Background : Cone‑beam computed tomography (CBCT) enables three‑dimensional assessment of craniofacial structures, yet interpretation of many inter‑correlated measurements remains challenging. Objective : To present an explainable workflow combining principal component analysis (PCA) and k‑means clustering to identify skeletal Class II phenotypes. Methods : Sixty‑three CBCT‑derived variables from 120 Yemeni adults with skeletal Class II malocclusion were standardized, reduced via PCA, and clustered (k = 2–8). Internal validity was evaluated using silhouette and Davies–Bouldin indices with multiple random initializations. Results : Seven principal components explained approximately 60% of variance; a five‑cluster solution balanced cohesion, separation, and clinical interpretability, delineating deep‑bite and open‑bite tendencies, mandibular retrusion patterns, and incisor‑protrusive phenotypes. Conclusions : The PCA→k‑means pipeline provides a transparent, reproducible framework for artificial intelligence (AI)-assisted orthodontic diagnosis and phenotype‑informed treatment planning. Trial registration: Not applicable (retrospective observational study; no prospective assignment to interventions). Cone‑beam computed tomography three‑dimensional imaging principal component analysis k‑means clustering phenotyping orthodontic diagnosis explainable AI Figures Figure 1 Figure 2 Figure 3 1. Introduction CBCT has expanded orthodontic diagnostics by providing isotropic three‑dimensional datasets suitable for craniofacial assessment and quantitative measurement [ 7 – 9 ]. However, CBCT‑derived morphometric datasets are typically high‑dimensional and inter‑correlated, which can complicate clinical interpretation when dozens of variables are considered simultaneously. Data‑driven craniofacial phenotyping seeks to transform multivariate morphology into clinically interpretable subgroups. In orthodontics, three‑dimensional morphometric analyses—including geometric morphometrics—have been promoted to capture variation in form beyond traditional two‑dimensional summaries and to support modern phenotype concepts [ 21 ]. Cluster‑based phenotyping is also increasingly used to reveal within‑class heterogeneity (e.g., distinct vertical and dentoalveolar compensations within sagittal class labels) that may inform treatment strategy selection [ 22 , 23 ]. Unsupervised machine learning is especially appropriate when the primary goal is discovery of latent phenotypic structure, rather than prediction against predefined labels. PCA reduces dimensionality by transforming correlated variables into orthogonal components [ 1 , 2 ], while k‑means clustering partitions individuals into centroid‑defined prototypes that can be translated into diagnostic “templates” [ 6 , 11 , 12 ]. Compared with deep learning approaches that often require large labeled datasets and may provide limited transparency, PCA‑based clustering offers direct interpretability through loadings and cluster centroids and can serve as an explainable front‑end to broader AI pipelines [ 16 – 18 , 20 ]. Accordingly, this study presents an explainable PCA→k‑means workflow for identifying skeletal Class II phenotypes using CBCT‑derived variables, emphasizing transparent reporting, internal validation, and clinically meaningful translation. 2. Materials and methods Study design and ethics: This retrospective observational study evaluated CBCT records of Yemeni adults diagnosed with skeletal Class II malocclusion. The study was conducted in accordance with the Declaration of Helsinki and was approved by the Medical Ethics Committee of Sana’a University, Yemen (Approval ID: OR 19/11/2023), which granted a waiver of informed consent because the study used retrospective, de-identified CBCT records. Participants: CBCT records meeting predefined inclusion criteria (adult dentition; diagnostic indication consistent with skeletal Class II) and excluding syndromic conditions, craniofacial trauma, and prior orthognathic surgery were screened. Records with poor image quality or artifacts affecting landmarking were excluded. CBCT acquisition and measurements: Scans were acquired using a PaX‑Flex 3D P2 system with a medium field of view. Measurements were extracted in InVivo 6.0 using standardized protocols. A total of 63 variables were recorded, spanning sagittal and vertical skeletal relationships, cranial‑base dimensions, and dentoalveolar parameters including maxillary and mandibular incisor positions and inclinations [ 7 – 9 , 13 – 15 ]. Measurement domains are summarized in Table 1 . [Insert Table 1 here] Preprocessing and PCA: Variables were standardized using z‑scores. PCA was performed on the correlation matrix to reduce dimensionality and mitigate multicollinearity [ 1 – 3 ]. Component retention considered eigenvalues, the scree profile, cumulative variance, and anatomical interpretability of loadings. The overall AI–CBCT workflow is illustrated in Fig. 1 . [Insert Fig. 1 here] k‑means clustering and selection of k: k‑means clustering was applied to retained PCA scores with k ranging from 2 to 8 [ 6 , 11 , 12 ]. Multiple random initializations were used to reduce sensitivity to local minima. Internal validation and sensitivity: Silhouette coefficients and Davies–Bouldin indices were used to quantify within‑cluster cohesion and between‑cluster separation [ 4 , 5 ]. Sensitivity analyses examined the influence of the number of retained components and alternative k values, and stability was checked across multiple random starts and sex stratifications. 3. Results Dimensionality reduction: Seven principal components explained approximately 60% of the variance in the 63‑variable CBCT feature space (Fig. 2). The leading components reflected clinically recognizable axes of variation, including cranial‑base dimensions and vertical skeletal pattern (mandibular plane steepness and facial height), with additional components capturing maxillary sagittal position and dentoalveolar contributions. The cumulative and incremental variance explained by retained components is summarized in Table 2. [Insert Table 2 here] Clustering and phenotypes: Based on internal validity indices and clinical interpretability, a five‑cluster solution was selected. The clusters delineated phenotypes consistent with deep‑bite and open‑bite tendencies, mandibular retrusion‑dominant patterns, and incisor‑protrusive compensations; these patterns were mapped to phenotype‑informed orthodontic strategies. A high-level description of the five phenotypic clusters and their potential orthodontic implications is provided in Table 3. [Insert Table 3 here] Visualization: PC score plots (PC1 vs PC2) demonstrated separable but partially overlapping cluster distributions (Fig. 3), consistent with expected craniofacial heterogeneity within skeletal Class II. Tables and figures Table 1. Overview of measurement domains used for phenotyping. Domain (examples) Role in phenotype interpretation Sagittal skeletal relationships Differentiate maxillary vs mandibular contributions and inform sagittal correction strategy Vertical skeletal pattern Identify deep‑bite vs open‑bite tendencies and vertical control needs Cranial‑base dimensions Capture baseline craniofacial size/shape contributions affecting sagittal measures Dentoalveolar inclinations/positions (incisors) Guide torque control, anchorage planning, and camouflage decisions Table 2. Cumulative and incremental variance explained by retained PCA components (from Fig. 2). Component set Cumulative variance (%) Incremental variance (%) PC1–PC1 10.2 10.2 PC1–PC2 19.5 9.3 PC1–PC3 28.6 9.1 PC1–PC4 37.4 8.8 PC1–PC5 46.1 8.7 PC1–PC6 54.3 8.2 PC1–PC7 60.2 5.9 Table 3. High-level description of the five phenotypic clusters and potential orthodontic implications. Cluster tendency (phenotype label) Phenotype‑informed orthodontic considerations (examples) Deep‑bite / low‑angle tendency Bite opening mechanics; torque control; avoid deepening; consider anterior intrusion strategies as indicated Open‑bite / high‑angle tendency Vertical control; minimize extrusive mechanics; evaluate habits/airway factors when relevant Mandibular retrusion‑dominant pattern Assess growth status; camouflage vs orthognathic planning depending on severity and soft‑tissue goals Incisor‑protrusive compensation Anchorage planning; incisor torque control; extraction decisions guided by crowding and profile Mixed/intermediate centroid Individualize using centroid profile as a structured checklist alongside clinical exam and imaging 4. Discussion This study demonstrates an explainable unsupervised workflow that converts a high‑dimensional CBCT measurement set into clinically interpretable skeletal Class II phenotypes. PCA provides decorrelation and dimensionality reduction [ 1 – 3 ], while k‑means clustering yields centroid‑defined prototypes that can be translated into diagnostic templates [ 6 , 11 , 12 ]. Relation to previous morphometric and phenotyping literature: Three‑dimensional morphometric approaches—particularly geometric morphometrics—have been advocated in orthodontics to quantify 3D form and to study craniofacial variation beyond conventional measures [ 21 ]. Similarly, earlier phenotype work in skeletal Class II highlighted substantial heterogeneity within sagittal labels, including vertical and dentoalveolar compensations that influence treatment choice [ 22 ]. Recent studies also demonstrate that data‑driven clustering can classify craniofacial patterns and may complement conventional diagnostics [ 23 ]. Why unsupervised learning and how it complements deep learning: Deep learning has shown strong performance for supervised orthodontic imaging tasks (e.g., landmark detection) but generally requires large labeled datasets and may offer limited transparency for phenotype discovery [ 18 ]. Unsupervised clustering is appropriate when the objective is discovery of latent structure without assuming predefined labels, and when interpretability via component loadings and centroids is required for clinical adoption. Such interpretable clustering can also provide a structured phenotype taxonomy to support subsequent supervised modeling or outcome‑prediction studies, aligning with emerging reporting expectations for AI in dentistry [ 16 , 17 , 20 ]. Clinical relevance and novelty for an orthodontic audience: The clinical value is not to replace established Class II descriptors, but to provide a reproducible multivariate summary tool that supports (i) standardized case documentation, (ii) stratified treatment planning (vertical vs sagittal emphasis; torque/anchorage priorities), and (iii) hypothesis generation for phenotype‑to‑outcome research. Limitations include the single‑population retrospective design and the use of engineered measurement variables rather than full-image radiomics features. Radiomics-style extensions may be feasible but require rigorous standardization and radiation justification [ 19 , 25 ]. Future work should validate phenotypes across centers/devices and link phenotypes to treatment outcomes and stability to demonstrate clinical utility. 5. Conclusions A PCA→k‑means clustering pipeline provides a transparent and reproducible approach to identify skeletal Class II phenotypes from CBCT‑derived variables, supporting explainable AI-assisted orthodontic diagnostics and phenotype‑informed treatment planning. Declarations Ethics approval and consent to participate: The Medical Ethics Committee of Sana’a University, Yemen (Approval ID: OR 19/11/2023) approved the study and waived the requirement for informed consent because the research was retrospective and used de-identified data. Trial registration: Not applicable (retrospective observational study; no prospective assignment to interventions). Consent for publication: Not applicable. Availability of data and materials: Data are available from the corresponding author upon reasonable request, subject to ethical restrictions. Competing interests: The authors declare no competing interests. Funding: No external funding was received. Authors’ contributions: Conceptualization: S.M.B.H.; Methodology: S.M.B.H., R.A.I., F.L.A.; Data curation: S.M.B.H.; Formal analysis and visualization: S.M.B.H.; Writing—original draft: S.M.B.H.; Writing—review and editing: S.M.B.H., R.A.I., F.L.A. All authors read and approved the final manuscript. Acknowledgements: The authors thank the Medical Ethics Committee of Sana’a University for oversight. References Smith LI. A tutorial on principal components analysis. 2002. Available at: http://www.cs.otago.ac.nz/cosc453/student_tutorials/principal_components.pdf. Accessed 2026-01-02. Abdi H, Williams LJ. Principal component analysis. Wiley Interdiscip Rev Comput Stat. 2010;2(4):433–459. https://doi.org/10.1002/wics.101. Hastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning. 2nd ed. Springer; 2009. https://link.springer.com/book/10.1007/978-0-387-84858-7. Rousseeuw PJ. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J Comput Appl Math. 1987;20:53–65. https://doi.org/10.1016/0377-0427(87)90125-7. Davies DL, Bouldin DW. A cluster separation measure. IEEE Trans Pattern Anal Mach Intell. 1979;1(2):224–227. https://doi.org/10.1109/TPAMI.1979.4766909. Hartigan JA, Wong MA. Algorithm AS 136: a k‑means clustering algorithm. J R Stat Soc Ser C Appl Stat. 1979;28(1):100–108. https://doi.org/10.2307/2346830. Scarfe WC, Farman AG. What is cone‑beam CT and how does it work? Dent Clin North Am. 2008;52(4):707–730. https://doi.org/10.1016/j.cden.2008.05.005. Venkatesh E, Elluru SV. Cone beam computed tomography: basics and applications in dentistry. J Istanb Univ Fac Dent. 2017;51(Suppl 1):S102–S121. https://pmc.ncbi.nlm.nih.gov/articles/PMC5750833/. Pauwels R, Beinsberger J, Collaert B, et al. Effective dose range for dental cone beam computed tomography scanners. Eur J Radiol. 2012;81(2):267–271. https://doi.org/10.1016/j.ejrad.2010.11.028. Izenman AJ. Modern Multivariate Statistical Techniques. Springer; 2008. https://link.springer.com/book/10.1007/978-0-387-78189-1. MacQueen J. Some methods for classification and analysis of multivariate observations. In: Proc 5th Berkeley Symp Math Stat Prob. 1967:281–297. https://projecteuclid.org/journals/proceedings-of-the-fifth-berkeley-symposium-on-mathematical-statistics-and-probability/volume-1/issue-1/Some-methods-for-classification-and-analysis-of-multivariate-observations/bsmsp/1200512992.full. Lloyd SP. Least squares quantization in PCM. IEEE Trans Inf Theory. 1982;28(2):129–137. https://doi.org/10.1109/TIT.1982.1056489. Damstra J, Fourie Z, Huddleston Slater JJR, Ren Y. Accuracy of linear measurements from CBCT‑derived 3‑D surface models of different voxel sizes. Am J Orthod Dentofacial Orthop. 2010;137(1):16.e1–16.e6. https://pubmed.ncbi.nlm.nih.gov/20122425/. Gribel BF, Gribel MN, Manzi FR, Brooks SL, McNamara JA Jr. Accuracy and reliability of craniometric measurements on lateral cephalograms vs CBCT. Angle Orthod. 2011;81(1):26–35. https://pubmed.ncbi.nlm.nih.gov/20936948/. Scarfe WC, Aboelsaad N, Farman AG, et al. CBCT in orthodontics. Aust Dent J. 2017;62(Suppl 1):33–50. https://doi.org/10.1111/adj.12479. Schwendicke F, Samek W, Krois J. Artificial intelligence in dentistry: chances and challenges. J Dent Res. 2020;99(7):769–774. https://doi.org/10.1177/0022034520915714. Arsiwala‑Scheppach LT, Chaurasia A, Müller A, Krois J, Schwendicke F. Machine learning in dentistry: a scoping review. J Clin Med. 2023;12(3):937. https://pubmed.ncbi.nlm.nih.gov/36769585/. Schwendicke F, Chaurasia A, Arsiwala‑Scheppach L, et al. Deep learning for cephalometric landmark detection: systematic review and meta‑analysis. Clin Oral Investig. 2021;25:4299–4312. https://doi.org/10.1007/s00784-021-03990-w. Pauwels R. A new formula for converting dose‑area product to effective dose in dental cone‑beam computed tomography. Phys Med. 2023;112:102639. https://doi.org/10.1016/j.ejmp.2023.00116-3. Lakhotia S, et al. Machine learning in dentistry: scoping review. PLOS Digit Health. 2025;4(9):e0000940. https://doi.org/10.1371/journal.pdig.0000940. Ghislanzoni L, et al. Measuring 3D shape in orthodontics through geometric morphometrics. Prog Orthod. 2017;18:38. https://doi.org/10.1186/s40510-017-0194-9. Moreno Uribe LM, et al. Phenotypic diversity in skeletal Class II malocclusion. Orthod Craniofac Res. 2014;17(1):33–45. https://doi.org/10.1111/ocr.12026. de Frutos‑Valle L, et al. Craniofacial classification in skeletal Class III malocclusion using a clustering and deep learning approach. Sci Rep. 2020;10:18870. https://doi.org/10.1038/s41598-020-74488-w. Hou SY, et al. Soft tissue facial changes among adult females during orthodontic alignment: a 3D geometric morphometric study. BMC Oral Health. 2021;21:180. https://doi.org/10.1186/s12903-021-01425-2. Tarakçı ÖD, et al. Radiomics-based diagnosis in dentomaxillofacial radiology: a systematic review. Dentomaxillofac Radiol. 2024;53(7):20230328. https://doi.org/10.1259/dmfr.20230328. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 13 Apr, 2026 Reviews received at journal 16 Mar, 2026 Reviewers agreed at journal 16 Mar, 2026 Reviews received at journal 13 Mar, 2026 Reviewers agreed at journal 07 Mar, 2026 Reviews received at journal 05 Mar, 2026 Reviewers agreed at journal 05 Mar, 2026 Reviewers invited by journal 02 Mar, 2026 Editor invited by journal 02 Mar, 2026 Editor assigned by journal 19 Feb, 2026 Submission checks completed at journal 19 Feb, 2026 First submitted to journal 19 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8809759","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":599218329,"identity":"476caddb-e288-4ccf-874e-0ce6b191167b","order_by":0,"name":"Salah M. Bin Hafedh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAElEQVRIie3RMWrDMBSA4Sc8eHnQjioadAWFgqC0JVexMWRyaKfWo4LBXZRkzUm81kHQKQcwqENNIVOHQBfTZqhCuto4W6H6Qdo+pCcB+Hx/skAFABWECiTshhFyJFiBJKuTSYBDAH9azz7v4ZVjmJfvN18lB5pU0GZlJxGbOGcr2I40vjxeThd2pOgkInpjuwnEiiEY8kxTyabaRuM6FQEpuglfNvm3I2PNPyS7cgTo3a6XQB0Xh1NiTVEyaA8khV4i6qa4RmESjZOHi7lys+BWrPtm4cvEWMzMrQ5NSdu95RAmzVub9Vzs9xGOkcJt55H7puHt3To7Bfh8Pt9/6AeyMljrg5+bbAAAAABJRU5ErkJggg==","orcid":"","institution":"Sana'a University","correspondingAuthor":true,"prefix":"","firstName":"Salah","middleName":"M. Bin","lastName":"Hafedh","suffix":""},{"id":599218331,"identity":"5caff553-03fe-45b8-a63d-8a66b0dbc87b","order_by":1,"name":"Ramy Abdulrahman Ishaq","email":"","orcid":"","institution":"Sana'a University","correspondingAuthor":false,"prefix":"","firstName":"Ramy","middleName":"Abdulrahman","lastName":"Ishaq","suffix":""},{"id":599218332,"identity":"f49cd1e6-cbda-42d2-99be-82d94d9e2a34","order_by":2,"name":"Fuad Lutf Al Motareb","email":"","orcid":"","institution":"Sana'a University","correspondingAuthor":false,"prefix":"","firstName":"Fuad","middleName":"Lutf Al","lastName":"Motareb","suffix":""}],"badges":[],"createdAt":"2026-02-06 17:39:49","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8809759/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8809759/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104402558,"identity":"0949e773-1d2f-405e-af63-2d72e9845f36","added_by":"auto","created_at":"2026-03-11 12:15:43","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":123927,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow of PCA–k‑means phenotyping (AI–CBCT pipeline).\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8809759/v1/8af243c844209f74b33834d4.png"},{"id":104024757,"identity":"9d0a5716-bbf5-4eab-b0f2-4683594afe1e","added_by":"auto","created_at":"2026-03-05 20:02:15","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":474112,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePCA scree plot and cumulative variance (PC1–PC7).\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-8809759/v1/951480294b74dd018cfc0b80.png"},{"id":104024755,"identity":"7c22504f-80ac-4662-bc20-7cb6ff225c43","added_by":"auto","created_at":"2026-03-05 20:02:15","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":248348,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePCA score plot with cluster overlays (PC1 vs PC2); centroids marked.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-8809759/v1/b47049c44e417058d1b2a919.png"},{"id":104408460,"identity":"5d0b706b-56bd-400e-a445-929f5efbccff","added_by":"auto","created_at":"2026-03-11 12:42:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1480582,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8809759/v1/73dd93eb-62b8-49a8-80e4-93d4777b0176.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Explainable CBCT based morphometric clustering for orthodontic diagnostics using principal component analysis and k means","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eCBCT has expanded orthodontic diagnostics by providing isotropic three‑dimensional datasets suitable for craniofacial assessment and quantitative measurement [\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. However, CBCT‑derived morphometric datasets are typically high‑dimensional and inter‑correlated, which can complicate clinical interpretation when dozens of variables are considered simultaneously.\u003c/p\u003e \u003cp\u003eData‑driven craniofacial phenotyping seeks to transform multivariate morphology into clinically interpretable subgroups. In orthodontics, three‑dimensional morphometric analyses\u0026mdash;including geometric morphometrics\u0026mdash;have been promoted to capture variation in form beyond traditional two‑dimensional summaries and to support modern phenotype concepts [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Cluster‑based phenotyping is also increasingly used to reveal within‑class heterogeneity (e.g., distinct vertical and dentoalveolar compensations within sagittal class labels) that may inform treatment strategy selection [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eUnsupervised machine learning is especially appropriate when the primary goal is discovery of latent phenotypic structure, rather than prediction against predefined labels. PCA reduces dimensionality by transforming correlated variables into orthogonal components [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], while k‑means clustering partitions individuals into centroid‑defined prototypes that can be translated into diagnostic \u0026ldquo;templates\u0026rdquo; [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Compared with deep learning approaches that often require large labeled datasets and may provide limited transparency, PCA‑based clustering offers direct interpretability through loadings and cluster centroids and can serve as an explainable front‑end to broader AI pipelines [\u003cspan additionalcitationids=\"CR17\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAccordingly, this study presents an explainable PCA\u0026rarr;k‑means workflow for identifying skeletal Class II phenotypes using CBCT‑derived variables, emphasizing transparent reporting, internal validation, and clinically meaningful translation.\u003c/p\u003e"},{"header":"2. Materials and methods","content":"\u003cp\u003eStudy design and ethics: This retrospective observational study evaluated CBCT records of Yemeni adults diagnosed with skeletal Class II malocclusion. The study was conducted in accordance with the Declaration of Helsinki and was approved by the Medical Ethics Committee of Sana\u0026rsquo;a University, Yemen (Approval ID: OR 19/11/2023), which granted a waiver of informed consent because the study used retrospective, de-identified CBCT records.\u003c/p\u003e \u003cp\u003eParticipants: CBCT records meeting predefined inclusion criteria (adult dentition; diagnostic indication consistent with skeletal Class II) and excluding syndromic conditions, craniofacial trauma, and prior orthognathic surgery were screened. Records with poor image quality or artifacts affecting landmarking were excluded.\u003c/p\u003e \u003cp\u003eCBCT acquisition and measurements: Scans were acquired using a PaX‑Flex 3D P2 system with a medium field of view. Measurements were extracted in InVivo 6.0 using standardized protocols. A total of 63 variables were recorded, spanning sagittal and vertical skeletal relationships, cranial‑base dimensions, and dentoalveolar parameters including maxillary and mandibular incisor positions and inclinations [\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Measurement domains are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e[Insert Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e here]\u003c/p\u003e \u003cp\u003ePreprocessing and PCA: Variables were standardized using z‑scores. PCA was performed on the correlation matrix to reduce dimensionality and mitigate multicollinearity [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Component retention considered eigenvalues, the scree profile, cumulative variance, and anatomical interpretability of loadings. The overall AI\u0026ndash;CBCT workflow is illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e[Insert Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e here]\u003c/p\u003e \u003cp\u003ek‑means clustering and selection of k: k‑means clustering was applied to retained PCA scores with k ranging from 2 to 8 [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Multiple random initializations were used to reduce sensitivity to local minima.\u003c/p\u003e \u003cp\u003eInternal validation and sensitivity: Silhouette coefficients and Davies\u0026ndash;Bouldin indices were used to quantify within‑cluster cohesion and between‑cluster separation [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Sensitivity analyses examined the influence of the number of retained components and alternative k values, and stability was checked across multiple random starts and sex stratifications.\u003c/p\u003e"},{"header":"3. Results","content":"\u003cp\u003eDimensionality reduction: Seven principal components explained approximately 60% of the variance in the 63‑variable CBCT feature space (Fig. 2). The leading components reflected clinically recognizable axes of variation, including cranial‑base dimensions and vertical skeletal pattern (mandibular plane steepness and facial height), with additional components capturing maxillary sagittal position and dentoalveolar contributions. The cumulative and incremental variance explained by retained components is summarized in Table 2.\u003c/p\u003e\n\u003cp\u003e[Insert Table 2 here]\u003c/p\u003e\n\u003cp\u003eClustering and phenotypes: Based on internal validity indices and clinical interpretability, a five‑cluster solution was selected. The clusters delineated phenotypes consistent with deep‑bite and open‑bite tendencies, mandibular retrusion‑dominant patterns, and incisor‑protrusive compensations; these patterns were mapped to phenotype‑informed orthodontic strategies. A high-level description of the five phenotypic clusters and their potential orthodontic implications is provided in Table 3.\u003c/p\u003e\n\u003cp\u003e[Insert Table 3 here]\u003c/p\u003e\n\u003cp\u003eVisualization: PC score plots (PC1 vs PC2) demonstrated separable but partially overlapping cluster distributions (Fig. 3), consistent with expected craniofacial heterogeneity within skeletal Class II.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTables and figures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1. Overview of measurement domains used for phenotyping.\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"\"\u003e\n \u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eDomain (examples)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eRole in phenotype interpretation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eSagittal skeletal relationships\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eDifferentiate maxillary vs mandibular contributions and inform sagittal correction strategy\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eVertical skeletal pattern\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eIdentify deep‑bite vs open‑bite tendencies and vertical control needs\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eCranial‑base dimensions\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eCapture baseline craniofacial size/shape contributions affecting sagittal measures\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eDentoalveolar inclinations/positions (incisors)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eGuide torque control, anchorage planning, and camouflage decisions\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2. Cumulative and incremental variance explained by retained PCA components (from Fig. 2).\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"\"\u003e\n \u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003eComponent set\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003eCumulative variance (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003eIncremental variance (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e10.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e10.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e19.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e9.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e28.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e9.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e37.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e8.8\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e46.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e8.7\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e54.3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e8.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003ePC1\u0026ndash;PC7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e60.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 208px;\"\u003e\n \u003cp\u003e5.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3. High-level description of the five phenotypic clusters and potential orthodontic implications.\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"\"\u003e\n \u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eCluster tendency (phenotype label)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003ePhenotype‑informed orthodontic considerations (examples)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eDeep‑bite / low‑angle tendency\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eBite opening mechanics; torque control; avoid deepening; consider anterior intrusion strategies as indicated\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eOpen‑bite / high‑angle tendency\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eVertical control; minimize extrusive mechanics; evaluate habits/airway factors when relevant\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eMandibular retrusion‑dominant pattern\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eAssess growth status; camouflage vs orthognathic planning depending on severity and soft‑tissue goals\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eIncisor‑protrusive compensation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eAnchorage planning; incisor torque control; extraction decisions guided by crowding and profile\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eMixed/intermediate centroid\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 312px;\"\u003e\n \u003cp\u003eIndividualize using centroid profile as a structured checklist alongside clinical exam and imaging\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis study demonstrates an explainable unsupervised workflow that converts a high‑dimensional CBCT measurement set into clinically interpretable skeletal Class II phenotypes. PCA provides decorrelation and dimensionality reduction [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], while k‑means clustering yields centroid‑defined prototypes that can be translated into diagnostic templates [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eRelation to previous morphometric and phenotyping literature: Three‑dimensional morphometric approaches\u0026mdash;particularly geometric morphometrics\u0026mdash;have been advocated in orthodontics to quantify 3D form and to study craniofacial variation beyond conventional measures [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Similarly, earlier phenotype work in skeletal Class II highlighted substantial heterogeneity within sagittal labels, including vertical and dentoalveolar compensations that influence treatment choice [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Recent studies also demonstrate that data‑driven clustering can classify craniofacial patterns and may complement conventional diagnostics [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWhy unsupervised learning and how it complements deep learning: Deep learning has shown strong performance for supervised orthodontic imaging tasks (e.g., landmark detection) but generally requires large labeled datasets and may offer limited transparency for phenotype discovery [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Unsupervised clustering is appropriate when the objective is discovery of latent structure without assuming predefined labels, and when interpretability via component loadings and centroids is required for clinical adoption. Such interpretable clustering can also provide a structured phenotype taxonomy to support subsequent supervised modeling or outcome‑prediction studies, aligning with emerging reporting expectations for AI in dentistry [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eClinical relevance and novelty for an orthodontic audience: The clinical value is not to replace established Class II descriptors, but to provide a reproducible multivariate summary tool that supports (i) standardized case documentation, (ii) stratified treatment planning (vertical vs sagittal emphasis; torque/anchorage priorities), and (iii) hypothesis generation for phenotype‑to‑outcome research.\u003c/p\u003e \u003cp\u003eLimitations include the single‑population retrospective design and the use of engineered measurement variables rather than full-image radiomics features. Radiomics-style extensions may be feasible but require rigorous standardization and radiation justification [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Future work should validate phenotypes across centers/devices and link phenotypes to treatment outcomes and stability to demonstrate clinical utility.\u003c/p\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eA PCA\u0026rarr;k‑means clustering pipeline provides a transparent and reproducible approach to identify skeletal Class II phenotypes from CBCT‑derived variables, supporting explainable AI-assisted orthodontic diagnostics and phenotype‑informed treatment planning.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate: \u003c/strong\u003eThe Medical Ethics Committee of Sana\u0026rsquo;a University, Yemen (Approval ID: OR 19/11/2023) approved the study and waived the requirement for informed consent because the research was retrospective and used de-identified data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTrial registration: \u003c/strong\u003eNot applicable (retrospective observational study; no prospective assignment to interventions).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials:\u0026nbsp;\u003c/strong\u003eData are available from the corresponding author upon reasonable request, subject to ethical restrictions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests:\u0026nbsp;\u003c/strong\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u0026nbsp;\u003c/strong\u003eNo external funding was received.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions:\u0026nbsp;\u003c/strong\u003eConceptualization: S.M.B.H.; Methodology: S.M.B.H., R.A.I., F.L.A.; Data curation: S.M.B.H.; Formal analysis and visualization: S.M.B.H.; Writing\u0026mdash;original draft: S.M.B.H.; Writing\u0026mdash;review and editing: S.M.B.H., R.A.I., F.L.A. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u0026nbsp;\u003c/strong\u003eThe authors thank the Medical Ethics Committee of Sana\u0026rsquo;a University for oversight.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSmith LI. A tutorial on principal components analysis. 2002. Available at: http://www.cs.otago.ac.nz/cosc453/student_tutorials/principal_components.pdf. Accessed 2026-01-02.\u003c/li\u003e\n\u003cli\u003eAbdi H, Williams LJ. Principal component analysis. Wiley Interdiscip Rev Comput Stat. 2010;2(4):433\u0026ndash;459. https://doi.org/10.1002/wics.101.\u003c/li\u003e\n\u003cli\u003eHastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning. 2nd ed. Springer; 2009. https://link.springer.com/book/10.1007/978-0-387-84858-7.\u003c/li\u003e\n\u003cli\u003eRousseeuw PJ. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J Comput Appl Math. 1987;20:53\u0026ndash;65. https://doi.org/10.1016/0377-0427(87)90125-7.\u003c/li\u003e\n\u003cli\u003eDavies DL, Bouldin DW. A cluster separation measure. IEEE Trans Pattern Anal Mach Intell. 1979;1(2):224\u0026ndash;227. https://doi.org/10.1109/TPAMI.1979.4766909.\u003c/li\u003e\n\u003cli\u003eHartigan JA, Wong MA. Algorithm AS 136: a k‑means clustering algorithm. J R Stat Soc Ser C Appl Stat. 1979;28(1):100\u0026ndash;108. https://doi.org/10.2307/2346830.\u003c/li\u003e\n\u003cli\u003eScarfe WC, Farman AG. What is cone‑beam CT and how does it work? Dent Clin North Am. 2008;52(4):707\u0026ndash;730. https://doi.org/10.1016/j.cden.2008.05.005.\u003c/li\u003e\n\u003cli\u003eVenkatesh E, Elluru SV. Cone beam computed tomography: basics and applications in dentistry. J Istanb Univ Fac Dent. 2017;51(Suppl 1):S102\u0026ndash;S121. https://pmc.ncbi.nlm.nih.gov/articles/PMC5750833/.\u003c/li\u003e\n\u003cli\u003ePauwels R, Beinsberger J, Collaert B, et al. Effective dose range for dental cone beam computed tomography scanners. Eur J Radiol. 2012;81(2):267\u0026ndash;271. https://doi.org/10.1016/j.ejrad.2010.11.028.\u003c/li\u003e\n\u003cli\u003eIzenman AJ. Modern Multivariate Statistical Techniques. Springer; 2008. https://link.springer.com/book/10.1007/978-0-387-78189-1.\u003c/li\u003e\n\u003cli\u003eMacQueen J. Some methods for classification and analysis of multivariate observations. In: Proc 5th Berkeley Symp Math Stat Prob. 1967:281\u0026ndash;297. https://projecteuclid.org/journals/proceedings-of-the-fifth-berkeley-symposium-on-mathematical-statistics-and-probability/volume-1/issue-1/Some-methods-for-classification-and-analysis-of-multivariate-observations/bsmsp/1200512992.full.\u003c/li\u003e\n\u003cli\u003eLloyd SP. Least squares quantization in PCM. IEEE Trans Inf Theory. 1982;28(2):129\u0026ndash;137. https://doi.org/10.1109/TIT.1982.1056489.\u003c/li\u003e\n\u003cli\u003eDamstra J, Fourie Z, Huddleston Slater JJR, Ren Y. Accuracy of linear measurements from CBCT‑derived 3‑D surface models of different voxel sizes. Am J Orthod Dentofacial Orthop. 2010;137(1):16.e1\u0026ndash;16.e6. https://pubmed.ncbi.nlm.nih.gov/20122425/.\u003c/li\u003e\n\u003cli\u003eGribel BF, Gribel MN, Manzi FR, Brooks SL, McNamara JA Jr. Accuracy and reliability of craniometric measurements on lateral cephalograms vs CBCT. Angle Orthod. 2011;81(1):26\u0026ndash;35. https://pubmed.ncbi.nlm.nih.gov/20936948/.\u003c/li\u003e\n\u003cli\u003eScarfe WC, Aboelsaad N, Farman AG, et al. CBCT in orthodontics. Aust Dent J. 2017;62(Suppl 1):33\u0026ndash;50. https://doi.org/10.1111/adj.12479.\u003c/li\u003e\n\u003cli\u003eSchwendicke F, Samek W, Krois J. Artificial intelligence in dentistry: chances and challenges. J Dent Res. 2020;99(7):769\u0026ndash;774. https://doi.org/10.1177/0022034520915714.\u003c/li\u003e\n\u003cli\u003eArsiwala‑Scheppach LT, Chaurasia A, M\u0026uuml;ller A, Krois J, Schwendicke F. Machine learning in dentistry: a scoping review. J Clin Med. 2023;12(3):937. https://pubmed.ncbi.nlm.nih.gov/36769585/.\u003c/li\u003e\n\u003cli\u003eSchwendicke F, Chaurasia A, Arsiwala‑Scheppach L, et al. Deep learning for cephalometric landmark detection: systematic review and meta‑analysis. Clin Oral Investig. 2021;25:4299\u0026ndash;4312. https://doi.org/10.1007/s00784-021-03990-w.\u003c/li\u003e\n\u003cli\u003ePauwels R. A new formula for converting dose‑area product to effective dose in dental cone‑beam computed tomography. Phys Med. 2023;112:102639. https://doi.org/10.1016/j.ejmp.2023.00116-3.\u003c/li\u003e\n\u003cli\u003eLakhotia S, et al. Machine learning in dentistry: scoping review. PLOS Digit Health. 2025;4(9):e0000940. https://doi.org/10.1371/journal.pdig.0000940.\u003c/li\u003e\n\u003cli\u003eGhislanzoni L, et al. Measuring 3D shape in orthodontics through geometric morphometrics. Prog Orthod. 2017;18:38. https://doi.org/10.1186/s40510-017-0194-9.\u003c/li\u003e\n\u003cli\u003eMoreno Uribe LM, et al. Phenotypic diversity in skeletal Class II malocclusion. Orthod Craniofac Res. 2014;17(1):33\u0026ndash;45. https://doi.org/10.1111/ocr.12026.\u003c/li\u003e\n\u003cli\u003ede Frutos‑Valle L, et al. Craniofacial classification in skeletal Class III malocclusion using a clustering and deep learning approach. Sci Rep. 2020;10:18870. https://doi.org/10.1038/s41598-020-74488-w.\u003c/li\u003e\n\u003cli\u003eHou SY, et al. Soft tissue facial changes among adult females during orthodontic alignment: a 3D geometric morphometric study. BMC Oral Health. 2021;21:180. https://doi.org/10.1186/s12903-021-01425-2.\u003c/li\u003e\n\u003cli\u003eTarak\u0026ccedil;ı \u0026Ouml;D, et al. Radiomics-based diagnosis in dentomaxillofacial radiology: a systematic review. Dentomaxillofac Radiol. 2024;53(7):20230328. https://doi.org/10.1259/dmfr.20230328.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"discover-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Medicine](https://link.springer.com/journal/44337)","snPcode":"44337","submissionUrl":"https://submission.springernature.com/new-submission/44337/3","title":"Discover Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Cone‑beam computed tomography, three‑dimensional imaging, principal component analysis, k‑means clustering, phenotyping, orthodontic diagnosis, explainable AI","lastPublishedDoi":"10.21203/rs.3.rs-8809759/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8809759/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e: Cone‑beam computed tomography (CBCT) enables three‑dimensional assessment of craniofacial structures, yet interpretation of many inter‑correlated measurements remains challenging.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eObjective\u003c/strong\u003e: To present an explainable workflow combining principal component analysis (PCA) and k‑means clustering to identify skeletal Class II phenotypes.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e: Sixty‑three CBCT‑derived variables from 120 Yemeni adults with skeletal Class II malocclusion were standardized, reduced via PCA, and clustered (k = 2–8). Internal validity was evaluated using silhouette and Davies–Bouldin indices with multiple random initializations.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e: Seven principal components explained approximately 60% of variance; a five‑cluster solution balanced cohesion, separation, and clinical interpretability, delineating deep‑bite and open‑bite tendencies, mandibular retrusion patterns, and incisor‑protrusive phenotypes.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e: The PCA→k‑means pipeline provides a transparent, reproducible framework for artificial intelligence (AI)-assisted orthodontic diagnosis and phenotype‑informed treatment planning.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTrial registration:\u003c/strong\u003e Not applicable (retrospective observational study; no prospective assignment to interventions).\u003c/p\u003e","manuscriptTitle":"Explainable CBCT based morphometric clustering for orthodontic diagnostics using principal component analysis and k means","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-05 20:02:10","doi":"10.21203/rs.3.rs-8809759/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"184538578911391455006763741714528006169","date":"2026-04-14T01:53:19+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-16T08:30:43+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"158816006034467469542583098575921649142","date":"2026-03-16T08:12:45+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-13T07:41:14+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"288748308949824292675448060456081895267","date":"2026-03-07T17:29:18+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-05T09:23:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"231264315216996975676467151465797678968","date":"2026-03-05T06:47:29+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-02T10:24:49+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-02T08:56:52+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-19T19:34:04+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-19T18:47:32+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Medicine","date":"2026-02-19T18:43:01+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"discover-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Medicine](https://link.springer.com/journal/44337)","snPcode":"44337","submissionUrl":"https://submission.springernature.com/new-submission/44337/3","title":"Discover Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7dcea1dc-ce91-422f-beb1-15dac1fd91da","owner":[],"postedDate":"March 5th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-03-05T20:02:11+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-05 20:02:10","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8809759","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8809759","identity":"rs-8809759","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.