Explainable AI to predict a complex multifactorial outcome, childhood obesity: Application to clinical epidemiology

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study used Kolmogorov-Arnold Networks (KAN) to predict childhood obesity at age 8, outperforming traditional models and identifying key predictors like Year 5 BMI z-score, mid-arm circumference, maternal occupation, and polygenic risk scores.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-17 · read from full text

This study applied Kolmogorov-Arnold Networks (KAN) and several conventional machine-learning models to predict BMI at age 8 in participants of the Raine Study Gen2 cohort (n=2,868) using perinatal, early-life, and polygenic risk score (PGS) data collected before age 5. The authors report that KAN achieved higher predictive performance (R2=0.81) than models such as Random Forest, Gradient Boosting, Lasso, and a Multi-Layer Perceptron, identifying predictors including Year 5 BMI z-score, mid-arm circumference, mother’s occupation, and PGS. A publicly accessible online calculator was developed, and the performance was reported to remain at R2=0.81 even without using PGS. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background Childhood obesity, driven by genetic and epidemiological factors, poses significant health risks, yet traditional machine learning models lack interpretability for clinical use. Objective This study aims to apply Kolmogorov-Arnold Networks (KAN), an explainable machine learning model, to predict body mass index (BMI) at age 8 as an indicator of obesity risk and to develop a publicly accessible prediction tool. Methods We utilized the Raine Study Gen2 cohort (n=2,868) to train KAN and traditional models (such as Random Forest, Gradient Boosting, Lasso, and Multi-Layer Perceptron) using perinatal, early-life, and polygenic risk score (PGS) data collected before age 5. Feature importance was analyzed across all the models. A publicly accessible online calculator was developed for practical use. Results KAN achieved an R 2 of 0.81, outperforming traditional models. Key predictors included Year 5 BMI z-score, mid-arm circumference, occupation of mother, and PGS. The online calculator supports predictions without PGS, maintaining an R 2 of 0.81. Conclusions KAN’s transparent formulas enhance interpretability, offering a practical approach to predicting childhood obesity. The freely accessible tool enables clinicians to implement personalized prevention strategies, advancing precision medicine. Graphical abstract KAN model predicts childhood obesity (BMI at age 8), showcasing key features, top performance, and accurate formularised results with epidemiological and genetic factors. Online calculator is available at https://bmi-y8-calc.onrender.com/ .
Full text 4,430 characters · extracted from oa-doi-fallback · 5 sections · click to expand

Abstract

Background Childhood obesity, driven by genetic and epidemiological factors, poses significant health risks, yet traditional machine learning models lack interpretability for clinical use.

Objective

This study aims to apply Kolmogorov-Arnold Networks (KAN), an explainable machine learning model, to predict body mass index (BMI) at age 8 as an indicator of obesity risk and to develop a publicly accessible prediction tool.

Methods

We utilized the Raine Study Gen2 cohort (n=2,868) to train KAN and traditional models (such as Random Forest, Gradient Boosting, Lasso, and Multi-Layer Perceptron) using perinatal, early-life, and polygenic risk score (PGS) data collected before age 5. Feature importance was analyzed across all the models. A publicly accessible online calculator was developed for practical use.

Results

KAN achieved an R2 of 0.81, outperforming traditional models. Key predictors included Year 5 BMI z-score, mid-arm circumference, occupation of mother, and PGS. The online calculator supports predictions without PGS, maintaining an R2 of 0.81.

Conclusions

KAN’s transparent formulas enhance interpretability, offering a practical approach to predicting childhood obesity. The freely accessible tool enables clinicians to implement personalized prevention strategies, advancing precision medicine. Competing Interest Statement The authors have declared no competing interest. Funding Statement This study did not receive any funding. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Our Ref: 2025/ET000396 Monday, 28 April, 2025 Dr Fuling Chen The University of Western Australia Dear Dr Chen HUMAN RESEARCH ETHICS EXEMPTION FROM REVIEW. Early Life Origins of Cardiovascular Health: A Machine Learning-Driven Risk Prediction Algorithm for Pediatric Populations Based on the information you have provided to the Human Ethics office in relation to the above project, the described activity has been assessed as exempt from ethics review at the University of Western Australia. However, should there be any significant changes to the project, you must contact the HREO to determine whether your exempt status remains valid or whether you will be required to submit an application for ethics approval. If you have any queries please contact the Human Ethics office at humanethics{at}uwa.edu.au. Please ensure that you quote the file reference 2025/ET000396 and the associated project title in all future correspondence. Yours sincerely Senior Officer Human Ethics & Clinical Trials I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Footnotes Update the comparative models and results. Add the online accessible tool. Data Availability The datasets generated and analyzed during this study are not available. The Raine study is committed to a high level of confidentiality of the data in line with the informed consent provided by participants. Requests for data should be directed to the Raine Study Executive. Abbreviations - EN - Elastic Net - ERF - Extreme Random Forest - GBM - Gradient Boosting Machine - KAN - Kolmogorov Arnold Network - Lasso - Least Absolute Shrinkage and Selection Operator - MLP - Multi-Layer Perceptron - PGS - Polygenic Score - RFE - Recursive Feature Elimination - XGB - Extreme Gradient Boosting

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00