Predictors of Cardiovascular Disease Among Adults: A Multivariable Logistic Regression Analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predictors of Cardiovascular Disease Among Adults: A Multivariable Logistic Regression Analysis OlaKunle Daniel Olorunkemi This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8889390/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Cardiovascular disease remains the leading cause of morbidity and mortality worldwide. Early identification of individuals at high risk of heart disease is essential for effective prevention and clinical intervention. Routinely collected demographic and clinical variables may offer valuable predictive insight when evaluated using appropriate statistical models. This study aimed to examine whether age, sex, cholesterol level, and exercise-induced angina significantly predict the likelihood of heart disease among adults using a publicly available heart failure dataset. A secondary analysis was conducted using data from 918 adults obtained from the Kaggle Heart Failure Dataset. Heart disease status (yes/no) was modeled as a binary outcome. Multivariable logistic regression was used to assess associations between predictors and heart disease. Model assumptions were evaluated using variance inflation factors, Box–Tidwell tests for linearity of the logit, influence diagnostics, and goodness-of-fit statistics. Discriminatory performance was assessed using receiver operating characteristic analysis. The logistic regression model significantly improved prediction of heart disease compared with the intercept-only model (likelihood ratio χ²₄=370.14, P <.001). The model demonstrated strong discrimination (area under the curve = 0.846) and good calibration (Hosmer–Lemeshow P =.667). All predictors were statistically significant ( P <.001). Increasing age was associated with higher odds of heart disease (odds ratio [OR] = 1.051, 95% CI 1.032–1.070). Males had substantially higher odds than females (OR = 3.494, 95% CI 2.299–5.309). Exercise-induced angina was the strongest predictor (OR = 9.911, 95% CI 6.873–14.293). Cholesterol level showed a statistically significant but inverse association with heart disease (OR = 0.995, 95% CI 0.994–0.997). Age, sex, cholesterol level, and exercise-induced angina were significant predictors of heart disease in this dataset. Exercise-induced angina and male sex demonstrated particularly strong associations. These findings highlight the value of simple, routinely collected clinical indicators for heart disease risk stratification and support the use of logistic regression as an effective analytical approach in population-level cardiovascular research. Epidemiology INTRODUCTION Cardiovascular diseases remain the leading cause of morbidity and mortality worldwide, accounting for more than 17 million deaths annually (World Health Organization, 2023). Among these conditions, coronary heart disease is one of the most prevalent and serious forms, driven by a combination of demographic, behavioral, and clinical risk factors. Early identification of individuals at high risk is essential for prevention, targeted intervention, and reduction of long-term complications. A substantial body of epidemiologic and clinical research has established that age , sex , cholesterol levels , and exercise-induced angina are among the most prominent predictors of heart disease. Age is consistently associated with increased vascular stiffness, endothelial dysfunction, and cumulative exposure to risk factors, making it a strong non-modifiable determinant of cardiovascular risk (Benjamin et al., 2019). Sex differences have also been well documented, with men showing significantly higher heart disease prevalence and earlier onset compared to women (Virani et al., 2021). Lipid abnormalities particularly elevated serum cholesterol play a central role in atherosclerosis development. Numerous studies have demonstrated that high cholesterol levels are associated with plaque formation, impaired arterial flow, and increased likelihood of ischemic events (Grundy et al., 2019; Yusuf et al., 2020). Exercise-induced angina, reflecting myocardial oxygen supply-demand imbalance during exertion, is considered a strong clinical indicator of existing or impending coronary artery disease (Lanza & Crea, 2010). Given the public health significance of these factors, statistical modeling provides a powerful approach for identifying which predictors are most strongly associated with heart disease outcomes, especially in diverse populations. Logistic regression, in particular, is widely used to estimate the probability of disease presence based on individual-level characteristics (Harrell, 2015). The present study analyzes a heart disease dataset containing 918 adult observations , examining key demographic and clinical variables. Prior to analysis, the dataset was reviewed for completeness, variable types, outliers, and duplicate records. No missing data were present for the variables of interest, and descriptive statistics were used to evaluate distributions and identify potential anomalies. The data were subsequently prepared for regression modeling and checked for multicollinearity, linearity of the logit, and influential observations to ensure compliance with analytic assumptions. R esearch Question D o age, sex, cholesterol level, and exercise-induced angina significantly predict the likelihood of heart disease among adults in the heart failure dataset? Hypotheses A ge • H₀ : Age is not associated with the likelihood of heart disease (β₁ = 0). • H₁ : Age is associated with the likelihood of heart disease (β₁ ≠ 0). Sex • H₀ : Sex (male vs female) is not associated with the likelihood of heart disease (β₂ = 0). • H₁ : Sex is associated with the likelihood of heart disease (β₂ ≠ 0). C holesterol • H₀ : Cholesterol level is not associated with the likelihood of heart disease (β₃ = 0). • H₁ : Cholesterol level is associated with the likelihood of heart disease (β₃ ≠ 0). Exercise-Induced Angina • H₀ : Exercise-induced angina is not associated with the likelihood of heart disease (β₄ = 0). • H₁ : Exercise-induced angina is associated with the likelihood of heart disease (β₄ ≠ 0). This analysis provides an evidence-based evaluation of key clinical predictors using logistic regression, contributing to practical insights for risk assessment and early detection of heart disease. METHODOLOGY Study Design and Data Source This analysis used a secondary dataset obtained from Kaggle titled Heart Failure Dataset (N = 918). The dataset contains demographic and clinical variables commonly used to assess cardiovascular health, including age, sex, cholesterol level, and indicators of exercise-induced angina. Because the dataset includes one observation per individual and no repeated measures or longitudinal information, it functions as cross-sectional observational data , although the original study design and data c ollection procedures were not explicitly documented by the source. All statistical analyses were performed using SAS OnDemand for Academics (version 9.4). D ata Preparation and Cleaning The dataset was examined for missing values, incorrect entries, and outliers. No missing values were present for the variables used in the regression model. Continuous variables (Age, Cholesterol) were assessed for outliers using descriptive statistics and visual diagnostics. Categorical variables (Sex, ExerciseAngina, HeartDisease) were checked for coding accuracy and consistency. The dependent variable, HeartDisease , was coded as binary (0 = no heart disease, 1 = heart disease). Sex was recoded as a reference-coded categorical predictor (0 = Female [reference], 1 = Male). Exercise-induced angina was coded as a binary categorical variable (0 = No, 1 = Yes). V ariables Included in the Model • Dependent Variable: • HeartDisease (0 = No, 1 = Yes) • I ndependent Variables: • Age (continuous) • Sex (0 = Female, 1 = Male) • Cholesterol (continuous) • ExerciseAngina (0 = No angina, 1 = Yes angina) These predictors were selected based on prior literature documenting their significant clinical relevance in cardiovascular risk assessment. StatisticalApproach Descriptive statistics were computed to summarize the characteristics of the sample. Means and standard deviations were reported for continuous variables, and frequencies and percentages were reported for categorical variables. A multivariable binary logistic regression model was used to assess whether age, sex, cholesterol level, and exercise-induced angina significantly predicted the likelihood of heart disease. Logistic regression was chosen because the outcome variable was binary. Assu mption Checks Several diagnostic procedures were conducted to verify that the model met the assumptions of logistic regression: 1. Multicollinearity: Variance Inflation Factors (VIFs) were examined using PROC REG. All VIF values were < 2, indicating no multicollinearity concerns. 2. Linearity of the Logit: The Box–Tidwell test was performed by creating log-transformed interaction terms for Age and Cholesterol. Non-significant Wald tests indicated that the linearity assumption was reasonably met. 3. Influential Observations: Leverage, deviance residuals, Cook’s distance, and DFBETAs were inspected using diagnostic plots generated from PROC LOGISTIC. No extreme influential points were detected that would threaten model stability. 4. Goodness of Fit: The Hosmer–Lemeshow test was used to assess overall model fit. A non-significant p-value (p = 0.667) indicated good model fit. 5. PredictiveAccuracy: A Receiver Operating Characteristic (ROC) curve was generated to assess discriminative ability. The model achieved an Area Under the Curve (AUC) of 0.846 , reflecting strong predictive performance. Statistical Significance All tests were two-tailed with statistical significance set at α = 0.05. Odds ratios (ORs) and 95% confidence intervals (CI) were reported to quantify the magnitude and direction of associations. RESULTS D escriptive Statistics A total of 918 adults were included in the dataset. The mean age of participants was 53.51 years (SD = 9.43). Average cholesterol level was 198.80 mg/dL (SD = 109.38), with considerable variability. Most participants were male (78.98%), 40.41% reported exercise-induced angina, and 55.34% were classified as having heart disease. M odel Fit and Diagnostics The logistic regression model significantly predicted the likelihood of heart disease when compared to the intercept-only model, Likelihood Ratio χ²(4) = 370.14, p < .0001 , indicating the overall model was statistically significant. Model fit statistics suggested strong performance: • A I C decreased from 1264.14 (null model) to 901.99 (final model), • −2 Log Likelihood also improved markedly (891.99), • Max-rescaled R² = 0.4441 , indicating moderate explanatory power. The Hosmer–Lemeshow test was non-significant, χ²(8) = 5.82, p = .667 , suggesting the model fits the data well. The ROC curve yielded an AUC of 0.846 , demonstrating excellent discrimination between individuals with and without heart disease. I ndividual Predictor Effects All predictors were statistically significant at α = .05. A ge • β = 0.0498, Wald χ²(1) = 29.47, p < .0001 • OR = 1.051 (95% CI: 1.032–1.070) I nterpretation: Each additional year of age increases the odds of heart disease by 5.1% , holding other variables constant. Sex (Male vs Female) • β = 1.2509, Wald χ²(1) = 34.31, p < .0001 • OR = 3.494 (95% CI: 2.299–5.309) I nterpretation: Males have 3.5 times higher odds of heart disease compared to females. C holesterol • β = –0.00468, Wald χ²(1) = 31.47, p < .0001 • OR = 0.995 (95% CI: 0.994–0.997) I nterpretation: Higher cholesterol values were associated with slightly reduced odds of heart disease. Although counterintuitive, this likely reflects characteristics of this specific dataset. Exercise-Induced Angina • β = 2.2937, Wald χ² = 150.80, p < .0001 • OR = 9.911 (95% CI: 6.873–14.293) I nterpretation: Individuals with exercise-induced angina have nearly 10 times higher odds of heart disease than those without angina, making this the strongest predictor. I nfluential Observations Examination of leverage values, deviance residuals, and Cook’s distance showed no extreme influential points. Thus, all observations were retained , and the final model was considered robust. DISCUSSION The purpose of this study was to examine whether age, sex, cholesterol level, and exercise-induced angina predict the likelihood of heart disease among adults in the Heart Failure Dataset. Consistent with prior epidemiological research, all four predictors were statistically significant, and the final logistic regression model demonstrated strong discriminatory ability (AUC = 0.846) and good overall fit (Hosmer–Lemeshow p = .667). Age emerged as a significant predictor, with older adults showing greater odds of heart disease. This finding aligns with well-established evidence that cardiovascular risk increases with vascular aging, cumulative exposure to risk factors, and declining endothelial function (Benjamin et al., 2019 ). Theassociation between sex and heart disease was also robust: males had 3.5 times higher odds of heart disease compared to females. This is consistent with prior studies showing earlier onset and higher prevalence of coronary artery disease in men, likely due to biological, hormonal, and behavioral contributors (Virani et al., 2021 ). Exercise-induced angina was the strongest predictor in the model, associated with nearly a tenfold increase in the odds of heart disease. This aligns with clinical literature identifying exertional chest pain as a key marker of myocardial oxygen imbalance and coronary ischemia (Lanza & Crea, 2010 ). This variable’s large odds ratio underscores its diagnostic value and the importance of screening for exercise-induced symptoms in clinical settings. Interestingly, higher cholesterol levels were associated with slightly lower odds of heart disease in this dataset. Although this finding contrasts with extensive literature linking hyperlipidemia to atherosclerosis (Grundy et al., 2019 ; Yusuf et al., 2020 ), several explanations are plausible. First, cholesterol measurement in this dataset may reflect non-fasting values, treatment effects (e.g., statin therapy), or reverse causation where diagnosed heart disease patients may be more aggressively managed and therefore exhibit lower cholesterol readings. This inverse relationship should not be interpreted as protective, but rather as dataset-specific bias. The model as a whole performed well, with moderate explanatory power (Max-rescaled R² = 0.4441), excellent discrimination (AUC = 0.846), and no evidence of poor fit. Diagnostic checks confirmed adherence to logistic regression assumptions, and no influential observations were identified. These findings support the validity of the analytic approach and confirm that the chosen predictors meaningfully contribute to heart disease risk classification in this sample. STRENGTHSAND LIMITATIONS Strengths 1. Use of a Multivariable Logistic Regression Model The analysis employed an appropriate statistical method for a binary outcome, allowing simultaneous evaluation of multiple clinically meaningful predictors (Age, Sex, Cholesterol, and Exercise-InducedAngina). Odds ratios provided clear interpretation of each predictor’s effect. 2. Reproducible and Systematic Workflow The project followed a transparent sequence of steps: importing the data, conducting descriptive statistics, performing assumption checks, running the logistic model, and interpreting results. All analysis was conducted using SAS, and the complete syntax can be replicated, strengthening reproducibility. 3. ComprehensiveAssumption Diagnostics Several key logistic regression assumptions were checked, including: • Multicollinearity using VIF • Linearity of the logit using Box–Tidwell • Influence diagnostics using Cook’s distance, deviance residuals, and leverage • Model fit using the Hosmer–Lemeshow test • Predictive accuracy using ROC/AUC These steps increase the credibility and robustness of the final model. 4. Strong Predictive Performance The model achieved an AUC of 0.846 , indicating excellent discriminatory ability. The selected predictors meaningfully contributed to identifying individuals at higher risk of heart disease. 5. Dataset Consistency for Key Variables The dataset contained no missing values for the variables used in the regression model, which allowed the full sample to be used without imputation in the final analysis. Limitations 1. Limited Set of Predictors Included in the Final Model Although the dataset contained additional clinically relevant variables (such as MaxHR, Oldpeak, ST_Slope, RestingBP, and ChestPainType), only four predictors were included to maintain simplicity. This may restrict the model’s explanatory power.Additionally, widely recognized cardiovascular risk factors such as smoking status, medication use, diabetes history, and hypertension were not available at all. 2. Potential Issues With Cholesterol Measurements The dataset includes zero cholesterol values, which are biologically unlikely and may represent measurement or data entry issues. These zero values caused SAS to exclude 172 observations during the Box–Tidwell assumption test due to undefined log calculations.Although this did not affect the final logistic model, it may have influenced the linearity assessment. 3. Uncertain Study Design and Temporal Limitations The dataset provides a single observation per individual and includes no time-based variables. Because the original data collection method is unknown, temporal relationships cannot be determined. Therefore, associations identified in this model should not be interpreted causally. The dataset appears cross-sectional, but the true study design cannot be confirmed. 4. Gender Imbalance in the Sample The dataset is heavily male-dominant (approximately 79% male), which may limit generalizability and may inflate the apparent effect of Sex in the model. 5. Secondary Dataset With Unknown Sampling Procedures Because the dataset was obtained from Kaggle and no metadata describing recruitment, measurement protocols, or population characteristics was provided, external validity is limited. The findings should be interpreted as exploratory rather than representative of a broader population. 6. Unexpected Direction of Cholesterol–DiseaseAssociation Cholesterol showed a statistically significant but inverse association with heart disease in this dataset. This contrasts with established clinical evidence and may reflect treatment effects (e.g., statin therapy), misclassification, or other unmeasured confounders. Therefore, this result should be interpreted cautiously. CONCLUSION This study examined whether age, sex, cholesterol level, and exercise-induced angina predict the likelihood of heart disease among adults using a publicly available Heart Failure Dataset. Logistic regression analysis demonstrated that the overall model significantly predicted heart disease, with strong model fit indicators, including anAUC of 0.846, high concordance (84.6%), and a non-significant Hosmer–Lemeshow test, suggesting good calibration. The findings revealed that sex and exercise-induced angina were the strongest predictors of heart disease. Males had significantly higher odds of heart disease compared to females, and individuals who experienced exercise-induced angina had almost ten times higher odds of heart disease relative to those without angina. Age was also positively associated with heart disease risk, although its effect was smaller in magnitude. Cholesterol level showed a statistically significant but modest negative association with heart disease in the final model. These results align with existing clinical literature showing that age, male sex, and symptoms of cardiac stress (such as exercise-induced angina) are well-established predictors of cardiovascular disease. The strong predictive accuracy of the model highlights the value of simple, routinely collected clinical variables in identifying individuals at elevated risk. Overall, this analysis provides evidence supporting the importance of demographic and clinical markers in heart disease prediction and emphasizes the usefulness of logistic regression as a practical tool for risk stratification. Future research using longitudinal or more detailed clinical datasets may further clarify causal pathways and enhance predictive performance. RECOMMENDATIONS Based on the study findings and in line with JMIR’s emphasis on digital health relevance and clinical applicability, the following recommendations are proposed: Clinical Screening and Risk Stratification Exercise-induced angina should be prioritized as a key screening variable in both clinical and digital health risk-assessment tools, given its strong association with heart disease. Incorporating symptom-based indicators alongside demographic factors may improve early detection. Integration Into Digital Health Tools The predictors examined in this study are routinely collected and well-suited for integration into electronic health records, clinical decision support systems, and mobile health applications aimed at cardiovascular risk assessment. Interpretation of Cholesterol Findings The inverse association between cholesterol level and heart disease observed in this dataset should be interpreted cautiously. Future studies should account for medication use (eg, statins), fasting status, and disease management history to clarify this relationship. Expansion of Predictive Models Future research should incorporate additional clinical and behavioral risk factors such as smoking status, diabetes, hypertension, physical activity, and medication use to improve explanatory power and predictive accuracy. Longitudinal and Diverse Populations Longitudinal datasets are needed to assess causal relationships and temporal risk trajectories. Additionally, more balanced samples with respect to sex and broader population representation would enhance generalizability. Validation and External Testing External validation using independent datasets is recommended before clinical or digital implementation to ensure robustness and transportability of the predictive model. References Benjamin, E. J., Muntner, P., Alonso, A., Bittencourt, M. S., Callaway, C. W., Carson, A. P., … & Virani, S. S. (2019). Heart disease and stroke statistics—2019 update:A report from theAmerican HeartAssociation. Circulation, 139 (10), e56–e528. https://doi.org/10.1161/CIR.0000000000000659 Grundy, S. M., Stone, N. J., Bailey, A. L., Beam, C., Birtcher, K. K., Blumenthal, R. S., … & Yeboah, J. (2019). 2018 AHA/ACC guideline on the management of blood cholesterol. Journal of the American College of Cardiology, 73 (24), e285–e350. https://doi.org/10.1016/j.jacc.2018.11.003 Harrell, F. E. (2015). Regression modeling strategies: With applications to linear models, logistic regression, and survival analysis (2nd ed.). Springer. https://doi.org/10.1007/978-3-319-19425-7 Lanza, G. A., & Crea, F. (2010). Primary coronary microvascular dysfunction: Clinical presentation, pathophysiology, and management. Circulation, 121 (21), 2317–2325. https://doi.org/10.1161/CIRCULATIONAHA.109.900191 Virani, S. S., Alonso, A., Aparicio, H. J., Benjamin, E. J., Bittencourt, M. S., Callaway, C. W., … & Tsao, C. W. (2021). Heart disease and stroke statistics—2021 update:A report from theAmerican HeartAssociation. Circulation, 143 (8), e254–e743. https://doi.org/10.1161/CIR.0000000000000950 World Health Organization. (2023). Cardiovascular diseases (CVDs). https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds) Yusuf, S., Joseph, P., Rangarajan, S., Islam, S., Mente, A., Ramesh, D., … & Teo, K. K. (2020). Modifiable risk factors, cardiovascular disease, and mortality in 155,722 individuals from 21 high-, middle-, and low-income countries (PURE):A prospective cohort study. The Lancet, 395 (10226), 795– 808. https://doi.org/10.1016/S0140-6736(19)32008-2 Tables Table 1 and 2 are available in the Supplementary Files section. Additional Declarations The authors declare no competing interests. Supplementary Files Tablea.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8889390","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":591832422,"identity":"dd923105-4c5a-4d77-bbae-67e0df1a388f","order_by":0,"name":"OlaKunle Daniel Olorunkemi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABK0lEQVRIie2PMUsDMRTH3xG4LHfeJgkF/Qo5AkXxwA/icuHALu0kSIWiOQLpIrgKxU/h4mgJdDr9AucihW6FgktHc6Uc2KsHbg75wQsv/7wf5AE4HP8RDMieJ7aQrO4MwBuvV7Y7gM1TE7TJiS2vViR9tJH/BwWgE7QoEUJzHr6QI8B5Pr/RHzw6NJInQ3PrY2UYjJKLHYUqn2dhQTgEUxW/6UWXToTM+oUhfjDLUphdDuRPhRngJtRESCI0zbVJWCmkGWirkD5/9WzfUPBXpdzVynkpcnXaqgT2Y5qksFW6rCMU8rZKukehKriKnzSJdbWLfDeclEJ790WParsLS5u7RHj8TJY6OY6wmn7KaxM/THoLWA/PIpsYsholu0qNvz9Ofxl3OBwORyvfQSVprTry8/cAAAAASUVORK5CYII=","orcid":"","institution":"Georgia Southern University","correspondingAuthor":true,"prefix":"","firstName":"OlaKunle","middleName":"Daniel","lastName":"Olorunkemi","suffix":""}],"badges":[],"createdAt":"2026-02-16 03:47:31","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-8889390/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8889390/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102897382,"identity":"1d65796c-23a8-4d06-bca5-3212bdf2be41","added_by":"auto","created_at":"2026-02-18 06:56:43","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":17787,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cbr\u003e\u003c/p\u003e","description":"","filename":"Tablea.docx","url":"https://assets-eu.researchsquare.com/files/rs-8889390/v1/0189eadad6f64d1385094797.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003ePredictors of Cardiovascular Disease Among Adults: A Multivariable Logistic Regression Analysis\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eCardiovascular diseases remain the leading cause of morbidity and mortality worldwide, accounting for more than 17 million deaths annually (World Health Organization, 2023). Among these conditions, \u003cstrong\u003ecoronary heart disease\u0026nbsp;\u003c/strong\u003eis one of the most prevalent and serious forms, driven by a combination of demographic, behavioral, and clinical risk factors. Early identification of individuals at high risk is essential for prevention, targeted intervention, and reduction of long-term complications.\u003c/p\u003e\n\u003cp\u003eA substantial body of epidemiologic and clinical research has established that \u003cstrong\u003eage\u003c/strong\u003e, \u003cstrong\u003esex\u003c/strong\u003e, \u003cstrong\u003echolesterol levels\u003c/strong\u003e, and \u003cstrong\u003eexercise-induced angina\u0026nbsp;\u003c/strong\u003eare among the most prominent predictors of heart disease. Age is consistently associated with increased vascular stiffness, endothelial dysfunction, and cumulative exposure to risk factors, making it a strong non-modifiable determinant of cardiovascular risk (Benjamin et al., 2019). Sex differences have also been well documented, with \u003cstrong\u003emen showing significantly higher heart disease prevalence and earlier onset\u0026nbsp;\u003c/strong\u003ecompared to women (Virani et al., 2021).\u003c/p\u003e\n\u003cp\u003eLipid abnormalities particularly elevated serum cholesterol play a central role in atherosclerosis development. Numerous studies have demonstrated that high cholesterol levels are associated with plaque formation, impaired arterial flow, and increased likelihood of ischemic events (Grundy et al., 2019; Yusuf et al., 2020). Exercise-induced angina, reflecting myocardial oxygen supply-demand imbalance during exertion, is considered a strong clinical indicator of existing or impending coronary artery disease (Lanza \u0026amp; Crea, 2010).\u003c/p\u003e\n\u003cp\u003eGiven the public health significance of these factors, statistical modeling provides a powerful approach for identifying which predictors are most strongly associated with heart disease outcomes, especially in diverse populations. Logistic regression, in particular, is widely used to estimate the probability of disease presence based on individual-level characteristics (Harrell, 2015).\u003c/p\u003e\n\u003cp\u003eThe present study analyzes a \u003cstrong\u003eheart disease dataset containing 918 adult observations\u003c/strong\u003e, examining key\u0026nbsp;demographic and clinical variables. Prior to\u0026nbsp;analysis, the dataset\u0026nbsp;was reviewed for completeness, variable types, outliers, and duplicate records. No missing data were present for the variables of interest, and\u0026nbsp;descriptive statistics were used to\u0026nbsp;evaluate distributions and\u0026nbsp;identify\u0026nbsp;potential\u0026nbsp;anomalies. The data\u0026nbsp;were subsequently prepared for regression modeling and checked\u0026nbsp;for multicollinearity, linearity of the logit, and influential observations to ensure compliance with analytic assumptions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eR\u003c/strong\u003e\u003cstrong\u003eesearch Question\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD\u003c/strong\u003e\u003cstrong\u003eo age, sex, cholesterol level, and exercise-induced angina significantly predict the likelihood of heart disease among adults in the heart failure dataset?\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHypotheses\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA\u003c/strong\u003e\u003cstrong\u003ege\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cstrong\u003eH₀\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eAge is \u003cstrong\u003enot\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₁ = 0). \u0026bull; \u003cstrong\u003eH₁\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eAge \u003cstrong\u003eis\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₁ \u0026ne; 0).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSex\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cstrong\u003eH₀\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eSex (male vs female) is \u003cstrong\u003enot\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₂ = 0). \u0026bull; \u003cstrong\u003eH₁\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eSex \u003cstrong\u003eis\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₂ \u0026ne; 0).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC\u003c/strong\u003e\u003cstrong\u003eholesterol\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cstrong\u003eH₀\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eCholesterol level is \u003cstrong\u003enot\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₃ = 0). \u0026bull; \u003cstrong\u003eH₁\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eCholesterol level \u003cstrong\u003eis\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₃ \u0026ne; 0).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eExercise-Induced Angina\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cstrong\u003eH₀\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eExercise-induced angina is \u003cstrong\u003enot\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₄ = 0). \u0026bull; \u003cstrong\u003eH₁\u003c/strong\u003e\u003cstrong\u003e:\u0026nbsp;\u003c/strong\u003eExercise-induced angina \u003cstrong\u003eis\u0026nbsp;\u003c/strong\u003eassociated with the likelihood of heart disease (\u0026beta;₄ \u0026ne; 0).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThis analysis provides an evidence-based evaluation of key clinical predictors using logistic regression, contributing to practical insights for risk assessment and early detection of heart disease.\u003c/p\u003e"},{"header":"METHODOLOGY","content":"\u003cp\u003e\u003cstrong\u003eStudy Design and Data Source\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThis analysis used a secondary dataset obtained from Kaggle titled \u003cem\u003eHeart Failure Dataset\u0026nbsp;\u003c/em\u003e(N = 918). The dataset contains demographic and clinical variables commonly used to assess cardiovascular health, including age, sex, cholesterol level, and indicators of exercise-induced angina. Because the dataset includes one observation per individual and no repeated measures or longitudinal information, it functions as \u003cstrong\u003ecross-sectional observational data\u003c/strong\u003e, although the \u003cstrong\u003eoriginal study design and data\u003c/strong\u003e\u003cstrong\u003e\u003cbr\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ec\u003c/strong\u003e\u003cstrong\u003eollection procedures were not explicitly documented\u0026nbsp;\u003c/strong\u003eby the source.\u003c/p\u003e\n\u003cp\u003eAll statistical analyses were performed using SAS OnDemand for Academics (version 9.4).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD\u003c/strong\u003e\u003cstrong\u003eata Preparation and Cleaning\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe dataset was examined for missing values, incorrect entries, and outliers. No missing values were present for the variables used in the regression model. Continuous variables (Age, Cholesterol) were assessed for outliers using descriptive statistics and visual diagnostics. Categorical variables (Sex, ExerciseAngina, HeartDisease) were checked for coding accuracy and consistency.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe dependent variable, \u003cstrong\u003eHeartDisease\u003c/strong\u003e, was coded as binary (0 = no heart disease, 1 = heart disease). Sex was recoded as a reference-coded categorical predictor (0 = Female [reference], 1 = Male). Exercise-induced angina was coded as a binary categorical variable (0 = No, 1 = Yes).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eV\u003c/strong\u003e\u003cstrong\u003eariables Included in the Model\u0026nbsp;\u003c/strong\u003e\u0026bull; \u003cstrong\u003eDependent Variable:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cem\u003eHeartDisease\u0026nbsp;\u003c/em\u003e(0 = No, 1 = Yes) \u0026bull; \u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003endependent Variables:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cem\u003eAge\u0026nbsp;\u003c/em\u003e(continuous)\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cem\u003eSex\u0026nbsp;\u003c/em\u003e(0 = Female, 1 = Male) \u0026bull; \u003cem\u003eCholesterol\u0026nbsp;\u003c/em\u003e(continuous)\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cem\u003eExerciseAngina\u0026nbsp;\u003c/em\u003e(0 = No angina, 1 = Yes angina)\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThese predictors were selected based on prior literature documenting their significant clinical relevance in cardiovascular risk assessment.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatisticalApproach\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDescriptive statistics were computed to summarize the characteristics of the sample. Means and standard deviations were reported for continuous variables, and frequencies and percentages were reported for categorical variables.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eA multivariable \u003cstrong\u003ebinary logistic regression\u0026nbsp;\u003c/strong\u003emodel was used to assess whether age, sex, cholesterol level, and exercise-induced angina significantly predicted the likelihood of heart disease. Logistic regression was chosen because the outcome variable was binary.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAssu\u003c/strong\u003e\u003cstrong\u003emption Checks\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSeveral diagnostic procedures were conducted to verify that the model met the assumptions of logistic regression:\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e1. \u003cstrong\u003eMulticollinearity:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eVariance Inflation Factors (VIFs) were examined using PROC REG. All VIF values were \u0026lt; 2, indicating no multicollinearity concerns.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e2. \u003cstrong\u003eLinearity of the Logit:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Box\u0026ndash;Tidwell test was performed by creating log-transformed interaction terms for Age and Cholesterol. Non-significant Wald tests indicated that the linearity assumption was reasonably met.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e3. \u003cstrong\u003eInfluential Observations:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLeverage, deviance residuals, Cook\u0026rsquo;s distance, and DFBETAs were inspected using diagnostic plots generated from PROC LOGISTIC. No extreme influential points were detected that would threaten model stability.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e4. \u003cstrong\u003eGoodness of Fit:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe \u003cstrong\u003eHosmer\u0026ndash;Lemeshow test\u0026nbsp;\u003c/strong\u003ewas used to assess overall model fit. A non-significant p-value (p = 0.667) indicated good model fit.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e5. \u003cstrong\u003ePredictiveAccuracy:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA Receiver Operating Characteristic (ROC) curve was generated to assess discriminative ability. The model achieved an Area Under the Curve (AUC) of \u003cstrong\u003e0.846\u003c/strong\u003e, reflecting strong predictive performance.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Significance\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAll tests were two-tailed with statistical significance set at \u0026alpha; = 0.05. Odds ratios (ORs) and 95% confidence intervals (CI) were reported to quantify the magnitude and direction of associations.\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cp\u003e\u003cstrong\u003eD\u003c/strong\u003e\u003cstrong\u003eescriptive Statistics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of 918 adults were included in the dataset. The mean age of participants was \u003cstrong\u003e53.51 years\u0026nbsp;\u003c/strong\u003e(SD = 9.43). Average cholesterol level was \u003cstrong\u003e198.80 mg/dL\u0026nbsp;\u003c/strong\u003e(SD = 109.38), with considerable variability. Most participants were male (78.98%), 40.41% reported exercise-induced angina, and 55.34% were classified as having heart disease.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eM\u003c/strong\u003e\u003cstrong\u003eodel Fit and Diagnostics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe logistic regression model significantly predicted the likelihood\u0026nbsp;of heart disease\u0026nbsp;when compared to the intercept-only model,\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLikelihood Ratio \u0026chi;\u0026sup2;(4) =\u0026nbsp;370.14, p \u0026lt; .0001\u003c/strong\u003e, indicating the overall model was statistically significant.\u003c/p\u003e\n\u003cp\u003eModel fit statistics suggested strong performance:\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cstrong\u003eA\u003c/strong\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003eC decreased\u0026nbsp;\u003c/strong\u003efrom 1264.14 (null model) to \u003cstrong\u003e901.99\u0026nbsp;\u003c/strong\u003e(final model), \u0026bull; \u003cstrong\u003e\u0026minus;2 Log Likelihood\u0026nbsp;\u003c/strong\u003ealso improved markedly (891.99),\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u003cstrong\u003eMax-rescaled R\u0026sup2; =\u0026nbsp;0.4441\u003c/strong\u003e, indicating moderate explanatory power.\u003c/p\u003e\n\u003cp\u003eThe \u003cstrong\u003eHosmer\u0026ndash;Lemeshow test\u0026nbsp;\u003c/strong\u003ewas non-significant,\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026chi;\u0026sup2;(8)\u0026nbsp;= 5.82, p = .667\u003c/strong\u003e, suggesting the model fits the data well.\u003c/p\u003e\n\u003cp\u003eThe \u003cstrong\u003eROC curve yielded an AUC of 0.846\u003c/strong\u003e, demonstrating excellent discrimination between individuals with and without heart disease.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003endividual Predictor Effects\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll predictors were statistically significant at \u0026alpha; = .05.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA\u003c/strong\u003e\u003cstrong\u003ege\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u0026beta; = 0.0498, Wald \u0026chi;\u0026sup2;(1) = 29.47, \u003cstrong\u003ep \u0026lt; .0001\u0026nbsp;\u003c/strong\u003e\u0026bull; OR = \u003cstrong\u003e1.051\u0026nbsp;\u003c/strong\u003e(95% CI: 1.032\u0026ndash;1.070)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003enterpretation:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEach additional year of age increases the odds of heart disease by \u003cstrong\u003e5.1%\u003c/strong\u003e, holding other variables constant.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSex (Male vs Female)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u0026beta; = 1.2509, Wald \u0026chi;\u0026sup2;(1) = 34.31, \u003cstrong\u003ep \u0026lt; .0001\u0026nbsp;\u003c/strong\u003e\u0026bull; OR = \u003cstrong\u003e3.494\u0026nbsp;\u003c/strong\u003e(95% CI: 2.299\u0026ndash;5.309)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003enterpretation:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMales have \u003cstrong\u003e3.5 times higher odds\u0026nbsp;\u003c/strong\u003eof heart disease compared to females.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC\u003c/strong\u003e\u003cstrong\u003eholesterol\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u0026beta; = \u0026ndash;0.00468, Wald \u0026chi;\u0026sup2;(1) = 31.47, \u003cstrong\u003ep \u0026lt; .0001\u0026nbsp;\u003c/strong\u003e\u0026bull; OR = \u003cstrong\u003e0.995\u0026nbsp;\u003c/strong\u003e(95% CI: 0.994\u0026ndash;0.997)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003enterpretation:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHigher cholesterol values were associated with slightly reduced odds of heart disease. Although counterintuitive, this likely reflects characteristics of this specific dataset.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eExercise-Induced Angina\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026bull; \u0026beta; = 2.2937, Wald \u0026chi;\u0026sup2; = 150.80, \u003cstrong\u003ep \u0026lt; .0001\u0026nbsp;\u003c/strong\u003e\u0026bull; OR = \u003cstrong\u003e9.911\u0026nbsp;\u003c/strong\u003e(95% CI: 6.873\u0026ndash;14.293)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003enterpretation:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIndividuals with exercise-induced angina have \u003cstrong\u003enearly 10 times higher odds\u0026nbsp;\u003c/strong\u003eof heart disease than those without angina, making this the strongest predictor.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003cstrong\u003enfluential Observations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eExamination of leverage values, deviance residuals, and Cook\u0026rsquo;s distance showed no extreme influential points.\u003c/p\u003e\n\u003cp\u003eThus, \u003cstrong\u003eall observations were retained\u003c/strong\u003e, and the final model was considered robust.\u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThe purpose of this study was to examine whether age, sex, cholesterol level, and exercise-induced angina predict the likelihood of heart disease among adults in the Heart Failure Dataset. Consistent with prior epidemiological research, all four predictors were statistically significant, and the final logistic regression model demonstrated strong discriminatory ability (AUC\u0026thinsp;=\u0026thinsp;0.846) and good overall fit (Hosmer\u0026ndash;Lemeshow p = .667).\u003c/p\u003e \u003cp\u003eAge emerged as a significant predictor, with older adults showing greater odds of heart disease. This finding aligns with well-established evidence that cardiovascular risk increases with vascular aging, cumulative exposure to risk factors, and declining endothelial function (Benjamin et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Theassociation between sex and heart disease was also robust: males had 3.5 times higher odds of heart disease compared to females. This is consistent with prior studies showing earlier onset and higher prevalence of coronary artery disease in men, likely due to biological, hormonal, and behavioral contributors (Virani et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eExercise-induced angina was the strongest predictor in the model, associated with nearly a tenfold increase in the odds of heart disease. This aligns with clinical literature identifying exertional chest pain as a key marker of myocardial oxygen imbalance and coronary ischemia (Lanza \u0026amp; Crea, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2010\u003c/span\u003e). This variable\u0026rsquo;s large odds ratio underscores its diagnostic value and the importance of screening for\u003c/p\u003e \u003cp\u003eexercise-induced symptoms in clinical settings.\u003c/p\u003e \u003cp\u003eInterestingly, higher cholesterol levels were associated with \u003cem\u003eslightly\u003c/em\u003e lower odds of heart disease in this dataset. Although this finding contrasts with extensive literature linking hyperlipidemia to atherosclerosis (Grundy et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Yusuf et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), several explanations are plausible. First, cholesterol measurement in this dataset may reflect non-fasting values, treatment effects (e.g., statin therapy), or reverse causation where diagnosed heart disease patients may be more aggressively managed and therefore exhibit lower cholesterol readings. This inverse relationship should not be interpreted as protective, but rather as dataset-specific bias.\u003c/p\u003e \u003cp\u003eThe model as a whole performed well, with moderate explanatory power (Max-rescaled R\u0026sup2; = 0.4441), excellent discrimination (AUC\u0026thinsp;=\u0026thinsp;0.846), and no evidence of poor fit. Diagnostic checks confirmed adherence to logistic regression assumptions, and no influential observations were identified. These findings support the validity of the analytic approach and confirm that the chosen predictors meaningfully contribute to heart disease risk classification in this sample.\u003c/p\u003e"},{"header":"STRENGTHSAND LIMITATIONS","content":"\u003cp\u003e\u003cstrong\u003eStrengths\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e1. \u003cstrong\u003eUse of a Multivariable Logistic Regression Model\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe analysis employed an appropriate statistical method for a binary outcome, allowing simultaneous evaluation of multiple clinically meaningful predictors (Age, Sex, Cholesterol, and Exercise-InducedAngina). Odds ratios provided clear interpretation of each predictor\u0026rsquo;s effect.\u003c/p\u003e\n\u003cp\u003e2. \u003cstrong\u003eReproducible and Systematic Workflow\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe project followed a transparent sequence of steps: importing the data, conducting descriptive statistics, performing assumption checks, running the logistic model, and interpreting results. All analysis was conducted using SAS, and the complete syntax can be replicated, strengthening reproducibility.\u003c/p\u003e\n\u003cp\u003e3. \u003cstrong\u003eComprehensiveAssumption Diagnostics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSeveral key logistic regression assumptions were checked, including:\u003c/p\u003e\n\u003cp\u003e\u0026bull; Multicollinearity using VIF\u003c/p\u003e\n\u003cp\u003e\u0026bull; Linearity of the logit using Box\u0026ndash;Tidwell\u003c/p\u003e\n\u003cp\u003e\u0026bull; Influence diagnostics using Cook\u0026rsquo;s distance, deviance residuals, and leverage\u0026nbsp;\u0026bull; Model fit using the Hosmer\u0026ndash;Lemeshow test\u003c/p\u003e\n\u003cp\u003e\u0026bull; Predictive accuracy using ROC/AUC\u003c/p\u003e\n\u003cp\u003eThese steps increase the credibility and robustness of the final model.\u003c/p\u003e\n\u003cp\u003e4. \u003cstrong\u003eStrong Predictive Performance\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe model achieved an AUC of \u003cstrong\u003e0.846\u003c/strong\u003e, indicating excellent discriminatory ability. The selected predictors meaningfully contributed to identifying individuals at higher risk of heart disease.\u003c/p\u003e\n\u003cp\u003e5. \u003cstrong\u003eDataset Consistency for Key Variables\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset contained \u003cstrong\u003eno missing values\u0026nbsp;\u003c/strong\u003efor the variables used in the regression model, which allowed the full sample to be used without imputation in the final analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLimitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e1. \u003cstrong\u003eLimited Set of Predictors Included in the Final Model\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAlthough the dataset contained additional clinically relevant variables (such as MaxHR, Oldpeak, ST_Slope, RestingBP, and ChestPainType), only four predictors were included to maintain simplicity. This may restrict the model\u0026rsquo;s explanatory power.Additionally, widely recognized cardiovascular risk factors such as smoking status, medication use, diabetes history, and hypertension were not available at all.\u003c/p\u003e\n\u003cp\u003e2. \u003cstrong\u003ePotential Issues With Cholesterol Measurements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset includes zero cholesterol values, which are biologically unlikely and may represent measurement or data entry issues. These zero values caused SAS to exclude 172 observations during the Box\u0026ndash;Tidwell assumption test due to undefined log calculations.Although this did not affect the final logistic model, it may have influenced the linearity assessment.\u003c/p\u003e\n\u003cp\u003e3. \u003cstrong\u003eUncertain Study Design and Temporal Limitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset provides a single observation per individual and includes no time-based variables. Because the original data collection method is unknown, temporal relationships cannot be determined. Therefore, associations identified in this model should not be interpreted causally. The dataset \u003cem\u003eappears\u0026nbsp;\u003c/em\u003ecross-sectional, but the true study design cannot be confirmed.\u003c/p\u003e\n\u003cp\u003e4. \u003cstrong\u003eGender Imbalance in the Sample\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset is heavily male-dominant (approximately 79% male), which may limit generalizability and may inflate the apparent effect of Sex in the model.\u003c/p\u003e\n\u003cp\u003e5. \u003cstrong\u003eSecondary Dataset With Unknown Sampling Procedures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBecause the dataset was obtained from Kaggle and no metadata describing recruitment, measurement protocols, or population characteristics was provided, external validity is limited. The findings should be interpreted as exploratory rather than representative of a broader population.\u003c/p\u003e\n\u003cp\u003e6. \u003cstrong\u003eUnexpected Direction of Cholesterol\u0026ndash;DiseaseAssociation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCholesterol showed a statistically significant but inverse association with heart disease in this dataset. This contrasts with established clinical evidence and may reflect treatment effects (e.g., statin therapy), misclassification, or other unmeasured confounders. Therefore, this result should be interpreted cautiously.\u003c/p\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eThis study examined whether age, sex, cholesterol level, and exercise-induced angina predict the likelihood of heart disease among adults using a publicly available Heart Failure Dataset. Logistic regression analysis demonstrated that the overall model significantly predicted heart disease, with strong model fit indicators, including anAUC of 0.846, high concordance (84.6%), and a non-significant Hosmer\u0026ndash;Lemeshow test, suggesting good calibration.\u003c/p\u003e \u003cp\u003eThe findings revealed that \u003cb\u003esex\u003c/b\u003e and \u003cb\u003eexercise-induced angina\u003c/b\u003e were the strongest predictors of heart disease. Males had significantly higher odds of heart disease compared to females, and individuals who experienced exercise-induced angina had almost \u003cb\u003eten times higher odds\u003c/b\u003e of heart disease relative to those without angina. \u003cb\u003eAge\u003c/b\u003e was also positively associated with heart disease risk, although its effect was smaller in magnitude. Cholesterol level showed a statistically significant but modest negative association with heart disease in the final model.\u003c/p\u003e \u003cp\u003eThese results align with existing clinical literature showing that age, male sex, and symptoms of cardiac stress (such as exercise-induced angina) are well-established predictors of cardiovascular disease. The strong predictive accuracy of the model highlights the value of simple, routinely collected clinical variables in identifying individuals at elevated risk.\u003c/p\u003e \u003cp\u003eOverall, this analysis provides evidence supporting the importance of demographic and clinical markers in heart disease prediction and emphasizes the usefulness of logistic regression as a practical tool for risk stratification. Future research using longitudinal or more detailed clinical datasets may further clarify causal pathways and enhance predictive performance.\u003c/p\u003e"},{"header":"RECOMMENDATIONS","content":"\u003cp\u003eBased on the study findings and in line with JMIR\u0026rsquo;s emphasis on digital health relevance and clinical applicability, the following recommendations are proposed:\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003eClinical Screening and Risk Stratification\u003cbr\u003e\u0026nbsp;Exercise-induced angina should be prioritized as a key screening variable in both clinical and digital health risk-assessment tools, given its strong association with heart disease. Incorporating symptom-based indicators alongside demographic factors may improve early detection.\u003c/li\u003e\n \u003cli\u003eIntegration Into Digital Health Tools\u003cbr\u003e\u0026nbsp;The predictors examined in this study are routinely collected and well-suited for integration into electronic health records, clinical decision support systems, and mobile health applications aimed at cardiovascular risk assessment.\u003c/li\u003e\n \u003cli\u003eInterpretation of Cholesterol Findings\u003cbr\u003e\u0026nbsp;The inverse association between cholesterol level and heart disease observed in this dataset should be interpreted cautiously. Future studies should account for medication use (eg, statins), fasting status, and disease management history to clarify this relationship.\u003c/li\u003e\n \u003cli\u003eExpansion of Predictive Models\u003cbr\u003e\u0026nbsp;Future research should incorporate additional clinical and behavioral risk factors such as smoking status, diabetes, hypertension, physical activity, and medication use to improve explanatory power and predictive accuracy.\u003c/li\u003e\n \u003cli\u003eLongitudinal and Diverse Populations\u003cbr\u003e\u0026nbsp;Longitudinal datasets are needed to assess causal relationships and temporal risk trajectories. Additionally, more balanced samples with respect to sex and broader population representation would enhance generalizability.\u003c/li\u003e\n \u003cli\u003eValidation and External Testing\u003cbr\u003e\u0026nbsp;External validation using independent datasets is recommended before clinical or digital implementation to ensure robustness and transportability of the predictive model.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eBenjamin, E. J., Muntner, P., Alonso, A., Bittencourt, M. S., Callaway, C. W., Carson, A. P., \u0026hellip; \u0026amp; Virani, S. S. (2019). Heart disease and stroke statistics\u0026mdash;2019 update:A report from theAmerican HeartAssociation. \u003cem\u003eCirculation, 139\u003c/em\u003e(10), e56\u0026ndash;e528. https://doi.org/10.1161/CIR.0000000000000659\u003c/li\u003e\n \u003cli\u003eGrundy, S. M., Stone, N. J., Bailey, A. L., Beam, C., Birtcher, K. K., Blumenthal, R. S., \u0026hellip; \u0026amp; Yeboah, J. (2019). 2018 AHA/ACC guideline on the management of blood cholesterol. \u003cem\u003eJournal of the American College of Cardiology, 73\u003c/em\u003e(24), e285\u0026ndash;e350. https://doi.org/10.1016/j.jacc.2018.11.003\u003c/li\u003e\n \u003cli\u003eHarrell, F. E. (2015). \u003cem\u003eRegression modeling strategies: With applications to linear models, logistic regression, and survival analysis\u0026nbsp;\u003c/em\u003e(2nd ed.). Springer. https://doi.org/10.1007/978-3-319-19425-7\u003c/li\u003e\n \u003cli\u003eLanza, G. A., \u0026amp; Crea, F. (2010). Primary coronary microvascular dysfunction: Clinical presentation, pathophysiology, and management. \u003cem\u003eCirculation, 121\u003c/em\u003e(21), 2317\u0026ndash;2325. https://doi.org/10.1161/CIRCULATIONAHA.109.900191\u003c/li\u003e\n \u003cli\u003eVirani, S. S., Alonso, A., Aparicio, H. J., Benjamin, E. J., Bittencourt, M. S., Callaway, C. W., \u0026hellip; \u0026amp; Tsao, C. W. (2021). Heart disease and stroke statistics\u0026mdash;2021 update:A report from theAmerican HeartAssociation. \u003cem\u003eCirculation, 143\u003c/em\u003e(8), e254\u0026ndash;e743. https://doi.org/10.1161/CIR.0000000000000950\u003c/li\u003e\n \u003cli\u003eWorld Health Organization. (2023). \u003cem\u003eCardiovascular diseases (CVDs).\u0026nbsp;\u003c/em\u003ehttps://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds)\u003c/li\u003e\n \u003cli\u003eYusuf, S., Joseph, P., Rangarajan, S., Islam, S., Mente, A., Ramesh, D., \u0026hellip; \u0026amp; Teo, K. K. (2020). Modifiable risk factors, cardiovascular disease, and mortality in 155,722 individuals from 21 high-, middle-, and low-income countries (PURE):A prospective cohort study. \u003cem\u003eThe Lancet, 395\u003c/em\u003e(10226), 795\u0026ndash; 808. https://doi.org/10.1016/S0140-6736(19)32008-2\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1 and 2 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8889390/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8889390/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eCardiovascular disease remains the leading cause of morbidity and mortality worldwide. Early identification of individuals at high risk of heart disease is essential for effective prevention and clinical intervention. Routinely collected demographic and clinical variables may offer valuable predictive insight when evaluated using appropriate statistical models. This study aimed to examine whether age, sex, cholesterol level, and exercise-induced angina significantly predict the likelihood of heart disease among adults using a publicly available heart failure dataset.\u003c/p\u003e \u003cp\u003eA secondary analysis was conducted using data from 918 adults obtained from the Kaggle Heart Failure Dataset. Heart disease status (yes/no) was modeled as a binary outcome. Multivariable logistic regression was used to assess associations between predictors and heart disease. Model assumptions were evaluated using variance inflation factors, Box\u0026ndash;Tidwell tests for linearity of the logit, influence diagnostics, and goodness-of-fit statistics. Discriminatory performance was assessed using receiver operating characteristic analysis.\u003c/p\u003e \u003cp\u003eThe logistic regression model significantly improved prediction of heart disease compared with the intercept-only model (likelihood ratio χ\u0026sup2;₄=370.14, \u003cem\u003eP\u003c/em\u003e\u0026lt;.001). The model demonstrated strong discrimination (area under the curve\u0026thinsp;=\u0026thinsp;0.846) and good calibration (Hosmer\u0026ndash;Lemeshow \u003cem\u003eP\u003c/em\u003e=.667). All predictors were statistically significant (\u003cem\u003eP\u003c/em\u003e\u0026lt;.001). Increasing age was associated with higher odds of heart disease (odds ratio [OR]\u0026thinsp;=\u0026thinsp;1.051, 95% CI 1.032\u0026ndash;1.070). Males had substantially higher odds than females (OR\u0026thinsp;=\u0026thinsp;3.494, 95% CI 2.299\u0026ndash;5.309). Exercise-induced angina was the strongest predictor (OR\u0026thinsp;=\u0026thinsp;9.911, 95% CI 6.873\u0026ndash;14.293). Cholesterol level showed a statistically significant but inverse association with heart disease (OR\u0026thinsp;=\u0026thinsp;0.995, 95% CI 0.994\u0026ndash;0.997).\u003c/p\u003e \u003cp\u003eAge, sex, cholesterol level, and exercise-induced angina were significant predictors of heart disease in this dataset. Exercise-induced angina and male sex demonstrated particularly strong associations. These findings highlight the value of simple, routinely collected clinical indicators for heart disease risk stratification and support the use of logistic regression as an effective analytical approach in population-level cardiovascular research.\u003c/p\u003e","manuscriptTitle":"Predictors of Cardiovascular Disease Among Adults: A Multivariable Logistic Regression Analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-18 06:55:33","doi":"10.21203/rs.3.rs-8889390/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c381e2dc-dc84-449a-ae15-95452e61d0ee","owner":[],"postedDate":"February 18th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":63111113,"name":"Epidemiology"}],"tags":[],"updatedAt":"2026-02-18T06:55:33+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-18 06:55:33","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8889390","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8889390","identity":"rs-8889390","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.