⚙
AI-generated deep summary
by claude@2026-07, 2026-07-03
· read from full text
ⓘ
The paper proposes a Hybrid Stacked Sparse Autoencoder (HSSAE) model to predict type 2 diabetes using sparse numerical healthcare data, integrating L1 and L2 regularization with binary cross-entropy loss, dropout, and batch normalization to improve feature selection and training stability. The authors compare HSSAE against traditional classifiers (Decision Tree, Random Forest, K-Nearest Neighbors, Naïve Bayes) and a baseline stacked sparse autoencoder (SSAE), reporting that HSSAE achieved the highest accuracy of 88.73% on a sparse diabetic dataset. A major caveat explicitly stated is that the work is a preprint and has not been peer reviewed, with data described as potentially preliminary. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.
Abstract
Sparse numerical datasets are common in applied mathematics, astronomy, finance, and healthcare, posing challenges due to their high dimensionality and sparsity. Most values are zero, complicating optimal feature selection. To address this, the Hybrid Stacked Sparse Autoencoder (HSSAE) integrates L1 and L2 regularization with binary cross-entropy loss to enhance feature selection efficiency. L1 regularization penalizes large weights, simplifying data representations, while L2 regularization prevents overfitting by limiting the total weight size. Additionally, the dropout technique improves model performance by randomly deactivating neurons during training, ensuring the model relies on the active neurons. Batch normalization stabilizes weight distributions, reducing computational time and accelerating convergence. The proposed HSSAE model was compared with traditional classifiers such as Decision Tree, Random Forest, K-Nearest Neighbors, Naïve Bayes and a deep learning algorithm Stacked Sparse Autoencoder model (SSAE) using a sparse diabetic dataset. Quantitively, the proposed HSSAE produced the highest accuracy (88.73%) compared to the traditional classifiers and SSAE Model. The ability of the proposed HSSAE model to generate good feature selection makes the model robust and suitable for any kind of sparse data applications, especially for sensitive applications such as the healthcare diabetes dataset which requires high accuracy in the prediction.
Full text
2,511 characters
· extracted from
oa-doi-fallback
· click to expand
This is a preprint and has not been peer reviewed. Data may be preliminary.
A Hybrid Stacked Sparse Autoencoder (HSSAE) Model for Predicting Type 2 Diabetes
Abstract
Sparse numerical datasets are common in applied mathematics, astronomy, finance, and healthcare, posing challenges due to their high dimensionality and sparsity. Most values are zero, complicating optimal feature selection. To address this, the Hybrid Stacked Sparse Autoencoder (HSSAE) integrates L1 and L2 regularization with binary cross-entropy loss to enhance feature selection efficiency. L1 regularization penalizes large weights, simplifying data representations, while L2 regularization prevents overfitting by limiting the total weight size. Additionally, the dropout technique improves model performance by randomly deactivating neurons during training, ensuring the model relies on the active neurons. Batch normalization stabilizes weight distributions, reducing computational time and accelerating convergence. The proposed HSSAE model was compared with traditional classifiers such as Decision Tree, Random Forest, K-Nearest Neighbors, Naïve Bayes and a deep learning algorithm Stacked Sparse Autoencoder model (SSAE) using a sparse diabetic dataset. Quantitively, the proposed HSSAE produced the highest accuracy (88.73%) compared to the traditional classifiers and SSAE Model. The ability of the proposed HSSAE model to generate good feature selection makes the model robust and suitable for any kind of sparse data applications, especially for sensitive applications such as the healthcare diabetes dataset which requires high accuracy in the prediction.
Supplementary Material
File (manuscript.docx)
- Download
- 1.23 MB
Information & Authors
Information
Version history
Copyright
This work is licensed under a Non Exclusive No Reuse License.
Collection
Authors
Metrics & Citations
Metrics
Article Usage
559views
204downloads
Citations
Download citation
Abdussamad, Hanita Daud, Rajalingam Sokkalingam, et al.
A Hybrid Stacked Sparse Autoencoder (HSSAE) Model for Predicting Type 2 Diabetes. Authorea. 21 February 2025.
DOI: https://doi.org/10.22541/au.174013925.50853039/v1
DOI: https://doi.org/10.22541/au.174013925.50853039/v1
If you have the appropriate software installed, you can download article citation data to the citation manager of your choice. Simply select your manager software from the list below and click Download.
For more information or tips please see 'Downloading to a citation manager' in the Help menu.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.