REVEAL-MM: Retrospective Evaluation of Variables in Early Assessment and Landmark trends in Multiple Myeloma – a US Claims-Based Case-Control Study

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-06, 2026-06-24 · read from full text

REVEAL-MM is a retrospective 1:1 matched case-control study using Optum Clinformatics administrative claims (9,466 adults ≥50) to characterize healthcare utilization patterns in the 24 months before multiple myeloma (MM) diagnosis and to build predictive models for identifying at-risk individuals 12, 9, and 6 months pre-diagnosis. Using features derived from diagnostic, procedure, prescription, and physician-visit claims (with clinician grouping and expert review), LASSO and Random Forest models found distinct encounter patterns associated with MM as early as 12 months before diagnosis, with predictive performance improving as diagnosis neared (peak AUC 0.826 at 6 months). The authors note that the study uses claims data and is subject to limitations inherent to retrospective claims-based identification, including reliance on code capture and model generalizability beyond the dataset. This paper is centrally about endometriosis and/or adenomyosis—however, it does not discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Multiple myeloma (MM) frequently presents with non-specific symptoms that overlap other common conditions, often leading to diagnostic delays and poorer clinical outcomes. Although diagnostic delays in MM are well recognized, the healthcare utilization patterns that precede MM diagnosis are not well defined. This retrospective, 1:1 matched case-control study used US administrative claims data from 9,466 patients to identify early signals of MM and evaluate whether such data could be used to predict individuals at risk of MM prior to their formal diagnosis. All diagnostic, procedural, prescription, and physician-visit claims were evaluated at 12, 9, and 6 months within the two years before diagnosis. Interpretable predictive models (LASSO and Random Forest) identified distinct encounter patterns associated with MM as early as 12-months before diagnosis. Predictive performance across all models increased as diagnosis approached, with the machine learning model reaching a peak area under the curve (AUC) of 0.826 at 6 months. Claims consistent with typical MM manifestations, including anemia-, musculoskeletal-, and M-protein–related testing, were more common prior to MM diagnosis. These findings suggest that routinely collected claims data could support earlier identification and evaluation of individuals at risk of MM, enabling more timely diagnosis and improved outcomes.
Full text 108,130 characters · extracted from preprint-html · click to expand
REVEAL-MM: Retrospective Evaluation of Variables in Early Assessment and Landmark trends in Multiple Myeloma – a US Claims-Based Case-Control Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article REVEAL-MM: Retrospective Evaluation of Variables in Early Assessment and Landmark trends in Multiple Myeloma – a US Claims-Based Case-Control Study Faith Davies, Beth Faiman, Hayley Beer, Anne Quinn Young, Katie Joyner, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8849252/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 9 You are reading this latest preprint version Abstract Multiple myeloma (MM) frequently presents with non-specific symptoms that overlap other common conditions, often leading to diagnostic delays and poorer clinical outcomes. Although diagnostic delays in MM are well recognized, the healthcare utilization patterns that precede MM diagnosis are not well defined. This retrospective, 1:1 matched case-control study used US administrative claims data from 9,466 patients to identify early signals of MM and evaluate whether such data could be used to predict individuals at risk of MM prior to their formal diagnosis. All diagnostic, procedural, prescription, and physician-visit claims were evaluated at 12, 9, and 6 months within the two years before diagnosis. Interpretable predictive models (LASSO and Random Forest) identified distinct encounter patterns associated with MM as early as 12-months before diagnosis. Predictive performance across all models increased as diagnosis approached, with the machine learning model reaching a peak area under the curve (AUC) of 0.826 at 6 months. Claims consistent with typical MM manifestations, including anemia-, musculoskeletal-, and M-protein–related testing, were more common prior to MM diagnosis. These findings suggest that routinely collected claims data could support earlier identification and evaluation of individuals at risk of MM, enabling more timely diagnosis and improved outcomes. Health sciences/Risk factors Biological sciences/Cancer/Haematological cancer/Myeloma Multiple myeloma Diagnostic delay Healthcare utilization patterns Predictive modelling Claims data Figures Figure 1 Figure 2 Figure 3 Introduction Multiple myeloma (MM) poses a significant diagnostic challenge due to its non-specific symptoms such as bone pain, fatigue, and recurrent infections, which frequently overlap with other conditions including osteoporosis, chronic kidney disease, and degenerative spinal disorders. 1 – 3 The wide spectrum of non-specific symptoms seen in MM may lead to diagnostic delays or mis-referrals to non-hematology specialists, including, nephrologists, orthopedic surgeons, or rheumatologists; causing further delays. 4 Compared with other cancers, MM has one of the highest proportions of patients requiring three or more consultations before specialist referral. 5 Such delays are associated with a poorer disease stage at diagnosis which is subsequently associated with poorer survival. 1 In addition, patients with longer diagnostic intervals are also reported to have reduced disease-free survival, and more treatment-associated secondary complications. 4 These delays disproportionately affect racial and ethnic minority populations; African American and Hispanic/Latino patients with MM experience longer intervals from symptom onset to treatment initiation compared with their White counterparts. 6 Addressing these disparities is crucial to achieving equitable outcomes for all MM patients and reducing preventable morbidity. Precursor conditions such as monoclonal gammopathy of undetermined significance (MGUS) and smoldering multiple myeloma (SMM) represent opportunities for earlier monitoring and intervention, however these states are identified in less than 10% of patients. 7 , 8 As a result, most individuals progress to symptomatic MM without prior surveillance, contributing to missed opportunities for timely diagnosis and downstream delays. 7 Although diagnostic delays in MM are well documented, the specific healthcare utilization patterns associated with these delays remain poorly characterized. Existing evidence points to healthcare system-, provider-, and patient-related factors, 9 yet how these factors translate in real-world diagnostic pathways is not well defined. Large-scale claims databases provide an opportunity to examine diagnostic patterns at scale and to identify recurring encounter trends that may represent early signals of MM prior to a formal diagnosis. Statistical and machine learning (artificial intelligence) models can be applied to test this hypothesis. 10 , 11 When implemented within real-world health systems, such models may be able to flag high-risk individuals months before symptom onset, supporting earlier referral and improved patient outcomes. In addition, timely diagnosis may decrease preventable complications and downstream costs associated with delayed diagnosis. 12 Furthermore, by generating actionable risk signals in real time, these models can support primary care and non-hematology specialists in recognizing patterns suggestive of MM, helping bridge awareness gaps and accelerate timely referral to hematology and diagnosis. Methods Study Overview REVEAL-MM is a retrospective, 1:1 matched case-control study designed to characterize pre-diagnostic healthcare utilization patterns and to develop predictive models that identify individuals at increased risk of MM prior to their formal diagnosis. This study compares longitudinal diagnostic, procedural, prescription, and physician-visit patterns between patients who later developed MM and matched controls without MM throughout the study period. Data Source Data were extracted from the Optum® Clinformatics® Data Mart (CDM), a large, geographically diverse, US administrative claims database containing de-identified longitudinal medical and prescription claims from commercially insured and Medicare Advantage beneficiaries. CDM includes demographics, diagnostic, procedure, prescription, and inpatient and outpatient encounters with associated ICD-9/10 codes. Study Population The pre-MM cohort included adults aged ≥ 50 years with a confirmed MM diagnosis between January 1, 2020, and January 31, 2024, defined by at least two MM-coded claims (ICD-10-CM: C90.x; ICD-9-CM: 203.x) occurring ≥ 30 days apart. Patients needed at least 24 months of continuous data coverage prior to their first MM diagnosis (the index date) and were excluded if they had any malignant cancer diagnosis or MGUS during the study period. The control cohort comprised of adults aged ≥ 50 years without an MM diagnosis during the study period, matched 1:1 via propensity score matching on age, gender, race, geographic region, insurance type, Charlson Comorbidity Index (CCI) and National Cancer Institute (NCI) comorbidity scores assessed three months prior to their last medical encounter date (index date). For controls, the index date was assigned from the distribution of MM cohort index dates to ensure temporal alignment. Covariate balance was achieved with standardized mean differences below 0.1. Study Design All available healthcare encounter data (including diagnostic, procedure, and prescription codes, and physician visits by specialty) were extracted for the 24-month period prior to the index date for both cohorts. Encounter information was summarized at three pre-diagnostic timepoints: T-12 months (12 months before index date), T-9 months (9 months before index date), and T-6 months (6 months before index date). The study design is shown in Fig. 1 . Data Preprocessing At each timepoint, both frequency variables and binary indicators were generated for all encounter categories. The full analysis set (FAS) had 1,287, 1,451 and 1,619 independent variables for T-12, T-9, and T-6 periods, respectively. Diagnosis codes were encoded as binary indicators reflecting the presence or absence of each diagnosis. Procedure and prescription codes were represented as frequency counts within the corresponding time window. No missing values were observed, as the absence of a code was interpreted as zero. All variables were first screened using Bonferroni-adjusted univariate non-parametric tests against the outcome variable (Pearson’s chi-square test for diagnosis codes and Wilcoxon rank-sum tests for all other numerical variables). Variables meeting statistical significance thresholds were subsequently reviewed by clinical experts. Claims codes were not presented verbatim; instead, they were translated into expanded clinical descriptors to enhance interpretability. To generate clinically meaningful features, subject-matter experts grouped diagnostic, procedure, and prescription codes into predefined clinical categories (e.g., “Anemias,” “Renal complications,” “Musculoskeletal pain”). Physician visits were categorized by specialty. The grouping framework is provided in Supplementary Table 1 . Following clinical review and grouping, the final FAS used for model development for T-12, T-9, T-6 had 43, 52, and 66 independent variables respectively. Four complementary modelling strategies were implemented to balance interpretability, variable selection, and predictive accuracy. LASSO provided an interpretable model; Random Forest and Deep Learning captured non-linear relationships; and PLS-DA addressed multicollinearity. Model 1: Least Absolute Shrinkage and Selection Operator (LASSO) Regression (Penalized linear modelling) LASSO regression was employed for sparse feature selection and risk prediction. Performance was evaluated on the hold-out test set using Area Under the receiver operating characteristic Curve (AUC). Non-zero coefficients were interpreted as clinically actionable risk factors. Model 2: Random Forest (Ensemble Tree-Based Modelling) Random Forest was implemented to capture non-linear relationships and interactions among predictors. Final model performance was assessed on the test set using AUC. Variable importance was computed via mean decrease in impurity. Model 3: Partial Least Squares Discriminant Analysis (PLS-DA) PLS-DA was applied to address potential multicollinearity in the data (e.g., correlated lab panels). This supervised dimension-reduction technique maximizes covariance between predictors and the outcome. The number of latent components was determined by 10-fold cross-validation. Model discrimination was evaluated using AUC on the test set and loading plot visualized relationships between original variables and latent components. Model 4: Deep Learning with Sequential Neural Network A feed-forward sequential neural network was constructed using Keras/TensorFlow to model the complex, non-linear patterns inaccessible to traditional methods. 10-fold cross-validation was used to optimize the model structure. All model fittings were conducted in R v4.3 (LASSO, RF, PLS-DA and DL via glmnet, random Forest, mixOmics and Kera3 packages). Ethical Considerations This study utilized a commercially available, de-identified database, compliant with the Health Insurance Portability and Accountability Act (HIPAA). As the research did not involve human subjects or access to personally identifiable information, this study was exempt from Institutional Review Board (IRB) oversight and did not require Ethics approval. Results Baseline characteristics A total of 4,733 patients met the inclusion criteria for the pre-MM cohort and were matched (1:1) with 4,733 controls, resulting in a total study population of 9,466 patients. Baseline demographic and clinical characteristics of the matched cohorts are shown in Table 1 . The mean age of the population was 74.1 years (SD 8.14), and 50% were female. Racial distribution (51% White) and geographic region (45% residing in the US South Census Region) were comparable between cohorts. Medicare Advantage was the primary insurer for 86% of patients. Comorbidity burden was well balanced following propensity score matching, with similar mean CCI scores (2.29 for pre-MM vs. 2.18 for controls) and NCI comorbidity scores (0.71 vs. 0.67, respectively). All standardized mean differences for matching variables were < 0.1, confirming excellent covariate balance. Table 1 Baseline characteristics of the study population. CCI = Charlson Comorbidity Index; NCI = National Cancer Institute Comorbidity Index. COM = Commercial health insurance. MCR = Medicare health insurance. SD = Standard deviation. Overall Pre-MM cohort Non-MM cohort n 9466 4733 4733 Mean (SD) Mean (SD) Mean (SD) Age 73.96 (8.25) 74.10 (8.14) 73.81 (8.36) CCI 2.24 (2.33) 2.29 (2.29) 2.18 (2.36) NCI* 0.69 (0.73) 0.71 (0.72) 0.67 (0.73) Count (%) Count (%) Count (%) Gender Female 4,717 (50) 2,360 (50) 2,357 (50) Male 4,744 (50) 2,370 (50) 2,374 (50) Undisclosed 5 (< 0.1) 3 (< 0.1) 2 (< 0.1) Race White 4,847 (51) 2,419 (51) 2,428 (51) Asian 205 (2.2) 104 (2.2) 101 (2.1) Black 1,579 (17) 786 (17) 793 (17) Undisclosed 298 (3.1) 149 (3.1) 149 (3.1) Other 2,537 (27) 1,275 (27) 1,262 (27) Region Midwest 1,858 (20) 938 (20) 920 (19) Northeast 1,360 (14) 673 (14) 687 (15) South 4,278 (45) 2,149 (45) 2,129 (45) West 1,956 (21) 966 (20) 990 (21) Other 14 (0.1) 7 (0.1) 7 (0.1) Insurance Type COM 1,290 (14) 624 (13) 666 (14) MCR 8,176 (86) 4,109 (87) 4,067 (86) LASSO Regression Model Performance LASSO regression was used to interpret which variables most strongly linked with an MM diagnosis. The model demonstrated modest and progressively improving discrimination as the index date approached. AUC increased from 0.6727 at T-12M to 0.6928 at T-9M and 0.6966 at T-6M (Fig. 2 ). Classification accuracy At T-12M, overall accuracy was 62.0%, with the model correctly identifying 2,127 pre-MM cases (sensitivity 44.9%) and 3,743 controls (specificity 79.1%). Accuracy improved to 64.0% at T-9M and 64.3% at T-6M, driven by increased sensitivity (44.9% to 59.3%) and reduced specificity (79.1% to 69.2%) ( Supplementary Table 2 ). Variable interpretation LASSO coefficients provided insight into the relative importance and direction of associations (Table 2 ). At T-12M, variables associated with increased MM risk included diagnostic codes for anemias and neutropenia, and procedure codes related to M-protein and immunoglobulin evaluation. Additional high-risk indicators included gastroesophageal reflux disease (GERD) with esophagitis, cardiovascular disorders, and vaccination- and infection-related encounters (e.g., influenza and pneumococcal vaccination, vaccination administration, other infection concerns). Variables consistently associated with lower MM risk included dementia, depression (unspecified), altered mental status, antiviral prescriptions (e.g., nirmatrelvir/ritonavir), and visits to non-physician providers. Table 2 LASSO coefficients for codes across T-12M associated with higher or lower MM diagnosis risk. Positive coefficients are labelled in green. Negative coefficients are labelled in red. DX = diagnosis code; HCP = Healthcare professional visit; PROC = procedure code; GNM = Prescription code. Claim code Variable description Coefficient GNM_G2 Flu vaccination 0.360882741 DXG_1 Anemias 0.169056209 PROC_G1 M-protein / immunoglobulin quantification 0.078367363 DXG_7 Cough 0.071624292 PROC_G12 Routine follow-up office visit 0.063930757 HCP_00030 Haematology & oncology visit 0.062831215 DXG_11 Other infection concern 0.050016439 DXG_2 Neutropenia 0.049724713 GNM_G3 Pneumococcal vaccination 0.043819732 DX_K210 GERD with esophagitis 0.041424254 PROC_G13 Observation/short-term hospital care 0.039689921 DXG_4 Cardiovascular disorders 0.035108252 PROC_G8 Vaccination administration 0.032322284 PROC_G10 COVID-19 diagnostic testing 0.029408403 PROC_92012 Interim ophthalmic exam, established pt 0.026418157 PROC_88342 Immunochem/cytochem first antibody 0.017946621 PROC_88313 Special stains (group 2) 0.017527554 DXG_9 COVID or viral exposure 0.015766174 DXG_6 Musculoskeletal pain 0.011627152 DX_D61818 Other pancytopenia 0.007789679 PROC_83883 Nephelometry assay (not specified) 0.006681015 PROC_G9 Metabolic panel 0.00425867 PROC_G14 Nursing facility/long-term care visit -0.0339063 GNM_00798 Nirmatrelvir/ritonavir -0.044641046 HCP_00077 Registered nurse practitioner visit -0.048882152 DX_R4182 Altered mental status (unspecified) -0.067758286 PROC_G15 Routine monitoring -0.068581005 DX_F32A Depression (unspecified) -0.073602922 PROC_1159F Medication list documented -0.076745669 HCP_00055 Other non-physician provider visit -0.08223409 Dementia Dementia -0.113692015 Temporal Trends Longitudinal patterns from T-12M to T-6M demonstrated strengthening associations for several high-risk variables. For example, the coefficient for M-protein/immunoglobulin testing increased from 0.0784 (T-12M) to 0.0995 (T-9M) and 0.1599 (T-6M), while neutropenia rose from 0.0497 to 0.0797 and 0.0965, respectively. Some early signals, such as cough, were not retained at later timepoints. Variables associated with lower MM risk, particularly dementia, depression, and altered mental status, remained consistently negative across all periods ( Supplementary Tables 3 and 4 ). Random Forest A Random Forest model combines many small decision rules from patient history, giving an indication of which variables most often point toward early MM. The model showed modest discriminatory performance, with AUC values increasing slightly closer to diagnosis: 0.6316 at T-12M, 0.6544 at T-9M, and 0.6537 at T-6M ( Supplementary Fig. 1 ). Classification accuracy ranged from 62.4% at T-12M to 63.7% at T-6M ( Supplementary Table 5 ). Variable importance rankings were consistent across time points and consistent with the LASSO findings. Hematologic codes, infection- and vaccination-related codes, and routine care encounters were among the top contributors ( Supplementary Tables 6, 7, and 8 ). Partial Least Squares Discriminant Analysis (PLS-DA) PLS-DA groups related clinical codes into clusters to identify whether patients who later develop MM have distinct healthcare utilization patterns. PLS-DA AUC values improved as the index date approached: 0.6758 at T-12M, 0.6947 at T-9M, and 0.692 at T-6M ( Supplementary Fig. 2 ). Loading plots demonstrated consistent variable clustering patterns across all time periods ( Supplementary Fig. 3 ). Hematologic evaluation codes (e.g., anemias, neutropenia, bone marrow, cytogenetic testing, peripheral blood analysis) formed stable, contiguous clusters. Routine care codes (e.g., established patient visits, routine monitoring) also appeared in adjacent regions. Comorbidity-related codes, including dementia, depression (unspecified), and altered mental status, clustered separately and consistently across all timepoints. Vaccine- and infection-related codes behaved less uniformly. For example, influenza vaccine codes were outliers at T-12M and T-9M, while COVID-19 vaccine codes showed similar behavior at T-6M. Sequential Neural Network (SNN) Model Performance A Sequential Neural Network uses artificial intelligence to learn complex, changing patterns in patients’ medical histories, allowing it to detect early signs of MM that simpler models may miss. The model demonstrated the strongest overall performance among all models, with AUC improving steadily as diagnosis approached: 0.7654 at T-12M, 0.8001 at T-9M, and 0.8257 at T-6M (Fig. 3 ). Classification Accuracy Consistent with AUC results, classification accuracy improved at each timepoint. At T-12M, the model achieved 68.3% accuracy, correctly identifying 2,943 pre-MM cases and 3,541 controls. Accuracy increased to 70.7% at T-9M and 72.8% at T-6M, with corresponding increases in correct MM classifications (3,581 and 3,734, respectively) ( Supplementary Table 9 ). Discussion In this large, US claims-based matched case-control study, clinical signals of MM were identified as early as 12 months before MM diagnosis using routinely collected administrative claims data. Across all four models used to analyze claims data, discrimination improved as the index date approached, reflecting the clear clinical manifestations of pre-diagnostic MM over time. 13 , 14 The SNN model using artificial intelligence, achieved the strongest performance, with an AUC of 0.8257 at T-6M, outperforming traditional linear and tree-based models. These findings highlight the potential of deep learning techniques to capture complex, nonlinear patterns in high-dimensional healthcare data. 15 The SNN model’s strong prediction ability shows it could be used at the population level for screening, helping to identify high-risk individuals earlier, before complications develop. Across all modelling approaches, consistent hematological abnormality codes were observed, particularly anemia, which emerged among the strongest predictors of a MM diagnosis. These findings align with known early manifestations of MM. 13,14 Similarly, claims related to diagnostic evaluation for monoclonal gammopathy (e.g., M-protein or immunoglobulin testing) showed progressively stronger associations closer to diagnosis, suggesting that clinicians may be responding to subtle or unexplained laboratory abnormalities well before MM is formally confirmed. Unexpected claims codes associated with gastro-esophageal reflux disease (GERD) and cardiovascular disorders also emerged, warranting further investigation as potential early indicators or markers of diagnostic complexity. Vaccination- (including COVID-19, influenza, and pneumococcal vaccines) and infection-related codes were also frequently selected by LASSO and featured prominently in importance measures for Random Forest and PLS-DA models. In the PLS-DA loading plots, however, these codes showed less consistent clustering across timepoints, with influenza vaccine codes appearing as outliers at T-12M and T-9M and COVID-19 vaccine codes showing similar outlier behavior at T-6M. This may reflect increased healthcare encounters prompted by nonspecific symptoms or immune dysregulation that precede MM diagnosis. 3 Alternatively, these patterns may capture confounding healthcare utilization patterns, as patients with underlying hematologic or constitutional symptoms may seek care more frequently. 16 Moreover, the observed associations may partly reflect temporary effects, such as seasonal vaccination patterns. 17 Further investigation is needed to understand these patterns. Interestingly, several comorbidity-related codes, including dementia, depression, and altered mental status, were consistently associated with a reduced likelihood of developing MM. These inverse associations may represent differences in healthcare utilization, competing management of existing comorbidities, or under-evaluation of hematologic abnormalities/ concerns in these patients. This phenomenon has been described in other cancer detection studies, where lower diagnostic intensity is applied to medically complex populations. 18 , 19 The relative performance of the models provides important insights for clinical translation. While LASSO and Random Forest demonstrated modest discriminatory ability (AUC ~ 0.63–0.70), they provided interpretable variable importance patterns that can inform clinical hypotheses. In contrast, the SNN, though less interpretable, achieved substantially higher discrimination, suggesting that deep neural architectures may be better suited to capturing the nonlinear, evolving patterns characteristic of MM’s pre-diagnostic phase. While these results are promising, several limitations warrant consideration. First, models were trained on administrative claims data, which lack granular laboratory values (e.g., specific hemoglobin and M-protein levels) and may miss subtle clinical cues detectable in structured electronic healthcare records data or clinician notes. Second, although propensity score matching effectively balanced patient characteristics, residual confounding and misclassification inherent in claims-based data analysis remain possible. Third, performance metrics of models, though encouraging, may not be sufficient for unsupervised population-wide screening. Instead, these models may be more appropriate as decision-support tools to prompt targeted evaluation in high-risk subgroups. Finally, external validation in broader populations, particularly in integrated health systems with richer clinical data across geographies is needed to assess generalizability and real-world applicability. Despite these limitations, this study demonstrates the potential of administrative claims data to support earlier identification of individuals at risk for MM, particularly at the health system level. The importance of earlier identification has never been stronger, with emerging clinical evidence suggesting that earlier treatment, particularly among patients with high-risk precursor conditions, may delay or prevent progression to overt MM and its associated complications. 20 While claims data are not universally available at the point of care, they can be analyzed centrally within health systems to identify patterns of healthcare use that signal increased risk. Integrating these system-level insights with laboratory results and electronic health record data offers a practical path to improve risk identification at the point of care. Together, this approach supports a shift toward earlier evaluation, timely referral, and more proactive care, with the potential to improve outcomes and reduce avoidable complications for patients with MM. Conclusion Overall, our findings support the feasibility of using real-world administrative claims data to detect patterns indicative of emerging MM up to a year before diagnosis. Hematologic abnormalities, diagnostic evaluation patterns, and infection-related encounters emerged as consistent early signals. Among the tested models, deep learning SNN demonstrated the strongest predictive ability, emphasizing its potential role in future risk-stratification initiatives within population-level health systems. Further research is required to explore some of the trends highlighted in this paper, with the potential to reduce diagnostic delays and improve clinical outcomes. Declarations Competing Interests statement Dr Davies has received consultancy fees from GSK, Sanofi, Bristol Myers Squibb, Regeneron, Johnson & Johnson, and Takeda. Dr Mikhael has received consultancy fees from Sanofi, Bristol Myers Squibb, Johnson & Johnson, and Menarini. Ms Joyner and Ms Morgan have received consultancy fees, donated to Myeloma Patients Europe, from Johnson & Johnson, Pfizer and GSK. Ms Young has received consultancy fees, donated to Multiple Myeloma Research Foundation, from Johnson & Johnson, AstraZeneca and Bristol Myers Squibb. Dr Faiman and Ms Beer have received consultancy fees from Johnson & Johnson. Dr Huo and Dr Bartlett are employees of Johnson & Johnson and hold stock options in the company. Ms Attfield, Dr. Han and Dr. Egbase are employees of VML Health. Author Contributions SH, JBB and YH were responsible for designing the study. YH was responsible for data extraction from the from the Optum® CDM and analyzing data. FED, BF, HB, AQY, KJ, KM, SH, JBB, YH, GA, DE, JM were responsible for interpreting the results. GA and DE were responsible for writing the report. FED, BF, HB, AQY, KJ, KM, SH, JBB, YH and JM provided feedback on the report. Acknowledgements This study was funded by J&J Innovative Medicine, LLC. Additional writing support was provided by VML Health. References Koshiaris C, Oke J, Abel L, Nicholson BD, Ramasamy K, Van den Bruel A. Quantifying intervals to diagnosis in myeloma: a systematic review and meta-analysis. BMJ open. 2018;8(6):e019758. Rajkumar SV. Multiple myeloma: 2022 update on diagnosis, risk stratification, and management. American journal of hematology. 2022;97(8):1086–1107. Friese CR, Abel GA, Magazu LS, Neville BA, Richardson LC, Earle CC. Diagnostic delay and complications for older adults with multiple myeloma. Leukemia & lymphoma. 2009;50(3):392–400. Kariyawasan C, Hughes D, Jayatillake M, Mehta A. Multiple myeloma: causes and consequences of delay in diagnosis. QJM: An International Journal of Medicine. 2007;100(10):635–640. Lyratzopoulos G, Neal RD, Barbiere JM, Rubin GP, Abel GA. Variation in number of general practitioner consultations before hospital referral for cancer: findings from the 2010 National Cancer Patient Experience Survey in England. The lancet oncology. 2012;13(4):353–365. Bhutani M, Blue BJ, Cole C, Badros AZ, Usmani SZ, Nooka AK, et al. Addressing the disparities: the approach to the African American patient with multiple myeloma. Blood Cancer J. 2023;13(1):189. doi: 10.1038/s41408-023-00961-0 . PMID: 38110338; PMCID: PMC10728116. Smith L, Carmichael J, Cook G, Shinkins B, Neal RD. Diagnosing myeloma in general practice: how might earlier diagnosis be achieved? Br J Gen Pract. 2022;72(723):462–3. Wong A, Jimenez-Zepeda VH, Rankin K, Sandhu I, Chu M, Delluc A, et al. Incidence and prevalence of clinically detected smoldering multiple myeloma within the general population: a retrospective observational cohort study. Blood Cancer J. 2025;15(1):149. Greene JA, Lea AS. Digital Futures Past - The Long Arc of Big Data in Medicine. N Engl J Med. 2019;381(5):480–485. Mittelman M, Shantsila A, Sagy I, et al. Prediction of multiple myeloma development using a machine learning model. Br J Haematol. 2024. (This is the key reference paper, justifying the use of ML models for MM risk prediction.) Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347–1358. Porteous A, Gibson S, Eddowes LA, Drayson M, Pratt G, Bowcock S, et al. An Economic Model to Establish the Costs Associated With Routes to Presentation for Patients With Multiple Myeloma in the United Kingdom. Value Health Reg Issues. 2023;35:27–33. Seesaghur A, Petruski-Ivleva N, Banks VL, et al. Clinical features and diagnosis of multiple myeloma: a population-based cohort study in primary care. BMJ Open. 2021;11(10):e052759. Rajkumar SV, Dimopoulos MA, Palumbo A, Blade J, Merlini G, Mateos MV, et al. International Myeloma Working Group updated criteria for the diagnosis of multiple myeloma. Lancet Oncol. 2014;15(12):e538-48. Miotto R, Wang F, Wang S, Jiang X, Dudley JT. Deep learning for healthcare: review, opportunities and challenges. Brief Bioinform. 2018;19(6):1236–1246. Virgilsen LF, Vedsted P, Jensen H, Frederiksen H, El-Galaly TC, Rasmussen LA. Diagnostic Window Prior to a Haematological Cancer Diagnosis and the Association With Patient Pathways: A Nationwide Register-Based Cohort Study on Healthcare Utilization in Denmark. Eur J Haematol. 2025;114(2):353–364. Padhi A, Bhatt P, Chauhan J, Rajyaguru B, Agarwal S, Chaudhary A, et al. Seasonality of Influenza and Optimizing Timing of Vaccination: Systematic Review and Meta-Analysis. Cureus. 2025;17(9):e93607. Renzi C, Kaushal A, Emery J, Hamilton W, Neal RD, Rachet B, Rubin G et al. Comorbid chronic diseases and cancer diagnosis: disease-specific effects and underlying mechanisms. Nat Rev Clin Oncol. 2019;16(12):746–761 Selie AP, van der Willik KD, Ikram MA, Labrecque JA, Schagen SB. Dementia and Cancer: Unravelling Methodological Biases in a Population-Based Cohort. Neuroepidemiology. 2025 Oct 7:1–10. Dimopoulos MA, Voorhees PM, Schjesvold F, Cohen YC, Hungria V, Sandhu I, et al. Daratumumab or Active Monitoring for High-Risk Smoldering Multiple Myeloma. N Engl J Med. 2025;392(18):1777–1788. Additional Declarations Yes there is potential conflict of interest. Dr Davies has received consultancy fees from GSK, Sanofi, Bristol Myers Squibb, Regeneron, Johnson & Johnson, and Takeda. Dr Mikhael has received consultancy fees from Sanofi, Bristol Myers Squibb, Johnson & Johnson, and Menarini. Ms Joyner and Ms Morgan have received consultancy fees, donated to Myeloma Patients Europe, from Johnson & Johnson, Pfizer and GSK. Ms Young has received consultancy fees, donated to Multiple Myeloma Research Foundation, from Johnson & Johnson, AstraZeneca and Bristol Myers Squibb. Dr Faiman and Ms Beer have received consultancy fees from Johnson & Johnson. Dr Huo and Dr Bartlett are employees of Johnson & Johnson and hold stock options in the company. Ms Attfield, Dr. Han and Dr. Egbase are employees of VML Health. Supplementary Files REVEALMMManuscriptSupplementaryTable1FINAL03FEB26.xlsx Supplementary Table 1 REVEALMMManuscriptSupplementaryMaterialsFINAL03FEB26.pdf Supplementary materials Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: revise 06 Mar, 2026 Review # 2 received at journal 02 Mar, 2026 Reviewer # 2 agreed at journal 18 Feb, 2026 Review # 1 received at journal 14 Feb, 2026 Reviewer # 1 agreed at journal 11 Feb, 2026 Reviewers invited by journal 11 Feb, 2026 Editor assigned by journal 11 Feb, 2026 Submission checks completed at journal 11 Feb, 2026 First submitted to journal 11 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8849252","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":589863950,"identity":"070bccb5-c3cc-4538-9417-ad35c54a966c","order_by":0,"name":"Faith Davies","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJUlEQVRIie3QsUrDQBjA8e84SJav3poQ1FeIBAIl1r6KoZAsqXR0kobAuahZG3wJx46BG7pE8BFaujh0SMnkpHcNikNqHQXvz+VIAr/7SAB0uj+YOSVTCtAuqOV1dIhgSdLsk5CZfGMcJkBaIqP4K2KK2/UEglN3QUUzmIuQm8/idTs/B+YsXjoJhmk2g/jsSRiRM64kwauoX1QR2A/JpIsMQRIEQYoMfTrmkkDiez0uwK3wsnMKW+3IsMhY0/QVYRtF3vcTq50S5vLbHaKIlXjrHi8lMctuskof0Y1HOTV8+57HHrc2Pin4CO077P5jLF42eB0MOBPr+o0HxzlLvHrLL04YmstOs8v9usvUZljqMLncfeB7N2qjdfvw0xSdTqf7R30AMkFcHKq49HQAAAAASUVORK5CYII=","orcid":"","institution":"NYU Langone","correspondingAuthor":true,"prefix":"","firstName":"Faith","middleName":"","lastName":"Davies","suffix":""},{"id":589863951,"identity":"a2e48ee8-4207-4b0b-ad85-15c5f5837cbe","order_by":1,"name":"Beth Faiman","email":"","orcid":"https://orcid.org/0000-0002-8153-0438","institution":"Cleveland Clinic","correspondingAuthor":false,"prefix":"","firstName":"Beth","middleName":"","lastName":"Faiman","suffix":""},{"id":589863952,"identity":"59107bd7-1323-4b91-b61c-e14918a60448","order_by":2,"name":"Hayley Beer","email":"","orcid":"","institution":"Peter MacCallum Cancer Centre","correspondingAuthor":false,"prefix":"","firstName":"Hayley","middleName":"","lastName":"Beer","suffix":""},{"id":589863953,"identity":"f06ff114-b942-4503-8c06-a56d6e267979","order_by":3,"name":"Anne Quinn Young","email":"","orcid":"","institution":"Multiple Myeloma Research Foundation","correspondingAuthor":false,"prefix":"","firstName":"Anne","middleName":"Quinn","lastName":"Young","suffix":""},{"id":589863954,"identity":"648e1e38-fd0d-419e-89e0-e1188b8ad3a1","order_by":4,"name":"Katie Joyner","email":"","orcid":"","institution":"Myeloma Patients Europe","correspondingAuthor":false,"prefix":"","firstName":"Katie","middleName":"","lastName":"Joyner","suffix":""},{"id":589863955,"identity":"e1095090-f157-404a-9d04-0b7104e6bb05","order_by":5,"name":"Kate Morgan","email":"","orcid":"","institution":"Myeloma Patients Europe","correspondingAuthor":false,"prefix":"","firstName":"Kate","middleName":"","lastName":"Morgan","suffix":""},{"id":589863956,"identity":"fe9a5f9e-7ef0-40c5-8557-d9e4318eba92","order_by":6,"name":"Stephen Huo","email":"","orcid":"","institution":"Janssen Research \u0026 Development, LLC","correspondingAuthor":false,"prefix":"","firstName":"Stephen","middleName":"","lastName":"Huo","suffix":""},{"id":589863957,"identity":"35feda3d-b590-4ad4-9967-c4daf1154b01","order_by":7,"name":"J. Blake Bartlett","email":"","orcid":"","institution":"Janssen Research \u0026 Development, LLC","correspondingAuthor":false,"prefix":"","firstName":"J.","middleName":"Blake","lastName":"Bartlett","suffix":""},{"id":589863960,"identity":"b9854441-1b37-4ab8-a224-8ab654486399","order_by":8,"name":"Yi Han","email":"","orcid":"","institution":"VML Health","correspondingAuthor":false,"prefix":"","firstName":"Yi","middleName":"","lastName":"Han","suffix":""},{"id":589863962,"identity":"00f20c42-7e9f-413f-88fe-522a2876f3c5","order_by":9,"name":"Georgia Attfield","email":"","orcid":"","institution":"VML Health","correspondingAuthor":false,"prefix":"","firstName":"Georgia","middleName":"","lastName":"Attfield","suffix":""},{"id":589863963,"identity":"a5d5df07-a703-46d1-98cc-748e677577f3","order_by":10,"name":"Daniel Egbase","email":"","orcid":"","institution":"VML Health","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"","lastName":"Egbase","suffix":""},{"id":589863965,"identity":"a49f18a5-63d9-42a8-9ddc-56486e42ef6f","order_by":11,"name":"Joseph Mikhael","email":"","orcid":"https://orcid.org/0000-0001-9670-2864","institution":"Translational Genomics Research Institute, Phoenix","correspondingAuthor":false,"prefix":"","firstName":"Joseph","middleName":"","lastName":"Mikhael","suffix":""}],"badges":[],"createdAt":"2026-02-11 08:46:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8849252/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8849252/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102839660,"identity":"a0a8b4e1-6c91-4904-bb88-71f7c128f7f3","added_by":"auto","created_at":"2026-02-17 11:46:52","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":53261,"visible":true,"origin":"","legend":"\u003cp\u003eREVEAL-MM Study Design. PSM = propensity score matching\u003c/p\u003e","description":"","filename":"REVEALMMManuscriptFigure1FINAL03FEB26.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8849252/v1/02f6f71c69abee1ced2fde0d.jpg"},{"id":102839659,"identity":"4cf76f45-8f94-436c-8d05-f0af484be2bf","added_by":"auto","created_at":"2026-02-17 11:46:52","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":42632,"visible":true,"origin":"","legend":"\u003cp\u003eReceiver operating characteristic (ROC) curve performance of the LASSO regression model, evaluated using claims data at T-12M(A), T-9M(B), and T-6M(C). The AUC for each model is shown.\u003c/p\u003e","description":"","filename":"REVEALMMManuscriptFigure2FINAL03FEB26.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8849252/v1/4bc31eebead7a84af254a92b.jpg"},{"id":102839662,"identity":"1939d3c0-bee8-402f-b9a5-38a50268a26c","added_by":"auto","created_at":"2026-02-17 11:46:52","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":42818,"visible":true,"origin":"","legend":"\u003cp\u003eReceiver operating characteristic (ROC) curve performance of the SNN model, evaluated using claims data at T-12M(A), T-9(B), and T-6M(C). The AUC for each model is shown.\u003c/p\u003e","description":"","filename":"REVEALMMManuscriptFigure3FINAL03FEB26.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8849252/v1/054ef3a5e1264317ff30bacd.jpg"},{"id":103056924,"identity":"71bc79a7-3a27-4a67-9d2b-3d383a4dbb2a","added_by":"auto","created_at":"2026-02-20 09:25:55","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1186364,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8849252/v1/0919c648-54eb-47e0-a6c1-72073e943147.pdf"},{"id":103056390,"identity":"d0af11c2-86e0-423b-91cc-9c41bacf61e5","added_by":"auto","created_at":"2026-02-20 09:08:48","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":14801,"visible":true,"origin":"","legend":"Supplementary Table 1","description":"","filename":"REVEALMMManuscriptSupplementaryTable1FINAL03FEB26.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8849252/v1/652e985910f7d44a63769e1e.xlsx"},{"id":102839663,"identity":"3dee2713-4091-4f45-8753-5b302f5f49f4","added_by":"auto","created_at":"2026-02-17 11:46:52","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":385903,"visible":true,"origin":"","legend":"Supplementary materials","description":"","filename":"REVEALMMManuscriptSupplementaryMaterialsFINAL03FEB26.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8849252/v1/1b2bab9cf21206f82d05badd.pdf"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential conflict of interest.\nDr Davies has received consultancy fees from GSK, Sanofi, Bristol Myers Squibb, Regeneron, Johnson \u0026 Johnson, and Takeda. Dr Mikhael has received consultancy fees from Sanofi, Bristol Myers Squibb, Johnson \u0026 Johnson, and Menarini. Ms Joyner and Ms Morgan have received consultancy fees, donated to Myeloma Patients Europe, from Johnson \u0026 Johnson, Pfizer and GSK. Ms Young has received consultancy fees, donated to Multiple Myeloma Research Foundation, from Johnson \u0026 Johnson, AstraZeneca and Bristol Myers Squibb. Dr Faiman and Ms Beer have received consultancy fees from Johnson \u0026 Johnson. Dr Huo and Dr Bartlett are employees of Johnson \u0026 Johnson and hold stock options in the company. Ms Attfield, Dr. Han and Dr. Egbase are employees of VML Health.","formattedTitle":"REVEAL-MM: Retrospective Evaluation of Variables in Early Assessment and Landmark trends in Multiple Myeloma – a US Claims-Based Case-Control Study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMultiple myeloma (MM) poses a significant diagnostic challenge due to its non-specific symptoms such as bone pain, fatigue, and recurrent infections, which frequently overlap with other conditions including osteoporosis, chronic kidney disease, and degenerative spinal disorders.\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e The wide spectrum of non-specific symptoms seen in MM may lead to diagnostic delays or mis-referrals to non-hematology specialists, including, nephrologists, orthopedic surgeons, or rheumatologists; causing further delays.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e Compared with other cancers, MM has one of the highest proportions of patients requiring three or more consultations before specialist referral.\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e Such delays are associated with a poorer disease stage at diagnosis which is subsequently associated with poorer survival.\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e In addition, patients with longer diagnostic intervals are also reported to have reduced disease-free survival, and more treatment-associated secondary complications.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e These delays disproportionately affect racial and ethnic minority populations; African American and Hispanic/Latino patients with MM experience longer intervals from symptom onset to treatment initiation compared with their White counterparts.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e Addressing these disparities is crucial to achieving equitable outcomes for all MM patients and reducing preventable morbidity.\u003c/p\u003e \u003cp\u003ePrecursor conditions such as monoclonal gammopathy of undetermined significance (MGUS) and smoldering multiple myeloma (SMM) represent opportunities for earlier monitoring and intervention, however these states are identified in less than 10% of patients.\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e As a result, most individuals progress to symptomatic MM without prior surveillance, contributing to missed opportunities for timely diagnosis and downstream delays.\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e Although diagnostic delays in MM are well documented, the specific healthcare utilization patterns associated with these delays remain poorly characterized. Existing evidence points to healthcare system-, provider-, and patient-related factors,\u003csup\u003e9\u003c/sup\u003e yet how these factors translate in real-world diagnostic pathways is not well defined.\u003c/p\u003e \u003cp\u003eLarge-scale claims databases provide an opportunity to examine diagnostic patterns at scale and to identify recurring encounter trends that may represent early signals of MM prior to a formal diagnosis. Statistical and machine learning (artificial intelligence) models can be applied to test this hypothesis.\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e When implemented within real-world health systems, such models may be able to flag high-risk individuals months before symptom onset, supporting earlier referral and improved patient outcomes. In addition, timely diagnosis may decrease preventable complications and downstream costs associated with delayed diagnosis.\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e Furthermore, by generating actionable risk signals in real time, these models can support primary care and non-hematology specialists in recognizing patterns suggestive of MM, helping bridge awareness gaps and accelerate timely referral to hematology and diagnosis.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Overview\u003c/h2\u003e \u003cp\u003eREVEAL-MM is a retrospective, 1:1 matched case-control study designed to characterize pre-diagnostic healthcare utilization patterns and to develop predictive models that identify individuals at increased risk of MM prior to their formal diagnosis. This study compares longitudinal diagnostic, procedural, prescription, and physician-visit patterns between patients who later developed MM and matched controls without MM throughout the study period.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eData Source\u003c/h3\u003e\n\u003cp\u003eData were extracted from the Optum\u0026reg; Clinformatics\u0026reg; Data Mart (CDM), a large, geographically diverse, US administrative claims database containing de-identified longitudinal medical and prescription claims from commercially insured and Medicare Advantage beneficiaries. CDM includes demographics, diagnostic, procedure, prescription, and inpatient and outpatient encounters with associated ICD-9/10 codes.\u003c/p\u003e\n\u003ch3\u003eStudy Population\u003c/h3\u003e\n\u003cp\u003eThe pre-MM cohort included adults aged\u0026thinsp;\u0026ge;\u0026thinsp;50 years with a confirmed MM diagnosis between January 1, 2020, and January 31, 2024, defined by at least two MM-coded claims (ICD-10-CM: C90.x; ICD-9-CM: 203.x) occurring\u0026thinsp;\u0026ge;\u0026thinsp;30 days apart. Patients needed at least 24 months of continuous data coverage prior to their first MM diagnosis (the index date) and were excluded if they had any malignant cancer diagnosis or MGUS during the study period.\u003c/p\u003e \u003cp\u003eThe control cohort comprised of adults aged\u0026thinsp;\u0026ge;\u0026thinsp;50 years without an MM diagnosis during the study period, matched 1:1 via propensity score matching on age, gender, race, geographic region, insurance type, Charlson Comorbidity Index (CCI) and National Cancer Institute (NCI) comorbidity scores assessed three months prior to their last medical encounter date (index date). For controls, the index date was assigned from the distribution of MM cohort index dates to ensure temporal alignment. Covariate balance was achieved with standardized mean differences below 0.1.\u003c/p\u003e\n\u003ch3\u003eStudy Design\u003c/h3\u003e\n\u003cp\u003eAll available healthcare encounter data (including diagnostic, procedure, and prescription codes, and physician visits by specialty) were extracted for the 24-month period prior to the index date for both cohorts. Encounter information was summarized at three pre-diagnostic timepoints: T-12 months (12 months before index date), T-9 months (9 months before index date), and T-6 months (6 months before index date). The study design is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eData Preprocessing\u003c/h3\u003e\n\u003cp\u003eAt each timepoint, both frequency variables and binary indicators were generated for all encounter categories. The full analysis set (FAS) had 1,287, 1,451 and 1,619 independent variables for T-12, T-9, and T-6 periods, respectively. Diagnosis codes were encoded as binary indicators reflecting the presence or absence of each diagnosis. Procedure and prescription codes were represented as frequency counts within the corresponding time window. No missing values were observed, as the absence of a code was interpreted as zero.\u003c/p\u003e \u003cp\u003eAll variables were first screened using Bonferroni-adjusted univariate non-parametric tests against the outcome variable (Pearson\u0026rsquo;s chi-square test for diagnosis codes and Wilcoxon rank-sum tests for all other numerical variables). Variables meeting statistical significance thresholds were subsequently reviewed by clinical experts. Claims codes were not presented verbatim; instead, they were translated into expanded clinical descriptors to enhance interpretability. To generate clinically meaningful features, subject-matter experts grouped diagnostic, procedure, and prescription codes into predefined clinical categories (e.g., \u0026ldquo;Anemias,\u0026rdquo; \u0026ldquo;Renal complications,\u0026rdquo; \u0026ldquo;Musculoskeletal pain\u0026rdquo;). Physician visits were categorized by specialty. The grouping framework is provided in \u003cb\u003eSupplementary Table\u0026nbsp;1\u003c/b\u003e. Following clinical review and grouping, the final FAS used for model development for T-12, T-9, T-6 had 43, 52, and 66 independent variables respectively.\u003c/p\u003e \u003cp\u003eFour complementary modelling strategies were implemented to balance interpretability, variable selection, and predictive accuracy. LASSO provided an interpretable model; Random Forest and Deep Learning captured non-linear relationships; and PLS-DA addressed multicollinearity.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eModel 1: Least Absolute Shrinkage and Selection Operator (LASSO) Regression (Penalized linear modelling)\u003c/h2\u003e \u003cp\u003eLASSO regression was employed for sparse feature selection and risk prediction. Performance was evaluated on the hold-out test set using Area Under the receiver operating characteristic Curve (AUC). Non-zero coefficients were interpreted as clinically actionable risk factors.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eModel 2: Random Forest (Ensemble Tree-Based Modelling)\u003c/h3\u003e\n\u003cp\u003eRandom Forest was implemented to capture non-linear relationships and interactions among predictors. Final model performance was assessed on the test set using AUC. Variable importance was computed via mean decrease in impurity.\u003c/p\u003e\n\u003ch3\u003eModel 3: Partial Least Squares Discriminant Analysis (PLS-DA)\u003c/h3\u003e\n\u003cp\u003ePLS-DA was applied to address potential multicollinearity in the data (e.g., correlated lab panels). This supervised dimension-reduction technique maximizes covariance between predictors and the outcome. The number of latent components was determined by 10-fold cross-validation. Model discrimination was evaluated using AUC on the test set and loading plot visualized relationships between original variables and latent components.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eModel 4: Deep Learning with Sequential Neural Network\u003c/h2\u003e \u003cp\u003eA feed-forward sequential neural network was constructed using Keras/TensorFlow to model the complex, non-linear patterns inaccessible to traditional methods. 10-fold cross-validation was used to optimize the model structure. All model fittings were conducted in R v4.3 (LASSO, RF, PLS-DA and DL via glmnet, random Forest, mixOmics and Kera3 packages).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eEthical Considerations\u003c/h2\u003e \u003cp\u003eThis study utilized a commercially available, de-identified database, compliant with the Health Insurance Portability and Accountability Act (HIPAA). As the research did not involve human subjects or access to personally identifiable information, this study was exempt from Institutional Review Board (IRB) oversight and did not require Ethics approval.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eBaseline characteristics\u003c/h2\u003e \u003cp\u003eA total of 4,733 patients met the inclusion criteria for the pre-MM cohort and were matched (1:1) with 4,733 controls, resulting in a total study population of 9,466 patients. Baseline demographic and clinical characteristics of the matched cohorts are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The mean age of the population was 74.1 years (SD 8.14), and 50% were female. Racial distribution (51% White) and geographic region (45% residing in the US South Census Region) were comparable between cohorts. Medicare Advantage was the primary insurer for 86% of patients. Comorbidity burden was well balanced following propensity score matching, with similar mean CCI scores (2.29 for pre-MM vs. 2.18 for controls) and NCI comorbidity scores (0.71 vs. 0.67, respectively). All standardized mean differences for matching variables were \u0026lt;\u0026thinsp;0.1, confirming excellent covariate balance.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBaseline characteristics of the study population. CCI\u0026thinsp;=\u0026thinsp;Charlson Comorbidity Index; NCI\u0026thinsp;=\u0026thinsp;National Cancer Institute Comorbidity Index. COM\u0026thinsp;=\u0026thinsp;Commercial health insurance. MCR\u0026thinsp;=\u0026thinsp;Medicare health insurance. SD\u0026thinsp;=\u0026thinsp;Standard deviation.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eOverall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePre-MM cohort\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNon-MM cohort\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003en\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9466\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e4733\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4733\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eMean (SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eMean (SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eMean (SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAge\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73.96 (8.25)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e74.10 (8.14)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e73.81 (8.36)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eCCI\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2.24 (2.33)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2.29 (2.29)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2.18 (2.36)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eNCI*\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.69 (0.73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.71 (0.72)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.67 (0.73)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eCount (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eCount (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eCount (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e\u003cb\u003eGender\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eFemale\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,717 (50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2,360 (50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2,357 (50)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMale\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,744 (50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2,370 (50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2,374 (50)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eUndisclosed\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 (\u0026lt;\u0026thinsp;0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3 (\u0026lt;\u0026thinsp;0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2 (\u0026lt;\u0026thinsp;0.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e \u003cp\u003e\u003cb\u003eRace\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eWhite\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,847 (51)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2,419 (51)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2,428 (51)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eAsian\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e205 (2.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e104 (2.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e101 (2.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eBlack\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,579 (17)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e786 (17)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e793 (17)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eUndisclosed\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e298 (3.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e149 (3.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e149 (3.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eOther\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2,537 (27)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1,275 (27)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1,262 (27)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e \u003cp\u003e\u003cb\u003eRegion\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMidwest\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,858 (20)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e938 (20)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e920 (19)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eNortheast\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,360 (14)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e673 (14)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e687 (15)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eSouth\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,278 (45)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2,149 (45)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2,129 (45)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eWest\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,956 (21)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e966 (20)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e990 (21)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eOther\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14 (0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7 (0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7 (0.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eInsurance Type\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCOM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,290 (14)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e624 (13)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e666 (14)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMCR\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8,176 (86)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e4,109 (87)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4,067 (86)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eLASSO Regression\u003c/h2\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003eModel Performance\u003c/h2\u003e \u003cp\u003eLASSO regression was used to interpret which variables most strongly linked with an MM diagnosis. The model demonstrated modest and progressively improving discrimination as the index date approached. AUC increased from 0.6727 at T-12M to 0.6928 at T-9M and 0.6966 at T-6M (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eClassification accuracy\u003c/h2\u003e \u003cp\u003eAt T-12M, overall accuracy was 62.0%, with the model correctly identifying 2,127 pre-MM cases (sensitivity 44.9%) and 3,743 controls (specificity 79.1%). Accuracy improved to 64.0% at T-9M and 64.3% at T-6M, driven by increased sensitivity (44.9% to 59.3%) and reduced specificity (79.1% to 69.2%) (\u003cb\u003eSupplementary Table\u0026nbsp;2\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eVariable interpretation\u003c/h2\u003e \u003cp\u003eLASSO coefficients provided insight into the relative importance and direction of associations (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). At T-12M, variables associated with increased MM risk included diagnostic codes for anemias and neutropenia, and procedure codes related to M-protein and immunoglobulin evaluation. Additional high-risk indicators included gastroesophageal reflux disease (GERD) with esophagitis, cardiovascular disorders, and vaccination- and infection-related encounters (e.g., influenza and pneumococcal vaccination, vaccination administration, other infection concerns).\u003c/p\u003e \u003cp\u003eVariables consistently associated with lower MM risk included dementia, depression (unspecified), altered mental status, antiviral prescriptions (e.g., nirmatrelvir/ritonavir), and visits to non-physician providers.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eLASSO coefficients for codes across T-12M associated with higher or lower MM diagnosis risk. Positive coefficients are labelled in green. Negative coefficients are labelled in red. DX\u0026thinsp;=\u0026thinsp;diagnosis code; HCP\u0026thinsp;=\u0026thinsp;Healthcare professional visit; PROC\u0026thinsp;=\u0026thinsp;procedure code; GNM\u0026thinsp;=\u0026thinsp;Prescription code.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClaim code\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVariable description\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCoefficient\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGNM_G2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFlu vaccination\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.360882741\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAnemias\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.169056209\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eM-protein / immunoglobulin quantification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.078367363\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCough\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.071624292\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRoutine follow-up office visit\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.063930757\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHCP_00030\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHaematology \u0026amp; oncology visit\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.062831215\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOther infection concern\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.050016439\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNeutropenia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.049724713\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGNM_G3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePneumococcal vaccination\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.043819732\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDX_K210\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGERD with esophagitis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.041424254\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eObservation/short-term hospital care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.039689921\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCardiovascular disorders\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.035108252\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVaccination administration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.032322284\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCOVID-19 diagnostic testing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.029408403\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_92012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eInterim ophthalmic exam, established pt\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.026418157\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_88342\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eImmunochem/cytochem first antibody\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.017946621\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_88313\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSpecial stains (group 2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.017527554\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCOVID or viral exposure\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.015766174\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDXG_6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMusculoskeletal pain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.011627152\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDX_D61818\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOther pancytopenia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.007789679\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_83883\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNephelometry assay (not specified)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.006681015\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMetabolic panel\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.00425867\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNursing facility/long-term care visit\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.0339063\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGNM_00798\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNirmatrelvir/ritonavir\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.044641046\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHCP_00077\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRegistered nurse practitioner visit\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.048882152\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDX_R4182\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAltered mental status (unspecified)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.067758286\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_G15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRoutine monitoring\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.068581005\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDX_F32A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDepression (unspecified)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.073602922\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePROC_1159F\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMedication list documented\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.076745669\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHCP_00055\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOther non-physician provider visit\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.08223409\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDementia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDementia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-0.113692015\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eTemporal Trends\u003c/h2\u003e \u003cp\u003eLongitudinal patterns from T-12M to T-6M demonstrated strengthening associations for several high-risk variables. For example, the coefficient for M-protein/immunoglobulin testing increased from 0.0784 (T-12M) to 0.0995 (T-9M) and 0.1599 (T-6M), while neutropenia rose from 0.0497 to 0.0797 and 0.0965, respectively. Some early signals, such as cough, were not retained at later timepoints. Variables associated with lower MM risk, particularly dementia, depression, and altered mental status, remained consistently negative across all periods (\u003cb\u003eSupplementary Tables\u0026nbsp;3 and 4\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eRandom Forest\u003c/h2\u003e \u003cp\u003eA Random Forest model combines many small decision rules from patient history, giving an indication of which variables most often point toward early MM. The model showed modest discriminatory performance, with AUC values increasing slightly closer to diagnosis: 0.6316 at T-12M, 0.6544 at T-9M, and 0.6537 at T-6M (\u003cb\u003eSupplementary Fig.\u0026nbsp;1\u003c/b\u003e). Classification accuracy ranged from 62.4% at T-12M to 63.7% at T-6M (\u003cb\u003eSupplementary Table\u0026nbsp;5\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eVariable importance rankings were consistent across time points and consistent with the LASSO findings. Hematologic codes, infection- and vaccination-related codes, and routine care encounters were among the top contributors (\u003cb\u003eSupplementary Tables\u0026nbsp;6, 7, and 8\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003ePartial Least Squares Discriminant Analysis (PLS-DA)\u003c/h2\u003e \u003cp\u003ePLS-DA groups related clinical codes into clusters to identify whether patients who later develop MM have distinct healthcare utilization patterns. PLS-DA AUC values improved as the index date approached: 0.6758 at T-12M, 0.6947 at T-9M, and 0.692 at T-6M (\u003cb\u003eSupplementary Fig.\u0026nbsp;2\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eLoading plots demonstrated consistent variable clustering patterns across all time periods (\u003cb\u003eSupplementary Fig.\u0026nbsp;3\u003c/b\u003e). Hematologic evaluation codes (e.g., anemias, neutropenia, bone marrow, cytogenetic testing, peripheral blood analysis) formed stable, contiguous clusters. Routine care codes (e.g., established patient visits, routine monitoring) also appeared in adjacent regions.\u003c/p\u003e \u003cp\u003eComorbidity-related codes, including dementia, depression (unspecified), and altered mental status, clustered separately and consistently across all timepoints.\u003c/p\u003e \u003cp\u003eVaccine- and infection-related codes behaved less uniformly. For example, influenza vaccine codes were outliers at T-12M and T-9M, while COVID-19 vaccine codes showed similar behavior at T-6M.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eSequential Neural Network (SNN)\u003c/h2\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eModel Performance\u003c/h2\u003e \u003cp\u003eA Sequential Neural Network uses artificial intelligence to learn complex, changing patterns in patients\u0026rsquo; medical histories, allowing it to detect early signs of MM that simpler models may miss. The model demonstrated the strongest overall performance among all models, with AUC improving steadily as diagnosis approached: 0.7654 at T-12M, 0.8001 at T-9M, and 0.8257 at T-6M (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eClassification Accuracy\u003c/h2\u003e \u003cp\u003eConsistent with AUC results, classification accuracy improved at each timepoint. At T-12M, the model achieved 68.3% accuracy, correctly identifying 2,943 pre-MM cases and 3,541 controls. Accuracy increased to 70.7% at T-9M and 72.8% at T-6M, with corresponding increases in correct MM classifications (3,581 and 3,734, respectively) (\u003cb\u003eSupplementary Table\u0026nbsp;9\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this large, US claims-based matched case-control study, clinical signals of MM were identified as early as 12 months before MM diagnosis using routinely collected administrative claims data. Across all four models used to analyze claims data, discrimination improved as the index date approached, reflecting the clear clinical manifestations of pre-diagnostic MM over time.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e The SNN model using artificial intelligence, achieved the strongest performance, with an AUC of 0.8257 at T-6M, outperforming traditional linear and tree-based models. These findings highlight the potential of deep learning techniques to capture complex, nonlinear patterns in high-dimensional healthcare data.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e The SNN model\u0026rsquo;s strong prediction ability shows it could be used at the population level for screening, helping to identify high-risk individuals earlier, before complications develop.\u003c/p\u003e \u003cp\u003eAcross all modelling approaches, consistent hematological abnormality codes were observed, particularly anemia, which emerged among the strongest predictors of a MM diagnosis. These findings align with known early manifestations of MM.\u003csup\u003e13,14\u003c/sup\u003e Similarly, claims related to diagnostic evaluation for monoclonal gammopathy (e.g., M-protein or immunoglobulin testing) showed progressively stronger associations closer to diagnosis, suggesting that clinicians may be responding to subtle or unexplained laboratory abnormalities well before MM is formally confirmed. Unexpected claims codes associated with gastro-esophageal reflux disease (GERD) and cardiovascular disorders also emerged, warranting further investigation as potential early indicators or markers of diagnostic complexity.\u003c/p\u003e \u003cp\u003eVaccination- (including COVID-19, influenza, and pneumococcal vaccines) and infection-related codes were also frequently selected by LASSO and featured prominently in importance measures for Random Forest and PLS-DA models. In the PLS-DA loading plots, however, these codes showed less consistent clustering across timepoints, with influenza vaccine codes appearing as outliers at T-12M and T-9M and COVID-19 vaccine codes showing similar outlier behavior at T-6M. This may reflect increased healthcare encounters prompted by nonspecific symptoms or immune dysregulation that precede MM diagnosis.\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e Alternatively, these patterns may capture confounding healthcare utilization patterns, as patients with underlying hematologic or constitutional symptoms may seek care more frequently.\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e Moreover, the observed associations may partly reflect temporary effects, such as seasonal vaccination patterns.\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e Further investigation is needed to understand these patterns.\u003c/p\u003e \u003cp\u003eInterestingly, several comorbidity-related codes, including dementia, depression, and altered mental status, were consistently associated with a reduced likelihood of developing MM. These inverse associations may represent differences in healthcare utilization, competing management of existing comorbidities, or under-evaluation of hematologic abnormalities/ concerns in these patients. This phenomenon has been described in other cancer detection studies, where lower diagnostic intensity is applied to medically complex populations.\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThe relative performance of the models provides important insights for clinical translation. While LASSO and Random Forest demonstrated modest discriminatory ability (AUC\u0026thinsp;~\u0026thinsp;0.63\u0026ndash;0.70), they provided interpretable variable importance patterns that can inform clinical hypotheses. In contrast, the SNN, though less interpretable, achieved substantially higher discrimination, suggesting that deep neural architectures may be better suited to capturing the nonlinear, evolving patterns characteristic of MM\u0026rsquo;s pre-diagnostic phase.\u003c/p\u003e \u003cp\u003eWhile these results are promising, several limitations warrant consideration. First, models were trained on administrative claims data, which lack granular laboratory values (e.g., specific hemoglobin and M-protein levels) and may miss subtle clinical cues detectable in structured electronic healthcare records data or clinician notes. Second, although propensity score matching effectively balanced patient characteristics, residual confounding and misclassification inherent in claims-based data analysis remain possible. Third, performance metrics of models, though encouraging, may not be sufficient for unsupervised population-wide screening. Instead, these models may be more appropriate as decision-support tools to prompt targeted evaluation in high-risk subgroups. Finally, external validation in broader populations, particularly in integrated health systems with richer clinical data across geographies is needed to assess generalizability and real-world applicability.\u003c/p\u003e \u003cp\u003eDespite these limitations, this study demonstrates the potential of administrative claims data to support earlier identification of individuals at risk for MM, particularly at the health system level. The importance of earlier identification has never been stronger, with emerging clinical evidence suggesting that earlier treatment, particularly among patients with high-risk precursor conditions, may delay or prevent progression to overt MM and its associated complications.\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e While claims data are not universally available at the point of care, they can be analyzed centrally within health systems to identify patterns of healthcare use that signal increased risk. Integrating these system-level insights with laboratory results and electronic health record data offers a practical path to improve risk identification at the point of care. Together, this approach supports a shift toward earlier evaluation, timely referral, and more proactive care, with the potential to improve outcomes and reduce avoidable complications for patients with MM.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eOverall, our findings support the feasibility of using real-world administrative claims data to detect patterns indicative of emerging MM up to a year before diagnosis. Hematologic abnormalities, diagnostic evaluation patterns, and infection-related encounters emerged as consistent early signals. Among the tested models, deep learning SNN demonstrated the strongest predictive ability, emphasizing its potential role in future risk-stratification initiatives within population-level health systems. Further research is required to explore some of the trends highlighted in this paper, with the potential to reduce diagnostic delays and improve clinical outcomes.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCompeting Interests statement\u003c/h2\u003e \u003cp\u003eDr Davies has received consultancy fees from GSK, Sanofi, Bristol Myers Squibb, Regeneron, Johnson \u0026amp; Johnson, and Takeda. Dr Mikhael has received consultancy fees from Sanofi, Bristol Myers Squibb, Johnson \u0026amp; Johnson, and Menarini. Ms Joyner and Ms Morgan have received consultancy fees, donated to Myeloma Patients Europe, from Johnson \u0026amp; Johnson, Pfizer and GSK. Ms Young has received consultancy fees, donated to Multiple Myeloma Research Foundation, from Johnson \u0026amp; Johnson, AstraZeneca and Bristol Myers Squibb. Dr Faiman and Ms Beer have received consultancy fees from Johnson \u0026amp; Johnson. Dr Huo and Dr Bartlett are employees of Johnson \u0026amp; Johnson and hold stock options in the company. Ms Attfield, Dr. Han and Dr. Egbase are employees of VML Health.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eAuthor Contributions\u003c/h2\u003e \u003cp\u003eSH, JBB and YH were responsible for designing the study. YH was responsible for data extraction from the from the Optum\u0026reg; CDM and analyzing data. FED, BF, HB, AQY, KJ, KM, SH, JBB, YH, GA, DE, JM were responsible for interpreting the results. GA and DE were responsible for writing the report. FED, BF, HB, AQY, KJ, KM, SH, JBB, YH and JM provided feedback on the report.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eThis study was funded by J\u0026amp;J Innovative Medicine, LLC. Additional writing support was provided by VML Health.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eKoshiaris C, Oke J, Abel L, Nicholson BD, Ramasamy K, Van den Bruel A. Quantifying intervals to diagnosis in myeloma: a systematic review and meta-analysis. BMJ open. 2018;8(6):e019758.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajkumar SV. Multiple myeloma: 2022 update on diagnosis, risk stratification, and management. American journal of hematology. 2022;97(8):1086\u0026ndash;1107.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFriese CR, Abel GA, Magazu LS, Neville BA, Richardson LC, Earle CC. Diagnostic delay and complications for older adults with multiple myeloma. Leukemia \u0026amp; lymphoma. 2009;50(3):392\u0026ndash;400.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKariyawasan C, Hughes D, Jayatillake M, Mehta A. Multiple myeloma: causes and consequences of delay in diagnosis. QJM: An International Journal of Medicine. 2007;100(10):635\u0026ndash;640.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLyratzopoulos G, Neal RD, Barbiere JM, Rubin GP, Abel GA. Variation in number of general practitioner consultations before hospital referral for cancer: findings from the 2010 National Cancer Patient Experience Survey in England. The lancet oncology. 2012;13(4):353\u0026ndash;365.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBhutani M, Blue BJ, Cole C, Badros AZ, Usmani SZ, Nooka AK, et al. Addressing the disparities: the approach to the African American patient with multiple myeloma. Blood Cancer J. 2023;13(1):189. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41408-023-00961-0\u003c/span\u003e\u003cspan address=\"10.1038/s41408-023-00961-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 38110338; PMCID: PMC10728116.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith L, Carmichael J, Cook G, Shinkins B, Neal RD. Diagnosing myeloma in general practice: how might earlier diagnosis be achieved? Br J Gen Pract. 2022;72(723):462\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWong A, Jimenez-Zepeda VH, Rankin K, Sandhu I, Chu M, Delluc A, et al. Incidence and prevalence of clinically detected smoldering multiple myeloma within the general population: a retrospective observational cohort study. Blood Cancer J. 2025;15(1):149.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGreene JA, Lea AS. Digital Futures Past - The Long Arc of Big Data in Medicine. N Engl J Med. 2019;381(5):480\u0026ndash;485.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMittelman M, Shantsila A, Sagy I, et al. Prediction of multiple myeloma development using a machine learning model. Br J Haematol. 2024. (This is the key reference paper, justifying the use of ML models for MM risk prediction.)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347\u0026ndash;1358.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePorteous A, Gibson S, Eddowes LA, Drayson M, Pratt G, Bowcock S, et al. An Economic Model to Establish the Costs Associated With Routes to Presentation for Patients With Multiple Myeloma in the United Kingdom. Value Health Reg Issues. 2023;35:27\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeesaghur A, Petruski-Ivleva N, Banks VL, et al. Clinical features and diagnosis of multiple myeloma: a population-based cohort study in primary care. BMJ Open. 2021;11(10):e052759.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajkumar SV, Dimopoulos MA, Palumbo A, Blade J, Merlini G, Mateos MV, et al. International Myeloma Working Group updated criteria for the diagnosis of multiple myeloma. Lancet Oncol. 2014;15(12):e538-48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiotto R, Wang F, Wang S, Jiang X, Dudley JT. Deep learning for healthcare: review, opportunities and challenges. Brief Bioinform. 2018;19(6):1236\u0026ndash;1246.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVirgilsen LF, Vedsted P, Jensen H, Frederiksen H, El-Galaly TC, Rasmussen LA. Diagnostic Window Prior to a Haematological Cancer Diagnosis and the Association With Patient Pathways: A Nationwide Register-Based Cohort Study on Healthcare Utilization in Denmark. Eur J Haematol. 2025;114(2):353\u0026ndash;364.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePadhi A, Bhatt P, Chauhan J, Rajyaguru B, Agarwal S, Chaudhary A, et al. Seasonality of Influenza and Optimizing Timing of Vaccination: Systematic Review and Meta-Analysis. Cureus. 2025;17(9):e93607.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRenzi C, Kaushal A, Emery J, Hamilton W, Neal RD, Rachet B, Rubin G et al. Comorbid chronic diseases and cancer diagnosis: disease-specific effects and underlying mechanisms. Nat Rev Clin Oncol. 2019;16(12):746\u0026ndash;761\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSelie AP, van der Willik KD, Ikram MA, Labrecque JA, Schagen SB. Dementia and Cancer: Unravelling Methodological Biases in a Population-Based Cohort. Neuroepidemiology. 2025 Oct 7:1\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDimopoulos MA, Voorhees PM, Schjesvold F, Cohen YC, Hungria V, Sandhu I, et al. Daratumumab or Active Monitoring for High-Risk Smoldering Multiple Myeloma. N Engl J Med. 2025;392(18):1777\u0026ndash;1788.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"blood-cancer-journal","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"bcj","sideBox":"Learn more about [Blood Cancer Journal](http://www.nature.com/bcj/)","snPcode":"41408","submissionUrl":"https://mts-bcj.nature.com/cgi-bin/main.plex","title":"Blood Cancer Journal","twitterHandle":"@bloodcancerjnl","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Multiple myeloma, Diagnostic delay, Healthcare utilization patterns, Predictive modelling, Claims data","lastPublishedDoi":"10.21203/rs.3.rs-8849252/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8849252/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMultiple myeloma (MM) frequently presents with non-specific symptoms that overlap other common conditions, often leading to diagnostic delays and poorer clinical outcomes. Although diagnostic delays in MM are well recognized, the healthcare utilization patterns that precede MM diagnosis are not well defined.\u003c/p\u003e \u003cp\u003eThis retrospective, 1:1 matched case-control study used US administrative claims data from 9,466 patients to identify early signals of MM and evaluate whether such data could be used to predict individuals at risk of MM prior to their formal diagnosis. All diagnostic, procedural, prescription, and physician-visit claims were evaluated at 12, 9, and 6 months within the two years before diagnosis.\u003c/p\u003e \u003cp\u003eInterpretable predictive models (LASSO and Random Forest) identified distinct encounter patterns associated with MM as early as 12-months before diagnosis. Predictive performance across all models increased as diagnosis approached, with the machine learning model reaching a peak area under the curve (AUC) of 0.826 at 6 months.\u003c/p\u003e \u003cp\u003eClaims consistent with typical MM manifestations, including anemia-, musculoskeletal-, and M-protein\u0026ndash;related testing, were more common prior to MM diagnosis. These findings suggest that routinely collected claims data could support earlier identification and evaluation of individuals at risk of MM, enabling more timely diagnosis and improved outcomes.\u003c/p\u003e","manuscriptTitle":"REVEAL-MM: Retrospective Evaluation of Variables in Early Assessment and Landmark trends in Multiple Myeloma – a US Claims-Based Case-Control Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-17 11:46:47","doi":"10.21203/rs.3.rs-8849252/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"revise","date":"2026-03-06T11:58:04+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"This content is not available.","date":"2026-03-02T23:28:24+00:00","index":2,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2026-02-18T16:27:32+00:00","index":2,"fulltext":"This content is not available."},{"type":"editorInvitedReview","content":"This content is not available.","date":"2026-02-14T15:33:17+00:00","index":1,"fulltext":"This content is not available."},{"type":"reviewerAgreed","content":"This content is not available.","date":"2026-02-11T23:59:27+00:00","index":1,"fulltext":"This content is not available."},{"type":"reviewersInvited","content":"","date":"2026-02-11T21:19:50+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-11T15:41:14+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-11T15:39:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"Blood Cancer Journal","date":"2026-02-11T08:41:23+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"blood-cancer-journal","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"bcj","sideBox":"Learn more about [Blood Cancer Journal](http://www.nature.com/bcj/)","snPcode":"41408","submissionUrl":"https://mts-bcj.nature.com/cgi-bin/main.plex","title":"Blood Cancer Journal","twitterHandle":"@bloodcancerjnl","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8f952bdb-dbe8-4417-845b-42452839f6c3","owner":[],"postedDate":"February 17th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":62768179,"name":"Health sciences/Risk factors"},{"id":62768180,"name":"Biological sciences/Cancer/Haematological cancer/Myeloma"}],"tags":[],"updatedAt":"2026-03-06T12:01:37+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-17 11:46:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8849252","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8849252","identity":"rs-8849252","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0