Machine learning approach identifies miRNA signatures for breast cancer detection and classification from patient urine samples

preprint OA: closed
Full text JSON View at publisher
⚙ AI-generated deep summary by qwen3.7-flash, 2026-09-07 · read from full text ⓘ

This study utilized miRNA sequencing on urine samples from 32 breast cancer patients and 50 healthy controls to identify non-invasive biomarkers for disease detection. Random Forest analysis revealed a specific signature of 275 miRNAs capable of distinguishing invasive breast cancer cases from healthy individuals, as well as differentiating between intrinsic subtypes such as luminal A, luminal B, HER2-enriched, and triple-negative breast cancers. The authors note that while this approach offers a painless alternative to mammography, larger multicenter studies are required to validate clinical utility and establish sensitivity and specificity. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Introduction Breast cancer is the most common cancer in women, with one in eight women suffering from this disease in her lifetime. The implementation of centrally organized mammography screening for women between 50 and 69 years of age was a major step in the direction of early detection and lead to a significant improvement in cure rates. Within the screening program, women undergo a mammogram every two years with an implemented centralized quality-controlled review process. However, the participation rate reached only approximately 50% of the eligible women. In addition to several others, the technical aspects of mammography, including painful compression of the breast, are cited as a reason for not participating in this very important program. Therefore, focusing current research on less painful and less invasive techniques for the detection of breast cancer seems to be highly clinically useful. Liquid biopsies offer this option with distinct molecules or cells in line with the research. Blood-based tests of circulating tumor cells (CTCs), cell-free DNA (ctDNA) and cell-free miRNA have been performed in a variety of studies and tumor entities. Methods We performed miRNA sequencing on 82 urine samples, 32 samples from breast cancer patients (9× luminal A, 8× luminal B, 9× triple-negative and 6× HER2) and 50 healthy control samples. Data were analyzed and interpreted using Random Forest analysis. Results We identified a signature of 275 miRNAs that allows the detection of invasive breast cancer in urine from breast cancer patients. Furthermore, we identified distinct miRNA expression patterns for the major intrinsic subtypes of breast cancer, specifically luminal A, luminal B, HER2-enriched and triple-negative breast cancer. Conclusions Here, we present the first approach for sequencing miRNAs in female urine to detect breast cancer and, subsequently, intrinsic subtype-specific miRNA patterns. This experimental approach specifically validates miRNA sequencing as a technique for breast cancer detection in urine samples and opens the door to a new, easy and painless procedure for regular breast cancer screening.
Full text 98,084 characters · extracted from preprint-html · click to expand
Machine learning approach identifies miRNA signatures for breast cancer detection and classification from patient urine samples | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine learning approach identifies miRNA signatures for breast cancer detection and classification from patient urine samples Jochen Maurer, Matthias Rübner, Chao-Chung Kuo, Birgit Klein, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3993094/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Introduction Breast cancer is the most common cancer in women, with one in eight women suffering from this disease in her lifetime. The implementation of centrally organized mammography screening for women between 50 and 69 years of age was a major step in the direction of early detection and lead to a significant improvement in cure rates. Within the screening program, women undergo a mammogram every two years with an implemented centralized quality-controlled review process. However, the participation rate reached only approximately 50% of the eligible women. In addition to several others, the technical aspects of mammography, including painful compression of the breast, are cited as a reason for not participating in this very important program. Therefore, focusing current research on less painful and less invasive techniques for the detection of breast cancer seems to be highly clinically useful. Liquid biopsies offer this option with distinct molecules or cells in line with the research. Blood-based tests of circulating tumor cells (CTCs), cell-free DNA (ctDNA) and cell-free miRNA have been performed in a variety of studies and tumor entities. Methods We performed miRNA sequencing on 82 urine samples, 32 samples from breast cancer patients (9× luminal A, 8× luminal B, 9× triple-negative and 6× HER2) and 50 healthy control samples. Data were analyzed and interpreted using Random Forest analysis. Results We identified a signature of 275 miRNAs that allows the detection of invasive breast cancer in urine from breast cancer patients. Furthermore, we identified distinct miRNA expression patterns for the major intrinsic subtypes of breast cancer, specifically luminal A, luminal B, HER2-enriched and triple-negative breast cancer. Conclusions Here, we present the first approach for sequencing miRNAs in female urine to detect breast cancer and, subsequently, intrinsic subtype-specific miRNA patterns. This experimental approach specifically validates miRNA sequencing as a technique for breast cancer detection in urine samples and opens the door to a new, easy and painless procedure for regular breast cancer screening. Breast cancer miRNA sequencing urine luminal A luminal B HER2 TNBC screening patient classification Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Background Breast cancer (BC) is still the leading cause of cancer-related deaths among the female population, affecting 1 in 8 women in her lifetime ( 1 ). Moreover, this disease still places an enormous burden on healthcare systems worldwide. In addition to tumor biology, which guides treatment options and therapeutic needs in general, tumor stage is still an important risk factor for treatment decisions. The implementation of mammography-based early detection programs consecutively led to a decreased tumor size (T stage) and a decreased rate of axillary lymph node involvement (N stage), thereby improving survival rates in patients with breast cancer ( 2 ). Almost 100% of German women aged 50 to 69 years are invited to undergo mammography screening every 2 years. However, since its implementation in 2009, the participation rate has stagnated at approximately 50% ( 3 ). Therefore, it appears reasonable that these programs could dramatically profit from new early detection methodologies, which are less invasive than the current standards. MicroRNAs (miRNAs) are small noncoding RNA molecules that regulate gene expression by binding to specific target mRNAs, affecting their translation. They are found in many different cell types and tissues and have been shown to play a role in a wide range of biological processes, including cell growth and proliferation, differentiation, and apoptosis. Recent studies have shown that miRNAs can also be found in body fluids, such as blood and urine ( 4 ), where they can be easily isolated and quantified and can act as biomarkers for various diseases, including cancer ( 5 , 6 ). As biomarkers, miRNAs possess high potential for cancer detection because they are stable in body fluids, which allows easy collection and storage of samples ( 7 ). They are also present and detectable in very small samples, which makes them ideal biomarkers ( 8 ). Additionally, miRNA-based diagnostic tests can be noninvasive, which is especially important for the early detection of cancer. Several studies have investigated the use of miRNAs as biomarkers for breast cancer and other gynecological malignancies. In breast cancer, several miRNAs, including miR-21, miR-155, and miR-205, have been identified as biomarkers for this disease( 9 – 11 ). Studies have shown that the levels of these miRNAs are elevated in the blood of breast cancer patients compared to healthy individuals. Furthermore, an earlier study using qPCR analysis identified miR-424 and miR-423 as BC markers in urine as source materials( 12 , 13 ). Researchers have also shown that miRNAs can be used to distinguish between different types of breast cancer, such as invasive ductal carcinoma and invasive lobular carcinoma. It is worth noting that while these studies have been promising, additional research is needed to fully establish the clinical utility of miRNA-based diagnostic tests for breast and ovarian cancer. There is still much to be learned about the potential of miRNA-based diagnostic tests. However, additional research is needed to fully understand the biology of miRNAs and how they are regulated in cancer cells. Additionally, additional studies are needed to validate the use of miRNA-based diagnostic tests in clinical settings, including larger, multicenter studies that will help to establish the specificity and sensitivity of these tests. To investigate the feasibility of miRNA detection as a diagnostic tool for breast cancer, we investigated a cohort of 82 urine samples from BC patients and healthy individuals using miRNA sequencing for the first time. The goal was to identify all currently known miRNAs regulated in BC in a completely unbiased setting without prior selection of BC-specific miRNAs from the literature. Here, we present the first step toward the clinical use of miRNA sequencing in BC detection and provide insight into the stratification of BC patients utilizing this information. Methods Patient Selection and Urine Sample Preparation Samples were collected at the University Hospital Aachen (ethics vote 206/09) and at the University Hospital Erlangen. Participants in Erlangen were recruited within the iMODE-B study (Imaging and Molecular Detection of Breast Cancer; ethical approval by the ethics committee of the Friedrich-Alexander-Universität Erlangen-Nuremberg; #325_19 B). Patients were eligible for inclusion if they had an indication for a diagnostic biopsy due to a suspicious breast lesion. The main aim of the iMODE-B study was to identify molecular markers that are predictive of patient prognosis and treatment response at the time of the first diagnosis of breast cancer. After the participants provided written informed consent in accordance with the Declaration of Helsinki, biospecimen sampling was performed (blood draw and urine). A total of 355 urine samples were collected between November 2019 and July 2020. The urine samples were centrifuged at 944 × g for 10 minutes at room temperature to separate the cell pellet and cell-free supernatant, which were stored separately at -80°C until further use. For miRNA extraction, at least 7 ml of cell-free supernatant was available, 4 ml was used for miRNA extraction, and the remaining volume was used as a backup. From these 282 participants, 82 were randomly selected (n = 50 healthy controls and n = 32 cancer patients) for final analysis. RNA isolation RNA isolation was performed with the miRNeasy Mini Kit by Qiagen (#217004) following the instructions of the user manual. One 5 ml aliquot was isolated to ensure the same volume for each sample. After isolation, the RNA was stored at -80°C until cDNA transcription. Quantitative PCR Transcribed miRNA samples were analyzed using a TaqMan Advanced miRNA Assay (#A25576 Applied Biosystems) in combination with TaqMan Fast Advanced Master Mix (#4444557 Applied Biosystems). A Roche LightCycler 480 Instrument II (#05015243001) was used for detection. The samples and master mix were added to 384-well plates (#04729749001 LightCycler 480 Multiwell Plate white, Roche). For correct preparation following the manufacturer’s instructions, the samples were diluted with 0.1x TE. Samples were pipetted in triplicate on each plate per assay as well as the exogenous control ath-mir-159a. qPCR statistics The LightCycler 480 data were exported as MS-EXCEL files and analyzed. The resulting Ct values were analyzed using the ∆∆Ct method. The microRNA ath-mir-159a was used as a reference gene to normalize the data to the ΔCt, and samples from healthy donors served as the second reference to calculate the miRNA fold change. GraphPad Prism software was used for statistical evaluation. Student´s t test was used to determine significant differences. miRNA sequencing and statistical analysis Sequencing libraries were prepared with the QIASeq miRNA UDI Library Kit (Qiagen, Hilde, Germany) according to the manufacturer’s instructions. To the recommended 4 µl sample input for biofluids, 1 µl of synthetic miRNAs from the QIASeq miRNA Library QC Kit was added as additional quality control. The quality of the libraries was checked on a Bioanalyzer or Tapestation (both Agilent, Waldbronn, Germany), and the libraries were quantified by a Quantus fluorometer (Promega, Madison, WI, USA). All the samples were sequenced on an Illumina NextSeq 500 instrument (Illumina, San Diego, CA, USA) in 72 bp single-end mode. Sequencing yielded a mean coverage of approximately five million reads per sample. FASTQ files were generated using bcl2fastq (Illumina). To facilitate reproducible analysis, samples were processed using the publicly available nf-core/smRNAseq pipeline version 2.2.1 (Ewels et al. 2020) implemented in Nextflow 23.04 (Di Tommaso et al. 2017) using Docker 24.0.2 (Merkel 2014) with the minimal command. All analyses were performed using custom scripts in R version 4.2.2 using the DESeq2 v.1.38.3 framework (Love, Huber, and Anders 2014). One sample was removed due to poor quality, and the remaining samples (from 32 cancer patients and 49 healthy individuals) were used in the downstream analyses. After normalizing the read counts using DESeq2, we applied cross-validation with five data folds for robust evaluation. Our approach comprises two strategies: first, the random forest algorithm was used on all 4039 miRNAs to capture broad interactions, and second, random forest-based feature selection was employed to narrow down relevant miRNAs. We utilized Python tools such as scikit-learn, numpy, pandas, matplotlib, and seaborn for analysis and visualization, ensuring a comprehensive exploration of miRNA‒target dynamics. Results Analysis of 82 urine samples from breast cancer patients and healthy female individuals We implemented miRNA sequencing as an unbiased method to evaluate the currently known miRNA genome from human urine samples. The first aim was to implement reliable and consistent detection of miRNAs in urine. Since this study should evaluate the feasibility of miRNA sequencing from reasonably small sample sizes, which can be obtained during a regular visit in an outpatient setting, we tested a sample size of 4 ml of urine in this sequencing approach. Eighty-two individual patient samples were utilized in this study, and the miRNAs were extracted from 4 ml urine samples as described in the Methods section. The 82 samples were classified into 50 healthy tumor-free control, 6 HER2-enriched, 9 luminal A, 8 luminal B and 9 TNBC tumor-bearing patient urine samples (Table 1 ). Table 1 Description of the patient cohort used for analysis. Description all healthy cancer N = 82 N = 50 N = 32 mean (sd) mean (sd) mean (sd) or or or n (%) n (%) n (%) age at urine sampling 54.5 (13.2) 50.7 (11.3) 60.4 (13.9) age at urine sampling by group 50 years 52 (63.4) 27 (54.0) 25 (78.1) age at first diagnosis na na 59.7 (14.1) BMI 26.5 (6.6) 25.6 (5.9) 27.6 (7.6) Tumor size T1 na na 16 (50.0) T2-4 na na 16 (50.0) Tumor grade G1/2 na na 12 (37.5) G3 na na 20 (62.5) Distant metastasis status cM0 na na 29 (90.6) cM1 na na 3 (9.4) Histology ductal na na 25 (78.1) lobular na na 6 (18.8) others na na 1 (3.1) molecular-like subtype Luminal A-like na na 9 (28.1) Luminal B-like na na 8 (25.0) HER2 positive na na 6 (18.8) TNBC na na 9 (28.1) miRNA sequencing of urine samples enables the detection of more than 4000 target miRNAs We analyzed the expression of more than 4000 miRNAs and detected the consistent expression of 4039 distinct miRNAs. The mean absolute miRNA expression (normalized read counts) of all the samples varied between 6.5 and 14.1, with the majority of the samples showing an average expression of 7–8 +/- 1 (Fig. 1 A). Using a 1.5-fold log change up- or downregulation of expression with an adjusted p value of 0.05 or less as a cutoff, we identified several miRNAs exhibiting differential expression between healthy and tumor-bearing individuals (Fig. 1 B). Further stratification of patient groups yielded different numbers of differentially expressed miRNAs for luminal A vs healthy (161 miRNAs), HER2 vs healthy (30 miRNAs), luminal B vs healthy (19 miRNAs) and TNBC vs healthy (12 miRNAs) patients. Several of these differentially expressed miRNAs were found in more than one of the subgroup comparisons and could therefore not be used as biomarkers to stratify individual patients. Confirmation of miRNA expression profiles by qPCR To validate our initial results, we analyzed the expression of the most highly differentially expressed miRNA markers using qPCR and confirmed the differential regulation of some (Fig. 2 A-I) but not others (Supplemental Fig. 1A-H). Overall, it must be stated that the variability in the qPCR results far exceeded the variability found in the sequencing results. Nevertheless, some subgroup-specific expression patterns, such as high expression of miR-30a-5p in TNBC, were validated (Fig. 2 F). In general, the detection of cancer in urine compared to urine from healthy individuals was more consistent even though it varied distinctly among the whole patient population. The random forest approach for data modeling The miRNA expression across the groups determined by qPCR analysis varied strongly, and there was no absolute consistency pattern within each cancer subtype for stratifying a distinct number of patients using this method. This caused the differentially expressed miRNAs to be useful as potential biomarkers for stratifying cancer subtypes. An analysis based on the miRNA sequencing data showed that the unsupervised clustering of the whole dataset did not match distinct patterns within the cancer group or a subtype (Fig. 3 A). Moreover, PCA did not reveal any distinct patterns (Fig. 3 B). To investigate the data in greater detail, we applied several machine learning algorithms to detect hidden patterns of miRNA expression. Figure 3 C shows the initial benchmark of the shallow learning methods. The random forest (RF) outperforms logistic regression, decision tree and SVM due to its ensemble approach, which reduces overfitting, captures complex relationships, handles high-dimensional data, and balances bias-variance trade-offs by aggregating diverse decision trees. Sequencing data indicating the ability of 275 individual miRNAs to distinguish BC patients from healthy women We trained the RF model with two approaches: one with 4039 miRNAs and the other with feature selection via random forest. The prediction with 275 miRNAs selected by the RF algorithm (mean AUC = 0.67) (Fig. 4 A, right) performed much better than the prediction with the whole miRNA dataset (mean AUC = 0.58) (Fig. 4 A, left). Figure 4 B shows the heatmap showing the expression of the filtered miRNAs. The ensemble approach involving the random forest algorithm produced better results than did the generic statistical approach. This is probably due to the complicated or combinatorial expression patterns of miRNAs in urine. The miRNAs in urine are a mixture of different tissues and organs at their final stop, so their expression patterns are no longer obvious and are detectable by generic statistical methods. Random forest algorithms with filtered miRNAs identify all intrinsic subtypes of breast cancer Following the approach to distinguish healthy controls from women with breast cancer, we applied random forest analysis to identify miRNA patterns that would allow us to substratify patients into the distinct breast cancer subtypes of our patient cohort: luminal A, luminal B, Her2-enriched and TNBC. For HER2-enriched BCs, 175 miRNAs out of the 4039 miRNAs were sufficient to increase the AUC on average from 0.55+/-38 to 0.68+/- 0.32 (Fig. 5 A, compare left to right). In the case of LumA-type cancer, differential expression of 195 miRNAs was associated with an increase in the AUC from 0.7+/-0.32 to 0.78+/-0.26 on average (Fig. 5 B, compared left to right). As one can easily derive from the figures with ever-increasing ROCs, an increased number of runs with more samples will increase the true positive rate dramatically. For luminal B-type breast cancer, we found 191 miRNAs that distinguish Lum B-carrying patients from healthy individuals, the difference in which increased the area under the curve (AUC) from 0.57+/-0.2 to 0.71+/-0.24 (Fig. 6 A, left/right). Finally, TNBC was detected by RF analysis of 189 miRNAs, for which the area under the curve (AUC) was 0.65 ± 0.31 and the AUC was 0.39 ± 0.22 for all the detectable miRNAs (Fig. 6 B). We found this to be the most dramatic increase in sensitivity among all the subgroups. The filtered miRNA subgroups exhibited no overlap Among the filtered miRNAs above, there were very few overlapping miRNAs (Fig. 7 A). There were no common miRNAs among the four subtypes. This finding suggested that the associated miRNAs might be distinct across these four subtypes. Most of these filtered miRNAs were not significantly or differentially expressed according to DGEA. Discussion Breast cancer treatment is based on tumor biology and tumor stage. Therefore, early detection has been an important step toward improving the curation rates observed over the last several decades. Today, it is common practice in industrialized countries to screen the female population for breast cancer on a regular basis in national programs based on mammography. Improved mammography approaches using machine learning for deeper and more accurate image analysis are therefore the next logical step in an effort to detect breast cancer as early as possible to improve treatment and curation options ( 14 , 15 ). Nevertheless, the tremendous technical and timely effort, physical discomfort during the procedure and monetary aspects of this technique could lead to the use of an easy, fast, and cost-effective prescreening method, which in the case of a positive finding would lead to an additional imaging method. Furthermore, early information about tumor biology would likely be useful for stratifying consecutive imaging and work-up procedures. Stratifying cancer patients based on noninvasive methods is currently a tremendous challenge. Especially in breast cancer, the diagnosis of TNBC has much more severe implications for the patient than a luminal A type tumor. Therefore, detecting this disease noninvasively and obtaining further information on the type of tumor would be extremely beneficial. This approach would give the treating physician a distinct advantage for subsequent work-up and treatment decisions. Here, we present the first tightly controlled miRNA sequencing effort of urine samples from breast cancer patients to gain insight into how the miRNA genome is regulated in this disease and its intrinsic subtypes. Earlier efforts from our group focused on specific miRNAs known to be regulated in breast cancer using a proprietary miRNA amplification paradigm ( 9 ). Nevertheless, in the current approach, we implemented miRNA sequencing as an innovative approach for urinary analysis to understand how many miRNAs in the currently known genome are regulated in breast cancer and whether consecutively identified signatures might represent specific subclasses of BC, allowing their detection from noninvasive urine samples. We found the let-7-miRNA family to be strongly represented in the cancer cohort, as would be expected from studies on other cancer entities using different methods of detection. The Let7-miRNAs are dysregulated in lung ( 16 ), pancreatic ( 17 ), colorectal ( 18 ), and papillary thyroid ( 19 ) cancers and, as recently described, in breast cancer ( 20 ). Let7 was further shown to regulate cancer stemness ( 21 ). Apart from these initial findings, we also detected considerable variability among the top regulated miRNAs in some samples (e.g., variability of let-7c expression in healthy individuals [Figure 2 A]), making an individual diagnosis of breast cancer or its subclasses less reliable. We therefore applied a machine learning approach to the sequencing data to investigate whether the patterns of multiple miRNAs would be more informative than those of several strongly differentially regulated miRNAs. Interestingly, the random forest approach outclassed the decision tree, logistic regression and SVM so dramatically, making it the method of choice for future analysis of miRNA sequencing data from urine samples. An increase or decrease in a single given miRNA did not seem to have as much impact as the whole “signature” of miRNA expression changes (Fig. 2 vs. Figure 4 ). The detection of very specific subsets of miRNA patterns, specifically identifying both breast cancer patients and even their specific intrinsic subtypes, is innovative and, thus far, not known. Nevertheless, more surprisingly, these patterns of miRNAs overlap very little with each other; on average, only 10–15% of miRNAs are commonly regulated, whereas most miRNAs clearly identify a subgroup or breast cancer in general. This, to our knowledge, has not been shown before and raises the question of whether previous data should be reanalyzed with a more unbiased approach to possibly identify yet unknown patterns. However, only a machine learning approach can unravel this issue, as has been shown in other fields of research ( 22 – 24 ). Consecutively, an important focus of further research should be the reduction and minimization of miRNAs included in our identified distinct miRNA pattern. The applicability of our technology for screening or early detection also relies on the sensitivity, specificity, false positive and false negative rates. The optimization of these pertinent parameters relies on large cohorts of patient and healthy control samples, which have been analyzed for this purpose. Conclusions In summary, our study represents an innovative approach and “proof of principle” concept for a sensitive noninvasive, urine-based, liquid biopsy test to detect breast cancer and its distinct intrinsic subtypes with a wide variety of application options. Declarations Ethics approval and consent to participate Samples were collected at the University Hospital Aachen (ethics vote 206/09) and at the University Hospital Erlangen (ethics vote #325_19 B). Participants in Erlangen were recruited within the iMODE-B study (Imaging and Molecular Detection of Breast Cancer; ethical approval by the ethics committee of the Friedrich-Alexander-Universität Erlangen-Nuremberg; #325_19 B). Patients were eligible for inclusion if they had an indication for a diagnostic biopsy due to a suspicious breast lesion. The main aim of the iMODE-B study was to identify molecular markers that are predictive of patient prognosis and treatment response at the time of the first diagnosis of breast cancer. After the participants provided written informed consent in accordance with the Declaration of Helsinki, biospecimen sampling was performed (blood draw and urine). Availability of data and materials The datasets generated and/or analyzed during the current study are not publicly available due an ongoing licensing process but are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests. Funding This work was strongly supported by Dr. Pommer-Jung-Stiftung. Authors' contributions JM analyzed and interpreted patient data and was a major contributor to the writing and editing of the manuscript. MR provided patient urine samples and information on the patients. CK worked on the machine learning algorithms and provided the random forest analysis. BK performed the qPCR analysis. JF provided sequencing data. JW interpreted patient data. TK analyzed the correlations of patient data with subtypes of breast cancer. LN edited the manuscript and provided writing support. PF contributed to writing and editing the manuscript. ES provided funding and the original idea for the research and supported the execution and writing and editing of the manuscript. All the authors read and approved the final manuscript. Acknowledgments This work was supported by the Genomics Facility, a core facility of the Interdisciplinary Center for Clinical Research (IZKF) Aachen within the Faculty of Medicine at RWTH Aachen University. We thank Lothar Häberle for advice and support. References Nolan E, Lindeman GJ, Visvader JE. Deciphering breast cancer: from biology to the clinic. Cell. 2023;186(8):1708–28. Luu XQ, Lee K, Jun JK, Suh M, Jung KW, Choi KS. Effect of mammography screening on the long-term survival of breast cancer patients: results from the National Cancer Screening Program in Korea. Epidemiol Health. 2022;44:e2022094. Lowry KP, Callaway KA, Lee JM, Zhang F, Ross-Degnan D, Wharam JF, et al. Trends in Annual Surveillance Mammography Participation Among Breast Cancer Survivors From 2004 to 2016. J Natl Compr Canc Netw. 2022;20(4):379–86. e9. Kupec T, Bleilevens A, Klein B, Hansen T, Najjari L, Wittenborn J et al. Comparison of Serum and Urine as Sources of miRNA Markers for the Detection of Ovarian Cancer. Biomedicines. 2023;11(9). Duque G, Manterola C, Otzen T, Arias C, Palacios D, Mora M, et al. Cancer Biomarkers in Liquid Biopsy for Early Detection of Breast Cancer: A Systematic Review. Clin Med Insights Oncol. 2022;16:11795549221134831. Shiao MS, Chang JM, Lertkhachonsuk AA, Rermluk N, Jinawath N. Circulating Exosomal miRNAs as Biomarkers in Epithelial Ovarian Cancer. Biomedicines. 2021;9(10). Kupec T, Bleilevens A, Iborra S, Najjari L, Wittenborn J, Maurer J, Stickeler E. Stability of circulating microRNAs in serum. PLoS ONE. 2022;17(8):e0268958. Hulstaert E, Morlion A, Levanon K, Vandesompele J, Mestdagh P. Candidate RNA biomarkers in biofluids for early diagnosis of ovarian cancer: A systematic review. Gynecol Oncol. 2021;160(2):633–42. Erbes T, Hirschfeld M, Rucker G, Jaeger M, Boas J, Iborra S, et al. Feasibility of urinary microRNA detection in breast cancer patients and its potential as an innovative noninvasive biomarker. BMC Cancer. 2015;15:193. Kumar S, Keerthana R, Pazhanimuthu A, Perumal P. Overexpression of circulating miRNA-21 and miRNA-146a in plasma samples of breast cancer patients. Indian J Biochem Biophys. 2013;50(3):210–4. Rama K, Bitla AR, Hulikal N, Yootla M, Yadagiri LA, Asha T et al. Assessment of serum microRNA-21 and miRNA-205 as diagnostic markers for stage I and II breast cancer in Indian population. Indian J Cancer. 2023. Zhang L, Xu Y, Jin X, Wang Z, Wu Y, Zhao D, et al. A circulating miRNA signature as a diagnostic biomarker for noninvasive early detection of breast cancer. Breast Cancer Res Treat. 2015;154(2):423–34. Zhao H, Gao A, Zhang Z, Tian R, Luo A, Li M, et al. Genetic analysis and preliminary function study of miR-423 in breast cancer. Tumor Biol. 2015;36(6):4763–71. Houssami N, Marinovich ML. AI for mammography screening: enter evidence from prospective trials. Lancet Digit Health. 2023;5(10):e641–e2. Ng AY, Oberije CJG, Ambrozay E, Szabo E, Serfozo O, Karpati E, et al. Prospective implementation of AI-assisted screen reading to improve early detection of breast cancer. Nat Med. 2023;29(12):3044–9. Takamizawa J, Konishi H, Yanagisawa K, Tomida S, Osada H, Endoh H, et al. Reduced expression of the let-7 microRNAs in human lung cancers in association with shortened postoperative survival. Cancer Res. 2004;64(11):3753–6. Xiong G, Liu C, Yang G, Feng M, Xu J, Zhao F, et al. Long noncoding RNA GSTM3TV2 upregulates LAT2 and OLR1 by competitively sponging let-7 to promote gemcitabine resistance in pancreatic cancer. J Hematol Oncol. 2019;12(1):97. Langevin SM, Christensen BC. Let-7 microRNA-binding-site polymorphism in the 3'UTR of KRAS and colorectal cancer outcome: a systematic review and meta-analysis. Cancer Med. 2014;3(5):1385–95. Perdas E, Stawski R, Kaczka K, Zubrzycka M. Analysis of Let-7 Family miRNA in Plasma as Potential Predictive Biomarkers of Diagnosis for Papillary Thyroid Cancer. Diagnostics (Basel). 2020;10(3). Chiu SC, Chung HY, Cho DY, Chan TM, Liu MC, Huang HM, et al. Therapeutic potential of microRNA let-7: tumor suppression or impeding normal stemness. Cell Transpl. 2014;23(4–5):459–69. Ma Y, Shen N, Wicha MS, Luo M. The Roles of the Let-7 Family of MicroRNAs in the Regulation of Cancer Stemness. Cells. 2021;10(9). Lee E, Jung SY, Hwang HJ, Jung J. Patient-Level Cancer Prediction Models From a Nationwide Patient Cohort: Model Development and Validation. JMIR Med Inf. 2021;9(8):e29807. Zhang K, Liu C, Sha X, Yao S, Li Z, Yu Y, et al. Development and validation of a prediction model to predict major adverse cardiovascular events in elderly patients undergoing noncardiac surgery: A retrospective cohort study. Atherosclerosis. 2023;376:71–9. Awad A, Bader-El-Den M, McNicholas J, Briggs J. Early hospital mortality prediction of intensive care unit patients using an ensemble learning approach. Int J Med Inf. 2017;108:185–95. Additional Declarations No competing interests reported. Supplementary Files FigureLayoutmiRNAPaperfinalJMSupplFig1.tif Supplemental Figure 1 Quantitative PCR of several differentially expressed (healthy vs. tumor) miRNAs. QPCR-based expression of the depicted miRNAs in healthy individuals (healthy) from all tumor entities combined (tumor) and from the subclasses luminal A (Lum A), luminal B (Lum B), triple-negative breast cancer (TNBC) and HER2-enriched tumors (Her2). The expression of let-7e-5p (A), let-7f-5p (B), 451b (C), 21-3p (D), 21-5p (E), 451a (F), 125b-5p (G) and 26a-5p (H) was analyzed. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3993094","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":275319560,"identity":"c05529a2-6073-4e18-8f4f-841be48cfe12","order_by":0,"name":"Jochen Maurer","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIiWNgGAWjYBACxgYILcfAA+clEKfFmHgtMJDYQLQW5mmHn0lXVNxL33Dm8MMPP3fY5DGwJx/A77DZaWaSZ84U524422Ys2XsmrZiB5xl+axhnJ5hJNrYl5G44z2Agzdh2OLFBIseAgJb0b5KN/xLSDc6zf/7N2PYfqCX/AwEtOUBbGhISDM72mAFtOQCyBa8OkJZiy4ZjCYYzz5wps+xtS05s43mG32GGs9M33myoSZDnO5O++cbPNrvEfvbkB/i1NKCLsOF3FgODPCEFo2AUjIJRMAoYALQUSwqzb9kmAAAAAElFTkSuQmCC","orcid":"","institution":"University Hospital RWTH Aachen","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jochen","middleName":"","lastName":"Maurer","suffix":""},{"id":275319561,"identity":"66890358-340d-4da3-a86c-112ce17d35b9","order_by":1,"name":"Matthias Rübner","email":"","orcid":"","institution":"Erlangen University Hospital, Friedrich-Alexander University Erlangen-Nuremberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Matthias","middleName":"","lastName":"Rübner","suffix":""},{"id":275319562,"identity":"185bf0ec-97c2-4881-b67f-dff3c8103e82","order_by":2,"name":"Chao-Chung Kuo","email":"","orcid":"","institution":"RWTH Aachen University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Chao-Chung","middleName":"","lastName":"Kuo","suffix":""},{"id":275319563,"identity":"8272006b-5637-4688-a9d7-728fe172fc17","order_by":3,"name":"Birgit Klein","email":"","orcid":"","institution":"University Hospital RWTH Aachen","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Birgit","middleName":"","lastName":"Klein","suffix":""},{"id":275319564,"identity":"f8815500-5248-41d0-916d-e7577f88bc48","order_by":4,"name":"Julia Franzen","email":"","orcid":"","institution":"RWTH Aachen University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Julia","middleName":"","lastName":"Franzen","suffix":""},{"id":275319565,"identity":"bb0eec9d-794c-4dd3-8011-12026396b161","order_by":5,"name":"Julia Wittenborn","email":"","orcid":"","institution":"University Hospital RWTH Aachen","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Julia","middleName":"","lastName":"Wittenborn","suffix":""},{"id":275319566,"identity":"711dbea5-d3f4-44dd-b089-a94c7d77973d","order_by":6,"name":"Tomas Kupec","email":"","orcid":"","institution":"University Hospital RWTH Aachen","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tomas","middleName":"","lastName":"Kupec","suffix":""},{"id":275319567,"identity":"c795e464-435c-417a-a36a-ee9b70d8bc3f","order_by":7,"name":"Laila Najjari","email":"","orcid":"","institution":"University Hospital RWTH Aachen","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Laila","middleName":"","lastName":"Najjari","suffix":""},{"id":275319568,"identity":"f0452d1c-b7a7-4560-9f28-e8c9ae634123","order_by":8,"name":"Peter Fasching","email":"","orcid":"","institution":"Erlangen University Hospital, Friedrich-Alexander University Erlangen-Nuremberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Peter","middleName":"","lastName":"Fasching","suffix":""},{"id":275319569,"identity":"aca3a53e-c5ef-4ec5-a95c-362495ca5d0d","order_by":9,"name":"Elmar Stickeler","email":"","orcid":"","institution":"University Hospital RWTH Aachen","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Elmar","middleName":"","lastName":"Stickeler","suffix":""}],"badges":[],"createdAt":"2024-02-27 06:44:49","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3993094/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3993094/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":51822193,"identity":"69704218-0b30-46c4-b458-35f0e1d28a9a","added_by":"auto","created_at":"2024-02-29 16:12:36","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1535094,"visible":true,"origin":"","legend":"\u003cp\u003eInitial analysis and evaluation of the sequencing data. (A) Mean normalized expression data (read counts). (B) Volcano plot showing a 1.5-fold log change up- or downregulation of expression with an adjusted p value of 0.05 or less as a cutoff (miRNAs identified as red dots were considered significant; those identified as not significant are shown in blue). (C) Subclasses of breast cancer could be identified by groups of miRNAs. LumA (Luminal A), LumB (Luminal B), HER2 (Her2-enriched) and TNBC (triple-negative breast cancer).\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig1.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/648edb94399990b1da6a7b5a.png"},{"id":51822189,"identity":"802bc32d-372a-4371-8658-c10e2807ade7","added_by":"auto","created_at":"2024-02-29 16:12:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":362867,"visible":true,"origin":"","legend":"\u003cp\u003eQuantitative PCR of several differentially expressed (healthy vs. tumor) miRNAs. QPCR-based expression of the depicted miRNAs in healthy individuals (healthy) from all tumor entities combined (tumor) and from the subclasses luminal A (Lum A), luminal B (Lum B), triple-negative breast cancer (TNBC) and HER2-enriched tumors (Her2). The expression of let-7c-5p (A), 30a-3p (B), let-7a-5p (C), 151-3p (D), let-7i-5p (E), 30a-5p (F), 30e-5p (G), 30d-5p (H) and 196a-5p (I) was analyzed.\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig2.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/fc01820de02326fa8ef6d868.png"},{"id":51822188,"identity":"445153d9-0787-4da9-b950-d3542d810069","added_by":"auto","created_at":"2024-02-29 16:12:36","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":2274888,"visible":true,"origin":"","legend":"\u003cp\u003eData analysis (A). Unsupervised clustering of the whole dataset of miRNA sequencing data. (B) PCA. (C) Analysis of several machine learning algorithms to detect hidden patterns of miRNA expression. Theinitial benchmarks of the learning methods used are as follows: logistic regression, decision tree, random forest and support vector machine(SVM).\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig3.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/36e9d726f71edb34b7128859.png"},{"id":51822191,"identity":"e793b0d5-0db3-4a63-95e9-9e2c1587feac","added_by":"auto","created_at":"2024-02-29 16:12:36","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1881580,"visible":true,"origin":"","legend":"\u003cp\u003eReceiver operating characteristic (ROC) curves were generated using sets of miRNAs to distinguish cancer patients from healthy individuals. (A) ROC curve analysis of all 4039 detectable miRNAs (left side) and a small subset of 275 filtered miRNAs (right side). (B) Cluster dendrogram of all samples for the275 filtered miRNAs.\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig4.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/802abc14dce48931b7c80a1b.png"},{"id":51822195,"identity":"7bc246f3-4744-46f5-8600-8d1d3cb1ee13","added_by":"auto","created_at":"2024-02-29 16:12:37","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1762348,"visible":true,"origin":"","legend":"\u003cp\u003eReceiver operating characteristic (ROC) curves were generated using sets of miRNAs to distinguish subtypes of breast cancer. (A) ROC curve for HER2 patient samples vs. healthy individuals using all 4039 detectable miRNAs (left side) and with only a small subset of 175 filtered miRNAs (right side). (B) ROC curve for luminal A (Lum A)-type breast cancer patient samples vs. healthy individuals using all 4039 detectable miRNAs (left side) and with only a small subset of 195 filtered miRNAs (right side).\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig5.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/8c9f15eaf4579aab1bda01d1.png"},{"id":51822701,"identity":"8cf69316-13e5-4b26-a735-d0e3f50f1eac","added_by":"auto","created_at":"2024-02-29 16:20:37","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1583821,"visible":true,"origin":"","legend":"\u003cp\u003eReceiver operating characteristic (ROC) curves were generatedusing sets of miRNAs to distinguish subtypes of breast cancer. (A) ROC curve for luminal B (Lum B) breast cancer patient samples vs. healthy individuals using all 4039 detectable miRNAs (left side) and with only a small subset of 191 filtered miRNAs (right side). (B) ROC curve for triple-negative breast cancer patient samples vs. healthy individuals using all 4039 detectable miRNAs (left side) and with only a small subset of 189 filtered miRNAs (right side).\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig6.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/49a9f0fb405602807e6b18c0.png"},{"id":51822194,"identity":"bd00f16a-9600-42e0-845a-56d2892d0729","added_by":"auto","created_at":"2024-02-29 16:12:36","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":424188,"visible":true,"origin":"","legend":"\u003cp\u003eOverlapping and distinct miRNAs identified in cancer patients and the indicated subgroups.\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMFig7.png","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/e15cabace99c7ad2b4de8209.png"},{"id":56808891,"identity":"58085f31-1abe-484b-a9f0-41da57c0aca5","added_by":"auto","created_at":"2024-05-20 18:42:10","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":10039605,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/ecc8aaaf-f219-45c2-bf63-a2c2636ba532.pdf"},{"id":51822190,"identity":"456318af-59b9-4745-9163-41832cd66f49","added_by":"auto","created_at":"2024-02-29 16:12:36","extension":"tif","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2086096,"visible":true,"origin":"","legend":"\u003cp\u003eSupplemental Figure 1\u003c/p\u003e\n\u003cp\u003eQuantitative PCR of several differentially expressed (healthy vs. tumor) miRNAs. QPCR-based expression of the depicted miRNAs in healthy individuals (healthy) from all tumor entities combined (tumor) and from the subclasses luminal A (Lum A), luminal B (Lum B), triple-negative breast cancer (TNBC) and HER2-enriched tumors (Her2). The expression of let-7e-5p (A), let-7f-5p (B), 451b (C), 21-3p (D), 21-5p (E), 451a (F), 125b-5p (G) and 26a-5p (H) was analyzed.\u003c/p\u003e","description":"","filename":"FigureLayoutmiRNAPaperfinalJMSupplFig1.tif","url":"https://assets-eu.researchsquare.com/files/rs-3993094/v1/51a690787b43fbb83a62576e.tif"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine learning approach identifies miRNA signatures for breast cancer detection and classification from patient urine samples","fulltext":[{"header":"Background","content":"\u003cp\u003eBreast cancer (BC) is still the leading cause of cancer-related deaths among the female population, affecting 1 in 8 women in her lifetime (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Moreover, this disease still places an enormous burden on healthcare systems worldwide. In addition to tumor biology, which guides treatment options and therapeutic needs in general, tumor stage is still an important risk factor for treatment decisions. The implementation of mammography-based early detection programs consecutively led to a decreased tumor size (T stage) and a decreased rate of axillary lymph node involvement (N stage), thereby improving survival rates in patients with breast cancer (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). Almost 100% of German women aged 50 to 69 years are invited to undergo mammography screening every 2 years. However, since its implementation in 2009, the participation rate has stagnated at approximately 50% (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTherefore, it appears reasonable that these programs could dramatically profit from new early detection methodologies, which are less invasive than the current standards.\u003c/p\u003e \u003cp\u003eMicroRNAs (miRNAs) are small noncoding RNA molecules that regulate gene expression by binding to specific target mRNAs, affecting their translation. They are found in many different cell types and tissues and have been shown to play a role in a wide range of biological processes, including cell growth and proliferation, differentiation, and apoptosis. Recent studies have shown that miRNAs can also be found in body fluids, such as blood and urine (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e), where they can be easily isolated and quantified and can act as biomarkers for various diseases, including cancer (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAs biomarkers, miRNAs possess high potential for cancer detection because they are stable in body fluids, which allows easy collection and storage of samples (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). They are also present and detectable in very small samples, which makes them ideal biomarkers (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). Additionally, miRNA-based diagnostic tests can be noninvasive, which is especially important for the early detection of cancer.\u003c/p\u003e \u003cp\u003eSeveral studies have investigated the use of miRNAs as biomarkers for breast cancer and other gynecological malignancies. In breast cancer, several miRNAs, including miR-21, miR-155, and miR-205, have been identified as biomarkers for this disease(\u003cspan additionalcitationids=\"CR10\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). Studies have shown that the levels of these miRNAs are elevated in the blood of breast cancer patients compared to healthy individuals. Furthermore, an earlier study using qPCR analysis identified miR-424 and miR-423 as BC markers in urine as source materials(\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Researchers have also shown that miRNAs can be used to distinguish between different types of breast cancer, such as invasive ductal carcinoma and invasive lobular carcinoma.\u003c/p\u003e \u003cp\u003eIt is worth noting that while these studies have been promising, additional research is needed to fully establish the clinical utility of miRNA-based diagnostic tests for breast and ovarian cancer. There is still much to be learned about the potential of miRNA-based diagnostic tests. However, additional research is needed to fully understand the biology of miRNAs and how they are regulated in cancer cells. Additionally, additional studies are needed to validate the use of miRNA-based diagnostic tests in clinical settings, including larger, multicenter studies that will help to establish the specificity and sensitivity of these tests.\u003c/p\u003e \u003cp\u003eTo investigate the feasibility of miRNA detection as a diagnostic tool for breast cancer, we investigated a cohort of 82 urine samples from BC patients and healthy individuals using miRNA sequencing for the first time. The goal was to identify all currently known miRNAs regulated in BC in a completely unbiased setting without prior selection of BC-specific miRNAs from the literature. Here, we present the first step toward the clinical use of miRNA sequencing in BC detection and provide insight into the stratification of BC patients utilizing this information.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePatient Selection and Urine Sample Preparation\u003c/h2\u003e \u003cp\u003e Samples were collected at the University Hospital Aachen (ethics vote 206/09) and at the University Hospital Erlangen. Participants in Erlangen were recruited within the iMODE-B study (Imaging and Molecular Detection of Breast Cancer; ethical approval by the ethics committee of the Friedrich-Alexander-Universit\u0026auml;t Erlangen-Nuremberg; #325_19 B). Patients were eligible for inclusion if they had an indication for a diagnostic biopsy due to a suspicious breast lesion. The main aim of the iMODE-B study was to identify molecular markers that are predictive of patient prognosis and treatment response at the time of the first diagnosis of breast cancer. After the participants provided written informed consent in accordance with the Declaration of Helsinki, biospecimen sampling was performed (blood draw and urine).\u003c/p\u003e \u003cp\u003eA total of 355 urine samples were collected between November 2019 and July 2020. The urine samples were centrifuged at 944 \u0026times; g for 10 minutes at room temperature to separate the cell pellet and cell-free supernatant, which were stored separately at -80\u0026deg;C until further use. For miRNA extraction, at least 7 ml of cell-free supernatant was available, 4 ml was used for miRNA extraction, and the remaining volume was used as a backup. From these 282 participants, 82 were randomly selected (n\u0026thinsp;=\u0026thinsp;50 healthy controls and n\u0026thinsp;=\u0026thinsp;32 cancer patients) for final analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eRNA isolation\u003c/h2\u003e \u003cp\u003eRNA isolation was performed with the miRNeasy Mini Kit by Qiagen (#217004) following the instructions of the user manual. One 5 ml aliquot was isolated to ensure the same volume for each sample. After isolation, the RNA was stored at -80\u0026deg;C until cDNA transcription.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eQuantitative PCR\u003c/h2\u003e \u003cp\u003eTranscribed miRNA samples were analyzed using a TaqMan Advanced miRNA Assay (#A25576 Applied Biosystems) in combination with TaqMan Fast Advanced Master Mix (#4444557 Applied Biosystems). A Roche LightCycler 480 Instrument II (#05015243001) was used for detection. The samples and master mix were added to 384-well plates (#04729749001 LightCycler 480 Multiwell Plate white, Roche). For correct preparation following the manufacturer\u0026rsquo;s instructions, the samples were diluted with 0.1x TE. Samples were pipetted in triplicate on each plate per assay as well as the exogenous control ath-mir-159a.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eqPCR statistics\u003c/h2\u003e \u003cp\u003eThe LightCycler 480 data were exported as MS-EXCEL files and analyzed. The resulting Ct values were analyzed using the ∆∆Ct method. The microRNA ath-mir-159a was used as a reference gene to normalize the data to the ΔCt, and samples from healthy donors served as the second reference to calculate the miRNA fold change.\u003c/p\u003e \u003cp\u003eGraphPad Prism software was used for statistical evaluation. Student\u0026acute;s t test was used to determine significant differences.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003emiRNA sequencing and statistical analysis\u003c/h2\u003e \u003cp\u003eSequencing libraries were prepared with the QIASeq miRNA UDI Library Kit (Qiagen, Hilde, Germany) according to the manufacturer\u0026rsquo;s instructions. To the recommended 4 \u0026micro;l sample input for biofluids, 1 \u0026micro;l of synthetic miRNAs from the QIASeq miRNA Library QC Kit was added as additional quality control. The quality of the libraries was checked on a Bioanalyzer or Tapestation (both Agilent, Waldbronn, Germany), and the libraries were quantified by a Quantus fluorometer (Promega, Madison, WI, USA). All the samples were sequenced on an Illumina NextSeq 500 instrument (Illumina, San Diego, CA, USA) in 72 bp single-end mode. Sequencing yielded a mean coverage of approximately five million reads per sample.\u003c/p\u003e \u003cp\u003eFASTQ files were generated using bcl2fastq (Illumina). To facilitate reproducible analysis, samples were processed using the publicly available nf-core/smRNAseq pipeline version 2.2.1 (Ewels et al. 2020) implemented in Nextflow 23.04 (Di Tommaso et al. 2017) using Docker 24.0.2 (Merkel 2014) with the minimal command. All analyses were performed using custom scripts in R version 4.2.2 using the DESeq2 v.1.38.3 framework (Love, Huber, and Anders 2014).\u003c/p\u003e \u003cp\u003eOne sample was removed due to poor quality, and the remaining samples (from 32 cancer patients and 49 healthy individuals) were used in the downstream analyses. After normalizing the read counts using DESeq2, we applied cross-validation with five data folds for robust evaluation. Our approach comprises two strategies: first, the random forest algorithm was used on all 4039 miRNAs to capture broad interactions, and second, random forest-based feature selection was employed to narrow down relevant miRNAs. We utilized Python tools such as scikit-learn, numpy, pandas, matplotlib, and seaborn for analysis and visualization, ensuring a comprehensive exploration of miRNA‒target dynamics.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of 82 urine samples from breast cancer patients and healthy female individuals\u003c/h2\u003e \u003cp\u003eWe implemented miRNA sequencing as an unbiased method to evaluate the currently known miRNA genome from human urine samples. The first aim was to implement reliable and consistent detection of miRNAs in urine. Since this study should evaluate the feasibility of miRNA sequencing from reasonably small sample sizes, which can be obtained during a regular visit in an outpatient setting, we tested a sample size of 4 ml of urine in this sequencing approach.\u003c/p\u003e \u003cp\u003eEighty-two individual patient samples were utilized in this study, and the miRNAs were extracted from 4 ml urine samples as described in the \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003eMethods\u003c/span\u003e section. The 82 samples were classified into 50 healthy tumor-free control, 6 HER2-enriched, 9 luminal A, 8 luminal B and 9 TNBC tumor-bearing patient urine samples (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDescription of the patient cohort used for analysis.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ehealthy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ecancer\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eN\u0026thinsp;=\u0026thinsp;82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eN\u0026thinsp;=\u0026thinsp;50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eN\u0026thinsp;=\u0026thinsp;32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emean (sd)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003emean (sd)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003emean (sd)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eor\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003en (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003en (%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eage at urine sampling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e54.5 (13.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50.7 (11.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e60.4 (13.9)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eage at urine sampling by group\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;50 years\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e30 (36.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e23 (46.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7 (21.9)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026gt;\u0026thinsp;50 years\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e52 (63.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e27 (54.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e25 (78.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eage at first diagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e59.7 (14.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBMI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e26.5 (6.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e25.6 (5.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e27.6 (7.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTumor size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eT1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16 (50.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eT2-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16 (50.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTumor grade\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eG1/2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e12 (37.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eG3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e20 (62.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDistant metastasis status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ecM0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e29 (90.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ecM1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3 (9.4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHistology\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eductal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e25 (78.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elobular\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6 (18.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eothers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 (3.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emolecular-like subtype\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLuminal A-like\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9 (28.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLuminal B-like\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e8 (25.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHER2 positive\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6 (18.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTNBC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ena\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9 (28.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003emiRNA sequencing of urine samples enables the detection of more than 4000 target miRNAs\u003c/h2\u003e \u003cp\u003eWe analyzed the expression of more than 4000 miRNAs and detected the consistent expression of 4039 distinct miRNAs. The mean absolute miRNA expression (normalized read counts) of all the samples varied between 6.5 and 14.1, with the majority of the samples showing an average expression of 7\u0026ndash;8 +/- 1 (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). Using a 1.5-fold log change up- or downregulation of expression with an adjusted p value of 0.05 or less as a cutoff, we identified several miRNAs exhibiting differential expression between healthy and tumor-bearing individuals (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB). Further stratification of patient groups yielded different numbers of differentially expressed miRNAs for luminal A vs healthy (161 miRNAs), HER2 vs healthy (30 miRNAs), luminal B vs healthy (19 miRNAs) and TNBC vs healthy (12 miRNAs) patients. Several of these differentially expressed miRNAs were found in more than one of the subgroup comparisons and could therefore not be used as biomarkers to stratify individual patients.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eConfirmation of miRNA expression profiles by qPCR\u003c/h2\u003e \u003cp\u003eTo validate our initial results, we analyzed the expression of the most highly differentially expressed miRNA markers using qPCR and confirmed the differential regulation of some (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA-I) but not others (Supplemental Fig.\u0026nbsp;1A-H). Overall, it must be stated that the variability in the qPCR results far exceeded the variability found in the sequencing results. Nevertheless, some subgroup-specific expression patterns, such as high expression of miR-30a-5p in TNBC, were validated (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eF). In general, the detection of cancer in urine compared to urine from healthy individuals was more consistent even though it varied distinctly among the whole patient population.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eThe random forest approach for data modeling\u003c/h2\u003e \u003cp\u003eThe miRNA expression across the groups determined by qPCR analysis varied strongly, and there was no absolute consistency pattern within each cancer subtype for stratifying a distinct number of patients using this method. This caused the differentially expressed miRNAs to be useful as potential biomarkers for stratifying cancer subtypes.\u003c/p\u003e \u003cp\u003eAn analysis based on the miRNA sequencing data showed that the unsupervised clustering of the whole dataset did not match distinct patterns within the cancer group or a subtype (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). Moreover, PCA did not reveal any distinct patterns (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). To investigate the data in greater detail, we applied several machine learning algorithms to detect hidden patterns of miRNA expression. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC shows the initial benchmark of the shallow learning methods. The random forest (RF) outperforms logistic regression, decision tree and SVM due to its ensemble approach, which reduces overfitting, captures complex relationships, handles high-dimensional data, and balances bias-variance trade-offs by aggregating diverse decision trees.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eSequencing data indicating the ability of 275 individual miRNAs to distinguish BC patients from healthy women\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe trained the RF model with two approaches: one with 4039 miRNAs and the other with feature selection via random forest. The prediction with 275 miRNAs selected by the RF algorithm (mean AUC\u0026thinsp;=\u0026thinsp;0.67) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA, right) performed much better than the prediction with the whole miRNA dataset (mean AUC\u0026thinsp;=\u0026thinsp;0.58) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA, left).\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB shows the heatmap showing the expression of the filtered miRNAs. The ensemble approach involving the random forest algorithm produced better results than did the generic statistical approach. This is probably due to the complicated or combinatorial expression patterns of miRNAs in urine. The miRNAs in urine are a mixture of different tissues and organs at their final stop, so their expression patterns are no longer obvious and are detectable by generic statistical methods.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eRandom forest algorithms with filtered miRNAs identify all intrinsic subtypes of breast cancer\u003c/h2\u003e \u003cp\u003eFollowing the approach to distinguish healthy controls from women with breast cancer, we applied random forest analysis to identify miRNA patterns that would allow us to substratify patients into the distinct breast cancer subtypes of our patient cohort: luminal A, luminal B, Her2-enriched and TNBC.\u003c/p\u003e \u003cp\u003eFor HER2-enriched BCs, 175 miRNAs out of the 4039 miRNAs were sufficient to increase the AUC on average from 0.55+/-38 to 0.68+/- 0.32 (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA, compare left to right). In the case of LumA-type cancer, differential expression of 195 miRNAs was associated with an increase in the AUC from 0.7+/-0.32 to 0.78+/-0.26 on average (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB, compared left to right). As one can easily derive from the figures with ever-increasing ROCs, an increased number of runs with more samples will increase the true positive rate dramatically.\u003c/p\u003e \u003cp\u003eFor luminal B-type breast cancer, we found 191 miRNAs that distinguish Lum B-carrying patients from healthy individuals, the difference in which increased the area under the curve (AUC) from 0.57+/-0.2 to 0.71+/-0.24 (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, left/right). Finally, TNBC was detected by RF analysis of 189 miRNAs, for which the area under the curve (AUC) was 0.65\u0026thinsp;\u0026plusmn;\u0026thinsp;0.31 and the AUC was 0.39\u0026thinsp;\u0026plusmn;\u0026thinsp;0.22 for all the detectable miRNAs (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eB). We found this to be the most dramatic increase in sensitivity among all the subgroups.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eThe filtered miRNA subgroups exhibited no overlap\u003c/h2\u003e \u003cp\u003eAmong the filtered miRNAs above, there were very few overlapping miRNAs (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA). There were no common miRNAs among the four subtypes. This finding suggested that the associated miRNAs might be distinct across these four subtypes. Most of these filtered miRNAs were not significantly or differentially expressed according to DGEA.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eBreast cancer treatment is based on tumor biology and tumor stage. Therefore, early detection has been an important step toward improving the curation rates observed over the last several decades. Today, it is common practice in industrialized countries to screen the female population for breast cancer on a regular basis in national programs based on mammography. Improved mammography approaches using machine learning for deeper and more accurate image analysis are therefore the next logical step in an effort to detect breast cancer as early as possible to improve treatment and curation options (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). Nevertheless, the tremendous technical and timely effort, physical discomfort during the procedure and monetary aspects of this technique could lead to the use of an easy, fast, and cost-effective prescreening method, which in the case of a positive finding would lead to an additional imaging method.\u003c/p\u003e \u003cp\u003eFurthermore, early information about tumor biology would likely be useful for stratifying consecutive imaging and work-up procedures.\u003c/p\u003e \u003cp\u003eStratifying cancer patients based on noninvasive methods is currently a tremendous challenge. Especially in breast cancer, the diagnosis of TNBC has much more severe implications for the patient than a luminal A type tumor. Therefore, detecting this disease noninvasively and obtaining further information on the type of tumor would be extremely beneficial. This approach would give the treating physician a distinct advantage for subsequent work-up and treatment decisions.\u003c/p\u003e \u003cp\u003eHere, we present the first tightly controlled miRNA sequencing effort of urine samples from breast cancer patients to gain insight into how the miRNA genome is regulated in this disease and its intrinsic subtypes. Earlier efforts from our group focused on specific miRNAs known to be regulated in breast cancer using a proprietary miRNA amplification paradigm (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). Nevertheless, in the current approach, we implemented miRNA sequencing as an innovative approach for urinary analysis to understand how many miRNAs in the currently known genome are regulated in breast cancer and whether consecutively identified signatures might represent specific subclasses of BC, allowing their detection from noninvasive urine samples.\u003c/p\u003e \u003cp\u003eWe found the let-7-miRNA family to be strongly represented in the cancer cohort, as would be expected from studies on other cancer entities using different methods of detection. The Let7-miRNAs are dysregulated in lung (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e), pancreatic (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e), colorectal (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e), and papillary thyroid (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e) cancers and, as recently described, in breast cancer (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). Let7 was further shown to regulate cancer stemness (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eApart from these initial findings, we also detected considerable variability among the top regulated miRNAs in some samples (e.g., variability of let-7c expression in healthy individuals [Figure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA]), making an individual diagnosis of breast cancer or its subclasses less reliable. We therefore applied a machine learning approach to the sequencing data to investigate whether the patterns of multiple miRNAs would be more informative than those of several strongly differentially regulated miRNAs. Interestingly, the random forest approach outclassed the decision tree, logistic regression and SVM so dramatically, making it the method of choice for future analysis of miRNA sequencing data from urine samples.\u003c/p\u003e \u003cp\u003eAn increase or decrease in a single given miRNA did not seem to have as much impact as the whole \u0026ldquo;signature\u0026rdquo; of miRNA expression changes (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e vs. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). The detection of very specific subsets of miRNA patterns, specifically identifying both breast cancer patients and even their specific intrinsic subtypes, is innovative and, thus far, not known. Nevertheless, more surprisingly, these patterns of miRNAs overlap very little with each other; on average, only 10\u0026ndash;15% of miRNAs are commonly regulated, whereas most miRNAs clearly identify a subgroup or breast cancer in general. This, to our knowledge, has not been shown before and raises the question of whether previous data should be reanalyzed with a more unbiased approach to possibly identify yet unknown patterns. However, only a machine learning approach can unravel this issue, as has been shown in other fields of research (\u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eConsecutively, an important focus of further research should be the reduction and minimization of miRNAs included in our identified distinct miRNA pattern. The applicability of our technology for screening or early detection also relies on the sensitivity, specificity, false positive and false negative rates. The optimization of these pertinent parameters relies on large cohorts of patient and healthy control samples, which have been analyzed for this purpose.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn summary, our study represents an innovative approach and \u0026ldquo;proof of principle\u0026rdquo; concept for a sensitive noninvasive, urine-based, liquid biopsy test to detect breast cancer and its distinct intrinsic subtypes with a wide variety of application options.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSamples were collected at the University Hospital Aachen (ethics vote 206/09)\u0026nbsp;and\u0026nbsp;at the University Hospital Erlangen (ethics vote #325_19 B). Participants in Erlangen were recruited within the iMODE-B study (Imaging and Molecular Detection of Breast Cancer; ethical\u0026nbsp;approval\u0026nbsp;by the ethics committee of the Friedrich-Alexander-Universität Erlangen-Nuremberg; #325_19 B). Patients were eligible for inclusion if they had an indication for a diagnostic biopsy due to a suspicious breast lesion. The main aim of the iMODE-B study\u0026nbsp;was\u0026nbsp;to identify molecular markers\u0026nbsp;that are predictive of patient prognosis and treatment response at the time of the first diagnosis of breast cancer. After the participants provided\u0026nbsp;written informed consent in accordance with the Declaration of Helsinki, biospecimen sampling was\u0026nbsp;performed\u0026nbsp;(blood draw and urine).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analyzed during the current study are not publicly available due an ongoing licensing process but are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was strongly supported by Dr. Pommer-Jung-Stiftung.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors' contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJM\u0026nbsp;analyzed\u0026nbsp;and interpreted patient data\u0026nbsp;and\u0026nbsp;was a major contributor\u0026nbsp;to the\u0026nbsp;writing and editing\u0026nbsp;of\u0026nbsp;the manuscript.\u0026nbsp;MR provided patient urine samples and information on\u0026nbsp;the\u0026nbsp;patients. CK worked on the machine learning algorithms and provided the random forest analysis. BK\u0026nbsp;performed\u0026nbsp;the qPCR analysis. JF provided sequencing data. JW interpreted patient data. TK analyzed\u0026nbsp;the\u0026nbsp;correlations of patient data with subtypes of breast cancer. LN edited the manuscript and provided writing support. PF contributed to writing and editing the manuscript. ES provided funding and the original idea for the research\u0026nbsp;and\u0026nbsp;supported\u0026nbsp;the\u0026nbsp;execution and writing and editing\u0026nbsp;of the manuscript.\u0026nbsp;All\u0026nbsp;the\u0026nbsp;authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the Genomics Facility, a core facility of the Interdisciplinary Center for Clinical Research (IZKF) Aachen within the Faculty of Medicine at RWTH Aachen University. We thank Lothar Häberle for advice and support.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eNolan E, Lindeman GJ, Visvader JE. Deciphering breast cancer: from biology to the clinic. Cell. 2023;186(8):1708\u0026ndash;28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuu XQ, Lee K, Jun JK, Suh M, Jung KW, Choi KS. Effect of mammography screening on the long-term survival of breast cancer patients: results from the National Cancer Screening Program in Korea. Epidemiol Health. 2022;44:e2022094.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLowry KP, Callaway KA, Lee JM, Zhang F, Ross-Degnan D, Wharam JF, et al. Trends in Annual Surveillance Mammography Participation Among Breast Cancer Survivors From 2004 to 2016. J Natl Compr Canc Netw. 2022;20(4):379\u0026ndash;86. e9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKupec T, Bleilevens A, Klein B, Hansen T, Najjari L, Wittenborn J et al. Comparison of Serum and Urine as Sources of miRNA Markers for the Detection of Ovarian Cancer. Biomedicines. 2023;11(9).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuque G, Manterola C, Otzen T, Arias C, Palacios D, Mora M, et al. Cancer Biomarkers in Liquid Biopsy for Early Detection of Breast Cancer: A Systematic Review. Clin Med Insights Oncol. 2022;16:11795549221134831.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShiao MS, Chang JM, Lertkhachonsuk AA, Rermluk N, Jinawath N. Circulating Exosomal miRNAs as Biomarkers in Epithelial Ovarian Cancer. Biomedicines. 2021;9(10).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKupec T, Bleilevens A, Iborra S, Najjari L, Wittenborn J, Maurer J, Stickeler E. Stability of circulating microRNAs in serum. PLoS ONE. 2022;17(8):e0268958.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHulstaert E, Morlion A, Levanon K, Vandesompele J, Mestdagh P. Candidate RNA biomarkers in biofluids for early diagnosis of ovarian cancer: A systematic review. Gynecol Oncol. 2021;160(2):633\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eErbes T, Hirschfeld M, Rucker G, Jaeger M, Boas J, Iborra S, et al. Feasibility of urinary microRNA detection in breast cancer patients and its potential as an innovative noninvasive biomarker. BMC Cancer. 2015;15:193.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKumar S, Keerthana R, Pazhanimuthu A, Perumal P. Overexpression of circulating miRNA-21 and miRNA-146a in plasma samples of breast cancer patients. Indian J Biochem Biophys. 2013;50(3):210\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRama K, Bitla AR, Hulikal N, Yootla M, Yadagiri LA, Asha T et al. Assessment of serum microRNA-21 and miRNA-205 as diagnostic markers for stage I and II breast cancer in Indian population. Indian J Cancer. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang L, Xu Y, Jin X, Wang Z, Wu Y, Zhao D, et al. A circulating miRNA signature as a diagnostic biomarker for noninvasive early detection of breast cancer. Breast Cancer Res Treat. 2015;154(2):423\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao H, Gao A, Zhang Z, Tian R, Luo A, Li M, et al. Genetic analysis and preliminary function study of miR-423 in breast cancer. Tumor Biol. 2015;36(6):4763\u0026ndash;71.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoussami N, Marinovich ML. AI for mammography screening: enter evidence from prospective trials. Lancet Digit Health. 2023;5(10):e641\u0026ndash;e2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNg AY, Oberije CJG, Ambrozay E, Szabo E, Serfozo O, Karpati E, et al. Prospective implementation of AI-assisted screen reading to improve early detection of breast cancer. Nat Med. 2023;29(12):3044\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakamizawa J, Konishi H, Yanagisawa K, Tomida S, Osada H, Endoh H, et al. Reduced expression of the let-7 microRNAs in human lung cancers in association with shortened postoperative survival. Cancer Res. 2004;64(11):3753\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiong G, Liu C, Yang G, Feng M, Xu J, Zhao F, et al. Long noncoding RNA GSTM3TV2 upregulates LAT2 and OLR1 by competitively sponging let-7 to promote gemcitabine resistance in pancreatic cancer. J Hematol Oncol. 2019;12(1):97.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLangevin SM, Christensen BC. Let-7 microRNA-binding-site polymorphism in the 3'UTR of KRAS and colorectal cancer outcome: a systematic review and meta-analysis. Cancer Med. 2014;3(5):1385\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerdas E, Stawski R, Kaczka K, Zubrzycka M. Analysis of Let-7 Family miRNA in Plasma as Potential Predictive Biomarkers of Diagnosis for Papillary Thyroid Cancer. Diagnostics (Basel). 2020;10(3).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChiu SC, Chung HY, Cho DY, Chan TM, Liu MC, Huang HM, et al. Therapeutic potential of microRNA let-7: tumor suppression or impeding normal stemness. Cell Transpl. 2014;23(4\u0026ndash;5):459\u0026ndash;69.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa Y, Shen N, Wicha MS, Luo M. The Roles of the Let-7 Family of MicroRNAs in the Regulation of Cancer Stemness. Cells. 2021;10(9).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee E, Jung SY, Hwang HJ, Jung J. Patient-Level Cancer Prediction Models From a Nationwide Patient Cohort: Model Development and Validation. JMIR Med Inf. 2021;9(8):e29807.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang K, Liu C, Sha X, Yao S, Li Z, Yu Y, et al. Development and validation of a prediction model to predict major adverse cardiovascular events in elderly patients undergoing noncardiac surgery: A retrospective cohort study. Atherosclerosis. 2023;376:71\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAwad A, Bader-El-Den M, McNicholas J, Briggs J. Early hospital mortality prediction of intensive care unit patients using an ensemble learning approach. Int J Med Inf. 2017;108:185\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Breast cancer, miRNA sequencing, urine, luminal A, luminal B, HER2, TNBC, screening, patient classification","lastPublishedDoi":"10.21203/rs.3.rs-3993094/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3993094/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIntroduction\u003c/p\u003e\n\u003cp\u003eBreast cancer is the most common cancer in women, with one in eight women suffering from this disease in her lifetime. The implementation of centrally organized mammography screening for women between 50 and 69 years of age was a major step in the direction of early detection and lead to a significant improvement in cure rates. Within the screening program, women undergo a mammogram every two years with an implemented centralized quality-controlled review process. However, the participation rate reached only approximately 50% of the eligible women. In addition to several others, the technical aspects of mammography, including painful compression of the breast, are cited as a reason for not participating in this very important program. Therefore, focusing current research on less painful and less invasive techniques for the detection of breast cancer seems to be highly clinically useful. Liquid biopsies offer this option with distinct molecules or cells in line with the research. Blood-based tests of circulating tumor cells (CTCs), cell-free DNA (ctDNA) and cell-free miRNA have been performed in a variety of studies and tumor entities.\u003c/p\u003e\n\u003cp\u003eMethods\u003c/p\u003e\n\u003cp\u003eWe performed miRNA sequencing on 82 urine samples, 32 samples from breast cancer patients (9× luminal A, 8× luminal B, 9× triple-negative and 6× HER2) and 50 healthy control samples. Data were analyzed and interpreted using Random Forest analysis.\u003c/p\u003e\n\u003cp\u003eResults\u003c/p\u003e\n\u003cp\u003eWe identified a signature of 275 miRNAs that allows the detection of invasive breast cancer in urine from breast cancer patients. Furthermore, we identified distinct miRNA expression patterns for the major intrinsic subtypes of breast cancer, specifically luminal A, luminal B, HER2-enriched and triple-negative breast cancer.\u003c/p\u003e\n\u003cp\u003eConclusions\u003c/p\u003e\n\u003cp\u003eHere, we present the first approach for sequencing miRNAs in female urine to detect breast cancer and, subsequently, intrinsic subtype-specific miRNA patterns. This experimental approach specifically validates miRNA sequencing as a technique for breast cancer detection in urine samples and opens the door to a new, easy and painless procedure for regular breast cancer screening.\u003c/p\u003e","manuscriptTitle":"Machine learning approach identifies miRNA signatures for breast cancer detection and classification from patient urine samples","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-29 16:12:31","doi":"10.21203/rs.3.rs-3993094/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d284208a-53a0-46a2-bfd5-1926e983474f","owner":[],"postedDate":"February 29th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-05-20T18:33:55+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-29 16:12:31","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3993094","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3993094","identity":"rs-3993094","version":["v1"]},"buildId":"GqpaHPwrfC8PjnIFayRh5","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: preprint-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00