Results
Clinical characteristics of the study populations are shown in Table 1 . The total study population comprised 678 subjects among whom 286 subjects were subsequently diagnosed with malignancy, and 392 subjects had benign findings. Overall, 290 (43%) of the subjects were premenopausal. Among the external validation set, 33% of the cancer cases were non-serous histologies, and 39% were Stage I or II disease.
Table 2 displays the miRNA chosen by forward regression from the training set panel of 180 miRNA. While we did perform significance testing on the training data to avoid spurious correlation, Table 2 provides the p -values for univariate comparisons calculated using the combined internal and external validation sets, as this a better indicator of true performance. After 10-fold validation on the training set, k = 10 miRNA offered the optimal AUC. For comparison, the proteins and metadata variables with their corresponding p- values are presented as well. There is no significant difference detected in age ( p = .35 ) and menopausal status ( p = .06 ) when comparing cancer and control patients. However, 6 of 7 proteins resulted in significant p -values, underscoring the relevance of these markers.
A feature correlation matrix of the input variables used in the model shows their linear relationships ( Figure 1A ). The first 10 rows (and columns) of the correlation matrix correspond to the 10 miRNA chosen by forward regression, and the remaining rows correspond to proteins and metadata. The feature correlation matrix has an approximate block diagonal structure, i.e., many of the miRNA are inter-correlated (e.g., miR-19a-3p and miR-424–5p), and so are the proteins and metadata (e.g., age and FSH), but the miRNA and proteins do not exhibit a high degree of inter-correlation. Furthermore, many of the miRNA (e.g., miR-200c) and proteins (e.g., HE4) have high correlation to the cancer labels ( Figure 1B ). All correlation scores were calculated using only the training data. Thus, the miRNA and proteins have good synergy and are well-suited to train a linear classification model. The t-distributed stochastic neighbor embedding (t-SNE) plots of the validation set shown in Figure 2A - C further emphasize this and illustrate that the combined features offer the best linear separation (least overlap) between cases and controls.
We used ROC curves to compare ovarian cancer classification models trained on miRNA data, protein with metadata, or a combination of all data types ( Figure 3A - B ; Table 3 ). The combined model offered the best AUC on the internal (AUC=0.9; 95% CI 0.81–0.95) and external validation sets (AUC=0.95; 95% CI 0.90–0.98). The combined model also offered the highest overall sensitivity (82% internal; 92% external), and specificity (91% internal; 80% external). The specificity score on the external set (80%) was tied with the protein + metadata model, so the combined model is joint best in terms of specificity in this case. The confusion matrices for each data type considered (e.g., using proteins + metadata) on both the internal and external validation sets appear in the supplementary data ( Figures S1 and S2 ).
We performed subset analyses looking at the performance of the models by stage and histology ( Table 3 ). Detailed histology and stage data were not available for a large portion of the internal validation set, so we restricted our subset analyses to the external set. The combined model had an overall sensitivity of 80% for early-stage ovarian cancers, and 100% for late-stage ovarian cancers. By histology the sensitivity was 88% for non-serous cancers and 94% for serous cancers. For ovarian cancers that were both early-stage and serous, the combined model had a sensitivity of 78%.
Materials
Serum samples for model training (n=468) were provided by Aspira Women’s Health (Turnbull, CT) using archival samples collected from prior prospective studies validating the MIA and MIA2G assays using WIRB protocol numbers OVA1–001-CO1, OVA2–001-CO3, and PS110001(the last of these being unpublished but identical in design as the others, initiated as a 522 order from the FDA) [ 10 , 13 ]. Once the model was designed, Aspira provided an independent sample set for internal validation (n=100). The investigators were blinded to the diagnoses among the samples sent for independent internal validation. An external validation set (n=110) comprised archival serum samples collected through the Pelvic Mass Protocol (Brigham and Women’s Hospital (BWH) Institutional Review Board Protocol 2000-P-001678) and the New England Case Control Study (Dana-Farber Cancer Institute Institutional Review Board Protocol 05–060) [ 18 ]. Samples from Aspira Women’s Health were sent to Brigham and Women’s Hospital for miRNA analysis, and samples from Brigham and Women’s Hospital were sent to Aspira Women’s Health for protein analysis. All study subjects provided written informed consent using locally approved IRB documents. Full study population statistics are summarized in Table 1 .
Serum samples were collected fresh in 13 × 75 mm BD Vacutainer Plus Plastic Serum tubes (BD Life Sciences, Franklin Lakes, NJ) with spray-coated silica. Samples were allowed to clot 1 hour at room temperature before processing, then spun down by centrifugation at 1300 x g x 10 min, aliquoted into 1.5 ml vials and stored at – 80 °C. Serum samples were stored at – 80 °C until use, then brought to room temperature and gently vortexed prior to analysis.
The seven biomarker concentrations were determined on the Cobas 6000 using the Cobas assays (Roche Diagnostics, Indianapolis IN). Assays for apolipoprotein A-1, β2-microglobulin, transthyretin (prealbumin), and TRF are immunoturbidimetric assays; CA125-II, HE4, and FSH assays use electrochemiluminescent detection. All clinical specimens were run in singlicate, with at least 2 independent level controls per run. All data were determined in a Clinical Laboratory Improvements Amendments-certified laboratory (Aspira Labs, Austin TX). Notably, raw protein measurements were used for model construction, and this analysis utilized the individual biomarker serum concentration determinations as inputs, not calculated MIA or MIA2G scores.
All samples were analyzed in the BWH Gynecologic Oncology Laboratory using a custom panel of 180 miRNA probes produced by Abcam, Inc. using the Fireplex ® assay (Abcam, Inc, Cambridge, MA) [ 19 ]. The probe list appears in Supplemental Table 1 and was empirically optimized for detection of serum miRNA based on prior studies looking at normal ranges for miRNA expression using this assay among large, diverse populations [ 15 , 17 , 20 ]. The 180 miRNA probes were distributed across three panels for each sample and were all processed on one plate. A reference panel of three positive control miRNA probes provided by Abcam was included in all plates, and a miRNA-like target, X-Control, was included in the FirePlex Hybridization Buffer as an internal standard. Blank particles bearing no probe were included in each well to correct for background fluorescence, as well as two off-species controls, to test for non-specific signals. The miRNA probe list also included seven overlapping miRNA (hsa-let-7a-5p, hsa-let-7d-5p, hsa-mir-17–5p, hsa-20b-5p, hsa-mir-93–5p, hsa-mir-16–5p, and hsa-mir-122–5p) across all three panels for signal normalization. Each complete miRNA profile required 25 µl of human serum. Samples were processed 29 at a time using a liquid handling robot (Starlet, Hamilton Robotics, Franklin, MA). Each plate also included a pooled human serum control to serve as an interplate calibrator. The samples were scanned using Guava Easycyte 5HT flow cytometers (Luminex, Austin, TX). Data were recorded in .fcs files and analyzed using the Fireplex Analysis Workbench software. The raw data are available in a supplemental data file ( Data_File S1 ).
The data consisted of 7 proteins, 2 metadata variables (age and menopausal status), and the miRNA panel. To reduce the dimension of the miRNA data, we used a forward regression algorithm as follows:
Let X = x 1 , … , x m ∈ ℝ n × m denote a data matrix (e.g., of miRNA data) and Y ∈ 0 , 1 n , a binary vector of labels, where 0 indicates a control, and 1, a case. Then, the forward regression algorithm can be summarized as follows 29 :
Initialize S = ø and the number of features, k .
for every integer j ∈ 0 , k − 1
do
Let x i j be a variable maximizing R Y , S j ∪ i j 2 , and set S j + 1 = S j ∪ i j .
Output S k
In the algorithm above, R Y , S 2 denotes the squared multiple correlation, which is defined
R Y , S 2 = Var Y − E Y − Y ′ 2 Var Y ,
and Y ′ = ∑ j = 0 k − 1 v i j x i j , where v = v i 0 , … , v i k − 1 is defined
v = argmin v x i 0 , … , x i k − 1 v − Y 2 2 .
The output, S k , is a set of k indices identifying the optimal features within the index range 1 ≤ j ≤ m . The forward regression algorithm starts by selecting the feature, x j , most highly correlated to Y , and then adds more features in an iterative sense to maximize linear separation of the control and case group.
Once k < 180 miRNA features were chosen by forward regression, we combined the miRNA, protein, and metadata variables to form a new data matrix X ∈ ℝ n × k + 9 . X was formed simply by concatenating the selected k miRNA features, 7 proteins, and 2 metadata variables. Then, we trained a linear classifier to map X to the cancer labels, Y
P y = j x = e x T w j + b j ∑ i = 0 n c − 1 e x T w i + b i ,
where j ∈ 0 , 1 is the class label, x ∈ ℝ k + 9 is a patient sample (i.e., one row of X ), and the ( w j , b j ) are weights and biases to be trained. Here y denotes the class label assigned to x , and n c = 2 denotes the number of classes. The class with the highest probability P is then chosen for membership. A linear classification model is appropriate in this application as, by Figure 1 , many of the miRNA and protein features are linearly correlated to cancer, and there is little correlation between the miRNA and protein predictors.
To explain how these models work, linear classifiers have a simple geometric interpretation. That is, in binary classification (as is considered here), a linear classifier fits a plane, x ∈ ℝ k + 9 : f x = 0 , between the case and control group, where, in this case, the classification plane has defining equation f x = x T w 1 − w 0 + b 1 − b 0 . When f x = 0 , P y = 0 x = P y = 1 x = 0.5 , and the sample ( x ) is equally likely to be a case or control. When f x > 0 it is more likely that x is a case and more likely a control when f x < 0 .
In model training, the most highly correlated features contribute more to the weight magnitudes (i.e., the magnitude of the w j ) and vice-versa. When the miRNA are combined with proteins, they offer a clearer linear separation of cases and controls when compared to proteins alone, and this is reflected in the model performance (see also Figure 2 which illustrates this effect in 2-D using t-SNE plots).
Our data were split into three groups, one for training the classification model (sample size n = 468 , with 191 cancer patients and 277 controls), one internal validation set (44 cancers and 56 controls) separately analyzed but drawn from the same source (Aspira Women’s Health) as the training data, and one external validation set (51 cancers and 59 controls), which was collected independently of the training data and from a different population (BWH). All model fitting and hyperparameter selection was done using only the training data. Then, the internal and external validation sets were inputted directly into the trained model for evaluation. To choose k , i.e., the number of miRNA features, we used 10-fold cross-validation on the training data. Specifically, we chose the k ∈ 1 , … , 20 which maximized the Receiver Operator Characteristic Area Under the Curve 28 (ROC AUC) score after 10-fold validation.
Sensitivity and specificity were calculated by setting a cancer probability threshold, t ∈ 0 , 1 . That is, we classified a subject as having cancer if their predicted cancer probability was greater than t . We set t as the largest t which guarantees at least 80% sensitivity on the training data after 10-fold cross-validation. We chose t in this way to ensure high model sensitivity (i.e., at least 80%) without sacrificing too much specificity (i.e., so that t is not too small). The t values were chosen using purely the training ROC curves after 10-fold cross-validation and kept consistent on the internal and external validation sets. The t values used for each data set (e.g., proteins + metadata) are given in Table 3 .
Univariate p -value analyses were also conducted. For each miRNA selected by forward regression using the training data, and protein and metadata feature, we performed a two-tailed t-test based on the validation data samples to evaluate if the mean case and control values were significantly different on the held-out data. In the t-test parameters, we assumed the case and control groups had non-equal variance and sample size. Multiple hypothesis correction was not required, as the t-tests were calculated using the validation data, whereas the features were selected using the training data. Moreover, the number of selected features is also quite small. All statistical tests, machine learning model training, and feature selection were conducted in Matlab, version 2020b (The MathWorks Inc., Natick, MA).
Discussion
Ovarian cancer patients have better clinical outcomes when their initial surgical procedures are performed by a gynecologic oncologist [ 3 , 21 ]. However, large numbers of surgical procedures are still performed by general surgeons or general gynecologists, typically at low-volume centers [ 2 ]. This is particularly true for rural, marginalized, and underserved populations [ 22 – 24 ]. As more than 98% of gynecologic oncologists practice in urban areas, there remains a tremendous need for clinical decision-making tools to help guide referral of patients at highest risk for ovarian cancer to appropriate surgical care [ 7 ].
Currently, three assays are available to assist with clinical triage of women with an adnexal mass [ 25 ]. The risk of malignancy index (RMI) combines CA125, ultrasound findings, and menopausal status [ 26 ]. The risk of ovarian malignancy algorithm (ROMA) combines CA125 and human epididymis protein 4 (HE4) [ 27 ]. MIA and its related tests combine CA125 with additional protein biomarkers and menopausal status [ 10 ]. However, each of these tests remains heavily dependent on CA125 and patient age to drive clinical performance, and these tests have had comparatively lower success identifying early-stage and non-serous cancers. Moreover, the addition of more protein biomarkers has minimally improved the accuracy of these biomarkers, especially when patients have early-stage disease [ 28 ]. For this reason, new classes of circulating biomarkers, such as miRNA, have attracted interest for possible synergy with these approaches [ 14 ].
In this work, we examine the utility of combining serum miRNA with protein biomarkers. Our principal finding is that a relatively small set of ten miRNA, when combined with protein expression and metadata, provides a more accurate assessment of an adnexal mass compared to protein biomarkers or miRNA alone. In particular, this integrated model offers improved sensitivity for early-stage and non-serous ovarian cancers, while preserving specificity. The improvement in accuracy for non-serous cancers, where protein-based models typically underperform, is particularly notable, as the combined model offered 88% sensitivity for non-serous cancers, compared to 65% offered by the protein + metadata model.
The results highlight the importance of using advanced statistical techniques when analyzing miRNA. For example, 5 of the 10 listed miRNA have non-significant p -values ( p > .05 ) when univariate analyses are performed on the validation data. Nonetheless, these markers still provide useful information. For example, cel-miR-39–3p, one of the off-species controls, was chosen as part of the k = 10 miRNA set using forward regression. The corresponding p -value on the validation set (0.59) is quite high, which likely indicates a spurious relation was detected in the training data. This makes sense as cel-miR-39–3p is an off-species target and was included in our panel for normalization purposes. cel-miR-39–3p was likely also favored by the forward regression algorithm given its weak correlation to human miRNA (see Figure 1(A) ). The remaining 5 miRNA listed in Table 2(a) yield significant p -values ( p ≤ 0.05 ) on the validation data. Some of the miRNA listed in Table 2(a) are common to miRNA identified in the literature as being predictive of cancer. For example, we previously reported miR-320c and miR-200c-3p as part of an ovarian cancer signature [ 15 ].
The current study has strengths and weaknesses. Strengths of our study include the use of internal and external validation sets as well as the blinded analysis of the samples by the collaborating investigators. The AWH data were collected from multi-site studies, and thus have strong representation of the US population. We were also able to include a large number of early-stage samples, a broad range of histologies, and both premenopausal and postmenopausal subjects. Weaknesses in the study include the fact that the external validation set had a larger percentage of post-menopausal women than the training and internal validation sets. Given the importance of menopausal status in the model, this might inflate the accuracy, however the predictive value of menopausal status alone was minimal. This effect was likely attenuated because the external validation set also had a greater proportion of early-stage cases. As noted, the intended use of the model is for the evaluation of an adnexal mass to guide appropriate surgical triage to a general gynecologist or gynecologic oncologist. Applications of this model to an unselected, asymptomatic population for cancer screening would require additional data and hypothesis testing.
We also note that the informative set of miRNA selected for the triage model in this study only partially overlaps the set described in our prior reports or by other authors [ 14 – 17 ]. This stems from three factors. First, the earlier studies relied on a different platform, next-generation sequencing, which considers a larger panel of potential miRNA inputs and has greater sensitivity for miRNA with very low absolute expression levels in serum. The current study focused on miRNA more likely to be detected within a measurable range for both cases and controls using PCR-based techniques. Second, while our earlier works focused on models wherein miRNA alone would be used for cancer detection, the current work aimed to identify miRNA that would provide optimal complementarity to the protein biomarkers. This tended to lower the specificity of the miRNA-only model in favor of sensitivity. Finally, miRNA profiles may vary among different populations based on medical comorbidities and demographic factors. While we have examined miRNA profiles from various populations in prior studies, we cannot exclude inherent differences between this study population and other cohorts.
When comparing features for model building, the model trained on miRNA alone is highly sensitive (98%) but at the cost of specificity (25%). However, the sensitivity of the combined model (92%) was 10% higher than the protein + metadata model (82%), with equal specificity (80%). The improvement in sensitivity is not uniform across cancer stage and histology and is weighted towards early-stage and non-serous cancers, where protein-based models typically underperform. Thus, the inclusion of miRNA achieved our goal to improve the accuracy of a triage test for these groups. Notably, imaging features were not included in our model. The integration of imaging data into multianalyte testing is an active area of investigation in our group and will be an interesting area for future work.
In conclusion, clinical assays which can estimate the risk of ovarian cancer in an indeterminate pelvic mass remain an important tool in surgical triage. Expanding the diversity of circulating biomarkers to include miRNA increases the robustness of these assays and has the potential to ensure that more women receive optimal surgical management. These findings will be important to validate in new prospective clinical trials.
Introduction
Ovarian cancer remains a leading cause of gynecologic cancer death in the United States [ 1 ]. Therapy usually consists of surgical resection and platinum-based chemotherapy. Oncologic outcomes are better if the initial surgery is performed by a gynecologic oncologist in a high-volume center rather than with a general surgeon or general gynecologist [ 2 , 3 ]. However, approximately 1 in 6 women will develop a pelvic mass in her lifetime, and the overwhelming majority of these pelvic masses are benign and could be safely managed by general gynecologists [ 4 – 6 ]. As more than 57 million American women live in a county without a practicing gynecologic oncologist, limiting over-referral is an important consideration for nationwide resource utilization [ 7 ]. At the same time, triage algorithms must have sufficiently high sensitivity for malignancy to ensure that as many cancers as possible receive appropriate initial surgical management. While CA125 is commonly used to inform the risk of malignancy of an adnexal mass, this assay is not approved by the Food and Drug Administration (FDA) for this purpose and has insufficient sensitivity for non-serous ovarian cancers or early-stage ovarian cancers to be used alone as a diagnostic biomarker [ 8 ]. This is an increasing problem as the incidence of high-grade serous ovarian cancers is declining, while the incidence of non-serous, endometriosis-associated ovarian cancers is increasing [ 9 ].
The Multivariate Index Assay (MIA) (commercially sold as OVA1 ® , Aspira Women's Health Inc.) was designed as an aid to clinical decision making for women being evaluated for an adnexal mass [ 10 , 11 ]. MIA combines five serum proteins: apolipoprotein A-1 (APO-A1), transthyretin (TT), beta 2-microglobulin (B2M), transferrin (TRF), and Cancer Antigen 125-II (CA125). The test favors sensitivity and accurately refers more than 90% of ovarian cancers to a gynecologic oncologist, but approximately half of benign cases are inaccurately classified as potentially malignant. MIA2G (Overa ® Aspira Women's Health Inc.), a second-generation multivariate index assay, was developed to improve the specificity of MIA [ 12 ]. MIA2G combines the serum concentrations of APO-A1, TRF, CA125, Human Epididymis Protein 4 (HE4) and follicle-stimulating hormone (FSH). Both algorithms utilize a modified support vector machine unified maximum separability algorithm (UMSA-SVM) to define a linear solution for classification of cancer versus non-cancer. In an independent dataset, MIA2G as a reflex test for adnexal masses with an indeterminate MIA result improved the assay specificity to 73% while preserving sensitivity, a two-step assay called OVA1plus ® [ 13 ]. However, the test was still less accurate for pre-menopausal women, early-stage tumors, and non-serous histologies.
Serum microRNA (miRNA) have emerged as a new class of non-invasive biomarkers with potential for ovarian cancer diagnosis [ 14 – 16 ]. Among women presenting for surgery, the sensitivity of an artificial neural network model comprising fourteen miRNA was 90% [ 15 ]. The model outperformed CA125 alone, particularly among premenopausal women and non-serous histologies. In a subsequent study comparing miRNA expression patterns among healthy women, women with benign adnexal masses, or early-stage ovarian cancers, the model continued to perform well with 73% sensitivity at 91% specificity [ 17 ]. In this study, we investigated the potential complementarity of serum miRNA and serum protein biomarkers for triage of an adnexal mass and tested the hypothesis that a multimodal assay would improve the accuracy of diagnosing early-stage and non-serous ovarian cancers.