Real
Endometrios is a disorder with tissues that are normally inside the uterus growing outside. The NICHD Physician Reliability Study (PRS) examined accuracy of diagnosing endometriosis under settings of different amount of clinical information and by physician of different level of experiences. The PRS consisted of four successive settings and involved three groups of physicians. We focus on a subset of the PRS data in this analysis, in particular the first two settings and consider only scores of the regional expert physicians. Of note, setting 1 in PRS provides intrauteral digital images, while setting 2 adds surgeon notes on top of setting 1. This gives rise to the final subset of the data that contains information of 113 participants in setting 1 and 126 participants in setting 2. Detailed descriptions of the study design and some main findings can be found elsewhere. 15 , 17 The regional experts consist of four OB/GYN physicians. Following Hwang and Chen, 17 we use the average rASRM scores of them as the test score in our analysis.
An interesting feature of the PRS design is that clinical information is sequentially augmented, and can be treated as a prior belief that diagnostic accuracy in a later setting (say, setting 2) is higher than an early one (say, setting 1). Let ROC 1 and ROC 2 be the ROC curves of Settings 1 and 2, respectively. This prior belief can be formulated as ROC 1 ( t ) ≤ ROC 2 ( t ) for all t ∈ (0, 1). We fit the SO model to the PRS data with this prior belief as a constraint and compare the results with those from the NO model that is free of this a priori constraint. To obtain PVs, we estimate the CDFs of healthy test scores from the two settings using equation (7) . Then the PVs are calculated as
(11)
z j = 1 − ∑ k = 1 K π ^ k Φ y j 1 − θ ^ j k * τ ^ j , j = 1 , 2 ,
where π ^ k are common mixing weights, θ ^ j k * is the posterior mean from the k th cluster and j th setting for the healthy score distribution, and τ ^ j is the posterior variance from the j th setting for healthy distribution. The NO and SO models are then fit to the estimated PVs, and the corresponding posterior estimates of AUCs and ROC curves are reported in Figure 1 .
Results from Figure 1 suggest that, overall, the regional expert physicians are mostly accurate in diagnosing endometriosis in both settings 1 and 2, with the estimated AUCs from 0.793 to 0.826. While the regional experts are slightly better in setting 1 than 2 (0.826 vs. 0.802) according to the NO model, they are slightly worse in setting 1 than 2 (0.793 vs. 0.818) by the SO model. The reversal of direction in AUCs is also illustrated in the ROC curves. In the NO model, the two curves cross, with the ROC curve of setting 1 about the same as that of setting 2 at high specificity level, rising above afterwards until about 0.5 specificity, then decreasing below. In contrast, the ROC curve of setting 1 always stay below that of setting 2 in the SO model.
The NICHD Fetal Growth Studies (FGS) is an epidemiologic study that aimed to establish a standard for normal fetal growth and size for gestational age in the U.S. population and improve estimation accuracy of fetal weight. 51 The cohort had five ultrasound examination visits during pregnancy at pre-defined gestational weeks. Using Hadlock’s formula, 52 measurements (abdominal circumference, head circumference and femur length) from these ultrasound examinations were converted into the estimated fetal weight (EFW), a useful score for predicting abnormal birth weight categories, such as small-for-gestational age (SGA). SGA birth corresponds to the birth weight below the 10 th percentile at a given gestation, 53 and carries potential neonatal risks of respiratory distress syndrome, necrotizing enterocolitis, chronic lung disease, neurologic disease, and death from infection. 54 , 55 The FGS data consist of 2054 pregnancies with 191 SGA births. It has been previously demonstrated that an EFW from ultrasounds closer to delivery has better discriminatory capacity for the abnormal birth weights than one farther. 14 This information can be translated into a stochastic order constraint that the diagnostic accuracy of an EFW at a later visit (say visit 4) is higher than that at an early one (say visit 3). To that end, we fit both NO and SO models to the FGS data from visits 3 and 4 and examine the diagnostic parameter estimates of EFWs in discriminating SGA births from normal ones. PVs were estimated in a similar fashion as described in equation (11) for the PRS data ( Section 4.1 ).
The posterior estimates of AUCs and ROC curves are reported in Figure 2 . Unlike in the PRS analysis, NO and SO models applied to FGS data produce similar AUC estimates at both visits. This similarity is a reflection that the FGS data agree with the a priori constraint that visit 4 provides better discriminating capacity than visit 3. The ROC curve estimates also confirm this finding. Note that the two ROC curves are already ordered in the NO model. As expected, we observe some statistical efficiency gains when the constraint is considered, as reflected in the narrower credible intervals in the SO model. For example, both NO (0.726 and 0.818) and SO (0.728 and 0.816) models estimate similar AUCs for visits 3 and 4, respectively, and the SO model achieves 7% and 3% efficiency gains in AUC estimates at visits 3 and 4, respectively.
It is of interest to assess whether the diagnostic accuracy of EFW is associated with important covariates such as maternal body mass index (BMI). Maternal obesity might impede the performance of ultrasound measures, which in turn could affect the accuracy of EFW in predicting SGA. To obtain PVs, we estimate covariate-specific CDFs of the non-SGA EFWs at the two visits following procedures in Section 2.5 . In particular, the PVs are calculated as
z j = 1 − ∑ k = 1 K π ^ k Φ y j 1 − θ ^ j k * − γ ^ j 0 B M I τ ^ j , j = 1 , 2 ,
where γ ^ j 0 is estimated coefficient for BMI in setting j . We then fit both NO and SO models to the estimated PVs and report AUC and ROC estimates at BMI levels of 20, 25, 30, and 35 kg/m 2 in Figure 3 and Table 3 , respectively. In both steps (estimation of PV and modeling PV), we used linear function of BMI due to its better goodness-of-fit compared to models with quadratic structure in BMI and to models adjusting for heteroskedasticity.
As in the above analysis of FGS data without covariates, when BMI is considered, the consideration of the a priori constraint does not change the direction of AUCs and ROC curves between visits 3 and 4. However, it does lead to improved statistical efficiency in SO estimates at all the covariate levels, as is evident in the narrower credible intervals. For example, at BMI level of 20 kg/m 2 , the estimated AUCs at visits 3 and 4 under SO model have 20% and 16% narrower CIs than those under NO model. A similar pattern is observed at other levels of BMI. It is interesting to note that the diagnostic capacity of EFW increases with BMI, possibly because variability in ultrasound measurements decreases with BMI.
Methods
Let Y 0 and Y 1 represent continuous test scores from a healthy and a diseased population, respectively, with corresponding CDFs F 0 and F 1 . The ROC curve is a graphical relationship between the true positivity rate, or sensitivity, and the false positivity rate, or one minus specificity, at all possible cut-off points. Given data y 0 = y 1 0 , , y m 0 and y 1 = y 1 1 , , y n 1 from F 0 and F 1 , respectively, where m and n are the corresponding sample sizes, one can estimate F 0 and F 1 and subsequently estimate ROC using R O C t = 1 − F 1 F 0 − 1 1 − t , t ∈ 0 , 1 , where F −1 ( t ) is the inverse function of F ( t ).
Alternatively, DeLong et al. 23 suggests estimating F 0 and obtaining placement values z i = P Y 0 > y i 1 = 1 − F 0 y i 1 , i = 1, , n ; see also Pepe 37 and Pepe and Cai. 26 Let Z be the random variable for the PV. It is easy to show that the ROC curve is simply the CDF of Z . 26 This alternative way of estimating ROC curves provides a new perspective in considering a priori constraints, as we develop below.
Assume for now placement values z = ( z 1 , , z n ) are available; Section 2.4 will describe how to use Bayesian semiparametric methods to estimate PVs. We model the CDF F of z semiparametrically using DPM by specifying, for i = 1, , n ,
(1)
F z i ∣ μ i , σ 2 = ∫ K z i ; μ i , σ 2 d H μ i , H μ i ∼ 𝒟 α H 0 μ i ,
where H is a mixing probability measure, 𝒟 α H 0 denotes a Dirichlet process (DP) 38 ) with base measure H 0 and precision α , and kernel K ⋅ ; μ , σ 2 is a distribution function with parameters μ and σ 2 .
Following the constructive representation of a DP as a stick-breaking process 39 and using a Gaussian kernel and a Gaussian base measure, we can model z as
(2)
η z i ~ N μ i , σ 2 , μ i ~ H , H = ∑ k p k δ μ k * , p k = V k ∏ l < k 1 − V l , V k ~ B e t a ( 1 , α ) , μ k * ~ N μ 0 , σ 0 2 ,
where η is a pre-specified monotone increasing transformation (e.g. logit) that maps the placement value on (0, 1) onto the real line ℜ , N μ , σ 2 denotes a normal distribution with mean μ and variance σ 2 , Beta ( α 1 , α 2 ) denotes a beta distribution with shape parameters α 1 and α 2 , δ c is a point mass at c , μ 0 , and σ 0 2 are unknown hyperparameters, V k ’s are the “stick” lengths, p k ’s are the mixing weights, and k indexes clusters. When restricting the model dimension to finite so that k = 1, . . ., K < ∞, we obtain the truncated stick-breaking process. It has been shown that the truncated stick-breaking process provides accurate approximation when K is moderate, say K = 30. 40 A well-established strength of the above Dirichlet process mixture is its ability to model flexible distributions. This feature suits the modeling of placement values particularly well. It is our experience that ROC curve and AUC estimates are sensitive to tail behaviors of the PV distributions and that PVs after η transformation still exhibit skewness and heavy tails. 41 , 42 In (2) , we have chosen to mix with the location parameter only. As pointed out in the literature, 43 , 44 any density on the real line can be approximated using a Dirichlet process location mixture of normals.
The model in (2) can be fit straightforwardly with a Markov chain Monte Carlo (MCMC) algorithm. 45 With a sample p ^ 1 ( l ) , … , p ^ K l , μ ^ 1 * l , , μ ^ K * l , and σ ^ l at the l th MCMC iteration, the ROC curve can be estimated as
R O C ^ l t = ∑ k = 1 K p ^ k l Φ η t − μ ^ k * l σ ^ l , t ∈ 0 , 1 ,
where Φ(·) is the CDF of the standard normal distribution. With L iterations, the posterior mean of the ROC curve can be obtained as the corresponding averages. AUC estimates can be consequently estimated by the Trapezoidal rule. Variability measures (posterior standard deviations or credible intervals) of these estimates can be obtained easily through the MCMC sample as well.
We now consider J tests, whose ROC curves are believed to be ordered a priori as
(3)
R O C 1 t ≤ ⋯ ≤ R O C J t , for all t ∈ 0 , 1 .
Suppose for the j th test, healthy scores y j 0 = y j 1 0 , … , y j m j 0 and diseased scores y j 1 = y j 1 1 , … , y j n j 1 are available, where m j and n j are the corresponding sample sizes, respectively. Let the PVs of the J tests be Z 1 , . . ., Z J , where Z j = Z j 1 , … , Z j n j , j = 1, . . ., J . The a priori constraint in (3) implies P ( Z 1 k ≤ z ) ≤ . . . ≤ P ( Z Jk ≤ z ) for all z ∈ (0, 1), k . For this reason, we call (3) stochastically ordered ROC curves.
Similar to Section 2.2 , assume for now PVs z j = z j 1 , … , z j n j , j = 1, . . ., J are available. In the absence of any a priori order, the J ROC curves can be estimated through the model formulation
(4)
η z j i ~ N μ j i , σ j 2 , μ j i ~ H j , H j = ∑ k p k δ μ j k * , p k = V k ∏ l < k 1 − V l , V k ~ B e t a 1 , α ,
with the base measure specified as
(5)
μ j k * ~ N μ j 0 , σ j 0 2 ,
for j = 1, ..., J . Here, all notations are defined similarly as in (2) , and all have the j subscript to index test, except the mixing weights p 1 , . . ., p K . The shared mixing weights across the J tests induce dependence among the J ROC curves ROC 1 , . . ., ROC J , hence effectively account for the possible correlation in the observed test scores.
To incorporate the stochastic order constraint (3) , we require μ 1 k ∗ ≥ … ≥ μ J k ∗ at all k and σ 1 2 = … = σ J 2 = σ 2 . The former can be achieved by replacing (5) with
(6)
μ j k * = ∑ l = 1 J β l k I l ≤ j , β 1 k ~ N a 1 , b 1 2 , β j k ~ N − a j , b j 2 , j = 2 , … , J .
Here I ( c ) is an indicator function that takes value 1 if condition c is true and 0 otherwise, N_ ( a , b 2 ) is the normal distribution with mean a and variance b 2 truncated to be negative, and a j ’s and b j ’s are unknown hyperparameters. The specifications in (6) ensure the ordering of the μ j k ∗ ’s. For example, in the case of J = 2, μ 1 k ∗ = β 1 k and μ 2 k ∗ = β 1 k + β 2 k . In this case, a negative β 2 k ensures μ 1 k ∗ ≥ μ 2 k ∗ , which in turn ensures P ( Z 1 < t ) ≤ P ( Z 2 < t ), hence ROC 1 ( t ) ≤ ROC 2 ( t ).
The following summarizes the result. See also Kottas and Gelfand, 46 Dunson and Peddada, 47 and Hoff. 48
Result 2.1. Suppose η (·) is a known monotone increasing function from (0, 1) onto ℜ . Then model formulations (4) and (6) are sufficient for (3) .
Proof. A brief proof is provided in Section 1 of Supplemental Material , available online. □
We have so far assumed that placement values z j ’s are available and known. In reality, however, they have to be estimated by using test scores of both healthy (for estimating the F 0 ’s) and diseased (for computing PVs) populations. To be consistent with using DPM to model the placement values, we also use DPM to model the healthy test scores y j 0 in order to estimate the placement values z j . Specifically, for i = 1, . . ., m j , j = 1, . . ., J , we assume the following truncated stick breaking process
(7)
y j i 0 ~ N θ j i , τ j 2 , θ j i ~ Q j , Q j = ∑ k = 1 K π k δ θ j k * , π k = ν k ∏ l < k 1 − ν l , ν k ~ B e t a 1 , α 0 , θ j k * ~ N θ j 0 , λ j 2 , k = 1 , … , K ,
where Q j ’s are mixing probability measures, π k ’s are mixing weights, ν k ’s are stick lengths and θ
j 0 and λ j 2 are mean and variance of the normal base measures.
Again, the shared mixing weights π k ’s effectively account for the possible correlations between y j i 0 and y j ′ i 0 for j ≠ j ′ . Using an MCMC algorithm similar to that in Section 2.2 , we can obtain a sample from the posterior distribution of the model parameters. Given the l th sample from L 0 iterations θ ^ j 1 * l , … , θ ^ j K * l , π ^ 1 * l , … , π ^ K * l , and τ ^ j l , the PVs can be estimated as
z ^ j i l = 1 − F ^ j 0 l y j i 1 = 1 − ∑ k = 1 K π ^ k l Φ y j i 1 − θ ^ j k * l τ ^ j l ,
where F ^ j 0 y is the estimate of F j 0 y , the CDF of y j i 0 .
In practice, covariates can impact the prediction ability of the test. Let W 0 and W 1 be vectors of covariates for diseased and healthy populations, respectively. The covariate-specific ROC curve given a covariate value w is defined as
(8)
R O C t ∣ w = 1 − F F 0 − 1 1 − t ∣ W 0 = w ∣ W 1 = w ,
where F ( t | W = w ) is a CDF conditional on w . The covariate-specific AUC is similarly defined as
A U C ( w ) = ∫ 0 1 R O C t ∣ w d t .
Let z w be the covariate-specific placement value so that
z w = P Y 0 > y 1 ∣ w = 1 − F 0 y 1 ∣ w .
Then (8) can be re-written as
(9)
R O C ( t ∣ w ) = P z w < t ∣ w .
Accordingly, estimation of covariate-specific ROC curves can be implemented by incorporating w in both the PV estimation stage (Stage 1) and the ROC estimation stage (Stage 2). The former will produce z w and the latter ROC ( t | w ). Using DPM, Stage 1 is a simple extension of (7) , with the base measure replaced by θ j k * ~ N θ j 0 + γ 0 T W 0 , λ j 2 . Similarly, Stage 2 extends (4) and (6) with the new base measure μ j k * = ∑ l = 1 J β l k I l ≤ j + γ 1 T W 1 . As such, the stochastic order constraint (3) applies to ROC curves at a given covariate value w :
R O C 1 t ∣ w ≤ ⋯ ≤ R O C J t ∣ w , for all t ∈ 0 , 1 .
We use proper but vague priors for all model parameters. In particular, we use univariate normal distribution N (0, 1000) for regression coefficients θ
j 0 , a j , j = 1, . . ., J , multivariate normal distribution N ( 0 , 1000 I ) for γ 0 and γ 1 , where 0 and I are vector of zeros and identity matrix, respectively, of length or dimension conforming to W 0 and W 1 , Gamma distribution G (0.01, 0.01) for the reciprocals of τ j , σ j 2 , λ j 2 and b j 2 j = 1 , … , J , and G (1, 1) for α 0 and α . Here G ( α , β ) denotes the gamma distribution with shape α and rate β . We assumed that these parameters are all independent a priori .
We develop an MCMC algorithm to generate samples from the posterior distribution of the model parameters given the data. Both visual inspection of the trace plots and diagnostic tools 49 are used to ensure convergence of the MCMC chain. After convergence, we thin the iterations to obtain a sample of 5000 iterations to produce posterior means, standard deviations and 95% credible intervals. The algorithm is implemented in R. 50
Discussion
In this paper, we proposed a semiparametric approach to stochastically ordered ROC curves using placement values. Although stochastic orders have been considered in ROC analysis before, they were imposed on the distributions of the healthy and diseased test scores. Thus, the constraints were not directly reflected on the ROC curves. The use of placement values allows the assessment of the direct covariate impact on the ROC curves, which facilitates estimating diagnostic accuracy for a particular sub-population of interest. For instance, the obstetrics example in Section 4.2 assessed the capacity of EFW as a biomarker to predict SGA for populations with different levels of maternal BMI.
The objective of the proposed model was to impose stochastic order constraint on ROC curves given a priori belief. If the data do not agree with the constraint, the proposed model ensures that the estimated diagnostic parameters reflect the constraint. On the other hand, when the data agree with the belief, the proposed model estimates the diagnostic parameters with higher statistical efficiency than the un-constrained counterpart. Both our simulations and real data examples confirmed these. Additionally, the semiparametric approach provides a flexible framework to model the ROC curves, providing a robust method against unusual distributions in placement values.
The a priori constraint in this paper requires ROC 1 ( t ) ≤ . . . ≤ ROC J ( t ) for all t ∈ (0, 1). This is a strong assumption. A less stringent constraint is the partial AUC ordering: ∫ 0 s R O C 1 t d t ≤ … ≤ ∫ 0 s R O C J t d t for all s , t ∈ 0 , 1 . This constraint requires ordering in all partial AUCs but allows the overall ROC curves to cross. We are currently developing estimation procedures to incorporate the partial AUC order constraint in ROC analysis. It is also of interest to consider the a priori order constraint on covariate-adjusted ROC curves. 56 Despite its novel advantages, the proposed model has a limitation common to any 2-stage approach. In stage 2 of the proposed model, we used the placement value as data and disregarded the uncertainty associated with them when they were estimated in stage 1. This in turn under-estimates the variability of the diagnostic accuracy measures in stage 2 and produces coverage probabilities that are below the nominal level. One possible remedy is a brutal-force procedure that a) generates a sample (say, L 0 ) of PVs from their posterior distributions in the first stage, b) repeats the second stage (estimating AUCs) using each of the PV samples in a) to obtain a sample (say, L ) of AUCs, and c) quantifies the variabilities in the AUC estimates using the L 0 × L posterior samples. Clearly this approach is prohibitively expensive in computation. An alternative possibility is to consider Bayesian bootstrap similar to Inacio de Carvalho & Rodrìguez-Àlvarez 57 . However, it is not immediately clear how the strategy can be implemented in the framework of ordered ROC curves, and some new developments are needed to adequately solve the problem with reasonable computational costs. This is a future work in consideration.
Simulations
In this section, we present simulation studies to illustrate the performance of the proposed modeling framework. These simulations are carried out without covariate in Section 3.1 and with covariates in Section 3.2 .
We consider two tests (Test 1 and Test 2) with healthy ( Y 1 0 and Y 2 0 ) and diseased ( Y 1 1 and Y 2 1 ) scores from the following data generating mechanism
(10)
Y j 0 ~ N μ j 0 , σ j 0 2 Y j 1 ~ N μ j 1 , σ j 1 2 ,
where j = 1, 2 indexes the tests, μ j 0 , μ j 1 and σ j 0 , σ j 1 are pre-specified means and variances of healthy and diseased test scores distributions, respectively. Under this setup, PVs are computed as
z j = 1 − Φ y j 1 − μ j 0 σ j 0 ,
where y j 1 is a realization of Y j 1 .
By varying the true values of μ j 0 , σ j 0 , μ j 1 , and σ j 1 in ( 10 ), we create cases where the two ROC curves are either stochastically ordered ( ROC 1 ≤ ROC 2 ) or unordered. For each case, we allow different levels of separation between ROC curves (low, medium, high), and generate n 1 = n 2 = n PVs with n = 150, 250, 500. Table 1 of Supplemental Material displays these six scenarios and the associated values of μ ’s and σ ’s. We fit our proposed model (SO) and its unconstrained counterpart (NO) to the generated PV data. We create 1000 data replicates and report average posterior means, biases, and average posterior standard deviations of AUC estimates. We also report the posterior means, biases, and average posterior standard deviations of difference of AUCs from test 2 and test 1. We also plot posterior mean estimates of ROC curves from 200 randomly selected data replicates. The simulation results with n = 150 are tabulated in Table 1 , with the corresponding ROC curve estimates shown in Figure 1 of Supplemental Material , available online.
When the true ROC curves are ordered, NO and SO models produce AUC estimates that are similar in magnitude and of the same direction (Test 1 less accurate than Test 2; Table 1 ). However, SO results in lower SDs, suggesting improved statistical efficiency when ordering constraint is incorporated. For example, when ROC separation is “Low,” Test 1 has lower AUC estimates than Test 2 in both NO (0.638 vs. 0.661) and SO (0.630 vs. 0.670) models, but, SO produces smaller SDs in both tests, resulting in efficiency gains of around 9% and 14% in Tests 1 and 2, respectively. The efficiency gain in estimation of the differences of test 2 and test 1 AUCs are also visible in the last column of the Table 1 . On the other hand, when the true ROC curves are unordered, the proposed SO model enforces the a priori constraint so that the estimates conform to the prior belief. For example, when the ROC separation is “High,” NO model produces higher AUC estimates in Test 1 than 2 (0.942 vs. 0.924) while SO reverses the direction (0.918 vs. 0.945). The last row of Figure 1 of Supplemental Material illustrates this reversal phenomenon clearly.
Simulation results with sample sizes n = 250 and 500 are provided in the Supplemental Material , in Tables 2 and 3 (AUCs) and Figures 2 and 3 (ROCs). Similar conclusions from the primary case ( n = 150) hold. We note that as sample size increases, the statistical efficiency gains under SO become more pronounced.
In this setting, we generate a covariate W from the uniform distribution on unit interval and generate diseased scores from
y j 1 ~ N 1 α j 1 α j 0 + β j W , 1 α j 1 2 ,
where j = 1, 2 indexes tests 1 and 2, respectively. Assuming that the healthy scores follow standard normal distribution, we generate the PVs as
z j = 1 − Φ y j 1 , j = 1 , 2.
Similar to Section 3.1 , we consider two ROC relations (stochastically ordered and unordered). Table 4 of Supplemental Material provides the values of α j 0 , α j 1 , and β
j used in this simulation study. In each relation, we consider covariate values (0.2, 0.5, and 0.8) to provide various levels of separation between the ROC curves. Similar to Section 3.1 , we simulate n PVs with n = 150, 250, 500. We generate 1000 data based on the above setup and for each of the 1000 data, we fit both NO and SO models. Table 2 of this paper and Figure 4 of Supplemental Material contain results for sample size n = 150.
Overall, the results in Table 2 and Figure 4 of Supplemental Material track those in Table 1 of this paper and Figure 1 of Supplemental Material . In particular, when the true ROCs are unordered, NO and SO produce similar estimates with elevated efficiency gain in the SO model. For example, when the covariate value is 0.2, Test 1 has lower AUC estimates than Test 2 for both NO (0.797 vs. 0.864) and SO (0.802 vs. 0.860) models. However, SO model produces AUCs with lower SDs as compared to its NO counterpart, resulting in efficiency gains of around 17% and 16% in Test 1 and Test 2, respectively. When the true ROCs are unordered, SO model produces estimates of AUCs and ROCs that satisfy the a priori constraint. For example, when covariate value is 0.8, although NO model produces higher AUC for Test 2 (0.935 vs. 0.945), the corresponding ROC estimates intersect. On the other hand, the SO model not only produces higher AUC estimate for Test 2 (0.918 vs. 0.926), it ensures that the corresponding ROC estimates are strictly stochastically ordered. The last row of Figure 4 of Supplemental Material illustrates this.
Results are similar for the other two sample sizes ( n = 250 and 500) and are reported in Tables 5 and 6 and Figures 5 and 6 of Supplemental Material online .
Introduction
The receiver operating characteristic (ROC) curve is a graphical tool to illustrate diagnostic ability of a continuous test in predicting a binary disease status, for example diseased and healthy. 1 A useful index of diagnostic accuracy is the area under the ROC curve (AUC) that can be interpreted as the probability that a randomly selected diseased subject has a higher test score than a randomly selected healthy one. 2 A large body of literature exists on ROC curves in diverse fields including engineering, 3 finance, 4 economics, 5 biomedical science, 6 , 7 environmental science 8 and machine learning. 9 Pepe 10 and Zhou et al. 11 provide comprehensive reviews of ROC analysis and its applications.
In diagnostic studies, a priori information are often available. In obstetrics, estimated fetal weight (EFW), a measure of fetal size derived from ultrasound examinations during pregnancy, is usually used to predict certain birth weight outcomes such as small-for-gestational age (SGA). 12 , 13 Prior literature has shown that an EFW from an examination late in pregnancy (e.g. in third trimester) should have higher discriminatory capacity than one earlier (e.g. in second trimester). 14 Such an a priori belief can be formulated as a constraint that the ROC curve of the late EFW should dominate that of the early one. As another example, in the NICHD Physician Reliability Study, 15 clinical information was successively made available to physicians to diagnose endometriosis, a women’s disorder in which endometrium tissue grows outside of the uterus, using the revised American Society for Reproductive Medicine (rASRM) score. 16 It is reasonable to believe that a setting with more clinical information should render higher diagnostic accuracy than one with less. 17 It is well established that, when such a priori information are available, it is beneficial to incorporate them in the modeling process as they usually lead to improved statistical efficiency in estimates. 18 The consideration of a priori constraints in ROC curve modeling is not new. Hanson et al. 19 developed a nonparametric Bayesian approach to ROC analysis to impose the stochastic order constraint between healthy and diseased test score distributions. Kottas 20 considered stochastic precedence, a constraint that is less restrictive than the stochastic order. Both Hanson et al. 19 and Kottas 20 were concerned with a single test (hence a single ROC curve) and used the constraints on the relationship between healthy and diseased test score distributions. To consider constraints in multiple tests (hence more than one ROC curves), Hwang and Chen 17 proposed an integrated method to impose both stochastic and variability orders 21 , 22 with the former applied to the relationship between healthy and diseased test score distributions within each test and the latter to the relationship between the multiple test score distributions within either the healthy or diseased population. While useful, these existing approaches formulated a priori constraints between test score distributions, and hence are not desirable when constraints are between ROC curves. In the aforementioned obstetrics study, the a priori belief concerns the relationship between ROC curves of an early and a late ultrasound EFW, not the relationship between EFW distributions of SGA and non-SGA cohort within the early or late ultrasound examination. Similarly, the a priori constraint in the Physician Reliability Study is formulated as the relationship between two settings with different clinical information, not as the relationship between rASRM score distributions of women with and without endometriosis within each setting. Clearly, it is desirable to have a framework where a priori constraints can be considered directly on the ROC curves.
In this paper, we exploit the idea of placement value (PV). 23 Briefly, PV is a standardization of the diseased test score with respect to the healthy test score distribution. A nice and important feature of the PV-based ROC analysis is that the ROC curve of a test is simply the cumulative distribution function (CDF) of the PV random variable associated with the test. As such, two tests with ordered ROC curves can be viewed as two tests with their PV random variables stochastically ordered. With this novel connection, the complicated and indirect process of imposing constraints on ROC curves through test score distributions, as implemented in Hwang and Chen, 17 can be replaced by working with stochastically ordered random variables, a task that is direct and straightforward. PV-based ROC analytical approaches have been considered by various authors. 24 – 28 None of them, however, considered ordered ROC curves.
The adoption of the PV framework for modeling ordered ROC curves also makes it possible to include covariates and assess their effects on ROC curves. In this article, we take a Bayesian semiparametric approach to ROC analysis using Dirichlet process mixture (DPM) models 29 to provide flexibility in the estimations. DPM has been proven useful to estimate distribution functions 30 – 32 and has been extended to adjust for covariates. 33 – 36 The rest of the article is organized as follows. Section 2 provides the detailed development of the proposed method on stochastically ordered ROC curves with and without covariates using placement values. We demonstrate the performance of the developed methodology through simulations in Section 3 . In Section 4 , we illustrate the utility of the proposed framework using both an EFW dataset in obstetrics and a dataset from the Physician Reliability Study. We conclude with a brief discussion in Section 5 .
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.