A
In this section, we consider the optimal subsample allocation problem in the
ODS design under a fixed total budget and a fixed underlying cohort. To fix
notation, suppose that the unit cost for observing ( T, δ,
Z C ) is $ C 1 and for exposures
is $ C 2 , respectively. Let $ B denote the
total budget at the disposal of investigators. Under the proposed ODS design,
(4.1) m C 1 + ( n 0 + n 1 + n 3 ) C 2 = B , which is equal to C 1 +
ρ V C 2 = B/m
with n = n 0 +
n 1 + n 3 and
ρ V = n/m . Due to the
fact that m , B , C 1 and
C 2 are fixed, the constraint (4.1) implies that
ρ V is fixed.
Suppose the primary hypothesis in the ODS study is to test whether
Z E is related to the failure time
T ˜ , i.e., we are interested in testing: (4.2) H 0 : β 0 = 0 V S H 1 : β 0 = d , where d is a non-zero constant vector and
β 0 is the regression coefficient of exposures
in model (2.1) . We would like to
clarify that the d in the hypothesis is not a tuning parameter,
rather it is the value specified in the alternative hypothesis. The choice of the
value for d is based on the scientific question of interest of the
biomedical investigators. Note the power of hypothesis test in (4.2) is a function of subsamples allocation
( n 0 , n 1 ,
n 3 ). Our proposed optimal power-based ODS design is
to find the optimal subsample allocation of n 0 ,
n 1 , n 3 such that the
power of test of β 0 in (4.2) is maximized under the overall cost
constraint (4.1) . This is also
equivalent to an optimal allocation of ρ 0 ,
ρ 1 , ρ 3 to
maximize the power function.
To implement the above idea of optimal criteria, we outline it as follows.
By the asymptotic normality of β ^ m in Theorem 1 ,
the rejection region of testing β 0 = 0 with the
Wald statistic and type I error α is W = { β ^ m : β ^ m ′ Σ − 1 ( θ ^ m ) p × p β ^ m ≥ χ α / 2 2 ( p ) } ∪ { β ^ m : β ^ m ′ Σ − 1 ( θ ^ m ) p × p β ^ m ≤ χ 1 − α / 2 2 ( p ) } , where χ α 2 ( p ) denotes the upper α
quantile of chi-squared distribution with p degrees of freedom,
A p × p is the
upper-left p × p submatrix of matrix
A , and Σ ( θ ) = 1 m [ Σ A − 1 ( θ ) Σ F ( θ ) Σ A − 1 ( θ ) ] + 1 − ρ 0 ρ V m ρ 0 ρ V [ Σ A − 1 ( θ ) Σ B ( θ ) Σ A − 1 ( θ ) ] + ( 1 − ρ 0 ρ V ) ( π 1 ( 1 − ρ 0 ρ V ) − ρ 1 ρ V ) m ρ 1 ρ V [ Σ A − 1 ( θ ) Σ C ( θ ) Σ A − 1 ( θ ) ] + ( 1 − ρ 0 ρ V ) ( π 3 ( 1 − ρ 0 ρ V ) − ρ 3 ρ V ) m ρ 3 ρ V [ Σ A − 1 ( θ ) Σ D ( θ ) Σ A − 1 ( θ ) ] + 1 − ρ 0 ρ V m ρ 0 ρ V [ Σ A − 1 ( θ ) Σ E ( θ ) Σ A − 1 ( θ ) ] , and Σ B ( θ ) = E [ ( 1 − δ 1 ) H 1 ( θ ) ⊗ 2 ] , Σ C ( θ ) = E [ δ 1 ζ 1 , 1 H 1 ( θ ) ⊗ 2 ] , Σ D ( θ ) = E [ δ 1 ζ 1 , 3 H 1 ( θ ) ⊗ 2 ] , Σ E ( θ ) = E [ δ 1 ( 1 − ζ 1 ) H 1 ( θ ) ⊗ 2 ] with H 1 ( θ ) = ∫ − ∞ ∞ s ( 0 ) ( θ ; t ) [ X 1 − e X ( θ ; t ) ] d M 1 ( θ ; t ) .
The power function for testing (4.2) is P o w e r = P r ( β ^ m ∈ W ∣ H 1 : β 0 = d ) , which is a function of ρ 0 ,
ρ 1 and ρ 3 .
We consider the following optimization problem to maximize the power function under
the constraint (4.1) . Let
Power ( ρ 0 ,
ρ 1 , ρ 3 )
denote the power function and deriving the Lagrange multiplier argument as follows:
(4.3) L m ( ρ 0 , ρ 1 , ρ 3 , λ ) = P o w e r ( ρ 0 , ρ 1 , ρ 3 ) − λ { 1 − ρ 0 − ρ 1 − ρ 3 } . It is usually assumed that ρ 1 =
ρ 3 in the ODS literature ( Zhou et al., 2007 ). Therefore, taking derivatives of
(4.3) with respect to
ρ 0 and λ will result
in two equations and the optimal sampling proportion ρ 0 * = arg max 0 < ρ 0 < 1 L m ( ρ 0 , ρ 1 , ρ 3 , λ ) can be obtained by the nonlinear optimization
algorithm. To design an ODS study with an optimal allocation in practice, one
approach is to conduct a pilot study first with a small simple random sample. Based
on the pilot study, θ can be estimated by
θ ^ S R S and Σ A ( θ ) , Σ F ( θ ) , Σ B ( θ ) , Σ C ( θ ) , Σ D ( θ ) and Σ E ( θ ) can be estimated by their respective sample mean
evaluated at θ ^ S R S . Based on these estimators, the optimal allocation
( ρ 0 , ρ 1 ,
ρ 3 ) can be obtained by solving the
optimization problem (4.3) .
Analysis
In this section, we analyze a data set from the Norwegian Mother and Child
Cohort Study (MoBa) by the proposed method ( Ding et
al., 2014 ). The MoBa study is an ongoing long-term prospective cohort
study of 110,000 pregnant Norwegian women and their children. The study is conducted
by the U.S. National Institute of Environmental Health Sciences in collaboration
with the Norwegian National Public Health Institute to study the causes of disease
among mothers and children. The study began to recruit pregnant women from 1999 to
2008 who were around 17 weeks of gestation. The women completed a questionnaire
about demographic, lifestyle factors, and medical and reproductive history at the
time of enrollment. Women were also asked if their pregnancy was planned and
reported the time to pregnancy (TTP). Our goal is to assess the relationship between
exposure PFOA, the concentration of which is measured by high performance liquid
chromatography/tandem mass spectrometry based on 150 μl of
plasma and women’s subfecundity which is defined as TTP
> 12 months. In this analysis, we focus on the subset of
the women who enrolled in the study during 2003 to 2004 and delivered a live born
child. The total number of such women is 23,437 and the number of cases (defined by
TTP > 12 months) is 1,330 in the entire cohort ( Magnus et al. 2006 ). The size of the SRS is
550. Eleven subjects were excluded because 9 subjects did not report planning their
pregnancy and 2 subjects inconsistently reported their time to pregnancy. Four
hundreds supplemental subjects were selected from women with TTP greater than 12
months.
There are a total of 31 cases in the subcohort (SRS) and the blood samples
were collected from above eligible women around gestational week 17. We consider the
following six variables as confounders: (I) pre-pregnancy body mass index (BMI);
(II) maternal age (MotherAge), which is a categorical variable and the values are 1,
2, 3, 4 denoting the age belongs to the intervals [35, ∞), [30, 35), [25, 30)
and [0, 25), respectively; (III) maternal education (MotherEdu), which is a
categorical variable (1: less than high school, 2: high school and other; 3: some
college; 4: 4+ Years of college); (VI) paternal age (FatherAge); (V) SexFreq, which
denotes the frequency of sexual intercourse 1 month before pregnancy and (VI)
maternal diseases endometriosis (Endo).
Because there are 10 subjects in SRS and 10 subjects in supplemental samples
who did not report their BMI, the sample size of SRS and supplemental samples for
the final fitted model are 529 and 390. The number of cases in SRS is 30. We also
conduct the analysis based only on SRS. All the results are presented in Table 4 .
From Table 4 , we observe that results
from both designs indicate that PFOA is significantly associated with TTP. The
proposed method for ODS is able to identify MotherAge, MotherEdu and Endo to be
significantly associated with TTP while the SRS design is not. BMI is not
significant from either design. The same data set was analyzed by Ding et al. (2014) under the framework of the Cox
proportional hazards model. Their results confirmed that all the variable are
significant. Our results show the variables BMI and SexFreq are not significant. In
general, the proposed method provides a more efficient estimate because it utilizes
more information from the available survival data and takes advantage of the nature
of the ODS design. We conducted a goodness-of-fit test for the accelerated failure
time model based on the real data, which was proposed by Novák (2013) . The corresponding
p -value is 0.469, which means we can not reject the assumption of
the accelerated failure time model.
Asymptotic
Let s (0) ( θ ;
t ), s (1) ( θ ;
t ) and
e X ( θ ; t ) be the
corresponding limits of S (0) ( θ ;
t ), S (1) ( θ ;
t ) and X ¯ ( θ ; t ) , respectively. By the strong law of large numbers
( Pollard, 1990 ), the above convergence is
almost surely. Define Λ i ( θ ; t ) = ∫ − ∞ t Y i ( θ ; u ) λ ( u ) d u , i = 1 , … , m , where λ (·) is the
baseline hazards function of the error terms and
M i ( θ ; t ) =
N i ( θ ;
t ) −
Λ i ( θ ;
t ). It is easy to show that
M i ( θ 0 ;
t ) is a martingale ( Fleming and
Harrington, 1991 ). In order to establish the asymptotic properties of the
proposed estimator, we assume the following regularity conditions hold:
The parameter space Θ containing
θ 0 is a compact subset of
R p + q .
The exposure Z E and covariates
Z C are bounded almost surely by a
nonrandom constant.
The density function f (·) of the random
error and its derivative f ′(·) are bounded on
real space R and ∫ { f ′ ( t ) f ( t ) } 2 f ( t ) d t < ∞ .
The distribution of C i is absolutely
continuous and has a bounded density
g i (·) on real space
R , for i = 1, …,
m .
The matrix Σ A ( θ 0 ) is nonsingular, where
Σ A ( θ ) is the limit of m − 1 ∂ U ¯ m , G ( θ ) / ∂ θ as m goes to infinity.
n/m → ρ V ,
n k /n →
ρ k in probability, for
k = 0, 1, 3 as m goes to infinity, and
0 < ρ V < 1, 0
< ρ k < 1, for
k = 0, 1 and 3.
Remarks: (C1)–(C5) are standard regularity conditions to ensure the
consistency and asymptotic normality of the solution to the estimating equations (2.2) . Condition (C6) is
needed to ensure the desired asymptotic convergence of the validation samples, SRS,
and supplemental failures.
Under Conditions (C1)–(C6),
θ ^ m
has strong consistency and
m ( θ ^ m − θ 0 )
converges in distribution to zero-mean normal with covariance
Σ A − 1 ( θ 0 ) ( Σ F ( θ 0 ) + Σ O ( θ 0 ) ) ( Σ A − 1 ( θ 0 ) ) ′ .
We defer the proof and the definition of Σ F ( θ 0 ) and Σ O ( θ 0 ) to the Appendix .
Concluding
In this article, we studied the inference and optimal study allocation
problem for the outcome-dependent sampling design under the accelerated failure time
model. We proposed weighted estimating equation to account for the biased sampling
mechanism. To overcome the computational difficulty with non-smooth functions, we
further adopted the induced smoothing procedure to the weighted Gehan estimating
equation that leads to continuously differentiable estimating equations. These
induced smoothed estimating equations can be solved by the standard numerical
methods. Simulation studies show the proposed method performed well.
To facilitate the practical use of the proposed design and method, we
proposed and investigated the formula for the optimal study size allocation derived
by maximizing the power of a significance test for the exposures. This is an
especially useful tool in aiding investigators to design a cost-effective study and
it is the first time to consider the power based optimal subsamples allocation in
the ODS allocation literature. We recommend using
ρ 0 to be around 0.6 with the cutpoint
a 1 being 0.15 − 0.30 quantile of failure time
and the cutpoint a 2 being 0.70 − 0.85 quantile of
failure time in practice.
We only used the validation data set through the weight method in this
article. There are much potentially useful information from the nonvalidation data
set that we did not use. In our future work, we will develop a two-stage
outcome-dependent sampling design for AFT model using more information from the
nonvalidation data set to enhance the efficiency of the studies.
In this article, we considered time-invariant covariates. We are currently
investigating methods for handling time-dependent covariates in the accelerated
failure time model under the ODS design by the kernel-smoothed profile likelihood
function method ( Zeng and Lin, 2007 ). The
results will be reported in a future communication. We used Bernoulli sampling for
the subsamples in the ODS design. It will be of interest to investigate the
performance of stratified sampling for the subcohort to enhance the study
efficiency. Studies along the above mentioned directions are currently under
way.
Simulation
In this section, we conduct simulation studies to investigate the finite
sample performances of the proposed method. The failure time is generated by the
following accelerated failure time model: (5.1) log T ˜ = β 0 Z E + γ 0 Z C + ϵ , where Z E follows the standard normal
distribution, Z C follows the Bernoulli distribution with
a success probability of 0.5 and error term follows the standard normal distribution
which will result in a log-normal distribution for the failure time
T ˜ . The parameters
β 0 and γ 0
are set to be 0 and 0.5, respectively. The censoring time is generated from the
mixed uniform distribution over [0, c 1 ] and
c 2 with mixed probability α ,
which is chosen to yield around 80% or 70% censoring rate. we adopt
c 1 = 0.2, c 2 = 6 and
α are set by 0.795 and 0.682 to generate the censoring
rate being 80% and 70%, respectively.
For the ODS design, we first select the SRS of size
n 0 , then partition the domain of failure time into
three stratum through the quantiles of the observed failure times under different
scenarios. We select the supplemental failures of size
n 1 and n 3 from the low
stratum and the high stratum, respectively. We choose two pairs of the cutpoints
(0.30 and 0.70 quantiles, and 0.15 and 0.85 quantiles) to investigate the impact of
different cutpoints for the proposed ODS design. We compare the proposed estimator
θ ^ O D S with two estimators: θ ^ S R S and θ ^ G C C , which are the estimators based on the simple
random sampling design with the same sample size as the ODS design and the
generalized case–cohort design with the same subcohort as the ODS design and
the same size of supplemental failures that are randomly selected from the remaining
subjects outside of subcohort, respectively.
We investigate different scenarios for different sampling sizes, censoring
rate and cutpoints. For each configuration, we generate 1000 simulated data sets
with the number of total cohort being 4000. The sample standard deviation of the
1000 estimators is given in the corresponding column “SD”. The column
“SE” gives the average of the estimated standard error and
“CP” is the nominal 95% confidence interval coverage using the
estimated standard error. We use the regular bootstrap method to obtain the
estimated standard error with the number of bootstrap being 100 in this article. The
simulation results are given in Table 1 .
From Table 1 , we can make the
following observations. First, under all of the situations considered here, the
estimators θ ^ O D S , θ ^ S R S and θ ^ G C C are all approximately unbiased. Second, the average
of the standard error estimates is close to the sample standard deviation and the
confidence intervals attain coverage rate close to the nominal 95% level. Third, the
proposed estimator θ ^ O D S is more efficient than θ ^ S R S and θ ^ G C C which indicates that the proposed ODS design for
the survival analysis can be a more efficient alternative to the simple random
sampling design and generalized case–cohort design with the same sample size.
Finally, we note that the proposed estimator ( θ ^ O D S ) tends to be more efficient when the cutpoints are
further out ((0.15,0.85) vs. (0.30,0.70)) when there is a bigger allocation to
n 0 and the efficiency of the proposed estimator is
increasing when the censoring rate is decreasing.
We extend above simulation to evaluate the performance of the proposed
estimator θ ^ O D S under the different cutpoints and supplemental
samples allocation with the fixed validation sample size. In addition, consider
model (5.1) for generating the
failure time but the error term ϵ follows the standard extreme value
distribution which will result in an exponential distribution for the failure time.
We consider the situation that the censoring rate is 80% and 70%, and the cutpoints
are set to be (0.15,0.85) and (0.30,0.70) quantiles of failure time, respectively.
We generate 1000 simulated data sets with the number of total cohort being 6000 and
the number of ODS samples is 300. The results of simulation are summarized in Table 2 .
Simulation results from Table 2
confirm that the proposed estimator is approximately unbiased, the proposed variance
estimator perform well, and the confidence intervals attain coverage rate close to
the nominal 95% level. The results from Table
2 also show that the efficiency of the estimators depend on sample size
allocation, censoring rates, and cutpoints. For example, when the cutpoints are
(0.15, 0.85) quantiles of failure time, the efficiency of the proposed estimator
increases as the fraction of subcohort ( ρ 0 )
increases. However, when the cutpoints are (0.30,0.70), the proposed estimator is
more efficient under the situation where the fraction of subcohort is one-third.
Thus it is of interest to investigate how to best allocate resources. We will
consider an optimal power-based ODS design.
To compare the performance of the induced smoothing method and the standard
linear programming-based approach, we consider the following model to generate the
failure time: (5.2) log T ˜ = β 1 Z E , 1 + β 2 Z E , 2 + β 3 Z C , 1 + β 4 Z C , 2 + ϵ , where Z E, 1 follows the
standard normal distribution, Z E, 2
follows Bernoulli distribution with a success probability of 0.5,
Z C, 1 follows the uniform
distribution over [0, 1], Z C, 2 follows
Bernoulli distribution with a success probability of 0.7, error term follows the
standard normal distribution, and the parameters
( β 1 , β 2 ,
β 3 , β 4 )
are set to be (0, 0, 1, −0.5). The censoring time is generated from the mixed
uniform distribution over [0, c 1 ] and
c 2 with mixed probability α .
We adopt c 1 = 0.2, c 2 = 6
and α are set by 0.795 and 0.682 to generate the censoring
rate being 80% and 70%, respectively. The size of the underlying cohort is
m = 5000 and the total number of ODS samples is
n = 300 with n 0 = 200,
n 1 = n 3 = 50. The
results of simulation are presented in Table
3 .
Simulation results from Table 3 show
that the estimators by the induced smoothing method and the standard linear
programming-based approach have the similar bias, standard deviation, standard
deviation estimator and coverage rate. Hence, both the induced smoothing and linear
programming-based methods can be used to obtain the estimators.
With the same parameter settings as the above simulation, we consider a
hypothesis test with the significant level being α =
0.05, which is given as H 0 : β 1 = 0 and β 2 = 0 VS H 1 : β 1 = 0.3 and β 2 = 0.3. We investigate different scenarios for the censoring rate,
cutpoints and the fraction of the subcohort
ρ 0 ( ρ 0
= n 0 /n ), among which,
ρ 0 would be the most interested quality
for a practical study design in a given scenario. The size of the underlying
cohort is m = 5000 and the total number of ODS samples is
n = 300. For each configuration, we generate 1000 simulated
data sets and the simulation results are summarized in Figure 2 .
The results in Figure 2 demonstrate
that the various factors considered play a role in determining the optimal power
of the above test. In summary, when the censoring rate is 80% and cutpoints
being (0.15, 0.85) quantiles of the failure time, the maximized powers are 0.874
under ρ 0 = 0.733. When the censoring rate is
80% and cutpoints being (0.30,0.70) quantiles of the failure time, the maximized
power is 0.907 under ρ 0 = 0.467. When the
censoring rate is 70% and cutpoints being (0.15,0.85) quantiles of the failure
time, the maximized power is 0.897 under ρ 0 =
0.800. When the censoring rate is 70% and cutpoints being (0.30,0.70) quantiles
of the failure time, the maximized power is 0.932 under
ρ 0 = 0.467. Therefore, those results
suggest that for relatively high censoring rates, a good subcohort fraction
ρ 0 starting point is about 0.467 with
cutpoints being (0.30,0.70) quantiles of the failure time for a real study
design.
Introduction
Cost of exposure assessment is often a limiting factor for the size of the
study in many epidemiological and biomedical studies. As such, cost-effective study
designs are desirable by investigators with a limited budget. Those designs provide
investigators an alternative to the simple random sampling design for the same
overall cost but with bigger statistical power. Outcome-dependent sampling (ODS)
refers to such a cost-effective sampling scheme where one observes the exposure with
a probability that depends on the value of the outcome (e.g., Zhou et al., 2002 ; Weaver
and Zhou, 2005 ). The case–control study is the most well-known
outcome-dependent sampling design with the outcome being a binary variable (e.g.,
Prentice and Pyke, 1979 ; Breslow and Cain, 1988 ; Weinberg and Wacholder, 1993 ; Breslow and
Holubskov, 1997 ; Wang and Zhou,
2010 ). While a case–control study over-samples the cases, recently
developed ODS with a continuous outcome over-samples from the two tail regions of
the distribution of the response variable (e.g., Song, Zhou and Kosorok, 2009 ; Zhou, Qin
and Longnecker, 2011 ; Zhou et al.,
2014 ; Tan et al., 2016 ). Schildcrout et al. (2013) and ( 2015 ) studied the outcome-dependent sampling design with
the response being longitudinal. The principal idea of such ODS design is to
concentrate resources on the subjects who are believed to have more information
about the relationship between the exposure and the outcome ( Zhou et al., 2007 ).
The case–cohort design ( Prentice,
1986 ) is the most well-known outcome-dependent sampling design for the
time-to-event data especially when event is rare. In case–cohort design, the
expensive exposure is only observed for a subcohort randomly selected from the
underlying cohort and additional failures outside the subcohort. When the rate of
failures is not rare, an improved design where a subset of failures instead of all
the failures in case–cohort design is selected. This is referred to as the
generalized case–cohort design (e.g., Chen,
2001 ; Cai and Zeng, 2007 ; Kang and Cai, 2009 ; Kang, Cai and Chambless, 2013 ). Inspired by the fact that
the subjects who have high and low outcomes usually carry more information about the
exposure-response relationship than those in the middle ( Zhou et al., 2002 , 2007 ), recent research on ODS with failure time data has considered to
select the supplemental failures from those subset who failed early or late in order
to enhance the efficiency. For example, Ding et al.
(2014) presented the estimated maximum semiparametric empirical
likelihood estimation under the Cox proportional hazards model with
outcome-dependent sampling. Yu et al. (2016)
considered a weighted parital likelihood estimating equation for the Cox
proportional hazards model under outcome-dependent sampling and also studied the
optimal outcome-dependent sampling design by the asymptotic relative efficiency
between the proposed estimator and the estimator based on a simple random sampling
design.
In many cases, it may be more attractive to relate the failure time to
covariates directly as an alternative to the Cox proportional hazards model and
additive hazards model. The accelerated failure time (AFT) model linearly relates
the logarithm of the failure time to the covariates without specifying the error
distribution. The popular statistical methods for AFT model are based on the
rank-based estimating equation, which are derived from the weighted log-rank test
( Prentice, 1978 ). The asymptotic
properties of the rank-based estimator have been studied by many authors (e.g.,
Tsiatis, 1990 ; Ying, 1993 ; Jin et al.,
2003 ). Recently, the statistical methods for the case–cohort
design under the AFT model has been also studied in the literature (e.g., Nan, Yu and Kalbfleisch, 2006 ; Kong and Cai, 2009 ; Chiou,
Kang and Yan, 2014 ; Kim, Sit and Ying,
2016 ).
Our research is motivated by the need to evaluate the potential effects of
perfluorinated and polyfluorinated compounds (PFCs) on womens subfecundity in the
Norwegian Mother and Child Cohort Study. The PFCs are man-made chemicals and widely
used as industrial surfactant and emulsifiers ( Whitworth et al., 2012 ). The most widely studied PFCs are
perfluorooctane sulfonate (PFOS) and perfluorooctanoic acid (PFOA). The toxicity of
PFCs has been extensively studied in experimental animals. However, the relationship
of PFCs with health outcomes has been less attended. The aim of the present study is
to examine the impact of exposure PFCs on women’s subfecundity, measured by
the time to pregnancy (TTP). Two groups of women are chosen to measure the PFCs
levels: 550 women from the underlying study cohort by way of simple random sampling
and a supplemental sample of 400 women randomly sampled from those who delivered a
child and whose time to pregnancy (TTP) is greater than 12 months. Clearly, this
study design is an enhanced failure time ODS design where supplemental failures are
coming from the high end only (i.e., those TTP> 12).
In this article, we consider the statistical inference for the accelerated
failure time model under the outcome-dependent sampling design. The innovations of
the proposed approach can be summarized in the following three folds. First, we
adopt the weighted Gehan rank-based estimating equation to adjust the biased
sampling mechanism in ODS design, which is also the monotone estimating equation in
each component of parameters. Second, we take the induced smoothing approach to
smooth the weighted Gehan rank-based estimating equation, which leads to a
continuously differentiable estimating equation and can be solved by the standard
numerical methods. Third, it is the first time to consider the primary optimal
subsample allocation for ODS design through maximizing the power function for the
hypothesis test of the main exposure effect. Practical recommendations on using the
proposed subsamples allocation algorithm is also provided.
The rest of the article is organized as follows. In Section 2 , we introduce the accelerated failure time
model, the failure time outcome-dependent sampling scheme, and the smoothed weighted
Gehan estimating equation approach for the unknown regression parameters in the AFT
model. The asymptotic properties of the proposed estimator are established in Section 3 . In Section 4 , we discuss the optimal subsample allocation by maximizing the
power function of a significance test of the exposures in AFT model. In Section 5 , we present simulation studies to
evaluate the finite sample performances of the proposed method. Analysis of data
from the Norwegian Mother and Child Cohort Study is presented in Section 6 . We provide some concluding remarks in Section 7 . The proofs for theoretical results
are outlined in the Appendix .
Outcome Dependent
To fix notations, let T ˜ and C denote the potential
failure time and censoring time, respectively. Due to right censoring, we only
can observe T = min ( T ˜ , C ) and δ = I ( T ˜ ≤ C ) , where I (·) is the
indicator function for censoring. Let Z E =
( Z 1 , …,
Z p )′ be a p -dimensional
vector of primary exposures that are hard to obtain due to expense or difficult
to measure (e.g., bioassay or genetic analysis of blood specimens), and
Z C = ( G 1 ,
…, G q )′ be a
q -dimensional vector of covariates which are easy or cheap to
obtain (e.g., age, sex, race, smoking), respectively. Given variables
Z E and Z C ,
T ˜ and C are assumed independent.
We consider the following accelerated failure time model: (2.1) log ( T ˜ ) = β 0 ′ Z E + γ 0 ′ Z C + ϵ , where β 0 and
γ 0 are unknown regression parameters and
the dimensions of β 0 and
γ 0 are p and
q , respectively, and ϵ is the random error with an
unspecified distribution function F (·) and the density
function f (·).
Suppose there are m subjects in the underlying cohort
and ( T i , δ i , Z E,i ,
Z C,i , i = 1, …, m ) are the
independent copies of ( T, δ, Z E ,
Z C ). We propose the following outcome-dependent sampling
scheme with the outcome being the failure time subject to right censoring. This
is a retrospective design and the exposure ( Z E ) is
only observed for the selected subjects. First, we draw a simple random
subcohort (SRS) from the underlying cohort. Second, we partition the domain of
failure time T ˜ into a union of 3 mutually exclusive intervals,
A k =
[ a k −1 , a k ),
k = 1, 2, 3, where { a k :
k = 0, 1, 2, 3} are known constants satisfying:
a 0 ≡ 0 <
a 1
< a 2
< a 3 ≡ ∞. Then, the
supplemental failures are only selected from the intervals
A 1 and A 3 . The
sampling mechanism of ODS can be explained by the following figure:
We assume the sample size of SRS is n 0 and
let ξ i indicate, by the values 1 or 0,
whether or not the i -th subject is selected into SRS. Hence,
n 0 = ∑ i = 1 m ξ i . Let n 1 and
n 3 denote the sample size of the supplemental
failures from the intervals A 1 and
A 3 , respectively. Define
ζ i , k = I ( T ˜ i ∈ A k ) and let η i,k
(0–1 variable) denote whether or not the i -th subject
from the stratum A k and outside of SRS is selected
into supplemental failures, for k = 1, 3. Hence,
n 1 = ∑ i = 1 m η i , 1 and n 3 = ∑ i = 1 m η i , 3 . The samples with the exposures being
completely observed are referred to as validation sample, and the remaining
subjects are referred to as nonvalidation sample. Therefore, the observed data
structure for the proposed ODS design can be summarized as: Validation sample : { SRS : ( T i , δ i , Z E , i , Z C , i ) , i ∈ V 0 , Suppl: ( T j , δ j = 1 , Z E , j , Z C , j ) , j ∈ V k , k = 1 , 3 ;
Nonvalidation sample : ( T l , δ l , Z C , l ) , l ∈ V ¯ ; where V 0 ,
V 1 , V 3 and
V ¯ are the indexes for the SRS, supplemental
failures from the stratum A 1 ,
A 3 and the nonvalidation sample.
In order to simplify the notation, define X i = ( Z E , i ′ , Z C , i ′ ) ′ , i = 1 , … , m , θ 0 = ( β 0 ′ , γ 0 ′ ) ′ , θ =
( β ′, γ ′)′,
and the dimension of θ is
p + q . Define
e i ( θ ) =
log( T i )− θ ′ X i ,i
= 1, …, m . Let
N i ( θ ; t )
= I ( e i ( θ )
≤ t, δ i = 1) and
Y i ( θ ; t )
= I ( e i ( θ )
≥ t ) denote the counting process and at risk process,
respectively. When the data { T i , δ i ,
Z E,i , Z C,i , i =
1,…, m } are completely observed, the regression
parameters in the model (2.1) can
be estimated through the following rank-based estimating equation: (2.2) U m , ψ ( θ ) = ∑ i = 1 m ∫ − ∞ ∞ ψ ( θ ; t ) [ X i − X ¯ ( θ ; t ) ] d N i ( θ ; t ) = 0 , where ψ is a data-dependent weight
function and X ¯ ( θ ; t ) = S ( 1 ) ( θ ; t ) / S ( 0 ) ( θ ; t ) with S ( d ) ( θ ; t ) = m − 1 ∑ j = 1 m Y j ( θ ; t ) X j d for d = 0, 1. The weight
function ψ ( θ ;
t ) = 1 and ψ ( θ ;
t ) =
S (0) ( θ ;
t ) are corresponding to the log-rank statistics and Gehan
statistic, respectively. The asymptotic properties of the estimator from equation (2.2) have been well
studied by Tsiatis (1990) and Ying (1993) . In the ODS design, the
exposure Z E is not completely available for each
subject in the underlying cohort and the distribution of the supplemental
failures is different from that of the underlying population. Therefore, we
propose the following inverse probability weight to adjust the biased sampling
mechanism in the ODS design: W i = ξ i δ i ζ i + ξ i δ i ( 1 − ζ i ) ρ 0 ρ V + ξ i ( 1 − δ i ) ρ 0 ρ V + ( 1 − ξ i ) δ i ∑ k = { 1 , 3 } π ^ k ( 1 − ρ 0 ρ V ) ζ i , k η i , k ρ k ρ V , i = 1 , … , m , where ρ V ,
ρ 0 , ρ 1 ,
ρ 3 are the limits of
n/m , n 0 /n ,
n 1 /n ,
n 3 /n in probability with
n = n 0 +
n 1 + n 3 ,
ζ i =
ζ i, 1 +
ζ i, 3 , and
π ^ k being the empirical estimator of
π k =
Pr ( T ∈ A k ,
δ = 1) for k = 1, 3. The above weights have
the following properties: (i) the nonvalidation samples are excluded by setting
W = 0; (ii) the selected failures in the subcohort are
weighted by 1 if their failure time belong to A 1 or
A 3 , and by
( ρ 0 ρ V
) −1 otherwise; (iii) the selected censored samples have
the inverse of the sampling probability,
( ρ 0 ρ V
) −1 ; (iv) the selected supplemental failures are weighted
by π ^ 1 ( 1 − ρ 0 ρ V ) ( ρ 1 ρ V ) − 1 and π ^ 3 ( 1 − ρ 0 ρ V ) ( ρ 3 ρ V ) − 1 with their failure time belonging to
A 1 and A 3 ,
respectively.
In the ODS design, the regression coefficients
θ 0 in the model (2.1) can be estimated by solving the
following weighted estimating equations: U ˜ m , ψ ˜ ( θ ) = ∑ i = 1 m W i ∫ − ∞ ∞ ψ ˜ ( θ ; t ) [ X i − X ˜ ( θ ; t ) ] d N i ( θ ; t ) = 0 , where ψ ˜ is a possible data-dependent weight function
and X ˜ ( θ ; t ) = S ˜ ( 1 ) ( θ ; t ) / S ˜ ( 0 ) ( θ ; t ) with S ˜ ( d ) ( θ ; t ) = m − 1 ∑ j = 1 m W j Y j ( θ ; t ) X j d for d = 0, 1. In this article,
we adopt ψ ˜ ( θ ; t ) = S ˜ ( 0 ) ( θ ; t ) . Hence, the weighted Gehan estimating equations
can be written as: (2.3) U ˜ m , G ( θ ) = m − 1 ∑ i = 1 m ∑ j = 1 m δ i W i W j ( X i − X j ) I ( e i ( θ ) ≤ e j ( θ ) ) = 0 , which is monotone in each component of θ
( Fygenson and Ritov, 1994 ). Let
θ ˜ m denote the estimator by solving the estimating
equation (2.3) , which can be
obtained by minimizing the convex objective function with a standard linear
programming technique ( Jin et al.,
2003 ).
Although the standard linear programming technique can be used to solve
the estimating equations, another effective way is to use some suitable smooth
functions to approximate the non-smooth estimating function. In the section of
simulation study, we will show the inducing smoothing method and the standard
linear programming method have the similar performances. In this article, we
adopted the induced smoothing procedure to smooth the weighted Gehan estimating
equations that could lead to continuously differentiable estimating equations
and can be solved through the standard numerical algorithms ( Brown and Wang, 2007 ; Johnson and Strawderman, 2009 ; Chiou,
Kang and Yan, 2014 ). By the induced smoothing method, the estimating
equations (2.3) can be
replaced with U ¯ m , G ( θ ) = E V [ U ˜ m , G ( θ + m − 1 / 2 V ) ] = 0 , where the expectation is taken with respect to
V , which is a p +
q -dimensional independent standard normal random vector.
Therefore, the smoothed weighted Gehan estimating equations can be written as:
(2.4) U ¯ m , G ( θ ) = m − 1 ∑ i = 1 m ∑ j = 1 m δ i W i W j ( X i − X j ) Φ ( e j ( θ ) − e i ( θ ) r i j ) = 0 , where r i j 2 = m − 1 ( X j − X i ) ′ ( X j − X i ) and Φ(·) is the standard normal
distribution function. Let θ ^ m denote the estimator by solving equations (2.4) . In the following Section, we
will show the solution θ ^ m is a consistent estimator for
θ 0 and is also asymptotically normal.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.