Accelerated failure time model for data from outcome-dependent sampling.

OA: closed
AI-generated summary by qwen3.7-flash, 2026-08-14

This paper develops a smoothed weighted Gehan estimating equation for accelerated failure time models under outcome-dependent sampling, demonstrating superior efficiency and optimal subsample allocation through simulations and an analysis of perfluoroalkyl substances exposure.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by qwen3.7-flash, 2026-08-14 · read from full text

This methodological paper develops a statistical framework for analyzing accelerated failure time models using outcome-dependent sampling designs to improve efficiency in exposure assessment. The authors propose a weighted Gehan rank-based estimating equation with induced smoothing and determine optimal subsample allocation strategies by maximizing the power of hypothesis tests for exposure effects. These methods are applied to data from the Norwegian Mother and Child Cohort Study to evaluate the impact of perfluorinated compounds on women’s subfecundity, measured as time to pregnancy. Relevance to endometriosis: listed as one indication for GnRH antagonists, though the paper's main focus is uterine fibroids.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Outcome-dependent sampling designs such as the case-control or case-cohort design are widely used in epidemiological studies for their outstanding cost-effectiveness. In this article, we propose and develop a smoothed weighted Gehan estimating equation approach for inference in an accelerated failure time model under a general failure time outcome-dependent sampling scheme. The proposed estimating equation is continuously differentiable and can be solved by the standard numerical methods. In addition to developing asymptotic properties of the proposed estimator, we also propose and investigate a new optimal power-based subsamples allocation criteria in the proposed design by maximizing the power function of a significant test. Simulation results show that the proposed estimator is more efficient than other existing competing estimators and the optimal power-based subsamples allocation will provide an ODS design that yield improved power for the test of exposure effect. We illustrate the proposed method with a data set from the Norwegian Mother and Child Cohort Study to evaluate the relationship between exposure to perfluoroalkyl substances and women's subfecundity.
Full text 34,890 characters · extracted from pmc-nxml · 7 sections · click to expand

A

In this section, we consider the optimal subsample allocation problem in the ODS design under a fixed total budget and a fixed underlying cohort. To fix notation, suppose that the unit cost for observing ( T, δ, Z C ) is $ C 1 and for exposures is $ C 2 , respectively. Let $ B denote the total budget at the disposal of investigators. Under the proposed ODS design, (4.1) m C 1 + ( n 0 + n 1 + n 3 ) C 2 = B , which is equal to C 1 + ρ V C 2 = B/m with n = n 0 + n 1 + n 3 and ρ V = n/m . Due to the fact that m , B , C 1 and C 2 are fixed, the constraint (4.1) implies that ρ V is fixed. Suppose the primary hypothesis in the ODS study is to test whether Z E is related to the failure time T ˜ , i.e., we are interested in testing: (4.2) H 0 : β 0 = 0 V S H 1 : β 0 = d , where d is a non-zero constant vector and β 0 is the regression coefficient of exposures in model (2.1) . We would like to clarify that the d in the hypothesis is not a tuning parameter, rather it is the value specified in the alternative hypothesis. The choice of the value for d is based on the scientific question of interest of the biomedical investigators. Note the power of hypothesis test in (4.2) is a function of subsamples allocation ( n 0 , n 1 , n 3 ). Our proposed optimal power-based ODS design is to find the optimal subsample allocation of n 0 , n 1 , n 3 such that the power of test of β 0 in (4.2) is maximized under the overall cost constraint (4.1) . This is also equivalent to an optimal allocation of ρ 0 , ρ 1 , ρ 3 to maximize the power function. To implement the above idea of optimal criteria, we outline it as follows. By the asymptotic normality of β ^ m in Theorem 1 , the rejection region of testing β 0 = 0 with the Wald statistic and type I error α is W = { β ^ m : β ^ m ′ Σ − 1 ( θ ^ m ) p × p β ^ m ≥ χ α / 2 2 ( p ) } ∪ { β ^ m : β ^ m ′ Σ − 1 ( θ ^ m ) p × p β ^ m ≤ χ 1 − α / 2 2 ( p ) } , where χ α 2 ( p ) denotes the upper α quantile of chi-squared distribution with p degrees of freedom, A p × p is the upper-left p × p submatrix of matrix A , and Σ ( θ ) = 1 m [ Σ A − 1 ( θ ) Σ F ( θ ) Σ A − 1 ( θ ) ] + 1 − ρ 0 ρ V m ρ 0 ρ V [ Σ A − 1 ( θ ) Σ B ( θ ) Σ A − 1 ( θ ) ] + ( 1 − ρ 0 ρ V ) ( π 1 ( 1 − ρ 0 ρ V ) − ρ 1 ρ V ) m ρ 1 ρ V [ Σ A − 1 ( θ ) Σ C ( θ ) Σ A − 1 ( θ ) ] + ( 1 − ρ 0 ρ V ) ( π 3 ( 1 − ρ 0 ρ V ) − ρ 3 ρ V ) m ρ 3 ρ V [ Σ A − 1 ( θ ) Σ D ( θ ) Σ A − 1 ( θ ) ] + 1 − ρ 0 ρ V m ρ 0 ρ V [ Σ A − 1 ( θ ) Σ E ( θ ) Σ A − 1 ( θ ) ] , and Σ B ( θ ) = E [ ( 1 − δ 1 ) H 1 ( θ ) ⊗ 2 ] , Σ C ( θ ) = E [ δ 1 ζ 1 , 1 H 1 ( θ ) ⊗ 2 ] , Σ D ( θ ) = E [ δ 1 ζ 1 , 3 H 1 ( θ ) ⊗ 2 ] , Σ E ( θ ) = E [ δ 1 ( 1 − ζ 1 ) H 1 ( θ ) ⊗ 2 ] with H 1 ( θ ) = ∫ − ∞ ∞ s ( 0 ) ( θ ; t ) [ X 1 − e X ( θ ; t ) ] d M 1 ( θ ; t ) . The power function for testing (4.2) is P o w e r = P r ( β ^ m ∈ W ∣ H 1 : β 0 = d ) , which is a function of ρ 0 , ρ 1 and ρ 3 . We consider the following optimization problem to maximize the power function under the constraint (4.1) . Let Power ( ρ 0 , ρ 1 , ρ 3 ) denote the power function and deriving the Lagrange multiplier argument as follows: (4.3) L m ( ρ 0 , ρ 1 , ρ 3 , λ ) = P o w e r ( ρ 0 , ρ 1 , ρ 3 ) − λ { 1 − ρ 0 − ρ 1 − ρ 3 } . It is usually assumed that ρ 1 = ρ 3 in the ODS literature ( Zhou et al., 2007 ). Therefore, taking derivatives of (4.3) with respect to ρ 0 and λ will result in two equations and the optimal sampling proportion ρ 0 * = arg max 0 < ρ 0 < 1 L m ( ρ 0 , ρ 1 , ρ 3 , λ ) can be obtained by the nonlinear optimization algorithm. To design an ODS study with an optimal allocation in practice, one approach is to conduct a pilot study first with a small simple random sample. Based on the pilot study, θ can be estimated by θ ^ S R S and Σ A ( θ ) , Σ F ( θ ) , Σ B ( θ ) , Σ C ( θ ) , Σ D ( θ ) and Σ E ( θ ) can be estimated by their respective sample mean evaluated at θ ^ S R S . Based on these estimators, the optimal allocation ( ρ 0 , ρ 1 , ρ 3 ) can be obtained by solving the optimization problem (4.3) .

Analysis

In this section, we analyze a data set from the Norwegian Mother and Child Cohort Study (MoBa) by the proposed method ( Ding et al., 2014 ). The MoBa study is an ongoing long-term prospective cohort study of 110,000 pregnant Norwegian women and their children. The study is conducted by the U.S. National Institute of Environmental Health Sciences in collaboration with the Norwegian National Public Health Institute to study the causes of disease among mothers and children. The study began to recruit pregnant women from 1999 to 2008 who were around 17 weeks of gestation. The women completed a questionnaire about demographic, lifestyle factors, and medical and reproductive history at the time of enrollment. Women were also asked if their pregnancy was planned and reported the time to pregnancy (TTP). Our goal is to assess the relationship between exposure PFOA, the concentration of which is measured by high performance liquid chromatography/tandem mass spectrometry based on 150 μl of plasma and women’s subfecundity which is defined as TTP > 12 months. In this analysis, we focus on the subset of the women who enrolled in the study during 2003 to 2004 and delivered a live born child. The total number of such women is 23,437 and the number of cases (defined by TTP > 12 months) is 1,330 in the entire cohort ( Magnus et al. 2006 ). The size of the SRS is 550. Eleven subjects were excluded because 9 subjects did not report planning their pregnancy and 2 subjects inconsistently reported their time to pregnancy. Four hundreds supplemental subjects were selected from women with TTP greater than 12 months. There are a total of 31 cases in the subcohort (SRS) and the blood samples were collected from above eligible women around gestational week 17. We consider the following six variables as confounders: (I) pre-pregnancy body mass index (BMI); (II) maternal age (MotherAge), which is a categorical variable and the values are 1, 2, 3, 4 denoting the age belongs to the intervals [35, ∞), [30, 35), [25, 30) and [0, 25), respectively; (III) maternal education (MotherEdu), which is a categorical variable (1: less than high school, 2: high school and other; 3: some college; 4: 4+ Years of college); (VI) paternal age (FatherAge); (V) SexFreq, which denotes the frequency of sexual intercourse 1 month before pregnancy and (VI) maternal diseases endometriosis (Endo). Because there are 10 subjects in SRS and 10 subjects in supplemental samples who did not report their BMI, the sample size of SRS and supplemental samples for the final fitted model are 529 and 390. The number of cases in SRS is 30. We also conduct the analysis based only on SRS. All the results are presented in Table 4 . From Table 4 , we observe that results from both designs indicate that PFOA is significantly associated with TTP. The proposed method for ODS is able to identify MotherAge, MotherEdu and Endo to be significantly associated with TTP while the SRS design is not. BMI is not significant from either design. The same data set was analyzed by Ding et al. (2014) under the framework of the Cox proportional hazards model. Their results confirmed that all the variable are significant. Our results show the variables BMI and SexFreq are not significant. In general, the proposed method provides a more efficient estimate because it utilizes more information from the available survival data and takes advantage of the nature of the ODS design. We conducted a goodness-of-fit test for the accelerated failure time model based on the real data, which was proposed by Novák (2013) . The corresponding p -value is 0.469, which means we can not reject the assumption of the accelerated failure time model.

Asymptotic

Let s (0) ( θ ; t ), s (1) ( θ ; t ) and e X ( θ ; t ) be the corresponding limits of S (0) ( θ ; t ), S (1) ( θ ; t ) and X ¯ ( θ ; t ) , respectively. By the strong law of large numbers ( Pollard, 1990 ), the above convergence is almost surely. Define Λ i ( θ ; t ) = ∫ − ∞ t Y i ( θ ; u ) λ ( u ) d u , i = 1 , … , m , where λ (·) is the baseline hazards function of the error terms and M i ( θ ; t ) = N i ( θ ; t ) − Λ i ( θ ; t ). It is easy to show that M i ( θ 0 ; t ) is a martingale ( Fleming and Harrington, 1991 ). In order to establish the asymptotic properties of the proposed estimator, we assume the following regularity conditions hold: The parameter space Θ containing θ 0 is a compact subset of R p + q . The exposure Z E and covariates Z C are bounded almost surely by a nonrandom constant. The density function f (·) of the random error and its derivative f ′(·) are bounded on real space R and ∫ { f ′ ( t ) f ( t ) } 2 f ( t ) d t < ∞ . The distribution of C i is absolutely continuous and has a bounded density g i (·) on real space R , for i = 1, …, m . The matrix Σ A ( θ 0 ) is nonsingular, where Σ A ( θ ) is the limit of m − 1 ∂ U ¯ m , G ( θ ) / ∂ θ as m goes to infinity. n/m → ρ V , n k /n → ρ k in probability, for k = 0, 1, 3 as m goes to infinity, and 0 < ρ V < 1, 0 < ρ k < 1, for k = 0, 1 and 3. Remarks: (C1)–(C5) are standard regularity conditions to ensure the consistency and asymptotic normality of the solution to the estimating equations (2.2) . Condition (C6) is needed to ensure the desired asymptotic convergence of the validation samples, SRS, and supplemental failures. Under Conditions (C1)–(C6), θ ^ m has strong consistency and m ( θ ^ m − θ 0 ) converges in distribution to zero-mean normal with covariance Σ A − 1 ( θ 0 ) ( Σ F ( θ 0 ) + Σ O ( θ 0 ) ) ( Σ A − 1 ( θ 0 ) ) ′ . We defer the proof and the definition of Σ F ( θ 0 ) and Σ O ( θ 0 ) to the Appendix .

Concluding

In this article, we studied the inference and optimal study allocation problem for the outcome-dependent sampling design under the accelerated failure time model. We proposed weighted estimating equation to account for the biased sampling mechanism. To overcome the computational difficulty with non-smooth functions, we further adopted the induced smoothing procedure to the weighted Gehan estimating equation that leads to continuously differentiable estimating equations. These induced smoothed estimating equations can be solved by the standard numerical methods. Simulation studies show the proposed method performed well. To facilitate the practical use of the proposed design and method, we proposed and investigated the formula for the optimal study size allocation derived by maximizing the power of a significance test for the exposures. This is an especially useful tool in aiding investigators to design a cost-effective study and it is the first time to consider the power based optimal subsamples allocation in the ODS allocation literature. We recommend using ρ 0 to be around 0.6 with the cutpoint a 1 being 0.15 − 0.30 quantile of failure time and the cutpoint a 2 being 0.70 − 0.85 quantile of failure time in practice. We only used the validation data set through the weight method in this article. There are much potentially useful information from the nonvalidation data set that we did not use. In our future work, we will develop a two-stage outcome-dependent sampling design for AFT model using more information from the nonvalidation data set to enhance the efficiency of the studies. In this article, we considered time-invariant covariates. We are currently investigating methods for handling time-dependent covariates in the accelerated failure time model under the ODS design by the kernel-smoothed profile likelihood function method ( Zeng and Lin, 2007 ). The results will be reported in a future communication. We used Bernoulli sampling for the subsamples in the ODS design. It will be of interest to investigate the performance of stratified sampling for the subcohort to enhance the study efficiency. Studies along the above mentioned directions are currently under way.

Simulation

In this section, we conduct simulation studies to investigate the finite sample performances of the proposed method. The failure time is generated by the following accelerated failure time model: (5.1) log T ˜ = β 0 Z E + γ 0 Z C + ϵ , where Z E follows the standard normal distribution, Z C follows the Bernoulli distribution with a success probability of 0.5 and error term follows the standard normal distribution which will result in a log-normal distribution for the failure time T ˜ . The parameters β 0 and γ 0 are set to be 0 and 0.5, respectively. The censoring time is generated from the mixed uniform distribution over [0, c 1 ] and c 2 with mixed probability α , which is chosen to yield around 80% or 70% censoring rate. we adopt c 1 = 0.2, c 2 = 6 and α are set by 0.795 and 0.682 to generate the censoring rate being 80% and 70%, respectively. For the ODS design, we first select the SRS of size n 0 , then partition the domain of failure time into three stratum through the quantiles of the observed failure times under different scenarios. We select the supplemental failures of size n 1 and n 3 from the low stratum and the high stratum, respectively. We choose two pairs of the cutpoints (0.30 and 0.70 quantiles, and 0.15 and 0.85 quantiles) to investigate the impact of different cutpoints for the proposed ODS design. We compare the proposed estimator θ ^ O D S with two estimators: θ ^ S R S and θ ^ G C C , which are the estimators based on the simple random sampling design with the same sample size as the ODS design and the generalized case–cohort design with the same subcohort as the ODS design and the same size of supplemental failures that are randomly selected from the remaining subjects outside of subcohort, respectively. We investigate different scenarios for different sampling sizes, censoring rate and cutpoints. For each configuration, we generate 1000 simulated data sets with the number of total cohort being 4000. The sample standard deviation of the 1000 estimators is given in the corresponding column “SD”. The column “SE” gives the average of the estimated standard error and “CP” is the nominal 95% confidence interval coverage using the estimated standard error. We use the regular bootstrap method to obtain the estimated standard error with the number of bootstrap being 100 in this article. The simulation results are given in Table 1 . From Table 1 , we can make the following observations. First, under all of the situations considered here, the estimators θ ^ O D S , θ ^ S R S and θ ^ G C C are all approximately unbiased. Second, the average of the standard error estimates is close to the sample standard deviation and the confidence intervals attain coverage rate close to the nominal 95% level. Third, the proposed estimator θ ^ O D S is more efficient than θ ^ S R S and θ ^ G C C which indicates that the proposed ODS design for the survival analysis can be a more efficient alternative to the simple random sampling design and generalized case–cohort design with the same sample size. Finally, we note that the proposed estimator ( θ ^ O D S ) tends to be more efficient when the cutpoints are further out ((0.15,0.85) vs. (0.30,0.70)) when there is a bigger allocation to n 0 and the efficiency of the proposed estimator is increasing when the censoring rate is decreasing. We extend above simulation to evaluate the performance of the proposed estimator θ ^ O D S under the different cutpoints and supplemental samples allocation with the fixed validation sample size. In addition, consider model (5.1) for generating the failure time but the error term ϵ follows the standard extreme value distribution which will result in an exponential distribution for the failure time. We consider the situation that the censoring rate is 80% and 70%, and the cutpoints are set to be (0.15,0.85) and (0.30,0.70) quantiles of failure time, respectively. We generate 1000 simulated data sets with the number of total cohort being 6000 and the number of ODS samples is 300. The results of simulation are summarized in Table 2 . Simulation results from Table 2 confirm that the proposed estimator is approximately unbiased, the proposed variance estimator perform well, and the confidence intervals attain coverage rate close to the nominal 95% level. The results from Table 2 also show that the efficiency of the estimators depend on sample size allocation, censoring rates, and cutpoints. For example, when the cutpoints are (0.15, 0.85) quantiles of failure time, the efficiency of the proposed estimator increases as the fraction of subcohort ( ρ 0 ) increases. However, when the cutpoints are (0.30,0.70), the proposed estimator is more efficient under the situation where the fraction of subcohort is one-third. Thus it is of interest to investigate how to best allocate resources. We will consider an optimal power-based ODS design. To compare the performance of the induced smoothing method and the standard linear programming-based approach, we consider the following model to generate the failure time: (5.2) log T ˜ = β 1 Z E , 1 + β 2 Z E , 2 + β 3 Z C , 1 + β 4 Z C , 2 + ϵ , where Z E, 1 follows the standard normal distribution, Z E, 2 follows Bernoulli distribution with a success probability of 0.5, Z C, 1 follows the uniform distribution over [0, 1], Z C, 2 follows Bernoulli distribution with a success probability of 0.7, error term follows the standard normal distribution, and the parameters ( β 1 , β 2 , β 3 , β 4 ) are set to be (0, 0, 1, −0.5). The censoring time is generated from the mixed uniform distribution over [0, c 1 ] and c 2 with mixed probability α . We adopt c 1 = 0.2, c 2 = 6 and α are set by 0.795 and 0.682 to generate the censoring rate being 80% and 70%, respectively. The size of the underlying cohort is m = 5000 and the total number of ODS samples is n = 300 with n 0 = 200, n 1 = n 3 = 50. The results of simulation are presented in Table 3 . Simulation results from Table 3 show that the estimators by the induced smoothing method and the standard linear programming-based approach have the similar bias, standard deviation, standard deviation estimator and coverage rate. Hence, both the induced smoothing and linear programming-based methods can be used to obtain the estimators. With the same parameter settings as the above simulation, we consider a hypothesis test with the significant level being α = 0.05, which is given as H 0 : β 1 = 0 and β 2 = 0 VS H 1 : β 1 = 0.3 and β 2 = 0.3. We investigate different scenarios for the censoring rate, cutpoints and the fraction of the subcohort ρ 0 ( ρ 0 = n 0 /n ), among which, ρ 0 would be the most interested quality for a practical study design in a given scenario. The size of the underlying cohort is m = 5000 and the total number of ODS samples is n = 300. For each configuration, we generate 1000 simulated data sets and the simulation results are summarized in Figure 2 . The results in Figure 2 demonstrate that the various factors considered play a role in determining the optimal power of the above test. In summary, when the censoring rate is 80% and cutpoints being (0.15, 0.85) quantiles of the failure time, the maximized powers are 0.874 under ρ 0 = 0.733. When the censoring rate is 80% and cutpoints being (0.30,0.70) quantiles of the failure time, the maximized power is 0.907 under ρ 0 = 0.467. When the censoring rate is 70% and cutpoints being (0.15,0.85) quantiles of the failure time, the maximized power is 0.897 under ρ 0 = 0.800. When the censoring rate is 70% and cutpoints being (0.30,0.70) quantiles of the failure time, the maximized power is 0.932 under ρ 0 = 0.467. Therefore, those results suggest that for relatively high censoring rates, a good subcohort fraction ρ 0 starting point is about 0.467 with cutpoints being (0.30,0.70) quantiles of the failure time for a real study design.

Introduction

Cost of exposure assessment is often a limiting factor for the size of the study in many epidemiological and biomedical studies. As such, cost-effective study designs are desirable by investigators with a limited budget. Those designs provide investigators an alternative to the simple random sampling design for the same overall cost but with bigger statistical power. Outcome-dependent sampling (ODS) refers to such a cost-effective sampling scheme where one observes the exposure with a probability that depends on the value of the outcome (e.g., Zhou et al., 2002 ; Weaver and Zhou, 2005 ). The case–control study is the most well-known outcome-dependent sampling design with the outcome being a binary variable (e.g., Prentice and Pyke, 1979 ; Breslow and Cain, 1988 ; Weinberg and Wacholder, 1993 ; Breslow and Holubskov, 1997 ; Wang and Zhou, 2010 ). While a case–control study over-samples the cases, recently developed ODS with a continuous outcome over-samples from the two tail regions of the distribution of the response variable (e.g., Song, Zhou and Kosorok, 2009 ; Zhou, Qin and Longnecker, 2011 ; Zhou et al., 2014 ; Tan et al., 2016 ). Schildcrout et al. (2013) and ( 2015 ) studied the outcome-dependent sampling design with the response being longitudinal. The principal idea of such ODS design is to concentrate resources on the subjects who are believed to have more information about the relationship between the exposure and the outcome ( Zhou et al., 2007 ). The case–cohort design ( Prentice, 1986 ) is the most well-known outcome-dependent sampling design for the time-to-event data especially when event is rare. In case–cohort design, the expensive exposure is only observed for a subcohort randomly selected from the underlying cohort and additional failures outside the subcohort. When the rate of failures is not rare, an improved design where a subset of failures instead of all the failures in case–cohort design is selected. This is referred to as the generalized case–cohort design (e.g., Chen, 2001 ; Cai and Zeng, 2007 ; Kang and Cai, 2009 ; Kang, Cai and Chambless, 2013 ). Inspired by the fact that the subjects who have high and low outcomes usually carry more information about the exposure-response relationship than those in the middle ( Zhou et al., 2002 , 2007 ), recent research on ODS with failure time data has considered to select the supplemental failures from those subset who failed early or late in order to enhance the efficiency. For example, Ding et al. (2014) presented the estimated maximum semiparametric empirical likelihood estimation under the Cox proportional hazards model with outcome-dependent sampling. Yu et al. (2016) considered a weighted parital likelihood estimating equation for the Cox proportional hazards model under outcome-dependent sampling and also studied the optimal outcome-dependent sampling design by the asymptotic relative efficiency between the proposed estimator and the estimator based on a simple random sampling design. In many cases, it may be more attractive to relate the failure time to covariates directly as an alternative to the Cox proportional hazards model and additive hazards model. The accelerated failure time (AFT) model linearly relates the logarithm of the failure time to the covariates without specifying the error distribution. The popular statistical methods for AFT model are based on the rank-based estimating equation, which are derived from the weighted log-rank test ( Prentice, 1978 ). The asymptotic properties of the rank-based estimator have been studied by many authors (e.g., Tsiatis, 1990 ; Ying, 1993 ; Jin et al., 2003 ). Recently, the statistical methods for the case–cohort design under the AFT model has been also studied in the literature (e.g., Nan, Yu and Kalbfleisch, 2006 ; Kong and Cai, 2009 ; Chiou, Kang and Yan, 2014 ; Kim, Sit and Ying, 2016 ). Our research is motivated by the need to evaluate the potential effects of perfluorinated and polyfluorinated compounds (PFCs) on womens subfecundity in the Norwegian Mother and Child Cohort Study. The PFCs are man-made chemicals and widely used as industrial surfactant and emulsifiers ( Whitworth et al., 2012 ). The most widely studied PFCs are perfluorooctane sulfonate (PFOS) and perfluorooctanoic acid (PFOA). The toxicity of PFCs has been extensively studied in experimental animals. However, the relationship of PFCs with health outcomes has been less attended. The aim of the present study is to examine the impact of exposure PFCs on women’s subfecundity, measured by the time to pregnancy (TTP). Two groups of women are chosen to measure the PFCs levels: 550 women from the underlying study cohort by way of simple random sampling and a supplemental sample of 400 women randomly sampled from those who delivered a child and whose time to pregnancy (TTP) is greater than 12 months. Clearly, this study design is an enhanced failure time ODS design where supplemental failures are coming from the high end only (i.e., those TTP> 12). In this article, we consider the statistical inference for the accelerated failure time model under the outcome-dependent sampling design. The innovations of the proposed approach can be summarized in the following three folds. First, we adopt the weighted Gehan rank-based estimating equation to adjust the biased sampling mechanism in ODS design, which is also the monotone estimating equation in each component of parameters. Second, we take the induced smoothing approach to smooth the weighted Gehan rank-based estimating equation, which leads to a continuously differentiable estimating equation and can be solved by the standard numerical methods. Third, it is the first time to consider the primary optimal subsample allocation for ODS design through maximizing the power function for the hypothesis test of the main exposure effect. Practical recommendations on using the proposed subsamples allocation algorithm is also provided. The rest of the article is organized as follows. In Section 2 , we introduce the accelerated failure time model, the failure time outcome-dependent sampling scheme, and the smoothed weighted Gehan estimating equation approach for the unknown regression parameters in the AFT model. The asymptotic properties of the proposed estimator are established in Section 3 . In Section 4 , we discuss the optimal subsample allocation by maximizing the power function of a significance test of the exposures in AFT model. In Section 5 , we present simulation studies to evaluate the finite sample performances of the proposed method. Analysis of data from the Norwegian Mother and Child Cohort Study is presented in Section 6 . We provide some concluding remarks in Section 7 . The proofs for theoretical results are outlined in the Appendix .

Outcome Dependent

To fix notations, let T ˜ and C denote the potential failure time and censoring time, respectively. Due to right censoring, we only can observe T = min ( T ˜ , C ) and δ = I ( T ˜ ≤ C ) , where I (·) is the indicator function for censoring. Let Z E = ( Z 1 , …, Z p )′ be a p -dimensional vector of primary exposures that are hard to obtain due to expense or difficult to measure (e.g., bioassay or genetic analysis of blood specimens), and Z C = ( G 1 , …, G q )′ be a q -dimensional vector of covariates which are easy or cheap to obtain (e.g., age, sex, race, smoking), respectively. Given variables Z E and Z C , T ˜ and C are assumed independent. We consider the following accelerated failure time model: (2.1) log ( T ˜ ) = β 0 ′ Z E + γ 0 ′ Z C + ϵ , where β 0 and γ 0 are unknown regression parameters and the dimensions of β 0 and γ 0 are p and q , respectively, and ϵ is the random error with an unspecified distribution function F (·) and the density function f (·). Suppose there are m subjects in the underlying cohort and ( T i , δ i , Z E,i , Z C,i , i = 1, …, m ) are the independent copies of ( T, δ, Z E , Z C ). We propose the following outcome-dependent sampling scheme with the outcome being the failure time subject to right censoring. This is a retrospective design and the exposure ( Z E ) is only observed for the selected subjects. First, we draw a simple random subcohort (SRS) from the underlying cohort. Second, we partition the domain of failure time T ˜ into a union of 3 mutually exclusive intervals, A k = [ a k −1 , a k ), k = 1, 2, 3, where { a k : k = 0, 1, 2, 3} are known constants satisfying: a 0 ≡ 0 < a 1 < a 2 < a 3 ≡ ∞. Then, the supplemental failures are only selected from the intervals A 1 and A 3 . The sampling mechanism of ODS can be explained by the following figure: We assume the sample size of SRS is n 0 and let ξ i indicate, by the values 1 or 0, whether or not the i -th subject is selected into SRS. Hence, n 0 = ∑ i = 1 m ξ i . Let n 1 and n 3 denote the sample size of the supplemental failures from the intervals A 1 and A 3 , respectively. Define ζ i , k = I ( T ˜ i ∈ A k ) and let η i,k (0–1 variable) denote whether or not the i -th subject from the stratum A k and outside of SRS is selected into supplemental failures, for k = 1, 3. Hence, n 1 = ∑ i = 1 m η i , 1 and n 3 = ∑ i = 1 m η i , 3 . The samples with the exposures being completely observed are referred to as validation sample, and the remaining subjects are referred to as nonvalidation sample. Therefore, the observed data structure for the proposed ODS design can be summarized as: Validation sample : { SRS : ( T i , δ i , Z E , i , Z C , i ) , i ∈ V 0 , Suppl: ( T j , δ j = 1 , Z E , j , Z C , j ) , j ∈ V k , k = 1 , 3 ; Nonvalidation sample : ( T l , δ l , Z C , l ) , l ∈ V ¯ ; where V 0 , V 1 , V 3 and V ¯ are the indexes for the SRS, supplemental failures from the stratum A 1 , A 3 and the nonvalidation sample. In order to simplify the notation, define X i = ( Z E , i ′ , Z C , i ′ ) ′ , i = 1 , … , m , θ 0 = ( β 0 ′ , γ 0 ′ ) ′ , θ = ( β ′, γ ′)′, and the dimension of θ is p + q . Define e i ( θ ) = log( T i )− θ ′ X i ,i = 1, …, m . Let N i ( θ ; t ) = I ( e i ( θ ) ≤ t, δ i = 1) and Y i ( θ ; t ) = I ( e i ( θ ) ≥ t ) denote the counting process and at risk process, respectively. When the data { T i , δ i , Z E,i , Z C,i , i = 1,…, m } are completely observed, the regression parameters in the model (2.1) can be estimated through the following rank-based estimating equation: (2.2) U m , ψ ( θ ) = ∑ i = 1 m ∫ − ∞ ∞ ψ ( θ ; t ) [ X i − X ¯ ( θ ; t ) ] d N i ( θ ; t ) = 0 , where ψ is a data-dependent weight function and X ¯ ( θ ; t ) = S ( 1 ) ( θ ; t ) / S ( 0 ) ( θ ; t ) with S ( d ) ( θ ; t ) = m − 1 ∑ j = 1 m Y j ( θ ; t ) X j d for d = 0, 1. The weight function ψ ( θ ; t ) = 1 and ψ ( θ ; t ) = S (0) ( θ ; t ) are corresponding to the log-rank statistics and Gehan statistic, respectively. The asymptotic properties of the estimator from equation (2.2) have been well studied by Tsiatis (1990) and Ying (1993) . In the ODS design, the exposure Z E is not completely available for each subject in the underlying cohort and the distribution of the supplemental failures is different from that of the underlying population. Therefore, we propose the following inverse probability weight to adjust the biased sampling mechanism in the ODS design: W i = ξ i δ i ζ i + ξ i δ i ( 1 − ζ i ) ρ 0 ρ V + ξ i ( 1 − δ i ) ρ 0 ρ V + ( 1 − ξ i ) δ i ∑ k = { 1 , 3 } π ^ k ( 1 − ρ 0 ρ V ) ζ i , k η i , k ρ k ρ V , i = 1 , … , m , where ρ V , ρ 0 , ρ 1 , ρ 3 are the limits of n/m , n 0 /n , n 1 /n , n 3 /n in probability with n = n 0 + n 1 + n 3 , ζ i = ζ i, 1 + ζ i, 3 , and π ^ k being the empirical estimator of π k = Pr ( T ∈ A k , δ = 1) for k = 1, 3. The above weights have the following properties: (i) the nonvalidation samples are excluded by setting W = 0; (ii) the selected failures in the subcohort are weighted by 1 if their failure time belong to A 1 or A 3 , and by ( ρ 0 ρ V ) −1 otherwise; (iii) the selected censored samples have the inverse of the sampling probability, ( ρ 0 ρ V ) −1 ; (iv) the selected supplemental failures are weighted by π ^ 1 ( 1 − ρ 0 ρ V ) ( ρ 1 ρ V ) − 1 and π ^ 3 ( 1 − ρ 0 ρ V ) ( ρ 3 ρ V ) − 1 with their failure time belonging to A 1 and A 3 , respectively. In the ODS design, the regression coefficients θ 0 in the model (2.1) can be estimated by solving the following weighted estimating equations: U ˜ m , ψ ˜ ( θ ) = ∑ i = 1 m W i ∫ − ∞ ∞ ψ ˜ ( θ ; t ) [ X i − X ˜ ( θ ; t ) ] d N i ( θ ; t ) = 0 , where ψ ˜ is a possible data-dependent weight function and X ˜ ( θ ; t ) = S ˜ ( 1 ) ( θ ; t ) / S ˜ ( 0 ) ( θ ; t ) with S ˜ ( d ) ( θ ; t ) = m − 1 ∑ j = 1 m W j Y j ( θ ; t ) X j d for d = 0, 1. In this article, we adopt ψ ˜ ( θ ; t ) = S ˜ ( 0 ) ( θ ; t ) . Hence, the weighted Gehan estimating equations can be written as: (2.3) U ˜ m , G ( θ ) = m − 1 ∑ i = 1 m ∑ j = 1 m δ i W i W j ( X i − X j ) I ( e i ( θ ) ≤ e j ( θ ) ) = 0 , which is monotone in each component of θ ( Fygenson and Ritov, 1994 ). Let θ ˜ m denote the estimator by solving the estimating equation (2.3) , which can be obtained by minimizing the convex objective function with a standard linear programming technique ( Jin et al., 2003 ). Although the standard linear programming technique can be used to solve the estimating equations, another effective way is to use some suitable smooth functions to approximate the non-smooth estimating function. In the section of simulation study, we will show the inducing smoothing method and the standard linear programming method have the similar performances. In this article, we adopted the induced smoothing procedure to smooth the weighted Gehan estimating equations that could lead to continuously differentiable estimating equations and can be solved through the standard numerical algorithms ( Brown and Wang, 2007 ; Johnson and Strawderman, 2009 ; Chiou, Kang and Yan, 2014 ). By the induced smoothing method, the estimating equations (2.3) can be replaced with U ¯ m , G ( θ ) = E V [ U ˜ m , G ( θ + m − 1 / 2 V ) ] = 0 , where the expectation is taken with respect to V , which is a p + q -dimensional independent standard normal random vector. Therefore, the smoothed weighted Gehan estimating equations can be written as: (2.4) U ¯ m , G ( θ ) = m − 1 ∑ i = 1 m ∑ j = 1 m δ i W i W j ( X i − X j ) Φ ( e j ( θ ) − e i ( θ ) r i j ) = 0 , where r i j 2 = m − 1 ( X j − X i ) ′ ( X j − X i ) and Φ(·) is the standard normal distribution function. Let θ ^ m denote the estimator by solving equations (2.4) . In the following Section, we will show the solution θ ^ m is a consistent estimator for θ 0 and is also asymptotically normal.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-08-16T09:21:09.727480+00:00