To pool or not to pool, from whether to when: applications of pooling to biospecimens subject to a limit of detection.

OA: closed

Abstract

Pooling of biological specimens has been utilised as a cost-efficient sampling strategy, but cost is not the unique limiting factor in biomarker development and evaluation. We examine the effect of different sampling strategies of biospecimens for exposure assessment that cannot be detected below a detection threshold (DT). The paper compares use of pooled samples to a randomly selected sample from a cohort in order to evaluate the efficiency of parameter estimates. The proposed approach shows that a pooling design is more efficient than a random sample strategy under certain circumstances. Moreover, because pooling minimises the amount of information lost below the DT, the use of pooled data is preferable (in a context of a parametric estimation) to using all available individual measurements, for certain values of the DT. We propose a combined design, which applies pooled and unpooled biospecimens, in order to capture the strengths of the different sampling strategies and overcome instrument limitations (i.e. DT). Several Monte Carlo simulations and an example based on actual biomarker data illustrate the results of the article.
Full text 23,245 characters · extracted from pmc-nxml · 4 sections · click to expand

Section 4

Having introduced the pooled–unpooled hybrid design, we can utilise maximum likelihood to estimate unknown parameters of a biomarker’s distribution. We can consider a biomarker that follows X i ~ N (μ, σ 2 ), i = 1, …, N , where estimation of the mean of X i (the variance is assumed to be unknown) based on Z or Z ( p ) is of interest. Note that Z i is not normally distributed because there will be a mass of data at the DT. Therefore, the direct classical maximum likelihood method based on normal data is unsuitable, and naively applying the least squares method to the available data leads to an estimator of E ( X i , X i ≥ d ) ≠ μ. Instead, we utilise maximum likelihood estimators proposed by Gupta that are appropriate for this case because they yield asymptotically unbiased estimators of the mean and variance (see Appendix for estimation details). 15 The target likelihood function [ Appendix formula (A.1) ] has two parts: a component related to the number of unobserved test results and a component for the numerically observed data (this part is similar to the classic MLE). The statistical properties and rate of convergence to unbiasedness of the maximum likelihood estimators are functions of the location of the DT. These estimators are based on likelihood theory; they maximise the Fisher information and are the most efficient. However, their accuracy and efficiency are direct results of the study design. We propose the obtained efficiency of (μ̂, σ̂) in formula ( A.2 ) of the Appendix . Suppose X i has the standard normal distribution N (0,1) Figure 4 displays the efficiency of the maximum likelihood estimators. The estimations of μ and σ based on the pooled data are more efficient than those based upon the random sample up to d < μ. However, if d ≫ μ, then the pooling strategy is not recommended. We can estimate parameters based on data following a gamma distribution in a similar manner. 16 The gamma-shape parameter of the pooled data is p × the shape parameter of unpooled data and the scale parameter of the pooled data is 1/ p × the scale parameter of unpooled data. 3 The conclusions for the gamma case are similar to the normal. The likelihood function in Section 4.1 is composed of two parts: one related to N / A -observed data (where X < d ) and a second for numerically observed data (where X ≥ d ); estimation for pooled–unpooled resampled data has three kinds of data. The first sample ( S 1 ) has only N / A elements. Test results in this sample were initially below the DT and have not been reconstructed by applying the pooling resampling. Thus, for all k = 1, …, N , we have The second sample ( S 2 ) has reconstructed elements. Test results in this sample were initially below the DT and have been reconstructed by applying the pooling resampling. Therefore, elements of set S 2 have distribution function The last sample ( S 3 ), as in Section 4.1, includes the numerically observed data. The likelihood function is a product of the densities that correspond to ( 3.1 ), ( 3.2 ) and the case, where numerical results were initially observed. We describe the likelihood in detail in Appendix formula (A.3) . To illustrate the proposed method, we generated a random sample { X 1 ,…, X N =300 } 10 000 times from N ( μ = 1 , σ X 2 = 4 ) . We applied the one-step resampling to each simulated sample for d = 0.5. The average number of observed numerical values in the simulation (#{ X i > d }) was about 179.58 (which agrees with the estimate n × ( 1 − ∫ − ∞ d 1 2 π σ X exp ( − 1 2 ( u − μ σ X ) 2 ) d u ) = 179.6119 . After applying the resampling strategy, the average number of numerical elements in the Monte Carlo (MC) data was about 221.63 (which agrees with the estimate n × (1 − Pr{ X k ∈ S 1 }) = 221.569). Thus, on average, the numerical information increased by approximately 25%. As would be expected, the MC variances of the μ and σ-estimators improved due to the resampling by about 24%. Similar results are attained for different values of d , where d < μ. Cholesterol measurements were collected for 10 normal volunteers at a medical centre. The mean and standard deviation for total cholesterol were estimated to be (μ̂ = 200.73) and (σ̂ = 51.72) respectively. The specimens were then randomly paired and the pooled specimens were assayed. For the purpose of demonstration, we artificially created a threshold (DT = 150) such that some numerical values could not be observed. In Table 1 , we show the individual and pooled cholesterol values with and without the DT. In this example, 20% of the individual measurements are below the threshold, whereas no pooled observations are below the DT. Applying the maximum likelihood method, the asymptotically unbiased mean and standard deviation were estimated to be (μ̂ = 196.13) and (σ̂ = 56.38), respectively, from unpooled data with the DT. Although more costly, by assaying both pooled and unpooled specimens, we can reasonably estimate values below the DT ( Table 1 ). Moreover, using both the reconstructed data and the unpooled data above the DT, the mean and the variance are estimated to be μ̂ = 198.99 and σ̂ = 53.44

Intro

The use of biomarkers for exposure assessment is common in epidemiology. The power gained by using a large sample of individuals must be weighed against the cost of performing many assays. After reproducibility and variability are established for the biomarker, financial constraints usually limit further evaluation to small sets of samples. For example, the cost of a single assay measuring polychlorinated biphenyls (PCBs) is up to $1000 so only small studies have been able to examine, for example, whether PCBs are associated with cancer or endometriosis. 1 , 2 However, the imprecision of the results limits the conclusions that can be drawn for the suggested association. Currently, two different approaches have been suggested to evaluate expensive biomarkers. Suppose we have biological specimens from a patient population A of size N , A = { A 1 , A 2 , …, A N }, with test results X = { X 1 , X 2 , …, X N }. One approach selects a random sample of the patient population A ( r ) = { A k 1 , A k 2 , …, A k n } ∈ A, where n (≤ N ) is determined by a power calculation and { k i , i = 1, …, n } is a subsequence of set {1, 2, 3, …, N } where assays are performed on the subset of specimens with observed results { X k 1 , X k 2 , …, X k n }. Alternatively, a pooling strategy may be employed where two or more specimens are physically combined into a single ‘pooled’ unit for analysis. Thus, a greater portion of the population is assayed for the same price compared with the random sampling approach. The amount of information per assay increases so the number of assays needed to achieve equivalent information decreases. 3 – 6 Formally, the samples from patient population A are randomly combined into n = N / p pooled specimens of size p . The n pooled assays are considered the average of the contributing individual results, i.e. (1.1) X ( p ) = { X 1 ( p ) , X 2 ( p ) , … , X n ( p ) } = { 1 p ( X k 11 + … + X k 1 p ) , 1 p ( X k 21 + … + X k 2 p ) , … , 1 p ( X k n 1 + … + X k n p ) } , where { k 1 i , i = 1, …, p }, …, { k ni , i = 1, …, p } are some disjoint subsequences of set {1, 2, 3, …, N }. Note that, this formal definition of pooled data is commonly accepted for methodological analyses and practical applications of the pooling design. 3 – 6 The concept of pooling biospecimens can be utilised in population-based epidemiological studies to explore the relationship between biomarker levels and outcome. The method’s primary goal is in establishing distributional parameters for a specific biomarker. Consequently, pooling can be seen as a primary tool for case–control and cohort studies exploring discrete outcomes. The technique has been explored extensively in the literature starting with publications related to cost-efficient syphilis testing of World War II recruits. 7 Weinberg and Umbach introduced pooling to estimate odds ratios for case–control studies. 8 Faraggi et al. 3 and Liu and Schisterman 4 , 5 examined the inference of the effect of pooling on the area under the Receiver Operating Characteristic curve. Cost is not the only limiting factor in biomarker evaluation. Instrument sensitivity may also be problematic. Another common complexity arises when some participants have levels below the detection threshold (DT). 9 Under these circumstances, biomarker values at or above the DT are measured and reported, but values below the DT are unobservable. Formally, instead of X , we observe Z = { Z 1 , Z 2 , …}, such that (1.2) Z i = { X i , if X i ≥ d ; Not Available ( N / A ) , if X i < d , where d is a value of the DT. Similarly, for the pooling design, we observed Z ( p ) = { Z 1 ( p ) , … , Z n ( p ) } , where A variety of approaches have been used to analyse data with a lower DT. Substitution of d / 2 or d / 2 for observations below the DT has been previously described. 9 – 11 These values are based on the assumption of a normal ( d / 2 ) or lognormal ( d / 2 ) distribution. 12 Lubin and colleagues proposed multiple imputation based on bootstrapping when the exposure distribution function is known. 13 Recent work shows that substitution of E ( X | X < d ) for data below the DT allows for unbiased estimation of linear and, under certain conditions, logistic regression parameters. 12 Schisterman and colleagues have shown that unbiased estimates may also be obtained non-parametrically if data below the DT are replaced by zero for no intercept models and by an estimator of E ( X | X ≥ d ) for intercept models. 12 The main objective of this paper is to examine parameter estimation and efficiency of the pooling approach compared with the random sampling approach for assays with a DT. In Section 2, we compare the numerical (quantifiable) information available and efficiency of each sampling scheme. In Section 3, we propose a mixed (unpooling–pooling) design, which takes advantage of the strengths of each approach. In Section 4, we present maximum likelihood techniques to be utilised with the different designs. In Section 5 we illustrate methods to account for pooling and random measurement error and in Section 6 we present our conclusions.

Methods

Although definition ( 1.1 ) shows the theoretical notation for pooled data, practically, pooling biological specimens can lead to additive pooling errors. In this section we use the maximum likelihood method from Section 4 and revise definition ( 1.1 ) to (5.1) X ( p ) = { X 1 ( p ) , X 2 ( p ) , … , X n ( p ) } = { 1 p ( X k 11 + … + X k 1 p ) + ε 1 , 1 p ( X k 21 + … + X k 2 p ) + ε 2 , … 1 p ( X k n 1 + … + X k n p ) + ε n } , where pooling errors ε 1 ,…, ε n are independent N ( 0 , σ ε 2 ) distributed random variables and X ∼ N ( μ , σ X 2 ) . Definition ( 5.1 ) accounts for the pooling errors which were ignored by definition ( 1.1 ). In order to investigate the robustness of our approach for addressing pooling errors, we executed MC simulations. Formally, we assumed that only n = N / p measurements can be performed and compare: Random sampling: We randomly choose X 1 ,…, X N/p from the full sample and observed Z 1 ,…, Z N/p because of the DT. The mean of X i was estimated using the likelihood approach on the truncated data { Z 1 ,…, Z N/p }. Pooling: We randomly choose biospecimen sets of size p with X i ( p ) = 1 p ∑ j = p ( i − 1 ) + 1 p i X j + ε i , i = 1 , … , N p and observed Z i ( p ) , i = 1 , … , N p by ( 1.3 ). Again, the mean of X i was estimated using the MLE based on Z 1 ( p ) , … , Z N / p ( p ) . The accuracy of estimators (μ̂, σ̂) of ( μ , σ X ) is indicated by their MC variances. We assumed a biomarker distribution X i ∼ N ( μ = 1 , σ X 2 = 4 ) , i = 1,…, N , and generated a random sample X 1 ,…, X N =300 10 000 times for pool sizes p = 2,4 and for various DTs. Figure 5 presents the MC estimators of E (μ̂ − EX ) 2 and E (σ̂ − σ X ) 2 , where μ̂, the estimator of EX , is based on full data { Z 1 ,…, Z N ), the random sample { Z 1 ,…, Z N/p } and the pooled data { Z 1 ( p ) , … , Z N / p ( p ) } with σ ε = σ X /10, σ ε = σ X /5, σ ε = σ X /3, σ ε = σ X /2 and σ ε = σ X . The figure suggests that the conclusions in Sections 2 and 4 are correct for μ-estimation up to σ ε ≤ σ X /2 [ Figs 5(a.1)–(a.4) ]. However, above σ ε ≥ σ X /5, the σ X -estimator using the pooled data has the largest variance. The likelihood function ( A.1 ) should be changed because of pooling errors as shown in ( 5.1 ). The statistical properties of ε 1 ,…, ε n can be evaluated in a similar manner to a study by Schisterman et al . 17 In addition to pooling error, studies are also subjected to random measurement error as a function of instrument calibration. Random measurement error occurs as a result of random instrument variability. One can account for random measurement error in the pooled or random sample designs through the use of standard techniques previously developed in the literature. 18 , 19 These techniques include utilising error models, regression calibration models, validation studies or replication data to estimate and adjust for random measurement error. In addition, while not explicitly described here, standard information reported by a laboratory such as the coefficient of variation for the biomarker and reliability of the assay can be included in these models.

Discussion

In this paper, we examine pooling and random sampling as strategies to evaluate biospecimens with a DT. These types of data are common in epidemiological research and include two types of values: numerical and non-numerical (i.e. N/A). Because numerical values yield more information than missing data, it is a goal of any researcher to minimise the number of N/A observations. Accordingly, we have explored theoretical methods as well as simulations where a pooling design is more efficient than a random sample. In addition, we show that the efficiency of the pooling design is dictated by the location of the DT but is independent of the distributional assumptions (e.g. gamma, t-distribution, Lognormal). For all distributions, there is a range of DTs where the pooling strategy is more efficient than a random sample because the inference-based pooling design provides more numerical information. In fact, in some cases pooling is more efficient than using the full sample. We showed that whenever EX > d (i.e. >50%) pooling is always the most efficient sampling strategy, but other factors, such as the underlying distribution, must be considered when EX d , is appropriate. Towards this end, the unpooled–pooled strategy proposed in Section 3 is not only helpful for the evaluation of pooling errors but can also be applied to a first-stage data study. In addition, the efficiency of MLEs under each design can be evaluated. Cost has been the main motivation for pooling biological specimens or to randomly select a subset of individuals to be assayed. However, we have shown that in some cases, even when the full data are available, estimations based on pooled data increase efficiency over the use of individual measurements when the assay has a DT. This is because of the greater number of observations above the DT under pooling, which can then be used in the estimation procedure. However, using unpooled data allows, for example, distributional assumptions to be tested, the location of the DT to be estimated and the expected number of observations below the DT. In addition, one is able to stratify the pooled samples by confounders in order to retain confounding and covariate information in the pooled samples. To take advantage of the strengths of each of these approaches, we proposed a pooled–unpooled resampling design. According to this design, in the first stage all the patient population (or a random sample of them) is measured individually, and in the second stage, the patient population is pooled in groups of size p and these pooled samples are assayed. By employing this approach, we are able to reconstruct data that were unobserved in the first stage due to the DT. This simple approach that we propose captures the strengths of the statistical properties of the distribution function of the averages by physically grouping biological specimens in order to overcome the instrument limitations.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-08-04T06:16:37.499272+00:00