Introduction
Colorectal cancer (CRC) is the second most common cancer in women and third most common cancer
in men with an estimated 500 000 new cancer c ases (1). Recent work has elucidated the importance of
genetic events leading to colorectal cancer based on clinical and molecular studies of colorect al tumors
(2, 3).
CRC prevalence is likely to grow in the future so it is necessary to develop approaches to pre -
diagnosis screening (4). The current programs that use both invasive and non -invasive methods mostly
target the population in ages 50 β74 years (5, 6) . General g uidelines continue to evolve and current
population screening in Europe is mostly performed using a fecal occult blood test (FOBT) and
coloscopy (7) but participation remains quite low even in groups with a family history of CRC (8). In most
countries tha t use the FOBT, screening is available in every 2 years. The screening schedule is less
frequent with coloscopy and flexible sigmoidoscopy, generally every 10 years (9).
Colorectal cancer is known to have a significant genomic component with an estimated heritability of
around 12-35% (10). Family -based studies have identified rare high -penetrance mutations in at least a
dozen genes, but collectively, these account for only a small fraction of familial risk - 5% of colorectal
cancers arise in the setting of a well -established Mendelian inherited disorder (11). Frequently
described CRC -predisposition genes include APC, MLH1, MSH2, MSH6, PMS2, STK11, MUTYH,
SMAD4, BMPR1A, PTEN, TP53, CHEK2, POLD1, POLE (12-14). Individuals with specific genetics linked
hereditary CRC syndromes typically involve much more intense screening regimens often with a nnual
coloscopy starting in young adulthood (15). Testing for genes associated with highly penetrant
hereditary CRC syndromes is likely to provide significant clinical benefits in a cost -effective manner (16).
CRC risk factors have shown promise to stratify uniform screening programs (17, 18) . Risk prediction
models are attractive as they are non -invasive and are easier to implement in a general population or
primary care screening setting . Known conventional risk factors include age, obesity, a diet high in fat
and low in fibre, alcohol consumption, smoking, type II diabetes, and a family history of CRC (19).
Published models include data routinely available from electronic health records such as age, gender,
and body mass index (BMI), to more complex models containing detailed information about lifestyle
factors and genetic biomarkers (20). QCancer10 (21) performs well in men, others include Tao et al.
(2014) (22), Driver et al. (23), Ma et al. (2010) (24), Wells et al . (2014) (25). However, sporadic cancers
derived from a large r number of common, low -penetrance genetic variants with individually small
effects account for 70% of all CRC cases (26) . Therefore, it would be important for risk
stratification to apply information from the low pene trance genetic variants.
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Genome -wide association st udies (GWASs) for sporadic CRC have been extensively applied to
identify independent signals associated with modifications in CRC risk (27-39). A recent study by
Huyghe et al. enhanced the number of known independent signals for CRC to around 100 (11).
Individual SNPs can be aggregated into a polygenic risk score (P RS). The combination of a large
number of such SNPs in genetic risk scores has been demonstrated to enable relevant risk stratification
(40-44). As such, PRS complements other factors that identify groups with modified risk and could serve
as a stand-alone method for risk stratification before the diagnostic screening (45).
The aim of this study is to evaluate the CRC risk prediction performance of several published PRS and to
assess its use as a risk stratification approach in the context of Estonia. Concretely, we aim to develop a
model to input information into a PRS risk -stratified CRC screening regimen.
Methods
Participant data of UK Biobank
This study used genotypes from the UK Biobank cohort ( obtained 07.11.2019) and made available to
Antegenes under application reference number 53602. The data was collected, genotyped using either
the UK BiLEVE or Affymetrix UK Biobank Axiom Array. Colorectal cancer cases in the UK Biobank cohort
were retrieved by the status of ICD -10 codes C18, C19, and C20. We additionally included cases with
self-reported UK Biobank code β1020β.
Quality control steps and in detail methods applied in imputation data prepara tion have been
described by the UKBB and made available at
http://www.ukbiobank.ac.uk/wp -
content/uploads/2014/04/UKBiobank_genotyping_QC_documentation -web.pdf. We applied additional
quality controls on autosomal chromosomes. First, we removed all variants with allele frequencies
outside 0.1% and 99.9%, genotyping call rate <0.1, imputation (INFO) score <0.4 and Hardy -Weinberg
equilibrium test p -value < 1E -6. Sample quality control filters were based on several pre -defined UK
Biobank filters. We removed samples with excessive heterozygosity, individuals with sex ch romosome
aneuploidy, and excess relatives (> 10). Additionally, we only kept individuals for whom the submitted
gender matched the inferred gender, and the genotyping missingness rate was below 5%.
Quality controlled samples were divided into prevalent an d incident datasets for females and males
separately. The prevalent datasets included CRC cases diagnosed before Biobank recruitment with 5
times as many controls without the diagnosis. Incident datasets included cases diagnosed in any of the
linked databases after recruitment to the Biobank and all controls not included in the prevalent dataset.
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Prevalent datasets were used for identifying the best sex -specific candidate model and the incident
datasets were used to obtain an independent PRS effect estimate on CRC status.
Model selection from candidate risk models
We searched the literature for PRS models in the public domain. The requirements for inclusion to the
candidate set were the availability of the chromosomal location, reference and alternative allele, minor
allele frequency, and an estimator for the effect size either as odds ratio (OR) or its logarithm (log -OR)
specified for each genetic variant. In cases of iterative model developments on the same underlying
base data, we retained chronologica lly newer ones. The search was performed with Google Scholar and
PubMed web search engines by working through a list of publications using the search [βPolygenic risk
scoreβ or βgenetic risk scoreβ and βcolorectal cancerβ], and then manually checking the r esults for the
inclusion criteria. We addit ionally pruned the PRS from multi-allelic, non-autosomal, non -retrievable
variants based on bioinformatics re -analysis with Illumina GSA -24v1 genotyping.
PRSs were calculated as ππ
π = π½! π!"π₯π!
!!!
!
! , where π!" is the probability of observing genotype j,
where j β{0,1,2) for the i -th SNP; m is the number of SNPs; and π½! is the effect size of the i -th SNP
estimated in the PRS. The mean and standard deviation of PRS in the cohort were extracted to
standardize in dividual risk scores to Gaussian. We tested the assumption of normality with the mean of
1000 Shapiro-Wilks test replications on a random subsample of 1000 standardized PRS values.
Next, we evaluated the relationshi p between CRC status and s tandardized PRS in the two sex-specific
prevalent datasets with a logistic regression model to estimate the logistic regression -based odds ratio
per 1 standard deviation of PRS ( ORsd), its p -value, model Akaike information criterion (AIC) and Area
Under the ROC Curve (AUC) . The logistic regression model was compared to the null model using the
likelihood ratio test and to estimate the Nagelkerke and McFadden pseudo -R2. We selected the
candidate model with the highest AUC to independently assess risk strat ification in the incident dataset.
Independent performance evaluation of a polygenic risk score model
Firstly, we repeated the main an alyses in the prevalent dataset. The main aim was to derive a primary
risk stratification estimate, hazard ratio per 1 unit of standardized PRS (HRsd), using a right -censored and
left-truncated Cox -regression survival model. The start of time interval was defined as the age of
recruitment; follow-up time was fixed as the time of diagnosis. Scaled PRS was fixed as the only
independent variable of CRC diagnosis status. 95% confidence intervals were created using the
standard error of the log -hazard ratio.
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Further, we evaluated the concordance between theoretical hazard ratio estimates derive d with the
continuous per unit PRS (HR sd) estimate and the hazard ratio estimates inferred empirically from data.
For this, we binned the individuals by PRS to 5%-percentiles and estimated the empiric hazard ratio of
CRC directly between those classified in each bin and those within 40 -60 PRS percentile . Theoretical
estimates were derived from the relationship π»π
!"
!!!(!,!), where the exponent is the expected Gaussian
value between two arbitrary percentiles a and b (bounded between 0 and 1, a<b) of the Gaussian
distribution, Ξ¦!!(π, π)=(f(Q b ) β f(Q a ))
(π β π), where Q(b) is the Gaussian quantile function on a
percentile b and f(Q(b)) is the Gaussian probability density function value at a quantile function value.
We compared the two approaches by using the Spearman correlation coefficient and the proportion of
distribution-based π»π
!"
!!!(!,!) estimates in empirical confidence intervals.
Absolute risk estimation
Individual π-year (eg. 10-year) a bsolute risk calculations are based on the risk model developed by Pal
Choudhury et al (46). Individual absolute risks are estimated for currently a-year old individuals in the
presence of known risk factors ( Z) and their relative log hazard -ratio parameters ( π½). 95% uncertainty
intervals for the hazard ratio were derived using the standard error and z -statistic 95% quantiles
CIHR=exp(π½ Β±1.96*se(HRsd)), where se( HRsd) is the standard error of the log-hazard ratio estimate. Risk
factors have a multiplicative effect on the baseline hazard function. The model specifies the next π-year
absolute risk for a currently a-year old individual as
π! π‘
!!!
!
exp π½!π ππ₯π β π! u exp π½!π + π π’
!
!
ππ’ ππ‘,
where m(t) is age-specific mortality rate function and π!(t) is the baseline-hazard function , π‘ β₯ π and T is
the time to onset of the disease. The baseline -hazard function is deri ved from marginal age -specific
CRC incidence rates ( π!(t)) and distribution of risk factors Z in the general population (F(z)).
π!(π‘) β π! π‘ exp π½!π ππΉ(π§)
This absolute risk model allows disease background data from any country. In this analysis, we used
Estonian background information. We calculated average cumulative risks using data from the National
Institute of Health Development of Estonia (47) that provides population average disease rates in age
groups of 5 -year intervals. Sample sizes for each age group were acquired from Statistics Estonia for
2013-2016. Next, we assumed constant incidence rates for eac h year in the 5 -year groups. Thus,
incidence rates for each age group were calculated as IR=Xt/Nt , where X t is the number of first -time
cases at age t and Nt is the total number of individuals in this age group. Final per -year incidences were
averaged over time range 2013 -2016. Age-and sex-specific mortality data were retrieved from the World
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Health Organization (48) and competing mortality rates were constructed by subtracting yearly age - and
sex-specific disease mortality rates fro m general mortality rates. Colorectal cancer mortality estimates
were derived from the Global Cancer Observatory (49).
We applied this model to estimate absolute risks for individuals in the 1 st, 10th, 25th, 50th, 75th, 90th and
99th PRS quantiles, eg. an individual on the 50 th percentile would have a standar dized PRS of 0.
Confidence intervals for the absolute risk are estimated with the upper and lower confidence intervals of
the continuous per unit log -hazard ratio . Similarly, we used the absolute risk model to estimate lifetime
risks (between ages 0 and 85) for the individuals in the same risk p ercentiles.
A polygenic risk score based screening recommendations
Next, we simulate cumulative PRS stratified lifetime risks using sex -specific CRC incidences in the
Estonian population by evaluating lifetime absolute risks for individuals in various P RS risk percentiles.
Additionally, we assessed the differences in ages where individuals in various PRS risk percentiles attain
1 to 3-fold increases of risk compared to the 10 -year risk of an average individual of the same sex.
Lastly, we combine PRS risk -based screening recommendations from Naber et al. (50) with PRS -based
relative risk. They optimized the screening intervals against a hypothetical cost scenario under various
AUC levels of PRS models. We adapt these recommendation s to support relative risks and estimate the
proportion of individuals given different screening recommendations with our best performing PRS
model for a basis to provide individualized CRC screening recommendations. We use the ratio between
the population stratified with our best performing PRS and relative risks of 10 -year CRC to estimate the
proportion of individuals in each available recommendation category.
Results
Polygenic risk score re-validation in a UKBB population cohort
There was a total of 487,410 quality -controlled genotypes available for males and females in the
complete UK Biobank cohort. Phenotype data was available for a total of 458 696 individuals. This
included 242 832 (241 022 controls, 1810 cases) cases for inc ident females and 6230 (5198 controls, 1032
cases) for prevalent females, and additional 8117 prevalent (6769 controls, 1348 cases) males and 201
517 (199 107 controls, 2410 cases) incident males.
Altogether, 5 PRS models from 3 different publications were revalidated (11, 40, 51) . Three different
models adapted from Huyghe et al . included different subsets of identified variants: CRC7 was
composed of loci previously reported at genome -wide significance, CRC4 of the variants presented in
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
the meta-analysis of known and novel CRC risk loci and CRC6 was combined of variants that the authors
used in PRS analyses . Normality assumption of the standardized PRS was not violated with any tested
models (Shapiro-Wilks test p -values in CRC3 = 0.36, CRC4 = 0.48, CRC5 = 0.48 , CRC6 =0.49,
CRC7=0.48). The best performing model was selected based on AUC, ORsd,, AIC and pseudo -R2 metrics
for both females and males . The CRC6 model that was based on Huyghe et al. (11) performed the best
(Table 1). The modelβs AUC under the ROC curve for the a ssociation between the PRS and CR C
diagnosis was 0.626 (SE = 0.018) for males (Figure 1) and 0.622 (SE = 0.021) for females .
Table 1. Comparison metrics of colorectal cancer PRS models based on the prevalent UK Biobank
dataset.
CRC3 (40) CRC4 (11) CRC5 (51) CRC6 (11)
CRC7 (11)
Variants in original PRS
31 72 74 95 61
Variants included in our
model
31 70 74 91 57
Males
AUC (SE) 0.578
(0.019)
0.609
(0.019)
0.567
(0.019)
0.626
(0.018)
0.601
(0.019)
ORsd ((SE [log
ORsd])) 1.32 (0.03) 1.49 (0.03) 1.27(0.03) 1.58 (0.03) 1.44 (0.03)
AIC 7214.4 7132.8 7237.8 7072.2 7152.1
McFadden /
Nagelkerke
Pseudo-R2
1.2% / 1.8% 2.3% / 3.5% 0.9% / 1.3% 3.2% / 4.7% 2.1% /
3.1%
Females
AUC (SE) 0.564
(0.022)
0.608
(0.022)
0.560
(0.022)
0.622
(0.021)
0.600
(0.021)
ORsd ((SE [log
ORsd])) 1.25 (0.03) 1.48 (0.04) 1.25 (0.03) 1.56 (0.04) 1.42 (0.03)
AIC 5553.1 5468.6 5556.2 5432.3 5491.5
McFadden /
Nagelkerke
Pseudo-R2
0.8% / 1.2% 2.3% / 3.5% 0.7% / 1.1% 2.9% / 4.4%
1.9% /
2.8%
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Figure 1. ROC plot of CRC cases and con trols in the UK Biobank prevalent dataset of males.
Next, we evaluated the performance of the best performing CRC6 model in an independent UK
Biobank incident dataset with the main aim of estimating the hazard ratio per unit of PRS. Table 2
presents the performance estimation metrics in the incident dataset . Hazard ratio per 1 unit of standard
deviation (HR sd) in model CRC6 was 1.53 with standard error (log ( HR)) = 0.02) for males. The
concordance index (C -index) of the survival model testing th e relationship between PRS and CRC
diagnosis status in the femal esβ dataset was 0.617 (0.006) and highly similar in males.
Table 2. Performance metrics of the CRC6 model in the incident UK Biobank dataset.
HRsd (95%
confidence interval)
C-index (SE) -2 x log
likelihood
Likelihood ratio
test p-value
CRC6 Males
1.53 (1.47 β 1.59) 0.617 (0.006) 423.1 < 2e-16
Specificity
Sensitivity
1.0 0.8 0.6 0.4 0.2 0.0
0.0 0.2 0.4 0.6 0.8 1.0
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
(11) Females 1.50 (1.44 - 1.58) 0.613 (0.007) 298.8 < 2e-16
Hazard ratio estimates compared to individuals in 40 -60 percentile of PRS for females and males are
visualized on Figure 2. For both females (panel A) and males (panel B), the theoretical hazard ratio
matched in empirical estimateβs confidence intervals in 16 out of 16 comparisons. Alternatively, the
Spearman correlation coefficient between the empiric and theoretical hazard ratio estimates was 0.992
for males and 0.995 for females indicating near -perfect correspondence.
Polygenic risk score in colorectal cancer screening stratification
We used a model by Pal Choudhury et al. to derive individual 10 -year risks (46) and specified F(z) as the
distribution of PRS estimates in the whole UKBB cohort. The log -hazard ratio ( π½) is based on the sex-
specific estimate of the log -hazard ratio in the CRC6 model of the incident UKBB dataset. Age -specific
CRC incidence and competing mortality rates provided the background for CRC incidences in the
Estonian population.
In the Estonian population, the absolute risk of developing colorectal cancer in the next 10 years among
50-year old men in the 1st percentile is 0.16% (0.14% - 0.18%) and 1.15% (1.07% - 1.24%) in the 99 th
percentile (0.16% [0.14% - 0.18%] and 1.06% [0.97% - 1.16%] for women , respectively ). At age 70,
corresponding risks for the same percentiles of men become 1.07% ( 0.96% - 1.19%) and 7.4 % (6.85% -
7.96%). Considerably lower esti mates were found for women: 0.71% (0.63% - 0.81%) and 4.67% (4.27% -
5.09%) respectively. The relative risks between the most ext reme percentiles are therefore around 6.7 -
fold. Similarly, c ompeting risks accounted cumulative risks of females in the 99th percentile surpass
11.4% (10.6% - 12.2%) by age 85 but remain at 1.71% (1.53% - 1.91%) in the 1 st percentile (Figure 3).
Equivalent values for women are somewhat lower: 10.2% (9.33% β 11.1%) and 1.6% (1.41% - 2.6%).
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Figure 2. Hazard ratio estimates between quantiles 40 -60 of the PRS and categorized 5% bins in the
UKBB incident dataset. White dots and blue lines represent empirically estimated hazard ratio
β
β
β
β β
β β β
β
β
β
β β
β
β
β
0.5
1.0
1.5
2.0
2.5
q1 q2 q3 q4 q5 q6 q7 q8 q9 q10 q11 q12 q13 q14 q15 q16
HR
Hazard Ratio
2
1
Polygenic risk score quantiles
0-5
5-10
10-15
15-20
20-25
25-30
30β35
35β40
60-65
65-70
70-75
75-80
80-85
85-90
90-95
95-100
A
+
β
β β
β β
β
β
β
β β
β
β
β
β
β
β
1
2
q1 q2 q3 q4 q5 q6 q7 q8 q9 q10 q11 q12 q13 q14 q15 q16
HR
Hazard Ratio
2
1
Polygenic risk score quantiles
0-5
5-10
10-15
15-20
20-25
25-30
30β35
35β40
60-65
65-70
70-75
75-80
80-85
85-90
90-95
95-100
B
+
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
estimates and corresponding confidence intervals. Black dashe s represent the theoretical hazard ratio
for the 5%-quantile bins derived from the hazard ratio of per unit PRS. (A) Females (B) Males
0
3
6
9
40 60 80
Age (years)
Risk of colorectal cancer for Estonian females (%)
AnteCRC risk percentile
1%
10%
25%
50%
75%
90%
99%
Cumulative risks by AnteCRC risk group
Risk percentile
A
0.0
2.5
5.0
7.5
10.0
12.5
40 60 80
Age (years)
Risk of colorectal cancer for Estonian males (%)
AnteCRC risk percentile
1%
10%
25%
50%
75%
90%
99%
Cumulative risks by AnteCRC risk group
Risk percentile
B
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Figure 3. Cumulative risks (%) of colorectal cancer between ages 20 to 85 in various risk percentile s. (A)
Females (B) Males
Genetically average 50 -year-old males have a 10-year risk of CRC equal ing 0.432% and for females it is
0.412%. A 41-year-old male in the 99th percentile (42 in females) has a larger risk than an average 50 -
year-old. At the same time, males in the 1st percentile attain this risk by age 58 (61 in females). Males
above the 95th percentile (96th percentile in females) have a more than 2 -fold risk increase compared to
the average. A genetically average female doubles her risk at age 50 by age 58 (55 in males) and triples
it by age 63 (59 in males). The 1 st percentile of females only attains the double of average 50 -year oldβs
risk by age 72 (66 in males) (Figure 4).
A
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Figure 4. Ages when Estonian individuals in different risk percent iles attain 1-3 fold multiples of 10-year
risk compared to 50-year old individuals with population average PRS (Risk level: β averageβ). (A)
Females (B) Males
Relative risk stratification with PRS provided a basis for personal CRC screening intervals .
Recommendations for coloscopy based screening presented below are adapted from Naber et al. (50)
for a model with AUC=0.6 and required individual relative risks as input . CRC screening is individualized
by differences in the intervals of planned colos copies. Relative risks are derived from 10 -year risk
differences compared to average PRS. As an alternative to coloscopy, we recommend annual fecal
immunochemical test ing, with individualized starting from the age at which the patient attains the 10-
year risk of the average 50-year-old person (52-54).
Alternative A. Coloscopy
1. Relative risk less than 0.6
β’ 1 coloscopy at age 60*
2. Relative risk between 0.6 and 0.7
β’ Coloscopies at ages 55 and 70**
3. Relative risk between 0.7 and 0.9
β’ Coloscopies at ages 50 and 65**
4. Relative risk between 0.9 and 1.4
B
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
β’ Coloscopies every 10 years, ages 50 -75***
5. Relative risk between 1.4 and 1.6
β’ Coloscopies every 10 years, ages 50 -75***
6. Relative risk between 1.6 and 3.1
β’ Coloscopies every 7 years, ages 45 -75***
7. Relative risk larger than 3.1
β’ Coloscopy every 7 years, ages 40 -80***
Alternative B. Fecal immunochemical test
1. Relative risk higher than 1
β’ Annual fecal immunochemical test from the age at which the patient 10-
year risk reaches that of the average 50-year-old person*
2. Relative risk lower than 1
β’ Annual fecal immunochemical test from age 50*
* - If the recommended age is below the individual 's current age, then recommend current age
** - If the patient is older than first coloscopy age then the first coloscopy is suggested at current
age and the second coloscopy is suggested after 15 years (before age 75) only if the patient is
currently below 60
*** - If the patient is older than start of coloscopy interval then start at current age and use the
coloscopy interval to plan visits until the end of age interval.
Next, we estimated the proportion of individuals with relative risks to an average individual in the
Estonian population. We estimated that relative risks as more than 3.1, compared to an individual with
median population PRS, in around 0.3% of females, between 1.6 and 3.1 in 11.9%, between 1.4 and 1.6
in 10.3%, between 0.9 and 1.4 in 37.7%, between 0.7 and 0.9 in 20.8%, between 0.6 and 0.7 in 9.4% and
below 0.6 in 9.5%. The equivalent proportions for males were highly similar.
References
1. Ferlay J, Colombet M, Soerjomataram I, Dyba T, Randi G, Bettio M, et al. Cancer incidence and
mortality patterns in Europe: Estimates for 40 countries and 25 major cancers in 2018. European journal
of cancer (Oxford, England : 1990). 2018;103:356-87.
2. Genetic testing for colon cancer: joint statement of the American College of Medical Genetics
and American Society of Human Genetics. Joint Test and Technology Transfer Committee Working
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
Group. Genetics in medicine : official journal of the American College of Medical Genetics.
2000;2(6):362-6.
3. Stanesby O, Jenkins M. Comparison of the efficiency of colorectal cancer screening programs
based on age and genetic risk for reduction of colorectal cancer mortality. European journa l of human
genetics : EJHG. 2017;25(7):832 -8.
4. Arnold M, Sierra MS, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global patterns and
trends in colorectal cancer incidence and mortality. Gut. 2017;66(4):683 -91.
5. Senore C, Basu P, Anttila A, Ponti A, Tomatis M, Vale DB, et al. Performance of colorectal
cancer screening in the European Union Member States: data from the second European screening
report. Gut. 2018.
6. Schreuders EH, Ruco A, Rabeneck L, Schoen RE, Sung JJ, Young GP, et al. Colorectal canc er
screening: a global overview of existing programmes. Gut. 2015;64(10):1637 -49.
7. Zavoral M, Suchanek S, Zavada F, Dusek L, Muzik J, Seifert B, et al. Colorectal cancer screening
in Europe. World journal of gastroenterology: WJG. 2009;15(47):5907.
8. Tsai M-H, Xirasagar S, Li Y -J, De Groen PC. Peer Reviewed: Colonoscopy Screening Among US
Adults Aged 40 or Older With a Family History of Colorectal Cancer. Preventing chronic disease.
2015;12.
9. Rex DK. Screening Tests for Colon Cancer. Gastroenterology & hepatology. 2016;12(3):197 -9.
10. Mucci LA, Hjelmborg JB, Harris JR, Czene K, Havelick DJ, Scheike T, et al. Familial Risk and
Heritability of Cancer Among Twins in Nordic Countries. Jama. 2016;315(1):68-76.
11. Huyghe JR, Bien SA, Harrison TA, Kang HM, Chen S, Schmit SL, et al. Discovery of common
and rare genetic risk variants for colorectal cancer. Nature genetics. 2019;51(1):76 -87.
12. Fishel R, Lescoe MK, Rao MR, Copeland NG, Jenkins NA, Garber J, et al. The human mutator
gene homolog MSH2 and its a ssociation with hereditary nonpolyposis colon cancer. Cell.
1993;75(5):1027-38.
13. Bronner CE, Baker SM, Morrison PT, Warren G, Smith LG, Lescoe MK, et al. Mutation in the
DNA mismatch repair gene homologue hMLH1 is associated with hereditary non -polyposis colon
cancer. Nature. 1994;368(6468):258 -61.
14. Al-Tassan N, Chmiel NH, Maynard J, Fleming N, Livingston AL, Williams GT, et al. Inherited
variants of MYH associated with somatic G:C -->T:A mutations in colorectal tumors. Nature genetics.
2002;30(2):227-32.
15. Syngal S, Brand RE, Church JM, Giardiello FM, Hampel HL, Burt RW. ACG clinical guideline:
genetic testing and management of hereditary gastrointestinal cancer syndromes. The American journal
of gastroenterology. 2015;110(2):223.
16. Gallego CJ, Shi rts BH, Bennette CS, Guzauskas G, Amendola LM, Horike -Pyne M, et al. Next -
Generation Sequencing Panels for the Diagnosis of Colorectal Cancer and Polyposis Syndromes: A
Cost-Effectiveness Analysis. Journal of clinical oncology : official journal of the Ame rican Society of
Clinical Oncology. 2015;33(18):2084 -91.
17. Win AK, Macinnis RJ, Hopper JL, Jenkins MA. Risk prediction models for colorectal cancer: a
review. Cancer epidemiology, biomarkers & prevention : a publication of the American Association for
Cancer Research, cosponsored by the American Society of Preventive Oncology. 2012;21(3):398 -410.
18. Read TE, Kodner IJ. Colorectal cancer: risk factors and recommendations for early detection.
American Family Physician. 1999;59(11):3083.
19. Le Marchand L, Wilkens LR, Kolonel LN, Hankin JH, Lyu L -C. Associations of sedentary lifestyle,
obesity, smoking, alcohol use, and diabetes with the risk of colorectal cancer. Cancer research.
1997;57(21):4787-94.
20. Usher-Smith JA, Harshfield A, Saunders CL, Sharp SJ, Emery J, Walter FM, et al. External
validation of risk prediction models for incident colorectal cancer using UK Biobank. British journal of
cancer. 2018;118(5):750 -9.
21. Hippisley-Cox J, Coupland C. Development and validation of risk prediction algorithm s to
estimate future risk of common cancers in men and women: prospective cohort study. BMJ open.
2015;5(3):e007825.
22. Sharara AI, Harb AH. Development and validation of a scoring system to identify individuals at
high risk for advanced colorectal neopla sms who should undergo colonoscopy screening. Clinical
gastroenterology and hepatology : the official clinical practice journal of the American
Gastroenterological Association. 2014;12(12):2135 -6.
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
23. Driver JA, Gaziano JM, Gelber RP, Lee IM, Buring JE, Ku rth T. Development of a risk score for
colorectal cancer in men. The American journal of medicine. 2007;120(3):257 -63.
24. Ma E, Sasazuki S, Iwasaki M, Sawada N, Inoue M. 10 -Year risk of colorectal cancer:
development and validation of a prediction model i n middle -aged Japanese men. Cancer
epidemiology. 2010;34(5):534 -41.
25. Wells BJ, Kattan MW, Cooper GS, Jackson L, Koroukian S. Colorectal cancer predicted risk
online (CRC -PRO) calculator using data from the multi -ethnic cohort study. Journal of the Ameri can
Board of Family Medicine : JABFM. 2014;27(1):42 -55.
26. Fearon ER, Vogelstein B. A genetic model for colorectal tumorigenesis. Cell. 1990;61(5):759 -67.
27. Broderick P, Carvajal -Carmona L, Pittman AM, Webb E, Howarth K, Rowan A, et al. A genome -
wide as sociation study shows that common alleles of SMAD7 influence colorectal cancer risk. Nature
genetics. 2007;39(11):1315 -7.
28. Tomlinson I, Webb E, Carvajal -Carmona L, Broderick P, Kemp Z, Spain S, et al. A genome -wide
association scan of tag SNPs identifie s a susceptibility variant for colorectal cancer at 8q24.21. Nature
genetics. 2007;39(8):984 -8.
29. Zanke BW, Greenwood CM, Rangrej J, Kustra R, Tenesa A, Farrington SM, et al. Genome -wide
association scan identifies a colorectal cancer susceptibility locu s on chromosome 8q24. Nature
genetics. 2007;39(8):989 -94.
30. Berndt SI, Potter JD, Hazra A, Yeager M, Thomas G, Makar KW, et al. Pooled analysis of
genetic variation at chromosome 8q24 and colorectal neoplasia risk. Human molecular genetics.
2008;17(17):2665-72.
31. Houlston RS, Webb E, Broderick P, Pittman AM, Di Bernardo MC, Lubbe S, et al. Meta -analysis
of genome -wide association data identifies four new susceptibility loci for colorectal cancer. Nature
genetics. 2008;40(12):1426 -35.
32. Jaeger E, Webb E, Howarth K, Carvajal -Carmona L, Rowan A, Broderick P, et al. Common
genetic variants at the CRAC1 (HMPS) locus on chromosome 15q13.3 influence colorectal cancer risk.
Nature genetics. 2008;40(1):26 -8.
33. Tenesa A, Farrington SM, Prendergast JG, Porteous ME, Walker M, Haq N, et al. Genome -wide
association scan identifies a colorectal cancer susceptibility locus on 11q23 and replicates risk loci at
8q24 and 18q21. Nature genetics. 2008;40(5):631 -7.
34. Tomlinson IP, Webb E, Carvajal -Carmona L, Broderick P, Howarth K, Pittman AM, et al. A
genome-wide association study identifies colorectal cancer susceptibility loci on chromosomes 10p14
and 8q23.3. Nature genetics. 2008;40(5):623 -30.
35. Houlston RS, Cheadle J, Dobbins SE, Tenesa A, Jones AM, Howarth K, et a l. Meta-analysis of
three genome-wide association studies identifies susceptibility loci for colorectal cancer at 1q41, 3q26.2,
12q13.13 and 20q13.33. Nature genetics. 2010;42(11):973 -7.
36. Tomlinson IP, Carvajal -Carmona LG, Dobbins SE, Tenesa A, Jones AM , Howarth K, et al.
Multiple common susceptibility variants near BMP pathway loci GREM1, BMP4, and BMP2 explain part
of the missing heritability of colorectal cancer. PLoS genetics. 2011;7(6):e1002105.
37. Peters U, Jiao S, Schumacher FR, Hutter CM, Aragaki AK, Baron JA, et al. Identification of
Genetic Susceptibility Loci for Colorectal Tumors in a Genome -Wide Meta -analysis. Gastroenterology.
2013;144(4):799-807.e24.
38. Whiffin N, Hosking FJ, Farrington SM , Palles C, Dobbins SE, Zgaga L, et al. Identification of
susceptibility loci for colorectal cancer in a genome -wide meta -analysis. Human molecular genetics.
2014;23(17):4729-37.
39. Al-Tassan NA, Whiffin N, Hosking FJ, Palles C, Farrington SM, Dobbins SE, et al. A new GWAS
and meta -analysis with 1000Genomes imputation identifies novel risk variants for colorectal cancer.
Scientific reports. 2015;5:10442.
40. Hsu L, Jeon J, Brenner H, Gruber SB, Schoen RE, Berndt SI, et al. A model to determine
colorectal c ancer risk using common genetic susceptibility loci. Gastroenterology. 2015;148(7):1330 -
9.e14.
41. Jenkins MA, Makalic E, Dowty JG, Schmidt DF, Dite GS, MacInnis RJ, et al. Quantifying the
utility of single nucleotide polymorphisms to guide colorectal canc er screening. Future oncology
(London, England). 2016;12(4):503 -13.
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
42. Zhang B, Shrubsole MJ, Li G, Cai Q, Edwards T, Smalley WE, et al. Association of genetic
variants for colorectal cancer differs by subtypes of polyps in the colorectum. Carcinogenesis.
2012;33(12):2417-23.
43. Abuli A, Castells A, Bujanda L, Lozano JJ, Bessa X, Hernandez C, et al. Genetic Variants
Associated with Colorectal Adenoma Susceptibility. PloS one. 2016;11(4):e0153084.
44. Xin J, Chu H, Ben S, Ge Y, Shao W, Zhao Y, et al. Evalu ating the effect of multiple genetic risk
score models on colorectal cancer risk prediction. Gene. 2018;673:174 -80.
45. Frampton MJ, Law P, Litchfield K, Morris EJ, Kerr D, Turnbull C, et al. Implications of polygenic
risk for personalised colorectal cance r screening. Annals of oncology : official journal of the European
Society for Medical Oncology. 2016;27(3):429 -34.
46. Pal Choudhury P, Maas P, Wilcox A, Wheeler W, Brook M, Check D, et al. iCARE: An R package
to build, validate and apply absolute risk mo dels. PloS one. 2020;15(2):e0228198.
47. Database HSaHR. 2020.
48. Organization WH. Global Health Observatory indicator views 2018 [Available from:
https://apps.who.int/gho/data /node.imr.LIFE_0000000029?lang=en .
49. Ferlay J, Colombet M, Soerjomataram I, Mathers C, Parkin DM, Pineros M, et al. Estimating the
global cancer incidence and mortality in 2018: GLOBOCAN sources and methods. International journal
of cancer. 2019;144(8):1 941-53.
50. Naber SK, Kundu S, Kuntz KM, Dotson WD, Williams MS, Zauber AG, et al. Cost -Effectiveness
of risk-stratified colorectal cancer screening based on polygenic risk: current status and future potential.
JNCI cancer spectrum. 2020;4(1):pkz086.
51. Schmit SL, Edlund CK, Schumacher FR, Gong J, Harrison TA, Huyghe JR, et al. Novel Common
Genetic Susceptibility Loci for Colorectal Cancer. Journal of the National Cancer Institute.
2019;111(2):146-57.
52. Armaroli P, Villain P, Suonio E, Almonte M, Anttila A, Atkin WS, et al. European Code against
Cancer, 4th Edition: Cancer screening. Cancer epidemiology. 2015;39 Suppl 1:S139 -52.
53. Autier P. Personalised and risk based cancer screening. BMJ (Clinical research ed). 2019;367.
54. Doubeni C. Screening for c olorectal cancer: Strategies in patients at average risk 2019
[Available from: https://www.uptodate.com/contents/screening -for-colorectal-cancer-strategies-in-
patients-at-average-risk.
55. Iwasaki M, Tanaka -Mizuno S, Kuchiba A, Yamaji T, Sawada N, Goto A, et al. Inclusion of a
genetic risk score into a validated risk prediction model for colorectal cancer in Japanese men improves
performance. Cancer prevention research. 2017;10(9):535 -41.
56. Weigl K, Thomsen H, Balavarca Y, Hellwege JN, Shrubsole MJ, Brenner H. Genetic Risk Score
Is Associated With Prevalence of Advanced Neoplasms in a Colorectal Cancer Screening Population.
Gastroenterology. 2018;155(1):88-98.e10.
57. Weigl K, Chang -Claude J, Knebel P, Hsu L, Hoffmeister M, Brenner H. Strongly enhanced
colorectal cancer risk stratification by combining family history and genetic risk score. Clinical
epidemiology. 2018;10:143.
58. Archambault AN, Su Y -R, Jeon J, Thomas M, Lin Y, Conti DV, et al. Cumulative Burden of
Colorectal Cancer βAssociated Genetic Variants Is More Strongly Associated With Early -Onset vs Late -
Onset Cancer. Gastroenterology. 2020;158(5):1274 -86. e12.
59. Oncology ESfM. D ecline in colorectal cancer deaths in Europe is a 'major success' story 2018
[Available from: https://www.eurekalert.org/pub_releases/2018 -03/esfm-dic031518.php.
60. Hall N, Birt L, Rees CJ, Walter FM, Elliot S, Ritchie M, et al. Concerns, perceived need and
competing priorities: a qualitative exploration of decision -making and non-participation in a population -
based flexible sigmoidoscopy screening programme to prevent color ectal cancer. BMJ open.
2016;6(11):e012304.
61. Issa IA, Noureddine M. Colorectal cancer screening: An updated review of the available
options. World journal of gastroenterology. 2017;23(28):5086.
62. Li B, Gan A, Chen X, Wang X, He W, Zhang X, et al. Diagnostic Performance of DNA
Hypermethylation Markers in Peripheral Blood for the Detection of Colorectal Cancer: A Meta -Analysis
and Systematic Review. PloS one. 2016;11(5):e0155095.
63. Lieberman DA, Rex DK, Winawer SJ, Giardiello FM, Johnson DA, Levin TR. Guidelines for
colonoscopy surveillance after screening and polypectomy: a consensus update by the US Multi -Society
Task Force on Colorectal Cancer. Gastroenterology. 2012;143(3):844 -57.
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted August 22, 2020. ; https://doi.org/10.1101/2020.08.19.20177931doi: medRxiv preprint