Full text
57,958 characters
· extracted from
oa-pdf
· click to expand
Inferring simple but precise quantitative models of human oocyte and early embryo
development
Brian D. Leahy, 1, 2 Catherine Racowsky, 3, 4 and Daniel Needleman 1, 2, 5
1Dept. of Molecular and Cellular Biology, Harvard University, Cambridge MA
2SEAS, Harvard University, Cambridge, MA
3Brigham Women’s Hospital, Boston, MA
4Harvard Medical School, Boston, MA
5Center for Computational Biology, Flatiron Institute, New York, NY
(Dated: March 24, 2021)
Macroscopic, phenomenological models have proven useful as concise framings of our understand-
ings in fields from statistical physics to economics to biology. Constructing a phenomenological
model for development would provide a framework for understanding the complicated, regulatory
nature of oogenesis and embryogenesis. Here, we use a data-driven approach to infer quantitative,
precise models of human oocyte maturation and pre-implantation embryo development, by analyzing
existing clinical In-Vitro Fertilization (IVF) data on 7,399 IVF cycles resulting in 57,827 embryos.
Surprisingly, we find that both oocyte maturation and early embryo development are quantitatively
described by simple models with minimal interactions. This simplicity suggests that oogenesis and
embryogenesis are composed of modular processes that are relatively siloed from one another. In
particular, our analysis provides strong evidence that (i) pre-antral follicles produce anti-M¨ ullerian
hormone independently of effects from other follicles, (ii) oocytes mature to metaphase-II indepen-
dently of the woman’s age, her BMI, and other factors, (iii) early embryo development is memoryless
for the variables assessed here, in that the probability of an embryo transitioning from its current
developmental stage to the next is independent of its previous stage. Our results both provide
insight into the fundamentals of oogenesis and embryogenesis and have implications for the clinical
practice of IVF.
I. INTRODUCTION
Understanding the manner by which a multicellular
organism develops from a single cell is one of the grand
challenges of biology. In mammals, this process begins
with oogenesis inside the female, which results in an egg
that becomes an embryo after fertilization. Early em-
bryo development in mammals, including humans, is self-
organized [1, 2]: the course of events that unfold are gov-
erned by the embryo’s internal dynamics and can pro-
ceed without external signals. Oogenesis and early em-
bryogenesis have been studied from diverse perspectives,
including molecular genetic, cell biological, chemical, and
mechanical [3–13]. Despite the vast amount of knowledge
that has been obtained, many basic questions remain, in-
cluding: what determines which oocytes are selected for
ovulation? How is the timing of embryonic events regu-
lated? How are oogenesis and embryogenesis negatively
impacted by age and disease? Answering these will pro-
vide fundamental insight and have strong implications for
evolution and medical treatments of infertility. However,
these issues are difficult to study using the molecular ap-
proaches that are the mainstay of current research, be-
cause of the integrated nature of the problems they pose
concerning the overall trajectory of development.
An alternative to the microscopic, molecular perspec-
tive is to attempt to develop a macroscopic, phenomeno-
logical understanding. Such an approach has been pro-
ductive in diverse areas from statistical physics [14] to
economics [15] to some fields of biology [16, 17], includ-
ing protein evolution [18] and cell-size control in bac-
teria [19–21]. One significant concern is that the great
complexity of oogenesis and embryogenesis might make
simple, phenomenological descriptions inapplicable. Fur-
thermore, the validity of the phenomenological approach
can only be determined by developing models and rigor-
ously testing them. This requires a large amount of quan-
titative data, which is difficult to obtain from oocytes and
embryos in model organisms.
Here, we overcome this challenge by leveraging a large
data set from 7,399 routine clinical In-Vitro Fertilization
(IVF) treatment cycles, resulting in 98,264 oocytes and
57,827 embryos. We show that the statistical structure
present in this data can be quantitatively described us-
ing simple, phenomenological models. The models we
develop are Bayesian networks, a form of probabilistic
graphical model which represent conditional dependen-
cies by a directed, acyclic graph. We infer models di-
rectly from the data, making little use of prior knowl-
edge. Despite this, the resulting models recapitulate
well-established aspects of oocyte and embryo develop-
ment. Moreover, the resultant models are sparse: only
one or two factors directly impact physiological processes.
This implies that human oogenesis and embryogenesis
are highly modular. Our analysis leads to a number of
additional, surprising conclusions. We present strong ev-
idence that:
i) each pre-antral follicle produces anti-M¨ ullerian hor-
mone (“AMH”) independently of effects from other folli-
cles. This argues that AMH is a faithful indicator of the
number of pre-antral follicles, consistent with its physio-
logical role in regulating follicle recruitment and further
arXiv:2103.12187v1 [q-bio.CB] 22 Mar 2021
2
supporting its clinical use as a measure of ovarian re-
serves [22–25].
ii) while the number of oocytes released from folli-
cles depends on many factors, the probability that a re-
leased oocyte matures to metaphase-II is independent
of the patient’s age, BMI, and other, external factors.
This argues that physiological processes that are corre-
lated with these external factors, such as mitochondrial
metabolism and aneuploidy [26–28], do not significantly
impact oocyte maturation.
iii) after oocytes are fertilized, the probability of suc-
cessfully transitioning from one embryonic developmen-
tal stage to the next depends on the embryo’s present
state, but not on its state at earlier stages. Thus, em-
bryo development is Markovian, at least for the variables
examined here. This argues that clinical embryo selection
procedures need only consider the state of the embryos
immediately before transfer, as the state of the embryo
at earlier times provides no additional information.
Taken together, our results show that the development
of oocytes and embryos emerges as a simple process, de-
spite the underlying molecular complexities of the biology
and despite the plethora of disease etiologies and treat-
ment protocols presenting in a clinic. More broadly, this
work validates the use of phenomenological models of oo-
genesis and embryogenesis by demonstrating that simple
models can be constructed without sacrificing quantita-
tive accuracy. Although we infer the models using data
drawn from controlled ovarian stimulation and not from
natural menstrual cycles, the models provide insight into
the principles that govern oogenesis and embryogenesis,
and may be useful in guiding clinical IVF treatments.
II. RESULTS
A. Oocyte Development
The formation of a healthy embryo depends on the suc-
cessful progression of an oocyte from prophase-I arrest
to metaphase-II. Thus, we begin by examining what de-
termines the number of metaphase-II oocytes retrieved
during an IVF ovarian stimulation. The clinical data
contain 69 variables that describe either the patient or
the ovarian stimulation. We start by examining four of
the variables that are strongly correlated with the num-
ber of metaphase-II oocytes: (1) the number of total
eggs retrieved during an ovarian stimulation (“Eggs”),
(2) the number of eggs in metaphase-II arrest (“MII”),
(3) the patient’s maximum serum estradiol concentra-
tion during the treatment cycle (“E2”), and (4) the pa-
tient’s serum anti-M¨ ullerian hormone before the cycle
(“AMH”). Estradiol is a hormone produced by the ovaries
during both natural and stimulated ovulatory cycles [29];
AMH is considered a measure of the patient’s ovarian re-
serve [25]. Each of these variables varies widely across
the 4,910 cycles for which all four variables are recorded,
with coefficients of variation of 0.5 – 1.2 (Figure 1a, left).
All four variables are strongly correlated with one an-
other, with correlation coefficients between 0.26 – 0.91
and P-values between 10−300 – 10−155 (Figure 1a, right).
We start by searching for conditional independencies
among the variables AMH, MII, and Eggs. A condi-
tional independency between two variables implies that
one variable can be completely described without direct
knowledge of the other, suggesting that a simple, phe-
nomenological model of the data exists. To search for
conditional independencies, we nonlinearly regress MII
on Eggs and AMH on Eggs, by finding the best-fit poly-
nomial that maximizes the Bayesian posterior evidence.
This method allows for capturing complex dependencies
without overfitting the data [30, 31] (see SI Section 1);
we also split the data into separate train and test sets
as a further check against overfitting. We then take the
residuals from the two regressions and evaluate their cor-
relation. We denote this procedure as Corr(AMH, MII |
Eggs). We find that, although there is a strong correla-
tion between AMH and MII (Figure 1b, left), that corre-
lation disappears after conditioning on Eggs: Corr(AMH,
MII, | Eggs) = -0.02 ( P = 0.19; Figure 1b, center). This
correlation is both consistent with zero and smaller than
an effect size threshold of 0.05, suggesting that AMH and
MII are conditionally independent given Eggs.
We encode this conditional independency using a class
of graphical models known as Bayesian networks. These
have found usage in causal inference [32–35]; here, we use
them to construct phenomenological models that corre-
spond to mechanistic descriptions of biology. Briefly, for
a given factorization of a probability distribution, these
graphs contain a directed edge from one variable to an-
other if the probability of the second variable depends on
the first. Two variables are conditionally independent if
all paths from one variable to the other are “blocked”,
where paths are blocked by head-to-tail or tail-to-tail
nodes meeting at a variable that is conditioned on, or
head-to-head nodes meeting at an edge that is not con-
ditioned on [33, 36]. Both the observed correlation be-
tween MII and AMH and their conditional independency
given Eggs can be captured by any of the graphs AMH
→ Eggs → MII, AMH ← Eggs → MII, or AMH ← Eggs
← MII; we denote this ambiguity by AMH — Eggs —
MII, with an as-yet undetermined orientation of the ar-
rows (Figure 1b, right). If the data are described by one
of these graphs, then Eggs and AMH should remain cor-
related given MII, which is indeed the case: Corr(Eggs,
AMH | MII) = 0.24 ( P = 10 −47; SI Figure 3a). Like-
wise, MII and Eggs should remain correlated given AMH,
which is also the case: Corr(MII, Eggs | AMH) = 0.87
(P < 10−300; SI Figure 3b).
Next, we examine the variables AMH, E2, and Eggs.
There is a strong association between AMH and E2 (Fig-
ure 1c, left), but regressing both AMH and E2 on Eggs
shows that Corr(AMH, E2| Eggs) = 0.03, which is consis-
tent with no conditional correlation (P = 0.11; Figure 1c,
center). This suggests a graph of the form AMH — Eggs
— E2 (Figure 1c, right). Finally, we examine Eggs, MII,
3
Eggs MIIAMH
Eggs
MIIAMH
E2
0.49 0.88
0.08*
0.50
Eggs E2AMH
MII E2
Eggs
Names
Eggs, MII
Eggs, AMH
Eggs, E2
MII, AMH
MII, E2
AMH, E2
Corr.
0.91
0.49
0.50
0.43
0.48
0.26
P-value
<10-300
10-204
10-212
10-157
10-197
10-155
AMH
E2
Metaphase II
Egg
...contain prophase-I-arrested eggs.
...produce
into antral
...mature
Pre-antral
follicles...
Antral follicles...
...produce E2.
...contain eggs
which progress
AMH.
Ovarian Reserve
to MII.
follicles.
...induces a
residual LH
surge through
hypothalamus
& pituitary.
E2 | EggsE2
Counts
0
400
slope = -(2.0 ± 1.6)×10-2
slope = 7.5 ± 4.7
slope = -(2.7 ± 0.6)×10-4
-1000 0 1000
AMH
10 150 5
1000
-1000
0
1000 2000 30000
3000
200 300 300
0 0 0
0 000 20 60 40 6000
AMH Eggs MII E2
10 30 20 3000
MII
0
20
5
15
10
AMH
10 150 5
0
E2
E2 | Eggs-500
500
AMH | Eggs
0 4-4 2 -2
AMH | Eggs
0 4-4 2 -2
1000
2000
4
-4
0
MII | Eggs
2
-2
MII
0
20
5
15
10
500-500
4
-4
0
MII | Eggs
2
-2
P = 0.19
P = 0.11
P = 2 × 10-6
a)
b)
e)
c)
d)
f)
FIG. 1: (a) Distributions and correlations of the four variables AMH, Eggs, MII, and E2. (b) Left: The number of metaphase-II
oocytes retrieved (MII) is strongly correlated with the patient’s serum AMH. Gray dots: raw data, red circles: raw data binned
into 20 separate bins with equal counts, green line and shaded region: nonlinear regression and errors. Center: That correlation
disappears after regressing against the total number of retrieved oocytes (Eggs). Gray dots: residuals after regressing against
Eggs, red circles: residuals binned into 20 separate bins with equal counts, green line and shaded region: linear fit to the
residuals, with slope and standard error shown at top of plot. Right: This conditional independency suggests a graph of the
form AMH – Eggs – MII. (c) Left: The patient’s serum estradiol concentration (E2) is strongly correlated with AMH. Center:
That correlation disappears after regressing against Eggs. Right: This conditional independency suggests a graph of the form
AMH – Eggs – MII. (d) Left: MII is strongly correlated with E2. Center: While regressing against Eggs greatly weakens
that correlation, E2 and MII remain correlated after conditioning on Eggs. This suggests a fully connected graph is needed to
describe these three variables (right). (e) A graphical model that is consistent with the data. Edge labels show the conditional
correlation coefficients after conditioning on all other incoming edges; the data is consistent with the arrow marked with a *
oriented in either direction. (f) The graphical model expected from prior knowledge of ovarian stimulation.
4
and E2. While regressing on Eggs greatly weakens the
correlation between MII and E2 (compare Figure 1d left
and center), the measured conditional correlation is in-
consistent with no correlation: Corr(E2,MII — Eggs) =
0.08, P ≈ 10−6. Thus, an edge must connect each of
Eggs, E2, and MII (1d, right).
Of the 543 graphical models that describe 4 variables,
only 8 graphical models capture exactly the two con-
ditional independencies in the data that are described
above. The data alone cannot distinguish between these
graphs. However, the patient’s AMH is measured before
the ovarian stimulation starts, whereas the other three
variables are measured during the treatment. Thus, any
graph with an edge pointing into AMH cannot corre-
spond to a mechanistic description of the biology. Rul-
ing these graphs out leaves only two graphs consistent
with both the data and a mechanistic interpretation (Fig-
ure 1e).
The quantitative phenomenological model encoded by
these graphs recapitulates our qualitative understanding
of ovarian stimulation (Figure 1f): 1) Pre-antral folli-
cles, which contain immature oocytes and associated so-
matic cells, produce the hormone AMH [22, 24, 37]. Since
pre-antral follicles have the potential to grow into large
antral follicles with prophase-I-arrested eggs, AMH is a
measure of the potential number of oocytes that could
develop. This is captured by the inferred arrow in Fig-
ure 1e from AMH to Eggs, which indicates that the pa-
tient’s AMH determines how many eggs she will produce.
2) Antral follicles produce estradiol, captured by the in-
ferred arrow from Eggs to E2. 3) Some, but not all, eggs
progress from prophase-I arrest to metaphase-II arrest.
This is captured by the inferred arrow from Eggs to MII.
4) During natural ovulation, the estradiol produced by
antral follicles signals the pituitary and hypothalamus to
release hormones which modulate oocyte maturation. In
ovarian stimulation protocols, clinicians attempt to tem-
porarily disable this feedback between the hypothalamus
and the pituitary [29], suggesting that estradiol should
not impact oocyte maturation during ovarian stimula-
tion. However, the inferred arrow from E2 to MII sug-
gests that a weak feedback between the hypothalamus,
the pituitary, and estradiol is still present during an ovar-
ian stimulation cycle.
The inferred phenomenological model (Figure 1e) pro-
vides a quantitative representation of oocyte develop-
ment that allows direct and indirect effects to be disen-
tangled. For instance, a patient starting treatment with
a higher AMH is likely to produce more oocytes, via the
direct arrow AMH → Eggs. In addition, that patient is
likely to have a higher estradiol level during the cycle, as
the additional eggs she is likely to produce will on average
produce more estradiol, via the path AMH → Eggs →
E2. However, the graph states that this effect is indirect:
the patient’s E2 increases only through the associated
increase in the number of eggs for high-AMH patients.
This is borne out by the data. Likewise, a patient with a
larger number of MII oocytes retrieved is likely to have a
higher AMH, since following the arrows backwards shows
that higher MII implies that Eggs is higher, and higher
Eggs implies a higher AMH. However, once again, this
is an indirect effect; the patient’s AMH is more likely to
be high only because of the associated increase in Eggs
when many MII oocytes are retrieved.
B. Additional effects on Oocyte Development
How do factors other than these four variables affect
oocyte development? First, we examine the effect of the
woman’s age on an oocyte’s ability to reach metaphase-
II. The data show no conditional correlation between MII
and the woman’s age: Corr(Age, MII | Eggs, E2) = 0.02
(P = 0 .35, Figure 2a). In addition, the data constrain
the magnitude of any effect to be tiny. Fitting a line to
the residuals constrains the slope to be (1 .1 ±1.2) ×10−2
MII oocytes / year (point estimate ± standard error).
To place this in perspective, consider a typical treatment
cycle for a 33 year old and a 40 year old woman (the 25th
and 75th percentile in the data). If the treatment results
in the same number of total oocytes and the same max
estradiol for both women, then on average the number
of retrieved MII oocytes should differ by no more than
0.2. Since the median MII oocytes retrieved per cycle is
8, the data constrain the direct effect of age on MII to be
2% or less. Next, we examine the effect of the woman’s
BMI on the oocyte. The data show no conditional cor-
relation between MII and BMI: Corr(BMI, MII | Eggs,
E2) = -0.005 ( P = 0.78; Figure 2b). Moreover, the data
constrain the direct effect of BMI on MII to be 1% or
less. Finally, the data show no conditional correlation
between MII and the dose of either of the ovarian stimu-
lation drugs FSH or HMG: Corr(FSH, MII | Eggs, E2) =
0.01 (P = 0.49; Figure 2c), Corr(HMG, MII | Eggs, E2)
= -0.03 ( P = 0.08; Figure 2d). Likewise, the data con-
strain the direct effect of these stimulation drugs to be
1% or less for FSH, and 5% or less for HMG. These obser-
vations show that oocytes develop to MII independently
of a patient’s age, her BMI, or details of the ovarian stim-
ulation procedure.
Does the probability of an oocyte being in MII arrest
depend on the total number of eggs retrieved? While
Eggs is strongly predictive of MII, it provides no predic-
tive power for the fraction of eggs in MII arrest (MII /
Eggs): Corr(MII /Eggs, Eggs) = −0.01 (P = 0.48). Lin-
early regressing MII / Eggs on Eggs gives a slope tightly
constrained near zero (Figure 2e).
Combined, these observations suggest the following
simple picture for oocyte maturation: Each follicle in-
dependently triggers its oocyte to leave prophase-I, with
a probability that depends only on E2. The oocyte then
progresses to metaphase-II, with both processes indepen-
dent of interactions with other follicles or the aggressive-
ness of the ovarian stimulation.
To check whether this simple picture completely de-
scribes oocyte maturation, we examine the distribution
5
FSH | Eggs, E2
Age | Eggs, E2 BMI | Eggs, E2
HMG | Eggs, E2
slope = (1.1 ± 1.2)×10-2 slope = (-2.1 ± 7.7)×10-3
slope = (-6.4 ± 2.8)×10-3slope = (1.5 ± 3.0)×10-3
Fraction of MII Eggs
Eggs
0.0
0.75
1.0
0.5
0.25
0 10 30 20
slope = (-2.8 ± 4.0)×10-4
0
10
30
20
MII
0 6 17 12
Counts
MII | Eggs, E2 MII | Eggs, E2
MII | Eggs, E2MII | Eggs, E2
-4
0
4
2
-2
-10 0 10-5 5
-20 0 20-10 10
-10 0 10-5 5
-4
0
4
2
-2
-20 0 20-10 10
-4
0
4
2
-2
-4
0
4
2
-2
P = 0.35 P = 0.79
P = 0.62 P = 0.02
P = 0.48
MIIAge BMI
FSH HMG
MII
MII MII
a)
c)
e)
b)
d)
f)
FIG. 2: MII versus the patient’s age (a), BMI (b), and the doses of stimulation drugs FSH (c) and HMG (d), after regressing
against Eggs and E2. Gray dots show the residuals from the regressions, red circles and error bars show the mean and standard
error of the data binned into 20 bins with equal number of points, and green lines, shaded regions, and labeled slopes show
the mean and standard error of the best linear fit to the residuals. (e) The fraction of MII oocytes (MII / Eggs) vs Eggs. (f)
Histogram of observed MII (red circles) vs that expected from independently-triggering follicles (green line and shaded region
show the expected counts and their standard deviation), for the 116 cycles that have 17 eggs retrieved, which is where the
discrepancy between the two histograms is the largest as measured by a χ2 test (P <10−300, primarily due to the cycles with
MII of 1–5). On the scale of the plot, the expected histograms from a binomial distribution where the probability for an oocyte
being metaphase-II is constant is indistinguishable from one where the probability varies with E2.
of MII. If each follicle independently triggers its egg to
progress to metaphase-II with a probability p, then MII
for each cycle should be binomially distributed, denoted
as B(MII; Eggs, p). If that probability depends only on
E2, then for fixed Eggs the measured distribution of MII
across cycles should be the average of many binomial
distributions, each with a probability that depends on
E2: ⟨B(MII; Eggs, p(E2))⟩E2. Instead, the empirical dis-
tribution of MII is much broader than the expected one
(Figure 2f). This discrepancy suggests that additional
factors affect oocyte maturation. These additional fac-
tors could be biochemical processes within the oocyte
that are shared by multiple eggs from the same patient,
interactions between follicles beyond serum estradiol, or
simply other clinical factors that we have not accounted
for.
The data provide some insight into human meiosis, es-
pecially given prior knowledge of human ovulation. As a
woman ages, the oocytes she ovulates become much more
likely to be aneuploid, rising from an aneuploidy rate of
roughly 25% at age 30 to 80% at age 42 [26, 38, 39].
Chromosomal signatures show that aneuploidy in hu-
man oocytes can arise both during the oocyte’s pro-
gression from prophase-I to metaphase-II and immedi-
ately after fertilization [11, 40–42]. In mitotic cells, mis-
segregation of chromosomes is reduced by the spindle-
assembly checkpoint [43], which can arrest mitosis un-
til chromosomes are correctly lined up. If there were
a strong spindle assembly checkpoint in meiosis I, then
the typically aneuploid oocytes from older women would
reach metaphase-II at a reduced incidence than those
from younger women. Instead, oocytes from older and
younger women reach metaphase-II arrest at the same in-
cidence. This is consistent with experimental work that
shows that human oocytes have a weak meiotic spindle
assembly checkpoint [44–46].
Next, we ask how the rest of oocyte recruitment is af-
fected by the patient’s age, BMI, and the doses of ovarian
6
Patient
Oocytes
Hormones
Rank
1 20 40 60 80 99
P-Value
10-3
10-2
10-1
100
Rank
1 20 40 60 80 99
P-Value
10-3
10-2
10-1
100
Model in (a)
Observed, Train
Observed, Test
Fully-Connected Model
Observed, Train
Observed, Test
Eggs MII
AMH
E2
Age
BMI
FSH
HMG
0.07*
-0.29* 0.17,
0.26 0.110.07,
-0.32,
-0.24
0.05*
0.37
-0.12
-0.11
0.09
0.47 0.08*
0.88
Prognostic
Variables
Treatment
Variables
Response
Variables
—,
-0.07,
—
a) b)
c)
FIG. 3: (a) A graphical model of oogenesis that is consistent with a mechanistic interpretation of the data. Labels show
conditional correlation coefficients. For edges with two labels, the upper corresponds to FSH, the lower to HMG; dashes signify
no dependence. The data is consistent with marked arrows (*) oriented in either direction. (b) Rank plots of P-values for the
99 conditional correlations corresponding to the conditional independencies predicted by the model. The green line shows that
from the training data, the red line that from the test data. The black line and shaded regions show the median, 95%, and
99.9% centered percentile of rank plots from 3,000 datasets simulated according to the proposed model. (c) The same as (b),
but showing the distribution of rank plots for the 99 conditional correlations from fully-connected, linear Gaussian models.
stimulation drugs FSH and HMG. We examine the joint
distribution of all eight variables and construct a directed
acyclic graph that is consistent with the data. The num-
ber of possible graphs grows rapidly with the number of
variables: there are 543 possible graphs for four variables,
but 783,702,329,343 possible graphs with eight variables.
To deal with this inordinately large number of graphs,
we use prior knowledge and split the variables into three
groups: prognostic variables measured before the treat-
ment starts (Age, BMI, and AMH), treatment variables
(FSH and HMG), and response variables measured after
the drugs have been applied (Eggs, MII, E2). We then
search for graphs that are consistent with a mechanistic
interpretation, by excluding graphs with edges directed
from treatment to prognostic variables, from response to
prognostic variables, or from response to treatment vari-
ables.
We start by examining the prognostic variables. There
is one conditional independency among the prognostic
variables: Corr(AMH , BMI | Age) = −0.04 ( P = 0 .01;
SI Figure 4a). Physiologically, this implies that the like-
lihood that a primordial follicle develops into a pre-antral
follicles is independent of any hormonal effects of obesity.
Next, we examine the treatment variables. The FSH
dosage depends on all three prognostic variables, and the
HMG dosage depends on Age, AMH, and FSH, but not
BMI (SI Figure 4f). The complex dependencies of FSH
and HMG on other variables imply that clinicians cus-
tomize the treatment based on the patient, which is to
be expected. Nevertheless, the lack of conditional in-
dependencies between the treatment variables and the
prognostic variables demonstrates that the conditional
independencies we do see elsewhere are real and not an
artifact of our analysis technique.
Finally, we examine the response variables. Af-
ter conditioning on AMH and HMG, Eggs is
conditionally independent of both Age and BMI
(Corr(Age, Eggs | AMH, HMG) = −0.04, P = 0 .01;
Corr(BMI, Eggs | AMH, HMG) = −0.01, P = 0 .54;
SI Figure 4c,d). Physiologically, this suggests
that the ability of a pre-antral follicle to be re-
cruited does not worsen with age or obesity. Like-
wise, Eggs is conditionally independent of FSH:
Corr(FSH, Eggs | AMH, HMG) = −0.04 ( P = 0 .02;
SI Figure 4e). This suggests that clinicians pre-
scribe sufficient FSH to recruit all the follicles in the
cohort activated from the primordial pool in that
menstrual cycle. Interestingly, we observe a weak
negative correlation between Eggs and the dose of
HMG: Corr(HMG , Eggs | AMH) = −0.11 ( P = 10 −11).
7
Taken at face value, this implies that HMG is typically
supplied at more than the optimal dose. While there is
some evidence that excessive HMG can cause follicles to
degrade [47, 48], another possibility is that the negative
correlation is due to clinicians prescribing more HMG to
patients whom they know a priori to be poor responders
even after accounting for their age, BMI, and AMH – for
instance, patients who have had a poor response in pre-
vious treatments. However, the correlation remains neg-
ative even when excluding patients on repeat stimulation
cycles: Corr(HMG , Eggs | AMH, first cycle) = −0.09,
P ≈ 10−5. Combined, the data paint a simple picture
for oocyte recruitment: all available follicles are typically
recruited, independently of effects from age or obesity
but weakly affected by HMG. Finally, in contrast to
the simplicity of Eggs and MII, E2 depends on all of
Age, BMI, FSH, HMG, and Eggs. These combined
observed conditional independencies are captured by the
graphical model in Figure 3a.
The model in Figure 3a predicts 99 conditional inde-
pendencies among the 8 variables shown. We measure
the conditional correlations and associated P-values for
each of these 99 independencies in both the train and test
sets, and compare the distribution of these P-values to
those calculated from 3,000 datasets simulated according
to the model in Figure 3a. The distribution of the P-
values measured from the data is consistent with what is
expected if the model is true, although the lowest P-value
for the data is lower than that typical for the simulated
data (Figure 3b and SI Section 4). In contrast, datasets
generated according to fully-connected models display a
different distribution of these P-values (Figure 3c and SI
Section 4). Moreover, anything missing from the model
in Figure 3a must correspond to a small effect with small
explanatory power. Fitting the training data with a fully-
connected model explains 0.6% or less of each variable’s
variance in the test data, with the fully-connected model
actually performing worse on the test set for most of the
fits than the model in Figure 3a does. Combined, these
observations show that the graphical model accurately
describes human ovarian physiology and oogenesis.
Viewed holistically, the probabilistic graphical model
in Figure 3 has a simple mechanistic interpretation. The
patient’s age and BMI only act to determine the hor-
mone levels. Hormones other than estradiol determine
how many antral follicles develop. The follicles then pro-
duce estrogen and trigger eggs to progress to metaphase-
II, with a slight feedback between those two. Because
of this simplicity, changes due to age or obesity manifest
themselves in simple ways, after ignoring the physiolog-
ically irrelevant question of how clinicians choose drug
doses for ovarian stimulation. In particular, the only di-
rect effect of obesity on the oocyte maturation process is
to decrease E2. This is consistent with work showing that
obesity affects fertility by interfering with the hormonal
regulation of ovulation [28]. The primary effect of age on
the oocyte maturation process is to decrease the ovarian
reserve, as measured by the patient’s AMH. Physiologi-
Var(AMH)Var(E2) / 106
AMH
E2
Eggs
0 10 20 30
Eggs
0 10 20 30
0
2
4
6
8
0
10
20
30
0
1000
2000
3000
4000
0
1
2
Eggs
0 10 20 30
Eggs
0 10 20 30
a)
c) d)
b)
FIG. 4: (a) AMH vs Eggs. Gray dots show raw data, red
circles show the mean and standard error of AMH binned at
each value of Eggs, and green line shows the best linear fit to
the data, with slope of 0.20±0.01 and intercept 0.36±0.11. A
linear fit is the best-fit polynomial to the data, as determined
by the model evidence. (b) Variance and error estimate of
AMH vs Eggs, trimmed to the central 95% for each value
of Eggs (red dots, errors calculated using the variance of the
k-statistic). The green line shows the best linear fit to the
variance, with a slope of 0.28±0.01 and intercept−0.22±0.03.
(c) E2 vs Eggs. The green line shows the best linear fit to the
data. A quadratic model (not shown) provides the best fit
to E2 vs Eggs, as determined by the model evidence. (d)
Variance and error estimate of E2 vs Eggs, trimmed to the
central 95% for each value of Eggs.
cally, this is consistent with the well-known decrease in
a woman’s pool of primordial follicles as she ages [29].
AMH and E2 are produced by ovarian follicles at differ-
ent stages in follicular growth. The data provide insight
into both these processes. For a given number of retrieved
oocytes, the mean of AMH is linear in the number of re-
trieved oocytes, with an intercept near zero (Figure 4a,
also SI Figure 5). For Eggs not too large, the variance
of AMH is also linear in Eggs (Figure 4b); the devia-
tion from linearity at large Eggs is due to patients with
polycystic ovary syndrome, who tend to have high AMH.
Since means and variances add when summing indepen-
dent random variables, the linearity of both the mean and
the variance of AMH in the number of follicles suggests
that each follicle produces AMH independently from in-
teractions with other follicles. (The small but nonzero
intercept could arise if some pre-antral follicles do not
mature sufficiently to be retrieved during the IVF re-
trieval procedure.) In contrast, the mean E2 is not linear
in Eggs, systematically deviating from linearity and hav-
ing an intercept that is far from zero (Figure 4c; variance
in panel d). These variations from linearity show that
the E2 is not produced independently by each follicle,
8
perhaps due to additional production from outside fol-
licles (such as by adipose tissue or the adrenal glands),
due to inter-follicle feedback, or simply due to external
factors that affect E2 production, such as those shown in
in Figure 3.
C. Pre-implantation development
Once the oocyte is fertilized, the resulting embryo
starts to divide. By the third day after fertilization, a
human embryo typically has 8 cells. As the cells con-
tinue to divide, on the fifth day the embryo differen-
tiates into a blastocyst, composed of two distinct cell
lineages: the trophectoderm and the inner-cell mass.
In natural fertilizations, the blastocyst then attaches to
the woman’s endometrial epithelium and implants in her
uterus [10, 12, 29, 49]. In the IVF clinic, human embryos
are typically cultured for 3 or 5 days after fertilization,
at which point embryologists attempt to select the high-
est quality embryo(s), which is then transferred into the
patient’s uterus.
To construct a phenomenological model of develop-
ment, we first consider three variables that describe the
overall trajectory of development: the embryo’s number
of cells on day 3 after fertilization (Day 3 Cells), its de-
velopmental stage on day 5 (Day 5 Stage), and whether
it resulted in a fetal heartbeat after transfer (FH; see SI
for details).
However, not all embryos are cultured to day 5, and
thus not all embryos have data on both day 3 and day 5.
Since the decision to culture embryos to day 5 is made
based on patient prognosis and embryo quality, embryos
that are assessed on day 5 systematically differ from those
that are not. To avoid biases due to this missing data,
we treat missingness as an additional variable and model
both the variable and its missingness [50]. The clini-
cal data are in the “missing at random” regime, where
a variable’s missingness depends on other variables in
the dataset, but not on the missing variable itself. In
the language of probabilistic graphical models, there are
edges from some of the normal variables to the missing-
ness variables, but no edge from a variable to its own
missingness. In this regime, valid inferences require con-
ditioning on missingness and the variables on which the
missingness depends. Multiple transfers cause an addi-
tional problem for the measurement of fetal heartbeat –
if two embryos are transferred simultaneously and one fe-
tal heartbeat is observed, it is not obvious which embryo
formed the fetus. We solve this problem via generative
modeling. Briefly, we construct a parameterized model
that predicts the probability of one embryo implanting
from properties of the embryo and the woman. We then
fit the model to the data by finding the maximum a pos-
teriori parameters, using a Poisson-binomial likelihood
for multiple transfers. To check whether a variable is
conditionally independent of fetal heartbeat, we fit two
models, one with the additional variable and one without,
and perform Bayesian model selection to see if the addi-
tional variable is necessary (SI Section 1). We include all
available cycles with four or fewer embryos transferred
(95% of cycles).
We use this approach to construct a model
of human pre-implantation development.
The embryo’s number of cells on day 3 is
strongly correlated with its stage on day 5:
Corr(Day 3 Cells , Day 5 Stage | both measured, Age, BMI, MII) =
0.44 ( P < 10−300, Figure 5a). The embryo’s Day 3
Cells is also predictive of fetal heartbeat (Figure 5b), as
is appreciated in the literature [51, 52]. Likewise, the
embryo’s Day 5 Stage is predictive of fetal heartbeat
(Figure 5c) [53]. However, the probability that a
transferred embryo results in a fetal heartbeat depends
only on its stage on day 5; Day 3 Cells provides no
additional information whether the embryo will develop
(Figure 5d). This conditional independence implies a
model of the form Day 3 Cells → Day 5 Stage → FH;
this is the only graph with two or fewer edges that is
consistent with the data and the fact that day 3 happens
before day 5.
To understand how properties of the woman and
her ovarian stimulation affect embryo development, we
examine the effect of three other variables: the woman’s
age, her BMI, and the number of MII oocytes retrieved
in the ovarian stimulation. The woman’s age is weakly
correlated with the embryo’s number of cells on day
3 and its stage on day 5: Corr(Day 3 Cells , Age) =
−0.07 ( P = 10 −47; SI Figure 6a),
Corr(Day 5 Stage , Age | Day 3 Cells , measured) =
−0.11 ( P = 10 −78; SI Figure 6b). In contrast, the
woman’s age has a much stronger effect on the probabil-
ity that an embryo forms a fetal heartbeat (SI Figure 7c).
Surprisingly, the patient’s BMI is conditionally uncorre-
lated with all of Day 3 Cells (SI Figure 6c), Day 5 Stage
(SI Figure 6d), and FH (SI Figure 7e). Likewise, the
conditional correlation of MII with all of D3 Cells, D5
Stage, and FH is either consistent with zero or less than
0.05 in magnitude (SI Figure 6e,f and 7f). Combined,
these observations yield the simple graphical model for
development in Figure 5e. The woman’s age determines
the embryo’s cell number on day 3; her age and the
embryo’s cell number determine the embryo’s stage on
day 5; and her age and the embryo’s stage determine
whether it will continue to develop. Neither the patient’s
BMI nor the number of retrieved MII eggs significantly
affect the embryo’s development. In contrast, the data’s
missingness shows a much more complex distribution,
with almost all variables affecting the clinical decisions
regarding embryo transfer and culture duration. Once
again, the data paint a picture of simple biology but
complex clinical decisions.
The lack of a directed edge from the number of cells
on day 3 to fetal heartbeat is particularly striking. The
time-varying morphology (morphokinetics) of human em-
bryos is known to be predictive of developmental suc-
cess [29, 52]. In principle, morphokinetics up to day 3
9
could provide a different set of information than mor-
phokinetics between days 3 and 5. For example, since the
embryo’s genome activates on day 3 [54], one might rea-
sonably propose that the embryo’s progress before day 3
provides information about the oocyte’s cytoplasm, that
the embryo’s progress from day 3 to day 5 provides in-
formation about aneuploidy, and that both of these fac-
tors independently determine the embryo’s prognosis. In-
stead, the data show that the embryo’s stage at day 5
contains all the information about its viability that its
stage at day 3 contains.
One physiological interpretation consistent with the
data is that there are two distinct sets of mechanisms that
influence an embryo’s development (Figure 5f). One set
of mechanisms influences pre-implantation development
only through its overall rate, including the number of
cells on day 3 and the stage on day 5. This set of mech-
anisms depends weakly on age. Another set of mecha-
nisms influences an embryo’s post-implantation develop-
mental potential and depends strongly on age. Perhaps
surprisingly, meiotic aneuploidy of the oocyte cannot
strongly affect pre-implantation development, since mei-
otic aneuploidy is strongly associated with the woman’s
age. Instead, the data suggest that meiotic aneuploidy
primarily affects post-implantation development, rather
than pre-implantation development. Other works have
provided mixed evidence whether aneuploidy affects pre-
implantation development [55–58].
To see how robust this picture is, we augment the anal-
ysis with additional information on the embryo’s pre-
implantation development. In addition to the number
of cells, the dataset also describes the embryo on Day
3 with the presence of cytoplasmic fragments, the pres-
ence of multiple nuclei in individual cells, size asymme-
tries among cells within the embryo, the presence of large
vacuoles, and the granularity of the cell cytoplasm. On
day 5, the dataset also includes grades of the inner-cell
mass and the trophectoderm of embryos that have formed
blastocysts. Of the day 3 variables, embryo fragmenta-
tion and cell symmetry are predictive of fetal heartbeat,
along with Age and Day 3 Cells. Likewise, of the day 5
variables, the trophectoderm grade is predictive of fetal
heartbeat, along with Age and Day 5 Stage, in agreement
with recent studies [59, 60]. (The data are consistent with
the other day 3 variables providing no additional predic-
tive power once Age, D3 Cells, symmetry, and fragmen-
tation are known; and the data are consistent with the
inner-cell mass grade providing no additional predictive
power once Age, D5 stage, and the trophectoderm grade
are known.) Nevertheless, the day 3 variables provide no
additional predictive power for fetal heartbeat once the
day 5 variables are known (SI Section 4). These obser-
vations suggest a memoryless model of pre-implantation
development: provided the embryo makes it to the blas-
tocyst stage, what happened before is irrelevant for its vi-
ability. Thus, despite the molecular complexities of early
development and the complicated trajectory of human
development before 12 weeks, a simple, coarse-grained
view of embryonic viability may be possible without sac-
rificing quantitative accuracy.
III. DISCUSSION
Here, we have used clinical IVF data and minimal prior
knowledge to infer quantitative, phenomenological mod-
els of human oogenesis and embryogenesis. Not only does
constructing these models with a data-driven approach
give confidence in their validity, but the models reca-
pitulate known aspects of oogenesis and embryogenesis
as well. Surprisingly, the models that best describe the
data are sparse, with only one or two factors affecting
most physiological processes. This suggests that oogene-
sis and embryogenesis are modular processes. Moreover,
our analysis leads to three additional, surprising conclu-
sions which support this overall picture of modularity.
i) AMH production by one pre-antral follicle is inde-
pendent of the amount produced by others. This is in
stark contrast to other hormones produced by the folli-
cle. Hormones such as estradiol and inhibin participate
in feedback loops which regulate the formation of the
dominant follicle in a natural cycle. As a result, these
hormones are produced in a highly regulated manner,
and not independently by each follicle [29, 61]. More-
over, mathematical modeling suggests that, for these
feedback loops to function, the hormone production and
response needs to be highly nonlinear [62–64]. Thus, the
amount of these hormones produced by one follicle de-
pends strongly on the amount produced by other folli-
cles. In contrast, AMH appears to be produced without
regulatory feedback. This is particularly interesting be-
cause AMH also regulates the number of growing folli-
cles in a cohort, by regulating the growth of primordial
follicles into primary (pre-antral) follicles. Thus, while
feedback loops are needed to accurately control the num-
ber of mature, Graafian follicles recruited during natu-
ral ovulation, feedback loops appear to be unnecessary
to sufficiently control the number of primary follicles re-
cruited.
ii) Neither age, obesity, the ovarian stimulation, nor
even the number of recruited oocytes affects whether an
individual oocyte progresses to metaphase-II once trig-
gered to resume meiosis. These observations have sev-
eral implications for the biology of the oocyte. Defini-
tive evidence shows that age is strongly correlated with
aneuploidy in the oocyte [26, 38]. However, since our
analysis shows that the ability of an oocyte to progress
to metaphase-II is independent of age, we conclude that
this ability is the same for both euploid and aneuploid
oocytes. Thus, the spindle assembly checkpoint in human
oocytes must be weak, in agreement with recent exper-
imental work [44–46]. A similar argument can be made
regarding the effect of metabolism on meiosis. Evidence
suggests that oocyte mitochondrial metabolism worsens
with increasing obesity [28, 65]. Since oocytes progress
to metaphase-II independently of obesity, metabolic de-
10
Day 5 Stage
Day 5 Stagelogit P(FH | Day 5 Stage)
1 3 5 7 9
P(Fetal Heart)
1.0
0.75
0.5
0.25
0.0
P(Fetal Heart)
1.0
0.75
0.5
0.25
0.0
1
3
5
7
9
Day 3 Cells
0 4 8 12 16
Day 3 Cells
0 4 8 12 16
Day 3 Cells
0 4 8 12 16
-2
0
2
Cycle Variables Embryo Variables
Missingness Variables
Age
BMI
0.09 -0.07 -0.11
0.44MII
-0.25
Trans.D5
Rec.
D5
Stage Fetal
Heart
D3
Cells
0.14-0.04-0.641.49
1.21, -0.09 —, 1.61
0.45
-1.25, -0.64
0.92,
0.63
-0.40
metabolic defects?
oxidative damage?
DNA damage?
....
meiotic
aneuploidy?
Post-implantation
Development
Progress of
Pre-implantation Development
Woman'sAge
a) e)
f)
b)
c)
d)
FIG. 5: (a) Day 5 Stage vs Day 3 Cells. Gray dots show the raw data, red circles show the mean Day 5 Stage for each
separate value of Day 3 Cells, the green line and shaded region show the nonlinear model with the highest evidence and its
uncertainty. (b) Estimated probability of an embryo resulting in a fetal heartbeat (FH) as a function of Day 3 Cells alone, for
embryos recorded on day 3 and transferred. The red circles and error bars show the probability estimated by a model that
fits an independent probability of implantation for each number of cells; the green line and shaded region shows the nonlinear
model with the highest model evidence and its uncertainty. (c) The estimated probability of FH as a function of Day 5 Stage
alone, for embryos recorded on day 5 and transferred. (d) The logit of the estimated probability of FH as a function of Day
3 Cells, after regressing against Day 5 Stage. Red circles and error bars show the additional log probability estimated from a
model that fits an independent logit for each value of Day 3 Cells; green line shows the best linear model and uncertainty for
the logit. The data is consistent with Day 3 Cells having no additional predictive power on FH once Day 5 Stage is known. (e)
The graph with the minimal number of edges that is consistent with the data and a mechanistic interpretation. Black nodes
and arrows show the measured data; gray nodes and arrows show the data’s missingness. Only the arrow Age → BMI can be
re-oriented without breaking consistency with the data or a mechanistic interpretation. Edge labels for continuous variables are
conditional correlation coefficients (roman typeface). Edge labels for discrete variables are coefficients from logistic regression
(italic typeface), after treating the effects of other variables with edges into the discrete variable and after normalizing the input
variable by its mean and standard deviation. The missingness variables D5 Rec. and Trans. are 1 if the embryo is recorded
on day 5 or transferred, respectively, and 0 if that information is missing. The distribution of Trans. changes depending on
whether Day 5 Stage was recorded; the two labels on edges into Trans correspond to Day 5 Stage missing or recorded. (f) The
data is consistent with a picture where processes which control pre-implantation development are largely different from those
which control post-implantation development.
fects must not typically be enough to stop an oocyte from
progressing from prophase-I to metaphase-II.
iii) Early embryonic development is memoryless, in
that embryos with the same status on day 5 develop the
same, regardless of their status on day 3. This memo-
rylessness is reminiscent of the robustness of early mam-
malian embryos to damage to individual cells [66–68];
however, memorylessness is more than robustness. Ro-
bustness signifies that an embryo can recover from a set-
back. Memorylessness signifies that, once recovered, nei-
ther the setback nor what caused it has any impact on
the rest of development. The memorylessness implies a
modularity in development [69].
The models we present also have implications for clin-
11
ical IVF. For embryos transferred on day-5, embryo se-
lection can be based solely on how developed they are on
day 5, independent of their status on day 3. While day
3 information is correlated with implantation potential,
the effect is completely captured by the embryo’s sta-
tus on day 5. For ovarian stimulation, the data provide
no evidence that aggressive ovarian stimulation is detri-
mental to the oocyte, either in its ability to mature to
metaphase II, to develop as an embryo, or, if transferred,
to form a viable pregnancy (Figures 3 & 5); moreover,
the data constrain any of these effects to be small. Thus,
a clinic should not be concerned about a potential trade-
off between the quality and quantity of retrieved oocytes.
Conversely, the data show that, at the clinic from which
our data were derived, the ovarian stimulation drugs FSH
and HMG are typically applied at saturating or slightly
deleterious doses for follicular recruitment. Thus, the
hormone dosage could presumably be slightly reduced,
to mitigate side effects such as ovarian hyper-stimulation
syndrome or the high cost of stimulation drugs, without
a large decrease in the number of retrieved oocytes.
More broadly, our results provide fundamental insights
into the overall process of development. Unlike the
sparseness that arises in theoretically-motivated models
simply as a way to manage complexity, the sparseness
in our models is a property of the data. This sparse-
ness implies that the biology itself is simple, consisting
of modularized processes that are quantitatively siloed
from one another. The simplicity is particularly surpris-
ing given the many ways that oogenesis and embryogen-
esis are affected by diseases, including age-related infer-
tility, endometriosis, sperm malfunction, and polycystic
ovary syndrome, all of which are present in our dataset.
Overall, our results suggest that, despite their underlying
complexities, oogenesis and embryogenesis are modular
processes that result in simple, emergent behavior, and
that a concise, quantitative understanding of the rest of
development may be possible.
IV. ACKNOWLEDGMENTS
We would like to acknowledge K. Walden, S. Kou, V.
Manoharan, and the Needleman lab for useful discus-
sions. BL acknowledges the Harvard Quantitative Bi-
ology Initiative for support and useful discussions. DN
and CR acknowledge the National Institutes of Health
for financial support (R01HD092550-01).
[1] M. Zhu and M. Zernicka-Goetz, Cell 183, 1467 (2020).
[2] S. Wennekamp, S. Mesecke, F. N´ ed´ elec, and T. Hiiragi,
Nature reviews Molecular cell biology 14, 452 (2013).
[3] Y. Song and S. Y. Shvartsman, Trends in Genetics
(2020).
[4] P. Gross, K. V. Kumar, and S. W. Grill, Annual review
of biophysics 46, 337 (2017).
[5] C. J. Chan, C.-P. Heisenberg, and T. Hiiragi, Current
Biology 27, R1024 (2017).
[6] J. Fiorentino, M.-E. Torres-Padilla, and A. Scialdone,
Annual Review of Genetics 54, 167 (2020).
[7] D. Clift and M. Schuh, Nature reviews Molecular cell
biology 14, 549 (2013).
[8] M. N. Shahbazi, E. D. Siggia, and M. Zernicka-Goetz,
Science 364, 948 (2019).
[9] J. Rossant and P. P. Tam, Development136, 701 (2009).
[10] M. D. White, J. Zenker, S. Bissiere, and N. Plachta, De-
velopmental cell 45, 667 (2018).
[11] T. Hassold and P. Hunt, Nature Reviews Genetics 2, 280
(2001).
[12] K. K. Niakan, J. Han, R. A. Pedersen, C. Simon, and
R. A. R. Pera, Development 139, 829 (2012).
[13] L. A. Jaffe and J. R. Egbert, Annual review of Physiology
79, 237 (2017).
[14] J. Sethna et al., Statistical mechanics: entropy, order
parameters, and complexity , vol. 14 (Oxford University
Press, 2006).
[15] J.-P. Bouchaud and M. Potters, Theory of financial risk
and derivative pricing: from statistical physics to risk
management (Cambridge university press, 2003).
[16] W. Bialek, Biophysics: searching for principles (Prince-
ton University Press, 2012).
[17] D. Needleman and Z. Dogic, Nature Reviews Materials
2, 1 (2017).
[18] N. Halabi, O. Rivoire, S. Leibler, and R. Ranganathan,
Cell 138, 774 (2009).
[19] S. Taheri-Araghi, S. Bradde, J. T. Sauls, N. S. Hill, P. A.
Levin, J. Paulsson, M. Vergassola, and S. Jun, Current
biology 25, 385 (2015).
[20] J. T. Sauls, D. Li, and S. Jun, Current opinion in cell
biology 38, 38 (2016).
[21] M. Kohram, H. Vashistha, S. Leibler, B. Xue, and
H. Salman, Current Biology (2020).
[22] C. Weenen, J. S. Laven, A. R. von Bergh, M. Cranfield,
N. P. Groome, J. A. Visser, P. Kramer, B. C. Fauser,
and A. P. Themmen, MHR: Basic science of reproductive
medicine 10, 77 (2004).
[23] A. La Marca, F. Broekmans, A. Volpe, B. Fauser, and
N. Macklon, Human Reproduction 24, 2264 (2009).
[24] A. Durlinger, J. Visser, and A. Themmen, Reproduction
(2002).
[25] A. La Marca, G. Sighinolfi, D. Radi, C. Argento,
E. Baraldi, A. C. Artenisio, G. Stabile, and A. Volpe,
Human reproduction update 16, 113 (2010).
[26] J. M. Franasiak, E. J. Forman, K. H. Hong, M. D.
Werner, K. M. Upham, N. R. Treff, and R. T. Scott Jr,
Fertility and sterility 101, 656 (2014).
[27] A. Talmor and B. Dunphy, Best practice & research Clin-
ical obstetrics & gynaecology 29, 498 (2015).
[28] D. E. Broughton and K. H. Moley, Fertility and sterility
107, 840 (2017).
[29] K. Elder and B. Dale, In-vitro fertilization (Cambridge
University Press, 2020).
[30] D. J. MacKay and D. J. Mac Kay, Information theory,
12
inference and learning algorithms (Cambridge university
press, 2003).
[31] D. J. MacKay, Neural computation 4, 415 (1992).
[32] N. Friedman, M. Linial, I. Nachman, and D. Pe’er, Jour-
nal of computational biology 7, 601 (2000).
[33] J. Pearl, Causality (Cambridge university press, 2009).
[34] P. B¨ uhlmann, M. Kalisch, and L. Meier, Annu. Rev. Stat.
Appl. 1, 255 (2014).
[35] N. Friedman, Science 303, 799 (2004).
[36] C. M. Bishop, Pattern recognition and machine learning
(springer, 2006).
[37] L. Pellatt, S. Rice, and H. D. Mason, Reproduction 139,
825 (2010).
[38] J. D. Erickson, Annals of human genetics 41, 289 (1978).
[39] J. K. Morris, D. E. Mutton, and E. Alberman, Journal
of medical screening 9, 2 (2002).
[40] J. R. Gruhn, A. P. Zielinska, V. Shukla, R. Blanshard,
A. Capalbo, D. Cimadomo, D. Nikiforov, A. C.-H. Chan,
L. J. Newnham, I. Vogel, et al., Science365, 1466 (2019).
[41] Z. Holubcov´ a, M. Blayney, K. Elder, and M. Schuh, Sci-
ence 348, 1143 (2015).
[42] A. Webster and M. Schuh, Trends in cell biology 27, 55
(2017).
[43] A. Musacchio, Current biology 25, R1002 (2015).
[44] A. I. Mihajlovi´ c and G. FitzHarris, Current Biology 28,
R895 (2018).
[45] J. E. Holt and K. T. Jones, Molecular human reproduc-
tion 15, 139 (2009).
[46] K. Howe and G. FitzHarris, Biology of reproduction 89,
71 (2013).
[47] J. Hugues, J. Soussis, I. Calderon, J. Balasch, R. An-
derson, and A. Romeu, Human Reproduction 20, 629
(2005).
[48] M. Wikland, C. Bergh, K. Borg, T. Hillensj¨ o, C. Howles,
A. Knutsson, L. Nilsson, and M. Wood, Human Repro-
duction 16, 1676 (2001).
[49] J. Cha, X. Sun, and S. K. Dey, Nature medicine 18, 1754
(2012).
[50] R. J. Little and D. B. Rubin, Statistical analysis with
missing data, vol. 793 (John Wiley & Sons, 2019).
[51] C. Racowsky, J. E. Stern, W. E. Gibbons, B. Behr, K. O.
Pomeroy, and J. D. Biggers, Fertility and sterility 95,
1985 (2011).
[52] M. Vernon, J. E. Stern, G. D. Ball, D. Wininger,
J. Mayer, and C. Racowsky, Fertility and sterility 95,
2761 (2011).
[53] D. K. Gardner and B. Balaban, MHR: Basic science of
reproductive medicine 22, 704 (2016).
[54] P. Braude, V. Bolton, and S. Moore, Nature 332, 459
(1988).
[55] L. Rienzi, A. Capalbo, M. Stoppa, S. Romano, R. Mag-
giulli, L. Albricci, C. Scarica, A. Farcomeni, G. Vajta,
and F. Ubaldi, Reproductive biomedicine online 30, 57
(2015).
[56] M. G. Minasi, A. Colasante, T. Riccio, A. Ruberti,
V. Casciani, F. Scarselli, F. Spinella, F. Fiorentino, M. T.
Varricchio, and E. Greco, Human Reproduction 31, 2245
(2016).
[57] A. Campbell, S. Fishel, N. Bowman, S. Duffy, M. Sedler,
and C. F. L. Hickman, Reproductive biomedicine online
26, 477 (2013).
[58] D. J. Kaser and C. Racowsky, Human reproduction up-
date 20, 617 (2014).
[59] M. J. Hill, K. S. Richter, R. J. Heitmann, J. R. Graham,
M. J. Tucker, A. H. DeCherney, P. E. Browne, and E. D.
Levens, Fertility and sterility 99, 1283 (2013).
[60] A. Ahlstr¨ om, C. Westin, E. Reismer, M. Wikland, and
T. Hardarson, Human reproduction 26, 3289 (2011).
[61] N. S. Macklon, R. L. Stouffer, L. C. Giudice, and B. C.
Fauser, Endocrine reviews 27, 170 (2006).
[62] E. Akin and H. M. Lacker, Journal of mathematical bi-
ology 20, 113 (1984).
[63] H. Lacker, Biophysical journal 35, 433 (1981).
[64] H. M. Lacker and E. Akin, Mathematical Biosciences 90,
305 (1988).
[65] C. Leary, H. J. Leese, and R. G. Sturmey, Human repro-
duction 30, 122 (2015).
[66] H. Van de Velde, G. Cauffman, H. Tournaye, P. Devroey,
and I. Liebaers, Human reproduction 23, 1742 (2008).
[67] S. Willadsen, Reproduction 59, 357 (1980).
[68] W. Allen and R. Pashen, Reproduction 71, 607 (1984).
[69] L. H. Hartwell, J. J. Hopfield, S. Leibler, and A. W.
Murray, Nature 402, C47 (1999).
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.