Methods
to predict the interprovincial transmissions in mainland China, especially those from Hubei
Province, and predicted that the COVID-19 in China is likely to decelerate before Feb 18th and to end
before April 2020. Chen et al. (2020) made prediction based on epidemiological surveys and analyses,
which showed that the total number of diagnoses would be 2-3 times that of SARS, and the peak is
predicted to be in early or middle February. Yu et al. (2020) revised the SIR model based on the
characteristics of the COVID-19 epidemic development, and proposed a time-varying parameter-SIR
model to study the trend of the number of infected people. Peng et al. (2020) used the SEIR method
to predict the end of the epidemic in most cities in mainland China. Wu et al. (2020) used the Markov
chain Monte Carlo method to estimate R0, and inferred from the SEIR model that the peak COVID
in Wuhan would be reached in April, and other cities in China would be delayed by 1 to 2 weeks.
However, there are some obvious shortcomings of forecasting method based on epidemic model
in terms of outbreak prediction. For example, SEIR model is a mathematical method relying on
an assumption of epidemiological parameters for disease progression, which is absent for the novel
pathogen. For instance, the basic infection number R0, the daily recovery rate, the characteristics
of the disease itself (such as the infection rate and the conversion rate of the latent to the infected),
the daily exposure rate of the latent and infected, and their initial population infection status (total
3
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
population, infected, the initial value of the latent, the susceptible, the healer, etc.) and many other
key parameters need to be set. For infectious diseases that have already appeared in the past, or those
who have a large amount of data, it is not difficult to obtain these parameters. However, for unknown,
sudden and early infectious diseases, obtaining these parameters is full of difficulties, which leads to a
great uncertainty and limitations in the prediction of the epidemic situation using the SEIR model.
Moreover, there exist many challenges for the prediction of a new epidemic situation similar to
the COVID-19. First, little prior knowledge can be used to analogize or refer to for a brand new
epidemic; secondly, the existence of government management will make the development of the epi-
demic completely different from that under free development, thus how to incorporate the influence
of government measures into the fitting process of parameters and build a statistical model from this
should be taken into consideration; thirdly, in the early-outbreak the initial data often fluctuates vio-
lently and the data quality is low, thus many commonly used parameter estimation methods are not
applicable anymore; furthermore, the amount of data in early stage is too small, so it is difficult to
directly rely on the inertia of the data to make forward prediction. In summary, in the early stages of
brand new epidemics, how to use some low-quality and small data sets to make basic and relatively
accurate forecast judgments for the entire process of the epidemic, is a long-term pain point.
To cope with these challenges, we propose a simple and effective framework incorporating the
effectiveness of the government control to forecast the whole process of a new unknown infectious
disease in its early-outbreak, from which we emphasis on the prediction of meaningful milepost mo-
ments. Specifically, we first propose a series of iconic indicators to characterize the extent of epidemic
spread, and describe four periods of the whole process corresponding to the four meaningful milepost
moments: two turning points and two “zero” points; then we develop the proposed procedure with
mild and reasonable assumption, especially without relying on an assumption of epidemiological pa-
rameters for disease progression. Finally we apply it to analyze and evaluate the COVID-19 using
the public available data in mainland China beyond Hubei Province from the China CDC during the
period of Jan 29th, 2020, to Feb 29th, 2020, which shows the effectiveness of the proposed procedure.
From the empirical study, we can conjecture that the proposed method may cast a flexible frame-
work and perspective for early prediction of a sudden and unknown new infectious disease with effective
government control. Specifically, in the early stage of the epidemic when some regular information is
initially displayed, the proposed method can be used to predict the process of epidemic development
and to judge which stage of development the situation is at, when the peak will be reached, and when
the turning point will appear. Moreover, by continuously accumulating data and updating the model
during the development of the epidemic, we can also predict when the epidemic will basically end.
Finally, the proposed method enjoys great generalizability, which can be generalized to understand
4
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
the epidemiological trend of COVID-19 spread in other counties, which will provide useful guidance
for fighting against it.
The reminder of this paper is organized as follows. In Section 2, we proposed the main methodology,
where we defined the iconic indicators to characterize the extent of epidemic spread in Section 2.1,
yielding four periods of the whole process corresponding to the four meaningful milepost moments: two
turning points and two “zero” points in Section 2.2, then Section 2.3 presents the proposed procedure
with mild and reasonable assumption. Then we applied the proposed method to the COVID-19 using
the public available data in mainland China beyond Hubei Province from the China CDC during the
period of Jan 29th, 2020, to Feb 29th, 2020, and describe the trend of the COVID-19 spread in detail
in Section 3. Some conclusions and discussions are finally given in Section 4.
2 Methodology
In order to assess and predict the epidemic, we first define a set of necessary indicators that can reflect
the status of disease contagion. We then divide the cycle of epidemic into four stages, which divided
by the turning points of the proposed indicators. Finally, we propose a computational framework to
predict the turning points.
2.1 The iconic indicators to characterize a epidemic
It is obvious that the contagion process of an unknown virus in different regions would be diverse with
respect to the number of patients and the growth pattern of epidemic, because of population density,
population mobility, public health conditions, disease prevention and control measures. Therefore, we
first constructed a set of indicators to monitor the essential laws of the development of the disease.
There are several requirements for the monitoring indicators. Firstly, the scale of the data should
be eliminated so that the analysis methods and results are comparable across regions. Secondly, they
can well reflect the general laws and characteristics of the epidemic process as well as accurately
and coherently describe the entire process of the epidemic from the begin to the end. Especially, they
should be able to answer the question of when the turning point of the epidemic would appear. Thirdly,
they should be as simple and convenient as possible so that it can be applied with publicly available
data. Last but not the least, the indicators should have clear meaning and good interpretability.
Following the above, we first adopt three basic indicators that are published daily by the provincial
and municipal governments of China. That is, for time t, the daily confirmed cases Et, the daily
recovered ones Ot, the daily deaths Dt. Then we define a few monitoring indicators to characterize
the epidemic stages, that is the number of infectious cases in hospital Nt, the daily infection rate Kt
and the daily removed (the sum of recovered and deaths) rate It, which are defined as follows.
5
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
• The number of infectious cases in hospital Nt is defined as the cumulative confirmed cases with
recovered ones and deaths removed up to t, that is
Nt =
t∑
i=1
(Et− Oi− Di).
Note that Nt is essential for epidemic investigation, since it reflects the size of local patients and
the pressure of medical system.
• The daily infection rate Kt is defined as the ratio of the daily confirmed cases at time t and the
number of infectious cases in hospital at time t− 1, i.e.
Kt = Et
Nt−1
.
Obviously, Kt reflects the rate at which patients enter the treatment system. It is influenced
by many factors, including the property of infectious diseases, the average immune capacity of
the population, population density, climate condition, public health conditions, public health
awareness, the awareness of self-prevention of diseases and the efforts of epidemic prevention
and control.
• Similarly, the daily removed rate It is defined as the ratio of the daily removed cases at time t
and the number of infectious cases in hospital at time t− 1, i.e.
It = Ot + Dt
Nt−1
,
where It reflects the rate at which patients leave the medical system, that is, the rate at which
the pressure of medical resource is released.
Using the above indicators, we further define Rt as the outbreak status on day t as follow:
Rt = 1 + Kt− It.
Obviously, it holds that
Nt = Nt−1Rt = N0
t∏
l=1
(1 + Kl− Nl), (1)
where N0 denotes the initial number of patients in hospital at the beginning of the outbreak. In
particular, when the daily infection rate and removed rate are relatively stable, denoted as K and I
respectively, we have the constant epidemic status index R = 1 + K− I. Then (1) can be written as:
Nt = N0· Rt = N0· (1 + K− I)t, (2)
which shows that the epidemic situation is in the form of an exponential curve. And the epidemic
status indicator R can well reflect the rate of expansion or convergence of the population with infectious
capacity.
6
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
2.2 Four stages of a epidemic
In this section, we will describe the whole process of a epidemic under assumption that the government
has implemented effective control measures, which can be divided into four stages, i.e. “outbreak
period”, “controlled period”, “mitigation period” and “convergence period” successively. And we will
quantify the iconic features for each stage, which corresponds to the two turning points and two “zero”
points, respectively.
Stage 1: Outbreak Period
In the initial stage of an epidemic outbreak, there is delay of social response due to the limited
knowledge of the epidemic, and the power of contagion preventing and control is inevitably not enough.
Thus the daily infection rate Kt would be high. At the same time, the healing process in the initial
stage is relatively long, and the number of severe patients is small, leading the daily removed rate It
to be close to zero. Therefore, the outbreak status indicator Rt during this period is usually much
larger than 1, that is:
Kt≫ It, Rt = 1 + Kt− It > 1⇒ Nt > N t−1.
It can be seen that, during the outbreak period, the number of newly diagnosed patients increases
sharply, and the number of patients in the hospital will increase dramatically correspondingly, which
will pose great burden on medical institutions, especially for hospitals.
As the epidemic exacerbates, if the government began to intervene by a series of emergency mea-
sures, where a disease prevention and control system will be quickly established, the daily infection
rate Kt will significantly decrease. Usually, the new daily confirmed cases will begin to decline as
well. During the epidemic prevention and control process, once the situation improves, we will see
the emergence of the first turning point denoted as T1. Then after the data T1, the newly diagnosed
patients Et changes from a rapid rise in the outbreak period to a descending channel (E t < E t−1). In
summary, the emergence of the first turning point T1 indicates that the disease control measures have
begun to work, which implies the end of the “Outbreak Period”.
Stage 2: Controlled Period
The emergence of the first turning point is a very positive signal, indicating that the public health
management measures have obviously taken effect and the epidemic has entered the “controlled pe-
riod”. However, due to the fact that the completion rate It at this stage is still relatively low, the
number of patients treated in the hospital will continue to increase. The controlled period will continue
until the second turning point T2 appears, that is, patients in hospital Nt reaches the peak and starts
to decline. This is because the completion rate increase so significantly that Kt = It is fulfilled after a
long period of treatment in the previous stage. When the completion rate It surpasses infection rate
Kt, the number of patients treated in the hospital began to decline from peak.
7
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
Stage 3: Mitigation Period
The sign of the end of the controlled period is Kt = It. Thereafter, Kt will continue to fall with
the rise of It, which gives
Kt 1⇒ Nt < N t−1
This indicates that the daily completion rate It will start to be greater than the daily infection rate
Kt, that is, the value of the outbreak status indicator Rt becomes less than 1. The population size
with infectious capacity will be reduced, and the pressure of medical resources will be significantly
relieved, marking the beginning of “mitigation period”. The mitigation period will continue until the
appearance of zero report for newly confirmed cases, that is, Et = 0, which we call it as the first
“zero” point Z1. After the first zero point is reached, the intensity of prevention and control in the
entire society will be relieved except for the hospital, that is, the “mitigation period” ends and the
“convergence period” starts.
Stage 4: Convergence Period
The “convergence period” will end at the second “zero” point Z2, which means that the number
of people treated in the hospital is equal to or close to zero. After reaching the second zero point, the
epidemic is completely over.
For clarity, we summarize the iconic features and the corresponding milepost moments of each
stage in the whole process of the epidemic in Table 1.
T able 1: The four stage of an epidemic
Stage Outbreak Controlled Mitigation Convergence
Begin
with
the number of newly
diagnosed increases
the number of newly
diagnosed decreases
(first turning point)
the number of
patients in hospital
decreases (second
turning point)
the number of newly
diagnosed equals to
0 (first zero point)
End
with
the number of newly
diagnosed reached
peak (first turning
point)
the number of
patients in hospital
reached peak
(second turning
point)
the number of
patients in hospital
equals to 0 (first
zero point)
the number of
patients in hospital
equals to 0 (second
zero point)
K >> I , R >> 1 K > I , R > 1 K < I , R < 1 K = 0 , R << 1
8
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
2.3 Implementation: the proposed model
According to Section 2.2, the modeling and predicting of the epidemic need to be divided into two
parts. The first part corresponds to the outbreak period, where the intervention and disease curing
is not effective enough. The infection rate Kt increases rapidly and the completion rate It is small.
Thus, the number of newly diagnosed patients Et increases rapidly, and the number of patients treated
in hospital Nt increases. The pressure on medical resources will soon be overwhelmed. According to
equation (2), Nt will be in an exponential growth trend without forming a convex curve, nor will the
so-called two turning points or two “zero” points appear.
The second part, which is the focus of this article, is when the Kt starts to decrease and It starts
to increase due to effective intervention and improved healing level for individual patients. Only in
this situation will the turning points and zero points T1, T2, Z1, Z2 successively appear, and then the
epidemic could end. Therefore, we will model the development of the epidemic under the assumption
of effective intervention, then we can obtain the early prediction of two turning points and two “zero”
points based on the predicting modeling of Et and Nt.
Suppose that the infection rate Kt and the removed rate It change gently within a time window
m before time t0 with exponential growth, then given m and t0, denote VK|(t0,m) and VI|(t0,m) as the
average change rate of Kt and It respectively, that is,
VK|(t0,m) =
{ Kt0
Kt0−m+1
}1/(m−1)
, V I|(t0,m) =
{ It0
It0−m+1
}1/(m−1)
. (3)
For any t > t 0, the infection rate Kt and the removed rate It can be predicted as follows:
ˆKt|t0 := ˆKt0 (t− t0) = ˆKt0 (t− t0− 1)· VK|(t0,m) =··· = ˆKt0 (1)· V t−t0−1
K|(t0,m) = Kt0· V t−t0
K|(t0,m), (4)
ˆIt|t0 := ˆIt0 (t− t0) = ˆIt0 (t− t0− 1)· VI|(t0,m) =··· = ˆIt0 (1)· V t−t0−1
I|(t0,m) = It0· V t−t0
I|(t0,m). (5)
Thus, we can obtain the outbreak status Rt, the number of patients in the hospitalNt, and the number
of newly diagnosed Et as
ˆRt|t0 = 1 + ˆKt|t0− ˆIt|t0,
ˆNt|t0 = ˆNt−1|t0· ˆRt|t0,
ˆEt|t0 = ˆNt−1|t0· ˆKt|t0.
According to the prediction process, it can be seen that the prediction results mainly depend on
VK|(t0,m) and VI|(t0,m), whose value is up to the selection of time window m and starting point t0.
However, it is worth noticing that the selection of m and t0 is not arbitrary, which is suggested as in
the follow assumption.
9
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
Assumption 1. The time window m and the starting point t0 should be chosen satisfying VK|(t0,m) 1. Meanwhile, keeping It < 1 due to interpretability constraints, and the starting point
t0 should be close to the date of the latest published data as much as possible.
In summary, here we describe details of the proposed procedure in Algorithm 1.
Algorithm 1 Main Prediction Procedure.
1: Initial setting m and t0, which satisfying Assumption 1;
2: Compute VK and VI according to (3); Set t = t0 + 1.
3: Prediction: updating the predicted results at time t via the forecasting value ahead ofl = t−t0-step
as follows:
ˆRt|t0 = ˆRt0(l) = Kt0· V l
K|(t0,m)
ˆIt|t0 = ˆIt0(l) = It0· V l
I|(t0,m)
ˆRt|t0 = 1 + ˆRt|t0− ˆIt|t0
ˆNt|t0 = ˆNt−1|t0· ˆRt|t0
ˆEt|t0 = ˆNt−1|t0· ˆKt|t0
4: Prediction of the milepost moments: If ˆEt−1|t0 < ˆEt|tc, then T1 = t− 1; If ˆNt−1|t0 < ˆNt|t0, then
T2 = t− 1; If ˆEt−1|t0 < E 0 = 1, then Z1 = t− 1; If ˆNt−1|t0 < N 0 = 1, then Z2 = t− 1; If none of
the above is satisfied, turn to the next step.
5: Set t = t + 1, return to Step 2 until T1, T2, Z1, Z2 are obtained.
It is also worth noticing that in practice, the more data information we accumulate, the clearer
the underly law of the epidemic. Therefore, we can also continuously modify the iterative prediction
model according to the actual data, so that the prediction of the next stage and the prediction of the
long-term situation can be more accurate.
3 Application: Analysis of the COVID-19 in mainland China be-
yond Hubei Province
We apply it to analyze and evaluate the COVID-19 using the public available data in mainland China
beyond Hubei Province from the China CDC during the period of Jan 29th, 2020, to Feb 29th, 2020.
Here we first show the actual trend of the COVID-19, and then compared with the predicted ones via
the proposed method. Finally, we will show the effect of m on the predicted results.
10
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
3.1 The turning points and zero points observed
After the shutdown of most parts of Hubei province in Jan 23rd, other parts of China also immediately
launched prevention and control strategies, including regional isolation, admission of all confirmed
patients, isolating all suspected patients and so on. The effective implementation of these intervention
policies quickly controlled the rapid spread of the epidemic in these areas. As can be seen in Figure 1,
the parameter infectious rate Kt, which reflects the intensity of the spread of the epidemic, has shown
a significant downward trend since Jan 27th after severe fluctuations from Jan 22nd to 26th. As can
be seen in Fig. 1, we find out that the daily confirmed cases reached at peak on Jan 30th, 2020, with
761 confirmed cases and then continued to decline for two consecutive days.
Fig. 1: Trend of the daily confirmed cases from 01/22 to 02/01, 2020.
However, the migration raised from people returning to work after Chinese New Year on Feb 3rd
undermines the continuous decline of Et. Since Feb 2nd, the number of daily confirmed patients in
mainland China beyond Hubei Province has increased for two consecutive days, where the Et on Feb
3rd has increased by 23% compared to that on Feb 2nd. It can be concluded that these fluctuations are
caused by the resuming of social activities, which leads Et continue to decline since Feb 4th. In many
literature and media reports, Feb 3rd is used as the time point when the number of newly confirmed
patients starts to decline. But considering the fact that the epidemic was already under control, here
we still view Jan 30th as the first turning point .
After that, the second turning point T2, which is the time point when the number of infectious
cases in hospital Nt starts to decline, have also been observed. Fig. 2 shows the true curves of the daily
infection rate Kt, daily removed rate It, and Nt calculated based on the actual data from mainland
China beyond Hubei from Jan 22th, 2020 to Mar 13th, 2020. It can be seen that the second turning
point T2 appeared on Feb 11th, with the emergence of Kt < I t on that day, and the number of patients
in the hospital continued to decreases since then.
As for the first zero point Z1, the definition is the time when the number of daily confirmed
11
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
Fig. 2: Observed Kt, It and Nt of the COVID-19 from Jan 22 to Mar 13, 2020.
cases is equal to zero, which is too strict for the real situation Thus, in this article, we take the
criteria for cancelling travel warnings developed by the WTO during SARS as a reference, and make
some adjustments to the definition of the first zero point: the time when the daily confirmed cases
Et continues to be less than 5 for 3 days is revised to be Z1. Then, if we exclude confirmed cases
that originated from abroad, daily confirmed cases has already become less than 5 since Mar 3rd
in mainland China beyond Hubei Province, thus according to our revised definition, Mar 5th is Z1.
However, there were still 1,089 patients in hospital on that day. Therefore, it would still take some
extra time to reach the second zero point Z2.
3.2 Prediction results
Starting from Jan 29th, we use the proposed forecasting method to make real-time predictions on the
two turning points T1 and T2 and two “zero” points Z1 and Z2 with window size m = 5. The specific
and predicted results are as follows.
We first conducted the proposed prediction model on Jan 29th, which indicated that the first
turning point T1 would arrive on Jan 31st, i.e., Et < E t− 1. In reality, the first turning point did
arrive on Jan 30th, which is only one day away from our predicted result.
As for the second turning point, since the trueT2 occurred on Feb 11th, we summarize the frequency
of the prediction results obtained with t0 varying from Jan 29th to Feb 10th, 2020 andm = 5 in Figure
3(a). From it we can see that the prediction of second turning point mainly concentrated in the range
from Feb 9th to Feb 11th, which is consistent with the observed second turning point in reality. It
is worth mentioning that we got the general information of T2 at a very early stage: we predicted on
Feb 2nd that the second turning point T2 would arrive on Feb 11th, which is exactly the same as the
12
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
second turning point that observed in reality. Since then, we have continuously tracked the rolling
predictions, which have not yet changed much.
Fig. 3: The frequency of prediction results of turning points and zero points.
Similarly, Figure 3(b) and Figure 3(c) show the frequency of the prediction results for two “zero”
points obtained with t0 varying from Jan 29th to Feb 29th, 2020 and m = 5, respectively. Specifically,
for the predicted first zero point Z1 in Figure 3(b), we divide the prediction results from these days
into 5 intervals, which can be seen that the prediction results of the first zero point Z1 are mainly
concentrated on Mar 1st to 5th, which is consistent with the actual result. There is also a “pessimistic”
prediction as a result of the sudden fluctuation of data on Feb 3rd, which predicted that the first zero
point would arrive on Mar 17th. For the predicted second zero point Z2 in Figure 3(c), it can be
seen that the second zero point will be reached from mid-March to mid-April. However, there is a
prediction result that Z2 will appear on May 11th, which is far away from other results. The reason
for this uncommon result is that the starting point of this forecast is Jan 29th, when the epidemic
situation in mainland China beyond Hubei was still in the outbreak period with Et still rising, It very
small, so the prediction result about the finish of the epidemic may not be accurate.
Furthermore, we also present the forecast results of the four milepost moments together with
the trend of the cumulative number of infectious cases in hospital ˆNt and the cumulative number of
infectious ∑t
l=1 ˆEl in Figure 4 when the prediction starting point t0 fixed at Jan 29th, Jan 31st, Feb
12th and Feb 26th, 2020, respectively. As can be seen from Figure 4(a), in Jan 29th, which is the
very early stage of the epidemic, we predicted that the first turning point would appear on Jan 31st,
which is only one day behind the actual observation. Additionally, the time of the second turning
point result predicted on that day was Feb 14th, which is only 3 days away from the reality. The first
zero and second zero forecast results are Mar 7th and May 11th, respectively.
Figure 4(b) shows the prediction results when the first turning point have already appeared, from
which we can see that the prediction for T2 on Jan 31st is accurately with the second turning point
possible occurring on Feb 11th. Meanwhile the first zero point and the second zero point are predicted
13
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
Fig. 4: Forecasting results of the four milepost moments together with the trend of the cumulative
number of infectious cases in hospital ˆNt and the cumulative number of infectious ∑t
l=1 ˆEl compared
with their observed cases Nt and ∑t
l=1 El when the prediction starting point t0 fixed at Jan 29th (a),
Jan 31st (b), Feb 12th (c) and Feb 26 (d), 2020, respectively.
to appear around Mar 4th and Mar 23rd, respectively.
Similarly, after the arrival of the second zero point, Figure 4(c) shows the forecast results of the
first and second zero points predicted on Feb 12th, which show the forecast results for Z1 and Z2 are
on Mar 9th and Mar 25th, respectively. From the fitting results, we know that our prediction of the
cumulative number of patients in the hospital Nt and the total number of confirmed patients is very
similar to the actual situation, so our prediction results are highly reliable. Finally, we also give a
very recent (Feb 26th) forecast in Figure 4(d), which is similar to the results mentioned above.
3.3 Results with different window sizes m
Note that the number of m plays an important role in the proposed procedure, and all the results
we discussed in the section 3.2 is obtained with fixed m = 5. In this section, we will illustrate the
impact of different choice of m on the results, and give the empirical choice in real data analysis.
Parallel to Section 3.2, here we obtain the results for the second turning point and both zero points
14
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
via implementation of the proposed procedure with m =3, 4, and 6, respectively. And we summarize
all these results for the second turning point and both zero points in Fig. 5, respectively.
Before Feb 09Feb 09~11 Feb 12~14 After Feb 14
Date
m=3m=4m=5m=6
0 7 5 1
1 7 2 3
0 11 2 1
0 9 3 1
(a) Prediction of 2nd turning point
(a)
Before Feb 25Feb 25~29Mar 01~05Mar 05~10After Mar 10
Date
m=3m=4m=5m=6
0 12 12 6 2
1 11 13 4 3
0 5 18 8 1
0 7 16 8 1
(b) Prediction of 1st zero point (b)
Before Mar 01Mar 01~10Mar 11~20Mar 21~30
Mar 31~Apr 09After Apr 09
Date
m=3m=4m=5m=6
0 13 9 4 2 4
1 7 9 7 2 6
1 7 12 6 5 1
0 10 11 6 0 5
(c) Prediction of 2nd zero point (c)
Fig. 5: Summary of prediction for the second turning point (a), the first (b) and second (c) “zero”
points with different m.
From Fig. 5, we can see that the most possible of forecasts for the second turning point occur
around the period from Feb 9th to 11th for all choice of m; similar results hold for the forecast of
the first zero with the most likelihood of appearance around the early March. Both results show the
limited influence of m on the results. From Fig. 5(c), although the forecasts for the second “zero”
point with different m seems not as good as those for the second turning point and the first zero, it
vary slightly, with its occurrence from mid-March to mid-April.Overall, the choice of m seems not a
critical value for the forecasting results, and we recommend its empirical choice from 3 to 6.
4 Discussion and conclusion
Focusing on the four meaningful mileposts, we put forward a simple and effective framework incor-
porating the effectiveness of the government control to forecast the whole process of a new unknown
infectious disease in its early-outbreak. Specifically, we first propose a series of iconic indicators to
characterize the extent of epidemic spread, and describe four periods of the whole process correspond-
ing to the four meaningful milepost moments: two turning points and two “zero” points; then we
develop the proposed procedure with mild and reasonable assumption, especially without relying on
an assumption of epidemiological parameters for disease progression.
We examine our model with COVID-19 data in mainland China beyond Hubei province, which
can detect the gross process of the epidemic at its early-outbreak. Specifically, in the first predicting
task that conducted on Jan 29, the predicted date when the number of newly confirmed patients Et
would fall for the first time is only one day behind the observation in reality. On Feb 2nd, our model
15
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
predicted that the date when the number of patients in the hospital Nt reaches its peak is Feb 11th,
which is consistent with the real world situation. Later, the forecasting results fluctuated but were
overall stable and close to the true observation. Meanwhile, we predict that the first zero point Z1
will arrive between the end of Feb and the beginning of March. And the second zero point Z2 will
arrive at mid-March to mid-April. We also checked the robustness of our model under different time
windows and found that the selection of the time window has little effect on the prediction of turning
points. As a prediction model for the task of early warning of a new epidemic, our prediction model
is proved to be quite efficient.
At present, many countries around the world are overwhelmed by the COVID-19 epidemic, which
calls for global efforts. While our method is able to depict and predict the trend of an epidemic at
a very early stage, it can be used to predict the current COVID-19 epidemic internationally, or any
other new, unknown, explosive epidemic in the future. We believe that the prediction results of this
References
Cleo Anastassopoulou, Lucia Russo, Athanasios Tsakris, and Constantinos Siettos. Data-based anal-
ysis, modelling and forecasting of the novel coronavirus (2019-ncov) outbreak. medRxiv, 2020.
Domenico Benvenuto, Marta Giovanetti, Marco Salemi, Mattia Prosperi, Cecilia De Flora, Luiz Carlos
Junior Alcantara, Silvia Angeletti, and Massimo Ciccozzi. The global spread of 2019-ncov: a
molecular evolutionary analysis. Pathogens and Global Health , pages 1–4, 2020.
Zeliang Chen, Wenjun Zhang, Yi Lu, Cheng Guo, Zhongmin Guo, Conghui Liao, Xi Zhang, Yi Zhang,
Xiaohu Han, Qianlin Li, et al. From sars-cov to wuhan 2019-ncov outbreak: Similarity of early
epidemic and prediction of future trends. CELL-HOST-MICROBE-D-20-00063, 2020.
Matteo Chinazzi, Jessica T Davis, Marco Ajelli, Corrado Gioannini, Maria Litvinova, Stefano Merler,
16
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
Ana Pastore y Piontti, Kunpeng Mu, Luca Rossi, Kaiyuan Sun, et al. The effect of travel restrictions
on the spread of the 2019 novel coronavirus (covid-19) outbreak. Science, 2020.
Yi Fan, Kai Zhao, Zheng-Li Shi, and Peng Zhou. Bat coronaviruses in china. Viruses, 11(3):210, 2019.
Michelle L Holshue, Chas DeBolt, Scott Lindquist, Kathy H Lofy, John Wiesman, Hollianne Bruce,
Christopher Spitters, Keith Ericson, Sara Wilkerson, Ahmet Tural, et al. First case of 2019 novel
coronavirus in the united states. New England Journal of Medicine , 2020.
Chaolin Huang, Yeming Wang, Xingwang Li, Lili Ren, Jianping Zhao, Yi Hu, Li Zhang, Guohui Fan,
Jiuyang Xu, Xiaoying Gu, et al. Clinical features of patients infected with 2019 novel coronavirus
in wuhan, china. The Lancet, 395(10223):497–506, 2020.
David S Hui, Esam EI Azhar, Tariq A Madani, Francine Ntoumi, Richard Kock, Osman Dar, Giuseppe
Ippolito, Timothy D Mchugh, Ziad A Memish, Christian Drosten, et al. The continuing epidemic
threat of novel coronaviruses to global health-the latest novel coronavirus outbreak in wuhang,
china. International Journal of Infectious Diseases , 2020.
Qun Li, Xuhua Guan, Peng Wu, Xiaoye Wang, Lei Zhou, Yeqing Tong, Ruiqi Ren, Kathy SM Leung,
Eric HY Lau, Jessica Y Wong, et al. Early transmission dynamics in wuhan, china, of novel
coronavirus–infected pneumonia. New England Journal of Medicine , 2020.
Hayes KH Luk, Xin Li, Joshua Fung, Susanna KP Lau, and Patrick CY Woo. Molecular epidemiology,
evolution and phylogeny of sars coronavirus. Infection, Genetics and Evolution , 2019.
World Health Organization et al. Novel coronavirus (2019-ncov) situation reports. 2020.
Liangrong Peng, Wuyue Yang, Dongyan Zhang, Changjing Zhuge, and Liu Hong. Epidemic analysis
of covid-19 in china by dynamical modeling. arXiv preprint arXiv:2002.06563, 2020.
Bastian Prasse, Massimo A Achterberg, Long Ma, and Piet Van Mieghem. Network-based prediction
of the 2019-ncov epidemic outbreak in the chinese province hubei. arXiv preprint arXiv:2002.04482 ,
2020.
Biao Tang, Xia Wang, Qian Li, Nicola Luigi Bragazzi, Sanyi Tang, Yanni Xiao, and Jianhong Wu. Es-
timation of the transmission risk of the 2019-ncov and its implication for public health interventions.
Journal of Clinical Medicine , 9(2):462, 2020.
Joseph T Wu, Kathy Leung, and Gabriel M Leung. Nowcasting and forecasting the potential domestic
and international spread of the 2019-ncov outbreak originating in wuhan, china: a modelling study.
The Lancet, 2020.
17
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint
Wen-Bin Yu, Guang-Da Tang, Li Zhang, and Richard T Corlett. Decoding the evolution and trans-
missions of the novel pneumonia coronavirus (sars-cov-2) using whole genomic data. ChinaXiv,
202002:v2, 2020.
Tianyu Zeng, Yunong Zhang, Zhenyu Li, Xiao Liu, and Binbin Qiu. Predictions of 2019-ncov trans-
mission ending via comprehensive methods. arXiv preprint arXiv:2002.04945, 2020.
Hao Zhang, Zijian Kang, Haiyi Gong, Da Xu, Jing Wang, Zifu Li, Xingang Cui, Jianru Xiao, Tong
Meng, Wang Zhou, et al. The digestive system is a potential route of 2019-ncov infection: a
bioinformatics analysis based on single-cell transcriptomes. BioRxiv, 2020.
Shi Zhao, Qianyin Lin, Jinjun Ran, Salihu S Musa, Guangpu Yang, Weiming Wang, Yijun Lou,
Daozhou Gao, Lin Yang, Daihai He, et al. Preliminary estimation of the basic reproduction number
of novel coronavirus (2019-ncov) in china, from 2019 to 2020: A data-driven analysis in the early
phase of the outbreak. International Journal of Infectious Diseases , 92:214–217, 2020.
Tao Zhou, Quanhui Liu, Zimo Yang, Jingyi Liao, Kexin Yang, Wei Bai, Xin Lu, and Wei Zhang.
Preliminary prediction of the basic reproduction number of the wuhan novel coronavirus 2019-ncov.
Journal of Evidence-Based Medicine , 2020.
18
. CC-BY-NC-ND 4.0 International licenseIt is made available under a
perpetuity.
is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint
The copyright holder for thisthis version posted March 24, 2020. ; https://doi.org/10.1101/2020.03.21.20040139doi: medRxiv preprint