Conclusion
This review presents an overview of multimodal data fusion in the field of neuroimaging, including current developments and challenges. We first outlined the fundamental limitations of individual modalities, which can include distortion, non-quantitative nature, and limited temporal/spatial resolutions. These limitations are the general motivators for the development of multimodal neuroimaging and fusion. Multimodal neuroimaging provides more comprehensive information on pathology.
We have summarised the individual benefits and limitations of the current imaging technologies and modalities, including CT, PET, SPECT, MRI, fMRI, DWI, PWI, and MRF. Building upon the available individual techniques, we summarised current development and application of multimodal neuroimaging and fusion in terms of neurological disorders and brain diseases, with a focus in three areas: developing brains, degenerative brains, and psychiatric disorders. The utilization of multiple modalities helps in clinical diagnosis, prevention of misdiagnosis, progression analysis, and research-oriented studies that allow us to gain a more in-depth understanding of human brain pathologies. Nowadays the effects of COVID in the human body are not well-known and maybe we could find in the future some people affected in the brain structure (micro ictus or ischemia). Nevertheless, the fusion techniques and strategies discussed in this survey may be transferred to COVID-19 multimodality image analysis [459] . AI can contribute to the information fusion [460] .
The forms of multimodality fusion include multi-modal, multi-focus, multi-temporal, and multi-view. These forms combine images from different instruments/acquisitions, acquisition focal lengths, time of acquisitions and conditions, respectively. Fusion rules were specified with respect to its components and three levels of fusion and theoretical foundation, i.e. rules derived from fuzzy logic, statistic models, and the human vision system. Then we summarised conventional and novel image decomposition and reconstruction methods for the fusion process, including methods based on RGB-IHS, pyramid representations, multi-resolution analysis, sparse representation, and salient features. In addition, we summarised both subjective and objective methods for fusion quality assessment.
The major benefits of multimodal data fusion in neuroimaging include distortion correction, higher temporal/spatial resolution, the combination of structural and functional information. We summarised these benefits with the current applications of multimodal fusion, e.g. MR-PET, EEG-fMRI, and EPI correction. This review also adds a particular focus on the application of multimodal image fusion in standardization, via atlas-based anatomical brain segmentation and the use of multi-atlas fusion. In addition, we summarised the effect of multimodal data fusion on the shift of neuroimaging diagnosis from qualitative analysis to quantitative evaluation, with an example of the effect of multimodal data in photon attenuation, scatter, and partial volume effects of PET and SPECT quantification.
Modern neuroimaging has seen significant improvements in acquisition quality and a constant increase in the abundance of imaging modalities. The fusion of modalities combines complementary information, expands resolution limits, provides standardization, and improves data quality.
It is expected that the effect of multimodal imaging fusion could also effectively scale the amount and quality of information accessible to radiologists for more both precise diagnosis and higher quality research in multiple aspects.
Multimodal
The application and improvement of the quantification in medical image allows to increase characteristics such as sensitivity and specificity, among others, obtaining more accurate patient's diagnoses. To perform this analysis, the emission computed tomography (ECT) is used, since it is the most important medical imaging modality in nuclear medicine. More specifically, the two image modalities analyzed are positron emission tomography (PET) and single-photon emission computed tomography (SPECT), which differ in the radiotracer used and the nature of the emission measurement. It should be noted that the factors associated with quantification that affect PET and SPECT are also extrapolated to multimodal images. Furthermore, multimodal images can be of great interest to improve quantification. Both, PET and SPECT, have proven to be effective imaging techniques for the diagnosis and monitoring of treatments in different medical applications, offering information about biological processes, which is very important for the study of brain activity [333] . The level of intensity in these types of images is bound to physical parameters, such as cerebral blood flow, glucose or receptor binding among others [334,
335] .
Since these modalities were conceived in the 50’s and 60’s Jones and Townsend [336] , several improvements have been developed in terms of quantification. One of the most relevant improvements in this field is due to the development of dual-modality imaging in the 90′s [337] . This type of imaging allows the combination of complementing functional information (PET and SPECT) with structural information (computed tomography (CT) and magnetic resonance imaging (MRI)), leading to increases in sensitivity and specificity. The combination of different characteristics that can be observed from each type of image allows us to obtain a better understanding of the structure and function of the human body. For example, PET/MRI combination offers remarkable advantages compared to the use of PET/CT since CT radiation dose is avoided and soft tissue images from MR could be acquired at the same time as the PET ones [338] .
PET offers higher resolution and better quality than SPECT, particularly in PET/CT. Furthermore, it allows for easier quantitative measurements [335] . For example, sensitivity could be up to two orders of magnitude greater with comparable axial fields of view [339] . Nevertheless, the spatial resolution that can be obtained in both techniques tends to be low, due to its dependence on the radiation dose administrated to the patients. Also, ECT images are usually degraded by several factors, such as photon attenuation, partial volume effect or scatter radiation [337] . Moreover, spatial resolutions associated with structural images continue to be higher (≈ 1 mm) than in functional images (≈ 4–6 mm) [340] .
Regarding quantification, an increasing number of relevant research aims to propose new quantification techniques or improve previously existing ones is observed in recent years. The possibility of quantification in nuclear medicine imaging is one of its successes. Due to the increasing use of nuclear images in therapies, the way of measuring functional images is changing. While historically, the use of relative and semiquantitative measures was typical, the recent application of absolute quantification is gaining support. One of the first steps was considered activity concentration and normalized uptake using the standard uptake value (SUV) [341] . Nevertheless, to reach quantitative imaging as a useful and potential tool, several factors must be addressed. As Zaidi and Hasegawa [339] mentioned among these factors are the system sensitivity and spatial resolution, dead-time and pulse pile-up effects, the linear and angular sampling intervals of the projections, the size of the object, photon attenuation, scatter, partial volume effects, patient motion, kinetic effects, and filling of excretory routes (e.g., the bladder) by the radiopharmaceutical.
In neuroimaging, PET and SPECT are a popular option to detect biomarkers associated with degenerative diseases. For Alzheimer's disease (AD), the accumulation of Amyloid- β plaques and tau aggregates can be detected [342,
343] , while an increased diffusivity in the striatum and thalamus can be observed in Parkinson's disease (PD) patients. A well-defined quantification of these biomarkers would allow more precise and reliable diagnoses.
This section is organized as follows. Firstly, quantification both in PET and SPECT is analyzed considering the associated effects and their correction, especially those related to photon attenuation, scatter, and partial volume effects. Then, novel developments focused on cerebral medical imaging are highlighted.
Even though PET scans could be evaluated qualitatively through the visual examination of the tracer uptake in cortical regions by a trained radiologist [344] , the best way to analyze these images is quantitative. For that, automated or semi-automated localization methods can be used to evaluate regional levels of tracer uptake [345,
346] . Within automated studies, an essential step is to use spatial normalization to register subjects' brain to a standardized template space, so that all subjects in the study can be compared [347] . Another procedure of great importance is intensity normalization, which is the object of study in this analysis.
As Foster, Bagci [348] mentioned, several semi-quantitative and quantitative parameters that can be considered for intensity normalization exist, such as standardized uptake value (SUV), fractional uptake rate (FUR), tumor-to-background ratio (TBR), nonlinear regression techniques, total lesion evaluation (TLE) or the Patlak-derived methods. Nevertheless, the most widely used is SUV, which can also be used in SPECT [348] . The use of body weight (BW) in SUV has been discussed in the literature, and the general advice is to use more reliable measures such as body surface area (BSA) or lean body mass (LBM) [349,
350] . Moreover, a decay factor that depends on the particular radiotracer used may also be considered [351] .
Several physiological and physical factors can influence the standardized uptake value obtained [348,
352] . Regarding physiological factors, some of them are weight and fat of the subject or blood glucose concentration, etc. Physical factors include partial volume effect (PVE), image manipulation (reconstruction, smoothing), and artifacts related to involuntary movements of the patient. The relevance of these factors lies in the variability of SUV. Studies carried out in the literature suggest that these factors may alter SUV in the range of 10%–30% [353,
354] . Nevertheless, the importance of correcting these effects depends on the clinical trial performed, since it can be a complex task and is not always worthwhile [355] .
Moreover, these factors affect not only semi-quantitative measurements but also affect absolute quantification. There are several methods to consider within the absolute quantification, such as image reconstruction, effects correction or calibration to obtain the measured activity distribution [356,
357] .
Among the mentioned factors associated with imaging quantification, we analyzed three effects and their respective corrections, since they are popular bias sources for PET and SPECT quantification. These effects are PVE, scatter, and photon attenuation effects.
Regarding the quantitative accuracy of PET image, partial volume effect (PVE) is the most relevant effect [358] . PVE is related to the finite spatial resolution of PET scanners and its discrete nature. That discrete nature causes a voxel to could be composed of more than one tissue, which leads to a final signal that is an averaged mix of signals, which is also called tissue fraction effect. Moreover, the larger the voxel size, the more tissues can coexist in it. This effect, along with the point-spread effect (or spillover), is the cause of image blur [354,
359] . Also, it is an effect that often causes confusion among clinicians and researchers when analyzing the image, since it is necessary to distinguish between loss of radioactivity due to PVE and the true loss of tissue.
Strategies to address this effect are called partial volume correction (PVC) methods. Distribution of both signal and noise are the main factors to consider for the selection of algorithms for PVC to apply [359] . Several PVC techniques have been proposed in the literature for improving image quality and quantitative accuracy in PET [360] , [361] , [362] , [363] . Historically, these techniques have been classified in several forms. In this work, these methods are grouped as region-based (RB) and voxel-based (VB) techniques, based the most commonly used classifications [340,
355] . The most relevant differences between them are that VB methods produce PVE-recovered images while RB methods do not, and RB methods are associated with regions of interest (ROIs) while VB methods are applied at a voxel level considering the recovery of the spatial resolution of the system. Usually, RB techniques tend to be used more frequently than VB ones since they include regional homogeneity assumptions, simplifying the problem [359] .
Table 5
shows a summary of the different existing techniques. Among the several PVC methods, the one associated with recovery coefficients (RC) is the simplest and most popular in clinical trials. Nevertheless, the most common technique used in neuroimaging is PVE correction based on anatomical images, which was developed for neuroimaging. The first technique proposed by Videen, Perlmutter [364] in 2D and later extended to 3D by Meltzer, Leal [365] was a segmentation in anatomy into two classes: brain and non-brain. Mullergartner, Links [366] proposed an improved version of this technique (known as GM algorithm) to use three instead of two classes: grey matter (GM), white matter (WM) and cerebrospinal fluid (CSF). Meltzer, Zubieta [367] added a fourth class to the previous algorithms in order to compensate for the real heterogeneity in GM tissue. The importance of this method is that PVE can be corrected between high and low intensity brain structures. Regarding the iterative reconstruction algorithms associated to VB methods, as V. Bettinardi, Castiglioni [340] noted, some of the best known and used is the Van Cittert (VC) algorithm, the re-blurred Van Cittert (R-VC) and the Richardson–Lucy (RL) algorithm. For broader coverage, readers are referred to reviews related to PVE [340] and iterative reconstruction methods [368] . Table 5 Classification of PVC methods (RB = region-based; VB = voxel-based). Table 5 Category Method Definition Benefits Limitations Examples RB Recovery coefficients (RC) They are numerical calculated in the image domain as a ratio of measured radioactivity concentration and actual concentration using spheres filled with a known radioactivity concentration Fairly simple Practical Popular within clinical trials More related to oncology [369, 370] PET raw data Instead of using the image domain, the sinogram is used Low computational cost Some methods require segmentation [371] , [372] , [373] Geometric transfer matrix (GTM) method Average uptake is estimated considering multiple ROIs using co-registered anatomical images Commonly used Good accuracy Requires segmentation Bias [374] , [375] , [376] , [377] VB Image reconstruction Spatial resolution is recovered within the image reconstruction process. It is highly recommended to use iterative reconstruction algorithms. The most relevant drawback of this is the high number of iterations needed Increase the quality of reconstruction Computational cost [378] , [379] , [380] , [381] Image deconvolution Spatial resolution is recovered by using a post-reconstruction restoration technique (deconvolution) that applies the point spread function (PSF). Iterative algorithms are commonly used Compensation of spill-over effects Computational cost [382] Multi-resolution approach Spatial resolution is recovered by using information from spatially coregistered high-resolution anatomical images, which allows transfer high spatial frequencies from anatomical to functional images Adjust of radioactivity concentration Used in brain imaging In poorly correlated areas the amount of artifacts increases [383] Anatomical images First, a segmentation of a coregistered anatomical image is done. Then, PVE correction is applied to the set of voxels associated with each region/class (tissue) Highly used in brain imaging Need segmentation Two classes: [364, 365] Three classes: [366] Four classes: [367]
Classification of PVC methods (RB = region-based; VB = voxel-based).
Finally, data extracted from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu) is used in order to show some PVE correction methods. This data consists of a patient diagnosed with Alzheimer's disease from which FDG-PET and T1 weighted MR images were taken, as shown in Fig. 23
. Both modalities are registered to respective templates of T1 and PET, and PET image was not normalized in intensity due to use only one subject in this manuscript. PVC techniques are applied by using PVElab [384,
385] , a platform developed by the EU sponsored PVEOut project. PVElab can facilitate PVE correction through a graphical interface with several steps, including registration, segmentation, reslicing, optional application of an atlas (none in our case), and PVE correction. Three methods discussed above are those analyzed: 3D Meltzer technique [365] and GM algorithm [366] regarding VB methods, and the RB method proposed by Rousset, Ma [374] . Fig. 23 Original PET (left) and MR (right) images. Fig. 23
Original PET (left) and MR (right) images.
The large difference between methods is shown in Fig. 24
a. The difference in the result between considering two ( Fig. 24 a) and three tissues ( Fig. 24 b) is highlighted. The method related to GTM ( Fig. 24 c) reflects intermediate results between the two images already commented. Fig. 24 Evolution of PVC techniques. Fig. 24
Evolution of PVC techniques.
Another relevant effect to analyze is photon attenuation. This phenomenon occurs due to the interaction between the photon associated to the radiotracer and elements in the body such as tissue. Normally, this interaction leads to a scatter in the photon radiation. The probability that a photon experiences an interaction is represented by the linear attenuation coefficient [386] . The situation of this phenomenon differs between PET and SPECT imaging. In PET, two antiparallel photons are detected in the collimator, so the total tissue thickness crossed by them is the same body thickness that intersected the straight line between the two detections (line of response). Nevertheless, in SPECT imaging, the attenuation depends on the total tissue thickness crossed and its type (e.g., soft tissue, bone), which varies depending on where the points of emission and detection are [386] . Traditionally, this correction was much more commonly performed in PET imaging than in SPECT. Nowadays, it is done in both to be able to achieve much more accurate quantitative results, although it is more difficult to apply in SPECT.
Attenuation correction is necessary in order to obtain accurate quantitative results and is widely implemented. For this, an attenuation map is needed, which Zaidi and Hasegawa [386] defined as a representation of the spatial distribution of linear attenuation coefficients and delineates the body structures located in the image. Once the attenuation map is done, the reconstruction algorithm can be implemented with more information. Attenuation correction techniques can be classified into two major groups: transmission-less approaches and the ones based on transmission scanning. While the first group considers the measured emission data to develop the attenuation map or assumes a uniform distribution for the coefficients, the second one uses transmission data of external sources such as CT and MRI, i.e., the use of anatomical data. The second group offers a more accurate solution. Nevertheless, there are situations in which the correction made by the first group is sufficient, without the need to complicate the applied technique. A brief introduction to the various existing methods for performing the attenuation map is shown in Table 6
. For a more extensive reading of this topic, readers are referred to reviews about attenuation correction [386] , [387] , [388] . Table 6 Classification of attenuation correction methods. Table 6 Category Method Definition Benefit Limitations Examples Transmission-less based Uniform fit-ellipse method (UFEM) Object outline is approximated as an ellipse around its edges. In order to generate the attenuation map, uniform attenuation is assigned Then, uniform attenuation is assigned inside the figure. Quick and easy Functional for brain studies Low precision Limited to homogeneous areas [389] Automated contour detection method (ACDM) Edge-detection algorithms are used to generate the shape of the object. This allows convex shapes It is independent of the specialist Functional for brain studies Low precision but higher than UFEM Only for homogeneous areas [390, 391] Other methods Several techniques fit here, such as algebraic reconstruction–based techniques (MLAA or MLACF) or machine-learning techniques No anatomical information is needed Possibility of cross-talk between emission data and attenuation map Virtually not used in clinical trials [392] , [393] , [394] , [395] , [396] Transmission-based Radionuclide transmission Apply an external source (PET, SPECT or SPECT/PET) interleaving transmission and emission scanning Available in most systems. Existence of complementary methods to reduce noise Need to modify the obtained attenuation coefficients due to be energy-dependent Possibility of errors due to cross-talk of data [397] , [398] , [399] , [400] CT transmission It can be addressed by segmenting the CT regions and assigning linear attenuation coefficients to each tissue or transforming the CT image to the attenuation map associated with the radiotracer photon's energy (or a combination of both methods using different scale factors for bone and tissues) Quick Low noise Good spatial resolution Functional for brain studies Misregistration due to respiratory movements Erroneus uptake due to patient's possession of unnatural materials (e.g. prostheses) CT photonic energy usually differs from that of the radionuclide used for emission scanning [401] , [402] , [403] , [404] , [405] MRI transmission Segmentation-based techniques, where first the PET and MRI images are co-registered and then a segmentation technique is applied. Usually fuzzy is used to divide the image from two to five tissues, assigning to each tissue some attenuation coefficients Atlas-based techniques, where an MR template is used instead of multistep segmentation procedures High precision Popular for brain studies Total dependence on co-registration success in the patient's image [370, [405] , [406] , [407] , [408] , [409]
Classification of attenuation correction methods.
Transmission-less based methods are easily applied in brain imaging because the brain is considered a practically homogeneous region composed mainly of soft-tissue. Moreover, these methods allow results regarding the skull [390] and optical tracking systems to be used to obtain head contours [391] . Despite being a typology that has not been used frequently, research on these methods has increased over the last years due to shift away from anatomical information. One of the first methods was proposed by Nuyts, Dupont [392] , which consists of a maximum likelihood approach known as MLAA (maximum likelihood reconstruction of attenuation and activity). It is commonly used in non-time-of-flight (TOF) PET images, and several studies have tried to improve or test it [393,
395,
396] . Another popular algorithm is MLACF (maximum likelihood attenuation correction factors) [394] , mostly used in TOF PET, which jointly estimates the attenuation sinogram and the activity image. Less popular than the previous one, this method has also been tested in other studies [410] . Moreover, machine learning techniques for attenuation correction are emerging, such as deep convolutional encoder-decoders [334] or deep convolutional neural networks [411] have also been applied in attenuation correction. For this last type of technique, there is considerable interest in this field of research, and new algorithms are constantly presented.
Regarding transmission-based techniques, the added difficulty for MRI systems to obtain attenuation maps compared to CT systems should be highlighted. The reason for this is the connection between the attenuation coefficients and the electron density of the tissue. In CT, the data is associated with electron density and photon attenuation properties of the tissues. In contrast, MRI data is correlated to proton density and magnetic relaxation properties of the tissues [387] . As a challenge related to PET/MRI (or SPECT/MRI) systems, several methods have been proposed in literature regarding attenuation estimation. The MRI-based methods can be classified into two groups: segmentation-based and atlas-based methods.
Of the two groups, the segmentation-based one tends to be used more on the literature. One of the first methods published in this area was that of Le Goff-Rougetet, Frouin [412] . This method aims to reduce the patient dose associated with PET imaging without affecting the accurate quantification. This technique is based on a surface matching technique for coregistration of PET and MR images. Another relevant proposal was based on registered T1-weighted MRI. Zaidi, Montandon [407] used a supervised fuzzy C-means clustering segmentation technique, which depends on the density and composition of five tissues: air, skull, brain tissue, and nasal sinuses. Similar to the latter case, Wagenknecht, Kops [408] presented a method where tissue classification is done with neural networks. First, a voxel classification with five tissues (GM, WM, CFS, adipose tissue (AT) and background (BG)) is made, then a second classification is done depending on the previous classes to detect extracerebral tissue and finally, segmentation is obtained. The main drawback of this method is the possibility of mis-segmentation or over-segmentation of bones, especially in the presence of abnormal anatomy or pathology.
In order to improve bone detection, ultrashort echo time (UTE) MRI began to be investigated. Firstly, dual-echo ultra-short echo time and three tissue classification (bone, soft tissue and air) were used [370,
413] . In parallel, methods related to the Dixon technique were raised. Dixon-Water–Fat-segmentation (DWFS) allows the separation of soft and adipose tissues. The method was applied DWFS for the whole body in [414] . Considering both techniques (UTE and DWFS), Berker, Franke [415] proposed a method trying to reduce the drawbacks of these two techniques: time-consuming and complex image registration. This new technique applies UTE sampling for bone detection and uses gradient echoes for water-fat separation, obtaining a four-class PET attenuation map. An example of DWFS application in neuroimaging is the study conducted by Andersen, Ladefoged [416] , where this technique is used to generate the attenuation maps. Taking into account the time-consuming problem with UTE, short echo-time (STE) has been tested as a method to obtain attenuation maps of three (cortical bone, air and soft tissue) [417] and four (cortical bone, air, soft tissue and fat tissue) tissues [418] , both in combination with fuzzy C-means (FCM) clustering method. Besides, the second study also used two-point Dixon sequences in image acquisition. The latest research in this field has been the use of zero echo time (ZTE), initially proposed by Wiesinger, Sacolick [419] as part of a segmentation method for cranial bone structures. All studies found in the literature, specially the ones related to ZTE, indicate an improvement in results compared to the use of atlas-based methods [420] , [421] , [422] , [423] .
The basis of atlas-based approaches is the use of an MR template instead of multistep segmentation procedures, so atlas-based methods use a general map while segmentation-based methods generate a map for each individual. This template can be obtained from two sources: an image labeled as if it were a segmentation of the different tissues or a coregistered attenuation map from a PET or CT scan with continuous attenuation values [354] . As usual, it is obtained from an average of co-registered normal patients, and it has to be warped to the target patient image volume. In any case, most of the studies related to this approach combine atlas-based and segmentation-based methods since the use of an atlas can be an important source of information and can reduce computational cost [424] , [425] , [426] . Despite this, other studies have proposed methods purely based on atlas images [405,
409] , such as the one developed by Johanson, Garpebring [427] where linear attenuation coefficients are predicted from a Gaussian mixture regression algorithm.
Once the attenuation map is generated, attenuation correction is applied in the process. For that, two techniques are mainly used [386] . The first one multiplies the attenuation coefficients obtained by PET image data in the sinogram (or projection) space [428] . The other procedure is performed when PET image reconstruction is performed with an iterative algorithm, using attenuation coefficients as data weighting. While PET images can apply both methods, only iterative methods are used for SPECT [429] , [430] , [431] .
Scatter in PET and SPECT is another relevant effect for absolute quantification, especially in SPECT, and which is usually associated with Compton scattering. It consists of an energy loss and a change of direction of a photon after an interaction with surrounding atoms. Nevertheless, this effect is not highly relevant in the clinical environment, according to the literature. The reason for this is that the techniques implemented have little impact on the final result, and since the photon change of direction is practically zero, the energy loss is minimal. However, several researchers are convinced that in order to obtain a high accuracy quantitative image, this artifact should be removed [432,
433] .
Several methods have been proposed for scatter correction, but most of them were developed many years ago and are inefficient. However, a few recent methods are getting remarkable results. Regarding the categorization of these techniques, a similar classification could be made for both PET and SPECT, considering that PET techniques began to be developed much earlier. Following the review of Zaidi and Montandon [432] , the classification would consist of five groups: hardware approaches using coarse septa or beam stoppers, multiple-energy-window approaches, convolution/deconvolution-based approaches, approaches based on direct estimation of scatter distribution and approaches based on statistical reconstruction. Due to the low clinical implementation and the diversity of existing methods, this work will only highlight the most innovative and relevant methods. Therefore, readers interested in this effect are advised to read the following reviews [389,
433,
434] . Table 7
summarizes the different existing categories indicating some examples. Table 7 Classification of scatter correction methods. Table 7 Category Methods Description Benefit Limitations Examples Hardware-based techniques If coarse septa or beam stoppers are used. lines of response intercepted by the septa can be used to determine the scatter component No noise increase Unused [442, 443] Multiple-energy window techniques The energy spectrum is estimated by using windows below and above the photopeak window Highly used Simple Noise [444, 445] Convolution and deconvolution-based techniques In this case, the standard energy acquisition window is used. Data collected in it helps to estimate the distribution of scatter Good image contrast Good accuracy Not commonly used [446] , [447] , [448] Direct calculation techniques Extract information from emission data, or a combination of emission and transmission data for estimating scatter distribution. Monte Carlo technique and ToF information can achieve great progress The most popular High accuracy Computational cost [449] , [450] , [451] , [452] Iterative reconstruction-based scatter-correction techniques Scatter distribution is obtained and used during image reconstruction Parallel processing High contrast Low noise Computational cost [368, [453] , [454] , [455]
Classification of scatter correction methods.
The first method proposed to narrow the photopeak energy window to avoid the acceptance of scattered photons. However, this technique has significant drawbacks, such as the elimination of unscattered photons in the process, and therefore, a loss of intensity appears in the image [434,
435] . Thus, multiple-energy-window became much more popular, and it is one of the simplest and most used approaches. The techniques have been developed from two [436,
437] three [438,
439] or even multiple energy windows [440,
441] . Decades have passed since it is known that Monte Carlo techniques are ideal for scatter correction [433] , [434] , [435] . Nevertheless, it has recently become a viable solution in the clinical environment, as the computational cost is reducing and techniques are faster. Moreover, the Monte Carlo technique can be applied to several methods, either iterative reconstruction based or direct calculation methods.
Finally, another factor to consider in the quantification of functional images is the voluntary and involuntary movement of the patient (e.g. breathing). However, since it does not have high relevance in neuroimaging, its analysis has been dismissed in this study.
Functional imaging is very useful for the diagnosis of neurodegenerative diseases, such as Alzheimer's disease or Parkinson's disease due to differences in brain activity which can be observed in temporal and parietal lobes (AD) or the striatum (PD) with respect to healthy subjects. Therefore, the systems must allow a correct quantification in order to offer an accurate diagnosis of the patient's condition.
As already mentioned, the correction of the aforementioned effects improves the quantification of the images. For example, some studies associated with attenuation correction demonstrate this improvement. Delso, Kemp [421] achieved a bias reduction of −0 . 5% using a CT-based correction instead of a regular ZTE attenuation correction in a PET/MR image. Also associated with PET/MR images, Berker, Franke [415] compared 4-class tissue segmentation and 3-class tissue segmentation, concluding that better results are obtained with a 4-class tissue segmentation, and these results are also very similar to those that would be obtained with a PET/CT system. In the study developed by Sousa, Appel [422] , a comparison between using an attenuation correction method based on ZTE and atlas-based correction shows that the former produces less variability than the latter. However, bias correction is similar in analyzed brain regions. For example, the correlation coefficient associated with anterior cortical regions is 0.99 for ZTE-based correction and 0.92 for atlas-based correction. Similar results are obtained in the study by Sgard, Khalife [423] . Other interesting results are produced by those that apply machine learning methods. For example, the deep convolutional neural network used by Yang, Park [411] for both attenuation and scatter correction designed for situations where it is difficult to use a combined CT or transmission source. The results were similar to the ones obtain with an CT-based scatter and attenuation correction.
PVC methods are also used in brain images. For example, regarding PD studies, Du, Tsui [376] proved than using a modified GTM method in brain SPECT images, the underestimation of striatal activities could be reduced to 1 . 2% from an initial value of 30%.
Finally, a correct normalization of the images can also highly increase the accuracy of the study. Salas-Gonzalez, Gorriz [456] proposed a method for intensity normalization of FP-CIT SPECT brain image based on α -stable distribution. This method was tested by Castillo-Barnes, Arenas [457] , showing significant differences between the images before and after normalizing. For more information on intensity normalization, especially associated with PD, we recommend reading [458] .
In conclusion, the possibility of improving the techniques exposed in this manuscript to reduce unwanted effects on images is highlighted. Despite being applied in clinical systems, they are not considered of high relevance due to the limited improvement they provide or, sometimes, the increase in noise they produce. Therefore, the area of absolute quantification must continue to be investigated for more accurate and faster solutions.
Introduction
Neuroimaging has been playing pivotal roles in clinical diagnosis and basic biomedical research in the past decades. As described in the following section, the most widely used imaging modalities are magnetic resonance imaging (MRI), computerized tomography (CT), positron emission tomography (PET), and single-photon emission computed tomography (SPECT). Among them, MRI itself is a non-radioactive, non-invasive, and versatile technique that has derived many unique imaging modalities, such as diffusion-weighted imaging, diffusion tensor imaging, susceptibility-weighted imaging, and spectroscopic imaging. PET is also versatile, as it may use different radiotracers to target different molecules or to trace different biologic pathways of the receptors in the body.
Therefore, these individual imaging modalities (the use of one imaging modality), with their characteristics in signal sources, energy levels, spatial resolutions, and temporal resolutions, provide complementary information on anatomical structure, pathophysiology, metabolism, structural connectivity, functional connectivity, etc. Over the past decades, everlasting efforts have been made in developing individual modalities and improving their technical performance. Directions of improvements include data acquisition and data processing aspects to increase spatial and/or temporal resolutions, improve signal-to-noise ratio and contrast to noise ratio, and reduce scan time. On application aspects, individual modalities have been widely used to meet clinical and scientific challenges. At the same time, technical developments and biomedical applications of the concert, integrated use of multiple neuroimaging modalities are trending up in both research and clinical institutions. The driving force of this trend is twofold. First, all individual modalities have their limitations. For example, some lesions in MS can appear normal in T1-weighted or T2-weighted MR images but show pathological changes in DWI or SWI images [1] . Second, a disease, disorder, or lesion may manifest itself in different forms, symptoms, or etiology; or on the other hand, different diseases may share some common symptoms or appearances [2,
3] . Therefore, an individual image modality may not be able to reveal a complete picture of the disease; and multimodal imaging modality (the use of multiple imaging modalities) may lead to a more comprehensive understanding, identify factors, and develop biomarkers of the disease.
In the narrow sense, a multimodal imaging study would mean the use of multiple imaging devices such as PET and MRI scanners, different imaging modes such as structural MRI, diffusion-weighted imaging, and magnetic resonance spectroscopy, or even different contract mechanism such as with or without contract agents in a single examination or experiment of a subject. This practice has been widely used in clinical diagnosis and medical research. For example, a routine protocol of MRI examination of a stroke patient may include T1-weighted, T1-weighted high-resolution structural MRI scans, diffusion-weighted imaging, SWI, etc [4,
5] . A protocol of an MRI study of a psychiatric disorder may contain a combination of structural MRI, functional MRI, MR spectroscopic imaging, etc [6,
7] .
In the broad sense, a multimodal imaging study may mean the use of multimodal imaging data obtained separately, from different subjects, and/or from different clinical or research sites. This practice offers the advantages of large and diverse datasets. However, it also comes with challenges of sophisticated models, complicated data normalization (that includes correction of errors and variations imbedded in data from different institutions), data fusion, and data integration [8,
9] .
In recent years, the quantity of peer-reviewed journal articles on neuroimaging has been increasing steadily. A database (PubMed) query using the keywords in titles of “neuroimaging” OR “brain imaging” returned more than 39,000 articles from 2010 to the present time when this paper was drafted in Feb 2020 ( Fig. 1
). These publications include not only applications of multimodal neuroimaging in clinical examinations and biomedical research but also methodological studies in imaging processing and fusion of multimodal neuroimaging. Fig. 1 Numbers of peer-reviewed papers with the keywords of “neuroimaging” or “brain imaging” in titles (the numbers and the bar graph were generated by PubMed in Feb 2020). Fig. 1
Numbers of peer-reviewed papers with the keywords of “neuroimaging” or “brain imaging” in titles (the numbers and the bar graph were generated by PubMed in Feb 2020).
Therefore, this present paper will focus on the following two main aspects: (1) we will review some of the recent, typical papers that exhibit the strength and limitations of the neuroimaging modalities and the corresponding analysis methods, and in particular, the needs for improved image fusion methods and (2) we will review recent methodological development in data preprocessing and data fusion in multimodal neuroimaging. We note that although we tried to cover all neuroimaging modalities, we inevitably paid more attention to MRI modalities. This is not only due to the most practical application and versatility of the MRI but also due to the limitations of our expertise. Fig. 2
shows the taxonomy of this review. Fig. 2 Taxonomy of this review. Fig. 2
Taxonomy of this review.
The main contents of the paper are organized as follows. Chapter 2 will give a brief introduction to neuroimaging, and challenges of multimodal imaging; Chapter 3 introduces the commonly used neuroimaging modalities, which include computerized tomography, positron emission tomography, single-proton emission computed tomography, and magnetic resonance imaging, which has many modalities in its own right. For each modality, we will concisely describe its signal source, energy level, spatial resolution, temporal resolution, and major applications; Chapter 4 describe applications of neuroimaging in three major areas: the developing brains, the degenerative brains, and mental disorders. In each part, we will first briefly describe what the clinical and/or biomedical problems are, we then review recent papers on how neuroimaging has been used to address these problems, and we point out what the unmet needs and challenges;
Chapters 5 to 9 are devoted to the multimodal neuroimaging fusion, covering some important procedures in data fusion. The topics are not necessarily complete and their order of presentation is not necessarily coherent with the pipeline of fusion processing. Chapter 5 reviews the fundamental methods, which covers types, rules, atlas-based segmentation, decomposition, reconstruction, and quantification; Chapter 6 reviews subjective and objective assessment of data fusion in multimodal neuroimaging; Chapter 7 reviews the advantages of data fusion in improving the spatial/temporal resolution, distortion correction, and contrast; it also reviews the benefits of these advantages in fusing structural and functional images; Chapter 8 reviews atlas-based segmentations in multimodal imaging fusion; Chapter 9 reviews the quantification in multimodal neuroimaging fusion. While the focus of this part is given to PET and SPECT, some of the approaches and principles discussed here, such as partial volume correction and attenuation (relaxation), can be applied to quantitative MRI modalities, such as DTI, ASL, quantitative susceptibility mapping (QSM), etc. Chapter 10 concludes the paper.