Convolutional Encoder-Decoder Networks for Volumetric Computed Tomography Surviews from Single- and Dual-View Topograms | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Convolutional Encoder-Decoder Networks for Volumetric Computed Tomography Surviews from Single- and Dual-View Topograms Nadav Shapira, Siddharth Bharthulwar, Peter B. Noël This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2449089/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Computed tomography (CT) is an extensively used imaging modality capable of generating detailed images of a patient’s internal anatomy for diagnostic and interventional procedures. High-resolution volumes are created by measuring and combining information along many radiographic projection angles. In current medical practice, single and dual-view two-dimensional (2D) topograms are utilized for planning the proceeding diagnostic scans and for selecting favorable acquisition parameters, either manually or automatically, as well as for dose modulation calculations. In this study, we develop modified 2D to three-dimensional (3D) encoder-decoder neural network architectures to generate CT-like volumes from single and dual-view topograms. We validate the developed neural networks on synthesized topograms from publicly available thoracic CT datasets. Finally, we assess the viability of the proposed transformational encoder-decoder architecture on both common image similarity metrics and quantitative clinical use case metrics, a first for 2D-to-3D CT reconstruction research. According to our findings, both single-input and dual-input neural networks are able to provide accurate volumetric anatomical estimates. The proposed technology will allow for improved (i) planning of diagnostic CT acquisitions, (ii) input for various dose modulation techniques, and (iii) recommendations for acquisition parameters and/or automatic parameter selection. It may also provide for an accurate attenuation correction map for positron emission tomography (PET) with only a small fraction of the radiation dose utilized. Health sciences/Medical research Health sciences/Medical research/Translational research CT reconstruction radiography transformation networks encoder-decoders neural networks Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 I. Introduction X-ray computed tomography (CT) is an extensively used imaging modality that captures three-dimensional (3D) anatomical structures, primarily for diagnostic applications 1 . Clinical CT systems rely on measuring individual x-ray projections in axial or helical acquisition modes while rotating around the patient body, eliminating spatial disadvantages of prior two-dimensional (2D) radiographic x-ray modalities 2 , 3 . Through filtered back-projection (FBP) or state-of-the-art reconstruction techniques, such as iterative reconstruction (IR) or artificial intelligence (AI), modern CT scanners can produce high-resolution volumetric images of a patient’s anatomy 4 . However, by measuring hundreds of projections at varying angles, CT scans involve ionizing radiation exposure to patients, raising numerous health concerns such as damage to DNA and the associated increased risk of cancer 5 – 8 . Dose modulation techniques 9 – 11 and manual or automatic kVp selection 12 , 13 are examples of modern CT imaging methods that provide means to significantly reduce the amount of radiation required to generate datasets of diagnostic image quality. Dose modulation allows a reduction of radiation dose levels by adjusting the x-ray flux during a scan through current modulations of the accelerated electrons that bombard the anode. Such current modulations are based on patient-specific information that is derived from either coronal or sagittal 2D topograms, or both. Such 2D topograms, referred to by different CT vendors as scouts, surviews, or tomograms, provide input for both planning of the diagnostic scan and for estimating water-equivalent diameters (WED) that represent the patient size in every cross-section and serve as a basis for actively reducing/increasing the radiation dose in large/small cross sections 14 . Planning topograms are also being used to select the accelerating electron voltage (kVp), which can be determined either manually or automatically, in order to control the quality of the beam and ensure that a large proration of x-ray photons is utilized for imaging, e.g., higher kVp values for larger patient sizes 12 , 13 , 15 . Information in 3D-surview volumes that were generated from single- or dual-view topograms has the potential to improve current and future dose modulation techniques and acquisition parameter selections for reduced patient risk. For example, dynamic bowtie filter concepts for angle-dependent x-ray flux modulation 15 – 17 would also greatly benefit from the additional information contained within 3D-surviews. Finally, 3D-surviews could enable improved planning of proceeding diagnostic scans, e.g., through better organ visualization. In recent years the appeal of deep learning and AI in the medical imaging domain has inspired several studies investigating deep-learning facilitated CT reconstruction 18 – 20 . Notable studies involve transforming and denoising low-dose CT scans with convolutional autoencoders and convolutional neural network (CNN) facilitated limited-angle reconstruction 21 – 23 . Within the specific subfield of 2D-to-3D CT mapping, Shen et al. investigate deep learning methods for generating volumetric datasets from single-view radiographic projections 22 . However, few studies investigate the advantages of utilizing stereo 2D-to-3D CT. Katsen et al. developed a CNN to generate 3D knee bone segmentations from pairs of 2D knee radiographs 23 . However, the challenge of generating volumetric datasets for additional body sections, e.g., for thoracic CT scans, from more than one clinically obtainable projection remains. In this study, we develop and implement single-view and dual-view CNNs for volumetric CT mapping from 2D topogram projections. Both neural networks follow a modified encoder-decoder architecture, as shown in Fig. 1 , and are comprised of three individual convolutional subnetworks: a feature representation subnetwork, a latent transformation subnetwork, and a generation subnetwork. To determine the theoretical viability of these encoder-decoder architectures, we train and assess their performance on publicly available CT datasets from the Lung Image Data Consortium dataset 24 . Finally, we assess the clinical viability of the two neural networks with use case-specific metrics for common CT applications, including dose modulation techniques and improved clinical scan planning that is strengthened by superior anatomy-awareness. Ii. Methods A. Data Collection and Generation of Paired Topograms The publicly available Lung Image Data Consortium (LIDC) dataset includes images from 1050 helical thoracic CT examinations that is compiled from seven academic centers and eight different CT vendors 24 . Each exam contains both volumetric pixel information as well as relevant scan parameters (e.g., tube current and tube voltage). However, the database includes topograms for only a very limited number of cases, and typically only for a single view, i.e., sagittal or coronal. To accurately train deep neural networks that generate volumetric images from input 2D topograms, a large training dataset of paired 2D-3D data is required. For this, we implemented a dedicated synthetic topogram creation process 25 . The process, which is depicted in Fig. 2 , involves simulating multiple beams that travers through clinical CT volumes from a fixed x-ray point source. For each ray starting at the predefined point source location and entrance voxel, the next point of intersection is defined as the closest voxel border to the path of the ray. The process is repeated with the point of intersection representing the new entrance point, until the ray exits the volume. For each ray, the products of distance between entrance and exit point and the voxel attenuation are aggregated. Topogram intensities are then synthesized on a simulated \(1024\times 1024\) pixel x-ray detector by utilizing the Beer-Lambert Law of x-ray attenuation 5 : $$I\left({x}_{n}\right)=I\left({x}_{0}\right){e}^{-\sum _{i=0}^{n}A\left(i\right){\Delta }{x}_{i}} \left(1\right)$$ B. Encoder-Decoder Neural Networks Our formulated solution to the 2D-to-3D mapping problem is defined through a modified encoder-decoder neural network architecture with an embedded transformation subnetwork that is based on the autoencoder architecture 26 (Fig. 2 ). Consider an autoencoder \(A\) with input \(X\) and output \(Y\) such that \(A\left(X\right)=Y\) . Through stochastic backpropagation and a predefined loss function \(L(X, Y)\) , the autoencoder iteratively updates the parameters of the neural network so that, given a representative input from the input dataset, the neural network may produce a near-identical output. We developed two architectures of encoder-decoder models which follow a similar methodology, one for a single-view input and one for a dual-view (stereo) input. We consider the problem of volumetric CT mapping from single or double 2D projection as a traditional image translation task with an additional transformation layer to increase image dimensionality. Given a coronal and/or sagittal topogram projections \({X}_{1}\) and \({X}_{2}\) , the goal of the neural network is to generate a predicted output volume \({Y}_{pred}\) such that \({Y}_{pred}= {Y}_{truth}\) , where \({Y}_{truth}\) is the ground-truth clinical CT volume. We therefore define two deep learning mapping functions \({F}_{1}\) and \({F}_{2}\) , such that \({F}_{1}\left({X}_{1}\right)={Y}_{pred}\) and \({F}_{2}({X}_{1}, {X}_{2})={Y}_{pred}\) . Via stochastic gradient descent and convolutional backpropagation, we iteratively update the model weights of \({F}_{1}\) and \({F}_{2}\) according to the ground-truth CT volume \({Y}_{truth}\) through a loss function \(L({Y}_{pred}, {Y}_{truth})\) . C. Representation Subnetwork The representation subnetwork (shown with yellow background in Figs. 1 and 3 ) is tasked with reducing the dimensionality of the input 2D topogram or topograms. Given a single-view input topogram \({X}_{1}\) , or two dual-view topograms \({X}_{1}\) and \({X}_{2}\) , a representation subnetwork is trained to generate a single output latent tensor \(L\) for each of the input topograms. The data flow of the subnetwork in both the single-view and dual-view neural networks is given as: \(1024\times 1024\times 1 \to 1024\times 1024\times 32 \to 512\times 512\times 64 \to 256\times 256\times 128 \to 128\times 128\times 256 \to 64\times 64\times 512 \to 32\times 32\times 1024 \to 8\times 8\times 4096 \to 4\times 4\times 4096\) with each ’ \(\to\) ’ representing a convolutional block with batch normalization and the Rectified Linear Unit (ReLU) activation function 27 . We chose the ReLU activation function over alternatives due to its faster convergence, and its tendency to resolve issues such as vanishing or exploding gradient, as found during initial tests. D. Transformation Subnetwork The transformation subnetwork (shown with red background in Figs. 1 and 3 ) is tasked with combining the latent tensor representations of each radiographic projection and increasing their dimensionality. Considering transformation subnetwork \(T\) , latent tensor \(L\) , and reshaped latent tensor \(Z\) , the single-view subnetwork is invoked such that \(T\left(L\right)=Z\) . In the dual-view architecture, both previously generated latent tensors \({L}_{1}\) and \({L}_{2}\) are concatenated to form a higher-dimensionality latent tensor \(Z\) . Hence, we denote the operation as \(T\left(L1, L2\right)=Z\) . In both variants of the transformation subnetwork, a single convolution with kernel size \(1\times 1\times 1\) is invoked to learn the new reshaped spatial hierarchies. E. Generation Subnetwork The generation subnetwork is the final component of the developed set of encoder-decoder neural networks and is tasked with enlarging the reshaped latent tensor \(Z\) into a final volumetric output. Considering the reshaped latent tensor \(Z\) and final output CT volume \({Y}_{pred}\) , the generation subnetwork, abstracted as \(G\) , is invoked such that \(G\left(Z\right)={Y}_{pred}\) . The data flow of the hidden convolutional layers in the generation subnetwork is given as: \(4\times 4\times 4\times 1024 \to 8\times 8\times 8\times 512 \to 16\times 16\times 16\times 256 \to 32\times 32\times 32\ times 128 \to 64\times 64\times 64\times 64 \to 128\times 128\times 512 \to 128\times 128\times 128\times 1\) , with each ’ \(\to\) ’ representing a deconvolutional block with the ReLU activation function. F. Network Training To determine the accuracy of both single-view and dual-view architectures, both neural networks are trained on the paired synthesized topogram and volumetric CT dataset. For both models, an initial learning rate of 0.0002 on the Adam optimizer is used to minimize the mean-squared-error (MSE) loss function \(L({Y}_{pred}, {Y}_{truth})\) via stochastic gradient descent and backpropagation. All training was conducted on two SLI-connected NVIDIA Tesla P100 GPUs, each with 16 GB of VRAM. Model weights are saved locally every 10 epochs, and the training process automatically terminates after convergence. The weights with the lowest average loss are preserved and serialized. G. Evaluation Metrics For quantitative evaluation, four common image similarity metrics are calculated: mean-squared error (MSE), mean average error (MAE), structural similarity (SSIM), peak signal-to-noise ratio (PSNR). Additionally, accuracy of lung/tissue segmentations that are based on thresholding the volumes produced by both neural networks and the original (ground-truth) CT volume is used as an application-focused metric with DICE score calculations 28 , 29 . Finally, we adopt a methodology to determine the quality of the recovered volumes as input for dose modulation techniques. For this, we first remove extraneous objects from the CT volume through thresholding and connected component labeling 30 . Next, we calculate the water equivalent diameter ( \({D}_{w}\) ) of each slice on the ground-truth and estimate volume slices by: $${D}_{w}=2\sqrt{\frac{{A}_{w}}{\pi }} where {A}_{w}={A}_{pixel}\times \sum \left(\frac{\mu \left(x,y\right)}{{\mu }_{water}}\right)={A}_{pixel}\times \sum \left(\frac{CT\#\left(x,y\right)}{1000}\right) \left(2\right)$$ where \(CT\#\left(x,y\right)\) represents the water-normalized attenuation of a voxel situated at coordinates \((x, y)\) , \({A}_{pixel}\) represents the area of a pixel, and \({A}_{w}\) represents the water-equivalent area of the given CT slice. This method with its underlying approximations was demonstrated to be valid in Wang et al. with both analytical and Monte Carlo methods 31 . Similar calculations and are also included as part of the AAPM task group 220 Report 14 . In addition to these quantitative image similarity metrics, qualitative analysis was performed on typical generated samples. Iii. Results The results reported in the section were averaged over all 60 test dataset volumes, i.e., CT cases that were never seen by either of the neural networks, to determine the efficacy of the proposed deep learning framework for dose modulation and patient size estimation tasks. Figures 4 and 5 display example results of the two neural networks as compared to their corresponding ground-truth CT datasets. Error maps of intensity difference between original CT volumes, representing ground-truth intensities, and 3D volumes based on 2D inputs demonstrate considerably lower errors for volumes generated from dual-view inputs compared to those generated from single-view inputs. These observations are consistent with the lower MAE and MSE values of the dual-view neural network, as given in Table 1 . In addition, a significantly lower error is observed at the patient boundary in both error maps and true-positive/false-positive/false-negative maps. These results indicate that the dual-view neural network is more adept than the single-view neural network at recovering the boundaries of the thoracic anatomy of the patients, an important consideration for advanced scan planning applications. Specific examples of areas where the dual-view neural network outperforms the single-view neural network can be observed in both the left and right upper apicoposterior lobes, as well as the medial and anterior lower lobes. Additionally, the trachea and superior mediastinum regions shown are almost fully recovered by both models. The dual-view neural network recovered the trachea almost perfectly, while the single-view neural network leaves several boundary voxels omitted. While soft tissue anatomies are recovered accurately, we note a generally poor recovery of contrast and details of boney tissue that is most evident in the spine. That said, the dual-view neural network does demonstrate more promise in recovering both soft and boney tissues. Table 1 Image similarity metrics for single- and dual-view neural network architectures. While both networks demonstrate low errors and high correspondence with the ground-truth CT datasets, there is an apparent advantage for the dual-view network for all matrices. Architecture MAE MSE SSIM PSNR DICE Single-view 0.0013 0.0012 0.8893 29.4531 0.9401 Dual-view 0.0004 0.0009 0.9396 31.5240 0.9676 The lungs in a CT volume can be (naïvely) segmented by thresholding and applying connected component analysis 30 . Comparisons of lung segmentations of neural network generated volumes and of lung segmentations from ground-truth CT volumes reveal that both networks result in accurately segmented lungs, with DICE scores of 0.9555 and 0.9789 for the single and dual-view neural networks, respectively (Table 1 ). The high-fidelity lung/tissue-segmentations can also be perceived from true-positive (TP) / false-positive (FP) / false-negative (FN) error maps for both models, where the volumes generated from dual-view neural network present fewer false-negative and false-positive voxels than the volumes generated from the single-view neural network. Examples of patient sizes estimations, as represented by water equivalent areas ( \({A}_{w})\) distributions along the patient axis, are displayed in Fig. 6 . Note the similarities in mesh geometries between the ground truth representation (a), the single-view neural network representation (b), and the dual-view neural network representation (c). Calculated average water equivalent diameters ( \({D}_{w}\) ) resulted in accurate patient size estimations for both neural networks, with the single-view architecture scoring 0.911 and the dual-view architecture scoring 0.925. Based on these results, the two neural networks, and especially the dual-view variant, demonstrate a potential to improve dose modulation techniques by providing accurate volumetric information as input for calculating efficient dose-modulated CT acquisitions. Iv. Discussion In this investigation, we develop and refine two encoder-decoder neural network architectures for recovering volumetric CT data from single- or dual-view radiographic projections. The developed models were trained and tested on paired data of publicly available CT datasets and synthetically generated CT tomograms (Figure 2). While the generated 3D-surviews do not contain sufficient information for diagnostic tasks and are thus not considered as an alternative to volumetric CT datasets, both the single and the dual-view networks perform well on image similarity and use-case specific metrics. Moreover, our results indicate that, in most cases, the additional information that is available with the dual-view neural network enables improved accuracy as compared to the single-view neural network. Specifically, from qualitative observations and quantitative image similarity metrics across multiple samples, we acknowledge that the dual-view neural network is generally superior at recovering both the anatomy and soft-tissue/bone intensities than the single-view neural network and is thus more relevant for further clinical translation developments. However, both models perform comparably to the current state-of-the-art in deep learning CT recovery from limited information 21–23 , indicating the viability of both architectures in practice. Based on our qualitative evaluation, the DICE scores from Table 1, and the water-equivalent area and diameter comparisons (Figure 6), we deduce that there are only minor differences between patient size estimations that rely on the complete information available with the (retrospective) diagnostic scan and patient size estimations that rely on information that was extracted by the single or dual-view neural networks. Our results indicate that the developed deep neural networks can provide a promising method with potential applications for automatic acquisition parameter selection 32 , e.g., automatic kVp selection 12,13 , more efficient dose modulation techniques that utilize improved high-dimensional input, e.g., for various dynamic bowtie filter concepts 15–17 , and improved scan planning with better anatomy-awareness. Tube current modulation during diagnostic CT scans is crucial a technique that allows a significant reduction of ionizing radiation dose, and thus improved patient safety, while maintaining diagnostic image quality. Conventionally, metrics such as patient size, measured as water-equivalent diameter, or local (along the patient axis) cylindrical or elliptical cross-section representations, also measured in water-equivalent diameter, are used to calculate position/angle-dependent radiation dosages 14 . In addition, calculated patient sizes provide the basis for size-specific dose estimates 14 (SSDE). Finally, the additional information may be of sufficient resolution to enable sufficiently accurate tissue attenuation inputs for the generation of PET attenuation correction maps 33 with a significant reduction of ionizing radiation dose levels, i.e., use of single- or dual-view tomograms instead of a complete volumetric scan. In conclusion, we demonstrate an ability to utilize low-dose single- or dual-view CT topograms, which in the current clinical practice are acquired in almost all CT examinations, for estimating volumetric patient anatomies with high accuracy. Due to the abundance of matched CT-topogram data, such neural networks can be trained and integrated into all existing CT scanners relatively seemingly, i.e., without requiring any hardware changes, to enable acquisition parameter recommendations or automatic selections, a reduction in ionizing radiation dose levels through improved dose modulation calculations, and have a potential to enable PET attenuation correction maps with only a fraction of the currently utilized radiation dose. Declarations Acknowledgment We acknowledge support through the National Institutes of Health (R01EB030494) and Philips Healthcare. Availability of Data and Materials The datasets used and/or analyzed during the current study available from the corresponding author on reasonable request. References Bercovich, E. & Javitt, M. C. Medical Imaging: From Roentgen to the Digital Revolution, and Beyond. Rambam Maimonides Medical Journal 9 , e0034 (2018). Exadaktylos, A. K., Sclabas, G., Schmid, S. W., Schaller, B. & Zimmermann, H. Do We Really Need Routine Computed Tomographic Scanning in the Primary Evaluation of Blunt Chest Trauma in Patients with “Normal” Chest Radiograph? Journal of Trauma and Acute Care Surgery 51 , (2001). Gross, B. H. & Spizarny, D. L. Computed Tomography of the Chest in the Intensive Care Unit. Critical Care Clinics 10 , 267–275 (1994). Willemink, M. J. & Noël, P. B. The evolution of image reconstruction for CT—from filtered back projection to artificial intelligence. European Radiology 29 , 2185–2195 (2019). Bushberg, J. T., Seibert, J. Anthony., Leidholdt, E. Marion. & Boone, J. M. The essential physics of medical imaging . Cohen, M. D. ALARA, Image Gently and CT-induced cancer. Pediatric Radiology 45 , 465–470 (2015). International Commission on Radiological Protection. ICRP Publication 103 . ICRP 103 , (2007). McCollough, C. H., Primak, A. N., Braun, N., Kofler, J., Yu, L. & Christner, J. Strategies for Reducing Radiation Dose in CT. Radiologic Clinics of North America 47 , 27–40 (2009). Lee, C. H., Goo, J. M., Lee, H. J., Ye, S. J., Park, C. M., Chun, E. J. & Im, J. G. Radiation dose modulation techniques in the multidetector CT era: From basics to practice. Radiographics 28 , 1451–1459 (2008). McCollough, C. H., Bruesewitz, M. R. & Kofler, J. M. CT dose reduction and dose management tools: Overview of available options. Radiographics 26 , 503–512 (2006). Kalra, M. K., Maher, M. M., Toth, T. L., Schmidt, B., Westerman, B. L., Morgan, H. T. & Saini, S. Techniques and applications of automatic tube current modulation for CT. Radiology 233 , 649–657 (2004). Schindera, S. T., Winklehner, A., Alkadhi, H., Goetti, R., Fischer, M., Gnannt, R. & Szucs-Farkas, Z. Effect of automatic tube voltage selection on image quality and radiation dose in abdominal CT angiography of various body sizes: A phantom study. Clinical Radiology 68 , e79–e86 (2013). Niemann, T., Henry, S., Faivre, J. B., Yasunaga, K., Bendaoud, S., Simeone, A., Remy, J., Duhamel, A., Flohr, T. & Remy-Jardin, M. Clinical evaluation of automatic tube voltage selection in chest CT angiography. European Radiology 23 , 2643–2651 (2013). McCollough, C., Bakalyar, D. M., Bostani, M., Brady, S., Boedeker, K., Boone, J. M., Chen-Mayer, H. H., Christianson, O. I., Leng, S., Li, B., McNitt-Gray, M. F., Nilsen, R. A., Supanich, M. P. & Wang, J. Use of Water Equivalent Diameter for Calculating Patient Size and Size-Specific Dose Estimates (SSDE) in CT: The Report of AAPM Task Group 220. AAPM Rep 2014 , 6–23 (2014). Szczykutowicz, T. P. & Mistretta, C. A. Design of a digital beam attenuation system for computed tomography: Part I. System design and simulation framework. Medical Physics 40 , 021905 (2013). Hsieh, S. S. & Pelc, N. J. The feasibility of a piecewise‐linear dynamic bowtie filter. Medical Physics 40 , 031910 (2013). Gang, G. J., Mao, A., Wang, W., Siewerdsen, J. H., Mathews, A., Kawamoto, S., Levinson, R. & Stayman, J. W. Dynamic fluence field modulation in computed tomography using multiple aperture devices. Physics in Medicine & Biology 64 , 105024 (2019). Nakamura, Y., Higaki, T., Tatsugami, F., Honda, Y., Narita, K., Akagi, M. & Awai, K. Possibility of Deep Learning in Medical Imaging Focusing Improvement of Computed Tomography Image Quality. J Comput Assist Tomogr 44 , 161–167 (2020). Wang, G., Ye, J. C. & de Man, B. Deep learning for tomographic image reconstruction. Nature Machine Intelligence 2020 2:12 2 , 737–748 (2020). Baguer, D. O., Leuschner, J. & Schmidt, M. Computed tomography reconstruction using deep image prior and learned reconstruction methods. Inverse Problems 36 , 094004 (2020). Würfl, T., Hoffmann, M., Christlein, V., Breininger, K., Huang, Y., Unberath, M. & Maier, A. K. Deep Learning Computed Tomography: Learning Projection-Domain Weights From Image Domain in Limited Angle Problems. IEEE Transactions on Medical Imaging 37 , 1454–1463 (2018). Shen, L., Zhao, W. & Xing, L. Patient-specific reconstruction of volumetric computed tomography images from a single projection view via deep learning. Nature Biomedical Engineering 2019 3:11 3 , 880–888 (2019). Kasten, Y., Doktofsky, D. & Kovler, I. End-To-End Convolutional Neural Network for 3D Reconstruction of Knee Bones From Bi-Planar X-Ray Images. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 12450 LNCS , 123–133 (2020). Armato, S. G., McLennan, G., Bidaut, L., McNitt-Gray, M. F., Meyer, C. R., Reeves, A. P., Zhao, B., Aberle, D. R., Henschke, C. I., Hoffman, E. A., Kazerooni, E. A., MacMahon, H., van Beek, E. J. R., Yankelevitz, D., Biancardi, A. M., Bland, P. H., Brown, M. S., … Clarke, L. P. The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A Completed Reference Database of Lung Nodules on CT Scans. Medical Physics 38 , 915–931 (2011). Kengyelics, S. M., Treadgold, L. A. & Davies, A. G. X-ray system simulation software tools for radiology and radiography education. Computers in Biology and Medicine 93 , 175–183 (2018). Baldi, P. Autoencoders, unsupervised learning, and deep architectures. in Proceedings of ICML workshop on unsupervised and transfer learning 37–49 (2012). Fukushima. Visual Feature Extraction by a Multilayered Network of Analog Threshold Elements. IEEE Transactions on Systems Science and Cybernetics 5 , 322–333 (1969). Dice, L. R. Measures of the Amount of Ecologic Association Between Species. Ecology 26 , 297–302 (1945). Sørensen, T. A method of establishing groups of equal amplitude in plant sociology based on similarity of species content and its application to analyses of the vegetation on Danish commons. (I kommission hos E. Munksgaard, 1948). Samet, H. & Tamminen, M. K. Efficient Component Labeling of Images of Arbitrary Dimension Represented by Linear Bintrees. IEEE Transactions on Pattern Analysis and Machine Intelligence 10 , 579–586 (1988). Wang, J., Duan, X., Christner, J. A., Leng, S., Yu, L. & McCollough, C. H. Attenuation-based estimation of patient size for the purpose of size specific dose estimation in CT. Part I. Development and validation of methods using the CT image. Medical Physics 39 , 6764–6771 (2012). Mayer, C., Meyer, M., Fink, C., Schmidt, B., Sedlmair, M., Schoenberg, S. O. & Henzler, T. Potential for Radiation Dose Savings in Abdominal and Chest CT Using Automatic Tube Voltage Selection in Combination With Automatic Tube Current Modulation. http://dx.doi.org/10.2214/AJR.13.11628 203 , 292–299 (2014). Kinahan, P. E., Townsend, D. W., Beyer, T. & Sashin, D. Attenuation correction for a combined 3D PET/CT scanner. Medical Physics 25 , 2046–2053 (1998). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2449089","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":166777705,"identity":"45741bc5-a1e0-4ff8-bcac-546c66fd4e12","order_by":0,"name":"Nadav Shapira","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIiWNgGAWjYHACAwglwXwARMoQqwWIJdgSQFp4SNHCA7aOsBZz9sMbP/Mw/JHjn93z+dWNGgseBvbDRzfg02LZk1YszcNgYCxx5+w265xjQIfxpKXdwOuqAzkGkjMYDBI3SORuM85hA2qR4DHDr+X8G+OfQC31GyRynhnn/CNGy40cM4kPDAYJBhI5zI9z24jS8qzM4oOBseGMG2lmzLl9EjxsBP1yPnnzjYQKOXn+GcmPP+d8q5PjZz98DK8WqEYwySYBJgkrRwDmD6SoHgWjYBSMgpEDAGszQn01TdvjAAAAAElFTkSuQmCC","orcid":"","institution":"Perelman School of Medicine of the University of Pennsylvania","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Nadav","middleName":"","lastName":"Shapira","suffix":""},{"id":166777706,"identity":"9f5b28e2-e0e3-4be4-bac3-84db894c1da7","order_by":1,"name":"Siddharth Bharthulwar","email":"","orcid":"","institution":"Perelman School of Medicine of the University of Pennsylvania","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Siddharth","middleName":"","lastName":"Bharthulwar","suffix":""},{"id":166777707,"identity":"f2edf382-b725-4264-8e95-4e5a7aaaab2d","order_by":2,"name":"Peter B. Noël","email":"","orcid":"","institution":"Perelman School of Medicine of the University of Pennsylvania","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Peter","middleName":"B.","lastName":"Noël","suffix":""}],"badges":[],"createdAt":"2023-01-06 06:59:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2449089/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2449089/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":31563128,"identity":"db7eb502-651d-4893-8c74-92a5a480c492","added_by":"auto","created_at":"2023-01-13 22:44:54","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":139873,"visible":true,"origin":"","legend":"\u003cp\u003eSchematic workflow diagram of the implemented encoder-decoder neural network architecture for single- or dual-view topogram (2D) to volume (3D) mapping. The workflow uses single- or dual-view topograms as input and includes a representation subnetworks (yellow), a transformation subnetwork (red), and a generation subnetwork (blue) to estimate CT-like volumetric outputs.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/a68b3cae60c3e3fef9036847.png"},{"id":31563240,"identity":"a4bf5970-240b-43b2-9d7a-70f808e1f376","added_by":"auto","created_at":"2023-01-13 22:52:54","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":174970,"visible":true,"origin":"","legend":"\u003cp\u003eIllustration of the synthetic topogram creation process we applied to generate coronal (left) and sagittal (right) topograms from 1050 publicly available clinical CT volumes. The process involves simulating multiple x-ray beams that traverse through CT volumes from a fixed x-ray point source by utilizing the Beer-Lambert Law of attenuation, where final intensities are measured on a simulated detector with 1024 × 1024 pixels behind the patient body.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/e597f16fe6534a24e200ea93.png"},{"id":31563125,"identity":"a3852fe7-e20d-4162-8ce1-1f26ed1e2712","added_by":"auto","created_at":"2023-01-13 22:44:54","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":255159,"visible":true,"origin":"","legend":"\u003cp\u003eArchitecture of a dual-input neural network, featuring four distinct modules: two representation subnetworks (yellow), a transformation subnetwork (red), and a generation subnetwork (blue). The number of channels is indicated by the number below each hidden layer. A coronal input view topogram (C) and a sagittal input view topogram (S), both with dimensions 1024 × 1024 × 1 are passed through separate but identical representation subnetworks and transformed into a volumetric output with dimensions 128 x 128 x 128. A similar architecture, with only a single representation network, is utilized for the dual-input neural network.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/44420f2120c7b7b99064549d.png"},{"id":31563239,"identity":"eba37ad8-d061-42d5-98c1-54140e5b43df","added_by":"auto","created_at":"2023-01-13 22:52:54","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":543604,"visible":true,"origin":"","legend":"\u003cp\u003eVolumetric surview (3D-surview) results for an example CT dataset of a single patient viewed in axial (top row), coronal (middle row) and sagittal (bottom row) anatomical planes. Displayed columns are (a) ground-truth CT slices, (b) slices generated by the single-view neural network, (c) a true positive (green), false positive (blue), and false negative (red) map for single-view volume generation, (d) absolute HU error map for single-view volume generation, (e) slices generated by the single-view neural network, (f) a true positive (green), false positive (blue), and false negative (red) map for single-view volume generation, (g) absolute HU error map for single-view volume generation.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/839bda8f5c87588ae7057bc2.png"},{"id":31563241,"identity":"eb38bf22-5224-46ce-a356-0074ca819393","added_by":"auto","created_at":"2023-01-13 22:52:54","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":805352,"visible":true,"origin":"","legend":"\u003cp\u003eVolumetric surview (3D-surview) results for an example CT dataset of a single patient viewed in multiple axial slice depths (different rows). Displayed columns are (a) ground-truth CT slices, (b) slices generated by the single-view neural network, (c) a true positive (green), false positive (blue), and false negative (red) map for single-view volume generation, (d) absolute HU error map for single-view volume generation, (e) slices generated by the single-view neural network, (f) a true positive (green), false positive (blue), and false negative (red) map for single-view volume generation, (g) absolute HU error map for single-view volume generation.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/81f8d79b2827a6acddaaad99.png"},{"id":31563130,"identity":"6c1f5ac8-78d6-45f9-b8a8-74679c8a2122","added_by":"auto","created_at":"2023-01-13 22:44:54","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":506506,"visible":true,"origin":"","legend":"\u003cp\u003eAn example of volumetric external bodies and volumetric water equivalent areas \u003cem\u003e(A\u003c/em\u003e\u003csub\u003e\u003cem\u003eW \u003c/em\u003e\u003c/sub\u003e\u003cem\u003e)\u003c/em\u003e for ground-truth datasets and neural network generated 3D-surview volumes. Displayed are (a) meshes created from the water-equivalent-areas of a ground-truth CT volume, (b) a CT-like 3D-surview generated by the single-view neural network, and (c) a CT-like 3D-surview generated by the dual-view neural network. Our results demonstrate excellent recovery of body size representations for a variety of applications, e.g., radiation dose modulation.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/a890603364fab8c90998fb91.png"},{"id":41220473,"identity":"b39d2850-fefd-4acc-86ed-751f28dcca34","added_by":"auto","created_at":"2023-08-08 06:07:25","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2547831,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2449089/v1/33ea82e1-66a7-43c4-9da2-ccc2a17c3a01.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Convolutional Encoder-Decoder Networks for Volumetric Computed Tomography Surviews from Single- and Dual-View Topograms","fulltext":[{"header":"I. Introduction","content":"\u003cp\u003eX-ray computed tomography (CT) is an extensively used imaging modality that captures three-dimensional (3D) anatomical structures, primarily for diagnostic applications\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. Clinical CT systems rely on measuring individual x-ray projections in axial or helical acquisition modes while rotating around the patient body, eliminating spatial disadvantages of prior two-dimensional (2D) radiographic x-ray modalities\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Through filtered back-projection (FBP) or state-of-the-art reconstruction techniques, such as iterative reconstruction (IR) or artificial intelligence (AI), modern CT scanners can produce high-resolution volumetric images of a patient\u0026rsquo;s anatomy\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. However, by measuring hundreds of projections at varying angles, CT scans involve ionizing radiation exposure to patients, raising numerous health concerns such as damage to DNA and the associated increased risk of cancer\u003csup\u003e\u003cspan additionalcitationids=\"CR6 CR7\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Dose modulation techniques\u003csup\u003e\u003cspan additionalcitationids=\"CR10\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e and manual or automatic kVp selection\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e are examples of modern CT imaging methods that provide means to significantly reduce the amount of radiation required to generate datasets of diagnostic image quality. Dose modulation allows a reduction of radiation dose levels by adjusting the x-ray flux during a scan through current modulations of the accelerated electrons that bombard the anode. Such current modulations are based on patient-specific information that is derived from either coronal or sagittal 2D topograms, or both. Such 2D topograms, referred to by different CT vendors as scouts, surviews, or tomograms, provide input for both planning of the diagnostic scan and for estimating water-equivalent diameters (WED) that represent the patient size in every cross-section and serve as a basis for actively reducing/increasing the radiation dose in large/small cross sections\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. Planning topograms are also being used to select the accelerating electron voltage (kVp), which can be determined either manually or automatically, in order to control the quality of the beam and ensure that a large proration of x-ray photons is utilized for imaging, e.g., higher kVp values for larger patient sizes\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. Information in 3D-surview volumes that were generated from single- or dual-view topograms has the potential to improve current and future dose modulation techniques and acquisition parameter selections for reduced patient risk. For example, dynamic bowtie filter concepts for angle-dependent x-ray flux modulation\u003csup\u003e\u003cspan additionalcitationids=\"CR16\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e would also greatly benefit from the additional information contained within 3D-surviews. Finally, 3D-surviews could enable improved planning of proceeding diagnostic scans, e.g., through better organ visualization.\u003c/p\u003e \u003cp\u003eIn recent years the appeal of deep learning and AI in the medical imaging domain has inspired several studies investigating deep-learning facilitated CT reconstruction\u003csup\u003e\u003cspan additionalcitationids=\"CR19\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Notable studies involve transforming and denoising low-dose CT scans with convolutional autoencoders and convolutional neural network (CNN) facilitated limited-angle reconstruction\u003csup\u003e\u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. Within the specific subfield of 2D-to-3D CT mapping, Shen \u003cem\u003eet al.\u003c/em\u003e investigate deep learning methods for generating volumetric datasets from single-view radiographic projections\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. However, few studies investigate the advantages of utilizing stereo 2D-to-3D CT. Katsen \u003cem\u003eet al.\u003c/em\u003e developed a CNN to generate 3D knee bone segmentations from pairs of 2D knee radiographs\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. However, the challenge of generating volumetric datasets for additional body sections, e.g., for thoracic CT scans, from more than one clinically obtainable projection remains.\u003c/p\u003e \u003cp\u003eIn this study, we develop and implement single-view and dual-view CNNs for volumetric CT mapping from 2D topogram projections. Both neural networks follow a modified encoder-decoder architecture, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, and are comprised of three individual convolutional subnetworks: a feature representation subnetwork, a latent transformation subnetwork, and a generation subnetwork. To determine the theoretical viability of these encoder-decoder architectures, we train and assess their performance on publicly available CT datasets from the Lung Image Data Consortium dataset \u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Finally, we assess the clinical viability of the two neural networks with use case-specific metrics for common CT applications, including dose modulation techniques and improved clinical scan planning that is strengthened by superior anatomy-awareness.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Ii. Methods","content":"\u003cp\u003eA. Data Collection and Generation of Paired Topograms\u003c/p\u003e\n\u003cp\u003eThe publicly available Lung Image Data Consortium (LIDC) dataset includes images from 1050 helical thoracic CT examinations that is compiled from seven academic centers and eight different CT vendors\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Each exam contains both volumetric pixel information as well as relevant scan parameters (e.g., tube current and tube voltage). However, the database includes topograms for only a very limited number of cases, and typically only for a single view, i.e., sagittal or coronal. To accurately train deep neural networks that generate volumetric images from input 2D topograms, a large training dataset of paired 2D-3D data is required. For this, we implemented a dedicated synthetic topogram creation process\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. The process, which is depicted in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e, involves simulating multiple beams that travers through clinical CT volumes from a fixed x-ray point source. For each ray starting at the predefined point source location and entrance voxel, the next point of intersection is defined as the closest voxel border to the path of the ray. The process is repeated with the point of intersection representing the new entrance point, until the ray exits the volume. For each ray, the products of distance between entrance and exit point and the voxel attenuation are aggregated. Topogram intensities are then synthesized on a simulated \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(1024\\times 1024\\)\u003c/span\u003e\u003c/span\u003e pixel x-ray detector by utilizing the Beer-Lambert Law of x-ray attenuation\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e:\u003c/p\u003e\n\u003cdiv class=\"Equation\" id=\"Equa\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e$$I\\left({x}_{n}\\right)=I\\left({x}_{0}\\right){e}^{-\\sum _{i=0}^{n}A\\left(i\\right){\\Delta }{x}_{i}} \\left(1\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cspan style=\"text-align: inherit;\"\u003eB. Encoder-Decoder Neural Networks\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003eOur formulated solution to the 2D-to-3D mapping problem is defined through a modified encoder-decoder neural network architecture with an embedded transformation subnetwork that is based on the autoencoder architecture\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). Consider an autoencoder \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(A\\)\u003c/span\u003e\u003c/span\u003e with input \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(X\\)\u003c/span\u003e\u003c/span\u003e and output \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Y\\)\u003c/span\u003e\u003c/span\u003e such that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(A\\left(X\\right)=Y\\)\u003c/span\u003e\u003c/span\u003e. Through stochastic backpropagation and a predefined loss function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(L(X, Y)\\)\u003c/span\u003e\u003c/span\u003e, the autoencoder iteratively updates the parameters of the neural network so that, given a representative input from the input dataset, the neural network may produce a near-identical output. We developed two architectures of encoder-decoder models which follow a similar methodology, one for a single-view input and one for a dual-view (stereo) input. We consider the problem of volumetric CT mapping from single or double 2D projection as a traditional image translation task with an additional transformation layer to increase image dimensionality. Given a coronal and/or sagittal topogram projections \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X}_{2}\\)\u003c/span\u003e\u003c/span\u003e, the goal of the neural network is to generate a predicted output volume \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({Y}_{pred}\\)\u003c/span\u003e\u003c/span\u003e such that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({Y}_{pred}= {Y}_{truth}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({Y}_{truth}\\)\u003c/span\u003e\u003c/span\u003e is the ground-truth clinical CT volume. We therefore define two deep learning mapping functions \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{2}\\)\u003c/span\u003e\u003c/span\u003e, such that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{1}\\left({X}_{1}\\right)={Y}_{pred}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{2}({X}_{1}, {X}_{2})={Y}_{pred}\\)\u003c/span\u003e\u003c/span\u003e. Via stochastic gradient descent and convolutional backpropagation, we iteratively update the model weights of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{2}\\)\u003c/span\u003e\u003c/span\u003e according to the ground-truth CT volume \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({Y}_{truth}\\)\u003c/span\u003e\u003c/span\u003e through a loss function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(L({Y}_{pred}, {Y}_{truth})\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003eC. Representation Subnetwork\u003c/p\u003e\n\u003cp\u003eThe representation subnetwork (shown with yellow background in Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e) is tasked with reducing the dimensionality of the input 2D topogram or topograms. Given a single-view input topogram \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X}_{1}\\)\u003c/span\u003e\u003c/span\u003e, or two dual-view topograms \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X}_{2}\\)\u003c/span\u003e\u003c/span\u003e, a representation subnetwork is trained to generate a single output latent tensor \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(L\\)\u003c/span\u003e\u003c/span\u003e for each of the input topograms. The data flow of the subnetwork in both the single-view and dual-view neural networks is given as: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(1024\\times 1024\\times 1 \\to 1024\\times 1024\\times 32 \\to 512\\times 512\\times 64 \\to 256\\times 256\\times 128 \\to 128\\times 128\\times 256 \\to 64\\times 64\\times\u003cbr/\u003e\u0026nbsp;512 \\to 32\\times 32\\times 1024 \\to 8\\times 8\\times 4096 \\to 4\\times 4\\times 4096\\)\u003c/span\u003e\u003c/span\u003e with each \u0026rsquo;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\to\\)\u003c/span\u003e\u003c/span\u003e\u0026rsquo; representing a convolutional block with batch normalization and the Rectified Linear Unit (ReLU) activation function\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. We chose the ReLU activation function over alternatives due to its faster convergence, and its tendency to resolve issues such as vanishing or exploding gradient, as found during initial tests.\u003c/p\u003e\n\u003cp\u003eD. Transformation Subnetwork\u003c/p\u003e\n\u003cp\u003eThe transformation subnetwork (shown with red background in Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e) is tasked with combining the latent tensor representations of each radiographic projection and increasing their dimensionality. Considering transformation subnetwork \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(T\\)\u003c/span\u003e\u003c/span\u003e, latent tensor \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(L\\)\u003c/span\u003e\u003c/span\u003e, and reshaped latent tensor \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Z\\)\u003c/span\u003e\u003c/span\u003e, the single-view subnetwork is invoked such that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(T\\left(L\\right)=Z\\)\u003c/span\u003e\u003c/span\u003e. In the dual-view architecture, both previously generated latent tensors \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({L}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({L}_{2}\\)\u003c/span\u003e\u003c/span\u003e are concatenated to form a higher-dimensionality latent tensor \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Z\\)\u003c/span\u003e\u003c/span\u003e. Hence, we denote the operation as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(T\\left(L1, L2\\right)=Z\\)\u003c/span\u003e\u003c/span\u003e. In both variants of the transformation subnetwork, a single convolution with kernel size \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(1\\times 1\\times 1\\)\u003c/span\u003e\u003c/span\u003e is invoked to learn the new reshaped spatial hierarchies.\u003c/p\u003e\n\u003cp\u003e\u003cspan style=\"text-align: inherit;\"\u003eE. Generation Subnetwork\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003eThe generation subnetwork is the final component of the developed set of encoder-decoder neural networks and is tasked with enlarging the reshaped latent tensor \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Z\\)\u003c/span\u003e\u003c/span\u003e into a final volumetric output. Considering the reshaped latent tensor \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Z\\)\u003c/span\u003e\u003c/span\u003e and final output CT volume \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({Y}_{pred}\\)\u003c/span\u003e\u003c/span\u003e, the generation subnetwork, abstracted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(G\\)\u003c/span\u003e\u003c/span\u003e, is invoked such that \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(G\\left(Z\\right)={Y}_{pred}\\)\u003c/span\u003e\u003c/span\u003e. The data flow of the hidden convolutional layers in the generation subnetwork is given as: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(4\\times 4\\times 4\\times 1024 \\to 8\\times 8\\times 8\\times 512 \\to 16\\times 16\\times 16\\times 256 \\to 32\\times 32\\times 32\\\u003cbr/\u003etimes 128 \\to 64\\times 64\\times 64\\times 64 \\to 128\\times 128\\times 512 \\to 128\\times 128\\times 128\\times 1\\)\u003c/span\u003e\u003c/span\u003e, with each \u0026rsquo;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\to\\)\u003c/span\u003e\u003c/span\u003e\u0026rsquo; representing a deconvolutional block with the ReLU activation function.\u003c/p\u003e\n\u003cp\u003eF. Network Training\u003c/p\u003e\n\u003cp\u003eTo determine the accuracy of both single-view and dual-view architectures, both neural networks are trained on the paired synthesized topogram and volumetric CT dataset. For both models, an initial learning rate of 0.0002 on the Adam optimizer is used to minimize the mean-squared-error (MSE) loss function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(L({Y}_{pred}, {Y}_{truth})\\)\u003c/span\u003e\u003c/span\u003e via stochastic gradient descent and backpropagation. All training was conducted on two SLI-connected NVIDIA Tesla P100 GPUs, each with 16 GB of VRAM. Model weights are saved locally every 10 epochs, and the training process automatically terminates after convergence. The weights with the lowest average loss are preserved and serialized.\u003c/p\u003e\n\u003cp\u003eG. Evaluation Metrics\u003c/p\u003e\n\u003cp\u003eFor quantitative evaluation, four common image similarity metrics are calculated: mean-squared error (MSE), mean average error (MAE), structural similarity (SSIM), peak signal-to-noise ratio (PSNR). Additionally, accuracy of lung/tissue segmentations that are based on thresholding the volumes produced by both neural networks and the original (ground-truth) CT volume is used as an application-focused metric with DICE score calculations\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e,\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. Finally, we adopt a methodology to determine the quality of the recovered volumes as input for dose modulation techniques. For this, we first remove extraneous objects from the CT volume through thresholding and connected component labeling\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. Next, we calculate the water equivalent diameter (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({D}_{w}\\)\u003c/span\u003e\u003c/span\u003e) of each slice on the ground-truth and estimate volume slices by:\u003c/p\u003e\n\u003cdiv class=\"Equation\" id=\"Equb\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e$${D}_{w}=2\\sqrt{\\frac{{A}_{w}}{\\pi }} where {A}_{w}={A}_{pixel}\\times \\sum \\left(\\frac{\\mu \\left(x,y\\right)}{{\\mu }_{water}}\\right)={A}_{pixel}\\times \\sum \\left(\\frac{CT\\#\\left(x,y\\right)}{1000}\\right) \\left(2\\right)$$\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(CT\\#\\left(x,y\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the water-normalized attenuation of a voxel situated at coordinates \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\((x, y)\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({A}_{pixel}\\)\u003c/span\u003e\u003c/span\u003e represents the area of a pixel, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({A}_{w}\\)\u003c/span\u003e\u003c/span\u003e represents the water-equivalent area of the given CT slice. This method with its underlying approximations was demonstrated to be valid in Wang \u003cem\u003eet al.\u003c/em\u003e with both analytical and Monte Carlo methods\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Similar calculations and are also included as part of the AAPM task group 220 Report\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eIn addition to these quantitative image similarity metrics, qualitative analysis was performed on typical generated samples.\u003c/p\u003e"},{"header":"Iii. Results","content":"\u003cp\u003eThe results reported in the section were averaged over all 60 test dataset volumes, i.e., CT cases that were never seen by either of the neural networks, to determine the efficacy of the proposed deep learning framework for dose modulation and patient size estimation tasks.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigures \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e display example results of the two neural networks as compared to their corresponding ground-truth CT datasets. Error maps of intensity difference between original CT volumes, representing ground-truth intensities, and 3D volumes based on 2D inputs demonstrate considerably lower errors for volumes generated from dual-view inputs compared to those generated from single-view inputs. These observations are consistent with the lower MAE and MSE values of the dual-view neural network, as given in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. In addition, a significantly lower error is observed at the patient boundary in both error maps and true-positive/false-positive/false-negative maps. These results indicate that the dual-view neural network is more adept than the single-view neural network at recovering the boundaries of the thoracic anatomy of the patients, an important consideration for advanced scan planning applications. Specific examples of areas where the dual-view neural network outperforms the single-view neural network can be observed in both the left and right upper apicoposterior lobes, as well as the medial and anterior lower lobes. Additionally, the trachea and superior mediastinum regions shown are almost fully recovered by both models. The dual-view neural network recovered the trachea almost perfectly, while the single-view neural network leaves several boundary voxels omitted.\u003c/p\u003e \u003cp\u003eWhile soft tissue anatomies are recovered accurately, we note a generally poor recovery of contrast and details of boney tissue that is most evident in the spine. That said, the dual-view neural network does demonstrate more promise in recovering both soft and boney tissues.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eImage similarity metrics for single- and dual-view neural network architectures. While both networks demonstrate low errors and high correspondence with the ground-truth CT datasets, there is an apparent advantage for the dual-view network for all matrices.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eArchitecture\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSSIM\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePSNR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDICE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSingle-view\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.0013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.0012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.8893\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e29.4531\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9401\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDual-view\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.0004\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.0009\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.9396\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e31.5240\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9676\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe lungs in a CT volume can be (na\u0026iuml;vely) segmented by thresholding and applying connected component analysis \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. Comparisons of lung segmentations of neural network generated volumes and of lung segmentations from ground-truth CT volumes reveal that both networks result in accurately segmented lungs, with DICE scores of 0.9555 and 0.9789 for the single and dual-view neural networks, respectively (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The high-fidelity lung/tissue-segmentations can also be perceived from true-positive (TP) / false-positive (FP) / false-negative (FN) error maps for both models, where the volumes generated from dual-view neural network present fewer false-negative and false-positive voxels than the volumes generated from the single-view neural network.\u003c/p\u003e \u003cp\u003eExamples of patient sizes estimations, as represented by water equivalent areas (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({A}_{w})\\)\u003c/span\u003e\u003c/span\u003e distributions along the patient axis, are displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. Note the similarities in mesh geometries between the ground truth representation (a), the single-view neural network representation (b), and the dual-view neural network representation (c). Calculated average water equivalent diameters (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({D}_{w}\\)\u003c/span\u003e\u003c/span\u003e) resulted in accurate patient size estimations for both neural networks, with the single-view architecture scoring 0.911 and the dual-view architecture scoring 0.925. Based on these results, the two neural networks, and especially the dual-view variant, demonstrate a potential to improve dose modulation techniques by providing accurate volumetric information as input for calculating efficient dose-modulated CT acquisitions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Iv.\tDiscussion","content":"\u003cp\u003eIn this investigation, we develop and refine two encoder-decoder neural network architectures for recovering volumetric CT data from single- or dual-view radiographic projections. The developed models were trained and tested on paired data of publicly available CT datasets and synthetically generated CT tomograms (Figure 2). While the generated 3D-surviews do not contain sufficient information for diagnostic tasks and are thus not considered as an alternative to volumetric CT datasets, both the single and the dual-view networks perform well on image similarity and use-case specific metrics. Moreover, our results indicate that, in most cases, the additional information that is available with the dual-view neural network enables improved accuracy as compared to the single-view neural network. Specifically, from qualitative observations and quantitative image similarity metrics across multiple samples, we acknowledge that the dual-view neural network is generally superior at recovering both the anatomy and soft-tissue/bone intensities than the single-view neural network and is thus more relevant for further clinical translation developments. However, both models perform comparably to the current state-of-the-art in deep learning CT recovery from limited information\u003csup\u003e\u0026nbsp;21\u0026ndash;23\u003c/sup\u003e, indicating the viability of both architectures in practice.\u003c/p\u003e\n\u003cp\u003eBased on our qualitative evaluation, the DICE scores from Table 1, and the water-equivalent area and diameter comparisons (Figure 6), we deduce that there are only minor differences between patient size estimations that rely on the complete information available with the (retrospective) diagnostic scan and patient size estimations that rely on information that was extracted by the single or dual-view neural networks. Our results indicate that the developed deep neural networks can provide a promising method with potential applications for automatic acquisition parameter selection\u003csup\u003e32\u003c/sup\u003e, e.g., automatic kVp selection\u003csup\u003e\u0026nbsp;12,13\u003c/sup\u003e, more efficient dose modulation techniques that utilize improved high-dimensional input, e.g., for various dynamic bowtie filter concepts\u003csup\u003e15\u0026ndash;17\u003c/sup\u003e, and improved scan planning with better anatomy-awareness. Tube current modulation during diagnostic CT scans is crucial a technique that allows a significant reduction of ionizing radiation dose, and thus improved patient safety, while maintaining diagnostic image quality. Conventionally, metrics such as patient size, measured as water-equivalent diameter, or local (along the patient axis) cylindrical or elliptical cross-section representations, also measured in water-equivalent diameter, are used to calculate position/angle-dependent radiation dosages\u003csup\u003e14\u003c/sup\u003e. In addition, calculated patient sizes provide the basis for size-specific dose estimates\u003csup\u003e14\u003c/sup\u003e (SSDE). Finally, the additional information may be of sufficient resolution to enable sufficiently accurate tissue attenuation inputs for the generation of PET attenuation correction maps\u003csup\u003e33\u003c/sup\u003e with a significant reduction of ionizing radiation dose levels, i.e., use of single- or dual-view tomograms instead of a complete volumetric scan.\u003c/p\u003e\n\u003cp\u003eIn conclusion, we demonstrate an ability to utilize low-dose single- or dual-view CT topograms, which in the current clinical practice are acquired in almost all CT examinations, for estimating volumetric patient anatomies with high accuracy. Due to the abundance of matched CT-topogram data, such neural networks can be trained and integrated into all existing CT scanners relatively seemingly, i.e., without requiring any hardware changes, to enable acquisition parameter recommendations or automatic selections, a reduction in ionizing radiation dose levels through improved dose modulation calculations, and have a potential to enable PET attenuation correction maps with only a fraction of the currently utilized radiation dose.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eAcknowledgment\u003c/p\u003e\n\u003cp\u003eWe acknowledge support through the National Institutes of Health (R01EB030494) and Philips Healthcare. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAvailability of Data and Materials\u003c/p\u003e\n\u003cp\u003eThe datasets used and/or analyzed during the current study available from the corresponding author on reasonable request.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBercovich, E. \u0026amp; Javitt, M. C. Medical Imaging: From Roentgen to the Digital Revolution, and Beyond. \u003cem\u003eRambam Maimonides Medical Journal\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, e0034 (2018).\u003c/li\u003e\n\u003cli\u003eExadaktylos, A. K., Sclabas, G., Schmid, S. W., Schaller, B. \u0026amp; Zimmermann, H. Do We Really Need Routine Computed Tomographic Scanning in the Primary Evaluation of Blunt Chest Trauma in Patients with \u0026ldquo;Normal\u0026rdquo; Chest Radiograph? \u003cem\u003eJournal of Trauma and Acute Care Surgery\u003c/em\u003e \u003cstrong\u003e51\u003c/strong\u003e, (2001).\u003c/li\u003e\n\u003cli\u003eGross, B. H. \u0026amp; Spizarny, D. L. Computed Tomography of the Chest in the Intensive Care Unit. \u003cem\u003eCritical Care Clinics\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 267\u0026ndash;275 (1994).\u003c/li\u003e\n\u003cli\u003eWillemink, M. J. \u0026amp; No\u0026euml;l, P. B. The evolution of image reconstruction for CT\u0026mdash;from filtered back projection to artificial intelligence. \u003cem\u003eEuropean Radiology\u003c/em\u003e \u003cstrong\u003e29\u003c/strong\u003e, 2185\u0026ndash;2195 (2019).\u003c/li\u003e\n\u003cli\u003eBushberg, J. T., Seibert, J. Anthony., Leidholdt, E. Marion. \u0026amp; Boone, J. M. \u003cem\u003eThe essential physics of medical imaging\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eCohen, M. D. ALARA, Image Gently and CT-induced cancer. \u003cem\u003ePediatric Radiology\u003c/em\u003e \u003cstrong\u003e45\u003c/strong\u003e, 465\u0026ndash;470 (2015).\u003c/li\u003e\n\u003cli\u003eInternational Commission on Radiological Protection. \u003cem\u003eICRP Publication 103\u003c/em\u003e. \u003cem\u003eICRP\u003c/em\u003e \u003cstrong\u003e103\u003c/strong\u003e, (2007).\u003c/li\u003e\n\u003cli\u003eMcCollough, C. H., Primak, A. N., Braun, N., Kofler, J., Yu, L. \u0026amp; Christner, J. Strategies for Reducing Radiation Dose in CT. \u003cem\u003eRadiologic Clinics of North America\u003c/em\u003e \u003cstrong\u003e47\u003c/strong\u003e, 27\u0026ndash;40 (2009).\u003c/li\u003e\n\u003cli\u003eLee, C. H., Goo, J. M., Lee, H. J., Ye, S. J., Park, C. M., Chun, E. J. \u0026amp; Im, J. G. Radiation dose modulation techniques in the multidetector CT era: From basics to practice. \u003cem\u003eRadiographics\u003c/em\u003e \u003cstrong\u003e28\u003c/strong\u003e, 1451\u0026ndash;1459 (2008).\u003c/li\u003e\n\u003cli\u003eMcCollough, C. H., Bruesewitz, M. R. \u0026amp; Kofler, J. M. CT dose reduction and dose management tools: Overview of available options. \u003cem\u003eRadiographics\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e, 503\u0026ndash;512 (2006).\u003c/li\u003e\n\u003cli\u003eKalra, M. K., Maher, M. M., Toth, T. L., Schmidt, B., Westerman, B. L., Morgan, H. T. \u0026amp; Saini, S. Techniques and applications of automatic tube current modulation for CT. \u003cem\u003eRadiology\u003c/em\u003e \u003cstrong\u003e233\u003c/strong\u003e, 649\u0026ndash;657 (2004).\u003c/li\u003e\n\u003cli\u003eSchindera, S. T., Winklehner, A., Alkadhi, H., Goetti, R., Fischer, M., Gnannt, R. \u0026amp; Szucs-Farkas, Z. Effect of automatic tube voltage selection on image quality and radiation dose in abdominal CT angiography of various body sizes: A phantom study. \u003cem\u003eClinical Radiology\u003c/em\u003e \u003cstrong\u003e68\u003c/strong\u003e, e79\u0026ndash;e86 (2013).\u003c/li\u003e\n\u003cli\u003eNiemann, T., Henry, S., Faivre, J. B., Yasunaga, K., Bendaoud, S., Simeone, A., Remy, J., Duhamel, A., Flohr, T. \u0026amp; Remy-Jardin, M. Clinical evaluation of automatic tube voltage selection in chest CT angiography. \u003cem\u003eEuropean Radiology\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 2643\u0026ndash;2651 (2013).\u003c/li\u003e\n\u003cli\u003eMcCollough, C., Bakalyar, D. M., Bostani, M., Brady, S., Boedeker, K., Boone, J. M., Chen-Mayer, H. H., Christianson, O. I., Leng, S., Li, B., McNitt-Gray, M. F., Nilsen, R. A., Supanich, M. P. \u0026amp; Wang, J. Use of Water Equivalent Diameter for Calculating Patient Size and Size-Specific Dose Estimates (SSDE) in CT: The Report of AAPM Task Group 220. \u003cem\u003eAAPM Rep\u003c/em\u003e \u003cstrong\u003e2014\u003c/strong\u003e, 6\u0026ndash;23 (2014).\u003c/li\u003e\n\u003cli\u003eSzczykutowicz, T. P. \u0026amp; Mistretta, C. A. Design of a digital beam attenuation system for computed tomography: Part I. System design and simulation framework. \u003cem\u003eMedical Physics\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 021905 (2013).\u003c/li\u003e\n\u003cli\u003eHsieh, S. S. \u0026amp; Pelc, N. J. The feasibility of a piecewise‐linear dynamic bowtie filter. \u003cem\u003eMedical Physics\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 031910 (2013).\u003c/li\u003e\n\u003cli\u003eGang, G. J., Mao, A., Wang, W., Siewerdsen, J. H., Mathews, A., Kawamoto, S., Levinson, R. \u0026amp; Stayman, J. W. Dynamic fluence field modulation in computed tomography using multiple aperture devices. \u003cem\u003ePhysics in Medicine \u0026amp; Biology\u003c/em\u003e \u003cstrong\u003e64\u003c/strong\u003e, 105024 (2019).\u003c/li\u003e\n\u003cli\u003eNakamura, Y., Higaki, T., Tatsugami, F., Honda, Y., Narita, K., Akagi, M. \u0026amp; Awai, K. Possibility of Deep Learning in Medical Imaging Focusing Improvement of Computed Tomography Image Quality. \u003cem\u003eJ Comput Assist Tomogr\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, 161\u0026ndash;167 (2020).\u003c/li\u003e\n\u003cli\u003eWang, G., Ye, J. C. \u0026amp; de Man, B. Deep learning for tomographic image reconstruction. \u003cem\u003eNature Machine Intelligence 2020 2:12\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 737\u0026ndash;748 (2020).\u003c/li\u003e\n\u003cli\u003eBaguer, D. O., Leuschner, J. \u0026amp; Schmidt, M. Computed tomography reconstruction using deep image prior and learned reconstruction methods. \u003cem\u003eInverse Problems\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 094004 (2020).\u003c/li\u003e\n\u003cli\u003eW\u0026uuml;rfl, T., Hoffmann, M., Christlein, V., Breininger, K., Huang, Y., Unberath, M. \u0026amp; Maier, A. K. Deep Learning Computed Tomography: Learning Projection-Domain Weights From Image Domain in Limited Angle Problems. \u003cem\u003eIEEE Transactions on Medical Imaging\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 1454\u0026ndash;1463 (2018).\u003c/li\u003e\n\u003cli\u003eShen, L., Zhao, W. \u0026amp; Xing, L. Patient-specific reconstruction of volumetric computed tomography images from a single projection view via deep learning. \u003cem\u003eNature Biomedical Engineering 2019 3:11\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 880\u0026ndash;888 (2019).\u003c/li\u003e\n\u003cli\u003eKasten, Y., Doktofsky, D. \u0026amp; Kovler, I. End-To-End Convolutional Neural Network for 3D Reconstruction of Knee Bones From Bi-Planar X-Ray Images. \u003cem\u003eLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)\u003c/em\u003e \u003cstrong\u003e12450 LNCS\u003c/strong\u003e, 123\u0026ndash;133 (2020).\u003c/li\u003e\n\u003cli\u003eArmato, S. G., McLennan, G., Bidaut, L., McNitt-Gray, M. F., Meyer, C. R., Reeves, A. P., Zhao, B., Aberle, D. R., Henschke, C. I., Hoffman, E. A., Kazerooni, E. A., MacMahon, H., van Beek, E. J. R., Yankelevitz, D., Biancardi, A. M., Bland, P. H., Brown, M. S., \u0026hellip; Clarke, L. P. The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A Completed Reference Database of Lung Nodules on CT Scans. \u003cem\u003eMedical Physics\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, 915\u0026ndash;931 (2011).\u003c/li\u003e\n\u003cli\u003eKengyelics, S. M., Treadgold, L. A. \u0026amp; Davies, A. G. X-ray system simulation software tools for radiology and radiography education. \u003cem\u003eComputers in Biology and Medicine\u003c/em\u003e \u003cstrong\u003e93\u003c/strong\u003e, 175\u0026ndash;183 (2018).\u003c/li\u003e\n\u003cli\u003eBaldi, P. Autoencoders, unsupervised learning, and deep architectures. in \u003cem\u003eProceedings of ICML workshop on unsupervised and transfer learning\u003c/em\u003e 37\u0026ndash;49 (2012).\u003c/li\u003e\n\u003cli\u003eFukushima. Visual Feature Extraction by a Multilayered Network of Analog Threshold Elements. \u003cem\u003eIEEE Transactions on Systems Science and Cybernetics\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 322\u0026ndash;333 (1969).\u003c/li\u003e\n\u003cli\u003eDice, L. R. Measures of the Amount of Ecologic Association Between Species. \u003cem\u003eEcology\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e, 297\u0026ndash;302 (1945).\u003c/li\u003e\n\u003cli\u003eS\u0026oslash;rensen, T. \u003cem\u003eA method of establishing groups of equal amplitude in plant sociology based on similarity of species content and its application to analyses of the vegetation on Danish commons.\u003c/em\u003e (I kommission hos E. Munksgaard, 1948).\u003c/li\u003e\n\u003cli\u003eSamet, H. \u0026amp; Tamminen, M. K. Efficient Component Labeling of Images of Arbitrary Dimension Represented by Linear Bintrees. \u003cem\u003eIEEE Transactions on Pattern Analysis and Machine Intelligence\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 579\u0026ndash;586 (1988).\u003c/li\u003e\n\u003cli\u003eWang, J., Duan, X., Christner, J. A., Leng, S., Yu, L. \u0026amp; McCollough, C. H. Attenuation-based estimation of patient size for the purpose of size specific dose estimation in CT. Part I. Development and validation of methods using the CT image. \u003cem\u003eMedical Physics\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 6764\u0026ndash;6771 (2012).\u003c/li\u003e\n\u003cli\u003eMayer, C., Meyer, M., Fink, C., Schmidt, B., Sedlmair, M., Schoenberg, S. O. \u0026amp; Henzler, T. Potential for Radiation Dose Savings in Abdominal and Chest CT Using Automatic Tube Voltage Selection in Combination With Automatic Tube Current Modulation. \u003cem\u003ehttp://dx.doi.org/10.2214/AJR.13.11628\u003c/em\u003e \u003cstrong\u003e203\u003c/strong\u003e, 292\u0026ndash;299 (2014).\u003c/li\u003e\n\u003cli\u003eKinahan, P. E., Townsend, D. W., Beyer, T. \u0026amp; Sashin, D. Attenuation correction for a combined 3D PET/CT scanner. \u003cem\u003eMedical Physics\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 2046\u0026ndash;2053 (1998).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"CT reconstruction, radiography, transformation networks, encoder-decoders, neural networks","lastPublishedDoi":"10.21203/rs.3.rs-2449089/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2449089/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eComputed tomography (CT) is an extensively used imaging modality capable of generating detailed images of a patient\u0026rsquo;s internal anatomy for diagnostic and interventional procedures. High-resolution volumes are created by measuring and combining information along many radiographic projection angles. In current medical practice, single and dual-view two-dimensional (2D) topograms are utilized for planning the proceeding diagnostic scans and for selecting favorable acquisition parameters, either manually or automatically, as well as for dose modulation calculations. In this study, we develop modified 2D to three-dimensional (3D) encoder-decoder neural network architectures to generate CT-like volumes from single and dual-view topograms. We validate the developed neural networks on synthesized topograms from publicly available thoracic CT datasets. Finally, we assess the viability of the proposed transformational encoder-decoder architecture on both common image similarity metrics and quantitative clinical use case metrics, a first for 2D-to-3D CT reconstruction research. According to our findings, both single-input and dual-input neural networks are able to provide accurate volumetric anatomical estimates. The proposed technology will allow for improved (i) planning of diagnostic CT acquisitions, (ii) input for various dose modulation techniques, and (iii) recommendations for acquisition parameters and/or automatic parameter selection. It may also provide for an accurate attenuation correction map for positron emission tomography (PET) with only a small fraction of the radiation dose utilized.\u003c/p\u003e","manuscriptTitle":"Convolutional Encoder-Decoder Networks for Volumetric Computed Tomography Surviews from Single- and Dual-View Topograms","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-01-13 22:44:49","doi":"10.21203/rs.3.rs-2449089/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"07683838-aa14-426b-a359-5d215fedd0b1","owner":[],"postedDate":"January 13th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":18313683,"name":"Health sciences/Medical research"},{"id":18313684,"name":"Health sciences/Medical research/Translational research"}],"tags":[],"updatedAt":"2023-08-08T05:59:16+00:00","versionOfRecord":[],"versionCreatedAt":"2023-01-13 22:44:49","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2449089","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2449089","identity":"rs-2449089","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.