Comparison of AI software tools for automated detection, quantification and categorization of pulmonary nodules in the HANSE LCS trial

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study compared two AI tools for pulmonary nodule detection and quantification in CT scans, finding significant volume differences that could alter Lung-RADS classifications and patient management.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Abstract Purpose To compare the performance of two AI-based software tools for detection, quantification and categorization of pulmonary nodules in a lung cancer screening (LCS) program in Northern Germany (HANSE-trial). Method 946 low-dose baseline CT-examinations were analyzed by two AI software tools regarding lung nodule detection, quantification and categorization and compared to the final radiologist read. The relationship between detected nodule volumes by both software tools was assessed by Pearson correlation (r) and tested for significance using Wilcoxon signed-rank test. The consistency of Lung-RADS classifications was evaluated by Cohen’s kappa (κ) and percentual agreement (PA). Results 1032 (88%) and 782 (66%) of all (n = 1174, solid, semi-solid and ground-glass) lung nodules (volume ≥ 34mm3) were detected by Software tool 1 (S1) and Software tool 2 (S2), respectively. Although, the derived volumes of true positive nodules were strongly correlated (r > 0.95), the volume derived by S2 was significantly higher than by S1 (P < 0.0001, mean difference: 6mm3). Moderate PA (62%) between S1 and S2 was found in the assignment of Lung-RADS classification (κ = 0.45). The PA of Lung-RADS classification to final read was 75% and 55% for S1 and S2. Conclusion Participant management depends on the assigned Lung Imaging Reporting and Data System (Lung-RADS) category, which is based on reliable detection and volumetry of pulmonary nodules. Significant nodule volume differences between AI software tools lead to different Lung-RADS scores in 38% of cases, which may result in altered participant management. Therefore, high performance and agreement of accredited AI software tools are necessary for a future national LCS program.
Full text 125,912 characters · extracted from preprint-html · click to expand
Comparison of AI software tools for automated detection, quantification and categorization of pulmonary nodules in the HANSE LCS trial | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Comparison of AI software tools for automated detection, quantification and categorization of pulmonary nodules in the HANSE LCS trial Rimma Kondrashova, Filip Klimeš, Till Frederik Kaireit, Katharina May, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3392224/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Purpose To compare the performance of two AI-based software tools for detection, quantification and categorization of pulmonary nodules in a lung cancer screening (LCS) program in Northern Germany (HANSE-trial). Method 946 low-dose baseline CT-examinations were analyzed by two AI software tools regarding lung nodule detection, quantification and categorization and compared to the final radiologist read. The relationship between detected nodule volumes by both software tools was assessed by Pearson correlation ( r ) and tested for significance using Wilcoxon signed-rank test. The consistency of Lung-RADS classifications was evaluated by Cohen’s kappa ( κ ) and percentual agreement ( PA ). Results 1032 (88%) and 782 (66%) of all (n = 1174, solid, semi-solid and ground-glass) lung nodules (volume ≥ 34mm 3 ) were detected by Software tool 1 (S1) and Software tool 2 (S2), respectively. Although, the derived volumes of true positive nodules were strongly correlated ( r > 0.95), the volume derived by S2 was significantly higher than by S1 ( P < 0.0001, mean difference: 6mm 3 ). Moderate PA (62%) between S1 and S2 was found in the assignment of Lung-RADS classification ( κ = 0.45). The PA of Lung-RADS classification to final read was 75% and 55% for S1 and S2. Conclusion Participant management depends on the assigned Lung Imaging Reporting and Data System (Lung-RADS) category, which is based on reliable detection and volumetry of pulmonary nodules. Significant nodule volume differences between AI software tools lead to different Lung-RADS scores in 38% of cases, which may result in altered participant management. Therefore, high performance and agreement of accredited AI software tools are necessary for a future national LCS program. Biological sciences/Cancer/Lung cancer Biological sciences/Cancer/Cancer screening Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Lung cancer is the leading cause of cancer death worldwide. ( 1 – 2 ) If lung cancer is detected at an early stage, it is more likely to be treated curatively. According to the American Thoracic Society, screening for lung cancer is recommended for people between 50–80 years old and in fairly good health with at least a 20-pack-year smoking history who currently smoke or have quit in the past 15 years. ( 3 ) During the last years, multiple large multicenter lung cancer screening (LCS) studies worldwide such as The National Lung Screening Trial in the US, Dutch-Belgian Randomized Lung Cancer Screening Trial (NELSON), Italian Lung Cancer Screening Trial (ITALUNG) or German Lung Cancer Screening Intervention Trial (LUSI) concluded that lung cancer screening using low-dose computed tomography (LDCT) reduces lung cancer mortality. ( 4 – 8 ) Recently a holistic northern German interdisciplinary lung cancer screening study (HANSE) ( 9 ) has been started for optimal definition of the high-risk screening population and implementation of the screening workflow in the current national healthcare infrastructure. Lung nodule management in the HANSE study depends on the Lung Imaging Reporting and Data System (Lung-RADS 1.1) assessment. ( 9 – 11 ) Five categories to discriminate between high-risk and low-risk nodules are part of the Lung-RADS assessment. After the detection of a nodule and correct recognition of nodule type, an accurate measurement of the size of detected nodules is a prerequisite for subsequent successful nodule management. ( 12 ) Recent studies have shown that classification based on mean nodule diameter leads to a massive overestimation of true nodule size, therefore classification based on volume is recommended. ( 13 ) Manual 3D volume segmentation of pulmonary nodules is time-consuming and error-prone; therefore, radiologists are nowadays assisted by artificial intelligence (AI)-based software tools when reading LDCT examinations. ( 14 – 17 ) There are currently 17 products on the market that can detect, quantify, and categorize pulmonary nodules with distinct detection performance. ( 18 ) Accurate detection and quantification are essential for the correct categorization of detected nodules. If a different Lung-RADS grade is assigned to the same subject by different software, the participant may be treated differently. In some cases, variability in the volumetric assessment of pulmonary nodules may result in a false-positive or false-negative diagnosis. ( 19 ) Such inconsistency may occur, for example, when different software is used among different centers as parts of a national screening program. ( 20 ) With the wide range of AI software available, it is important to know if there are significant differences in lung nodule detection, quantification and classification between the different AI tools. The purpose of this study was to investigate and compare the performance of two AI-based software tools regarding lung nodule detection, quantification, and categorization in the HANSE trial population prior to the implementation of a national lung cancer screening program in Germany. ( 21 ) Method Study population The study participants were part of the HANSE LCS trial ( 9 ), a northern Germany multicenter LCS study. In this study, 946 LDCT randomly selected baseline participant examinations, performed (between July 2021 and July 2022) in a German Cancer Society-certified lung cancer center at Hannover Medical School, were retrospectively analyzed. The institutional review board of all participating institutions approved the HANSE study and all study participants provided written informed consent. The methods employed in this study conform to the principles outlined in the Declaration of Helsinki. The image acquisition details are listed in the Supplement. LDCT evaluation After the first reading of LDCTs by one experienced radiologist with specialized training in LCS, a computer-aided detection (CAD) technology of Software tool 1 (S1) automatically detected, segmented, and classified pulmonary nodules. The nodule candidates suggested by CAD technology, were flagged by the radiologist as either correct- or false-positive findings. Subsequently, the false positives were excluded from the report. Further, nodules were automatically classified into solid, non-solid, calcified and part-solid groups according to their attenuation. Inaccuracies in the measurement of volume and diameter as well as the nodule group affiliation were corrected using a threshold-based algorithm if necessary. Based on the measured volume and assigned nodule group, an individual nodule Lung-RADS score was determined. The overall Lung-RADS score for each participant was defined by the highest individual nodule Lung-RADS score. Cases classified as 3 or 4 according to Lung-RADS 1.1 scoring system ( 22 ) were reviewed by another experienced board-certified thoracic radiologist. In case of nodule classification discrepancy between both readers, the case was presented to the radiological conference for a consensus evaluation. The final decision, consisting of the first read by the radiologist, AI-assistance and the second read of Lung-RADS 3–4 cases, was considered as a final reading (FR) and defined the reference gold standard for comparison of AI software tools. Similarly to S1, software tool 2 (S2) was used for automated nodule detection, segmentation and classification. Both software tools were trained using deep learning and were dedicated for lung cancer screening. ( 23 ) While S1 was used prospectively, S2 analyzed the results of the LDCT retrospectively using the same criteria. True Positive (TP) nodules were identified as those detected in the FR dataset. Nodules detected by the S1 / S2 but not confirmed by the FR dataset were defined as false positives (FP). Nodules found in the FR dataset but not in the S1 / S2 dataset were designated as false negatives (FN). The matching of detected nodules from the S1 / S2 dataset to the FR dataset is described in the Supplement. Statistical Analysis Statistical analysis was performed with JMP Pro 16 software (SAS Institute, Cary, NC). Nodule Detection Nodule detection performance was assessed by sensitivity and positive predictive value (PPV) between FR and S1, FR and S2 of all detected nodules. An identical sub-analysis was performed for clinically relevant nodules between FR and S1 and FR and S2. Filters of 34mm 3 and 113mm 3 volume were applied. These thresholds correspond to the Lung-RADS 1.1 scoring system. Additionally, in the group of nodules ≥ 34mm 3 volume, the same analysis was executed for nodule subgroups (solid, non-solid, calcified and part-solid). Nodule quantification After excluding the FP nodules, both software tools (S1 and S2) and the FR dataset were compared regarding nodule diameter and volume. In the S1 vs S2 comparison, only nodules detected by both software tools were included. Nodule diameter and nodule volume parameters were tested for normality using the Shapiro-Wilk test. Since both parameters were not normally distributed, a non-parametric paired Wilcoxon signed-rank test was used to assess diameter and volume median differences in the following comparisons: FR vs S1, FR vs S2 and S1 vs S2. P values < 0.05 was considered as a statistically significant difference. Further, the mean differences were quantified by Bland-Altman analysis and the association between the measured diameters and volumes was explored by Pearson correlation analysis ( r ). The relative volumetric errors between the final reading and the two software tools were computed as a percentage difference between two measuring volumes divided by the mean of the two values. The same analysis was also conducted for TP nodules after the application of 34mm 3 and 113mm 3 volume thresholds. Nodule categorization On the participant and individual nodule level, the agreement of Lung-RADS classification of both software systems (S1, S2) with the ground truth (FR) and between each other (S1 vs S2) was determined by Cohen's kappa coefficient (κ) and percentual agreement (PA). PA was defined as a fraction of the number of Lung-RADS classifications in agreement and the total number of Lung-RADS classifications. Finally, the fraction was converted to PA by multiplication with 100. At the nodule level, when comparing the two software tools with each other, only the correctly positive nodules detected by both software tools were scored for fair comparison. Consistently, the identical analysis was conducted at the nodule level for nodules with volume ≥ 34mm 3 and ≥ 113mm 3 . Additionally, on the participant level, the whole analysis was repeated for participants with a total Lung-RADS score ≥ 3 and participants with a total Lung-RADS score < 3. Further details regarding the statistics of Lung-RADS misclassification at the individual nodule level and false-positive rate on the participant level are presented in the Supplement. Results To reflect the real world of LCS both nodules and masses (tumors with diameter larger than three cm) were examined. The volume of the smallest detected nodule and the largest mass was 0.3mm 3 and 18.9cm 3 respectively. In the study population of 946 participants, a total of 3345 lung nodules were found in the FR dataset and 765 out of 946 subjects had at least one nodule. The number of detected nodules ranged from 1 to 86 per subject. Of the 3345 nodules, 1174 were found clinically relevant (nodule volume ≥ 34mm 3 ). Besides 791 solid nodules, 258, 112, and 13 were non-solid, calcified, and part-solid, respectively. Of 1174 clinically relevant nodules, 282 nodules had volumes greater than 113mm 3 . Five masses were found in the final read of which three were detected by each software tool. Nodule Detection Examples of FP and FN pulmonary nodule detection are shown in Fig. 1 – 3 . Sensitivities and PPVs of the S1 software tool were higher when compared to the S2 software tool for all nodules, for nodules with volume ≥ 34mm 3 and nodules with volume ≥ 113mm 3 (see Table 1 ). Table 2 shows detection results for clinically relevant nodules ≥ 34mm 3 concerning their type. Sensitivities and PPVs were higher for S1 for all nodule types. The lowest sensitivities were observed for part-solid and non-solid using S1 and S2, respectively. For both software tools, the highest PPV was found for non-solid nodules. Table 1 Detection performance of both software tools in all nodules, in nodules with volume ≥ 34 mm 3 and nodules with volume ≥ 113 mm 3 . Software tool Nodule volume TP FP FN Sensitivity [%] PPV [%] S1 all 2142 316 1203 64 87 ≥ 34 mm 3 1032 210 142 88 83 ≥ 113 mm 3 234 105 48 83 69 S2 all 1526 538 1819 46 74 ≥ 34 mm 3 782 459 392 66 63 ≥ 113 mm 3 202 305 80 72 40 FN, false negative; FP, false positive; PPV, positive predictive value; S1, software tool 1 dataset; S2, software tool 2 dataset; TP, true positive. Table 2 Detection performance of both software in subgroups of clinically relevant nodules with volume ≥ 34 mm 3 . Software tool Nodule type TP FP FN Sensitivity [%] PPV [%] S1 Solid 720 175 71 91 80 Non-Solid 207 16 51 80 92 Calcified 97 18 15 87 84 Part-Solid 8 1 5 62 80 S2 Solid 570 321 221 72 63 Non-Solid 121 9 137 47 93 Calcified 83 57 29 74 59 Part-Solid 8 3 5 61 72 FN, false negative; FP, false positive; PPV, positive predictive value; S1, software tool 1 dataset; S2, software tool 2 dataset; TP, true positive. Nodule Quantification Although the derived volumes of all TP nodules from S1 and S2 datasets were strongly correlated (all r > 0.95), the volume derived by S2 was significantly larger than that obtained by S1 (both P < 0.0001, see Table 3 ), except for the nodules with volume ≥ 113mm 3 . In Fig. 4 , Bland-Altman analysis between S1 and S2 software tools is depicted. Table 3 Comparison of detected TP nodules regarding median and mean volume for FR, S1 and S2 datasets. Median (IQR) Mean (SD) Mean Bias (relative difference in %) P value a Pearson correlation r All S1 35.6 (52.6) 97.7 (418.4) -13.3 (-) < 0.0001* 0.95 S2 58.8 (52.8) 111.1 (375.8) ≥ 34 mm 3 S1 63.0 (73.0) 164.4 (566.9) -6.0 (-) < 0.0001* 0.95 S2 85.3 (73.9) 170.5 (509.2) ≥ 113 mm 3 S1 183.6 (204.9) 470.6 (1079.5) 31.6 (-) 0.31 0.95 S2 199.5 (159.2) 439.0 (971.8) All FR 35.3 (51.9) 106.9 (446.9) -15.7 (-46%) < 0.0001* 0.81 S2 56.0 (52.6) 122.5 (580.6) ≥ 34 mm 3 FR 68.9 (72.0) 191.2 (612.6) -7.6 (-15%) < 0.0001* 0.81 S2 86.1 (77.0) 198.8 (803.6) ≥ 113 mm 3 FR 198.6 (211.6) 565.1 (1125.9) 29.7 (11%) 0.14 0.80 S2 209.2 (177.1) 535.5 (1533.2) All FR 32.4 (47.8) 88.5 (527.3) 9.5 (4%) 0.24 0.59 S1 31.5 (48.5) 79.0 (343.4) ≥ 34 mm 3 FR 67.1 (63.6) 164.8 (752.3) 22.6 (10%) 0.0037* 0.58 S1 61.7 (69.7) 142.3 (486.6) ≥ 113 mm 3 FR 186.2 (168.3) 517.6 (1530.1) 106.4 (21%) 0.0083* 0.55 S1 177.4 (172.1) 411.2 (973.1) FR, final reading dataset; IQR, interquartile range; Mean bias, mean difference derived from Bland-Altman analysis; S1, software tool 1 dataset; S2, software tool 2 dataset; SD, standard deviation; r , Pearson correlation coefficient a paired Wilcoxon signed rank test, significantly different measurements ( P < 0.05) are marked with *. It should be noted that each comparison exclusively incorporates TP nodules detected either by both software tools or by software tool and final read. Consequently, this leads to changed volume numbers for the same software tool across various comparisons. In comparison between FR and S2 datasets (see Fig. 5 ), all volumes obtained by the S2 software tool were significantly larger (both P < 0.0001, Table 3 ), except for nodules with volume ≥ 113mm 3 . Comparing the S1 dataset with the FR dataset (Fig. 6 ), FR volume measurements were found significantly higher (both P < 0.0037, Table 3 ), except for the comparison of all nodules. The results regarding the diameter comparison are presented in the Supplement (Supporting Information Table S1 ). Nodule categorization Supporting Information Table S2 shows the individual nodule categorization comparison according to its size, measured by volume. There was a good agreement in comparison of S1 and S2 datasets (all PA > 67%, all κ > 0.61). Higher Lung-RADS consensus was achieved between FR vs S1 (all PA > 69%, κ > 0.58) when compared to FR vs S2 (all PA > 55%, κ > 0.44). For all comparisons, the Lung-RADS agreement was decreased with nodule size. Similarly, on a participant level, a moderate agreement of Lung-RADS categorization was observed in the comparison of S1 and S2 datasets (all PA > 54%, all κ > 0.41, see Table 4 ). As shown in Table 4 , the comparison of FR vs S1 reached a slightly higher agreement (all PA > 60%, all κ > 0.40) than the comparison of FR vs S2 (all PA > 55%, all κ > 0.28). Table 4 Comparison of Lung-RADS categorization on a patient level. Comparison Lung-RADS selection Cohen’s κ PA [%] S1 vs S2 Lung-RADS ≥ 3 0.45 54 Lung-RADS < 3 0.41 64 All 0.45 62 FR vs S2 Lung-RADS ≥ 3 0.41 60 Lung-RADS < 3 0.28 55 All 0.33 55 FR vs S1 Lung-RADS ≥ 3 0.40 60 Lung-RADS < 3 0.53 76 All 0.55 75 FR, final reading dataset; Lung-RADS, lung CT screening reporting & data system; PA, percent agreement; S1, software tool 1 dataset; S2, software tool 2 dataset. For both software tools, the volume difference (n = 17 (65.4%) and n = 18 (69.2%) for software tool S1 and S2, respectively) was the most common reason for false nodule categorization with Lung-RADS score ≥ 3. Less frequent causes of classification differences were: false positive nodules (n = 5 (19.2%) and n = 6 (23.1%) for software tool S1 and S2, respectively), false nodule type classification (n = 3 (11.5%) and n = 0 (0%) software tool S1 and S2, respectively) and false negative nodules (n = 1 (3.8%) and n = 2 (7.7%) for software tool S1 and S2, respectively). Incorrect categorization of nodules with LUNG-RADS score < 3 is presented in the Supplement. Discussion The main findings of the present study include percentage disagreement in the assignment of Lung-RADS classification in 38% of participants as well as significant differences in volumetric measurement and sensitivity of pulmonary nodule detection when comparing S1 and S2. There is currently a large body of data comparing computer-aided detection (CAD) in lung cancer screening with radiologist performance ( 23 – 25 ), demonstrating that AI improves the lung nodule detection performance of radiologists. In addition in this study, we found that sensitivity as well as PPV also differed when using the two examined software tools. Considering different nodule subgroups, both software tools showed a higher sensitivity concerning solid nodules than to non-solid nodules, whereas the sensitivity of clinically relevant part-solid ( 26 ) nodules was similar by both software tools (61% and 62%). The decreased sensitivity in the part-solid group is likely to be influenced by the limited number of part-solid nodules found in the FR dataset, only 13 part-solid nodules were present in the whole cohort. Therefore, further studies are needed in the area of non-solid as well as part-solid nodules to increase detection performance in both nodule subgroups. ( 27 ) False positive but also false negative rates represent a drawback of CAD algorithms, which may be crucial when choosing a software tool for an LCS program. Comparing both rates, S1 exhibited superior performance. This result is not surprising, since the FR dataset was influenced by the prospective read from the S1 software tool. Successful patient management requires correct nodule measurement. In this study, the mean axial diameter and the volume parameters were considered. For TP nodules, S2 measured both parameters larger than S1, except for the volumes of the TP nodules with volume ≥ 113mm 3 , which were not statistically different. Identical results were observed for the comparison with FR with statistically higher volumes and diameters of the S2 software tool and significantly lower volumes and diameters for the S1 tool even for larger nodules (volume ≥ 113mm 3 ). This finding is unexpected because the FR was influenced by the S1 software. This supports the fact, that due to manual adjustment of the nodule size by the radiologist significant changes to volume and diameter were made. This is often necessary due to adjacent anatomic structures of the nodules such as pulmonary vessels for example. This finding clearly shows that there is still a radiologist interaction needed for visual inspection and quality control with the current performance of AI tools in the LCS setting. In all comparisons, both software tools delivered significantly different mean diameter measurements. This finding is in agreement with previous publications, which propose semiautomated volumetry to avoid overestimation of true nodule diameter. ( 12 , 13 , 28 ) In the study by Zhao et al. ( 19 ), three software tools were compared with each other based on the median volume of the detected nodules on baseline scans in the LCS. The nodules were classified according to characteristics, such as location, attachment, shape and edge. Similarly to our results, the study showed significant differences in volumetry between software packages for nearly all nodule groups, except for non-smooth nodules. However, the clinical output regarding the Lung-RADS classification was not reported in detail. The volumetry differences in our study are also in agreement with the previously published results of Hoop et al ( 29 ). The authors evaluated six software packages for solid lung nodule volumetry and reported significant mean volume differences in 11 out of 15 possible pairs of software tools. Furthermore, Ashraf et al. ( 30 ) demonstrated in their study that the reliability of volume measurements significantly declines when different algorithms are used, which is also congruent with our results. The above-mentioned differences may lead to differences in Lung-RADS classification and thus potentially to different clinical treatments of the participants. In the study of van Riel et al. ( 22 ) interobserver disagreement in the Lung-RADS category, however without using artificial intelligence, was seen in one-third of the examined scan pairs and led to different patient management in 8%. Therefore, when using different AI software tools in a national screening program, high agreement of all accredited AI software tools is mandatory. In our study, the differences in total Lung-RADS categorization between the two software tools were clinically relevant. Both software tools assigned the same category to the same participant in just 62% of the nodules. Our finding is in agreement with a recent phantom study by Peters et al. ( 31 ), which found, that participant management may be influenced by an incorrectness of a CAD System. The authors showed the inconsistency in assigning a Lung-RADS score between two software tools in 14.9%, which is lower than our results, likely due to a different study design. Comparing the individual software tool regarding the clinically relevant score (Lung-RADS ≥ 3) in each case with the FR - the value agrees with both software tools in 60%. That means in 40% of the cases the radiologist changed the score. For Lung-RADS levels ≥ 3, S1 nodule size measurement values were significantly lower and S2 significantly higher compared to the final result of the radiologist. In both cases, the incorrect measurement of the volume by one of the software tools was the reason in about 65%. Looking at the Lung-RADS stages 1 and 2, in more than 90% of the cases the errors were due to detection (false positive and false negative nodules) issues and in the wrong subgroup classification (especially calcified, non-solid subgroups). Only in the remaining 10% of cases, volume measurement differences caused incorrect Lung-RADS classification. At the individual nodule level, the differences were less severe. When comparing both software tools, there was a difference of 33% for the nodules with volume > 113mm 3 and 23% for the nodules with volume > 34mm 3 . Comparing both software tools with FR, especially nodules with volume > 11mm 3 were incorrectly assigned to Lung-RADS scores 3 and 4a. If Lung-RADS score 3 entails a 6-month LDCT for control, 4a entails a control in 3 months or, if malignancy is highly probable, a PET CT. For category 4b or 4x, a tissue biopsy is performed if the probability of malignancy is high, or a PET CT or chest CT with or without contrast is performed if the probability of malignancy is low. ( 10 ) With such a high inconsistency of 46% in the Lung-RADS classification from Lung-RADS Score 3 onwards (comparison S1 vs S2 on participant level), participants may be treated differently at different centers using different software tools. In a recent publication, Hwang et al. analyzed an actual nationwide LCS situation in South Korea. They concluded: there is a high inter-institutional variability in the interpretation of LCS results partially explained by different usage of the same CAD-system (e.g., disagreement in the Lung-RADS category occurred in 50.6% of the participants). ( 20 ) If the Lung-RADS classification is falsely high, this may lead to unnecessary psychological strain on the participant, unnecessary radiation exposure and even invasive measures as well as higher costs. If the classification is too low, this may lead to overlooking (potentially) malignant findings. The automated decision of the software tool regarding Lung-RADS categorization may also influence the opinion of the radiologist, resulting in clinically relevant consequences. ( 32 ) This study had several limitations. Firstly, software tool S1 was used prospectively in the HANSE study itself and likely influenced the final decision (FR dataset) of the radiologist, which was used as a gold standard for comparisons. This fact could partly explain the bias in the comparison of the two software tools with FR. Undeniably, all nodule candidates detected by CAD systems, as well as volumes and Lung-RADS classification, are currently evaluated by the radiologist. However, AI may influence the radiologist's opinion. Secondly, the distribution of nodule types regarding attenuation was in favor of solid nodules. A low number of part-solid nodules were included in the FR dataset. Therefore, the analysis in this subgroup might be examined in future studies. Also, only two available software tools were evaluated in this study. S1 was commercially available. Currently, there are more software packages available and they hold the potential to be subjects to future research. Conclusion This study found that, despite the accurate detection of solid nodules, there were significant differences between AI software tools in measuring diameter and volume and as a consequence in assigning the Lung-RADS score. Significant nodule volume differences between AI software tools led to different Lung-RADS scores in 38% of cases, which may result in altered participant management. Our findings clearly show that there is still a radiologist interaction needed for visual inspection and quality control with the current performance of AI tools in the LCS setting. Therefore, high performance and agreement of accredited AI software tools and adequate user training are mandatory for successful integration in a future national LCS program. Declarations Acknowledgement The authors would like to express their gratitude to the radiographers from the Department of Radiology for their support with the MR measurements and patient care. Data availability All data was acquired at Hannover Medical School, Germany, and is available in de-identified form from the corresponding author upon request. A formal data sharing agreement is needed. The MATLAB code is available from the corresponding author upon request. A formal code sharing agreement is needed. Additional information This work was supported by Siemens Healthcare GmbH. Jonathan Sperl is an employee of Siemens Healthcare GmbH, Erlangen, Germany. For the remaining authors none were declared. Author contributions R. K. conceptualization, data curation, writing original draft, formal analysis; F.K. programming, formal analysis, writing original draft; T.F.K. programming, review & editing; K.M. review & editing; J.B. review & editing; S.S. review & editing; J.S. data curation, programming, review & editing; S.D. review & editing; F.W. review & editing; J.V.C. conceptualization, supervision, review & editing. All Authors revised and approved the submitted version. References Schwartz Ann G.and Cote ML. Epidemiology of Lung Cancer. In: Ahmad Aamirand Gadgeel S, ed. Lung Cancer and Personalized Medicine: Current Knowledge and Therapies . Cham: Springer International Publishing; 2016:21–41. Siegel RL, Miller KD, Wagle NS, et al. Cancer statistics, 2023. CA Cancer J Clin. 2023;73(1):17–48. Smith RA, Andrews KS, Brooks D, et al. Cancer screening in the United States, 2018: A review of current American Cancer Society guidelines and current issues in cancer screening. CA Cancer J Clin. 2018;68(4):297–316. The National Lung Screening Trial: Overview and Study Design 1 National Lung Screening Trial Research Team. Radiology, 2011. Zhao YR, Xie X, De Koning HJ, et al. NELSON lung cancer screening study. Cancer Imaging. 2011;11(SPEC. ISS. A). Reduced Lung-Cancer Mortality with Low-Dose Computed Tomographic Screening. N Engl J Med. 2011;365(5):395–409. Becker N, Motsch E, Trotter A, et al. Lung cancer mortality reduction by LDCT screening—Results from the randomized German LUSI trial. Int J Cancer. 2020;146(6):1503–1513. Paci E, Puliti D, Lopes Pegna A, et al. Mortality, survival and incidence rates in the ITALUNG randomised lung cancer screening trial. Thorax. 2017;72(9):825–831. Vogel-Claussen J, Lasch F, Bollmann B-A, et al. Design and Rationale of the HANSE Study: A Holistic German Lung Cancer Screening Trial Using Low-Dose Computed Tomography TT - Design und Rationale der HANSE-Studie: Eine ganzheitliche deutsche Lungenkrebs-Früherkennungs-Studie unter Verwendung von Niedr. Rofo. 2022;(EFirst). American College of Radiology Committee on Lung-RADS®. LungRADS Assessment Categories version1.1. Available at: https://www.acr.org/-/media/ACR/Files/RADS/Lung-RADS/LungRADSAssessmentCategoriesv1-1.pdf . Accessed 21 July, 2022. Nam JG, Goo JM. Evaluation and Management of Indeterminate Pulmonary Nodules on Chest Computed Tomography in Asymptomatic Subjects: The Principles of Nodule Guidelines. Semin Respir Crit Care Med. 2022. de Margerie-Mellon C, Heidinger BH, Bankier AA. 2D or 3D measurements of pulmonary nodules: Preliminary answers and more open questions. J Thorac Dis. 2018;10(2):547–549. Heuvelmans MA, Walter JE, Vliegenthart R, et al. Disagreement of diameter and volume measurements for pulmonary nodule size estimation in CT lung cancer screening. Thorax. 2018;73(8). Snoeckx A, Franck C, Silva M, et al. The radiologist’s role in lung cancer screening. Transl Lung Cancer Res. 2021;10(5):2356–2367. Lancaster HL, Zheng S, Aleshina OO, et al. Outstanding negative prediction performance of solid pulmonary nodule volume AI for ultra-LDCT baseline lung cancer screening risk stratification. Lung Cancer. 2022;165:133–140. Schreuder A, Scholten ET, van Ginneken B, et al. Artificial intelligence for detection and characterization of pulmonary nodules in lung cancer CT screening: ready for practice? Transl Lung Cancer Res. 2021;10(5):2378–2388. Mathew CJ, David AM, Mathew CMJ. Artificial intelligence and its future potential in lung cancer screening. EXCLI J. 2020;19:1552–1562. AI for Radiology. grand-challenge.org/aiforradiology/. Accessed September 28, 2022. Zhao YR, Ooijen PMA van, Dorrius MD, et al. Comparison of three software systems for semi-automatic volumetry of pulmonary nodules on baseline and follow-up CT examinations. Acta radiol. 2014;55(6):691–698. Hwang EJ, Goo JM, Kim HY, et al. Variability in interpretation of low-dose chest CT using computerized assessment in a nationwide lung cancer screening program: comparison of prospective reading at individual institutions and retrospective central reading. Eur Radiol. 2021;31(5):2845–2855. Herth FJF, Reinmuth N, Wormanns D, et al. Positionspapier der Deutschen Röntgengesellschaft und der Deutschen Gesellschaft für Pneumologie und Beatmungsmedizin zu einem qualitätsgesicherten Früherkennungsprogramm des Lungenkarzinoms mittels Niedrigdosis-CT TT - Joint Statement of the German Radi. Pneumologie. 2019;73(10):573–577. van Riel SJ, Jacobs C, Scholten ET, et al. Observer variability for Lung-RADS categorisation of lung cancer screening CTs: impact on patient management. Eur Radiol. 2019;29(2):924–931. Chamberlin J, Kocher MR, Waltz J, et al. Automated detection of lung nodules and coronary artery calcium using artificial intelligence on low-dose CT scans for lung cancer screening: accuracy and prognostic value. BMC Med. 2021;19(1). Zhang Y, Jiang B, Zhang L, et al. Lung Nodule Detectability of Artificial Intelligence-assisted CT Image Reading in Lung Cancer Screening. Curr Med Imaging Former Curr Med Imaging Rev. 2021;18(3):327–334. Murchison JT, Ritchie G, Senyszak D, et al. Validation of a deep learning computer aided system for CT based lung nodule detection, classification, and growth rate estimation in a routine clinical population. PLoS One. 2022;17(5 May). Lee SM, Park CM, Song YS, et al. CT assessment-based direct surgical resection of part-solid nodules with solid component larger than 5 mm without preoperative biopsy: experience at a single tertiary hospital. Eur Radiol. 2017;27(12):5119–5126. Benzakoun J, Bommart S, Coste J, et al. Computer-aided diagnosis (CAD) of subsolid nodules: Evaluation of a commercial CAD system. Eur J Radiol. 2016;85(10):1728–1734. Heuvelmans MA, Oudkerk M. Pulmonary nodules measurements in CT lung cancer screening. J Thorac Dis. 2018;10:S2100-S2102. Hoop B, Gietema H, Ginneken B, et al. A comparison of six software packages for evaluation of solid lung nodules using semi-automated volumetry: What is the minimum increase in size to detect growth in repeated CT examinations. Eur Radiol. 2009;19(4):800–808. Ashraf H, de Hoop B, Shaker SB, et al. Lung nodule volumetry: segmentation algorithms within the same software package cannot be used interchangeably. Eur Radiol. 2010;20(8):1878–1885. Peters AA, Christe A, von Stackelberg O, et al. “WPeters, A. A., Christe, A., von Stackelberg, O., Pohl, M., Kauczor, H. U., Heußel, C. P., Wielpütz, M. O., & Ebner, L. (2023). “Will I change nodule management recommendations if I change my CAD system?”—impact of volumetric deviation between different. Eur Radiol . 2023. Kim JH, Han SG, Cho A, et al. Effect of deep learning-based assistive technology use on chest radiograph interpretation by emergency department physicians: a prospective interventional simulation-based study. BMC Med Inform Decis Mak. 2021;21(1):311. Additional Declarations Competing interest reported. Jonathan Sperl is an employee of Siemens Healthcare GmbH, Erlangen, Germany. Supplementary Files SupportingInformationfin.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3392224","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":236931315,"identity":"e44e5cbc-d688-4901-825e-5fbf92b89a78","order_by":0,"name":"Rimma Kondrashova","email":"","orcid":"","institution":"Hannover Medical School","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rimma","middleName":"","lastName":"Kondrashova","suffix":""},{"id":236931316,"identity":"e373152d-4ea9-4d41-baf7-5a9fb851944d","order_by":1,"name":"Filip Klimeš","email":"","orcid":"","institution":"Hannover Medical School","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Filip","middleName":"","lastName":"Klimeš","suffix":""},{"id":236931317,"identity":"b5f16f79-9f72-447a-b696-0b373d4e57f5","order_by":2,"name":"Till Frederik Kaireit","email":"","orcid":"","institution":"Hannover Medical School","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Till","middleName":"Frederik","lastName":"Kaireit","suffix":""},{"id":236931318,"identity":"95c99fc4-399f-4803-9a49-9e87761c09fd","order_by":3,"name":"Katharina May","email":"","orcid":"","institution":"University Hospital Schleswig-Holstein","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Katharina","middleName":"","lastName":"May","suffix":""},{"id":236931319,"identity":"a2944e3f-55a3-489f-bcfb-97b95e642058","order_by":4,"name":"Jörg Barkhausen","email":"","orcid":"","institution":"University Hospital Schleswig-Holstein","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jörg","middleName":"","lastName":"Barkhausen","suffix":""},{"id":236931320,"identity":"9c4a399d-e622-475a-8ca1-55e7bad0e12a","order_by":5,"name":"Susanne Stiebeler","email":"","orcid":"","institution":"Hospital Grosshansdorf","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Susanne","middleName":"","lastName":"Stiebeler","suffix":""},{"id":236931321,"identity":"adc265b3-c127-422c-b4c5-388173e11991","order_by":6,"name":"Jonathan Sperl","email":"","orcid":"","institution":"MR Application Predevelopment, Siemens Healthcare GmbH","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jonathan","middleName":"","lastName":"Sperl","suffix":""},{"id":236931322,"identity":"05563c89-4c85-4b21-8769-2c141e990fce","order_by":7,"name":"Sabine Dettmer","email":"","orcid":"","institution":"Hannover Medical School","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sabine","middleName":"","lastName":"Dettmer","suffix":""},{"id":236931323,"identity":"3af19e0d-c91d-4ce4-8fd2-455a495e411e","order_by":8,"name":"Frank Wacker","email":"","orcid":"","institution":"Hannover Medical School","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Frank","middleName":"","lastName":"Wacker","suffix":""},{"id":236931324,"identity":"ca673d87-5712-4e03-a9a1-0c8c19db8be1","order_by":9,"name":"Jens Vogel-Claussen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABSUlEQVRIie2RMUvDQBTHXzlIl7NZ76jEr3ClkEniV7kSMEvikkXEIVMyda8o+hV061iXZjkEt5OKpgQ6OaR00cHqXUFsEhDcBPODwLtcfu/d/wLQ0PBXyQB2v2qknlbWioB/70/qitrF2wpiv1I0BvlJYWlyn3FwsDny8mI5hrbZTRYnD2M4Ms+Htzk+foJOGpUUIULGwcVE+jY9E4DohbBngYCQPN65fSxCoKI0ho78QzJYTzBI30A7scqhilkQfwyiTZOYA5O8onivHCZ4T3r56l0pB9JbhEEMg6uNslbKc7atmMSbglZUK9Zt6SmE20gr10ohy0hPKcU3sUBEZ+mJF5sOY4JUqH5XKzeqYMWUYypKBzPaybwowLGsVN3YW7zvqqubr7RyKf1exk+51UkrPwaz7RVxoQquvWmXsoFT+6ChoaHh3/MJnS96rrPg/D8AAAAASUVORK5CYII=","orcid":"","institution":"Hannover Medical School","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jens","middleName":"","lastName":"Vogel-Claussen","suffix":""}],"badges":[],"createdAt":"2023-09-27 12:14:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3392224/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3392224/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":44211179,"identity":"9cc1e551-a54b-4861-b51f-cfd59b3c28e3","added_by":"auto","created_at":"2023-10-06 20:43:52","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":106769,"visible":true,"origin":"","legend":"\u003cp\u003eExamples of FP detection for both software tools. In (a) a consolidation due to pneumonia, in (b) a metallic foreign body and in (c) an osteophyte cases were detected by both software tools as a pulmonary nodule.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/b561d59504620ed3018c62f1.png"},{"id":44209446,"identity":"1574ac73-df35-4103-bd1e-4582cf644333","added_by":"auto","created_at":"2023-10-06 20:35:52","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":147904,"visible":true,"origin":"","legend":"\u003cp\u003eExemplary FN detection of pulmonary solid nodule (volume 1898.2 mm\u003csup\u003e3\u003c/sup\u003e, Lung-RADS 4B category) for both software tools for a 67-year-old female patient. No malignancy was found in the biopsy.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/b2a9b581ea6a6ac2eb46544f.png"},{"id":44209448,"identity":"502a0fdf-af4f-4371-8fc8-2cd028f507c3","added_by":"auto","created_at":"2023-10-06 20:35:52","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":143538,"visible":true,"origin":"","legend":"\u003cp\u003eExemplary FN detection of pulmonary solid mass (volume 5976.7 mm\u003csup\u003e3\u003c/sup\u003e,\u003csup\u003e \u003c/sup\u003eLung-RADS 4X category) for S1 software tool for a 69-year-old patient. After CT-assisted transthoracic puncture, the histology showed adenocarcinoma.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/503d715076b87026efbe12b2.png"},{"id":44209445,"identity":"16b857ae-49cf-47c4-97f3-460de17264ab","added_by":"auto","created_at":"2023-10-06 20:35:52","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":36879,"visible":true,"origin":"","legend":"\u003cp\u003eBland-Altman analysis of volume for all nodules (A), for nodules with volume ≥ 34 mm\u003csup\u003e3 \u003c/sup\u003e(B) and nodules with volume ≥ 113 mm\u003csup\u003e3 \u003c/sup\u003e(C) between software tools S1 and S2. The full red line indicates the mean difference (-13.3 (A), -6.0 (B) and 31.6 (C) mm\u003csup\u003e3\u003c/sup\u003e) and blue dotted lines indicate 95 % limits of agreement.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/4deaa7a726d2b72ff929004b.png"},{"id":44209444,"identity":"dbf6d111-dc11-47e6-8c9e-06ce5fd969cc","added_by":"auto","created_at":"2023-10-06 20:35:52","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":32961,"visible":true,"origin":"","legend":"\u003cp\u003eBland-Altman analysis of volume for all nodules (A), for nodules with volume ≥ 34 mm\u003csup\u003e3 \u003c/sup\u003e(B) and for nodules with volume ≥ 113 mm\u003csup\u003e3 \u003c/sup\u003e(C) between FR and software tool S2. The full red line indicates the mean difference (-15.7 (A), -7.6 (B) and 29.7 (C) mm\u003csup\u003e3\u003c/sup\u003e) and blue dotted lines indicate 95 % limits of agreement.\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/e6b0b82c7fe4c37af30b191d.png"},{"id":44209443,"identity":"3a39fba5-43bd-4a09-a670-6e94ea7300ff","added_by":"auto","created_at":"2023-10-06 20:35:52","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":31126,"visible":true,"origin":"","legend":"\u003cp\u003eBland-Altman analysis of volume for all nodules (A), for nodules with volume ≥ 34 mm\u003csup\u003e3 \u003c/sup\u003e(B) and for nodules with volume ≥ 113 mm\u003csup\u003e3 \u003c/sup\u003e(C) between FR and software tool S1. The full red line indicates the mean difference (9.5 (A), 22.6 (B) and 106.4 (C) mm\u003csup\u003e3\u003c/sup\u003e) and blue dotted lines indicate 95 % limits of agreement.\u003c/p\u003e","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/da43b2f52f0be46f22878b34.png"},{"id":60183534,"identity":"1eae5295-7a7b-400a-86b8-3015b8e72d4c","added_by":"auto","created_at":"2024-07-12 18:17:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1081124,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/75137453-6a7e-4f33-a84f-262b6a3556a8.pdf"},{"id":44209450,"identity":"c0fd5337-3e02-41ac-a2e5-6677157d29b2","added_by":"auto","created_at":"2023-10-06 20:35:53","extension":"docx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":419598,"visible":true,"origin":"","legend":"","description":"","filename":"SupportingInformationfin.docx","url":"https://assets-eu.researchsquare.com/files/rs-3392224/v1/adc1a9ff0f0640b5294b660e.docx"}],"financialInterests":"Competing interest reported. Jonathan Sperl is an employee of Siemens Healthcare GmbH, Erlangen, Germany.","formattedTitle":"Comparison of AI software tools for automated detection, quantification and categorization of pulmonary nodules in the HANSE LCS trial","fulltext":[{"header":"Introduction","content":"\u003cp\u003eLung cancer is the leading cause of cancer death worldwide. (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eIf lung cancer is detected at an early stage, it is more likely to be treated curatively. According to the American Thoracic Society, screening for lung cancer is recommended for people between 50\u0026ndash;80 years old and in fairly good health with at least a 20-pack-year smoking history who currently smoke or have quit in the past 15 years. (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) During the last years, multiple large multicenter lung cancer screening (LCS) studies worldwide such as The National Lung Screening Trial in the US, Dutch-Belgian Randomized Lung Cancer Screening Trial (NELSON), Italian Lung Cancer Screening Trial (ITALUNG) or German Lung Cancer Screening Intervention Trial (LUSI) concluded that lung cancer screening using low-dose computed tomography (LDCT) reduces lung cancer mortality. (\u003cspan additionalcitationids=\"CR5 CR6 CR7\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e) Recently a holistic northern German interdisciplinary lung cancer screening study (HANSE) (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e) has been started for optimal definition of the high-risk screening population and implementation of the screening workflow in the current national healthcare infrastructure.\u003c/p\u003e \u003cp\u003eLung nodule management in the HANSE study depends on the Lung Imaging Reporting and Data System (Lung-RADS 1.1) assessment. (\u003cspan additionalcitationids=\"CR10\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e) Five categories to discriminate between high-risk and low-risk nodules are part of the Lung-RADS assessment. After the detection of a nodule and correct recognition of nodule type, an accurate measurement of the size of detected nodules is a prerequisite for subsequent successful nodule management. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) Recent studies have shown that classification based on mean nodule diameter leads to a massive overestimation of true nodule size, therefore classification based on volume is recommended. (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eManual 3D volume segmentation of pulmonary nodules is time-consuming and error-prone; therefore, radiologists are nowadays assisted by artificial intelligence (AI)-based software tools when reading LDCT examinations. (\u003cspan additionalcitationids=\"CR15 CR16\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e) There are currently 17 products on the market that can detect, quantify, and categorize pulmonary nodules with distinct detection performance. (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e) Accurate detection and quantification are essential for the correct categorization of detected nodules. If a different Lung-RADS grade is assigned to the same subject by different software, the participant may be treated differently. In some cases, variability in the volumetric assessment of pulmonary nodules may result in a false-positive or false-negative diagnosis. (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e) Such inconsistency may occur, for example, when different software is used among different centers as parts of a national screening program. (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e) With the wide range of AI software available, it is important to know if there are significant differences in lung nodule detection, quantification and classification between the different AI tools.\u003c/p\u003e \u003cp\u003eThe purpose of this study was to investigate and compare the performance of two AI-based software tools regarding lung nodule detection, quantification, and categorization in the HANSE trial population prior to the implementation of a national lung cancer screening program in Germany. (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e)\u003c/p\u003e"},{"header":"Method","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy population\u003c/h2\u003e \u003cp\u003eThe study participants were part of the HANSE LCS trial (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e), a northern Germany multicenter LCS study. In this study, 946 LDCT randomly selected baseline participant examinations, performed (between July 2021 and July 2022) in a German Cancer Society-certified lung cancer center at Hannover Medical School, were retrospectively analyzed. The institutional review board of all participating institutions approved the HANSE study and all study participants provided written informed consent. The methods employed in this study conform to the principles outlined in the Declaration of Helsinki. The image acquisition details are listed in the Supplement.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eLDCT evaluation\u003c/h2\u003e \u003cp\u003eAfter the first reading of LDCTs by one experienced radiologist with specialized training in LCS, a computer-aided detection (CAD) technology of Software tool 1 (S1) automatically detected, segmented, and classified pulmonary nodules. The nodule candidates suggested by CAD technology, were flagged by the radiologist as either correct- or false-positive findings. Subsequently, the false positives were excluded from the report. Further, nodules were automatically classified into solid, non-solid, calcified and part-solid groups according to their attenuation. Inaccuracies in the measurement of volume and diameter as well as the nodule group affiliation were corrected using a threshold-based algorithm if necessary. Based on the measured volume and assigned nodule group, an individual nodule Lung-RADS score was determined. The overall Lung-RADS score for each participant was defined by the highest individual nodule Lung-RADS score. Cases classified as 3 or 4 according to Lung-RADS 1.1 scoring system (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e) were reviewed by another experienced board-certified thoracic radiologist. In case of nodule classification discrepancy between both readers, the case was presented to the radiological conference for a consensus evaluation. The final decision, consisting of the first read by the radiologist, AI-assistance and the second read of Lung-RADS 3\u0026ndash;4 cases, was considered as a final reading (FR) and defined the reference gold standard for comparison of AI software tools.\u003c/p\u003e \u003cp\u003eSimilarly to S1, software tool 2 (S2) was used for automated nodule detection, segmentation and classification. Both software tools were trained using deep learning and were dedicated for lung cancer screening. (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e) While S1 was used prospectively, S2 analyzed the results of the LDCT retrospectively using the same criteria.\u003c/p\u003e \u003cp\u003eTrue Positive (TP) nodules were identified as those detected in the FR dataset. Nodules detected by the S1 / S2 but not confirmed by the FR dataset were defined as false positives (FP). Nodules found in the FR dataset but not in the S1 / S2 dataset were designated as false negatives (FN).\u003c/p\u003e \u003cp\u003eThe matching of detected nodules from the S1 / S2 dataset to the FR dataset is described in the Supplement.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eStatistical analysis was performed with JMP Pro 16 software (SAS Institute, Cary, NC).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eNodule Detection\u003c/h2\u003e \u003cp\u003eNodule detection performance was assessed by sensitivity and positive predictive value (PPV) between FR and S1, FR and S2 of all detected nodules.\u003c/p\u003e \u003cp\u003eAn identical sub-analysis was performed for clinically relevant nodules between FR and S1 and FR and S2. Filters of 34mm\u003csup\u003e3\u003c/sup\u003e and 113mm\u003csup\u003e3\u003c/sup\u003e volume were applied. These thresholds correspond to the Lung-RADS 1.1 scoring system.\u003c/p\u003e \u003cp\u003eAdditionally, in the group of nodules\u0026thinsp;\u0026ge;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e volume, the same analysis was executed for nodule subgroups (solid, non-solid, calcified and part-solid).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eNodule quantification\u003c/h2\u003e \u003cp\u003eAfter excluding the FP nodules, both software tools (S1 and S2) and the FR dataset were compared regarding nodule diameter and volume. In the S1 vs S2 comparison, only nodules detected by both software tools were included. Nodule diameter and nodule volume parameters were tested for normality using the Shapiro-Wilk test. Since both parameters were not normally distributed, a non-parametric paired Wilcoxon signed-rank test was used to assess diameter and volume median differences in the following comparisons: FR vs S1, FR vs S2 and S1 vs S2. \u003cem\u003eP\u003c/em\u003e values\u0026thinsp;\u0026lt;\u0026thinsp;0.05 was considered as a statistically significant difference. Further, the mean differences were quantified by Bland-Altman analysis and the association between the measured diameters and volumes was explored by Pearson correlation analysis (\u003cem\u003er\u003c/em\u003e). The relative volumetric errors between the final reading and the two software tools were computed as a percentage difference between two measuring volumes divided by the mean of the two values.\u003c/p\u003e \u003cp\u003eThe same analysis was also conducted for TP nodules after the application of 34mm\u003csup\u003e3\u003c/sup\u003e and 113mm\u003csup\u003e3\u003c/sup\u003e volume thresholds.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eNodule categorization\u003c/h2\u003e \u003cp\u003eOn the participant and individual nodule level, the agreement of Lung-RADS classification of both software systems (S1, S2) with the ground truth (FR) and between each other (S1 vs S2) was determined by Cohen's kappa coefficient (κ) and percentual agreement (PA). PA was defined as a fraction of the number of Lung-RADS classifications in agreement and the total number of Lung-RADS classifications. Finally, the fraction was converted to PA by multiplication with 100. At the nodule level, when comparing the two software tools with each other, only the correctly positive nodules detected by both software tools were scored for fair comparison.\u003c/p\u003e \u003cp\u003eConsistently, the identical analysis was conducted at the nodule level for nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e and \u0026ge;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAdditionally, on the participant level, the whole analysis was repeated for participants with a total Lung-RADS score\u0026thinsp;\u0026ge;\u0026thinsp;3 and participants with a total Lung-RADS score\u0026thinsp;\u0026lt;\u0026thinsp;3.\u003c/p\u003e \u003cp\u003eFurther details regarding the statistics of Lung-RADS misclassification at the individual nodule level and false-positive rate on the participant level are presented in the Supplement.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eTo reflect the real world of LCS both nodules and masses (tumors with diameter larger than three cm) were examined. The volume of the smallest detected nodule and the largest mass was 0.3mm\u003csup\u003e3\u003c/sup\u003e and 18.9cm\u003csup\u003e3\u003c/sup\u003e respectively. In the study population of 946 participants, a total of 3345 lung nodules were found in the FR dataset and 765 out of 946 subjects had at least one nodule. The number of detected nodules ranged from 1 to 86 per subject. Of the 3345 nodules, 1174 were found clinically relevant (nodule volume\u0026thinsp;\u0026ge;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e). Besides 791 solid nodules, 258, 112, and 13 were non-solid, calcified, and part-solid, respectively. Of 1174 clinically relevant nodules, 282 nodules had volumes greater than 113mm\u003csup\u003e3\u003c/sup\u003e. Five masses were found in the final read of which three were detected by each software tool.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eNodule Detection\u003c/h2\u003e \u003cp\u003eExamples of FP and FN pulmonary nodule detection are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Sensitivities and PPVs of the S1 software tool were higher when compared to the S2 software tool for all nodules, for nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e and nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e (see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows detection results for clinically relevant nodules\u0026thinsp;\u0026ge;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e concerning their type. Sensitivities and PPVs were higher for S1 for all nodule types. The lowest sensitivities were observed for part-solid and non-solid using S1 and S2, respectively. For both software tools, the highest PPV was found for non-solid nodules.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDetection performance of both software tools in all nodules, in nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e and nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;113 mm\u003csup\u003e3\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSoftware tool\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNodule volume\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSensitivity [%]\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ePPV [%]\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eall\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2142\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e316\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1203\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e87\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1032\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e210\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e142\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e83\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;113 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e234\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e105\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eall\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1526\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e538\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1819\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e74\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e782\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e459\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e392\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e63\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;113 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e202\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e305\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e \u003cp\u003eFN, false negative; FP, false positive; PPV, positive predictive value; S1, software tool 1 dataset; S2, software tool 2 dataset; TP, true positive.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDetection performance of both software in subgroups of clinically relevant nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSoftware tool\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNodule type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSensitivity [%]\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ePPV [%]\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSolid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e720\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e175\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e80\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNon-Solid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e207\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e92\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCalcified\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePart-Solid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e62\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e80\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSolid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e570\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e321\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e221\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e63\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNon-Solid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e121\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e137\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCalcified\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePart-Solid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e72\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e \u003cp\u003eFN, false negative; FP, false positive; PPV, positive predictive value; S1, software tool 1 dataset; S2, software tool 2 dataset; TP, true positive.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eNodule Quantification\u003c/h2\u003e \u003cp\u003eAlthough the derived volumes of all TP nodules from S1 and S2 datasets were strongly correlated (all \u003cem\u003er\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.95), the volume derived by S2 was significantly larger than that obtained by S1 (both \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001, see Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), except for the nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e. In Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, Bland-Altman analysis between S1 and S2 software tools is depicted.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of detected TP nodules regarding median and mean volume for FR, S1 and S2 datasets.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMedian (IQR)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMean (SD)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMean Bias (relative difference in %)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cem\u003eP\u003c/em\u003e value\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ePearson correlation \u003cem\u003er\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e35.6 (52.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e97.7 (418.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e-13.3 (-)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e58.8 (52.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e111.1 (375.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e63.0 (73.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e164.4 (566.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e-6.0 (-)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e85.3 (73.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e170.5 (509.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;113 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e183.6 (204.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e470.6 (1079.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e31.6 (-)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e199.5 (159.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e439.0 (971.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e35.3 (51.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e106.9 (446.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e-15.7 (-46%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e56.0 (52.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e122.5 (580.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e68.9 (72.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e191.2 (612.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e-7.6 (-15%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.0001*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e86.1 (77.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e198.8 (803.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;113 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e198.6 (211.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e565.1 (1125.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e29.7 (11%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e209.2 (177.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e535.5 (1533.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e32.4 (47.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e88.5 (527.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e9.5 (4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31.5 (48.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e79.0 (343.4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;34 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e67.1 (63.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e164.8 (752.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e22.6 (10%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.0037*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.58\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e61.7 (69.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e142.3 (486.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;113 mm\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e186.2 (168.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e517.6 (1530.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e106.4 (21%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.0083*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.55\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e177.4 (172.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e411.2 (973.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e \u003cp\u003eFR, final reading dataset; IQR, interquartile range; Mean bias, mean difference derived from Bland-Altman analysis; S1, software tool 1 dataset; S2, software tool 2 dataset; SD, standard deviation; \u003cem\u003er\u003c/em\u003e, Pearson correlation coefficient\u003c/p\u003e \u003cp\u003e\u003csup\u003ea\u003c/sup\u003epaired Wilcoxon signed rank test, significantly different measurements (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05) are marked with *. It should be noted that each comparison exclusively incorporates TP nodules detected either by both software tools or by software tool and final read. Consequently, this leads to changed volume numbers for the same software tool across various comparisons.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn comparison between FR and S2 datasets (see Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e), all volumes obtained by the S2 software tool were significantly larger (both \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001, Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), except for nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eComparing the S1 dataset with the FR dataset (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e), FR volume measurements were found significantly higher (both P\u0026thinsp;\u0026lt;\u0026thinsp;0.0037, Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), except for the comparison of all nodules.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe results regarding the diameter comparison are presented in the Supplement (Supporting Information Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eNodule categorization\u003c/h2\u003e \u003cp\u003eSupporting Information Table S2 shows the individual nodule categorization comparison according to its size, measured by volume. There was a good agreement in comparison of S1 and S2 datasets (all PA\u0026thinsp;\u0026gt;\u0026thinsp;67%, all κ\u0026thinsp;\u0026gt;\u0026thinsp;0.61). Higher Lung-RADS consensus was achieved between FR vs S1 (all PA\u0026thinsp;\u0026gt;\u0026thinsp;69%, κ\u0026thinsp;\u0026gt;\u0026thinsp;0.58) when compared to FR vs S2 (all PA\u0026thinsp;\u0026gt;\u0026thinsp;55%, κ\u0026thinsp;\u0026gt;\u0026thinsp;0.44). For all comparisons, the Lung-RADS agreement was decreased with nodule size.\u003c/p\u003e \u003cp\u003eSimilarly, on a participant level, a moderate agreement of Lung-RADS categorization was observed in the comparison of S1 and S2 datasets (all PA\u0026thinsp;\u0026gt;\u0026thinsp;54%, all κ\u0026thinsp;\u0026gt;\u0026thinsp;0.41, see Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). As shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, the comparison of FR vs S1 reached a slightly higher agreement (all PA\u0026thinsp;\u0026gt;\u0026thinsp;60%, all κ\u0026thinsp;\u0026gt;\u0026thinsp;0.40) than the comparison of FR vs S2 (all PA\u0026thinsp;\u0026gt;\u0026thinsp;55%, all κ\u0026thinsp;\u0026gt;\u0026thinsp;0.28).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of Lung-RADS categorization on a patient level.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eComparison\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS selection\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCohen\u0026rsquo;s κ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePA [%]\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e\u003cb\u003eS1 vs S2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS\u0026thinsp;\u0026ge;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS\u0026thinsp;\u0026lt;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e62\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e\u003cb\u003eFR vs S2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS\u0026thinsp;\u0026ge;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS\u0026thinsp;\u0026lt;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e\u003cb\u003eFR vs S1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS\u0026thinsp;\u0026ge;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLung-RADS\u0026thinsp;\u0026lt;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e76\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAll\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eFR, final reading dataset; Lung-RADS, lung CT screening reporting \u0026amp; data system; PA, percent agreement; S1, software tool 1 dataset; S2, software tool 2 dataset.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eFor both software tools, the volume difference (n\u0026thinsp;=\u0026thinsp;17 (65.4%) and n\u0026thinsp;=\u0026thinsp;18 (69.2%) for software tool S1 and S2, respectively) was the most common reason for false nodule categorization with Lung-RADS score\u0026thinsp;\u0026ge;\u0026thinsp;3. Less frequent causes of classification differences were: false positive nodules (n\u0026thinsp;=\u0026thinsp;5 (19.2%) and n\u0026thinsp;=\u0026thinsp;6 (23.1%) for software tool S1 and S2, respectively), false nodule type classification (n\u0026thinsp;=\u0026thinsp;3 (11.5%) and n\u0026thinsp;=\u0026thinsp;0 (0%) software tool S1 and S2, respectively) and false negative nodules (n\u0026thinsp;=\u0026thinsp;1 (3.8%) and n\u0026thinsp;=\u0026thinsp;2 (7.7%) for software tool S1 and S2, respectively).\u003c/p\u003e \u003cp\u003eIncorrect categorization of nodules with LUNG-RADS score\u0026thinsp;\u0026lt;\u0026thinsp;3 is presented in the Supplement.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe main findings of the present study include percentage disagreement in the assignment of Lung-RADS classification in 38% of participants as well as significant differences in volumetric measurement and sensitivity of pulmonary nodule detection when comparing S1 and S2.\u003c/p\u003e \u003cp\u003eThere is currently a large body of data comparing computer-aided detection (CAD) in lung cancer screening with radiologist performance (\u003cspan additionalcitationids=\"CR24\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e), demonstrating that AI improves the lung nodule detection performance of radiologists. In addition in this study, we found that sensitivity as well as PPV also differed when using the two examined software tools. Considering different nodule subgroups, both software tools showed a higher sensitivity concerning solid nodules than to non-solid nodules, whereas the sensitivity of clinically relevant part-solid (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e) nodules was similar by both software tools (61% and 62%). The decreased sensitivity in the part-solid group is likely to be influenced by the limited number of part-solid nodules found in the FR dataset, only 13 part-solid nodules were present in the whole cohort. Therefore, further studies are needed in the area of non-solid as well as part-solid nodules to increase detection performance in both nodule subgroups. (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eFalse positive but also false negative rates represent a drawback of CAD algorithms, which may be crucial when choosing a software tool for an LCS program. Comparing both rates, S1 exhibited superior performance. This result is not surprising, since the FR dataset was influenced by the prospective read from the S1 software tool.\u003c/p\u003e \u003cp\u003eSuccessful patient management requires correct nodule measurement. In this study, the mean axial diameter and the volume parameters were considered. For TP nodules, S2 measured both parameters larger than S1, except for the volumes of the TP nodules with volume\u0026thinsp;\u0026ge;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e, which were not statistically different. Identical results were observed for the comparison with FR with statistically higher volumes and diameters of the S2 software tool and significantly lower volumes and diameters for the S1 tool even for larger nodules (volume\u0026thinsp;\u0026ge;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e). This finding is unexpected because the FR was influenced by the S1 software. This supports the fact, that due to manual adjustment of the nodule size by the radiologist significant changes to volume and diameter were made. This is often necessary due to adjacent anatomic structures of the nodules such as pulmonary vessels for example. This finding clearly shows that there is still a radiologist interaction needed for visual inspection and quality control with the current performance of AI tools in the LCS setting.\u003c/p\u003e \u003cp\u003eIn all comparisons, both software tools delivered significantly different mean diameter measurements. This finding is in agreement with previous publications, which propose semiautomated volumetry to avoid overestimation of true nodule diameter. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e) In the study by Zhao et al. (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e), three software tools were compared with each other based on the median volume of the detected nodules on baseline scans in the LCS. The nodules were classified according to characteristics, such as location, attachment, shape and edge. Similarly to our results, the study showed significant differences in volumetry between software packages for nearly all nodule groups, except for non-smooth nodules. However, the clinical output regarding the Lung-RADS classification was not reported in detail. The volumetry differences in our study are also in agreement with the previously published results of Hoop et al (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). The authors evaluated six software packages for solid lung nodule volumetry and reported significant mean volume differences in 11 out of 15 possible pairs of software tools. Furthermore, Ashraf et al. (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e) demonstrated in their study that the reliability of volume measurements significantly declines when different algorithms are used, which is also congruent with our results.\u003c/p\u003e \u003cp\u003eThe above-mentioned differences may lead to differences in Lung-RADS classification and thus potentially to different clinical treatments of the participants. In the study of van Riel et al. (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e) interobserver disagreement in the Lung-RADS category, however without using artificial intelligence, was seen in one-third of the examined scan pairs and led to different patient management in 8%. Therefore, when using different AI software tools in a national screening program, high agreement of all accredited AI software tools is mandatory.\u003c/p\u003e \u003cp\u003eIn our study, the differences in total Lung-RADS categorization between the two software tools were clinically relevant. Both software tools assigned the same category to the same participant in just 62% of the nodules. Our finding is in agreement with a recent phantom study by Peters et al. (\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e), which found, that participant management may be influenced by an incorrectness of a CAD System. The authors showed the inconsistency in assigning a Lung-RADS score between two software tools in 14.9%, which is lower than our results, likely due to a different study design.\u003c/p\u003e \u003cp\u003eComparing the individual software tool regarding the clinically relevant score (Lung-RADS\u0026thinsp;\u0026ge;\u0026thinsp;3) in each case with the FR - the value agrees with both software tools in 60%. That means in 40% of the cases the radiologist changed the score. For Lung-RADS levels\u0026thinsp;\u0026ge;\u0026thinsp;3, S1 nodule size measurement values were significantly lower and S2 significantly higher compared to the final result of the radiologist. In both cases, the incorrect measurement of the volume by one of the software tools was the reason in about 65%. Looking at the Lung-RADS stages 1 and 2, in more than 90% of the cases the errors were due to detection (false positive and false negative nodules) issues and in the wrong subgroup classification (especially calcified, non-solid subgroups). Only in the remaining 10% of cases, volume measurement differences caused incorrect Lung-RADS classification.\u003c/p\u003e \u003cp\u003eAt the individual nodule level, the differences were less severe. When comparing both software tools, there was a difference of 33% for the nodules with volume\u0026thinsp;\u0026gt;\u0026thinsp;113mm\u003csup\u003e3\u003c/sup\u003e and 23% for the nodules with volume\u0026thinsp;\u0026gt;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e. Comparing both software tools with FR, especially nodules with volume\u0026thinsp;\u0026gt;\u0026thinsp;11mm\u003csup\u003e3\u003c/sup\u003e were incorrectly assigned to Lung-RADS scores 3 and 4a.\u003c/p\u003e \u003cp\u003eIf Lung-RADS score 3 entails a 6-month LDCT for control, 4a entails a control in 3 months or, if malignancy is highly probable, a PET CT. For category 4b or 4x, a tissue biopsy is performed if the probability of malignancy is high, or a PET CT or chest CT with or without contrast is performed if the probability of malignancy is low. (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e) With such a high inconsistency of 46% in the Lung-RADS classification from Lung-RADS Score 3 onwards (comparison S1 vs S2 on participant level), participants may be treated differently at different centers using different software tools. In a recent publication, Hwang et al. analyzed an actual nationwide LCS situation in South Korea. They concluded: there is a high inter-institutional variability in the interpretation of LCS results partially explained by different usage of the same CAD-system (e.g., disagreement in the Lung-RADS category occurred in 50.6% of the participants). (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e) If the Lung-RADS classification is falsely high, this may lead to unnecessary psychological strain on the participant, unnecessary radiation exposure and even invasive measures as well as higher costs. If the classification is too low, this may lead to overlooking (potentially) malignant findings. The automated decision of the software tool regarding Lung-RADS categorization may also influence the opinion of the radiologist, resulting in clinically relevant consequences. (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eThis study had several limitations. Firstly, software tool S1 was used prospectively in the HANSE study itself and likely influenced the final decision (FR dataset) of the radiologist, which was used as a gold standard for comparisons. This fact could partly explain the bias in the comparison of the two software tools with FR. Undeniably, all nodule candidates detected by CAD systems, as well as volumes and Lung-RADS classification, are currently evaluated by the radiologist. However, AI may influence the radiologist's opinion. Secondly, the distribution of nodule types regarding attenuation was in favor of solid nodules. A low number of part-solid nodules were included in the FR dataset. Therefore, the analysis in this subgroup might be examined in future studies. Also, only two available software tools were evaluated in this study. S1 was commercially available. Currently, there are more software packages available and they hold the potential to be subjects to future research.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study found that, despite the accurate detection of solid nodules, there were significant differences between AI software tools in measuring diameter and volume and as a consequence in assigning the Lung-RADS score. Significant nodule volume differences between AI software tools led to different Lung-RADS scores in 38% of cases, which may result in altered participant management. Our findings clearly show that there is still a radiologist interaction needed for visual inspection and quality control with the current performance of AI tools in the LCS setting. Therefore, high performance and agreement of accredited AI software tools and adequate user training are mandatory for successful integration in a future national LCS program.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to express their gratitude to the radiographers from the Department of Radiology for their support with the MR measurements and patient care.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e All data was acquired at Hannover Medical School, Germany, and is available in de-identified form from the corresponding author upon request. A formal data sharing agreement is needed. The MATLAB code is available from the corresponding author upon request. A formal code sharing agreement is needed.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional information\u003c/strong\u003e This work was supported by Siemens Healthcare GmbH. Jonathan Sperl is an employee of Siemens Healthcare GmbH, Erlangen, Germany. For the remaining authors none were declared.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eR. K. conceptualization, data curation, writing original draft, formal analysis; F.K. programming, formal analysis, writing original draft; T.F.K. programming, review \u0026amp; editing; K.M. review \u0026amp; editing; J.B. review \u0026amp; editing; S.S. review \u0026amp; editing; J.S. data curation, programming, review \u0026amp; editing; S.D. review \u0026amp; editing; F.W. review \u0026amp; editing; J.V.C. conceptualization, supervision, review \u0026amp; editing. All Authors revised and approved the submitted version.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSchwartz Ann G.and Cote ML. Epidemiology of Lung Cancer. In: Ahmad Aamirand Gadgeel S, ed. \u003cem\u003eLung Cancer and Personalized Medicine: Current Knowledge and Therapies\u003c/em\u003e. Cham: Springer International Publishing; 2016:21\u0026ndash;41.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiegel RL, Miller KD, Wagle NS, et al. Cancer statistics, 2023. CA Cancer J Clin. 2023;73(1):17\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith RA, Andrews KS, Brooks D, et al. Cancer screening in the United States, 2018: A review of current American Cancer Society guidelines and current issues in cancer screening. CA Cancer J Clin. 2018;68(4):297\u0026ndash;316.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThe National Lung Screening Trial: Overview and Study Design 1 National Lung Screening Trial Research Team. Radiology, 2011.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao YR, Xie X, De Koning HJ, et al. NELSON lung cancer screening study. Cancer Imaging. 2011;11(SPEC. ISS. A).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReduced Lung-Cancer Mortality with Low-Dose Computed Tomographic Screening. N Engl J Med. 2011;365(5):395\u0026ndash;409.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBecker N, Motsch E, Trotter A, et al. Lung cancer mortality reduction by LDCT screening\u0026mdash;Results from the randomized German LUSI trial. Int J Cancer. 2020;146(6):1503\u0026ndash;1513.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaci E, Puliti D, Lopes Pegna A, et al. Mortality, survival and incidence rates in the ITALUNG randomised lung cancer screening trial. Thorax. 2017;72(9):825\u0026ndash;831.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVogel-Claussen J, Lasch F, Bollmann B-A, et al. Design and Rationale of the HANSE Study: A Holistic German Lung Cancer Screening Trial Using Low-Dose Computed Tomography TT - Design und Rationale der HANSE-Studie: Eine ganzheitliche deutsche Lungenkrebs-Fr\u0026uuml;herkennungs-Studie unter Verwendung von Niedr. Rofo. 2022;(EFirst).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmerican College of Radiology Committee on Lung-RADS\u0026reg;. LungRADS Assessment Categories version1.1. Available at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.acr.org/-/media/ACR/Files/RADS/Lung-RADS/LungRADSAssessmentCategoriesv1-1.pdf\u003c/span\u003e\u003cspan address=\"https://www.acr.org/-/media/ACR/Files/RADS/Lung-RADS/LungRADSAssessmentCategoriesv1-1.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 21 July, 2022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNam JG, Goo JM. Evaluation and Management of Indeterminate Pulmonary Nodules on Chest Computed Tomography in Asymptomatic Subjects: The Principles of Nodule Guidelines. Semin Respir Crit Care Med. 2022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ede Margerie-Mellon C, Heidinger BH, Bankier AA. 2D or 3D measurements of pulmonary nodules: Preliminary answers and more open questions. J Thorac Dis. 2018;10(2):547\u0026ndash;549.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeuvelmans MA, Walter JE, Vliegenthart R, et al. Disagreement of diameter and volume measurements for pulmonary nodule size estimation in CT lung cancer screening. Thorax. 2018;73(8).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSnoeckx A, Franck C, Silva M, et al. The radiologist\u0026rsquo;s role in lung cancer screening. Transl Lung Cancer Res. 2021;10(5):2356\u0026ndash;2367.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLancaster HL, Zheng S, Aleshina OO, et al. Outstanding negative prediction performance of solid pulmonary nodule volume\u0026nbsp;AI for ultra-LDCT baseline lung cancer screening risk stratification. Lung Cancer. 2022;165:133\u0026ndash;140.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchreuder A, Scholten ET, van Ginneken B, et al. Artificial intelligence for detection and characterization of pulmonary nodules in lung cancer CT screening: ready for practice? Transl Lung Cancer Res. 2021;10(5):2378\u0026ndash;2388.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMathew CJ, David AM, Mathew CMJ. Artificial intelligence and its future potential in lung cancer screening. EXCLI J. 2020;19:1552\u0026ndash;1562.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAI for Radiology. grand-challenge.org/aiforradiology/. Accessed September 28, 2022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao YR, Ooijen PMA van, Dorrius MD, et al. Comparison of three software systems for semi-automatic volumetry of pulmonary nodules on baseline and follow-up CT examinations. Acta radiol. 2014;55(6):691\u0026ndash;698.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHwang EJ, Goo JM, Kim HY, et al. Variability in interpretation of low-dose chest CT using computerized assessment in a nationwide lung cancer screening program: comparison of prospective reading at individual institutions and retrospective central reading. Eur Radiol. 2021;31(5):2845\u0026ndash;2855.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHerth FJF, Reinmuth N, Wormanns D, et al. Positionspapier der Deutschen R\u0026ouml;ntgengesellschaft und der Deutschen Gesellschaft f\u0026uuml;r Pneumologie und Beatmungsmedizin zu einem qualit\u0026auml;tsgesicherten Fr\u0026uuml;herkennungsprogramm des Lungenkarzinoms mittels Niedrigdosis-CT TT - Joint Statement of the German Radi. Pneumologie. 2019;73(10):573\u0026ndash;577.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Riel SJ, Jacobs C, Scholten ET, et al. Observer variability for Lung-RADS categorisation of lung cancer screening CTs: impact on patient management. Eur Radiol. 2019;29(2):924\u0026ndash;931.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChamberlin J, Kocher MR, Waltz J, et al. Automated detection of lung nodules and coronary artery calcium using artificial intelligence on low-dose CT scans for lung cancer screening: accuracy and prognostic value. BMC Med. 2021;19(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y, Jiang B, Zhang L, et al. Lung Nodule Detectability of Artificial Intelligence-assisted CT Image Reading in Lung Cancer Screening. Curr Med Imaging Former Curr Med Imaging Rev. 2021;18(3):327\u0026ndash;334.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMurchison JT, Ritchie G, Senyszak D, et al. Validation of a deep learning computer aided system for CT based lung nodule detection, classification, and growth rate estimation in a routine clinical population. PLoS One. 2022;17(5 May).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee SM, Park CM, Song YS, et al. CT assessment-based direct surgical resection of part-solid nodules with solid component larger than 5 mm without preoperative biopsy: experience at a single tertiary hospital. Eur Radiol. 2017;27(12):5119\u0026ndash;5126.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBenzakoun J, Bommart S, Coste J, et al. Computer-aided diagnosis (CAD) of subsolid nodules: Evaluation of a commercial CAD system. Eur J Radiol. 2016;85(10):1728\u0026ndash;1734.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeuvelmans MA, Oudkerk M. Pulmonary nodules measurements in CT lung cancer screening. J Thorac Dis. 2018;10:S2100-S2102.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoop B, Gietema H, Ginneken B, et al. A comparison of six software packages for evaluation of solid lung nodules using semi-automated volumetry: What is the minimum increase in size to detect growth in repeated CT examinations. Eur Radiol. 2009;19(4):800\u0026ndash;808.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAshraf H, de Hoop B, Shaker SB, et al. Lung nodule volumetry: segmentation algorithms within the same software package cannot be used interchangeably. Eur Radiol. 2010;20(8):1878\u0026ndash;1885.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeters AA, Christe A, von Stackelberg O, et al. \u0026ldquo;WPeters, A. A., Christe, A., von Stackelberg, O., Pohl, M., Kauczor, H. U., Heu\u0026szlig;el, C. P., Wielp\u0026uuml;tz, M. O., \u0026amp; Ebner, L. (2023). \u0026ldquo;Will I change nodule management recommendations if I change my CAD system?\u0026rdquo;\u0026mdash;impact of volumetric deviation between different. \u003cem\u003eEur Radiol\u003c/em\u003e. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim JH, Han SG, Cho A, et al. Effect of deep learning-based assistive technology use on chest radiograph interpretation by emergency department physicians: a prospective interventional simulation-based study. BMC Med Inform Decis Mak. 2021;21(1):311.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3392224/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3392224/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003ePurpose\u003c/h2\u003e \u003cp\u003eTo compare the performance of two AI-based software tools for detection, quantification and categorization of pulmonary nodules in a lung cancer screening (LCS) program in Northern Germany (HANSE-trial).\u003c/p\u003e\u003ch2\u003eMethod\u003c/h2\u003e \u003cp\u003e946 low-dose baseline CT-examinations were analyzed by two AI software tools regarding lung nodule detection, quantification and categorization and compared to the final radiologist read. The relationship between detected nodule volumes by both software tools was assessed by Pearson correlation (\u003cem\u003er\u003c/em\u003e) and tested for significance using Wilcoxon signed-rank test. The consistency of Lung-RADS classifications was evaluated by Cohen\u0026rsquo;s kappa (\u003cem\u003eκ\u003c/em\u003e) and percentual agreement (\u003cem\u003ePA\u003c/em\u003e).\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003e1032 (88%) and 782 (66%) of all (n\u0026thinsp;=\u0026thinsp;1174, solid, semi-solid and ground-glass) lung nodules (volume\u0026thinsp;\u0026ge;\u0026thinsp;34mm\u003csup\u003e3\u003c/sup\u003e) were detected by Software tool 1 (S1) and Software tool 2 (S2), respectively. Although, the derived volumes of true positive nodules were strongly correlated (\u003cem\u003er\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.95), the volume derived by S2 was significantly higher than by S1 (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001, mean difference: 6mm\u003csup\u003e3\u003c/sup\u003e). Moderate PA (62%) between S1 and S2 was found in the assignment of Lung-RADS classification (\u003cem\u003eκ\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.45). The PA of Lung-RADS classification to final read was 75% and 55% for S1 and S2.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eParticipant management depends on the assigned Lung Imaging Reporting and Data System (Lung-RADS) category, which is based on reliable detection and volumetry of pulmonary nodules. Significant nodule volume differences between AI software tools lead to different Lung-RADS scores in 38% of cases, which may result in altered participant management. Therefore, high performance and agreement of accredited AI software tools are necessary for a future national LCS program.\u003c/p\u003e","manuscriptTitle":"Comparison of AI software tools for automated detection, quantification and categorization of pulmonary nodules in the HANSE LCS trial","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-10-06 20:35:47","doi":"10.21203/rs.3.rs-3392224/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"773a8789-a362-412b-8e39-4883018573bf","owner":[],"postedDate":"October 6th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":25086072,"name":"Biological sciences/Cancer/Lung cancer"},{"id":25086073,"name":"Biological sciences/Cancer/Cancer screening"}],"tags":[],"updatedAt":"2024-07-12T18:08:59+00:00","versionOfRecord":[],"versionCreatedAt":"2023-10-06 20:35:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3392224","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3392224","identity":"rs-3392224","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0