Head and Neck Tumour Segmentation in PET Images: Performance Evaluation of 3D U-Net with Maximum Voting-Based Surrogate Ground Truth

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Accurate segmentation of head and neck tumours in PET images is critical for effective treatment planning, disease progression monitoring, and radiotherapy. However, achieving reliable ground truth data remains challenging due to inter- and intra-observer variability. The U-Net, a Deep Convolutional Neural Network (DCNN), has demonstrated strong potential for automated segmentation, yet the lack of definitive ground truth for training limits its full effectiveness. This study investigates the performance of a 3D U-Net deep learning framework trained using surrogate ground truth masks created through maximum voting (MV) from manual and semi-automatic segmentations. The analysis utilized the QIN-HEADNECK dataset, comprising a total of 217 lesions of 59 PET/CT scans. Each lesion was annotated by three radiologists, with two trials conducted for each segmentation method, resulting in a total of twelve trials. Three MV-based surrogate masks: M-MV (manual), SA-MV (semi-automatic), and A-MV (combined) were then generated. Segmentation performance was assessed using the Dice Similarity Coefficient (DSC). The results revealed that manual and semi-automatic segmentations achieved average DSC scores of 0.75 and 0.91, respectively, when compared to each MV result for individual trials. The combined MV (A-MV) produced DSC scores of 0.74 and 0.85 when compared to the MV results from the manual and semi-automatic segmentations, respectively. Among the 3D U-Net models, the framework trained with SA-MV achieved the highest average DSC score of 0.86, performing similarly to A-MV (0.86) and surpassing the model trained with M-MV (0.83). While the 3D U-Net outperformed manual segmentation (DSC of 0.83 vs. 0.75), its performance was still lower than that of semi-automatic segmentation (DSC of 0.86 vs. 0.91). These findings highlight the reliability of semi-automatic segmentation methods in producing consistent results, despite the added time required for their implementation. The study also indicates that while deep learning models are effective in standardizing and automating processes, their performance can be further enhanced by refining training datasets and pre-processing techniques. Additionally, the incorporation of advanced ground truth generation methods could significantly improve segmentation accuracy and increase the clinical applicability of these models.
Full text 101,490 characters · extracted from preprint-html · click to expand
Head and Neck Tumour Segmentation in PET Images: Performance Evaluation of 3D U-Net with Maximum Voting-Based Surrogate Ground Truth | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Head and Neck Tumour Segmentation in PET Images: Performance Evaluation of 3D U-Net with Maximum Voting-Based Surrogate Ground Truth Mahbubunnabi Tamal, Maha Alshammari, Maryam AlHasim, Saleh I. Alzahrani, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6198985/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Accurate segmentation of head and neck tumours in PET images is critical for effective treatment planning, disease progression monitoring, and radiotherapy. However, achieving reliable ground truth data remains challenging due to inter- and intra-observer variability. The U-Net, a Deep Convolutional Neural Network (DCNN), has demonstrated strong potential for automated segmentation, yet the lack of definitive ground truth for training limits its full effectiveness. This study investigates the performance of a 3D U-Net deep learning framework trained using surrogate ground truth masks created through maximum voting (MV) from manual and semi-automatic segmentations. The analysis utilized the QIN-HEADNECK dataset, comprising a total of 217 lesions of 59 PET/CT scans. Each lesion was annotated by three radiologists, with two trials conducted for each segmentation method, resulting in a total of twelve trials. Three MV-based surrogate masks: M-MV (manual), SA-MV (semi-automatic), and A-MV (combined) were then generated. Segmentation performance was assessed using the Dice Similarity Coefficient (DSC). The results revealed that manual and semi-automatic segmentations achieved average DSC scores of 0.75 and 0.91, respectively, when compared to each MV result for individual trials. The combined MV (A-MV) produced DSC scores of 0.74 and 0.85 when compared to the MV results from the manual and semi-automatic segmentations, respectively. Among the 3D U-Net models, the framework trained with SA-MV achieved the highest average DSC score of 0.86, performing similarly to A-MV (0.86) and surpassing the model trained with M-MV (0.83). While the 3D U-Net outperformed manual segmentation (DSC of 0.83 vs. 0.75), its performance was still lower than that of semi-automatic segmentation (DSC of 0.86 vs. 0.91). These findings highlight the reliability of semi-automatic segmentation methods in producing consistent results, despite the added time required for their implementation. The study also indicates that while deep learning models are effective in standardizing and automating processes, their performance can be further enhanced by refining training datasets and pre-processing techniques. Additionally, the incorporation of advanced ground truth generation methods could significantly improve segmentation accuracy and increase the clinical applicability of these models. Head and Neck Tumour Segmentation PET 3D U-Net Maximum Voting Surrogate Ground Truth Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 INTRODUCTION Head and neck cancer (HNC) is the seventh most common cancer type and one of the leading causes of cancer deaths worldwide, resulting in 1.1 million new cases annually (1). Knowing that HNC malignancies have site-specific and histology-specific behaviour, treatment paradigm requires a multidisciplinary approach, and remain challenging to diagnosis and treatment [18]. Early-stage diagnosis of HNC is of primary importance for tailoring treatment and lowering patient morbidity and mortality. A multimodality imaging approach, including computed tomography (CT) and Positron emission tomography (PET), is highly useful for assessing the primary tumour and any extent of the disease (2). Both imaging and histologic data are integrated to evaluate staging and optimization of the treatment plan. Positron Emission Tomography (PET) is a minimally invasive functional imagining modality that captures three-dimensional images of administered positron-emitting radiopharmaceuticals (3). Image segmentation methods used in PET images play an important role in distinguishing abnormal tissue from its healthy surrounding area. Therefore, it is crucial to achieve accurate segmentation of the images to achieve successful detection of diseases, planning of treatments, and follow-ups. Precise image segmentation methods, apart from the commonly used region of interest (ROI) analysis is of great importance for accurate tumour recognition and delineation in PET. Tumour delineations are used for several clinical tasks, including PET-based radiation therapy planning and well-grounded quantification of volumetric and radiomic features (4–6). Although image segmentation is possible in PET images, their success can be affected by several factors. Poor spatial resolution reduces the contrast between different objects in the images, and often making unclear the boundaries between adjacent objects. Big variabilities in the position, shape, and texture of the pathologies present complexities in choosing the segmentation methods suitable for those cases. High noise in PET images leads to additional challenges in the segmentation methods that use the standardized uptake value (SUV) to adjust their parameters (4). Tumours in PET scans can be delineated through manual segmentation. However, it is labour-intensive and can be time-consuming, especially for large datasets or multiple scans. On top of that it can produce inaccurate results due to inter- and intra-reader variability (7). To tackle these challenges, various computer-aided segmentation techniques have been developed. The most common used technique is based on thresholding (fixed, adaptive, and intensive) due to its simplicity which transforms a grey-level image into a binary image (8). This is completed by labelling all voxels higher than a set value to be foreground, while all remaining voxels as background (9). Such method is prone to inaccurate segmentation due to presence of high noise in PET images (8,10). Texture present in the images also impact the segmentation (11). Region-based segmentation techniques use the intensities of PET images; however, are much more focused on the local distribution (homogeneity) of the intensities found in the image (12). In gradient-based methods, a big change in intensity values on the edges of a PET image is an indication of the boundary of an object. To detect the location of the change in intensities, the gradient of the image is typically calculated between a voxel and its surrounding voxels (10,13). This does not always give the best results due to low resolution and high partial volume effect (PVE) as well presence of texture in the images (14,15). Other computer-aided methods include stochastic and learning methods, and joint segmentation methods. Similarly to manual segmentation, computer-aided segmentation techniques also demonstrate a few setbacks, including the necessity of manual input (16), high sensitivity to PVEs (17), challenges when assumptions are not satisfied (18), and the required calibration for the various types of scanners (19). Due to these challenges, more accurate, strong, and automated segmentation techniques are required. Deep-learning (DL) methods, specifically those that are based on convolutional neural networks have been promising (20). However, many limitations have yet to be addressed. PET images have high noise and low resolution which makes the job of specifying ground-truth for training DL-based techniques challenging (21). DL-based methods commonly use manual delineation as substitute of ground truth, but the low accuracy and high variability of PET images can restrict the use of manual segmentation as ground truth (22). Additionally, DL-based methods need a substantial amount of training data to achieve optimal results, which is not easily obtainable because PET is not a widely used imaging modality (23). This study investigates the training effectiveness of U-Net, a widely used convolutional neural network (CNN) architecture, using different PET segmentation methods (manual and semi-automatic). METHODOLOGY A database of images from the Quantitative Imaging Network in Head and Neck Cancer (QIN-HEADNECK) was used in this study. The images were originally collected for a study comparing manual segmentation methods with semi-automated segmentation methods (24). The database consists of 59 PET/CT scans taken from 59 subjects, each containing multiple head and neck lesions. Three experts, a faculty professor and two residents, performed semi-automated and manual segmentations on 217 lesions of the 59 subjects, each repeated twice, resulting in 12 segmented binary masks (ROI) for each lesion. Intra and inter observer variability was observed in the segmentation masks. Maximum Voting (MV): Each of the 12 segmented mask of each tumour was first extracted as a 3D volume using the boundary coordinates of the corresponding ROI. In absence of gold standard, three different surrogate masks of the gold standard of each lesion were generated using a majority voting (MV) method to obtain three reference masks – i) using 6 manual segmentations (M-MV), ii) using 6 semiautomatic segmentations (SA-MV) and iii) combining all 12 segmentations (A-MV). The overall process is illustrated in Fig. 1 . In MV method, pixels that are common to most of the segments are considered as part of the tumour. For manual (M-MV) and semiautomatic (SA-MV), only those pixels were considered as part of the tumour that were common to more than three segments. On the other hand, for all segmentation (A-MV) case, a pixel needs to be common in more than six segments to be considered as part of the tumour. ROI Pre-processing and U-Net Training: Each tumour was extracted as a 3D volume from the original grey level images using the boundary coordinates of the corresponding ROI. The size of each 3D grey level volume was then increased to include a greater number of background voxels in such a way that the ratio of background-to-tumour volume were greater than 2.5. The size of each corresponding mask was also increased in a similar way. If there was any adjacent tumour, it was removed from the extracted volume without removing other high uptake tissues. Histogram equalization was applied on the extracted tumours to enhance contrast. A grey level normalization process was then carried. All the normalized grey level volumes and binary masks were resized to \(\:64\times\:64\times\:64\:\) 3D cubic volumes to meet the input layer size requirements of the 3D U-net which is a widely used deep convolution neural network (DCNN) based segmentation algorithm (20). The output of the U-net is 3 training models based on 3 different majority voting ROIs as described above. Figure 2 describes the process of generating three training models using U-net. The 3D U-Net model is an advanced deep learning architecture designed for volumetric (3D) medical image segmentation. An extension of the original 2D U-Net, this model adapts the original architecture to handle three-dimensional data, making it highly effective for segmenting complex structures in 3D medical imaging modalities such as MRI and CT scans. Architecture: Encoder (Down-sampling Path) : The encoder utilizes 3D convolutional layers to extract hierarchical features from the input volume. Max pooling operations reduce spatial dimensions while increasing feature depth. Bottleneck : At the deepest level, the bottleneck captures abstract features with a series of 3D convolutions, providing a rich representation of the volumetric data. Decoder (Up-sampling Path) : The decoder reconstructs the spatial resolution of the image through transposed convolutions, combining up-sampled feature maps with high-resolution features from the encoder using skip connections. Output Layer : The final layer consists of 3D convolutional operations with activation functions (softmax for multi-class or sigmoid for binary segmentation) to produce the segmentation maps. The 3D U-net was trained on a machine with and NVIDIA GeForce GTX 1060 GPU. All processing, training, and analysis were performed using MATLAB 2022 platform. For the training, 150 epochs using the Adam Optimizer with a learning rate of 0.001, a batch size of 3, and a binary cross-entropy loss function were selected. Data Splitting: For each extracted lesion, the metabolic tumour volume (MTV), contrast-to-noise ratio (CNR), and signal-to-noise ratio (SNR) was determined. The lesions were randomly divided into three sets (65% training, 10% validation, and 25% test) in such a way that they would have similar gaussian distributions of MTV, CNR and SNR to avoid bias in training, validation, and testing sets. In the testing phase, only the \(\:64\times\:64\times\:64\:\) 3D cubic grey level volumes were fed to the U-net. Finally, the dice similarity coefficient (DSC) between the predicted segmentation generated using three different training sets were compared with the three surrogate ground truth segmentations. Images used in the test set were not part of the training model. DSC is calculated as: $$\:\text{D}\text{S}\text{C}=\frac{2\left|{\text{S}}_{\text{P}\text{R}\text{E}}\cap\:{\text{S}}_{\text{S}\text{G}\text{T}}\right|}{{\text{S}}_{\text{P}\text{R}\text{E}}+{\text{S}}_{\text{S}\text{G}\text{T}}}$$ where, \(\:{\text{S}}_{\text{S}\text{G}\text{T}}\) is surrogate ground truth segmentation (M-MV, SA-MV and A-MV) \(\:{\text{S}}_{\text{P}\text{R}\text{E}}\) is predicted segmentation using 3D U-net and three different training sets (M-MV, SA-MV and A-MV) RESULTS The representative contours of six manual and six semi-automatic segmentation methods on a single PET slice illustrated in Fig. 2 . The figure reveals that manual segmentation exhibits significant uncertainty near the tumour boundary, whereas the semi-automatic methods demonstrate more consistency. The results of the Maximum Voting (MV) segmentation methods—M-MV, SA-MV, and A-MV compared with each individual manual segmentation provide valuable insights into the strengths and limitations of different segmentation approaches. The results are shown in Table 1 . M-MV, which aggregates manual segmentations using maximum voting, achieves the highest average Dice Similarity Coefficient (DSC) of 0.75. This is because the M-MV was derived from the manual segmentations. However, the variability in DSC scores, ranging from 0.73 to 0.79, suggests some inconsistency between trials, likely caused by human subjectivity in boundary delineation. SA-MV, which combines semi-automatic segmentation methods using maximum voting, achieves an average DSC of 0.72 (ranging from 0.70 to 0.74) when compared to manual segmentations. While semi-automatic methods are generally more efficient than fully manual approaches, their reliability tends to decline when directly compared to manual segmentation. Table 1 Comparison of the average DSC for each of the three MV segmentations with each manual segmentation Maximum Voting (MV) Segmentation User 1- Trail 1 User 1 - Trail12 User 2 - Trail 1 User 2 - Trail 2 User 3 - Trail 1 User 3 Trail − 2 Average M-MV 0.75 0.75 0.77 0.79 0.74 0.73 0.75 SA-MV 0.71 0.71 0.74 0.73 0.70 0.71 0.72 A-MV 0.73 0.73 0.75 0.76 0.72 0.73 0.74 On the other hand, A-MV, representing the combined MV segmentations of both approaches, achieves an average DSC of 0.74, nearly matching M-MV's performance. With scores ranging from 0.72 to 0.76, A-MV shows less variability compared to SA-MV, reflecting its independence from user input. While A-MV has not yet surpassed the accuracy of manual methods, it shows significant promise for refinement, especially in scenarios where efficiency and consistency are critical. The results of the Maximum Voting (MV) segmentation methods—M-MV, SA-MV, and A-MV compared with each individual semi-automatic segmentation is shown in Table 2 . The results of M-MV, which combines manual segmentations through maximum voting, achieves a consistent DSC score of 0.74 to 0.75 across all trials, with an average DSC of 0.74. Table 2 Comparison of the average DSC for each of the three MV segmentations with each semi-automatic segmentation Maximum Voting (MV) Segmentation User 1- Trail 1 User 1 - Trail12 User 2 - Trail 1 User 2 - Trail 2 User 3 - Trail 1 User 3 Trail − 2 Average M-MV 0.74 0.74 0.75 0.74 0.75 0.75 0.74 SA-MV 0.91 0.87 0.92 0.91 0.91 0.92 0.91 A-MV 0.85 0.84 0.86 0.84 0.86 0.86 0.85 SA-MV, which relies on semi-automatic segmentation methods combined through maximum voting, performs the best overall, with DSC values ranging from 0.87 to 0.92 and an average of 0.91. This high performance indicates that semi-automatic methods can produce a segmentation that aligns closely with the MV, likely due to their ability to standardize the segmentation process while maintaining some level of user control. A-MV achieves DSC scores ranging from 0.84 to 0.86, with an average DSC of 0.85. It falls between that of M-MV and SA-MV, indicating solid performance but not quite matching the high accuracy of semi-automatic methods. Table 3 presents the average DSC values between each manual segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately. The table shows how each training method performed in comparison to the individual manual segmentation. U-net with M-MV training model achieved DSC values ranging from 0.70 to 0.74, with an average of 0.71. The relatively consistent average DSC score of 0.71 indicates that, training with manual segmentation provides stable results. Table 3 Comparison of the average DSC between each manual segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately Training Model User 1- Trail 1 User 1 - Trail12 User 2 - Trail 1 User 2 - Trail 2 User 3 - Trail 1 User 3 Trail − 2 Average M-MV 0.71 0.71 0.73 0.74 0.70 0.70 0.71 SA-MV 0.70 0.70 0.73 0.72 0.69 0.70 0.71 A-MV 0.71 0.70 0.73 0.73 0.70 0.71 0.71 Segmentation performed using SA-MV resulted in DSC values ranging from 0.69 to 0.73, with an average of 0.71. These results are comparable to those of M-MV, with SA-MV showing slightly lower performance in some instances but remaining consistent across trials. The minimal differences in performance between M-MV and SA-MV indicate that when the segmentation map generated by the SA-MV U-Net model is compared to manual methods, no significant improvement in accuracy is observed. The same conclusion applies to the A-MV U-Net model. The average DSC values between each semi-automatic segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately is shown in Table 4 . Table 4 Comparison of the average DSC between each semi-automatic segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately Training Model User 1- Trail 1 User 1 - Trail12 User 2 - Trail 1 User 2 - Trail 2 User 3 - Trail 1 User 3 Trail − 2 Average M-MV 0.72 0.71 0.72 0.72 0.73 0.73 0.72 SA-MV 0.82 0.80 0.83 0.82 0.82 0.83 0.82 A-MV 0.79 0.78 0.80 0.79 0.80 0.80 0.79 M-MV U-net model achieved DSC values ranging from 0.71 to 0.73, with an average of 0.72. These results indicate that M-MV U-net model provides relatively consistent performance across different users and trials. On the other hand, SA-MV U-net model, which relies on semi-automatic segmentation methods combined through maximum voting, performs significantly better with DSC values ranging from 0.80 to 0.83, and an average of 0.82. This higher performance suggests that model trained with the MV generated with semi-automatic methods, when aggregated, can achieve a more accurate segmentation, likely due to their ability to standardize the segmentation process while still involving user input. The improved DSC values across all trials indicate that SA-MV outperforms M-MV, providing a more accurate segmentation solution. A-MV, achieved DSC values between 0.78 and 0.80, with an average of 0.79. These results show that A-MV delivers solid performance, with DSC scores consistently higher than M-MV but slightly lower than SA-MV. The results suggest that combined training model can provide reliable and consistent segmentations compared to the M-MV, but they still fall short of the accuracy achieved by semi-automatic methods. When M-MV U-Net is used, it achieves the highest DSC with the M-MV method (0.83). This suggests that manual segmentations, when aggregated through maximum voting, provide the most accurate results when paired with the M-MV U-Net model. However, the performance decreases slightly when M-MV U-Net is compared with the other two methods. For SA-MV U-Net, it achieves a DSC of 0.73, and for A-MV U-Net, it reaches a DSC of 0.76. This indicates that M-MV U-Net is not as effective when paired with semi-automatic or combined methods. When the SA-MV U-Net model is used, the method performs very well, particularly in comparison with the semi-automatic methods itself. SA-MV U-Net achieves the highest DSC of 0.86 when paired with the SA-MV method, showing that semi-automatic segmentation provides the most accurate results with this model. The DSC is slightly lower when using SA-MV U-Net with M-MV (0.75) and A-MV (0.83), suggesting that the SA-MV U-Net model performs better when using semi-automatic segmentation compared to the other methods. For the combined A-MV U-Net model, it performs the best with A-MV, achieving the highest DSC of 0.86. This indicates that the A-MV U-Net model provides the most accurate segmentation results when combined with manual and semi-automated segmentation methods. When using A-MV U-Net with M-MV, the DSC is 0.76, and with SA-MV, the DSC is 0.81, showing a generally good performance, although not as high as with its own A-MV method. DISCUSSION The study examines the application of a 3D U-Net deep learning framework for head and neck tumour segmentation in PET images. By utilizing surrogate ground truths derived from manual, semi-automatic, and combined segmentation approaches, the study provides critical insights into optimizing segmentation performance in the challenging context of PET imaging. Three distinct surrogate ground truths—M-MV (manual), SA-MV (semi-automatic), and A-MV (all)—were developed to evaluate their impact on segmentation accuracy. The manual-based M-MV demonstrated moderate Dice Similarity Coefficient (DSC) values (average 0.72), reflecting inherent limitations in manual delineation such as observer variability and subjective boundary definitions (25)​. In contrast, SA-MV exhibited the highest DSC values (average 0.82), indicating the potential of semi-automatic methods to achieve more consistent and reproducible segmentation (26). The inclusion of A-MV, combining manual and semi-automatic inputs, further demonstrated its utility by balancing robustness and accuracy with an average DSC of 0.79, surpassing M-MV while approaching the performance of SA-MV. The superior performance of SA-MV highlights the efficacy of semi-automatic methods in leveraging computational precision to minimize user-dependent variability (25,26). Semi-automatic methods standardize segmentation tasks, thereby enhancing reproducibility without entirely removing human oversight, which remains critical in clinical applications​. Furthermore, semi-automatic methods are better equipped to handle the low resolution and high noise levels of PET images (27), particularly when combined with deep learning architectures like 3D U-Net (28). The use of the 3D U-Net framework significantly enhanced segmentation performance across all surrogate ground truth models. The ability of the model to process volumetric data using an encoder-decoder structure with skip connections enabled detailed feature extraction and accurate tumour delineation (29). The SA-MV-trained 3D U-Net achieved the best performance when tested against the SA-MV ground truth, with DSC values consistently exceeding those of M-MV and A-MV models. This underscores the compatibility of semi-automatic segmentation data with deep learning-based refinement​. The A-MV-trained U-Net delivered comparable results to SA-MV while demonstrating greater versatility by integrating multiple segmentation strategies. However, the slightly lower DSC values compared to SA-MV suggest that introducing heterogeneity through combined inputs may add complexity that could hinder performance in specific scenarios​. Despite its promising results, the study highlights several challenges in PET image segmentation. First, the inherent properties of PET images, such as noise and low resolution, complicate the delineation of tumour boundaries (8). 3D U-Net's architecture addresses some of these challenges, further advancements in pre-processing techniques, such as enhanced denoising and contrast enhancement, could improve outcomes (13,30)​. Second, the reliance on a relatively small dataset (59 subjects) limits the generalizability of the findings. Future studies should explore larger and more diverse datasets to improve model robustness​. Finally, the absence of a definitive gold standard for segmentation necessitates reliance on surrogate ground truths. Developing new techniques for generating more accurate ground truths could further refine training and evaluation processes​. The findings underscore the potential of combining semi-automatic methods with deep learning for clinical applications. By reducing dependency on manual delineation and standardizing segmentation outputs, this approach can streamline workflows in radiation therapy planning and tumour monitoring. Moreover, the adoption of SA-MV in clinical settings could address inter-observer variability, improving the reliability of PET-based diagnostics and treatment​. CONCLUSION This study assesses the performance of a 3D U-Net deep learning framework for segmenting head and neck tumours in PET images, using surrogate ground truths generated through maximum voting (MV) from manual and semi-automatic segmentations. By comparing three distinct surrogate masks—M-MV (manual), SA-MV (semi-automatic), and A-MV (combined)—the study provides critical insights into the strengths and limitations of both manual and semi-automatic approaches in the context of deep learning-based tumour segmentation. The results demonstrate that the 3D U-Net model trained using the SA-MV surrogate mask consistently outperforms models trained with M-MV and A-MV, achieving the highest Dice Similarity Coefficient (DSC) of 0.82. This highlights the ability of semi-automatic methods to generate consistent and accurate ground truths, which are particularly well-suited for deep learning applications. In contrast, models trained on M-MV exhibited lower performance, likely due to the inherent inter- and intra-observer variability in manual delineation. The combined approach, A-MV, balanced accuracy and robustness, achieving a DSC of 0.79, making it a viable alternative where both manual and semi-automatic inputs are available. While the 3D U-Net model demonstrates its potential for automating tumour segmentation, its performance remains marginally lower than the standalone semi-automatic methods, which achieved a DSC of 0.91. These findings highlight the reliability of semi-automatic segmentation methods in producing consistent results, despite the added time required for their implementation. The study also indicates that while deep learning models are effective in standardizing and automating processes, their performance can be further enhanced by refining training datasets and pre-processing techniques. Additionally, the incorporation of advanced ground truth generation methods could significantly improve segmentation accuracy and increase the clinical applicability of these models. Statements and Declarations Funding: The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Competing Interests: The authors have no relevant financial or non-financial interests to disclose. Author Contributions: Mahbubunnabi Tamal and Maha Alshammari contributed to the study conception and design. Material preparation and data analysis were performed by Mahbubunnabi Tamal and Maha Alshammari. The manuscript was written by Mahbubunnabi Tamal, Maha Alshammari, and Maryam Alhashim. Validation, interpretation, and critical review of the results were conducted by Saleh I. Alzahraniand Murad Althobaiti. All authors read and approved the final manuscript. Ethics approval: Ethics approval was not required for this study, as it utilizes publicly available data. References Mody MD, Rocco JW, Yom SS, Haddad RI, Saba NF. Head and neck cancer. The Lancet. 2021 Dec;398(10318):2289–99. Pham N, Ju C, Kong T, Mukherji SK. Artificial Intelligence in Head and Neck Imaging. Seminars in Ultrasound, CT and MRI. 2022 Apr;43(2):170–5. Lameka K, Farwell MD, Ichise M. Positron Emission Tomography. In 2016. p. 209–27. Foster B, Bagci U, Mansoor A, Xu Z, Mollura DJ. A review on segmentation of positron emission tomography images. Comput Biol Med. 2014 Jul;50:76–96. Jha AK, Mena E, Caffo B, Ashrafinia S, Rahmim A, Frey E, et al. Practical no-gold-standard evaluation framework for quantitative imaging methods: application to lesion segmentation in positron emission tomography. Journal of Medical Imaging. 2017 Mar 3;4(1):011011. Mena E, Sheikhbahaei S, Taghipour M, Jha AK, Vicente E, Xiao J, et al. 18F-FDG PET/CT Metabolic Tumor Volume and Intratumoral Heterogeneity in Pancreatic Adenocarcinomas. Clin Nucl Med. 2017 Jan;42(1):e16–21. Angulakshmi M, Lakshmi Priya GG. Automated brain tumour segmentation techniques— A review. Int J Imaging Syst Technol. 2017 Mar 21;27(1):66–77. Tamal M. Intensity threshold based solid tumour segmentation method for Positron Emission Tomography (PET) images: A review. Heliyon. 2020;6(10). Saha PK, Udupa JK. Optimum image thresholding via class uncertainty and region homogeneity. IEEE Trans Pattern Anal Mach Intell. 2001 Jul;23(7):689–706. Tamal M. A phantom study to assess the reproducibility, robustness and accuracy of PET image segmentation methods against statistical fluctuations. PLoS One. 2019;14(7). Anan N, Zainon R, Tamal M. A review on advances in 18F-FDG PET/CT radiomics standardisation and application in lung disease management. Insights Imaging. 2022 Dec 5;13(1):22. Foster B, Bagci U, Mansoor A, Xu Z, Mollura DJ. A review on segmentation of positron emission tomography images. Comput Biol Med. 2014 Jul;50:76–96. Tamal M. A hybrid region growing tumour segmentation method for low contrast and high noise Nuclear Medicine (NM) images by combining a novel non-linear diffusion filter and global gradient measure (HNDF-GGM-RG). Heliyon. 2019;5(12). Tamal M. A Phantom Study to Investigate Robustness and Reproducibility of Grey Level Co-Occurrence Matrix (GLCM)-Based Radiomics Features for PET. Applied Sciences. 2021 Jan 7;11(2):535. Tamal M. Grey Level Co-occurrence Matrix (GLCM) as a Radiomics Feature for Artificial Intelligence (AI) Assisted Positron Emission Tomography (PET) Images Analysis. IOP Conf Ser Mater Sci Eng. 2019 Oct 1;646(1):012047. Aristophanous M, Penney BC, Martel MK, Pelizzari CA. A Gaussian mixture model for definition of lung tumor volumes in positron emission tomography. Med Phys. 2007 Nov 17;34(11):4223–35. Brambilla M, Matheoud R, Secco C, Loi G, Krengli M, Inglese E. Threshold segmentation for PET target volume delineation in radiation treatment planning: The role of target‐to‐background ratio and target size. Med Phys. 2008 Apr 7;35(4):1207–13. Belhassen S, Zaidi H. A novel fuzzy C‐means algorithm for unsupervised heterogeneous tumor quantification in PET. Med Phys. 2010 Mar 25;37(3):1309–24. Zaidi H, Abdoli M, Fuentes CL, El Naqa IM. Comparative methods for PET image segmentation in pharyngolaryngeal squamous cell carcinoma. Eur J Nucl Med Mol Imaging. 2012 May 31;39(5):881–91. Ronneberger O, Fischer P, Brox T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In 2015. p. 234–41. Leung KH, Marashdeh W, Wray R, Ashrafinia S, Pomper MG, Rahmim A, et al. A physics-guided modular deep-learning based automated framework for tumor segmentation in PET. Phys Med Biol. 2020 Dec 21;65(24):245032. Giraud P, Elles S, Helfre S, De Rycke Y, Servois V, Carette MF, et al. Conformal radiotherapy for lung cancer: different delineation of the gross tumor volume (GTV) by radiologists and radiation oncologists. Radiotherapy and Oncology. 2002 Jan;62(1):27–36. Shen D, Wu G, Suk HI. Deep Learning in Medical Image Analysis. Annu Rev Biomed Eng. 2017 Jun 21;19(1):221–48. Beichel RR, Van Tol M, Ulrich EJ, Bauer C, Chang T, Plichta KA, et al. Semiautomated segmentation of head and neck cancers in 18F‐FDG PET scans: A just‐enough‐interaction approach. Med Phys. 2016 Jun 18;43(6Part1):2948–64. Pfaehler E, Burggraaff C, Kramer G, Zijlstra J, Hoekstra OS, Jalving M, et al. PET segmentation of bulky tumors: Strategies and workflows to improve inter-observer variability. PLoS One. 2020 Mar 30;15(3):e0230901. Philip MM, Watts J, Moeini SNM, Musheb M, McKiddie F, Welch A, et al. Comparison of semi-automatic and manual segmentation methods for tumor delineation on head and neck squamous cell carcinoma (HNSCC) positron emission tomography (PET) images. Phys Med Biol. 2024 May 7;69(9):095005. Foster B, Bagci U, Mansoor A, Xu Z, Mollura DJ. A review on segmentation of positron emission tomography images. Comput Biol Med. 2014 Jul;50:76–96. Jiang H, Diao Z, Shi T, Zhou Y, Wang F, Hu W, et al. A review of deep learning-based multiple-lesion recognition from medical images: classification, detection and segmentation. Comput Biol Med. 2023 May;157:106726. Azad R, Aghdam EK, Rauland A, Jia Y, Avval AH, Bozorgpour A, et al. Medical Image Segmentation Review: The Success of U-Net. IEEE Trans Pattern Anal Mach Intell. 2024 Dec;46(12):10076–95. Tamal M. Nonlinear Diffusion Filter for Low Count Positron Emission Tomography Utilizing Orientation Information of Neighbouring Gradient Vectors. In: 2017 9th IEEE-GCC Conference and Exhibition (GCCCE). IEEE; 2017. p. 1–3. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6198985","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":430177421,"identity":"c914bcdf-4cbc-48a7-b053-fb12957e690b","order_by":0,"name":"Mahbubunnabi Tamal","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAArElEQVRIiWNgGAWjYNCCAhsGBgmiVbOBCIM00rUcJkGL7vzmpxt+GJxP7J/dfPABQ41NNEEtZsfYzG72GNxOnHHnWLIBw7G03AbCWhjMbvAAtTTcyDGTYGw4TIwW9m83/xicS5xPghYes9s8BgcSN5CgJafstoxBsvHGG2nJBglE+eXw8W0331TYyc67kXzwwYcaG8JaYMARrDKBWOUgYE+K4lEwCkbBKBhhAAANSENmQRATUQAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-3553-3531","institution":"Imam Abdulrahman Bin Faisal University","correspondingAuthor":true,"prefix":"","firstName":"Mahbubunnabi","middleName":"","lastName":"Tamal","suffix":""},{"id":430177422,"identity":"8bc2627e-4a1d-4a6c-a6f0-10e2ba9f1bcd","order_by":1,"name":"Maha Alshammari","email":"","orcid":"","institution":"Imam Abdulrahman Bin Faisal University","correspondingAuthor":false,"prefix":"","firstName":"Maha","middleName":"","lastName":"Alshammari","suffix":""},{"id":430177423,"identity":"428eeb0f-eac3-401c-ab91-cdba3d5a73c0","order_by":2,"name":"Maryam AlHasim","email":"","orcid":"","institution":"King Fahad Specialist Hospital Dammam","correspondingAuthor":false,"prefix":"","firstName":"Maryam","middleName":"","lastName":"AlHasim","suffix":""},{"id":430177424,"identity":"ef2df652-1b08-4924-bb87-ff80edc07094","order_by":3,"name":"Saleh I. Alzahrani","email":"","orcid":"","institution":"Imam Abdulrahman Bin Faisal University","correspondingAuthor":false,"prefix":"","firstName":"Saleh","middleName":"I.","lastName":"Alzahrani","suffix":""},{"id":430177425,"identity":"4d2d26bb-7d21-4940-bf07-d10b650e97fb","order_by":4,"name":"Murad Althobaiti","email":"","orcid":"","institution":"Imam Abdulrahman Bin Faisal University","correspondingAuthor":false,"prefix":"","firstName":"Murad","middleName":"","lastName":"Althobaiti","suffix":""}],"badges":[],"createdAt":"2025-03-11 00:21:58","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6198985/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6198985/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79260753,"identity":"757856f8-4403-4a93-91d6-4bdd24398c3d","added_by":"auto","created_at":"2025-03-26 09:25:59","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":136609,"visible":true,"origin":"","legend":"\u003cp\u003eIllustration of three maximum voting (manual, semiautomatic and all) reference masks generation as surrogate of the gold standard. There were six segmentation masks for each manual and semiautomatic method.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6198985/v1/703962495d6dfc5f51f24548.png"},{"id":79260752,"identity":"916f6f37-142e-4932-a06b-2084b3df93b5","added_by":"auto","created_at":"2025-03-26 09:25:59","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":141166,"visible":true,"origin":"","legend":"\u003cp\u003ePre-processing and resizing steps to generate three training models using majority voting (MV) segmentation masks using manual, semiautomatic and all segmentation masks.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6198985/v1/ff1bdc83da165fb729af84ad.png"},{"id":79259203,"identity":"98a274c3-f55b-48fa-ace6-deb9350140bb","added_by":"auto","created_at":"2025-03-26 09:17:59","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":131260,"visible":true,"origin":"","legend":"\u003cp\u003eFlow diagram of the segmentation procedure on test images. Segmented images with 3 trained models were compared with 3 surrogated ground truths derived using maximum voting (MV) from manual segmentations, semi-automatic segmentations, and combined segmentations.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-6198985/v1/af278d6cd7769ff778a219be.png"},{"id":79259204,"identity":"bec062bb-21e9-404d-b6e8-a581fe781c61","added_by":"auto","created_at":"2025-03-26 09:17:59","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":95018,"visible":true,"origin":"","legend":"\u003cp\u003eContours of six manual and six semi-automatic segmentations methods on one PET slice\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-6198985/v1/224eada69f3adcd617065bf5.png"},{"id":79259211,"identity":"3ff6a0f1-9dcc-45bc-8c75-20381c60d572","added_by":"auto","created_at":"2025-03-26 09:17:59","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":16588,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of segmentation output of each of the three models (M-MV U-Net, SA-MV U-Net and A-MV U-Net) with the three surrogated ground truth MV (M-MV, SA-MV and A-MV).\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-6198985/v1/593d7d7e81deac7f63769b7b.png"},{"id":107706918,"identity":"5bd1ad6b-8d26-4bc5-b457-667f609df4e6","added_by":"auto","created_at":"2026-04-24 09:19:04","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":745757,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6198985/v1/dbb05fcf-96df-4eca-ad73-0ab9545f19ab.pdf"}],"financialInterests":"","formattedTitle":"Head and Neck Tumour Segmentation in PET Images: Performance Evaluation of 3D U-Net with Maximum Voting-Based Surrogate Ground Truth","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eHead and neck cancer (HNC) is the seventh most common cancer type and one of the leading causes of cancer deaths worldwide, resulting in 1.1\u0026nbsp;million new cases annually (1). Knowing that HNC malignancies have site-specific and histology-specific behaviour, treatment paradigm requires a multidisciplinary approach, and remain challenging to diagnosis and treatment [18]. Early-stage diagnosis of HNC is of primary importance for tailoring treatment and lowering patient morbidity and mortality. A multimodality imaging approach, including computed tomography (CT) and Positron emission tomography (PET), is highly useful for assessing the primary tumour and any extent of the disease (2). Both imaging and histologic data are integrated to evaluate staging and optimization of the treatment plan.\u003c/p\u003e \u003cp\u003ePositron Emission Tomography (PET) is a minimally invasive functional imagining modality that captures three-dimensional images of administered positron-emitting radiopharmaceuticals (3). Image segmentation methods used in PET images play an important role in distinguishing abnormal tissue from its healthy surrounding area. Therefore, it is crucial to achieve accurate segmentation of the images to achieve successful detection of diseases, planning of treatments, and follow-ups.\u003c/p\u003e \u003cp\u003ePrecise image segmentation methods, apart from the commonly used region of interest (ROI) analysis is of great importance for accurate tumour recognition and delineation in PET. Tumour delineations are used for several clinical tasks, including PET-based radiation therapy planning and well-grounded quantification of volumetric and radiomic features (4\u0026ndash;6). Although image segmentation is possible in PET images, their success can be affected by several factors. Poor spatial resolution reduces the contrast between different objects in the images, and often making unclear the boundaries between adjacent objects. Big variabilities in the position, shape, and texture of the pathologies present complexities in choosing the segmentation methods suitable for those cases. High noise in PET images leads to additional challenges in the segmentation methods that use the standardized uptake value (SUV) to adjust their parameters (4).\u003c/p\u003e \u003cp\u003eTumours in PET scans can be delineated through manual segmentation. However, it is labour-intensive and can be time-consuming, especially for large datasets or multiple scans. On top of that it can produce inaccurate results due to inter- and intra-reader variability (7).\u003c/p\u003e \u003cp\u003eTo tackle these challenges, various computer-aided segmentation techniques have been developed. The most common used technique is based on thresholding (fixed, adaptive, and intensive) due to its simplicity which transforms a grey-level image into a binary image (8). This is completed by labelling all voxels higher than a set value to be foreground, while all remaining voxels as background (9). Such method is prone to inaccurate segmentation due to presence of high noise in PET images (8,10). Texture present in the images also impact the segmentation (11). Region-based segmentation techniques use the intensities of PET images; however, are much more focused on the local distribution (homogeneity) of the intensities found in the image (12). In gradient-based methods, a big change in intensity values on the edges of a PET image is an indication of the boundary of an object. To detect the location of the change in intensities, the gradient of the image is typically calculated between a voxel and its surrounding voxels (10,13). This does not always give the best results due to low resolution and high partial volume effect (PVE) as well presence of texture in the images (14,15). Other computer-aided methods include stochastic and learning methods, and joint segmentation methods. Similarly to manual segmentation, computer-aided segmentation techniques also demonstrate a few setbacks, including the necessity of manual input (16), high sensitivity to PVEs (17), challenges when assumptions are not satisfied (18), and the required calibration for the various types of scanners (19).\u003c/p\u003e \u003cp\u003eDue to these challenges, more accurate, strong, and automated segmentation techniques are required. Deep-learning (DL) methods, specifically those that are based on convolutional neural networks have been promising (20). However, many limitations have yet to be addressed. PET images have high noise and low resolution which makes the job of specifying ground-truth for training DL-based techniques challenging (21). DL-based methods commonly use manual delineation as substitute of ground truth, but the low accuracy and high variability of PET images can restrict the use of manual segmentation as ground truth (22). Additionally, DL-based methods need a substantial amount of training data to achieve optimal results, which is not easily obtainable because PET is not a widely used imaging modality (23).\u003c/p\u003e \u003cp\u003eThis study investigates the training effectiveness of U-Net, a widely used convolutional neural network (CNN) architecture, using different PET segmentation methods (manual and semi-automatic).\u003c/p\u003e"},{"header":"METHODOLOGY","content":"\u003cp\u003eA database of images from the Quantitative Imaging Network in Head and Neck Cancer (QIN-HEADNECK) was used in this study. The images were originally collected for a study comparing manual segmentation methods with semi-automated segmentation methods (24). The database consists of 59 PET/CT scans taken from 59 subjects, each containing multiple head and neck lesions. Three experts, a faculty professor and two residents, performed semi-automated and manual segmentations on 217 lesions of the 59 subjects, each repeated twice, resulting in 12 segmented binary masks (ROI) for each lesion. Intra and inter observer variability was observed in the segmentation masks.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eMaximum Voting (MV):\u003c/h2\u003e \u003cp\u003eEach of the 12 segmented mask of each tumour was first extracted as a 3D volume using the boundary coordinates of the corresponding ROI. In absence of gold standard, three different surrogate masks of the gold standard of each lesion were generated using a majority voting (MV) method to obtain three reference masks \u0026ndash; i) using 6 manual segmentations (M-MV), ii) using 6 semiautomatic segmentations (SA-MV) and iii) combining all 12 segmentations (A-MV). The overall process is illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. In MV method, pixels that are common to most of the segments are considered as part of the tumour. For manual (M-MV) and semiautomatic (SA-MV), only those pixels were considered as part of the tumour that were common to more than three segments. On the other hand, for all segmentation (A-MV) case, a pixel needs to be common in more than six segments to be considered as part of the tumour.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eROI Pre-processing and U-Net Training:\u003c/h3\u003e\n\u003cp\u003eEach tumour was extracted as a 3D volume from the original grey level images using the boundary coordinates of the corresponding ROI. The size of each 3D grey level volume was then increased to include a greater number of background voxels in such a way that the ratio of background-to-tumour volume were greater than 2.5. The size of each corresponding mask was also increased in a similar way. If there was any adjacent tumour, it was removed from the extracted volume without removing other high uptake tissues. Histogram equalization was applied on the extracted tumours to enhance contrast. A grey level normalization process was then carried. All the normalized grey level volumes and binary masks were resized to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:64\\times\\:64\\times\\:64\\:\\)\u003c/span\u003e\u003c/span\u003e3D cubic volumes to meet the input layer size requirements of the 3D U-net which is a widely used deep convolution neural network (DCNN) based segmentation algorithm (20). The output of the U-net is 3 training models based on 3 different majority voting ROIs as described above. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e describes the process of generating three training models using U-net.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe 3D U-Net model is an advanced deep learning architecture designed for volumetric (3D) medical image segmentation. An extension of the original 2D U-Net, this model adapts the original architecture to handle three-dimensional data, making it highly effective for segmenting complex structures in 3D medical imaging modalities such as MRI and CT scans.\u003c/p\u003e\n\u003ch3\u003eArchitecture:\u003c/h3\u003e\n\u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eEncoder (Down-sampling Path)\u003c/b\u003e: The encoder utilizes 3D convolutional layers to extract hierarchical features from the input volume. Max pooling operations reduce spatial dimensions while increasing feature depth.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eBottleneck\u003c/b\u003e: At the deepest level, the bottleneck captures abstract features with a series of 3D convolutions, providing a rich representation of the volumetric data.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eDecoder (Up-sampling Path)\u003c/b\u003e: The decoder reconstructs the spatial resolution of the image through transposed convolutions, combining up-sampled feature maps with high-resolution features from the encoder using skip connections.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eOutput Layer\u003c/b\u003e: The final layer consists of 3D convolutional operations with activation functions (softmax for multi-class or sigmoid for binary segmentation) to produce the segmentation maps.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe 3D U-net was trained on a machine with and NVIDIA GeForce GTX 1060 GPU. All processing, training, and analysis were performed using MATLAB 2022 platform. For the training, 150 epochs using the Adam Optimizer with a learning rate of 0.001, a batch size of 3, and a binary cross-entropy loss function were selected.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eData Splitting:\u003c/h3\u003e\n\u003cp\u003eFor each extracted lesion, the metabolic tumour volume (MTV), contrast-to-noise ratio (CNR), and signal-to-noise ratio (SNR) was determined. The lesions were randomly divided into three sets (65% training, 10% validation, and 25% test) in such a way that they would have similar gaussian distributions of MTV, CNR and SNR to avoid bias in training, validation, and testing sets.\u003c/p\u003e \u003cp\u003eIn the testing phase, only the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:64\\times\\:64\\times\\:64\\:\\)\u003c/span\u003e\u003c/span\u003e3D cubic grey level volumes were fed to the U-net. Finally, the dice similarity coefficient (DSC) between the predicted segmentation generated using three different training sets were compared with the three surrogate ground truth segmentations. Images used in the test set were not part of the training model. DSC is calculated as:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\text{D}\\text{S}\\text{C}=\\frac{2\\left|{\\text{S}}_{\\text{P}\\text{R}\\text{E}}\\cap\\:{\\text{S}}_{\\text{S}\\text{G}\\text{T}}\\right|}{{\\text{S}}_{\\text{P}\\text{R}\\text{E}}+{\\text{S}}_{\\text{S}\\text{G}\\text{T}}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere,\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{\\text{S}}_{\\text{S}\\text{G}\\text{T}}\\)\u003c/span\u003e \u003c/span\u003e is surrogate ground truth segmentation (M-MV, SA-MV and A-MV)\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{\\text{S}}_{\\text{P}\\text{R}\\text{E}}\\)\u003c/span\u003e \u003c/span\u003e is predicted segmentation using 3D U-net and three different training sets (M-MV, SA-MV and A-MV)\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cp\u003eThe representative contours of six manual and six semi-automatic segmentation methods on a single PET slice illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The figure reveals that manual segmentation exhibits significant uncertainty near the tumour boundary, whereas the semi-automatic methods demonstrate more consistency.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe results of the Maximum Voting (MV) segmentation methods\u0026mdash;M-MV, SA-MV, and A-MV compared with each individual manual segmentation provide valuable insights into the strengths and limitations of different segmentation approaches. The results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. M-MV, which aggregates manual segmentations using maximum voting, achieves the highest average Dice Similarity Coefficient (DSC) of 0.75. This is because the M-MV was derived from the manual segmentations. However, the variability in DSC scores, ranging from 0.73 to 0.79, suggests some inconsistency between trials, likely caused by human subjectivity in boundary delineation. SA-MV, which combines semi-automatic segmentation methods using maximum voting, achieves an average DSC of 0.72 (ranging from 0.70 to 0.74) when compared to manual segmentations. While semi-automatic methods are generally more efficient than fully manual approaches, their reliability tends to decline when directly compared to manual segmentation.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the average DSC for each of the three MV segmentations with each manual segmentation\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMaximum Voting (MV)\u003c/p\u003e \u003cp\u003eSegmentation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUser 1- Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUser 1 - Trail12\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUser 2 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUser 2 - Trail 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUser 3 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eUser 3 Trail \u0026minus;\u0026thinsp;2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eAverage\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eM-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eOn the other hand, A-MV, representing the combined MV segmentations of both approaches, achieves an average DSC of 0.74, nearly matching M-MV's performance. With scores ranging from 0.72 to 0.76, A-MV shows less variability compared to SA-MV, reflecting its independence from user input. While A-MV has not yet surpassed the accuracy of manual methods, it shows significant promise for refinement, especially in scenarios where efficiency and consistency are critical.\u003c/p\u003e \u003cp\u003eThe results of the Maximum Voting (MV) segmentation methods\u0026mdash;M-MV, SA-MV, and A-MV compared with each individual semi-automatic segmentation is shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The results of M-MV, which combines manual segmentations through maximum voting, achieves a consistent DSC score of 0.74 to 0.75 across all trials, with an average DSC of 0.74.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the average DSC for each of the three MV segmentations with each semi-automatic segmentation\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMaximum Voting (MV)\u003c/p\u003e \u003cp\u003eSegmentation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUser 1- Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUser 1 - Trail12\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUser 2 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUser 2 - Trail 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUser 3 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eUser 3 Trail \u0026minus;\u0026thinsp;2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eAverage\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eM-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.85\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eSA-MV, which relies on semi-automatic segmentation methods combined through maximum voting, performs the best overall, with DSC values ranging from 0.87 to 0.92 and an average of 0.91. This high performance indicates that semi-automatic methods can produce a segmentation that aligns closely with the MV, likely due to their ability to standardize the segmentation process while maintaining some level of user control. A-MV achieves DSC scores ranging from 0.84 to 0.86, with an average DSC of 0.85. It falls between that of M-MV and SA-MV, indicating solid performance but not quite matching the high accuracy of semi-automatic methods.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e presents the average DSC values between each manual segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately. The table shows how each training method performed in comparison to the individual manual segmentation. U-net with M-MV training model achieved DSC values ranging from 0.70 to 0.74, with an average of 0.71. The relatively consistent average DSC score of 0.71 indicates that, training with manual segmentation provides stable results.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the average DSC between each manual segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTraining Model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUser 1- Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUser 1 - Trail12\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUser 2 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUser 2 - Trail 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUser 3 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eUser 3 Trail \u0026minus;\u0026thinsp;2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eAverage\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eM-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.71\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eSegmentation performed using SA-MV resulted in DSC values ranging from 0.69 to 0.73, with an average of 0.71. These results are comparable to those of M-MV, with SA-MV showing slightly lower performance in some instances but remaining consistent across trials. The minimal differences in performance between M-MV and SA-MV indicate that when the segmentation map generated by the SA-MV U-Net model is compared to manual methods, no significant improvement in accuracy is observed. The same conclusion applies to the A-MV U-Net model.\u003c/p\u003e \u003cp\u003eThe average DSC values between each semi-automatic segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately is shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the average DSC between each semi-automatic segmentation and the U-net segmentation trained with each MV (M-MV, SA-MV and A-MV) separately\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTraining Model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUser 1- Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUser 1 - Trail12\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUser 2 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUser 2 - Trail 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUser 3 - Trail 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eUser 3 Trail \u0026minus;\u0026thinsp;2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eAverage\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eM-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.72\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.82\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eA-MV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.79\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eM-MV U-net model achieved DSC values ranging from 0.71 to 0.73, with an average of 0.72. These results indicate that M-MV U-net model provides relatively consistent performance across different users and trials. On the other hand, SA-MV U-net model, which relies on semi-automatic segmentation methods combined through maximum voting, performs significantly better with DSC values ranging from 0.80 to 0.83, and an average of 0.82. This higher performance suggests that model trained with the MV generated with semi-automatic methods, when aggregated, can achieve a more accurate segmentation, likely due to their ability to standardize the segmentation process while still involving user input. The improved DSC values across all trials indicate that SA-MV outperforms M-MV, providing a more accurate segmentation solution. A-MV, achieved DSC values between 0.78 and 0.80, with an average of 0.79. These results show that A-MV delivers solid performance, with DSC scores consistently higher than M-MV but slightly lower than SA-MV. The results suggest that combined training model can provide reliable and consistent segmentations compared to the M-MV, but they still fall short of the accuracy achieved by semi-automatic methods.\u003c/p\u003e \u003cp\u003eWhen M-MV U-Net is used, it achieves the highest DSC with the M-MV method (0.83). This suggests that manual segmentations, when aggregated through maximum voting, provide the most accurate results when paired with the M-MV U-Net model. However, the performance decreases slightly when M-MV U-Net is compared with the other two methods. For SA-MV U-Net, it achieves a DSC of 0.73, and for A-MV U-Net, it reaches a DSC of 0.76. This indicates that M-MV U-Net is not as effective when paired with semi-automatic or combined methods.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWhen the SA-MV U-Net model is used, the method performs very well, particularly in comparison with the semi-automatic methods itself. SA-MV U-Net achieves the highest DSC of 0.86 when paired with the SA-MV method, showing that semi-automatic segmentation provides the most accurate results with this model. The DSC is slightly lower when using SA-MV U-Net with M-MV (0.75) and A-MV (0.83), suggesting that the SA-MV U-Net model performs better when using semi-automatic segmentation compared to the other methods.\u003c/p\u003e \u003cp\u003eFor the combined A-MV U-Net model, it performs the best with A-MV, achieving the highest DSC of 0.86. This indicates that the A-MV U-Net model provides the most accurate segmentation results when combined with manual and semi-automated segmentation methods. When using A-MV U-Net with M-MV, the DSC is 0.76, and with SA-MV, the DSC is 0.81, showing a generally good performance, although not as high as with its own A-MV method.\u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThe study examines the application of a 3D U-Net deep learning framework for head and neck tumour segmentation in PET images. By utilizing surrogate ground truths derived from manual, semi-automatic, and combined segmentation approaches, the study provides critical insights into optimizing segmentation performance in the challenging context of PET imaging.\u003c/p\u003e \u003cp\u003eThree distinct surrogate ground truths\u0026mdash;M-MV (manual), SA-MV (semi-automatic), and A-MV (all)\u0026mdash;were developed to evaluate their impact on segmentation accuracy. The manual-based M-MV demonstrated moderate Dice Similarity Coefficient (DSC) values (average 0.72), reflecting inherent limitations in manual delineation such as observer variability and subjective boundary definitions (25)​. In contrast, SA-MV exhibited the highest DSC values (average 0.82), indicating the potential of semi-automatic methods to achieve more consistent and reproducible segmentation (26). The inclusion of A-MV, combining manual and semi-automatic inputs, further demonstrated its utility by balancing robustness and accuracy with an average DSC of 0.79, surpassing M-MV while approaching the performance of SA-MV.\u003c/p\u003e \u003cp\u003eThe superior performance of SA-MV highlights the efficacy of semi-automatic methods in leveraging computational precision to minimize user-dependent variability (25,26). Semi-automatic methods standardize segmentation tasks, thereby enhancing reproducibility without entirely removing human oversight, which remains critical in clinical applications​. Furthermore, semi-automatic methods are better equipped to handle the low resolution and high noise levels of PET images (27), particularly when combined with deep learning architectures like 3D U-Net (28).\u003c/p\u003e \u003cp\u003eThe use of the 3D U-Net framework significantly enhanced segmentation performance across all surrogate ground truth models. The ability of the model to process volumetric data using an encoder-decoder structure with skip connections enabled detailed feature extraction and accurate tumour delineation (29). The SA-MV-trained 3D U-Net achieved the best performance when tested against the SA-MV ground truth, with DSC values consistently exceeding those of M-MV and A-MV models. This underscores the compatibility of semi-automatic segmentation data with deep learning-based refinement​.\u003c/p\u003e \u003cp\u003eThe A-MV-trained U-Net delivered comparable results to SA-MV while demonstrating greater versatility by integrating multiple segmentation strategies. However, the slightly lower DSC values compared to SA-MV suggest that introducing heterogeneity through combined inputs may add complexity that could hinder performance in specific scenarios​.\u003c/p\u003e \u003cp\u003eDespite its promising results, the study highlights several challenges in PET image segmentation. First, the inherent properties of PET images, such as noise and low resolution, complicate the delineation of tumour boundaries (8). 3D U-Net's architecture addresses some of these challenges, further advancements in pre-processing techniques, such as enhanced denoising and contrast enhancement, could improve outcomes (13,30)​. Second, the reliance on a relatively small dataset (59 subjects) limits the generalizability of the findings. Future studies should explore larger and more diverse datasets to improve model robustness​. Finally, the absence of a definitive gold standard for segmentation necessitates reliance on surrogate ground truths. Developing new techniques for generating more accurate ground truths could further refine training and evaluation processes​.\u003c/p\u003e \u003cp\u003eThe findings underscore the potential of combining semi-automatic methods with deep learning for clinical applications. By reducing dependency on manual delineation and standardizing segmentation outputs, this approach can streamline workflows in radiation therapy planning and tumour monitoring. Moreover, the adoption of SA-MV in clinical settings could address inter-observer variability, improving the reliability of PET-based diagnostics and treatment​.\u003c/p\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eThis study assesses the performance of a 3D U-Net deep learning framework for segmenting head and neck tumours in PET images, using surrogate ground truths generated through maximum voting (MV) from manual and semi-automatic segmentations. By comparing three distinct surrogate masks\u0026mdash;M-MV (manual), SA-MV (semi-automatic), and A-MV (combined)\u0026mdash;the study provides critical insights into the strengths and limitations of both manual and semi-automatic approaches in the context of deep learning-based tumour segmentation. The results demonstrate that the 3D U-Net model trained using the SA-MV surrogate mask consistently outperforms models trained with M-MV and A-MV, achieving the highest Dice Similarity Coefficient (DSC) of 0.82. This highlights the ability of semi-automatic methods to generate consistent and accurate ground truths, which are particularly well-suited for deep learning applications. In contrast, models trained on M-MV exhibited lower performance, likely due to the inherent inter- and intra-observer variability in manual delineation. The combined approach, A-MV, balanced accuracy and robustness, achieving a DSC of 0.79, making it a viable alternative where both manual and semi-automatic inputs are available. While the 3D U-Net model demonstrates its potential for automating tumour segmentation, its performance remains marginally lower than the standalone semi-automatic methods, which achieved a DSC of 0.91. These findings highlight the reliability of semi-automatic segmentation methods in producing consistent results, despite the added time required for their implementation. The study also indicates that while deep learning models are effective in standardizing and automating processes, their performance can be further enhanced by refining training datasets and pre-processing techniques. Additionally, the incorporation of advanced ground truth generation methods could significantly improve segmentation accuracy and increase the clinical applicability of these models.\u003c/p\u003e"},{"header":"Statements and Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors have no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMahbubunnabi Tamal and\u0026nbsp;Maha Alshammari\u0026nbsp;contributed to the study conception and design. Material preparation and data analysis were performed by Mahbubunnabi Tamal and\u0026nbsp;Maha Alshammari. The manuscript was written by Mahbubunnabi Tamal,\u0026nbsp;Maha Alshammari, and Maryam Alhashim.\u0026nbsp;Validation, interpretation, and critical review of the results were conducted by\u0026nbsp;Saleh I. Alzahraniand Murad Althobaiti. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEthics approval was not required for this study, as it utilizes publicly available data.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eMody MD, Rocco JW, Yom SS, Haddad RI, Saba NF. Head and neck cancer. The Lancet. 2021 Dec;398(10318):2289\u0026ndash;99.\u003c/li\u003e\n \u003cli\u003ePham N, Ju C, Kong T, Mukherji SK. Artificial Intelligence in Head and Neck Imaging. Seminars in Ultrasound, CT and MRI. 2022 Apr;43(2):170\u0026ndash;5.\u003c/li\u003e\n \u003cli\u003eLameka K, Farwell MD, Ichise M. Positron Emission Tomography. In 2016. p. 209\u0026ndash;27.\u003c/li\u003e\n \u003cli\u003eFoster B, Bagci U, Mansoor A, Xu Z, Mollura DJ. A review on segmentation of positron emission tomography images. Comput Biol Med. 2014 Jul;50:76\u0026ndash;96.\u003c/li\u003e\n \u003cli\u003eJha AK, Mena E, Caffo B, Ashrafinia S, Rahmim A, Frey E, et al. Practical no-gold-standard evaluation framework for quantitative imaging methods: application to lesion segmentation in positron emission tomography. Journal of Medical Imaging. 2017 Mar 3;4(1):011011.\u003c/li\u003e\n \u003cli\u003eMena E, Sheikhbahaei S, Taghipour M, Jha AK, Vicente E, Xiao J, et al. 18F-FDG PET/CT Metabolic Tumor Volume and Intratumoral Heterogeneity in Pancreatic Adenocarcinomas. Clin Nucl Med. 2017 Jan;42(1):e16\u0026ndash;21.\u003c/li\u003e\n \u003cli\u003eAngulakshmi M, Lakshmi Priya GG. Automated brain tumour segmentation techniques\u0026mdash; A review. Int J Imaging Syst Technol. 2017 Mar 21;27(1):66\u0026ndash;77.\u003c/li\u003e\n \u003cli\u003eTamal M. Intensity threshold based solid tumour segmentation method for Positron Emission Tomography (PET) images: A review. Heliyon. 2020;6(10).\u003c/li\u003e\n \u003cli\u003eSaha PK, Udupa JK. Optimum image thresholding via class uncertainty and region homogeneity. IEEE Trans Pattern Anal Mach Intell. 2001 Jul;23(7):689\u0026ndash;706.\u003c/li\u003e\n \u003cli\u003eTamal M. A phantom study to assess the reproducibility, robustness and accuracy of PET image segmentation methods against statistical fluctuations. PLoS One. 2019;14(7).\u003c/li\u003e\n \u003cli\u003eAnan N, Zainon R, Tamal M. A review on advances in 18F-FDG PET/CT radiomics standardisation and application in lung disease management. Insights Imaging. 2022 Dec 5;13(1):22.\u003c/li\u003e\n \u003cli\u003eFoster B, Bagci U, Mansoor A, Xu Z, Mollura DJ. A review on segmentation of positron emission tomography images. Comput Biol Med. 2014 Jul;50:76\u0026ndash;96.\u003c/li\u003e\n \u003cli\u003eTamal M. A hybrid region growing tumour segmentation method for low contrast and high noise Nuclear Medicine (NM) images by combining a novel non-linear diffusion filter and global gradient measure (HNDF-GGM-RG). Heliyon. 2019;5(12).\u003c/li\u003e\n \u003cli\u003eTamal M. A Phantom Study to Investigate Robustness and Reproducibility of Grey Level Co-Occurrence Matrix (GLCM)-Based Radiomics Features for PET. Applied Sciences. 2021 Jan 7;11(2):535.\u003c/li\u003e\n \u003cli\u003eTamal M. Grey Level Co-occurrence Matrix (GLCM) as a Radiomics Feature for Artificial Intelligence (AI) Assisted Positron Emission Tomography (PET) Images Analysis. IOP Conf Ser Mater Sci Eng. 2019 Oct 1;646(1):012047.\u003c/li\u003e\n \u003cli\u003eAristophanous M, Penney BC, Martel MK, Pelizzari CA. A Gaussian mixture model for definition of lung tumor volumes in positron emission tomography. Med Phys. 2007 Nov 17;34(11):4223\u0026ndash;35.\u003c/li\u003e\n \u003cli\u003eBrambilla M, Matheoud R, Secco C, Loi G, Krengli M, Inglese E. Threshold segmentation for PET target volume delineation in radiation treatment planning: The role of target‐to‐background ratio and target size. Med Phys. 2008 Apr 7;35(4):1207\u0026ndash;13.\u003c/li\u003e\n \u003cli\u003eBelhassen S, Zaidi H. A novel fuzzy C‐means algorithm for unsupervised heterogeneous tumor quantification in PET. Med Phys. 2010 Mar 25;37(3):1309\u0026ndash;24.\u003c/li\u003e\n \u003cli\u003eZaidi H, Abdoli M, Fuentes CL, El Naqa IM. Comparative methods for PET image segmentation in pharyngolaryngeal squamous cell carcinoma. Eur J Nucl Med Mol Imaging. 2012 May 31;39(5):881\u0026ndash;91.\u003c/li\u003e\n \u003cli\u003eRonneberger O, Fischer P, Brox T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In 2015. p. 234\u0026ndash;41.\u003c/li\u003e\n \u003cli\u003eLeung KH, Marashdeh W, Wray R, Ashrafinia S, Pomper MG, Rahmim A, et al. A physics-guided modular deep-learning based automated framework for tumor segmentation in PET. Phys Med Biol. 2020 Dec 21;65(24):245032.\u003c/li\u003e\n \u003cli\u003eGiraud P, Elles S, Helfre S, De Rycke Y, Servois V, Carette MF, et al. Conformal radiotherapy for lung cancer: different delineation of the gross tumor volume (GTV) by radiologists and radiation oncologists. Radiotherapy and Oncology. 2002 Jan;62(1):27\u0026ndash;36.\u003c/li\u003e\n \u003cli\u003eShen D, Wu G, Suk HI. Deep Learning in Medical Image Analysis. Annu Rev Biomed Eng. 2017 Jun 21;19(1):221\u0026ndash;48.\u003c/li\u003e\n \u003cli\u003eBeichel RR, Van Tol M, Ulrich EJ, Bauer C, Chang T, Plichta KA, et al. Semiautomated segmentation of head and neck cancers in 18F‐FDG PET scans: A just‐enough‐interaction approach. Med Phys. 2016 Jun 18;43(6Part1):2948\u0026ndash;64.\u003c/li\u003e\n \u003cli\u003ePfaehler E, Burggraaff C, Kramer G, Zijlstra J, Hoekstra OS, Jalving M, et al. PET segmentation of bulky tumors: Strategies and workflows to improve inter-observer variability. PLoS One. 2020 Mar 30;15(3):e0230901.\u003c/li\u003e\n \u003cli\u003ePhilip MM, Watts J, Moeini SNM, Musheb M, McKiddie F, Welch A, et al. Comparison of semi-automatic and manual segmentation methods for tumor delineation on head and neck squamous cell carcinoma (HNSCC) positron emission tomography (PET) images. Phys Med Biol. 2024 May 7;69(9):095005.\u003c/li\u003e\n \u003cli\u003eFoster B, Bagci U, Mansoor A, Xu Z, Mollura DJ. A review on segmentation of positron emission tomography images. Comput Biol Med. 2014 Jul;50:76\u0026ndash;96.\u003c/li\u003e\n \u003cli\u003eJiang H, Diao Z, Shi T, Zhou Y, Wang F, Hu W, et al. A review of deep learning-based multiple-lesion recognition from medical images: classification, detection and segmentation. Comput Biol Med. 2023 May;157:106726.\u003c/li\u003e\n \u003cli\u003eAzad R, Aghdam EK, Rauland A, Jia Y, Avval AH, Bozorgpour A, et al. Medical Image Segmentation Review: The Success of U-Net. IEEE Trans Pattern Anal Mach Intell. 2024 Dec;46(12):10076\u0026ndash;95.\u003c/li\u003e\n \u003cli\u003eTamal M. Nonlinear Diffusion Filter for Low Count Positron Emission Tomography Utilizing Orientation Information of Neighbouring Gradient Vectors. In: 2017 9th IEEE-GCC Conference and Exhibition (GCCCE). IEEE; 2017. p. 1\u0026ndash;3.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Head and Neck Tumour Segmentation, PET, 3D U-Net, Maximum Voting, Surrogate Ground Truth","lastPublishedDoi":"10.21203/rs.3.rs-6198985/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6198985/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccurate segmentation of head and neck tumours in PET images is critical for effective treatment planning, disease progression monitoring, and radiotherapy. However, achieving reliable ground truth data remains challenging due to inter- and intra-observer variability. The U-Net, a Deep Convolutional Neural Network (DCNN), has demonstrated strong potential for automated segmentation, yet the lack of definitive ground truth for training limits its full effectiveness. This study investigates the performance of a 3D U-Net deep learning framework trained using surrogate ground truth masks created through maximum voting (MV) from manual and semi-automatic segmentations. The analysis utilized the QIN-HEADNECK dataset, comprising a total of 217 lesions of 59 PET/CT scans. Each lesion was annotated by three radiologists, with two trials conducted for each segmentation method, resulting in a total of twelve trials. Three MV-based surrogate masks: M-MV (manual), SA-MV (semi-automatic), and A-MV (combined) were then generated. Segmentation performance was assessed using the Dice Similarity Coefficient (DSC). The results revealed that manual and semi-automatic segmentations achieved average DSC scores of 0.75 and 0.91, respectively, when compared to each MV result for individual trials. The combined MV (A-MV) produced DSC scores of 0.74 and 0.85 when compared to the MV results from the manual and semi-automatic segmentations, respectively. Among the 3D U-Net models, the framework trained with SA-MV achieved the highest average DSC score of 0.86, performing similarly to A-MV (0.86) and surpassing the model trained with M-MV (0.83). While the 3D U-Net outperformed manual segmentation (DSC of 0.83 vs. 0.75), its performance was still lower than that of semi-automatic segmentation (DSC of 0.86 vs. 0.91). These findings highlight the reliability of semi-automatic segmentation methods in producing consistent results, despite the added time required for their implementation. The study also indicates that while deep learning models are effective in standardizing and automating processes, their performance can be further enhanced by refining training datasets and pre-processing techniques. Additionally, the incorporation of advanced ground truth generation methods could significantly improve segmentation accuracy and increase the clinical applicability of these models.\u003c/p\u003e","manuscriptTitle":"Head and Neck Tumour Segmentation in PET Images: Performance Evaluation of 3D U-Net with Maximum Voting-Based Surrogate Ground Truth","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-26 09:17:54","doi":"10.21203/rs.3.rs-6198985/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"143f10e2-62b9-40e7-adbd-70c4b0d94cc4","owner":[],"postedDate":"March 26th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-23T04:00:48+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-26 09:17:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6198985","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6198985","identity":"rs-6198985","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00