DeepPyramid+: medical image segmentation using Pyramid View Fusion and Deformable Pyramid Reception

other OA: gold CC-BY-4.0
AI-generated summary by claude@2026-07, 2026-07-14

DeepPyramid+ improves medical image segmentation by using Pyramid View Fusion and Deformable Pyramid Reception to overcome challenges with diverse features, achieving significant performance gains across multiple modalities.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

This paper studies semantic segmentation for medical images and surgical videos, focusing on heterogeneous, deformable, and scale-varying targets as well as distortions like motion blur, reflection, and blunt or transparent content. Using an encoder-decoder architecture based on U-Net with a VGG16 encoder, the authors introduce Pyramid View Fusion (PVF) to capture narrow-to-wide global feature views centered at each pixel and Deformable Pyramid Reception (DPR) using deformable dilated convolutions and shape/scale-adaptive extraction to improve robustness; experiments across five intra-domain and two cross-domain datasets report that DeepPyramid+ outperforms state-of-the-art baselines, with ablation studies supporting each module’s contribution. A stated caveat is that the work is evaluated on specific intra- and cross-domain datasets rather than explicitly characterizing performance across all possible medical imaging modalities. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

PURPOSE: Semantic segmentation plays a pivotal role in many applications related to medical image and video analysis. However, designing a neural network architecture for medical image and surgical video segmentation is challenging due to the diverse features of relevant classes, including heterogeneity, deformability, transparency, blunt boundaries, and various distortions. We propose a network architecture, DeepPyramid+, which addresses diverse challenges encountered in medical image and surgical video segmentation. METHODS: The proposed DeepPyramid+ incorporates two major modules, namely "Pyramid View Fusion" (PVF) and "Deformable Pyramid Reception" (DPR), to address the outlined challenges. PVF replicates a deduction process within the neural network, aligning with the human visual system, thereby enhancing the representation of relative information at each pixel position. Complementarily, DPR introduces shape- and scale-adaptive feature extraction techniques using dilated deformable convolutions, enhancing accuracy and robustness in handling heterogeneous classes and deformable shapes. RESULTS: Extensive experiments conducted on diverse datasets, including endometriosis videos, MRI images, OCT scans, and cataract and laparoscopy videos, demonstrate the effectiveness of DeepPyramid+ in handling various challenges such as shape and scale variation, reflection, and blur degradation. DeepPyramid+ demonstrates significant improvements in segmentation performance, achieving up to a 3.65% increase in Dice coefficient for intra-domain segmentation and up to a 17% increase in Dice coefficient for cross-domain segmentation. CONCLUSIONS: DeepPyramid+ consistently outperforms state-of-the-art networks across diverse modalities considering different backbone networks, showcasing its versatility. Accordingly, DeepPyramid+ emerges as a robust and effective solution, successfully overcoming the intricate challenges associated with relevant content segmentation in medical images and surgical videos. Its consistent performance and adaptability indicate its potential to enhance precision in computerized medical image and surgical video analysis applications.
Full text 36,919 characters · extracted from pmc-nxml · 5 sections · click to expand

Related

U-Net [ 7 ] was initially proposed for medical image segmentation and achieved succeeding performance being attributed to its skip connections. Many U-Net-based architectures have been proposed over the past years to improve the segmentation accuracy and address different flaws and restrictions in the previous architectures [ 8 – 14 ]. Attention mechanisms can be broadly described as the techniques to guide the network’s computational resources (i.e.,the convolutional operations) toward the most determinative features in the input feature map [ 9 , 15 , 16 ]. Such mechanisms have been especially proven to be gainful in the case of semantic segmentation. The scSE blocks [ 15 ] aim to recalibrate the feature maps based on pixel-wise and channel-wise global features. BARNet [ 12 ] adopts a bilinear-attention module to extract the cross-dependencies between the different channels of a convolutional feature map. PAANET [ 11 ] uses a double-attention module to model semantic dependencies between channels and spatial positions in the convolutional feature map. Fusion modules can be characterized as modules designed to improve semantic representation via combining several feature maps. The input feature maps could range from varying-level semantic features to the features coming from parallel operations. PSPNet [ 17 ] adopts a pyramid pooling module (PPM) containing parallel sub-region average pooling layers followed by upsampling to fuse the multi-scale sub-region representations. Atrous spatial pyramid pooling (ASPP) [ 18 , 19 ] was proposed to deal with objects’ scale variance by aggregating multi-scale features extracted using parallel varying-rate dilated convolutions. CPFNet [ 13 ] uses another fusion approach for scale-aware feature extraction. Fig. 1 Overall architecture of DeepPyramid+ consisting of encoder blocks of the VGG16 network, and the proposed PVF and DPR modules. The numbers in each block correspond to the output feature map’s dimensions Overall architecture of DeepPyramid+ consisting of encoder blocks of the VGG16 network, and the proposed PVF and DPR modules. The numbers in each block correspond to the output feature map’s dimensions

Conclusion

In recent years, considerable attention has been devoted to computerized medical image and surgical video analysis. A reliable relevant-instance-segmentation approach is a prerequisite for a majority of these applications. In this paper, we introduce a novel network architecture for semantic segmentation that addresses the challenges encountered in medical image and surgical video segmentation. Our proposed architecture, DeepPyramid+, incorporates two innovative modules, namely “Pyramid View Fusion” and “Deformable Pyramid Reception.” Experimental results demonstrate the effectiveness of DeepPyramid+ in capturing object features in challenging scenarios, including shape and scale variation, reflection and blur degradation, blunt edges, and deformability, resulting in competitive performance in cross-domain segmentation compared to state-of-the-art networks. The ablation study validates the efficacy of the proposed modules in DeepPyramid+, showcasing their performance across diverse datasets. The obtained promising results indicate the potential of DeepPyramid+ to enhance the precision in various computerized medical imaging and surgical video analysis applications.

Methodology

We present a segmentation network that focuses on (I) modeling heterogeneous classes featuring deformations, shape, scale, color, and context variation, (II) dealing with content distortion due to motion blur and reflection, and (III) handling objects’ transparency and blunt boundaries (Fig.  1 ). At its core, our network adopts the U-Net architecture, with the encoder part being set to VGG16. We develop two decoder modules specifically tailored to tackle the mentioned challenges: (1) Pyramid View Fusion (PVF) , which aims to replicate a deduction process within the neural network analogous to the functioning of the human visual system by enhancing the representation of relative information at each individual pixel position. (2) Deformable Pyramid Reception (DPR) , which addresses the limitations of regular convolutional layers by introducing deformable dilated convolutions and shape- and scale-adaptive feature extraction techniques. This module allows for handling the complexities of heterogeneous classes and deformable shapes, resulting in improved accuracy and robustness in the segmentation performance. We specify the functionality of each module in the following subsections. Additional discussions regarding the effectiveness of each module and an analysis of the complexity for each module are available in the supplementary material. Notations . Throughout this paper, we represent convolutional layers with a kernel size of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(k\times k)$$\end{document} ( k × k ) , dilation of d , m output channels, and g groups as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\circledast _{k,d}^{m,g}$$\end{document} ⊛ k , d m , g . For deformable convolutions, we use the symbol \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${\tilde{\circledast }}_{k,d}^{m,g}$$\end{document} ⊛ ~ k , d m , g . Additionally, we illustrate the average-pooling layer with a kernel size of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(k\times k)$$\end{document} ( k × k ) and a stride of s pixels as and global average pooling as . The symbol \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$+\!\!\!\!+\,_{D}$$\end{document} + + D denotes feature map concatenation over dimension D . Furthermore, we employ \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\Uparrow ^{(W_{out}, H_{out})}$$\end{document} ⇑ ( W out , H out ) and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\Downarrow ^{(W_{out}, H_{out})}$$\end{document} ⇓ ( W out , H out ) for upsampling and downsampling operations with a scale factor of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(W_{out}, H_{out})$$\end{document} ( W out , H out ) , respectively. We use \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma (\cdot )$$\end{document} σ ( · ) to represent the Softmax operation, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\Vert \cdot \Vert _{n}$$\end{document} ‖ · ‖ n for layer normalization over the last n dimensions, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {R}(\cdot )$$\end{document} R ( · ) for the ReLU nonlinearity function, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\tau (\cdot )$$\end{document} τ ( · ) for the hard tangent hyperbolic function. To optimize computational complexity, the initial step involves creating a bottleneck by employing a convolutional layer with a kernel size of one, as illustrated in Fig.  2 . Following this dimensionality reduction stage, the resulting convolutional feature map is fed into four parallel branches. The first branch features a global average pooling layer, which is subsequently followed by upsampling. The other three branches employ average pooling layers with progressively increasing filter sizes while maintaining a stride of one pixel. The use of a one-pixel stride is specifically important to achieve a pixel-wise centralized pyramid view, as opposed to the region-wise pyramid attention approach employed in PSPNet [ 17 ]. The output feature maps from all branches are then concatenated and fed into a convolutional layer with four groups, for extracting inter-channel dependencies during dimensionality reduction. Subsequently, a regular convolutional layer is applied to extract joint intra-channel and inter-channel dependencies. The resulting feature map is then passed through a layer-normalization function, which helps normalize the activations for improved stability and performance. Fig. 2 The detailed architecture of the PVF and DPR modules The detailed architecture of the PVF and DPR modules The architecture of the Deformable Pyramid Reception (DPR) module, as depicted in Fig.  2 , can be described as follows. Initially, the upsampled coarse-grained semantic feature map from the preceding layer is concatenated with its symmetric fine-grained feature map from the encoder. Subsequently, these concatenated features are passed through three parallel branches. The first branch employs a regular convolution operation, while the other two branches utilize deformable convolutions with different dilation rates of three and six. The structured convolution covers the immediate neighboring pixels up to one pixel to the central pixel. The deformable convolutions with the dilation rate of three and six cover an area from two to four and five to seven pixels far away from each central pixel, respectively. Accordingly, the DPR module forms a learnable sparse receptive field of size \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$15\times 15$$\end{document} 15 × 15 pixels by incorporating these layers. These layers share the weights to avoid imposing a huge number of trainable parameters. To compute the feature-map-adaptive offset field for each deformable convolution, a regular convolution operation is employed. Considering the target area of the two deformable convolutions, the offset field should be computed based on the internal content within four and seven pixels away from each central pixel ( \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k=9$$\end{document} k = 9 , \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k=15$$\end{document} k = 15 ). The computed offset values are then passed through a tangent hyperbolic function, which clips them within the range of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$[-1, 1]$$\end{document} [ - 1 , 1 ] , to ensure that each deformable convolution adaptively covers an area within the range of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$[k-1, k+1]$$\end{document} [ k - 1 , k + 1 ] . The offset field provides two values per element in the deformable convolutional kernel (horizontal and vertical offsets). Accordingly, the number of offset field’s output channels for a deformable convolution with a kernel of size \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$3\times 3$$\end{document} 3 × 3 is equal to 18. This enables the deformable convolution to spatially adjust its receptive field based on the learned offset values, improving its ability to capture contextually relevant information. The output feature maps of the parallel structured and deformable convolutions are then passed through a feature fusion decision (FFD) module [ 4 ]. This module determines the significance of each input feature map based on the spatial descriptors using pixel-wise convolutions. These descriptors are concatenated and subjected to a Softmax operation, resulting in normalized descriptors. The normalized descriptors determine the pixel-wise contribution or weight of each input convolutional feature map in the final fused feature map. The output feature map of the FFD module is obtained as a weighted sum of the input feature maps, where the normalized descriptors serve as pixel-wise weights. The resulting feature map from the FFD module goes through a series of additional operations for deeper feature extraction and normalization. Table 1 Specifications of the single-domain and cross-domain datasets Application Modality Objects Folds Train \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | test size (per fold) Reference Single domain Cataract Video Instruments 4 207 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 138 Cataract-1K [ 20 ] Laparoscopy Video Instruments 4 109 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 1179 Endovis [ 21 ] Endometriosis Video Endometrial implant 4 119 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 39 ENID [ 22 ] Prostate MR MRI Prostate 4 275 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 110 MS-Net [ 23 ] Retina OCT IRF Fluid 4 105 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 299 RETOUCH [ 24 ] (Spectralis) Application Modality Objects Source \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | target set Folds Train \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | test size (per fold) Reference Cross-domain Cataract Video Instruments Cataract-1K \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | CaDIS 4 207 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 458 Cataract-1K \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | CaDIS [ 25 ] Prostate MR MRI Prostate BMC \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | BIDMC 4 275 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\vert $$\end{document} | 64 MS-Net [ 23 ] This dataset is a small subset of Cataract-1K dataset that will be released upon the acceptance of this paper Specifications of the single-domain and cross-domain datasets This dataset is a small subset of Cataract-1K dataset that will be released upon the acceptance of this paper Fig. 3 Exemplary images from the different datasets along with their corresponding overlayed masks Exemplary images from the different datasets along with their corresponding overlayed masks

Experimental

Table  3 reports the segmentation performance of the proposed and state-of-the-art networks across three different modalities. DeepPyramid+ consistently demonstrates the highest average performance across all datasets with various backbones, while other methods, such as CPFNet, exhibit varying performance with different backbones and 2.22% compared to DeepPyramid+, respectively. Besides, DeepPyramid+ achieves the best results with all three backbones for endometrial implants and prostate segmentation and the best results with ResNet34 and ResNet50 backbones for IRF segmentation in OCT. Considering instrument segmentation performance (Table  4 ), DeepPyramid+ with VGG16 backbone shows more than 5.6% gain in segmentation compared to CPFNet as its main alternative (58.93% vs. 53.29%). Across all backbones, DeepPyramid+ with VGG16 backbone shows more than 2.7% higher performance compared to other methods. Besides, the best results for both datasets correspond to DeepPyramid+ with VGG16 backbone. Overall, DeepPyramid+ with our suggested backbone (VGG16) achieves the best segmentation performance in instrument and organ/disease segmentation. Table 3 Quantitative comparisons among the performance of DeepPyramid+ and alternative methods in organ and disease segmentation, with top two results shown in italic and bold, respectively Modality Endometriosis surgery MRI OCT Backbone Network IoU (%) Dice (%) IoU (%) Dice (%) IoU (%) Dice (%) Avg. IoU (%) VGG16 UNet+ 51.02 64.94 72.44 82.30 51 . 89 64 . 95 58 . 45 scSENet 48.95 62.95 72.31 82.23 52.12 65.18 57.79 FEDNet 50.38 64.16 71.03 81.82 48.70 62.31 56.70 CE-Net 44.40 58.48 70.84 81.27 47.33 61.10 54.19 CPFNet 52 . 82 66 . 09 73 . 28 82 . 80 47.90 61.17 58.00 UNetPP 51.02 64.83 72.49 82.32 51.58 64.60 58.36 DeepPyramid+ 53.22 66.37 76.02 85.36 50.79 63.99 60.01 ResNet34 scSENet 46 . 20 60 . 65 72 . 99 82 . 66 49 . 64 63 . 31 56 . 28 FEDNet 28.19 37.85 70.24 80.61 45.34 59.79 47.92 CE-Net 11.02 18.94 13.76 17.51 15.68 25.53 13.49 CPFNet 23.07 35.54 59.95 73.26 18.29 28.99 33.77 DeepPyramid+ 54.04 67.63 73.46 83.66 51.22 64.82 59.57 ResNet50 UPerNet 48 . 93 62 . 56 72 . 73 82 . 68 46 . 17 60 . 18 55 . 94 DeepLabV3+ 43.00 56.66 70.83 80.79 44.64 58.86 52.82 DeepPyramid+ 53.11 67.12 73.93 83.97 50.99 64.63 59.34 The best results are shown in bold and the second-best are shown in results italic Table 4 Quantitative comparisons among the performance of DeepPyramid+ and alternative methods in instrument segmentation, with top two results shown in italic and bold, respectively Modality Cataract surgery Laparoscopy surgery Backbone Network IoU (%) Dice (%) IoU (%) Dice (%) Avg. IoU (%) VGG16 UNet+ 45.79 56.73 57 . 74 70 . 29 51.76 scSENet 45.74 56.19 56.14 69.08 50.94 FEDNet 44.25 55.45 53.69 66.69 48.97 CE-Net 41.72 53.04 48.91 62.34 45.31 CPFNet 49 . 42 60 . 15 57.16 69.60 53 . 29 UNetPP 45.74 56.67 57.40 69.88 51.57 DeepPyramid+ 56.48 66.40 61.39 73.09 58.93 ResNet34 scSENet 50.00 60.91 50.71 63.88 50.35 FEDNet 47.10 58.68 23.34 32.44 35.22 CE-Net 16.46 26.52 30.73 44.48 23.59 CPFNet 28.70 41.14 49.28 63.40 38.99 RAUNet 43.36 55.36 1.62 2.97 22.49 BARNet 51.78 63.09 51 . 09 64 . 70 51 . 43 PAANet 51 . 14 58 . 68 48.91 61.85 50.02 DeepPyramid+ 49.11 59.72 57.14 69.80 53.12 ResNet50 UPerNet 56.27 66.93 56 . 08 68 . 70 56.17 DeepLabV3+ 36.98 49.16 48.80 62.04 42.89 DeepPyramid+ 49.28 59.90 57.12 70.06 53 . 20 The best results are shown in bold and the second-best are shown in results italic Quantitative comparisons among the performance of DeepPyramid+ and alternative methods in organ and disease segmentation, with top two results shown in italic and bold, respectively The best results are shown in bold and the second-best are shown in results italic Quantitative comparisons among the performance of DeepPyramid+ and alternative methods in instrument segmentation, with top two results shown in italic and bold, respectively The best results are shown in bold and the second-best are shown in results italic Table  5 compares the cross-domain segmentation performance of DeepPyramid+ and its best two alternatives for three backbones (considering single-domain results in Table  3 and Table  4 ). Overall, DeepPyramid+ consistently outperforms other methods across all backbones. Considering the MRI dataset, DeepPyramid+ with VGG16 backbone shows more than 4.8% gain in Dice compared to alternatives. For instrument segmentation in cataract surgery, DeepPyramid+ with the VGG16 backbone exhibits an impressive improvement of approximately 19.5% in Dice score compared to CPFNet with the same backbone (55.10% vs. 35.59%), and a 17% improvement compared to the best alternative across all backbones (55.10% vs. 38.10% achieved by UPerNet). This exceptional performance in dealing with cross-domain distribution gaps [ 28 ] can be attributed to the effectiveness of the proposed modules in incorporating multi-scale local and global features. Table 5 Quantitative comparisons of cross-domain performance among DeepPyramid+ and state-of-the-art methods, with top two results shown in italic and bold, respectively Modality MRI Cataract surgery Backbone Network IoU (%) Dice (%) Network IoU (%) Dice (%) VGG16 CPFNet 40 . 66 54 . 24 UNet+ 26 . 26 35 . 59 UNet++ 36.30 49.30 CPFNet 25.14 34.04 DeepPyramid+ 44.43 59.11 DeepPyramid+ 42.93 55.10 ResNet34 scSENet 40 . 23 53 . 19 BARNet 20 . 22 29 . 31 FEDNet 33.05 44.47 PAANet 14.01 20.48 DeepPyramid+ 41.52 56.14 DeepPyramid+ 32.87 43.96 ResNet50 UPerNet 38 . 45 51 . 78 UPerNet 28 . 40 38 . 10 DeepLabV3+ 37.55 49.62 DeepLabV3+ 9.14 14.16 DeepPyramid+ 38.89 53.00 DeepPyramid+ 29.76 40.54 The best results are shown in bold and the second-best are shown in results italic Quantitative comparisons of cross-domain performance among DeepPyramid+ and state-of-the-art methods, with top two results shown in italic and bold, respectively The best results are shown in bold and the second-best are shown in results italic Table  6 provides an ablation study of DeepPyramid+ components. The results suggest that both PVF and DPR modules contribute significantly to improvements in segmentation performance across all datasets. This impact is more prominent in the case of cataract surgery, where the addition of PVF and DPR modules lead to a 4.95% and 4.72% increase in the Dice coefficient, respectively. Table 6 Ablation study of DeepPyramid+ component across different datasets PVF DPR Endometriosis MRI Cataract surgery Laparoscopy surgery IoU (%) Dice (%) IoU (%) Dice (%) IoU (%) Dice (%) IoU (%) Dice (%) ✗ ✗ 51.02 64.94 72.44 82.30 45.79 56.73 57.74 70.29 ✔ ✗ 52.68 66.14 74.51 84.41 51.42 61.68 60.65 72.72 ✔ ✔ 53.22 66.37 76.02 85.36 56.48 66.40 61.39 73.09 Ablation study of DeepPyramid+ component across different datasets

Introduction

Semantic segmentation has emerged as a critical tool in computerized medical image and surgical video analysis, empowering numerous applications in various domains. In surgical videos, semantic segmentation is a prerequisite in several applications ranging from phase and action recognition, irregularity detection, surgical training, objective skill assessment, relevance-based compression, surgical planning, operation room organization, and so forth [ 1 – 4 ]. In the case of volumetric medical images, semantic segmentation can considerably aid in the diagnosis, treatment planning, and monitoring [ 5 ]. Automatic segmentation of medical images and videos can also reduce subjective errors caused by time constraints and workloads while enhancing treatment and surgical efficiency. Designing a neural network architecture for medical image and surgical video segmentation presents a challenge due to the diverse features exhibited by different relevant labels. Specifically, many classes of objects relevant to the medical image and surgical video analysis are heterogeneous, featuring deformable or amorphous instances, as well as color, texture, and scale variation. Besides in surgical videos, the problem of motion blur degradation becomes more critical due to the camera’s proximity to the surgical scene. Unlike general images, medical images and surgical videos may contain transparent relevant content (such as intraocular lens) or exhibit blunt boundaries, further complicating the task of semantic segmentation. Accordingly, an effective network for medical image and surgical video segmentation should be able to simultaneously deal with (I) heterogeneity and deformability in relevant objects, and (II) transparency, blunt edges, and distortions such as motion and defocus blur. This paper introduces a U-Net-based CNN for semantic segmentation, which effectively addresses the challenges associated with segmenting relevant content in medical images and surgical videos by adaptively capturing semantic information. 1 The proposed network, called DeepPyramid+, comprises two key modules: (i) Pyramid View Fusion (PVF) module, which offers a narrow-to-wide-angle global view of the feature map centering at each pixel position, and (ii) Deformable Pyramid Reception (DPR) module, responsible for performing shape-adaptive feature extraction on the input convolutional feature map 2 . We provide comprehensive experiments to compare the performance of DeepPyramid+ with state-of-the-art baselines for five intra-domain and two cross-domain datasets. Experimental results reveal the superiority of DeepPyramid+ compared to the baselines. Ablation studies confirm the effectiveness of each proposed module in boosting semantic segmentation performance. To support reproducibility and further investigations, we will release the PyTorch implementation of DeepPyramid+ and all dataset splits with the acceptance of this paper.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosis

MeSH descriptors

Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer Neural Networks, Computer

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-11T06:11:44.160905+00:00
pubmed
last seen: 2026-08-11T06:10:19.540980+00:00
unpaywall
last seen: 2026-05-14T19:30:52.867331+00:00
License: CC-BY-4.0 · commercial use OK · attribution required
Courtesy of the U.S. National Library of Medicine