Adnexal torsion diagnosis framework with CT-based adaptive preprocessing and deep neural networks.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by qwen3.7-flash, 2026-09-09

This study evaluated deep learning models for automated adnexal torsion diagnosis using CT scans, finding that a 3D EfficientNet architecture achieved the highest diagnostic performance among tested frameworks.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by qwen3.7-flash, 2026-09-09 · read from full text

This retrospective study evaluated the feasibility of using deep learning models, specifically 2D Multiple Instance Learning and 3D convolutional neural networks, to detect adnexal cystic torsion from abdominopelvic CT images. The researchers analyzed data from patients who underwent emergency or elective surgery for cystic adnexal lesions, comparing those with confirmed torsion against a control group including conditions like hemorrhagic cysts and ruptured tumors. The results demonstrated that these automated frameworks could effectively classify torsion cases based on imaging features, addressing diagnostic challenges where traditional signs like the whirl sign are often absent. Relevance to endometriosis: listed as one indication for GnRH antagonists, though the paper's main focus is uterine fibroids.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Adnexal cystic torsion is a gynecological emergency that requires prompt and accurate diagnosis followed by immediate surgical intervention to preserve ovarian function. Although ultrasound or computed tomography (CT) are commonly employed, their diagnostic accuracy is often limited by image quality and interobserver variability. This study aimed to evaluate the feasibility of applying deep learning models for the automated diagnosis of adnexal torsion using abdominopelvic CT. We retrospectively collected 1,191 CT scans from 514 women at two tertiary hospitals in Korea. The disease group included 259 surgically confirmed cases of adnexal torsion, while the control group comprised 255 cases of adnexal cysts without torsion, evaluated during diagnostic exploration for acute abdominal pain. Preprocessing included histogram-based clustering of Hounsfield unit distributions for adaptive windowing and anatomical region selection using the pretrained segmentation model TotalSegmentator. We trained both slice-based 2D multiple instance learning frameworks and fully volumetric 3D convolutional neural networks were using stratified five-fold cross-validation. Among the models tested, the 3D EfficientNet architecture demonstrated the best performance, with 79.25% AUC, 73.4% accuracy, 73.77% specificity. These findings highlight the potential of deep learning-assisted CT interpretation as a tool in gynecologic emergency settings.
Full text 69,108 characters · extracted from pmc-nxml · 6 sections · click to expand

Dataset

This retrospective study included consecutive patients who underwent diagnostic exploration or surgical resection for cystic adnexal lesions at Kyungpook National University Hospital (KNUH) and Kyungpook National University Chilgok Hospital (KNUCH) in Daegu, Republic of Korea, due to acute abdominal pain between January 1, 2011, and October 31, 2023. Most of the patients required emergency surgery after admission to the emergency departments of KNUH or KNUCH. All CT images were retrieved using the Picture Archiving and Communication System (PACS). The study was approved by the Institutional Review Boards of KNUH (KNUH 2024-07-020-001) and KNUCH (KNUCH 2024-07-022-001), with informed consent waived. The primary objective of the chart review was to classify patients into the adnexal cystic torsion (disease) group or the control group for the development of the deep learning algorithm. The disease group was further subdivided according to the initial clinical presentation department. A subgroup included patients with an ovarian or fallopian tubal cyst in whom a twisted vascular pedicle was identified during emergency surgery for diagnostic exploration. The other subgroup included patients whose torsion was discovered incidentally during elective surgeries, such as ovarian cystectomy. Diagnoses in this group included torsion of ovarian, para-tubal, or para-ovarian cysts or tumors, and hydrosalpinx. Most of the patients in this group underwent abdominopelvic CT due to acute abdominal pain, with the exception of cases discovered incidentally. The control group consisted of patients with adnexal cystic lesions who did not show torsion during emergency explorative surgery due to severe abdominal pain or discomfort, reflecting the clinical presentation of most cases of adnexal cystic torsion. All patients in this group underwent abdominopelvic CT due to acute abdominal pain. Diagnoses in the control group included hemorrhagic ovarian cysts, ectopic pregnancies in the uterine adnexa or uterine cornus, ruptured adnexal cysts or tumors, and cystic adenomyosis in the uterine cornus (Fig.  1 A). Only patients who had at least one abdominopelvic CT before surgery were included. For those with multiple scans, the most recent was selected. Patients who underwent only non-contrast CT scans were also included. The exclusion criteria were as follows: (1) no CT scan within 90 days prior to surgery, (2) absence of clearly identifiable whole adnexal cystic lesions on CT, or (3) poor image quality during preprocessing due to excessive artifacts. A total of 27 patients were excluded according to these criteria. The detailed inclusion and exclusion process was illustrated in Fig.  1 B. Fig. 1 ( A ) Diagnoses included in this study according to the anatomical structures in the female pelvic cavity. ( B ) Flowchart of the inclusion and exclusion criteria. ( A ) Diagnoses included in this study according to the anatomical structures in the female pelvic cavity. ( B ) Flowchart of the inclusion and exclusion criteria. In addition to the CT scans performed at KNUH and KNUCH, CT scans from other regional medical centers were also included. Due to differences in CT protocols and equipment, variability in image quality, imaging planes, contrast phases, and slice intervals was observed. Most scans included both non-contrast and contrast-enhanced phases. However, some patients only underwent non-contrast imaging due to misdiagnosis by the primary physician (e.g., urolithiasis) or certain medical conditions such as poor renal function or chronic kidney disease, preventing contrast administration. As a result, some patients had transverse, coronal, and sagittal images in both venous and arterial phases, while others had only partial or single-phase images. All CT findings and reports were verified by an experienced gynecologist (L.J.) and a gynecological radiologist (P.S.Y.). Although coronal or sagittal images were lacking in many cases, transverse images were available for all patients. Therefore, the deep learning model was developed using transverse images exclusively. The size of the adnexal cystic lesions was measured using the longest diameter on the transverse CT images, regardless of the shape of the lesion (e.g. oval, dumbbell-shaped, irregular). We recorded all cystic findings or tumors in the bilateral uterine adnexa measuring 1.5 cm or more on CT images, including functional cysts such as a follicle or luteal cyst in the ovary. Lesions or cystic findings smaller than 1.5 cm were recorded as 0 cm due to measurement ambiguity in young women where follicles 1 cm are common. Cystic lesions removed during surgery were sent for pathological examination. In premenopausal women, if functional cysts were suspected based on CT findings, resection was not always performed. In some cases of torsion, only surgical reduction was performed. The dataset includes both contrast-enhanced and non-contrast abdominal and pelvic CT scans. Each scan was annotated with a binary label indicating the presence or absence of adnexal torsion, resulting in a binary classification task. Each scan consists of approximately 50 to 200 axial slices, with a slice thickness ranging from 1 to 5 mm and an in-plane resolution of 512  \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times$$\end{document}  512 pixels. The Hounsfield Unit (HU) values roughly span from −1000 to 3000 To restrict the analysis to anatomically relevant regions, a depth-based cropping procedure was employed using TotalSegmentator 29 . This tool provides 3D segmentations of 104 anatomical structures based on the nnU-Net framework 30 . In this study, the left hip, right hip, and sacrum were selected as anatomical markers due to their consistent proximity to the uterus. Axial slices containing any of these structures were extracted, while those located outside this region were discarded. This step reduced the volume of irrelevant data and ensured that model training and inference focused on spatially consistent and clinically meaningful regions. In particular, it helped prevent the inclusion of upper abdominal or lower extremity slices that could introduce noise and increase computational cost. By limiting the input range to slices containing key pelvic structures, this approach preserved anatomical consistency across patients and enhanced the concentration of learning on regions likely to contain torsion-related features. Fig. 2 ( A ) Visibility enhancement and background masking based on centroid-derived windowing and morphological erosion. ( B ) Organ-focused CT image with enhanced visibility and background masking. ( A ) Visibility enhancement and background masking based on centroid-derived windowing and morphological erosion. ( B ) Organ-focused CT image with enhanced visibility and background masking. CT windowing is a technique used to enhance the visualization of specific anatomical structures by adjusting the displayed range of Hounsfield Unit (HU) values. It operates by setting a window center (WC) and window width (WW), which determine the HU range mapped to grayscale intensities. HU values outside this range are rendered as black or white, improving contrast and facilitating the interpretation of target tissues. To determine the optimal windowing parameters, we employed a histogram-based clustering approach. As illustrated in Fig.  2 B, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm 31 was applied to the HU histogram to identify distinct tissue density clusters. The centroids of the two largest clusters were extracted and used to define the window center (WC). The window width (WW) was optimized via grid search. A series of candidate values were systematically evaluated by applying windowing to the CT slices and computing the corresponding edge density. Edge density was defined as the ratio of edge pixels detected using the Canny edge detector to the total number of pixels. This metric served as a proxy objective function, based on the assumption that clearer anatomical boundaries result in sharper and denser edge maps. Among all candidate values, the window width corresponding to 20% of the HU value range produced the highest average edge density across samples and was selected for subsequent processing. Based on this configuration, windowing was performed using the lower cluster centroid as the reference intensity level. An erosion operation was then applied using the mean of all cluster centroids to suppress background noise and refine tissue boundaries. This visibility enhancement and masking process is illustrated in Fig.  2

Methods

All methods described in this study, including data collection, preprocessing, and model development, were performed in strict accordance with relevant institutional guidelines and regulations, including those approved by the Institutional Review Boards of Kyungpook National University Hospital (KNUH 2024-07-020-001) and Kyungpook National University Chilgok Hospital (KNUCH 2024-07-022-001) To classify adnexal cystic torsion from abdominal CT images, we implemented and evaluated two major categories of deep learning models: 2D multiple instance learning (MIL) 32 -based architectures and 3D volumetric models, including convolutional neural networks (CNNs) and Vision Transformers (ViTs). MIL is a weakly supervised learning framework in which labels are assigned at the bag level rather than the instance level, making it well suited for CT scans where only scan-level annotations are available and pathological signs, such as the whirl sign, appear in only a few slices. In this approach, each scan is treated as a bag \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$B = \{x_1, x_2, ..., x_n\}$$\end{document} of 2D axial slices \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$x_i$$\end{document} , and each slice is encoded into a feature vector \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$h_i = f_{\theta }(x_i)$$\end{document} using a shared convolutional encoder \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$f_{\theta }$$\end{document} . These feature vectors are subsequently aggregated into a bag-level representation through a pooling operation. Various pooling strategies can be employed for aggregation. In this study, we implemented several MIL variants based on this general structure, including: ABMIL 33 , which uses attention to compute weighted instance aggregations; CLAM 34 , which incorporates clustering constraints to enhance instance feature separability; DSMIL 35 , which introduces a dual-stream architecture for concurrent instance- and bag-level learning; TransMIL 36 , which employs Transformer-based global attention to model inter-slice dependencies; BayesMIL 37 , which introduces a Bayesian formulation to quantify uncertainty in instance contributions; and MRAN 38 , which enhances robustness to anatomical variability via multi-scale hierarchical attention. MRAN (Multi-scale Representation Attention Network) 38 introduces a hierarchical attention mechanism that aggregates features across multiple spatial scales in a weakly supervised learning setting. Originally designed for whole-slide image (WSI) classification, MRAN can be adapted to volumetric medical imaging by interpreting a 3D scan as a collection of spatially structured instances. The model constructs three levels of feature representations: cell-level, patch-level, and bag-level. First, each patch is divided into cell-level regions \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$z_{ijk} \in \mathbb {R}^d$$\end{document} , and attention weights are computed based on \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\ell _2$$\end{document} -norm and thresholding: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\eta _{ijk}' = \Vert z_{ijk} \Vert _2, \quad \eta _{ijk} = \frac{\max (\eta _{ijk}' - \tau _c, 0)}{\sum _{k'} \max (\eta _{ijk'}' - \tau _c, 0)}, \quad z^{\text {patch}}_{ij} = \sum _k \eta _{ijk} \cdot z_{ijk}$$\end{document} Next, the patch-level features are aggregated into a bag-level feature using a similar procedure: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\delta _{ij}' = \Vert z^{\text {patch}}_{ij} \Vert _2, \quad \delta _{ij} = \frac{\max (\delta _{ij}' - \tau _p, 0)}{\sum _{j'} \max (\delta _{ij'}' - \tau _p, 0)}, \quad z^{\text {bag}}_i = \sum _j \delta _{ij} \cdot z^{\text {patch}}_{ij}$$\end{document} Finally, all bag-level features are aggregated into a scan-level representation: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\gamma _i' = \Vert z^{\text {bag}}_i \Vert _2, \quad \gamma _i = \frac{\max (\gamma _i' - \tau _b, 0)}{\sum _{i'} \max (\gamma _{i'}' - \tau _b, 0)}, \quad z^{\text {final}} = \sum _i \gamma _i \cdot z^{\text {bag}}_i$$\end{document} This multi-level structure enables MRAN to emphasize salient features at fine (cell), intermediate (patch), and global (bag) scales. The model is jointly optimized using a multi-instance classification loss applied at all three levels, allowing it to improve performance and provide interpretable attention maps across spatial hierarchies. Such properties make MRAN particularly effective for detecting sparse, localized, and hierarchically structured patterns in complex medical data. These multi-level MIL architectures offer a principled way to emphasize salient patterns across multiple spatial resolutions under weak supervision. However, these models operate on discrete 2D slices or patches and inherently lack the capacity to model continuous volumetric context across adjacent slices. To overcome this limitation and more effectively capture 3D anatomical relationships, we investigated volumetric deep learning models that treat the entire CT scan as a unified 3D tensor. Unlike MIL-based approaches, which aggregate slice-level predictions, 3D convolutional neural networks (CNNs) apply volumetric kernels of size \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k \times k \times k$$\end{document} across spatial dimensions, thereby learning spatially coherent features in all three axes. We implemented 3D ResNet 39 , an extension of residual networks to volumetric inputs that facilitates the training of deep models through identity skip connections. In addition, we adopted 3D EfficientNet 40 , which applies a compound scaling strategy to balance network depth, width, and resolution for optimal performance. Each 3D EfficientNet block is constructed using the Mobile Inverted Bottleneck Convolution (MBConv) structure with the following components: A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$1 \times 1 \times 1$$\end{document} pointwise convolution for channel expansion, A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$3 \times 3 \times 3$$\end{document} depthwise 3D convolution applied separately to each channel, A Squeeze-and-Excitation (SE) module to recalibrate feature responses, and A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$1 \times 1 \times 1$$\end{document} projection layer to reduce channels back to the original dimension. The SE module performs channel-wise recalibration by first applying global average pooling, followed by two fully connected layers and a sigmoid activation: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$s = \sigma (W_2 \cdot \delta (W_1 \cdot \text {GAP}(x))), \quad x' = x \cdot s$$\end{document} where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\delta$$\end{document} is the ReLU activation, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma$$\end{document} is the sigmoid function, and GAP denotes global average pooling. The compound scaling method uniformly scales the network using a set of coefficients \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(\phi )$$\end{document} as follows: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {depth: } d = \alpha ^{\phi }, \quad \text {width: } w = \beta ^{\phi }, \quad \text {resolution: } r = \gamma ^{\phi } \quad \text {subject to: } \alpha \cdot \beta ^2 \cdot \gamma ^2 \approx 2, \quad \alpha , \beta , \gamma> 1$$\end{document} This ensures that model capacity is increased in a balanced way without unnecessary computational cost. We used the B6 variant of 3D EfficientNet, which provided the best trade-off between complexity and accuracy in our task. Its architecture enabled rich hierarchical representation learning over the full 3D CT volume, yielding the highest diagnostic performance across all models. A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$1 \times 1 \times 1$$\end{document} pointwise convolution for channel expansion, A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$3 \times 3 \times 3$$\end{document} depthwise 3D convolution applied separately to each channel, A Squeeze-and-Excitation (SE) module to recalibrate feature responses, and A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$1 \times 1 \times 1$$\end{document} projection layer to reduce channels back to the original dimension. Additionally, we investigated 3D Vision Transformers (ViTs) 41 , which extend the success of transformer-based architectures in natural language processing to volumetric medical imaging. Unlike CNNs that capture local patterns through convolutional kernels, ViTs operate on a sequence of non-overlapping 3D patches (or tokens) extracted from the input volume. Each patch is flattened and linearly projected into a token embedding, and positional encodings are added to retain spatial information across the 3D structure. Formally, given an input CT volume \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$x \in \mathbb {R}^{D \times H \times W \times C}$$\end{document} , it is divided into \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$N$$\end{document} non-overlapping 3D patches of size \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$P \times P \times P$$\end{document} . Each patch is then flattened and mapped to an embedding vector using a learnable projection matrix \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E \in \mathbb {R}^{(P^3 \cdot C) \times d}$$\end{document} , resulting in a sequence of patch embeddings \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\{z_1, z_2, ..., z_N\}$$\end{document} . A special classification token \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$[CLS]$$\end{document} is prepended to the sequence, and the full input is passed through \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$L$$\end{document} layers of multi-head self-attention and feed-forward networks. The self-attention mechanism in each layer allows every patch token to attend to all others, thereby capturing long-range dependencies across the entire volume: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {Attention}(Q,K,V) = \text {softmax}\left( \frac{QK^\top }{\sqrt{d_k}} \right) V$$\end{document} This global receptive field makes ViTs particularly appealing for modeling complex spatial relationships in 3D medical data, where distant regions may be diagnostically correlated. However, the application of ViTs to 3D CT volumes introduces several challenges. First, due to the cubic growth of token count with input resolution, computational and memory costs increase rapidly, making training difficult without substantial hardware resources or input downsampling. Second, unlike 2D ViTs pretrained on large-scale datasets like ImageNet, 3D ViTs currently lack widely available pretrained weights for volumetric medical data, making them prone to overfitting in limited-data regimes. Lastly, positional encoding in 3D is less mature, and the spatial continuity in volumetric data may not be effectively modeled by naïve extensions of 2D positional embeddings. In our experiments, the 3D ViT model underperformed relative to 3D CNN-based methods, likely due to these factors. Despite its theoretical strength in modeling the global context, its practical utility in CT-based diagnosis remains constrained by data efficiency and computational limitations. As noted previously, each patient had at least one image set and many had multiple sets depending on clinical conditions, for example, a non-enhanced phase or various contrast-enhanced phases. Each set of images was reconstructed into a single video image, which was then analyzed by the deep learning model. If the model classified the cystic adnexal torsion in any of the image sets of a patient, the patient was assigned to the torsion (disease) group. In contrast, if all image sets were classified as negative, the patient was assigned to the control group. Then these AI-based classifications were compared with the reference standard to assess diagnostic accuracy. The performance of individual image sets was not evaluated separately, as the model was designed for patient-wise interpretation 42 . The final classification results were aggregated to calculate diagnostic performance metrics. To provide a comprehensive evaluation of the model performance under clinical requirements, nine metrics were used: accuracy (ACC), area under receiver operating characteristic curve (AUC), sensitivity (SEN), specificity (SPE), positive predictive value (PPV), F1-score (F1), balanced accuracy (BAL_ACC), negative predictive value (NPV), and false positive rate (FPR). Let TP , TN , FP , and FN denote the number of true positives, true negatives, false positives, and false negatives, respectively. Accuracy (ACC) measures the proportion of correctly classified cases. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {ACC} = \frac{TP + TN}{TP + TN + FP + FN}$$\end{document} Sensitivity (SEN) (Recall) is the proportion of actual positive (disease) cases correctly identified: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {SEN} = \frac{TP}{TP + FN}$$\end{document} Specificity (SPE) represents the proportion of actual negative (control) cases correctly identified: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {SPE} = \frac{TN}{TN + FP}$$\end{document} Positive predictive value (PPV) quantifies the proportion of predicted positive cases that are true positives: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {PPV} = \frac{TP}{TP + FP}$$\end{document} F1-score (F1) is the harmonic mean of precision and sensitivity, providing a balanced measure of both: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {F1} = 2 \cdot \frac{\text {PPV} \cdot \text {SEN}}{\text {PPV} + \text {SEN}}$$\end{document} Balanced Accuracy (BAL_ACC) averages the recall obtained on each class, mitigating class imbalance effects: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {BAL\_ACC} = \frac{\text {SEN} + \text {SPE}}{2}$$\end{document} Negative Predictive Value (NPV) measures the proportion of negative predictions that are true negatives: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {NPV} = \frac{TN}{TN + FN}$$\end{document} False Positive Rate (FPR) is the proportion of negative cases incorrectly classified as positive: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {FPR} = \frac{FP}{FP + TN}$$\end{document} The area under curve (AUC) is defined as the area under the receiver operating characteristic (ROC) curve, which plots the true positive rate (TPR) against the false positive rate (FPR) across thresholds: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {AUC} = \int _0^1 \text {TPR}(t)\, d\text {FPR}(t)$$\end{document} where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {TPR}(t) = \frac{TP}{TP + FN}$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {FPR}(t) = \frac{FP}{FP + TN}$$\end{document} vary with the decision threshold t . To ensure a robust evaluation of model performance, we employed stratified 5-fold cross-validation. The dataset was divided at the CT volume. Each fold contained approximately 80% of the data for training and 20% for validation, maintaining class balance within each subset. All models were trained five times, using a different fold as the validation set in each iteration and the remaining four folds for training. Performance metrics, including AUC, ACC, SEN, and SPE, were calculated for each fold and then averaged to produce the final reported results. This strategy reduces variance and provides a more reliable estimate of model generalizability among different patient populations. For the analysis of continuous parameters, the Student t test was used to compare the mean value. For the analysis of categorical parameters, the Chi-square test, Fisher’s exact test. All statistical significances were estimated using proportions with 95% confidence intervals. A p-value less than 0.05 was considered statistically significant. All statistical analyses were performed using Python’s scikit-learn library (version 1.6.1).

Related

Deep learning, particularly convolutional neural networks (CNNs), has revolutionized medical image analysis by automatically learning hierarchical features from raw imaging data. Its ability to outperform traditional handcrafted feature-based methods has led to widespread adoption across multiple imaging modalities, including CT, MRI, and ultrasound 14 – 16 . In thoracic imaging, deep learning models have been effectively applied to detect pneumonia 9 , tuberculosis 13 , and pulmonary nodules 10 on chest radiographs, achieving diagnostic performance comparable to that of radiologists. In cardiovascular imaging, CNN-based models have been employed to identify coronary artery disease and myocardial infarction from cardiac CT and MRI scans 17 , 18 . In neuroimaging, deep learning approaches have been utilized for automated stroke detection, intracranial hemorrhage identification, and brain tumor classification 19 , 20 . Similarly, in abdominal imaging, deep learning models have demonstrated high sensitivity and specificity in detecting liver metastases and other acute abdominal conditions 16 , 20 . The choice of deep learning methodology has been primarily determined by imaging characteristics and diagnostic requirements. For two-dimensional imaging modalities such as chest X-rays and ultrasound, two-dimensional convolutional neural networks (CNNs) have been commonly used due to their computational efficiency and robust performance. In contrast, volumetric modalities like CT and MRI provide spatial continuity across slices, where three-dimensional CNNs are preferred to capture inter-slice dependencies and maintain anatomical context, which is critical for accurate interpretation of complex structures 6 , 7 . This ability to model volumetric features makes 3D CNNs particularly suitable for abdominal CT imaging. When explicit slice-level annotations are unavailable in volumetric data, Multiple Instance Learning (MIL) frameworks provide an alternative perspective by treating a scan as a collection of instances rather than a single holistic entity. MIL approaches have demonstrated success across various medical imaging tasks, such as hemorrhage detection from CT scans 21 , histopathology whole-slide image classification 22 , and diabetic retinopathy grading 23 . In particular, Chang et al 24 . proposed a deep MIL framework that leveraged pre-trained convolutional features from CT slices, aggregated through an attention-based pooling mechanism, to predict chemotherapy response in non-small cell lung cancer. Their work highlights the potential of MIL to effectively learn from weakly labeled volumetric data without requiring extensive manual annotation. More recently, transformer-based architectures have been introduced into medical imaging for their ability to model long-range dependencies. Vision Transformers (ViTs), in particular, have demonstrated strong performance across various tasks, including segmentation 25 , registration 26 , and classification 27 , 28 . Motivated by the demonstrated success of CNNs, MIL frameworks, and transformer-based models across diverse medical imaging applications, we incorporated both 2D MIL architectures and 3D volumetric deep learning frameworks in this study to evaluate their feasibility for classifying adnexal cystic torsion using abdominal CT images.

Results

In this study, a total of 514 women underwent abdominopelvic CT examination were included. The baseline characteristics of the patient are summarized in Table 1 . The disease group consisted of 259 cases (50.4%), while the control group included 255 cases (49.6%). Patients in the disease group were significantly older and their CT examinations were performed significantly earlier than those in the control group. There were no significant differences between the groups regarding the type of CT studies: contrast enhanced, non-contrast, or both. In the disease group, cystic findings were observed in the right adnexa in 195 cases (75.3%), the left adnexa in 150 cases (57.9%), and bilaterally in 86 cases (33.2%). In the control group, cystic findings were present in the right adnexa in 177 cases (69.4%), the left adnexa in 144 cases (56.5%), and bilaterally in 66 cases (25.9%). Although these distributions were not statistically different between the groups, the mean size of cystic lesions on each side was significantly larger in the disease group. Fig. 3 A and B showed the histolocial distribution of all cystic findings or lesions in the bilateral uterine adnexa shown on abdominopelvic CT scans in the disease group or the control group in this study. The clinical characteristics of the disease group are also presented in Table 1 . Among the 259 patients diagnosed with adnexal cystic torsion in the surgical field, 229 (88.4%) had ovarian cystic tumors and 30 (11. 6%) had extraovarian cystic lesions such as para-ovarian cysts, para-tubal cysts, or hydrosalpinx. Torsion occurred in the right adnexa in 155 cases (59.8%) and in the left adnexa in 104 cases (40.2%). The mean maximum diameter of the torsed cystic tumors on CT was 9.80 ± 3.73 cm. Histologically, the majority of cystic lesions (229 cases, 88.4%) were benign cystic tumors. Functional cysts, such as follicular or corpus luteal cysts, represented 13 cases (5.0%), and 17 cases (6.5%) were borderline or malignant tumors, regardless of histological subtypes. The most common single histological diagnosis was mature cystic teratoma, found in 72 cases (27.8%). In 34 cases (13.1%), a definitive pathological diagnosis could not be made due to an extensive infarction resulting from the torsion. The histological distribution of the adnexal tumors is shown in Fig. 3 C. Fig. 3 Pathologic diagnoses of adnexal cystic findings or torsed adnexal tumors in this study. ( A ) Cystic findings in the left adnexa shown on abdominopelvic CT. ( B ) Cystic findings in the right adnexa shown on abdominopelvic CT. ( C ) Torsed adnexal tumors proven surgically. Table 1 Basal characteristics of the patients in this study. Disease group (n=259) Control group (n=255) p -value Age (yrs) 43.03 ± 18.72 32.82 ± 11.81 <0.001 Features of abdominopelvic CT examinations Time interval from CT to surgery (d) 8.83 ± 15.92 1.86 ± 4.89 <0.001 Presence of non-contrast study (n) 251 (96.9%) 252 (98.8%) 0.222 Presence of contrast-enhanced study (n) 241 (93.1%) 224 (87.8%) 0.051 Presence of both studies (n) 233 (90.0%) 221 (86.7%) 0.273 Cystic findings shown in the CT * Presence in right side (n) 195 (75.3%) 177 (69.4%) 0.147 Size in right side (cm) 6.40 ± 4.56 3.57 ± 4.22 <0.001 Presence in left side (n) 150 (57.9%) 144 (56.5%) 0.789 Size in left side (cm) 4.73 ± 5.26 3.08 ± 4.18 <0.001 Presence in both sides (n) 86 (33.2%) 66 (25.9%) 0.082 Clinical features of torsed adnexal tumors Laterality (n) In the right adnexa 155 (59.8%) N/A \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^\dag$$\end{document} N/A In the left adnexa 104 (40.2%) N/A N/A Maximal diameter (cm) 9.80 ± 3.73 N/A N/A Pathologic diagnosis (n) Mature cystic teratoma 72 (27.8%) N/A N/A Unknown due to extensive infarction 34 (13.1%) N/A N/A Mucinous cystadenoma 27 (10.4%) N/A N/A Serous cystadenoma 18 (6.9%) N/A N/A No pathologic examination \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^\dag$$\end{document} 18 (6.9%) N/A N/A Borderline or malignant tumors 17 (6.5%) N/A N/A Functional cysts 16 (6.2%) N/A N/A Other benign tumors \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\S }$$\end{document} 57 (22.0%) N/A N/A Note: Data are expressed as numbers (%) and means ± standard deviations. Statistical significance was evaluated with student t test or chi-square test. * All adnexal cystic findings or tumors in the bilateral uterine adnexa, measuring 1.5 cm or more on transverse CT images. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\dag }$$\end{document} Not applicable. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\ddag }$$\end{document} The cases the surgeon omitted surgical removal of the cystic lesion according to medical conditions. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\S }$$\end{document} The histological distribution was presented in Figure 3 . Abbreviations: CT, computed tomography. Pathologic diagnoses of adnexal cystic findings or torsed adnexal tumors in this study. ( A ) Cystic findings in the left adnexa shown on abdominopelvic CT. ( B ) Cystic findings in the right adnexa shown on abdominopelvic CT. ( C ) Torsed adnexal tumors proven surgically. Basal characteristics of the patients in this study. Note: Data are expressed as numbers (%) and means ± standard deviations. Statistical significance was evaluated with student t test or chi-square test. * All adnexal cystic findings or tumors in the bilateral uterine adnexa, measuring 1.5 cm or more on transverse CT images. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\dag }$$\end{document} Not applicable. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\ddag }$$\end{document} The cases the surgeon omitted surgical removal of the cystic lesion according to medical conditions. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\S }$$\end{document} The histological distribution was presented in Figure 3 . Abbreviations: CT, computed tomography. Table  2 and Fig.  4 summarize the image-wise and patient-wise classification performance of 2D MIL-based and 3D volumetric models for diagnosing adnexal cystic torsion from abdominal CT scans. Among all evaluated models, 3D EfficientNet achieved the highest image-wise performance, with an AUC of 0.7925, ACC of 0.7340, SEN of 0.7306, and SPE of 0.7377, demonstrating balanced detection of positive and negative cases. In comparison, other 3D models such as 3D ResNet and 3D ViT showed limited performance, with AUCs of 0.3964 and 0.6616, respectively. Within 2D MIL-based models, MRAN outperformed others with an AUC of 0.7387, ACC of 0.6843, SEN of 0.6258, and SPE of 0.7482. While models like BayesMIL and ABMIL were also competitive, MRAN consistently performed best across key metrics. Simpler pooling-based approaches (e.g., Max Pooling, Average Pooling) yielded respectable AUCs of 0.7076 and 0.6729. Notably, ABMIL achieved near-perfect sensitivity (0.9968) but very low specificity (0.0123), indicating over-sensitivity, whereas CLAM prioritized specificity (0.6866) at the cost of modest sensitivity (0.6403). Under patient-wise evaluation, where a patient was considered positive if any scan was predicted positive, 3D EfficientNet again delivered the strongest results (AUC: 0.8015, ACC: 0.7417, SEN: 0.7408, SPE: 0.7923), exceeding its scan-level performance and demonstrating robustness in aggregated predictions. MRAN remained the top 2D model in this setting as well, showing balanced sensitivity and specificity, and maintaining consistent performance across evaluation paradigms. These patient-wise results further support the clinical relevance of the top-performing models. Performance generally improved when predictions were aggregated per patient, suggesting that multiple scan-level predictions provide synergistic diagnostic value. The complementary strengths of 3D EfficientNet and MRAN highlight their potential for integration in future hybrid or ensemble AI-based diagnostic frameworks. Table 2 Patient-wise classification performance for various models. A. Individual model results. B. Comparison between 2D MIL-based and 3D volumetric-based approaches. A Model AUC ACC SEN SPE PPV F1 BAL_ACC NPV FPR 3D EfficientNet 0.8015 0.7417 0.7923 0.6892 0.7254 0.7574 0.7408 0.7621 0.3108 MRAN 0.7348 0.6888 0.6885 0.6892 0.6965 0.6925 0.6889 0.6811 0.3108 BayesMIL 0.7132 0.6810 0.8115 0.5458 0.6492 0.7214 0.6787 0.7366 0.4542 ABMIL 0.7118 0.5108 1.0000 0.0040 0.5098 0.6753 0.5020 1.0000 0.9960 Max pool 0.7027 0.6301 0.7308 0.5259 0.6149 0.6678 0.6283 0.6535 0.4741 CLAM 0.6806 0.6595 0.7115 0.6056 0.6514 0.6801 0.6586 0.6696 0.3944 Avg pool 0.6673 0.6047 0.6885 0.5179 0.5967 0.6393 0.6032 0.6161 0.4821 3D ResNet 0.3964 0.4951 0.6885 0.2948 0.5028 0.5812 0.4916 0.4774 0.7052 DSMIL 0.3380 0.5040 0.5423 0.4496 0.5204 0.5078 0.4960 0.4872 0.5504 B Model Type Model AUC ACC SEN SPE PPV F1 BAL_ACC NPV FPR 2D Avg pool 0.6729 0.6153 0.6000 0.6320 0.6403 0.6195 0.6160 0.5914 0.3680 Max pool 0.7076 0.6448 0.6532 0.6356 0.6618 0.6575 0.6444 0.6267 0.3644 ABMIL 0.7263 0.5261 0.9968 0.0123 0.5242 0.6870 0.5045 0.7778 0.9877 CLAM 0.6979 0.6625 0.6403 0.6866 0.6904 0.6644 0.6635 0.6362 0.3134 DSMIL 0.4570 0.4630 0.5210 0.3996 0.4864 0.5031 0.4603 0.4332 0.6004 TransMIL 0.6618 0.6481 0.7097 0.5810 0.6490 0.6780 0.6453 0.6471 0.4190 BayesMIL 0.7170 0.6911 0.7500 0.6268 0.6869 0.7170 0.6884 0.6967 0.3732 MRAN 0.7387 0.6843 0.6258 0.7482 0.7307 0.6742 0.6870 0.6469 0.2518 3D 3D ResNet 0.3964 0.4951 0.4916 0.6885 0.2948 0.5028 0.5812 0.4774 0.7052 3D ViT 0.6616 0.5896 0.5403 0.6433 0.6164 0.5538 0.5918 0.5816 0.3567 3D EfficientNet 0.7925 0.7340 0.7306 0.7377 0.7525 0.7414 0.7342 0.7150 0.2623 Patient-wise classification performance for various models. A. Individual model results. B. Comparison between 2D MIL-based and 3D volumetric-based approaches. Fig. 4 Diagnostic performance of 3D EfficientNet under image-wise and patient-wise evaluations. ( A , B ) show image-wise confusion matrices and ROC curves computed per scan. ( C , D ) present patient-wise results, where patients are considered positive if any scan is predicted positive. Diagnostic performance of 3D EfficientNet under image-wise and patient-wise evaluations. ( A , B ) show image-wise confusion matrices and ROC curves computed per scan. ( C , D ) present patient-wise results, where patients are considered positive if any scan is predicted positive. To analyze the effectiveness of preprocessing, we conducted an ablation study. In Deep learning CT image analysis, preprocessing is one of the most important factors that affect model performance. In this section, we quantitatively validated the effect of the preprocessing method. The experiment used the 3D EfficientNet model, which achieved the best performance, and the input data was divided into two conditions for comparison. The first used the original CT images without any preprocessing, and the second used data processed by DBSCAN HU histogram windowing and pixel scaling to remove noise. As shown in Table  3 , the application of HU windowing improved all major evaluation metrics (AUC, ACC, SEN, SPE). Notably, the improvements in AUC and sensitivity were prominent, indicating that the preprocessing method helped clarify the visual features of lesions and guided the model to better learn their structural characteristics. Table 3 Effect of DBSCAN windowing and pixel scaling preprocessing on 3D EfficientNet performance. Setting AUC ACC SEN SPE Original (without preprocessing) 0.5967 0.5966 0.6370 0.5526 Proposed (with preprocessing) 0.7967 0.7321 0.7300 0.7343 Effect of DBSCAN windowing and pixel scaling preprocessing on 3D EfficientNet performance. To comprehensively assess model performance, we conducted evaluation at both the image level and the patient level. In the image-level evaluation, each CT volume was treated independently, allowing detailed analysis of the model’s response to specific cases. While this method is valuable for technical benchmarking, it may not fully represent clinical workflows. Alternatively, the patient-level evaluation classifies a patient as positive if any of their CT scans are predicted positive. This approach more accurately reflects real-world clinical decision-making, where identifying any sign of torsion warrants immediate action. Notably, most models showed improved AUC and specificity under patient-level evaluation. For instance, 3D EfficientNet achieved a higher AUC of 0.8015 and specificity of 0.7923 in the patient-level setting, compared to 0.7925 and 0.7377 at the image level. 3D ResNet adopts a conventional architecture with stacked 3D convolutional layers and residual connections. However, it lacks components that model inter-channel dependencies or contextual relationships, which limits its ability to capture complex anatomical features. This limitation was reflected in its poor performance, especially its low specificity (0.3032), which may limit its clinical applicability in scenarios that require reliable exclusion of non-torsion cases. 3D ViT, based on the Vision Transformer architecture, showed moderate performance (AUC: 0.6616) but suffered from unstable training and overfitting–especially in the absence of large-scale pretraining. Consistent with prior findings 43 , transformer-based models may face challenges in small-scale medical imaging tasks without architectural adaptation or data augmentation. In contrast, 3D EfficientNet demonstrated the highest classification performance, achieving an AUC of 0.7925, ACC of 0.7340, and SEN of 0.7306 at the image level. Its compound scaling strategy and efficient depth-resolution configuration enabled robust feature extraction across entire CT volumes. The model’s strong sensitivity and specificity balance is particularly critical in emergency scenarios requiring timely and accurate diagnosis. Its advantage was even more pronounced in the patient-level analysis, with improved metrics across the board. Among the 2D MIL-based models, MRAN emerged as the top performer, achieving an image-level AUC of 0.7387 and showing consistent results in the patient-wise evaluation. MRAN incorporates a hierarchical attention strategy that subdivides CT slices into smaller patches before feature encoding. This enables fine-grained localization of torsion-specific features–such as the whirl sign–that may only appear in a few slices. Its performance highlights the strength of region-aware learning in 2D-based analysis. Other MIL models, such as BayesMIL and ABMIL, also delivered competitive results, especially when paired with adaptive preprocessing strategies. Even simpler pooling-based models demonstrated reasonable performance, although they generally underperformed compared to 3D models. In summary, our results show that both 3D volumetric CNNs and 2D MIL frameworks are viable tools for the diagnosis of adnexal torsion from abdominal CT scans. While 3D EfficientNet offers superior volumetric comprehension, MRAN provides efficient regional focus with minimal computational overhead. These complementary strengths suggest that future work may benefit from hybrid or ensemble modeling strategies to combine the advantages of both paradigms.

Discussion

The AI algorithm developed in this study demonstrated the feasibility of detecting adnexal cystic torsion using abdominopelvic CT. The algorithm achieved a SEN of 79.23%, SPE of 68.92%, PPV of 72.54%, NPV of 76.21%, and an ACC of 74.17%, with an AUC of 0.80 in patient-wise evaluation–comparable to image-wise diagnostic performance. According to a recent meta-analysis, human investigators diagnosed adnexal torsion with approximately 79% sensitivity and 76% specificity using ultrasound and 81% sensitivity and 91% specificity using MRI 44 . Taking these reference values into account, the AI algorithm in our study shows promising potential in gynecologic emergency settings. To our knowledge, this is among the first studies to develop and validate a CT-based deep learning framework for detecting adnexal cystic torsion. The AI model was developed using a 3D extension of the EfficientNet architecture and was trained on retrospectively collected CT scans with adaptive Hounsfield unit windowing and organ-level cropping to emphasize clinically relevant regions. Technically, the framework was tailored to the volumetric nature of CT data by integrating a 3D convolutional backbone and a preprocessing pipeline that dynamically adjusts HU windows and filters anatomically relevant structures. This design minimized inter-scan heterogeneity and improved the model’s generalizability. Through extensive empirical evaluation, we identified a robust combination of preprocessing and architecture–specifically, adaptive HU windowing, organ-level cropping, and a lightweight 3D CNN–that consistently delivered high performance despite the absence of slice- or region-level annotations. This experimentally validated setup demonstrates that thoughtful design choices can partially compensate for weak supervision and provides a strong baseline for future studies incorporating more granular annotations. The development of the AI algorithm in this study faced several challenges and limitations. One major issue was the heterogeneity of CT scans, including variations in scan axis, timing, slice thickness, imaging phase, and use of contrast agents. Due to the lack of coronal and sagittal reconstructions in many cases, we used only transverse CT images–both contrast-enhanced and non-enhanced. To develop a more robust algorithm, a larger dataset with multi-planar reconstructions is essential. In particular, certain signs, such as the ’whirl sign’, may be more apparent in coronal or sagittal views than in transverse images 45 . Standardizing imaging protocols, particularly with respect to contrast phases, would enhance diagnostic consistency and AI reliability. Anatomical variations among patients also presented a challenge. The degree of torsion varied, ranging from approximately 360 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\circ }$$\end{document} to more than 720 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^{\circ }$$\end{document} , as confirmed during surgery. Furthermore, the algorithm did not control for variations in uterine or cyst size. Conditions such as enlarged uteri, multiple leiomyomas, or extensive adnexal cystic tumors can obscure the twisted pedicle. Some patients had undergone hysterectomy, making identification of the twisted pedicle more difficult due to the absence of bilateral round ligaments. Another limitation was the inclusion of CT examinations performed up to 90 days before surgery. This was necessary to increase the size of the data set, but may have introduced variability in disease presentation. In several cases, a suspicious whirl sign was visible on CT 60 to 90 days before surgery; in others, no definitive signs were present, complicating the validation of the algorithm detection ability. To improve the reliability of CT-based AI detection of adnexal torsion, future studies should control clinical variables more rigorously and address imaging artifacts to improve quality and reduce exclusions. In addition, performance could be improved by training the algorithm with annotated images showing specific diagnostic signs, marked by experienced gynecologists or radiologists. However, this introduces another challenge. A previous study identified four key CT indicators of adnexal torsion 46 : ipsilateral twisted pedicle (whirl sign), thickened fallopian tube, central afollicular stroma, and abnormal ovarian enhancement. Among these, the twisted pedicle is pathognomonic and can sometimes be detected by ultrasound 2 . Although we attempted to annotate these features in our dataset, the number of clearly identifiable cases was too small for effective training. In addition, currently there are no standard criteria available to define these signs. With a sufficient number of annotated cases, future studies can train algorithms to detect specific or nonspecific features or even develop scoring systems based on AI output. In this study, we present four representative cases with transverse CT and laparoscopic images, showing clear whirl signs (Fig.  5 ). Fig. 5 Four cases of their transverse CT images and the corresponding laparoscopic images, exhibiting the whirl sign on their transverse CT scans clearly (arrow head: whirl sign or twisted pedicle, arrow: ovarian tumor in the torsed side). Four cases of their transverse CT images and the corresponding laparoscopic images, exhibiting the whirl sign on their transverse CT scans clearly (arrow head: whirl sign or twisted pedicle, arrow: ovarian tumor in the torsed side). Multiple–instance learning (MIL) approaches treat each CT volume as a bag of independent 2D axial slices. Although MIL is attractive when slice-level labels are unavailable, this design inherently discards volumetric continuity: adjacent slices that together depict the twisted pedicle are presented to the network as unordered, context-free instances. As a result, inter–slice anatomical relationships e.g. the helical trajectory of the ’whirl sign’ across several millimetres cannot be faithfully captured. In our experiments, the best MIL variant (MRAN) reached an AUC of 0.74, noticeably lower than the fully volumetric 3D EfficientNet (AUC 0.80, Table  2 ), suggesting a performance ceiling imposed by this 2D decomposition. Future extensions may therefore require hybrid strategies that embed 3D positional encodings or token-level transformers to restore lost spatial coherence. All architectures slice based MIL as well as 3D CNN and ViT models were trained with binary, scan level labels (torsion present vs. absent). Without voxel or slice-level supervision, the networks must infer the ’whirl sign’ location implicitly from weak global gradients. This weak supervision is problematic for two reasons. First, the ’whirl sign’ occupies only a minute fraction of the volume; gradient signals are therefore highly diluted, yielding diffuse or mis-localised attention maps. Second, negative scans may still contain swirling vascular patterns unrelated to torsion, increasing the false-positive risk. Consequently, even our 3D models despite processing entire volumes occasionally misclassify cases in which the twisted pedicle is subtle or obscured by cyst burden. High-quality regional annotations, or at least coarse bounding boxes for a subset of cases, appear indispensable for teaching the algorithm where to “look” and for enabling clinically interpretable saliency maps. Ultrasound remains the first-line imaging modality for suspected adnexal cystic torsion, particularly in women of reproductive age, given its lack of ionizing radiation, accessibility, and established diagnostic performance. A recent meta-analysis reported pooled sensitivity and specificity of 0.79 (95% CI = 0.63–0.92) and 0.76 (95% CI = 0.50–0.93), respectively, for ultrasound. MRI demonstrated higher pooled sensitivity and specificity of 0.81 (95% CI = 0.63–0.91) and 0.91 (95% CI = 0.80–0.96) 44 . These findings underscore the central role of ultrasound and MRI in the diagnostic evaluation of adnexal cystic torsion. In our cohort, preoperative ultrasound was performed in 250 patients (96.5%) in the disease group, reflecting adherence to standard clinical practice. CT examinations were primarily obtained in emergency settings for evaluation of severe abdominal or flank pain with broad differential diagnoses, rather than as a dedicated first-line test for suspected torsion. In such acute-care workflows, CT is frequently performed to exclude other urgent gynecologic, gastrointestinal, urinary, or vascular conditions. Within this real-world context, the aim of the present study was not to replace established imaging pathways, but to investigate whether AI may assist in identifying adnexal cystic torsion on abdominopelvic CT scans that have already been acquired during routine clinical care. In this context, AI-based assistance may enhance interpretative confidence and reduce diagnostic oversight on CT, thereby complementing–but not replacing–established first-line imaging strategies. In this study, we evaluated diagnostic accuracy on a per patient basis rather than assessing each individual image set obtained from a single patient. Specifically, if the deep learning model identified an adnexal cystic torsion in at least one of the patient’s image sets, the patient was classified into the torsion group; if all the image sets were classified as negative, the patient was placed in the control group. A recent study reported that many deep learning models currently under development or already published do not adopt patient-wise separation of training and test datasets, to maximize accuracy scores 42 . This practice may lead to inaccurate accuracy estimates and compromise the applicability of the model in a clinical setting in the real world 42 . As these investigators noted, diagnoses using imaging modalities such as CT are typically made through the integrated analysis of multiple imaging phases in clinical practice. However, for the diagnosis of adnexal cystic torsion–the primary objective of this study–the diagnostic accuracy associated with specific individual imaging sequences has not yet been clearly established. This is a significant issue, as it directly affects the identification of the most effective imaging-based diagnostic approach for adnexal cystic torsion, a condition that remains challenging to diagnose. Therefore, more studies are needed to establish a larger dataset including a wider range of image sequences and reassess the diagnostic accuracy accordingly. In this study, we demonstrated the feasibility of developing an AI model to detect adnexal cystic torsion using abdominopelvic CT scans. Considering the diagnostic accuracy achieved by human experts in identifying this condition, the AI model developed in this study showed promising potential for application in gynecologic emergency settings. From a technical perspective, the proposed framework was designed to account for the volumetric nature of CT data by employing a 3D CNN and incorporating a preprocessing pipeline that adaptively adjusts the HU window and filters the anatomically relevant regions. This integration improved the robustness of the model by reducing input heterogeneity and emphasizing clinically meaningful features. However, further improvement is necessary to develop a more accurate and reliable model suitable for real-world clinical practice. A critical step toward this goal involves the construction of a more homogeneous dataset by minimizing variability in CT protocols and anatomical presentation. Future studies should also focus on training AI models with individually annotated CT images, particularly highlighting the presence of the ’whirl sign’, to improve interpretability, diagnostic sensitivity, and explainability.

Introduction

Adnexal cystic torsion is a gynecological emergency caused by twisting of vascular structures within the infundibulopelvic ligament, which connects the ovary to the ipsilateral pelvic wall. Patients with adnexal torsion usually present with acute pelvic or abdominal pain; however, clinical symptoms can be nonspecific. Some patients report vague discomfort rather than sharp pain, and others may experience systemic symptoms such as fever or vomiting 1 . Torsion leads to impaired venous and lymphatic outflow, resulting in ovarian edema or enlargement. This pathology is often precipitated by adnexal cystic lesions, such as cystic tumors of the ovary or solid masses. As the condition progresses, the arterial blood supply may become compromised, leading to thrombosis, ischemia, and hemorrhagic infarction 2 . In severe cases, prolonged torsion can cause irreversible ovarian damage, which can result in subfertility or premature menopause 3 . According to previous studies, ovarian torsion, a common form of adnexal torsion, accounts for approximately 2 to 3% of gynecological emergencies 2 . On imaging, the twisted vascular pedicle, referred to as the ’whirl sign’, is considered pathognomonic for adnexal torsion when visualized 2 . However, it appears in fewer than one third of cases on computed tomography (CT) or magnetic resonance imaging (MRI), which, despite their objectivity, still present diagnostic challenges 4 , 5 . Therefore, rapid and accurate detection of this subtle but critical sign is essential in improving patient outcomes. Vision-based deep learning algorithms have shown significant potential in medical image analysis, achieving high diagnostic performance in various imaging modalities, including CT, MRI, and ultrasound 6 – 8 . In particular, for pathologies such as pneumonia, tuberculosis, and pulmonary nodules, deep learning based approaches have demonstrated diagnostic accuracy comparable to that of radiologists 9 – 12 . Furthermore, several studies have highlighted the applicability of deep learning in time-sensitive clinical settings, including emergency department imaging 2 , 13 . Despite its numerous applications in medical imaging, the application of deep learning for diagnosing adnexal cystic torsion remains unexplored. In this study, our objective was to investigate the feasibility of a deep learning model to detect adnexal cystic torsion on abdominopelvic CT images. Thus, we evaluated diagnostic accuracy in various models and determined which one achieved the best performance at the patient level.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

SciLite annotations

organisms 5
noordeloos 2009062 noordeloos 2009062 human noordeloos 2009062 human

Source provenance

europepmc
last seen: 2026-09-13T09:25:22.628771+00:00
scilite
last seen: 2026-09-06T10:05:09.034756+00:00
unpaywall
last seen: 2026-09-07T06:27:18.705824+00:00
License: CC-BY-NC-ND-4.0