Background
Lack of anatomy recognition represents a clinically relevant risk factor in abdominal 35
surgery. While machine learning methods have the potential to aid in recognition of visible patterns 36
and structures, limited availability and diversity of (annotated) laparoscopic image data restrict the 37
clinical potential of such applications in practice. This study explores the potential of machine 38
learning algorithms to identify and delineate abdominal organs and anatomical structures using a 39
robust and comprehensive dataset, and compares algorithm performance to that of humans. 40
Methods
Based on the Dresden Surgical Anatomy Dataset providing 13195 laparoscopic images 41
with pixel-wise segmentations of eleven anatomical structures, two machine learning algorithms 42
were developed: individual segmentation algorithms for each structure, and a combined algorithm 43
with a common encoder and structure -specific decoders. Performance was assessed using F1 44
score, Intersection-over-Union ( IoU), prec ision, recall, and specificity. Using the example of 45
pancreas segmentation on a sample dataset of 35 images, algorithm performance was compared 46
to that of a cohort of 28 physicians, medical students, and medical laypersons. 47
Results
Mean IoU for segmentation of intraabdominal structures ranged from 0.28 to 0.83 and 48
from 0.22 to 0.78 for the structure -specific and the combined s emantic segmentation model, 49
respectively. Average inference for the structure -specific (one anatomical structure) and the 50
combined model (eleven anatomical structures) took 28 ms and 71 ms, respectively. Both models 51
outperformed 26 out of 28 human participants in pancreas segmentation. 52
Conclusions
Machine learning methods have the potential to provide relevant assistance in 53
anatomy recognition in minimally -invasive surgery in near -real-time. Future research should 54
investigate the educational value and subsequent clinical impact of respective assistance 55
systems. 56
57
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
3
Introduction
58
Computer vision describes the computerized analysis of digital images aiming at the automation 59
of human visual capabilities, most commonly using machine learning methods, in particular deep 60
learning. This approach has transformed medicine in recent years, with successful applications 61
including computer-aided diagnosis of colonic polyp dignity in endoscopy1,2, detection of clinically 62
actionable genetic alterations in histopathology 3, and melanoma detection in dermatology 4. 63
Availability of large amounts of training data is the defining prerequisite for successful application 64
of deep learning methods. With the establishment of laparoscopy as the gold standard for a variety 65
of surgical procedures 5–8 and the increasing availability of computing resources, these concepts 66
have gradually been applied to abdominal surgery. The overwhelming majority of research efforts 67
in the field of Artificial Intelligence (AI)-based analysis of intraoperative surgical imaging data (i.e. 68
video data from laparoscopic or open surgeries) has focused on classifying im ages with respect 69
to the presence and/or location of previously annotated surgical instruments or anatomical 70
structures9–13 or on analysis of surgical proficiency 14–16 based on recorded procedures. However, 71
almost all research endeavors in the field of computer vision in laparoscopic surgery have 72
concentrated on preclinical stages and to date, no AI model based on intraoperative surgical 73
imaging data could demonstrate a palpable clinical benefit.17 Among the studies closest to clinical 74
application are recent works on identification of instruments and hepatobiliary anatomy during 75
cholecystectomy for automated assessment of the critical view of safety 13, and on the automated 76
segmentation of safe and unsafe preparation zones during cholecystectomy18. 77
In surgery, patient outcome heavily depends on experience and performance of the surgical 78
team.19,20 In a recent analysis of Human Perform ance Deficiencies in major cardiothoracic, 79
vascular, abdominal transplant, surgical oncology, acute care, and general surgical operations , 80
more than half of the cases with postoperative complications were associated with identifiable 81
human error. Among these errors, lack of recognition (including misidentified anatomy) accounted 82
for 18.8%, making it the most common Human Performance Deficiency overall.21 While AI-based 83
systems identifying anatomical risk and target structures would theoretically have the potential to 84
alleviate this risk, limited availability and diversity of (annotated) laparoscopic image data 85
drastically restrict the clinical potential of such applications in practice. 86
To advance and diversify the applications of computer vision in laparoscopic surgery, we have 87
recently published the Dresden Surgical Anatomy Dataset22, providing 13195 laparoscopic images 88
with high-quality annotations of the presence and exact location of eleven intraabdominal 89
anatomical structures : abdominal wall, colon, intestinal vessels (inferior mesenteric artery and 90
inferior mesenteric vein w ith their subsidiary vessels), liver, pancreas, small intestine, spleen, 91
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
4
stomach, ureter and vesicular gland s. Here, we present the first study evaluating automated 92
detection and localization of organs and anatomical structures in laparoscopic view based on this 93
dataset, and, using the example of delineation of the pancreas, compare algorithm performance 94
to that of humans. 95
96
Methods
97
Patient cohort 98
Video data from 32 robot -assisted anterior rectal resections or rectal extirpations were gathered 99
at the University Hospital Carl Gustav Carus Dresden between February 2019 and February 2021. 100
All included patients had a clinical indication for the surgical procedure , recommended by an 101
interdisciplinary tumor board . The procedures were performed using the da V inci® Xi system 102
(Intuitive Surgical, Sunnyvale, CA, USA) with a standard Da Vinci® Xi/X Endoscope with Camera 103
(8 mm diameter, 30˚ angle, Intuitive Surgical, Sunnyvale, CA, USA, Item code 470057). Surgeries 104
were recorded using the CAST system (Orpheus Medical GmBH, Frankfurt a.M., Germany). Each 105
record was saved at a resolution of 1920 x 1080 pixels in MPEG-4 format. 106
All experiments were performed in accord ance with the ethical standards of the Declaration of 107
Helsinki and its later amendments. The local Ins titutional Review Board (ethics committee at the 108
Technical University Dresden) reviewed and approved this study (approval number: BO -EK-109
137042018). The trial was registered on clinicaltrials.gov (trial registration ID: NCT05268432). 110
Written informed consent to laparoscopic image data acquisition, data annotation, data analysis, 111
and anonymized data publication was obtained from all participants. Before publication, all data 112
was anonymized according to the general data protection regulation of the European Union. 113
114
Dataset 115
Based on the full-length surgery recordings and respective temporal annotations of organ visibility, 116
individual image frames were extracted and annotated as described previously. 22 The resulting 117
Dresden Surgical Anatomy Dataset comprises 13195 distinct images with pixel -wise 118
segmentations of eleven anatomical structures: abdominal wall, c olon, intestinal vessels (inferior 119
mesenteric artery and inferior mesenteric vein with their subsidiary vessels), liver, pancreas, small 120
intestine, spleen, stomach, ureter and vesicular glands. Moreover, the dataset comprises binary 121
annotations of the pres ence of each of these organs for each image. The dataset is publicly 122
available via the following link: https://www.nature.com/articles/s41597-022-01719-2. 123
124
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
5
For machine learning purposes, the Dresden Surgical Anatomy Dataset was split into training, 125
validation, and test data as follows (Figure 1): 126
— Training set (at least 12 surgeries per anatomical structure): surgeries 1, 4, 5, 6, 8, 9, 10, 127
12, 15, 16, 17, 19, 22, 23, 24, 25, 27, 28, 29, 30, 31. 128
— Validation set (3 surgeries per anatomical structure): surgeries 3, 21, 26. 129
— Test set (5 surgeries per anatomical structure): surgeries 2, 7, 11, 13, 14, 18, 20, 32. 130
This split is proposed for future works using the Dresden Surgical Anatomy Dataset to reproduce 131
the variance of the entire dataset within each subset, and to ensure comparability regarding clinical 132
variables between the training, the validation, and the test set . Surgeries for the test set were 133
selected to minimize variance regarding the number of frames over the segmented classes. Out 134
of the remaining surgeries, the validation set was separated from the training set using the same 135
criterion. 136
137
Structure-specific semantic segmentation model 138
To segment each anatomical structure, a separate convolutional neural network for segmentation 139
a DeeplabV323 model with a ResNet50 backbone with default PyTorch pretraining on the COCO 140
dataset24, was used. The networks were trained using cross -entropy loss and the AdamW 141
optimizer25 for 100 epochs with a starting learning rate of 10-4 and a linear learning rate scheduler 142
decreasing the learning rate by 0.9 every 10 epochs. For data augmentation, we applied random 143
scaling and rotation, as well as brightness adjustments. The final model for each organ was 144
selected via the Intersection-over-Union (IoU) on the validation dataset and evaluated using the 145
Dresden Surgical Anatomy Dataset with the abovementioned training -validation-test split (Figure 146
1). 147
Segmentation performance was assessed using F1 score, IoU, precision, recall , and specificity 148
on the test folds. These parameters are commonly used technical measures of prediction 149
exactness, ranging from 0 (least exact prediction) to 1 (entirely correct pr ediction without any 150
misprediction). 151
152
Combined semantic segmentation model 153
A convolutional neural network with a common encoder and eleven decoders for combined 154
segmentation of the eleven anatomical structures was trained. The used architecture is an 155
extension of DeepLabV323. A shared ResNet50 backbone with default PyTorch pretraining on the 156
COCO dataset24, was used. For each class, a DeepLabV3 decoder was then run on the features 157
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
6
extracted from a given image by the backbone. As the images are only annotated for binary 158
classes, the loss is only calculated for the decoder associated with the class annotated in a given 159
image. The remaining training procedure was identical to the structure-specific model. The model 160
was trained and evaluated using the Dresden Surgical Anatomy Dataset with the abovementioned 161
training-validation-test split (Figure 1). 162
Segmentation performance was assessed using F1 score, IoU, precision, recall, and specificity 163
on the test folds.26 164
165
Comparative evaluation of algorithmic and human performance 166
To determine the clinical potential of automated segmentation of anatomical risk structures, the 167
segmentation performance of 28 humans was compared to that of the structure -specific and the 168
combined semantic segmentation models using the examp le of the pancreas. The local 169
Institutional Review Board (ethics committee at the Technical University Dresden) reviewed and 170
approved this study (approval number: BO-EK-566122021). All participants provided written 171
informed consent to anonymous study participation, data acquisition and analysis, and publication. 172
In total, 28 participants (physician and non-physician medical staff, medical students, and medical 173
laypersons) marked the pancreas in 35 images from the Dresden Surgical Anatomy Dataset22 with 174
bounding boxes . These images originated from 26 different surgeries, and the pancreas was 175
visible in 16 of the 35 images . Each of the previously selected 35 images was shown once, the 176
order being arbitrarily chosen but identical for all participants. The open -source annotation 177
software Computer Vision Annotation Tool (CVAT) was used for annotations. In cases where the 178
pancreas was seen in multiple, non-connected locations in the image, participants were asked to 179
create separate bounding boxes for each area. 180
Based on the structure -specific and the combined semantic segmentation model s, axis-aligned 181
bounding boxes marking the pancreas were generated in the 35 images from the pixel -wise 182
segmentation. To guarantee that the respective images were not part of the training data, four -183
fold cross validation was used, i.e. the origin surgeries were split into four equal -sized batches, 184
and algorithms were trained on three bat ches that did not c ontain the respective origin image 185
before being applied to segmentation. 186
To compare human and algorithm performance, the bounding box es created by each participant 187
and the structure-specific as well as the combined semantic segmentation models were compared 188
to bounding boxes derived from the Dresden Surgical Anatomy Dataset, which were defined as 189
ground truth. IoU between the m anual or algorithmical bounding box and the ground truth was 190
used to compare segmentation accuracy. 191
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
7
Data Availability 192
The Dresden Surgical Anatomy Dataset is publicly available via the following link: 193
https://www.nature.com/articles/s41597-022-01719-2. All other data generated and analyz ed 194
during the current study are available from the corresponding authors on reasonable request. To 195
gain access, data requestors will need to sign a data access agreement. 196
197
Code Availability 198
The most relevant scripts used for dataset compilation are publicly available via the following link: 199
https://zenodo.org/record/6958337#.YzsBdnZBzOg. 200
201
Results
202
Machine Learning-based anatomical structure segmentation in structure-specific models 203
Structure-specific multi-layer convolutional neural networks (Figure 1) were trained to segment 204
the abdominal wall, the colon, intestinal vessels (inferior mesenteric artery and inferior mesenteric 205
vein with their subsidiary vessels), the liver, the pancreas, the small intestine, the spleen, the 206
stomach, the ureter, and vesicular glands. Table 1 displays mean F1 score, IoU, precision, recall, 207
and specificity for individual anatomical structures as predicted by the structure -specific 208
algorithms. Out of the analyzed segmentation models, performance was lowest for the vesicular 209
glands (mean IoU: 0.28 ± 0. 21) and the pancreas (mean IoU: 0.28 ± 0.27) , w hile excellent 210
predictions were achieved for the abdominal wall (mean IoU: 0.83 ± 0.14) and the small intestine 211
(mean IoU: 0.80 ± 0.18 ). In segmentation of the ureter, the vesicular glands and the intestinal 212
vessel structures , there was a relevant proportion of images with no detection or no overlap 213
between ground truth, while for all remaining anatomical structures, this proportion was minimal 214
(Figure 2 a). While the images, in which the highest IoUs were observed , mostly displayed large 215
organ segments that were clearly visible (Figure 2 b), the images with the lowest IoU were of 216
variable quality with confounding factors such as blood, smoke, soiling of the endoscope lens, or 217
pictures blurred by camera shake (Figure 2 c). 218
Inference on a single image with a resolution of 640 x 512 pixels required, on average, 28 ms on 219
an Nvidia A5000, resulting in a frame rate of almost 36 frames per second. This runtime includes 220
one decoder, meaning that only the segmentation for one anatomical class is included. 221
222
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
8
Table 1: Summary of performance metrics for anatomical structure segmentation using 223
structure-specific models based on the DeepLabv3 architecture. For each metric, mean and 224
standard deviation are displayed. 225
226
Anatomical structure F1 score IoU Precision Recall Specificity
Abdominal wall 0.90 ± 0.10 0.83 ± 0.14 0.89 ± 0.14 0.93 ± 0.07 0.97 ± 0.04
Colon 0.79 ± 0.20 0.69 ± 0.22 0.80 ± 0.21 0.82 ± 0.21 0.97 ± 0.05
Inferior mesenteric artery 0.54 ± 0.26 0.41 ± 0.22 0.55 ± 0.25 0.67 ± 0.33 0.99 ± 0.01
Intestinal veins 0.54 ± 0.33 0.44 ± 0.29 0.70 ± 0.26 0.56 ± 0.36 1.00 ± 0.00
Liver 0.80 ± 0.23 0.71 ± 0.25 0.85 ± 0.21 0.81 ± 0.24 0.98 ± 0.03
Pancreas 0.37 ± 0.32 0.28 ± 0.27 0.59 ± 0.37 0.37 ± 0.36 1.00 ± 0.01
Small intestine 0.87 ± 0.14 0.80 ± 0.18 0.87 ± 0.16 0.91 ± 0.15 0.97 ± 0.04
Spleen 0.79 ± 0.23 0.69 ± 0.24 0.74 ± 0.22 0.90 ± 0.24 0.99 ± 0.01
Stomach 0.71 ± 0.24 0.60 ± 0.25 0.65 ± 0.25 0.89 ± 0.21 0.98 ± 0.02
Ureter 0.47 ± 0.30 0.36 ± 0.25 0.53 ± 0.28 0.57 ± 0.39 1.00 ± 0.00
Vesicular glands 0.40 ± 0.25 0.28 ± 0.21 0.37 ± 0.28 0.62 ± 0.35 0.97 ± 0.03
227
Machine Learning-based anatomical structure segmentation in a combined model 228
Using all annotated images from the Dresden Surgical Anatomy Dataset, a combined model with 229
a mutual encoder and organ-specific decoders was trained (Figure 1). Table 2 displays mean F1 230
score, IoU, precision, recall, and specificity for anatomical structure segmentation in the combined 231
model. The performance of the combined model was overall similar to that of structure -specific 232
models (Table 1), with highest segmentation performance for the abdominal wall (IoU: 0.78 ± 0.14) 233
and the small intestine (IoU: 0.74 ± 0.18), and the lowest performance for the ureter (IoU: 0.22 ± 234
0.24) and the vesicular glands (IoU: 0.32 ± 0.20). In comparison to the respective structure-specific 235
models, the combined model performed notably weaker in liver, stomach, and ureter 236
segmentation, while performance for the other anatomical structures was similar. The proportion 237
of un- or entirely mispredicted images was largest in the ureter, the stomach, the abdominal vessel 238
structures, and the vesicular glands (Figure 3 a). Similar to the structure-specific models, trends 239
towards an impact of segment size, endoscope lens soiling, blurry images, and presence of blood 240
or smoke were seen when comparing image quality of well -predicted images (Figure 3 b) to 241
images with poor or no prediction (Figure 3 c). 242
Inference on a single image with a resolution of 640 x 512 pixels required, on average, 71 ms on 243
an Nvidia A5000, resulting in a frame rate of about 14 frames per second. This runtime includes 244
all 11 decoders, meaning that segmentations for all organ classes are included. 245
246
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
9
Table 2: Summary of performance metrics for anatomical structure segmentation using the 247
combined model (common encoder w ith structure -specific decoders ). For each metric, 248
mean and standard deviation are displayed. 249
250
Anatomical structure F1 score IoU Precision Recall Specificity
Abdominal wall 0.87 ± 0.10 0.78 ± 0.14 0.85 ± 0.12 0.92 ± 0.13 0.94 ± 0.05
Colon 0.75 ± 0.17 0.62 ± 0.19 0.71 ± 0.18 0.87 ± 0.20 0.96 ± 0.03
Inferior mesenteric artery 0.52 ± 0.29 0.40 ± 0.25 0.57 ± 0.26 0.61 ± 0.35 0.99 ± 0.01
Intestinal veins 0.61 ± 0.25 0.48 ± 0.23 0.61 ± 0.22 0.72 ± 0.31 1.00 ± 0.00
Liver 0.59 ± 0.33 0.49 ± 0.32 0.70 ± 0.30 0.59 ± 0.37 0.97 ± 0.03
Pancreas 0.63 ± 0.21 0.49 ± 0.19 0.65 ± 0.23 0.70 ± 0.23 0.99 ± 0.01
Small intestine 0.84 ± 0.15 0.74 ± 0.18 0.80 ± 0.16 0.91 ± 0.17 0.96 ± 0.04
Spleen 0.68 ± 0.28 0.57 ± 0.28 0.69 ± 0.24 0.78 ± 0.32 0.99 ± 0.01
Stomach 0.46 ± 0.35 0.37 ± 0.31 0.62 ± 0.30 0.48 ± 0.39 0.99 ± 0.02
Ureter 0.30 ± 0.31 0.22 ± 0.24 0.52 ± 0.30 0.36 ± 0.39 1.00 ± 0.01
Vesicular glands 0.45 ± 0.24 0.32 ± 0.20 0.53 ± 0.29 0.51 ± 0.31 0.98 ± 0.02
251
Performance of machine learning models in relation to human performance 252
To approximate the clinical value of the previously described algorithms for anatomical structure 253
segmentation, the performances of the structure-specific and the combined model were compared 254
to that of of a cohort of 28 physicians, medical students, and persons with no medical background 255
(Figure 4 a), and different degrees of experience in laparoscopic surgery (Figure 4 b). A vulnerable 256
anatomical structure with – measured by classical metrics of overlap (Tables 1 and 2) – 257
comparably weak segmentation performance of the trained algorithms, the pancreas was selected 258
as an example. 259
Comparing bounding box segmentat ions of the pa ncreas of human annotators, the medical and 260
laparoscopy-specific experience of participants was mirrored by the respective IoUs describing 261
the overlap between the pancreas annotation and the ground truth. The pancreas-specific 262
segmentation model (IoU: 0.29) and the combined segmentation model (IoU: 0.21) outperformed 263
26 out of the 28 human participants (Figures 4 c and d). Overall, these results demon strate that 264
the developed models have clinical potential to improve the recognition of vulnerable anatomical 265
structures in laparoscopy. 266
267
Discussion
268
In surgery , misinterpretation of visual cues can result in objectifiable errors with serious 269
consequences.21 Based on a robust public dataset providing 13195 laparoscopic images with 270
segmentations of eleven intra -abdominal anatomical structures, this study explores the potential 271
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
10
of machine learning for automated segmentation of these organs, and compares algorithmic 272
segmentation quality to that of humans with varying experience in minimally -invasive abdominal 273
surgery. 274
In summary, the presented findings suggest that machine learning -based segmentation of 275
intraabdominal organs and anatomical structures is possible and has the potential to provide 276
clinically valuable information . At an average runtime of 71 ms per image, corresponding to a 277
frame rate of 14 frames per second, the combined model would facilitate near -real-time 278
identification of eleven anatomical structures. These runtimes mirror the performance of a non -279
optimized version of the model, which can be significantly improved using methods such as 280
TensorRT from Nvidia. Measured by classical metrics of overla p between segmentation and 281
ground truth, predictions were, overall, better for large and similar -appearing organs such as the 282
abdominal wall, the liver, the stomach, and the spleen as compared to smaller and more diverse-283
appearing organs such as the pancreas, the ureter, or vesicular glands. Furthermore, poor image 284
quality (i.e. images blurred by camera movements, presence of blood or smoke in images) w as 285
linked to lower accuracy of machine learning -based segmentation s. These findings imply that 286
computer vision studies in laparoscopy should be carefully interpreted taking representativity and 287
potential selection of underlying training and validation data into consideration. 288
Measured by classical metrics of overlap (e.g. IoU, F1 score, precision, recall, specificity) that are 289
commonly used to evaluate segmentation performance , the structure -specific models and the 290
combined model performed similar ly for most anatomical structures with average IoUs ranging 291
from 0.28 to 0.83 and from 0. 22 to 0.78, respectively. The structure-specific model was superior 292
to the combined segmentation model w ith respect to liver, stomach, and ureter segmentation . 293
Interpretation of such metrics of overlap , however, represents a major challenge in computer 294
vision applications in medical domains such as dermatology and endoscopy27–29 as well as non-295
medical domains such as autonomous driving30. In the specific use case of laparoscopic surgery, 296
evidence suggests that such technical metrics alone are not sufficient to characterize the clinical 297
potential and utility of segmentation algorithms.31,32 In this context, the subjective clinical utility of 298
a bounding box-based detection system recognizing the common bile duct and the cystic duct at 299
average precisions of 0.32 and 0.07, respectively, demonstrated by Tokuyasu et al., supports this 300
hypothesis.12 In the presented analysis, the trained structure-specific and combined machine 301
learning algorithm s outperformed all human participants in the specific t ask of bounding box 302
segmentation of the pancreas except for two surgical specialists with over 10 years of experience. 303
This suggests that even for structures such as the pancreas with seemingly poor segmentation 304
quality (segmentation IoU of the best-performing model: 0.49 ± 0. 19 in the test set ) have the 305
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
11
potential to provide clinically valuable help in anatomy recognition. Notably, the best average IoUs 306
achieved in this comparative study were 0. 29 (for the structure-specific model) and 0.36 (for the 307
best human participan t), which w ould both be considered less reliable segmentation quality 308
measures on paper. This encourages further discussion about metrics for segmentation quality 309
assessment in clinical AI. In the future, the potential of the descr ibed dataset 22 and organ 310
segmentation algorithms could be exploited for educational purposes 33,34, for guidance systems 311
facilitating real -time detection of risk and target structures 18,32,35, or as an auxiliary function 312
integrated in more complex surgical assistance systems, such as guidance systems re lying on 313
automated liver registration36. 314
The limitations of this work are mostly related to the dataset and general limitations of machine 315
learning-based segmentation: First, the Dresden Surgical Anatomy Dataset is a monocentric 316
dataset based on 32 robot -assisted rectal surgeries. Therefore, the images used for algorithm 317
training and validation display organs from specific angles, which could limit gene ralizability and 318
transferability of the presented findings to other minimally -invasive abdominal surgeries, 319
particularly non-robotic procedures. Second, annotations were required for training of mac hine 320
learning algorithms, potentially inducing some bias towards the way that organs were annotated 321
in the algorithms, which may differ from individual healthcare professionals’ way of recognizing an 322
organ. This is particularly relevant for organs such as the ureters or the pancreas, which often 323
appear covered by layers of tissue. Here, computer vision -based algorithms that solely consider 324
the laparoscopic images provided by the Dresden Surgical Anatomy Dataset for identification of 325
risk structures will only be able to identify an organ once it is visible. For an earlier recognition of 326
such hidden risk structures, more training data with meaningful annotations would be necessary. 327
Importantly, the presented comparison to human performance focused on segmentation of visible 328
anatomy as well, neglecting that humans (and possibly computers, too) could already identify a 329
risk structure hidden underneath tissue layers. The existing limitations notwithstanding, the 330
presented study represent s an important addition to the growing body of research on medical 331
image analysis in laparoscopic surgery, particularly by linking technical metrics to human 332
performance. 333
In conclusion, this study demonstrates that m achine learning methods have the potential to 334
provide clinically relevant near-real-time assistance in anatomy recognition in minimally-invasive 335
surgery. This study is the first to use the recently published Dresden Surgical Anatomy Dataset, 336
providing baseline algorithms for organ segmentation and evaluating the clinical relevance of such 337
algorithms. Future research should investigate other segmentation methods, the transferability of 338
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
12
these results to other surgical procedures, and the clinical impact of real-time surgical assistance 339
systems and didactic applications based on automated segmentation algorithms. 340
341
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
13
References
342
1. Wang, P. et al. Effect of a deep -learning computer-aided detection system on adenoma 343
detection during colonoscopy (CADe -DB trial): a dou ble-blind randomised study. Lancet 344
Gastroenterol. Hepatol. 5, 343–351 (2020). 345
2. Wang, P. et al. Real-time automatic detection system increases colonoscopic polyp and 346
adenoma detection rates: a prospective randomised controlled study. Gut 68, 1813–1819 347
(2019). 348
3. Kather, J. N. et al. Pan-cancer image -based detection of clinically actionable genetic 349
alterations. Nat. Cancer 2020 18 1, 789–799 (2020). 350
4. Esteva, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. 351
Nature 542, 115–118 (2017). 352
5. Simillis, C. et al. Open Versus Laparoscopic Versus Robotic Versus Transanal Mesorectal 353
Excision for Rectal Cancer: A Systematic Review and Network Meta -analysis. Ann. Surg. 354
270, 59–68 (2019). 355
6. Zhao, J. J. et al. Comparative outcom es of needlescopic, single -incision laparoscopic, 356
standard laparoscopic, mini -laparotomy, and open cholecystectomy: A systematic review 357
and network meta-analysis of 96 randomized controlled trials with 11,083 patients. Surgery 358
170, 994–1003 (2021). 359
7. Luketich, J. D. et al. Outcomes after minimally invasive esophagectomy: review of over 360
1000 patients. Ann. Surg. 256, 95–103 (2012). 361
8. Thomson, J. E. et al. Laparoscopic versus open surgery for complicated appendicitis: a 362
randomized controlled trial to prove safety. Surg. Endosc. 29, 2027–2032 (2015). 363
9. Islam, M., Atputharuban, D. A., Ramesh, R. & Ren, H. Real -time instrument segmentation 364
in robotic surgery using auxiliary supervised deep adversarial learning. IEEE Robot. Autom. 365
Lett. 4, 2188–2195 (2019). 366
10. Roß, T. et al. Comparative validation of multi -instance instrument segmentation in 367
endoscopy: results of the ROBUST -MIS 2019 challenge. Med. Image Anal. 70, 101920 368
(2020). 369
11. Shvets, A. A., Rakhlin, A., Kalinin, A. A. & Iglovikov, V. I. Automatic Instrum ent 370
Segmentation in Robot-Assisted Surgery using Deep Learning. in Proceedings - 17th IEEE 371
International Conference on Machine Learning and Applications, ICMLA 2018 624–628 372
(Institute of Electrical and Electronics Engineers Inc., 2019). 373
doi:10.1109/ICMLA.2018.00100. 374
12. Tokuyasu, T. et al. Development of an artificial intelligence system using deep learning to 375
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
14
indicate anatomical landmarks during laparoscopic cholecy stectomy. Surg. Endosc. 2020 376
354 35, 1651–1658 (2020). 377
13. Mascagni, P. et al. Artificial Intelligence for Surgical Safety: Automatic Assessment of the 378
Critical View of Safety in Laparoscopic Cholecystectomy Using Deep Learning. Ann. Surg. 379
(2020) doi:10.1097/SLA.0000000000004351. 380
14. Jin, A. et al. Tool Detection and Operative Skill Assessment in Surgical Videos Using 381
Region-Based Convolutional Neural Networks. Proc. - 2018 IEEE Winter Conf. Appl. 382
Comput. Vision, WACV 2018 2018-January, 691–699 (2018). 383
15. Funke, I. et al. Using 3D Convolutional Neural Networks to Learn Spatiotemporal Features 384
for Automatic Surgical Gesture Recognition in Video. Med. Image Comput. Comput. Assist. 385
Interv. – MICCAI 2019. Lect. Notes Comput. Sci. 11768, 467–475 (2019). 386
16. Lavanchy, J. L. et al. Automation of surgical skill assessment using a three-stage machine 387
learning algorithm. Sci. Reports 2021 111 11, 1–9 (2021). 388
17. Maier-Hein, L. et al. Surgical data science – from concepts toward clinical translation. Med. 389
Image Anal. 76, 102306 (2022). 390
18. Madani, A. et al. Artificial Intelligence for Intraoperative Guidance. Ann. Surg. (2020) 391
doi:10.1097/sla.0000000000004594. 392
19. Fecso, A. B., Szasz, P., Kerezov, G. & Grantcharov, T. P. The effect of technical 393
performance on patient outcomes in surgery. Ann. Surg. 265, 492–501 (2017). 394
20. Mazzocco, K. et al. Surgical team behaviors and patient outcomes. Am. J. Surg. 197, 678–395
685 (2009). 396
21. Suliburk, J. W. et al. Analysis of Human Performance Deficiencies Associated With Surgical 397
Adverse Events. JAMA Netw. Open 2, e198067–e198067 (2019). 398
22. Carstens, M. et al. The Dresden Surgical Anatomy Dataset for abdominal organ 399
segmentation in surgical data science. Sci. Data (2023) doi:10.1038/s41597-022-01719-2. 400
23. Chen, L.-C., Papandreou, G., Schroff, F. & Adam, H. Rethinking Atrous Convolution for 401
Semantic Image Segmentation. arXiv (2017) doi:10.48550/arxiv.1706.05587. 402
24. Lin, T. Y. et al. Microsoft COCO: Common Objects in Context. Lect. Notes Comput. Sci. 403
(including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics) 8693 LNCS, 740–404
755 (2014). 405
25. Loshchilov, I. & Hutter, F. Decoupled Weight Decay Regularization. 7th Int. Conf. Learn. 406
Represent. ICLR 2019 (2017) doi:10.48550/arxiv.1711.05101. 407
26. Leger, S. et al. A comparative study of machine learning methods for time-to-event survival 408
data for radiomics risk modelling. Sci. Rep. 7, 11 (2017). 409
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
15
27. Renard, F., Guedria, S., Palma, N. De & Vuillerme, N. Variability and reproducibility in deep 410
learning for medical image segmentation. Sci. Rep. 10, 1–16 (2020). 411
28. Powers, D. M. W. & Ailab. Evaluation: from precision, recall and F -measure to ROC, 412
informedness, markedness and correlation. arXiv (2020) doi:10.48550/arxiv.2010.16061. 413
29. Parikh, R. B., Teeple, S. & Navathe, A. S. Addressing Bias in Artificial Intelligence in Health 414
Care. JAMA 322, 2377–2378 (2019). 415
30. Zhang, Y., Mehta, S. & Caspi, A. Rethinking Semantic Segmentation Evaluation for 416
Explainability and Model Selection. (2021). 417
31. Reinke, A. et al. Common Limitations of Image Processing Metrics: A Picture Story. arXiv 418
(2021) doi:10.48550/arxiv.2104.05642. 419
32. Hashimoto, D. A. et al. Computer Vision Analysis of Intraoperative Video: Automated 420
Recognition of Operative Steps in Laparoscopic Sleeve Gastrectomy. Ann. Surg. 270, 414–421
421 (2019). 422
33. Hu, Y. Y. et al. Complementing Operating Room Teaching With Video -Based Coaching. 423
JAMA Surg. 152, 318–325 (2017). 424
34. Mizota, T., Anton, N. E. & Stefanidis, D. Surgeons see anatomical structures faster and 425
more accurately compared to novices: Development of a pattern recognition skill 426
assessment platform. Am. J. Surg. 217, 222–227 (2019). 427
35. Ward, T. M. et al. Computer vision in surgery. Surgery 169, 1253–1256 (2021). 428
36. Docea, R. et al. Simultaneous localisation and mapping for laparoscopic liver navigation : a 429
comparative evaluation study. in Medical Imaging 2021: Image -Guided Procedures, 430
Robotic Interventions, and Modeling (eds. Linte, C. A. & Siewerdsen, J. H.) vol. 11598 8 431
(SPIE, 2021). 432
433
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
16
Figures and Figure Captions 434
435
Figure 1 436
Fig. 1 Schematic illustration of the structure-specific and combined machine learning 437
models used for semantic segmentation. The Dresden Surgical Anatomy Dataset was split into 438
a training, a validation, and a test set. For spatial segmentation, two machine learning models 439
were trained: A structure-specific model with individual encoders and decoders, and a combined 440
model with a common encoder and structure-specific decoders. 441
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
17
442
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
18
Figure 2 443
Fig. 2 Pixel-wise organ segmentation with structure-specific models trained on the 444
respective organ subsets of the Dresden Surgical Anatomy Dataset. (a) Violin plot 445
illustrations of performance metrics for structure-specific segmentation models. The median and 446
quartiles are illustrated as solid and dashed lines, respectively. (b) Example images with the 447
highest IoUs for liver, pancreas, stomach, and ureter segmentation with structure -specific 448
segmentation models. Ground truth is displayed as blue line (upper panel), model segmentations 449
are displayed as white overlay (lower panel). (c) Example images with the lowest IoUs for liver, 450
pancreas, stomach, and ureter segmentation with structu re-specific segmentation models. 451
Ground truth is displayed as blue line (upper panel), model segmentations are displayed as white 452
overlay (lower panel). 453
454
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
19
455
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
20
Figure 3 456
Fig. 3 Pixel-wise organ segmentation with the combined model trained on the entire 457
Dresden Surgical Anatomy Dataset with a common encoder and structure -specific 458
decoders. (a) Violin plot illustrations of performance metrics for the combined segmentation 459
model. The median and quartiles are illustrated as solid and dashed lines, respectively. (b) 460
Example images with the highest IoUs for liver, pancreas, stomach, and ureter segmentation with 461
the combined segmentation model. Ground truth is displayed as blue line (upper panel), model 462
segmentations are displayed as white overlay (lower panel). (c) Example images with the lowest 463
IoUs for liver, pancreas, stomach, and ureter segmentation with the combined segmentation 464
model. Ground truth is displayed as blue line (upper panel), model s egmentations are displayed 465
as white overlay (lower panel). 466
467
468
Figure 4 469
Fig. 4 Comparison of pancreas segmentation performance of the structure-specific and the 470
combined semantic segmentation models with a cohort of 28 human participants. (a) 471
Distribution of medical and non -medical professions among human participants. (b) Distribution 472
of laparoscopy experience among human participants. (c) Waterfall chart displaying the average 473
pancreas segmentation IoUs of participants with different professi ons as compared to the IoU 474
generated by the structure-specific and the combined semantic segmentation models. (d) 475
Waterfall chart displaying the average pancreas segmentation IoUs of participants with varying 476
laparoscopy experience as compared to the IoU generated by the structure-specific and the 477
combined semantic segmentation models. 478
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
21
Abbreviations 479
AI Artificial Intelligence 480
IoU Intersection-over-Union 481
SD Standard deviation 482
483
Acknowledgements
and Funding 484
This work has been funded by the project fundi ng within the Else Kröner Fresenius Center for 485
Digital Health (EKFZ), Dresden, Germany (project “CoBot”) and by the German Research 486
Foundation DFG within the Cluster of Excellence EXC 2050: “Center for Tactile Internet with 487
Human-in-the-Loop (CeTI)” (project number 390696704). FRK received funding from the Medical 488
Faculty of the Technical University Dresden within the MedDrive Start program (grant number 489
60487) and from the Joachim Herz Foundation (Add -On Fellowship for Interdisciplinary Life 490
Science). FMR received a doctoral student scholarship from the Carus Promotionskolleg Dresden. 491
The authors gratefully acknowledge excellent project coordination by Dr. Elisabeth Fischermeier 492
and Dr. Grit Krause-Jüttler. 493
494
Author contributions 495
FRK, JW, MD, SS, and SB conceptualized the study. FRK, FMR, and MC collected and annotated 496
clinical and video data and contributed to data analysis. ACJ, SL, and SB implemented and trained 497
the neural networks and contributed to data analysis. JW, MD, and SS supervised the projec t, 498
provided infrastructure and gave important scientific input. FRK drafted the initial manuscript text. 499
All authors reviewed, edited, and approved the final manuscript. 500
501
Competing interests 502
The authors declare no conflicts of interest. 503
504
All rights reserved. No reuse allowed without permission.
(which was not certified by peer review) is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity.
The copyright holder for this preprintthis version posted January 11, 2023. ; https://doi.org/10.1101/2022.11.11.22282215doi: medRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.