Objective assessment tools in laparoscopic or robotic-assisted gynecological surgery: A systematic review.

OA: gold CC-BY-NC-ND-4.0

Abstract

IntroductionThere is a growing emphasis on proficiency-based progression within surgical training. To enable this, clearly defined metrics for those newly acquired surgical skills are needed. These can be formulated in objective assessment tools. The aim of the present study was to systematically review the literature reporting on available tools for objective assessment of minimally invasive gynecological surgery (simulated) performance and evaluate their reliability and validity.Material and methodsA systematic search (1989-2022) was conducted in MEDLINE, Embase, PubMed, Web of Science in accordance with PRISMA. The trial was registered with the Prospective Register of Systematic Reviews (PROSPERO) ID: CRD42022376552. Randomized controlled trials, prospective comparative studies, prospective single-group (with pre- and post-training assessment) or consensus studies that reported on the development, validation or usage of assessment tools of surgical performance in minimally invasive gynecological surgery, were included. Three independent assessors assessed study setting and validity evidence according to a contemporary framework of validity, which was adapted from Messick's validity framework. Methodological quality of included studies was assessed using the modified medical education research study quality instrument (MERSQI) checklist. Heterogeneity in data reporting on types of tools, data collection, study design, definition of expertise (novice vs. experts) and statistical values prevented a meaningful meta-analysis.ResultsA total of 19 746 titles and abstracts were screened of which 72 articles met the inclusion criteria. A total of 37 different assessment tools were identified of which 13 represented manual global assessment tools, 13 manual procedure-specific assessment tools and 11 automated performance metrices. Only two tools showed substantive evidence of validity. Reliability and validity per tool were provided. No assessment tools showed direct correlation between tool scores and patient related outcomes.ConclusionsExisting objective assessment tools lack evidence on predicting patient outcomes and suffer from limitations in transferability outside of the research environment, particularly for automated performance metrics. Future research should prioritize filling these gaps while integrating advanced technologies like kinematic data and AI for robust, objective surgical skill assessment within gynecological advanced surgical training programs.
Full text 34,235 characters · extracted from pmc-nxml · 8 sections · click to expand

Author

Freweini Martha Tesfai: Concept, design, data collection, analysis and interpretation, manuscript preparation. Jasleen Nagi and Iona Morrison: Data collection, data analysis and interpretation, manuscript preparation. Matt Boal: Concept and design, data analysis and interpretation, manuscript preparation. Adeola Olaitan, Dhivya Chandrasekaran, Danail Stoyanov, Anne Lanceley and Nader Francis: Data analysis and interpretation, manuscript preparation.

Results

A total of 19 746 titles and abstracts were screened for their eligibility and four additional studies were identified through other sources (citation searching n  = 4, gray literature n  = 0). A total of 174 articles were included for full text review and 102 studies were deemed ineligible, primarily because these studies were not reporting solely on MIGS or because no assessment tools were reported in those studies. Finally, 72 studies were included for further analysis. A breakdown of inclusion and exclusion is shown in the PRISMA diagram (Figure  1 ). PRISMA flow diagram. Study characteristics are summarized in Table  2 . Studies were predominately conducted in the USA (48.6%) or Europe (25.0%); however, authors were represented from five different continents (all except South America and Antarctica). Studies were published between 2002 and 2023. Included studies consisted of manual assessment tools ( n  = 26) and APMs ( n  = 11). The later ones were tools from which the scoring system was directly derived from kinematic data and systems events data, usually in a VR setting. A total of 36 out of 72 (50%) studies were designed to address the utilization of previously validated tools in an educational intervention setting, followed by 31 (43.7%) studies aimed to either develop a new tool or assess the validity of existing ones. Characteristics of 72 included studies reporting on objective assessment tools in minimally invasive gynecological surgery. With the exception of one paper, scoring five points, 21 all other papers had a score ranging from 10 to 16.5 on the 18‐MERSQI checklist. The main limitations were lack of randomized controlled trials (study design), lack of multicenter studies (sampling) and the absence of correlating study outcomes with clinical outcomes. The risk of bias per tool can be found in Tables  3 , 4 , 5 , 6 , 7 and risk of bias per study can be found in Table  S1 . Validity evidence per tool, manual global assessment tools (1/2). Intraoperative Dry laboratory Intraoperative Wet laboratory Dry laboratory Intraoperative Wet laboratory Dry laboratory Virtual Reality Intraoperative Dry laboratory Virtual Reality Intraoperative Dry laboratory Robotic supracervical hysterectomy Laparoscopic colpotomy Total laparoscopic hysterectomy Laparoscopic myomectomy Laparoscopic colpotomy Laparoscopic vaginal cuff suturing Laparoscopic sacrocolpopexy Laparoscopic supracervical hysterectomy Laparoscopic myomectomy Laparoscopic oophorectomy, dissection and ligature of uterine artery Diagnostic laparoscopy Laparoscopic sacrocolpopexy Laparoscopic bilateral tubal ligation Laparoscopic salpingectomy Total laparoscopic hysterectomy Laparoscopic suturing and intracorporeal knot tying Basic dry laboratory tasks Laparoscopic ovarian cystectomy Total laparoscopic hysterectomy Laparoscopic bilateral tubal ligation Laparoscopic salpingo‐oophorectomy Laparoscopic vaginal cuff closure Bilateral tubal ligation Basic dry laboratory tasks Total laparoscopic hysterectomy Laparoscopic Salpingectomy/salpingo‐oophorectomy Delphi consensus developed tool Referred to previous literature for content validity Response process present in all studies except for one study Inter rater reliability: Strong–Excellent Intrarater reliability: Not assessed Item analysis: Not assessed Inter rater reliability: Good–Excellent Intrarater reliability: Good Item analysis: Not assessed Inter‐rater reliability: Good Intrarater reliability: Not assessed Inter item analysis: Excellent Inter‐rater reliability: Weak–excellent Intrarater reliability: Not assessed Item analysis: Not assessed Inter‐rater reliability: Acceptable Intrarater reliability: Excellent Inter item analysis: not assessed Inter‐rater reliability: Good Intrarater reliability: Very strong Inter item analysis: Excellent Relations to other variables Construct validity (training level or case experience) Concurrent validity (other performance scores) Predictive validity (relation to clinical outcomes) Construct validity Concurrent validity Construct validity Concurrent validity Pass mark score defined at 27/35 82/85 32/40 However, not applied for benchmarking/credentialing in training curriculum Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence. 103 Abbreviations: GOALS, global operative sssessment of laparoscopic skills; GRS, global rating skills; m, modified; MERSQI, medical education research study quality 16 ; OSATS, objective structured assessment of technical skills; R‐OSATS, robotic objective structured assessment of technical skills. Validity evidence per tool, manual global assessment tools (2/2). Intraoperative Dry laboratory Dry laboratory Virtual reality Laparoscopic salpingectomy/salpingo‐oophorectomy Laparoscopic resection of endometriosis, adhesiolysis and ovarian drilling Relations to other variables Construct (training level or case experience) Concurrent validity (other performance scores) Predictive validity (relation to clinical outcomes) Pass mark defined at 27/35 However, not applied for benchmarking/credentialing in training curriculum Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence. 103 Abbreviations: GEARS, global evaluative assessment of robotic skills; GERT, generic error rating tool; GRIT, global rating index of technical skills; GSAT, generic skills assessment tool; m, modified; MERSQI, medical education research study quality 16 ; OPRS, operating performance rating system; TCPE, time to correct performed exercise. Validity evidence per tool, manual procedure‐specific assessment tools. Intraoperative Virtual reality Delphi consensus to develop tool Referred to previous literature Experts developed task‐analysis Referred to previous literature for content validity Inter‐rater reliability: Excellent Fair–excellent Relations to other variables Construct validity (training level or case experience) Concurrent validity (other performance scores) Predictive validity (relation to clinical outcomes) Concurrent validity Construct validity Pass mark defined (90/150) However, not applied for benchmarking/credentialing in training curriculum Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence. 103 Abbreviations: CAT‐LSH, competency assessment tool for laparoscopic supracervical hysterectomy; DART, dissection of assessment of robotic technique; H‐OSATS, objective scale for assessment of technical skills of total laparoscopic hysterectomy (TLH); LSO‐OSATS, OSATS of laparoscopic salpingo‐oophorectomy; MERSQI, medical education research study quality 16 ; OSA‐LS, objective structured assessment of laparoscopic salpingectomy; RHAS, robotic hysterectomy assessment score. A total of 26 manual technical skills assessments were included. These consisted of 13 global rating tools, and 13 procedure‐specific tools. Furthermore 11 APMs were identified. The results are summarized under three categories: global, procedure‐specific and automated metrics tools. The OSATS, global operative assessment of laparoscopic skills (GOALS), global evaluative assessment of robotic skills (GEARS), global rating scale (GRS—not further specified) and modified versions of these tools were used most frequently ( n  = 32) in the included studies. With the exception of three tools (robotic‐OSATs, modified GEARS, operative performance rating system [OPRS]), all other tools were validated intraoperatively. Other settings included wet and dry models. The (modified) OSATS was the only manual tool assessed in a VR setting. 22 , 23 The only error rating tool identified was the generic error rating tool (GERT). 24 Kilani proposed the global rating index of technical skills (GRITS) to intraoperatively assess the correlation of surgical skill performance scores between expert assessment and self‐assessment in various laparoscopic gynecological procedures, concluding that self‐assessments have a higher evaluation than expert assessments. 25 Finally, the OPRS was used in a dry laboratory setting, assessing robotic‐assisted laparoscopic radical hysterectomy and pelvic lymphadenectomy performance. 26 A total of 10 procedure‐specific tools were assessed intraoperatively. The robotic sacrocolpopexy simulation model, 18 the “assessment tool for total laparoscopic hysterectomy” 27 and the OSATS for laparoscopic suturing and intracorporeal knot tying 28 were assessed in a laboratory‐based setting. The objective structured assessment of laparoscopic salpingectomy (OSA‐LS) was based on both the original OSATS and a modified rating scale for laparoscopic cholecystectomy developed by Grantcharov et al. 29 The myTIPreport is a smartphone application where both the trainee performing the procedure and a faculty member assessed the technical skills on a checklist immediately after the procedure. 30 The laparoscopic salpingo‐oophorectomy‐OSATS (LSO‐OSATS) was based on the OSA‐LS but consisted of fewer items (6 in the LSO‐OSATS vs. 10 in the OSA‐LS). Remarkably, six minimal invasive hysterectomy procedure‐specific tools were included: Objective scale for assessment of technical skills of TLH (H‐OSATS), the objective structured assessment of TLH (OSA‐TLH), laparoscopic hysterectomy‐OSATS (LH‐OSATS) and the assessment tool for TLH, competency assessment tool for laparoscopic supracervical hysterectomy (CAT‐LSH) and the robotic hysterectomy assessment score (RHAS). 22 , 27 , 31 , 32 , 33 , 34 All manual assessment tools and studies are summarized in Tables  3 , 4 , 5 , 6 . Validity evidence per tool, manual procedure‐specific assessment tools. Relations to other variables Construct validity (training level or case experience) Concurrent validity (other performance scores) Predictive validity (relation to clinical outcomes) Pass mark defined (29.3/55) However, not applied for benchmarking/credentialing in training curriculum Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence. 103 Abbreviations: MERSQI, medical education research study quality 16 ; OSA‐TLH, OSATS of TLH, LH‐OSATS is OSATS of TLH. A total of 11 APMs were identified in this systematic review. These include APMs in robotic‐assisted laparoscopic VR simulations; da Vinci Surgical Simulation, RobotiX Mentor Simulation, and laparoscopic VR simulations: LapSim, LapMentor, MIST, SurgicalSim, MISTELS, TRLCD05, FastTrack and LapVR simulator. No APMs were used intraoperatively. The following part of this review focuses on the sources of validity evidence of each included assessment tool, specified on the unitary framework for manual tools and APMs (Table  1 ). Given the heterogeneity of the interventions being investigated, each tool was categorized along the five dimensions of the contemporary framework (Tables  3 , 4 , 5 , 6 , 7 ; Table  S2 ). Validity evidence per tool, automated performance metrics. Basic laparoscopic VR modules Laparoscopic Salpingectomy Basic laparoscopic VR modules Laparoscopic Salpingectomy/Salpingo‐oophorectomy Relations to other variables Construct validity (training level or case experience) Concurrent validity (other performance scores) Predictive validity (relation to clinical outcomes) Both studies showed construct validity Concurrent validity Pass mark defined at 75/110 48 However, not applied for benchmarking/credentialing in training curriculum Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence. 103 Abbreviations: DvSS, Da Vinci surgical system; LapSim, laparoscopic simulator; LapVR simulator, laparoscopic virtual reality simulator; MERSQI, medical education research study quality 16 ; MIST, minimally invasive surgical trainer, McGill inanimate system for training and evaluation of laparoscopic skills; VBLAST‐PT, virtual basic laparoscopic skill trainer. A total of 8 out of 13 (61.5%) global assessment tools had content validity. 30 , 34 , 35 , 36 , 37 , 38 , 39 , 40 , 41 The studies reporting on the generic skills assessment tool, GERT and OPRS, did not demonstrate content validity. In contrast to the global rating tools, content validity for the procedure‐specific tools was provided 12/13 (92.3%) studies. A total of 10 (76.9%) tools underwent a (hierarchical) task analysis, 21 , 22 , 27 , 28 , 31 , 32 , 33 , 42 , 43 , 44 , 45 , 46 , 47 , 48 , 49 , 50 usually followed by a consensus study with experts. Content validity was demonstrated in 6 (54.5%) different APMs. 51 , 52 , 53 , 54 , 55 , 56 One APM study reported on consensus methodology to reach content validity. 56 Eight out of 13 global assessment tools (61.5%) provided evidence of rater training, either by a training session or by providing a manual for tool usage. 24 , 25 , 26 , 28 , 30 , 37 , 38 , 40 , 41 , 57 , 58 This was applicable to seven out of 12 (58.3%) procedure‐specific tools. 31 , 32 , 33 , 42 , 43 , 44 , 45 , 46 , 47 , 48 , 49 , 50 , 59 , 60 Addison et al. used crowd‐sourced assessment of technical skills (CSATS) for GEARS and raters were routinely trained and evaluated for their rating reliability. 41 All APMs inherently demonstrate response process, as they are automated, hence removing rater bias, and theoretically having 100% reliability. Internal structure was assessed in different ways. The most common reported form was inter‐rater reliability with 10 out of 13 (76.9%) global rating tools and 10 out of 13 (76.9%) procedure‐specific tools. All manual global tools report good to excellent inter‐rater reliability, 23 , 24 , 25 , 30 , 33 , 35 , 37 , 38 , 39 , 40 , 57 , 61 , 62 , 63 , 64 , 65 , 66 , 67 , 68 , 69 , 70 with the exception of one modified OSATS in a dry laboratory setting for a simple laparoscopic ovarian cystectomy by Chahine et al. 71 Inter‐rater reliability among procedure‐specific tools 31 , 32 , 33 , 42 , 43 , 44 , 46 , 47 , 59 , 60 showed good to excellent correlation except for specific domains for the H‐OSATS and the dissection assessment for robotic technique. 43 , 60 Intrarater reliability was reported for 4/13 (30.8%) global tools demonstrating excellent intrarater reliability. 28 , 30 , 38 , 40 The H‐OSATS was the only procedure‐specific tool reporting excellent intrarater reliability. 31 , 60 Only one APM study calculated internal structure, reporting poor internal consistency (Cronbach's alpha 0.58) on the RobotiX Mentor Simulator. 55 A total of 10 out of 13 (79.6%) global tools reported relationships to other variables by either comparing novices to experts (construct validity) or showing significant correlation between scoring outcomes and other performance assessment tools, considered the gold standard (concurrent validity). 22 , 23 , 24 , 27 , 28 , 30 , 35 , 36 , 38 , 39 , 40 , 57 , 58 , 61 , 62 , 63 , 64 , 65 , 66 , 67 , 68 , 69 , 70 , 71 , 72 , 73 , 74 , 75 , 76 , 77 Nine out of 13 (69.2%) procedure‐specific tool studies 31 , 32 , 33 , 42 , 43 , 44 , 46 , 47 , 48 , 49 , 50 , 60 showed construct or concurrent validity. Nine out of 11 (81.8%) APM showed construct or criterion validity. 51 , 52 , 53 , 55 , 56 , 67 , 72 , 78 , 79 , 80 , 81 , 82 , 83 , 84 , 85 , 86 , 87 , 88 , 89 , 90 None of the included studies reported on the association between intraoperative performance of practicing surgeons to clinical/postoperative outcomes of patients (predictive validity). Four out of 26 (15.4%) manual (global and procedure‐specific tools) provided benchmark scores. 30 , 32 , 35 , 57 , 60 , 64 , 67 One out of 11 (9.1%) APMs provided a benchmark score on the RobotiX Mentor providing a pass/fail score of 75/110. 55 Table  8 summarizes the evidence of validity of all tools based on the scoring tool from Table  1 . Only one manual tool showed substantial evidence of validity (score 11–15): the total laparoscopic hysterectomy procedure specific tools: OSA‐TLH. The RobotiX Mentor Simulator was the only APM showing substantive evidence. A total of 17 tools showed moderate evidence (score 6–10) of which 10 were global tools, six were procedure specific and one APMs. Finally, 18 tools showed limited evidence of validity, of which nine were APMs, three manual global tools and six procedure specific tools. Objective assessment tools arranged by strength of validity based on the validity evidence scoring list from Table  1 (substantial, moderate and limited evidence).

Discussion

This systematic review of 72 articles identified 37 surgical performance assessment tools that have been studied in a laparoscopic and robotic‐assisted gynecological surgery setting. This review provided a comprehensive evaluation of the validity and reliability of assessment tools, using a contemporary validity framework (Table  1 ). These included 26 manual tools and 11 APMs. Interestingly, none of the studies were able to show predictive validity (correlating the tool score with clinical outcomes). Tough achieving predictive validity often necessitates a more demanding endeavor, and there is still a significant opportunity to develop study settings correlating tool scores with clinical outcomes. 91 , 92 The General Medical Council (GMC) in the UK has even stated that in the absence of the gold standard, exploring the strength of the relationship between similar established assessment tools, from different surgical specialities, might offer itself as an alternative. 93 Furthermore, more granular analysis of surgical skills, such as the objective clinical human reliability analysis (OCHRA) could enhance the likelihood of achieving predictive validity, associating technical kills with clinical outcomes, regardless of level of expertise. 94 When looking at current training programs, such as the Royal College of Obstetrics and Gynecology (RCOG) in the UK, it interesting to see that the most frequent used objective assessment tool is the OSATS. 95 Global assessment tools are easily available for different procedures. However, this systematic review showed that the only manual tool showing substantive evidence was a procedure specific tool. It should also be noted that the exchange of constructive feedback within the trainer‐trainee dialogue often plays a greater role in shaping learning outcomes. Culligan et al. proposed a robotic surgery simulation training curriculum and established predictive validity by demonstrating a correlation between program completion and improved live surgery outcomes. 72 These included reduced estimated blood loss, shorter operating times, and enhanced intraoperative GOALS scores. However, the study's generalizability was limited by its restriction to board‐certified obstetrics and gynecology surgeons. Despite other studies reporting pass/fail scores for a modified GOALS, modified GEARS, H‐OSATS, OSA‐TLH (both TLH procedure specific tool), the RobotiX Mentor simulator and DvSS, none of them showed any evidence of successful implementation of curriculums for credentialing. Future research should not only focus on investigating other aspects of validity, but also on benchmarking already available objective assessment tools to make them useful additions to surgical national curriculums. This will ultimately enhance the standardization and effectiveness of resident and fellow training in MIGS. This systematic review had some limitations. First, it was limited by only including studies in English. Another limitation was that the majority of the studies were small, conducted once or in a nonrandomized single center setting, risking potential biases and compromising reproducibility of results. Often, different thresholds and definitions were used, producing heterogeneity and the subsequent inability to perform a meaningful meta‐analysis, highlighting that tools should be evaluated more thoroughly in large, well‐run studies. Furthermore, assessment tools for intrauterine and vaginal surgery were not included. However, a significant increase in numbers of assessment tools ( n  = 37) in MIGS were identified, making it, to our knowledge, the most comprehensive and detailed systematic review on the subject of minimal invasive gynecology surgery. It is important to inform the gynecological surgical community of all available tools that can be applied not only in the research settings but to support learning and teaching. Ferriss et al., published a systematic review of intraoperative assessment tools in MIGS, focusing mainly on manual performance metrics. 96 They concluded that procedure‐specific tools are more thoroughly evaluated, however described their limited use due to poor quality studies and borderline reliability. Another scoping review by Hennings et al. explained that most surgical assessment scales were validated in simulation settings, compromising transferability to the operating room. 12 However, comprehensive evaluations of the tools' validity were not reported, mainly lacking the consequence component. This review also provided a comprehensive review of APMs available in MIGS. One of the advantages of APMs compared to manual tools is the automated collection of the performances, preventing rater bias. Furthermore, it is less time consuming and manual tools require a degree of subjective scoring. Furthermore, research has been suggested that skills in laparoscopic surgery can be increased by proficiency‐based procedural VR simulator training. However, this review showed that the majority of APMs (81.8%) has limited validity evidence. This low level of validity hinders transferability outside of the research environment. Moreover, these metrics alone cannot be considered substitutes of experts' input towards surgical competencies. Until true objective assessment tools are in place to provide expert opinion on trainees' performances within the clinical context, APMs are useful adjunct to support objective assessments of surgical skills. This systematic review did not identify APMs using kinematic data from live surgery. However, recent studies in different specialities have been able to correlate kinematic data derived from recording devices with technical performance to create a scoring index in dry lab surgeries and live operations. Lyman et al., were able to correlate the operating robotic index model to level of experience in a dry laboratory robotic‐assisted laparoscopic hepaticojejunostomy reconstruction, showing construct validity. 97 Another example of appliance is the emerging interdisciplinary field of surgical data science (SDS) aiming to improve quality of interventional healthcare by capturing, organizing, analyzing, and modeling data. Mascagni et al., were able to assess the critical view of safety criteria in laparoscopic cholecystectomy through annotating anatomical segments and training a deep neural network to predict critical view of safety occurrence. 98 Utilization of these tools, including kinematic data from advanced computing devices and surgical systems and artificial neural networks will become essential to better understand factors in surgical performance and ultimately standardize safe operation. One key challenge for developing these approaches further is the current absence of large‐scale datasets that fully represent the domains of variation; for example, experience level, subtask, instrument type, in order to allow robust training of AI models with only limited clinical datasets currently available. 99 , 100 Future validation of APMs will support utilization while they are likely to expand in the future with artificial intelligence and machine learning.

Conclusions

This comprehensive review offers an up‐to‐date overview of existing assessment tools for MIGS. With 37 tools identified, including both established manual techniques and APMs, it provides a valuable resource for researchers, educators, and practitioners alike. While global assessment tools remain readily available, procedure‐specific tools hold great educational potential. Importantly, the review highlights the gap in evidence regarding predictive validity—linking assessment scores to patient outcomes. Additionally, it underscores the limitations of current APMs, mainly due to insufficient content validity assessments. Nonetheless, APMs show promise in their objective data collection and potential for reducing rater bias. Future research should focus on addressing these limitations while continuing to explore the integration of advanced technologies like kinematic data analysis and artificial intelligence.

Introduction

Minimally invasive surgery (MIS) in gynecology has a prominent role in the management of gynecological benign and oncological diagnoses. MIS reduces hospital stay and enhances postoperative recovery, making it one of the preferred routes of operation in many diagnoses. 1 In the last two decades, robotic‐assisted laparoscopic surgery has emerged as a new entity within MIS. 2 However, with the introduction of new medical techniques and devices comes the risk of increased errors and unknown consequences. 3 In addition and distinct from open surgery, laparoscopic surgery requires specific surgical skills and endoscopic psychomotor skills to ensure patient safety. 4 Especially in laparoscopic surgery, depth perception is hindered and tactile feedback is reduced. Minimal movements are amplified and range of motion is decreased due to fixation of the trocars. 5 There is increasing evidence that simulation‐based training and assessment such as lower fidelity physical/box video training and higher fidelity VR increase technical skills in the operating room, however, linkage to patient outcomes in minimally invasive gynecological surgery (MIGS) is lacking. 6 Furthermore, interpersonal skills, such teamwork and leadership, but also personal resourcefulness and advanced cognitive skills including error recognition and surgical planning play an important role in skills acquisition and intraoperative performance. 7 , 8 Moreover several studies have shown that surgical performance is associated with clinical outcomes and complication rates. 9 , 10 Recently, there has been a focus on proficiency‐based progression, dictating that the learner must meet specific performance benchmarks before progressing to the next stage in training. 11 To enable this, clearly defined metrics for those newly acquired surgical skills are needed. These can be formulated in objective assessment tools, defining and assessing the key steps of a specific procedure to support credentialing. Global tools, such as the objective structured assessment of technical skills (OSATS), lack specificity which limits their applications in accreditation for a specific procedure, such as a hysterectomy. 12 To address this issue, an increasing number of recent cohort studies focusing on procedure‐specific tools are being published. Furthermore, the emerging use of automated performance metrics (APMs) has not been reflected in previous systematic reviews assessing validity in MIGS. The aim of this study was therefore to provide a comprehensive evaluation and updated review of the literature, reporting on all available assessment tools in MIGS. This evidence synthesis also appraised the reliability and validity of all reported tools including manual and automated in both simulated and live surgery.

Coi Statement

The authors have no conflicts of interest.

Materials And Methods

The protocol for the study was developed in accordance with the Preferred Reporting Items for Systematic Reviews and Meta‐Analysis guidelines (PRISMA). 13 The trial was registered with Prospective Register of Systematic Reviews (PROSPERO) (ID: CRD42022376552), a database of ongoing systematic reviews, to avoid duplication. 14 We searched for papers in the following databases: PubMed, MEDLINE, Embase and Web of Science from their inception until 17/11/2022. A broad search strategy was used (see Appendix  S1 ). The search was performed capturing terms for minimally invasive including robotic‐assisted laparoscopic gynecological procedures and assessment of performance. Finally, titles and abstracts were screened and full text eligible articles were reviewed. Eligibility for inclusion was assessed by three independent assessors (FT, JN, IM). Articles reporting only on technical skills assessment in laparoscopic or robotic‐assisted gynecological surgery were included in this review. This included randomized controlled trials, cohort and case–control studies. Furthermore, studies reporting on piloting these tools in any intraoperative, animal/wet laboratory and virtual reality (VR) simulated settings were also included. Exclusion criteria consisted of research reporting on open surgery, intrauterine and vaginal surgery, nontechnical skills tools, nonfull text available articles, abstracts or conference proceedings, pediatric studies, narrative reviews, commentary, editorials and non‐English articles. Data were independently extracted by three independent assessors (FT, JN, IM) using the Covidence online platform to aid analysis. Disagreements were resolved by discussion and if consensus could not be reached, a final decision was taken by the primary reviewer (FT). Extracted data included: study aim, study design, multi‐ and single center, number of participants, levels of participants, assessor blinding and validity evidence according to a contemporary framework of validity, which was adapted from Messick's validity framework. 15 The quality of data and risk of bias for each included study were evaluated independently by the three assessors using the modified medical education research study quality instrument (MERSQI), a checklist appraising the methodological quality of medical education research studies. 16 The possibility of performing a meta‐analysis was considered and deemed unfeasible due to the heterogeneity in data reporting on the types of tools, data collection, study design and definition of expertise (novice vs. experts). All five aspects of the contemporary framework were used to assess the validity of the assessment tools. This included content validity: testing whether the items of the objective assessment tool were relevant and represented the procedure. This is usually achieved by performing a consensus study among experts. Response process: observing how well scores reflect the observed performance. This could be achieved by providing a manual for the objective assessment tool or making sure the raters were blinded from the assessed participant. In the case of APMs, response process was always achieved because rater‐bias was not present. Internal structure: testing whether scores are reliably reproducible. This was commonly achieved by providing inter‐rater reliability (degree of agreement among multiple raters who independently assess the same surgeon) or intrarater reliability (assesses the consistency of a single rater's judgments over time). APMs cannot demonstrate rater reliability but can demonstrate internal consistency: the degree to which different items of an objective assessment tool are able to measure the same skill. Furthermore, we assessed the relationship to other variables testing whether scores correlated to clinical outcomes (predictive validity), scores from other assessment tools (concurrent validity) or level of surgical experience (construct validity). Finally, we assessed the impact using the assessment tools (consequences). This could be represented in a pass–fail score or the development of a summative or formative assessment tool. We used a scoring system rating each tool from each study, provided initially by Beckman et al. and later adjusted by Ghaderi et al., Haug et al. and Grüter et al., but modified for this systematic review. 17 , 18 , 19 , 20 Each aspect of the validity framework would count for a score from 0 to 3. The maximum score was 15: A score of 1–5 was associated with limited validity, a score of 6–10 with moderate validity and 11–15 with substantial validity. The definitions, examples and scoring system for manual and APMs (simulation) are summarized in Table  1 . Framework of validity used in this study: Manual and APMs simulation. Task analysis/hierarchical task analysis References to a previously validated tool Reference to previous validated content/items of the APM. Well defined developing process, both theoretical basis for the chosen items and systematic review by experts Multiple sources of data examining response error through critical examination of response process and respondents Rater training Note : This table has been adapted and includes the modified framework of Messick‘s validity with evidence scoring list, adopted from Beckman et al., 17 Ghaderi et al., 18 Haug et al., 19 Grüter et al., 20 further adapted for this review. Abbreviation: APMS, automated performance metrics.

Supplementary Material

Appendix S1. Table S1: Table S2.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-07-26T06:08:39.051465+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-ND-4.0