Results
A total of 19 746 titles and abstracts were screened for their eligibility and four additional studies were identified through other sources (citation searching n = 4, gray literature n = 0). A total of 174 articles were included for full text review and 102 studies were deemed ineligible, primarily because these studies were not reporting solely on MIGS or because no assessment tools were reported in those studies. Finally, 72 studies were included for further analysis. A breakdown of inclusion and exclusion is shown in the PRISMA diagram (Figure 1 ).
PRISMA flow diagram.
Study characteristics are summarized in Table 2 . Studies were predominately conducted in the USA (48.6%) or Europe (25.0%); however, authors were represented from five different continents (all except South America and Antarctica). Studies were published between 2002 and 2023. Included studies consisted of manual assessment tools ( n = 26) and APMs ( n = 11). The later ones were tools from which the scoring system was directly derived from kinematic data and systems events data, usually in a VR setting. A total of 36 out of 72 (50%) studies were designed to address the utilization of previously validated tools in an educational intervention setting, followed by 31 (43.7%) studies aimed to either develop a new tool or assess the validity of existing ones.
Characteristics of 72 included studies reporting on objective assessment tools in minimally invasive gynecological surgery.
With the exception of one paper, scoring five points,
21
all other papers had a score ranging from 10 to 16.5 on the 18‐MERSQI checklist. The main limitations were lack of randomized controlled trials (study design), lack of multicenter studies (sampling) and the absence of correlating study outcomes with clinical outcomes. The risk of bias per tool can be found in Tables 3 , 4 , 5 , 6 , 7 and risk of bias per study can be found in Table S1 .
Validity evidence per tool, manual global assessment tools (1/2).
Intraoperative
Dry laboratory
Intraoperative
Wet laboratory
Dry laboratory
Intraoperative
Wet laboratory
Dry laboratory
Virtual Reality
Intraoperative
Dry laboratory
Virtual Reality
Intraoperative
Dry laboratory
Robotic supracervical hysterectomy
Laparoscopic colpotomy
Total laparoscopic hysterectomy
Laparoscopic myomectomy
Laparoscopic colpotomy
Laparoscopic vaginal cuff suturing
Laparoscopic sacrocolpopexy
Laparoscopic supracervical hysterectomy
Laparoscopic myomectomy
Laparoscopic oophorectomy, dissection and ligature of uterine artery
Diagnostic laparoscopy
Laparoscopic sacrocolpopexy
Laparoscopic bilateral tubal ligation
Laparoscopic salpingectomy
Total laparoscopic hysterectomy
Laparoscopic suturing and intracorporeal knot tying
Basic dry laboratory tasks
Laparoscopic ovarian cystectomy
Total laparoscopic hysterectomy
Laparoscopic bilateral tubal ligation
Laparoscopic salpingo‐oophorectomy
Laparoscopic vaginal cuff closure
Bilateral tubal ligation
Basic dry laboratory tasks
Total laparoscopic hysterectomy
Laparoscopic Salpingectomy/salpingo‐oophorectomy
Delphi consensus developed tool
Referred to previous literature for content validity
Response process present in all studies except for one study
Inter rater reliability:
Strong–Excellent
Intrarater reliability: Not assessed
Item analysis: Not assessed
Inter rater reliability: Good–Excellent
Intrarater reliability: Good
Item analysis: Not assessed
Inter‐rater reliability: Good
Intrarater reliability: Not assessed
Inter item analysis: Excellent
Inter‐rater reliability: Weak–excellent
Intrarater reliability: Not assessed
Item analysis: Not assessed
Inter‐rater reliability: Acceptable
Intrarater reliability: Excellent
Inter item analysis: not assessed
Inter‐rater reliability: Good
Intrarater reliability: Very strong
Inter item analysis: Excellent
Relations to other variables
Construct validity (training level or case experience)
Concurrent validity (other performance scores)
Predictive validity (relation to clinical outcomes)
Construct validity
Concurrent validity
Construct validity
Concurrent validity
Pass mark score defined at
27/35
82/85
32/40
However, not applied for benchmarking/credentialing in training curriculum
Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence.
103
Abbreviations: GOALS, global operative sssessment of laparoscopic skills; GRS, global rating skills; m, modified; MERSQI, medical education research study quality
16
; OSATS, objective structured assessment of technical skills; R‐OSATS, robotic objective structured assessment of technical skills.
Validity evidence per tool, manual global assessment tools (2/2).
Intraoperative
Dry laboratory
Dry laboratory
Virtual reality
Laparoscopic salpingectomy/salpingo‐oophorectomy
Laparoscopic resection of endometriosis, adhesiolysis and ovarian drilling
Relations to other variables
Construct (training level or case experience)
Concurrent validity (other performance scores)
Predictive validity (relation to clinical outcomes)
Pass mark defined at 27/35
However, not applied for benchmarking/credentialing in training curriculum
Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence.
103
Abbreviations: GEARS, global evaluative assessment of robotic skills; GERT, generic error rating tool; GRIT, global rating index of technical skills; GSAT, generic skills assessment tool; m, modified; MERSQI, medical education research study quality
16
; OPRS, operating performance rating system; TCPE, time to correct performed exercise.
Validity evidence per tool, manual procedure‐specific assessment tools.
Intraoperative
Virtual reality
Delphi consensus to develop tool
Referred to previous literature
Experts developed task‐analysis
Referred to previous literature for content validity
Inter‐rater reliability: Excellent
Fair–excellent
Relations to other variables
Construct validity (training level or case experience)
Concurrent validity (other performance scores)
Predictive validity (relation to clinical outcomes)
Concurrent validity
Construct validity
Pass mark defined (90/150)
However, not applied for benchmarking/credentialing in training curriculum
Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence.
103
Abbreviations: CAT‐LSH, competency assessment tool for laparoscopic supracervical hysterectomy; DART, dissection of assessment of robotic technique; H‐OSATS, objective scale for assessment of technical skills of total laparoscopic hysterectomy (TLH); LSO‐OSATS, OSATS of laparoscopic salpingo‐oophorectomy; MERSQI, medical education research study quality
16
; OSA‐LS, objective structured assessment of laparoscopic salpingectomy; RHAS, robotic hysterectomy assessment score.
A total of 26 manual technical skills assessments were included. These consisted of 13 global rating tools, and 13 procedure‐specific tools. Furthermore 11 APMs were identified. The results are summarized under three categories: global, procedure‐specific and automated metrics tools.
The OSATS, global operative assessment of laparoscopic skills (GOALS), global evaluative assessment of robotic skills (GEARS), global rating scale (GRS—not further specified) and modified versions of these tools were used most frequently ( n = 32) in the included studies. With the exception of three tools (robotic‐OSATs, modified GEARS, operative performance rating system [OPRS]), all other tools were validated intraoperatively. Other settings included wet and dry models. The (modified) OSATS was the only manual tool assessed in a VR setting.
22
,
23
The only error rating tool identified was the generic error rating tool (GERT).
24
Kilani proposed the global rating index of technical skills (GRITS) to intraoperatively assess the correlation of surgical skill performance scores between expert assessment and self‐assessment in various laparoscopic gynecological procedures, concluding that self‐assessments have a higher evaluation than expert assessments.
25
Finally, the OPRS was used in a dry laboratory setting, assessing robotic‐assisted laparoscopic radical hysterectomy and pelvic lymphadenectomy performance.
26
A total of 10 procedure‐specific tools were assessed intraoperatively. The robotic sacrocolpopexy simulation model,
18
the “assessment tool for total laparoscopic hysterectomy”
27
and the OSATS for laparoscopic suturing and intracorporeal knot tying
28
were assessed in a laboratory‐based setting. The objective structured assessment of laparoscopic salpingectomy (OSA‐LS) was based on both the original OSATS and a modified rating scale for laparoscopic cholecystectomy developed by Grantcharov et al.
29
The myTIPreport is a smartphone application where both the trainee performing the procedure and a faculty member assessed the technical skills on a checklist immediately after the procedure.
30
The laparoscopic salpingo‐oophorectomy‐OSATS (LSO‐OSATS) was based on the OSA‐LS but consisted of fewer items (6 in the LSO‐OSATS vs. 10 in the OSA‐LS). Remarkably, six minimal invasive hysterectomy procedure‐specific tools were included: Objective scale for assessment of technical skills of TLH (H‐OSATS), the objective structured assessment of TLH (OSA‐TLH), laparoscopic hysterectomy‐OSATS (LH‐OSATS) and the assessment tool for TLH, competency assessment tool for laparoscopic supracervical hysterectomy (CAT‐LSH) and the robotic hysterectomy assessment score (RHAS).
22
,
27
,
31
,
32
,
33
,
34
All manual assessment tools and studies are summarized in Tables 3 , 4 , 5 , 6 .
Validity evidence per tool, manual procedure‐specific assessment tools.
Relations to other variables
Construct validity (training level or case experience)
Concurrent validity (other performance scores)
Predictive validity (relation to clinical outcomes)
Pass mark defined (29.3/55)
However, not applied for benchmarking/credentialing in training curriculum
Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence.
103
Abbreviations: MERSQI, medical education research study quality
16
; OSA‐TLH, OSATS of TLH, LH‐OSATS is OSATS of TLH.
A total of 11 APMs were identified in this systematic review. These include APMs in robotic‐assisted laparoscopic VR simulations; da Vinci Surgical Simulation, RobotiX Mentor Simulation, and laparoscopic VR simulations: LapSim, LapMentor, MIST, SurgicalSim, MISTELS, TRLCD05, FastTrack and LapVR simulator. No APMs were used intraoperatively.
The following part of this review focuses on the sources of validity evidence of each included assessment tool, specified on the unitary framework for manual tools and APMs (Table 1 ). Given the heterogeneity of the interventions being investigated, each tool was categorized along the five dimensions of the contemporary framework (Tables 3 , 4 , 5 , 6 , 7 ; Table S2 ).
Validity evidence per tool, automated performance metrics.
Basic laparoscopic VR modules
Laparoscopic Salpingectomy
Basic laparoscopic VR modules
Laparoscopic Salpingectomy/Salpingo‐oophorectomy
Relations to other variables
Construct validity (training level or case experience)
Concurrent validity (other performance scores)
Predictive validity (relation to clinical outcomes)
Both studies showed construct validity
Concurrent validity
Pass mark defined at 75/110
48
However, not applied for benchmarking/credentialing in training curriculum
Note : Level of evidence according to the 2011 Oxford CEBM Levels of Evidence.
103
Abbreviations: DvSS, Da Vinci surgical system; LapSim, laparoscopic simulator; LapVR simulator, laparoscopic virtual reality simulator; MERSQI, medical education research study quality
16
; MIST, minimally invasive surgical trainer, McGill inanimate system for training and evaluation of laparoscopic skills; VBLAST‐PT, virtual basic laparoscopic skill trainer.
A total of 8 out of 13 (61.5%) global assessment tools had content validity.
30
,
34
,
35
,
36
,
37
,
38
,
39
,
40
,
41
The studies reporting on the generic skills assessment tool, GERT and OPRS, did not demonstrate content validity. In contrast to the global rating tools, content validity for the procedure‐specific tools was provided 12/13 (92.3%) studies. A total of 10 (76.9%) tools underwent a (hierarchical) task analysis,
21
,
22
,
27
,
28
,
31
,
32
,
33
,
42
,
43
,
44
,
45
,
46
,
47
,
48
,
49
,
50
usually followed by a consensus study with experts.
Content validity was demonstrated in 6 (54.5%) different APMs.
51
,
52
,
53
,
54
,
55
,
56
One APM study reported on consensus methodology to reach content validity.
56
Eight out of 13 global assessment tools (61.5%) provided evidence of rater training, either by a training session or by providing a manual for tool usage.
24
,
25
,
26
,
28
,
30
,
37
,
38
,
40
,
41
,
57
,
58
This was applicable to seven out of 12 (58.3%) procedure‐specific tools.
31
,
32
,
33
,
42
,
43
,
44
,
45
,
46
,
47
,
48
,
49
,
50
,
59
,
60
Addison et al. used crowd‐sourced assessment of technical skills (CSATS) for GEARS and raters were routinely trained and evaluated for their rating reliability.
41
All APMs inherently demonstrate response process, as they are automated, hence removing rater bias, and theoretically having 100% reliability.
Internal structure was assessed in different ways. The most common reported form was inter‐rater reliability with 10 out of 13 (76.9%) global rating tools and 10 out of 13 (76.9%) procedure‐specific tools. All manual global tools report good to excellent inter‐rater reliability,
23
,
24
,
25
,
30
,
33
,
35
,
37
,
38
,
39
,
40
,
57
,
61
,
62
,
63
,
64
,
65
,
66
,
67
,
68
,
69
,
70
with the exception of one modified OSATS in a dry laboratory setting for a simple laparoscopic ovarian cystectomy by Chahine et al.
71
Inter‐rater reliability among procedure‐specific tools
31
,
32
,
33
,
42
,
43
,
44
,
46
,
47
,
59
,
60
showed good to excellent correlation except for specific domains for the H‐OSATS and the dissection assessment for robotic technique.
43
,
60
Intrarater reliability was reported for 4/13 (30.8%) global tools demonstrating excellent intrarater reliability.
28
,
30
,
38
,
40
The H‐OSATS was the only procedure‐specific tool reporting excellent intrarater reliability.
31
,
60
Only one APM study calculated internal structure, reporting poor internal consistency (Cronbach's alpha 0.58) on the RobotiX Mentor Simulator.
55
A total of 10 out of 13 (79.6%) global tools reported relationships to other variables by either comparing novices to experts (construct validity) or showing significant correlation between scoring outcomes and other performance assessment tools, considered the gold standard (concurrent validity).
22
,
23
,
24
,
27
,
28
,
30
,
35
,
36
,
38
,
39
,
40
,
57
,
58
,
61
,
62
,
63
,
64
,
65
,
66
,
67
,
68
,
69
,
70
,
71
,
72
,
73
,
74
,
75
,
76
,
77
Nine out of 13 (69.2%) procedure‐specific tool studies
31
,
32
,
33
,
42
,
43
,
44
,
46
,
47
,
48
,
49
,
50
,
60
showed construct or concurrent validity. Nine out of 11 (81.8%) APM showed construct or criterion validity.
51
,
52
,
53
,
55
,
56
,
67
,
72
,
78
,
79
,
80
,
81
,
82
,
83
,
84
,
85
,
86
,
87
,
88
,
89
,
90
None of the included studies reported on the association between intraoperative performance of practicing surgeons to clinical/postoperative outcomes of patients (predictive validity).
Four out of 26 (15.4%) manual (global and procedure‐specific tools) provided benchmark scores.
30
,
32
,
35
,
57
,
60
,
64
,
67
One out of 11 (9.1%) APMs provided a benchmark score on the RobotiX Mentor providing a pass/fail score of 75/110.
55
Table 8 summarizes the evidence of validity of all tools based on the scoring tool from Table 1 . Only one manual tool showed substantial evidence of validity (score 11–15): the total laparoscopic hysterectomy procedure specific tools: OSA‐TLH. The RobotiX Mentor Simulator was the only APM showing substantive evidence. A total of 17 tools showed moderate evidence (score 6–10) of which 10 were global tools, six were procedure specific and one APMs. Finally, 18 tools showed limited evidence of validity, of which nine were APMs, three manual global tools and six procedure specific tools.
Objective assessment tools arranged by strength of validity based on the validity evidence scoring list from Table 1 (substantial, moderate and limited evidence).
Discussion
This systematic review of 72 articles identified 37 surgical performance assessment tools that have been studied in a laparoscopic and robotic‐assisted gynecological surgery setting. This review provided a comprehensive evaluation of the validity and reliability of assessment tools, using a contemporary validity framework (Table 1 ). These included 26 manual tools and 11 APMs. Interestingly, none of the studies were able to show predictive validity (correlating the tool score with clinical outcomes).
Tough achieving predictive validity often necessitates a more demanding endeavor, and there is still a significant opportunity to develop study settings correlating tool scores with clinical outcomes.
91
,
92
The General Medical Council (GMC) in the UK has even stated that in the absence of the gold standard, exploring the strength of the relationship between similar established assessment tools, from different surgical specialities, might offer itself as an alternative.
93
Furthermore, more granular analysis of surgical skills, such as the objective clinical human reliability analysis (OCHRA) could enhance the likelihood of achieving predictive validity, associating technical kills with clinical outcomes, regardless of level of expertise.
94
When looking at current training programs, such as the Royal College of Obstetrics and Gynecology (RCOG) in the UK, it interesting to see that the most frequent used objective assessment tool is the OSATS.
95
Global assessment tools are easily available for different procedures. However, this systematic review showed that the only manual tool showing substantive evidence was a procedure specific tool. It should also be noted that the exchange of constructive feedback within the trainer‐trainee dialogue often plays a greater role in shaping learning outcomes.
Culligan et al. proposed a robotic surgery simulation training curriculum and established predictive validity by demonstrating a correlation between program completion and improved live surgery outcomes.
72
These included reduced estimated blood loss, shorter operating times, and enhanced intraoperative GOALS scores. However, the study's generalizability was limited by its restriction to board‐certified obstetrics and gynecology surgeons. Despite other studies reporting pass/fail scores for a modified GOALS, modified GEARS, H‐OSATS, OSA‐TLH (both TLH procedure specific tool), the RobotiX Mentor simulator and DvSS, none of them showed any evidence of successful implementation of curriculums for credentialing. Future research should not only focus on investigating other aspects of validity, but also on benchmarking already available objective assessment tools to make them useful additions to surgical national curriculums. This will ultimately enhance the standardization and effectiveness of resident and fellow training in MIGS.
This systematic review had some limitations. First, it was limited by only including studies in English. Another limitation was that the majority of the studies were small, conducted once or in a nonrandomized single center setting, risking potential biases and compromising reproducibility of results. Often, different thresholds and definitions were used, producing heterogeneity and the subsequent inability to perform a meaningful meta‐analysis, highlighting that tools should be evaluated more thoroughly in large, well‐run studies. Furthermore, assessment tools for intrauterine and vaginal surgery were not included.
However, a significant increase in numbers of assessment tools ( n = 37) in MIGS were identified, making it, to our knowledge, the most comprehensive and detailed systematic review on the subject of minimal invasive gynecology surgery. It is important to inform the gynecological surgical community of all available tools that can be applied not only in the research settings but to support learning and teaching.
Ferriss et al., published a systematic review of intraoperative assessment tools in MIGS, focusing mainly on manual performance metrics.
96
They concluded that procedure‐specific tools are more thoroughly evaluated, however described their limited use due to poor quality studies and borderline reliability. Another scoping review by Hennings et al. explained that most surgical assessment scales were validated in simulation settings, compromising transferability to the operating room.
12
However, comprehensive evaluations of the tools' validity were not reported, mainly lacking the consequence component.
This review also provided a comprehensive review of APMs available in MIGS. One of the advantages of APMs compared to manual tools is the automated collection of the performances, preventing rater bias. Furthermore, it is less time consuming and manual tools require a degree of subjective scoring. Furthermore, research has been suggested that skills in laparoscopic surgery can be increased by proficiency‐based procedural VR simulator training. However, this review showed that the majority of APMs (81.8%) has limited validity evidence. This low level of validity hinders transferability outside of the research environment. Moreover, these metrics alone cannot be considered substitutes of experts' input towards surgical competencies. Until true objective assessment tools are in place to provide expert opinion on trainees' performances within the clinical context, APMs are useful adjunct to support objective assessments of surgical skills. This systematic review did not identify APMs using kinematic data from live surgery.
However, recent studies in different specialities have been able to correlate kinematic data derived from recording devices with technical performance to create a scoring index in dry lab surgeries and live operations. Lyman et al., were able to correlate the operating robotic index model to level of experience in a dry laboratory robotic‐assisted laparoscopic hepaticojejunostomy reconstruction, showing construct validity.
97
Another example of appliance is the emerging interdisciplinary field of surgical data science (SDS) aiming to improve quality of interventional healthcare by capturing, organizing, analyzing, and modeling data. Mascagni et al., were able to assess the critical view of safety criteria in laparoscopic cholecystectomy through annotating anatomical segments and training a deep neural network to predict critical view of safety occurrence.
98
Utilization of these tools, including kinematic data from advanced computing devices and surgical systems and artificial neural networks will become essential to better understand factors in surgical performance and ultimately standardize safe operation. One key challenge for developing these approaches further is the current absence of large‐scale datasets that fully represent the domains of variation; for example, experience level, subtask, instrument type, in order to allow robust training of AI models with only limited clinical datasets currently available.
99
,
100
Future validation of APMs will support utilization while they are likely to expand in the future with artificial intelligence and machine learning.
Materials And Methods
The protocol for the study was developed in accordance with the Preferred Reporting Items for Systematic Reviews and Meta‐Analysis guidelines (PRISMA).
13
The trial was registered with Prospective Register of Systematic Reviews (PROSPERO) (ID: CRD42022376552), a database of ongoing systematic reviews, to avoid duplication.
14
We searched for papers in the following databases: PubMed, MEDLINE, Embase and Web of Science from their inception until 17/11/2022. A broad search strategy was used (see Appendix S1 ). The search was performed capturing terms for minimally invasive including robotic‐assisted laparoscopic gynecological procedures and assessment of performance. Finally, titles and abstracts were screened and full text eligible articles were reviewed.
Eligibility for inclusion was assessed by three independent assessors (FT, JN, IM). Articles reporting only on technical skills assessment in laparoscopic or robotic‐assisted gynecological surgery were included in this review. This included randomized controlled trials, cohort and case–control studies. Furthermore, studies reporting on piloting these tools in any intraoperative, animal/wet laboratory and virtual reality (VR) simulated settings were also included. Exclusion criteria consisted of research reporting on open surgery, intrauterine and vaginal surgery, nontechnical skills tools, nonfull text available articles, abstracts or conference proceedings, pediatric studies, narrative reviews, commentary, editorials and non‐English articles.
Data were independently extracted by three independent assessors (FT, JN, IM) using the Covidence online platform to aid analysis. Disagreements were resolved by discussion and if consensus could not be reached, a final decision was taken by the primary reviewer (FT). Extracted data included: study aim, study design, multi‐ and single center, number of participants, levels of participants, assessor blinding and validity evidence according to a contemporary framework of validity, which was adapted from Messick's validity framework.
15
The quality of data and risk of bias for each included study were evaluated independently by the three assessors using the modified medical education research study quality instrument (MERSQI), a checklist appraising the methodological quality of medical education research studies.
16
The possibility of performing a meta‐analysis was considered and deemed unfeasible due to the heterogeneity in data reporting on the types of tools, data collection, study design and definition of expertise (novice vs. experts).
All five aspects of the contemporary framework were used to assess the validity of the assessment tools. This included content validity: testing whether the items of the objective assessment tool were relevant and represented the procedure. This is usually achieved by performing a consensus study among experts. Response process: observing how well scores reflect the observed performance. This could be achieved by providing a manual for the objective assessment tool or making sure the raters were blinded from the assessed participant. In the case of APMs, response process was always achieved because rater‐bias was not present. Internal structure: testing whether scores are reliably reproducible. This was commonly achieved by providing inter‐rater reliability (degree of agreement among multiple raters who independently assess the same surgeon) or intrarater reliability (assesses the consistency of a single rater's judgments over time). APMs cannot demonstrate rater reliability but can demonstrate internal consistency: the degree to which different items of an objective assessment tool are able to measure the same skill. Furthermore, we assessed the relationship to other variables testing whether scores correlated to clinical outcomes (predictive validity), scores from other assessment tools (concurrent validity) or level of surgical experience (construct validity). Finally, we assessed the impact using the assessment tools (consequences). This could be represented in a pass–fail score or the development of a summative or formative assessment tool. We used a scoring system rating each tool from each study, provided initially by Beckman et al. and later adjusted by Ghaderi et al., Haug et al. and Grüter et al., but modified for this systematic review.
17
,
18
,
19
,
20
Each aspect of the validity framework would count for a score from 0 to 3. The maximum score was 15: A score of 1–5 was associated with limited validity, a score of 6–10 with moderate validity and 11–15 with substantial validity. The definitions, examples and scoring system for manual and APMs (simulation) are summarized in Table 1 .
Framework of validity used in this study: Manual and APMs simulation.
Task analysis/hierarchical task analysis
References to a previously validated tool
Reference to previous validated content/items of the APM.
Well defined developing process, both theoretical basis for the chosen items and systematic review by experts
Multiple sources of data examining response error through critical examination of response process and respondents
Rater training
Note : This table has been adapted and includes the modified framework of Messick‘s validity with evidence scoring list, adopted from Beckman et al.,
17
Ghaderi et al.,
18
Haug et al.,
19
Grüter et al.,
20
further adapted for this review.
Abbreviation: APMS, automated performance metrics.