Inter-observer reproducibility of the 2021 AAGL Endometriosis Classification

other OA: gold CC-BY-NC-ND-4.0
AI-generated summary by gemini-2.5-flash-lite, 2026-06-12

This study found the AAGL 2021 Endometriosis Classification had poor initial inter-observer agreement, but improved to good agreement after clarifying ambiguous interpretation rules.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-06, 2026-06-12 · read from full text

The study evaluated inter-observer reproducibility of the 2021 AAGL Endometriosis Classification in women with suspected endometriosis undergoing laparoscopy, using a multicentre retrospective database of 379 cases (317 included after excluding negative diagnostic laparoscopies). Three expert minimally invasive gynecologic surgeons/fellows retrospectively assigned AAGL stages from coded operative findings in two runs: first using independent interpretation of the reference paper, then after a consensus meeting that produced rules to resolve ambiguities (notably around lesion size measurement). Initial staging showed poor inter-observer agreement, with strong evidence that observer 1 more often assigned higher stages, and disagreement was attributed to interpretive variability in lesion size and how it should be measured. After consensus rules—specifically deciding to ignore lesion size and use the maximum diameter assumption—agreement improved to “good,” with no evidence of differences between observers. This paper is centrally about endometriosis—specifically the reproducibility of the 2021 AAGL endometriosis staging system and how consensus interpretation affects observer agreement.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

BACKGROUND: Inter-observer agreement for the American Association of Gynecologic Laparoscopists (AAGL) 2021 Endometriosis Classification staging system has not been described. Its predecessor staging system, the revised American Society for Reproductive Medicine (rASRM), has historically demonstrated poor inter-observer agreement. AIMS: We aimed to determine the inter-observer agreement performance of the AAGL 2021 Endometriosis Classification staging system, and compare this with the rASRM staging system. MATERIALS AND METHODS: A database of 317 patients with coded surgical data was retrospectively analysed. Three independent observers allocated AAGL surgical stages (1-4), twice. Observers made their own interpretation of how to apply the tool in the first staging allocation. Consensus rules were then developed for a second staging allocation. RESULTS: First staging allocation: odds ratio (OR) (and 95% CI) for observer 1 to score higher than observer 2 was 8.08 (5.12-12.76). Observer 1 to score higher than observer 3 was 12.98 (7.99-21.11) and observer 2 to score higher than observer 3 was 1.61 (1.03-2.51). This represents poor agreement. Second staging allocation (after consensus): OR for observer 1 to score higher than observer 2 was 1.14 (0.64-2.03), observer 1 to score higher than observer 3 was 1.81 (0.99-3.28) and observer 2 to score higher than observer 3 was 1.59 (0.87-2.89). This represents good agreement. CONCLUSIONS: These findings suggest that in its current format the AAGL 2021 Endometriosis Classification staging system has poor inter-observer agreement, not superior to the rASRM staging system. However, performance improved when additional measures were taken to simplify and clarify areas of ambiguity in interpreting the staging system.
Full text 14,512 characters · extracted from pmc-nxml · 4 sections · click to expand

Results

The database contained 379 cases with complete data. Sixty‐two were excluded for having negative laparoscopies; therefore 317 were included in the final analysis. Summary data are presented in Table  1 . In the second staging allocation, stage 1 disease represented 63.1–63.4% of cases, stage 2 was 8.2–12.6%, stage 3 was 8.5–10.1% and stage 4 was 15.1–18.3% of the cohort. Summary data – American Association of Gynecologic Laparoscopists (AAGL) 2021 Endometriosis Classification points score For the first staging allocation run, raw inter‐observer agreement data and association of a higher stage expressed as OR with 95% confidence intervals (CIs) are shown in Table  2 . OR for observer 1 to score higher than observer 2 was 8.08 (5.12–12.76). OR for observer 1 to score higher than observer 3 was 12.98 (7.99–21.11). OR for observer 2 to score higher than observer 3 was 1.61 (1.03–2.51). Therefore, there is strong evidence that observer 1 rated higher than both observers 2 and 3, and moderate evidence that assessor 2 rated higher than assessor 3. This represents poor inter‐observer agreement in the first staging allocation. Inter‐observer agreement on first staging allocation Following the first staging allocation, a consensus meeting was held to develop rules for consistency of interpretation. During this process a number of examples of ambiguity in how the AAGL system could be interpreted and applied were identified. Consensus rules for interpretation of the AAGL system were developed and are shown in Figure  1 . These seven rules covered areas that the research group felt other users of the AAGL system might also have differences of opinion on, and where further clarification within the tool was not available. 4 Regarding estimation of lesion size, a number of potential areas of ambiguity were identified. For the purpose of this research, our group concluded that the only way to ensure consistency was to ignore lesion size and assume the maximum diameter. Points of inter‐observer disagreement included the following. (i) Lesion size could refer to the sum of individual lesions, or the size of the excision required to remove multiple lesions. For example, 29 separate 1‐mm lesions could be interpreted as being 29 mm if taken individually and would therefore score two points. However, the area of excision would be expected to be far >30 mm and would therefore score four points. (ii) There is no standardised method for size measurement. For example, size could be estimated by in vivo visual estimation, crudely measured by knowing the length of laparoscopic instruments, directly measured by introducing a ruler through a port, or in vitro by measuring specimens after retrieval, for example by the histopathologist. (iii) Regarding deep lesions, size is difficult to estimate and it was not clear whether the measured area should include the diseased area visible on the peritoneum prior to excision, or the area of abnormality defined by the surgery. Where deep infiltrating endometriosis obliterates the pouch of Douglas, little or no disease might be visible initially and the tissue specimen removed might be relatively small compared to the area of reconstructed pouch of Douglas by the end of the procedure. (iv) Regarding rectal lesions, the lesion that is visible on the peritoneal surface might include a broader area of induration that is excised. Likewise, the rectal lesion measured at the level of the serosa usually involves a larger rectal muscularis lesion that can only be measured by ultrasound, or by the surgeon or histopathologist transecting the rectal specimen in vitro . (v) Where endometriosis causes an adhesion, it is not clear whether the area of attachment of the adhesion should be included in the disease measurement. (vi) Regarding endometriomata, measurement on imaging would be more accurate; however, visual estimation at the time of laparoscopy might be the intended method for this tool. The accuracy of visual estimation of endometriosis lesion size at the time of laparoscopy, and cumulative size of multiple lesions has not been reported. Visual estimation and quantification by clinicians in other disease entities is consistently inaccurate and overestimation is usually observed. For example, estimating ligament length at arthroscopy, 11 polyp size at colonoscopy , 12 as well as visual quantification of obstetric blood loss and neonatal jaundice , 13 have all demonstrated low accuracy and overestimation. It is likely that these phenomena would also occur in laparoscopic estimation of endometriotic lesion size. In order to avoid these multiple points of ambiguity, a decision was made to ignore lesion size and assume the greater size, wherever required. By removing this variable, inter‐observer variability would be improved, and consistency maintained. This would optimise performance of the AAGL system in our study. This consensus rule on size estimation would also maintain consistency between our study results, and what would be the expected real‐world application of the AAGL system where the phenomenon of overestimation would most likely apply. The consensus rules shown in Figure  1 were applied during the second staging allocation run. Raw inter‐observer agreement data and association of a higher stage expressed as OR is shown in Table  3 . OR for observer 1 to score higher than observer 2 was 1.14 (0.64–2.03). OR for observer 1 to score higher than observer 3 was 1.81 (0.99–3.28). OR for observer 2 to score higher than observer 3 was 1.59 (0.87–2.89). There was no evidence of a difference in staging between the observers. This represents good inter‐observer agreement in the second staging allocation run. Observer 1's interpretation of the AAGL system was most different from observers 2 and 3 in the first staging allocation and this difference resolved with the development and application of consensus rules. Inter‐observer agreement on second staging allocation

Discussion

There are two key implications of this research. One highlights a strength of the AAGL tool and the other highlights a weakness. Regarding the primary objective of this study, the AAGL system showed good inter‐observer agreement; however, this was only under optimised conditions. Optimisation required the development of consensus rules. The secondary aim of this study was to determine if consensus rules changed inter‐observer agreement. A significant difference was demonstrated. Taken together, these results suggest that the AAGL system has the potential to be an improvement on its predecessor, the rASRM tool, which demonstrated poor inter‐observer variability. However, there are areas of ambiguity in the AAGL system that may need further clarification before it is implemented. A number of differences of opinion were identified between the three observers in this study, and so among the broader population of potential users of the tool one could expect even more heterogeneity. Estimating lesion size is particularly subjective and problematic. This retrospective study used descriptive data rather than laparoscopic images to simulate endometriosis staging, and therefore is not as representative as real‐time intraoperative staging. A prospective study where scoring and staging is performed in the usual contemporaneous fashion would overcome this. In order to overcome multiple dimensions of subjectivity in lesion size estimation, the greatest lesion size was assumed. This was to ensure consistency, to avoid underestimating disease severity and to ensure the AAGL system was given optimised conditions to perform. A similar methodology was used in a previously published paper on endometriosis staging, with the same rationale. 7 Another limitation is that three observers were utilised in the experiment. Using more than three observers would strengthen the generalisability of the study. A strength of this paper was the use of a consensus meeting to ensure optimal conditions for assessing the AAGL tool. While the process was thorough and robust, there are two important cautions in interpreting the results under these conditions. Firstly, these consensus rules do not reflect the current real‐world application of the AAGL system. Secondly, the consensus rules formulated by this research group were for the purposes of this research only. They are not endorsed by the authors of the AAGL system and might not necessarily reflect how the tool was intended to be used. Other users might interpret the tool in a different manner. It is possible that the performance of the older rASRM system would also be improved by developing similar consensus rules in a similar manner, although that was not the focus of this paper. To our knowledge, this is the first paper to assess the AAGL system for inter‐observer agreement. The AAGL system has the potential to be an improvement on the rASRM in these two areas, but not in its current format. Further assessment of the AAGL system should follow, ideally with a large prospective external validation.

Introduction

Until recently the most widely known endometriosis staging tool was the revised American Society for Reproductive Medicine (rASRM) classification. It has been criticised for failing to correlate with surgical complexity and with clinical outcomes relevant to the patient, like pain and fertility. 1 , 2 It has also demonstrated poor inter‐observer agreement. 3 The American Association of Gynecologic Laparoscopists (AAGL) 2021 Endometriosis Classification was announced in November 2021 with the stated aim of replacing the rASRM with an improved staging system. The primary objectives were correlation with surgical complexity, and usability. The secondary objective was correlation with baseline clinical factors. 4 The tool uses a schema of disease location, morphology and size that generates cumulative points that are then applied to predetermined thresholds for allocating stage. A separate four‐level surgical complexity scale was developed to validate the AAGL system. A large prospective cohort study demonstrated high concordance between AAGL stage and surgical complexity level, 4 while the only external validation to date failed to reproduce this result. 5 Inter‐observer agreement was not assessed. Herein we will refer to the 2021 AAGL Endometriosis Classification as a whole as the AAGL system, the cumulative raw point score as AAGL points, the 1–4 stages subsequently determined by those points as AAGL stage, and the original paper by Abrão et al . 4 as the reference paper. The primary aim of this study was to determine inter‐observer agreement for AAGL stage. The secondary aim was to see if this changes after consensus rules are developed for interpreting the AAGL system.

Materials And Methods

We performed a multicentre retrospective inter‐observer agreement study on women with suspected endometriosis. The study was approved by the Nepean Blue Mountains Local Health District ethics committee; 2022/ETH00000. A database of 379 cases from four previous published studies on endometriosis 6 , 7 , 8 , 9 was utilised in this study. Participants were seen between January 2016 and October 2021. The patients underwent laparoscopy with one of seven minimally invasive gynaecological surgeons in metropolitan Sydney, Australia. The gynaecological surgeons all had the highest level of surgical ability, as per the Royal Australian New Zealand College of Obstetricians and Gynaecologists/Australasian Gynaecological Endoscopy & Surgery Society (RANZCOG/AGES). 10 Systematic visual inspection of the pelvis, upper abdomen, and appendix was performed on each participant. Comprehensive, coded data mapping the location and morphology of disease, as well as the surgical procedures performed for each case were therefore available. This made retrospective allocation of AAGL stage and AAGL level possible. Women with negative diagnostic laparoscopy for biopsy‐proven endometriosis were excluded. Participants were de‐identified. Inclusion criteria from the original studies were women of reproductive age and either a history of chronic pelvic pain, or a history of endometriosis, or both. Exclusion criteria were women with suspected malignancy, pregnancy, premature ovarian failure or menopause. The database was interrogated to ensure there were no duplications or cases with incomplete data. All women included in the database had previously consented for their de‐identified surgical data to be used in research. Three expert observers were either qualified in minimally invasive gynaecologic surgery (MIGS) or Fellows in their final year of MIGS training. The three observers were asked to allocate an AAGL stage for each case, using the point scoring schema described in the reference paper. Staging of each case was not performed in real time, or with the aid of intraoperative images. Rather, the aforementioned database contained prospectively and systematically collected surgical findings. Detailed mapping and coding of each case was therefore available. This made it possible for staging using the AAGL system to be performed retrospectively, in a simulated manner. Staging was performed twice. In the first staging allocation run, the three observers were asked to read the reference paper and make their own interpretation of how to apply the AAGL system. Following this, a consensus meeting was held between the three observers and other contributing authors on the project. Wherever potential areas of ambiguity in the AAGL system were identified, consensus rules were agreed on (Fig.  1 ). The three observers were then asked to allocate an AAGL stage for each case, using the consensus rules. The second staging allocation run was performed one week after the first. The observers were blinded to each other's results and to their own previous results. Consensus rules for interpreting the 2021 American Association of Gynecologic Laparoscopists (AAGL) 2021 Endometriosis Classification. Inter‐assessor agreement for the ordinal AAGL stage ratings was by mixed linear model with each participant treated as a random effect. Odds ratio (OR) was calculated for estimating the probability of a higher versus lower rating, comparing the assessors and assessment points. SAS version 9.4 was used. This research was ethically approved by the Nepean Blue Mountains Local Health District Human Research Ethics Committee under reference number 2022/ETH00000, on 4/2/2022.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Outcome instruments

rASRM

Condition tags

endometriosis

MeSH descriptors

Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

SciLite annotations

organisms 1
human

Source provenance

europepmc
last seen: 2026-08-30T09:23:35.175841+00:00
pubmed
last seen: 2026-08-30T06:06:32.752556+00:00
scilite
last seen: 2026-05-18T04:26:01.642840+00:00
unpaywall
last seen: 2026-05-14T19:30:52.867331+00:00
License: CC-BY-NC-ND-4.0 · commercial use OK · attribution required
Courtesy of the U.S. National Library of Medicine