A strict infertility diagnosis has poor agreement with the clinical diagnosis entered into the Society for Assisted Reproductive Technology registry

other OA: green public-domain-us
AI-generated summary by gemini-2.5-flash-lite, 2026-07-09

This study found that clinical infertility diagnoses do not align well with strict diagnostic criteria, highlighting the need for standardized definitions for prognosis and research.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-09 · read from full text

This study reviewed clinical records of 590 IVF patients (treated at the University of Pennsylvania from Dec 2003 to Jun 2006) and adjudicated their infertility diagnoses using stricter, evidence-based criteria, then compared these “strict” diagnoses to the diagnoses entered into the Society for Assisted Reproductive Technology (SART) registry. Agreement between clinical and strict diagnostic categories was poor overall, with the weakest concordance for uterine factor and diminished ovarian reserve and only moderate agreement for endometriosis, tubal factor, and male factor, leaving ≥20% discordance in those groups; pregnancy rates also shifted meaningfully when strict criteria were applied (with success generally decreasing for most categories except diminished ovarian reserve). The authors note the study was designed to assess diagnostic agreement and was not powered to test differences in pregnancy outcomes between groups. The paper relates to endometriosis because it reports moderate agreement between clinical and strict endometriosis diagnosis coding and discusses how endometriosis may show inconsistent IVF prognostic findings depending on how the condition is diagnosed or coded in the SART database.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Based on a recent review of the medical literature, a clinical diagnosis of infertility may not agree with strict criteria. Standardized definitions of diagnostic categories are essential for accurate patient prognosis and future research.
Full text 8,390 characters · extracted from pmc-nxml · 4 sections · click to expand

Intro

Prior to attempting In Vitro Fertilization (IVF), patients almost always express a desire to know their chance of success. The Society for Assisted Reproductive Technology (SART) makes publicly available the self reported data of participating IVF clinics throughout the United States. Patients are able to access these data via the SART website ( www.sart.org ) and find specific success rates for their age group and diagnostic category. Clinicians often refer to these rates when counseling patients on their prognosis for pregnancy. Previous authors have demonstrated that age and infertility diagnosis are strong predictors of ultimate success ( 1 , 2 ). In one population based study, older patients were found to be more likely to have “unexplained” and tubal factor infertility, while younger women are more likely to have ovulatory dysfunction or endometriosis( 3 ). Secondary infertility has also been associated with an increased chance of becoming pregnant with IVF( 4 ). Each participating SART clinic defines the specific criteria for infertility diagnoses given to their patients. In clinical practice, patients may be given one diagnosis when in fact they do not meet the strict criteria for a specific condition. Correct characterization of a patient’s etiology is essential to provide them with their true prognosis for achieving pregnancy. Furthermore, some authors have questioned whether different clinics can be appropriately compared to each other because of differences in populations, number of cycles and methods to determine diagnoses. ( 5 ). We hypothesize that the clinical criteria by which many of these diagnoses are made affects the prognostic value of the success rate quoted to patients.

Results

Charts for 590 patients were adjudicated according to strict criteria. Live birth rates for each diagnostic criterion are presented in Table 1 . The degree of agreement, represented by Kappa coefficients, between clinical and strict diagnoses was poorest among patients with diagnosis of uterine factor and diminished ovarian reserve. Strict criteria for unexplained infertility and PCOS showed slightly improved agreement with clinical criteria. While there was moderate agreement for diagnoses of endometriosis, tubal factor and Male factor, there remained 20% or greater discordance between clinical and strict diagnoses. There was at least a 3% absolute change in pregnancy rate for every diagnostic criterion when strict and clinical criteria were compared. When pregnancy rates were calculated for each diagnostic category, success rates changed by more than 15 percent for patients with uterine factor, unexplained infertility and diminished ovarian reserve. Pregnancy rates decreased when strict criteria were applied for most diagnostic categories with the exception of diminished ovarian reserve. Patients with multiple factors were less likely to achieve a pregnancy regardless criteria of which were applied; however their likelihood of pregnancy was even lower with adjudicated diagnoses. By strict criteria, these patients were significantly less likely to have a live birth than those with a single diagnosis (OR=0.61, p=0.019). This finding was similar with clinical criteria (OR=0.68, p=0.06).

Conclusions

These data provide evidence that Dthere is poor agreement between clinical infertility diagnoses and evidence-based, strict infertility diagnosis. Diagnoses with objective criteria showed higher agreement than those with subjective criteria indicating variability in a clinician’s diagnosis. Furthermore, success rates in some diagnostic categories changed markedly when strict criteria were applied. With the exception of diminished ovarian reserve, success rates dropped in all other categories. This discrepancy with DOR patients might reflect that isolated diminished ovarian reserve is actually rare and that it may not carry the same implications as DOR associated with increased age or endometriosis. Furthermore, it is important to note that patients with multiple diagnoses may have lower success than those with a single diagnosis. Given such wide variation in pregnancy rates between clinical and adjudicated diagnoses, we feel it is therefore imperative that clinicians make the most accurate diagnosis when providing their patients with an estimate of their probability of achieving pregnancy. Previous studies have examined the prognosis for pregnancy associated with specific infertility diagnoses such as tubal factor, endometriosis and PCOS ( 7 - 10 ). Our results may help explain apparent inconsistencies of studies in the literature. According to SART and previous studies, endometriosis patients have no difference in IVF success compared with other groups( 12 ). However, there are conflicting studies which suggest that endometriosis may be associated with a lower chance of success. A meta analysis published from our group confirmed that these patients have a lower chance of pregnancy in IVF and that more severe forms of endometriosis resulted in lower success( 9 ). The discrepancy between previous studies may be due to differences in how endometriosis was diagnosed or coded in the SART database. While our own small sample did not allow for subdivision of endometriosis into minimal, mild, moderate and severe subcategories, instituting these criteria into SART may prove useful in determining patient specific prognosis. Furthermore, lack of precision in underlying diagnosis may affect the validity of past and future research. It is essential that investigators work towards a standardization of diagnostic criteria for all infertility diagnoses in the manner that the Rotterdam Conferences standardized the diagnosis of PCOS( 13 ). When examining SART clinic specific success rates, it is important to examine diagnosis specific rates( 11 ). The current SART database does not establish specific criteria for each of these diagnoses, but merely offers diagnostic guidelines to participating clinics. In order to accurately compare success rates between clinics, standardization of these criteria are necessary. Careful and critical inspection of published studies in a systematic review of the literature is called for in determining which criteria have the most evidence for affecting outcome. Consensus statements such as the Rotterdam Criteria for PCOS are particularly helpful in bringing together experts in the field to establish definitive criteria for specific diagnoses. We have demonstrated how merely standardizing these criteria in our own practice affected diagnosis specific prognosis. As patients also look at these success rates in order to choose a clinic, standardized and accurate reporting becomes more important. Accurate infertility diagnoses are important to provide patients with accurate prognosis and help them in deciding how and where to best pursue fertility.

Materials|Methods

Current criteria for specific infertility diagnoses that form SART diagnostic categories were reviewed in recent medical literature. Special emphasis was given to position statements from ASRM and ESHRE as well as systematic reviews which analyze the breadth of studies available. Objective criteria for each diagnostic category were determined. IVF patients enrolled for other studies at the University of Pennsylvania between December 2003 and June 2006 had their clinical records reviewed, abstracted and their infertility diagnosis adjudicated according to these strict criteria by trained personnel. Couples were permitted to have multiple diagnoses as long as they met the minimum criteria in each category. Institutional Review Board approval T was obtained prior to chart abstraction. Adjudicated “strict” diagnoses were then compared to clinical as entered into SART. The degree of agreement between clinical and “strict” diagnoses was calculated using Kappa statistics for specific diagnostic categories and evaluated according to the method of Landis, et al ( 6 ). Clinical pregnancy rates per transfer were calculated for each stratum and compared using generalized estimating equations models, an extension of logistic regression, which adjusts for repeated measures per subject. All calculations were performed using Stata v.10, College Station, Texas. This study was designed to assess agreement between diagnostic criteria and was not powered to assess for differences between pregnancy rates of the two groups.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosisinfertility

MeSH descriptors

Diagnosis-Related Groups Diagnosis-Related Groups Infertility, Female Registries Reproductive Techniques, Assisted Adult Diagnosis-Related Groups Endometriosis Endometriosis Fallopian Tube Diseases Fallopian Tube Diseases Female Humans Infertility, Female Male Polycystic Ovary Syndrome Polycystic Ovary Syndrome Pregnancy Pregnancy Rate Prognosis

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-08-16T09:21:09.727480+00:00
pubmed
last seen: 2026-05-13T22:13:59.677786+00:00
unpaywall
last seen: 2026-05-14T19:30:52.867331+00:00
License: public-domain-us · commercial use OK · attribution required
Courtesy of the U.S. National Library of Medicine