Full text
4,771 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Background: Data-extraction forms are tested before use in systematic reviews, but the guidance and error studies behind that practice concern numerical data. When the datum is a term transcribed verbatim from a source, as in evidence maps that inventory terminology, neither an agreement measure nor extraction conventions are established. We piloted, before Phase 1 of the ATLAS-ONTO project, a form built to extract operative-step terms from the surgical literature on deep endometriosis, to determine whether it produced acceptable inter-reviewer agreement and what revisions it required. Methods: Blinded inter-reviewer agreement study, reported according to GRRAS. Two reviewers independently extracted terms from 11 purposively heterogeneous articles in English and French. The primary measure was the mean per-article Jaccard index of normalized term sets, with an explicit rule for empty sets; secondary measures were Cohen's kappa for categorical attributes of exactly matching terms and simple agreement for anatomical structure. Thresholds were fixed before extraction (Jaccard >=0.70; kappa >=0.60; simple agreement >=0.80). Unmatched terms were classified retrospectively by mechanism, and the same articles were reassessed after a harmonization session. Results: Eight articles were extracted, seven with a defined index. The form failed its threshold: the mean Jaccard index was 0.581 (pooled 55/92 = 0.598). Of 92 terms, 55 matched exactly and 37 did not; the unmatched terms were attributed to source coverage (21), span extent (8), source-language comprehension (7) and one residual discrepancy. kappa conditional on the 55 matching terms was high (end definition 0.930; laterality 0.781) and anatomical-structure agreement was 0.873, so kappa alone did not detect the failure. Two record-integrity defects found during analysis are reported with their effect. Four written conventions and a split start-definition field followed; reassessment after harmonization reached 0.913, which is convergence, not independent reproducibility. Conclusions: A form that extracts terminology or verbatim text should be piloted with a set-overlap measure, a rule for empty sets and a defined admissible source, against a threshold fixed in advance. The mechanisms of discordance found here are properties of text extraction and are not specific to surgery. The revised form will be tested on new material in Phase 1.
Competing Interest Statement
The authors have declared no competing interest.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Data Availability
All data referred to in the manuscript are publicly available. The datasets supporting the conclusions of this article are deposited in the Open Science Framework registration of the study (https://doi.org/10.17605/OSF.IO/GMXZ5): the blank extraction forms (initial and revised), the extraction manual with its dated change log, the datasets of every round of the pilot (blinded round, the two exploratory exercises, the post-harmonization round, the corrected dataset and the directed audit), the audited pre-extraction and appraisal files, the data dictionary, the analysis scripts with the R version and locale guard, the Python cross-check, the consolidated results file and the complete analytical outputs before and after the two corrections described in the Methods. The corrected analysis script (version 1.1) and its output, produced after the registration and reproducing every estimate reported in the manuscript, are in the associated public OSF project (https://osf.io/vca2s/, folder "Post-registration corrections"). No individual patient or participant data were used.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.