End-to-End Reliability of Automated Systems for Diagnostic Evidence Extraction: A Prospective Benchmark Study

preprint OA: closed CC-BY-4.0
AI-generated deep summary by claude@2026-07, 2026-07-04 · read from full text

This prospective benchmarking study evaluated the end-to-end reliability of automated diagnostic evidence extraction systems (MedNuggetizer and three LLMs) using a locked corpus derived from a diagnostic accuracy meta-analysis of Uromonitor and urine cytology for bladder cancer. Across 16 datasets from 8 publications, systems were run 20 repeated times without browsing using a conservative adjudicated human benchmark, with primary evaluation based on run-level correctness defined as exact 2×2 contingency-table extraction for derivable datasets or explicit abstention for non-derivable datasets. MedNuggetizer (312/320 correct runs; 97.50%) and Claude Opus 4.5 (313/320; 97.81%) exceeded the prespecified reliability threshold, with MedNuggetizer showing perfect abstention safety on non-derivable datasets and high repeatability, while other LLMs did not meet the threshold. The paper does not address whether these findings generalize beyond the specific diagnostic-corpus and locked, no-browsing conditions, and it is a preprint that has not been peer reviewed. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Background Accurate data extraction remains a bottleneck in evidence synthesis. In diagnostic test accuracy research, valid quantitative synthesis depends on reconstruction of complete 2 × 2 contingency tables, and extraction errors can distort downstream estimates. Large language models (LLMs) have shown promise for literature-based extraction, but their reliability under locked, conservative benchmark conditions remains insufficiently defined. Methods In this prospective benchmarking study, the evaluation corpus was derived from a diagnostic accuracy meta-analysis of Uromonitor and urine cytology. The corpus comprised 16 datasets from 8 publications. Based on an adjudicated human benchmark, 11 datasets were classified as derivable and 5 as non-derivable. MedNuggetizer, ChatGPT-5.2, Claude Opus 4.5, and Gemini 3 Pro were each evaluated across 20 repeated runs under a locked prompt without browsing, yielding 320 dataset-run observations per system. The primary endpoint was run-level correctness, defined as exact extraction for derivable datasets or explicit abstention for non-derivable datasets. Secondary endpoints included cell-level exactness, safety on non-derivable datasets, repeatability, metric fidelity, and efficiency relative to human extraction. Results MedNuggetizer achieved 312 of 320 correct dataset-runs (97.50%) and Claude Opus 4.5 achieved 313 of 320 (97.81%); both exceeded the prespecified benchmark threshold of 0.95, with one-sided 97.5% lower confidence bounds of 95.13% and 95.55%, respectively. ChatGPT-5.2 achieved 308 of 320 correct dataset-runs (96.25%), and Gemini 3 Pro achieved 301 of 320 (94.06%); neither met the threshold criterion. Claude Opus 4.5 showed the highest exact-match rate on derivable data (859/880, 97.61%), followed by MedNuggetizer (853/880, 96.93%). MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro produced no hallucinated numeric outputs on non-derivable datasets, whereas Claude Opus 4.5 produced 1 such event in 100 non-derivable dataset-runs. Repeatability was high across automated systems (Gwet’s AC1 0.929–0.958); inter-operator agreement among 4 human raters was lower (AC1 0.608). Median execution times ranged from 6.7 to 14.0 minutes, compared with 41.9 minutes for human extraction. Conclusions Under locked, conservative benchmark conditions, MedNuggetizer and Claude Opus 4.5 showed high reliability for diagnostic data extraction in evidence synthesis. Claude Opus 4.5 was strongest on extraction correctness, whereas MedNuggetizer combined threshold-level correctness with perfect abstention safety, strong repeatability, and efficiency gains over manual extraction.
Full text 182,054 characters · extracted from preprint-html · click to expand
End-to-End Reliability of Automated Systems for Diagnostic Evidence Extraction: A Prospective Benchmark Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article End-to-End Reliability of Automated Systems for Diagnostic Evidence Extraction: A Prospective Benchmark Study Anton Kravchuk, Julio Ruben Rodas Garzaro, Gregor Donabauer, Samy Ateia, and 10 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9260490/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 8 You are reading this latest preprint version Abstract Background Accurate data extraction remains a bottleneck in evidence synthesis. In diagnostic test accuracy research, valid quantitative synthesis depends on reconstruction of complete 2 × 2 contingency tables, and extraction errors can distort downstream estimates. Large language models (LLMs) have shown promise for literature-based extraction, but their reliability under locked, conservative benchmark conditions remains insufficiently defined. Methods In this prospective benchmarking study, the evaluation corpus was derived from a diagnostic accuracy meta-analysis of Uromonitor and urine cytology. The corpus comprised 16 datasets from 8 publications. Based on an adjudicated human benchmark, 11 datasets were classified as derivable and 5 as non-derivable. MedNuggetizer, ChatGPT-5.2, Claude Opus 4.5, and Gemini 3 Pro were each evaluated across 20 repeated runs under a locked prompt without browsing, yielding 320 dataset-run observations per system. The primary endpoint was run-level correctness, defined as exact extraction for derivable datasets or explicit abstention for non-derivable datasets. Secondary endpoints included cell-level exactness, safety on non-derivable datasets, repeatability, metric fidelity, and efficiency relative to human extraction. Results MedNuggetizer achieved 312 of 320 correct dataset-runs (97.50%) and Claude Opus 4.5 achieved 313 of 320 (97.81%); both exceeded the prespecified benchmark threshold of 0.95, with one-sided 97.5% lower confidence bounds of 95.13% and 95.55%, respectively. ChatGPT-5.2 achieved 308 of 320 correct dataset-runs (96.25%), and Gemini 3 Pro achieved 301 of 320 (94.06%); neither met the threshold criterion. Claude Opus 4.5 showed the highest exact-match rate on derivable data (859/880, 97.61%), followed by MedNuggetizer (853/880, 96.93%). MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro produced no hallucinated numeric outputs on non-derivable datasets, whereas Claude Opus 4.5 produced 1 such event in 100 non-derivable dataset-runs. Repeatability was high across automated systems (Gwet’s AC1 0.929–0.958); inter-operator agreement among 4 human raters was lower (AC1 0.608). Median execution times ranged from 6.7 to 14.0 minutes, compared with 41.9 minutes for human extraction. Conclusions Under locked, conservative benchmark conditions, MedNuggetizer and Claude Opus 4.5 showed high reliability for diagnostic data extraction in evidence synthesis. Claude Opus 4.5 was strongest on extraction correctness, whereas MedNuggetizer combined threshold-level correctness with perfect abstention safety, strong repeatability, and efficiency gains over manual extraction. artificial intelligence benchmarking data extraction diagnostic test accuracy evidence synthesis large language models reproducibility systematic review reliability urinary biomarkers Figures Figure 1 Introduction Systematic reviews and meta-analyses depend on accurate study-level data extraction, yet this step remains among the most labor-intensive and error-prone components of evidence synthesis [ 1 , 2 ]. The challenge is particularly acute in diagnostic test accuracy research, where quantitative synthesis frequently requires reconstruction of complete 2 × 2 contingency tables comprising true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN). In this setting, extraction errors are not trivial clerical imperfections. A single incorrect cell may distort sensitivity, specificity, predictive values, overall accuracy, and any downstream pooled estimate. As a result, duplicate extraction, adjudication, and prespecified decision rules remain central safeguards, even in experienced review teams [ 1 ]. Large language models (LLMs) have recently attracted major interest as tools for extracting data from full-text biomedical articles, and several early evaluations suggest that their performance can be impressive under selected conditions [ 3 – 8 ]. At the same time, the emerging literature also shows that extraction accuracy is highly task-dependent and degrades when the target information is numerically dense, structurally fragmented, or methodologically constrained [ 3 – 8 ]. This distinction is critical for diagnostic evidence synthesis. Extracting an isolated number from a paper is fundamentally different from reconstructing a valid diagnostic 2 × 2 table. The latter requires exact alignment of reported numerators, denominators, index-test definitions, reference standards, and analytic populations. Outputs that appear numerically plausible may therefore still be methodologically unusable if they are incomplete, internally inconsistent, or unsupported by the source publication. A further challenge concerns datasets for which a complete 2 × 2 table is not derivable from the published article and its publicly available supplementary material alone. In diagnostic primary studies, sensitivity, specificity, or accuracy are often reported without the full underlying cell counts required for reconstruction. In such cases, back-calculation, approximation, or silent completion cannot be regarded as valid extraction. From the perspective of evidence synthesis, the correct behavior is explicit abstention rather than unsupported numeric completion. Accordingly, the evaluation of automated systems in this context should not be limited to correctness when extraction is possible. It must also address safety when extraction is not justified. A clinically meaningful benchmark must therefore capture both exactness on derivable datasets and abstention behavior on non-derivable datasets [ 3 , 4 , 8 ]. Reproducibility presents an additional unresolved problem. Even under nominally identical conditions, LLM outputs may vary across repeated runs because of stochastic generation, backend updates, or document-ingestion instability. In systematic review workflows, such variability directly affects auditability, confidence in extracted evidence, and the defensibility of subsequent quantitative synthesis. Recent reporting standards for generative artificial intelligence (AI) evaluation in health research have consequently emphasized transparent documentation of model identity, prompt development, evaluation procedures, reference standards, and failure modes [ 9 ]. Although not developed specifically for evidence extraction, the Chatbot Assessment Reporting Tool (CHART) is highly relevant in this setting because it foregrounds methodological transparency, reproducibility, and explicit definition of successful and unsafe model behavior [ 9 ]. MedNuggetizer was developed against this background as a confidence-based information extraction framework for long-form medical documents [ 10 ]. Its architecture is built around repeated extraction and confidence-based aggregation rather than reliance on a single response, thereby explicitly addressing output variability and supporting more stable evidence acquisition across runs [ 10 ]. However, although the system has been introduced conceptually, its performance for diagnostic data extraction has not yet been prospectively evaluated against a locked human benchmark under conservative extraction constraints. The same applies more broadly to contemporary standard-access LLMs, whose usefulness for diagnostic evidence acquisition remains insufficiently characterized under protocolized, audit-oriented conditions. To address this gap, we conducted a prospective evaluation of MedNuggetizer and contemporary LLMs for diagnostic data extraction in evidence synthesis using a locked corpus derived from a previously published systematic review and diagnostic accuracy meta-analysis on Uromonitor and urine cytology in bladder cancer detection [ 11 ]. The analytical framework, including derivability rules, endpoints, and statistical methods, was prespecified in a publicly available protocol and developed in close alignment with CHART principles [ 9 , 12 ]. The primary objective was to determine whether MedNuggetizer and the evaluated LLMs achieved run-level correctness within a prespecified benchmark threshold when judged against the adjudicated human benchmark. Secondary objectives were to assess exactness of extracted diagnostic cells on derivable datasets, abstention safety on benchmark-defined non-derivable datasets, repeatability across repeated runs, fidelity of derived diagnostic metrics, and operational efficiency relative to time-controlled human extraction. We hypothesized that MedNuggetizer would meet the prespecified performance threshold, that at least one contemporary LLM would show comparable end-to-end reliability under the same locked conditions, and that automated systems would complete extraction substantially faster than human raters while maintaining high abstention safety on non-derivable datasets. Methods 2.1. Study design and protocol framework This was a prospective, protocol-driven methodological benchmark study evaluating MedNuggetizer and three contemporary LLMs for diagnostic data extraction from published full-text uro-oncologic literature. The full analytical framework, including objectives, derivability rules, endpoints, hypotheses, run structure, and the statistical analysis plan, was prespecified before system execution and published as a study protocol on medRxiv [ 12 ]. The present manuscript follows that protocol closely and was developed in alignment with the Chatbot Assessment Reporting Tool (CHART) statement, particularly with regard to model specification, prompt transparency, reference-standard definition, reproducibility, and reporting of unsafe or misleading outputs [ 9 , 12 ]. The study was intentionally designed as a conservative end-to-end benchmark. All systems were evaluated exclusively on the basis of published full-text Portable Document Format (PDF) articles and publicly available supplementary material. External browsing, database searches, retrieval augmentation, author contact, and any form of post hoc information enrichment were prohibited during execution. This design was chosen to reflect the real extraction conditions encountered in systematic review practice and to prevent artificial completion of incompletely reported diagnostic data [ 1 , 12 ]. 2.2. Source corpus and benchmark construction The locked evaluation corpus was derived from a previously published systematic review and diagnostic accuracy meta-analysis on the Uromonitor assay and urine cytology for bladder cancer detection [ 11 ]. That review identified eight peer-reviewed primary studies and provided the adjudicated human benchmark used in the present investigation [ 11 , 13 – 20 ]. Across the source studies, two target diagnostic tests were queried whenever applicable, namely Uromonitor and urine cytology. At the level of study-by-test combinations, this yielded 16 evaluable diagnostic datasets for the present benchmark [ 11 , 12 ]. The human benchmark served as the sole reference standard for two distinct determinations: first, whether a complete 2 × 2 contingency table was derivable from the publication and its supplementary material alone; and second, whether extracted numeric values were correct when derivation was possible [ 11 , 12 ]. The benchmark therefore governed both correctness and non-derivability. This distinction was central to the study design because a system could be correct either by exact extraction of all four contingency-table cells or by explicit abstention when a complete table was not derivable from the available documents [ 11 , 12 ]. 2.3. Derivability classification and analysis sets All datasets were classified a priori against the human benchmark as either derivable or non-derivable [ 11 , 12 ]. Derivable datasets were defined as those for which a complete 2 × 2 contingency table comprising true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN) could be reconstructed unambiguously from the PDF and publicly available supplementary material. Non-derivable datasets were defined as those for which at least one required cell could not be established with certainty from the documents alone. Unsupported back-calculation, approximation, or silent imputation was not permitted [ 1 , 11 , 12 ]. The benchmark included 11 derivable datasets and 5 non-derivable datasets. These 5 non-derivable datasets originated from 3 publications prespecified as sentinel publications for safety evaluation [ 12 ]. Thus, the term sentinel referred to publication-level designation, whereas the corresponding safety analysis operated at the dataset level. Three prespecified analysis sets were used. The primary analysis set comprised all 16 datasets and treated each dataset-run as a binary correct or incorrect outcome under the end-to-end primary endpoint. The derivable analysis set was restricted to benchmark-derivable datasets and was used for exact numeric evaluation of contingency-table cells and derived diagnostic metrics. The safety analysis set comprised the benchmark-defined non-derivable datasets and was used to quantify unsafe numeric output, hereafter termed hallucination behavior [ 11 , 12 ]. 2.4. Automated systems under evaluation Four automated systems were evaluated. MedNuggetizer, an academically developed confidence-based extraction workflow previously described for long-form medical documents, served as the prespecified system under validation [ 10 ]. The comparator systems were three contemporary, standard-access LLMs representing distinct provider ecosystems and model lineages: ChatGPT-5.2, Claude Opus 4.5, and Gemini 3 Pro [ 10 , 21 – 23 ]. For every run, the exact model identifier, provider, access modality, and execution timestamp were documented. Any provider-side model change detected during the study period was treated as protocol-defined model drift and documented accordingly [ 12 ]. The protocol restricted the benchmark to these three LLMs to preserve interpretability under a fixed corpus and paired repeated-run design while still covering a small set of widely accessible frontier model families relevant to document-based medical extraction [ 21 – 23 ]. 2.5. Prompt development, locking, and execution constraints The primary analysis used a single canonical extraction prompt applied uniformly across all automated systems. Prompt development was completed before study initiation through iterative refinement by the study team to operationalize a conservative extraction strategy consistent with systematic review methodology [ 1 , 12 ]. The locked prompt instructed systems to rely only on explicitly reported absolute numbers and to classify a dataset as non-derivable if any required contingency-table cell could not be determined with certainty. No model-specific prompt adaptation was permitted in the primary analysis [ 12 ]. Additional details on prompt development, locking, exploratory optimization, and the final wording of both the locked and exploratory prompts are provided in the Supplementary Materials . All automated runs were performed under controlled conditions during a prespecified execution window from March 1 to March 12, 2026. Systems were queried in newly opened sessions for every run; no run was continued within a prior chat history, and prior sessions were not reused for subsequent queries. External browsing and tools were disabled, and only the PDF article and publicly available supplementary material were allowed as input sources. Exact execution timestamps were logged for every run. The protocol required each output to contain either a complete 2 × 2 contingency table or an explicit declaration of non-derivability [ 12 ]. Provenance information supporting extracted values was requested and evaluated descriptively as part of qualitative error interpretation, although provenance plausibility was not part of the confirmatory primary endpoint. In addition to the locked primary analysis, the study included exploratory prompt-optimization analyses based on a shared, more permissive prompt applied uniformly across all systems to examine upper-bound extraction performance under transparently relaxed conditions. These analyses were prespecified as exploratory and were not part of the confirmatory primary inference [ 12 ]. Automated execution was standardized across two study authors with fixed system assignment: Anton Kravchuk conducted the MedNuggetizer and ChatGPT-5.2 runs, and Julio Ruben Rodas Garzaro conducted the Claude Opus 4.5 and Gemini 3 Pro runs. 2.6. Run design and evaluation hierarchy Each automated system was applied independently to the full corpus in 20 repeated runs under identical nominal conditions. Consecutive runs of the same automated system were separated by a fixed 5-minute interval. Each run included all 16 datasets, yielding 320 dataset-run observations per system for the primary analysis [ 12 ]. Repeated execution was chosen prospectively to quantify within-system variability and to avoid overinterpreting single-pass outputs as stable system behavior [ 9 , 10 , 12 ]. Three hierarchical evaluation units were defined. The primary evaluation unit was the dataset-run. The secondary numeric evaluation unit was the individual contingency-table cell, restricted to derivable datasets. A third descriptive unit comprised dataset-level derived diagnostic metrics calculated from complete extracted contingency tables. This hierarchy was chosen to avoid conflating exact end-to-end correctness with partial numeric agreement [ 11 , 12 ]. 2.7. Human comparator extraction To contextualize automated performance against real-world manual evidence acquisition, time-controlled human extraction was performed independently by four clinical raters with medical training. Operators 1 and 2 were research-active residents in urology at an earlier stage of academic development and had not previously conducted a systematic review. Operators 3 and 4 were board-certified urologists with substantial scientific experience, including prior involvement in the conduct and methodological supervision of systematic reviews. None of the four human raters participated in the creation of the locked benchmark dataset used as the reference standard for the present comparison [ 11 , 12 ]. During the human comparator exercise, the raters had no access to the locked benchmark content and were not informed of the outputs or performance results of MedNuggetizer or any of the comparator LLMs. Human extraction served as an operational comparator for accuracy, inter-rater agreement, and efficiency, but not as a competing reference standard. Each human rater extracted every dataset once under time-controlled conditions using the same source documents and standardized capture templates; the human comparator exercise was conducted within the same March 1 to March 12, 2026 study window as the automated benchmark. Unlike the benchmark review, the human comparator exercise did not include duplicate adjudication and was intended to reflect practical manual extraction performance under bounded working conditions [ 11 , 12 ]. 2.8. Endpoints The primary endpoint was end-to-end dataset-run correctness. For derivable datasets, a dataset-run was classified as correct if and only if the system returned a complete 2 × 2 contingency table with all four cells exactly matching the human benchmark. For benchmark-defined non-derivable datasets, a dataset-run was classified as correct if and only if the system explicitly declared non-derivability and returned no numeric values. All 16 datasets contributed to the primary endpoint in every run [ 11 , 12 ]. Secondary endpoints were prespecified as follows. First, derivable-only numeric extraction performance comprised dataset-level exactness and cell-level exact-match accuracy for TP, FP, FN, and TN. Second, safety behavior on non-derivable datasets was assessed by hallucination rate, defined as any numeric output when the correct benchmark behavior was explicit abstention. Third, repeatability was evaluated across repeated runs at the level of output-category stability and numeric stability. Fourth, fidelity of derived diagnostic metrics was assessed by calculating sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and a prespecified binary area under the receiver operating characteristic curve (AUC_binary) from extracted contingency tables and comparing these values with benchmark-derived metrics. Fifth, operational efficiency was assessed by execution time per run for automated systems and by extraction time for human raters. Provenance plausibility and omission errors were recorded descriptively to support qualitative interpretation of failure modes [ 11 , 12 ]. 2.9. Statistical analysis The primary analysis tested whether system-level run correctness exceeded a prespecified reliability threshold of 0.95. Let p denote the probability of a correct dataset-run under the primary endpoint definition. For MedNuggetizer, the confirmatory hypothesis was framed as H0: p ≤ 0.95 versus H1: p > 0.95, corresponding to a threshold-based confirmatory analysis with a prespecified 5 percentage point margin relative to the adjudicated human benchmark [ 11 , 12 ]. The same framework was applied secondarily to each LLM. Primary and secondary threshold analyses used exact one-sided binomial testing and one-sided 97.5% lower confidence bounds. Multiplicity across the three LLM threshold tests was controlled using the Holm procedure with a family-wise alpha of 0.025 [ 12 ]. Safety analyses were restricted to the benchmark-defined non-derivable datasets. For each system, hallucination rate was calculated as the proportion of non-derivable dataset-runs with any numeric output, together with Clopper-Pearson confidence intervals. Between-system comparisons for safety were performed using Fisher’s exact test with Holm adjustment across pairwise comparisons against MedNuggetizer [ 12 ]. Repeatability was assessed on two levels. Output-category agreement across the three-category outcome space of correct numeric extraction, correct abstention, and incorrect or unsafe output was quantified per dataset and system using Gwet’s AC1 coefficient with 95% confidence intervals. Numeric repeatability on derivable datasets was defined as the proportion of runs reproducing the modal complete 2 × 2 table for a given dataset and system. Between-system comparisons of repeatability were performed using dataset-paired Wilcoxon signed-rank tests with Holm adjustment. Derived diagnostic metrics were computed from complete extracted contingency tables according to standard definitions and compared descriptively with benchmark-derived values [ 11 , 12 ]. For each dataset and system, the maximum absolute deviation across runs was reported. Operational efficiency was summarized descriptively by median, interquartile range, minimum, and maximum execution time. The protocol planned all confirmatory inference at the run level, whereas consensus-level summaries from repeated runs were reported only as supportive robustness analyses [ 12 ]. All statistical analyses were performed using IBM SPSS Statistics, version 31. 2.10. Sample size and number of runs Because the corpus of eligible diagnostic studies was fixed by the benchmark review, effective sample size for the primary analysis was determined by the number of repeated runs rather than by the number of publications alone [ 11 , 12 ]. The protocol therefore prespecified 20 independent runs per automated system. With 16 datasets per run, this yielded 320 primary evaluation units per system, allowing exact threshold testing of run-level correctness under the locked benchmark design [ 12 ]. 2.11. Ethics, transparency, and data governance The study used only publicly available published articles and publicly accessible supplementary material and did not involve patient contact or individual-level identifiable data. The protocol was made publicly available before execution, and the study was designed to support transparent reporting of prompts, model identity, endpoints, and failure modes in accordance with CHART principles [ 9 , 12 ]. Results 3.1 Study corpus, analysis sets, and run structure The locked benchmark corpus comprised eight peer-reviewed full-text publications reporting diagnostic accuracy data for Uromonitor and/or urine cytology in bladder cancer detection. At the level of study-by-test combinations, this yielded 16 evaluable diagnostic datasets, as summarized in Fig. 1 and Table 1. Of these, 11 datasets were classified by the adjudicated human benchmark as derivable and 5 as non-derivable. The 5 non-derivable datasets originated from 3 prespecified sentinel publications, including one study in which urine cytology was not assessed and two studies in which neither Uromonitor nor urine cytology permitted unambiguous reconstruction of a complete 2 × 2 contingency table from the publication and publicly available supplementary material alone. Table 1 Locked benchmark corpus, dataset structure, and derivability classification Dataset ID Study Study design Diagnostic test Derivability status Sentinel publication 1.1 Batista [13] Prospective multicenter observational diagnostic validation study Uromonitor® Derivable No 1.2 Urine cytology Derivable 2.1 Sieverink [14] Prospective case–control single-center study Uromonitor® Derivable No 2.2 Urine cytology Derivable 3.1 Ecke [15] Prospective multicenter diagnostic accuracy study Uromonitor® Non-derivable Yes 3.2 Urine cytology Non-derivable 4.1 Rabien [16] Prospective case–control single-center study Uromonitor® Derivable No 4.2 Urine cytology Derivable 5.1 Wolff [17] Prospective multicenter double-blind real-world diagnostic accuracy study Uromonitor® Derivable No 5.2 Urine cytology Derivable 6.1 Rubio-Briones [18] Prospective multicenter real-world observational validation study Uromonitor® Non-derivable Yes 6.2 Urine cytology Non-derivable 7.1 Ramos [19] Prospective single-center observational diagnostic validation study Uromonitor® Derivable No 7.2 Urine cytology Derivable 8.1 Azawi [20] Prospective multicenter observational surveillance study Uromonitor® Derivable Yes 8.2 Urine cytology (not reported) Non-derivable Note: Locked benchmark corpus derived from the previously published systematic review and diagnostic accuracy meta-analysis on Uromonitor and urine cytology in bladder cancer detection [11]. Datasets were defined at the level of study-by-test combinations, yielding 16 evaluable diagnostic datasets across 8 publications. Derivability status was determined a priori against the adjudicated human benchmark. A dataset was classified as derivable if a complete 2 × 2 contingency table comprising TP, FP, FN, and TN could be reconstructed unambiguously from the full-text publication and publicly available supplementary material alone. A dataset was classified as non-derivable if at least one required cell could not be established with certainty from the available documents. Five non-derivable datasets arose from 3 prespecified sentinel publications. Sentinel designation refers to publication-level safety relevance, whereas derivability was classified at the dataset level [11,12]. Abbreviations: FN, false negatives; FP, false positives; TN, true negatives; TP, true positives Each automated system was applied to the full locked corpus in 20 independent repeated runs under protocol-conform conditions, yielding 320 dataset-run observations per system for the primary analysis. All runs were completed without protocol deviations. The derivability structure of the corpus, including sentinel designation at the publication level and dataset-level classification, is detailed in Table 1. 3.2 Primary endpoint: run-level correctness and threshold analysis Under the prespecified primary endpoint, a dataset-run was classified as correct if the system returned a fully correct 2 × 2 contingency table for a derivable dataset or explicitly abstained without numeric output for a benchmark-defined non-derivable dataset. The corresponding run-level correctness results are summarized in Table 2. Table 2 Primary endpoint: run-level correctness and prespecified threshold analysis across automated systems System Correct dataset-runs (n/N) Observed proportion Lower bound, one-sided 97.5% CI One-sided p-value Prespecified threshold met MedNuggetizer 312/320 97.50% 95.13% 0.020 Yes ChatGPT-5.2 308/320 96.25% 93.54% 0.186 No Claude Opus 4.5 313/320 97.81% 95.55% 0.009 Yes Gemini 3 Pro 301/320 94.06% 90.88% 0.818 No Note: This table summarizes the prespecified primary analysis. Run-level correctness was defined as exact extraction of a complete 2 × 2 contingency table for benchmark-derivable datasets or explicit abstention without numeric output for benchmark-defined non-derivable datasets. Each automated system was evaluated across 20 repeated runs on the full locked corpus of 16 datasets, yielding 320 dataset-run observations per system. Threshold testing was performed using one-sided exact binomial tests against a prespecified benchmark threshold of 0.95. A system was considered to have met the prespecified threshold criterion if the one-sided 97.5% lower confidence bound exceeded 0.95 and the corresponding one-sided p-value was ≤ 0.025. Abbreviations: CI, confidence interval; N, total number of dataset-run observations MedNuggetizer achieved 312 correct dataset-runs out of 320, corresponding to an observed correctness proportion of 97.50%. The one-sided 97.5% lower confidence bound was 95.13%, exceeding the prespecified benchmark threshold of 0.95, and the exact one-sided binomial test yielded a p-value of 0.020. Claude Opus 4.5 achieved the highest overall run-level correctness, with 313 of 320 correct dataset-runs (97.81%), a lower confidence bound of 95.55%, and a p-value of 0.009. Both systems therefore met the prespecified threshold criterion. Among the remaining large language models (LLMs), ChatGPT-5.2 achieved 308 of 320 correct dataset-runs (96.25%), but its one-sided lower confidence bound of 93.54% did not exceed the prespecified threshold, and the corresponding p-value was 0.186. Gemini 3 Pro achieved 301 of 320 correct dataset-runs (94.06%), with a lower confidence bound of 90.88% and a p-value of 0.818. Accordingly, ChatGPT-5.2 and Gemini 3 Pro did not meet the prespecified threshold criterion in the confirmatory primary analysis. Taken together, the primary analysis showed high run-level correctness across all evaluated systems, but threshold confirmation was limited to MedNuggetizer and Claude Opus 4.5 (Table 2). 3.3 Derivable datasets: exact extraction performance Secondary analyses on derivable datasets were restricted to the 11 benchmark-derivable datasets, corresponding to 44 individual contingency-table cells per run and 880 cell-level observations per system across 20 repeated runs. Summary results are presented in Supplementary Table S1 . At the cell level, Claude Opus 4.5 achieved the highest overall exact-match rate, with 859 of 880 correct cells (97.61%), followed by MedNuggetizer with 853 of 880 (96.93%), ChatGPT-5.2 with 834 of 880 (94.77%), and Gemini 3 Pro with 813 of 880 (92.39%). When correctness was evaluated more stringently at the derivable dataset-run level, requiring all four contingency-table cells to be correct simultaneously, Claude Opus 4.5 again performed best, with 214 of 220 all-cell-correct runs (97.3%), followed by MedNuggetizer with 213 of 220 (96.8%), ChatGPT-5.2 with 208 of 220 (94.5%), and Gemini 3 Pro with 201 of 220 (91.4%) ( Supplementary Table S1 ). Performance was not fully homogeneous across datasets. Most derivable datasets were extracted with near-perfect or perfect fidelity across systems, whereas residual errors clustered in a small subset of more challenging datasets. In particular, performance decrements were concentrated in the Batista datasets and, for ChatGPT-5.2, also in the Azawi Uromonitor dataset. For example, Claude Opus 4.5 achieved 20 of 20 fully correct runs for Batista urine cytology and for the Azawi Uromonitor dataset, whereas ChatGPT-5.2 achieved 14 of 20 fully correct runs for Azawi Uromonitor and Gemini 3 Pro achieved 7 of 20 fully correct runs for Batista urine cytology ( Supplementary Table S1 ). These findings indicate that residual extraction errors were dataset-specific rather than broadly distributed across the corpus. 3.4 Safety performance on non-derivable datasets Safety analyses were restricted to the 5 benchmark-defined non-derivable datasets, yielding 100 non-derivable dataset-runs per system. Summary results are presented in Table 3, and the event-level taxonomy of unsafe outputs is shown in Supplementary Table S2 . Table 3 Secondary endpoints: safety, repeatability, and operational efficiency across automated systems System Hallucination events, n/N (%) Gwet’s AC1 (95% CI) Agreement, % Median time per dataset-run, s (IQR) Median time per full-corpus run, min (IQR) p-value MedNuggetizer 0/100 (0.0%) 0.952 (0.898, 0.994) 95.4% 38 (22, 66) 12.2 (11.9, 12.4) < 0.001 ChatGPT-5.2 0/100 (0.0%) 0.931 (0.847, 0.987) 93.6% 32 (19, 72) 14.0 (13.0, 15.9) < 0.001 Claude Opus 4.5 1/100 (1.0%) 0.958 (0.900, 1.000) 96.0% 20 (14, 30) 6.7 (6.0, 7.5) < 0.001 Gemini 3 Pro 0/100 (0.0%) 0.929 (0.817, 1.000) 93.7% 35 (18, 51) 12.7 (8.7, 15.3) < 0.001 Human Extraction Not applicable 0.608 (0.319, 0.876) 79.2% 98 (60, 256) 41.9 (28.1, 42.4) — Note : This table summarizes the three prespecified secondary endpoint domains across the evaluated automated systems: safety on benchmark-defined non-derivable datasets, repeatability across 20 repeated runs, and operational efficiency. Safety was quantified as hallucination events across 100 benchmark-defined non-derivable dataset-runs per system. Repeatability was assessed using Gwet’s AC1 with 95% confidence intervals for output-category agreement across repeated automated runs. Operational efficiency is reported at two levels: per individual dataset-run and per complete full-corpus run comprising all 16 datasets. The p-values refer to exploratory Mann-Whitney U comparisons of full-corpus execution time between each automated system and the human extraction comparator. Human extraction data are shown for contextual comparison only. Human agreement reflects between-rater variability in a single-pass manual extraction exercise and should not be interpreted as directly analogous to within-system repeatability across repeated automated runs. Human extraction was not part of the prespecified safety analysis. Abbreviations: AC1, Gwet’s AC1 coefficient; CI, confidence interval; IQR, interquartile range; min, minutes; s, seconds MedNuggetizer correctly abstained in all 100 of 100 non-derivable dataset-runs, with no numeric output observed. The same was true for ChatGPT-5.2 and Gemini 3 Pro, each of which also showed 0 hallucination events across 100 non-derivable dataset-runs. Claude Opus 4.5 produced 1 hallucination event in 100 non-derivable dataset-runs, corresponding to a hallucination rate of 1.0% (Table 3). The single unsafe event occurred in run 12 for dataset 6.1, a benchmark-defined non-derivable Uromonitor dataset from Rubio-Briones, in which numeric output was provided despite the correct behavior being explicit abstention ( Supplementary Table S2 ). No other unsafe numeric outputs were observed. Thus, abstention safety was perfect for MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro and near-perfect for Claude Opus 4.5. 3.5 Repeatability across repeated runs Repeatability of output-category classification across repeated executions was assessed using Gwet’s AC1 with 95% confidence intervals. These results are summarized in Table 3, with detailed run-by-run agreement matrices provided in Supplementary Tables S3a-S3f . All automated systems showed high within-system repeatability across the 20 repeated runs. Claude Opus 4.5 demonstrated the highest repeatability, with an AC1 of 0.958 and 96.0% agreement, followed by MedNuggetizer with an AC1 of 0.952 and 95.4% agreement, ChatGPT-5.2 with an AC1 of 0.931 and 93.6% agreement, and Gemini 3 Pro with an AC1 of 0.929 and 93.7% agreement. Across all four automated systems, the observed AC1 values were consistent with near-complete stability of output-category classification across repeated runs. For contextual comparison, inter-operator agreement among the four human raters was lower, with an AC1 of 0.608 and 79.2% agreement ( Supplementary Table S3f ). Because this analysis reflects between-rater variability in a single-pass manual extraction exercise rather than within-system stability across repeated automated runs, it should be interpreted as an operational context measure and not as a directly analogous repeatability estimate. 3.6 Fidelity of derived diagnostic metrics For derivable datasets, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and accuracy were calculated from extracted contingency tables and compared with benchmark-derived values. Metric-level deviations are summarized in Supplementary Tables S4a-S4b . Across all systems, the median absolute deviation from benchmark-derived diagnostic metrics was 0.00 percentage points, indicating that the typical extracted dataset-run reproduced downstream diagnostic metrics exactly. Deviations occurred only in dataset-runs with underlying cell-level extraction errors. The maximum absolute deviation observed across all runs was 25.0 percentage points for MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro, whereas the maximum deviation for Claude Opus 4.5 was 10.6 percentage points ( Supplementary Table S4a ). These deviations were confined to a narrow subset of derivable datasets. Specifically, the largest deviations clustered in the Batista datasets. MedNuggetizer showed a maximum deviation of 25.0 percentage points for Batista urine cytology, ChatGPT-5.2 showed maximum deviations of 10.6 percentage points for Batista Uromonitor and 25.0 percentage points for Batista urine cytology, Claude Opus 4.5 showed a maximum deviation of 10.6 percentage points for Batista Uromonitor and no deviation for Batista urine cytology, and Gemini 3 Pro showed maximum deviations of 10.6 percentage points for Batista Uromonitor and 25.0 percentage points for Batista urine cytology ( Supplementary Table S4a ). Outside these datasets, extracted contingency tables reproduced derived diagnostic metrics exactly. 3.7 Operational efficiency Operational efficiency results are summarized in Table 3, with detailed per-run timing distributions reported in Supplementary Table S5 . Median execution time per dataset-run was 38 seconds (interquartile range [IQR], 22–66) for MedNuggetizer, 32 seconds (IQR, 19–72) for ChatGPT-5.2, 20 seconds (IQR, 14–30) for Claude Opus 4.5, and 35 seconds (IQR, 18–51) for Gemini 3 Pro. At the level of a complete full-corpus run comprising all 16 datasets, median execution times were 12.2 minutes (IQR, 11.9–12.4) for MedNuggetizer, 14.0 minutes (IQR, 13.0-15.9) for ChatGPT-5.2, 6.7 minutes (IQR, 6.0-7.5) for Claude Opus 4.5, and 12.7 minutes (IQR, 8.7–15.3) for Gemini 3 Pro (Table 3). The corresponding median time for human extraction of the full corpus was 41.9 minutes (IQR not estimable at the same granularity in the main analysis; operator-specific totals shown in Supplementary Table S5 ). All automated systems were therefore substantially faster than the time-controlled human comparator. The largest time advantage was observed for Claude Opus 4.5. 3.8 Exploratory optimized-prompt analyses Exploratory prompt-optimization analyses under a shared optimized prompt are reported in Supplementary Tables S6a-S6b . Under these more permissive but transparently documented conditions, all three evaluated LLMs achieved 33 of 33 correct extractions across the 11 derivable datasets over three repeated runs, corresponding to 100% correctness for derivable datasets ( Supplementary Table S6a ). Performance on the absent non-derivable dataset was also uniformly correct, with all systems declaring non-derivability in 3 of 3 runs. Behavior remained more heterogeneous on the ambiguous non-derivable datasets. ChatGPT-5.2 attempted reconstruction in 6 of 12 ambiguous non-derivable evaluations and declared non-derivability in the remaining 6 of 12. Claude Opus 4.5 attempted reconstruction in 9 of 12 evaluations and documented inconsistency in 2 of these cases. Gemini 3 Pro attempted reconstruction in 7 of 12 evaluations and documented inconsistency in 3 cases ( Supplementary Tables S6a-S6b ). These exploratory findings indicate that benchmark performance on derivable datasets can improve under optimized prompting, but they also reinforce the importance of retaining explicit safety evaluation for non-derivable inputs. Discussion This prospective benchmark study was designed to test three linked hypotheses: that MedNuggetizer would meet a prespecified reliability threshold under locked extraction conditions, that at least one contemporary large language model (LLM) would achieve comparable end-to-end performance, and that automated systems would reduce extraction time relative to time-controlled human extraction while maintaining high safety on non-derivable datasets. All three hypotheses were supported, albeit not uniformly across systems. MedNuggetizer and Claude Opus 4.5 exceeded the prespecified threshold in the confirmatory primary analysis, whereas ChatGPT-5.2 and Gemini 3 Pro did not. At the same time, MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro showed perfect abstention safety on benchmark-defined non-derivable datasets, all automated systems showed high repeatability across repeated runs, and all completed the full corpus substantially faster than the human comparator exercise. The central message is therefore not simply that automated extraction can perform well, but that under prospectively locked and methodologically conservative conditions it can achieve a level of correctness, safety, and stability that is operationally relevant for structured diagnostic evidence synthesis. The distinctive strength of the present study lies in its integrated benchmark logic. Prior evaluations of automated extraction have often emphasized point accuracy on derivable material, with much less attention to unsafe completion, stochastic variability across repeated executions, or the effects of prompt permissiveness on apparent performance [ 3 – 8 ]. In contrast, the present study evaluated correctness, abstention safety, repeatability, and efficiency within a single prospectively defined framework. That distinction is particularly important in diagnostic test accuracy research. In such settings, a plausible but unsupported reconstruction of a 2 × 2 contingency table is not a minor technical imperfection. It can directly distort sensitivity, specificity, predictive values, and pooled estimates while remaining difficult to detect during downstream synthesis. By treating non-derivability as a correct benchmark outcome rather than as nuisance variation, the present design addresses a methodological problem that has been insufficiently confronted in the current extraction literature. Our findings align with the broader conclusion that automated extraction performance is strongly task-dependent and deteriorates as targets become more numerically dense, structurally fragmented, or incompletely reported. Jansen et al. evaluated LLM-based data extraction across 22 systematic review databases, encompassing 2,179 primary studies, 186 variables, and 312,329 extraction comparisons against human-coded reference data [ 3 ]. Using eight models, they showed that extraction accuracy varied far more by variable type than by model, with variables required for effect-size computation performing substantially worse than contextual or moderator variables [ 3 ]. This comparison is highly informative for the present benchmark, because it suggests that extraction success is not a stable attribute of a given model, but depends heavily on the structure and inferential demands of the target variable. In that context, the strong performance observed here likely reflects not only model capability, but also the fact that a narrowly defined diagnostic extraction task, while methodologically demanding, can be benchmarked with far greater endpoint clarity than broader evidence-synthesis abstractions. Caponio et al. similarly showed that, although contemporary LLMs can retrieve many numerical items correctly, performance declines as evaluation shifts from isolated values to higher aggregation levels [ 4 ]. In their comparative study of four models extracting quantitative outcome data from unstructured full-text PDFs across six dental specialties, sub-outcome-level accuracy was exceptionally high, yet errors increased at the outcome and study level, with omission errors emerging as a dominant limitation of full-text extraction [ 4 ]. That pattern is directly relevant to our results, because it underscores the distinction between retrieving individual data points and reconstructing an analyzable dataset end to end. By defining correctness at the dataset-run level and accepting explicit abstention as the only correct response for non-derivable material, our benchmark imposed a more conservative and operationally more meaningful criterion than item-wise numeric retrieval alone. Liu et al. reported strong performance for structured extraction of randomized trial data under carefully designed prompting conditions, achieving an overall correct rate of 94.8% across 1,873 extracted items from 10 randomized controlled trials spanning six Cochrane Handbook domains [ 8 ]. Their work supports the value of structured prompting for improving extraction quality, but it did not examine repeated-run stability, dataset-level coherence of interdependent values, or explicit abstention on non-derivable material [ 8 ]. In that respect, the present study extends the literature by showing that high performance can be maintained even when evaluation is anchored to a locked and deliberately conservative framework in which complete diagnostic table reconstruction, rather than broad item-wise abstraction, is the operative benchmark. Viewed together, these studies and the present benchmark point to the same conclusion: performance estimates become far more informative when the unit of evaluation is shifted from isolated values to synthesis-relevant structured outputs. The present study extends this literature in three important ways. First, it prospectively evaluated repeated executions rather than single-pass outputs, thereby making output stability measurable rather than assumed. Second, it incorporated explicit safety assessment on benchmark-defined non-derivable datasets rather than restricting evaluation to derivable material alone. Third, it tested an academically developed task-specific extraction framework against three contemporary standard-access LLMs under identical locked conditions. These features enhance interpretability and make the resulting performance profile more relevant to real-world evidence synthesis than a benchmark based solely on one-off extraction accuracy. MedNuggetizer warrants specific discussion because, unlike the comparator models, it was developed as an academically oriented, task-specific extraction workflow rather than as a general-purpose conversational system, with repeated extraction and confidence-based aggregation built into its design [ 10 ]. Within the present benchmark, this design was associated with a favorable overall performance profile. MedNuggetizer met the prespecified threshold in the confirmatory primary analysis, produced no hallucinated numeric outputs on benchmark-defined non-derivable datasets, showed high repeatability across 20 repeated runs, and substantially reduced extraction time relative to manual review. Claude Opus 4.5 achieved the numerically strongest primary-endpoint performance and the highest cell-level exactness, but it also produced the only unsafe numeric output on a non-derivable dataset. That distinction is methodologically consequential, because in evidence synthesis the most useful system is not defined by maximal yield alone, but by the degree to which correctness, abstention discipline, output stability, and auditability remain aligned under fixed conditions. Viewed in that light, the present findings support the potential value of academically developed, audit-ready extraction workflows for structured diagnostic evidence synthesis, with MedNuggetizer representing one such implementation. The repeated-run analyses further strengthen the practical relevance of the study. Strong performance on a single execution is of limited value if outputs shift materially across repeated runs. Here, all automated systems showed high repeatability, with Gwet’s AC1 values ranging from 0.929 to 0.958. By contrast, inter-operator agreement in the time-controlled human extraction exercise was substantially lower. This comparison should be interpreted with caution, because between-rater variability in a single-pass manual exercise is not directly analogous to within-system repeatability across repeated automated runs. Even so, it remains informative as an operational context measure. It suggests that variability persists even among medically trained raters when extraction is performed under practical, non-adjudicated conditions. This is likely to be one of the most realistic niches for carefully constrained automated systems: not as substitutes for methodological oversight, but as front-end tools that reduce first-pass burden, standardize initial extraction, and make disagreement visible earlier in the workflow. This also helps contextualize the findings of Khan et al., who used GPT-4-turbo and Claude-3-Opus in a collaborative two-reviewer workflow across 10 trials from 22 publications and 23 extraction variables [ 5 ]. In their held-out test set, 87% of responses were concordant and these concordant responses achieved an accuracy of 0.94, whereas discordant responses were substantially less reliable but improved after cross-critique [ 5 ]. Their results suggest that inter-model agreement can serve as a practical confidence signal, but they also show that disagreement marks a qualitatively different error regime rather than mere random noise. By contrast, the present benchmark assessed reliability more directly through adjudicated correctness, abstention behavior, and repeatability across 20 repeated runs, thereby reducing reliance on concordance as a surrogate for trustworthiness. The optimized-prompt analyses sharpen this interpretation further. Under a shared supplementary prompt that was more permissive but still explicitly documented, all three evaluated LLMs achieved perfect performance on the derivable datasets across three repeated runs, and all correctly abstained on the absent non-derivable dataset. However, this convergence at the apparent upper bound did not extend to the ambiguous non-derivable datasets, where reconstruction attempts remained heterogeneous across systems. That pattern is highly instructive. It indicates that the performance ceiling on clearly derivable material is higher than the locked primary benchmark alone might suggest, but it also shows that safety on non-derivable material remains the more discriminating challenge once prompting becomes more permissive. Prompt optimization can therefore narrow performance differences on derivable datasets, but it does not eliminate the need for explicit safety evaluation. This supports the decision to anchor confirmatory inference to a locked, uniform prompt and to interpret the optimized-prompt analyses as supplementary contextualization rather than as an alternative primary benchmark. The study also carries broader implications for how AI evaluation in evidence synthesis should be conducted and reported ( Supplementary Table S7 ). CHART emphasizes transparent reporting of model identity, prompt development, evaluators, protocol availability, failure modes, and reproducibility procedures [ 9 ]. In this context, these are not merely reporting niceties. They are part of the underlying methodology. Small degrees of freedom in prompting, derivability definitions, or endpoint formulation can materially alter performance impressions. The present benchmark therefore supports the use of CHART-aligned reporting as a minimum standard for future studies of generative AI in evidence synthesis [ 9 ]. More broadly, it suggests that prospective protocolization, explicit benchmark definitions, and repeated-run designs should become standard practice in this field. Several limitations warrant consideration. First, the benchmark was confined to a fixed uro-oncologic diagnostic domain, and generalizability to other specialties, document types, and extraction targets remains uncertain. Second, the task was intentionally narrow and focused on diagnostic 2 × 2 contingency tables; performance on other evidence synthesis targets, including time-to-event outcomes, adjusted effect estimates, adverse event data, and risk-of-bias judgments, was not assessed. Accordingly, the present findings should not be assumed to generalize unchanged to other clinical domains, document constellations, or extraction targets without independent validation of the workflow under comparably prespecified benchmark conditions. Third, although repeated runs were treated as independent realizations under nominally identical conditions, residual dependencies related to provider-side updates or document-ingestion behavior cannot be excluded completely. Fourth, the human comparator exercise was designed as an operational context benchmark rather than a fully adjudicated duplicate-extraction workflow and should be interpreted accordingly. Fifth, the optimized supplementary prompt is informative but also illustrates how rapidly methodological transparency can be diluted once prompt permissiveness increases and reconstruction rules become more elaborate. Finally, only a limited number of frontier LLMs were evaluated. As collaborative and agentic systems become more prominent, their evaluation will require even greater methodological discipline and transparency [ 5 , 7 ]. Future work should extend this framework in three directions. First, similar benchmarks should be applied to additional clinical domains and to extraction targets beyond diagnostic test accuracy. Second, hybrid workflows in which automated extraction is followed by targeted human verification should be evaluated formally, as this is likely to be the most realistic near-term implementation strategy. Third, academically developed task-specific systems such as MedNuggetizer should be compared not only with base LLMs, but also with collaborative and agentic architectures under prospectively locked conditions. The most plausible near-term role for MedNuggetizer is not fully autonomous evidence synthesis, but deployment as an audit-ready extraction layer within structured review workflows, where it may reduce time burden, improve consistency, and make uncertainty explicit before final human adjudication. Within that bounded remit, such systems may become genuinely useful methodological instruments for evidence synthesis rather than merely impressive demonstrations of technical capability. Conclusions Under prospectively locked and methodologically conservative benchmark conditions, automated systems achieved high performance in diagnostic data extraction from full-text publications, but meaningful differences emerged once correctness, abstention safety, repeatability, and efficiency were considered together. MedNuggetizer and Claude Opus 4.5 met the prespecified threshold criterion in the confirmatory primary analysis; Claude Opus 4.5 was numerically strongest on raw correctness, whereas MedNuggetizer combined perfect abstention safety on non-derivable datasets with high repeatability and substantial efficiency gains over manual extraction. These findings support a cautious but clearly optimistic view of automated extraction in evidence synthesis. Tightly constrained, prospectively evaluated, and audit-ready systems can now achieve performance levels that are operationally relevant for structured diagnostic review workflows. The most credible near-term use case is not fully autonomous evidence synthesis, but transparent hybrid pipelines in which automated extraction accelerates first-pass data acquisition while preserving human methodological control over adjudication, ambiguity, and final synthesis decisions. Abbreviations AC1: Gwet’s AC1 coefficient AI: artificial intelligence AUC_binary: binary area under the receiver operating characteristic curve CHART: Chatbot Assessment Reporting Tool FN: false negatives FP: false positives IQR: interquartile range LLMs: large language models NPV: negative predictive value PDF: Portable Document Format PPV: positive predictive value TN: true negatives TP: true positives Declarations Ethics approval and consent to participate Not applicable. This study did not involve human participants, human data, human tissue, or animals. The benchmark corpus was derived exclusively from publicly available full-text articles and publicly available supplementary materials. Consent for publication Not applicable. Availability of data and materials All data generated or analyzed during this study are included in this published article and its supplementary information files. The benchmark corpus was derived exclusively from publicly available published full-text articles and publicly available supplementary materials. Additional run-level logs and supporting benchmark materials are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests. Funding No specific funding was received for this work. Open access funding may be enabled and organized by the University of Regensburg through the applicable Springer Nature institutional agreement. Authors’ contributions AK, JRRG, CE, and MM conceived the study. GD and SA contributed to the development of MedNuggetizer and to the methodological design of the extraction workflow evaluated in this study. AK, JRRG, GD, SA, and UK contributed to benchmark development, prompt design, and benchmark operationalization. AK and JRRG conducted the automated system runs. MH, ER, SH, and CE performed the time-controlled human comparator extraction. SH, SS, CG, MB, VE, ER, MH, CE, and MM contributed to clinical interpretation, benchmark contextualization, and critical review of the study design and results. CE and MM supervised the study. AK drafted the first manuscript version. CE and MM critically revised the manuscript for important intellectual content. All authors read and approved the final manuscript. Acknowledgements Not applicable. Authors’ information Not applicable. References Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al. editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5, updated August 2024. Cochrane; 2024 [cited 2026 Mar 14]. Available from: Salari MA, Baghi Keshtan S, et al. Global Trends in Urological Evidence Synthesis: A Bibliometric Analysis of Systematic Reviews and Meta-analyses. Med J Islam Repub Iran. 2025;39:76. 10.47176/mjiri.39.76 . PMID: 41089610; PMCID: PMC12516444. Jansen T, Liebenow LW, Mertens U et al. Data extraction by generative artificial intelligence: Assessing determinants of accuracy using human-extracted data from systematic review databases. Psychol Bull. 2025;151(10):1280–1306. 10.1037/bul0000501 . PMID: 41396533. Caponio VCA, Lorenzo-Pouso AI, Magalhaes M, et al. Accuracy of LLMs to retrieve numeric data for meta-analysis in dentistry. J Dent. 2026;164:106245. 10.1016/j.jdent.2025.106245 . Epub 2025 Nov 19. PMID: 41265689. Khan MA, Ayub U, Naqvi SAA, et al. Collaborative large language models for automated data extraction in living systematic reviews. J Am Med Inf Assoc. 2025;32(4):638–47. 10.1093/jamia/ocae325 . PMID: 39836495; PMCID: PMC12005628. Purewal A, Fautsch K, Klasova J, Hussain N, D'Souza RS. Human versus artificial intelligence: evaluating ChatGPT's performance in conducting published systematic reviews with meta-analysis in chronic pain research. Reg Anesth Pain Med. 2025 Feb 16:rapm-2024-106358. 10.1136/rapm-2024-106358 . Epub ahead of print. PMID: 39956557. Gorenshtein A, Omar M, Glicksberg BS, Nadkarni GN, Klang E. AI Agents in Clinical Medicine: A Systematic Review. medRxiv [Preprint]. 2025 Aug 26:2025.08.22.25334232. doi: 10.1101/2025.08.22.25334232. PMID: 40909853; PMCID: PMC12407621. Liu J, Lai H, Zhao W, et al. ADVANCED Working Group. AI-driven evidence synthesis: data extraction of randomized controlled trials with large language models. Int J Surg. 2025;111(3):2722–6. PMID: 39903558; PMCID: PMC12372713. CHART Collaborative. Reporting guideline for chatbot health advice studies: the Chatbot Assessment Reporting Tool (CHART) statement. BMJ Med. 2025;4(1):e001632. 10.1136/bmjmed-2025-001632 . PMID: 40761518; PMCID: PMC12320030. Donabauer G, Ateia S, Kruschwitz U et al. MedNuggetizer: Confidence-Based Information Nugget Extraction from Medical Documents. ECIR. 2026. arXiv:2512.15384 [cs.IR]. Rodas Garzaro JR, Kravchuk A, Burger M, et al. Diagnostic Performance and Clinical Utility of the Uromonitor® Molecular Urine Assay for Urothelial Carcinoma of the Bladder: A Systematic Review and Diagnostic Accuracy Meta-Analysis. Diagnostics (Basel). 2026;16(2):285. 10.3390/diagnostics16020285 . PMID: 41594261; PMCID: PMC12840495. May M, Rodas Garzaro JR, Kravchuk A et al. End-to-End Reliability of Automated Systems for Diagnostic Data Extraction: A Benchmark Study in Uro-Oncologic Evidence Synthesis. medRxiv. MS ID#: MEDRXIV/2025/342959. Protocol v1.0, December 24, 2025. Batista R, Vinagre J, Prazeres H, et al. Validation of a Novel, Sensitive, and Specific Urine-Based Test for Recurrence Surveillance of Patients With Non-Muscle-Invasive Bladder Cancer in a Comprehensive Multicenter Study. Front Genet. 2019;10:1237. 10.3389/fgene.2019.01237 . PMID: 31921291; PMCID: PMC6930177. Sieverink CA, Batista RPM, Prazeres HJM, et al. Clinical Validation of a Urine Test (Uromonitor-V2®) for the Surveillance of Non-Muscle-Invasive Bladder Cancer Patients. Diagnostics (Basel). 2020;10(10):745. 10.3390/diagnostics10100745 . PMID: 32987933; PMCID: PMC7599569. Ecke TH, Meisl CJ, Schlomm T, et al. Performance of Urinary Markers in Patients With Suspicious Cystoscopy During Follow-up of Recurrent Non-muscle Invasive Bladder Cancer: BTA Stat, NMP22 BladderChek, UBC Rapid Test, CancerCheck UBC Rapid VISUAL, and Uromonitor in Comparison to Cytology. Urology. 2025;197:119–25. Epub 2024 Dec 1. PMID: 39626834. Rabien A, Rong D, Rabenhorst S, et al. Diagnostic performance of Uromonitor and TERTpm ddPCR urine tests for the non-invasive detection of bladder cancer. Sci Rep. 2024;14(1):30617. 10.1038/s41598-024-83976-2 . PMID: 39715826; PMCID: PMC11666540. Wolff I, Kravchuk AP, Wirtz RM, et al. Real-world performance of Uromonitor® in urothelial bladder cancer detection: a multicentric trial. BJU Int. 2024;134(6):992–1000. 10.1111/bju.16450 . Epub 2024 Jun 24. PMID: 38923777. Rubio-Briones J, Guerrero Ramos F, Mercadé Sánchez A, et al. External validation of the Uromonitor®-version 2 urine test as a biomarker for optimisation of non-muscle-invasive bladder cancer management. BJU Int. 2026;137(1):216–24. 10.1111/bju.70010 . Epub 2025 Oct 1. PMID: 41031577. Ramos P, Brás JP, Dias C, et al. Uromonitor: Clinical Validation and Performance Assessment of a Urinary Biomarker Within the Surveillance of Patients With Nonmuscle-Invasive Bladder Cancer. J Urol. 2025;213(3):304–12. Epub 2024 Nov 19. PMID: 39561374. Azawi N, Vásquez JL, Dreyer T, et al. Surveillance of Low-Grade Non-Muscle Invasive Bladder Tumors Using Uromonitor: SOLUSION Trial. Cancers (Basel). 2023;15(8):2341. 10.3390/cancers15082341 . PMID: 37190269; PMCID: PMC10137147. OpenAI. GPT-5.2 model documentation [Internet]. San Francisco (CA): OpenAI; 2025 [cited 2026 Mar 14]. Available from: https://platform.openai.com/docs/models Anthropic. Claude model documentation and version registry [Internet]. San Francisco (CA): Anthropic; 2025 [cited 2026 Mar 14]. Available from: https://docs.anthropic.com Google DeepMind. Gemini 3 Pro model card [Internet]. London (UK): Google DeepMind; 2025 [cited 2026 Mar 14]. Available from: https://deepmind.google/technologies/gemini Additional Declarations No competing interests reported. Supplementary Files SupplementBMCMedicalResearchMethodology.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 28 Apr, 2026 Reviews received at journal 24 Apr, 2026 Reviewers agreed at journal 22 Apr, 2026 Reviewers invited by journal 21 Apr, 2026 Editor invited by journal 31 Mar, 2026 Editor assigned by journal 30 Mar, 2026 Submission checks completed at journal 30 Mar, 2026 First submitted to journal 29 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9260490","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":631526279,"identity":"4de87929-cc0e-4910-899e-8667047a8ef0","order_by":0,"name":"Anton Kravchuk","email":"","orcid":"","institution":"St. Elisabeth Hospital Straubing, Medical Campus Lower Bavaria (MCN)","correspondingAuthor":false,"prefix":"","firstName":"Anton","middleName":"","lastName":"Kravchuk","suffix":""},{"id":631526280,"identity":"2270a828-9445-48ca-9d0b-3baf16474ade","order_by":1,"name":"Julio Ruben Rodas Garzaro","email":"","orcid":"","institution":"St. Elisabeth Hospital Straubing, Medical Campus Lower Bavaria (MCN)","correspondingAuthor":false,"prefix":"","firstName":"Julio","middleName":"Ruben Rodas","lastName":"Garzaro","suffix":""},{"id":631526281,"identity":"c1df887c-e85b-4b90-af88-d4b5238e500b","order_by":2,"name":"Gregor Donabauer","email":"","orcid":"","institution":"University of Regensburg","correspondingAuthor":false,"prefix":"","firstName":"Gregor","middleName":"","lastName":"Donabauer","suffix":""},{"id":631526282,"identity":"410b656f-ec29-4275-a4c3-02c1263366fd","order_by":3,"name":"Samy Ateia","email":"","orcid":"","institution":"University of Regensburg","correspondingAuthor":false,"prefix":"","firstName":"Samy","middleName":"","lastName":"Ateia","suffix":""},{"id":631526283,"identity":"e6705842-7f7d-453b-b4ca-8dc6b3a21d82","order_by":4,"name":"Udo Kruschwitz","email":"","orcid":"","institution":"University of Regensburg","correspondingAuthor":false,"prefix":"","firstName":"Udo","middleName":"","lastName":"Kruschwitz","suffix":""},{"id":631526284,"identity":"98bf59a1-c4e8-471c-b087-15ca8b653f35","order_by":5,"name":"Valentin Egert","email":"","orcid":"","institution":"University of Regensburg","correspondingAuthor":false,"prefix":"","firstName":"Valentin","middleName":"","lastName":"Egert","suffix":""},{"id":631526285,"identity":"cf046ecf-629d-4708-8a19-00724caa67cb","order_by":6,"name":"Christoph Eckl","email":"","orcid":"","institution":"University of Regensburg","correspondingAuthor":false,"prefix":"","firstName":"Christoph","middleName":"","lastName":"Eckl","suffix":""},{"id":631526288,"identity":"d1477281-7424-4419-a6c2-7d1cb1802cc8","order_by":7,"name":"Emily Rinderknecht","email":"","orcid":"","institution":"University of Augsburg","correspondingAuthor":false,"prefix":"","firstName":"Emily","middleName":"","lastName":"Rinderknecht","suffix":""},{"id":631526290,"identity":"ff6ecb08-cdea-4ffb-8fbf-b24d34b0601b","order_by":8,"name":"Stefanie Herrmann","email":"","orcid":"","institution":"St. Elisabeth Hospital Straubing, Medical Campus Lower Bavaria (MCN)","correspondingAuthor":false,"prefix":"","firstName":"Stefanie","middleName":"","lastName":"Herrmann","suffix":""},{"id":631526291,"identity":"c25172b9-ae20-4e48-a20b-e4b9582fc5a6","order_by":9,"name":"Stephan Siepmann","email":"","orcid":"","institution":"St. Elisabeth Hospital Straubing, Medical Campus Lower Bavaria (MCN)","correspondingAuthor":false,"prefix":"","firstName":"Stephan","middleName":"","lastName":"Siepmann","suffix":""},{"id":631526292,"identity":"858b0a2f-cb43-4136-a83e-037d580f82ae","order_by":10,"name":"Maximilian Burger","email":"","orcid":"","institution":"University of Regensburg","correspondingAuthor":false,"prefix":"","firstName":"Maximilian","middleName":"","lastName":"Burger","suffix":""},{"id":631526297,"identity":"7df395ff-1ecf-4884-b35d-9f9c37b062da","order_by":11,"name":"Christian Gilfrich","email":"","orcid":"","institution":"St. Elisabeth Hospital Straubing, Medical Campus Lower Bavaria (MCN)","correspondingAuthor":false,"prefix":"","firstName":"Christian","middleName":"","lastName":"Gilfrich","suffix":""},{"id":631526298,"identity":"7df79954-7120-4042-88d9-18ddd64407a4","order_by":12,"name":"Maximilian Haas","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAoUlEQVRIiWNgGAWjYDACCRBRwcDA2MBgQIqWMyRrYWwDM4nUoju7+eHHn/Ps7JkbmDc+IEqL2Z1jxtK825ITGxvYiomzxuxGgoE047YDCYwNPGYSRGpJ//zz55wD9kAt5j+I1JJjJsHbcICxEWgLUTpAWsqseY4B/dLMVky0wzbf/FFjZ2/Y3rzxA3HWwIBhM2nqgUCeZB2jYBSMglEwYgAAr/0usYkijK0AAAAASUVORK5CYII=","orcid":"","institution":"University of Regensburg","correspondingAuthor":true,"prefix":"","firstName":"Maximilian","middleName":"","lastName":"Haas","suffix":""},{"id":631526299,"identity":"f21f1c51-0fde-4542-bf99-25d320c39c41","order_by":13,"name":"Matthias May","email":"","orcid":"","institution":"St. Elisabeth Hospital Straubing, Medical Campus Lower Bavaria (MCN)","correspondingAuthor":false,"prefix":"","firstName":"Matthias","middleName":"","lastName":"May","suffix":""}],"badges":[],"createdAt":"2026-03-29 18:23:27","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9260490/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9260490/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108246457,"identity":"8eab14c1-2a8c-42c1-9295-70bd61c555c3","added_by":"auto","created_at":"2026-05-01 00:04:25","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":500329,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy flow, benchmark derivability classification, and repeated-run evaluation framework\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNote:\u003c/strong\u003e The locked corpus comprised 8 peer-reviewed publications, yielding 16 diagnostic datasets at the level of study-by-test combinations. Based on the adjudicated human benchmark, 11 datasets were classified as derivable and 5 as non-derivable, the latter arising from 3 prespecified sentinel publications. Four automated systems were evaluated under a locked uniform prompt across 20 repeated runs per system, yielding 320 dataset-run observations per system. The primary endpoint was run-level correctness, defined as exact extraction for derivable datasets or explicit abstention without numeric output for non-derivable datasets. Prespecified secondary endpoints included cell-level exactness, safety on non-derivable datasets, repeatability across repeated runs, fidelity of derived diagnostic metrics, and operational efficiency.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9260490/v1/00fed5f1f9cecc4d5ef314b0.png"},{"id":108491574,"identity":"04c8ccd8-dc62-4b6a-a28f-55db8b4c9908","added_by":"auto","created_at":"2026-05-05 09:54:40","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":907689,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9260490/v1/7f5b8e7a-0c07-49aa-afef-c3ee8af4f464.pdf"},{"id":108246456,"identity":"2e240000-25b4-40ea-92c2-14dcdd14faf9","added_by":"auto","created_at":"2026-05-01 00:04:25","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":111442,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementBMCMedicalResearchMethodology.docx","url":"https://assets-eu.researchsquare.com/files/rs-9260490/v1/d92c4fb07cd6f002176f7d05.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"End-to-End Reliability of Automated Systems for Diagnostic Evidence Extraction: A Prospective Benchmark Study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSystematic reviews and meta-analyses depend on accurate study-level data extraction, yet this step remains among the most labor-intensive and error-prone components of evidence synthesis [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. The challenge is particularly acute in diagnostic test accuracy research, where quantitative synthesis frequently requires reconstruction of complete 2 \u0026times; 2 contingency tables comprising true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN). In this setting, extraction errors are not trivial clerical imperfections. A single incorrect cell may distort sensitivity, specificity, predictive values, overall accuracy, and any downstream pooled estimate. As a result, duplicate extraction, adjudication, and prespecified decision rules remain central safeguards, even in experienced review teams [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eLarge language models (LLMs) have recently attracted major interest as tools for extracting data from full-text biomedical articles, and several early evaluations suggest that their performance can be impressive under selected conditions [\u003cspan additionalcitationids=\"CR4 CR5 CR6 CR7\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. At the same time, the emerging literature also shows that extraction accuracy is highly task-dependent and degrades when the target information is numerically dense, structurally fragmented, or methodologically constrained [\u003cspan additionalcitationids=\"CR4 CR5 CR6 CR7\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. This distinction is critical for diagnostic evidence synthesis. Extracting an isolated number from a paper is fundamentally different from reconstructing a valid diagnostic 2 \u0026times; 2 table. The latter requires exact alignment of reported numerators, denominators, index-test definitions, reference standards, and analytic populations. Outputs that appear numerically plausible may therefore still be methodologically unusable if they are incomplete, internally inconsistent, or unsupported by the source publication.\u003c/p\u003e \u003cp\u003eA further challenge concerns datasets for which a complete 2 \u0026times; 2 table is not derivable from the published article and its publicly available supplementary material alone. In diagnostic primary studies, sensitivity, specificity, or accuracy are often reported without the full underlying cell counts required for reconstruction. In such cases, back-calculation, approximation, or silent completion cannot be regarded as valid extraction. From the perspective of evidence synthesis, the correct behavior is explicit abstention rather than unsupported numeric completion. Accordingly, the evaluation of automated systems in this context should not be limited to correctness when extraction is possible. It must also address safety when extraction is not justified. A clinically meaningful benchmark must therefore capture both exactness on derivable datasets and abstention behavior on non-derivable datasets [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eReproducibility presents an additional unresolved problem. Even under nominally identical conditions, LLM outputs may vary across repeated runs because of stochastic generation, backend updates, or document-ingestion instability. In systematic review workflows, such variability directly affects auditability, confidence in extracted evidence, and the defensibility of subsequent quantitative synthesis. Recent reporting standards for generative artificial intelligence (AI) evaluation in health research have consequently emphasized transparent documentation of model identity, prompt development, evaluation procedures, reference standards, and failure modes [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Although not developed specifically for evidence extraction, the Chatbot Assessment Reporting Tool (CHART) is highly relevant in this setting because it foregrounds methodological transparency, reproducibility, and explicit definition of successful and unsafe model behavior [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMedNuggetizer was developed against this background as a confidence-based information extraction framework for long-form medical documents [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Its architecture is built around repeated extraction and confidence-based aggregation rather than reliance on a single response, thereby explicitly addressing output variability and supporting more stable evidence acquisition across runs [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. However, although the system has been introduced conceptually, its performance for diagnostic data extraction has not yet been prospectively evaluated against a locked human benchmark under conservative extraction constraints. The same applies more broadly to contemporary standard-access LLMs, whose usefulness for diagnostic evidence acquisition remains insufficiently characterized under protocolized, audit-oriented conditions.\u003c/p\u003e \u003cp\u003eTo address this gap, we conducted a prospective evaluation of MedNuggetizer and contemporary LLMs for diagnostic data extraction in evidence synthesis using a locked corpus derived from a previously published systematic review and diagnostic accuracy meta-analysis on Uromonitor and urine cytology in bladder cancer detection [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. The analytical framework, including derivability rules, endpoints, and statistical methods, was prespecified in a publicly available protocol and developed in close alignment with CHART principles [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The primary objective was to determine whether MedNuggetizer and the evaluated LLMs achieved run-level correctness within a prespecified benchmark threshold when judged against the adjudicated human benchmark. Secondary objectives were to assess exactness of extracted diagnostic cells on derivable datasets, abstention safety on benchmark-defined non-derivable datasets, repeatability across repeated runs, fidelity of derived diagnostic metrics, and operational efficiency relative to time-controlled human extraction. We hypothesized that MedNuggetizer would meet the prespecified performance threshold, that at least one contemporary LLM would show comparable end-to-end reliability under the same locked conditions, and that automated systems would complete extraction substantially faster than human raters while maintaining high abstention safety on non-derivable datasets.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Study design and protocol framework\u003c/h2\u003e \u003cp\u003eThis was a prospective, protocol-driven methodological benchmark study evaluating MedNuggetizer and three contemporary LLMs for diagnostic data extraction from published full-text uro-oncologic literature. The full analytical framework, including objectives, derivability rules, endpoints, hypotheses, run structure, and the statistical analysis plan, was prespecified before system execution and published as a study protocol on medRxiv [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The present manuscript follows that protocol closely and was developed in alignment with the Chatbot Assessment Reporting Tool (CHART) statement, particularly with regard to model specification, prompt transparency, reference-standard definition, reproducibility, and reporting of unsafe or misleading outputs [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe study was intentionally designed as a conservative end-to-end benchmark. All systems were evaluated exclusively on the basis of published full-text Portable Document Format (PDF) articles and publicly available supplementary material. External browsing, database searches, retrieval augmentation, author contact, and any form of post hoc information enrichment were prohibited during execution. This design was chosen to reflect the real extraction conditions encountered in systematic review practice and to prevent artificial completion of incompletely reported diagnostic data [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Source corpus and benchmark construction\u003c/h2\u003e \u003cp\u003eThe locked evaluation corpus was derived from a previously published systematic review and diagnostic accuracy meta-analysis on the Uromonitor assay and urine cytology for bladder cancer detection [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. That review identified eight peer-reviewed primary studies and provided the adjudicated human benchmark used in the present investigation [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan additionalcitationids=\"CR14 CR15 CR16 CR17 CR18 CR19\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Across the source studies, two target diagnostic tests were queried whenever applicable, namely Uromonitor and urine cytology. At the level of study-by-test combinations, this yielded 16 evaluable diagnostic datasets for the present benchmark [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe human benchmark served as the sole reference standard for two distinct determinations: first, whether a complete 2 \u0026times; 2 contingency table was derivable from the publication and its supplementary material alone; and second, whether extracted numeric values were correct when derivation was possible [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The benchmark therefore governed both correctness and non-derivability. This distinction was central to the study design because a system could be correct either by exact extraction of all four contingency-table cells or by explicit abstention when a complete table was not derivable from the available documents [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Derivability classification and analysis sets\u003c/h2\u003e \u003cp\u003eAll datasets were classified a priori against the human benchmark as either derivable or non-derivable [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Derivable datasets were defined as those for which a complete 2 \u0026times; 2 contingency table comprising true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN) could be reconstructed unambiguously from the PDF and publicly available supplementary material. Non-derivable datasets were defined as those for which at least one required cell could not be established with certainty from the documents alone. Unsupported back-calculation, approximation, or silent imputation was not permitted [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe benchmark included 11 derivable datasets and 5 non-derivable datasets. These 5 non-derivable datasets originated from 3 publications prespecified as sentinel publications for safety evaluation [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Thus, the term sentinel referred to publication-level designation, whereas the corresponding safety analysis operated at the dataset level.\u003c/p\u003e \u003cp\u003eThree prespecified analysis sets were used. The primary analysis set comprised all 16 datasets and treated each dataset-run as a binary correct or incorrect outcome under the end-to-end primary endpoint. The derivable analysis set was restricted to benchmark-derivable datasets and was used for exact numeric evaluation of contingency-table cells and derived diagnostic metrics. The safety analysis set comprised the benchmark-defined non-derivable datasets and was used to quantify unsafe numeric output, hereafter termed hallucination behavior [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Automated systems under evaluation\u003c/h2\u003e \u003cp\u003eFour automated systems were evaluated. MedNuggetizer, an academically developed confidence-based extraction workflow previously described for long-form medical documents, served as the prespecified system under validation [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The comparator systems were three contemporary, standard-access LLMs representing distinct provider ecosystems and model lineages: ChatGPT-5.2, Claude Opus 4.5, and Gemini 3 Pro [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. For every run, the exact model identifier, provider, access modality, and execution timestamp were documented. Any provider-side model change detected during the study period was treated as protocol-defined model drift and documented accordingly [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe protocol restricted the benchmark to these three LLMs to preserve interpretability under a fixed corpus and paired repeated-run design while still covering a small set of widely accessible frontier model families relevant to document-based medical extraction [\u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5. Prompt development, locking, and execution constraints\u003c/h2\u003e \u003cp\u003eThe primary analysis used a single canonical extraction prompt applied uniformly across all automated systems. Prompt development was completed before study initiation through iterative refinement by the study team to operationalize a conservative extraction strategy consistent with systematic review methodology [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The locked prompt instructed systems to rely only on explicitly reported absolute numbers and to classify a dataset as non-derivable if any required contingency-table cell could not be determined with certainty. No model-specific prompt adaptation was permitted in the primary analysis [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Additional details on prompt development, locking, exploratory optimization, and the final wording of both the locked and exploratory prompts are provided in the \u003cb\u003eSupplementary Materials\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eAll automated runs were performed under controlled conditions during a prespecified execution window from March 1 to March 12, 2026. Systems were queried in newly opened sessions for every run; no run was continued within a prior chat history, and prior sessions were not reused for subsequent queries. External browsing and tools were disabled, and only the PDF article and publicly available supplementary material were allowed as input sources. Exact execution timestamps were logged for every run. The protocol required each output to contain either a complete 2 \u0026times; 2 contingency table or an explicit declaration of non-derivability [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Provenance information supporting extracted values was requested and evaluated descriptively as part of qualitative error interpretation, although provenance plausibility was not part of the confirmatory primary endpoint.\u003c/p\u003e \u003cp\u003eIn addition to the locked primary analysis, the study included exploratory prompt-optimization analyses based on a shared, more permissive prompt applied uniformly across all systems to examine upper-bound extraction performance under transparently relaxed conditions. These analyses were prespecified as exploratory and were not part of the confirmatory primary inference [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Automated execution was standardized across two study authors with fixed system assignment: Anton Kravchuk conducted the MedNuggetizer and ChatGPT-5.2 runs, and Julio Ruben Rodas Garzaro conducted the Claude Opus 4.5 and Gemini 3 Pro runs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6. Run design and evaluation hierarchy\u003c/h2\u003e \u003cp\u003eEach automated system was applied independently to the full corpus in 20 repeated runs under identical nominal conditions. Consecutive runs of the same automated system were separated by a fixed 5-minute interval. Each run included all 16 datasets, yielding 320 dataset-run observations per system for the primary analysis [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Repeated execution was chosen prospectively to quantify within-system variability and to avoid overinterpreting single-pass outputs as stable system behavior [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThree hierarchical evaluation units were defined. The primary evaluation unit was the dataset-run. The secondary numeric evaluation unit was the individual contingency-table cell, restricted to derivable datasets. A third descriptive unit comprised dataset-level derived diagnostic metrics calculated from complete extracted contingency tables. This hierarchy was chosen to avoid conflating exact end-to-end correctness with partial numeric agreement [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7. Human comparator extraction\u003c/h2\u003e \u003cp\u003eTo contextualize automated performance against real-world manual evidence acquisition, time-controlled human extraction was performed independently by four clinical raters with medical training. Operators 1 and 2 were research-active residents in urology at an earlier stage of academic development and had not previously conducted a systematic review. Operators 3 and 4 were board-certified urologists with substantial scientific experience, including prior involvement in the conduct and methodological supervision of systematic reviews. None of the four human raters participated in the creation of the locked benchmark dataset used as the reference standard for the present comparison [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. During the human comparator exercise, the raters had no access to the locked benchmark content and were not informed of the outputs or performance results of MedNuggetizer or any of the comparator LLMs.\u003c/p\u003e \u003cp\u003eHuman extraction served as an operational comparator for accuracy, inter-rater agreement, and efficiency, but not as a competing reference standard. Each human rater extracted every dataset once under time-controlled conditions using the same source documents and standardized capture templates; the human comparator exercise was conducted within the same March 1 to March 12, 2026 study window as the automated benchmark. Unlike the benchmark review, the human comparator exercise did not include duplicate adjudication and was intended to reflect practical manual extraction performance under bounded working conditions [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e2.8. Endpoints\u003c/h2\u003e \u003cp\u003eThe primary endpoint was end-to-end dataset-run correctness. For derivable datasets, a dataset-run was classified as correct if and only if the system returned a complete 2 \u0026times; 2 contingency table with all four cells exactly matching the human benchmark. For benchmark-defined non-derivable datasets, a dataset-run was classified as correct if and only if the system explicitly declared non-derivability and returned no numeric values. All 16 datasets contributed to the primary endpoint in every run [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSecondary endpoints were prespecified as follows. First, derivable-only numeric extraction performance comprised dataset-level exactness and cell-level exact-match accuracy for TP, FP, FN, and TN. Second, safety behavior on non-derivable datasets was assessed by hallucination rate, defined as any numeric output when the correct benchmark behavior was explicit abstention. Third, repeatability was evaluated across repeated runs at the level of output-category stability and numeric stability. Fourth, fidelity of derived diagnostic metrics was assessed by calculating sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and a prespecified binary area under the receiver operating characteristic curve (AUC_binary) from extracted contingency tables and comparing these values with benchmark-derived metrics. Fifth, operational efficiency was assessed by execution time per run for automated systems and by extraction time for human raters. Provenance plausibility and omission errors were recorded descriptively to support qualitative interpretation of failure modes [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e2.9. Statistical analysis\u003c/h2\u003e \u003cp\u003eThe primary analysis tested whether system-level run correctness exceeded a prespecified reliability threshold of 0.95. Let p denote the probability of a correct dataset-run under the primary endpoint definition. For MedNuggetizer, the confirmatory hypothesis was framed as H0: p\u0026thinsp;\u0026le;\u0026thinsp;0.95 versus H1: p\u0026thinsp;\u0026gt;\u0026thinsp;0.95, corresponding to a threshold-based confirmatory analysis with a prespecified 5 percentage point margin relative to the adjudicated human benchmark [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The same framework was applied secondarily to each LLM. Primary and secondary threshold analyses used exact one-sided binomial testing and one-sided 97.5% lower confidence bounds. Multiplicity across the three LLM threshold tests was controlled using the Holm procedure with a family-wise alpha of 0.025 [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSafety analyses were restricted to the benchmark-defined non-derivable datasets. For each system, hallucination rate was calculated as the proportion of non-derivable dataset-runs with any numeric output, together with Clopper-Pearson confidence intervals. Between-system comparisons for safety were performed using Fisher\u0026rsquo;s exact test with Holm adjustment across pairwise comparisons against MedNuggetizer [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eRepeatability was assessed on two levels. Output-category agreement across the three-category outcome space of correct numeric extraction, correct abstention, and incorrect or unsafe output was quantified per dataset and system using Gwet\u0026rsquo;s AC1 coefficient with 95% confidence intervals. Numeric repeatability on derivable datasets was defined as the proportion of runs reproducing the modal complete 2 \u0026times; 2 table for a given dataset and system. Between-system comparisons of repeatability were performed using dataset-paired Wilcoxon signed-rank tests with Holm adjustment.\u003c/p\u003e \u003cp\u003eDerived diagnostic metrics were computed from complete extracted contingency tables according to standard definitions and compared descriptively with benchmark-derived values [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. For each dataset and system, the maximum absolute deviation across runs was reported. Operational efficiency was summarized descriptively by median, interquartile range, minimum, and maximum execution time. The protocol planned all confirmatory inference at the run level, whereas consensus-level summaries from repeated runs were reported only as supportive robustness analyses [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAll statistical analyses were performed using IBM SPSS Statistics, version 31.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e2.10. Sample size and number of runs\u003c/h2\u003e \u003cp\u003eBecause the corpus of eligible diagnostic studies was fixed by the benchmark review, effective sample size for the primary analysis was determined by the number of repeated runs rather than by the number of publications alone [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The protocol therefore prespecified 20 independent runs per automated system. With 16 datasets per run, this yielded 320 primary evaluation units per system, allowing exact threshold testing of run-level correctness under the locked benchmark design [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e2.11. Ethics, transparency, and data governance\u003c/h2\u003e \u003cp\u003eThe study used only publicly available published articles and publicly accessible supplementary material and did not involve patient contact or individual-level identifiable data. The protocol was made publicly available before execution, and the study was designed to support transparent reporting of prompts, model identity, endpoints, and failure modes in accordance with CHART principles [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec15\"\u003e\n \u003ch2\u003e3.1 Study corpus, analysis sets, and run structure\u003c/h2\u003e\n \u003cp\u003eThe locked benchmark corpus comprised eight peer-reviewed full-text publications reporting diagnostic accuracy data for Uromonitor and/or urine cytology in bladder cancer detection. At the level of study-by-test combinations, this yielded 16 evaluable diagnostic datasets, as summarized in Fig. 1 and Table 1. Of these, 11 datasets were classified by the adjudicated human benchmark as derivable and 5 as non-derivable. The 5 non-derivable datasets originated from 3 prespecified sentinel publications, including one study in which urine cytology was not assessed and two studies in which neither Uromonitor nor urine cytology permitted unambiguous reconstruction of a complete 2 × 2 contingency table from the publication and publicly available supplementary material alone.\u003c/p\u003e\n \u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 1\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eLocked benchmark corpus, dataset structure, and derivability classification\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eDataset ID\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eStudy\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eStudy design\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eDiagnostic test\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivability status\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eSentinel publication\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e1.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eBatista [13]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective multicenter observational diagnostic validation study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e1.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e2.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eSieverink [14]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective case–control single-center study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e2.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e3.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eEcke [15]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective multicenter diagnostic accuracy study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eNon-derivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e3.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eNon-derivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e4.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eRabien [16]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective case–control single-center study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e4.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e5.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eWolff [17]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective multicenter double-blind real-world diagnostic accuracy study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e5.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e6.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eRubio-Briones [18]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective multicenter real-world observational validation study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eNon-derivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e6.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eNon-derivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e7.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eRamos [19]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective single-center observational diagnostic validation study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e7.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e8.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eAzawi [20]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eProspective multicenter observational surveillance study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUromonitor®\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDerivable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003e8.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUrine cytology (not reported)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eNon-derivable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eNote:\u003c/strong\u003e Locked benchmark corpus derived from the previously published systematic review and diagnostic accuracy meta-analysis on Uromonitor and urine cytology in bladder cancer detection [11]. Datasets were defined at the level of study-by-test combinations, yielding 16 evaluable diagnostic datasets across 8 publications. Derivability status was determined a priori against the adjudicated human benchmark. A dataset was classified as derivable if a complete 2 × 2 contingency table comprising TP, FP, FN, and TN could be reconstructed unambiguously from the full-text publication and publicly available supplementary material alone. A dataset was classified as non-derivable if at least one required cell could not be established with certainty from the available documents. Five non-derivable datasets arose from 3 prespecified sentinel publications. Sentinel designation refers to publication-level safety relevance, whereas derivability was classified at the dataset level [11,12].\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eAbbreviations:\u003c/strong\u003e FN, false negatives; FP, false positives; TN, true negatives; TP, true positives\u003cbr clear=\"all\"\u003e\u0026nbsp;Each automated system was applied to the full locked corpus in 20 independent repeated runs under protocol-conform conditions, yielding 320 dataset-run observations per system for the primary analysis. All runs were completed without protocol deviations. The derivability structure of the corpus, including sentinel designation at the publication level and dataset-level classification, is detailed in Table 1.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec16\"\u003e\n \u003ch2\u003e3.2 Primary endpoint: run-level correctness and threshold analysis\u003c/h2\u003e\n \u003cp\u003eUnder the prespecified primary endpoint, a dataset-run was classified as correct if the system returned a fully correct 2 × 2 contingency table for a derivable dataset or explicitly abstained without numeric output for a benchmark-defined non-derivable dataset. The corresponding run-level correctness results are summarized in Table 2. \u0026nbsp;\u003c/p\u003e\n \u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 2\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003ePrimary endpoint: run-level correctness and prespecified threshold analysis across automated systems\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eSystem\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eCorrect dataset-runs (n/N)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eObserved proportion\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eLower bound, one-sided 97.5% CI\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eOne-sided p-value\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003ePrespecified threshold met\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eMedNuggetizer\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e312/320\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e97.50%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e95.13%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e0.020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eChatGPT-5.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e308/320\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e96.25%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e93.54%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e0.186\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eClaude Opus 4.5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e313/320\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e97.81%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e95.55%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e0.009\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eGemini 3 Pro\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e301/320\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e94.06%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e90.88%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e0.818\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eNote:\u003c/strong\u003e This table summarizes the prespecified primary analysis. Run-level correctness was defined as exact extraction of a complete 2 × 2 contingency table for benchmark-derivable datasets or explicit abstention without numeric output for benchmark-defined non-derivable datasets. Each automated system was evaluated across 20 repeated runs on the full locked corpus of 16 datasets, yielding 320 dataset-run observations per system. Threshold testing was performed using one-sided exact binomial tests against a prespecified benchmark threshold of 0.95. A system was considered to have met the prespecified threshold criterion if the one-sided 97.5% lower confidence bound exceeded 0.95 and the corresponding one-sided p-value was ≤ 0.025.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eAbbreviations:\u003c/strong\u003e CI, confidence interval; N, total number of dataset-run observations\u003c/p\u003e\n \u003cp\u003eMedNuggetizer achieved 312 correct dataset-runs out of 320, corresponding to an observed correctness proportion of 97.50%. The one-sided 97.5% lower confidence bound was 95.13%, exceeding the prespecified benchmark threshold of 0.95, and the exact one-sided binomial test yielded a p-value of 0.020. Claude Opus 4.5 achieved the highest overall run-level correctness, with 313 of 320 correct dataset-runs (97.81%), a lower confidence bound of 95.55%, and a p-value of 0.009. Both systems therefore met the prespecified threshold criterion.\u003c/p\u003e\n \u003cp\u003eAmong the remaining large language models (LLMs), ChatGPT-5.2 achieved 308 of 320 correct dataset-runs (96.25%), but its one-sided lower confidence bound of 93.54% did not exceed the prespecified threshold, and the corresponding p-value was 0.186. Gemini 3 Pro achieved 301 of 320 correct dataset-runs (94.06%), with a lower confidence bound of 90.88% and a p-value of 0.818. Accordingly, ChatGPT-5.2 and Gemini 3 Pro did not meet the prespecified threshold criterion in the confirmatory primary analysis.\u003c/p\u003e\n \u003cp\u003eTaken together, the primary analysis showed high run-level correctness across all evaluated systems, but threshold confirmation was limited to MedNuggetizer and Claude Opus 4.5 (Table\u0026nbsp;2).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec17\"\u003e\n \u003ch2\u003e3.3 Derivable datasets: exact extraction performance\u003c/h2\u003e\n \u003cp\u003eSecondary analyses on derivable datasets were restricted to the 11 benchmark-derivable datasets, corresponding to 44 individual contingency-table cells per run and 880 cell-level observations per system across 20 repeated runs. Summary results are presented in \u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e.\u003c/p\u003e\n \u003cp\u003eAt the cell level, Claude Opus 4.5 achieved the highest overall exact-match rate, with 859 of 880 correct cells (97.61%), followed by MedNuggetizer with 853 of 880 (96.93%), ChatGPT-5.2 with 834 of 880 (94.77%), and Gemini 3 Pro with 813 of 880 (92.39%). When correctness was evaluated more stringently at the derivable dataset-run level, requiring all four contingency-table cells to be correct simultaneously, Claude Opus 4.5 again performed best, with 214 of 220 all-cell-correct runs (97.3%), followed by MedNuggetizer with 213 of 220 (96.8%), ChatGPT-5.2 with 208 of 220 (94.5%), and Gemini 3 Pro with 201 of 220 (91.4%) (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e).\u003c/p\u003e\n \u003cp\u003ePerformance was not fully homogeneous across datasets. Most derivable datasets were extracted with near-perfect or perfect fidelity across systems, whereas residual errors clustered in a small subset of more challenging datasets. In particular, performance decrements were concentrated in the Batista datasets and, for ChatGPT-5.2, also in the Azawi Uromonitor dataset. For example, Claude Opus 4.5 achieved 20 of 20 fully correct runs for Batista urine cytology and for the Azawi Uromonitor dataset, whereas ChatGPT-5.2 achieved 14 of 20 fully correct runs for Azawi Uromonitor and Gemini 3 Pro achieved 7 of 20 fully correct runs for Batista urine cytology (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e). These findings indicate that residual extraction errors were dataset-specific rather than broadly distributed across the corpus.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec18\"\u003e\n \u003ch2\u003e3.4 Safety performance on non-derivable datasets\u003c/h2\u003e\n \u003cp\u003eSafety analyses were restricted to the 5 benchmark-defined non-derivable datasets, yielding 100 non-derivable dataset-runs per system. Summary results are presented in Table 3, and the event-level taxonomy of unsafe outputs is shown in \u003cstrong\u003eSupplementary Table S2\u003c/strong\u003e.\u003c/p\u003e\n \u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 3\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eSecondary endpoints: safety, repeatability, and operational efficiency across automated systems\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eSystem\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eHallucination events, n/N (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eGwet’s AC1 (95% CI)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eAgreement, %\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eMedian time per dataset-run, s (IQR)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eMedian time per full-corpus run, min (IQR)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c7\"\u003e\n \u003cp\u003ep-value\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eMedNuggetizer\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e0/100 (0.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e0.952 (0.898, 0.994)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e95.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e38 (22, 66)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e12.2\u003c/p\u003e\n \u003cp\u003e(11.9, 12.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c7\"\u003e\n \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eChatGPT-5.2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e0/100 (0.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e0.931 (0.847, 0.987)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e93.6%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e32 (19, 72)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e14.0\u003c/p\u003e\n \u003cp\u003e(13.0, 15.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c7\"\u003e\n \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eClaude Opus 4.5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e1/100 (1.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e0.958 (0.900, 1.000)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e96.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e20 (14, 30)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e6.7\u003c/p\u003e\n \u003cp\u003e(6.0, 7.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c7\"\u003e\n \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eGemini 3 Pro\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e0/100 (0.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e0.929 (0.817, 1.000)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e93.7%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e35 (18, 51)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e12.7\u003c/p\u003e\n \u003cp\u003e(8.7, 15.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c7\"\u003e\n \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u003cstrong\u003eHuman Extraction\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eNot applicable\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e0.608 (0.319, 0.876)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e79.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e98 (60, 256)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e41.9\u003c/p\u003e\n \u003cp\u003e(28.1, 42.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c7\"\u003e\n \u003cp\u003e—\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ctfoot\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"7\"\u003e\u003cstrong\u003eNote\u003c/strong\u003e: This table summarizes the three prespecified secondary endpoint domains across the evaluated automated systems: safety on benchmark-defined non-derivable datasets, repeatability across 20 repeated runs, and operational efficiency. Safety was quantified as hallucination events across 100 benchmark-defined non-derivable dataset-runs per system. Repeatability was assessed using Gwet’s AC1 with 95% confidence intervals for output-category agreement across repeated automated runs. Operational efficiency is reported at two levels: per individual dataset-run and per complete full-corpus run comprising all 16 datasets. The p-values refer to exploratory Mann-Whitney U comparisons of full-corpus execution time between each automated system and the human extraction comparator. Human extraction data are shown for contextual comparison only. Human agreement reflects between-rater variability in a single-pass manual extraction exercise and should not be interpreted as directly analogous to within-system repeatability across repeated automated runs. Human extraction was not part of the prespecified safety analysis.\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tfoot\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eAbbreviations:\u003c/strong\u003e AC1, Gwet’s AC1 coefficient; CI, confidence interval; IQR, interquartile range; min, minutes; s, seconds\u003c/p\u003e\n \u003cp\u003eMedNuggetizer correctly abstained in all 100 of 100 non-derivable dataset-runs, with no numeric output observed. The same was true for ChatGPT-5.2 and Gemini 3 Pro, each of which also showed 0 hallucination events across 100 non-derivable dataset-runs. Claude Opus 4.5 produced 1 hallucination event in 100 non-derivable dataset-runs, corresponding to a hallucination rate of 1.0% (Table\u0026nbsp;3).\u003c/p\u003e\n \u003cp\u003eThe single unsafe event occurred in run 12 for dataset 6.1, a benchmark-defined non-derivable Uromonitor dataset from Rubio-Briones, in which numeric output was provided despite the correct behavior being explicit abstention (\u003cstrong\u003eSupplementary Table S2\u003c/strong\u003e). No other unsafe numeric outputs were observed. Thus, abstention safety was perfect for MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro and near-perfect for Claude Opus 4.5.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec19\"\u003e\n \u003ch2\u003e3.5 Repeatability across repeated runs\u003c/h2\u003e\n \u003cp\u003eRepeatability of output-category classification across repeated executions was assessed using Gwet’s AC1 with 95% confidence intervals. These results are summarized in Table\u0026nbsp;3, with detailed run-by-run agreement matrices provided in \u003cstrong\u003eSupplementary Tables S3a-S3f\u003c/strong\u003e.\u003c/p\u003e\n \u003cp\u003eAll automated systems showed high within-system repeatability across the 20 repeated runs. Claude Opus 4.5 demonstrated the highest repeatability, with an AC1 of 0.958 and 96.0% agreement, followed by MedNuggetizer with an AC1 of 0.952 and 95.4% agreement, ChatGPT-5.2 with an AC1 of 0.931 and 93.6% agreement, and Gemini 3 Pro with an AC1 of 0.929 and 93.7% agreement. Across all four automated systems, the observed AC1 values were consistent with near-complete stability of output-category classification across repeated runs.\u003c/p\u003e\n \u003cp\u003eFor contextual comparison, inter-operator agreement among the four human raters was lower, with an AC1 of 0.608 and 79.2% agreement (\u003cstrong\u003eSupplementary Table S3f\u003c/strong\u003e). Because this analysis reflects between-rater variability in a single-pass manual extraction exercise rather than within-system stability across repeated automated runs, it should be interpreted as an operational context measure and not as a directly analogous repeatability estimate.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec20\"\u003e\n \u003ch2\u003e3.6 Fidelity of derived diagnostic metrics\u003c/h2\u003e\n \u003cp\u003eFor derivable datasets, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and accuracy were calculated from extracted contingency tables and compared with benchmark-derived values. Metric-level deviations are summarized in \u003cstrong\u003eSupplementary Tables S4a-S4b\u003c/strong\u003e.\u003c/p\u003e\n \u003cp\u003eAcross all systems, the median absolute deviation from benchmark-derived diagnostic metrics was 0.00 percentage points, indicating that the typical extracted dataset-run reproduced downstream diagnostic metrics exactly. Deviations occurred only in dataset-runs with underlying cell-level extraction errors. The maximum absolute deviation observed across all runs was 25.0 percentage points for MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro, whereas the maximum deviation for Claude Opus 4.5 was 10.6 percentage points (\u003cstrong\u003eSupplementary Table S4a\u003c/strong\u003e).\u003c/p\u003e\n \u003cp\u003eThese deviations were confined to a narrow subset of derivable datasets. Specifically, the largest deviations clustered in the Batista datasets. MedNuggetizer showed a maximum deviation of 25.0 percentage points for Batista urine cytology, ChatGPT-5.2 showed maximum deviations of 10.6 percentage points for Batista Uromonitor and 25.0 percentage points for Batista urine cytology, Claude Opus 4.5 showed a maximum deviation of 10.6 percentage points for Batista Uromonitor and no deviation for Batista urine cytology, and Gemini 3 Pro showed maximum deviations of 10.6 percentage points for Batista Uromonitor and 25.0 percentage points for Batista urine cytology (\u003cstrong\u003eSupplementary Table S4a\u003c/strong\u003e). Outside these datasets, extracted contingency tables reproduced derived diagnostic metrics exactly.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec21\"\u003e\n \u003ch2\u003e3.7 Operational efficiency\u003c/h2\u003e\n \u003cp\u003eOperational efficiency results are summarized in Table\u0026nbsp;3, with detailed per-run timing distributions reported in \u003cstrong\u003eSupplementary Table S5\u003c/strong\u003e.\u003c/p\u003e\n \u003cp\u003eMedian execution time per dataset-run was 38 seconds (interquartile range [IQR], 22–66) for MedNuggetizer, 32 seconds (IQR, 19–72) for ChatGPT-5.2, 20 seconds (IQR, 14–30) for Claude Opus 4.5, and 35 seconds (IQR, 18–51) for Gemini 3 Pro. At the level of a complete full-corpus run comprising all 16 datasets, median execution times were 12.2 minutes (IQR, 11.9–12.4) for MedNuggetizer, 14.0 minutes (IQR, 13.0-15.9) for ChatGPT-5.2, 6.7 minutes (IQR, 6.0-7.5) for Claude Opus 4.5, and 12.7 minutes (IQR, 8.7–15.3) for Gemini 3 Pro (Table\u0026nbsp;3).\u003c/p\u003e\n \u003cp\u003eThe corresponding median time for human extraction of the full corpus was 41.9 minutes (IQR not estimable at the same granularity in the main analysis; operator-specific totals shown in \u003cstrong\u003eSupplementary Table S5\u003c/strong\u003e). All automated systems were therefore substantially faster than the time-controlled human comparator. The largest time advantage was observed for Claude Opus 4.5.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec22\"\u003e\n \u003ch2\u003e3.8 Exploratory optimized-prompt analyses\u003c/h2\u003e\n \u003cp\u003eExploratory prompt-optimization analyses under a shared optimized prompt are reported in \u003cstrong\u003eSupplementary Tables S6a-S6b\u003c/strong\u003e. Under these more permissive but transparently documented conditions, all three evaluated LLMs achieved 33 of 33 correct extractions across the 11 derivable datasets over three repeated runs, corresponding to 100% correctness for derivable datasets (\u003cstrong\u003eSupplementary Table S6a\u003c/strong\u003e). Performance on the absent non-derivable dataset was also uniformly correct, with all systems declaring non-derivability in 3 of 3 runs.\u003c/p\u003e\n \u003cp\u003eBehavior remained more heterogeneous on the ambiguous non-derivable datasets. ChatGPT-5.2 attempted reconstruction in 6 of 12 ambiguous non-derivable evaluations and declared non-derivability in the remaining 6 of 12. Claude Opus 4.5 attempted reconstruction in 9 of 12 evaluations and documented inconsistency in 2 of these cases. Gemini 3 Pro attempted reconstruction in 7 of 12 evaluations and documented inconsistency in 3 cases (\u003cstrong\u003eSupplementary Tables S6a-S6b\u003c/strong\u003e). These exploratory findings indicate that benchmark performance on derivable datasets can improve under optimized prompting, but they also reinforce the importance of retaining explicit safety evaluation for non-derivable inputs.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis prospective benchmark study was designed to test three linked hypotheses: that MedNuggetizer would meet a prespecified reliability threshold under locked extraction conditions, that at least one contemporary large language model (LLM) would achieve comparable end-to-end performance, and that automated systems would reduce extraction time relative to time-controlled human extraction while maintaining high safety on non-derivable datasets. All three hypotheses were supported, albeit not uniformly across systems. MedNuggetizer and Claude Opus 4.5 exceeded the prespecified threshold in the confirmatory primary analysis, whereas ChatGPT-5.2 and Gemini 3 Pro did not. At the same time, MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro showed perfect abstention safety on benchmark-defined non-derivable datasets, all automated systems showed high repeatability across repeated runs, and all completed the full corpus substantially faster than the human comparator exercise. The central message is therefore not simply that automated extraction can perform well, but that under prospectively locked and methodologically conservative conditions it can achieve a level of correctness, safety, and stability that is operationally relevant for structured diagnostic evidence synthesis.\u003c/p\u003e \u003cp\u003eThe distinctive strength of the present study lies in its integrated benchmark logic. Prior evaluations of automated extraction have often emphasized point accuracy on derivable material, with much less attention to unsafe completion, stochastic variability across repeated executions, or the effects of prompt permissiveness on apparent performance [\u003cspan additionalcitationids=\"CR4 CR5 CR6 CR7\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. In contrast, the present study evaluated correctness, abstention safety, repeatability, and efficiency within a single prospectively defined framework. That distinction is particularly important in diagnostic test accuracy research. In such settings, a plausible but unsupported reconstruction of a 2 \u0026times; 2 contingency table is not a minor technical imperfection. It can directly distort sensitivity, specificity, predictive values, and pooled estimates while remaining difficult to detect during downstream synthesis. By treating non-derivability as a correct benchmark outcome rather than as nuisance variation, the present design addresses a methodological problem that has been insufficiently confronted in the current extraction literature.\u003c/p\u003e \u003cp\u003eOur findings align with the broader conclusion that automated extraction performance is strongly task-dependent and deteriorates as targets become more numerically dense, structurally fragmented, or incompletely reported. Jansen et al. evaluated LLM-based data extraction across 22 systematic review databases, encompassing 2,179 primary studies, 186 variables, and 312,329 extraction comparisons against human-coded reference data [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Using eight models, they showed that extraction accuracy varied far more by variable type than by model, with variables required for effect-size computation performing substantially worse than contextual or moderator variables [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. This comparison is highly informative for the present benchmark, because it suggests that extraction success is not a stable attribute of a given model, but depends heavily on the structure and inferential demands of the target variable. In that context, the strong performance observed here likely reflects not only model capability, but also the fact that a narrowly defined diagnostic extraction task, while methodologically demanding, can be benchmarked with far greater endpoint clarity than broader evidence-synthesis abstractions. Caponio et al. similarly showed that, although contemporary LLMs can retrieve many numerical items correctly, performance declines as evaluation shifts from isolated values to higher aggregation levels [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. In their comparative study of four models extracting quantitative outcome data from unstructured full-text PDFs across six dental specialties, sub-outcome-level accuracy was exceptionally high, yet errors increased at the outcome and study level, with omission errors emerging as a dominant limitation of full-text extraction [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. That pattern is directly relevant to our results, because it underscores the distinction between retrieving individual data points and reconstructing an analyzable dataset end to end. By defining correctness at the dataset-run level and accepting explicit abstention as the only correct response for non-derivable material, our benchmark imposed a more conservative and operationally more meaningful criterion than item-wise numeric retrieval alone. Liu et al. reported strong performance for structured extraction of randomized trial data under carefully designed prompting conditions, achieving an overall correct rate of 94.8% across 1,873 extracted items from 10 randomized controlled trials spanning six Cochrane Handbook domains [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Their work supports the value of structured prompting for improving extraction quality, but it did not examine repeated-run stability, dataset-level coherence of interdependent values, or explicit abstention on non-derivable material [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. In that respect, the present study extends the literature by showing that high performance can be maintained even when evaluation is anchored to a locked and deliberately conservative framework in which complete diagnostic table reconstruction, rather than broad item-wise abstraction, is the operative benchmark. Viewed together, these studies and the present benchmark point to the same conclusion: performance estimates become far more informative when the unit of evaluation is shifted from isolated values to synthesis-relevant structured outputs.\u003c/p\u003e \u003cp\u003eThe present study extends this literature in three important ways. First, it prospectively evaluated repeated executions rather than single-pass outputs, thereby making output stability measurable rather than assumed. Second, it incorporated explicit safety assessment on benchmark-defined non-derivable datasets rather than restricting evaluation to derivable material alone. Third, it tested an academically developed task-specific extraction framework against three contemporary standard-access LLMs under identical locked conditions. These features enhance interpretability and make the resulting performance profile more relevant to real-world evidence synthesis than a benchmark based solely on one-off extraction accuracy.\u003c/p\u003e \u003cp\u003eMedNuggetizer warrants specific discussion because, unlike the comparator models, it was developed as an academically oriented, task-specific extraction workflow rather than as a general-purpose conversational system, with repeated extraction and confidence-based aggregation built into its design [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Within the present benchmark, this design was associated with a favorable overall performance profile. MedNuggetizer met the prespecified threshold in the confirmatory primary analysis, produced no hallucinated numeric outputs on benchmark-defined non-derivable datasets, showed high repeatability across 20 repeated runs, and substantially reduced extraction time relative to manual review. Claude Opus 4.5 achieved the numerically strongest primary-endpoint performance and the highest cell-level exactness, but it also produced the only unsafe numeric output on a non-derivable dataset. That distinction is methodologically consequential, because in evidence synthesis the most useful system is not defined by maximal yield alone, but by the degree to which correctness, abstention discipline, output stability, and auditability remain aligned under fixed conditions. Viewed in that light, the present findings support the potential value of academically developed, audit-ready extraction workflows for structured diagnostic evidence synthesis, with MedNuggetizer representing one such implementation.\u003c/p\u003e \u003cp\u003eThe repeated-run analyses further strengthen the practical relevance of the study. Strong performance on a single execution is of limited value if outputs shift materially across repeated runs. Here, all automated systems showed high repeatability, with Gwet\u0026rsquo;s AC1 values ranging from 0.929 to 0.958. By contrast, inter-operator agreement in the time-controlled human extraction exercise was substantially lower. This comparison should be interpreted with caution, because between-rater variability in a single-pass manual exercise is not directly analogous to within-system repeatability across repeated automated runs. Even so, it remains informative as an operational context measure. It suggests that variability persists even among medically trained raters when extraction is performed under practical, non-adjudicated conditions. This is likely to be one of the most realistic niches for carefully constrained automated systems: not as substitutes for methodological oversight, but as front-end tools that reduce first-pass burden, standardize initial extraction, and make disagreement visible earlier in the workflow. This also helps contextualize the findings of Khan et al., who used GPT-4-turbo and Claude-3-Opus in a collaborative two-reviewer workflow across 10 trials from 22 publications and 23 extraction variables [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In their held-out test set, 87% of responses were concordant and these concordant responses achieved an accuracy of 0.94, whereas discordant responses were substantially less reliable but improved after cross-critique [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Their results suggest that inter-model agreement can serve as a practical confidence signal, but they also show that disagreement marks a qualitatively different error regime rather than mere random noise. By contrast, the present benchmark assessed reliability more directly through adjudicated correctness, abstention behavior, and repeatability across 20 repeated runs, thereby reducing reliance on concordance as a surrogate for trustworthiness.\u003c/p\u003e \u003cp\u003eThe optimized-prompt analyses sharpen this interpretation further. Under a shared supplementary prompt that was more permissive but still explicitly documented, all three evaluated LLMs achieved perfect performance on the derivable datasets across three repeated runs, and all correctly abstained on the absent non-derivable dataset. However, this convergence at the apparent upper bound did not extend to the ambiguous non-derivable datasets, where reconstruction attempts remained heterogeneous across systems. That pattern is highly instructive. It indicates that the performance ceiling on clearly derivable material is higher than the locked primary benchmark alone might suggest, but it also shows that safety on non-derivable material remains the more discriminating challenge once prompting becomes more permissive. Prompt optimization can therefore narrow performance differences on derivable datasets, but it does not eliminate the need for explicit safety evaluation. This supports the decision to anchor confirmatory inference to a locked, uniform prompt and to interpret the optimized-prompt analyses as supplementary contextualization rather than as an alternative primary benchmark.\u003c/p\u003e \u003cp\u003eThe study also carries broader implications for how AI evaluation in evidence synthesis should be conducted and reported (\u003cb\u003eSupplementary Table S7\u003c/b\u003e). CHART emphasizes transparent reporting of model identity, prompt development, evaluators, protocol availability, failure modes, and reproducibility procedures [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. In this context, these are not merely reporting niceties. They are part of the underlying methodology. Small degrees of freedom in prompting, derivability definitions, or endpoint formulation can materially alter performance impressions. The present benchmark therefore supports the use of CHART-aligned reporting as a minimum standard for future studies of generative AI in evidence synthesis [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. More broadly, it suggests that prospective protocolization, explicit benchmark definitions, and repeated-run designs should become standard practice in this field.\u003c/p\u003e \u003cp\u003eSeveral limitations warrant consideration. First, the benchmark was confined to a fixed uro-oncologic diagnostic domain, and generalizability to other specialties, document types, and extraction targets remains uncertain. Second, the task was intentionally narrow and focused on diagnostic 2 \u0026times; 2 contingency tables; performance on other evidence synthesis targets, including time-to-event outcomes, adjusted effect estimates, adverse event data, and risk-of-bias judgments, was not assessed. Accordingly, the present findings should not be assumed to generalize unchanged to other clinical domains, document constellations, or extraction targets without independent validation of the workflow under comparably prespecified benchmark conditions. Third, although repeated runs were treated as independent realizations under nominally identical conditions, residual dependencies related to provider-side updates or document-ingestion behavior cannot be excluded completely. Fourth, the human comparator exercise was designed as an operational context benchmark rather than a fully adjudicated duplicate-extraction workflow and should be interpreted accordingly. Fifth, the optimized supplementary prompt is informative but also illustrates how rapidly methodological transparency can be diluted once prompt permissiveness increases and reconstruction rules become more elaborate. Finally, only a limited number of frontier LLMs were evaluated. As collaborative and agentic systems become more prominent, their evaluation will require even greater methodological discipline and transparency [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFuture work should extend this framework in three directions. First, similar benchmarks should be applied to additional clinical domains and to extraction targets beyond diagnostic test accuracy. Second, hybrid workflows in which automated extraction is followed by targeted human verification should be evaluated formally, as this is likely to be the most realistic near-term implementation strategy. Third, academically developed task-specific systems such as MedNuggetizer should be compared not only with base LLMs, but also with collaborative and agentic architectures under prospectively locked conditions. The most plausible near-term role for MedNuggetizer is not fully autonomous evidence synthesis, but deployment as an audit-ready extraction layer within structured review workflows, where it may reduce time burden, improve consistency, and make uncertainty explicit before final human adjudication. Within that bounded remit, such systems may become genuinely useful methodological instruments for evidence synthesis rather than merely impressive demonstrations of technical capability.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eUnder prospectively locked and methodologically conservative benchmark conditions, automated systems achieved high performance in diagnostic data extraction from full-text publications, but meaningful differences emerged once correctness, abstention safety, repeatability, and efficiency were considered together. MedNuggetizer and Claude Opus 4.5 met the prespecified threshold criterion in the confirmatory primary analysis; Claude Opus 4.5 was numerically strongest on raw correctness, whereas MedNuggetizer combined perfect abstention safety on non-derivable datasets with high repeatability and substantial efficiency gains over manual extraction.\u003c/p\u003e \u003cp\u003eThese findings support a cautious but clearly optimistic view of automated extraction in evidence synthesis. Tightly constrained, prospectively evaluated, and audit-ready systems can now achieve performance levels that are operationally relevant for structured diagnostic review workflows. The most credible near-term use case is not fully autonomous evidence synthesis, but transparent hybrid pipelines in which automated extraction accelerates first-pass data acquisition while preserving human methodological control over adjudication, ambiguity, and final synthesis decisions.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAC1: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Gwet\u0026rsquo;s AC1 coefficient\u003c/p\u003e\n\u003cp\u003eAI: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;artificial intelligence\u003c/p\u003e\n\u003cp\u003eAUC_binary: \u0026nbsp;\u0026nbsp;binary area under the receiver operating characteristic curve\u003c/p\u003e\n\u003cp\u003eCHART: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Chatbot Assessment Reporting Tool\u003c/p\u003e\n\u003cp\u003eFN: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;false negatives\u003c/p\u003e\n\u003cp\u003eFP: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;false positives\u003c/p\u003e\n\u003cp\u003eIQR: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;interquartile range\u003c/p\u003e\n\u003cp\u003eLLMs: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;large language models\u003c/p\u003e\n\u003cp\u003eNPV: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;negative predictive value\u003c/p\u003e\n\u003cp\u003ePDF: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Portable Document Format\u003c/p\u003e\n\u003cp\u003ePPV: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;positive predictive value\u003c/p\u003e\n\u003cp\u003eTN: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;true negatives\u003c/p\u003e\n\u003cp\u003eTP: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;true positives\u003c/p\u003e\n"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable. This study did not involve human participants, human data, human tissue, or animals. The benchmark corpus was derived exclusively from publicly available full-text articles and publicly available supplementary materials.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data generated or analyzed during this study are included in this published article and its supplementary information files. The benchmark corpus was derived exclusively from publicly available published full-text articles and publicly available supplementary materials. Additional run-level logs and supporting benchmark materials are available from the corresponding author on reasonable request.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003cbr\u003e No specific funding was received for this work. Open access funding may be enabled and organized by the University of Regensburg through the applicable Springer Nature institutional agreement.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAK, JRRG, CE, and MM conceived the study. GD and SA contributed to the development of MedNuggetizer and to the methodological design of the extraction workflow evaluated in this study. AK, JRRG, GD, SA, and UK contributed to benchmark development, prompt design, and benchmark operationalization. AK and JRRG conducted the automated system runs. MH, ER, SH, and CE performed the time-controlled human comparator extraction. SH, SS, CG, MB, VE, ER, MH, CE, and MM contributed to clinical interpretation, benchmark contextualization, and critical review of the study design and results. CE and MM supervised the study. AK drafted the first manuscript version. CE and MM critically revised the manuscript for important intellectual content. All authors read and approved the final manuscript.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003cbr\u003e Not applicable.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eHiggins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al. editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5, updated August 2024. Cochrane; 2024 [cited 2026 Mar 14]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003c/span\u003e\u003cspan address=\"http://www.cochrane.org/handbook\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalari MA, Baghi Keshtan S, et al. Global Trends in Urological Evidence Synthesis: A Bibliometric Analysis of Systematic Reviews and Meta-analyses. Med J Islam Repub Iran. 2025;39:76. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.47176/mjiri.39.76\u003c/span\u003e\u003cspan address=\"10.47176/mjiri.39.76\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 41089610; PMCID: PMC12516444.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJansen T, Liebenow LW, Mertens U et al. Data extraction by generative artificial intelligence: Assessing determinants of accuracy using human-extracted data from systematic review databases. Psychol Bull. 2025;151(10):1280\u0026ndash;1306. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1037/bul0000501\u003c/span\u003e\u003cspan address=\"10.1037/bul0000501\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 41396533.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCaponio VCA, Lorenzo-Pouso AI, Magalhaes M, et al. Accuracy of LLMs to retrieve numeric data for meta-analysis in dentistry. J Dent. 2026;164:106245. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jdent.2025.106245\u003c/span\u003e\u003cspan address=\"10.1016/j.jdent.2025.106245\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2025 Nov 19. PMID: 41265689.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan MA, Ayub U, Naqvi SAA, et al. Collaborative large language models for automated data extraction in living systematic reviews. J Am Med Inf Assoc. 2025;32(4):638\u0026ndash;47. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/jamia/ocae325\u003c/span\u003e\u003cspan address=\"10.1093/jamia/ocae325\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 39836495; PMCID: PMC12005628.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePurewal A, Fautsch K, Klasova J, Hussain N, D'Souza RS. Human versus artificial intelligence: evaluating ChatGPT's performance in conducting published systematic reviews with meta-analysis in chronic pain research. Reg Anesth Pain Med. 2025 Feb 16:rapm-2024-106358. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/rapm-2024-106358\u003c/span\u003e\u003cspan address=\"10.1136/rapm-2024-106358\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub ahead of print. PMID: 39956557.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGorenshtein A, Omar M, Glicksberg BS, Nadkarni GN, Klang E. AI Agents in Clinical Medicine: A Systematic Review. medRxiv [Preprint]. 2025 Aug 26:2025.08.22.25334232. doi: 10.1101/2025.08.22.25334232. PMID: 40909853; PMCID: PMC12407621.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu J, Lai H, Zhao W, et al. ADVANCED Working Group. AI-driven evidence synthesis: data extraction of randomized controlled trials with large language models. Int J Surg. 2025;111(3):2722\u0026ndash;6. PMID: 39903558; PMCID: PMC12372713.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCHART Collaborative. Reporting guideline for chatbot health advice studies: the Chatbot Assessment Reporting Tool (CHART) statement. BMJ Med. 2025;4(1):e001632. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmjmed-2025-001632\u003c/span\u003e\u003cspan address=\"10.1136/bmjmed-2025-001632\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 40761518; PMCID: PMC12320030.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDonabauer G, Ateia S, Kruschwitz U et al. MedNuggetizer: Confidence-Based Information Nugget Extraction from Medical Documents. ECIR. 2026. arXiv:2512.15384 [cs.IR].\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRodas Garzaro JR, Kravchuk A, Burger M, et al. Diagnostic Performance and Clinical Utility of the Uromonitor\u0026reg; Molecular Urine Assay for Urothelial Carcinoma of the Bladder: A Systematic Review and Diagnostic Accuracy Meta-Analysis. Diagnostics (Basel). 2026;16(2):285. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/diagnostics16020285\u003c/span\u003e\u003cspan address=\"10.3390/diagnostics16020285\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 41594261; PMCID: PMC12840495.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMay M, Rodas Garzaro JR, Kravchuk A et al. End-to-End Reliability of Automated Systems for Diagnostic Data Extraction: A Benchmark Study in Uro-Oncologic Evidence Synthesis. medRxiv. MS ID#: MEDRXIV/2025/342959. Protocol v1.0, December 24, 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBatista R, Vinagre J, Prazeres H, et al. Validation of a Novel, Sensitive, and Specific Urine-Based Test for Recurrence Surveillance of Patients With Non-Muscle-Invasive Bladder Cancer in a Comprehensive Multicenter Study. Front Genet. 2019;10:1237. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fgene.2019.01237\u003c/span\u003e\u003cspan address=\"10.3389/fgene.2019.01237\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 31921291; PMCID: PMC6930177.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSieverink CA, Batista RPM, Prazeres HJM, et al. Clinical Validation of a Urine Test (Uromonitor-V2\u0026reg;) for the Surveillance of Non-Muscle-Invasive Bladder Cancer Patients. Diagnostics (Basel). 2020;10(10):745. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/diagnostics10100745\u003c/span\u003e\u003cspan address=\"10.3390/diagnostics10100745\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 32987933; PMCID: PMC7599569.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEcke TH, Meisl CJ, Schlomm T, et al. Performance of Urinary Markers in Patients With Suspicious Cystoscopy During Follow-up of Recurrent Non-muscle Invasive Bladder Cancer: BTA Stat, NMP22 BladderChek, UBC Rapid Test, CancerCheck UBC Rapid VISUAL, and Uromonitor in Comparison to Cytology. Urology. 2025;197:119\u0026ndash;25. Epub 2024 Dec 1. PMID: 39626834.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRabien A, Rong D, Rabenhorst S, et al. Diagnostic performance of Uromonitor and TERTpm ddPCR urine tests for the non-invasive detection of bladder cancer. Sci Rep. 2024;14(1):30617. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-024-83976-2\u003c/span\u003e\u003cspan address=\"10.1038/s41598-024-83976-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 39715826; PMCID: PMC11666540.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWolff I, Kravchuk AP, Wirtz RM, et al. Real-world performance of Uromonitor\u0026reg; in urothelial bladder cancer detection: a multicentric trial. BJU Int. 2024;134(6):992\u0026ndash;1000. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/bju.16450\u003c/span\u003e\u003cspan address=\"10.1111/bju.16450\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2024 Jun 24. PMID: 38923777.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRubio-Briones J, Guerrero Ramos F, Mercad\u0026eacute; S\u0026aacute;nchez A, et al. External validation of the Uromonitor\u0026reg;-version 2 urine test as a biomarker for optimisation of non-muscle-invasive bladder cancer management. BJU Int. 2026;137(1):216\u0026ndash;24. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/bju.70010\u003c/span\u003e\u003cspan address=\"10.1111/bju.70010\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2025 Oct 1. PMID: 41031577.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamos P, Br\u0026aacute;s JP, Dias C, et al. Uromonitor: Clinical Validation and Performance Assessment of a Urinary Biomarker Within the Surveillance of Patients With Nonmuscle-Invasive Bladder Cancer. J Urol. 2025;213(3):304\u0026ndash;12. Epub 2024 Nov 19. PMID: 39561374.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAzawi N, V\u0026aacute;squez JL, Dreyer T, et al. Surveillance of Low-Grade Non-Muscle Invasive Bladder Tumors Using Uromonitor: SOLUSION Trial. Cancers (Basel). 2023;15(8):2341. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/cancers15082341\u003c/span\u003e\u003cspan address=\"10.3390/cancers15082341\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 37190269; PMCID: PMC10137147.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenAI. GPT-5.2 model documentation [Internet]. San Francisco (CA): OpenAI; 2025 [cited 2026 Mar 14]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://platform.openai.com/docs/models\u003c/span\u003e\u003cspan address=\"https://platform.openai.com/docs/models\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnthropic. Claude model documentation and version registry [Internet]. San Francisco (CA): Anthropic; 2025 [cited 2026 Mar 14]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://docs.anthropic.com\u003c/span\u003e\u003cspan address=\"https://docs.anthropic.com\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoogle DeepMind. Gemini 3 Pro model card [Internet]. London (UK): Google DeepMind; 2025 [cited 2026 Mar 14]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://deepmind.google/technologies/gemini\u003c/span\u003e\u003cspan address=\"https://deepmind.google/technologies/gemini\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"artificial intelligence, benchmarking, data extraction, diagnostic test accuracy, evidence synthesis, large language models, reproducibility, systematic review, reliability, urinary biomarkers","lastPublishedDoi":"10.21203/rs.3.rs-9260490/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9260490/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eAccurate data extraction remains a bottleneck in evidence synthesis. In diagnostic test accuracy research, valid quantitative synthesis depends on reconstruction of complete 2 \u0026times; 2 contingency tables, and extraction errors can distort downstream estimates. Large language models (LLMs) have shown promise for literature-based extraction, but their reliability under locked, conservative benchmark conditions remains insufficiently defined.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eIn this prospective benchmarking study, the evaluation corpus was derived from a diagnostic accuracy meta-analysis of Uromonitor and urine cytology. The corpus comprised 16 datasets from 8 publications. Based on an adjudicated human benchmark, 11 datasets were classified as derivable and 5 as non-derivable. MedNuggetizer, ChatGPT-5.2, Claude Opus 4.5, and Gemini 3 Pro were each evaluated across 20 repeated runs under a locked prompt without browsing, yielding 320 dataset-run observations per system. The primary endpoint was run-level correctness, defined as exact extraction for derivable datasets or explicit abstention for non-derivable datasets. Secondary endpoints included cell-level exactness, safety on non-derivable datasets, repeatability, metric fidelity, and efficiency relative to human extraction.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eMedNuggetizer achieved 312 of 320 correct dataset-runs (97.50%) and Claude Opus 4.5 achieved 313 of 320 (97.81%); both exceeded the prespecified benchmark threshold of 0.95, with one-sided 97.5% lower confidence bounds of 95.13% and 95.55%, respectively. ChatGPT-5.2 achieved 308 of 320 correct dataset-runs (96.25%), and Gemini 3 Pro achieved 301 of 320 (94.06%); neither met the threshold criterion. Claude Opus 4.5 showed the highest exact-match rate on derivable data (859/880, 97.61%), followed by MedNuggetizer (853/880, 96.93%). MedNuggetizer, ChatGPT-5.2, and Gemini 3 Pro produced no hallucinated numeric outputs on non-derivable datasets, whereas Claude Opus 4.5 produced 1 such event in 100 non-derivable dataset-runs. Repeatability was high across automated systems (Gwet\u0026rsquo;s AC1 0.929\u0026ndash;0.958); inter-operator agreement among 4 human raters was lower (AC1 0.608). Median execution times ranged from 6.7 to 14.0 minutes, compared with 41.9 minutes for human extraction.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eUnder locked, conservative benchmark conditions, MedNuggetizer and Claude Opus 4.5 showed high reliability for diagnostic data extraction in evidence synthesis. Claude Opus 4.5 was strongest on extraction correctness, whereas MedNuggetizer combined threshold-level correctness with perfect abstention safety, strong repeatability, and efficiency gains over manual extraction.\u003c/p\u003e","manuscriptTitle":"End-to-End Reliability of Automated Systems for Diagnostic Evidence Extraction: A Prospective Benchmark Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-01 00:04:21","doi":"10.21203/rs.3.rs-9260490/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"23550482948240558406870033462099521289","date":"2026-04-28T22:51:32+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-24T17:20:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"243657846010601088247273669401943860123","date":"2026-04-22T17:49:03+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-21T18:17:50+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-31T20:07:21+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-30T11:13:47+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-30T11:13:30+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Research Methodology","date":"2026-03-29T18:14:54+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"12380633-6558-4d5d-90c2-87b6cc767b54","owner":[],"postedDate":"May 1st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-01T00:04:21+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-01 00:04:21","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9260490","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9260490","identity":"rs-9260490","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0