Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Diagnostic decision support systems (DDSS) are increasingly deployed at scale, yet their diagnostic accuracy is insufficiently monitored once integrated into care. Traditional post-market surveillance relies on clinician review, which is costly, slow, and difficult to sustain. Large language models (LLMs) may offer a scalable and potentially automatable solution, but their performance in real-world monitoring remains unknown. Methods We conducted a diagnostic accuracy substudy within ESSENCE, a prospective evaluation of Ada Health’s DDSS integrated into Portugal’s largest private healthcare network. Clinical notes and ICD-10 diagnoses from 498 encounters were anonymised and classified using a filter–map–match framework. Manual clinician review served as the reference standard. We compared eligibility classification and condition mapping between clinicians and GPT-4.1 and GPT-5, and assessed diagnostic accuracy of two DDSS versions using both reference sets. Findings Manual review classified 385 of 498 encounters (77·3%) as eligible for diagnostic comparison. GPT-5 reproduced these classifications with 84·7% accuracy (κ = 0·57), showing high sensitivity but only moderate specificity. Among 347 encounters judged eligible by both approaches, GPT-5 exactly matched clinician-assigned diagnoses in 93·6% and proposed clinically plausible alternatives in 3·5%. Diagnostic accuracy estimates based on manual versus GPT-5 mappings were statistically indistinguishable at Top-1 and Top-3 across the full analyzable sets, with one significant difference at Top-5. In the overlapping 346 cases, no statistical differences were observed. Across both reference sets, the experimental DDSS version outperformed the original only at the Top-5 threshold. Interpretation LLMs can reproduce clinician review of real-world diagnostic encounters with close agreement. While GPT-5 performed comparably to clinicians for condition mapping, the eligibility filtering step - deciding which encounters should enter the diagnostic-accuracy analysis - remains the main source of divergence and is the priority for improvement. Embedding such approaches into health systems could enable automated and continuous performance and safety monitoring and support regulatory compliance. Broader evaluations across diverse care settings are needed to establish generalisability and equity impact.
Full text 156,116 characters · extracted from preprint-html · click to expand
Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study Fabienne Cotte, Marcel Schmude, Philipp Bode, Oula Suliman, Filipa Dias Lourenço, and 15 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8022874/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Background Diagnostic decision support systems (DDSS) are increasingly deployed at scale, yet their diagnostic accuracy is insufficiently monitored once integrated into care. Traditional post-market surveillance relies on clinician review, which is costly, slow, and difficult to sustain. Large language models (LLMs) may offer a scalable and potentially automatable solution, but their performance in real-world monitoring remains unknown. Methods We conducted a diagnostic accuracy substudy within ESSENCE, a prospective evaluation of Ada Health’s DDSS integrated into Portugal’s largest private healthcare network. Clinical notes and ICD-10 diagnoses from 498 encounters were anonymised and classified using a filter–map–match framework. Manual clinician review served as the reference standard. We compared eligibility classification and condition mapping between clinicians and GPT-4.1 and GPT-5, and assessed diagnostic accuracy of two DDSS versions using both reference sets. Findings Manual review classified 385 of 498 encounters (77·3%) as eligible for diagnostic comparison. GPT-5 reproduced these classifications with 84·7% accuracy (κ = 0·57), showing high sensitivity but only moderate specificity. Among 347 encounters judged eligible by both approaches, GPT-5 exactly matched clinician-assigned diagnoses in 93·6% and proposed clinically plausible alternatives in 3·5%. Diagnostic accuracy estimates based on manual versus GPT-5 mappings were statistically indistinguishable at Top-1 and Top-3 across the full analyzable sets, with one significant difference at Top-5. In the overlapping 346 cases, no statistical differences were observed. Across both reference sets, the experimental DDSS version outperformed the original only at the Top-5 threshold. Interpretation LLMs can reproduce clinician review of real-world diagnostic encounters with close agreement. While GPT-5 performed comparably to clinicians for condition mapping, the eligibility filtering step - deciding which encounters should enter the diagnostic-accuracy analysis - remains the main source of divergence and is the priority for improvement. Embedding such approaches into health systems could enable automated and continuous performance and safety monitoring and support regulatory compliance. Broader evaluations across diverse care settings are needed to establish generalisability and equity impact. Biological sciences/Computational biology and bioinformatics Health sciences/Diseases Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Figures Figure 1 Figure 2 Background Artificial intelligence (AI) has moved from proof-of-concept studies to approved tools in routine care across nearly all domains of medicine. AI now interprets radiology, guides ultrasound, supports pathology, triages symptoms, predicts deterioration, and optimises operations( 1 – 5 ). Every stage of care delivery is being re-examined through the lens of AI( 6 ). Robust evidence is needed to validate benefits, ensure safe use, and meet regulatory requirements for demonstrations of performance and clinical value, including methods in the postmarket phase, i.e., post-market surveillance and post-market studies (known as post-market clinical follow-up)( 7 , 8 ). Yet most evaluations rely on curated test sets, reader studies, or single-centre trials( 9 ). These establish feasibility but rarely capture the heterogeneity of practice, where performance can deteriorate once exposed to diverse patients and workflows( 10 – 12 ). Silent-mode trials help bridge this gap by embedding AI in clinical settings without influencing care, exposing biases and operational challenges before full deployment. In Kenya, an LLM tool for error detection was run in shadow mode to uncover workflow issues in primary care( 4 ). In the UK, a machine-learning framework for predicting hospital-acquired infections in patients with neurological impairments underwent a year-long silent-mode evaluation, demonstrating superior risk stratification compared with early warning scores( 13 ). Pre-deployment quality control is also essential. Asgari et al (2025) describe a clinician-in-the-loop framework (CREOLA) that stress-tests documentation assistants with a medical-device-style risk scheme, permitting only versions that meet safety thresholds to progress( 14 ). Once implemented, AI generates continuous real-world data - an opportunity and obligation to monitor performance, safety, and equity at scale( 15 ). Surveillance should be embedded in continuous improvement cycles using process control, A/B testing, and governance-linked monitoring, yet few organisations have operationalised such systems( 7 ). Traditional post-market evaluation relies on clinician review of AI outputs( 16 ): reliable but costly, poorly scalable and almost impossible to automate. Alternatives link AI predictions to registry-confirmed diagnoses, readmissions, or mortality( 12 , 17 ), or track whether clinicians accept or reject AI outputs (“passive labelling”)( 7 , 11 , 18 ). Patient- and clinician-reported outcomes highlight usability and trust but remain subjective and intermittent( 7 ). An emerging strategy is “AI to monitor AI.” LLMs have been tested for automated chart review and adverse drug reaction detection( 19 , 20 ). These systems replicate tasks previously reliant on clinicians, offering scale, consistency and automatability at lower cost. Yet they are not equivalent to humans: LLMs excel at structured extraction but struggle with ambiguity or missing data, risking error amplification. Human reviewers, by contrast, interpret context and detect what is absent( 19 ). The most promising models therefore combine both, with LLMs surfacing structured outputs and clinicians adjudicating uncertain or high-risk cases. A key domain where this approach could be transformative is monitoring diagnostic decision support systems (DDSS). For systems generating condition suggestions, a central challenge is assessing alignment with confirmed diagnoses( 21 ). Until now, this has relied on panels of physicians judging concordance with clinical diagnoses( 16 ). The task is further complicated by the nature of clinical documentation: recorded diagnoses are often ambiguous or coded at a high level of generality, may reflect only part of a multimorbid presentation, or remain provisional while investigations are ongoing. Objectives This study aimed to ( 1 ) develop a reproducible methodology for assessing DDSS outputs against confirmed diagnoses using a framework that accounts for diagnostic ambiguity and context, and ( 2 ) evaluate whether LLMs can replicate or scale this process by automating diagnostic-context classification and condition mapping with sufficient reliability to support post-market surveillance. Methods Study Design and Setting This diagnostic substudy was nested within ESSENCE, a prospective quality-improvement evaluation of a digital diagnostic decision support system (DDSS) developed by Ada Health (Berlin, Germany) and implemented in Portugal’s largest private healthcare network (CUF). Adults (≥ 18 years) who completed a Portuguese-language symptom assessment in the myCUF mobile application between Nov 1, 2023, and Oct 31, 2024, provided electronic consent, and shared their report with a CUF clinician were eligible. For this substudy, we included all encounters with both an ICD-10–coded diagnosis and clinical notes available in the electronic health record (EHR). Data Collection Procedures Data collection procedures are outlined in Fig. 1 . The full ESSENCE protocol has been reported previously (Cotte et al 2025, preprint). In brief, participants completed a symptom assessment and pre- and post-assessment care intentions, and shared the report with CUF physicians as part of routine care. Healthcare-seeking behavior was tracked through EHR review and follow-up email surveys. A research assistant extracted unstructured clinical notes, ICD-10-coded diagnoses, and study-related metadata into a secure electronic system. All free-text notes and diagnoses were translated from Portuguese to English using a locally run Mistral-Nemo-Instruct-2407, following manual anonymisation and approval by the data protection officer. Anonymisation followed standardized rules: removal of names, dates, occupations, and location identifiers, and conversion of exact ages into bands. A summary of all actions and affected cases is provided in Supplementary Table 1. Manual filtering and mapping (reference set) Clinicians employed by the DDSS and familiar with its ontology reviewed each case to establish a reference standard. They determined whether the clinic encounter addressed the same problem as the DDSS assessment and whether a final diagnosis had been reached. Cases with a different complaint, only a symptom recorded, or unspecific ICD-10 codes without clarifying details were excluded. Definitive diagnoses were mapped to one or more DDSS conditions; broad codes were refined using consultation notes. If no equivalent existed, cases were recorded as “condition not covered.” Reviews were performed independently, with disagreements resolved by consensus and a final reviewer ensuring consistency. Definitions of all diagnostic-context labels are provided in Supplementary Table 2. LLM-based filtering and mapping We next evaluated whether a LLM could reproduce this workflow. The model received the same anonymised and translated notes, ICD-10 codes, and user-entered symptoms from the DDSS assessment. Prompts instructed the LLM to act as a medical coding specialist and to follow the same four-step logic: ( 1 ) filtering cases based on the match between the user's complaint and the clinical encounter, ( 2 ) assessing the specificity of the final diagnosis, ( 3 ) refining unspecific diagnoses using details from the clinical notes, and ( 4 ) mapping the final, specified diagnosis to the DDSS's condition ontology. Few-shot examples of each diagnostic-context label were included to standardise reasoning, and outputs were returned in structured JSON format. Because the full DDSS ontology exceeded the model’s context window, we created vector embeddings for each DDSS condition (BAAI/bge-large-en-v1.5). For each case, the ICD-10 description and note text were used to query this index, and the top 100 most relevant conditions were inserted into the prompt as the candidate set. To ensure reliability, the model’s temperature was fixed at 0, and ten outputs were generated per case. Final labels and mappings were determined by majority vote. The process was fully automated in Python (OpenAI API). Condition Matching Two case sets were created: one from manual review and one from the LLM. These overlapped but were not identical, reflecting differences in eligibility and mappings. Each set was compared with condition suggestions from two DDSS versions. The original system, based on Ada’s probabilistic reasoning, was the version deployed during ESSENCE (Miller et al. 2020). The experimental version was a compound AI system under development utilising an LLM from the Llama 3 family. It was included in line with principles of post-market surveillance, where monitoring should establish whether newer versions provide measurable improvements. For fair comparison, it was restricted to the same user inputs without generating extra questions. Data Analysis Analyses were conducted in Python (v3.12) and Google Sheets. Diagnostic-context labelling was assessed by overall agreement and Cohen’s κ, with precision, recall, and F1 reported per label. Systematic differences between manual and LLM distributions were tested using χ². Eligibility for diagnostic-accuracy analysis was assessed as a binary outcome, with accuracy, sensitivity, specificity, precision, and F1 reported. For condition mapping, LLM-assigned diagnoses were compared with clinician mappings. Outputs were classified as exact match (LLM selected one of the clinician-assigned conditions), plausible additional (clinically valid but overlooked by clinicians), or mismatch (clinically inappropriate). Comparative analyses were performed for GPT-4.1 (gpt-4.1-2025-04-14) and GPT-5 (gpt-5-2025-08-07, reasoning effort “high”). Finally, the diagnostic accuracy of the DDSS was assessed by comparing reference diagnoses - derived from manual clinician review or GPT-5 mapping - with the ranked condition suggestions generated by the system. Accuracy was evaluated at the Top-1, Top-3, and Top-5 levels, defined as the proportion of cases in which the reference diagnosis appeared among the first one, three, or five system outputs. Diagnostic accuracy was reported with 95% CIs. Two-sample tests for proportions were used for unpaired comparisons, and McNemar’s test was applied for paired analyses. Comparisons between original and experimental DDSS versions were conducted within each reference set using paired methods. Cases in which the real condition was not modelled by the DDSS were still included in the diagnostic accuracy analysis. Missing Data Handling Encounters without ICD-10–coded diagnoses or consultation notes were excluded, as these data formed the reference standard and could not be imputed. Reasons for exclusion are listed in Supplementary Table 3. Ethics and Data Governance The ESSENCE study received ethical approval from the Comissão de Ética para a Investigação Clínica (2204JJ351, 2309JJ660) and was registered on ClinicalTrials.gov (NCT06846957) in 2025-02-21. All participants provided electronic informed consent. Data were anonymised and stored securely in Teamscope, and only de-identified records were used for LLM analysis. Procedures complied with the Declaration of Helsinki and ISO 14155:2020. Funding The Federal Ministry of Education and Research (Germany) funded the study but had no role in design, data collection, analysis, interpretation, or writing. Findings Participant characteristics and inclusions A total of 512 participants were included. The mean age was 39·2 years (SD 12·4; median 37 years [IQR 30-47]); 80·3% were aged 18-49 years. Females accounted for 57·6% (295/512). Females accounted for 57·6% (295/512). Participants reported a median of 2 presenting complaints (IQR 1–3; mean 2·21 [SD 1·43]), higher in females (2·40) than males (1·95; p=0·00036). The most common specialties were otorhinolaryngology/ENT (24·0%), orthopedics/trauma surgery (18·2%), gynecology (14·6%), and gastroenterology (13·5%). Details are provided in Table 1. Table 1 . Participant characteristics Characteristic Value Total N 512 Age, mean (SD) 39·2 (12·4) Median (IQR) 37 (30–47) Age range 18–79 Age distribution, n (% of 512) 18–29 118 (23·0%) 30–39 167 (32·6%) 40–49 126 (24·6%) 50–59 66 (12·9%) 60–69 25 (4·9%) 70–79 10 (2·0%) 80+ 0 (0%) Sex, n (% of 512) Female 295 (57·6%) Male 217 (42·4%) Presenting complaints (PCs) Mean (SD) overall 2·21 (1·43) Median (IQR) 2 (1–3) Female, mean 2·4 Male, mean 1·95 t-test p-value (F vs M) 0·00036 Specialty assigned (n, %, mean #PC, t-test p-value) Otorhinolaryngology / ENT 123 (24·0%), 2·54 Orthopedics / Trauma Surgery 93 (18·2%), 1·46 Gynecology 75 (14·6%), 1·91 Gastroenterology 69 (13·5%), 2·54 Pulmonology 40 (7·8%), 3·10 Dermatology 28 (5·5%), 1·71 Neurology 17 (3·3%), 2·47 Psychiatry 17 (3·3%), 2·88 Ophthalmology 14 (2·7%), 1·71 Urology 13 (2·5%), 1·69 Nephrology 8 (1·6%), 2·00 Cardiology 3 (0·6%), 5·00 Family Medicine 3 (0·6%), 1·67 Dentistry 2 (0·4%), 2·50 Rheumatology 2 (0·4%), 2·00 Angiology 1 (0·2%), 2·00 Endocrinology 1 (0·2%), 2·00 Internal Medicine 1 (0·2%), 6·00 Oncology 1 (0.2%), 1.00 Of 1,470 participants in the main cohort, confirmed healthcare-seeking behavior was available for 721 (49·0%). Within this group, 512 cases had both an ICD-10 code and consultation notes. Fourteen illustrative cases were set aside as examples for prompt development (Supplementary Table 4), leaving 498 cases for which labels were assigned by both manual review and the LLM (Figure 2). Anonymisation Outcomes All free-text notes were anonymised before LLM mapping. Of 512 notes, 235 (45·9%) required modification, most commonly conversion of exact ages to bands (n=186) and removal of dates (n=59), occupations (n=35), or locations (n=32). Names were removed in 18 cases. Over half of notes (283/512, 55·3%) required no changes (Supplementary Table 1). Filtering and Mapping Manual review classified 385/498 encounters (77·3%) as eligible and distributed the remaining 113 (22·7%) across exclusion categories: 46 (9·2%) “no diagnosis–symptom,” 12 (2·4%) “no diagnosis–unspecific code,” 25 (5·0%) “multimorbidity,” 21 (4·2%) “different complaint,” and 9 (1·8%) “not-covered condition” (Table 2). GPT-5 reproduced this overall distribution, also assigning 385/498 (77·3%) encounters as eligible. Among these, 347/385 (90·1%) overlapped with manual review. The model missed 38/385 (9·9%) manually eligible cases, while conversely assigning diagnoses in 38/113 (33·6%) cases that clinicians had excluded. At the binary level (eligible/ineligible), GPT-5 achieved 84·7% accuracy (95% CI 81·3–87·6), with sensitivity 90·1% (95% CI 86·7–92·7) and specificity 66·4% (95% CI 57·3–74·4). Agreement beyond chance was moderate (κ=0·57, 95% CI 0·50–0·64) (Table 3). Performance was robust for eligible cases (precision 0·90, recall 0·90, F1=0·90) but less consistent for exclusion categories, where unspecific codes were frequently over-assigned and multimorbidity under-recognised. Model comparison: GPT-4.1 vs GPT-5.0 Both models reproduced manual eligibility classifications with high overall accuracy (Table 3). GPT-4.1 achieved higher sensitivity (97·1% vs 90·1%) but much lower specificity (37·2% vs 66·4%). GPT-5·0 provided a more balanced profile with higher precision (90·1% vs 84·0%) and stronger agreement (κ=0·57 vs 0·42). Running this workflow required $5.08 for a single pass with GPT-4.1 and $44.40 for GPT-5. Condition mapping performance Among jointly eligible cases, GPT-5 exactly matched one of the clinician-assigned diagnoses in 324/347 (93·6%) and proposed an additional clinically plausible condition in 12 (3·5%). Disagreement occurred in 10 cases (2·9%), and no output in one (0·3%). GPT-4.1 showed a similar profile: 354/374 (94·7%) exact matches, 8 (2·1%) plausible additions, and 12 (3·2%) mismatches. When exact matches and plausible additions were combined, overall alignment was 97·1% for GPT-5 and 96·8% for GPT-4.1. Full distributions are shown in Table 3. Consistency To assess the internal consistency of GPT-5’s label assignment, we generated eight independent completions per case and quantified the distribution of distinct labels. In 385/498 cases (77.3%), all completions converged on the same label. Among the remaining 113 cases, two distinct labels were suggested in 63/498 (12.7%), three in 38/498 (7.6%), and four in 12/498 (2.4%). Even when multiple labels appeared, the majority label typically dominated: for two-label cases, 70% (44/63) showed a 6–2 or 7–1 split. In supplementary table 5, we present a confusion matrix of LLM vs. manual labels, with some examples provided in supplementary table 6. Table 2. Diagnostic-context label distribution and per-label performance Label Manual n (%) GPT-5 n (%) Both Missed Precision Recall F1 Mappable 385 (77·3) 385 (77·3) 347 38 0·9 0·9 0·9 No Dx – Symptom 46 (9·2) 41 (8·2) 30 16 0·73 0·65 0·69 No Dx – Unspecific code 12 (2·4) 27 (5·4) 5 7 0·19 0·42 0·26 Multimorbidity 25 (5·0) 11 (2·2) 10 15 0·91 0·4 0·56 Different complaint 21 (4·2) 22 (4·4) 11 10 0·5 0·52 0·51 Not-covered condition 9 (1·8) 12 (2·4) 6 3 0·5 0·67 0·57 Total 498 (100) 498 (100) — — — — — Table 3. Performance of GPT-4.1 and GPT-5.0 for eligibility classification and condition mapping (n=498) GPT-4.1 GPT-5.0 Eligibility classification True Positives (TP) 374 347 False Negatives (FN) 11 38 False Positives (FP) 71 38 True Negatives (TN) 42 75 Manual eligible cases 385 385 Model eligible cases 445 385 Accuracy (TP+TN / N) 83·5% 84·7% Sensitivity (recall for eligible) 97·1% 90·1% Specificity (recall for ineligible) 37·2% 66·4% Precision (PPV for eligible) 84·0% 90·1% F1 score (for eligible) 90·1% 90·1% Cohen’s κ (95% CI) 0·42 (0·33–0·52) 0·57 (0·50–0·64) Condition mapping Exact match 354/374 (94·7%) 324/347 (93·6%) Should be considered (plausible addition) 8/374 (2·1%) 12/347 (3·5%) No overlap (mismatch) 12/374 (3·2%) 10/347 (2·9%) No output – 1/347 (0·3%) Matching comparison For the diagnostic accuracy analyses, we included all cases, including those where the DDSS did not explicitly model the underlying condition, yielding 394 manual and 397 GPT-5–mapped cases (Table 4). With manual mapping, Top-1 accuracy was 47·2% (42·3–52·1) for the original and 48·0% (47·5–52·9) for the experimental DDSS version; Top-3 accuracy was 71·6% (66·5–76·3) versus 67·5% (62·4–72·2); and Top-5 accuracy was 78·4% (73·2–83·1) versus 81·2% (76·0–85·8). With GPT-5 mapping, Top-1 accuracy was 43·1% (38·3–48·0) versus 44·1% (39·2–49·0); Top-3 accuracy was 65·2% (60·3–69·8) versus 67·0% (62·2–71·5); and Top-5 accuracy was 71·3% (66·6–75·6) versus 75·1% (70·6–79·1). Across the 385 analyzable cases per set, manual and GPT-5 mapping performed similarly (supplementary table 7). Differences ranged from –4 to +7 percentage points, with only one significant finding: manual mapping achieved higher Top-5 accuracy under the original DDSS version (309/385 [80·3%] vs 283/385 [73·5%]; Δ=+6·8 pp, 95% CI +0·8 to +12·7; p=0·026). After GPT5 did not provide an output in one case, 346 cases overlapped. For these, results were almost identical (supplementary table 8). McNemar tests showed no significant differences at any threshold, with small differences (–5 to +2 pp) and confidence intervals crossing zero, indicating that GPT-5 mapping was statistically equivalent to manual mapping when applied to the same cases. Paired comparisons of the original and experimental DDSS versions (supplementary table 9) showed no significant differences at the Top-1 or Top-3 thresholds. At Top-5, however, the experimental version was consistently more accurate: manual 309/396 (78·0%) vs 320/396 (80·8%; Δ=+3·1 pp, 95% CI +0·4 to +5·8; p=0·038) and GPT-5 283/397 (71·3%) vs 298/397 (75·1%; Δ=+3·9 pp, 95% CI +1·2 to +6·6; p=0·009). Table 4 . Diagnostic accuracy of DDSS using manual and GPT-5 reference mapping. Reference set N Version Top-1 % (95% CI) Top-3 % (95% CI) Top-5 % (95% CI) Manual- mapped 394 Original 47·2 % (42·3–52·1) 71·6 % (66·5–76·3) 78·4 % (73·2–83·1) 394 Experimental 48·0% (47·5–52·9) 67·5 % (62·4–72·2) 81·2 % (76·0–85·8) GPT-5- mapped 397 Original 43·1 % (38·3–48·0) 65·2 % (60·3–69·8) 71·3 % (66·6–75·6) 397 Experimental 44·1 % (39·2–49·0) 67·0 % (62·2–71·5) 75·1 % (70·6–79·1) Manual- mapped 346 Original 46·8 % (41·6–52·1) 72·8 % (67·9–77·2) 80·3 % (75·8–84·2) 346 Experimental 48·3 % (43·1–53·5) 68·8 % (63·7–73·4) 83·5 % (79·3–87·1) GPT-5- mapped 346 Original 48·0 % (42·8–53·2) 71·4 % (66·4–75·9) 78·0 % (73·4–82·1) 346 Experimental 49·1 % (43·9–54·4) 74·0 % (69·1–78·3) 81·2 % (76·8–85·0) Interpretation Summary of results We evaluated whether large language models (LLMs) can automate real-world monitoring of diagnostic accuracy in digital diagnostic decision support systems (DDSS) using a filter–map–match framework. Of 498 cases, both manual and LLM review classified 385 as eligible for diagnostic accuracy analysis. Within this subset, agreement was high: both approaches classified the same 347 cases (90·1%) as eligible. The LLM additionally marked 38 cases as eligible that clinicians had excluded, while excluding 38 that clinicians had included. GPT-5 mapped the same clinician diagnoses in 93·6% of overlapping cases and suggested plausible alternatives in 3·5%. Diagnostic accuracy was statistically indistinguishable between manual and GPT-5 mapping at Top-1 and Top-3, with one difference at Top-5. In overlapping cases, no significant differences remained, confirming equivalence when eligibility filtering is held constant. Across reference sets, both approaches identified that the experimental DDSS was more accurate than the original only at the Top-5 threshold. Filtering–mapping–matching Comparing DDSS outputs with real-world outcomes requires identifying truly comparable encounters. In practice, the digital assessment and clinic visit may cover different complaints, diagnoses may still be under investigation, or multimorbidity can complicate documentation. In our cohort, only 77·3% of encounters were suitable for diagnostic comparison. Analysing all cases together would conflate incomparable scenarios and distort estimates. A structured filter–map–match workflow - (i) filtering to comparable encounters, (ii) mapping confirmed diagnoses to a shared ontology, and (iii) matching mapped diagnoses to DDSS outputs - offers a more rigorous basis for monitoring. For filtering, the model achieved high sensitivity and precision but only moderate specificity. This asymmetry is acceptable in surveillance: retaining most eligible cases preserves sample size, while extra inclusions can be resolved downstream. Similar frameworks are being explored for case identification in electronic records. Cheligeer et al, for example, applied a multi-stage method to detect hospital-acquired pulmonary embolism from 10,066 inpatients, of whom only 40 (0·4%) had true events. Models achieved sensitivity 87·5–100% and specificity 94·9–98·9%, but positive predictive values remained low (≈7–17%) because of rarity, with F1 scores peaking at 28·1%(22). Their method efficiently ruled out negatives at scale, but most flagged cases still needed manual review. One key objective of this research was to identify whether automating the entire filter-map-match workflow with AI would distort estimates of diagnostic accuracy. When comparing the full sets of analyzable cases (385 each), GPT-5 mapping produced results that were statistically indistinguishable from manual review at the Top-1 and Top-3 thresholds, with a single significant difference observed at Top-5. When restricting the analysis to the 347 overlapping cases, where the task was limited to mapping rather than eligibility filtering, no statistical differences were observed at any threshold. This indicates that the mapping step can already be automated with high reliability. The only residual divergence arose from imperfect eligibility filtering, which accounted for the single significant difference in one of the three comparisons. Importantly, both manual and GPT-5 mapping consistently identified that the experimental DDSS version was significantly more accurate at the Top-5 threshold only. Methodological considerations ICD-10 codes were initially considered as the main comparator, but have well-described limitations: agreement is only moderate at chapter level and poor for specific codes, with high inter-rater variability(23–25). Coding complexity and incomplete documentation often yield generic labels(26,27). We therefore combined ICD-10 review with free-text note analysis, which clarified unspecific codes in several cases. LLMs now make it feasible to screen documentation at scale, with high sensitivity and specificity for phenotyping but only moderate positive predictive value and limited causal attribution, requiring clinician oversight(28). Strengths This study has several strengths. It was conducted in a real-world clinical environment within Portugal’s largest private healthcare network, capturing the heterogeneity of routine practice beyond curated test sets. The inclusion of multiple specialties increases the generalisability of findings. A clinician-adjudicated reference standard ensured reproducibility of comparisons. Finally, the filter–map–match framework was explicitly aligned with clinical reality, and the privacy-preserving pipeline systematically anonymised notes, ensuring GDPR compliance. Finally, we also compared downstream diagnostic accuracy across both manual and GPT-5 mapping approaches and between original and experimental DDSS versions. This allowed us to test how the approaches would behave under real-world performance monitoring conditions. Limitations This study has several limitations. Manual review is an imperfect reference, and in some instances the LLM proposed plausible alternatives not documented by clinicians, highlighting the constraints of the comparator. Case review was conducted by clinicians employed by the DDSS and familiar with its ontology, which may have introduced bias. The analysis was a proof-of-concept, one-off evaluation rather than an implemented surveillance system, and cases without a final diagnosis at data collection were excluded, though some might have become eligible later. Performance was also assessed within a single healthcare network, so results may differ in other settings with different documentation practices or diagnostic distributions. Embedding the workflow in live systems will be essential to establish longitudinal performance, feasibility, and clinical impact. Implementation in Clinical Practice Several lessons emerge for real-world use. First, the core decision is binary- whether a case is eligible for diagnostic comparison. Although a multi-label framework was useful for benchmarking, rare categories performed inconsistently and added little operational value. Prioritising this binary decision offers a more pragmatic foundation for implementation, particularly as the mapping step showed no statistically significant differences between manual and GPT-5 review. Second, costs and cascade design matter. On our dataset, end-to-end inference cost US$44·40 with GPT-5 (≈US$0·09 per case) and US$5·08 with GPT-4·1 (≈US$0·01 per case). Given that divergence arose chiefly in eligibility filtering rather than mapping, a cost-sensitive routing strategy is feasible: use a smaller model for high-throughput filtering, escalate low-confidence or high-risk cases to GPT-5, and audit a sample with clinicians. This balances accuracy, scale, and spend, and can be tuned by confidence thresholds and sampling rates. Third, diagnosis is often a longitudinal process rather than a single clinical event. Patients may await test results or referrals before a final diagnosis is reached. Systems should therefore re-analyse such cases once definitive information becomes available, converting encounters initially excluded into eligible cases for analysis. Fourth, in-workflow one-click feedback controls offer a practical way to capture clinician judgement at the point of care - for example, by accepting a suggested condition or selecting a better match from a type-ahead ontology. Structured reporting in radiology has shown that even a single click can return machine-readable feedback for continual learning (Fuchs 2024). Although clinicians are unlikely to provide feedback for every consultation, such mechanisms are a scalable and efficient addition, delivering targeted corrections without disrupting workflow. Finally, privacy must be preserved. Advances in locally deployable anonymisation models now achieve near-perfect sensitivity and character-level accuracy for de-identifying clinical text without data leaving the system(29). Integrating such tools could replace manual redaction and enable GDPR-compliant note review at scale. Conclusion This study shows that large language models can automate key steps in monitoring the diagnostic accuracy of decision support systems. GPT-5 produced results that were highly consistent with clinician review, with the few differences explained by eligibility filtering - deciding which real-world cases could be included in the diagnostic accuracy analysis- rather than by the mapping step itself. Both manual and LLM-based approaches led to the same conclusions about the performance improvements from the original to the experimental version of the DDSS, indicating that LLM-based monitoring can be trusted to detect meaningful changes in system performance without continuous manual review. While these findings establish feasibility, further work is required to embed the workflow in live systems and to extend it to automated continuous monitoring. Focusing on improving eligibility filtering while including re-evaluation of provisional cases, one-click feedback, and local anonymisation will be critical to test scalability and clinical impact. If operationalised, such pipelines could enable continuous, learning health system monitoring of diagnostic AI. Research in Context: Evidence before this study Monitoring the diagnostic accuracy of digital decision support systems (DDSS) has traditionally relied on manual clinician review. Although accurate, this approach is labour-intensive and difficult to scale. Previous work shows that large language models (LLMs) can extract information from clinical notes and map diagnoses to standard ontologies, but their ability to fully replace clinical review in continuous surveillance of DDSS performance had not been established. Added value of this study This study tested whether an LLM could automate the filter-map-match workflow used to monitor DDSS accuracy in real-world practice. GPT-5 achieved high agreement with manual review, reproducing clinician-adjudicated diagnoses in most cases and generating comparable diagnostic accuracy estimates. Importantly, both manual and LLM-based workflows led to the same conclusions about relative performance of the original versus experimental DDSS versions. Differences between methods arose mainly from eligibility filtering rather than mapping itself. Implications of all the available evidence These findings show that LLMs can reliably replicate the core elements of clinician-led surveillance of DDSS diagnostic accuracy. The mapping step appears nearly ready for practice, whereas eligibility filtering requires further refinement to ensure consistency. If implemented, such an approach could substantially reduce reliance on manual practitioner panels and enable scalable, near real-time monitoring of DDSS in routine care. Declarations Acknowledgments This work was supported by the Federal Ministry of Education and Research (Bundesministerium für Bildung und Forschung) through the European Union-financed NextGenerationEU program under grant number 16KISA100K, project PATH—“Personal Mastery of Health and Wellness Data.” Author Contributions FC and SG were responsible for conceptualisation and methodology. Data collection and investigation were carried out by FDL, MP, MSM, PF, NK, MS, TM, and FC. Data anonymisation was performed by MS. NK, OS, FC, and KG conducted the manual filtering and mapping of cases to establish the clinician-adjudicated reference set. PB, FC, and OS performed formal analysis and visualisation. FC led writing the original draft in collaboration with all co-authors. VHa, AM, LS, VHe, OS, SK, VM, HH, and PE led the development of the experimental version of the app tested in the study and contributed to review of the manuscript. All authors had full access to aggregated data in the study and contributed to the decision to submit the manuscript for publication. S.G. is a News and Views Editor for npj Digital Medicine. S.G. played no role in the internal review or decision to publish this Original Research article. Corresponding author Correspondence to Stephen Gilbert [email protected] Pedro Flores Competing Interests PB, OS, VHa, AM, VHe, SK, HH, PE, KG and TM are employed by Ada Health. MS, NK, LS and VM are former employees of Ada Health. SG and FC are consultants for Ada Health. FC, PB, VHa, LS, VHe, SK, VM, HH, PE, SG, KG and TM hold share options in Ada Health. SG declares a nonfinancial interest as an Advisory Group member of the EY-coordinated “Study on Regulatory Governance and Innovation in the field of Medical Devices” conducted on behalf of the Directorate-General for Health and Food Safety (SANTE) of the European Commission. SG declares the following competing financial interests: he has or has had consulting relationships with Una Health GmbH, Lindus Health Ltd., Flo Ltd, ICURA ApS, Rock Health Inc., Thymia Ltd., FORUM Institut für Management GmbH, High-Tech Gründerfonds Management GmbH, Prova Health Ltd, Directorate-General for Research and Innovation Of the European Commission. TM declares the following competing financial interests: he has or has had consulting relationships with Suvera & iPlato Healthcare. FC declares the following competing financial interests: she has or has had consulting relationships with Flo Ltd. Furthermore, VHa, AM, LS, SK, VM, HH, and PE are co-inventors on a patent application titled "Computer system and method for supporting medical diagnosis of a human with symptoms associated with the human's medical condition" (Application No. EP24188258), with rights held by the co-inventors' employer Ada Health. This patent application relates to the compound AI system described in this paper. Data Sharing Data collected for the study, including individual participant data and a data dictionary defining each field in the set, will be made available to others after publication upon reasonable request, subject to approval. Requests for access should be made to the study team at Ada Health ( [email protected] ). After approval, a signed data sharing agreement will be required before data release. References Mohseni A, Ghotbi E, Kazemi F, Shababi A, Jahan SC, Mohseni A, et al. Artificial Intelligence in Radiology. Radiol Clin North Am. 2024;62(6):935–47. Lu MY, Chen B, Williamson DFK, Chen RJ, Zhao M, Chow AK, et al. A multimodal generative AI copilot for human pathology. Nature. 2024;634(8033):466–73. Fernandes M, Vieira SM, Leite F, Palos C, Finkelstein S, Sousa JMC. Clinical Decision Support Systems for Triage in the Emergency Department using Intelligent Systems: a Review. Artif Intell Med. 2020;102:101762. Korom R, Kiptinness S, Adan N, Said K, Ithuli C, Rotich O, et al. AI-based Clinical Decision Support for Primary Care: A Real-World Study [Internet]. arXiv; 2025 [cited 2025 July 28]. Available from: https://arxiv.org/abs/2507.16947 Pennisi F, Pinto A, Ricciardi GE, Signorelli C, Gianfredi V. Artificial intelligence in antimicrobial stewardship: a systematic review and meta-analysis of predictive performance and diagnostic accuracy. Eur J Clin Microbiol Infect Dis. 2025;44(3):463–513. Faiyazuddin Md, Rahman SJQ, Anand G, Siddiqui RK, Mehta R, Khatib MN, et al. The Impact of Artificial Intelligence on Healthcare: A Comprehensive Review of Advancements in Diagnostics, Treatment, and Operational Efficiency. Health Sci Rep. 2025;8(1):e70312. Gilbert S, Pimenta A, Stratton-Powell A, Welzel C, Melvin T. Continuous Improvement of Digital Health Applications Linked to Real-World Performance Monitoring: Safe Moving Targets? Mayo Clin Proc Digit Health. 2023 Sept;1(3):276–87. Fraser AG, Butchart EG, Szymański P, Caiani EG, Crosby S, Kearney P, et al. The need for transparency of clinical evidence for medical devices in Europe. The Lancet. 2018;392(10146):521–30. Han R, Acosta JN, Shakeri Z, Ioannidis JPA, Topol EJ, Rajpurkar P. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit Health. 2024;6(5):e367–73. Gilbert S, Mathias R, Schönfelder A, Wekenborg M, Steinigen-Fuchs J, Dillenseger A, et al. A roadmap for safe, regulation-compliant Living Labs for AI and digital health development. Sci Adv. 2025;11(20):eadv7719. Park SH, Han K, Jang HY, Park JE, Lee JG, Kim DW, et al. Methods for Clinical Evaluation of Artificial Intelligence Algorithms for Medical Diagnosis. Radiology. 2023;306(1):20–31. McKinney SM, Sieniek M, Godbole V, Godwin J, Antropova N, Ashrafian H, et al. International evaluation of an AI system for breast cancer screening. Nature. 2020;577(7788):89–94. Creagh AP, Pease T, Ashworth P, Bradley L, Duport S. Explainable machine learning to identify patients at risk of developing hospital acquired infections [Internet]. Health Informatics; 2024 [cited 2025 Sept 5]. Available from: http://medrxiv.org/lookup/doi/ 10.1101/2024.11.13.24317108 Asgari E, Montaña-Brown N, Dubois M, Khalil S, Balloch J, Yeung JA, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. Npj Digit Med. 2025;8(1):274. El Arab RA, Abu-Mahfouz MS, Abuadas FH, Alzghoul H, Almari M, Ghannam A, et al. Bridging the Gap: From AI Success in Clinical Trials to Real-World Healthcare Implementation—A Narrative Review. Healthcare. 2025;13(7):701. Semigran HL, Linder JA, Gidengil C, Mehrotra A. Evaluation of symptom checkers for self diagnosis and triage: audit study. BMJ. 2015 July 8;h3480. Andersen ES, Birk-Korch JB, Hansen RS, Fly LH, Röttger R, Arcani DMC, et al. Monitoring performance of clinical artificial intelligence in health care: a scoping review. JBI Evid Synth. 2024;22(12):2423–46. Fuchs M, Gonzalez C, Frisch Y, Hahn P, Matthies P, Gruening M, et al. Closing the loop for AI-ready radiology. RöFo - Fortschritte Auf Dem Geb Röntgenstrahlen Bildgeb Verfahr. 2024;196(02):154–62. Agatstein K. Chart Review Is Dead; Long Live Chart Review: How Artificial Intelligence Will Make Human Review of Medical Records Obsolete, One Day. Popul Health Manag. 2023;26(6):438–40. Zitu MM, Owen D, Manne A, Wei P, Li L. Large Language Models for Adverse Drug Events: A Clinical Perspective. J Clin Med. 2025;14(15):5490. Cook C. Challenges with diagnoses: sketchy reference standards. J Man Manip Ther. 2012;20(3):111–2. Cheligeer C, Southern DA, Yan J, Wu G, Pan J, Lee S, et al. Utilizing large language models for detecting hospital-acquired conditions: an empirical study on pulmonary embolism. J Am Med Inform Assoc. 2025;32(5):876–84. Wockenfuss R, Frese T, Herrmann K, Claussnitzer M, Sandholzer H. Three- and four-digit ICD-10 is not a reliable classification system in primary care. Scand J Prim Health Care. 2009;27(3):131–6. Stausberg J, Lehmann N, Kaczmarek D, Stein M. Reliability of diagnoses coding with ICD-10. Int J Med Inf. 2008;77(1):50–7. Osterhage KP, Hser Y, Mooney LJ, Sherman S, Saxon AJ, Ledgerwood M, et al. Identifying patients with opioid use disorder using International Classification of Diseases (ICD) codes: Challenges and opportunities. Addiction. 2024;119(1):160–8. Almagro M, Unanue RM, Fresno V, Montalvo S. ICD-10 Coding of Spanish Electronic Discharge Summaries: An Extreme Classification Problem. IEEE Access. 2020;8:100073–83. Zhou L, Cheng C, Ou D, Huang H. Construction of a semi-automatic ICD-10 coding system. BMC Med Inform Decis Mak. 2020;20(1):67. Bejan CA, Wang M, Venkateswaran S, Bergmann EA, Hiles L, Xu Y, et al. irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets [Internet]. Health Informatics; 2025 [cited 2025 Sept 3]. Available from: http://medrxiv.org/lookup/doi/ 10.1101/2025.03.05.25323445 Wiest IC, Leßmann ME, Wolf F, Ferber D, Treeck MV, Zhu J, et al. Deidentifying Medical Documents with Local, Privacy-Preserving Large Language Models: The LLM-Anonymizer. NEJM AI [Internet]. 2025 Mar 27 [cited 2025 Sept 3];2(4). Available from: https://ai.nejm .org/doi/10.1056/AIdbp2400537 Additional Declarations Competing interest reported. PB, OS, VHa, AM, VHe, SK, HH, PE, KG and TM are employed by Ada Health. MS, NK, LS and VM are former employees of Ada Health. SG and FC are consultants for Ada Health. FC, PB, VHa, LS, VHe, SK, VM, HH, PE, SG, KG and TM hold share options in Ada Health. SG declares a nonfinancial interest as an Advisory Group member of the EY-coordinated “Study on Regulatory Governance and Innovation in the field of Medical Devices” conducted on behalf of the Directorate-General for Health and Food Safety (SANTE) of the European Commission. SG declares the following competing financial interests: he has or has had consulting relationships with Una Health GmbH, Lindus Health Ltd., Flo Ltd, ICURA ApS, Rock Health Inc., Thymia Ltd., FORUM Institut für Management GmbH, High-Tech Gründerfonds Management GmbH, Prova Health Ltd, Directorate-General for Research and Innovation Of the European Commission. TM declares the following competing financial interests: he has or has had consulting relationships with Suvera & iPlato Healthcare. FC declares the following competing financial interests: she has or has had consulting relationships with Flo Ltd. Furthermore, VHa, AM, LS, SK, VM, HH, and PE are co-inventors on a patent application titled "Computer system and method for supporting medical diagnosis of a human with symptoms associated with the human's medical condition" (Application No. EP24188258), with rights held by the co-inventors' employer Ada Health. This patent application relates to the compound AI system described in this paper. Supplementary Files PATHLLMPaperAppendixV1.01withetechnicalcorrections.docx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 07 Jan, 2026 Reviews received at journal 02 Jan, 2026 Reviews received at journal 20 Dec, 2025 Reviewers agreed at journal 11 Dec, 2025 Reviewers agreed at journal 21 Nov, 2025 Reviewers agreed at journal 16 Nov, 2025 Reviewers invited by journal 10 Nov, 2025 Editor assigned by journal 10 Nov, 2025 Submission checks completed at journal 09 Nov, 2025 First submitted to journal 03 Nov, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8022874","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":545787634,"identity":"918cfcf7-76b3-46fe-82bb-8d217d837fe5","order_by":0,"name":"Fabienne Cotte","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Fabienne","middleName":"","lastName":"Cotte","suffix":""},{"id":545787635,"identity":"197d1142-0f79-4e6e-92da-c8911ddd76aa","order_by":1,"name":"Marcel Schmude","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Marcel","middleName":"","lastName":"Schmude","suffix":""},{"id":545787636,"identity":"47e89df4-1e5a-4071-b912-2299f061dd74","order_by":2,"name":"Philipp Bode","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Philipp","middleName":"","lastName":"Bode","suffix":""},{"id":545787637,"identity":"cfd619bb-48ea-4896-9220-f90602b82d51","order_by":3,"name":"Oula Suliman","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Oula","middleName":"","lastName":"Suliman","suffix":""},{"id":545787638,"identity":"0beb78cb-8ea5-4ea6-b193-4ba173e1b00c","order_by":4,"name":"Filipa Dias Lourenço","email":"","orcid":"","institution":"CUF","correspondingAuthor":false,"prefix":"","firstName":"Filipa","middleName":"Dias","lastName":"Lourenço","suffix":""},{"id":545787639,"identity":"d4676f6f-2a00-4be4-a7af-244a1ec7e6d0","order_by":5,"name":"Miguel Paiva Pereira","email":"","orcid":"","institution":"CUF","correspondingAuthor":false,"prefix":"","firstName":"Miguel","middleName":"Paiva","lastName":"Pereira","suffix":""},{"id":545787640,"identity":"88cff02a-7983-44f3-9c55-5c4d430b3263","order_by":6,"name":"Nisha Kini","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Nisha","middleName":"","lastName":"Kini","suffix":""},{"id":545787641,"identity":"f722aff2-1649-41c5-a12b-505f5a090f9c","order_by":7,"name":"Vera Hartenstein","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Vera","middleName":"","lastName":"Hartenstein","suffix":""},{"id":545787642,"identity":"12c50a44-4062-44c4-aee2-d5cb5cb1cc33","order_by":8,"name":"Alesandro Muscoloni","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Alesandro","middleName":"","lastName":"Muscoloni","suffix":""},{"id":545787643,"identity":"6c1e75d8-a9d2-470c-85b4-f05a02bb091f","order_by":9,"name":"Lisa Stroux","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Lisa","middleName":"","lastName":"Stroux","suffix":""},{"id":545787644,"identity":"a2c93725-cd36-4f7f-92c0-5ab4873188e9","order_by":10,"name":"Victor Hertz","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Victor","middleName":"","lastName":"Hertz","suffix":""},{"id":545787645,"identity":"35f30aa7-5596-44a0-b327-99e791fccd91","order_by":11,"name":"Sebastian Köhler","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Sebastian","middleName":"","lastName":"Köhler","suffix":""},{"id":545787646,"identity":"2ab72970-840a-4ba4-ba5f-72c8c92b0334","order_by":12,"name":"Valerio Morelli","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Valerio","middleName":"","lastName":"Morelli","suffix":""},{"id":545787647,"identity":"3e9ffac0-77b6-4bd7-8472-01dd0420ed04","order_by":13,"name":"Henry Hoffmann","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Henry","middleName":"","lastName":"Hoffmann","suffix":""},{"id":545787648,"identity":"929f07ca-b87a-4b3a-96ea-645b491e395d","order_by":14,"name":"Peter Engerer","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Peter","middleName":"","lastName":"Engerer","suffix":""},{"id":545787649,"identity":"86fdbfc4-c192-4e65-ba5d-07d50924e972","order_by":15,"name":"Stephen Gilbert","email":"","orcid":"","institution":"Else Kröner-Fresenius Center for Digital Health, TUD Dresden University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Stephen","middleName":"","lastName":"Gilbert","suffix":""},{"id":545787650,"identity":"26327f26-9722-4667-91b1-88187bfd9f85","order_by":16,"name":"Kirsten Gray","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Kirsten","middleName":"","lastName":"Gray","suffix":""},{"id":545787651,"identity":"d63047ff-b2ca-4db0-bab3-388791ded36e","order_by":17,"name":"Tauseef Mehrali","email":"","orcid":"","institution":"Ada Health GmbH","correspondingAuthor":false,"prefix":"","firstName":"Tauseef","middleName":"","lastName":"Mehrali","suffix":""},{"id":545787652,"identity":"7fc6f5f9-35d1-4583-8007-ae5a657569fe","order_by":18,"name":"Micaela Seemann Monteiro","email":"","orcid":"","institution":"CUF","correspondingAuthor":false,"prefix":"","firstName":"Micaela","middleName":"Seemann","lastName":"Monteiro","suffix":""},{"id":545787653,"identity":"962e0ed6-9a49-4d22-8891-31358075376c","order_by":19,"name":"Pedro Flores","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6ElEQVRIiWNgGAWjYHACAwjFzNzAwFDBwMDeQLwWRqDaMwwMPAcYGAnpgmoBKWRsI0KLOXvztg8fGA4zyLcztkn8nGeX2MPAe/wBPi2WPceKZ84AajE4zNgm2bstGaiFLxGvLQY3coyZeRjSGAyYGZsNeLcdSNzPwGOIX8v9NxAt8s2MzYZ/5xwA2kJIyw0ekBYbBobDjI2PeRuI0GLZk1bMOMPAhscApEXmWLJxDzNf4gx8WszZD29m+FAhISfff/jAwTc1drI97L0HPuB1GJTkQQgx8+BQi6IFFRDQMgpGwSgYBSMOAAC+BkKYkhbX9AAAAABJRU5ErkJggg==","orcid":"","institution":"CUF","correspondingAuthor":true,"prefix":"","firstName":"Pedro","middleName":"","lastName":"Flores","suffix":""}],"badges":[],"createdAt":"2025-11-03 23:08:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8022874/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8022874/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96454417,"identity":"470fdea8-9178-4384-98e1-df128583f1ca","added_by":"auto","created_at":"2025-11-21 10:02:43","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3428248,"visible":true,"origin":"","legend":"","description":"","filename":"PATHLLMPaperV1.0November3finalNature1withtechnicalcorrections.docx","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/c9f9defbc8bc13b3e23d85c8.docx"},{"id":96453569,"identity":"67c18cfd-8b1d-4b2c-a170-4b8be613db99","added_by":"auto","created_at":"2025-11-21 10:00:49","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":21281,"visible":true,"origin":"","legend":"","description":"","filename":"a7345d8cde9348d193c852204b6e5d89.json","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/8adb18c3cf30a29052205b0b.json"},{"id":96400037,"identity":"d46e30d6-0386-4e1b-a85e-1f60586ccfe5","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2048316,"visible":true,"origin":"","legend":"","description":"","filename":"PATHLLMPaperAppendixV1.01withetechnicalcorrections.docx","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/358c22f141140cc925c0c396.docx"},{"id":96400032,"identity":"66a31fc8-1082-427f-ba8a-7454758e977b","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":117547,"visible":true,"origin":"","legend":"","description":"","filename":"a7345d8cde9348d193c852204b6e5d891enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/c818f9aede2c98fc4e34b109.xml"},{"id":96400031,"identity":"4262efd7-a411-4d28-aa8f-2c0387d9400d","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":155024,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/41fd37a2cb5b8d2cd0395cf1.png"},{"id":96400041,"identity":"8570b2dd-bcc2-41f4-ba2f-267f2136a4f4","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":186155,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/327bb645035b70e9e0ade0da.png"},{"id":96400040,"identity":"2d58e4ee-344a-4843-b539-fd1853feb0b8","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"xml","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":119263,"visible":true,"origin":"","legend":"","description":"","filename":"a7345d8cde9348d193c852204b6e5d891structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/17b34104ae4fd55e8a0bd5a7.xml"},{"id":96453434,"identity":"5bca9328-f58c-4810-8a32-b6c5a2ca1dbe","added_by":"auto","created_at":"2025-11-21 09:59:50","extension":"html","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":129101,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/b68f832044d9279dd8fbb4bb.html"},{"id":96400030,"identity":"4cbc8627-30a4-41fb-ae74-b0f13c5ff547","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":759404,"visible":true,"origin":"","legend":"\u003cp\u003eData collection procedures\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/538e65c2db1acee18b97dc61.png"},{"id":96453335,"identity":"ba35f65a-92a7-4e1f-b2e4-fa9f5b06cc51","added_by":"auto","created_at":"2025-11-21 09:59:19","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":784826,"visible":true,"origin":"","legend":"\u003cp\u003eSTROBE Inclusion flowchart\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/25de7cd291ab007fd2fba8b9.jpeg"},{"id":96602834,"identity":"8e5eeebb-09ae-418a-941f-ef34bf477a0a","added_by":"auto","created_at":"2025-11-24 09:02:54","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2481901,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/81d65df7-52a3-48f3-addd-50e739ff1c2c.pdf"},{"id":96400038,"identity":"0e8486da-6a55-40ed-8d8a-180da0001bfb","added_by":"auto","created_at":"2025-11-20 16:03:14","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":2048316,"visible":true,"origin":"","legend":"","description":"","filename":"PATHLLMPaperAppendixV1.01withetechnicalcorrections.docx","url":"https://assets-eu.researchsquare.com/files/rs-8022874/v1/a345c1927bdf6c38362643d3.docx"}],"financialInterests":"Competing interest reported. PB, OS, VHa, AM, VHe, SK, HH, PE, KG and TM are employed by Ada Health. MS, NK, LS and VM are former employees of Ada Health. SG and FC are consultants for Ada Health. FC, PB, VHa, LS, VHe, SK, VM, HH, PE, SG, KG and TM hold share options in Ada Health. SG declares a nonfinancial interest as an Advisory Group member of the EY-coordinated “Study on Regulatory Governance and Innovation in the field of Medical Devices” conducted on behalf of the Directorate-General for Health and Food Safety (SANTE) of the European Commission. SG declares the following competing financial interests: he has or has had consulting relationships with Una Health GmbH, Lindus Health Ltd., Flo Ltd, ICURA ApS, Rock Health Inc., Thymia Ltd., FORUM Institut für Management GmbH, High-Tech Gründerfonds Management GmbH, Prova Health Ltd, Directorate-General for Research and Innovation Of the European Commission. TM declares the following competing financial interests: he has or has had consulting relationships with Suvera \u0026 iPlato Healthcare. FC declares the following competing financial interests: she has or has had consulting relationships with Flo Ltd. Furthermore, VHa, AM, LS, SK, VM, HH, and PE are co-inventors on a patent application titled \"Computer system and method for supporting medical diagnosis of a human with symptoms associated with the human's medical condition\" (Application No. EP24188258), with rights held by the co-inventors' employer Ada Health. This patent application relates to the compound AI system described in this paper.","formattedTitle":"Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study","fulltext":[{"header":"Background","content":"\u003cp\u003eArtificial intelligence (AI) has moved from proof-of-concept studies to approved tools in routine care across nearly all domains of medicine. AI now interprets radiology, guides ultrasound, supports pathology, triages symptoms, predicts deterioration, and optimises operations(\u003cspan additionalcitationids=\"CR2 CR3 CR4\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). Every stage of care delivery is being re-examined through the lens of AI(\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eRobust evidence is needed to validate benefits, ensure safe use, and meet regulatory requirements for demonstrations of performance and clinical value, including methods in the postmarket phase, i.e., post-market surveillance and post-market studies (known as post-market clinical follow-up)(\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). Yet most evaluations rely on curated test sets, reader studies, or single-centre trials(\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). These establish feasibility but rarely capture the heterogeneity of practice, where performance can deteriorate once exposed to diverse patients and workflows(\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Silent-mode trials help bridge this gap by embedding AI in clinical settings without influencing care, exposing biases and operational challenges before full deployment. In Kenya, an LLM tool for error detection was run in shadow mode to uncover workflow issues in primary care(\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). In the UK, a machine-learning framework for predicting hospital-acquired infections in patients with neurological impairments underwent a year-long silent-mode evaluation, demonstrating superior risk stratification compared with early warning scores(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e).\u003c/p\u003e\u003cp\u003ePre-deployment quality control is also essential. Asgari et al (2025) describe a clinician-in-the-loop framework (CREOLA) that stress-tests documentation assistants with a medical-device-style risk scheme, permitting only versions that meet safety thresholds to progress(\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eOnce implemented, AI generates continuous real-world data - an opportunity and obligation to monitor performance, safety, and equity at scale(\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). Surveillance should be embedded in continuous improvement cycles using process control, A/B testing, and governance-linked monitoring, yet few organisations have operationalised such systems(\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). Traditional post-market evaluation relies on clinician review of AI outputs(\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e): reliable but costly, poorly scalable and almost impossible to automate. Alternatives link AI predictions to registry-confirmed diagnoses, readmissions, or mortality(\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e), or track whether clinicians accept or reject AI outputs (\u0026ldquo;passive labelling\u0026rdquo;)(\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). Patient- and clinician-reported outcomes highlight usability and trust but remain subjective and intermittent(\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eAn emerging strategy is \u0026ldquo;AI to monitor AI.\u0026rdquo; LLMs have been tested for automated chart review and adverse drug reaction detection(\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). These systems replicate tasks previously reliant on clinicians, offering scale, consistency and automatability at lower cost. Yet they are not equivalent to humans: LLMs excel at structured extraction but struggle with ambiguity or missing data, risking error amplification. Human reviewers, by contrast, interpret context and detect what is absent(\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e). The most promising models therefore combine both, with LLMs surfacing structured outputs and clinicians adjudicating uncertain or high-risk cases.\u003c/p\u003e\u003cp\u003eA key domain where this approach could be transformative is monitoring diagnostic decision support systems (DDSS). For systems generating condition suggestions, a central challenge is assessing alignment with confirmed diagnoses(\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Until now, this has relied on panels of physicians judging concordance with clinical diagnoses(\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). The task is further complicated by the nature of clinical documentation: recorded diagnoses are often ambiguous or coded at a high level of generality, may reflect only part of a multimorbid presentation, or remain provisional while investigations are ongoing.\u003c/p\u003e\n\u003ch3\u003eObjectives\u003c/h3\u003e\n\u003cp\u003eThis study aimed to (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) develop a reproducible methodology for assessing DDSS outputs against confirmed diagnoses using a framework that accounts for diagnostic ambiguity and context, and (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) evaluate whether LLMs can replicate or scale this process by automating diagnostic-context classification and condition mapping with sufficient reliability to support post-market surveillance.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eStudy Design and Setting\u003c/h2\u003e\u003cp\u003eThis diagnostic substudy was nested within ESSENCE, a prospective quality-improvement evaluation of a digital diagnostic decision support system (DDSS) developed by Ada Health (Berlin, Germany) and implemented in Portugal\u0026rsquo;s largest private healthcare network (CUF). Adults (\u0026ge;\u0026thinsp;18 years) who completed a Portuguese-language symptom assessment in the myCUF mobile application between Nov 1, 2023, and Oct 31, 2024, provided electronic consent, and shared their report with a CUF clinician were eligible. For this substudy, we included all encounters with both an ICD-10\u0026ndash;coded diagnosis and clinical notes available in the electronic health record (EHR).\u003c/p\u003e\u003c/div\u003e\n\u003ch2\u003eData Collection Procedures\u003c/h2\u003e\u003cp\u003eData collection procedures are outlined in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The full ESSENCE protocol has been reported previously (Cotte et al 2025, preprint). In brief, participants completed a symptom assessment and pre- and post-assessment care intentions, and shared the report with CUF physicians as part of routine care. Healthcare-seeking behavior was tracked through EHR review and follow-up email surveys. A research assistant extracted unstructured clinical notes, ICD-10-coded diagnoses, and study-related metadata into a secure electronic system.\u003c/p\u003e\u003cp\u003eAll free-text notes and diagnoses were translated from Portuguese to English using a locally run Mistral-Nemo-Instruct-2407, following manual anonymisation and approval by the data protection officer. Anonymisation followed standardized rules: removal of names, dates, occupations, and location identifiers, and conversion of exact ages into bands. A summary of all actions and affected cases is provided in Supplementary Table\u0026nbsp;1.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eManual filtering and mapping (reference set)\u003c/h3\u003e\n\u003cp\u003eClinicians employed by the DDSS and familiar with its ontology reviewed each case to establish a reference standard. They determined whether the clinic encounter addressed the same problem as the DDSS assessment and whether a final diagnosis had been reached. Cases with a different complaint, only a symptom recorded, or unspecific ICD-10 codes without clarifying details were excluded. Definitive diagnoses were mapped to one or more DDSS conditions; broad codes were refined using consultation notes. If no equivalent existed, cases were recorded as \u0026ldquo;condition not covered.\u0026rdquo; Reviews were performed independently, with disagreements resolved by consensus and a final reviewer ensuring consistency. Definitions of all diagnostic-context labels are provided in Supplementary Table\u0026nbsp;2.\u003c/p\u003e\n\u003ch3\u003eLLM-based filtering and mapping\u003c/h3\u003e\n\u003cp\u003eWe next evaluated whether a LLM could reproduce this workflow. The model received the same anonymised and translated notes, ICD-10 codes, and user-entered symptoms from the DDSS assessment. Prompts instructed the LLM to act as a medical coding specialist and to follow the same four-step logic: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) filtering cases based on the match between the user's complaint and the clinical encounter, (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) assessing the specificity of the final diagnosis, (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) refining unspecific diagnoses using details from the clinical notes, and (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e) mapping the final, specified diagnosis to the DDSS's condition ontology. Few-shot examples of each diagnostic-context label were included to standardise reasoning, and outputs were returned in structured JSON format.\u003c/p\u003e\u003cp\u003eBecause the full DDSS ontology exceeded the model\u0026rsquo;s context window, we created vector embeddings for each DDSS condition (BAAI/bge-large-en-v1.5). For each case, the ICD-10 description and note text were used to query this index, and the top 100 most relevant conditions were inserted into the prompt as the candidate set. To ensure reliability, the model\u0026rsquo;s temperature was fixed at 0, and ten outputs were generated per case. Final labels and mappings were determined by majority vote. The process was fully automated in Python (OpenAI API).\u003c/p\u003e\n\u003ch3\u003eCondition Matching\u003c/h3\u003e\n\u003cp\u003e Two case sets were created: one from manual review and one from the LLM. These overlapped but were not identical, reflecting differences in eligibility and mappings. Each set was compared with condition suggestions from two DDSS versions. The original system, based on Ada\u0026rsquo;s probabilistic reasoning, was the version deployed during ESSENCE (Miller et al. 2020). The experimental version was a compound AI system under development utilising an LLM from the Llama 3 family. It was included in line with principles of post-market surveillance, where monitoring should establish whether newer versions provide measurable improvements. For fair comparison, it was restricted to the same user inputs without generating extra questions.\u003c/p\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003eData Analysis\u003c/h2\u003e\u003cp\u003eAnalyses were conducted in Python (v3.12) and Google Sheets. Diagnostic-context labelling was assessed by overall agreement and Cohen\u0026rsquo;s κ, with precision, recall, and F1 reported per label. Systematic differences between manual and LLM distributions were tested using χ\u0026sup2;. Eligibility for diagnostic-accuracy analysis was assessed as a binary outcome, with accuracy, sensitivity, specificity, precision, and F1 reported.\u003c/p\u003e\u003cp\u003eFor condition mapping, LLM-assigned diagnoses were compared with clinician mappings. Outputs were classified as exact match (LLM selected one of the clinician-assigned conditions), plausible additional (clinically valid but overlooked by clinicians), or mismatch (clinically inappropriate). Comparative analyses were performed for GPT-4.1 (gpt-4.1-2025-04-14) and GPT-5 (gpt-5-2025-08-07, reasoning effort \u0026ldquo;high\u0026rdquo;).\u003c/p\u003e\u003cp\u003e Finally, the diagnostic accuracy of the DDSS was assessed by comparing reference diagnoses - derived from manual clinician review or GPT-5 mapping - with the ranked condition suggestions generated by the system. Accuracy was evaluated at the Top-1, Top-3, and Top-5 levels, defined as the proportion of cases in which the reference diagnosis appeared among the first one, three, or five system outputs. Diagnostic accuracy was reported with 95% CIs. Two-sample tests for proportions were used for unpaired comparisons, and McNemar\u0026rsquo;s test was applied for paired analyses. Comparisons between original and experimental DDSS versions were conducted within each reference set using paired methods. Cases in which the real condition was not modelled by the DDSS were still included in the diagnostic accuracy analysis.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eMissing Data Handling\u003c/h3\u003e\n\u003cp\u003eEncounters without ICD-10\u0026ndash;coded diagnoses or consultation notes were excluded, as these data formed the reference standard and could not be imputed. Reasons for exclusion are listed in Supplementary Table\u0026nbsp;3.\u003c/p\u003e\n\u003ch3\u003eEthics and Data Governance\u003c/h3\u003e\n\u003cp\u003e The ESSENCE study received ethical approval from the Comiss\u0026atilde;o de \u0026Eacute;tica para a Investiga\u0026ccedil;\u0026atilde;o Cl\u0026iacute;nica (2204JJ351, 2309JJ660) and was registered on ClinicalTrials.gov (NCT06846957) in 2025-02-21. All participants provided electronic informed consent. Data were anonymised and stored securely in Teamscope, and only de-identified records were used for LLM analysis. Procedures complied with the Declaration of Helsinki and ISO 14155:2020.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Federal Ministry of Education and Research (Germany) funded the study but had no role in design, data collection, analysis, interpretation, or writing.\u003c/p\u003e"},{"header":"Findings","content":"\u003cp\u003e\u003cstrong\u003eParticipant characteristics and inclusions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of 512 participants were included. The mean age was 39\u0026middot;2 years (SD 12\u0026middot;4; median 37 years [IQR 30-47]); 80\u0026middot;3% were aged 18-49 years. Females accounted for 57\u0026middot;6% (295/512). Females accounted for 57\u0026middot;6% (295/512). Participants reported a median of 2 presenting complaints (IQR 1\u0026ndash;3; mean 2\u0026middot;21 [SD 1\u0026middot;43]), higher in females (2\u0026middot;40) than males (1\u0026middot;95; p=0\u0026middot;00036). The most common specialties were otorhinolaryngology/ENT (24\u0026middot;0%), orthopedics/trauma surgery (18\u0026middot;2%), gynecology (14\u0026middot;6%), and gastroenterology (13\u0026middot;5%). Details are provided in Table 1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1\u003c/strong\u003e. Participant characteristics\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"569\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCharacteristic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eValue\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eTotal N\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e512\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eAge, mean (SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e39\u0026middot;2 (12\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eMedian (IQR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e37 (30\u0026ndash;47)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eAge range\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e18\u0026ndash;79\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAge distribution, n (% of 512)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e18\u0026ndash;29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e118 (23\u0026middot;0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e30\u0026ndash;39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e167 (32\u0026middot;6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e40\u0026ndash;49\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e126 (24\u0026middot;6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e50\u0026ndash;59\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e66 (12\u0026middot;9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e60\u0026ndash;69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e25 (4\u0026middot;9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e70\u0026ndash;79\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e10 (2\u0026middot;0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e80+\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e0 (0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSex, n (% of 512)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eFemale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e295 (57\u0026middot;6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eMale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e217 (42\u0026middot;4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003ePresenting complaints (PCs)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eMean (SD) overall\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e2\u0026middot;21 (1\u0026middot;43)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eMedian (IQR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e2 (1\u0026ndash;3)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eFemale, mean\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e2\u0026middot;4\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eMale, mean\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e1\u0026middot;95\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003et-test p-value (F vs M)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e0\u0026middot;00036\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSpecialty assigned (n, %, mean #PC, t-test p-value)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eOtorhinolaryngology / ENT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e123 (24\u0026middot;0%), 2\u0026middot;54\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eOrthopedics / Trauma Surgery\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e93 (18\u0026middot;2%), 1\u0026middot;46\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eGynecology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e75 (14\u0026middot;6%), 1\u0026middot;91\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eGastroenterology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e69 (13\u0026middot;5%), 2\u0026middot;54\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003ePulmonology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e40 (7\u0026middot;8%), 3\u0026middot;10\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eDermatology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e28 (5\u0026middot;5%), 1\u0026middot;71\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eNeurology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e17 (3\u0026middot;3%), 2\u0026middot;47\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003ePsychiatry\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e17 (3\u0026middot;3%), 2\u0026middot;88\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eOphthalmology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e14 (2\u0026middot;7%), 1\u0026middot;71\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eUrology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e13 (2\u0026middot;5%), 1\u0026middot;69\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eNephrology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e8 (1\u0026middot;6%), 2\u0026middot;00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eCardiology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e3 (0\u0026middot;6%), 5\u0026middot;00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eFamily Medicine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e3 (0\u0026middot;6%), 1\u0026middot;67\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eDentistry\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e2 (0\u0026middot;4%), 2\u0026middot;50\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eRheumatology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e2 (0\u0026middot;4%), 2\u0026middot;00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eAngiology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e1 (0\u0026middot;2%), 2\u0026middot;00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eEndocrinology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e1 (0\u0026middot;2%), 2\u0026middot;00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eInternal Medicine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e1 (0\u0026middot;2%), 6\u0026middot;00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.5308%;\"\u003e\n \u003cp\u003eOncology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 57.4692%;\"\u003e\n \u003cp\u003e1 (0.2%), 1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eOf 1,470 participants in the main cohort, confirmed healthcare-seeking behavior was available for 721 (49\u0026middot;0%). Within this group, 512 cases had both an ICD-10 code and consultation notes. Fourteen illustrative cases were set aside as examples for prompt development (Supplementary Table 4), leaving 498 cases for which labels were assigned by both manual review and the LLM (Figure 2).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAnonymisation Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll free-text notes were anonymised before LLM mapping. Of 512 notes, 235 (45\u0026middot;9%) required modification, most commonly conversion of exact ages to bands (n=186) and removal of dates (n=59), occupations (n=35), or locations (n=32). Names were removed in 18 cases. Over half of notes (283/512, 55\u0026middot;3%) required no changes (Supplementary Table 1).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFiltering and Mapping\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eManual review classified 385/498 encounters (77\u0026middot;3%) as eligible and distributed the remaining 113 (22\u0026middot;7%) across exclusion categories: 46 (9\u0026middot;2%) \u0026ldquo;no diagnosis\u0026ndash;symptom,\u0026rdquo; 12 (2\u0026middot;4%) \u0026ldquo;no diagnosis\u0026ndash;unspecific code,\u0026rdquo; 25 (5\u0026middot;0%) \u0026ldquo;multimorbidity,\u0026rdquo; 21 (4\u0026middot;2%) \u0026ldquo;different complaint,\u0026rdquo; and 9 (1\u0026middot;8%) \u0026ldquo;not-covered condition\u0026rdquo; (Table 2).\u003c/p\u003e\n\u003cp\u003eGPT-5 reproduced this overall distribution, also assigning 385/498 (77\u0026middot;3%) encounters as eligible. Among these, 347/385 (90\u0026middot;1%) overlapped with manual review. The model missed 38/385 (9\u0026middot;9%) manually eligible cases, while conversely assigning diagnoses in 38/113 (33\u0026middot;6%) cases that clinicians had excluded.\u003c/p\u003e\n\u003cp\u003eAt the binary level (eligible/ineligible), GPT-5 achieved 84\u0026middot;7% accuracy (95% CI 81\u0026middot;3\u0026ndash;87\u0026middot;6), with sensitivity 90\u0026middot;1% (95% CI 86\u0026middot;7\u0026ndash;92\u0026middot;7) and specificity 66\u0026middot;4% (95% CI 57\u0026middot;3\u0026ndash;74\u0026middot;4). Agreement beyond chance was moderate (\u0026kappa;=0\u0026middot;57, 95% CI 0\u0026middot;50\u0026ndash;0\u0026middot;64) (Table 3). Performance was robust for eligible cases (precision 0\u0026middot;90, recall 0\u0026middot;90, F1=0\u0026middot;90) but less consistent for exclusion categories, where unspecific codes were frequently over-assigned and multimorbidity under-recognised.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel comparison: GPT-4.1 vs GPT-5.0\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBoth models reproduced manual eligibility classifications with high overall accuracy (Table 3). GPT-4.1 achieved higher sensitivity (97\u0026middot;1% vs 90\u0026middot;1%) but much lower specificity (37\u0026middot;2% vs 66\u0026middot;4%). GPT-5\u0026middot;0 provided a more balanced profile with higher precision (90\u0026middot;1% vs 84\u0026middot;0%) and stronger agreement (\u0026kappa;=0\u0026middot;57 vs 0\u0026middot;42). Running this workflow required $5.08 for a single pass with GPT-4.1 and $44.40 for GPT-5.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCondition mapping performance\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAmong jointly eligible cases, GPT-5 exactly matched one of the clinician-assigned diagnoses in 324/347 (93\u0026middot;6%) and proposed an additional clinically plausible condition in 12 (3\u0026middot;5%). Disagreement occurred in 10 cases (2\u0026middot;9%), and no output in one (0\u0026middot;3%). GPT-4.1 showed a similar profile: 354/374 (94\u0026middot;7%) exact matches, 8 (2\u0026middot;1%) plausible additions, and 12 (3\u0026middot;2%) mismatches. When exact matches and plausible additions were combined, overall alignment was 97\u0026middot;1% for GPT-5 and 96\u0026middot;8% for GPT-4.1. Full distributions are shown in Table 3.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsistency\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo assess the internal consistency of GPT-5\u0026rsquo;s label assignment, we generated eight independent completions per case and quantified the distribution of distinct labels. In 385/498 cases (77.3%), all completions converged on the same label. Among the remaining 113 cases, two distinct labels were suggested in 63/498 (12.7%), three in 38/498 (7.6%), and four in 12/498 (2.4%). Even when multiple labels appeared, the majority label typically dominated: for two-label cases, 70% (44/63) showed a 6\u0026ndash;2 or 7\u0026ndash;1 split. In supplementary table 5, we present a confusion matrix of LLM vs. manual labels, with some examples provided in supplementary table 6.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2.\u0026nbsp;\u003c/strong\u003eDiagnostic-context label distribution and per-label performance\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"615\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLabel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eManual n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eGPT-5 n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBoth\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMissed\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrecision\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRecall\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eF1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMappable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e385 (77\u0026middot;3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e385 (77\u0026middot;3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e347\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e0\u0026middot;9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e0\u0026middot;9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e0\u0026middot;9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNo Dx \u0026ndash; Symptom\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e46 (9\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e41 (8\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e30\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e16\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e0\u0026middot;73\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e0\u0026middot;65\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e0\u0026middot;69\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNo Dx \u0026ndash; Unspecific code\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e12 (2\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e27 (5\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e0\u0026middot;19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e0\u0026middot;42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e0\u0026middot;26\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMultimorbidity\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e25 (5\u0026middot;0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e11 (2\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e0\u0026middot;91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e0\u0026middot;4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e0\u0026middot;56\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDifferent complaint\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e21 (4\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e22 (4\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e11\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e0\u0026middot;5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e0\u0026middot;52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e0\u0026middot;51\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNot-covered condition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e9 (1\u0026middot;8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e12 (2\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e0\u0026middot;5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e0\u0026middot;67\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e0\u0026middot;57\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 15.4472%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.0081%;\"\u003e\n \u003cp\u003e498 (100)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.2195%;\"\u003e\n \u003cp\u003e498 (100)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.5935%;\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 14.1463%;\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 11.7073%;\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.5203%;\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 12.3577%;\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3.\u0026nbsp;\u003c/strong\u003ePerformance of GPT-4.1 and GPT-5.0 for eligibility classification and condition mapping (n=498)\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"587\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eGPT-4.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eGPT-5.0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"bottom\" style=\"width: 100%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEligibility classification\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eTrue Positives (TP)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e374\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e347\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eFalse Negatives (FN)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e11\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eFalse Positives (FP)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e71\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eTrue Negatives (TN)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e75\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eManual eligible cases\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e385\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e385\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eModel eligible cases\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e445\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e385\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eAccuracy (TP+TN / N)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e83\u0026middot;5%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e84\u0026middot;7%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eSensitivity (recall for eligible)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e97\u0026middot;1%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e90\u0026middot;1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eSpecificity (recall for ineligible)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e37\u0026middot;2%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e66\u0026middot;4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003ePrecision (PPV for eligible)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e84\u0026middot;0%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e90\u0026middot;1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eF1 score (for eligible)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e90\u0026middot;1%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e90\u0026middot;1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eCohen\u0026rsquo;s \u0026kappa; (95% CI)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e0\u0026middot;42 (0\u0026middot;33\u0026ndash;0\u0026middot;52)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e0\u0026middot;57 (0\u0026middot;50\u0026ndash;0\u0026middot;64)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"bottom\" style=\"width: 100%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCondition mapping\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eExact match\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e354/374 (94\u0026middot;7%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e324/347 (93\u0026middot;6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eShould be considered (plausible addition)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e8/374 (2\u0026middot;1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e12/347 (3\u0026middot;5%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eNo overlap (mismatch)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e12/374 (3\u0026middot;2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e10/347 (2\u0026middot;9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 42.0784%;\"\u003e\n \u003cp\u003eNo output\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e\u0026ndash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 28.9608%;\"\u003e\n \u003cp\u003e1/347 (0\u0026middot;3%)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eMatching comparison\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor the diagnostic accuracy analyses, we included all cases, including those where the DDSS did not explicitly model the underlying condition, yielding 394 manual and 397 GPT-5\u0026ndash;mapped cases (Table 4).\u003c/p\u003e\n\u003cp\u003eWith manual mapping, Top-1 accuracy was 47\u0026middot;2% (42\u0026middot;3\u0026ndash;52\u0026middot;1) for the original and 48\u0026middot;0% (47\u0026middot;5\u0026ndash;52\u0026middot;9) for the experimental DDSS version; Top-3 accuracy was 71\u0026middot;6% (66\u0026middot;5\u0026ndash;76\u0026middot;3) versus 67\u0026middot;5% (62\u0026middot;4\u0026ndash;72\u0026middot;2); and Top-5 accuracy was 78\u0026middot;4% (73\u0026middot;2\u0026ndash;83\u0026middot;1) versus 81\u0026middot;2% (76\u0026middot;0\u0026ndash;85\u0026middot;8). With GPT-5 mapping, Top-1 accuracy was 43\u0026middot;1% (38\u0026middot;3\u0026ndash;48\u0026middot;0) versus 44\u0026middot;1% (39\u0026middot;2\u0026ndash;49\u0026middot;0); Top-3 accuracy was 65\u0026middot;2% (60\u0026middot;3\u0026ndash;69\u0026middot;8) versus 67\u0026middot;0% (62\u0026middot;2\u0026ndash;71\u0026middot;5); and Top-5 accuracy was 71\u0026middot;3% (66\u0026middot;6\u0026ndash;75\u0026middot;6) versus 75\u0026middot;1% (70\u0026middot;6\u0026ndash;79\u0026middot;1).\u003c/p\u003e\n\u003cp\u003eAcross the 385 analyzable cases per set, manual and GPT-5 mapping performed similarly (supplementary table 7). Differences ranged from \u0026ndash;4 to +7 percentage points, with only one significant finding: manual mapping achieved higher Top-5 accuracy under the original DDSS version (309/385 [80\u0026middot;3%] vs 283/385 [73\u0026middot;5%]; \u0026Delta;=+6\u0026middot;8 pp, 95% CI +0\u0026middot;8 to +12\u0026middot;7; p=0\u0026middot;026).\u003c/p\u003e\n\u003cp\u003eAfter GPT5 did not provide an output in one case, 346 cases overlapped. For these, results were almost identical (supplementary table 8). McNemar tests showed no significant differences at any threshold, with small differences (\u0026ndash;5 to +2 pp) and confidence intervals crossing zero, indicating that GPT-5 mapping was statistically equivalent to manual mapping when applied to the same cases.\u003c/p\u003e\n\u003cp\u003ePaired comparisons of the original and experimental DDSS versions (supplementary table 9) showed no significant differences at the Top-1 or Top-3 thresholds. At Top-5, however, the experimental version was consistently more accurate: manual 309/396 (78\u0026middot;0%) vs 320/396 (80\u0026middot;8%; \u0026Delta;=+3\u0026middot;1 pp, 95% CI +0\u0026middot;4 to +5\u0026middot;8; p=0\u0026middot;038) and GPT-5 283/397 (71\u0026middot;3%) vs 298/397 (75\u0026middot;1%; \u0026Delta;=+3\u0026middot;9 pp, 95% CI +1\u0026middot;2 to +6\u0026middot;6; p=0\u0026middot;009).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4\u003c/strong\u003e. Diagnostic accuracy of DDSS using manual and GPT-5 reference mapping.\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"634\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 73px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eReference set\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eN\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eVersion\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTop-1 % (95% CI)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTop-3 % (95% CI)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTop-5 % (95% CI)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"bottom\" style=\"width: 73px;\"\u003e\n \u003cp\u003eManual-\u003c/p\u003e\n \u003cp\u003emapped\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e394\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eOriginal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e47\u0026middot;2 % (42\u0026middot;3\u0026ndash;52\u0026middot;1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e71\u0026middot;6 % (66\u0026middot;5\u0026ndash;76\u0026middot;3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e78\u0026middot;4 % (73\u0026middot;2\u0026ndash;83\u0026middot;1)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e394\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e48\u0026middot;0% (47\u0026middot;5\u0026ndash;52\u0026middot;9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e67\u0026middot;5 % (62\u0026middot;4\u0026ndash;72\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e81\u0026middot;2 % (76\u0026middot;0\u0026ndash;85\u0026middot;8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"bottom\" style=\"width: 73px;\"\u003e\n \u003cp\u003eGPT-5-\u003c/p\u003e\n \u003cp\u003emapped\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e397\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eOriginal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e43\u0026middot;1 % (38\u0026middot;3\u0026ndash;48\u0026middot;0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e65\u0026middot;2 % (60\u0026middot;3\u0026ndash;69\u0026middot;8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e71\u0026middot;3 % (66\u0026middot;6\u0026ndash;75\u0026middot;6)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e397\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e44\u0026middot;1 % (39\u0026middot;2\u0026ndash;49\u0026middot;0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e67\u0026middot;0 % (62\u0026middot;2\u0026ndash;71\u0026middot;5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e75\u0026middot;1 % (70\u0026middot;6\u0026ndash;79\u0026middot;1)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"bottom\" style=\"width: 73px;\"\u003e\n \u003cp\u003eManual-\u003c/p\u003e\n \u003cp\u003emapped\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e346\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eOriginal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e46\u0026middot;8 % (41\u0026middot;6\u0026ndash;52\u0026middot;1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e72\u0026middot;8 % (67\u0026middot;9\u0026ndash;77\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e80\u0026middot;3 % (75\u0026middot;8\u0026ndash;84\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e346\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e48\u0026middot;3 % (43\u0026middot;1\u0026ndash;53\u0026middot;5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e68\u0026middot;8 % (63\u0026middot;7\u0026ndash;73\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e83\u0026middot;5 % (79\u0026middot;3\u0026ndash;87\u0026middot;1)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"bottom\" style=\"width: 73px;\"\u003e\n \u003cp\u003eGPT-5-\u003c/p\u003e\n \u003cp\u003emapped\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e346\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eOriginal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e48\u0026middot;0 % (42\u0026middot;8\u0026ndash;53\u0026middot;2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e71\u0026middot;4 % (66\u0026middot;4\u0026ndash;75\u0026middot;9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e78\u0026middot;0 % (73\u0026middot;4\u0026ndash;82\u0026middot;1)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 45px;\"\u003e\n \u003cp\u003e346\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 93px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 141px;\"\u003e\n \u003cp\u003e49\u0026middot;1 % (43\u0026middot;9\u0026ndash;54\u0026middot;4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 148px;\"\u003e\n \u003cp\u003e74\u0026middot;0 % (69\u0026middot;1\u0026ndash;78\u0026middot;3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 134px;\"\u003e\n \u003cp\u003e81\u0026middot;2 % (76\u0026middot;8\u0026ndash;85\u0026middot;0)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"},{"header":"Interpretation","content":"\u003cp\u003e\u003cstrong\u003eSummary of results\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe evaluated whether large language models (LLMs) can automate real-world monitoring of diagnostic accuracy in digital diagnostic decision support systems (DDSS) using a filter\u0026ndash;map\u0026ndash;match framework. Of 498 cases, both manual and LLM review classified 385 as eligible for diagnostic accuracy analysis. Within this subset, agreement was high: both approaches classified the same 347 cases (90\u0026middot;1%) as eligible. \u0026nbsp;The LLM additionally marked 38 cases as eligible that clinicians had excluded, while excluding 38 that clinicians had included. GPT-5 mapped the same clinician diagnoses in 93\u0026middot;6% of overlapping cases and suggested plausible alternatives in 3\u0026middot;5%. Diagnostic accuracy was statistically indistinguishable between manual and GPT-5 mapping at Top-1 and Top-3, with one difference at Top-5. In overlapping cases, no significant differences remained, confirming equivalence when eligibility filtering is held constant. Across reference sets, both approaches identified that the experimental DDSS was more accurate than the original only at the Top-5 threshold.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFiltering\u0026ndash;mapping\u0026ndash;matching\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eComparing DDSS outputs with real-world outcomes requires identifying truly comparable encounters. In practice, the digital assessment and clinic visit may cover different complaints, diagnoses may still be under investigation, or multimorbidity can complicate documentation. In our cohort, only 77\u0026middot;3% of encounters were suitable for diagnostic comparison. Analysing all cases together would conflate incomparable scenarios and distort estimates. A structured filter\u0026ndash;map\u0026ndash;match workflow - (i) filtering to comparable encounters, (ii) mapping confirmed diagnoses to a shared ontology, and (iii) matching mapped diagnoses to DDSS outputs - offers a more rigorous basis for monitoring.\u003c/p\u003e\n\u003cp\u003eFor filtering, the model achieved high sensitivity and precision but only moderate specificity. This asymmetry is acceptable in surveillance: retaining most eligible cases preserves sample size, while extra inclusions can be resolved downstream. Similar frameworks are being explored for case identification in electronic records. Cheligeer et al, for example, applied a multi-stage method to detect hospital-acquired pulmonary embolism from 10,066 inpatients, of whom only 40 (0\u0026middot;4%) had true events. Models achieved sensitivity 87\u0026middot;5\u0026ndash;100% and specificity 94\u0026middot;9\u0026ndash;98\u0026middot;9%, but positive predictive values remained low (\u0026asymp;7\u0026ndash;17%) because of rarity, with F1 scores peaking at 28\u0026middot;1%(22). Their method efficiently ruled out negatives at scale, but most flagged cases still needed manual review.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOne key objective of this research was to identify whether automating the entire filter-map-match workflow with AI would distort estimates of diagnostic accuracy. When comparing the full sets of analyzable cases (385 each), GPT-5 mapping produced results that were statistically indistinguishable from manual review at the Top-1 and Top-3 thresholds, with a single significant difference observed at Top-5.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWhen restricting the analysis to the 347 overlapping cases, where the task was limited to mapping rather than eligibility filtering, no statistical differences were observed at any threshold. This indicates that the mapping step can already be automated with high reliability. The only residual divergence arose from imperfect eligibility filtering, which accounted for the single significant difference in one of the three comparisons. Importantly, both manual and GPT-5 mapping consistently identified that the experimental DDSS version was significantly more accurate at the Top-5 threshold only.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethodological considerations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eICD-10 codes were initially considered as the main comparator, but have well-described limitations: agreement is only moderate at chapter level and poor for specific codes, with high inter-rater variability(23\u0026ndash;25). Coding complexity and incomplete documentation often yield generic labels(26,27). We therefore combined ICD-10 review with free-text note analysis, which clarified unspecific codes in several cases. LLMs now make it feasible to screen documentation at scale, with high sensitivity and specificity for phenotyping but only moderate positive predictive value and limited causal attribution, requiring clinician oversight(28).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStrengths\u003cbr\u003e\u003c/strong\u003eThis study has several strengths. It was conducted in a real-world clinical environment within Portugal\u0026rsquo;s largest private healthcare network, capturing the heterogeneity of routine practice beyond curated test sets. The inclusion of multiple specialties increases the generalisability of findings. A clinician-adjudicated reference standard ensured reproducibility of comparisons. Finally, the filter\u0026ndash;map\u0026ndash;match framework was explicitly aligned with clinical reality, and the privacy-preserving pipeline systematically anonymised notes, ensuring GDPR compliance. Finally, we also compared downstream diagnostic accuracy across both manual and GPT-5 mapping approaches and between original and experimental DDSS versions. This allowed us to test how the approaches would behave under real-world performance monitoring conditions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLimitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study has several limitations. Manual review is an imperfect reference, and in some instances the LLM proposed plausible alternatives not documented by clinicians, highlighting the constraints of the comparator. Case review was conducted by clinicians employed by the DDSS and familiar with its ontology, which may have introduced bias. The analysis was a proof-of-concept, one-off evaluation rather than an implemented surveillance system, and cases without a final diagnosis at data collection were excluded, though some might have become eligible later. Performance was also assessed within a single healthcare network, so results may differ in other settings with different documentation practices or diagnostic distributions. Embedding the workflow in live systems will be essential to establish longitudinal performance, feasibility, and clinical impact.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eImplementation in Clinical Practice\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSeveral lessons emerge for real-world use. First, the core decision is binary- whether a case is eligible for diagnostic comparison. Although a multi-label framework was useful for benchmarking, rare categories performed inconsistently and added little operational value. Prioritising this binary decision offers a more pragmatic foundation for implementation, particularly as the mapping step showed no statistically significant differences between manual and GPT-5 review.\u003c/p\u003e\n\u003cp\u003eSecond, costs and cascade design matter. On our dataset, end-to-end inference cost US$44\u0026middot;40 with GPT-5 (\u0026asymp;US$0\u0026middot;09 per case) and US$5\u0026middot;08 with GPT-4\u0026middot;1 (\u0026asymp;US$0\u0026middot;01 per case). Given that divergence arose chiefly in eligibility filtering rather than mapping, a cost-sensitive routing strategy is feasible: use a smaller model for high-throughput filtering, escalate low-confidence or high-risk cases to GPT-5, and audit a sample with clinicians. This balances accuracy, scale, and spend, and can be tuned by confidence thresholds and sampling rates.\u003c/p\u003e\n\u003cp\u003eThird, diagnosis is often a longitudinal process rather than a single clinical event. Patients may await test results or referrals before a final diagnosis is reached. Systems should therefore re-analyse such cases once definitive information becomes available, converting encounters initially excluded into eligible cases for analysis.\u003c/p\u003e\n\u003cp\u003eFourth, in-workflow one-click feedback controls offer a practical way to capture clinician judgement at the point of care - for example, by accepting a suggested condition or selecting a better match from a type-ahead ontology. Structured reporting in radiology has shown that even a single click can return machine-readable feedback for continual learning (Fuchs 2024). Although clinicians are unlikely to provide feedback for every consultation, such mechanisms are a scalable and efficient addition, delivering targeted corrections without disrupting workflow.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFinally, privacy must be preserved. Advances in locally deployable anonymisation models now achieve near-perfect sensitivity and character-level accuracy for de-identifying clinical text without data leaving the system(29). Integrating such tools could replace manual redaction and enable GDPR-compliant note review at scale.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study shows that large language models can automate key steps in monitoring the diagnostic accuracy of decision support systems. GPT-5 produced results that were highly consistent with clinician review, with the few differences explained by eligibility filtering - deciding which real-world cases could be included in the diagnostic accuracy analysis- rather than by the mapping step itself. Both manual and LLM-based approaches led to the same conclusions about the performance improvements from the original to the experimental version of the DDSS, indicating that LLM-based monitoring can be trusted to detect meaningful changes in system performance without continuous manual review.\u003c/p\u003e\n\u003cp\u003eWhile these findings establish feasibility, further work is required to embed the workflow in live systems and to extend it to automated continuous monitoring. Focusing on improving eligibility filtering while including re-evaluation of provisional cases, one-click feedback, and local anonymisation will be critical to test scalability and clinical impact. If operationalised, such pipelines could enable continuous, learning health system monitoring of diagnostic AI.\u003c/p\u003e\n\u003ch2\u003eResearch in Context:\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eEvidence before this study\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMonitoring the diagnostic accuracy of digital decision support systems (DDSS) has traditionally relied on manual clinician review. Although accurate, this approach is labour-intensive and difficult to scale. Previous work shows that large language models (LLMs) can extract information from clinical notes and map diagnoses to standard ontologies, but their ability to fully replace clinical review in continuous surveillance of DDSS performance had not been established.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdded value of this study\u003cbr\u003e\u003c/strong\u003eThis study tested whether an LLM could automate the filter-map-match workflow used to monitor DDSS accuracy in real-world practice. GPT-5 achieved high agreement with manual review, reproducing clinician-adjudicated diagnoses in most cases and generating comparable diagnostic accuracy estimates. Importantly, both manual and LLM-based workflows led to the same conclusions about relative performance of the original versus experimental DDSS versions. Differences between methods arose mainly from eligibility filtering rather than mapping itself.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eImplications of all the available evidence\u003cbr\u003e\u003c/strong\u003eThese findings show that LLMs can reliably replicate the core elements of clinician-led surveillance of DDSS diagnostic accuracy. The mapping step appears nearly ready for practice, whereas eligibility filtering requires further refinement to ensure consistency. If implemented, such an approach could substantially reduce reliance on manual practitioner panels and enable scalable, near real-time monitoring of DDSS in routine care.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAcknowledgments\u003c/h2\u003e\n\u003cp\u003eThis work was supported by the Federal Ministry of Education and Research (Bundesministerium f\u0026uuml;r Bildung und Forschung) through the European Union-financed NextGenerationEU program under grant number 16KISA100K, project PATH\u0026mdash;\u0026ldquo;Personal Mastery of Health and Wellness Data.\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eAuthor Contributions\u003cbr\u003e\u0026nbsp;FC and SG were responsible for conceptualisation and methodology. Data collection and investigation were carried out by FDL, MP, MSM, PF, NK, MS, TM, and FC. Data anonymisation was performed by MS. NK, OS, FC, and KG conducted the manual filtering and mapping of cases to establish the clinician-adjudicated reference set. PB, FC, and OS performed formal analysis and visualisation. FC led writing the original draft in collaboration with all co-authors. VHa, AM, LS, VHe, OS, SK, VM, HH, and PE led the development of the experimental version of the app tested in the study and contributed to review of the manuscript. All authors had full access to aggregated data in the study and contributed to the decision to submit the manuscript for publication. S.G. is a News and Views Editor for npj Digital Medicine. S.G. played no role in the internal review or decision to publish this Original Research article.\u003c/p\u003e\n\u003cp\u003eCorresponding author\u003cbr\u003e\u0026nbsp;Correspondence to Stephen Gilbert [email protected]\u003c/p\u003e\n\u003cp\u003ePedro Flores\u003c/p\u003e\n\u003cp\u003eCompeting Interests\u003cbr\u003e\u0026nbsp;PB, OS, VHa, AM, VHe, SK, HH, PE, KG and TM are employed by Ada Health. MS, NK, LS and VM are former employees of Ada Health. SG and FC are consultants for Ada Health. FC, PB, VHa, LS, VHe, SK, VM, HH, PE, SG, KG and TM hold share options in Ada Health. SG declares a nonfinancial interest as an Advisory Group member of the EY-coordinated \u0026ldquo;Study on Regulatory Governance and Innovation in the field of Medical Devices\u0026rdquo; conducted on behalf of the Directorate-General for Health and Food Safety (SANTE) of the European Commission. SG declares the following competing financial interests: he has or has had consulting relationships with Una Health GmbH, Lindus Health Ltd., Flo Ltd, ICURA ApS, Rock Health Inc., Thymia Ltd., FORUM Institut für Management GmbH, High-Tech Gründerfonds Management GmbH, Prova Health Ltd, Directorate-General for Research and Innovation Of the European Commission. TM declares the following competing financial interests: he has or has had consulting relationships with Suvera \u0026amp; iPlato Healthcare. FC declares the following competing financial interests: she has or has had consulting relationships with Flo Ltd. Furthermore, VHa, AM, LS, SK, VM, HH, and PE are co-inventors on a patent application titled \u0026quot;Computer system and method for supporting medical diagnosis of a human with symptoms associated with the human\u0026apos;s medical condition\u0026quot; (Application No. EP24188258), with rights held by the co-inventors\u0026apos; employer Ada Health. This patent application relates to the compound AI system described in this paper.\u003c/p\u003e\n\u003ch2\u003eData Sharing\u003c/h2\u003e\n\u003cp\u003eData collected for the study, including individual participant data and a data dictionary defining each field in the set, will be made available to others after publication upon reasonable request, subject to approval. Requests for access should be made to the study team at Ada Health ([email protected]). After approval, a signed data sharing agreement will be required before data release.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eMohseni A, Ghotbi E, Kazemi F, Shababi A, Jahan SC, Mohseni A, et al. Artificial Intelligence in Radiology. Radiol Clin North Am. 2024;62(6):935\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLu MY, Chen B, Williamson DFK, Chen RJ, Zhao M, Chow AK, et al. A multimodal generative AI copilot for human pathology. Nature. 2024;634(8033):466\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFernandes M, Vieira SM, Leite F, Palos C, Finkelstein S, Sousa JMC. Clinical Decision Support Systems for Triage in the Emergency Department using Intelligent Systems: a Review. Artif Intell Med. 2020;102:101762.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKorom R, Kiptinness S, Adan N, Said K, Ithuli C, Rotich O, et al. AI-based Clinical Decision Support for Primary Care: A Real-World Study [Internet]. arXiv; 2025 [cited 2025 July 28]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://arxiv.org/abs/2507.16947\u003c/span\u003e\u003cspan address=\"https://arxiv.org/abs/2507.16947\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePennisi F, Pinto A, Ricciardi GE, Signorelli C, Gianfredi V. Artificial intelligence in antimicrobial stewardship: a systematic review and meta-analysis of predictive performance and diagnostic accuracy. Eur J Clin Microbiol Infect Dis. 2025;44(3):463\u0026ndash;513.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFaiyazuddin Md, Rahman SJQ, Anand G, Siddiqui RK, Mehta R, Khatib MN, et al. The Impact of Artificial Intelligence on Healthcare: A Comprehensive Review of Advancements in Diagnostics, Treatment, and Operational Efficiency. Health Sci Rep. 2025;8(1):e70312.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGilbert S, Pimenta A, Stratton-Powell A, Welzel C, Melvin T. Continuous Improvement of Digital Health Applications Linked to Real-World Performance Monitoring: Safe Moving Targets? Mayo Clin Proc Digit Health. 2023 Sept;1(3):276\u0026ndash;87.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFraser AG, Butchart EG, Szymański P, Caiani EG, Crosby S, Kearney P, et al. The need for transparency of clinical evidence for medical devices in Europe. The Lancet. 2018;392(10146):521\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHan R, Acosta JN, Shakeri Z, Ioannidis JPA, Topol EJ, Rajpurkar P. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit Health. 2024;6(5):e367\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGilbert S, Mathias R, Sch\u0026ouml;nfelder A, Wekenborg M, Steinigen-Fuchs J, Dillenseger A, et al. A roadmap for safe, regulation-compliant Living Labs for AI and digital health development. Sci Adv. 2025;11(20):eadv7719.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePark SH, Han K, Jang HY, Park JE, Lee JG, Kim DW, et al. Methods for Clinical Evaluation of Artificial Intelligence Algorithms for Medical Diagnosis. Radiology. 2023;306(1):20\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMcKinney SM, Sieniek M, Godbole V, Godwin J, Antropova N, Ashrafian H, et al. International evaluation of an AI system for breast cancer screening. Nature. 2020;577(7788):89\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCreagh AP, Pease T, Ashworth P, Bradley L, Duport S. Explainable machine learning to identify patients at risk of developing hospital acquired infections [Internet]. Health Informatics; 2024 [cited 2025 Sept 5]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://medrxiv.org/lookup/doi/\u003c/span\u003e\u003cspan address=\"http://medrxiv.org/lookup/doi/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2024.11.13.24317108\u003c/span\u003e\u003cspan address=\"10.1101/2024.11.13.24317108\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAsgari E, Monta\u0026ntilde;a-Brown N, Dubois M, Khalil S, Balloch J, Yeung JA, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. Npj Digit Med. 2025;8(1):274.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEl Arab RA, Abu-Mahfouz MS, Abuadas FH, Alzghoul H, Almari M, Ghannam A, et al. Bridging the Gap: From AI Success in Clinical Trials to Real-World Healthcare Implementation\u0026mdash;A Narrative Review. Healthcare. 2025;13(7):701.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSemigran HL, Linder JA, Gidengil C, Mehrotra A. Evaluation of symptom checkers for self diagnosis and triage: audit study. BMJ. 2015 July 8;h3480.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAndersen ES, Birk-Korch JB, Hansen RS, Fly LH, R\u0026ouml;ttger R, Arcani DMC, et al. Monitoring performance of clinical artificial intelligence in health care: a scoping review. JBI Evid Synth. 2024;22(12):2423\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFuchs M, Gonzalez C, Frisch Y, Hahn P, Matthies P, Gruening M, et al. Closing the loop for AI-ready radiology. R\u0026ouml;Fo - Fortschritte Auf Dem Geb R\u0026ouml;ntgenstrahlen Bildgeb Verfahr. 2024;196(02):154\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAgatstein K. Chart Review Is Dead; Long Live Chart Review: How Artificial Intelligence Will Make Human Review of Medical Records Obsolete, One Day. Popul Health Manag. 2023;26(6):438\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZitu MM, Owen D, Manne A, Wei P, Li L. Large Language Models for Adverse Drug Events: A Clinical Perspective. J Clin Med. 2025;14(15):5490.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCook C. Challenges with diagnoses: sketchy reference standards. J Man Manip Ther. 2012;20(3):111\u0026ndash;2.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCheligeer C, Southern DA, Yan J, Wu G, Pan J, Lee S, et al. Utilizing large language models for detecting hospital-acquired conditions: an empirical study on pulmonary embolism. J Am Med Inform Assoc. 2025;32(5):876\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWockenfuss R, Frese T, Herrmann K, Claussnitzer M, Sandholzer H. Three- and four-digit ICD-10 is not a reliable classification system in primary care. Scand J Prim Health Care. 2009;27(3):131\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStausberg J, Lehmann N, Kaczmarek D, Stein M. Reliability of diagnoses coding with ICD-10. Int J Med Inf. 2008;77(1):50\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOsterhage KP, Hser Y, Mooney LJ, Sherman S, Saxon AJ, Ledgerwood M, et al. Identifying patients with opioid use disorder using International Classification of Diseases (ICD) codes: Challenges and opportunities. Addiction. 2024;119(1):160\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAlmagro M, Unanue RM, Fresno V, Montalvo S. ICD-10 Coding of Spanish Electronic Discharge Summaries: An Extreme Classification Problem. IEEE Access. 2020;8:100073\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhou L, Cheng C, Ou D, Huang H. Construction of a semi-automatic ICD-10 coding system. BMC Med Inform Decis Mak. 2020;20(1):67.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBejan CA, Wang M, Venkateswaran S, Bergmann EA, Hiles L, Xu Y, et al. irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets [Internet]. Health Informatics; 2025 [cited 2025 Sept 3]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://medrxiv.org/lookup/doi/\u003c/span\u003e\u003cspan address=\"http://medrxiv.org/lookup/doi/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2025.03.05.25323445\u003c/span\u003e\u003cspan address=\"10.1101/2025.03.05.25323445\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWiest IC, Le\u0026szlig;mann ME, Wolf F, Ferber D, Treeck MV, Zhu J, et al. Deidentifying Medical Documents with Local, Privacy-Preserving Large Language Models: The LLM-Anonymizer. NEJM AI [Internet]. 2025 Mar 27 [cited 2025 Sept 3];2(4). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ai.nejm\u003c/span\u003e\u003cspan address=\"https://ai.nejm\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e.org/doi/10.1056/AIdbp2400537\u003c/span\u003e\u003cspan address=\".doi/10.1056/AIdbp2400537\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8022874/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8022874/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e\u003cp\u003eDiagnostic decision support systems (DDSS) are increasingly deployed at scale, yet their diagnostic accuracy is insufficiently monitored once integrated into care. Traditional post-market surveillance relies on clinician review, which is costly, slow, and difficult to sustain. Large language models (LLMs) may offer a scalable and potentially automatable solution, but their performance in real-world monitoring remains unknown.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e\u003cp\u003eWe conducted a diagnostic accuracy substudy within ESSENCE, a prospective evaluation of Ada Health\u0026rsquo;s DDSS integrated into Portugal\u0026rsquo;s largest private healthcare network. Clinical notes and ICD-10 diagnoses from 498 encounters were anonymised and classified using a filter\u0026ndash;map\u0026ndash;match framework. Manual clinician review served as the reference standard. We compared eligibility classification and condition mapping between clinicians and GPT-4.1 and GPT-5, and assessed diagnostic accuracy of two DDSS versions using both reference sets.\u003c/p\u003e\u003ch2\u003eFindings\u003c/h2\u003e\u003cp\u003eManual review classified 385 of 498 encounters (77\u0026middot;3%) as eligible for diagnostic comparison. GPT-5 reproduced these classifications with 84\u0026middot;7% accuracy (κ\u0026thinsp;=\u0026thinsp;0\u0026middot;57), showing high sensitivity but only moderate specificity. Among 347 encounters judged eligible by both approaches, GPT-5 exactly matched clinician-assigned diagnoses in 93\u0026middot;6% and proposed clinically plausible alternatives in 3\u0026middot;5%. Diagnostic accuracy estimates based on manual versus GPT-5 mappings were statistically indistinguishable at Top-1 and Top-3 across the full analyzable sets, with one significant difference at Top-5. In the overlapping 346 cases, no statistical differences were observed. Across both reference sets, the experimental DDSS version outperformed the original only at the Top-5 threshold.\u003c/p\u003e\u003ch2\u003eInterpretation\u003c/h2\u003e\u003cp\u003eLLMs can reproduce clinician review of real-world diagnostic encounters with close agreement. While GPT-5 performed comparably to clinicians for condition mapping, the eligibility filtering step - deciding which encounters should enter the diagnostic-accuracy analysis - remains the main source of divergence and is the priority for improvement. Embedding such approaches into health systems could enable automated and continuous performance and safety monitoring and support regulatory compliance. Broader evaluations across diverse care settings are needed to establish generalisability and equity impact.\u003c/p\u003e","manuscriptTitle":"Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-20 16:03:09","doi":"10.21203/rs.3.rs-8022874/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-01-08T03:48:49+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-02T14:51:41+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-20T23:47:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"94718594676560344991849655691603156356","date":"2025-12-11T09:58:59+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"178072837441239907023328546165936946585","date":"2025-11-22T00:15:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"173816128465983927900448918641529808340","date":"2025-11-16T13:29:11+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-11-11T02:54:30+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-11-10T20:31:35+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-11-10T03:33:22+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Digital Medicine","date":"2025-11-03T22:56:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"85c31f07-cdbb-430a-af2b-1a0ccb810d0a","owner":[],"postedDate":"November 20th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":58064698,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":58064699,"name":"Health sciences/Diseases"},{"id":58064700,"name":"Health sciences/Health care"},{"id":58064701,"name":"Physical sciences/Mathematics and computing"},{"id":58064702,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-05-19T22:23:16+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-20 16:03:09","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8022874","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8022874","identity":"rs-8022874","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00