Automated Extraction of Patient Demographics from Clinical Vignettes Using a Multi-Agent Large Language Model Framework

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background. Clinical vignettes are widely used in medical education to teach clinical reasoning, evaluate diagnostic skills, and contextualize patient care. These vignettes often include demographic information such as age, sex, race, ethnicity, and other identity markers that may influence learners’ understanding of patient diversity. Inclusion of underrepresented populations is particularly important when demographic factors may affect diagnosis and care. Prior research shows that demographic details in clinical educational materials may be presented inconsistently or selectively, reflecting elements of hidden curriculum and reinforcing stereotypes. However, manually extracting demographic information from clinical vignettes in curricular materials for audits is resource-intensive, while existing automated approaches show limited reliability. Methods. We propose a multi-agentic framework to automatically extract demographic information from clinical vignettes. Our framework consists of three stages: (1) a few-shot learning step to identify clinical vignettes, (2) a few-shot infused demographic information extraction step in which multiple parallel agents extract demographic data from identified vignettes; and (3) an LLM-as-a-judge validation step in which extracted information is autonomously reviewed to correct potential errors. We employed this framework on a corpus of 156 instructional materials containing 264 clinical vignettes across 11 medical specialties from two US medical schools. A dataset of 56 clinical vignettes from 25 curricular material units was annotated by two clinicians across 16 demographic categories for in-context learning and model evaluation using F 1 score. Results. Inter-annotator agreement between clinicians was high (average Krippendorff's alpha: 0.83). Clinical vignette identification achieved F 1 score of 1.0 and demographic information extraction across all categories achieved average F 1 score of 0.87. Few-shot extraction with autonomous error correction improved performance by 8% compared to zero-shot demographic extraction. The validation stage corrected 27.78% of errors from the initial extraction. Conclusion. A multi-agent LLM framework can effectively automate the identification of vignettes and extraction of demographic information from medical curricular materials. Multi-agent few-shot detection and autonomous correction significantly improves over zero-shot or single-agent methods. The ability to accurately audit large-scale educational corpora provides a scalable solution for medical educators to identify and address demographic representation gaps.
Full text 120,230 characters · extracted from preprint-html · click to expand
Automated Extraction of Patient Demographics from Clinical Vignettes Using a Multi-Agent Large Language Model Framework | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Automated Extraction of Patient Demographics from Clinical Vignettes Using a Multi-Agent Large Language Model Framework Sudeshna Das, Rand Ibrahim, Nkele Davis, Yao Ge, Yuting Guo, Abeed Sarker, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9254050/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract Background. Clinical vignettes are widely used in medical education to teach clinical reasoning, evaluate diagnostic skills, and contextualize patient care. These vignettes often include demographic information such as age, sex, race, ethnicity, and other identity markers that may influence learners’ understanding of patient diversity. Inclusion of underrepresented populations is particularly important when demographic factors may affect diagnosis and care. Prior research shows that demographic details in clinical educational materials may be presented inconsistently or selectively, reflecting elements of hidden curriculum and reinforcing stereotypes. However, manually extracting demographic information from clinical vignettes in curricular materials for audits is resource-intensive, while existing automated approaches show limited reliability. Methods. We propose a multi-agentic framework to automatically extract demographic information from clinical vignettes. Our framework consists of three stages: ( 1 ) a few-shot learning step to identify clinical vignettes, ( 2 ) a few-shot infused demographic information extraction step in which multiple parallel agents extract demographic data from identified vignettes; and ( 3 ) an LLM-as-a-judge validation step in which extracted information is autonomously reviewed to correct potential errors. We employed this framework on a corpus of 156 instructional materials containing 264 clinical vignettes across 11 medical specialties from two US medical schools. A dataset of 56 clinical vignettes from 25 curricular material units was annotated by two clinicians across 16 demographic categories for in-context learning and model evaluation using F 1 score. Results. Inter-annotator agreement between clinicians was high (average Krippendorff's alpha: 0.83). Clinical vignette identification achieved F 1 score of 1.0 and demographic information extraction across all categories achieved average F 1 score of 0.87. Few-shot extraction with autonomous error correction improved performance by 8% compared to zero-shot demographic extraction. The validation stage corrected 27.78% of errors from the initial extraction. Conclusion. A multi-agent LLM framework can effectively automate the identification of vignettes and extraction of demographic information from medical curricular materials. Multi-agent few-shot detection and autonomous correction significantly improves over zero-shot or single-agent methods. The ability to accurately audit large-scale educational corpora provides a scalable solution for medical educators to identify and address demographic representation gaps. Demographic extraction Clinical vignettes Large Language Models Multi-agent Artificial Intelligence Figures Figure 1 Figure 2 Background Clinical vignettes are a foundational component of medical education, widely used to teach clinical reasoning, assess diagnostic skills, and situate biomedical knowledge within patient-centered narratives ( 1 , 2 ). These narratives often include demographic descriptors such as age, sex, race, ethnicity, and other socioeconomic descriptors. When used appropriately, such information helps learners understand disease epidemiology, recognize health disparities, and develop skills for providing care to diverse patient populations, including those from underrepresented populations. At the same time, the inclusion of demographic characteristics may contribute to unintended educational consequences. Prior work has shown that demographic details are sometimes presented selectively or in association with specific conditions, potentially reinforcing stereotypes ( 3 ). This phenomenon has been described as part of the medical education “hidden curriculum” in which implicit messages shape learners’ assumptions about patients and disease ( 4 ). Ensuring that demographic information is included thoughtfully, consistently, and representatively is, therefore, an important consideration for curricular quality. Several studies have documented variability and imbalance in the representation of race, ethnicity, gender, and social determinants of health within clinical case materials ( 5 , 6 ). However, systematic evaluation of demographic representation at the curriculum level remains difficult. Medical school curricula consist of thousands of lecture slides, case discussions, and assessment items distributed across multiple courses and years. Manual review of such large and heterogeneous educational corpora is resource-intensive and difficult to sustain, limiting institutions’ ability to conduct comprehensive or longitudinal audits of representation. Automated text analysis offers a potential solution, but existing approaches face important limitations. Traditional natural language processing methods often struggle with the diverse and unstructured formats of educational materials. More recently, large language models (LLMs) have demonstrated strong performance in information extraction tasks; however, single-agent and zero-shot approaches may produce inconsistent results, including hallucinated ( 7 ) or inferred demographic attributes when information is ambiguous or incomplete. These reliability concerns are particularly salient when extracting sensitive identity-related characteristics. Multi-agent and iterative inference architectures have been proposed as strategies to improve extraction accuracy and reduce error propagation. In such frameworks, specialized agents perform targeted tasks in parallel ( 8 ), while secondary review stages aid in validation and correction. This approach allows high-granularity extraction across multiple dimensions while reducing the risks associated with single-pass inference ( 9 ). To address the challenges associated with demographic extraction from clinical vignettes at scale, we developed and evaluated a multi-agent LLM pipeline for the (i) automated identification of clinical vignettes and (ii) extraction of demographic information across 16 dimensions. Each demographic attribute was extracted by a dedicated agent, followed by a correction step performed by a higher-capacity model. We applied this framework to a corpus of 156 lecture slide decks and other related curricular materials from two medical schools in the United States (US) and evaluated performance using a clinician-annotated gold standard dataset. Our objective was to assess the viability of this approach in providing a reliable and scalable method to support large-scale audits of demographic representation in medical education materials. Methods Study Aim, Design & Setting The aim of this study was to develop and evaluate a multi-agentic LLM-based framework to identify clinical vignettes and extract the associated demographic characteristics from medical school curricular materials. This study was designed as a multi-site computational validation study in which automatically extracted information was evaluated against a clinician-annotated gold standard dataset. The research was conducted as a multi-site collaboration involving two US medical schools, referred to as Site A and Site B to maintain institutional anonymity. These institutions represent distinct medical education programs located in two different US states. The primary objective was to overcome the resource-intensive limitations of manual curricular review by implementing a scalable, multi-agent LLM framework capable of extracting clinical vignettes and associated patient demographic characteristics. Data Collection & Annotation Instructional materials, including Powerpoint slide decks, assigned readings, and assessment materials were contributed by faculty of two medical schools, each located in a different US state. A total of 156 unique files were collected, across clinical tracks and pre-clinical courses, including advanced electives and breadth courses from the medical education curriculum. Courses represented in the data were Cardiology, Global Health and Underserved Populations (GHUP), Neurology, Immunology, Hematology, Obstetrics and Gynecology (OB-GYN) Clerkship, Pulmonology, Endocrinology, Emergency Medicine (EM) Clerkship, Aging and Dying, and Reproductive Physiology. Courses with fewer than 10 file contributions were grouped together as ‘Others’ (Fig. 1 ). Instructional materials for subjects with greater than 10 files was contributed by multiple faculty members. All instructional materials were converted to machine-readable plain text format prior to analysis using automated text extraction. We created a gold standard annotated dataset from 25 randomly selected files. Files were randomly selected using a computer-generated stratified random sampling procedure, stratified by site to ensure representation from both institutions. Annotation was carried out in three rounds by three clinicians with advanced medical training. In the first stage, two annotators were tasked with: 1. Identifying clinical vignettes within the instructional material 2. Extracting demographic characteristics for each identified clinical vignette across 16 pre-defined demographic dimensions. A clinical vignette was defined as a narrative describing a specific patient with identifiable clinical context, including presentation, history, findings, or management. Following this initial annotation round, annotation feedback was collected and the annotation guidelines were refined to produce the final annotation protocol (Appendix 1). In the second round of annotation, the two annotators re-annotated vignettes where disagreement had been noted during the first stage. Any remaining cases of disagreement were resolved through adjudication by a third clinician annotator, resulting in the final gold-standard labelled dataset. Automated Demographic Extraction System We developed a multi-agent LLM framework for automated demographic extraction from clinical vignettes. The system consists of three sequential modules: (i) clinical vignette extraction; (ii) demographic information extraction; and (iii) demographic information validation. Each module processed the output of the preceding module. A schematic overview of the framework is shown in Fig. 2 . Clinical Vignette Extraction The clinical vignette extraction module utilizes an LLM-based agentic framework designed to identify vignette narratives within instructional materials. To facilitate the processing of comprehensive slide decks and lengthy documents, we deployed a local instance of an open-weight transformer model with 120 billion parameters (GPT-OSS-120B). Leveraging the 128K-token input context and a 64K-token output capacity ( 10 ) of the base model allowed the agent to process entire slide decks and lengthy instructional materials without context truncation. Demographic Information Extraction We developed a parallelized 16-agent extraction module for demographic profiling. Each agent was prompted to extract one specific demographic attribute from the clinical text. The attributes extracted included age group, sex, gender, race, ethnicity, religion, sexual orientation, socioeconomic status (SES), immigration status, housing status, addiction status, cigarette smoking status, mental health status, suicidality, disability, and obesity. To enable efficient parallel inference while maintaining high extraction accuracy, the module used Meta’s Llama3-70B model ( 11 ) to facilitate rapid, concurrent inference across vignettes. Demographic Information Validation The final stage of the framework involved a validation and correction pass performed by an agent powered by OpenAI’s GPT-OSS-120B model. This stage is conceptually similar to the LLM-as-a-judge paradigm, in which LLMs are used to evaluate the quality or correctness of model-generated outputs ( 12 ). In our framework, this idea is extended by assigning the LLM a corrective role in addition to the evaluative role. The model first audits the previously generated predictions and revises them when it detects inconsistencies. This agent ingests: ( 1 ) the clinical vignette extracted by the Clinical Vignette Extraction module, and ( 2 ) the preliminary demographic information extracted by the Demographic Information Extraction module. The agent is instructed to perform automated error detection and correction by auditing the initial demographics against the source vignette narrative. This step was designed to reduce extraction errors and improve overall prediction reliability. All models were executed using deterministic decoding settings (temperature = 0) to aid reproducibility of model outputs. Few-shot prompting examples were derived from the clinician-annotated gold standard dataset and were kept constant across model runs to maintain consistent inference conditions. The vignettes and demographic attributes extracted by the framework were compared against the clinician-annotated gold standard dataset to evaluate performance. We report the performance of our framework using attribute-level F 1 scores, which represents the harmonic mean of precision (positive predictive value) and recall (sensitivity). 95% confidence intervals for F 1 scores were estimated using nonparametric bootstrapping with repeated resampling of the dataset. Results Data Collection & Annotation Site A contributed 40.38% (N = 63) of the collected instruction material with Site B contributing the remaining 59.62% (N = 93). Instructional material on cardiology had the highest frequency, accounting for 30.77% (N = 48) of the total files, followed by GHUP at 17.31% (N = 27) and Neurology at 16.03% (N = 25). Perfect inter-annotator agreement (IAA) was observed between the annotators for clinical vignette extraction. For demographic information extraction, the IAA ranged from 0.65 (low reliability) to 1.0 (perfect agreement), as measured by Krippendorff’s alpha, to account for three annotators ( 13 ). The mean IAA shows moderate agreement (α = 0.83) (Table 1 ). Table 1 Krippendorff’s alpha demonstrating inter-annotator agreement for all 16 demographic attributes across all annotated clinical vignettes (N = 56). Mean alpha is seen to be 0.83 (moderate agreement). Dimension Krippendorff's Alpha Age 0.82 Sex 0.86 Gender 1.00 Race 0.66 Ethnicity 1.00 Religion 1.00 Sexual orientation 1.00 Socioeconomic status 0.66 Immigration status 1.00 Housing 0.65 Addiction 0.66 Cigarette smoking 0.66 Mental health 1.00 Suicidality 0.66 Disability 0.80 Obesity 0.79 Mean (SD) 0.83 (± 0.15) Automated Demographic Extraction System Clinical Vignette Extraction Of the 47 clinical vignettes identified by the clinicians, our clinical vignette extraction agent was able to identify all 47. No false positives were observed on the gold standard dataset, suggesting strong performance under evaluation conditions. A total of 5 files were marked by clinician annotators as not containing any clinical vignettes. The clinical vignette extraction agent was also able to identify these true negatives, achieving an overall F 1 score of 1.0. Examples of clinical vignettes extracted by the framework are provided in Appendix 3. Demographic Information Extraction For the demographic attributes of Gender , Ethnicity , Religion , Sexual orientation , Immigration status , and Mental health , the demographic extraction module achieved an F 1 score of 1.0. For the remaining ten demographic attributes, F 1 scores were in the range 0.66–0.85 (Table 2 ). Demographic Information Correction The demographic information correction module corrected a total of 27.8% of the errors made by the demographic information extraction module. Improvements were noted for the demographic attributes of Age (F 1 score increased from 0.83 to 0.85), Sex (F 1 score increased from 0.78 to 0.80), Socioeconomic status (F 1 score increased from 0.74 to 0.80), Cigarette smoking (F 1 score increased from 0.66 to 0.75), and Suicidality (F 1 score increased from 0.71 to 0.81). For the demographic attributes of Race , Housing , Addiction , Disability , and Obesity , F 1 scores remained the same after passing through the correction module. Demographic attributes which had achieved an F 1 score of 1.0 ( Gender , Ethnicity , Religion , Sexual orientation , Immigration status , and Mental health ) also remained unchanged. Zero-shot extraction achieved a mean F 1 score of 0.80, demonstrating the incremental benefit of few-shot prompting and iterative correction. Table 2 Automated demographic extraction system performance, as observed via F 1 scores. Our proposed method exhibits the best performance, with a mean F 1 score of 0.87 across all 16 demographic dimensions. Dimension F 1 score Zero-shot Learning Few-shot Learning Few-shot Learning + Validation Age 0.75; 95% CI [0.67, 0.80] 0.83; 95% CI [0.78, 0.87] 0.85; 95% CI [0.79, 0.90] Sex 0.66; 95% CI [0.57, 0.75] 0.78; 95% CI [0.69, 0.87] 0.80; 95% CI [0.71, 0.89] Gender 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] Race 0.66; 95% CI [0.58, 0.74] 0.66; 95% CI [0.59, 0.73] 0.66; 95% CI [0.58, 0.74] Ethnicity 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] Religion 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] Sexual orientation 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] Socioeconomic status 0.74; 95% CI [0.65, 0.83] 0.74; 95% CI [0.67, 0.81] 0.80; 95% CI [0.73, 0.87] Immigration status 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] Housing 0.78; 95% CI [0.71, 0.84] 0.80; 95% CI [0.71, 0.88] 0.80; 95% CI [0.71, 0.87] Addiction 0.85; 95% CI [0.80, 0.89] 0.85; 95% CI [0.78, 0.91] 0.85; 95% CI [0.79, 0.90] Cigarette smoking 0.00; 95% CI [0.00, 0.00] 0.66; 95% CI [0.60, 0.73] 0.75; 95% CI [0.68, 0.83] Mental health 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] 1.00; 95% CI [1.0, 1.0] Suicidality 0.71; 95% CI [0.62, 0.81] 0.71; 95% CI [0.63, 0.79] 0.81; 95% CI [0.73, 0.90] Disability 0.84; 95% CI [0.76, 0.91] 0.84; 95% CI [0.77, 0.91] 0.84; 95% CI [0.74, 0.90] Obesity 0.80; 95% CI [0.72, 0.87] 0.80; 95% CI [0.72, 0.89] 0.80; 95% CI [0.73, 0.87] Mean (SD) 0.80 ( ± 0.25 ) 0.85 ( ± 0.13 ) 0.87 (± 0.11) Demographic Distribution A total of 264 clinical vignettes were identified across 106 instructional files, while 50 files contained no clinical vignettes. Site B contributed a larger proportion of vignettes (66.3%, n = 175) compared with Site A (33.7%, n = 89), despite contributing fewer files proportionally. Across specialties, Neurology had the highest number of clinical vignettes (n = 77; 29.1%), followed by Cardiology (n = 65; 24.6%) and GHUP (n = 29; 11.0%). The number of vignettes per file varied by specialty, with Hematology, Neurology, and Immunology averaging approximately 3–4 vignettes per file, whereas OB-GYN materials contained fewer cases per file. The clinical vignette extraction module achieved perfect agreement with clinician annotations (F 1 = 1.0). For demographic extraction, the initial few-shot multi-agent approach achieved a mean F 1 score of 0.85 across 16 attributes. Following the correction stage, overall performance improved to a mean F 1 score of 0.87. The validation module aided in resolving initial errors, with the largest improvements observed for cigarette smoking, suicidality, and socioeconomic status. Attribute-level performance varied, with perfect F 1 scores for gender, ethnicity, religion, sexual orientation, immigration status, and mental health, whereas race and socioeconomic status remained more challenging. Across the corpus, most demographic attributes were unspecified, with “Unknown” labels exceeding 90% for race, ethnicity, religion, and socioeconomic status (Table 3 ). Table 3 Distribution of explicitly reported vs unknown demographic attributes across clinical vignettes (N = 264) Demographic Attribute Explicitly reported (%) Unknown (%) Age 233 (88.3%) 31 (11.7%) Sex 233 (88.3%) 31 (11.7%) Gender 7 (2.6%) 257 (97.4%) Race 11 (4.2%) 253 (95.8%) Ethnicity 14 (5.3%) 250 (94.7%) Religion 0 (0%) 264 (100%) Sexual orientation 16 (6.1%) 248 (93.9%) SES 12 (4.5%) 252 (95.5%) Immigration 6 (2.3%) 258 (97.7%) Housing 19 (7.2%) 245 (92.8%) Addiction 10 (3.8%) 254 (96.2%) Cigarette smoking 16 (6.1%) 248 (93.9%) Mental health 15 (5.7%) 249 (94.3%) Suicidality 1 (0.4%) 263 (99.6%) Disability 51 (19.3%) 213 (80.7%) Obesity 8 (3.0%) 256 (97.0%) Discussion Principal Findings This study demonstrates the feasibility of using a multi-agent LLM framework to automatically identify clinical vignettes and extract demographic information from large volumes of medical education materials, also showing the utility of LLMs in medical education ( 14 , 15 ). The system achieved near-perfect performance for vignette identification and strong overall performance for demographic information extraction, with measurable improvements over both zero-shot and few-shot approaches. The addition of a correction stage resolved approximately 27.8% of initial extraction errors, suggesting that iterative review mechanisms may improve the reliability of automated information extraction systems ( 9 , 16 ). Attributes with perfect F 1 scores were characterized by low prevalence and predominantly ‘Unknown’ labels, which may partially contribute to high agreement due to class imbalance. Importantly, the framework was designed to extract only explicitly stated demographic information and to avoid inference of protected characteristics. This design choice prioritizes precision and may reduce the risk of introducing algorithmic bias into curricular analyses. During annotation, reviewers occasionally inferred characteristics that were not explicitly stated in the vignette text. For example, diagnoses such as dementia were sometimes interpreted as implying disability status, even when disability was not directly mentioned. While such inferences may be reasonable for experienced clinicians, the framework was designed to reflect the perspective of the learner, who may not automatically infer these relationships. This distinction highlights a potential source of bias in expert annotation and suggests that there may be differences between what educators intend to teach and what learners explicitly encounter in curricular materials ( 4 ). Beyond technical performance, the extracted corpus provides insight into patterns of demographic representation in educational materials. Most demographic dimensions were marked as ‘Unknown’ in the majority of vignettes, particularly for race, ethnicity, religion, socioeconomic status, immigration status, and sexual orientation. This suggests that many social and identity characteristics are either rarely included or selectively reported, consistent with prior research demonstrating that demographic context and social determinants of health are unevenly represented in medical education materials ( 5 , 17 ). Age and binary sex were the most frequently specified characteristics, reflecting their traditional role in clinical reasoning ( 18 ). In contrast, non-binary gender identities were not represented, and explicit race and ethnicity were rarely stated. These patterns may reflect efforts to avoid stereotyping, but they may also indicate missed opportunities to contextualize health disparities and social determinants of health within clinical teaching cases, which has been emphasized as an important component of contemporary medical education ( 18 , 19 ). The high performance observed for clinical vignette identification likely reflects the structured narrative characteristics of vignette text, which differ from surrounding explanatory or didactic content. The modular separation of vignette detection from demographic extraction likely reduced task complexity at each stage, and may have contributed to overall system stability ( 20 ). The validation module played a critical role in improving performance. By re-evaluating initial outputs against the source text, the system reduced inconsistencies and mitigated reasoning errors that have been reported in single-pass LLM pipelines. This iterative validation step accounted for a 2.47% improvement over standard few-shot extraction. Clinician annotation provided a reliable reference standard, with a mean Krippendorff’s alpha of 0.83, indicating moderate to high agreement. Publicly available datasets for demographic extraction from medical education materials are currently limited, particularly for tasks involving structured identification of clinical vignettes and associated demographic attributes. As a result, the creation of the clinician-annotated dataset in this study can aid further research. Performance variability across demographic attributes likely reflects differences in how such information is expressed in educational materials ( 6 , 21 ). Attributes such as gender identity or immigration status, when present, were typically stated explicitly, whereas race, socioeconomic status, and housing context were often implicit or omitted. From a systems perspective, the use of a high-parameter model for context-intensive tasks alongside smaller parallel extraction agents allowed a balance between computational efficiency and accuracy. Future work could extend this framework to support intersectional analyses of demographic representation and explore real-time integration within curriculum development workflows, enabling faculty to receive automated feedback during content creation. Error Analysis Error analysis suggests that most errors occurred in attributes requiring interpretation of nuanced language rather than direct extraction (e.g., socioeconomic status). Attributes with sparse positive cases (e.g., suicidality, addiction, housing insecurity) showed greater performance variability due to class imbalance. The correction module produced the largest gains in these low-frequency categories. A small number of structural challenges were observed. One vignette presented multiple patients in tabular format, which required segmentation beyond standard narrative extraction. Additionally, approximately ten potential false-positive vignette candidates were identified in the broader corpus, typically where clinical descriptions were embedded within explanatory text rather than presented as standalone cases. No false positives were observed within the gold standard dataset, providing additional support for the robustness of the vignette identification module under evaluation conditions used in this study. Implications for Medical Education The ability to automatically extract demographic information from curricular materials has several potential applications for medical education research and curriculum oversight. First, it enables large-scale curricular audits to identify gaps or imbalances and unintentional omissions in demographic representation across courses, sites, or specialties. Such analyses can inform targeted instructional revisions to ensure that learners are exposed to diverse patient populations. Because many educational materials are developed by individual instructors over time, course directors may have limited visibility into the full set of cases presented across a curriculum. Tools that enable systematic review of case content may therefore help ensure that core topics and patient populations remain consistently represented. For example, our analysis identified no clinical cases involving menopause within the OB-GYN Clerkship curriculum. While omissions may reflect local curricular design choices, they may also occur with changes in lecture or course leadership. This example illustrates how automated case auditing can surface potential gaps in topic coverage that might otherwise go unnoticed. Omission of either common or high acuity-low frequency presentations from clerkship-level cases may contribute to gaps in learners' clinical pattern recognition, underscoring the downstream educational relevance of such findings. Second, automated monitoring may help identify patterned associations between demographic characteristics and specific conditions, allowing educators to detect potential reinforcement of stereotypes or implicit bias ( 6 , 22 ). This capability is particularly relevant given prior evidence that demographic details in clinical materials are often presented selectively, contributing to the hidden curriculum through which learners develop implicit assumptions about patients and disease ( 3 , 4 ). Identifying these patterns at scale is a necessary first step toward intentional curricular revision. Third, the framework supports longitudinal tracking, enabling institutions to evaluate the impact of educational initiatives over time and to monitor whether revisions to instructional materials produce measurable changes in representation. Fourth, automated demographic profiling can serve as a concrete, data-driven foundation for faculty development. Rather than relying on abstract discussions of representation in case design, course directors and instructors can engage with evidence drawn directly from their own curricular materials, supporting more targeted and actionable conversations about equitable vignette construction. Finally, this approach provides a scalable pathway to evidence-based curricular quality assurance that was not previously achievable through manual review alone. Manual audits are time-intensive, difficult to sustain across large and evolving educational repositories, and limited in their capacity to detect aggregate patterns. By enabling systematic, reproducible analysis at the curriculum level, this framework supports a mode of oversight that complements rather than replaces the judgment of faculty and enables full curricular evaluation. Limitations Although our proposed framework performs well, there are some limitations of this study. First, the study included materials from only two US medical schools. Curricular design, institutional priorities, and regional patient populations may influence demographic representation, limiting generalizability. Future validation across larger and more diverse corpora is warranted to assess generalizability. Second, the analysis was restricted to instructional materials and does not capture demographic exposure that occurs during clinical training, clerkships, or residency. Third, many demographic attributes were rare or absent in the source materials, resulting in class imbalance that may inflate performance estimates for categories dominated by “Unknown.” Finally, some demographic domains were limited to single-category extraction (e.g., one mental health condition), which may not reflect the complexity of comorbidity or multidimensional social factors. Conclusion This study demonstrates that a multi-agent LLM framework can reliably identify clinical vignettes and extract demographic information from medical education materials. These findings indicate that automated methods can be used to systematically audit demographic representation in curricular content. Beyond methodological contribution, the approach enables large-scale analysis of demographic representation across curricular repositories, supporting educators in identifying gaps, monitoring trends, and promoting more balanced educational materials. As medical education increasingly emphasizes equity, social determinants of health, and bias mitigation, automated auditing frameworks offer a practical mechanism to support these priorities. Future work should extend validation across additional institutions and explore integration of such tools into routine curriculum review processes. The automated pipeline demonstrates a scalable approach for evaluating whether clinical vignettes reflect the diversity of patient populations, with potential to support the development of more equitable clinical reasoning and diagnostic training in future physicians. Declarations Ethics approval and consent to participate This study did not include identifiable patient data or involve human subjects research, and was therefore exempt from IRB review by Emory University. Consent for publication Not applicable. Funding Research reported in this publication was supported by Emory Women's Impact Circle (WIC), Society for Academic Emergency Medicine (SAEM) and Rymer Innovation Grant. The content is solely the responsibility of the authors and does not necessarily represent the official views of the funding agencies. Author Contribution SD designed the methodology, conducted the analyses, and drafted the manuscript. RI and ND performed data annotation. YGe and YGuo contributed to data curation and analysis. AS contributed to methodology and supervised the study. MR conceptualized and supervised the study. All authors contributed to, reviewed and approved the final manuscript. Acknowledgement The authors thank the faculty members who contributed instructional materials for this study. Data Availability Extracted clinical vignettes and demographic information shall be made available upon reasonable request and signing of a data use agreement, for research purposes. References Eva KW. What every teacher needs to know about clinical reasoning. Med Educ. 2005;39(1):98–106. 10.1111/j.1365-2929.2004.01972.x . Plaum P, Visser LN, de Groot B, Morsink MEB, Duijst WLJM, Candel BGJ. Using case vignettes to study the presence of outcome, hindsight, and implicit bias in acute unplanned medical care: a cross-sectional study. Eur J Emerg Med. 2024;31(4):260. 10.1097/MEJ.0000000000001127 . Chapman EN, Kaatz A, Carnes M. Physicians and Implicit Bias: How Doctors May Unwittingly Perpetuate Health Care Disparities. J GEN INTERN MED. 2013;28(11):1504–10. 10.1007/s11606-013-2441-1 . Hafferty FW. Beyond curriculum reform: confronting medicine’s hidden curriculum. Acad Med. 1998;73(4):403–7. 10.1097/00001888-199804000-00013 . Lee CR, Gilliland KO, Beck Dallaghan GL, Tolleson-Rinehart S. Race, ethnicity, and gender representation in clinical case vignettes: a 20-year comparison between two institutions. BMC Med Educ. 2022;22(1):585. 10.1186/s12909-022-03665-4 . Tsai J, Ucik L, Baldwin N, Hasslinger C, George P. Race Matters? Examining and Rethinking Race Portrayal in Preclinical Medical Education. Acad Med. 2016;91(7):916–20. 10.1097/ACM.0000000000001232 . Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of Hallucination in Natural Language Generation. ACM Comput Surv. 2023;55(12):2481–248. 10.1145/3571730 . Park JS, O’Brien J, Cai CJ, Morris MR, Liang P, Bernstein MS. Generative Agents: Interactive Simulacra of Human Behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology [Internet]. New York, NY, USA: Association for Computing Machinery; 2023 [cited 2026 Feb 19]. pp. 1–22. (UIST ’23). Available from: https://dl.acm.org/doi/10.1145/3586183.3606763 doi:10.1145/3586183.3606763 Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, et al. Self-refine: Iterative refinement with self-feedback. Adv Neural Inf Process Syst. 2023;36:46534–94. OpenAI, Agarwal S, Ahmad L, Ai J, Altman S, Applebaum A et al. gpt-oss-120b & gpt-oss-20b Model Card [Internet]. arXiv; 2025 [cited 2026 Mar 10]. Available from: http://arxiv.org/abs/2508.10925 10.48550/arXiv.2508.10925 Grattafiori A, Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A et al. The Llama 3 herd of models. In: Neural Information Processing Systems [Internet]. Curran Associates; 2024 [cited 2026 Mar 10]. Available from: https://hal.science/hal-05412319/ Li J, Sun S, Yuan W, Fan RZ, Liu P. Generative Judge for Evaluating Alignment. In: The Twelfth International Conference on Learning Representations. 2024. Krippendorff K. Computing Krippendorff’s alpha-reliability [Internet]. 2011 [cited 2026 Mar 10]. Available from: https://repository.upenn.edu/handle/20.500.14332/2089 Abd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large language models in medical education: opportunities, challenges, and future directions. JMIR Med Educ. 2023;9(1):e48291. Vrdoljak J, Boban Z, Vilović M, Kumrić M, Božić J. A review of large language models in medical education, clinical decision support, and healthcare administration. In: Healthcare [Internet]. MDPI; 2025 [cited 2026 Mar 10]. p. 603. Available from: https://www.mdpi.com/2227-9032/13/6/603 Weng Y, Zhu M, Xia F, Li B, He S, Liu S et al. Large language models are better reasoners with self-verification. In: Findings of the Association for Computational Linguistics: EMNLP 2023 [Internet]. 2023 [cited 2026 Mar 10]. pp. 2550–75. Available from: https://aclanthology.org/2023.findings-emnlp.167/ Louie P, Wilkes R. Representations of race and skin tone in medical textbook imagery. Soc Sci Med. 2018;202:38–42. Ross PT, Hart-Johnson T, Santen SA, Zaidi NLB. Considerations for using race and ethnicity as quantitative variables in medical education research. Perspect Med Educ. 2020;9(5):318–23. 10.1007/S40037-020-00602-3 . Krishnan A, Rabinowitz M, Ziminsky A, Scott SM, Chretien KC. Addressing race, culture, and structural inequality in medical education: a guide for revising teaching cases. Acad Med. 2019;94(4):550–5. Khot T, Trivedi H, Finlayson M, Fu Y, Richardson K, Clark P et al. Decomposed Prompting: A Modular Approach for Solving Complex Tasks. In: The Eleventh International Conference on Learning Representations [Internet]. [cited 2026 Mar 10]. Available from: https://openreview.net/forum?id=_nGgzQjzaRy Carey-Ewend K, Feinberg A, Flen A, Williamson C, Gutierrez C, Cykert S, et al. Use of Sociodemographic Information in Clinical Vignettes of Multiple-Choice Questions for Preclinical Medical Students. MedSciEduc. 2023;33(3):659–67. 10.1007/s40670-023-01778-z . Amutah C, Greenidge K, Mante A, Munyikwa M, Surya SL, Higginbotham E Misrepresenting Race — The Role of Medical Schools in Propagating Physician Bias. New England Journal of Medicine., Amutah C, Greenidge K, Mante A, Munyikwa M, Surya SL, Higginbotham E et al. Misrepresenting Race — The Role of Medical Schools in Propagating Physician Bias. New England Journal of Medicine. 2021 Mar 3;384(9):872–8. doi:10.1056/NEJMms2025768. Additional Declarations No competing interests reported. Supplementary Files bmcmesupp.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 17 May, 2026 Reviewers agreed at journal 04 May, 2026 Reviewers invited by journal 24 Apr, 2026 Editor invited by journal 04 Apr, 2026 Editor assigned by journal 02 Apr, 2026 Submission checks completed at journal 02 Apr, 2026 First submitted to journal 28 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9254050","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":633979963,"identity":"17caf1df-0a92-4e8d-80e8-92ea7716d7ae","order_by":0,"name":"Sudeshna Das","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9UlEQVRIiWNgGAWjYHACAwbGBhCdwMDwASbGQ6wWxhkka2GGq8SnRX5G8saHP3fckWdgT3782aZim5z5jATGB2/b8FhxI63YmPfMM8MGnmdm0jlnbhvL3EhgNpyLT4tEjpk0Y9thxgaJBDPm3LbbiTMkEtikefFokZ+RY/7zZ9th+waJ9M+fLf+BtbD/xqeF4UaOGQNv2+HEBokcA2nGBogtzPi0GJx5Vgx0xuHkNp43ZZI9x24bS/A8bJaccw6Pw9qTN34EOsy2nz1984cfNbflJNiTD354U4bHYTDABmcJJDYQoR4F8B8gVccoGAWjYBQMcwAAq9hTrFRgwHgAAAAASUVORK5CYII=","orcid":"","institution":"Emory University","correspondingAuthor":true,"prefix":"","firstName":"Sudeshna","middleName":"","lastName":"Das","suffix":""},{"id":633979964,"identity":"344156c1-92c8-42c3-be57-417bee3fdd68","order_by":1,"name":"Rand Ibrahim","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Rand","middleName":"","lastName":"Ibrahim","suffix":""},{"id":633979965,"identity":"add9ead3-bffe-4ba4-a79b-bfadc3a58142","order_by":2,"name":"Nkele Davis","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Nkele","middleName":"","lastName":"Davis","suffix":""},{"id":633979977,"identity":"f615fdaf-4255-4fd9-9277-216a19905df8","order_by":3,"name":"Yao Ge","email":"","orcid":"","institution":"National Institutes of Health","correspondingAuthor":false,"prefix":"","firstName":"Yao","middleName":"","lastName":"Ge","suffix":""},{"id":633979979,"identity":"2b3404f0-0575-4d94-805b-9225efe7b222","order_by":4,"name":"Yuting Guo","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Yuting","middleName":"","lastName":"Guo","suffix":""},{"id":633979980,"identity":"5777c112-67a0-4625-9138-d81fa0695458","order_by":5,"name":"Abeed Sarker","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Abeed","middleName":"","lastName":"Sarker","suffix":""},{"id":633979981,"identity":"8eb6ba72-5f9e-4def-8cd2-b12b40f7b964","order_by":6,"name":"Marta Rowh","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Marta","middleName":"","lastName":"Rowh","suffix":""}],"badges":[],"createdAt":"2026-03-28 16:08:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9254050/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9254050/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108819952,"identity":"57314e8a-ff2f-4439-b9d9-5c47befed211","added_by":"auto","created_at":"2026-05-08 16:39:31","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":336494,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of data by site and topic. (a) Frequency of files (b) Frequency of clinical vignettes extracted. ‘\u003cem\u003eOthers\u003c/em\u003e’ comprises ‘\u003cem\u003eEndocrine\u003c/em\u003e’, ‘\u003cem\u003eEM Clerkship\u003c/em\u003e’, ‘\u003cem\u003eAging and Dying\u003c/em\u003e’, and ‘\u003cem\u003eReproductive Physiology\u003c/em\u003e’.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-9254050/v1/a7fb301a4cd443b5c77b9c33.jpeg"},{"id":108820119,"identity":"afbf0266-424a-4408-a141-253a6551763d","added_by":"auto","created_at":"2026-05-08 16:40:03","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":481478,"visible":true,"origin":"","legend":"\u003cp\u003eMethodology used for demographic information extraction from medical school curriculum. The first stage consists of data collection and annotation. The second stage shows the automated demographic extraction system; a multi-agent framework with three modules: clinical vignette extraction, demographic information extraction, and demographic information validation.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-9254050/v1/3241b2560809ef584e4ba82c.png"},{"id":108822566,"identity":"2d3a44d2-7804-4962-b568-24108401a25e","added_by":"auto","created_at":"2026-05-08 16:49:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":980017,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9254050/v1/014a02b2-6c25-47b2-b05a-6fe91d1bf6c0.pdf"},{"id":108819960,"identity":"c30137b1-3a6a-4ca7-ba20-a72145730025","added_by":"auto","created_at":"2026-05-08 16:39:34","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":20879,"visible":true,"origin":"","legend":"","description":"","filename":"bmcmesupp.docx","url":"https://assets-eu.researchsquare.com/files/rs-9254050/v1/1fb9bb4776a1ba83b60ab427.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Automated Extraction of Patient Demographics from Clinical Vignettes Using a Multi-Agent Large Language Model Framework","fulltext":[{"header":"Background","content":"\u003cp\u003eClinical vignettes are a foundational component of medical education, widely used to teach clinical reasoning, assess diagnostic skills, and situate biomedical knowledge within patient-centered narratives (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). These narratives often include demographic descriptors such as age, sex, race, ethnicity, and other socioeconomic descriptors. When used appropriately, such information helps learners understand disease epidemiology, recognize health disparities, and develop skills for providing care to diverse patient populations, including those from underrepresented populations.\u003c/p\u003e \u003cp\u003eAt the same time, the inclusion of demographic characteristics may contribute to unintended educational consequences. Prior work has shown that demographic details are sometimes presented selectively or in association with specific conditions, potentially reinforcing stereotypes (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). This phenomenon has been described as part of the medical education \u0026ldquo;hidden curriculum\u0026rdquo; in which implicit messages shape learners\u0026rsquo; assumptions about patients and disease (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). Ensuring that demographic information is included thoughtfully, consistently, and representatively is, therefore, an important consideration for curricular quality.\u003c/p\u003e \u003cp\u003eSeveral studies have documented variability and imbalance in the representation of race, ethnicity, gender, and social determinants of health within clinical case materials (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e). However, systematic evaluation of demographic representation at the curriculum level remains difficult. Medical school curricula consist of thousands of lecture slides, case discussions, and assessment items distributed across multiple courses and years. Manual review of such large and heterogeneous educational corpora is resource-intensive and difficult to sustain, limiting institutions\u0026rsquo; ability to conduct comprehensive or longitudinal audits of representation.\u003c/p\u003e \u003cp\u003eAutomated text analysis offers a potential solution, but existing approaches face important limitations. Traditional natural language processing methods often struggle with the diverse and unstructured formats of educational materials. More recently, large language models (LLMs) have demonstrated strong performance in information extraction tasks; however, single-agent and zero-shot approaches may produce inconsistent results, including hallucinated (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e) or inferred demographic attributes when information is ambiguous or incomplete. These reliability concerns are particularly salient when extracting sensitive identity-related characteristics. Multi-agent and iterative inference architectures have been proposed as strategies to improve extraction accuracy and reduce error propagation. In such frameworks, specialized agents perform targeted tasks in parallel (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e), while secondary review stages aid in validation and correction. This approach allows high-granularity extraction across multiple dimensions while reducing the risks associated with single-pass inference (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo address the challenges associated with demographic extraction from clinical vignettes at scale, we developed and evaluated a multi-agent LLM pipeline for the (i) automated identification of clinical vignettes and (ii) extraction of demographic information across 16 dimensions. Each demographic attribute was extracted by a dedicated agent, followed by a correction step performed by a higher-capacity model. We applied this framework to a corpus of 156 lecture slide decks and other related curricular materials from two medical schools in the United States (US) and evaluated performance using a clinician-annotated gold standard dataset. Our objective was to assess the viability of this approach in providing a reliable and scalable method to support large-scale audits of demographic representation in medical education materials.\u003c/p\u003e"},{"header":"Methods","content":" \u003cp\u003eStudy Aim, Design \u0026amp; Setting\u003c/p\u003e \u003cp\u003eThe aim of this study was to develop and evaluate a multi-agentic LLM-based framework to identify clinical vignettes and extract the associated demographic characteristics from medical school curricular materials. This study was designed as a multi-site computational validation study in which automatically extracted information was evaluated against a clinician-annotated gold standard dataset.\u003c/p\u003e \u003cp\u003eThe research was conducted as a multi-site collaboration involving two US medical schools, referred to as Site A and Site B to maintain institutional anonymity. These institutions represent distinct medical education programs located in two different US states. The primary objective was to overcome the resource-intensive limitations of manual curricular review by implementing a scalable, multi-agent LLM framework capable of extracting clinical vignettes and associated patient demographic characteristics.\u003c/p\u003e \u003cp\u003eData Collection \u0026amp; Annotation\u003c/p\u003e \u003cp\u003eInstructional materials, including Powerpoint slide decks, assigned readings, and assessment materials were contributed by faculty of two medical schools, each located in a different US state. A total of 156 unique files were collected, across clinical tracks and pre-clinical courses, including advanced electives and breadth courses from the medical education curriculum. Courses represented in the data were Cardiology, Global Health and Underserved Populations (GHUP), Neurology, Immunology, Hematology, Obstetrics and Gynecology (OB-GYN) Clerkship, Pulmonology, Endocrinology, Emergency Medicine (EM) Clerkship, Aging and Dying, and Reproductive Physiology. Courses with fewer than 10 file contributions were grouped together as \u0026lsquo;Others\u0026rsquo; (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Instructional materials for subjects with greater than 10 files was contributed by multiple faculty members.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAll instructional materials were converted to machine-readable plain text format prior to analysis using automated text extraction.\u003c/p\u003e \u003cp\u003eWe created a gold standard annotated dataset from 25 randomly selected files. Files were randomly selected using a computer-generated stratified random sampling procedure, stratified by site to ensure representation from both institutions. Annotation was carried out in three rounds by three clinicians with advanced medical training. In the first stage, two annotators were tasked with:\u003c/p\u003e \u003cp\u003e1. Identifying clinical vignettes within the instructional material\u003c/p\u003e \u003cp\u003e2. Extracting demographic characteristics for each identified clinical vignette across 16 pre-defined demographic dimensions.\u003c/p\u003e \u003cp\u003eA clinical vignette was defined as a narrative describing a specific patient with identifiable clinical context, including presentation, history, findings, or management.\u003c/p\u003e \u003cp\u003eFollowing this initial annotation round, annotation feedback was collected and the annotation guidelines were refined to produce the final annotation protocol (Appendix 1). In the second round of annotation, the two annotators re-annotated vignettes where disagreement had been noted during the first stage. Any remaining cases of disagreement were resolved through adjudication by a third clinician annotator, resulting in the final gold-standard labelled dataset.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAutomated Demographic Extraction System\u003c/p\u003e \u003cp\u003eWe developed a multi-agent LLM framework for automated demographic extraction from clinical vignettes. The system consists of three sequential modules: (i) clinical vignette extraction; (ii) demographic information extraction; and (iii) demographic information validation. Each module processed the output of the preceding module. A schematic overview of the framework is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eClinical Vignette Extraction\u003c/p\u003e \u003cp\u003eThe clinical vignette extraction module utilizes an LLM-based agentic framework designed to identify vignette narratives within instructional materials. To facilitate the processing of comprehensive slide decks and lengthy documents, we deployed a local instance of an open-weight transformer model with 120\u0026nbsp;billion parameters (GPT-OSS-120B). Leveraging the 128K-token input context and a 64K-token output capacity (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e) of the base model allowed the agent to process entire slide decks and lengthy instructional materials without context truncation.\u003c/p\u003e \u003cp\u003eDemographic Information Extraction\u003c/p\u003e \u003cp\u003eWe developed a parallelized 16-agent extraction module for demographic profiling. Each agent was prompted to extract one specific demographic attribute from the clinical text. The attributes extracted included age group, sex, gender, race, ethnicity, religion, sexual orientation, socioeconomic status (SES), immigration status, housing status, addiction status, cigarette smoking status, mental health status, suicidality, disability, and obesity. To enable efficient parallel inference while maintaining high extraction accuracy, the module used Meta\u0026rsquo;s Llama3-70B model (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e) to facilitate rapid, concurrent inference across vignettes.\u003c/p\u003e \u003cp\u003eDemographic Information Validation\u003c/p\u003e \u003cp\u003eThe final stage of the framework involved a validation and correction pass performed by an agent powered by OpenAI\u0026rsquo;s GPT-OSS-120B model. This stage is conceptually similar to the LLM-as-a-judge paradigm, in which LLMs are used to evaluate the quality or correctness of model-generated outputs (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). In our framework, this idea is extended by assigning the LLM a corrective role in addition to the evaluative role. The model first audits the previously generated predictions and revises them when it detects inconsistencies.\u003c/p\u003e \u003cp\u003eThis agent ingests: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) the clinical vignette extracted by the Clinical Vignette Extraction module, and (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) the preliminary demographic information extracted by the Demographic Information Extraction module. The agent is instructed to perform automated error detection and correction by auditing the initial demographics against the source vignette narrative. This step was designed to reduce extraction errors and improve overall prediction reliability.\u003c/p\u003e \u003cp\u003eAll models were executed using deterministic decoding settings (temperature\u0026thinsp;=\u0026thinsp;0) to aid reproducibility of model outputs. Few-shot prompting examples were derived from the clinician-annotated gold standard dataset and were kept constant across model runs to maintain consistent inference conditions.\u003c/p\u003e \u003cp\u003eThe vignettes and demographic attributes extracted by the framework were compared against the clinician-annotated gold standard dataset to evaluate performance. We report the performance of our framework using attribute-level F\u003csub\u003e1\u003c/sub\u003e scores, which represents the harmonic mean of precision (positive predictive value) and recall (sensitivity). 95% confidence intervals for F\u003csub\u003e1\u003c/sub\u003e scores were estimated using nonparametric bootstrapping with repeated resampling of the dataset.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eData Collection \u0026amp; Annotation\u003c/p\u003e \u003cp\u003eSite A contributed 40.38% (N\u0026thinsp;=\u0026thinsp;63) of the collected instruction material with Site B contributing the remaining 59.62% (N\u0026thinsp;=\u0026thinsp;93). Instructional material on cardiology had the highest frequency, accounting for 30.77% (N\u0026thinsp;=\u0026thinsp;48) of the total files, followed by GHUP at 17.31% (N\u0026thinsp;=\u0026thinsp;27) and Neurology at 16.03% (N\u0026thinsp;=\u0026thinsp;25).\u003c/p\u003e \u003cp\u003ePerfect inter-annotator agreement (IAA) was observed between the annotators for clinical vignette extraction. For demographic information extraction, the IAA ranged from 0.65 (low reliability) to 1.0 (perfect agreement), as measured by Krippendorff\u0026rsquo;s alpha, to account for three annotators (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). The mean IAA shows moderate agreement (α\u0026thinsp;=\u0026thinsp;0.83) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eKrippendorff\u0026rsquo;s alpha demonstrating inter-annotator agreement for all 16 demographic attributes across all annotated clinical vignettes (N\u0026thinsp;=\u0026thinsp;56). Mean alpha is seen to be 0.83 (moderate agreement).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKrippendorff's Alpha\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGender\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRace\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEthnicity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eReligion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSexual orientation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSocioeconomic status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImmigration status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHousing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAddiction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCigarette smoking\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMental health\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSuicidality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDisability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eObesity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMean (SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.83 (\u0026plusmn;\u0026thinsp;0.15)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAutomated Demographic Extraction System\u003c/p\u003e \u003cp\u003eClinical Vignette Extraction\u003c/p\u003e \u003cp\u003eOf the 47 clinical vignettes identified by the clinicians, our clinical vignette extraction agent was able to identify all 47. No false positives were observed on the gold standard dataset, suggesting strong performance under evaluation conditions. A total of 5 files were marked by clinician annotators as not containing any clinical vignettes. The clinical vignette extraction agent was also able to identify these true negatives, achieving an overall F\u003csub\u003e1\u003c/sub\u003e score of 1.0. Examples of clinical vignettes extracted by the framework are provided in Appendix 3.\u003c/p\u003e \u003cp\u003eDemographic Information Extraction\u003c/p\u003e \u003cp\u003eFor the demographic attributes of \u003cem\u003eGender\u003c/em\u003e, \u003cem\u003eEthnicity\u003c/em\u003e, \u003cem\u003eReligion\u003c/em\u003e, \u003cem\u003eSexual orientation\u003c/em\u003e, \u003cem\u003eImmigration status\u003c/em\u003e, and \u003cem\u003eMental health\u003c/em\u003e, the demographic extraction module achieved an F\u003csub\u003e1\u003c/sub\u003e score of 1.0. For the remaining ten demographic attributes, F\u003csub\u003e1\u003c/sub\u003e scores were in the range 0.66\u0026ndash;0.85 (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eDemographic Information Correction\u003c/p\u003e \u003cp\u003eThe demographic information correction module corrected a total of 27.8% of the errors made by the demographic information extraction module. Improvements were noted for the demographic attributes of \u003cem\u003eAge\u003c/em\u003e (F\u003csub\u003e1\u003c/sub\u003e score increased from 0.83 to 0.85), \u003cem\u003eSex\u003c/em\u003e (F\u003csub\u003e1\u003c/sub\u003e score increased from 0.78 to 0.80), \u003cem\u003eSocioeconomic status\u003c/em\u003e (F\u003csub\u003e1\u003c/sub\u003e score increased from 0.74 to 0.80), \u003cem\u003eCigarette smoking\u003c/em\u003e (F\u003csub\u003e1\u003c/sub\u003e score increased from 0.66 to 0.75), and \u003cem\u003eSuicidality\u003c/em\u003e (F\u003csub\u003e1\u003c/sub\u003e score increased from 0.71 to 0.81). For the demographic attributes of \u003cem\u003eRace\u003c/em\u003e, \u003cem\u003eHousing\u003c/em\u003e, \u003cem\u003eAddiction\u003c/em\u003e, \u003cem\u003eDisability\u003c/em\u003e, and \u003cem\u003eObesity\u003c/em\u003e, F\u003csub\u003e1\u003c/sub\u003e scores remained the same after passing through the correction module. Demographic attributes which had achieved an F\u003csub\u003e1\u003c/sub\u003e score of 1.0 (\u003cem\u003eGender\u003c/em\u003e, \u003cem\u003eEthnicity\u003c/em\u003e, \u003cem\u003eReligion\u003c/em\u003e, \u003cem\u003eSexual orientation\u003c/em\u003e, \u003cem\u003eImmigration status\u003c/em\u003e, and \u003cem\u003eMental health\u003c/em\u003e) also remained unchanged. Zero-shot extraction achieved a mean F\u003csub\u003e1\u003c/sub\u003e score of 0.80, demonstrating the incremental benefit of few-shot prompting and iterative correction.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAutomated demographic extraction system performance, as observed via F\u003csub\u003e1\u003c/sub\u003e scores. Our proposed method exhibits the best performance, with a mean F\u003csub\u003e1\u003c/sub\u003e score of 0.87 across all 16 demographic dimensions.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eDimension\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003eF\u003csub\u003e1\u003c/sub\u003e score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eZero-shot Learning\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFew-shot Learning\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFew-shot Learning\u0026thinsp;+\u0026thinsp;Validation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.75; 95% CI [0.67, 0.80]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.83; 95% CI [0.78, 0.87]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.85; 95% CI [0.79, 0.90]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.66; 95% CI [0.57, 0.75]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.78; 95% CI [0.69, 0.87]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.80; 95% CI [0.71, 0.89]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGender\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRace\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.66; 95% CI [0.58, 0.74]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.66; 95% CI [0.59, 0.73]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.66; 95% CI [0.58, 0.74]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEthnicity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eReligion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSexual orientation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSocioeconomic status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.74; 95% CI [0.65, 0.83]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.74; 95% CI [0.67, 0.81]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.80; 95% CI [0.73, 0.87]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImmigration status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHousing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.78; 95% CI [0.71, 0.84]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.80; 95% CI [0.71, 0.88]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.80; 95% CI [0.71, 0.87]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAddiction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.85; 95% CI [0.80, 0.89]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.85; 95% CI [0.78, 0.91]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.85; 95% CI [0.79, 0.90]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCigarette smoking\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.00; 95% CI [0.00, 0.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.66; 95% CI [0.60, 0.73]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.75; 95% CI [0.68, 0.83]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMental health\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.00; 95% CI [1.0, 1.0]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSuicidality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.71; 95% CI [0.62, 0.81]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.71; 95% CI [0.63, 0.79]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.81; 95% CI [0.73, 0.90]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDisability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.84; 95% CI [0.76, 0.91]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.84; 95% CI [0.77, 0.91]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.84; 95% CI [0.74, 0.90]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eObesity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.80; 95% CI [0.72, 0.87]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.80; 95% CI [0.72, 0.89]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.80; 95% CI [0.73, 0.87]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMean (SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.80 (\u003c/b\u003e\u0026plusmn;\u0026thinsp;0.25\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.85 (\u003c/b\u003e\u0026plusmn;\u0026thinsp;0.13\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.87 (\u0026plusmn;\u0026thinsp;0.11)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eDemographic Distribution\u003c/p\u003e \u003cp\u003eA total of 264 clinical vignettes were identified across 106 instructional files, while 50 files contained no clinical vignettes. Site B contributed a larger proportion of vignettes (66.3%, n\u0026thinsp;=\u0026thinsp;175) compared with Site A (33.7%, n\u0026thinsp;=\u0026thinsp;89), despite contributing fewer files proportionally.\u003c/p\u003e \u003cp\u003eAcross specialties, Neurology had the highest number of clinical vignettes (n\u0026thinsp;=\u0026thinsp;77; 29.1%), followed by Cardiology (n\u0026thinsp;=\u0026thinsp;65; 24.6%) and GHUP (n\u0026thinsp;=\u0026thinsp;29; 11.0%). The number of vignettes per file varied by specialty, with Hematology, Neurology, and Immunology averaging approximately 3\u0026ndash;4 vignettes per file, whereas OB-GYN materials contained fewer cases per file.\u003c/p\u003e \u003cp\u003eThe clinical vignette extraction module achieved perfect agreement with clinician annotations (F\u003csub\u003e1\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;1.0). For demographic extraction, the initial few-shot multi-agent approach achieved a mean F\u003csub\u003e1\u003c/sub\u003e score of 0.85 across 16 attributes. Following the correction stage, overall performance improved to a mean F\u003csub\u003e1\u003c/sub\u003e score of 0.87.\u003c/p\u003e \u003cp\u003eThe validation module aided in resolving initial errors, with the largest improvements observed for cigarette smoking, suicidality, and socioeconomic status. Attribute-level performance varied, with perfect F\u003csub\u003e1\u003c/sub\u003e scores for gender, ethnicity, religion, sexual orientation, immigration status, and mental health, whereas race and socioeconomic status remained more challenging. Across the corpus, most demographic attributes were unspecified, with \u0026ldquo;Unknown\u0026rdquo; labels exceeding 90% for race, ethnicity, religion, and socioeconomic status (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDistribution of explicitly reported vs unknown demographic attributes across clinical vignettes (N\u0026thinsp;=\u0026thinsp;264)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDemographic Attribute\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eExplicitly reported (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUnknown (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e233 (88.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31 (11.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e233 (88.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31 (11.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGender\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 (2.6%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e257 (97.4%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRace\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e11 (4.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e253 (95.8%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEthnicity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e14 (5.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e250 (94.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eReligion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0 (0%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e264 (100%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSexual orientation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16 (6.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e248 (93.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSES\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e12 (4.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e252 (95.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImmigration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e6 (2.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e258 (97.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHousing\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e19 (7.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e245 (92.8%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAddiction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10 (3.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e254 (96.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCigarette smoking\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16 (6.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e248 (93.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMental health\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e15 (5.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e249 (94.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSuicidality\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 (0.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e263 (99.6%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDisability\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e51 (19.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e213 (80.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eObesity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8 (3.0%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e256 (97.0%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003ePrincipal Findings\u003c/p\u003e \u003cp\u003eThis study demonstrates the feasibility of using a multi-agent LLM framework to automatically identify clinical vignettes and extract demographic information from large volumes of medical education materials, also showing the utility of LLMs in medical education (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). The system achieved near-perfect performance for vignette identification and strong overall performance for demographic information extraction, with measurable improvements over both zero-shot and few-shot approaches. The addition of a correction stage resolved approximately 27.8% of initial extraction errors, suggesting that iterative review mechanisms may improve the reliability of automated information extraction systems (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAttributes with perfect F\u003csub\u003e1\u003c/sub\u003e scores were characterized by low prevalence and predominantly \u0026lsquo;Unknown\u0026rsquo; labels, which may partially contribute to high agreement due to class imbalance. Importantly, the framework was designed to extract only explicitly stated demographic information and to avoid inference of protected characteristics. This design choice prioritizes precision and may reduce the risk of introducing algorithmic bias into curricular analyses. During annotation, reviewers occasionally inferred characteristics that were not explicitly stated in the vignette text. For example, diagnoses such as dementia were sometimes interpreted as implying disability status, even when disability was not directly mentioned. While such inferences may be reasonable for experienced clinicians, the framework was designed to reflect the perspective of the learner, who may not automatically infer these relationships. This distinction highlights a potential source of bias in expert annotation and suggests that there may be differences between what educators intend to teach and what learners explicitly encounter in curricular materials (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eBeyond technical performance, the extracted corpus provides insight into patterns of demographic representation in educational materials. Most demographic dimensions were marked as \u0026lsquo;Unknown\u0026rsquo; in the majority of vignettes, particularly for race, ethnicity, religion, socioeconomic status, immigration status, and sexual orientation. This suggests that many social and identity characteristics are either rarely included or selectively reported, consistent with prior research demonstrating that demographic context and social determinants of health are unevenly represented in medical education materials (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAge and binary sex were the most frequently specified characteristics, reflecting their traditional role in clinical reasoning (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). In contrast, non-binary gender identities were not represented, and explicit race and ethnicity were rarely stated. These patterns may reflect efforts to avoid stereotyping, but they may also indicate missed opportunities to contextualize health disparities and social determinants of health within clinical teaching cases, which has been emphasized as an important component of contemporary medical education (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe high performance observed for clinical vignette identification likely reflects the structured narrative characteristics of vignette text, which differ from surrounding explanatory or didactic content. The modular separation of vignette detection from demographic extraction likely reduced task complexity at each stage, and may have contributed to overall system stability (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe validation module played a critical role in improving performance. By re-evaluating initial outputs against the source text, the system reduced inconsistencies and mitigated reasoning errors that have been reported in single-pass LLM pipelines. This iterative validation step accounted for a 2.47% improvement over standard few-shot extraction.\u003c/p\u003e \u003cp\u003eClinician annotation provided a reliable reference standard, with a mean Krippendorff\u0026rsquo;s alpha of 0.83, indicating moderate to high agreement. Publicly available datasets for demographic extraction from medical education materials are currently limited, particularly for tasks involving structured identification of clinical vignettes and associated demographic attributes. As a result, the creation of the clinician-annotated dataset in this study can aid further research.\u003c/p\u003e \u003cp\u003ePerformance variability across demographic attributes likely reflects differences in how such information is expressed in educational materials (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Attributes such as gender identity or immigration status, when present, were typically stated explicitly, whereas race, socioeconomic status, and housing context were often implicit or omitted.\u003c/p\u003e \u003cp\u003eFrom a systems perspective, the use of a high-parameter model for context-intensive tasks alongside smaller parallel extraction agents allowed a balance between computational efficiency and accuracy.\u003c/p\u003e \u003cp\u003eFuture work could extend this framework to support intersectional analyses of demographic representation and explore real-time integration within curriculum development workflows, enabling faculty to receive automated feedback during content creation.\u003c/p\u003e \u003cp\u003eError Analysis\u003c/p\u003e \u003cp\u003eError analysis suggests that most errors occurred in attributes requiring interpretation of nuanced language rather than direct extraction (e.g., socioeconomic status). Attributes with sparse positive cases (e.g., suicidality, addiction, housing insecurity) showed greater performance variability due to class imbalance. The correction module produced the largest gains in these low-frequency categories.\u003c/p\u003e \u003cp\u003eA small number of structural challenges were observed. One vignette presented multiple patients in tabular format, which required segmentation beyond standard narrative extraction. Additionally, approximately ten potential false-positive vignette candidates were identified in the broader corpus, typically where clinical descriptions were embedded within explanatory text rather than presented as standalone cases.\u003c/p\u003e \u003cp\u003eNo false positives were observed within the gold standard dataset, providing additional support for the robustness of the vignette identification module under evaluation conditions used in this study.\u003c/p\u003e \u003cp\u003eImplications for Medical Education\u003c/p\u003e \u003cp\u003eThe ability to automatically extract demographic information from curricular materials has several potential applications for medical education research and curriculum oversight.\u003c/p\u003e \u003cp\u003eFirst, it enables large-scale curricular audits to identify gaps or imbalances and unintentional omissions in demographic representation across courses, sites, or specialties. Such analyses can inform targeted instructional revisions to ensure that learners are exposed to diverse patient populations. Because many educational materials are developed by individual instructors over time, course directors may have limited visibility into the full set of cases presented across a curriculum. Tools that enable systematic review of case content may therefore help ensure that core topics and patient populations remain consistently represented. For example, our analysis identified no clinical cases involving menopause within the OB-GYN Clerkship curriculum. While omissions may reflect local curricular design choices, they may also occur with changes in lecture or course leadership. This example illustrates how automated case auditing can surface potential gaps in topic coverage that might otherwise go unnoticed. Omission of either common or high acuity-low frequency presentations from clerkship-level cases may contribute to gaps in learners' clinical pattern recognition, underscoring the downstream educational relevance of such findings.\u003c/p\u003e \u003cp\u003eSecond, automated monitoring may help identify patterned associations between demographic characteristics and specific conditions, allowing educators to detect potential reinforcement of stereotypes or implicit bias (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e). This capability is particularly relevant given prior evidence that demographic details in clinical materials are often presented selectively, contributing to the hidden curriculum through which learners develop implicit assumptions about patients and disease (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). Identifying these patterns at scale is a necessary first step toward intentional curricular revision.\u003c/p\u003e \u003cp\u003eThird, the framework supports longitudinal tracking, enabling institutions to evaluate the impact of educational initiatives over time and to monitor whether revisions to instructional materials produce measurable changes in representation.\u003c/p\u003e \u003cp\u003eFourth, automated demographic profiling can serve as a concrete, data-driven foundation for faculty development. Rather than relying on abstract discussions of representation in case design, course directors and instructors can engage with evidence drawn directly from their own curricular materials, supporting more targeted and actionable conversations about equitable vignette construction.\u003c/p\u003e \u003cp\u003eFinally, this approach provides a scalable pathway to evidence-based curricular quality assurance that was not previously achievable through manual review alone. Manual audits are time-intensive, difficult to sustain across large and evolving educational repositories, and limited in their capacity to detect aggregate patterns. By enabling systematic, reproducible analysis at the curriculum level, this framework supports a mode of oversight that complements rather than replaces the judgment of faculty and enables full curricular evaluation.\u003c/p\u003e \u003cp\u003eLimitations\u003c/p\u003e \u003cp\u003eAlthough our proposed framework performs well, there are some limitations of this study. First, the study included materials from only two US medical schools. Curricular design, institutional priorities, and regional patient populations may influence demographic representation, limiting generalizability. Future validation across larger and more diverse corpora is warranted to assess generalizability.\u003c/p\u003e \u003cp\u003eSecond, the analysis was restricted to instructional materials and does not capture demographic exposure that occurs during clinical training, clerkships, or residency.\u003c/p\u003e \u003cp\u003eThird, many demographic attributes were rare or absent in the source materials, resulting in class imbalance that may inflate performance estimates for categories dominated by \u0026ldquo;Unknown.\u0026rdquo;\u003c/p\u003e \u003cp\u003eFinally, some demographic domains were limited to single-category extraction (e.g., one mental health condition), which may not reflect the complexity of comorbidity or multidimensional social factors.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study demonstrates that a multi-agent LLM framework can reliably identify clinical vignettes and extract demographic information from medical education materials. These findings indicate that automated methods can be used to systematically audit demographic representation in curricular content. Beyond methodological contribution, the approach enables large-scale analysis of demographic representation across curricular repositories, supporting educators in identifying gaps, monitoring trends, and promoting more balanced educational materials.\u003c/p\u003e \u003cp\u003eAs medical education increasingly emphasizes equity, social determinants of health, and bias mitigation, automated auditing frameworks offer a practical mechanism to support these priorities. Future work should extend validation across additional institutions and explore integration of such tools into routine curriculum review processes.\u003c/p\u003e \u003cp\u003eThe automated pipeline demonstrates a scalable approach for evaluating whether clinical vignettes reflect the diversity of patient populations, with potential to support the development of more equitable clinical reasoning and diagnostic training in future physicians.\u003c/p\u003e"},{"header":"Declarations","content":" \u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003eThis study did not include identifiable patient data or involve human subjects research, and was therefore exempt from IRB review by Emory University.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eConsent for publication\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eResearch reported in this publication was supported by Emory Women's Impact Circle (WIC), Society for Academic Emergency Medicine (SAEM) and Rymer Innovation Grant. The content is solely the responsibility of the authors and does not necessarily represent the official views of the funding agencies.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eSD designed the methodology, conducted the analyses, and drafted the manuscript. RI and ND performed data annotation. YGe and YGuo contributed to data curation and analysis. AS contributed to methodology and supervised the study. MR conceptualized and supervised the study. All authors contributed to, reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe authors thank the faculty members who contributed instructional materials for this study.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eExtracted clinical vignettes and demographic information shall be made available upon reasonable request and signing of a data use agreement, for research purposes.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eEva KW. What every teacher needs to know about clinical reasoning. Med Educ. 2005;39(1):98\u0026ndash;106. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/j.1365-2929.2004.01972.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1365-2929.2004.01972.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePlaum P, Visser LN, de Groot B, Morsink MEB, Duijst WLJM, Candel BGJ. Using case vignettes to study the presence of outcome, hindsight, and implicit bias in acute unplanned medical care: a cross-sectional study. Eur J Emerg Med. 2024;31(4):260. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/MEJ.0000000000001127\u003c/span\u003e\u003cspan address=\"10.1097/MEJ.0000000000001127\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChapman EN, Kaatz A, Carnes M. Physicians and Implicit Bias: How Doctors May Unwittingly Perpetuate Health Care Disparities. J GEN INTERN MED. 2013;28(11):1504\u0026ndash;10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s11606-013-2441-1\u003c/span\u003e\u003cspan address=\"10.1007/s11606-013-2441-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHafferty FW. Beyond curriculum reform: confronting medicine\u0026rsquo;s hidden curriculum. Acad Med. 1998;73(4):403\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/00001888-199804000-00013\u003c/span\u003e\u003cspan address=\"10.1097/00001888-199804000-00013\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee CR, Gilliland KO, Beck Dallaghan GL, Tolleson-Rinehart S. Race, ethnicity, and gender representation in clinical case vignettes: a 20-year comparison between two institutions. BMC Med Educ. 2022;22(1):585. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12909-022-03665-4\u003c/span\u003e\u003cspan address=\"10.1186/s12909-022-03665-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsai J, Ucik L, Baldwin N, Hasslinger C, George P. Race Matters? Examining and Rethinking Race Portrayal in Preclinical Medical Education. Acad Med. 2016;91(7):916\u0026ndash;20. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/ACM.0000000000001232\u003c/span\u003e\u003cspan address=\"10.1097/ACM.0000000000001232\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of Hallucination in Natural Language Generation. ACM Comput Surv. 2023;55(12):2481\u0026ndash;248. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/3571730\u003c/span\u003e\u003cspan address=\"10.1145/3571730\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark JS, O\u0026rsquo;Brien J, Cai CJ, Morris MR, Liang P, Bernstein MS. Generative Agents: Interactive Simulacra of Human Behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology [Internet]. New York, NY, USA: Association for Computing Machinery; 2023 [cited 2026 Feb 19]. pp. 1\u0026ndash;22. (UIST \u0026rsquo;23). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dl.acm.org/doi/10.1145/3586183.3606763 doi:10.1145/3586183.3606763\u003c/span\u003e\u003cspan address=\"https://dl.acm.doi/10.1145/3586183.3606763 doi:10.1145/3586183.3606763\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMadaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, et al. Self-refine: Iterative refinement with self-feedback. Adv Neural Inf Process Syst. 2023;36:46534\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenAI, Agarwal S, Ahmad L, Ai J, Altman S, Applebaum A et al. gpt-oss-120b \u0026amp; gpt-oss-20b Model Card [Internet]. arXiv; 2025 [cited 2026 Mar 10]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2508.10925\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2508.10925\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2508.10925\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2508.10925\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrattafiori A, Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A et al. The Llama 3 herd of models. In: Neural Information Processing Systems [Internet]. Curran Associates; 2024 [cited 2026 Mar 10]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://hal.science/hal-05412319/\u003c/span\u003e\u003cspan address=\"https://hal.science/hal-05412319/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi J, Sun S, Yuan W, Fan RZ, Liu P. Generative Judge for Evaluating Alignment. In: The Twelfth International Conference on Learning Representations. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrippendorff K. Computing Krippendorff\u0026rsquo;s alpha-reliability [Internet]. 2011 [cited 2026 Mar 10]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://repository.upenn.edu/handle/20.500.14332/2089\u003c/span\u003e\u003cspan address=\"https://repository.upenn.edu/handle/20.500.14332/2089\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large language models in medical education: opportunities, challenges, and future directions. JMIR Med Educ. 2023;9(1):e48291.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVrdoljak J, Boban Z, Vilović M, Kumrić M, Božić J. A review of large language models in medical education, clinical decision support, and healthcare administration. In: Healthcare [Internet]. MDPI; 2025 [cited 2026 Mar 10]. p. 603. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.mdpi.com/2227-9032/13/6/603\u003c/span\u003e\u003cspan address=\"https://www.mdpi.com/2227-9032/13/6/603\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeng Y, Zhu M, Xia F, Li B, He S, Liu S et al. Large language models are better reasoners with self-verification. In: Findings of the Association for Computational Linguistics: EMNLP 2023 [Internet]. 2023 [cited 2026 Mar 10]. pp. 2550\u0026ndash;75. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aclanthology.org/2023.findings-emnlp.167/\u003c/span\u003e\u003cspan address=\"https://aclanthology.org/2023.findings-emnlp.167/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLouie P, Wilkes R. Representations of race and skin tone in medical textbook imagery. Soc Sci Med. 2018;202:38\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoss PT, Hart-Johnson T, Santen SA, Zaidi NLB. Considerations for using race and ethnicity as quantitative variables in medical education research. Perspect Med Educ. 2020;9(5):318\u0026ndash;23. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/S40037-020-00602-3\u003c/span\u003e\u003cspan address=\"10.1007/S40037-020-00602-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrishnan A, Rabinowitz M, Ziminsky A, Scott SM, Chretien KC. Addressing race, culture, and structural inequality in medical education: a guide for revising teaching cases. Acad Med. 2019;94(4):550\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhot T, Trivedi H, Finlayson M, Fu Y, Richardson K, Clark P et al. Decomposed Prompting: A Modular Approach for Solving Complex Tasks. In: The Eleventh International Conference on Learning Representations [Internet]. [cited 2026 Mar 10]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://openreview.net/forum?id=_nGgzQjzaRy\u003c/span\u003e\u003cspan address=\"https://openreview.net/forum?id=_nGgzQjzaRy\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarey-Ewend K, Feinberg A, Flen A, Williamson C, Gutierrez C, Cykert S, et al. Use of Sociodemographic Information in Clinical Vignettes of Multiple-Choice Questions for Preclinical Medical Students. MedSciEduc. 2023;33(3):659\u0026ndash;67. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s40670-023-01778-z\u003c/span\u003e\u003cspan address=\"10.1007/s40670-023-01778-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmutah C, Greenidge K, Mante A, Munyikwa M, Surya SL, Higginbotham E Misrepresenting Race \u0026mdash; The Role of Medical Schools in Propagating Physician Bias. New England Journal of Medicine., Amutah C, Greenidge K, Mante A, Munyikwa M, Surya SL, Higginbotham E et al. Misrepresenting Race \u0026mdash; The Role of Medical Schools in Propagating Physician Bias. New England Journal of Medicine. 2021 Mar 3;384(9):872\u0026ndash;8. doi:10.1056/NEJMms2025768.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Demographic extraction, Clinical vignettes, Large Language Models, Multi-agent Artificial Intelligence","lastPublishedDoi":"10.21203/rs.3.rs-9254050/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9254050/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground.\u003c/h2\u003e \u003cp\u003eClinical vignettes are widely used in medical education to teach clinical reasoning, evaluate diagnostic skills, and contextualize patient care. These vignettes often include demographic information such as age, sex, race, ethnicity, and other identity markers that may influence learners\u0026rsquo; understanding of patient diversity. Inclusion of underrepresented populations is particularly important when demographic factors may affect diagnosis and care. Prior research shows that demographic details in clinical educational materials may be presented inconsistently or selectively, reflecting elements of hidden curriculum and reinforcing stereotypes. However, manually extracting demographic information from clinical vignettes in curricular materials for audits is resource-intensive, while existing automated approaches show limited reliability.\u003c/p\u003e\u003ch2\u003eMethods.\u003c/h2\u003e \u003cp\u003eWe propose a multi-agentic framework to automatically extract demographic information from clinical vignettes. Our framework consists of three stages: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) a few-shot learning step to identify clinical vignettes, (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) a few-shot infused demographic information extraction step in which multiple parallel agents extract demographic data from identified vignettes; and (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) an LLM-as-a-judge validation step in which extracted information is autonomously reviewed to correct potential errors. We employed this framework on a corpus of 156 instructional materials containing 264 clinical vignettes across 11 medical specialties from two US medical schools. A dataset of 56 clinical vignettes from 25 curricular material units was annotated by two clinicians across 16 demographic categories for in-context learning and model evaluation using F\u003csub\u003e1\u003c/sub\u003e score.\u003c/p\u003e\u003ch2\u003eResults.\u003c/h2\u003e \u003cp\u003eInter-annotator agreement between clinicians was high (average Krippendorff's alpha: 0.83). Clinical vignette identification achieved F\u003csub\u003e1\u003c/sub\u003e score of 1.0 and demographic information extraction across all categories achieved average F\u003csub\u003e1\u003c/sub\u003e score of 0.87. Few-shot extraction with autonomous error correction improved performance by 8% compared to zero-shot demographic extraction. The validation stage corrected 27.78% of errors from the initial extraction.\u003c/p\u003e\u003ch2\u003eConclusion.\u003c/h2\u003e \u003cp\u003eA multi-agent LLM framework can effectively automate the identification of vignettes and extraction of demographic information from medical curricular materials. Multi-agent few-shot detection and autonomous correction significantly improves over zero-shot or single-agent methods. The ability to accurately audit large-scale educational corpora provides a scalable solution for medical educators to identify and address demographic representation gaps.\u003c/p\u003e","manuscriptTitle":"Automated Extraction of Patient Demographics from Clinical Vignettes Using a Multi-Agent Large Language Model Framework","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-08 16:20:00","doi":"10.21203/rs.3.rs-9254050/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-05-17T08:39:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"19375677464472601341815145319696877215","date":"2026-05-04T08:31:08+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-24T06:29:51+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-04-04T16:19:31+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-02T10:48:07+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-02T10:47:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Education","date":"2026-03-28T15:55:17+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3fa0bed3-2f90-494c-aadf-7a66e271dcd5","owner":[],"postedDate":"May 8th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-17T08:39:08+00:00","index":75,"fulltext":""},{"type":"reviewerAgreed","content":"19375677464472601341815145319696877215","date":"2026-05-04T08:31:08+00:00","index":72,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-08T16:48:11+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-08 16:20:00","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9254050","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9254050","identity":"rs-9254050","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00