Artificial Intelligence Based Assessment of Clinical Reasoning Documentation: An Observational Study of the Impact of the Clinical Learning Environment on Resident Performance

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-06, 2026-06-24 · read from full text

This observational retrospective cross-sectional study used a supervised machine learning model to classify the quality of clinical reasoning documentation in 37,750 hospital admission notes written by 474 categorical internal medicine residents at two NYU hospital sites across 2018–2023. Using EHR-derived patient factors (age, sex, Charlson Comorbidity Index, and primary diagnosis category) and clinical learning environment markers (academic year, day vs night shift, and note index within shift), the authors analyzed associations with documentation quality using generalized estimating equations with resident-level clustering. Patients who were older, had higher comorbidity, and had certain diagnosis categories were linked to higher documentation quality; after adjustment, high-quality documentation was associated with later academic year, night shift, and a lower note index. A key limitation is that it relied on documentation quality as a proxy for clinical reasoning and excluded ICU notes because they were not part of the AI model’s validation set. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Background Objective measures and large datasets are needed to determine aspects of the Clinical Learning Environment (CLE) impacting resident performance. Artificial Intelligence (AI) offers a solution. Here, the authors sought to determine what aspects of the CLE might be impacting resident performance as measured by clinical reasoning documentation quality assessed by AI. Methods In this observational, retrospective cross-sectional analysis of hospital admission notes from the Electronic Health Record (EHR), all categorical internal medicine (IM) residents who wrote at least one admission note during the study period July 1, 2018 – June 30, 2023 at two sites of NYU Grossman School of Medicine’s IM residency program were included. Clinical reasoning documentation quality of admission notes was determined to be low or high-quality using a supervised machine learning model. From note-level data, the shift (day or night) and note index within shift (if a note was first, second, etc. within shift) were calculated. These aspects of the CLE were included as potential markers of workload, which have been shown to have a strong relationship with resident performance. Patient data was also captured, including age, sex, Charlson Comorbidity Index, and primary diagnosis. The relationship between these variables and clinical reasoning documentation quality was analyzed using generalized estimating equations accounting for resident-level clustering. Results Across 37,750 notes authored by 474 residents, patients who were older, had more pre-existing comorbidities, and presented with certain primary diagnoses (e.g., infectious and pulmonary conditions) were associated with higher clinical reasoning documentation quality. When controlling for these and other patient factors, variables associated with clinical reasoning documentation quality included academic year (adjusted odds ratio, aOR, for high-quality: 1.10; 95% CI 1.06-1.15; P<.001), night shift (aOR 1.21; 95% CI 1.13-1.30; P<.001), and note index (aOR 0.93; 95% CI 0.90-0.95; P<.001). Conclusions AI can be used to assess complex skills such as clinical reasoning in authentic clinical notes that can help elucidate the potential impact of the CLE on resident performance. Future work should explore residency program and systems interventions to optimize the CLE.
Full text 105,220 characters · extracted from preprint-html · click to expand
Artificial Intelligence Based Assessment of Clinical Reasoning Documentation: An Observational Study of the Impact of the Clinical Learning Environment on Resident Performance | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Artificial Intelligence Based Assessment of Clinical Reasoning Documentation: An Observational Study of the Impact of the Clinical Learning Environment on Resident Performance Verity Schaye, David J DiTullio, Daniel J Sartori, Kevin Hauck, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4427373/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 22 Apr, 2025 Read the published version in BMC Medical Education → Version 1 posted 12 You are reading this latest preprint version Abstract Background Objective measures and large datasets are needed to determine aspects of the Clinical Learning Environment (CLE) impacting resident performance. Artificial Intelligence (AI) offers a solution. Here, the authors sought to determine what aspects of the CLE might be impacting resident performance as measured by clinical reasoning documentation quality assessed by AI. Methods In this observational, retrospective cross-sectional analysis of hospital admission notes from the Electronic Health Record (EHR), all categorical internal medicine (IM) residents who wrote at least one admission note during the study period July 1, 2018 – June 30, 2023 at two sites of NYU Grossman School of Medicine’s IM residency program were included. Clinical reasoning documentation quality of admission notes was determined to be low or high-quality using a supervised machine learning model. From note-level data, the shift (day or night) and note index within shift (if a note was first, second, etc. within shift) were calculated. These aspects of the CLE were included as potential markers of workload, which have been shown to have a strong relationship with resident performance. Patient data was also captured, including age, sex, Charlson Comorbidity Index, and primary diagnosis. The relationship between these variables and clinical reasoning documentation quality was analyzed using generalized estimating equations accounting for resident-level clustering. Results Across 37,750 notes authored by 474 residents, patients who were older, had more pre-existing comorbidities, and presented with certain primary diagnoses (e.g., infectious and pulmonary conditions) were associated with higher clinical reasoning documentation quality. When controlling for these and other patient factors, variables associated with clinical reasoning documentation quality included academic year (adjusted odds ratio, aOR, for high-quality: 1.10; 95% CI 1.06-1.15; P <.001), night shift (aOR 1.21; 95% CI 1.13-1.30; P <.001), and note index (aOR 0.93; 95% CI 0.90-0.95; P <.001). Conclusions AI can be used to assess complex skills such as clinical reasoning in authentic clinical notes that can help elucidate the potential impact of the CLE on resident performance. Future work should explore residency program and systems interventions to optimize the CLE. Artificial Intelligence Clinical Reasoning Documentation Clinical Learning Environment Figures Figure 1 Background Direct patient care provides residents with clinical experiences central to the development of competence ( 1 ). However, learning in the authentic clinical environment can be challenging due to the unpredictability, pace, and complexity inherent to caring for sick patients. The clinical learning environment (CLE) is the interplay between the clinical context where trainees participate in patient care and education (including the expected learning outcomes and assessment practices) ( 2 , 3 ). The CLE is a component of residency accreditation and, as a mediator of attainment of clinical competence, has the potential for long-lasting impact on future practice patterns ( 4 – 8 ). Studies across different medical specialties have investigated how the CLE impacts resident performance using various outcomes. In the procedural specialties of obstetrics and gynecology and surgery, researchers have used procedural outcomes to examine the interplay between CLE and resident performance such as reviewing EHR data for maternal birth complications and postoperative complications ( 9 , 10 ). In internal medicine (IM), research around the CLE has focused on clinical workload, shift timing, and work hours with noted negative impacts on learning as assessed through trainee surveys; however, objective, attributable, and scalable measurement of such relationships proves difficult in practice ( 11 – 13 ). Those studies that do focus on objective, EHR-derived outcomes have suggested that increased resident workload, as measured by higher patient censuses and numbers of admissions, is associated with increased resource use, readmission rates, and potentially increased patient mortality ( 14 , 15 ). Similarly, reducing resident workload by lowering patient census has been shown to result in an improvement in the quality of discharge summaries as assessed by manual human review ( 16 ). However, more automated and scalable data measures, including better utilization of EHR data, are needed to determine what specific, modifiable aspects of the CLE might be impacting resident performance and patient outcomes ( 17 ). Artificial intelligence (AI) and big data analytics offer approaches for generating the necessary scalable, large datasets of resident-specific, objective measures of resident performance that – if melded with patient, team, and contextual factors – could further unpack the potential impact of the CLE on resident performance ( 18 ). Despite this opportunity, AI-based assessments of resident performance are not yet widely implemented ( 19 – 22 ). In earlier phases of this work, we developed an AI-based assessment of clinical reasoning documentation to increase the frequency and quality of feedback on this important skill ( 23 ). Clinical reasoning is a core skill that is essential to patient care and clinical documentation serves as an important measure of resident clinical reasoning ( 24 – 26 ). While the absence of high-quality documentation of clinical reasoning does not preclude that reasoning was performed, higher-quality clinical reasoning documentation has been hypothesized to be associated with improved patient outcomes including reduced diagnostic errors ( 27 – 29 ). Furthermore, the EHR and clinical documentation are integral to the CLE and can serve as an important objective measure of resident performance to better explore the impact of the CLE ( 30 ). Historically, there have been limitations in using quality of clinical documentation as a more systemic measurement; while multiple human rating tools exist to assess the quality of clinical reasoning documentation, they can be time-intensive and therefore difficult to generate insights at scale ( 18 , 26 , 31 – 33 ). AI-based automated assessment of clinical reasoning documentation enables the creation of nearly instantaneous large datasets measuring resident clinical reasoning documentation practices; such tools have the potential to assess clinical reasoning at the individual and aggregate level at a scale that is not feasible with human rating. Here, we report a cross-sectional analysis of the relationship between the CLE (e.g., markers of workload such as volume of admissions and time of shift) and AI-based assessment of clinical reasoning documentation quality among a large retrospective cohort of IM residents at two distinct residency program hospital sites within one quaternary academic health system. Methods Study Sites This study was conducted at NYU Grossman School of Medicine’s IM residency program, a large academic urban training program in New York City. The study was conducted at two of the residency program’s hospital sites—NYU Langone Hospital (NYULH)-Manhattan and NYULH-Brooklyn—which utilize a single EHR. The Manhattan site (181 residents overall) is a university-based hospital, quaternary referral and transplant center. These residents also rotate at two other affiliated sites with distinct EHRs not included in this study but do not rotate at the Brooklyn site. The Brooklyn site (41 residents overall) is a community-based tertiary hospital whose residency program became affiliated with NYU in 2016. At each site, residents staff 6 medicine inpatient teams. Residents are divided into day teams (7 am to 7 pm) and night teams (7 pm to 7 am), both responsible for caring for existing patients while also admitting patients daily, including writing admission notes. The day team structure typically comprises a supervising attending, a senior resident, two interns, and medical students. The night team is a resident-intern dyad. Study Population We included all categorical IM residents who wrote at least one admission note on a general IM service from July 1, 2018 – June 30, 2023 (five academic years), at either NYULH-Manhattan (n = 356 residents) or NYULH-Brooklyn (n = 118 residents). Preliminary (e.g., Neurology) and rotating residents (e.g., Psychiatry) were excluded. Notes authored in the intensive care unit were excluded, as such notes were not included in the initial validation work of the AI algorithm ( 23 ). Data Retrieval Admission notes were retrieved using a custom query of the EHR data warehouse in July 2023. Note-level data included the note creation date and time. From this, the shift (day or night) and creation time within the shift were calculated. Additionally, a note index was calculated at the shift level indicating if a given note was the first, second, third, etc. note authored within a shift. These aspects of the CLE were included as potential markers of workload, which have been shown to have a strong relationship with resident performance ( 14 – 16 ). Patient-level data at the time of each inpatient hospitalization was also captured, such as patient age, sex, insurance type, and Charlson Comorbidity Index (CCI). These variables were included as they are all patient characteristics that can influence the clinical reasoning process in the CLE ( 34 ). CCI encompasses 17 disease categories (e.g., congestive heart failure, renal disease) that were weighted using an age-independent adaptation revised for International Classification of Disease, 10th edition (ICD-10-CM) codes present before the date of hospital admission (score range 0 to 29, with higher indicating more comorbidity) ( 35 – 37 ). Hospitalizations were categorized by diagnostic area using the ICD-10-CM chapter of hospitalization-level primary diagnoses, as done previously (Supplement Digital Appendix 1) ( 38 ). Hospitalizations include a primary diagnosis (selected by a non-physician coder based on chart review) and a principal diagnosis (selected by clinicians). A chart review was conducted on a random sample of 100 hospitalizations to determine a rule-based approach for selecting the optimal diagnostic area. Primary diagnosis was used except when overly general codes were included (“sepsis, unspecified organism” or a sign/symptom ICD-10-CM code), where the principal diagnosis was used instead. Author-level data on each resident’s primary training site and post-graduate year (PGY) at the time of note creation were retrieved from an educational data warehouse ( 39 ). AI Algorithm Each note was analyzed using a previously developed supervised machine learning model that classifies clinical reasoning documentation quality in IM resident admissions notes as low- or high-quality ( 23 ). The human rating gold standard for training the model was the Revised-DEA assessment tool which defines a high-quality hospital admission note in three domains as a note that has a differential diagnosis (D) that is explicitly prioritized with specific diagnoses (rather than diagnostic categories, e.g. infection) and provides clear explanation of reasoning (E) for the lead and alternative (A) diagnoses in the differential ( 25 ). The model uses the Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES, v4.0.0), a pre-large language model (LLM) open-source natural language processing (NLP) tool that was trained on clinical notes and performs named entity recognition (i.e., extracting diagnoses and other concepts). The algorithm has excellent note-level performance with an area under the receiver operating characteristic curve of 0.88 ( 23 ). Statistical Analysis Descriptive statistics, chi-squared testing, regression analyses, and visualizations were performed with the R statistical software (v4.2.0) using a variety of libraries ( tidyverse , data.table , gtsummary , ggplot2 ). Hypothesis testing was two-sided at α = .05. Regression models for high-quality clinical reasoning were built both without adjustment for clustering (logistic regression using glm package of stats library) and with adjustment for clustering by resident using generalized estimating equations (GEE) with an exchangeable correlation structure ( geeglm package of geepack library) and robust standard error estimates. Interaction terms were evaluated for each of the CLE variables; significant interactions were included in the final model. The within-cluster intraclass correlation was approximately 0.08, indicating only modest within-resident clustering of performance. The study was approved by the NYU Grossman School of Medicine institutional review board. Results Descriptive statistics were computed for resident characteristics, CLE characteristics, and patient characteristics stratified by residency program site (Table 1 and Supplement Digital Appendix 2) and post-graduate year (PGY) (Supplement Digital Appendix 3). There were 37,750 admission notes (16,180 Manhattan and 21,570 Brooklyn) authored by 474 distinct residents (356 Manhattan residents and 118 Brooklyn residents) with a mean 79.6 notes per resident (mean 45 per Manhattan resident and 183 per Brooklyn resident) across 28,782 distinct patients between both sites from July 2018 to June 2023. Overall, 55% of the notes were high-quality (64% high-quality at Manhattan and 48% high-quality at Brooklyn). At both sites PGY-2 or PGY-3 residents accounted for the majority of high-quality notes written as compared to PGY-1 residents (overall n = 15,823 (76%) vs 5,034 (24%), P < .001). Senior residents authored more high-quality notes at Brooklyn (PGY2 or PGY3, n = 8,992 (86%) vs PGY1, n = 1,448 (14%), P < .001) compared to Manhattan (PGY2 or PGY3, n = 6,831 (66%) vs PGY1 n = 3,586 (34%), P < .001). At both sites more of the high-quality notes were written during the night shift as compared to day shift (overall n = 13,518 (65%) vs n = 7,339 (35%), P < .001) and notes with an earlier note index in the shift accounted for the majority of high-quality notes (overall 1st note in shift n = 11,890 (57%) vs 2nd note in shift n = 4,657 (22%) vs 3rd note in shift n = 2,130 (10%) vs 4th note in shift n = 1,119 (5.4%) vs 5th or more note in shift n = 1,061 (5.1%), P < .001). Controlling for covariates and accounting for resident-level clustering, patient characteristics associated with complex admitting presentations were associated with higher quality clinical reasoning documentation, including older age (adjusted odds ratio, aOR 1.13 for age > = 80.1; 95% CI 1.05–1.22; P < .001; aOR 1.17 for age 68.1–80; 95% CI 1.09–1.25; P < .001; aOR 1.05 for age 54.1–68; 95% CI 0.99–1.11; P = .10) and higher comorbidity index (aOR 2.18 for CCI ≥ 5 referenced to CCI 0; 95% CI 2.02–2.34; P < .001) (Table 2 ). The primary diagnosis areas most strongly associated with higher clinical reasoning documentation quality included pulmonary and infectious diseases (aOR 1.53; 95% CI 1.40–1.66; P < .001; aOR 1.58; 95% CI 1.45–1.73; P < .001, respectively). Several aspects of the CLE were significantly associated with clinical reasoning documentation quality (Table 2 ). Academic year was positively associated with clinical reasoning documentation quality (aOR 1.10; 95% CI 1.06–1.15; P < .001). Notably, there was a decline in clinical reasoning documentation quality during AY 2019–2020 driven by notes authored during the peak of the initial COVID-19 surge in New York City (March to June 2020), with subsequent recovery in documentation quality (47.7% high-quality AY 2019–2020 (49.9% high-quality July-February, 42.4% high-quality March-June) vs 56.8% high-quality all other AYs, P < .001, Fig. 1 a). The NYULH-Manhattan site experienced a bigger decline in clinical reasoning documentation quality during the COVID-19 surge than the NYULH-Brooklyn site (62.2% high-quality July-February AY 2019–2020 to 50.5% high-quality March-June AY 2019–2020 at Manhattan vs 44.5% high-quality July-February AY 2019–2020 to 37.0% high-quality March-June AY 2019–2020 at Brooklyn, P < .001). Overall, performance differed by site: those at NYULH-Brooklyn were less likely to write high-quality notes than at NYULH-Manhattan (aOR 0.46; 95% CI 0.39–0.55; P < .001). Additionally, only at the NYULH-Brooklyn site was an increase in clinical reasoning documentation quality observed by PGY. PGY 2 and PGY3 residents (i.e., senior residents) at the NYULH-Brooklyn site were more likely to write high-quality clinical reasoning notes as compared to their PGY1 peers (aOR 1.39; 95% CI 1.17–1.64; P < .001). Sensitivity analysis with PGY2 and PGY3 separated showed the same result. Notes written during the night shift were more likely to be high-quality than those in the day shift (aOR 1.21; 95% CI 1.13–1.30; P < .001). Although notes written later in the night shift appeared associated with lower note quality (Fig. 1 b), this relationship disappeared in the multivariable regression (aOR 1.00; 95% CI 0.99–1.01; P = 0.80). Rather, within each shift every additional note (i.e., higher note index) was associated with an 7% lower odds of high-quality clinical reasoning documentation (aOR 0.93; 95% CI 0.90–0.95; P < .001) (Fig. 1 c). Discussion We utilized AI-based assessment of resident clinical reasoning documentation to measure resident performance across five academic years, two residency program sites, and tens of thousands of hospitalizations to better understand the relationship between the CLE and resident performance. The size of the datasets generated and this exploration was only feasible due to the power of AI. A single person working continuously for eight hours per day spending two minutes per note would have taken over 150 days to rate all notes. Furthermore, no other validated AI-based assessment tools of clinical reasoning documentation in the authentic CLE have been described ( 20 , 40 – 44 ). Melding EHR-derived measures of the CLE with this AI-based assessment revealed clinically meaningful performance differences by several CLE factors. We extracted note-level data of creation time within shift and note index to explore markers of workload which can be a major barrier to learning in the CLE ( 2 , 11 ). We found that residents were more likely to write high-quality notes during their night shifts and less likely to write high-quality notes with an increasing number of notes within a shift. For the former, we hypothesize that note quality was higher during the night shift as residents typically can focus more on new admissions without the workload of routine non-emergent care at night. While there is consensus that busyness of work and service pressures can cause cognitive overload and be a barrier to learning and performance in the CLE, both findings provide objective measures that could inform structural changes ( 11 , 12 , 14 , 16 , 45 ). Such structural changes could include further limits to the number of admissions than already exists ( 46 ) or team structures that could minimize cognitive overload like the addition of advanced practice providers (APPs) to support resident teams or the creation of swing shifts to focus on admitting without cross-coverage obligations ( 47 – 49 ). These changes would need to be balanced with the tradeoffs including requirements for increased staffing and ensuring residents having enough clinical exposure to develop competence. Unsurprisingly the major disruptor to the CLE of the COVID-19 surge in New York City in the spring of 2020 impacted the educational outcome of clinical reasoning documentation quality at both sites. NYULH-Manhattan experienced a bigger impact than NYULH-Brooklyn which correlates with the higher patient volume experienced at this site during the surge ( 50 ). At both sites clinical reasoning documentation quality has since recovered. While clinical reasoning documentation quality at NYULH-Manhattan has mostly plateaued recently, NYULH-Brooklyn experienced a continued significant rise over the ensuing two academic years with a leveling off in the most recent AY 2022–2023. We hypothesize this is a result of a major change in the CLE that occurred in 2016 with the merger of NYULH with Lutheran Medical Center which is now NYULH-Brooklyn. A similar trend in improvement was seen in clinical outcomes (such as reduction in mortality rate and fewer central line and catheter-associated urinary tract infections) after the merger as described by Wang et al ( 51 ). Using aggregate performance of AI-based assessment of clinical reasoning documentation can be an important strategy to assess for any unintended or intended impact of systems changes like mergers or EHR changes. While likely not an impact of the CLE, another difference seen in educational outcomes between NYULH-Brooklyn and NYULH-Manhattan was an increase in clinical reasoning documentation quality by PGY year only at the former site. Explanations for this might be a ceiling effect at NYULH-Manhattan where the PGY1 residents started at a higher percentage of high-quality notes (66%) than NYULH-Brooklyn (42%). Despite these differences, both sites experienced similar impacts of the CLE when controlling for covariates. Lastly, we explored clinical reasoning documentation quality by primary diagnosis area to further understand if residents were meeting educational outcomes by content domain and the potential impact clinical context might have. After controlling for patient characteristics, including preexisting comorbidities, we found that residents were more likely to write high-quality clinical reasoning documentation in the primary diagnosis domains of infectious disease and pulmonary. There could be several explanations for this finding such as increased exposure to these diagnoses in particular infectious diseases during COVID-19 ( 52 ), residents having the opportunity to rotate on pulmonary and infectious disease subspecialty-led teams, or certain presenting diagnoses not lending themselves to requiring as much diagnostic reasoning (e.g. a hematology-oncology patient admitted for chemotherapy). Ultimately, these data could be used both for feedback for individual resident practice habits or rotation planning, scheduling residents for specialty rotations, and for improvements at the program level to inform curricular planning, such as focusing on diagnostic categories with overall lower quality clinical reasoning in noon conference series. Use for feedback on practice habits could be applicable not only to residents but also to attending physicians and APPs, and has the potential to reduce diagnostic errors and improve patient outcomes by improving clinical reasoning documentation quality ( 27 , 53 , 54 ). Validating the AI-based assessment tool for attending physician and APP clinical reasoning documentation quality and exploring the relationship between clinical reasoning documentation quality and diagnostic errors are both important next areas to explore with future studies. Limitations The AI model in this study was developed with older generation NLP technology, however it did have good performance ( 23 ). LLMs will only further improve the ability to analyze EHR-based note text, which we are currently integrating into an updated model. Additionally, in the current phase of our work, we are validating the AI-based assessment of clinical reasoning documentation at a second institution, University of Cincinnati, and creating a pipeline for dissemination of the tool. The current phase of the work will enhance generalizability of this study allowing other clinical sites to explore the potential impact of CLE on resident performance using AI-based assessment of clinical reasoning documentation ( 55 ). Additionally, in terms of generalizability, although this study was conducted at a single academic institution, it was at two distinct hospital sites that have distinct CLEs. Another limitation to consider is that a resident facing dashboard to provide feedback on clinical reasoning documentation generated from output of the AI-based assessment was implemented during the study period in November 2020 ( 23 ). This could have had some influence on resident clinical reasoning documentation practices in the later part of the study period in particular accounting for the trend in improvement seen by academic year which is an ideal outcome after implementing feedback via the dashboard. Lastly, while we controlled for many key variables we determined could influence clinical reasoning documentation quality in the CLE, residual confounding may exist and the retrospective cross-sectional design precludes assessment of causality. Conclusions In this retrospective cross-sectional analysis of IM resident clinical reasoning documentation quality, we demonstrated how AI-based assessment of resident performance derived from the EHR can provide insights to unpack the impact of the CLE. We found several specific factors in the CLE associated with clinically meaningful differences in resident performance. These findings can serve as a call to action to further expand to other institutions, specialties, and healthcare providers in future studies and inform important recommended systems changes. List Of Abbreviations AI Artificial Intelligence APP Advanced Practice Provider CLE Clinical Learning Environment EHR Electronic Health Record IM Internal Medicine LLMs Large Language Models NLP Natural Language Processing NYULH NYU Langone Health Declarations Ethics approval: The study was approved by the NYU Grossman School of Medicine institutional review board on 12/9/2023 i19-00280. As this was a retrospective, observational study of EHR data review informed consent from each participant was waived for model development and retrospective data analysis by the NYU Grossman School of Medicine Institutional Review Board. Consent for publication: Not applicable. Availability of data and materials: Given the sensitive nature of EHR data, additional data cannot be readily accessible. Please email the authors [email protected] with any inquiries about the data or machine learning model. Competing interests: No authors have conflicts of interest to disclose relevant to this manuscript. Funding: None Authors’ contributions: VS made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; DT made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; DS made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; KH made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; MH made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; IR made substantial contributions to the data acquisition and analysis; BG made substantial contributions to the data acquisition and analysis; JBR made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work, and the data acquisition and analysis. Acknowledgements: The authors would like to acknowledge Helen Finkelstein, data warehouse and business intelligence developer at the Institute for Innovations in Medical Education, for her assistance in calculating the CCI. References Teunissen P, Scheele F, Scherpbier A, Van Der Vleuten C, Boor K, Van Luijk S, et al. How residents learn: qualitative evidence for the pivotal role of clinical activities. Med Educ. 2007;41(8):763–70. Nordquist J, Hall J, Caverzagie K, Snell L, Chan MK, Thoma B, et al. The clinical learning environment. Med Teach. 2019;41(4):366–72. Gruppen LD. Context and complexity in the clinical learning environment. Med Teach. 2019;41(4):373–4. Weiss KB, Bagian JP, Nasca TJ. The clinical learning environment: the foundation of graduate medical education. JAMA. 2013;309(16):1687–8. Wagner R, Patow C, Newton R, Casey BR, Koh NJ, Weiss KB. The Overview of the CLER program: CLER national report of findings 2016. J Grad Med Educ. 2016;8(2 Suppl 1):11–3. Nasca TJ, Wagner R, Weiss KB. Introduction to the CLER national report of findings 2022: The COVID-19 pandemic and its impact on the clinical learning environment. J Grad Med Educ. 2023;15(1):140–2. Chen C, Petterson S, Phillips R, Bazemore A, Mullan F. Spending patterns in region of residency training and subsequent expenditures for care provided by practicing physicians for Medicare beneficiaries. JAMA. 2014;312(22):2385–93. Asch DA, Epstein A, Nicholson S. Evaluating medical training programs by the quality of care delivered by their alumni. JAMA. 2007;298(9):1049–51. Asch DA, Nicholson S, Srinivas S, Herrin J, Epstein AJ. Evaluating obstetrical residency programs using patient outcomes. JAMA. 2009;302(12):1277–83. Bansal N, Simmons KD, Epstein AJ, Morris JB, Kelz RR. Using patient outcomes to evaluate general surgery residency program performance. JAMA Surg. 2016;151(2):111–9. Kilty C, Wiese A, Bergin C, Flood P, Fu N, Horgan M, et al. A national stakeholder consensus study of challenges and priorities for clinical learning environments in postgraduate medical education. BMC Med Educ. 2017;17(1):226. Haney EM, Nicolaidis C, Hunter A, Chan BK, Cooney TG, Bowen JL. Relationship between resident workload and self-perceived learning on inpatient medicine wards: a longitudinal study. BMC Med Educ. 2006;6:35. Arora VM, Georgitis E, Siddique J, Vekhter B, Woodruff JN, Humphrey HJ, et al. Association of workload of on-call medical interns with on-call sleep duration, shift duration, and participation in educational activities. JAMA. 2008;300(10):1146–53. Ong M, Bostrom A, Vidyarthi A, McCulloch C, Auerbach A. House staff team workload and organization effects on patient outcomes in an academic general internal medicine inpatient service. Arch Intern Med. 2007;167(1):47–52. Averbukh Y, Southern W. The impact of the number of admissions to the inpatient medical teaching team on patient safety outcomes. J Grad Med Educ. 2012;4(3):307–11. Coit MH, Katz JT, McMahon GT. The effect of workload reduction on the quality of residents' discharge summaries. J Gen Intern Med. 2011;26(1):28–32. Burk-Rafel J, Sebok-Syer SS, Santen SA, Jiang J, Caretta-Weyer HA, Iturrate E, et al. TRainee attributable & automatable care evaluations in real-time (TRACERs): A scalable approach for linking education to patient care. Perspect Med Educ. 2023;12(1):149. Arora VM. Harnessing the power of big data to improve graduate medical education: big idea or bust? Acad Med. 2018;93(6):833–4. Turner L, Hashimoto D, Vasisht S, Schaye V, Demystifying AI. Current state and future role in precision medical education assessment. Acad Med. 2024;99(4S Suppl 1):S42–7. Boscardin CK, Gin B, Golde PB, Hauer KE. ChatGPT and generative artificial intelligence for medical education: potential impact and opportunity. Acad Med. 2023:101097. Abd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large language models in medical education: Opportunities, challenges, and future directions. JMIR Med Educ. 2023;9(1):e48291. Preiksaitis C, Rose C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: Scoping review. JMIR Med Educ. 2023;9(1):e48785. Schaye V, Guzman B, Burk-Rafel J, Marin M, Reinstein I, Kudlowitz D, et al. Development and validation of a machine learning model for automated assessment of resident clinical reasoning documentation. J Gen Intern Med. 2022;37(9):2230–8. Connor DM, Durning SJ, Rencic JJ. Clinical reasoning as a core competency. Acad Med. 2020;95(8):1166–71. Schaye V, Miller L, Kudlowitz D, Chun J, Burk-Rafel J, Cocks P, et al. Development of a clinical reasoning documentation assessment tool for resident and fellow admission notes: a shared mental model for feedback. J Gen Intern Med. 2022;37(3):507–12. Baker EA, Ledford CH, Fogg L, Way DP, Park YS. The IDEA assessment tool: Assessing the reporting, diagnostic reasoning, and decision-making skills demonstrated in medical students' hospital admission notes. Teach Learn Med. 2015;27(2):163–73. Kulkarni D, Heath J, Kosack A, Jackson NJ, Crummey A. An educational intervention to improve inpatient documentation of high-risk diagnoses by pediatric residents. Hosp Pediatr. 2018;8(7):430–5. Schiff GD, Bates DW. Can electronic clinical documentation help prevent diagnostic errors? N Engl J Med. 2010;362(12):1066–9. Singh H, Giardina TD, Meyer AN, Forjuoh SN, Reis MD, Thomas EJ. Types and origins of diagnostic errors in primary care settings. JAMA Intern Med. 2013;173(6):418–25. Thoma B, Turnquist A, Zaver F, Hall AK, Chan TM. Communication, learning and assessment: Exploring the dimensions of the digital learning environment. Med Teach. 2019;41(4):385–90. Hung H, Kueh LL, Tseng CC, Huang HW, Wang SY, Hu YN, et al. Assessing the quality of electronic medical records as a platform for resident education. BMC Med Educ. 2021;21(1):577. Kogan JR, Hess BJ, Conforti LN, Holmboe ES. What drives faculty ratings of residents' clinical skills? The impact of faculty's own clinical skills. Acad Med. 2010;85(10 Suppl):S25–8. Edwards ST, Neri PM, Volk LA, Schiff GD, Bates DW. Association of note quality and quality of care: a cross-sectional study. BMJ Qual Saf. 2014;23(5):406–13. Stolper E, Van Royen P, Jack E, Uleman J, Olde Rikkert M. Embracing complexity with systems thinking in general practitioners' clinical reasoning helps handling uncertainty. J Eval Clin Pract. 2021;27(5):1175–81. Quan H, Sundararajan V, Halfon P, Fong A, Burnand B, Luthi J-C et al. Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data. Med Care. 2005:1130–9. Deyo RA, Cherkin DC, Ciol MA. Adapting a clinical comorbidity index for use with ICD-9-CM administrative databases. J Clin Epidemiol. 1992;45(6):613–9. Charlson ME, Pompei P, Ales KL, MacKenzie CR. A new method of classifying prognostic comorbidity in longitudinal studies: development and validation. J Chronic Dis. 1987;40(5):373–83. König S, Pellissier V, Hohenstein S, Leiner J, Hindricks G, Meier-Hellmann A, et al. A comparative analysis of in-hospital mortality per disease groups in Germany before and during the COVID-19 pandemic from 2016 to 2020. JAMA Netw Open. 2022;5(2):e2148649–e. Triola MM, Pusic MV. The education data warehouse: a transformative tool for health education research. J Grad Med Educ. 2012;4(1):113–5. Lin SY, Shanafelt TD, Asch SM. Reimagining clinical documentation with artificial intelligence. Mayo Clin Proc. 2018;93(5):563-5. Salt J, Harik P, Barone MA. Leveraging natural language processing: Toward computer-assisted scoring of patient notes in the USMLE Step 2 clinical skills exam. Acad Med. 2019;94(3):314–6. Kanjee Z, Crowe B, Rodman A. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA. 2023;330(1):78–80. Strong E, DiGiammarino A, Weng Y, Kumar A, Hosamani P, Hom J, et al. Chatbot vs medical student performance on free-response clinical reasoning examinations. JAMA Intern Med. 2023;183(9):1028–30. Liu J, Wang C, Liu S. Utility of ChatGPT in clinical practice. J Med Internet Res. 2023;25:e48568. Fletcher KE, Reed DA, Arora VM. Doing the dirty work: measuring and optimizing resident workload. J Gen Intern Med. 2011;26(1):8–9. Accreditation Council for Graduate Medical Education. ACGME Program Requirements for Graduate Medical Education in Internal Medicine. 2022. https://www.acgme.org/globalassets/pfassets/programrequirements/140_internalmedicine_2023.pdf . Accessed 15 May 2024. Thanarajasingam U, McDonald FS, Halvorsen AJ, Naessens JM, Cabanela RL, Johnson MG et al. Service census caps and unit-based admissions: resident workload, conference attendance, duty hour compliance, and patient safety. Mayo Clin Proc. 2012;87(4):320-7. Chandra R, Farah F, Munoz-Lobato F, Bokka A, Benedetti KL, Brueggemann C, et al. Sleep is required to consolidate odor memory and remodel olfactory synapses. Cell. 2023;186(13):2911–e2820. Yu AT, Jepsen N, Prasad S, Klein JP, Doughty C. Adding nocturnal advanced practice providers to an academic inpatient neurology service improves residents’ educational experience. Neurohospitalist. 2023;13(2):130–6. Schaye VE, Reich JA, Bosworth BP, Stern DT, Volpicelli F, Shapiro NM et al. Collaborating across private, public, community, and federal hospital systems: Lessons learned from the Covid-19 pandemic response in NYC. NEJM Catal Innov Care Deliv. 2020;1(6). Wang E, Arnold S, Jones S, Zhang Y, Volpicelli F, Weisstuch J, et al. Quality and safety outcomes of a hospital merger following a full integration at a safety net hospital. JAMA Netw Open. 2022;5(1):e2142382. Rhee DW, Pendse J, Chan H, Stern DT, Sartori DJ. Mapping the clinical experience of a new york city residency program during the COVID-19 pandemic. J Hosp Med. 2021;16(6):353–6. Graber I, John M. Eisenberg Patient Safety and Quality Awards: An Interview with Gordon D. Schiff. Jt Comm J Qual Patient Saf. 2020;46(7):371 – 80. Schiff GD. Diagnosis and diagnostic errors: time for a new paradigm. BMJ Qual Saf. 2014;23(1):1–3. National Board of Medical Examiners. 2022 Stemmler Grant Projects. Development and Validation of a Machine Learning Model for Automated Workplace-Based Assessment of Resident Clinical Reasoning Documentation. 2022. https://www.nbme.org/sites/default/files/2024-03/2022_NBME_Stemmler%20_Grants_Projects.pdf . Accessed 15 May 2024. Tables Tables 1 and 2 are available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files Tables.docx SupplementalDigitalAppendix.docx Cite Share Download PDF Status: Published Journal Publication published 22 Apr, 2025 Read the published version in BMC Medical Education → Version 1 posted Editorial decision: Revision requested 07 Mar, 2025 Reviews received at journal 24 Feb, 2025 Reviewers agreed at journal 19 Feb, 2025 Reviewers agreed at journal 19 Feb, 2025 Reviews received at journal 31 Jul, 2024 Reviewers agreed at journal 09 Jul, 2024 Reviewers agreed at journal 22 May, 2024 Reviewers invited by journal 16 May, 2024 Editor assigned by journal 16 May, 2024 Editor invited by journal 16 May, 2024 Submission checks completed at journal 16 May, 2024 First submitted to journal 15 May, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4427373","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":306047634,"identity":"b71d424f-40b6-4b7c-8c03-f641ed8e2a9e","order_by":0,"name":"Verity Schaye","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABL0lEQVRIie3RPUvEMBwG8JRIukS6ptSXr1ApRAWh+E3apbfcJhwdBAtCXDznHop+hXPpXAm0S3G+4QaPg+J4BTk6KJhk0KWFjoJ5SJ4h8PunpADo6PzR5KqJrFhscwjJf0glNhx8jSQGG0DccrriDQO+dX9dk+ZxGc45rFcf8fIAmNP1exepSjd/YSBMlwW1Z1ktCDr29qraA7ikJx3ETiMgSeCSgDq7GRcEUMdmPExIhNwu8lQr4rtktHW+HiQxt5Jc9RGLIEWMORlTx0gkwdRuGA8AieBbF8Hiw6pXEqaL8cXpbcG9GccTB1T8iOECdb0YMgu4iSdnvpWOnhftJd+/K28yu435oWUyuOl7aQOp/7gj+zwRBbGcJhbpI+BTtZrpqxkt+D3R0dHR+ff5BmeTawOUZhkFAAAAAElFTkSuQmCC","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":true,"prefix":"","firstName":"Verity","middleName":"","lastName":"Schaye","suffix":""},{"id":306047635,"identity":"d8297d7d-dc29-4d36-ac90-1f3ee7367107","order_by":1,"name":"David J DiTullio","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"J","lastName":"DiTullio","suffix":""},{"id":306047636,"identity":"2c10fbd4-d0eb-471b-b6c9-9b669f903d50","order_by":2,"name":"Daniel J Sartori","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"J","lastName":"Sartori","suffix":""},{"id":306047637,"identity":"4d93e685-4f17-440d-b7e7-419ddaad660c","order_by":3,"name":"Kevin Hauck","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Kevin","middleName":"","lastName":"Hauck","suffix":""},{"id":306047638,"identity":"fb8cb713-0240-446e-b355-488b7c2aa884","order_by":4,"name":"Matthew Haller","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"","lastName":"Haller","suffix":""},{"id":306047639,"identity":"845dea0e-9088-4325-8ca3-9099efc295ad","order_by":5,"name":"Ilan Reinstein","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Ilan","middleName":"","lastName":"Reinstein","suffix":""},{"id":306047640,"identity":"6309476f-c81a-4ebe-a27c-c6faaffe54a2","order_by":6,"name":"Benedict Guzman","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Benedict","middleName":"","lastName":"Guzman","suffix":""},{"id":306047641,"identity":"1ab1e0e9-d136-4fd7-8839-78be7d69e3ba","order_by":7,"name":"Jesse Burk-Rafel","email":"","orcid":"","institution":"New York University Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Jesse","middleName":"","lastName":"Burk-Rafel","suffix":""}],"badges":[],"createdAt":"2024-05-15 21:53:26","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4427373/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4427373/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12909-025-07191-x","type":"published","date":"2025-04-22T15:58:01+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":57448187,"identity":"3c9029f7-32f6-430f-ad65-b64b51fc9dad","added_by":"auto","created_at":"2024-05-30 19:54:20","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":97010,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA:\u003c/strong\u003e High-quality clinical reasoning documentation by academic year (AY) demonstrating a decline AY 2019-2020 driven by notes authored during the peak of the initial COVID-19 surge in New York City (March to June 2020).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB:\u003c/strong\u003e High-quality clinical Reasoning documentation by shift and hour demonstrating notes written during the night shift were more likely to be high-quality than those in the day shift and notes written later in the night shift appeared associated with lower note quality but this relationship disappeared in the multivariable regression.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC:\u003c/strong\u003e High-quality clinical reasoning documentation by note index within shift demonstrating that within each shift every additional note (i.e., higher note index) was associated with lower quality clinical reasoning documentation.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-4427373/v1/bb6d0d11167e3dc71b77666d.png"},{"id":81569872,"identity":"38a4f454-faea-4daf-937d-b0c581ffdedf","added_by":"auto","created_at":"2025-04-28 16:12:11","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":648651,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4427373/v1/a129123a-5fcb-45a1-96f2-f6928a863634.pdf"},{"id":57448188,"identity":"3f7db35e-be17-4d8d-9c89-0f0621cc4a3c","added_by":"auto","created_at":"2024-05-30 19:54:20","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":693948,"visible":true,"origin":"","legend":"","description":"","filename":"Tables.docx","url":"https://assets-eu.researchsquare.com/files/rs-4427373/v1/9a90265a8c766bdcb66aae09.docx"},{"id":57448189,"identity":"3d9cc466-5dd9-42cf-8424-b67d24a8b210","added_by":"auto","created_at":"2024-05-30 19:54:20","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":542624,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementalDigitalAppendix.docx","url":"https://assets-eu.researchsquare.com/files/rs-4427373/v1/f09fa9a666ce54b6fc07c84f.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Artificial Intelligence Based Assessment of Clinical Reasoning Documentation: An Observational Study of the Impact of the Clinical Learning Environment on Resident Performance","fulltext":[{"header":"Background","content":"\u003cp\u003eDirect patient care provides residents with clinical experiences central to the development of competence (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). However, learning in the authentic clinical environment can be challenging due to the unpredictability, pace, and complexity inherent to caring for sick patients. The clinical learning environment (CLE) is the interplay between the clinical context where trainees participate in patient care and education (including the expected learning outcomes and assessment practices) (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). The CLE is a component of residency accreditation and, as a mediator of attainment of clinical competence, has the potential for long-lasting impact on future practice patterns (\u003cspan additionalcitationids=\"CR5 CR6 CR7\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eStudies across different medical specialties have investigated how the CLE impacts resident performance using various outcomes. In the procedural specialties of obstetrics and gynecology and surgery, researchers have used procedural outcomes to examine the interplay between CLE and resident performance such as reviewing EHR data for maternal birth complications and postoperative complications (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). In internal medicine (IM), research around the CLE has focused on clinical workload, shift timing, and work hours with noted negative impacts on learning as assessed through trainee surveys; however, objective, attributable, and scalable measurement of such relationships proves difficult in practice (\u003cspan additionalcitationids=\"CR12\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Those studies that do focus on objective, EHR-derived outcomes have suggested that increased resident workload, as measured by higher patient censuses and numbers of admissions, is associated with increased resource use, readmission rates, and potentially increased patient mortality (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). Similarly, reducing resident workload by lowering patient census has been shown to result in an improvement in the quality of discharge summaries as assessed by manual human review (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). However, more automated and scalable data measures, including better utilization of EHR data, are needed to determine what specific, modifiable aspects of the CLE might be impacting resident performance and patient outcomes (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eArtificial intelligence (AI) and big data analytics offer approaches for generating the necessary scalable, large datasets of resident-specific, objective measures of resident performance that \u0026ndash; if melded with patient, team, and contextual factors \u0026ndash; could further unpack the potential impact of the CLE on resident performance (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). Despite this opportunity, AI-based assessments of resident performance are not yet widely implemented (\u003cspan additionalcitationids=\"CR20 CR21\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e). In earlier phases of this work, we developed an AI-based assessment of clinical reasoning documentation to increase the frequency and quality of feedback on this important skill (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). Clinical reasoning is a core skill that is essential to patient care and clinical documentation serves as an important measure of resident clinical reasoning (\u003cspan additionalcitationids=\"CR25\" citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e). While the absence of high-quality documentation of clinical reasoning does not preclude that reasoning was performed, higher-quality clinical reasoning documentation has been hypothesized to be associated with improved patient outcomes including reduced diagnostic errors (\u003cspan additionalcitationids=\"CR28\" citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFurthermore, the EHR and clinical documentation are integral to the CLE and can serve as an important objective measure of resident performance to better explore the impact of the CLE (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e). Historically, there have been limitations in using quality of clinical documentation as a more systemic measurement; while multiple human rating tools exist to assess the quality of clinical reasoning documentation, they can be time-intensive and therefore difficult to generate insights at scale (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan additionalcitationids=\"CR32\" citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e). AI-based automated assessment of clinical reasoning documentation enables the creation of nearly instantaneous large datasets measuring resident clinical reasoning documentation practices; such tools have the potential to assess clinical reasoning at the individual and aggregate level at a scale that is not feasible with human rating.\u003c/p\u003e \u003cp\u003eHere, we report a cross-sectional analysis of the relationship between the CLE (e.g., markers of workload such as volume of admissions and time of shift) and AI-based assessment of clinical reasoning documentation quality among a large retrospective cohort of IM residents at two distinct residency program hospital sites within one quaternary academic health system.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Sites\u003c/h2\u003e \u003cp\u003eThis study was conducted at NYU Grossman School of Medicine\u0026rsquo;s IM residency program, a large academic urban training program in New York City. The study was conducted at two of the residency program\u0026rsquo;s hospital sites\u0026mdash;NYU Langone Hospital (NYULH)-Manhattan and NYULH-Brooklyn\u0026mdash;which utilize a single EHR. The Manhattan site (181 residents overall) is a university-based hospital, quaternary referral and transplant center. These residents also rotate at two other affiliated sites with distinct EHRs not included in this study but do not rotate at the Brooklyn site. The Brooklyn site (41 residents overall) is a community-based tertiary hospital whose residency program became affiliated with NYU in 2016. At each site, residents staff 6 medicine inpatient teams. Residents are divided into day teams (7 am to 7 pm) and night teams (7 pm to 7 am), both responsible for caring for existing patients while also admitting patients daily, including writing admission notes. The day team structure typically comprises a supervising attending, a senior resident, two interns, and medical students. The night team is a resident-intern dyad.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eStudy Population\u003c/h2\u003e \u003cp\u003eWe included all categorical IM residents who wrote at least one admission note on a general IM service from July 1, 2018 \u0026ndash; June 30, 2023 (five academic years), at either NYULH-Manhattan (n\u0026thinsp;=\u0026thinsp;356 residents) or NYULH-Brooklyn (n\u0026thinsp;=\u0026thinsp;118 residents). Preliminary (e.g., Neurology) and rotating residents (e.g., Psychiatry) were excluded. Notes authored in the intensive care unit were excluded, as such notes were not included in the initial validation work of the AI algorithm (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eData Retrieval\u003c/h2\u003e \u003cp\u003eAdmission notes were retrieved using a custom query of the EHR data warehouse in July 2023. Note-level data included the note creation date and time. From this, the shift (day or night) and creation time within the shift were calculated. Additionally, a note index was calculated at the shift level indicating if a given note was the first, second, third, etc. note authored within a shift. These aspects of the CLE were included as potential markers of workload, which have been shown to have a strong relationship with resident performance (\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). Patient-level data at the time of each inpatient hospitalization was also captured, such as patient age, sex, insurance type, and Charlson Comorbidity Index (CCI). These variables were included as they are all patient characteristics that can influence the clinical reasoning process in the CLE (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eCCI encompasses 17 disease categories (e.g., congestive heart failure, renal disease) that were weighted using an age-independent adaptation revised for International Classification of Disease, 10th edition (ICD-10-CM) codes present before the date of hospital admission (score range 0 to 29, with higher indicating more comorbidity) (\u003cspan additionalcitationids=\"CR36\" citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eHospitalizations were categorized by diagnostic area using the ICD-10-CM chapter of hospitalization-level primary diagnoses, as done previously (Supplement Digital Appendix 1) (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e). Hospitalizations include a primary diagnosis (selected by a non-physician coder based on chart review) and a principal diagnosis (selected by clinicians). A chart review was conducted on a random sample of 100 hospitalizations to determine a rule-based approach for selecting the optimal diagnostic area. Primary diagnosis was used except when overly general codes were included (\u0026ldquo;sepsis, unspecified organism\u0026rdquo; or a sign/symptom ICD-10-CM code), where the principal diagnosis was used instead.\u003c/p\u003e \u003cp\u003eAuthor-level data on each resident\u0026rsquo;s primary training site and post-graduate year (PGY) at the time of note creation were retrieved from an educational data warehouse (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eAI Algorithm\u003c/h2\u003e \u003cp\u003eEach note was analyzed using a previously developed supervised machine learning model that classifies clinical reasoning documentation quality in IM resident admissions notes as low- or high-quality (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). The human rating gold standard for training the model was the Revised-DEA assessment tool which defines a high-quality hospital admission note in three domains as a note that has a differential diagnosis (D) that is explicitly prioritized with specific diagnoses (rather than diagnostic categories, e.g. infection) and provides clear explanation of reasoning (E) for the lead and alternative (A) diagnoses in the differential (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). The model uses the Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES, v4.0.0), a pre-large language model (LLM) open-source natural language processing (NLP) tool that was trained on clinical notes and performs named entity recognition (i.e., extracting diagnoses and other concepts). The algorithm has excellent note-level performance with an area under the receiver operating characteristic curve of 0.88 (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eDescriptive statistics, chi-squared testing, regression analyses, and visualizations were performed with the R statistical software (v4.2.0) using a variety of libraries (\u003cem\u003etidyverse\u003c/em\u003e, \u003cem\u003edata.table\u003c/em\u003e, \u003cem\u003egtsummary\u003c/em\u003e, \u003cem\u003eggplot2\u003c/em\u003e). Hypothesis testing was two-sided at α\u0026thinsp;=\u0026thinsp;.05. Regression models for high-quality clinical reasoning were built both without adjustment for clustering (logistic regression using \u003cem\u003eglm\u003c/em\u003e package of \u003cem\u003estats\u003c/em\u003e library) and with adjustment for clustering by resident using generalized estimating equations (GEE) with an exchangeable correlation structure (\u003cem\u003egeeglm\u003c/em\u003e package of \u003cem\u003egeepack\u003c/em\u003e library) and robust standard error estimates. Interaction terms were evaluated for each of the CLE variables; significant interactions were included in the final model. The within-cluster intraclass correlation was approximately 0.08, indicating only modest within-resident clustering of performance. The study was approved by the NYU Grossman School of Medicine institutional review board.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eDescriptive statistics were computed for resident characteristics, CLE characteristics, and patient characteristics stratified by residency program site (Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplement Digital Appendix 2) and post-graduate year (PGY) (Supplement Digital Appendix 3). There were 37,750 admission notes (16,180 Manhattan and 21,570 Brooklyn) authored by 474 distinct residents (356 Manhattan residents and 118 Brooklyn residents) with a mean 79.6 notes per resident (mean 45 per Manhattan resident and 183 per Brooklyn resident) across 28,782 distinct patients between both sites from July 2018 to June 2023. Overall, 55% of the notes were high-quality (64% high-quality at Manhattan and 48% high-quality at Brooklyn). At both sites PGY-2 or PGY-3 residents accounted for the majority of high-quality notes written as compared to PGY-1 residents (overall n\u0026thinsp;=\u0026thinsp;15,823 (76%) vs 5,034 (24%), P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Senior residents authored more high-quality notes at Brooklyn (PGY2 or PGY3, n\u0026thinsp;=\u0026thinsp;8,992 (86%) vs PGY1, n\u0026thinsp;=\u0026thinsp;1,448 (14%),\u003c/p\u003e\n\u003cp\u003eP\u0026thinsp;\u0026lt;\u0026thinsp;.001) compared to Manhattan (PGY2 or PGY3, n\u0026thinsp;=\u0026thinsp;6,831 (66%) vs PGY1 n\u0026thinsp;=\u0026thinsp;3,586 (34%), P\u0026thinsp;\u0026lt;\u0026thinsp;.001). At both sites more of the high-quality notes were written during the night shift as compared to day shift (overall n\u0026thinsp;=\u0026thinsp;13,518 (65%) vs n\u0026thinsp;=\u0026thinsp;7,339 (35%), P\u0026thinsp;\u0026lt;\u0026thinsp;.001) and notes with an earlier note index in the shift accounted for the majority of high-quality notes (overall 1st note in shift n\u0026thinsp;=\u0026thinsp;11,890 (57%) vs 2nd note in shift n\u0026thinsp;=\u0026thinsp;4,657 (22%) vs 3rd note in shift n\u0026thinsp;=\u0026thinsp;2,130 (10%) vs 4th note in shift n\u0026thinsp;=\u0026thinsp;1,119 (5.4%) vs 5th or more note in shift n\u0026thinsp;=\u0026thinsp;1,061 (5.1%), P\u0026thinsp;\u0026lt;\u0026thinsp;.001).\u003c/p\u003e\n\u003cp\u003eControlling for covariates and accounting for resident-level clustering, patient characteristics associated with complex admitting presentations were associated with higher quality clinical reasoning documentation, including older age (adjusted odds ratio, aOR 1.13 for age\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;80.1; 95% CI 1.05\u0026ndash;1.22; P\u0026thinsp;\u0026lt;\u0026thinsp;.001; aOR 1.17 for age 68.1\u0026ndash;80; 95% CI 1.09\u0026ndash;1.25; P\u0026thinsp;\u0026lt;\u0026thinsp;.001; aOR 1.05 for age 54.1\u0026ndash;68; 95% CI 0.99\u0026ndash;1.11; P\u0026thinsp;=\u0026thinsp;.10) and higher comorbidity index (aOR 2.18 for CCI\u0026thinsp;\u0026ge;\u0026thinsp;5 referenced to CCI 0; 95% CI 2.02\u0026ndash;2.34; P\u0026thinsp;\u0026lt;\u0026thinsp;.001) (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). The primary diagnosis areas most strongly associated with higher clinical reasoning documentation quality included pulmonary and infectious diseases (aOR 1.53; 95% CI 1.40\u0026ndash;1.66; P\u0026thinsp;\u0026lt;\u0026thinsp;.001; aOR 1.58; 95% CI 1.45\u0026ndash;1.73; P\u0026thinsp;\u0026lt;\u0026thinsp;.001, respectively).\u003c/p\u003e\n\u003cp\u003eSeveral aspects of the CLE were significantly associated with clinical reasoning documentation quality (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). Academic year was positively associated with clinical reasoning documentation quality (aOR 1.10; 95% CI 1.06\u0026ndash;1.15; P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Notably, there was a decline in clinical reasoning documentation quality during AY 2019\u0026ndash;2020 driven by notes authored during the peak of the initial COVID-19 surge in New York City (March to June 2020), with subsequent recovery in documentation quality (47.7% high-quality AY 2019\u0026ndash;2020 (49.9% high-quality July-February, 42.4% high-quality March-June) vs 56.8% high-quality all other AYs, P\u0026thinsp;\u0026lt;\u0026thinsp;.001, Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea). The NYULH-Manhattan site experienced a bigger decline in clinical reasoning documentation quality during the COVID-19 surge than the NYULH-Brooklyn site (62.2% high-quality July-February AY 2019\u0026ndash;2020 to 50.5% high-quality March-June AY 2019\u0026ndash;2020 at Manhattan vs 44.5% high-quality July-February AY 2019\u0026ndash;2020 to 37.0% high-quality March-June AY 2019\u0026ndash;2020 at Brooklyn, P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Overall, performance differed by site: those at NYULH-Brooklyn were less likely to write high-quality notes than at NYULH-Manhattan (aOR 0.46; 95% CI 0.39\u0026ndash;0.55; P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Additionally, only at the NYULH-Brooklyn site was an increase in clinical reasoning documentation quality observed by PGY. PGY 2 and PGY3 residents (i.e., senior residents) at the NYULH-Brooklyn site were more likely to write high-quality clinical reasoning notes as compared to their PGY1 peers (aOR 1.39; 95% CI 1.17\u0026ndash;1.64; P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Sensitivity analysis with PGY2 and PGY3 separated showed the same result.\u003c/p\u003e\n\u003cp\u003eNotes written during the night shift were more likely to be high-quality than those in the day shift (aOR 1.21; 95% CI 1.13\u0026ndash;1.30; P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Although notes written later in the night shift appeared associated with lower note quality (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eb), this relationship disappeared in the multivariable regression (aOR 1.00; 95% CI 0.99\u0026ndash;1.01; P\u0026thinsp;=\u0026thinsp;0.80). Rather, within each shift every additional note (i.e., higher note index) was associated with an 7% lower odds of high-quality clinical reasoning documentation (aOR 0.93; 95% CI 0.90\u0026ndash;0.95; P\u0026thinsp;\u0026lt;\u0026thinsp;.001) (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ec).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eWe utilized AI-based assessment of resident clinical reasoning documentation to measure resident performance across five academic years, two residency program sites, and tens of thousands of hospitalizations to better understand the relationship between the CLE and resident performance. The size of the datasets generated and this exploration was only feasible due to the power of AI. A single person working continuously for eight hours per day spending two minutes per note would have taken over 150 days to rate all notes. Furthermore, no other validated AI-based assessment tools of clinical reasoning documentation in the authentic CLE have been described (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan additionalcitationids=\"CR41 CR42 CR43\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eMelding EHR-derived measures of the CLE with this AI-based assessment revealed clinically meaningful performance differences by several CLE factors. We extracted note-level data of creation time within shift and note index to explore markers of workload which can be a major barrier to learning in the CLE (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). We found that residents were more likely to write high-quality notes during their night shifts and less likely to write high-quality notes with an increasing number of notes within a shift. For the former, we hypothesize that note quality was higher during the night shift as residents typically can focus more on new admissions without the workload of routine non-emergent care at night. While there is consensus that busyness of work and service pressures can cause cognitive overload and be a barrier to learning and performance in the CLE, both findings provide objective measures that could inform structural changes (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e). Such structural changes could include further limits to the number of admissions than already exists (\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e) or team structures that could minimize cognitive overload like the addition of advanced practice providers (APPs) to support resident teams or the creation of swing shifts to focus on admitting without cross-coverage obligations (\u003cspan additionalcitationids=\"CR48\" citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e). These changes would need to be balanced with the tradeoffs including requirements for increased staffing and ensuring residents having enough clinical exposure to develop competence.\u003c/p\u003e \u003cp\u003eUnsurprisingly the major disruptor to the CLE of the COVID-19 surge in New York City in the spring of 2020 impacted the educational outcome of clinical reasoning documentation quality at both sites. NYULH-Manhattan experienced a bigger impact than NYULH-Brooklyn which correlates with the higher patient volume experienced at this site during the surge (\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e). At both sites clinical reasoning documentation quality has since recovered. While clinical reasoning documentation quality at NYULH-Manhattan has mostly plateaued recently, NYULH-Brooklyn experienced a continued significant rise over the ensuing two academic years with a leveling off in the most recent AY 2022\u0026ndash;2023. We hypothesize this is a result of a major change in the CLE that occurred in 2016 with the merger of NYULH with Lutheran Medical Center which is now NYULH-Brooklyn. A similar trend in improvement was seen in clinical outcomes (such as reduction in mortality rate and fewer central line and catheter-associated urinary tract infections) after the merger as described by Wang et al (\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e). Using aggregate performance of AI-based assessment of clinical reasoning documentation can be an important strategy to assess for any unintended or intended impact of systems changes like mergers or EHR changes.\u003c/p\u003e \u003cp\u003eWhile likely not an impact of the CLE, another difference seen in educational outcomes between NYULH-Brooklyn and NYULH-Manhattan was an increase in clinical reasoning documentation quality by PGY year only at the former site. Explanations for this might be a ceiling effect at NYULH-Manhattan where the PGY1 residents started at a higher percentage of high-quality notes (66%) than NYULH-Brooklyn (42%). Despite these differences, both sites experienced similar impacts of the CLE when controlling for covariates.\u003c/p\u003e \u003cp\u003eLastly, we explored clinical reasoning documentation quality by primary diagnosis area to further understand if residents were meeting educational outcomes by content domain and the potential impact clinical context might have. After controlling for patient characteristics, including preexisting comorbidities, we found that residents were more likely to write high-quality clinical reasoning documentation in the primary diagnosis domains of infectious disease and pulmonary. There could be several explanations for this finding such as increased exposure to these diagnoses in particular infectious diseases during COVID-19 (\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e), residents having the opportunity to rotate on pulmonary and infectious disease subspecialty-led teams, or certain presenting diagnoses not lending themselves to requiring as much diagnostic reasoning (e.g. a hematology-oncology patient admitted for chemotherapy). Ultimately, these data could be used both for feedback for individual resident practice habits or rotation planning, scheduling residents for specialty rotations, and for improvements at the program level to inform curricular planning, such as focusing on diagnostic categories with overall lower quality clinical reasoning in noon conference series. Use for feedback on practice habits could be applicable not only to residents but also to attending physicians and APPs, and has the potential to reduce diagnostic errors and improve patient outcomes by improving clinical reasoning documentation quality (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e, \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e). Validating the AI-based assessment tool for attending physician and APP clinical reasoning documentation quality and exploring the relationship between clinical reasoning documentation quality and diagnostic errors are both important next areas to explore with future studies.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eThe AI model in this study was developed with older generation NLP technology, however it did have good performance (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). LLMs will only further improve the ability to analyze EHR-based note text, which we are currently integrating into an updated model. Additionally, in the current phase of our work, we are validating the AI-based assessment of clinical reasoning documentation at a second institution, University of Cincinnati, and creating a pipeline for dissemination of the tool. The current phase of the work will enhance generalizability of this study allowing other clinical sites to explore the potential impact of CLE on resident performance using AI-based assessment of clinical reasoning documentation (\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e). Additionally, in terms of generalizability, although this study was conducted at a single academic institution, it was at two distinct hospital sites that have distinct CLEs. Another limitation to consider is that a resident facing dashboard to provide feedback on clinical reasoning documentation generated from output of the AI-based assessment was implemented during the study period in November 2020 (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). This could have had some influence on resident clinical reasoning documentation practices in the later part of the study period in particular accounting for the trend in improvement seen by academic year which is an ideal outcome after implementing feedback via the dashboard. Lastly, while we controlled for many key variables we determined could influence clinical reasoning documentation quality in the CLE, residual confounding may exist and the retrospective cross-sectional design precludes assessment of causality.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn this retrospective cross-sectional analysis of IM resident clinical reasoning documentation quality, we demonstrated how AI-based assessment of resident performance derived from the EHR can provide insights to unpack the impact of the CLE. We found several specific factors in the CLE associated with clinically meaningful differences in resident performance. These findings can serve as a call to action to further expand to other institutions, specialties, and healthcare providers in future studies and inform important recommended systems changes.\u003c/p\u003e"},{"header":"List Of Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial Intelligence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAPP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAdvanced Practice Provider\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCLE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eClinical Learning Environment\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eEHR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eElectronic Health Record\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eIM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInternal Medicine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLLMs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLarge Language Models\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNLP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNatural Language Processing\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNYULH\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNYU Langone Health\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval:\u0026nbsp;\u003c/strong\u003eThe study was approved by the NYU Grossman School of Medicine institutional review board on 12/9/2023 i19-00280. As this was a retrospective, observational study of EHR data review informed consent from each participant was waived for model development and retrospective data analysis by the NYU Grossman School of Medicine Institutional Review Board.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials:\u0026nbsp;\u003c/strong\u003eGiven the sensitive nature of EHR data, additional data cannot be readily accessible. Please email the authors [email protected] with any inquiries about the data or machine learning model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests:\u0026nbsp;\u003c/strong\u003eNo authors have conflicts of interest to disclose relevant to this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u0026nbsp;\u003c/strong\u003eNone\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions:\u0026nbsp;\u003c/strong\u003eVS made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; DT made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; DS made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; KH made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; MH made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work; IR made substantial contributions to the data acquisition and analysis; BG made substantial contributions to the data acquisition and analysis; JBR made substantial contributions to the conception of the work, interpretation of data, draft and revisions of the work, and the data acquisition and analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u0026nbsp;\u003c/strong\u003eThe authors would like to acknowledge Helen Finkelstein, data warehouse and business intelligence developer at the Institute for Innovations in Medical Education, for her assistance in calculating the CCI.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTeunissen P, Scheele F, Scherpbier A, Van Der Vleuten C, Boor K, Van Luijk S, et al. How residents learn: qualitative evidence for the pivotal role of clinical activities. Med Educ. 2007;41(8):763\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNordquist J, Hall J, Caverzagie K, Snell L, Chan MK, Thoma B, et al. The clinical learning environment. Med Teach. 2019;41(4):366\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGruppen LD. Context and complexity in the clinical learning environment. Med Teach. 2019;41(4):373\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeiss KB, Bagian JP, Nasca TJ. The clinical learning environment: the foundation of graduate medical education. JAMA. 2013;309(16):1687\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWagner R, Patow C, Newton R, Casey BR, Koh NJ, Weiss KB. The Overview of the CLER program: CLER national report of findings 2016. J Grad Med Educ. 2016;8(2 Suppl 1):11\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNasca TJ, Wagner R, Weiss KB. Introduction to the CLER national report of findings 2022: The COVID-19 pandemic and its impact on the clinical learning environment. J Grad Med Educ. 2023;15(1):140\u0026ndash;2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen C, Petterson S, Phillips R, Bazemore A, Mullan F. Spending patterns in region of residency training and subsequent expenditures for care provided by practicing physicians for Medicare beneficiaries. JAMA. 2014;312(22):2385\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAsch DA, Epstein A, Nicholson S. Evaluating medical training programs by the quality of care delivered by their alumni. JAMA. 2007;298(9):1049\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAsch DA, Nicholson S, Srinivas S, Herrin J, Epstein AJ. Evaluating obstetrical residency programs using patient outcomes. JAMA. 2009;302(12):1277\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBansal N, Simmons KD, Epstein AJ, Morris JB, Kelz RR. Using patient outcomes to evaluate general surgery residency program performance. JAMA Surg. 2016;151(2):111\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKilty C, Wiese A, Bergin C, Flood P, Fu N, Horgan M, et al. A national stakeholder consensus study of challenges and priorities for clinical learning environments in postgraduate medical education. BMC Med Educ. 2017;17(1):226.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaney EM, Nicolaidis C, Hunter A, Chan BK, Cooney TG, Bowen JL. Relationship between resident workload and self-perceived learning on inpatient medicine wards: a longitudinal study. BMC Med Educ. 2006;6:35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArora VM, Georgitis E, Siddique J, Vekhter B, Woodruff JN, Humphrey HJ, et al. Association of workload of on-call medical interns with on-call sleep duration, shift duration, and participation in educational activities. JAMA. 2008;300(10):1146\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOng M, Bostrom A, Vidyarthi A, McCulloch C, Auerbach A. House staff team workload and organization effects on patient outcomes in an academic general internal medicine inpatient service. Arch Intern Med. 2007;167(1):47\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAverbukh Y, Southern W. The impact of the number of admissions to the inpatient medical teaching team on patient safety outcomes. J Grad Med Educ. 2012;4(3):307\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCoit MH, Katz JT, McMahon GT. The effect of workload reduction on the quality of residents' discharge summaries. J Gen Intern Med. 2011;26(1):28\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBurk-Rafel J, Sebok-Syer SS, Santen SA, Jiang J, Caretta-Weyer HA, Iturrate E, et al. TRainee attributable \u0026amp; automatable care evaluations in real-time (TRACERs): A scalable approach for linking education to patient care. Perspect Med Educ. 2023;12(1):149.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArora VM. Harnessing the power of big data to improve graduate medical education: big idea or bust? Acad Med. 2018;93(6):833\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTurner L, Hashimoto D, Vasisht S, Schaye V, Demystifying AI. Current state and future role in precision medical education assessment. Acad Med. 2024;99(4S Suppl 1):S42\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoscardin CK, Gin B, Golde PB, Hauer KE. ChatGPT and generative artificial intelligence for medical education: potential impact and opportunity. Acad Med. 2023:101097.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, et al. Large language models in medical education: Opportunities, challenges, and future directions. JMIR Med Educ. 2023;9(1):e48291.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePreiksaitis C, Rose C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: Scoping review. JMIR Med Educ. 2023;9(1):e48785.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchaye V, Guzman B, Burk-Rafel J, Marin M, Reinstein I, Kudlowitz D, et al. Development and validation of a machine learning model for automated assessment of resident clinical reasoning documentation. J Gen Intern Med. 2022;37(9):2230\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eConnor DM, Durning SJ, Rencic JJ. Clinical reasoning as a core competency. Acad Med. 2020;95(8):1166\u0026ndash;71.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchaye V, Miller L, Kudlowitz D, Chun J, Burk-Rafel J, Cocks P, et al. Development of a clinical reasoning documentation assessment tool for resident and fellow admission notes: a shared mental model for feedback. J Gen Intern Med. 2022;37(3):507\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaker EA, Ledford CH, Fogg L, Way DP, Park YS. The IDEA assessment tool: Assessing the reporting, diagnostic reasoning, and decision-making skills demonstrated in medical students' hospital admission notes. Teach Learn Med. 2015;27(2):163\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKulkarni D, Heath J, Kosack A, Jackson NJ, Crummey A. An educational intervention to improve inpatient documentation of high-risk diagnoses by pediatric residents. Hosp Pediatr. 2018;8(7):430\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchiff GD, Bates DW. Can electronic clinical documentation help prevent diagnostic errors? N Engl J Med. 2010;362(12):1066\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSingh H, Giardina TD, Meyer AN, Forjuoh SN, Reis MD, Thomas EJ. Types and origins of diagnostic errors in primary care settings. JAMA Intern Med. 2013;173(6):418\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThoma B, Turnquist A, Zaver F, Hall AK, Chan TM. Communication, learning and assessment: Exploring the dimensions of the digital learning environment. Med Teach. 2019;41(4):385\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHung H, Kueh LL, Tseng CC, Huang HW, Wang SY, Hu YN, et al. Assessing the quality of electronic medical records as a platform for resident education. BMC Med Educ. 2021;21(1):577.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKogan JR, Hess BJ, Conforti LN, Holmboe ES. What drives faculty ratings of residents' clinical skills? The impact of faculty's own clinical skills. Acad Med. 2010;85(10 Suppl):S25\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEdwards ST, Neri PM, Volk LA, Schiff GD, Bates DW. Association of note quality and quality of care: a cross-sectional study. BMJ Qual Saf. 2014;23(5):406\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStolper E, Van Royen P, Jack E, Uleman J, Olde Rikkert M. Embracing complexity with systems thinking in general practitioners' clinical reasoning helps handling uncertainty. J Eval Clin Pract. 2021;27(5):1175\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQuan H, Sundararajan V, Halfon P, Fong A, Burnand B, Luthi J-C et al. Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data. Med Care. 2005:1130\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeyo RA, Cherkin DC, Ciol MA. Adapting a clinical comorbidity index for use with ICD-9-CM administrative databases. J Clin Epidemiol. 1992;45(6):613\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCharlson ME, Pompei P, Ales KL, MacKenzie CR. A new method of classifying prognostic comorbidity in longitudinal studies: development and validation. J Chronic Dis. 1987;40(5):373\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK\u0026ouml;nig S, Pellissier V, Hohenstein S, Leiner J, Hindricks G, Meier-Hellmann A, et al. A comparative analysis of in-hospital mortality per disease groups in Germany before and during the COVID-19 pandemic from 2016 to 2020. JAMA Netw Open. 2022;5(2):e2148649\u0026ndash;e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTriola MM, Pusic MV. The education data warehouse: a transformative tool for health education research. J Grad Med Educ. 2012;4(1):113\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin SY, Shanafelt TD, Asch SM. Reimagining clinical documentation with artificial intelligence. Mayo Clin Proc. 2018;93(5):563-5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalt J, Harik P, Barone MA. Leveraging natural language processing: Toward computer-assisted scoring of patient notes in the USMLE Step 2 clinical skills exam. Acad Med. 2019;94(3):314\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKanjee Z, Crowe B, Rodman A. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA. 2023;330(1):78\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStrong E, DiGiammarino A, Weng Y, Kumar A, Hosamani P, Hom J, et al. Chatbot vs medical student performance on free-response clinical reasoning examinations. JAMA Intern Med. 2023;183(9):1028\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu J, Wang C, Liu S. Utility of ChatGPT in clinical practice. J Med Internet Res. 2023;25:e48568.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFletcher KE, Reed DA, Arora VM. Doing the dirty work: measuring and optimizing resident workload. J Gen Intern Med. 2011;26(1):8\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAccreditation Council for Graduate Medical Education. ACGME Program Requirements for Graduate Medical Education in Internal Medicine. 2022. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.acgme.org/globalassets/pfassets/programrequirements/140_internalmedicine_2023.pdf\u003c/span\u003e\u003cspan address=\"https://www.acgme.org/globalassets/pfassets/programrequirements/140_internalmedicine_2023.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 15 May 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThanarajasingam U, McDonald FS, Halvorsen AJ, Naessens JM, Cabanela RL, Johnson MG et al. Service census caps and unit-based admissions: resident workload, conference attendance, duty hour compliance, and patient safety. Mayo Clin Proc. 2012;87(4):320-7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChandra R, Farah F, Munoz-Lobato F, Bokka A, Benedetti KL, Brueggemann C, et al. Sleep is required to consolidate odor memory and remodel olfactory synapses. Cell. 2023;186(13):2911\u0026ndash;e2820.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu AT, Jepsen N, Prasad S, Klein JP, Doughty C. Adding nocturnal advanced practice providers to an academic inpatient neurology service improves residents\u0026rsquo; educational experience. Neurohospitalist. 2023;13(2):130\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchaye VE, Reich JA, Bosworth BP, Stern DT, Volpicelli F, Shapiro NM et al. Collaborating across private, public, community, and federal hospital systems: Lessons learned from the Covid-19 pandemic response in NYC. NEJM Catal Innov Care Deliv. 2020;1(6).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang E, Arnold S, Jones S, Zhang Y, Volpicelli F, Weisstuch J, et al. Quality and safety outcomes of a hospital merger following a full integration at a safety net hospital. JAMA Netw Open. 2022;5(1):e2142382.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRhee DW, Pendse J, Chan H, Stern DT, Sartori DJ. Mapping the clinical experience of a new york city residency program during the COVID-19 pandemic. J Hosp Med. 2021;16(6):353\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGraber I, John M. Eisenberg Patient Safety and Quality Awards: An Interview with Gordon D. Schiff. Jt Comm J Qual Patient Saf. 2020;46(7):371\u0026thinsp;\u0026ndash;\u0026thinsp;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchiff GD. Diagnosis and diagnostic errors: time for a new paradigm. BMJ Qual Saf. 2014;23(1):1\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNational Board of Medical Examiners. 2022 Stemmler Grant Projects. Development and Validation of a Machine Learning Model for Automated Workplace-Based Assessment of Resident Clinical Reasoning Documentation. 2022. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.nbme.org/sites/default/files/2024-03/2022_NBME_Stemmler%20_Grants_Projects.pdf\u003c/span\u003e\u003cspan address=\"https://www.nbme.org/sites/default/files/2024-03/2022_NBME_Stemmler%20_Grants_Projects.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 15 May 2024.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTables 1 and 2 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Artificial Intelligence, Clinical Reasoning, Documentation, Clinical Learning Environment","lastPublishedDoi":"10.21203/rs.3.rs-4427373/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4427373/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eObjective measures and large datasets are needed to determine aspects of the Clinical Learning Environment (CLE) impacting resident performance. Artificial Intelligence (AI) offers a solution. Here, the authors sought to determine what aspects of the CLE might be impacting resident performance as measured by clinical reasoning documentation quality assessed by AI.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this observational, retrospective cross-sectional analysis of hospital admission notes from the Electronic Health Record (EHR), all categorical internal medicine (IM) residents who wrote at least one admission note during the study period July 1, 2018 – June 30, 2023 at two sites of NYU Grossman School of Medicine’s IM residency program were included.\u003cstrong\u003e \u003c/strong\u003eClinical reasoning documentation quality of admission notes was determined to be low or high-quality using a supervised machine learning model. From note-level data, the shift (day or night) and note index within shift (if a note was first, second, etc. within shift) were calculated. These aspects of the CLE were included as potential markers of workload, which have been shown to have a strong relationship with resident performance. Patient data was also captured, including age, sex, Charlson Comorbidity Index, and primary diagnosis. The relationship between these variables and clinical reasoning documentation quality was analyzed using generalized estimating equations accounting for resident-level clustering.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAcross 37,750 notes authored by 474 residents, patients who were older, had more pre-existing comorbidities, and presented with certain primary diagnoses (e.g., infectious and pulmonary conditions) were associated with higher clinical reasoning documentation quality. When controlling for these and other patient factors, variables associated with clinical reasoning documentation quality included academic year (adjusted odds ratio, aOR, for high-quality: 1.10; 95% CI 1.06-1.15; \u003cem\u003eP\u003c/em\u003e\u0026lt;.001), night shift (aOR 1.21; 95% CI 1.13-1.30; \u003cem\u003eP\u003c/em\u003e\u0026lt;.001), and note index (aOR 0.93; 95% CI 0.90-0.95; \u003cem\u003eP\u003c/em\u003e\u0026lt;.001).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAI can be used to assess complex skills such as clinical reasoning in authentic clinical notes that can help elucidate the potential impact of the CLE on resident performance. Future work should explore residency program and systems interventions to optimize the CLE.\u003c/p\u003e","manuscriptTitle":"Artificial Intelligence Based Assessment of Clinical Reasoning Documentation: An Observational Study of the Impact of the Clinical Learning Environment on Resident Performance","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-05-30 19:54:15","doi":"10.21203/rs.3.rs-4427373/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-03-07T09:13:02+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-02-24T17:27:05+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"125100816479020451436605981000282409653","date":"2025-02-19T18:34:52+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"252796152866929316320259524280145294995","date":"2025-02-19T14:45:31+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-07-31T22:14:49+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"182433191441227315755818582195184607258","date":"2024-07-09T14:02:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"291870109680706282778016453768959913634","date":"2024-05-22T20:41:15+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-05-16T14:23:39+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-05-16T14:17:44+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-05-16T12:00:49+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-05-16T11:58:17+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Education","date":"2024-05-15T21:49:14+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"96727b67-6425-4ab4-945c-a18ef9eb3243","owner":[],"postedDate":"May 30th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-04-28T16:06:09+00:00","versionOfRecord":{"articleIdentity":"rs-4427373","link":"https://doi.org/10.1186/s12909-025-07191-x","journal":{"identity":"bmc-medical-education","isVorOnly":false,"title":"BMC Medical Education"},"publishedOn":"2025-04-22 15:58:01","publishedOnDateReadable":"April 22nd, 2025"},"versionCreatedAt":"2024-05-30 19:54:15","video":"","vorDoi":"10.1186/s12909-025-07191-x","vorDoiUrl":"https://doi.org/10.1186/s12909-025-07191-x","workflowStages":[]},"version":"v1","identity":"rs-4427373","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4427373","identity":"rs-4427373","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00