Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

IMPORTANCE Large language models (LLMs) have been investigated for clinical documentation, with concerns about hallucinations and factual errors. Clinician review and revision of LLM-generated drafts are therefore considered essential, yet the impact of such workflow on both documentation time and quality remains unknown. OBJECTIVE To assess whether physician review and editing of LLM-generated drafts improves the time and quality of clinical documentation compared with clinician-only drafting in a randomized controlled trial setting. DESIGN Single-center, parallel-group, prospective, randomized, open-label, blinded-endpoint (PROBE) pilot trial conducted from February 18 to March 14, 2025. SETTING Kyoto University Hospital, Department of Ophthalmology. PARTICIPANTS Twenty-one ophthalmology physicians were randomized; 17 completed the study, and 4 withdrew before initiating intervention. INTERVENTIONS Participants were randomized to either the Clinician-in-the-loop group or the Clinician-only group. All participants created discharge summaries and referrals for six simulated patient records. In the Clinician-in-the-loop group, drafts were generated with an LLM assistant and then reviewed and edited by participants, whereas in the Clinician-only group, documents were drafted from scratch using matched templates. Unedited LLM drafts were additionally analyzed as the LLM-only group. MAIN OUTCOMES AND MEASURES Document creation time (primary) and expert-rated document quality across six domains plus overall quality (secondary). RESULTS Seventeen physicians submitted 48 discharge summaries and 48 discharge referrals in the Clinician-in-the-loop group, 54 of each document type in the Clinician-only group, and 48 of each in the LLM-only group. For summaries, Clinician-in-the-loop was associated with shorter creation time versus Clinician-only (β = −59.4 seconds; 95% CI, −118.1 to −0.8; P = .047). For referrals, clinician-in-the-loop required more time (β = 94.8 seconds; 95% CI, 40.4 to 149.3; P<.001). In most quality domains for both document types, the clinician-in-the-loop workflow outperformed clinician-only drafting. LLM-only drafts were fastest but had the lowest quality. CONCLUSIONS AND RELEVANCE A clinician-in-the-loop approach improved document quality and accelerated documentation. Active clinician review of LLM-generated drafts is essential for clinical documentation, and such workflows may help enhance working conditions and patient care. TRIAL REGISTRATION ClinicalTrials.gov Identifier NCT07187050 Key Points Question Does a workflow in which clinicians review and edit LLM-generated drafts (“clinician-in-the-loop”) improve the efficiency and quality of clinical documentation compared with clinician-only drafting? Findings In this single-center randomized controlled trial including 17 ophthalmology physicians, discharge summaries were significantly faster with the clinician-in-the-loop. Across both document types, clinician-in-the-loop drafts achieved higher quality scores than clinician-only documents. Meaning A clinician-in-the-loop workflow using LLMs can simultaneously enhance the efficiency and quality of clinical documentation; accordingly, active clinician review remains essential for improving working conditions and patient care overall.
Full text 53,245 characters · extracted from preprint-html · click to expand
Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality | medRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-P4HH5NV'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality View ORCID Profile Tomohiro Takayama , View ORCID Profile Keina Sado , View ORCID Profile Kenji Suda , View ORCID Profile Hiroshi Tamura , Naoko Ueda-Arakawa , Kenji Ishihara , Yuta Ueda , Kazuya Okamoto , Luciano Henrique de Oliveira Santos , Tetsuro Oshika , View ORCID Profile Tomohiro Kuroda , View ORCID Profile Akitaka Tsujikawa , View ORCID Profile Masahiro Miyake doi: https://doi.org/10.1101/2025.10.06.25337211 Tomohiro Takayama 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan 2 Kyoto University Hospital , Kyoto, Japan M.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Tomohiro Takayama Keina Sado 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan M.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Keina Sado Kenji Suda 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan M.D., Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Kenji Suda Hiroshi Tamura 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan 3 Center for Innovative Research and Education in Data Science, Institute for Liberal Arts and Sciences, Kyoto University , Kyoto, Japan M.D., Ph.D., ScM. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Hiroshi Tamura Naoko Ueda-Arakawa 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan M.D., Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site Kenji Ishihara 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan M.D., Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site Yuta Ueda 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan M.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site Kazuya Okamoto 4 Fitting Cloud Inc. , Kyoto, Japan Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site Luciano Henrique de Oliveira Santos 4 Fitting Cloud Inc. , Kyoto, Japan Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site Tetsuro Oshika 5 Japanese Society of Artificial Intelligence in Ophthalmology , 3-10-5, Nihonbashi, Chuoku, Tokyo, 1038276, Japan 6 Department of Ophthalmology, Faculty of Medicine, University of Tsukuba , Ibaraki, Japan M.D., Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site Tomohiro Kuroda 7 Division of Medical Information Technology and Administration Planning, Kyoto University Hospital, Kyoto University , Kyoto, Japan Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Tomohiro Kuroda Akitaka Tsujikawa 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan M.D., Ph.D. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Akitaka Tsujikawa Masahiro Miyake 1 Department of Ophthalmology and Visual Sciences, Kyoto University Graduate School of Medicine , Kyoto, Japan 5 Japanese Society of Artificial Intelligence in Ophthalmology , 3-10-5, Nihonbashi, Chuoku, Tokyo, 1038276, Japan M.D., Ph.D., M.P.H. Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Masahiro Miyake For correspondence: miyakem{at}kuhp.kyoto-u.ac.jp Abstract Full Text Info/History Metrics Supplementary material Data/Code Preview PDF Abstract IMPORTANCE Large language models (LLMs) have been investigated for clinical documentation, with concerns about hallucinations and factual errors. Clinician review and revision of LLM-generated drafts are therefore considered essential, yet the impact of such workflow on both documentation time and quality remains unknown. OBJECTIVE To assess whether physician review and editing of LLM-generated drafts improves the time and quality of clinical documentation compared with clinician-only drafting in a randomized controlled trial setting. DESIGN Single-center, parallel-group, prospective, randomized, open-label, blinded-endpoint (PROBE) pilot trial conducted from February 18 to March 14, 2025. SETTING Kyoto University Hospital, Department of Ophthalmology. PARTICIPANTS Twenty-one ophthalmology physicians were randomized; 17 completed the study, and 4 withdrew before initiating intervention. INTERVENTIONS Participants were randomized to either the Clinician-in-the-loop group or the Clinician-only group. All participants created discharge summaries and referrals for six simulated patient records. In the Clinician-in-the-loop group, drafts were generated with an LLM assistant and then reviewed and edited by participants, whereas in the Clinician-only group, documents were drafted from scratch using matched templates. Unedited LLM drafts were additionally analyzed as the LLM-only group. MAIN OUTCOMES AND MEASURES Document creation time (primary) and expert-rated document quality across six domains plus overall quality (secondary). RESULTS Seventeen physicians submitted 48 discharge summaries and 48 discharge referrals in the Clinician-in-the-loop group, 54 of each document type in the Clinician-only group, and 48 of each in the LLM-only group. For summaries, Clinician-in-the-loop was associated with shorter creation time versus Clinician-only (β = −59.4 seconds; 95% CI, −118.1 to −0.8; P = .047). For referrals, clinician-in-the-loop required more time (β = 94.8 seconds; 95% CI, 40.4 to 149.3; P<.001). In most quality domains for both document types, the clinician-in-the-loop workflow outperformed clinician-only drafting. LLM-only drafts were fastest but had the lowest quality. CONCLUSIONS AND RELEVANCE A clinician-in-the-loop approach improved document quality and accelerated documentation. Active clinician review of LLM-generated drafts is essential for clinical documentation, and such workflows may help enhance working conditions and patient care. TRIAL REGISTRATION ClinicalTrials.gov Identifier NCT07187050 Question Does a workflow in which clinicians review and edit LLM-generated drafts (“clinician-in-the-loop”) improve the efficiency and quality of clinical documentation compared with clinician-only drafting? Findings In this single-center randomized controlled trial including 17 ophthalmology physicians, discharge summaries were significantly faster with the clinician-in-the-loop. Across both document types, clinician-in-the-loop drafts achieved higher quality scores than clinician-only documents. Meaning A clinician-in-the-loop workflow using LLMs can simultaneously enhance the efficiency and quality of clinical documentation; accordingly, active clinician review remains essential for improving working conditions and patient care overall. Introduction Physicians devote a substantial proportion of their workday to clerical tasks, primarily electronic health record (EHR) documentation. A large U.S. time-motion study revealed that approximately 49% of physicians’ total working time is spent on EHR use and other desk work, with documentation alone accounting for roughly one-third of each patient encounter. 1 Similarly, a survey of physicians at 80 Japanese hospitals found that EHR-related documentation consumed 13% of weekly working hours and represented one of the strongest predictors of overtime. 2 This documentation workload not only extends working hours but also contributes to decreased job satisfaction and heightened risk of burnout among clinicians. 3 – 6 Large language models (LLMs) have been proposed to mitigate this burden, 7 which have received increasing attention because they can help simplify administrative tasks in healthcare. 8 , 9 Recent investigations have applied models such as GPT 4 and Claude to the automated drafting of discharge summaries, 10 – 16 discharge referrals, 16 , 17 radiology reports, 18 , 19 or other types of medical documents. 18 , 20 – 23 However, comparative evaluations consistently report substantial limitations. Hallucinated findings, 11 , 13 , 21 factual omissions, 21 factual errors, 11 – 13 , 21 , 22 and occasionally harmful recommendations, 13 , 19 , 22 leading reviewers to conclude that LLM-generated documents are not sufficiently reliable for clinical practice. 13 , 19 , 22 Because hallucinations and other factual inaccuracies are clinically unacceptable, LLM-generated drafts must be checked and corrected by clinicians. 15 , 20 , 24 Accordingly, evaluation of the clinical utility of LLMs for document creation should consider the entire documentation process—a clinician-in-the-loop workflow that includes clinician review and editing after generation. 24 However, most prior studies assess model output in isolation; to our knowledge, none has jointly evaluated time and quality within the workflow. In the current study, to evaluate this clinician-in-the-loop workflow, we conducted a pilot randomized controlled trial using a novel LLM assistant, CocktailAI. This trial assessed documentation time and quality by comparing three approaches: documents drafted by our LLM assistant and then modified by clinicians, documents written entirely from scratch by clinicians, and documents generated by the LLM assistant with no human review. Methods This study was reported following Consolidated Standards of Reporting Trials (CONSORT) reporting guideline for randomized clinical trials. 25 The protocol for this study is provided in Supplement 1. Study Design and Setting This single-center, parallel-group, prospective, randomized, open-label, blinded-endpoint (PROBE) trial was conducted at Kyoto University Hospital to evaluate the impact of the clinician-in-the-loop workflow. Participants were randomized to either an LLM-assisted documentation group (Clinician-in-the-loop group) or a manual, from-scratch documentation group (Clinician-only group). Randomization was performed using permuted block randomization with a block size of two and a 1:1 allocation ratio, stratified by participants’ position. The randomization list was generated in advance using an online tool ( https://www.sealedenvelope.com/simple-randomiser/v1/lists ), and participants were assigned sequentially according to this list. The study was conducted from February 18 to March 14, 2025. Participants were blinded to their group assignments prior to the initiation of the intervention; however, both participants and investigators were unblinded during the intervention phase. Outcome assessors evaluating the quality of the clinical documents were blinded to group assignments. Participants and Recruitment Eligible participants included junior residents, senior residents, and attending ophthalmologists in the Department of Ophthalmology at Kyoto University Hospital. Physicians were eligible if they were actively working in the department during the study period and provide the confirmation that they did not routinely use CocktailAI for clinical documentation. The study was conducted from February 18 to March 14, 2025. The investigators explained the study overview and eligibility criteria to the participants, and obtained written informed consent before any study procedures. Participants then completed a baseline questionnaire capturing demographic and professional characteristics, including age, gender, position, years of clinical experience, average daily time spent using EHRs, average number of days per week using EHRs, and the frequency of generative AI use. Simulated Medical Record Six simulated patient records were developed in Japanese by Keina S. and Kenji S., representing diverse ophthalmology cases: cataract, ocular trauma, retinal detachment, corneal ulcer, optic neuritis, and glaucoma. Each record comprised comprehensive inpatient demographic data, pre-admission outpatient SOAP notes, and daily inpatient SOAP notes. The simulated records were formally approved by the Japanese Society of Artificial Intelligence in Ophthalmology as suitable for evaluation. English translations of all cases are provided in Supplement 2. Evaluation System To precisely measure the time required for drafting and subsequent clinician modifications, we developed a purpose-built web application system that reproduces the clinical documentation workflow. In this system, CocktailAI, co-developed by the Department of Ophthalmology and Visual Sciences at Kyoto University Graduate School of Medicine, and Fitting Cloud Inc. (Kyoto, Japan), served as the LLM assistant. It assists in the creation of clinical documents by extracting relevant information from EHRs using LLMs and embedding the extracted content into predefined templates ( Figure 1 ). In this study, the system employed the generative AI model Gemini-2.0-flash-lite for text generation. Download figure Open in new tab Figure 1. Example of Template-Based Document Generation An LLM assistant was developed collaboratively by the Department of Ophthalmology and Visual Sciences at Kyoto University Graduate School of Medicine, and Fitting Cloud Inc. It extracts relevant information from EHRs and generates draft texts using predefined templates. Segments in {…} act as machine-readable directives, guiding the LLM assistant to retrieve the specified information directly from the EHR. The templates used in this study were created by T.T., who was not involved in developing the simulated patient records. They were based on recent discharge summaries and discharge referrals from the Department of Ophthalmology. The complete content of each template is provided in Supplement 3. Study Arms We tasked participants with creating two documents—discharge summary and discharge referral—for each of six simulated patient cases under two groups: Clinician-in-the-loop group and Clinician-only group. For each simulated case, participants created a discharge summary first and then a discharge referral; both the task order and the sequence of the six cases were fixed and identical for all participants. In the Clinician-in-the-loop group, participants first generated a draft with the LLM assistant based on the predefined templates. Then, they reviewed the LLM-generated text and edited it. Copy-and-paste from the patient record was permitted. No re-generation of LLM drafts was allowed. In the Clinician-only group, participants worked from the same six patient records and completed the same two documents but had no access to the LLM assistant. They used templates with the same structure as those in the Clinician-in-the-loop group, although all LLM instruction prompts had been removed in advance to avoid the additional time needed to delete them for each use. Copy-and-paste from the record was also permitted. Furthermore, an LLM-only group was included among the study groups. We retained the unedited drafts automatically produced by the LLM assistant before any human review or modification in the Clinician-in-the-loop group. These outputs represent LLM-generated documents without clinician review and were analyzed as a separate group. The above information is presented in Figure 2 . Download figure Open in new tab Figure 2. Study Arms and Clinical-Documentation Workflow across Six Simulated Cases Participants were randomized to the Clinician-in-the-loop group or Clinician-only group. Clinician-in-the-loop group: participants used the LLM assistant to generate an initial draft, then edited the draft and submitted the final documents (discharge summary and discharge referral); in parallel, the unedited LLM drafts produced before any clinician review were saved and analyzed as an LLM-only group. Clinician-only group: For each of six simulated EHR cases, participants created and submitted the clinical documents directly from the EHR without the LLM assistant. Outcome Measures The primary outcome was the average time required to complete each clinical document. Document creation time was automatically recorded within the study web application. For the secondary outcome, document quality was independently evaluated by five ophthalmologists—Keina S. (Case 1), K.I. (Case 2), M.M. (Case 3), N.U.-A. (Case 4), and Kenji S. (Cases 5 and 6). Cases were matched to each evaluator’s clinical expertise. Evaluators were blinded to participant identity and study-arm. Document quality was assessed across six domains: Medical Accuracy: The comprehensiveness of the information provided in the medical record Language: Text style and vocabulary fit the situation and the intended target group that will read the text Conciseness: The text contains only the strictly necessary information, and follows a logical or chronological order Presence of Hallucinations: Any text that is factually incorrect or inconsistent, or unrelated to the medical record Validity for Clinical Use: The text is deemed as usable in a true clinical scenario dealing with real patient information Possibility of Harm: The severity and likelihood of potential harm based on people acting on the text Domains 1–5 followed the framework of Sánchez-Rosenberg et al. (2024). 16 Although Sánchez-Rosenberg et al. defined Domain 6 as “possibility of bias,” we instead applied the rubric of Singhal et al. (2023) 26 and adapted Domain 6 to “possibility of harm,” since the former may suggest medical demographic bias, whereas our focus was on harmful medical misinformation. Detailed evaluation criteria —following Sánchez-Rosenberg et al. (2024) for Domains 1–5 and Singhal et al. (2023) for Domain 6—are provided in Supplement 4. In addition to these domain-based evaluations, each document was subjectively rated on an overall quality scale from 0 to 10. Furthermore, evaluators were asked to guess which study group each document belonged to. Analysis Although all participants were randomized, four participants withdrew before initiating any study procedures due to lack of time and provided no evaluable data; they had no exposure to the intervention and were unaware of their assignment. Consequently, analyses were conducted in the evaluable population, and an ITT analysis was not feasible because of missing data. For the primary outcome, the mean and 95% confidence intervals were calculated separately for discharge summaries and discharge referrals. Separate linear mixed models (LMMs) were constructed for each document type (i.e., discharge summaries and discharge referrals) to analyze the primary outcome of document creation time. Specifically, group (the Clinician-only group vs. the Clinician-in-the-loop group) and clinical experience level (junior residents vs. all other participants) were included as fixed effects, and a random intercept was specified to account for the nested structure of cases within participants. As an exploratory analysis in this pilot study, an additional model including the group × clinical experience level interaction was fitted, and its estimates were reported alongside those of the primary models. For the secondary outcome, descriptive statistics, including the mean for each evaluation domain, were calculated separately for each simulated patient case. In addition, for the overall quality scale, the mean and 95% confidence intervals were calculated for each study group, and for the task in which evaluators were asked to guess the group assignment of each document, corresponding correct identification rates were computed. All analysis was implemented using SAS Studio version 3.81. Ethics The Institutional Review Board of FINDEX Inc. reviewed this study and concluded that it was exempt from formal approval. Results A total of 21 physicians were randomized (11 to Clinician-in-the-loop, 10 to Clinician-only). In the Clinician-in-the-loop group, 8 participants completed documentation tasks and 3 withdrew due to lack of time; in the Clinician-only group, 9 completed and 1 withdrew for the same reason. Because group assignment was revealed only after login to the study system, dropout occurred before participants knew their assignment. In total, 17 physicians completed the study and were included in the analysis. The Clinician-in-the-loop group submitted 48 discharge summaries and 48 referrals, and the Clinician-only group submitted 54 of each; 48 unedited LLM drafts were stored as the LLM-only group. The recruitment and allocation process is shown in Figure 3 . Download figure Open in new tab Figure 3. Participant Flowchart A CONSORT-style flow diagram illustrating the enrollment, allocation, post-intervention, and analysis of participants. Among the 17 participants, 35.3% were junior residents, 41.2% senior residents, and 23.5% attending ophthalmologists, with a mean clinical experience of 5.8 years. Other demographic information is provided in Table 1 . View this table: View inline View popup Download powerpoint Table 1. Baseline Participant Demographics Primary Outcome The average mean time required to complete discharge summaries and discharge referrals (measured in seconds) is illustrated in Figure 4 . For discharge summaries, the Clinician-in-the-loop group demonstrated shorter mean times than the Clinician-only group in five of the six simulated cases. However, for discharge referrals, the Clinician-in-the-loop group required more time than the Clinician-only group in five of the six cases. A shorter mean time was observed only in Case 1. Detailed results are summarized in Table 2 . Download figure Open in new tab Figure 4. Time to Complete Clinical Documents Across Six Simulated Cases Bar charts show the mean time (seconds) required to complete each document for six simulated patient cases, comparing the Clinician-in-the-loop group (orange) with Clinician-only group (blue). Left: Discharge summaries. Right: Discharge referrals. Error bars indicate 95% confidence intervals. View this table: View inline View popup Download powerpoint Table 2. Mean Document Creation Time by Case and Group Separate LMMs were constructed for discharge summaries and referrals to evaluate the effects of group assignment and clinical experience level on document creation time, with a random intercept for each participant-case combination. Full summaries of estimates are presented in Table 3 . View this table: View inline View popup Download powerpoint Table 3. Linear Mixed Model Estimates for Document Creation Time In the model for discharge summaries, the Clinician-in-the-loop group assignment was associated with a significantly shorter document creation time compared to the Clinician-only group assignment (β = −59.4, 95% CI: −118.1 to −0.8, p = .047). No significant effect was observed for clinical experience level (β = −2.3, 95% CI: −63.6 to 59.0, p = .94). In the additional model for discharge referrals, the Clinician-in-the-loop group assignment was significantly associated with longer document creation time (β = 94.8, 95% CI: 40.4 to 149.3, p<.001), while clinical experience level remained non-significant (p = .70). Although the interaction term between group and clinical experience level was not statistically significant (p = .060), the estimated value was −64.9 seconds (95% CI: −132.8 to 2.9). Secondary Outcome The submitted clinical documents were evaluated across six predefined quality domains. The Clinician-in-the-loop group showed higher average scores than the Clinician-only group across nearly all domains in discharge summaries and discharge referrals. The LLM-only group consistently received the lowest scores across most domains for both discharge summaries and discharge referrals; however, in the discharge summaries the Possibility of Harm and Validity for Clinical Use domains, and in the discharge referrals the Possibility of Harm domain, achieved scores comparable to those of the Clinician-only group. ( Figure 5 ). Download figure Open in new tab Figure 5. Average Evaluation Scores Across Document Types by Domain Radar charts displaying the average evaluation scores for each quality domain across three groups (Clinician-in-the-loop, Clinician-only and LLM-only) for discharge summaries (left) and discharge referrals (right). Scores represent the mean values across all simulated cases, aggregated for six domains: Medical Accuracy, Language, Conciseness, Presence of Hallucinations, Validity for Clinical Use, and Possibility of Harm. Higher scores indicate better performance. Each document was also rated on an overall quality scale from 0 to 10. The Clinician-in-the-loop group received the highest mean scores for both document types: 7.8 (95% CI: 7.4 to 8.2) for discharge summaries and 8.0 (95% CI: 7.6 to 8.4) for discharge referrals, respectively. Detailed results are summarized in Table 4 . View this table: View inline View popup Download powerpoint Table 4. Overall Quality Ratings of Generated Documents by Group and Document Type For discharge summaries, the Clinician-in-the-loop group was correctly identified in 60.4% of cases, whereas the Clinician-only group was correctly identified in 75.9% of cases. For discharge referrals, the rates were 56.3% for the Clinician-in-the-loop group and 63.0% for the Clinician-only group ( Table 5 ). View this table: View inline View popup Download powerpoint Table 5. Distribution of Predicted Group Assignments by Actual Study Group (%) Discussion This randomized controlled trial is the first to assess the real-world utility of LLMs for clinical documentation while explicitly evaluating both time and quality outcomes. We compared a clinician-in-the-loop workflow with manual drafting because LLM-only output remains difficult to implement safely in routine clinical practice. In the Clinician-in-the-loop group, physicians completed discharge summary more quickly, and clinician-edited LLM drafts consistently outperformed manually drafted documents across most quality domains and on overall expert-review scores. Exploratory analyses further suggested a potential interaction between clinician seniority and the effect of LLM use, offering a concrete implication for how clinician-in-the-loop approaches may deliver value in practice. Several studies have compared AI-generated documents with those written by clinicians ( Table 6 ). For relatively simple, short-form clinical texts—such as radiology reports and progress notes—LLM outputs have achieved performance comparable to human-written documents. 18 However, as both the input prompts and target outputs lengthen (e.g., discharge referrals and discharge summaries), known issues such as hallucinations and other errors become more prominent. 11 , 12 , 15 , 16 Consistent with these reports, our LLM-only group produced drafts far more quickly yet with inferior quality. In routine clinical care, such deficiencies are unacceptable; accordingly, any LLM-enabled documentation workflow should make clinician review and revision an essential step rather than an optional safeguard. This physician-led use of LLM is also preferred from the patient perspective, as shown by recent large-scale international survey data. 27 View this table: View inline View popup Download powerpoint Table 6. Overview of published studies evaluating LLM interventions for clinical document generation Two studies have evaluated documentation time in clinician-in-the-loop workflows, one focused on discharge summaries and one on supervisory notes. 15 , 23 Both reported reduced completion time which aligns with our findings. However, neither study assessed document quality, leaving unresolved whether efficiency gains were achieved at the expense of accuracy or completeness. This is the first study to jointly evaluate documentation time and quality, and the first to assess the real-world clinical use of LLM-assisted documentation in a randomized controlled trial. Across multiple quality domains, the clinician-in-the-loop approach outperformed manual drafting, indicating that this method is not merely “quick and dirty” but instead improves quality while reducing documentation burden. These results suggest that clinician-in-the-loop deployment of LLMs can simultaneously raise document quality and relieve workload, a combination that may translate into better patient care overall. In this study, we implemented an extraction-based system and deployed it as an LLM assistant, reflecting the evidence that her data extraction can be highly accurate 28 – 30 . The observed difference in completion times between discharge summaries and discharge referrals likely reflects differences in document structure. For highly standardized documents such as discharge summaries, direct extraction into fixed templates may be sufficient. By contrast, discharge referrals involve greater narrative variability, and documents created from rigid templates often require additional modification to fit individual physician preferences. When creating documents with such narrative diversity, providing support for flexible, clinician-authored template creation alongside structured data extraction may be more appropriate and is likely to facilitate broader clinical adoption and accelerate real-world implementation. We also examined heterogeneity by clinician seniority during LLM use. The interaction between clinician experience and the effects of AI assistance has been investigated in the recognition of radiological and pathological images, 31 – 36 whereas, to our knowledge, no such evidence has been available in the context of LLMs. A sizable estimated seniority-group interaction in our study provides the first evidence in clinical LLM applications that greater clinical expertise may amplify the benefits of assistance. This finding suggests that LLM-specific training alone is unlikely to be sufficient; deeper clinical expertise may be required to maximize the benefits of LLMs in documentation workflows. Beyond documentation, AI in medicine now extends to areas such as clinical reasoning, 37 image recognition, 38 , 39 and basic research, 40 and it is important to investigate whether similar patterns are observed in these domains. This study has several limitations. First, our results reflect the capabilities of current language models. Future models with enhanced capabilities may produce different outcomes; nonetheless, our trial establishes a baseline for current model performance herEHR-focused applications. Second, we tested the system only in ophthalmology. Further research across other medical specialties is needed, including hospital-wide assessments of time savings, documentation quality, and downstream clinical outcomes. Third, templates were authored by an investigator rather than participants, so they may not fully reflect individual writing styles, as discussed earlier. Although we derived templates from recent discharge summaries and referrals in our Ophthalmology department to minimize stylistic mismatch, inter-physician variability likely persisted and may have influenced results. Future studies should evaluate the LLM system using clinician-authored, personalized templates prior to testing. In conclusion, a clinician-in-the-loop workflow both accelerated documentation and improved overall document quality. Accordingly, successful implementation of AI tools in clinical practice requires active clinician supervision. When hospitals combine accurate AI-generated drafts with thorough human review, they can achieve higher quality documentation, shorter writing times, and, as a consequence, improved working conditions and patient care overall. Author Contribution T.T. and Keina S. had full access to all the data in the study and took responsibility for the integrity of the data and the accuracy of the data analysis. Conceptualization and Study design: T.T., Keina S., Kenji S., H.T., M.M. Development of a purpose-built web application system: T.T., K.O., L.S. Development of the simulated medical record: Keina S., Kenji S., T.O., M.M. Data Collection: T.T. Assessment of document quality: Keina S., Kenji S., N.U.-A., K.I., M.M. Statistical analysis: T.T., Keina S., Y.U. Interpretation of Results: T.T., Keina S., Kenji S., H.T., M.M. Drafting of the manuscript: T.T. Critical revision of the manuscript for important intellectual content: all authors. Supervision: T.K., A.T., M.M. Conflict of Interest T.T. reports part-time employment with Fitting Cloud Inc. until March 2024; no equity holdings. Keina S. has received honoraria from FINDEX Inc. and Fitting Cloud Inc. Kenji S. has received a research grant from Fitting Cloud Inc. and Google Cloud Research Credits for the use of Gemini from Google Cloud, outside the submitted work. H.T. reports grants from FINDEX Inc. and grants from Google, outside the submitted work. K.O. reports receiving remuneration as a board member of Fitting Cloud Inc. L.S. reports being employed by Fitting Cloud Inc. A.T. reports receiving research funding from FINDEX Inc., outside the submitted work. M.M. has received honorarium from FINDEX Inc., outside the submitted work. Financial Support The application programming interface fee for Gemini was supported by Fitting Cloud Inc. Role of the Funder/Sponsor The funder had no organizational role in the design and conduct of the study; collection, management, analysis, or interpretation of the data; the preparation, review, or approval of the manuscript; or the decision to submit the manuscript for publication. The authors had full autonomy in all aspects of the study. T.T., while employed part-time by the funder, contributed to the design, conduct, and data collection in his capacity as an author. K.O. and L.S., employees of the funder, developed the LLM assistant (CocktailAI) and the study web application and participated in manuscript review and approval and in the decision to submit the manuscript for publication in their capacity as authors. Disclaimer Responsibility for the study’s analyses, findings, and opinions rests entirely with the authors. The content does not reflect the funder’s official policy or position, and no endorsement by the funder is implied. Declaration of generative AI use in the writing process During the preparation of this work, the authors used ChatGPT (GPT-4o and GPT-5 by OpenAI) for translation, readability improvement, and proofreading of the English text. After using these services, the authors reviewed and edited the content as needed and took full responsibility for the content of the publication. Data Availability The simulated patient records created for this study are publicly available in Supplement 2. Templates for the discharge summary and the discharge referral are provided in Supplement 3. Detailed evaluation criteria—following Sánchez-Rosenberg et al. (2024) for Domains 1–5 and Singhal et al. (2023) for Domain 6—are provided in Supplement 4. Acknowledgements Abbreviations LLM Large Language Model EHR Electronic Health Record CI Confidence Intervals LMM Linear Mixed Model References 1. ↵ Sinsky C , Colligan L , Li L , et al. Allocation of Physician Time in Ambulatory Practice: A Time and Motion Study in 4 Specialties . Ann Intern Med . 2016 ; 165 ( 11 ): 753 – 760 . doi: 10.7326/M16-0961 OpenUrl CrossRef PubMed 2. ↵ Yamaguchi N , Mochida Y , Matsubara S , Sumitani S . Factors contributing to physician overwork in 80 Saiseikai hospitals of Japan . Preprint posted online March 2 , 2020 . doi: 10.21203/rs.3.rs-15702/v1 OpenUrl CrossRef 3. ↵ West CP , Dyrbye LN , Shanafelt TD . Physician burnout: contributors, consequences and solutions . J Intern Med . 2018 ; 283 ( 6 ): 516 – 529 . doi: 10.1111/joim.12752 OpenUrl CrossRef PubMed 4. Rao SK , Kimball AB , Lehrhoff SR , et al. The Impact of Administrative Burden on Academic Physicians: Results of a Hospital-Wide Physician Survey . Acad Med J Assoc Am Med Coll . 2017 ; 92 ( 2 ): 237 – 243 . doi: 10.1097/ACM.0000000000001461 OpenUrl CrossRef PubMed 5. Gardner RL , Cooper E , Haskell J , et al. Physician stress and burnout: the impact of health information technology . J Am Med Inform Assoc . 2019 ; 26 ( 2 ): 106 – 114 . doi: 10.1093/jamia/ocy145 OpenUrl CrossRef PubMed 6. ↵ Gesner E , Dykes PC , Zhang L , Gazarian P . Documentation Burden in Nursing and Its Role in Clinician Burnout Syndrome . Appl Clin Inform . 2022 ; 13 ( 5 ): 983 – 990 . doi: 10.1055/s-0042-1757157 OpenUrl CrossRef PubMed 7. ↵ Omiye JA , Gui H , Rezaei SJ , Zou J , Daneshjou R . Large Language Models in Medicine: The Potentials and Pitfalls: A Narrative Review . Ann Intern Med . 2024 ; 177 ( 2 ): 210 – 220 . doi: 10.7326/m23-2772 OpenUrl CrossRef PubMed 8. ↵ Zheng Y , Gan W , Chen Z , Qi Z , Liang Q , Yu PS . Large Language Models for Medicine: A Survey . arXiv . Preprint posted online May 19, 2024. Accessed September 12, 2024 . http://arxiv.org/abs/2405.13055 9. ↵ Bedi S , Liu Y , Orr-Ewing L , et al. A Systematic Review of Testing and Evaluation of Healthcare Applications of Large Language Models (LLMs) . Preprint posted online April 16 , 2024 . doi: 10.1101/2024.04.15.24305869 OpenUrl Abstract / FREE Full Text 10. ↵ Koh MCY , Ngiam JN , Oon JEL , Lum LHW , Smitasin N , Archuleta S . Using ChatGPT for writing hospital inpatient discharge summaries – perspectives from an inpatient infectious diseases service . BMC Health Serv Res . 2025 ; 25 ( 1 ): 221 . doi: 10.1186/s12913-025-12373-w OpenUrl CrossRef PubMed 11. ↵ Ganzinger M , Kunz N , Fuchs P , et al. Automated generation of discharge summaries: leveraging large language models with clinical data . Sci Rep . 2025 ; 15 ( 1 ). doi: 10.1038/s41598-025-01618-7 OpenUrl CrossRef 12. ↵ Karino M , So C , Jinta T , et al. Exploring the Potential of ChatGPT for the Summarization of Patient Medical Histories: A Pilot Study . Cureus. Published online May 14 , 2025 . doi: 10.7759/cureus.84133 OpenUrl CrossRef 13. ↵ Schwieger A , Angst K , De Bardeci M , et al. Large language models can support generation of standardized discharge summaries – A retrospective study utilizing ChatGPT-4 and electronic health records . Int J Med Inf . 2024 ; 192 : 105654 . doi: 10.1016/j.ijmedinf.2024.105654 OpenUrl CrossRef PubMed 14. Clough RAJ , Sparkes WA , Clough OT , Sykes JT , Steventon AT , King K . Transforming healthcare documentation: harnessing the potential of AI to generate discharge summaries . BJGP Open . 2024 ; 8 ( 1 ):BJGPO.2023.0116. doi: 10.3399/bjgpo.2023.0116 OpenUrl CrossRef 15. ↵ Chua CE , Lee Ying Clara N , Furqan MS , et al. Integration of customised LLM for discharge summary generation in real-world clinical settings: a pilot study on RUSSELL GPT . Lancet Reg Health - West Pac . 2024 ; 51 : 101211 . doi: 10.1016/j.lanwpc.2024.101211 OpenUrl CrossRef 16. ↵ Sánchez-Rosenberg G , Magnéli M , Barle N , et al. ChatGPT-4 generates orthopedic discharge documents faster than humans maintaining comparable quality: a pilot study of 6 cases . Acta Orthop . 2024 ; 95 : 152 – 156 . doi: 10.2340/17453674.2024.40182 OpenUrl CrossRef PubMed 17. ↵ Tung JYM , Gill SR , Sng GGR , et al. Comparison of the Quality of Discharge Letters Written by Large Language Models and Junior Clinicians: Single-Blinded Study . J Med Internet Res . 2024 ; 26 : e57721 . doi: 10.2196/57721 OpenUrl CrossRef PubMed 18. ↵ Van Veen D , Van Uden C , Blankemeier L , et al. Adapted large language models can outperform medical experts in clinical text summarization . Nat Med . 2024 ; 30 ( 4 ): 1134 – 1142 . doi: 10.1038/s41591-024-02855-5 OpenUrl CrossRef PubMed 19. ↵ Chung EM , Zhang SC , Nguyen AT , Atkins KM , Sandler HM , Kamrava M . Feasibility and acceptability of ChatGPT generated radiology report summaries for cancer patients . Digit Health . 2023 ; 9 . doi: 10.1177/20552076231221620 OpenUrl CrossRef 20. ↵ Zaretsky J , Kim JM , Baskharoun S , et al. Generative Artificial Intelligence to Transform Inpatient Discharge Summaries to Patient-Friendly Language and Format . JAMA Netw Open . 2024 ; 7 ( 3 ): e240357 . doi: 10.1001/jamanetworkopen.2024.0357 OpenUrl CrossRef PubMed 21. ↵ Williams CYK , Bains J , Tang T , et al. Evaluating large language models for drafting emergency department encounter summaries . Navarro DF, ed. PLOS Digit Health . 2025 ; 4 ( 6 ): e0000899 . doi: 10.1371/journal.pdig.0000899 OpenUrl CrossRef 22. ↵ Sun Z , Ong H , Kennedy P , et al. Evaluating GPT4 on Impressions Generation in Radiology Reports . Radiology . 2023 ; 307 ( 5 ). doi: 10.1148/radiol.231259 OpenUrl CrossRef PubMed 23. ↵ Barak-Corren Y , Wolf R , Rozenblum R , et al. Harnessing the Power of Generative AI for Clinical Summaries: Perspectives From Emergency Physicians . Ann Emerg Med . 2024 ; 84 ( 2 ): 128 – 138 . doi: 10.1016/j.annemergmed.2024.01.039 OpenUrl CrossRef PubMed 24. ↵ Al-Garadi M , Mungle T , Ahmed A , Sarker A , Miao Z , Matheny ME . Large Language Models in Healthcare . arXiv . Preprint posted online 2025 . doi: 10.48550/ARXIV.2503.04748 OpenUrl CrossRef 25. ↵ Hopewell S , Chan AW , Collins GS , et al. CONSORT 2025 explanation and elaboration: updated guideline for reporting randomised trials . BMJ . Published online April 14 , 2025 : e081124 . doi: 10.1136/bmj-2024-081124 OpenUrl CrossRef 26. ↵ Singhal K , Azizi S , Tu T , et al. Large language models encode clinical knowledge . Nature . 2023 ; 620 ( 7972 ):172-180. doi: 10.1038/s41586-023-06291-2 OpenUrl CrossRef PubMed 27. ↵ Busch F , Hoffmann L , Xu L , et al. Multinational Attitudes Toward AI in Health Care and Diagnostics Among Hospital Patients . JAMA Netw Open . 2025 ; 8 ( 6 ): e2514452 . doi: 10.1001/jamanetworkopen.2025.14452 OpenUrl CrossRef PubMed 28. ↵ Huang J , Yang DM , Rong R , et al. A critical assessment of using ChatGPT for extracting structured data from clinical notes . Npj Digit Med . 2024 ; 7 ( 1 ): 106 . doi: 10.1038/s41746-024-01079-8 OpenUrl CrossRef PubMed 29. Ntinopoulos V , Rodriguez Cetina Biefer H , Tudorache I , et al. Large language models for data extraction from unstructured and semi-structured electronic health records: a multiple model performance evaluation . BMJ Health Care Inform . 2025 ; 32 ( 1 ): e101139 . doi: 10.1136/bmjhci-2024-101139 OpenUrl Abstract / FREE Full Text 30. ↵ Keloth VK , Selek S , Chen Q , et al. Social determinants of health extraction from clinical notes across institutions using large language models . Npj Digit Med . 2025 ; 8 ( 1 ). doi: 10.1038/s41746-025-01645-8 OpenUrl CrossRef 31. ↵ Yu F , Moehring A , Banerjee O , Salz T , Agarwal N , Rajpurkar P . Heterogeneity and predictors of the effects of AI assistance on radiologists . Nat Med . 2024 ; 30 ( 3 ): 837 – 849 . doi: 10.1038/s41591-024-02850-w OpenUrl CrossRef PubMed 32. Ahn JS , Ebrahimian S , McDermott S , et al. Association of Artificial Intelligence–Aided Chest Radiograph Interpretation With Reader Performance and Efficiency . JAMA Netw Open . 2022 ; 5 ( 8 ): e2229289 . doi: 10.1001/jamanetworkopen.2022.29289 OpenUrl CrossRef 33. Gaube S , Suresh H , Raue M , et al. Non-task expert physicians benefit from correct explainable AI advice when reviewing X-rays . Sci Rep . 2023 ; 13 ( 1 ): 1383 . doi: 10.1038/s41598-023-28633-w OpenUrl CrossRef PubMed 34. Bulten W , Balkenhol M , Belinga JJA , et al. Artificial intelligence assistance significantly improves Gleason grading of prostate biopsies by pathologists . Mod Pathol . 2021 ; 34 ( 3 ): 660 – 671 . doi: 10.1038/s41379-020-0640-y OpenUrl CrossRef PubMed 35. Kim HJ , Choi WJ , Gwon HY , et al. Improving mammography interpretation for both novice and experienced readers: a comparative study of two commercial artificial intelligence software . Eur Radiol . 2023 ; 34 ( 6 ): 3924 – 3934 . doi: 10.1007/s00330-023-10422-8 OpenUrl CrossRef PubMed 36. ↵ Dratsch T , Chen X , Rezazade Mehrizi M , et al. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance . Radiology . 2023 ; 307 ( 4 ): e222176 . doi: 10.1148/radiol.222176 OpenUrl CrossRef PubMed 37. ↵ Goh E , Gallo R , Hom J , et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial . JAMA Netw Open . 2024 ; 7 ( 10 ): e2440969 . doi: 10.1001/jamanetworkopen.2024.40969 OpenUrl CrossRef PubMed 38. ↵ Tomita K , Nishida T , Kitaguchi Y , Kitazawa K , Miyake M . Image Recognition Performance of GPT-4V(ision) and GPT-4o in Ophthalmology: Use of Images in Clinical Questions . Clin Ophthalmol . 2025 ;Volume 19 : 1557 – 1564 . doi: 10.2147/OPTH.S494480 OpenUrl CrossRef PubMed 39. ↵ Kamei T , Miyake M , Sado K , et al. Association between the retinal age gap and systemic diseases in the Japanese population: the Nagahama study . Jpn J Ophthalmol . 2025 ; 69 ( 4 ): 616 – 623 . doi: 10.1007/s10384-025-01205-3 OpenUrl CrossRef PubMed 40. ↵ Hosoda Y , Miyake M , Meguro A , et al. Keratoconus-susceptibility gene identification by corneal thickness genome-wide association study and artificial intelligence IBM Watson . Commun Biol . 2020 ; 3 ( 1 ): 410 . doi: 10.1038/s42003-020-01137-3 OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted October 07, 2025. Download PDF Supplementary Material Data/Code Email Thank you for your interest in spreading the word about medRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality Message Subject (Your Name) has forwarded a page to you from medRxiv Message Body (Your Name) thought you would like to see this page from the medRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality Tomohiro Takayama , Keina Sado , Kenji Suda , Hiroshi Tamura , Naoko Ueda-Arakawa , Kenji Ishihara , Yuta Ueda , Kazuya Okamoto , Luciano Henrique de Oliveira Santos , Tetsuro Oshika , Tomohiro Kuroda , Akitaka Tsujikawa , Masahiro Miyake medRxiv 2025.10.06.25337211; doi: https://doi.org/10.1101/2025.10.06.25337211 Share This Article: Copy Citation Tools Evaluating an LLM-Assisted Workflow for Clinical Documentation: A Pilot Randomized Controlled Trial on Time and Quality Tomohiro Takayama , Keina Sado , Kenji Suda , Hiroshi Tamura , Naoko Ueda-Arakawa , Kenji Ishihara , Yuta Ueda , Kazuya Okamoto , Luciano Henrique de Oliveira Santos , Tetsuro Oshika , Tomohiro Kuroda , Akitaka Tsujikawa , Masahiro Miyake medRxiv 2025.10.06.25337211; doi: https://doi.org/10.1101/2025.10.06.25337211 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Health Informatics Subject Areas All Articles Addiction Medicine (570) Allergy and Immunology (863) Anesthesia (300) Cardiovascular Medicine (4442) Dentistry and Oral Medicine (444) Dermatology (383) Emergency Medicine (609) Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1511) Epidemiology (15231) Forensic Medicine (30) Gastroenterology (1126) Genetic and Genomic Medicine (6610) Geriatric Medicine (668) Health Economics (998) Health Informatics (4542) Health Policy (1370) Health Systems and Quality Improvement (1613) Hematology (543) HIV/AIDS (1266) Infectious Diseases (except HIV/AIDS) (15924) Intensive Care and Critical Care Medicine (1103) Medical Education (623) Medical Ethics (147) Nephrology (668) Neurology (6607) Nursing (346) Nutrition (999) Obstetrics and Gynecology (1146) Occupational and Environmental Health (957) Oncology (3338) Ophthalmology (974) Orthopedics (369) Otolaryngology (420) Pain Medicine (436) Palliative Medicine (130) Pathology (665) Pediatrics (1693) Pharmacology and Therapeutics (692) Primary Care Research (712) Psychiatry and Clinical Psychology (5448) Public and Global Health (9239) Radiology and Imaging (2203) Rehabilitation Medicine and Physical Therapy (1370) Respiratory Medicine (1196) Rheumatology (596) Sexual and Reproductive Health (714) Sports Medicine (530) Surgery (712) Toxicology (99) Transplantation (289) Urology (265) (function(){function c(){var b=a.contentDocument||a.contentWindow.document;if(b){var d=b.createElement('script');d.innerHTML="window.__CF$cv$params={r:'a0202e4fef6533e5',t:'MTc3OTgzNDE3MA=='};var a=document.createElement('script');a.src='/cdn-cgi/challenge-platform/scripts/jsd/main.js';document.getElementsByTagName('head')[0].appendChild(a);";b.getElementsByTagName('head')[0].appendChild(d)}}if(document.body){var a=document.createElement('iframe');a.height=1;a.width=1;a.style.position='absolute';a.style.top=0;a.style.left=0;a.style.border='none';a.style.visibility='hidden';document.body.appendChild(a);if('loading'!==document.readyState)c();else if(window.addEventListener)document.addEventListener('DOMContentLoaded',c);else{var e=document.onreadystatechange||function(){};document.onreadystatechange=function(b){e(b);'loading'!==document.readyState&&(document.onreadystatechange=e,c())}}}})();

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00