Benchmarking eye health advice from generative artificial intelligence in terms of factual accuracy, safety, comprehensiveness and readability | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Benchmarking eye health advice from generative artificial intelligence in terms of factual accuracy, safety, comprehensiveness and readability Aleksander Stupnicki, Bernardo Mendes, Maxwell Reinstein, Ariel Ong, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9383173/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Generative artificial intelligence (genAI) chatbots are increasingly used for health advice despite lacking regulatory approval, raising concerns about their output quality and safety. This study assesses eye health advice from leading genAI platforms, benchmarking their quality against patient information leaflets. Methods We compared outputs from GPT-5 (OpenAI) and Gemini 3 (Google DeepMind) with clinical leaflets across nine eye conditions (41 questions, 123 texts total). Reference benchmark (547 items) was derived from patient materials produced by the Royal College of Ophthalmologists. Chatbot outputs were generated using verbatim leaflet subsection headings as prompts with word-count restrictions to match corresponding leaflet sections. All texts were evaluated using the Comprehensiveness, Accuracy, and Safety Evaluation Framework (CASEF). Two blinded ophthalmologists assessed genAI outputs for safety concerns. Readability was measured using Flesch-Kincaid Grade Level. Results Both genAI models showed higher factual alignment than clinical leaflets (GPT-5 = 37.4%, Gemini 3 = 36.2%, leaflets = 30.7%; both p < 0.001) and fewer omissions (GPT-5 = 6.7, Gemini 3 = 7.0, leaflets = 7.9; both p < 0.001). Safety scores were comparable across sources, but both models underreported treatment complications, exhibited guideline inconsistencies, and failed to include appropriate safety-netting. Moreover, genAI outputs required 2–3 more years of education to understand. High inter-rater reliability (ICC = 0.843 (95%CI:0.800–0.880)) validated the scoring methodology. Conclusions GenAI eye health advice matches clinical leaflets in accuracy and comprehensiveness. However, subtle yet clinically consequential errors remain, limiting the application of general-purpose genAI chatbots as a safe, standalone information source for ophthalmology patients. Health sciences/Medical research Scientific community and society/Scientific community/Education Figures Figure 1 Figure 2 Figure 3 Glossary Generative Artificial Intelligence (GenAI): An umbrella term capturing computational systems that are trained and fine-tuned to produce material in response to a user which may comprise text, images, or audio. Large Language Models (LLMs): A subset of generative artificial intelligence systems trained with large volumes of textual data and exhibiting abilities to interpret and produce text and thereby engage in conversation or question-answering. General-purpose models: GenAI models designed to perform a wide range of tasks across diverse-domains, in contrast to domain-specific models, which are trained and tuned for narrower applications such as healthcare. Training corpus: The large, curated dataset of texts (drawn from sources such as books, websites, and scientific literature) used to train LLMs. Hallucinations: Phenomenon where genAI text contains information that is fabricated (not present in the training dataset). Prompt: The input submitted by a user to a genAI model to elicit a response, typically an instruction or a question. Prompt engineering: Techniques of designing and refining prompts to encourage desired behaviour or outputs from genAI. What was known before Generative artificial intelligence chatbots are widely used by patients for gathering health information despite lacking regulatory approval for medical use. Early models were associated with high rates of hallucinations and factual inaccuracies, raising concerns about patient reliance on AI-generated health advice. Prior evaluations of LLMs in ophthalmology focused predominantly on performance in knowledge examination tasks, where latest models match or exceed expert performance. Whether that capability translates to production of reliable patient-facing materials on common eye conditions is unclear. What this study adds : This study demonstrates that current general-purpose AI chatbots can match or exceed clinical patient education materials in factual accuracy and comprehensiveness. However, current genAI models exhibit clinically-relevant safety concerns, including systematic underreporting of treatment complications, defaulting to US-centric guidelines, inadequate safety-netting for sight-threatening conditions, and high linguistic complexity. These findings highlight that before general-purpose genAI models can serve as reliable alternatives to formal patient education materials, robust regulatory oversight and patient education on the risks of genAI are essential. 1. Introduction Generative artificial intelligence (genAI) describes a set of computational systems encompassing large language models (LLMs), which are capable of generating human-like text in response to natural language prompts ( 1 ). Since the public release of conversational LLM-based chatbots in late 2022, these tools have been rapidly adopted as sources of health information or advice: recent reports suggest up to 32% of adults turn to genAI for healthcare advice, ( 2 ) and 25% of ChatGPT users submit health-related questions on a weekly basis ( 3 ). By generating tailored responses on-demand, LLMs offer personalised care and have the potential to improve patient health literacy, engagement and ownership of health decisions ( 4 , 5 ). However significant risks associated with LLM use have been raised across medical literature, including inaccurate advice, clinically important omissions, hallucinations, agreement with user misconceptions, and unsafe risk communication ( 6 – 12 ). The lack of formal regulatory approval of general-purpose LLMs for medical use further limits oversight and accountability ( 13 ). In addition to the quality of genAI advice, linguistic features are an important consideration ( 14 ). In ophthalmology, poor health literacy is associated with delayed diagnosis, medication non-adherence, and greater carer dependency, contributing to adverse outcomes in conditions such as glaucoma and diabetic retinopathy ( 15 – 17 ). In eye clinics, leaflets are frequently used to provide patients with advice that is evidence-based and pitched at an understandable level. However, patients have long been obtaining medical advice from sources outside of those directly provided in the healthcare setting, primarily through online search engines, despite reportedly inferior quality and readability ( 18 , 19 ). With the rise in genAI adoption, this behaviour is now rapidly evolving, and patients frequently encounter AI-generated advice before consulting formal sources ( 20 ). In ophthalmology, leading genAI models now surpass human-expert performance on examination-style questions ( 21 – 23 ), but whether this capability extends to patient-facing advice, where accuracy must be balanced with nuanced safety communication, accessibility, and coverage, is less well established ( 6 , 7 , 23 , 24 ). Although publicly accessible genAI chatbots were neither designed for clinical applications nor approved by regulators for medical use, patients employ these systems in ways directly impacting their healthcare interactions. Acknowledging this real-world usage, we assessed whether leading general-purpose LLM platforms (GPT-5, OpenAI and Gemini 3, Google DeepMind) produce eye health advice suitable for patients. Using traditional patient information leaflets as reference standards, we evaluated outputs from leading genAI models across four domains: comprehensiveness, factual accuracy, safety, and readability. 2. Methods 2.1 Control arm leaflet selection The United Kingdom’s National Health Service (NHS) patient information leaflets served as control comparators, representing real-world materials routinely used in clinical practice. Online patient education repositories from four large UK tertiary ophthalmology centres (Moorfields Eye Hospital, Manchester University, Imperial College Healthcare, and Oxford University Hospitals NHS Foundation Trusts) were evaluated for content alignment with the nine benchmark conditions (see Section 2.2 ) and informational depth. Manchester University NHS Foundation Trust was selected as the control source, as it was the only repository that provided complete coverage of all nine conditions, used question-format subheadings, and also demonstrated the greatest informational depth ( 25 ). 2.2 Benchmark curation A checklist-based benchmark was developed to evaluate AI-generated text against expert-authored patient information. Source material comprised all nine patient information leaflets produced by the Royal College of Ophthalmologists (RCOphth) in collaboration with the Royal National Institute of Blind People (RNIB), available at the time of study. The leaflets covered age-related macular degeneration (AMD), cataract, Charles Bonnet syndrome (CBS), diabetic retinopathy (DR), dry eye disease, glaucoma, nystagmus, posterior vitreous detachment (PVD), and retinal detachment (RD). Two researchers (AS, BSM) independently mapped NHS leaflet subsections to corresponding RCOphth/RNIB benchmark domains, identifying 41 subsections across the nine conditions that aligned with benchmark content. Each subsection was then converted into discrete factual checklist items, with discrepancies resolved through discussion and duplicate statements consolidated to prevent score inflation. This yielded 547 benchmark items across 41 sections, which formed the basis for subsequent AI output generation. Each question was retrospectively categorised into one of eight thematic groups for subgroup analysis: disease definition, causes, symptoms, diagnosis, risk factors, treatment, complications, and ‘other’. The full benchmark is available on request. 2.3 AI advice generation GenAI advice was elicited from two general-purpose chatbot platforms: GPT-5 (accessed October 4, 2025, via chat.openai.com; OpenAI, San Francisco, CA, USA) and Gemini 3.0 (accessed November 21, 2025, via gemini.google.com; Google DeepMind, London, UK). Both are instruction-tuned conversational variants selected due to highest public adoption at the time of the study ( 26 ). Both models were accessed from London, UK, via web-based interfaces using newly registered accounts and default generation parameters without modification. Outputs were generated using matched NHS leaflet subsection headings as verbatim prompts with an explicit word count constraint matching the corresponding NHS text length ( e.g. "What is age-related macular degeneration (AMD)? Restrict your answer to 154 words"). This controlled for response length as a confounder. In four instances where NHS subsection headings were phrased as declarative statements ( e.g. "Your operation"), these were rephrased into interrogative or imperative formats ( e.g. "Explain the operation") to align with standard LLM query conventions. 2.4 Evaluation framework 2.4.1 Comprehensiveness, Accuracy and Safety Evaluation Framework (CASEF) A novel scoring framework was developed to quantify the concordance between text outputs (NHS leaflets and genAI advice) and the gold-standard benchmark across three core dimensions: comprehensiveness, accuracy, and safety. The CASEF framework (Supplementary Table 2) employs a five-point bidirectional scale that captures both errors of commission and omission, similar to validated Likert-type frameworks. CASEF draws on the elements of the S.C.O.R.E. framework ( 27 ) most relevant to patient-facing health information: safety and consensus (corresponding to negative CASEF scores and total CASEF score respectively), with comprehensiveness representing an extension of the consensus domain. Each text output was evaluated against the benchmark items corresponding to its topic (Supplementary Table 1), with each item independently scored by a multi-evaluator panel. Item-level scores were summed and normalised as a percentage of the maximum attainable score to account for variation in benchmark item counts across questions, yielding a range from − 100% (total contradiction) to 100% (complete concordance). Safety and comprehensiveness metrics were derived directly from CASEF. Safety was quantified as the sum of negative CASEF scores (-1 or -2) for a given text sample, and comprehensiveness as the total count of items scored 0 (omissions), with values closer to zero indicating better performance with respect to each metric. 2.4.2 Qualitative safety evaluation Quantitative safety appraisal via CASEF was supplemented with a manual safety evaluation by a senior ophthalmology resident (AYO) and a consultant ophthalmologist (AM) who were blinded to the source of each text. All genAI outputs (n = 82) were independently reviewed for factual accuracy, guideline alignment, risk communication, and other safety concerns not captured by the scoring framework. 2.4.3 Readability Readability was assessed using three validated metrics: the Simple Measure of Gobbledygook (SMOG) Index, Flesch-Kincaid Grade Level (FKGL), and Automated Readability Index (ARI) ( 28 ). SMOG estimates reading level from polysyllabic word frequency, while FKGL and ARI derive US school grade-level scores from sentence and word length. All 123 text samples were analysed per source (NHS, GPT-5, Gemini 3.0) using the Python textstat library (v0.7.10). 2.5 Evaluation process A hybrid multi-evaluator panel comprising five genAI models and one human researcher (AS) was employed for text scoring. This 'LLM-as-a-judge' approach addresses a key limitation of traditional AI evaluation studies, where reliance on human evaluators alone reduces reproducibility and introduces subjectivity ( 29 ). Recent validation studies demonstrated scoring consistency and strong inter-rater reliability between human expert raters and LLM evaluators in specialised medical contexts ( 30 , 31 ). Five latest public models (at the time of evaluation (November 2025)) were selected to average out model-specific biases: GPT-5 (OpenAI, San Francisco, CA, USA), Gemini 2.5 Pro (Google DeepMind, London, UK / Mountain View, CA, USA), Deepseek V3.2 (Deepseek AI, Hangzhou, China), Claude Sonnet 4.5 (Anthropic, San Francisco, CA, USA), and Llama 4 (Meta, Menlo Park, CA, USA). One human evaluator (AS, a student doctor) independently assessed all outputs to validate that genAI assessments aligned with human judgement. Each of the 123 text samples (3 sources with 41 outputs each) was independently evaluated by all six panel members against the applicable benchmark section. All genAI models operated using default generation parameters and received identical, standardised scoring prompts (Supplementary Fig. 1), applied prior to scoring each condition. Conversation memory was manually cleaned between evaluations to prevent cross-contamination. Individual item scores from all six raters were aggregated as arithmetic means to generate final CASEF scores for each text sample. Reporting adhered to the Chatbot Assessment Reporting Tool (CHART) guidelines (Supplementary Fig. 3) ( 32 ). No patient data was used in this study, and ethical approval is not applicable. 2.5.1 Scoring prompt A structured evaluation prompt was engineered to ensure reproducible and systematic scoring by all LLMs in the evaluation panel. The prompt was iteratively refined through pilot testing across all five LLM evaluators using a random sample of NHS texts (n = 5) not included in the study dataset. Established prompt engineering strategies were employed, including role prompting, task decomposition, explicit output constraints and quantitative scaffolding ( 33 , 34 ). Instructions to disregard stylistic differences and maintain task adherence throughout the session were included to minimise bias and hallucination. All genAI evaluators were supervised with a human-in-the-loop system for errors and adherence to prompt structure. The complete prompt is provided in Supplementary Fig. 1. 2.6 Statistical Analysis Analyses were performed using Python 3.11.5 (Python Software Foundation, Beaverton, USA). All comparisons between information sources (genAI versus NHS and between genAI models) used paired samples t-tests with observations matched at the rater-question level to account for the repeated measures. Comparisons of overall CASEF scores, safety, comprehensiveness, and readability were conducted with Bonferroni correction (adjusted α = 0.0167). A significance threshold of p = 0.01 was used for subgroup analysis stratified by condition and question category. Descriptive analyses were presented as total mean scores and a mean difference from NHS scores, both with standard deviations (SD). Inter-rater reliability was quantified using the two-way, random-effects, single-measurement intraclass correlation coefficient model (ICC( 2 , 1 )) with 95% confidence intervals ( 35 ). Potential preferential scoring by GPT-5 and Gemini 2.5 Pro for outputs from their respective model family was assessed by comparing each evaluator’s deviation from the all-rater consensus with the deviations observed in individual raters (Wilcoxon signed-rank test). 3. Results 3.1. Overall performance A total of 123 text outputs (82 genAI-generated and 41 NHS leaflet-derived) were evaluated against 547 checklist items, comprising frequent queries on nine ophthalmic conditions. Both genAI models demonstrated significantly higher factual alignment with the ground-truth benchmark compared to NHS leaflets (GPT-5: 37.36 ± 16.91% versus NHS: 30.70 ± 17.40%, p < 0.001; Gemini 3: 36.24 ± 18.03% vs NHS, p < 0.001), with no significant difference between them ( p = 0.940). Similarly, both LLMs demonstrated greater comprehensiveness than NHS leaflets as evidenced by lower omission counts (GPT-5: 6.74 ± 4.33 versus NHS: 7.87 ± 5.14, p < 0.001; Gemini 3: 6.97 ± 5.06 versus NHS, p < 0.001), with comparable comprehensiveness between models ( p = 0.215). Safety scores did not differ significantly across all three sources (GPT-5: -0.07 ± 0.28; Gemini 3: -0.14 ± 0.39; NHS: -0.10 ± 0.32; all p > 0.01). 3.2 Subgroup analysis: When stratified by condition (Fig. 3 A; Table 1), both genAI models significantly outperformed NHS leaflets for dry eye disease, nystagmus, and PVD ( p < 0.01 for all comparisons). Gemini 3 additionally achieved significantly higher CASEF scores for AMD ( p < 0.01). Diabetic retinopathy was the only condition where NHS leaflets significantly outperformed both genAI models ( p < 0.01). No significant differences were observed for the remaining four conditions (cataract, CBS, glaucoma, and RD). Comprehensiveness mirrored this pattern, with GPT-5 demonstrating significantly fewer omissions across six of nine conditions (dry eye, glaucoma, nystagmus, PVD, CBS, and retinal detachment, p < 0.01 for all) and Gemini 3 across two out of nine conditions (AMD and nystagmus, both p < 0.001). When stratified by question category (Fig. 3 B; Table 1), both genAI models achieved significantly higher CASEF scores for disease definition and treatment content ( p < 0.001 for all comparisons). Conversely, NHS leaflets demonstrated higher concordance for treatment complication content (Gemini 3 vs NHS, p < 0.01), though the GPT-5 comparison did not reach adjusted significance threshold ( p = 0.012). This pattern was mirrored in the comprehensiveness analysis, with both models demonstrating significantly fewer omissions for disease definitions and treatment content (all p < 0.001), but not for complications. No significant differences were observed for the remaining categories. When stratified by condition and question category, safety scores remained non-significant across all subgroup comparisons. 3.3 Qualitative safety evaluation Independent qualitative review by two clinicians identified several clinically relevant safety concerns not captured by quantitative scoring. Both models received a comparable number of safety-related comments (35 and 20 for GPT-5; 38 and 24 for Gemini 3). First, both models systematically provided US-centric advice, recommending treatments without UK regulatory approval ( e.g. lifitegrast for dry eye disease), referencing US-based driving standards ("Department of Motor Vehicles standards", instead of the British Driver and Vehicle Licensing Agency (DVLA)), and assuming procedural norms inconsistent with NHS practice ( e.g. routine sedation for cataract surgery, stopping anticoagulants pre-operatively for cataract surgery, and multifocal lens availability). Factual inaccuracies, although infrequent, included incorrect statements regarding the prevalence of oscillopsia in congenital nystagmus, the mechanism of neovascular glaucoma, and the pathophysiology of posterior vitreous detachment. In one instance, GPT-5 recommended vision therapy for nystagmus which was flagged as potentially misleading, as no high-quality evidence exists to support its clinical effectiveness ( 36 ). Furthermore, outputs from both models occasionally demonstrated inappropriate calibration of urgency. Some recommendations were overly optimistic ( e.g. implying full visual recovery after macula-off retinal detachment or omitting surgical failure rates), whilst others used alarmist language ( e.g. describing macula-off retinal detachment repair as requiring "immediate" intervention, potentially generating unnecessary patient anxiety). Both models omitted key safety-netting information, particularly post-operative warning signs for endophthalmitis and the need for emergency review should retinal re-detachment symptoms develop. 3.4 Readability analysis Overall, both genAI models yielded significantly more complex texts compared to the NHS materials, with mean FKGL scores of 12.05 (Gemini 3) and 12.61 (GPT-5), compared to 9.63 for NHS materials ( p < 0.001 for both comparisons). These scores correspond to 12th- and 9th-grade reading levels (US standards), equivalent to A-level and GCSE standards for genAI and NHS materials, respectively ( 37 ). Similarly, both ARI and SMOG metrics were substantially higher for Gemini 3 (13.46 and 13.83, respectively) and GPT-5 (14.04 and 14.19, respectively), compared with NHS materials (10.14 and 12.22; p < 0.001 for all comparisons). There were no statistically significant differences between the two genAI models across all readability metrics. Overall, genAI advice required an additional 2–3 years of education for comprehension. Importantly, all three sources exceeded the sixth-to-eighth-grade readability recommendation for all patient-facing materials ( 38 ). 3.5 Inter-rater reliability Scoring consistency across all six assessors was evaluated using the intraclass correlation coefficient. For total scores, the combined ICC was 0.843 (95% CI: 0.800–0.880), indicating very good inter-rater reliability ( 35 ). For individual information sources, ICC values ranged from good to excellent: 0.828 for GPT-5 (95% CI: 0.75–0.890), 0.814 for Gemini 3 (95% CI: 0.730–0.880), and 0.880 for NHS materials (95% CI: 0.780–0.860). Furthermore, inter-rater reliability between the human evaluator and each LLM was good to excellent with a mean ICC of 0.811 (range: 0.768–0.855), indicating that genAI assessments aligned with human judgement (Supplementary Table 3). Analysis of evaluator bias revealed no evidence for preferential scoring of content generated by respective model families by GPT-5 or Gemini 2.5 Pro (mean deviations from consensus: -0.40% for Gemini 2.5 Pro ( p = 0.918) and 0.40% for GPT-5 ( p = 0.372)). Collectively, these findings indicate consistent agreement among raters, as well as human-genAI agreement, validating the semi-automated scoring methodology. 4. Discussion In this benchmarked comparison of two general-purpose genAI models against NHS ophthalmology leaflets, genAI produced health advice that was factually accurate and frequently more comprehensive across nine common eye conditions. However, this informational depth came at the cost of readability, with both models generating text at a higher reading level than the NHS materials. Although overall safety profiles were comparable between sources, qualitative review revealed infrequent, subtle, but clinically relevant concerns such as under-reporting of treatment complications, misalignment with regional guidelines, imprecision, and lack of appropriate safety-netting. 4.1 Main findings and contextualisation Our findings align with earlier, condition-specific studies demonstrating the capacity of genAI to generate medically accurate and comprehensive eye health advice ( 7 , 22 , 24 , 39 , 40 ). The greater comprehensiveness of AI-generated materials was a key driver of its superior CASEF scores, reflecting a higher density of information relative to NHS leaflets. Although information-dense content can lead to cognitive overload, actually hindering patient comprehension ( 41 ), genAI platforms offer an advantage over static leaflets in their ability to address follow-up questions in real time, facilitating deeper patient engagement with health information ( 40 ). Greater comprehensiveness of genAI outputs was also reflected in substantially higher textual complexity: at the default reading level identified in our study, only 14% of UK adults would be able to fully comprehend these materials ( 42 ). This readability gap is consistent with the broader literature, in which healthcare-related genAI outputs typically require university-level literacy ( 40 , 43 – 45 ). Nevertheless, successive model releases show progressive improvements in default readability: FKGL scores for GPT-generated cataract materials decreased from 14.0 (GPT-3.5 ( 8 )) to 12.7 (GPT-4o ( 45 )) and 12.6 (GPT-5, this study). Moreover, evidence across multiple specialties including general surgery ( 46 ), cardiology ( 47 ), oncology ( 48 ), and plastic surgery ( 5 , 49 ) consistently demonstrates that targeted prompt engineering can substantially improve readability without compromising accuracy. However, employing these strategies requires genAI literacy which may disadvantage those with the greatest need for accessible health information. Despite excelling at disease definitions and treatment explanations, genAI models underperformed in communicating treatment complications, which carries important safety implications. Incomplete risk disclosure may bias patients towards treatment acceptance, compromising informed consent. We hypothesise this may reflect the relative scarcity of high-quality complications data in training corpora compared to disease definitions or treatments ( 50 ). Consistent with the reported overrepresentation of North American materials in the genAI training corpora ( 51 , 52 ), both models defaulted to recommendations reflecting US-centric healthcare pathways, medication availability, and clinical guidelines. Although technically accurate, such geographic misalignments constitute contextual safety errors with potential for direct patient harm, through pernicious self-management decisions ( e.g. following advice to discontinue anticoagulation prior to cataract surgery), delayed diagnosis (arising from misunderstood care pathways), and erosion of trust when AI-generated advice contradicts locally received recommendations. Both models also failed to provide adequate safety-netting, omitting post-operative red flag symptoms for sight-threatening complications such as endophthalmitis and retinal re-detachment, where delayed presentation directly determines visual prognosis ( 53 , 54 ). This is particularly consequential given that approximately 70% of health-related LLM conversations occur outside clinic hours ( 55 ), and when patients experience new or worsening symptoms, with no immediate access to direct clinical guidance ( 56 ). This absence of safety-netting aligns with a broader trend of progressively declining medical disclaimer rates ( 57 ), suggesting that successive model iterations have prioritised conversational authority and factual comprehensiveness over safety communication. Despite the safety limitations described above, LLMs have been shown to offer more accurate and reliable guidance than other online health information sources ( 8 , 58 , 59 ). As genAI becomes an established entry point for patient health queries ( 55 ), regulatory frameworks should mandate minimum safety-netting standards, and patients should be educated on safe prompting strategies and careful output interpretation ( 60 , 61 ). As contemporary genAI chatbots approach factual parity with clinical materials, safety concerns shift from overt hallucinations or contradictions to contextual, regional, or linguistic errors, which standard accuracy-focused frameworks are not designed to detect ( 62 ). This underscores the necessity of hybrid, ‘clinician-in-the-loop’ evaluation frameworks to identify the subtle but clinically consequential limitations that neither quantitative metrics alone nor human evaluators alone can reliably capture ( 29 ). 4.2 Limitations Our study has several limitations. First, each prompt was executed once per model without regeneration, providing a cross-sectional snapshot of inherently stochastic technology, which might affect reproducibility. Second, prompts based on precise clinical terminology used in patient leaflets may have optimised GenAI output quality, potentially overestimating the accuracy seen with patients’ more informal queries. Moreover, our methodology employed single-turn prompts which may not reflect typical patient interaction patterns. Third, using control leaflets sourced from a single NHS Trust and all evaluations conducted in English limits generalisability outside a UK context. Finally, our work focused on evaluating textual properties of outputs, but did not measure implications on patient understanding and behaviour. Future studies should assess performance in realistic multi-turn dialogue scenarios, evaluate emerging health-specialised models ( e.g . ChatGPT Health or Claude for Healthcare), employ authentic patient queries (with misspellings, incorrect grammar, colloquial terminology) as prompts, and incorporate patient- and behaviour-centred designs. 5. Conclusions Contemporary genAI can provide eye health advice on common ophthalmic conditions that matches or exceeds clinical leaflets in terms of factual accuracy and scope. However, limitations in readability, regional guideline alignment, and risk communication preclude their endorsement as safe, standalone resources for patients. Further development and regulation may help ensure that genAI health advice meets established healthcare standards. As this technology continues to proliferate, clinicians and patients should be encouraged to appreciate the strengths and limitations of genAI health advice, and verify information carefully. Declarations Acknowledgements AYO is supported by a National Institute for Health Research (NIHR) - Moorfields Eye Charity (MEC) Doctoral Fellowship (NIHR303691). PAK is supported by a UK Research & Innovation Future Leaders Fellowship (MR/T019050/1), Moorfields Eye Charity with The Rubin Foundation Charitable Trust (GR001753), and an Alcon Research Institute Senior Investigator Award. AJT is supported by an NIHR Academic Clinical Fellowship (ACF-2025-20-001). The views expressed in this publication are those of the authors and not necessarily those of the abovementioned funding bodies. Conflicts of interest PAK is a cofounder and equity owner of Cascader Ltd. He is an equity owner in Big Picture Medical. He has acted as a consultant for Skleo Health, insitro, Retina Consultants of America, Roche, Boehringer-Ingleheim, and Bitfount. He has received speaker fees from Zeiss, Bayer, Boehringer-Ingleheim, Thea, Alimera, Topcon and Roche, and grant funding from Roche. He has received travel support from Bayer and Roche. He has attended advisory boards for Topcon, Bayer, Boehringer-Ingleheim, and Roche. AJT has received grant funding from Théa Pharmaceuticals. None of the other authors report any conflicts of interest. Author contribution statement: AS and AJT conceived and designed the study. AS and BSM curated the benchmark, mapped leaflet subsections to benchmark domains and developed the checklist items. AS generated all AI outputs, and conducted the formal scoring and statistical analyses. AYO and AM performed the blinded qualitative safety evaluation. AS, BSM, MJBR, AYO and AJT drafted the manuscript. AS, BSM, MJBR, AYO, PAK, and AJT reviewed and revised the manuscript. AJT and PAK supervised the study. All authors approved the final version of the manuscript. Data availability statement: All data supporting the conclusions of this study are available within the article and its supplementary materials. The benchmark checklist developed for this study is available upon request. No new clinical or patient data were generated in this study. References Teo ZL, Thirunavukarasu AJ, Elangovan K, Cheng H, Moova P, Soetikno B, et al. Generative artificial intelligence in medicine. Nat Med [Internet]. 2025;31(10):3270–82. Available from: http://dx.doi.org/10.1038/s41591-025-03983-2 Montero A, Montalvo J III, Kearney A, Valdes I, Kirzinger A, Hamel L. KFF. 2026 [cited 2026 Mar 30]. KFF Tracking Poll on Health Information and Trust: Use of AI For Health Information and Advice. Available from: https://www.kff.org/public-opinion/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/ Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature [Internet]. 2023 Aug 12 [cited 2025 Nov 23];620(7972):172–80. Available from: http://dx.doi.org/10.1038/s41586-023-06291-2 Kim SI, Park J, Kim T, Seo W, Kim T, Cha WC, et al. Enhancing patient participation in emergency department through patient-friendly clinical notes generated by large language models. Sci Rep [Internet]. 2025 Dec 5 [cited 2026 Mar 29];16(1):1409. Available from: http://dx.doi.org/10.1038/s41598-025-31113-y Swisher AR, Wu AW, Liu GC, Lee MK, Carle TR, Tang DM. Enhancing health literacy: Evaluating the readability of patient handouts revised by ChatGPT’s large language model. Otolaryngol Head Neck Surg [Internet]. 2024;171(6):1751–7. Available from: http://dx.doi.org/10.1002/ohn.927 Tan TF, Thirunavukarasu AJ, Campbell JP, Keane PA, Pasquale LR, Abramoff MD, et al. Generative artificial intelligence through ChatGPT and other large language models in ophthalmology: Clinical applications and challenges. Ophthalmol Sci [Internet]. 2023;3(4):100394. Available from: http://dx.doi.org/10.1016/j.xops.2023.100394 Momenaei B, Wakabayashi T, Shahlaee A, Durrani AF, Pandit SA, Wang K, et al. Appropriateness and readability of ChatGPT-4-generated responses for surgical treatment of retinal diseases. Ophthalmol Retina [Internet]. 2023;7(10):862–8. Available from: http://dx.doi.org/10.1016/j.oret.2023.05.022 Cohen SA, Brant A, Fisher AC, Pershing S, Do D, Pan C. Dr. Google vs. Dr. ChatGPT: Exploring the use of artificial intelligence in ophthalmology by comparing the accuracy, safety, and readability of responses to frequently asked patient questions regarding cataracts and cataract surgery. Semin Ophthalmol [Internet]. 2024;39(6):472–9. Available from: http://dx.doi.org/10.1080/08820538.2024.2326058 Bean AM, Payne RE, Parsons G, Kirk HR, Ciro J, Mosquera-Gómez R, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat Med [Internet]. 2026 Feb 9 [cited 2026 Mar 29];32(2):609–15. Available from: http://dx.doi.org/10.1038/s41591-025-04074-y Chen S, Gao M, Sasse K, Hartvigsen T, Anthony B, Fan L, et al. When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior. NPJ Digit Med [Internet]. 2025 Oct 17 [cited 2026 Mar 29];8(1):605. Available from: http://dx.doi.org/10.1038/s41746-025-02008-z Zack T, Lehman E, Suzgun M, Rodriguez JA, Celi LA, Gichoya J, et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit Health [Internet]. 2024 Jan 1 [cited 2026 Mar 29];6(1):e12–22. Available from: http://dx.doi.org/10.1016/S2589-7500(23)00225-X Teo ZL, Quek CWN, Wong JLY, Ting DSW. Cybersecurity in the generative artificial intelligence era. Asia Pac J Ophthalmol (Phila) [Internet]. 2024 Jul 1 [cited 2026 Mar 29];13(4):100091. Available from: http://dx.doi.org/10.1016/j.apjo.2024.100091 Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digit Med [Internet]. 2023;6(1):120. Available from: http://dx.doi.org/10.1038/s41746-023-00873-0 Muir KW, Lee PP. Health literacy and ophthalmic patient education. Surv Ophthalmol [Internet]. 2010 Sep 21 [cited 2026 Feb 8];55(5):454–9. Available from: http://dx.doi.org/10.1016/j.survophthal.2010.03.005 Schillinger D, Grumbach K, Piette J, Wang F, Osmond D, Daher C, et al. Association of health literacy with diabetes outcomes. JAMA [Internet]. 2002 Jul 24 [cited 2026 Feb 8];288(4):475–82. Available from: http://dx.doi.org/10.1001/jama.288.4.475 Juzych MS, Randhawa S, Shukairy A, Kaushal P, Gupta A, Shalauta N. Functional health literacy in patients with glaucoma in urban settings. Arch Ophthalmol [Internet]. 2008 May [cited 2026 Feb 8];126(5):718–24. Available from: http://dx.doi.org/10.1001/archopht.126.5.718 Muir KW, Santiago-Turla C, Stinnett SS, Herndon LW, Allingham RR, Challa P, et al. Health literacy and adherence to glaucoma therapy. Am J Ophthalmol [Internet]. 2006 Aug [cited 2026 Feb 8];142(2):223–6. Available from: http://dx.doi.org/10.1016/j.ajo.2006.03.018 Kloosterboer A, Yannuzzi NA, Patel NA, Kuriyan AE, Sridhar J. Assessment of the quality, content, and readability of freely available online information for patients regarding diabetic retinopathy. JAMA Ophthalmol [Internet]. 2019 Nov 1 [cited 2026 Feb 8];137(11):1240–5. Available from: http://dx.doi.org/10.1001/jamaophthalmol.2019.3116 Kloosterboer A, Yannuzzi N, Topilow N, Patel N, Kuriyan A, Sridhar J. Assessing the quality, content, and readability of freely available online information for patients regarding age-related macular degeneration. Semin Ophthalmol [Internet]. 2021 Aug 18 [cited 2026 Feb 8];36(5–6):400–5. Available from: http://dx.doi.org/10.1080/08820538.2021.1893761 Yun HS, Bickmore T. Online health information-seeking in the era of large language models: Cross-sectional web-based survey study. J Med Internet Res [Internet]. 2025;27(1):e68560. Available from: http://dx.doi.org/10.2196/68560 Abbas ASH, Ong AY, Antaki F, Akhtar HN, Shehab M, Keane PA. An updated analysis of large language model performance on ophthalmology speciality examinations. Eye (Lond) [Internet]. 2026;1–3. Available from: http://dx.doi.org/10.1038/s41433-026-04262-1 Thirunavukarasu AJ, Mahmood S, Malem A, Foster WP, Sanghera R, Hassan R, et al. Large language models approach expert-level clinical knowledge and reasoning in ophthalmology: A head-to-head cross-sectional study. PLOS Digit Health [Internet]. 2024;3(4):e0000341. Available from: http://dx.doi.org/10.1371/journal.pdig.0000341 Agnihotri AP, Nagel ID, Artiaga JCM, Guevarra MCB, Sosuan GMN, Kalaw FGP. Large language models in ophthalmology: A review of publications from top ophthalmology journals. Ophthalmol Sci [Internet]. 2025 May 1 [cited 2025 Nov 23];5(3):100681. Available from: http://dx.doi.org/10.1016/j.xops.2024.100681 Cheong KX, Zhang C, Tan TE, Fenner BJ, Wong WM, Teo KY, et al. Comparing generative and retrieval-based chatbots in answering patient questions regarding age-related macular degeneration and diabetic retinopathy. Br J Ophthalmol [Internet]. 2024;108(10):1443–9. Available from: http://dx.doi.org/10.1136/bjo-2023-324533 Manchester Royal Eye Hospital [Internet]. 2018 [cited 2025 Nov 9]. Patient Leaflets. Available from: https://mft.nhs.uk/royal-eye/patients-visitors/patient-leaflets/?pl_category_id=44 StatCounter Global Stats [Internet]. [cited 2025 Nov 9]. AI Chatbot Market Share Worldwide. Available from: https://gs.statcounter.com/ai-chatbot-market-share#monthly-202503-202508 Tan TF, Elangovan K, Ong J, Shah N, Sung J, Wong TY, et al. A proposed S.c.o.r.e. evaluation framework for large language models: Safety, Consensus, Objectivity, Reproducibility and Explainability [Internet]. arXiv [cs.CL]. 2024. Available from: http://arxiv.org/abs/2407.07666 Shedlosky-Shoemaker R, Sturm AC, Saleem M, Kelly KM. Tools for assessing readability and quality of health-related Web sites. J Genet Couns [Internet]. 2009;18(1):49–59. Available from: http://dx.doi.org/10.1007/s10897-008-9181-0 Tan TF, Elangovan K, Pollreisz A, Dy KB, Ng WY, Wong JLY, et al. Clinical validation of medical-based large language model chatbots on ophthalmic patient queries with LLM-based evaluation [Internet]. arXiv [cs.AI]. 2026. Available from: http://dx.doi.org/10.48550/arXiv.2602.05381 Gu J, Jiang X, Shi Z, Tan H, Zhai X, Xu C, et al. A Survey on LLM-as-a-Judge [Internet]. arXiv [cs.CL]. 2025. Available from: http://dx.doi.org/10.48550/arXiv.2411.15594 Croxford E, Gao Y, First E, Pellegrino N, Schnier M, Caskey J, et al. Automating evaluation of AI text generation in healthcare with a large language model (LLM)-as-a-Judge [Internet]. medRxiv. 2025. Available from: http://medrxiv.org/lookup/doi/ 10.1101/2025.04.22.25326219 CHART Collaborative, Huo B, Collins G, Chartash D, Thirunavukarasu A, Flanagin A, et al. Reporting guideline for chatbot health advice studies: The CHART statement. Artif Intell Med [Internet]. 2025 Oct [cited 2026 Feb 8];168(103222):103222. Available from: http://dx.doi.org/10.1016/j.artmed.2025.103222 White J, Fu Q, Hays S, Sandborn M, Olea C, Gilbert H, et al. A prompt pattern catalog to enhance prompt engineering with ChatGPT [Internet]. arXiv [cs.SE]. 2023. Available from: http://arxiv.org/abs/2302.11382 Liu P, Yuan W, Fu J, Jiang Z, Hayashi H, Neubig G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing [Internet]. arXiv [cs.CL]. 2021. Available from: http://arxiv.org/abs/2107.13586 Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med [Internet]. 2016;15(2):155–63. Available from: http://dx.doi.org/10.1016/j.jcm.2016.02.012 [cited 2026 Apr 10]. Available from: https://www.cosprc.ca/wp-content/uploads/2022/08/Vision-Therapy-Position-Statement_FINAL_ENGLISH.pdf Spadaro DC, Robinson LA, Smith LT. Assessing readability of patient information materials. Am J Health Syst Pharm [Internet]. 1980;37(2):215–21. Available from: https://dx.doi.org/10.1093/ajhp/37.2.215 Powell M. Health information: are you getting your message across? [Internet]. National Institute for Health Research; 2022. Available from: http://dx.doi.org/10.3310/nihrevidence_51109 Wang J, Shi R, Le Q, Shan K, Chen Z, Zhou X, et al. Evaluating the effectiveness of large language models in patient education for conjunctivitis. Br J Ophthalmol [Internet]. 2025;109(2):185–91. Available from: http://dx.doi.org/10.1136/bjo-2024-325599 Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: a scoping review of applications in medicine. Front Med (Lausanne) [Internet]. 2024;11:1477898. Available from: http://dx.doi.org/10.3389/fmed.2024.1477898 Klerings I, Weinhandl AS, Thaler KJ. Information overload in healthcare: too much of a good thing? Z Evid Fortbild Qual Gesundhwes [Internet]. 2015;109(4–5):285–90. Available from: http://dx.doi.org/10.1016/j.zefq.2015.06.005 [cited 2025 Dec 14]. Available from: https://assets.publishing.service.gov.uk/media/675330e020bcf083762a6d48/Survey_of_Adult_Skills_2023__PIAAC__National_Report_for_England.pdf Campbell DJ, Estephan LE, Mastrolonardo EV, Amin DR, Huntley CT, Boon MS. Evaluating ChatGPT responses on obstructive sleep apnea for patient education. J Clin Sleep Med [Internet]. 2023;19(12):1989–95. Available from: http://dx.doi.org/10.5664/jcsm.10728 Kianian R, Sun D, Crowell EL, Tsui E. The use of large language models to generate education materials about uveitis. Ophthalmol Retina [Internet]. 2024;8(2):195–201. Available from: http://dx.doi.org/10.1016/j.oret.2023.09.008 Bondok M, Selvakumar R, Law C, Ing EB, Bakshi NK, Felfeli T. Comparing ophthalmologist and Artificial Intelligence Chatbot responses to patient questions. Clin Ophthalmol [Internet]. 2025;19:4293–300. Available from: http://dx.doi.org/10.2147/OPTH.S549820 Srinivasan N, Samaan JS, Rajeev ND, Kanu MU, Yeo YH, Samakar K. Large language models and bariatric surgery patient education: a comparative readability analysis of GPT-3.5, GPT-4, Bard, and online institutional resources. Surg Endosc [Internet]. 2024;38(5):2522–32. Available from: http://dx.doi.org/10.1007/s00464-024-10720-2 King RC, Samaan JS, Haquang J, Bharani V, Margolis S, Srinivasan N, et al. Improving the readability of institutional heart failure-related patient education materials using GPT-4: Observational study. JMIR Cardio [Internet]. 2025;9( v9i8e68817 ):e68817. Available from: http://dx.doi.org/10.2196/68817 Hershenhouse JS, Mokhtar D, Eppler MB, Rodler S, Storino Ramacciotti L, Ganjavi C, et al. Accuracy, readability, and understandability of large language models for prostate cancer information to the public. Prostate Cancer Prostatic Dis [Internet]. 2025;28(2):394–9. Available from: http://dx.doi.org/10.1038/s41391-024-00826-y Cohen SA, Yadlapalli N, Tijerina JD, Alabiad CR, Chang JR, Kinde B, et al. Comparing the Ability of Google and ChatGPT to Accurately Respond to Oculoplastics-Related Patient Questions and Generate Customized Oculoplastics Patient Education Materials. Clin Ophthalmol [Internet]. 2024;18:2647–55. Available from: http://dx.doi.org/10.2147/OPTH.S480222 Schwarz K, Liao Y, Geiger A. On the frequency bias of generative models [Internet]. arXiv [cs.CV]. 2021. Available from: http://arxiv.org/abs/2111.02447 Ahsan H, Sharma AS, Amir S, Bau D, Wallace BC. Elucidating mechanisms of demographic bias in LLMs for healthcare [Internet]. arXiv [cs.CL]. 2025. Available from: http://arxiv.org/abs/2502.13319 Lynn-Green EE, Ofoje AA, Lynn-Green RH, Jones DS. Variations in how medical researchers report patient demographics: a retrospective analysis of published articles. EClinicalMedicine [Internet]. 2023;58(101903):101903. Available from: http://dx.doi.org/10.1016/j.eclinm.2023.101903 Mirzania D, Fleming TL, Robbins CB, Feng HL, Fekrat S. Time to presentation after symptom onset in endophthalmitis: Clinical features and visual outcomes. Ophthalmol Retina [Internet]. 2021;5(4):324–9. Available from: http://dx.doi.org/10.1016/j.oret.2020.07.027 Hazelwood JE, Mitry D, Singh J, Bennett HGB, Khan AA, Goudie CR. The Scottish Retinal Detachment Study: 10-year outcomes after retinal detachment repair. Eye (Lond) [Internet]. 2025;39(7):1318–21. Available from: http://dx.doi.org/10.1038/s41433-025-03613-8 OpenAI. cdn.openai.com. 2026. AI as a Healthcare Ally: How Americans are navigating the system with ChatGPT. Available from: https://cdn.openai.com/pdf/2cb29276-68cd-4ec6-a5f4-c01c5e7a36e9/OpenAI-AI-as-a-Healthcare-Ally-Jan-2026 .pdf#:~:text=Americans%20are%20using%20AI%20and,to%20navigate%20and%20makes%20decisions Ramaswamy A, Tyagi A, Hugo H, Jiang J, Jayaraman P, Jangda M, et al. ChatGPT Health performance in a structured test of triage recommendations. Nat Med [Internet]. 2026;1–5. Available from: http://dx.doi.org/10.1038/s41591-026-04297-7 Sharma S, Alaa AM, Daneshjou R. A longitudinal analysis of declining medical safety messaging in generative AI models. NPJ Digit Med [Internet]. 2025;8(1):592. Available from: http://dx.doi.org/10.1038/s41746-025-01943-1 Cohen SA, Fisher AC, Xu BY, Song BJ. Comparing the accuracy and readability of glaucoma-related Question Responses and Educational Materials by Google and ChatGPT. J Curr Glaucoma Pract [Internet]. 2024;18(3):110–6. Available from: http://dx.doi.org/10.5005/jp-journals-10078-1448 Trillo-Domínguez M, Martin-Neira JI, Olvera-Lobo MD. Dr. Google vs. Dr. ChatGPT in online health self-consultation: A scoping review of accuracy, bias, and actionability (2023–2025). Informatics (MDPI) [Internet]. 2026;13(3):41. Available from: http://dx.doi.org/10.3390/informatics13030041 Menz BD, Kuderer NM, Bacchi S, Modi ND, Chin-Yee B, Hu T, et al. Current safeguards, risk mitigation, and transparency measures of large language models against the generation of health disinformation: repeated cross sectional analysis. BMJ [Internet]. 2024;384:e078538. Available from: http://dx.doi.org/10.1136/bmj-2023-078538 Khair DO, Kale AU, Agbakoba R, Goh E, Mateen BA, Ong AY, et al. Building the health chatbot users’ guide. Nat Health [Internet]. 2026;1–2. Available from: http://dx.doi.org/10.1038/s44360-026-00074-5 Asgari E, Montaña-Brown N, Dubois M, Khalil S, Balloch J, Yeung JA, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digit Med [Internet]. 2025;8(1):274. Available from: http://dx.doi.org/10.1038/s41746-025-01670-7 Tables Table 1 is available in the Supplementary Files section. Additional Declarations There is no conflict of interest Supplementary Files SupplementaryTable3.xlsx Supplementary Table 3: Intraclass correlation coefficients (ICC (2,1)) between human assessor and large language model (LLM) evaluators across three patient information sources. Supplementaryfigure2.png Supplementary Figure 2: Breakdown CASEF scores stratified by assessor. Table1.xlsx Table 1 SupplementaryTable2.xlsx Supplementary Table 2: Comprehensiveness, Accuracy and Safety Evaluation Framework (CASEF) scoring scale. SupplementaryFigure1.pdf Supplementary Figure 1: Complete scoring prompt. SupplementaryTable1.pdf Supplementary Table 1: Text outputs from generative AI chatbots and clinical patient information leaflets across nine ophthalmology conditions. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9383173","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":624208361,"identity":"3327b9bf-1318-4a39-8520-4016aa7e16ac","order_by":0,"name":"Aleksander Stupnicki","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABR0lEQVRIie2RMWvCQBTHXzg4l1PXK636FSK3+mFyCDrZpVAcHA6EdBFdDYX2K+Qb9MKDuITOHbKUQqYOilCyaHup1tZI6FpofjccvLvf3f/xAEpK/iIVINlmqa8CNUvDsNMEILsizynkp6IPStQTh3d+VTIJLBelKlDqY4LtdBQ3ZuearFdufFmrhFIDJX1/EShIRyA9daRwpN0uCxPhTR3KAze5oqynNTA68COprEkI8jYXDJlAoCj9yAQKTB6XV5QGzgb+kwlWVSDvjo0W1tdBukX5EAFZfis279tGsTanio2MOFVz02cm5U6hJphjO5lCsl9ywdpIhahOUcwjy+XRYyJd1nO0WW3P9IIXIRe59puL8ctZ+oaN2YTgcngdy/ubUCxX2/dWbYHB8+uo05hrKOBzCvtTZ1/Tp4M8ofDBkpKSkn/MBztzgHfwOQmKAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0002-4840-5413","institution":"University College London","correspondingAuthor":true,"prefix":"","firstName":"Aleksander","middleName":"","lastName":"Stupnicki","suffix":""},{"id":624208362,"identity":"37f09f09-30e7-499d-8576-5bb1e3a92815","order_by":1,"name":"Bernardo Mendes","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Bernardo","middleName":"","lastName":"Mendes","suffix":""},{"id":624208363,"identity":"e37ab23c-2af8-4299-a8e2-04a730dabf02","order_by":2,"name":"Maxwell Reinstein","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Maxwell","middleName":"","lastName":"Reinstein","suffix":""},{"id":624208364,"identity":"f08e9288-9283-447e-af11-e5f4485c71e4","order_by":3,"name":"Ariel Ong","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Ariel","middleName":"","lastName":"Ong","suffix":""},{"id":624208365,"identity":"49efcd24-3288-4220-9038-cbc0ed469359","order_by":4,"name":"Andrew Malem","email":"","orcid":"","institution":"Cleveland Clinic Abu Dhabi","correspondingAuthor":false,"prefix":"","firstName":"Andrew","middleName":"","lastName":"Malem","suffix":""},{"id":624208366,"identity":"2fdccab4-06eb-432a-aeff-0e7894a2a0fa","order_by":5,"name":"Pearse Keane","email":"","orcid":"https://orcid.org/0000-0002-9239-745X","institution":"UCL Institute of Ophthalmology","correspondingAuthor":false,"prefix":"","firstName":"Pearse","middleName":"","lastName":"Keane","suffix":""},{"id":624208367,"identity":"e7f187b9-ce89-48cd-b237-dd528356bec7","order_by":6,"name":"Arun Thirunavukarasu","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Arun","middleName":"","lastName":"Thirunavukarasu","suffix":""}],"badges":[],"createdAt":"2026-04-10 21:50:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9383173/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9383173/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107487277,"identity":"649bf19a-caf1-4fc7-82b1-76bbc0367f9e","added_by":"auto","created_at":"2026-04-22 02:40:19","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1488880,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy Methodology Overview. \u003c/strong\u003eWorkflow for generating and evaluating genAI health advice against an established gold standard. \u003cstrong\u003e(A) \u003c/strong\u003eDefining NHS patient information leaflets as the control for patient information materials and producing subheading-matched AI-generated patient information materials. \u003cstrong\u003e(B)\u003c/strong\u003eBenchmark curation assigning combined leaflets from the Royal College of Ophthalmologists (RCOphth) and Royal National Institute of Blind People (RNIB) as the gold-standard guidelines for nine ophthalmic conditions. \u003cstrong\u003e(C)\u003c/strong\u003eEvaluation pathway comparing NHS and LLM-generated patient information materials, using CASEF (Comprehensiveness, Accuracy, and Safety Evaluation Framework) and multiple readability scoring systems, with a final qualitative safety review performed by a senior ophthalmology resident (AYO) and a consultant ophthalmologist (AM). Abbreviations: AMD, age-related macular degeneration; CBS, Charles Bonnet syndrome; DR, diabetic retinopathy; PVD, posterior vitreous detachment; RD, retinal detachment; SMOG, Simple Measure of Gobbledygook; FKGL, Flesch-Kincaid Grade Level; ARI, Automated Readability Index.\u003c/p\u003e","description":"","filename":"Figure1FINAL.png","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/0e0752f3f4aeebcb176275b0.png"},{"id":107378137,"identity":"c0ed310e-ab1e-4b75-8179-0d82a909e03c","added_by":"auto","created_at":"2026-04-21 01:33:30","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1139542,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverall quality of genAI advice and NHS material across four evaluation domains.\u003c/strong\u003e GPT-5 (purple), Gemini 3 (blue), and NHS leaflets (orange) were compared on total accuracy, safety, comprehensiveness and readability. \u003cstrong\u003e(A)\u003c/strong\u003e Both genAI models demonstrated CASEF (Comprehensiveness, Accuracy, Safety Evaluation Framework) scores superior to NHS leaflets, representing significantly higher factual alignment with the benchmark (\u003cem\u003ep\u0026lt;\u003c/em\u003e0.001). \u003cstrong\u003e(B)\u003c/strong\u003e Both genAI models demonstrated statistically comparable safety scores to NHS leaflets (values closer to zero indicate superior safety). \u003cstrong\u003e(C)\u003c/strong\u003e Both genAI models demonstrated significantly fewer omissions than NHS leaflets (\u003cem\u003ep\u0026lt;\u003c/em\u003e0.001), indicating higher comprehensiveness. \u003cstrong\u003e(D) \u003c/strong\u003eReadability assessment using three validated indices: Flesch-Kincaid Grade Level (FKGL), Automated Readability Index (ARI) and Simple Measure of Gobbledygook (SMOG), where higher scores indicate more complex texts requiring additional years of education for comprehension. The green horizontal band denotes the reading level recommended by the National Institutes of Health for patient-facing materials (sixth-eighth-grade reading level), which was not met by any of the tested materials. Both genAI models produced significantly less accessible texts compared with NHS materials across all three indices. All values are presented as mean ± standard error of the mean. Statistical comparisons were performed using paired t-tests with Bonferroni correction (adjusted α = 0.0167 for all analyses). *\u003cem\u003ep\u0026lt;\u003c/em\u003e0.05; ***\u003cem\u003ep\u0026lt;\u003c/em\u003e0.001; ns, not significant.\u003c/p\u003e","description":"","filename":"Figure2FINAL.png","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/83555ddf1ab3e42436aff162.png"},{"id":107485908,"identity":"6960bc5c-4d7b-4102-9778-0e72b56145da","added_by":"auto","created_at":"2026-04-22 02:36:45","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":267461,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStratified analysis of total CASEF scores by condition and question category. \u003c/strong\u003eGPT-5 outputs (purple), Gemini 3 outputs (blue) and NHS leaflets were compared across nine conditions and eight question categories. Yellow shading highlights instances where at least one genAI model demonstrated significantly higher scores compared with NHS leaflets. (A) Difference in total CASEF scores stratified by condition. Both genAI models significantly outperformed NHS leaflets for dry eye disease and nystagmus content (\u003cem\u003ep\u0026lt;\u003c/em\u003e0.001), whilst no significant differences were observed for the remaining seven conditions. (B) Difference in total CASEF score between each genAI model and NHS leaflets, stratified by question category. Positive values (rightward) indicate genAI superiority; negative values (leftward) indicate NHS superiority. Both genAI models achieved significantly higher scores for disease definition content (\u003cem\u003ep\u0026lt;\u003c/em\u003e0.001). NHS leaflets demonstrated significantly higher concordance for treatment complication content (Gemini 3 versus NHS, \u003cem\u003ep\u0026lt;\u003c/em\u003e0.001; GPT-5 versus NHS, p = 0.003). Statistical comparisons were performed using paired t-tests. **\u003cem\u003ep\u0026lt;\u003c/em\u003e0.01; ***\u003cem\u003ep\u0026lt;\u003c/em\u003e0.001.\u003c/p\u003e","description":"","filename":"Figure3FINAL.png","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/73c867b6258800fcf32b6d55.png"},{"id":107378135,"identity":"bc023491-e415-4f6b-a036-d1c53588c9c2","added_by":"auto","created_at":"2026-04-21 01:33:30","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":9718,"visible":true,"origin":"","legend":"Supplementary Table 3: Intraclass correlation coefficients (ICC (2,1)) between human assessor and large language model (LLM) evaluators across three patient information sources.","description":"","filename":"SupplementaryTable3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/2a68ebce37cc2a2f82bca13b.xlsx"},{"id":107487791,"identity":"a0b7e7bc-cc33-46c3-813d-da03172d9029","added_by":"auto","created_at":"2026-04-22 02:42:48","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":353856,"visible":true,"origin":"","legend":"Supplementary Figure 2: Breakdown CASEF scores stratified by assessor.","description":"","filename":"Supplementaryfigure2.png","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/f426bf9e8b040edfb3926439.png"},{"id":107378139,"identity":"cbb85024-5803-46e2-b9f7-1d47ff3664f1","added_by":"auto","created_at":"2026-04-21 01:33:30","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":11188,"visible":true,"origin":"","legend":"Table 1","description":"","filename":"Table1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/33366b5c5f878ce10452ce62.xlsx"},{"id":107488546,"identity":"e3aa4a0f-a7be-4ad5-bcbc-7412bda916f0","added_by":"auto","created_at":"2026-04-22 02:45:03","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":9665,"visible":true,"origin":"","legend":"Supplementary Table 2: Comprehensiveness, Accuracy and Safety Evaluation Framework (CASEF) scoring scale.","description":"","filename":"SupplementaryTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/a7633e2531989afc41611bc4.xlsx"},{"id":107378142,"identity":"9e8e53e8-1960-4ec1-879f-7d9e2100983a","added_by":"auto","created_at":"2026-04-21 01:33:30","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":55099,"visible":true,"origin":"","legend":"Supplementary Figure 1: Complete scoring prompt.","description":"","filename":"SupplementaryFigure1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/9b0e572d43e1c3013e870244.pdf"},{"id":107487364,"identity":"49e72c13-c64d-45cc-9f6a-f982cdce243a","added_by":"auto","created_at":"2026-04-22 02:41:06","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":695309,"visible":true,"origin":"","legend":"Supplementary Table 1: Text outputs from generative AI chatbots and clinical patient information leaflets across nine ophthalmology conditions.","description":"","filename":"SupplementaryTable1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9383173/v1/21cf8baf83bb4322e08b8747.pdf"}],"financialInterests":"There is no conflict of interest","formattedTitle":"Benchmarking eye health advice from generative artificial intelligence in terms of factual accuracy, safety, comprehensiveness and readability","fulltext":[{"header":"Glossary","content":"\u003cul\u003e\n \u003cli\u003e\u003cstrong\u003eGenerative Artificial Intelligence (GenAI):\u003c/strong\u003e An umbrella term capturing computational systems that are trained and fine-tuned to produce material in response to a user which may comprise text, images, or audio.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003eLarge Language Models (LLMs):\u003c/strong\u003e\u0026nbsp; \u0026nbsp;A subset of generative artificial intelligence systems trained with large volumes of textual data and exhibiting abilities to interpret and produce text and thereby engage in conversation or question-answering.\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003eGeneral-purpose models:\u003c/strong\u003e GenAI models designed to perform a wide range of tasks across diverse-domains, in contrast to domain-specific models, which are trained and tuned for narrower applications such as healthcare.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003eTraining corpus:\u003c/strong\u003e The large, curated dataset of texts (drawn from sources such as books, websites, and scientific literature) used to train LLMs.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003eHallucinations:\u003c/strong\u003e Phenomenon where genAI text contains information that is fabricated (not present in the training dataset). \u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003ePrompt:\u003c/strong\u003e The input submitted by a user to a genAI model to elicit a response, typically an instruction or a question.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003ePrompt engineering:\u003c/strong\u003e Techniques of designing and refining prompts to encourage desired behaviour or outputs from genAI.\u0026nbsp;\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"What was known before","content":"\u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eGenerative artificial intelligence chatbots are widely used by patients for gathering health information despite lacking regulatory approval for medical use. Early models were associated with high rates of hallucinations and factual inaccuracies, raising concerns about patient reliance on AI-generated health advice.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePrior evaluations of LLMs in ophthalmology focused predominantly on performance in knowledge examination tasks, where latest models match or exceed expert performance. Whether that capability translates to production of reliable patient-facing materials on common eye conditions is unclear.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eWhat this study adds\u003c/b\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eThis study demonstrates that current general-purpose AI chatbots can match or exceed clinical patient education materials in factual accuracy and comprehensiveness.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eHowever, current genAI models exhibit clinically-relevant safety concerns, including systematic underreporting of treatment complications, defaulting to US-centric guidelines, inadequate safety-netting for sight-threatening conditions, and high linguistic complexity.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThese findings highlight that before general-purpose genAI models can serve as reliable alternatives to formal patient education materials, robust regulatory oversight and patient education on the risks of genAI are essential.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e"},{"header":"1. Introduction","content":"\u003cp\u003eGenerative artificial intelligence (genAI) describes a set of computational systems encompassing large language models (LLMs), which are capable of generating human-like text in response to natural language prompts (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Since the public release of conversational LLM-based chatbots in late 2022, these tools have been rapidly adopted as sources of health information or advice: recent reports suggest up to 32% of adults turn to genAI for healthcare advice, (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) and 25% of ChatGPT users submit health-related questions on a weekly basis (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eBy generating tailored responses on-demand, LLMs offer personalised care and have the potential to improve patient health literacy, engagement and ownership of health decisions (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). However significant risks associated with LLM use have been raised across medical literature, including inaccurate advice, clinically important omissions, hallucinations, agreement with user misconceptions, and unsafe risk communication (\u003cspan additionalcitationids=\"CR7 CR8 CR9 CR10 CR11\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). The lack of formal regulatory approval of general-purpose LLMs for medical use further limits oversight and accountability (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn addition to the quality of genAI advice, linguistic features are an important consideration (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). In ophthalmology, poor health literacy is associated with delayed diagnosis, medication non-adherence, and greater carer dependency, contributing to adverse outcomes in conditions such as glaucoma and diabetic retinopathy (\u003cspan additionalcitationids=\"CR16\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). In eye clinics, leaflets are frequently used to provide patients with advice that is evidence-based and pitched at an understandable level. However, patients have long been obtaining medical advice from sources outside of those directly provided in the healthcare setting, primarily through online search engines, despite reportedly inferior quality and readability (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e). With the rise in genAI adoption, this behaviour is now rapidly evolving, and patients frequently encounter AI-generated advice before consulting formal sources (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn ophthalmology, leading genAI models now surpass human-expert performance on examination-style questions (\u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e), but whether this capability extends to patient-facing advice, where accuracy must be balanced with nuanced safety communication, accessibility, and coverage, is less well established (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAlthough publicly accessible genAI chatbots were neither designed for clinical applications nor approved by regulators for medical use, patients employ these systems in ways directly impacting their healthcare interactions. Acknowledging this real-world usage, we assessed whether leading general-purpose LLM platforms (GPT-5, OpenAI and Gemini 3, Google DeepMind) produce eye health advice suitable for patients. Using traditional patient information leaflets as reference standards, we evaluated outputs from leading genAI models across four domains: comprehensiveness, factual accuracy, safety, and readability.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Control arm leaflet selection\u003c/h2\u003e \u003cp\u003eThe United Kingdom\u0026rsquo;s National Health Service (NHS) patient information leaflets served as control comparators, representing real-world materials routinely used in clinical practice. Online patient education repositories from four large UK tertiary ophthalmology centres (Moorfields Eye Hospital, Manchester University, Imperial College Healthcare, and Oxford University Hospitals NHS Foundation Trusts) were evaluated for content alignment with the nine benchmark conditions (see Section \u003cspan refid=\"Sec4\" class=\"InternalRef\"\u003e2.2\u003c/span\u003e) and informational depth. Manchester University NHS Foundation Trust was selected as the control source, as it was the only repository that provided complete coverage of all nine conditions, used question-format subheadings, and also demonstrated the greatest informational depth (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Benchmark curation\u003c/h2\u003e \u003cp\u003eA checklist-based benchmark was developed to evaluate AI-generated text against expert-authored patient information. Source material comprised all nine patient information leaflets produced by the Royal College of Ophthalmologists (RCOphth) in collaboration with the Royal National Institute of Blind People (RNIB), available at the time of study. The leaflets covered age-related macular degeneration (AMD), cataract, Charles Bonnet syndrome (CBS), diabetic retinopathy (DR), dry eye disease, glaucoma, nystagmus, posterior vitreous detachment (PVD), and retinal detachment (RD).\u003c/p\u003e \u003cp\u003eTwo researchers (AS, BSM) independently mapped NHS leaflet subsections to corresponding RCOphth/RNIB benchmark domains, identifying 41 subsections across the nine conditions that aligned with benchmark content. Each subsection was then converted into discrete factual checklist items, with discrepancies resolved through discussion and duplicate statements consolidated to prevent score inflation. This yielded 547 benchmark items across 41 sections, which formed the basis for subsequent AI output generation.\u003c/p\u003e \u003cp\u003eEach question was retrospectively categorised into one of eight thematic groups for subgroup analysis: disease definition, causes, symptoms, diagnosis, risk factors, treatment, complications, and \u0026lsquo;other\u0026rsquo;. The full benchmark is available on request.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 AI advice generation\u003c/h2\u003e \u003cp\u003eGenAI advice was elicited from two general-purpose chatbot platforms: GPT-5 (accessed October 4, 2025, via chat.openai.com; OpenAI, San Francisco, CA, USA) and Gemini 3.0 (accessed November 21, 2025, via gemini.google.com; Google DeepMind, London, UK). Both are instruction-tuned conversational variants selected due to highest public adoption at the time of the study (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e). Both models were accessed from London, UK, via web-based interfaces using newly registered accounts and default generation parameters without modification.\u003c/p\u003e \u003cp\u003eOutputs were generated using matched NHS leaflet subsection headings as verbatim prompts with an explicit word count constraint matching the corresponding NHS text length (\u003cem\u003ee.g.\u003c/em\u003e \"What is age-related macular degeneration (AMD)? Restrict your answer to 154 words\"). This controlled for response length as a confounder. In four instances where NHS subsection headings were phrased as declarative statements (\u003cem\u003ee.g.\u003c/em\u003e \"Your operation\"), these were rephrased into interrogative or imperative formats (\u003cem\u003ee.g.\u003c/em\u003e \"Explain the operation\") to align with standard LLM query conventions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Evaluation framework\u003c/h2\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.4.1 Comprehensiveness, Accuracy and Safety Evaluation Framework (CASEF)\u003c/h2\u003e \u003cp\u003eA novel scoring framework was developed to quantify the concordance between text outputs (NHS leaflets and genAI advice) and the gold-standard benchmark across three core dimensions: comprehensiveness, accuracy, and safety. The CASEF framework (Supplementary Table\u0026nbsp;2) employs a five-point bidirectional scale that captures both errors of commission and omission, similar to validated Likert-type frameworks. CASEF draws on the elements of the S.C.O.R.E. framework (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e) most relevant to patient-facing health information: safety and consensus (corresponding to negative CASEF scores and total CASEF score respectively), with comprehensiveness representing an extension of the consensus domain.\u003c/p\u003e \u003cp\u003eEach text output was evaluated against the benchmark items corresponding to its topic (Supplementary Table\u0026nbsp;1), with each item independently scored by a multi-evaluator panel. Item-level scores were summed and normalised as a percentage of the maximum attainable score to account for variation in benchmark item counts across questions, yielding a range from \u0026minus;\u0026thinsp;100% (total contradiction) to 100% (complete concordance).\u003c/p\u003e \u003cp\u003eSafety and comprehensiveness metrics were derived directly from CASEF. Safety was quantified as the sum of negative CASEF scores (-1 or -2) for a given text sample, and comprehensiveness as the total count of items scored 0 (omissions), with values closer to zero indicating better performance with respect to each metric.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e2.4.2 Qualitative safety evaluation\u003c/h2\u003e \u003cp\u003eQuantitative safety appraisal via CASEF was supplemented with a manual safety evaluation by a senior ophthalmology resident (AYO) and a consultant ophthalmologist (AM) who were blinded to the source of each text. All genAI outputs (n\u0026thinsp;=\u0026thinsp;82) were independently reviewed for factual accuracy, guideline alignment, risk communication, and other safety concerns not captured by the scoring framework.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e2.4.3 Readability\u003c/h2\u003e \u003cp\u003eReadability was assessed using three validated metrics: the Simple Measure of Gobbledygook (SMOG) Index, Flesch-Kincaid Grade Level (FKGL), and Automated Readability Index (ARI) (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e). SMOG estimates reading level from polysyllabic word frequency, while FKGL and ARI derive US school grade-level scores from sentence and word length. All 123 text samples were analysed per source (NHS, GPT-5, Gemini 3.0) using the Python textstat library (v0.7.10).\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Evaluation process\u003c/h2\u003e \u003cp\u003eA hybrid multi-evaluator panel comprising five genAI models and one human researcher (AS) was employed for text scoring. This 'LLM-as-a-judge' approach addresses a key limitation of traditional AI evaluation studies, where reliance on human evaluators alone reduces reproducibility and introduces subjectivity (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). Recent validation studies demonstrated scoring consistency and strong inter-rater reliability between human expert raters and LLM evaluators in specialised medical contexts (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFive latest public models (at the time of evaluation (November 2025)) were selected to average out model-specific biases: GPT-5 (OpenAI, San Francisco, CA, USA), Gemini 2.5 Pro (Google DeepMind, London, UK / Mountain View, CA, USA), Deepseek V3.2 (Deepseek AI, Hangzhou, China), Claude Sonnet 4.5 (Anthropic, San Francisco, CA, USA), and Llama 4 (Meta, Menlo Park, CA, USA). One human evaluator (AS, a student doctor) independently assessed all outputs to validate that genAI assessments aligned with human judgement.\u003c/p\u003e \u003cp\u003eEach of the 123 text samples (3 sources with 41 outputs each) was independently evaluated by all six panel members against the applicable benchmark section. All genAI models operated using default generation parameters and received identical, standardised scoring prompts (Supplementary Fig.\u0026nbsp;1), applied prior to scoring each condition. Conversation memory was manually cleaned between evaluations to prevent cross-contamination. Individual item scores from all six raters were aggregated as arithmetic means to generate final CASEF scores for each text sample. Reporting adhered to the Chatbot Assessment Reporting Tool (CHART) guidelines (Supplementary Fig.\u0026nbsp;3) (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e). No patient data was used in this study, and ethical approval is not applicable.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e2.5.1 Scoring prompt\u003c/h2\u003e \u003cp\u003eA structured evaluation prompt was engineered to ensure reproducible and systematic scoring by all LLMs in the evaluation panel. The prompt was iteratively refined through pilot testing across all five LLM evaluators using a random sample of NHS texts (n\u0026thinsp;=\u0026thinsp;5) not included in the study dataset. Established prompt engineering strategies were employed, including role prompting, task decomposition, explicit output constraints and quantitative scaffolding (\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e). Instructions to disregard stylistic differences and maintain task adherence throughout the session were included to minimise bias and hallucination.\u003c/p\u003e \u003cp\u003eAll genAI evaluators were supervised with a human-in-the-loop system for errors and adherence to prompt structure. The complete prompt is provided in Supplementary Fig.\u0026nbsp;1.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Statistical Analysis\u003c/h2\u003e \u003cp\u003eAnalyses were performed using Python 3.11.5 (Python Software Foundation, Beaverton, USA). All comparisons between information sources (genAI versus NHS and between genAI models) used paired samples t-tests with observations matched at the rater-question level to account for the repeated measures. Comparisons of overall CASEF scores, safety, comprehensiveness, and readability were conducted with Bonferroni correction (adjusted α\u0026thinsp;=\u0026thinsp;0.0167). A significance threshold of \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.01 was used for subgroup analysis stratified by condition and question category. Descriptive analyses were presented as total mean scores and a mean difference from NHS scores, both with standard deviations (SD). Inter-rater reliability was quantified using the two-way, random-effects, single-measurement intraclass correlation coefficient model (ICC(\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e)) with 95% confidence intervals (\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e). Potential preferential scoring by GPT-5 and Gemini 2.5 Pro for outputs from their respective model family was assessed by comparing each evaluator\u0026rsquo;s deviation from the all-rater consensus with the deviations observed in individual raters (Wilcoxon signed-rank test).\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Overall performance\u003c/h2\u003e \u003cp\u003eA total of 123 text outputs (82 genAI-generated and 41 NHS leaflet-derived) were evaluated against 547 checklist items, comprising frequent queries on nine ophthalmic conditions.\u003c/p\u003e \u003cp\u003eBoth genAI models demonstrated significantly higher factual alignment with the ground-truth benchmark compared to NHS leaflets (GPT-5: 37.36\u0026thinsp;\u0026plusmn;\u0026thinsp;16.91% versus NHS: 30.70\u0026thinsp;\u0026plusmn;\u0026thinsp;17.40%, \u003cem\u003ep\u0026thinsp;\u0026lt;\u003c/em\u003e\u0026thinsp;0.001; Gemini 3: 36.24\u0026thinsp;\u0026plusmn;\u0026thinsp;18.03% vs NHS, \u003cem\u003ep\u0026thinsp;\u0026lt;\u003c/em\u003e\u0026thinsp;0.001), with no significant difference between them (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.940). Similarly, both LLMs demonstrated greater comprehensiveness than NHS leaflets as evidenced by lower omission counts (GPT-5: 6.74\u0026thinsp;\u0026plusmn;\u0026thinsp;4.33 versus NHS: 7.87\u0026thinsp;\u0026plusmn;\u0026thinsp;5.14, \u003cem\u003ep\u0026thinsp;\u0026lt;\u003c/em\u003e\u0026thinsp;0.001; Gemini 3: 6.97\u0026thinsp;\u0026plusmn;\u0026thinsp;5.06 versus NHS, \u003cem\u003ep\u0026thinsp;\u0026lt;\u003c/em\u003e\u0026thinsp;0.001), with comparable comprehensiveness between models (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.215). Safety scores did not differ significantly across all three sources (GPT-5: -0.07\u0026thinsp;\u0026plusmn;\u0026thinsp;0.28; Gemini 3: -0.14\u0026thinsp;\u0026plusmn;\u0026thinsp;0.39; NHS: -0.10\u0026thinsp;\u0026plusmn;\u0026thinsp;0.32; all \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.01).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Subgroup analysis:\u003c/h2\u003e \u003cp\u003eWhen stratified by condition (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA; Table\u0026nbsp;1), both genAI models significantly outperformed NHS leaflets for dry eye disease, nystagmus, and PVD (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01 for all comparisons). Gemini 3 additionally achieved significantly higher CASEF scores for AMD (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01). Diabetic retinopathy was the only condition where NHS leaflets significantly outperformed both genAI models (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01). No significant differences were observed for the remaining four conditions (cataract, CBS, glaucoma, and RD). Comprehensiveness mirrored this pattern, with GPT-5 demonstrating significantly fewer omissions across six of nine conditions (dry eye, glaucoma, nystagmus, PVD, CBS, and retinal detachment, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01 for all) and Gemini 3 across two out of nine conditions (AMD and nystagmus, both \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWhen stratified by question category (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB; Table\u0026nbsp;1), both genAI models achieved significantly higher CASEF scores for disease definition and treatment content (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001 for all comparisons). Conversely, NHS leaflets demonstrated higher concordance for treatment complication content (Gemini 3 vs NHS, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01), though the GPT-5 comparison did not reach adjusted significance threshold (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.012). This pattern was mirrored in the comprehensiveness analysis, with both models demonstrating significantly fewer omissions for disease definitions and treatment content (all \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001), but not for complications. No significant differences were observed for the remaining categories.\u003c/p\u003e \u003cp\u003eWhen stratified by condition and question category, safety scores remained non-significant across all subgroup comparisons.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Qualitative safety evaluation\u003c/h2\u003e \u003cp\u003eIndependent qualitative review by two clinicians identified several clinically relevant safety concerns not captured by quantitative scoring. Both models received a comparable number of safety-related comments (35 and 20 for GPT-5; 38 and 24 for Gemini 3).\u003c/p\u003e \u003cp\u003eFirst, both models systematically provided US-centric advice, recommending treatments without UK regulatory approval (\u003cem\u003ee.g.\u003c/em\u003e lifitegrast for dry eye disease), referencing US-based driving standards (\"Department of Motor Vehicles standards\", instead of the British Driver and Vehicle Licensing Agency (DVLA)), and assuming procedural norms inconsistent with NHS practice (\u003cem\u003ee.g.\u003c/em\u003e routine sedation for cataract surgery, stopping anticoagulants pre-operatively for cataract surgery, and multifocal lens availability). Factual inaccuracies, although infrequent, included incorrect statements regarding the prevalence of oscillopsia in congenital nystagmus, the mechanism of neovascular glaucoma, and the pathophysiology of posterior vitreous detachment. In one instance, GPT-5 recommended vision therapy for nystagmus which was flagged as potentially misleading, as no high-quality evidence exists to support its clinical effectiveness (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFurthermore, outputs from both models occasionally demonstrated inappropriate calibration of urgency. Some recommendations were overly optimistic (\u003cem\u003ee.g.\u003c/em\u003e implying full visual recovery after macula-off retinal detachment or omitting surgical failure rates), whilst others used alarmist language (\u003cem\u003ee.g.\u003c/em\u003e describing macula-off retinal detachment repair as requiring \"immediate\" intervention, potentially generating unnecessary patient anxiety). Both models omitted key safety-netting information, particularly post-operative warning signs for endophthalmitis and the need for emergency review should retinal re-detachment symptoms develop.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Readability analysis\u003c/h2\u003e \u003cp\u003eOverall, both genAI models yielded significantly more complex texts compared to the NHS materials, with mean FKGL scores of 12.05 (Gemini 3) and 12.61 (GPT-5), compared to 9.63 for NHS materials (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001 for both comparisons). These scores correspond to 12th- and 9th-grade reading levels (US standards), equivalent to A-level and GCSE standards for genAI and NHS materials, respectively (\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e). Similarly, both ARI and SMOG metrics were substantially higher for Gemini 3 (13.46 and 13.83, respectively) and GPT-5 (14.04 and 14.19, respectively), compared with NHS materials (10.14 and 12.22; \u003cem\u003ep\u0026thinsp;\u0026lt;\u003c/em\u003e\u0026thinsp;0.001 for all comparisons). There were no statistically significant differences between the two genAI models across all readability metrics.\u003c/p\u003e \u003cp\u003eOverall, genAI advice required an additional 2\u0026ndash;3 years of education for comprehension. Importantly, all three sources exceeded the sixth-to-eighth-grade readability recommendation for all patient-facing materials (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Inter-rater reliability\u003c/h2\u003e \u003cp\u003eScoring consistency across all six assessors was evaluated using the intraclass correlation coefficient. For total scores, the combined ICC was 0.843 (95% CI: 0.800\u0026ndash;0.880), indicating very good inter-rater reliability (\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e). For individual information sources, ICC values ranged from good to excellent: 0.828 for GPT-5 (95% CI: 0.75\u0026ndash;0.890), 0.814 for Gemini 3 (95% CI: 0.730\u0026ndash;0.880), and 0.880 for NHS materials (95% CI: 0.780\u0026ndash;0.860). Furthermore, inter-rater reliability between the human evaluator and each LLM was good to excellent with a mean ICC of 0.811 (range: 0.768\u0026ndash;0.855), indicating that genAI assessments aligned with human judgement (Supplementary Table\u0026nbsp;3).\u003c/p\u003e \u003cp\u003eAnalysis of evaluator bias revealed no evidence for preferential scoring of content generated by respective model families by GPT-5 or Gemini 2.5 Pro (mean deviations from consensus: -0.40% for Gemini 2.5 Pro (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.918) and 0.40% for GPT-5 (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.372)). Collectively, these findings indicate consistent agreement among raters, as well as human-genAI agreement, validating the semi-automated scoring methodology.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eIn this benchmarked comparison of two general-purpose genAI models against NHS ophthalmology leaflets, genAI produced health advice that was factually accurate and frequently more comprehensive across nine common eye conditions. However, this informational depth came at the cost of readability, with both models generating text at a higher reading level than the NHS materials. Although overall safety profiles were comparable between sources, qualitative review revealed infrequent, subtle, but clinically relevant concerns such as under-reporting of treatment complications, misalignment with regional guidelines, imprecision, and lack of appropriate safety-netting.\u003c/p\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Main findings and contextualisation\u003c/h2\u003e \u003cp\u003eOur findings align with earlier, condition-specific studies demonstrating the capacity of genAI to generate medically accurate and comprehensive eye health advice (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e). The greater comprehensiveness of AI-generated materials was a key driver of its superior CASEF scores, reflecting a higher density of information relative to NHS leaflets. Although information-dense content can lead to cognitive overload, actually hindering patient comprehension (\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e), genAI platforms offer an advantage over static leaflets in their ability to address follow-up questions in real time, facilitating deeper patient engagement with health information (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eGreater comprehensiveness of genAI outputs was also reflected in substantially higher textual complexity: at the default reading level identified in our study, only 14% of UK adults would be able to fully comprehend these materials (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e). This readability gap is consistent with the broader literature, in which healthcare-related genAI outputs typically require university-level literacy (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan additionalcitationids=\"CR44\" citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e). Nevertheless, successive model releases show progressive improvements in default readability: FKGL scores for GPT-generated cataract materials decreased from 14.0 (GPT-3.5 (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e)) to 12.7 (GPT-4o (\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e)) and 12.6 (GPT-5, this study). Moreover, evidence across multiple specialties including general surgery (\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e), cardiology (\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e), oncology (\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e), and plastic surgery (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e) consistently demonstrates that targeted prompt engineering can substantially improve readability without compromising accuracy. However, employing these strategies requires genAI literacy which may disadvantage those with the greatest need for accessible health information.\u003c/p\u003e \u003cp\u003eDespite excelling at disease definitions and treatment explanations, genAI models underperformed in communicating treatment complications, which carries important safety implications. Incomplete risk disclosure may bias patients towards treatment acceptance, compromising informed consent. We hypothesise this may reflect the relative scarcity of high-quality complications data in training corpora compared to disease definitions or treatments (\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eConsistent with the reported overrepresentation of North American materials in the genAI training corpora (\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e, \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e), both models defaulted to recommendations reflecting US-centric healthcare pathways, medication availability, and clinical guidelines. Although technically accurate, such geographic misalignments constitute contextual safety errors with potential for direct patient harm, through pernicious self-management decisions (\u003cem\u003ee.g.\u003c/em\u003e following advice to discontinue anticoagulation prior to cataract surgery), delayed diagnosis (arising from misunderstood care pathways), and erosion of trust when AI-generated advice contradicts locally received recommendations.\u003c/p\u003e \u003cp\u003eBoth models also failed to provide adequate safety-netting, omitting post-operative red flag symptoms for sight-threatening complications such as endophthalmitis and retinal re-detachment, where delayed presentation directly determines visual prognosis (\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e, \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e). This is particularly consequential given that approximately 70% of health-related LLM conversations occur outside clinic hours (\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e), and when patients experience new or worsening symptoms, with no immediate access to direct clinical guidance (\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e). This absence of safety-netting aligns with a broader trend of progressively declining medical disclaimer rates (\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e), suggesting that successive model iterations have prioritised conversational authority and factual comprehensiveness over safety communication. Despite the safety limitations described above, LLMs have been shown to offer more accurate and reliable guidance than other online health information sources (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e, \u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e). As genAI becomes an established entry point for patient health queries (\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e), regulatory frameworks should mandate minimum safety-netting standards, and patients should be educated on safe prompting strategies and careful output interpretation (\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e, \u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAs contemporary genAI chatbots approach factual parity with clinical materials, safety concerns shift from overt hallucinations or contradictions to contextual, regional, or linguistic errors, which standard accuracy-focused frameworks are not designed to detect (\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e). This underscores the necessity of hybrid, \u0026lsquo;clinician-in-the-loop\u0026rsquo; evaluation frameworks to identify the subtle but clinically consequential limitations that neither quantitative metrics alone nor human evaluators alone can reliably capture (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Limitations\u003c/h2\u003e \u003cp\u003eOur study has several limitations. First, each prompt was executed once per model without regeneration, providing a cross-sectional snapshot of inherently stochastic technology, which might affect reproducibility. Second, prompts based on precise clinical terminology used in patient leaflets may have optimised GenAI output quality, potentially overestimating the accuracy seen with patients\u0026rsquo; more informal queries. Moreover, our methodology employed single-turn prompts which may not reflect typical patient interaction patterns. Third, using control leaflets sourced from a single NHS Trust and all evaluations conducted in English limits generalisability outside a UK context. Finally, our work focused on evaluating textual properties of outputs, but did not measure implications on patient understanding and behaviour. Future studies should assess performance in realistic multi-turn dialogue scenarios, evaluate emerging health-specialised models (\u003cem\u003ee.g\u003c/em\u003e. ChatGPT Health or Claude for Healthcare), employ authentic patient queries (with misspellings, incorrect grammar, colloquial terminology) as prompts, and incorporate patient- and behaviour-centred designs.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eContemporary genAI can provide eye health advice on common ophthalmic conditions that matches or exceeds clinical leaflets in terms of factual accuracy and scope. However, limitations in readability, regional guideline alignment, and risk communication preclude their endorsement as safe, standalone resources for patients. Further development and regulation may help ensure that genAI health advice meets established healthcare standards. As this technology continues to proliferate, clinicians and patients should be encouraged to appreciate the strengths and limitations of genAI health advice, and verify information carefully.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch3\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/h3\u003e\n\u003cp\u003eAYO is supported by a National Institute for Health Research (NIHR) - Moorfields Eye Charity (MEC) Doctoral Fellowship (NIHR303691). PAK is supported by a UK Research \u0026amp; Innovation Future Leaders Fellowship (MR/T019050/1), Moorfields Eye Charity with The Rubin Foundation Charitable Trust (GR001753), and an Alcon Research Institute Senior Investigator Award. AJT is supported by an NIHR Academic Clinical Fellowship (ACF-2025-20-001). The views expressed in this publication are those of the authors and not necessarily those of the abovementioned funding bodies.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePAK is a cofounder and equity owner of Cascader Ltd. He is an equity owner in Big Picture Medical. He has acted as a consultant for Skleo Health, insitro, Retina Consultants of America, Roche, Boehringer-Ingleheim, and Bitfount. He has received speaker fees from Zeiss, Bayer, Boehringer-Ingleheim, Thea, Alimera, Topcon and Roche, and grant funding from Roche. He has received travel support from Bayer and Roche. He has attended advisory boards for Topcon, Bayer, Boehringer-Ingleheim, and Roche. AJT has received grant funding from Théa Pharmaceuticals. None of the other authors report any conflicts of interest.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contribution statement:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAS and AJT conceived and designed the study. AS and BSM curated the benchmark, mapped leaflet subsections to benchmark domains and developed the checklist items. AS generated all AI outputs, and conducted the formal scoring and statistical analyses. AYO and AM performed the blinded qualitative safety evaluation. AS, BSM, MJBR, AYO and AJT drafted the manuscript. AS, BSM, MJBR, AYO, PAK, and AJT reviewed and revised the manuscript. AJT and PAK supervised the study. All authors approved the final version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability statement:\u003c/strong\u003e\u003cbr\u003e\u0026nbsp;All data supporting the conclusions of this study are available within the article and its supplementary materials. The benchmark checklist developed for this study is available upon request. No new clinical or patient data were generated in this study.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTeo ZL, Thirunavukarasu AJ, Elangovan K, Cheng H, Moova P, Soetikno B, et al. Generative artificial intelligence in medicine. Nat Med [Internet]. 2025;31(10):3270\u0026ndash;82. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41591-025-03983-2\u003c/span\u003e\u003cspan address=\"10.1038/s41591-025-03983-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMontero A, Montalvo J III, Kearney A, Valdes I, Kirzinger A, Hamel L. KFF. 2026 [cited 2026 Mar 30]. KFF Tracking Poll on Health Information and Trust: Use of AI For Health Information and Advice. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kff.org/public-opinion/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/\u003c/span\u003e\u003cspan address=\"https://www.kff.org/public-opinion/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinghal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature [Internet]. 2023 Aug 12 [cited 2025 Nov 23];620(7972):172\u0026ndash;80. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41586-023-06291-2\u003c/span\u003e\u003cspan address=\"10.1038/s41586-023-06291-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim SI, Park J, Kim T, Seo W, Kim T, Cha WC, et al. Enhancing patient participation in emergency department through patient-friendly clinical notes generated by large language models. Sci Rep [Internet]. 2025 Dec 5 [cited 2026 Mar 29];16(1):1409. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41598-025-31113-y\u003c/span\u003e\u003cspan address=\"10.1038/s41598-025-31113-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwisher AR, Wu AW, Liu GC, Lee MK, Carle TR, Tang DM. Enhancing health literacy: Evaluating the readability of patient handouts revised by ChatGPT\u0026rsquo;s large language model. Otolaryngol Head Neck Surg [Internet]. 2024;171(6):1751\u0026ndash;7. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1002/ohn.927\u003c/span\u003e\u003cspan address=\"10.1002/ohn.927\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan TF, Thirunavukarasu AJ, Campbell JP, Keane PA, Pasquale LR, Abramoff MD, et al. Generative artificial intelligence through ChatGPT and other large language models in ophthalmology: Clinical applications and challenges. Ophthalmol Sci [Internet]. 2023;3(4):100394. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.xops.2023.100394\u003c/span\u003e\u003cspan address=\"10.1016/j.xops.2023.100394\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMomenaei B, Wakabayashi T, Shahlaee A, Durrani AF, Pandit SA, Wang K, et al. Appropriateness and readability of ChatGPT-4-generated responses for surgical treatment of retinal diseases. Ophthalmol Retina [Internet]. 2023;7(10):862\u0026ndash;8. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.oret.2023.05.022\u003c/span\u003e\u003cspan address=\"10.1016/j.oret.2023.05.022\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCohen SA, Brant A, Fisher AC, Pershing S, Do D, Pan C. Dr. Google vs. Dr. ChatGPT: Exploring the use of artificial intelligence in ophthalmology by comparing the accuracy, safety, and readability of responses to frequently asked patient questions regarding cataracts and cataract surgery. Semin Ophthalmol [Internet]. 2024;39(6):472\u0026ndash;9. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1080/08820538.2024.2326058\u003c/span\u003e\u003cspan address=\"10.1080/08820538.2024.2326058\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBean AM, Payne RE, Parsons G, Kirk HR, Ciro J, Mosquera-G\u0026oacute;mez R, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat Med [Internet]. 2026 Feb 9 [cited 2026 Mar 29];32(2):609\u0026ndash;15. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41591-025-04074-y\u003c/span\u003e\u003cspan address=\"10.1038/s41591-025-04074-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S, Gao M, Sasse K, Hartvigsen T, Anthony B, Fan L, et al. When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior. NPJ Digit Med [Internet]. 2025 Oct 17 [cited 2026 Mar 29];8(1):605. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41746-025-02008-z\u003c/span\u003e\u003cspan address=\"10.1038/s41746-025-02008-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZack T, Lehman E, Suzgun M, Rodriguez JA, Celi LA, Gichoya J, et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit Health [Internet]. 2024 Jan 1 [cited 2026 Mar 29];6(1):e12\u0026ndash;22. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/S2589-7500(23)00225-X\u003c/span\u003e\u003cspan address=\"10.1016/S2589-7500(23)00225-X\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTeo ZL, Quek CWN, Wong JLY, Ting DSW. Cybersecurity in the generative artificial intelligence era. Asia Pac J Ophthalmol (Phila) [Internet]. 2024 Jul 1 [cited 2026 Mar 29];13(4):100091. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.apjo.2024.100091\u003c/span\u003e\u003cspan address=\"10.1016/j.apjo.2024.100091\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMesk\u0026oacute; B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digit Med [Internet]. 2023;6(1):120. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41746-023-00873-0\u003c/span\u003e\u003cspan address=\"10.1038/s41746-023-00873-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMuir KW, Lee PP. Health literacy and ophthalmic patient education. Surv Ophthalmol [Internet]. 2010 Sep 21 [cited 2026 Feb 8];55(5):454\u0026ndash;9. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.survophthal.2010.03.005\u003c/span\u003e\u003cspan address=\"10.1016/j.survophthal.2010.03.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchillinger D, Grumbach K, Piette J, Wang F, Osmond D, Daher C, et al. Association of health literacy with diabetes outcomes. JAMA [Internet]. 2002 Jul 24 [cited 2026 Feb 8];288(4):475\u0026ndash;82. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1001/jama.288.4.475\u003c/span\u003e\u003cspan address=\"10.1001/jama.288.4.475\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJuzych MS, Randhawa S, Shukairy A, Kaushal P, Gupta A, Shalauta N. Functional health literacy in patients with glaucoma in urban settings. Arch Ophthalmol [Internet]. 2008 May [cited 2026 Feb 8];126(5):718\u0026ndash;24. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1001/archopht.126.5.718\u003c/span\u003e\u003cspan address=\"10.1001/archopht.126.5.718\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMuir KW, Santiago-Turla C, Stinnett SS, Herndon LW, Allingham RR, Challa P, et al. Health literacy and adherence to glaucoma therapy. Am J Ophthalmol [Internet]. 2006 Aug [cited 2026 Feb 8];142(2):223\u0026ndash;6. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.ajo.2006.03.018\u003c/span\u003e\u003cspan address=\"10.1016/j.ajo.2006.03.018\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKloosterboer A, Yannuzzi NA, Patel NA, Kuriyan AE, Sridhar J. Assessment of the quality, content, and readability of freely available online information for patients regarding diabetic retinopathy. JAMA Ophthalmol [Internet]. 2019 Nov 1 [cited 2026 Feb 8];137(11):1240\u0026ndash;5. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1001/jamaophthalmol.2019.3116\u003c/span\u003e\u003cspan address=\"10.1001/jamaophthalmol.2019.3116\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKloosterboer A, Yannuzzi N, Topilow N, Patel N, Kuriyan A, Sridhar J. Assessing the quality, content, and readability of freely available online information for patients regarding age-related macular degeneration. Semin Ophthalmol [Internet]. 2021 Aug 18 [cited 2026 Feb 8];36(5\u0026ndash;6):400\u0026ndash;5. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1080/08820538.2021.1893761\u003c/span\u003e\u003cspan address=\"10.1080/08820538.2021.1893761\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYun HS, Bickmore T. Online health information-seeking in the era of large language models: Cross-sectional web-based survey study. J Med Internet Res [Internet]. 2025;27(1):e68560. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.2196/68560\u003c/span\u003e\u003cspan address=\"10.2196/68560\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbbas ASH, Ong AY, Antaki F, Akhtar HN, Shehab M, Keane PA. An updated analysis of large language model performance on ophthalmology speciality examinations. Eye (Lond) [Internet]. 2026;1\u0026ndash;3. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41433-026-04262-1\u003c/span\u003e\u003cspan address=\"10.1038/s41433-026-04262-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThirunavukarasu AJ, Mahmood S, Malem A, Foster WP, Sanghera R, Hassan R, et al. Large language models approach expert-level clinical knowledge and reasoning in ophthalmology: A head-to-head cross-sectional study. PLOS Digit Health [Internet]. 2024;3(4):e0000341. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1371/journal.pdig.0000341\u003c/span\u003e\u003cspan address=\"10.1371/journal.pdig.0000341\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgnihotri AP, Nagel ID, Artiaga JCM, Guevarra MCB, Sosuan GMN, Kalaw FGP. Large language models in ophthalmology: A review of publications from top ophthalmology journals. Ophthalmol Sci [Internet]. 2025 May 1 [cited 2025 Nov 23];5(3):100681. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.xops.2024.100681\u003c/span\u003e\u003cspan address=\"10.1016/j.xops.2024.100681\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheong KX, Zhang C, Tan TE, Fenner BJ, Wong WM, Teo KY, et al. Comparing generative and retrieval-based chatbots in answering patient questions regarding age-related macular degeneration and diabetic retinopathy. Br J Ophthalmol [Internet]. 2024;108(10):1443\u0026ndash;9. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1136/bjo-2023-324533\u003c/span\u003e\u003cspan address=\"10.1136/bjo-2023-324533\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eManchester Royal Eye Hospital [Internet]. 2018 [cited 2025 Nov 9]. Patient Leaflets. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://mft.nhs.uk/royal-eye/patients-visitors/patient-leaflets/?pl_category_id=44\u003c/span\u003e\u003cspan address=\"https://mft.nhs.uk/royal-eye/patients-visitors/patient-leaflets/?pl_category_id=44\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStatCounter Global Stats [Internet]. [cited 2025 Nov 9]. AI Chatbot Market Share Worldwide. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://gs.statcounter.com/ai-chatbot-market-share#monthly-202503-202508\u003c/span\u003e\u003cspan address=\"https://gs.statcounter.com/ai-chatbot-market-share#monthly-202503-202508\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan TF, Elangovan K, Ong J, Shah N, Sung J, Wong TY, et al. A proposed S.c.o.r.e. evaluation framework for large language models: Safety, Consensus, Objectivity, Reproducibility and Explainability [Internet]. arXiv [cs.CL]. 2024. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2407.07666\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2407.07666\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShedlosky-Shoemaker R, Sturm AC, Saleem M, Kelly KM. Tools for assessing readability and quality of health-related Web sites. J Genet Couns [Internet]. 2009;18(1):49\u0026ndash;59. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1007/s10897-008-9181-0\u003c/span\u003e\u003cspan address=\"10.1007/s10897-008-9181-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan TF, Elangovan K, Pollreisz A, Dy KB, Ng WY, Wong JLY, et al. Clinical validation of medical-based large language model chatbots on ophthalmic patient queries with LLM-based evaluation [Internet]. arXiv [cs.AI]. 2026. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.48550/arXiv.2602.05381\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2602.05381\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGu J, Jiang X, Shi Z, Tan H, Zhai X, Xu C, et al. A Survey on LLM-as-a-Judge [Internet]. arXiv [cs.CL]. 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.48550/arXiv.2411.15594\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2411.15594\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCroxford E, Gao Y, First E, Pellegrino N, Schnier M, Caskey J, et al. Automating evaluation of AI text generation in healthcare with a large language model (LLM)-as-a-Judge [Internet]. medRxiv. 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://medrxiv.org/lookup/doi/\u003c/span\u003e\u003cspan address=\"http://medrxiv.org/lookup/doi/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2025.04.22.25326219\u003c/span\u003e\u003cspan address=\"10.1101/2025.04.22.25326219\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCHART Collaborative, Huo B, Collins G, Chartash D, Thirunavukarasu A, Flanagin A, et al. Reporting guideline for chatbot health advice studies: The CHART statement. Artif Intell Med [Internet]. 2025 Oct [cited 2026 Feb 8];168(103222):103222. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.artmed.2025.103222\u003c/span\u003e\u003cspan address=\"10.1016/j.artmed.2025.103222\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWhite J, Fu Q, Hays S, Sandborn M, Olea C, Gilbert H, et al. A prompt pattern catalog to enhance prompt engineering with ChatGPT [Internet]. arXiv [cs.SE]. 2023. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2302.11382\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2302.11382\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu P, Yuan W, Fu J, Jiang Z, Hayashi H, Neubig G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing [Internet]. arXiv [cs.CL]. 2021. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2107.13586\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2107.13586\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med [Internet]. 2016;15(2):155\u0026ndash;63. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.jcm.2016.02.012\u003c/span\u003e\u003cspan address=\"10.1016/j.jcm.2016.02.012\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e[cited 2026 Apr 10]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cosprc.ca/wp-content/uploads/2022/08/Vision-Therapy-Position-Statement_FINAL_ENGLISH.pdf\u003c/span\u003e\u003cspan address=\"https://www.cosprc.ca/wp-content/uploads/2022/08/Vision-Therapy-Position-Statement_FINAL_ENGLISH.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSpadaro DC, Robinson LA, Smith LT. Assessing readability of patient information materials. Am J Health Syst Pharm [Internet]. 1980;37(2):215\u0026ndash;21. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dx.doi.org/10.1093/ajhp/37.2.215\u003c/span\u003e\u003cspan address=\"10.1093/ajhp/37.2.215\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePowell M. Health information: are you getting your message across? [Internet]. National Institute for Health Research; 2022. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.3310/nihrevidence_51109\u003c/span\u003e\u003cspan address=\"10.3310/nihrevidence_51109\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang J, Shi R, Le Q, Shan K, Chen Z, Zhou X, et al. Evaluating the effectiveness of large language models in patient education for conjunctivitis. Br J Ophthalmol [Internet]. 2025;109(2):185\u0026ndash;91. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1136/bjo-2024-325599\u003c/span\u003e\u003cspan address=\"10.1136/bjo-2024-325599\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: a scoping review of applications in medicine. Front Med (Lausanne) [Internet]. 2024;11:1477898. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.3389/fmed.2024.1477898\u003c/span\u003e\u003cspan address=\"10.3389/fmed.2024.1477898\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKlerings I, Weinhandl AS, Thaler KJ. Information overload in healthcare: too much of a good thing? Z Evid Fortbild Qual Gesundhwes [Internet]. 2015;109(4\u0026ndash;5):285\u0026ndash;90. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.zefq.2015.06.005\u003c/span\u003e\u003cspan address=\"10.1016/j.zefq.2015.06.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e[cited 2025 Dec 14]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://assets.publishing.service.gov.uk/media/675330e020bcf083762a6d48/Survey_of_Adult_Skills_2023__PIAAC__National_Report_for_England.pdf\u003c/span\u003e\u003cspan address=\"https://assets.publishing.service.gov.uk/media/675330e020bcf083762a6d48/Survey_of_Adult_Skills_2023__PIAAC__National_Report_for_England.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCampbell DJ, Estephan LE, Mastrolonardo EV, Amin DR, Huntley CT, Boon MS. Evaluating ChatGPT responses on obstructive sleep apnea for patient education. J Clin Sleep Med [Internet]. 2023;19(12):1989\u0026ndash;95. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.5664/jcsm.10728\u003c/span\u003e\u003cspan address=\"10.5664/jcsm.10728\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKianian R, Sun D, Crowell EL, Tsui E. The use of large language models to generate education materials about uveitis. Ophthalmol Retina [Internet]. 2024;8(2):195\u0026ndash;201. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.oret.2023.09.008\u003c/span\u003e\u003cspan address=\"10.1016/j.oret.2023.09.008\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBondok M, Selvakumar R, Law C, Ing EB, Bakshi NK, Felfeli T. Comparing ophthalmologist and Artificial Intelligence Chatbot responses to patient questions. Clin Ophthalmol [Internet]. 2025;19:4293\u0026ndash;300. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.2147/OPTH.S549820\u003c/span\u003e\u003cspan address=\"10.2147/OPTH.S549820\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSrinivasan N, Samaan JS, Rajeev ND, Kanu MU, Yeo YH, Samakar K. Large language models and bariatric surgery patient education: a comparative readability analysis of GPT-3.5, GPT-4, Bard, and online institutional resources. Surg Endosc [Internet]. 2024;38(5):2522\u0026ndash;32. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1007/s00464-024-10720-2\u003c/span\u003e\u003cspan address=\"10.1007/s00464-024-10720-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKing RC, Samaan JS, Haquang J, Bharani V, Margolis S, Srinivasan N, et al. Improving the readability of institutional heart failure-related patient education materials using GPT-4: Observational study. JMIR Cardio [Internet]. 2025;9(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ev9i8e68817\u003c/span\u003e\u003cspan address=\"http://v9i8e68817\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e):e68817. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.2196/68817\u003c/span\u003e\u003cspan address=\"10.2196/68817\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHershenhouse JS, Mokhtar D, Eppler MB, Rodler S, Storino Ramacciotti L, Ganjavi C, et al. Accuracy, readability, and understandability of large language models for prostate cancer information to the public. Prostate Cancer Prostatic Dis [Internet]. 2025;28(2):394\u0026ndash;9. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41391-024-00826-y\u003c/span\u003e\u003cspan address=\"10.1038/s41391-024-00826-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCohen SA, Yadlapalli N, Tijerina JD, Alabiad CR, Chang JR, Kinde B, et al. Comparing the Ability of Google and ChatGPT to Accurately Respond to Oculoplastics-Related Patient Questions and Generate Customized Oculoplastics Patient Education Materials. Clin Ophthalmol [Internet]. 2024;18:2647\u0026ndash;55. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.2147/OPTH.S480222\u003c/span\u003e\u003cspan address=\"10.2147/OPTH.S480222\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchwarz K, Liao Y, Geiger A. On the frequency bias of generative models [Internet]. arXiv [cs.CV]. 2021. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2111.02447\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2111.02447\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhsan H, Sharma AS, Amir S, Bau D, Wallace BC. Elucidating mechanisms of demographic bias in LLMs for healthcare [Internet]. arXiv [cs.CL]. 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2502.13319\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2502.13319\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLynn-Green EE, Ofoje AA, Lynn-Green RH, Jones DS. Variations in how medical researchers report patient demographics: a retrospective analysis of published articles. EClinicalMedicine [Internet]. 2023;58(101903):101903. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.eclinm.2023.101903\u003c/span\u003e\u003cspan address=\"10.1016/j.eclinm.2023.101903\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMirzania D, Fleming TL, Robbins CB, Feng HL, Fekrat S. Time to presentation after symptom onset in endophthalmitis: Clinical features and visual outcomes. Ophthalmol Retina [Internet]. 2021;5(4):324\u0026ndash;9. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.oret.2020.07.027\u003c/span\u003e\u003cspan address=\"10.1016/j.oret.2020.07.027\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHazelwood JE, Mitry D, Singh J, Bennett HGB, Khan AA, Goudie CR. The Scottish Retinal Detachment Study: 10-year outcomes after retinal detachment repair. Eye (Lond) [Internet]. 2025;39(7):1318\u0026ndash;21. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41433-025-03613-8\u003c/span\u003e\u003cspan address=\"10.1038/s41433-025-03613-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenAI. cdn.openai.com. 2026. AI as a Healthcare Ally: How Americans are navigating the system with ChatGPT. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cdn.openai.com/pdf/2cb29276-68cd-4ec6-a5f4-c01c5e7a36e9/OpenAI-AI-as-a-Healthcare-Ally-Jan-2026\u003c/span\u003e\u003cspan address=\"https://cdn.openai.com/pdf/2cb29276-68cd-4ec6-a5f4-c01c5e7a36e9/OpenAI-AI-as-a-Healthcare-Ally-Jan-2026\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.pdf#:~:text=Americans%20are%20using%20AI%20and,to%20navigate%20and%20makes%20decisions\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamaswamy A, Tyagi A, Hugo H, Jiang J, Jayaraman P, Jangda M, et al. ChatGPT Health performance in a structured test of triage recommendations. Nat Med [Internet]. 2026;1\u0026ndash;5. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41591-026-04297-7\u003c/span\u003e\u003cspan address=\"10.1038/s41591-026-04297-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSharma S, Alaa AM, Daneshjou R. A longitudinal analysis of declining medical safety messaging in generative AI models. NPJ Digit Med [Internet]. 2025;8(1):592. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41746-025-01943-1\u003c/span\u003e\u003cspan address=\"10.1038/s41746-025-01943-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCohen SA, Fisher AC, Xu BY, Song BJ. Comparing the accuracy and readability of glaucoma-related Question Responses and Educational Materials by Google and ChatGPT. J Curr Glaucoma Pract [Internet]. 2024;18(3):110\u0026ndash;6. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.5005/jp-journals-10078-1448\u003c/span\u003e\u003cspan address=\"10.5005/jp-journals-10078-1448\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTrillo-Dom\u0026iacute;nguez M, Martin-Neira JI, Olvera-Lobo MD. Dr. Google vs. Dr. ChatGPT in online health self-consultation: A scoping review of accuracy, bias, and actionability (2023\u0026ndash;2025). Informatics (MDPI) [Internet]. 2026;13(3):41. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.3390/informatics13030041\u003c/span\u003e\u003cspan address=\"10.3390/informatics13030041\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMenz BD, Kuderer NM, Bacchi S, Modi ND, Chin-Yee B, Hu T, et al. Current safeguards, risk mitigation, and transparency measures of large language models against the generation of health disinformation: repeated cross sectional analysis. BMJ [Internet]. 2024;384:e078538. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1136/bmj-2023-078538\u003c/span\u003e\u003cspan address=\"10.1136/bmj-2023-078538\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhair DO, Kale AU, Agbakoba R, Goh E, Mateen BA, Ong AY, et al. Building the health chatbot users\u0026rsquo; guide. Nat Health [Internet]. 2026;1\u0026ndash;2. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s44360-026-00074-5\u003c/span\u003e\u003cspan address=\"10.1038/s44360-026-00074-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAsgari E, Monta\u0026ntilde;a-Brown N, Dubois M, Khalil S, Balloch J, Yeung JA, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digit Med [Internet]. 2025;8(1):274. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1038/s41746-025-01670-7\u003c/span\u003e\u003cspan address=\"10.1038/s41746-025-01670-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1 is available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9383173/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9383173/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eGenerative artificial intelligence (genAI) chatbots are increasingly used for health advice despite lacking regulatory approval, raising concerns about their output quality and safety. This study assesses eye health advice from leading genAI platforms, benchmarking their quality against patient information leaflets.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe compared outputs from GPT-5 (OpenAI) and Gemini 3 (Google DeepMind) with clinical leaflets across nine eye conditions (41 questions, 123 texts total). Reference benchmark (547 items) was derived from patient materials produced by the Royal College of Ophthalmologists. Chatbot outputs were generated using verbatim leaflet subsection headings as prompts with word-count restrictions to match corresponding leaflet sections. All texts were evaluated using the Comprehensiveness, Accuracy, and Safety Evaluation Framework (CASEF). Two blinded ophthalmologists assessed genAI outputs for safety concerns. Readability was measured using Flesch-Kincaid Grade Level.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eBoth genAI models showed higher factual alignment than clinical leaflets (GPT-5\u0026thinsp;=\u0026thinsp;37.4%, Gemini 3\u0026thinsp;=\u0026thinsp;36.2%, leaflets\u0026thinsp;=\u0026thinsp;30.7%; both p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and fewer omissions (GPT-5\u0026thinsp;=\u0026thinsp;6.7, Gemini 3\u0026thinsp;=\u0026thinsp;7.0, leaflets\u0026thinsp;=\u0026thinsp;7.9; both p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Safety scores were comparable across sources, but both models underreported treatment complications, exhibited guideline inconsistencies, and failed to include appropriate safety-netting. Moreover, genAI outputs required 2\u0026ndash;3 more years of education to understand. High inter-rater reliability (ICC\u0026thinsp;=\u0026thinsp;0.843 (95%CI:0.800\u0026ndash;0.880)) validated the scoring methodology.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eGenAI eye health advice matches clinical leaflets in accuracy and comprehensiveness. However, subtle yet clinically consequential errors remain, limiting the application of general-purpose genAI chatbots as a safe, standalone information source for ophthalmology patients.\u003c/p\u003e","manuscriptTitle":"Benchmarking eye health advice from generative artificial intelligence in terms of factual accuracy, safety, comprehensiveness and readability","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-21 01:33:25","doi":"10.21203/rs.3.rs-9383173/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"fb0c496c-0965-4078-8390-edb06fc6b49d","owner":[],"postedDate":"April 21st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":66437581,"name":"Health sciences/Medical research"},{"id":66437582,"name":"Scientific community and society/Scientific community/Education"}],"tags":[],"updatedAt":"2026-04-21T01:33:26+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-21 01:33:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9383173","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9383173","identity":"rs-9383173","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.