Can Large Language Models Aid Pre-surgical Epileptogenic Zone Localization? A Multi-source Text Analysis Performance Study in Drug-Resistant Epilepsy

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Large language models (LLMs) show promise for biomedical text analysis but are underused in pre-surgical epileptogenic zone (EZ) localization. This study evaluated LLMs' ability to analyze multi-source clinical text (medical records, EEG, MRI, FDG-PET, MEG) in 157 drug-resistant epilepsy patients (2020–2024) at Beijing Tiantan Hospital. Three LLMs (GPT-4.1, Deepseek-R1, Claude 3.7 Sonnet) performed (1) EZ laterality classification, (2) probabilistic lobar localization (top 3 most likely lobes), and (3) SEEG recommendation scoring (0-100), benchmarked against multidisciplinary team (MDT)-defined resection sites. Modality ablation and stability analyses were conducted. Results showed GPT-4.1 and Deepseek-R1 achieved 98.1% laterality accuracy, vs. 97.5% for Claude 3.7 Sonnet (p > 0.05). GPT-4.1 and Claude 3.7 Sonnet had median lobar localization scores of 70 vs. 60 for Deepseek-R1 (p < 0.001). Textual information source ablation study revealed MRI reports and medical records were critical for GPT-4.1’s localization. GPT-4.1’s stability analysis using a two-way mixed-effects model showed excellent consistency for localization scores: single-measure ICC = 0.960 (95% CI: 0.948–0.969, p < 0.001) and average-measure ICC = 0.986 (95% CI: 0.982–0.990, p < 0.001). Additionally, we used GPT-4.1 for the SEEG recommendation task: its SEEG scores were significantly higher in patients undergoing SEEG (90.00 [IQR: 85.00–90.00] vs. 25.00 [IQR: 10.00–90.00] for non-SEEG, p < 0.001). LLMs accurately inferred EZ laterality/lobar localization, aligning with MDT consensus, and their SEEG stratification potential may reduce invasive monitoring. Future research should focus on fine-tuning and multimodal fusion to optimize drug-resistant epilepsy outcomes.
Full text 124,385 characters · extracted from preprint-html · click to expand
Can Large Language Models Aid Pre-surgical Epileptogenic Zone Localization? A Multi-source Text Analysis Performance Study in Drug-Resistant Epilepsy | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Can Large Language Models Aid Pre-surgical Epileptogenic Zone Localization? A Multi-source Text Analysis Performance Study in Drug-Resistant Epilepsy Yueqian Sun, Shihao Ge, Yangyang Wang, Xiaoqiu Shao, Kai Zhang, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6853779/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Large language models (LLMs) show promise for biomedical text analysis but are underused in pre-surgical epileptogenic zone (EZ) localization. This study evaluated LLMs' ability to analyze multi-source clinical text (medical records, EEG, MRI, FDG-PET, MEG) in 157 drug-resistant epilepsy patients (2020–2024) at Beijing Tiantan Hospital. Three LLMs (GPT-4.1, Deepseek-R1, Claude 3.7 Sonnet) performed (1) EZ laterality classification, (2) probabilistic lobar localization (top 3 most likely lobes), and (3) SEEG recommendation scoring (0-100), benchmarked against multidisciplinary team (MDT)-defined resection sites. Modality ablation and stability analyses were conducted. Results showed GPT-4.1 and Deepseek-R1 achieved 98.1% laterality accuracy, vs. 97.5% for Claude 3.7 Sonnet (p > 0.05). GPT-4.1 and Claude 3.7 Sonnet had median lobar localization scores of 70 vs. 60 for Deepseek-R1 (p < 0.001). Textual information source ablation study revealed MRI reports and medical records were critical for GPT-4.1’s localization. GPT-4.1’s stability analysis using a two-way mixed-effects model showed excellent consistency for localization scores: single-measure ICC = 0.960 (95% CI: 0.948–0.969, p < 0.001) and average-measure ICC = 0.986 (95% CI: 0.982–0.990, p < 0.001). Additionally, we used GPT-4.1 for the SEEG recommendation task: its SEEG scores were significantly higher in patients undergoing SEEG (90.00 [IQR: 85.00–90.00] vs. 25.00 [IQR: 10.00–90.00] for non-SEEG, p < 0.001). LLMs accurately inferred EZ laterality/lobar localization, aligning with MDT consensus, and their SEEG stratification potential may reduce invasive monitoring. Future research should focus on fine-tuning and multimodal fusion to optimize drug-resistant epilepsy outcomes. Drug-resistant epilepsy Epileptogenic zone localization Large language models Multi-source clinical text analysis SEEG recommendation Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Epilepsy affects over 70 million individuals globally, with approximately 30% progressing to drug-resistant focal epilepsy (DRE)¹. Surgical resection of the epileptogenic zone (EZ) remains the most effective treatment for seizure control in DRE patients. The primary goal of presurgical evaluation—a cornerstone of preoperative assessment for epilepsy surgery—is to precisely localize the EZ while minimizing adverse effects, ideally through noninvasive methods²⁻⁴. Due to the inherent complexity of epilepsy, multidisciplinary preoperative evaluations and meticulous surgical decision-making are essential⁵. Phase 1 noninvasive assessment typically includes prolonged scalp video-electroencephalography (EEG) monitoring⁶, structural and functional neuroimaging⁷ , ⁸, and neuropsychological testing, with protocols tailored to individual patient needs. When these noninvasive investigations yield concordant results, they often provide a robust foundation for localizing the seizure onset zone (SOZ) and formulating a personalized resection plan⁹. However, in 25–50% of cases, noninvasive evaluations fail to delineate the SOZ¹⁰⁻¹². In such scenarios, invasive Phase 2 assessments—such as intracranial EEG recording—are required to precisely identify the SOZ and its spatial relationship to eloquent cortex⁴ , ⁵ , ¹³. This invasive approach enables definitive evaluation of resectability while mitigating risks to critical neurological functions. Common modalities include electrocorticography, subdural electrodes (strips/grids), and stereo-EEG (SEEG), with SEEG now serving as the primary intracranial EEG modality in most North American and Chinese epilepsy centers¹⁴ , ¹⁵. Despite the relatively high seizure-free rates achieved post-surgery, approximately one-third of patients experience persistent or recurrent seizures, primarily due to incomplete EZ resection¹⁶. This can stem from insufficient presurgical identification of the target resection volume or anatomical constraints near eloquent cortex. Additionally, challenges in epilepsy surgery evaluations include subjective expert judgments, inconsistent ancillary test results, difficulties in multimodal data integration, and time-consuming, empirical manual interpretations. Furthermore, SEEG carries risks such as limited brain coverage, intracranial infections from prolonged monitoring for spontaneous seizures, and the potential failure to capture habitual seizures³. Reducing unnecessary SEEG procedures through improved presurgical assessments not only alleviates patient burdens—such as physical discomfort, psychological stress, and economic costs—but also minimizes the risks associated with invasive interventions. Thus, there is an urgent need for innovative approaches to enhance the accuracy of presurgical evaluations and improve treatment outcomes for epilepsy patients. Large language models (LLMs), which have emerged as transformative tools in medicine, demonstrate groundbreaking capabilities in biomedical text analysis. Their advanced reasoning and universal adaptability far exceed traditional machine learning methods, offering promise for enhancing clinical decision-making, automating administrative tasks, and improving patient care¹⁷⁻¹⁹. In the field of epilepsy, LLMs have shown notable performance, including passing professional examinations²⁰ and aiding in differential diagnosis²¹. A recent study further developed an LLM-based method to automatically extract seizure frequencies from epilepsy monitoring unit (EMU) reports for Sudden Unexpected Death in Epilepsy (SUDEP) risk assessment²². However, systematic research focusing on the application of LLMs to localize the EZ through the comprehensive analysis of multi-source text data—a task that involves inferring spatial features from diverse and abstract textual sources—is still in its early stages. As a major innovation, this study constructs an LLM-driven analysis framework that integrates medical records, EEG, magnetic resonance imaging (MRI), positron emission tomography (PET), and magnetoencephalography (MEG) preoperative reports from 157 DRE patients. We compare the performance of three leading LLMs (GPT-4.1, Deepseek-R1, and Claude 3.7 Sonnet) without fine-tuning. The study aims to (1) determine seizure laterality, (2) localize the EZ at the lobar level, and (3) recommendation for Phase 2 preoperative evaluation (SEEG). By exploring the feasibility of text-driven EZ localization using LLMs, this study assesses their potential as an auxiliary tool in noninvasive preoperative evaluation. The aim is to provide initial insights that could support the complex localization process and contribute to informed clinical decision-making, particularly concerning the use of SEEG. Methods Patient Collection This retrospective study included patients with DRE who underwent surgical intervention (resection or thermal ablation) at the Beijing Tiantan Hospital Epilepsy Center between January 2020 and April 2024. Inclusion criteria were: (1) Confirmed DRE diagnosis. (2) Multidisciplinary team (MDT)-determined EZ localization and surgical strategy. (3) Complete baseline clinical data (clinical history, seizure semiology, in-hospital notes) and full preoperative evaluations (prolonged scalp EEG, structural MRI, ¹⁸F-FDG PET). (4) ≥12-month post-surgical follow-up. Exclusion criteria were: (1) Non-structural etiologies (e.g., genetic/metabolic disorders, encephalitis). (2) Missing data (incomplete EEG/MRI/PET records). (3) Involvement of three or more lobes, as determined by postoperative CT or MRI scans. The study was approved by the institutional ethics committee, with written informed consent obtained from all participants. The detailed patient selection process is illustrated in Figure S1. Standard Presurgical Evaluation and Gold Standard Definition An overview of the study workflow for EZ assessment using LLMs is provided in Figure 1. Standard presurgical evaluation included: (1) Seizure history and semiology. (2) Long-term scalp video-EEG monitoring. (3) 3-T thin-slice MRI. (4) Interictal FDG-PET/MRI co-registration (visually analyzed via standardized color scales). (5) Neuropsychological testing. (6) MEG was selectively performed. MRI-negative status was defined as the absence of clinically relevant structural abnormalities, as determined by consensus among epileptologists. All data were reviewed at MDT conferences to determine surgical eligibility and intervention plans, with a focus on identifying candidates for SEEG—primarily those with MRI-negative or subtle lesions. The gold standard for assessing the LLM's EZ localization accuracy was the set of surgically resected lobe(s), determined by MDT consensus as the target for surgical intervention. EZs were categorized as unilobar (single lobe) or multilobar (involving two distinct lobes). Data Preparation for Large Language Models For each patient, textual data from the following five sources were collected: a. Medical Records: This included admission history (detailing seizure history, comorbidities, etc.) and in-hospital progress notes (documenting dynamic seizure manifestations such as auras, automatisms, and other relevant clinical observations during hospitalization). b. MRI Reports: Textual descriptions and diagnostic conclusions provided by neuroradiologists. c. Scalp EEG Reports: Textual descriptions and interpretations of prolonged scalp video-EEG findings provided by neurophysiologists. d. PET Reports: Textual descriptions and diagnostic conclusions from nuclear medicine physicians. e. MEG Reports: Textual descriptions and analytical conclusions from MEG data analysts (available for a subset of patients). Data preprocessing included anonymization (removal of identifiers) and text data cleaning (formatting/special characters). For each patient, texts from these five sources were concatenated into one single text input to the LLM. The input size is constrained by the LLM’s prompt limit. Large Language Model Configuration, Tasks, and Analytical Framework Three LLMs—GPT-4.1 (OpenAI), Deepseek-R1 (Deepseek AI), and Claude 3.7 Sonnet (Anthropic)—were evaluated via official APIs with temperature = 0 (default parameters otherwise). A standardized prompt (Figure 2; Supplementary Figure S2) directed models to perform three tasks: (1) EZ Laterality Classification: determine the "Left" or "Right" hemisphere; (2) Lobar-Level Localization: select the most likely top 3 lobes based on the predicted probabilities (with a total sum of 100%). (3) SEEG Recommendation Scoring: assign a score from 0 to 100 to indicate SEEG necessity. Performance metrics included: (1) Laterality Accuracy: proportion of the hemisphere predictions matching MDT consensus (coded as 1 for correct, 0 for incorrect). (2) Localization Score: predicted probability score of the gold standard lobe (if ranked within top 3) or 0 for single-lobe epilepsy cases. For combined-lobe cases (e.g., fronto-temporal), the score was the sum of the probability scores assigned by the LLM to those MDT-defined epileptogenic lobes that were present in the LLM's top three predictions (range: 0-100%). (3) SEEG Recommendation Analysis: comparison of LLM-predicted scores between patients who underwent SEEG and those who did not. The analytical framework consisted of five stages: (1) Inter-Model Comparison: Evaluated laterality accuracy and localization scores across all LLMs using all multi-source data for the entire cohort and MRI-negative subgroup. (2) GPT-4.1 Subgroup Analysis: Assessed performance stratified by epilepsy type (frontal/parietal/temporal/occipital/insula/combined-lobe) and MRI status (negative/positive). (3) Textual Information Source Ablation Study: We compared GPT-4.1’s performance when the textual information collected from each of the five sources (i.e., clinical notes, MRI reports, EEG reports, PET reports, and MEG reports) was individually removed. The baseline for this comparison was the performance achieved using textual input from all five sources. The ablation analysis for the textual data from MEG reports was restricted to the 64 patients who had these reports. (4) Stability Analysis: Conducted three runs of GPT-4.1 to calculate the Intraclass Correlation Coefficient (ICC) for localization score consistency. (5) SEEG Necessity Stratification: Compared GPT-4.1’s SEEG recommendation scores between patients who underwent the procedure and those who did not. Statistical Analysis All statistical analyses were performed using SPSS Statistics version 27.0 (IBM Corp., Armonk, NY, USA). Descriptive statistics included frequencies and percentages [n (%)] for categorical variables. Continuous variables were assessed for normality using the Shapiro-Wilk test; normally distributed data were presented as mean ± standard deviation (Mean ± SD), while non-normally distributed data (e.g., localization scores) were presented as median and interquartile range [Median (IQR)]. For inferential statistics: (1) McNemar test was used for pairwise comparisons of laterality accuracy (e.g., inter-model, ablation effects). (2) Fisher’s exact test was used for independent subgroup comparisons of laterality accuracy (e.g., MRI-negative vs. -positive). (3) For non-normally distributed continuous variables (such as localization scores, SEEG recommendation scores, and age at assessment), the Wilcoxon signed-rank test was used for pairwise comparisons, the Mann-Whitney U test for independent two-group comparisons, and the Kruskal-Wallis H test for independent multi-group comparisons. (4) ICC with its 95% confidence interval (CI) was used to assess stability. All statistical tests were two-sided, with a p-value < 0.05 considered statistically significant. Results Patients Characteristics A total of 157 patients with DRE were enrolled in this study. The median age at assessment was 25.0 years (IQR: 19.0-32.0), with 97 (61.8%) male participants. Patients were stratified into epilepsy subtypes based on postoperative CT/MRI findings, as detailed in Table 1. The median duration from initial seizure onset to surgical intervention was 10.0 years (IQR: 4.0-18.0). Significant differences were observed among subtypes in age at assessment (p = 0.004), current antiseizure medication (ASM) count (p = 0.025), and frequency of focal impaired awareness seizures (FIAS) (p = 0.032). Pathological findings suggested a trend toward differences among subtypes, but this did not reach statistical significance (p = 0.088). Epileptogenic Zone Localization Using the surgical resection site determined by the MDT as the gold standard, GPT-4.1 and Deepseek-R1 demonstrated identical laterality accuracy (98.1%) for EZ localization, while Claude 3.7 Sonnet achieved 97.5% (p = 1.00; Figure 3a, Table 2, Table 3). For lobar-level localization, GPT-4.1 and Claude 3.7 Sonnet attained median scores of 70 versus 60 for Deepseek-R1 (p < 0.001; Figure 3b, Table 2). Notably, even in diagnostically challenging MRI-negative cases (n = 71), all models achieved high lateralization accuracy. Furthermore, both GPT-4.1 and Claude 3.7 Sonnet performed well in lobar localization within this subgroup (Supplementary Figure S3, Table S2). This ability to provide valuable localizing information for such patients underscores the potential utility of LLMs. Further analysis of GPT-4.1 across epilepsy subtypes revealed distinct patterns (Table 4). The model achieved perfect lateralization accuracy (100%) for frontal, parietal, occipital, and insular lobe epilepsy, with 98.6% accuracy in temporal lobe epilepsy (TLE) and 88.9% in multilobar cases. Localization scores varied significantly (Kruskal-Wallis test, p < 0.001): multilobar epilepsy yielded the highest median score (80.0 [IQR: 50.0-90.0]), followed by TLE (70.0 [IQR: 70.0-75.5]) and frontal lobe epilepsy (FLE) (65.0 [IQR: 60.0-70.0]). Parietal (60.0 [IQR: 35.0-60.0]) and occipital (55.0 [IQR: 12.5-60.0]) lobes showed moderate performance, while insular lobe epilepsy (ILE) had the lowest median score (15.0 [IQR: 2.5-45.0]), reflecting inherent challenges in localizing this subtype. Figure 4 visualizes these distributional differences across epilepsy types. GPT-4.1 demonstrated no significant differences in lateralization accuracy or localization scores between MRI-negative and MRI-positive patients (Supplementary Table S1). A modality ablation study (Supplementary Table S3) revealed that removal of medical records or MRI reports significantly impaired GPT-4.1’s EZ localization performance, whereas exclusion of EEG, PET, or MEG reports had minimal impact. To assess GPT-4.1’s stability, the lobar localization task with all five modalities was repeated three times. Using a two-way mixed-effects model, the ICC for localization scores demonstrated excellent consistency: single measures ICC = 0.960 (95% confidence interval [CI]: 0.948-0.969, p < 0.001) and average measures ICC = 0.986 (95% CI: 0.982-0.990, p < 0.001). SEEG Recommendation This study evaluated the capacity of LLMs to recommend SEEG monitoring, focusing on GPT-4.1. SEEG recommendation scores differed significantly between patients who underwent SEEG (n = 65) and those who did not (n = 92; p < 0.001; Figure 5). GPT-4.1 assigned higher median scores to SEEG-treated patients (90.00 [IQR: 85.00-90.00]) compared to non-SEEG patients (25.00 [IQR: 10.00-90.00]). These findings suggest that GPT-4.1 can stratify SEEG necessity with high discriminative power, potentially guiding clinical decisions to prioritize invasive monitoring in complex cases. Discussion This study systematically evaluates the utility of LLMs in presurgical epilepsy assessment, focusing on three major tasks: EZ laterality determination, lobar-level localization, and SEEG recommendation stratification. By leveraging multi-source clinical textual data—including medical records, EEG, MRI, PET, and MEG reports—without requiring direct access to electrophysiological or imaging datasets, our framework demonstrates the potential to enhance diagnostic accessibility across diverse clinical settings and streamline presurgical workflows. Notably, GPT-4.1 achieved decision-making accuracy comparable to MDT consensus, while Deepseek-R1 and Claude 3.7 Sonnet exhibited strong but suboptimal performance. These findings position advanced LLMs as scalable decision-support tools for epilepsy surgery, particularly in regions with limited MDT resources. Consistency among diagnostic modalities is a well-established predictor of favorable postoperative outcomes in epilepsy surgery 23 . Non-eloquent, MRI-visible epileptogenic lesions paired with congruent seizure semiology and EEG findings typically yield optimal surgical results 5 . However, conflicting evaluations, inconclusive data, or the presence of multiple MRI lesions often complicate EZ identification. Currently, no definitive preoperative gold standard exists for EZ localization beyond standardized initial assessments (e.g., semiology analysis, video-EEG, MRI, neuropsychological testing) 24 . Our study demonstrates that LLMs can achieve high accuracy in EZ lateralization and lobar localization, closely aligning with the EZ defined by MDT consensus, offering a novel approach to address diagnostic ambiguity. Unlike traditional workflows reliant on time-consuming, expert-driven evaluations, LLMs rapidly integrate heterogeneous data sources, providing consistent, evidence-based recommendations. This efficiency and reliability underscore their potential to revolutionize epilepsy diagnosis, particularly in settings with limited access to MDT expertise. Prior research has highlighted the educational and basic diagnostic capabilities of LLMs in epilepsy care. For instance, Kim et al. demonstrated high educational fidelity in addressing epilepsy FAQs but noted limitations in mechanistic insights into EZ localization 25 . Wu et al. further identified gaps in prognostic inference, with only 46.8% of outcome-related queries receiving comprehensive answers 26 . Recent studies have begun exploring LLMs for EZ localization: Zhang et al. (2025) showed that ChatGPT could infer EZ locations from seizure semiology descriptions alone, achieving 80–90% regional sensitivity for frontal/temporal lobes and ≥ 67% weighted sensitivity across imbalanced datasets 27 . However, this work focused exclusively on unimodal text interpretation (semiology reports) and excluded patients with multilobar involvement. In contrast, our study establishes a multimodal LLM framework that integrates structural MRI, EEG, PET, MEG, and clinical narratives to mirror real-world presurgical workflows. Compared to Zhang et al.’s semiology-centric approach, our model matches their laterality accuracy (98.1% vs. 80–90% regional sensitivity) while introducing probability-weighted lobar localization, a critical advancement for complex cases like combined-lobe epilepsy, where quantifying seizure burden guides surgical planning. This evolution extends beyond semiology-based localization to hierarchical clinical reasoning, aligning more closely with MDT decision-making processes. SEEG is a critical invasive tool for mapping the SOZ and epileptogenic network via multi-contact depth electrodes 28 . Its hypothesis-driven nature, guided by seizure semiology and anatomical-electro-clinical correlations, necessitates careful electrode placement to balance diagnostic yield and procedural risks 29 . While LLMs do not fully replicate MDT decisions in second-stage SEEG assessments, our findings reveal a notable alignment in certain scenarios. Specifically, GPT-4.1 assigned markedly higher SEEG recommendation scores to patients who ultimately underwent the procedure compared to those who did not. However, it is crucial to interpret these findings with caution. This statistical correlation, while significant, primarily suggests the LLM's potential to discern patterns indicative of SEEG necessity from the provided patient cohort's data, rather than a comprehensive replication of the complex, individualized MDT decision-making process. While prior benchmarks were often limited to categorical EZ localization, this exploratory study indicates LLMs' potential to contribute to hierarchical presurgical reasoning, perhaps by offering a preliminary quantification of SEEG consideration. Nonetheless, substantial further research is required to validate these initial observations and to understand how such tools might responsibly augment, rather than replace, expert clinical judgment in real-world workflows, ensuring that any potential benefits in stratifying SEEG necessity do not overlook critical nuances or introduce unforeseen biases. The prospect of reducing unnecessary invasive procedures remains an important goal, but one that necessitates rigorous validation of these emerging technologies. The promising yet variable performance of LLMs in EZ localization and SEEG recommendation is intrinsically tied to the current capabilities and limitations of language model technology. A key strength lies in their extensive context windows (e.g., GPT-4.1, Claude 3.7, and Deepseek-R1 supporting 128K-200K tokens), enabling comprehensive integration of complex patient histories, including detailed seizure evolution narratives and multifaceted examination reports. This capacity may allow LLMs to detect subtle patterns overlooked by clinicians due to time constraints or information overload 30 . Furthermore, prompting models to output their “reasoning process” alongside predictions provides partial transparency, mitigating the “black box” perception of deep learning models and offering insights into how textual evidence is weighted. Limitations This study has several limitations. Firstly, inherent constraints of current LLM technology impact our approach. These models rely exclusively on textual data, unable to interpret raw imaging or electrophysiological signals 30 , rendering them susceptible to the quality of input reports. While future multimodal LLMs, Retrieval-Augmented Generation (RAG), or domain-specific fine-tuning hold promise 31 , 32 , these current technological gaps are relevant. Secondly, our study design introduces further limitations. The single-center Chinese cohort may restrict generalizability. The small number of participating epileptologists could affect the robustness of our gold standard. Finally, our focus on MDT results without incorporating SEEG data — crucial for refining EZ localization in complex cases — might underestimate LLM accuracy, particularly its potential alignment with definitive postoperative outcomes when invasive imaging is considered. Conclusion These findings highlight the potential of LLMs in epilepsy care: In specialized centers, they serve as intelligent co-pilots, augmenting epileptologists’ diagnostic efficiency. In resource-limited settings, they address critical knowledge gaps, enabling clinicians to perform initial seizure classification and make informed clinical decisions, thereby improving access to care. Crucially, unlike clinicians who rely on individual experience, LLMs trained on extensive datasets can generate data-driven predictions of the EZ prior to surgery, offering scalable opportunities to optimize surgical outcomes beyond the scope of traditional workflows. Declarations Ethics approval and consent to participate The present study was approved by the Institutional Review Board of the Beijing Tiantan Hospital affiliated with Capital Medical University (Beijing, China) (KY2023-079-01). Availability of data and material The datasets generated and/or analyzed during the current study are available from the corresponding author on reasonable request. Declaration of interest statement The authors report no competing interests. Consent for publication All participants or their caregivers provided written informed consent for the publication. Funding The study was financially supported by the National Key R&D Program of China grant (2022YFC2503800), the Capital Health Research and Development of Special grants (2024-1-2041), National Natural Science Foundation of China (U24A20695 and 82371449) and China Postdoctoral Science Foundation (2024M762180). Authors' contributions QW, YQS and SHG concepted, designed, and supervised the study. SHG acquired the data. YQS, SHG, YYW, XQS and KZ analyzed and interpreted the data, provided statistical analysis, had full access to all of the data in the study, and are responsible for the integrity of the data and the accuracy of the data analysis. YQS and SHG drafted the manuscript, QW and ZHH critically revised the manuscript for important intellectual content. All authors read and approved the final manuscript. Acknowledgements We would like to express gratitude to all the participants for their cooperation in our study. We would also like to thank the researchers for their contributions in collecting and following up with the patients. References Thijs R.D., Surges R., O'Brien T.J., et al. Epilepsy in adults. Lancet. 2019;393(10172):689-701. Rosenow F., Lüders H. Presurgical evaluation of epilepsy. Brain. 2001;124(Pt 9):1683-700. Cossu M., Cardinale F., Castana L., et al. Stereoelectroencephalography in the presurgical evaluation of focal epilepsy: a retrospective analysis of 215 procedures. Neurosurgery. 2005;57(4):706-18; discussion 706-18. Mullin J.P., Shriver M., Alomar S., et al. Is SEEG safe? A systematic review and meta-analysis of stereo-electroencephalography-related complications. Epilepsia. 2016;57(3):386-401. Vakharia V.N., Duncan J.S., Witt J.A., et al. Getting the best outcomes from epilepsy surgery. Ann Neurol. 2018;83(4):676-690. Engel J., Jr., Driver M.V., Falconer M.A. Electrophysiological correlates of pathology and surgical results in temporal lobe epilepsy. Brain. 1975;98(1):129-56. Jack C.R., Jr., Sharbrough F.W., Marsh W.R. Use of MR imaging for quantitative evaluation of resection for temporal lobe epilepsy. Radiology. 1988;169(2):463-8. Mazziotta J.C., Engel J., Jr. The use and impact of positron computed tomography scanning in epilepsy. Epilepsia. 1984;25 Suppl 2S86-104. Lado F.A., Ahrens S.M., Riker E., et al. Guidelines for Specialized Epilepsy Centers: Executive Summary of the Report of the National Association of Epilepsy Centers Guideline Panel. Neurology. 2024;102(4):e208087. Thadani V.M., Williamson P.D., Berger R., et al. Successful epilepsy surgery without intracranial EEG recording: criteria for patient selection. Epilepsia. 1995;36(1):7-15. Kilpatrick C., Cook M., Kaye A., et al. Non-invasive investigations successfully select patients for temporal lobe surgery. J Neurol Neurosurg Psychiatry. 1997;63(3):327-33. Diehl B., Lüders H.O. Temporal lobe epilepsy: when are invasive recordings needed? Epilepsia. 2000;41 Suppl 3S61-74. Rodionov R., O'Keeffe A., Nowell M., et al. Increasing the accuracy of 3D EEG implantations. J Neurosurg. 2020;133(1):35-42. Gavvala J., Zafar M., Sinha S.R., et al. Stereotactic EEG Practices: A Survey of United States Tertiary Referral Epilepsy Centers. J Clin Neurophysiol. 2022;39(6):474-480. Lin Y., Hu S., Hao X., et al. Epilepsy centers in China: Current status and ways forward. Epilepsia. 2021;62(11):2640-2650. Spencer S., Huh L. Outcomes of epilepsy surgery in adults and children. Lancet Neurol. 2008;7(6):525-37. Quer G., Topol E.J. The potential for large language models to transform cardiovascular medicine. Lancet Digit Health. 2024;6(10):e767-e771. Bellini V., Bignami E.G. Generative Pre-trained Transformer 4 (GPT-4) in clinical settings. Lancet Digit Health. 2025;7(1):e6-e7. Boussina A., Krishnamoorthy R., Quintero K., et al. Large Language Models for More Efficient Reporting of Hospital Quality Measures. Nejm ai. 2024;1(11): Habib S., Butt H., Goldenholz S.R., et al. Large Language Model Performance on Practice Epilepsy Board Examinations. JAMA Neurol. 2024;81(6):660-661. Ford J., Pevy N., Grunewald R., et al. Can artificial intelligence diagnose seizures based on patients' descriptions? A study of GPT-4. Epilepsia. 2025; Abeysinghe R., Tao S., Lhatoo S.D., et al. Leveraging pretrained language models for seizure frequency extraction from epilepsy evaluation reports. NPJ Digit Med. 2025;8(1):208. Kankirawatana P., Mohamed I.S., Lauer J., et al. Relative contribution of individual versus combined functional imaging studies in predicting seizure freedom in pediatric epilepsy surgery: an area under the curve analysis. Neurosurg Focus. 2020;48(4):E13. Ryvlin P., Rheims S. Epilepsy surgery: eligibility criteria and presurgical evaluation. Dialogues Clin Neurosci. 2008;10(1):91-103. Kim H.W., Shin D.H., Kim J., et al. Assessing the performance of ChatGPT's responses to questions related to epilepsy: A cross-sectional study on natural language processing and medical information retrieval. Seizure. 2024;1141-8. Wu Y., Zhang Z., Dong X., et al. Evaluating the performance of the language model ChatGPT in responding to common questions of people with epilepsy. Epilepsy Behav. 2024;151109645. Luo Y., Jiao M., Fotedar N., et al. Clinical Value of ChatGPT for Epilepsy Presurgical Decision-Making: Systematic Evaluation of Seizure Semiology Interpretation. J Med Internet Res. 2025;27e69173. Lhatoo SD K.P., Lüders HO. Invasive Studies of the Human Epileptic Brain: Principles and Practice. Oxford: Oxford University Press. 2018; McGovern R.A., Ruggieri P., Bulacio J., et al. Risk analysis of hemorrhage in stereo-electroencephalography procedures. Epilepsia. 2019;60(3):571-580. Thirunavukarasu, A. J., et al. Large language models in medicine. Nature Medicine. 2023;29(8):1930-1940. AlSaad, R., Abd-alrazaq, A., Boughorbel, S., et al. Multimodal Large Language Models in Health Care: Applications, Challenges, and Future Outlook. Journal of Medical Internet Research. 2024;26:e59505. Gidaro, A., et al. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. npj Digital Medicine. 2024;7(1):119. Tables Table 1. Baseline characteristics of the patients Overall (n=157) FLE (n=39) PLE (n=15) TLE (n=73) OLE (n=4) ILE (n=8) Multilobar Epilepsy (n=18) P-Value Male, n (%) 97 (61.8) 22 (56.4) 12 (80.0) 42 (57.5) 3 (75.0) 7 (87.5) 11 (61.1) 0.338 a Age at assessment, y 25.0 [19.0,32.0] 19.0 [16.0,27.5] 25.0 [16.5,28.5] 29.0 [24.0,35.0] 25.5 [21.5,32.5] 24.5 [20.0,30.5] 23.5 [18.0,29.0] 0.004 b Duration of illness, y 10.0 [4.0,18.0] 9.0 [3.5,17.0] 12.0 [7.0,16.5] 10.0 [6.0,20.0] 12.5 [9.2,14.8] 10.5 [3.2,14.2] 11.0 [5.2,21.8] 0.806 b No. of past ASMs, n 4.0 [2.0,5.0] 4.0 [3.0,5.0] 4.0 [3.0,5.5] 3.0 [2.0,5.0] 3.0 [2.0,4.0] 4.5 [3.5,7.2] 4.0 [3.0,5.0] 0.448 b No. of current ASMs, n 2.0 [2.0,3.0] 3.0 [2.0,3.0] 2.0 [2.0,3.0] 2.0 [2.0,2.0] 2.0 [1.8,2.0] 3.0 [2.8,3.0] 2.5 [2.0,3.8] 0.025 b FIAS, n (%) 0.032 a 0 88 (56.1) 19 (48.7) 6 (40.0) 51 (69.9) 2 (50.0) 2 (25.0) 8 (44.4) 1 69 (43.9) 20 (51.3) 9 (60.0) 22 (30.1) 2 (50.0) 6 (75.0) 10 (55.6) FAS, n (%) 0.134 a 0 22 (14.0) 8 (20.5) 3 (20.0) 4 (5.5) 1 (25.0) 2 (25.0) 4 (22.2) 1 135 (86.0) 31 (79.5) 12 (80.0) 69 (94.5) 3 (75.0) 6 (75.0) 14 (77.8) FBTCS, n (%) 0.391 a 0 37 (23.6) 10 (25.6) 2 (13.3) 20 (27.4) 0 (0.0) 3 (37.5) 2 (11.1) 1 120 (76.4) 29 (74.4) 13 (86.7) 53 (72.6) 4 (100.0) 5 (62.5) 16 (88.9) Pathological results, n (%) 0.088 a FCD 88 (56.1) 21 (53.8) 9 (60.0) 44 (60.3) 3 (75.0) 2 (25.0) 9 (50.0) HS 6 (3.8) 1 (2.6) 0 (0.0) 5 (6.8) 0 (0.0) 0 (0.0) 0 (0.0) Tumor 25 (15.9) 2 (5.1) 1 (6.7) 17 (23.3) 1 (25.0) 1 (12.5) 3 (16.7) Cavernous hemangioma 3 (1.9) 1 (2.6) 0 (0.0) 2 (2.7) 0 (0.0) 0 (0.0) 0 (0.0) Others 12 (7.6) 4 (10.3) 2 (13.3) 3 (4.1) 0 (0.0) 1 (12.5) 2 (11.1) Deficiency 23 (14.6) 10 (25.6) 3 (20.0) 2 (2.7) 0 (0.0) 4 (50.0) 4 (22.2) FU duration, y 20.0 [13.0,25.0] 23.0 [16.5,29.0] 16.0 [13.5,23.5] 18.0 [12.0,25.0] 13.5 [12.0,15.0] 22.5 [18.5,26.5] 16.0 [13.0,20.5] 0.085 b Engel, n (%) 0.360 a I 117 (74.5) 30 (76.9) 10 (66.7) 56 (76.7) 3 (75.0) 6 (75.0) 12 (66.7) II 19 (12.1) 1 (2.6) 3 (20.0) 10 (13.7) 0 (0.0) 1 (12.5) 4 (22.2) III 12 (7.6) 4 (10.3) 1 (6.7) 6 (8.2) 0 (0.0) 1 (12.5) 0 (0.0) IV 9 (5.7) 4 (10.3) 1 (6.7) 1 (1.4) 1 (25.0) 0 (0.0) 2 (11.1) Continuous variables were recorded as median (IQR). P a Categorical variables were compared using the Chi-square test/Fisher's exact test. P b Continuous variables were compared using the Mann-Whitney U test (for binary categories) or the Kruskal-Wallis test (for multiple categories). All variables were confirmed to be non-normally distributed by the Shapiro-Wilk test with Bonferroni correction. Abbreviations: ASMs, Anti-seizure medications; FAS, Focal Aware Seizure (presence/absence); FBTCS, Focal to Bilateral Tonic-Clonic Seizure (presence/absence); Focal cortical dysplasia, FCD; FIAS, Focal Impaired Awareness Seizure (presence/absence); FLE, Frontal Lobe Epilepsy; Follow-up, FU; Hippocampal sclerosis, HS; ILE, Insular Lobe Epilepsy; OLE, Occipital Lobe Epilepsy; PLE, Parietal Lobe Epilepsy; TLE, Temporal Lobe Epilepsy; y, year. Table 2. Performance Comparison of Three Large Language Models on Epileptogenic Zone Lateralization and Localization Model Lateralization accuracy, n (%) P-Value Localization score, median [Q1, Q3] P-Value GPT-4.1 154 (98.1) -- 70 [60,70] -- Deepseek-R1 154 (98.1) 1.00 60 [50,70] < 0.001 Claude 3.7 Sonnet 153 (97.5) 1.00 70 [60,70] 0.104 Data are presented as n (%) for lateralization accuracy and median [Q1, Q3] for localization score. P-values represent comparisons to the GPT-4.1 model. Table 3. Performance of Large Language Models in Epileptogenic Zone Lateralization Model Accuracy(%) Sensitivity/Recall(%) Specificity(%) Precision/PPV(%) F1-Score(%) ChatGPT-4.1 98.1 98.8 97.4 97.6 98.2 Deepseek-R1 98.1 98.8 97.4 97.6 98.2 Claude 3.7 97.5 98.8 96.1 96.4 97.6 Metrics are presented as decimal value (percentage %). For binary classification metrics (Sensitivity, Specificity, Precision, NPV, MCC), 'Left side' (n=81) was defined as the positive class and 'Right side' (n=76) as the negative class. The total number of patients was 157. Abbreviations: PPV, Positive Predictive Value; NPV, Negative Predictive Value; MCC, Matthews Correlation Coefficient. Table 4. Performance of GPT-4.1 in Lateralization and Localization Across Different Epilepsy Types Types of Epilepsy Lateralization accuracy, n (%) Localization score, median [Q1, Q3] FLE (n=39) 39 (100.0) 65.0 [60.0,70.0] PLE (n=15) 15 (100.0) 60.0 [35.0,60.0] TLE (n=73) 72 (98.6) 70.0 [70.0,75.5] OLE (n=4) 4 (100.0) 55.0 [12.5,60.0] ILE (n=8) 8 (100.0) 15.0 [2.5,45.0] Multilobar Epilepsy (n=18) 16 (88.9) 80.0 [50.0,90.0] Data are presented as n (%) for lateralization accuracy and median [Q1, Q3] for localization score. Abbreviations: FLE: Frontal Lobe Epilepsy; ILE: Insular Lobe Epilepsy; OLE: Occipital Lobe Epilepsy; TLE: Temporal Lobe Epilepsy; PLE: Parietal Lobe Epilepsy. Additional Declarations No competing interests reported. Supplementary Files Supplementary.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6853779","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":470183358,"identity":"d42ce370-d3dc-467e-8741-7f9fa79b212e","order_by":0,"name":"Yueqian Sun","email":"","orcid":"","institution":"Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yueqian","middleName":"","lastName":"Sun","suffix":""},{"id":470183359,"identity":"95100ff0-2985-4bef-96c2-042203b1820e","order_by":1,"name":"Shihao Ge","email":"","orcid":"","institution":"Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Shihao","middleName":"","lastName":"Ge","suffix":""},{"id":470183360,"identity":"5f9cd883-31ed-4dad-973b-917a73d06edb","order_by":2,"name":"Yangyang Wang","email":"","orcid":"","institution":"Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yangyang","middleName":"","lastName":"Wang","suffix":""},{"id":470183361,"identity":"69a2384c-a8c0-4f62-a6b5-8712715f3859","order_by":3,"name":"Xiaoqiu Shao","email":"","orcid":"","institution":"Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Xiaoqiu","middleName":"","lastName":"Shao","suffix":""},{"id":470183362,"identity":"68cbc6f5-118b-493c-83b5-70c82e40b169","order_by":4,"name":"Kai Zhang","email":"","orcid":"","institution":"Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Kai","middleName":"","lastName":"Zhang","suffix":""},{"id":470183363,"identity":"be9f92eb-d2c0-4f76-a578-6292b4293fd4","order_by":5,"name":"Zhihai He","email":"","orcid":"","institution":"Southern University of Science and technology","correspondingAuthor":false,"prefix":"","firstName":"Zhihai","middleName":"","lastName":"He","suffix":""},{"id":470183364,"identity":"ace0f5b0-26f4-42dc-a1c9-799ed4f2198f","order_by":6,"name":"Qun Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAv0lEQVRIiWNgGAWjYNCCCijNQ7yWMyRrYWwjRYvB8bOHX7ydZyevOyOB8cHbNgZ5c4JazuSlWc7dlmy47UYCs+HcNgbDnQ0EtJgdyDEz5t12IMHsRgKbNG8bQ4LBAUJazr8BapkD1sL+mzgtN3KMH/M2QGxhJkqL/Y03ZoxzjgH9cuZhs+SccxKGGwhpkezPMf7wpsZO3ux48sEPb8ps5AnaAgRsEpDoYGwAEhKE1QMB8wcS0skoGAWjYBSMRAAAD/lAsHAUhfEAAAAASUVORK5CYII=","orcid":"","institution":"Capital Medical University","correspondingAuthor":true,"prefix":"","firstName":"Qun","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2025-06-09 11:08:20","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6853779/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6853779/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":84704138,"identity":"9eaf1d2b-6d82-4990-a215-162f245a203c","added_by":"auto","created_at":"2025-06-16 12:05:24","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":149821,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSchematic overview of the study workflow for evaluating Large Language Models in aiding epileptogenic zone assessment and SEEG recommendation.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/fa79ae5fbebe4a029ab514ce.png"},{"id":84704144,"identity":"0251f747-8db5-481d-bd66-91cc3c743577","added_by":"auto","created_at":"2025-06-16 12:05:25","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":226182,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEnglish translation of prompts guiding Large Language Models for epileptogenic zone assessment and SEEG recommendation.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/66a64c8a40814d425bc0fc6a.png"},{"id":84704139,"identity":"05dfcb23-125d-4498-b1f1-7dab5dfc141e","added_by":"auto","created_at":"2025-06-16 12:05:25","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":50206,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePerformance comparison of different Large Language Models in epileptogenic zone lateralization and localization.\u003c/strong\u003e (a) Lateralization prediction accuracy; (b) Distribution of epileptogenic zone localization scores for GPT-4.1, Deepseek-R1, and Claude 3.7 Sonnet.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/33e67eab42f62f05afd16d40.png"},{"id":84705188,"identity":"690f2b4e-85dc-4302-8f6f-fcdefed8c5a9","added_by":"auto","created_at":"2025-06-16 12:13:25","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":48661,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLocalization performance of the GPT-4.1 model across different epilepsy types.\u003c/strong\u003e Abbreviations: FLE, frontal lobe epilepsy; PLE, parietal lobe epilepsy; TLE, temporal lobe epilepsy; OLE, occipital lobe epilepsy; ILE, insular lobe epilepsy.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/dc516dc674bce8f2748c4802.png"},{"id":84705189,"identity":"71a3ae01-a682-4680-8cdb-5fe166303cd6","added_by":"auto","created_at":"2025-06-16 12:13:25","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":28127,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistribution of SEEG recommendation scores by GPT-4.1 based on actual SEEG application.\u003c/strong\u003e The figure illustrates the frequency distribution of SEEG recommendation scores (0-100) assigned by GPT-4.1, categorized by whether patients actually underwent SEEG (Yes, n=65) or not (No, n=92).\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/a2d419549e72c37d78d871e5.png"},{"id":86159230,"identity":"223290ab-2077-4976-8ea1-d77813333498","added_by":"auto","created_at":"2025-07-07 12:02:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1912165,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/c55025dc-c32b-45bc-aa07-b268533c4f13.pdf"},{"id":84704142,"identity":"dc35fd64-7531-4454-ad00-a6336440bffa","added_by":"auto","created_at":"2025-06-16 12:05:25","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":541802,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementary.docx","url":"https://assets-eu.researchsquare.com/files/rs-6853779/v1/ca03cb3193acaf2539dbd59f.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eCan Large Language Models Aid Pre-surgical Epileptogenic Zone Localization? A Multi-source Text Analysis Performance Study in Drug-Resistant Epilepsy\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eEpilepsy affects over 70\u0026nbsp;million individuals globally, with approximately 30% progressing to drug-resistant focal epilepsy (DRE)\u0026sup1;. Surgical resection of the epileptogenic zone (EZ) remains the most effective treatment for seizure control in DRE patients. The primary goal of presurgical evaluation\u0026mdash;a cornerstone of preoperative assessment for epilepsy surgery\u0026mdash;is to precisely localize the EZ while minimizing adverse effects, ideally through noninvasive methods\u0026sup2;⁻⁴.\u003c/p\u003e \u003cp\u003eDue to the inherent complexity of epilepsy, multidisciplinary preoperative evaluations and meticulous surgical decision-making are essential⁵. Phase 1 noninvasive assessment typically includes prolonged scalp video-electroencephalography (EEG) monitoring⁶, structural and functional neuroimaging⁷\u003csup\u003e,\u003c/sup\u003e⁸, and neuropsychological testing, with protocols tailored to individual patient needs. When these noninvasive investigations yield concordant results, they often provide a robust foundation for localizing the seizure onset zone (SOZ) and formulating a personalized resection plan⁹. However, in 25\u0026ndash;50% of cases, noninvasive evaluations fail to delineate the SOZ\u0026sup1;⁰⁻\u0026sup1;\u0026sup2;. In such scenarios, invasive Phase 2 assessments\u0026mdash;such as intracranial EEG recording\u0026mdash;are required to precisely identify the SOZ and its spatial relationship to eloquent cortex⁴\u003csup\u003e,\u003c/sup\u003e⁵\u003csup\u003e,\u003c/sup\u003e\u0026sup1;\u0026sup3;. This invasive approach enables definitive evaluation of resectability while mitigating risks to critical neurological functions. Common modalities include electrocorticography, subdural electrodes (strips/grids), and stereo-EEG (SEEG), with SEEG now serving as the primary intracranial EEG modality in most North American and Chinese epilepsy centers\u0026sup1;⁴\u003csup\u003e,\u003c/sup\u003e\u0026sup1;⁵.\u003c/p\u003e \u003cp\u003eDespite the relatively high seizure-free rates achieved post-surgery, approximately one-third of patients experience persistent or recurrent seizures, primarily due to incomplete EZ resection\u0026sup1;⁶. This can stem from insufficient presurgical identification of the target resection volume or anatomical constraints near eloquent cortex. Additionally, challenges in epilepsy surgery evaluations include subjective expert judgments, inconsistent ancillary test results, difficulties in multimodal data integration, and time-consuming, empirical manual interpretations. Furthermore, SEEG carries risks such as limited brain coverage, intracranial infections from prolonged monitoring for spontaneous seizures, and the potential failure to capture habitual seizures\u0026sup3;. Reducing unnecessary SEEG procedures through improved presurgical assessments not only alleviates patient burdens\u0026mdash;such as physical discomfort, psychological stress, and economic costs\u0026mdash;but also minimizes the risks associated with invasive interventions. Thus, there is an urgent need for innovative approaches to enhance the accuracy of presurgical evaluations and improve treatment outcomes for epilepsy patients.\u003c/p\u003e \u003cp\u003eLarge language models (LLMs), which have emerged as transformative tools in medicine, demonstrate groundbreaking capabilities in biomedical text analysis. Their advanced reasoning and universal adaptability far exceed traditional machine learning methods, offering promise for enhancing clinical decision-making, automating administrative tasks, and improving patient care\u0026sup1;⁷⁻\u0026sup1;⁹. In the field of epilepsy, LLMs have shown notable performance, including passing professional examinations\u0026sup2;⁰ and aiding in differential diagnosis\u0026sup2;\u0026sup1;. A recent study further developed an LLM-based method to automatically extract seizure frequencies from epilepsy monitoring unit (EMU) reports for Sudden Unexpected Death in Epilepsy (SUDEP) risk assessment\u0026sup2;\u0026sup2;. However, systematic research focusing on the application of LLMs to localize the EZ through the comprehensive analysis of multi-source text data\u0026mdash;a task that involves inferring spatial features from diverse and abstract textual sources\u0026mdash;is still in its early stages.\u003c/p\u003e \u003cp\u003eAs a major innovation, this study constructs an LLM-driven analysis framework that integrates medical records, EEG, magnetic resonance imaging (MRI), positron emission tomography (PET), and magnetoencephalography (MEG) preoperative reports from 157 DRE patients. We compare the performance of three leading LLMs (GPT-4.1, Deepseek-R1, and Claude 3.7 Sonnet) without fine-tuning. The study aims to (1) determine seizure laterality, (2) localize the EZ at the lobar level, and (3) recommendation for Phase 2 preoperative evaluation (SEEG). By exploring the feasibility of text-driven EZ localization using LLMs, this study assesses their potential as an auxiliary tool in noninvasive preoperative evaluation. The aim is to provide initial insights that could support the complex localization process and contribute to informed clinical decision-making, particularly concerning the use of SEEG.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003ePatient Collection\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis retrospective study included patients with DRE who underwent surgical intervention (resection or thermal ablation) at the Beijing Tiantan Hospital Epilepsy Center between January 2020 and April 2024. Inclusion criteria were: (1) Confirmed DRE diagnosis. (2) Multidisciplinary team (MDT)-determined EZ localization and surgical strategy. (3) Complete baseline clinical data (clinical history, seizure semiology, in-hospital notes) and full preoperative evaluations (prolonged scalp EEG, structural MRI, \u0026sup1;⁸F-FDG PET). (4) \u0026ge;12-month post-surgical follow-up. Exclusion criteria were: (1) Non-structural etiologies (e.g., genetic/metabolic disorders, encephalitis). (2) Missing data (incomplete EEG/MRI/PET records). (3) Involvement of three or more lobes, as determined by postoperative CT or MRI scans. The study was approved by the institutional ethics committee, with written informed consent obtained from all participants. The detailed patient selection process is illustrated in Figure S1.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStandard Presurgical Evaluation and Gold Standard Definition\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAn overview of the study workflow for EZ assessment using LLMs is provided in Figure 1. Standard presurgical evaluation included: (1) Seizure history and semiology. (2) Long-term scalp video-EEG monitoring. (3) 3-T thin-slice MRI. (4) Interictal FDG-PET/MRI co-registration (visually analyzed via standardized color scales). (5) Neuropsychological testing. (6) MEG was selectively performed. MRI-negative status was defined as the absence of clinically relevant structural abnormalities, as determined by consensus among epileptologists.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAll data were reviewed at MDT conferences to determine surgical eligibility and intervention plans, with a focus on identifying candidates for SEEG\u0026mdash;primarily those with MRI-negative or subtle lesions. The gold standard for assessing the LLM\u0026apos;s EZ localization accuracy was the set of surgically resected lobe(s), determined by MDT consensus as the target for surgical intervention. EZs were categorized as unilobar (single lobe) or multilobar (involving two distinct lobes).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Preparation for Large Language Models\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor each patient, textual data from the following five sources were collected:\u003c/p\u003e\n\u003cp\u003ea. Medical Records: This included admission history (detailing seizure history, comorbidities, etc.) and in-hospital progress notes (documenting dynamic seizure manifestations such as auras, automatisms, and other relevant clinical observations during hospitalization).\u003c/p\u003e\n\u003cp\u003eb. MRI Reports: Textual descriptions and diagnostic conclusions provided by neuroradiologists.\u003c/p\u003e\n\u003cp\u003ec. Scalp EEG Reports: Textual descriptions and interpretations of prolonged scalp video-EEG findings provided by neurophysiologists.\u003c/p\u003e\n\u003cp\u003ed. PET Reports: Textual descriptions and diagnostic conclusions from nuclear medicine physicians.\u003c/p\u003e\n\u003cp\u003ee. MEG Reports: Textual descriptions and analytical conclusions from MEG data analysts (available for a subset of patients).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eData preprocessing included anonymization (removal of identifiers) and text data cleaning (formatting/special characters). For each patient, texts from these five sources were concatenated into one single text input to the LLM. The input size is constrained by the LLM\u0026rsquo;s prompt limit.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLarge Language Model Configuration, Tasks, and Analytical Framework\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThree LLMs\u0026mdash;GPT-4.1 (OpenAI), Deepseek-R1 (Deepseek AI), and Claude 3.7 Sonnet (Anthropic)\u0026mdash;were evaluated via official APIs with temperature = 0 (default parameters otherwise). A standardized prompt (Figure 2; Supplementary Figure S2) directed models to perform three tasks: (1) EZ Laterality Classification: determine the \u0026quot;Left\u0026quot; or \u0026quot;Right\u0026quot; hemisphere; (2) Lobar-Level Localization: select the most likely top 3 lobes based on the predicted probabilities (with a total sum of 100%). (3) SEEG Recommendation Scoring: assign a score from 0 to 100 to indicate SEEG necessity. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePerformance metrics included: (1) Laterality Accuracy: proportion of the hemisphere predictions matching MDT consensus (coded as 1 for correct, 0 for incorrect). (2) Localization Score: predicted probability score of the gold standard lobe (if ranked within top 3) or 0 for single-lobe epilepsy cases. For combined-lobe cases (e.g., fronto-temporal), the score was the sum of the probability scores assigned by the LLM to those MDT-defined epileptogenic lobes that were present in the LLM\u0026apos;s top three predictions (range: 0-100%). (3) SEEG Recommendation Analysis: comparison of LLM-predicted scores between patients who underwent SEEG and those who did not.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe analytical framework consisted of five stages: (1) Inter-Model Comparison: Evaluated laterality accuracy and localization scores across all LLMs using all multi-source data for the entire cohort and MRI-negative subgroup. (2) GPT-4.1 Subgroup Analysis: Assessed performance stratified by epilepsy type (frontal/parietal/temporal/occipital/insula/combined-lobe) and MRI status (negative/positive). (3) Textual Information Source Ablation Study: We compared GPT-4.1\u0026rsquo;s performance when the textual information collected from each of the five sources (i.e., clinical notes, MRI reports, EEG reports, PET reports, and MEG reports) was individually removed. The baseline for this comparison was the performance achieved using textual input from all five sources. The ablation analysis for the textual data from MEG reports was restricted to the 64 patients who had these reports. (4) Stability Analysis: Conducted three runs of GPT-4.1 to calculate the Intraclass Correlation Coefficient (ICC) for localization score consistency. (5) SEEG Necessity Stratification: Compared GPT-4.1\u0026rsquo;s SEEG recommendation scores between patients who underwent the procedure and those who did not.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll statistical analyses were performed using SPSS Statistics version 27.0 (IBM Corp., Armonk, NY, USA). Descriptive statistics included frequencies and percentages [n (%)] for categorical variables. Continuous variables were assessed for normality using the Shapiro-Wilk test; normally distributed data were presented as mean \u0026plusmn; standard deviation (Mean \u0026plusmn; SD), while non-normally distributed data (e.g., localization scores) were presented as median and interquartile range [Median (IQR)].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor inferential statistics: (1) McNemar test was used for pairwise comparisons of laterality accuracy (e.g., inter-model, ablation effects). (2) Fisher\u0026rsquo;s exact test was used for independent subgroup comparisons of laterality accuracy (e.g., MRI-negative vs. -positive). (3) For non-normally distributed continuous variables (such as localization scores, SEEG recommendation scores, and age at assessment), the Wilcoxon signed-rank test was used for pairwise comparisons, the Mann-Whitney U test for independent two-group comparisons, and the Kruskal-Wallis H test for independent multi-group comparisons. (4) ICC with its 95% confidence interval (CI) was used to assess stability.\u003c/p\u003e\n\u003cp\u003eAll statistical tests were two-sided, with a p-value \u0026lt; 0.05 considered statistically significant.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003ePatients Characteristics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of 157 patients with DRE were enrolled in this study. The median age at assessment was 25.0 years (IQR: 19.0-32.0), with 97 (61.8%) male participants. Patients were stratified into epilepsy subtypes based on postoperative CT/MRI findings, as detailed in Table 1. The median duration from initial seizure onset to surgical intervention was 10.0 years (IQR: 4.0-18.0). Significant differences were observed among subtypes in age at assessment (p = 0.004), current antiseizure medication (ASM) count (p = 0.025), and frequency of focal impaired awareness seizures (FIAS) (p = 0.032). Pathological findings suggested a trend toward differences among subtypes, but this did not reach statistical significance (p = 0.088).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEpileptogenic Zone Localization\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUsing the surgical resection site determined by the MDT as the gold standard, GPT-4.1 and Deepseek-R1 demonstrated identical laterality accuracy (98.1%) for EZ localization, while Claude 3.7 Sonnet achieved 97.5% (p = 1.00; Figure 3a, Table 2, Table 3). For lobar-level localization, GPT-4.1 and Claude 3.7 Sonnet attained median scores of 70 versus 60 for Deepseek-R1 (p \u0026lt; 0.001; Figure 3b, Table 2). Notably, even in diagnostically challenging MRI-negative cases (n = 71), all models achieved high lateralization accuracy. Furthermore, both GPT-4.1 and Claude 3.7 Sonnet performed well in lobar localization within this subgroup (Supplementary Figure S3, Table S2). This ability to provide valuable localizing information for such patients underscores the potential utility of LLMs.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFurther analysis of GPT-4.1 across epilepsy subtypes revealed distinct patterns (Table 4). The model achieved perfect lateralization accuracy (100%) for frontal, parietal, occipital, and insular lobe epilepsy, with 98.6% accuracy in temporal lobe epilepsy (TLE) and 88.9% in multilobar cases. Localization scores varied significantly (Kruskal-Wallis test, p \u0026lt; 0.001): multilobar epilepsy yielded the highest median score (80.0 [IQR: 50.0-90.0]), followed by TLE (70.0 [IQR: 70.0-75.5]) and frontal lobe epilepsy (FLE) (65.0 [IQR: 60.0-70.0]). Parietal (60.0 [IQR: 35.0-60.0]) and occipital (55.0 [IQR: 12.5-60.0]) lobes showed moderate performance, while insular lobe epilepsy (ILE) had the lowest median score (15.0 [IQR: 2.5-45.0]), reflecting inherent challenges in localizing this subtype. Figure 4 visualizes these distributional differences across epilepsy types. GPT-4.1 demonstrated no significant differences in lateralization accuracy or localization scores between MRI-negative and MRI-positive patients (Supplementary Table S1).\u003c/p\u003e\n\u003cp\u003eA modality ablation study (Supplementary Table S3) revealed that removal of medical records or MRI reports significantly impaired GPT-4.1\u0026rsquo;s EZ localization performance, whereas exclusion of EEG, PET, or MEG reports had minimal impact.\u003c/p\u003e\n\u003cp\u003eTo assess GPT-4.1\u0026rsquo;s stability, the lobar localization task with all five modalities was repeated three times. Using a two-way mixed-effects model, the ICC for localization scores demonstrated excellent consistency: single measures ICC = 0.960 (95% confidence interval [CI]: 0.948-0.969, p \u0026lt; 0.001) and average measures ICC = 0.986 (95% CI: 0.982-0.990, p \u0026lt; 0.001).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSEEG Recommendation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study evaluated the capacity of LLMs to recommend SEEG monitoring, focusing on GPT-4.1. SEEG recommendation scores differed significantly between patients who underwent SEEG (n = 65) and those who did not (n = 92; p \u0026lt; 0.001; Figure 5). GPT-4.1 assigned higher median scores to SEEG-treated patients (90.00 [IQR: 85.00-90.00]) compared to non-SEEG patients (25.00 [IQR: 10.00-90.00]). These findings suggest that GPT-4.1 can stratify SEEG necessity with high discriminative power, potentially guiding clinical decisions to prioritize invasive monitoring in complex cases.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study systematically evaluates the utility of LLMs in presurgical epilepsy assessment, focusing on three major tasks: EZ laterality determination, lobar-level localization, and SEEG recommendation stratification. By leveraging multi-source clinical textual data\u0026mdash;including medical records, EEG, MRI, PET, and MEG reports\u0026mdash;without requiring direct access to electrophysiological or imaging datasets, our framework demonstrates the potential to enhance diagnostic accessibility across diverse clinical settings and streamline presurgical workflows. Notably, GPT-4.1 achieved decision-making accuracy comparable to MDT consensus, while Deepseek-R1 and Claude 3.7 Sonnet exhibited strong but suboptimal performance. These findings position advanced LLMs as scalable decision-support tools for epilepsy surgery, particularly in regions with limited MDT resources.\u003c/p\u003e \u003cp\u003eConsistency among diagnostic modalities is a well-established predictor of favorable postoperative outcomes in epilepsy surgery\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. Non-eloquent, MRI-visible epileptogenic lesions paired with congruent seizure semiology and EEG findings typically yield optimal surgical results\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. However, conflicting evaluations, inconclusive data, or the presence of multiple MRI lesions often complicate EZ identification. Currently, no definitive preoperative gold standard exists for EZ localization beyond standardized initial assessments (e.g., semiology analysis, video-EEG, MRI, neuropsychological testing)\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Our study demonstrates that LLMs can achieve high accuracy in EZ lateralization and lobar localization, closely aligning with the EZ defined by MDT consensus, offering a novel approach to address diagnostic ambiguity. Unlike traditional workflows reliant on time-consuming, expert-driven evaluations, LLMs rapidly integrate heterogeneous data sources, providing consistent, evidence-based recommendations. This efficiency and reliability underscore their potential to revolutionize epilepsy diagnosis, particularly in settings with limited access to MDT expertise.\u003c/p\u003e \u003cp\u003ePrior research has highlighted the educational and basic diagnostic capabilities of LLMs in epilepsy care. For instance, Kim et al. demonstrated high educational fidelity in addressing epilepsy FAQs but noted limitations in mechanistic insights into EZ localization\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. Wu et al. further identified gaps in prognostic inference, with only 46.8% of outcome-related queries receiving comprehensive answers\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Recent studies have begun exploring LLMs for EZ localization: Zhang et al. (2025) showed that ChatGPT could infer EZ locations from seizure semiology descriptions alone, achieving 80\u0026ndash;90% regional sensitivity for frontal/temporal lobes and \u0026ge;\u0026thinsp;67% weighted sensitivity across imbalanced datasets\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. However, this work focused exclusively on unimodal text interpretation (semiology reports) and excluded patients with multilobar involvement. In contrast, our study establishes a multimodal LLM framework that integrates structural MRI, EEG, PET, MEG, and clinical narratives to mirror real-world presurgical workflows. Compared to Zhang et al.\u0026rsquo;s semiology-centric approach, our model matches their laterality accuracy (98.1% vs. 80\u0026ndash;90% regional sensitivity) while introducing probability-weighted lobar localization, a critical advancement for complex cases like combined-lobe epilepsy, where quantifying seizure burden guides surgical planning. This evolution extends beyond semiology-based localization to hierarchical clinical reasoning, aligning more closely with MDT decision-making processes.\u003c/p\u003e \u003cp\u003eSEEG is a critical invasive tool for mapping the SOZ and epileptogenic network via multi-contact depth electrodes\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Its hypothesis-driven nature, guided by seizure semiology and anatomical-electro-clinical correlations, necessitates careful electrode placement to balance diagnostic yield and procedural risks\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. While LLMs do not fully replicate MDT decisions in second-stage SEEG assessments, our findings reveal a notable alignment in certain scenarios. Specifically, GPT-4.1 assigned markedly higher SEEG recommendation scores to patients who ultimately underwent the procedure compared to those who did not. However, it is crucial to interpret these findings with caution. This statistical correlation, while significant, primarily suggests the LLM's potential to discern patterns indicative of SEEG necessity from the provided patient cohort's data, rather than a comprehensive replication of the complex, individualized MDT decision-making process. While prior benchmarks were often limited to categorical EZ localization, this exploratory study indicates LLMs' potential to contribute to hierarchical presurgical reasoning, perhaps by offering a preliminary quantification of SEEG consideration. Nonetheless, substantial further research is required to validate these initial observations and to understand how such tools might responsibly augment, rather than replace, expert clinical judgment in real-world workflows, ensuring that any potential benefits in stratifying SEEG necessity do not overlook critical nuances or introduce unforeseen biases. The prospect of reducing unnecessary invasive procedures remains an important goal, but one that necessitates rigorous validation of these emerging technologies.\u003c/p\u003e \u003cp\u003eThe promising yet variable performance of LLMs in EZ localization and SEEG recommendation is intrinsically tied to the current capabilities and limitations of language model technology. A key strength lies in their extensive context windows (e.g., GPT-4.1, Claude 3.7, and Deepseek-R1 supporting 128K-200K tokens), enabling comprehensive integration of complex patient histories, including detailed seizure evolution narratives and multifaceted examination reports. This capacity may allow LLMs to detect subtle patterns overlooked by clinicians due to time constraints or information overload\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. Furthermore, prompting models to output their \u0026ldquo;reasoning process\u0026rdquo; alongside predictions provides partial transparency, mitigating the \u0026ldquo;black box\u0026rdquo; perception of deep learning models and offering insights into how textual evidence is weighted.\u003c/p\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eThis study has several limitations. Firstly, inherent constraints of current LLM technology impact our approach. These models rely exclusively on textual data, unable to interpret raw imaging or electrophysiological signals\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e, rendering them susceptible to the quality of input reports. While future multimodal LLMs, Retrieval-Augmented Generation (RAG), or domain-specific fine-tuning hold promise\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e,\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e, these current technological gaps are relevant. Secondly, our study design introduces further limitations. The single-center Chinese cohort may restrict generalizability. The small number of participating epileptologists could affect the robustness of our gold standard. Finally, our focus on MDT results without incorporating SEEG data \u0026mdash; crucial for refining EZ localization in complex cases \u0026mdash; might underestimate LLM accuracy, particularly its potential alignment with definitive postoperative outcomes when invasive imaging is considered.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThese findings highlight the potential of LLMs in epilepsy care: In specialized centers, they serve as intelligent co-pilots, augmenting epileptologists\u0026rsquo; diagnostic efficiency. In resource-limited settings, they address critical knowledge gaps, enabling clinicians to perform initial seizure classification and make informed clinical decisions, thereby improving access to care. Crucially, unlike clinicians who rely on individual experience, LLMs trained on extensive datasets can generate data-driven predictions of the EZ prior to surgery, offering scalable opportunities to optimize surgical outcomes beyond the scope of traditional workflows.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe present study was approved by the Institutional Review Board of the Beijing Tiantan Hospital affiliated with Capital Medical University (Beijing, China) (KY2023-079-01).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and material\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of interest statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors report no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll participants or their caregivers provided written informed consent for the publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study was financially supported by the National Key R\u0026amp;D Program of China grant (2022YFC2503800), the Capital Health Research and Development of Special grants (2024-1-2041), National Natural Science Foundation of China (U24A20695 and 82371449) and China Postdoctoral Science Foundation (2024M762180).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eQW, YQS and SHG concepted, designed, and supervised the study. SHG acquired the data. YQS, SHG, YYW, XQS and KZ analyzed and interpreted the data, provided statistical analysis, had full access to all of the data in the study, and are responsible for the integrity of the data and the accuracy of the data analysis. YQS and SHG drafted the manuscript, QW and ZHH critically revised the manuscript for important intellectual content. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe would like to express gratitude to all the participants for their cooperation in our study. We would also like to thank the researchers for their contributions in collecting and following up with the patients.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eThijs R.D., Surges R., O\u0026apos;Brien T.J., et al. Epilepsy in adults. Lancet. 2019;393(10172):689-701.\u003c/li\u003e\n\u003cli\u003eRosenow F., L\u0026uuml;ders H. Presurgical evaluation of epilepsy. Brain. 2001;124(Pt 9):1683-700.\u003c/li\u003e\n\u003cli\u003eCossu M., Cardinale F., Castana L., et al. Stereoelectroencephalography in the presurgical evaluation of focal epilepsy: a retrospective analysis of 215 procedures. Neurosurgery. 2005;57(4):706-18; discussion 706-18.\u003c/li\u003e\n\u003cli\u003eMullin J.P., Shriver M., Alomar S., et al. Is SEEG safe? A systematic review and meta-analysis of stereo-electroencephalography-related complications. Epilepsia. 2016;57(3):386-401.\u003c/li\u003e\n\u003cli\u003eVakharia V.N., Duncan J.S., Witt J.A., et al. Getting the best outcomes from epilepsy surgery. Ann Neurol. 2018;83(4):676-690.\u003c/li\u003e\n\u003cli\u003eEngel J., Jr., Driver M.V., Falconer M.A. Electrophysiological correlates of pathology and surgical results in temporal lobe epilepsy. Brain. 1975;98(1):129-56.\u003c/li\u003e\n\u003cli\u003eJack C.R., Jr., Sharbrough F.W., Marsh W.R. Use of MR imaging for quantitative evaluation of resection for temporal lobe epilepsy. Radiology. 1988;169(2):463-8.\u003c/li\u003e\n\u003cli\u003eMazziotta J.C., Engel J., Jr. The use and impact of positron computed tomography scanning in epilepsy. Epilepsia. 1984;25 Suppl 2S86-104.\u003c/li\u003e\n\u003cli\u003eLado F.A., Ahrens S.M., Riker E., et al. Guidelines for Specialized Epilepsy Centers: Executive Summary of the Report of the National Association of Epilepsy Centers Guideline Panel. Neurology. 2024;102(4):e208087.\u003c/li\u003e\n\u003cli\u003eThadani V.M., Williamson P.D., Berger R., et al. Successful epilepsy surgery without intracranial EEG recording: criteria for patient selection. Epilepsia. 1995;36(1):7-15.\u003c/li\u003e\n\u003cli\u003eKilpatrick C., Cook M., Kaye A., et al. Non-invasive investigations successfully select patients for temporal lobe surgery. J Neurol Neurosurg Psychiatry. 1997;63(3):327-33.\u003c/li\u003e\n\u003cli\u003eDiehl B., L\u0026uuml;ders H.O. Temporal lobe epilepsy: when are invasive recordings needed? Epilepsia. 2000;41 Suppl 3S61-74.\u003c/li\u003e\n\u003cli\u003eRodionov R., O\u0026apos;Keeffe A., Nowell M., et al. Increasing the accuracy of 3D EEG implantations. J Neurosurg. 2020;133(1):35-42.\u003c/li\u003e\n\u003cli\u003eGavvala J., Zafar M., Sinha S.R., et al. Stereotactic EEG Practices: A Survey of United States Tertiary Referral Epilepsy Centers. J Clin Neurophysiol. 2022;39(6):474-480.\u003c/li\u003e\n\u003cli\u003eLin Y., Hu S., Hao X., et al. Epilepsy centers in China: Current status and ways forward. Epilepsia. 2021;62(11):2640-2650.\u003c/li\u003e\n\u003cli\u003eSpencer S., Huh L. Outcomes of epilepsy surgery in adults and children. Lancet Neurol. 2008;7(6):525-37.\u003c/li\u003e\n\u003cli\u003eQuer G., Topol E.J. The potential for large language models to transform cardiovascular medicine. Lancet Digit Health. 2024;6(10):e767-e771.\u003c/li\u003e\n\u003cli\u003eBellini V., Bignami E.G. Generative Pre-trained Transformer 4 (GPT-4) in clinical settings. Lancet Digit Health. 2025;7(1):e6-e7.\u003c/li\u003e\n\u003cli\u003eBoussina A., Krishnamoorthy R., Quintero K., et al. Large Language Models for More Efficient Reporting of Hospital Quality Measures. Nejm ai. 2024;1(11):\u003c/li\u003e\n\u003cli\u003eHabib S., Butt H., Goldenholz S.R., et al. Large Language Model Performance on Practice Epilepsy Board Examinations. JAMA Neurol. 2024;81(6):660-661.\u003c/li\u003e\n\u003cli\u003eFord J., Pevy N., Grunewald R., et al. Can artificial intelligence diagnose seizures based on patients\u0026apos; descriptions? A study of GPT-4. Epilepsia. 2025;\u003c/li\u003e\n\u003cli\u003eAbeysinghe R., Tao S., Lhatoo S.D., et al. Leveraging pretrained language models for seizure frequency extraction from epilepsy evaluation reports. NPJ Digit Med. 2025;8(1):208.\u003c/li\u003e\n\u003cli\u003eKankirawatana P., Mohamed I.S., Lauer J., et al. Relative contribution of individual versus combined functional imaging studies in predicting seizure freedom in pediatric epilepsy surgery: an area under the curve analysis. Neurosurg Focus. 2020;48(4):E13.\u003c/li\u003e\n\u003cli\u003eRyvlin P., Rheims S. Epilepsy surgery: eligibility criteria and presurgical evaluation. Dialogues Clin Neurosci. 2008;10(1):91-103.\u003c/li\u003e\n\u003cli\u003eKim H.W., Shin D.H., Kim J., et al. Assessing the performance of ChatGPT\u0026apos;s responses to questions related to epilepsy: A cross-sectional study on natural language processing and medical information retrieval. Seizure. 2024;1141-8.\u003c/li\u003e\n\u003cli\u003eWu Y., Zhang Z., Dong X., et al. Evaluating the performance of the language model ChatGPT in responding to common questions of people with epilepsy. Epilepsy Behav. 2024;151109645.\u003c/li\u003e\n\u003cli\u003eLuo Y., Jiao M., Fotedar N., et al. Clinical Value of ChatGPT for Epilepsy Presurgical Decision-Making: Systematic Evaluation of Seizure Semiology Interpretation. J Med Internet Res. 2025;27e69173.\u003c/li\u003e\n\u003cli\u003eLhatoo SD K.P., L\u0026uuml;ders HO. Invasive Studies of the Human Epileptic Brain: Principles and Practice. Oxford: Oxford University Press. 2018;\u003c/li\u003e\n\u003cli\u003eMcGovern R.A., Ruggieri P., Bulacio J., et al. Risk analysis of hemorrhage in stereo-electroencephalography procedures. Epilepsia. 2019;60(3):571-580.\u003c/li\u003e\n\u003cli\u003eThirunavukarasu, A. J., et al. Large language models in medicine. Nature Medicine. 2023;29(8):1930-1940.\u003c/li\u003e\n\u003cli\u003eAlSaad, R., Abd-alrazaq, A., Boughorbel, S., et al. Multimodal Large Language Models in Health Care: Applications, Challenges, and Future Outlook. Journal of Medical Internet Research. 2024;26:e59505.\u003c/li\u003e\n\u003cli\u003eGidaro, A., et al. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. npj Digital Medicine. 2024;7(1):119.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003e\u003cstrong\u003eTable 1. Baseline characteristics of the patients\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"1017\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOverall (n=157)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFLE (n=39)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePLE (n=15)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTLE (n=73)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOLE (n=4)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eILE (n=8)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMultilobar Epilepsy (n=18)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eP-Value\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMale, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e97 (61.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e22 (56.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e12 (80.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e42 (57.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e3 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e7 (87.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e11 (61.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.338\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAge at assessment, y\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e25.0 [19.0,32.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e19.0 [16.0,27.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e25.0 [16.5,28.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e29.0 [24.0,35.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e25.5 [21.5,32.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e24.5 [20.0,30.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e23.5 [18.0,29.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.004\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDuration of illness, y\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e10.0 [4.0,18.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e9.0 [3.5,17.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e12.0 [7.0,16.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e10.0 [6.0,20.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e12.5 [9.2,14.8]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e10.5 [3.2,14.2]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e11.0 [5.2,21.8]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.806\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNo. of past ASMs, n\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e4.0 [2.0,5.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e4.0 [3.0,5.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e4.0 [3.0,5.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e3.0 [2.0,5.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e3.0 [2.0,4.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e4.5 [3.5,7.2]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e4.0 [3.0,5.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.448\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNo. of current ASMs, n\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e2.0 [2.0,3.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e3.0 [2.0,3.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e2.0 [2.0,3.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e2.0 [2.0,2.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e2.0 [1.8,2.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e3.0 [2.8,3.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e2.5 [2.0,3.8]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.025\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFIAS, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.032\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e88 (56.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e19 (48.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e6 (40.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e51 (69.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e2 (50.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e2 (25.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e8 (44.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e69 (43.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e20 (51.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e9 (60.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e22 (30.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e2 (50.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e6 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e10 (55.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFAS, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.134\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e22 (14.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e8 (20.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e3 (20.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e4 (5.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e1 (25.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e2 (25.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e4 (22.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e135 (86.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e31 (79.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e12 (80.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e69 (94.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e3 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e6 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e14 (77.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFBTCS, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.391\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e37 (23.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e10 (25.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e2 (13.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e20 (27.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e3 (37.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e2 (11.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e120 (76.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e29 (74.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e13 (86.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e53 (72.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e4 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e5 (62.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e16 (88.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePathological results, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.088\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFCD\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e88 (56.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e21 (53.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e9 (60.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e44 (60.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e3 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e2 (25.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e9 (50.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eHS\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e6 (3.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e1 (2.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e5 (6.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTumor\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e25 (15.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e2 (5.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e1 (6.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e17 (23.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e1 (25.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e1 (12.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e3 (16.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCavernous hemangioma\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e3 (1.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e1 (2.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e2 (2.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOthers\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e12 (7.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e4 (10.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e2 (13.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e3 (4.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e1 (12.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e2 (11.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDeficiency\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e23 (14.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e10 (25.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e3 (20.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e2 (2.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e4 (50.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e4 (22.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFU duration, y\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e20.0 [13.0,25.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e23.0 [16.5,29.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e16.0 [13.5,23.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e18.0 [12.0,25.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e13.5 [12.0,15.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e22.5 [18.5,26.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e16.0 [13.0,20.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.085\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEngel, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\n \u003cp\u003e0.360\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e117 (74.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e30 (76.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e10 (66.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e56 (76.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e3 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e6 (75.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e12 (66.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eII\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e19 (12.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e1 (2.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e3 (20.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e10 (13.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e1 (12.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e4 (22.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eIII\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e12 (7.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e4 (10.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e1 (6.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e6 (8.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e1 (12.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 196px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eIV\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 126px;\"\u003e\n \u003cp\u003e9 (5.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 104px;\"\u003e\n \u003cp\u003e4 (10.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 94px;\"\u003e\n \u003cp\u003e1 (6.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e1 (1.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 98px;\"\u003e\n \u003cp\u003e1 (25.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 109px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 123px;\"\u003e\n \u003cp\u003e2 (11.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 72px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eContinuous variables were recorded as median (IQR). P\u003csup\u003ea\u003c/sup\u003e Categorical variables were compared using the Chi-square test/Fisher\u0026apos;s exact test. P\u003csup\u003eb\u003c/sup\u003e Continuous variables were compared using the Mann-Whitney U test (for binary categories) or the Kruskal-Wallis test (for multiple categories). All variables were confirmed to be non-normally distributed by the Shapiro-Wilk test with Bonferroni correction.\u003c/p\u003e\n\u003cp\u003eAbbreviations:\u0026nbsp;ASMs, Anti-seizure medications; FAS, Focal Aware Seizure (presence/absence); FBTCS, Focal to Bilateral Tonic-Clonic Seizure (presence/absence); Focal cortical dysplasia, FCD; FIAS, Focal Impaired Awareness Seizure (presence/absence); FLE, Frontal Lobe Epilepsy; Follow-up, FU; Hippocampal sclerosis, HS;\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eILE, Insular Lobe Epilepsy; OLE, Occipital Lobe Epilepsy; PLE, Parietal Lobe Epilepsy; TLE, Temporal Lobe Epilepsy; y, year.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2. Performance Comparison of Three Large Language Models on Epileptogenic Zone Lateralization and Localization\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"881\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 246px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLateralization accuracy, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eP-Value\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 327px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLocalization score, median [Q1, Q3]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eP-Value\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eGPT-4.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 246px;\"\u003e\n \u003cp\u003e154 (98.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e--\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 327px;\"\u003e\n \u003cp\u003e70 [60,70]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e--\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDeepseek-R1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 246px;\"\u003e\n \u003cp\u003e154 (98.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 327px;\"\u003e\n \u003cp\u003e60 [50,70]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e\u0026lt; 0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eClaude 3.7 Sonnet\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 246px;\"\u003e\n \u003cp\u003e153 (97.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 327px;\"\u003e\n \u003cp\u003e70 [60,70]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 78px;\"\u003e\n \u003cp\u003e0.104\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eData are presented as n (%) for lateralization accuracy and median [Q1, Q3] for localization score. P-values represent comparisons to the GPT-4.1 model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3. Performance of Large Language Models in Epileptogenic Zone Lateralization\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"891\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 141px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy(%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSensitivity/Recall(%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSpecificity(%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrecision/PPV(%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eF1-Score(%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 141px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eChatGPT-4.1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e98.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003e98.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e97.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e97.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e98.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 141px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDeepseek-R1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e98.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003e98.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e97.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e97.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e98.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 141px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eClaude 3.7\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 151px;\"\u003e\n \u003cp\u003e97.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 159px;\"\u003e\n \u003cp\u003e98.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e96.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e96.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 146px;\"\u003e\n \u003cp\u003e97.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eMetrics are presented as decimal value (percentage %). For binary classification metrics (Sensitivity, Specificity, Precision, NPV, MCC), \u0026apos;Left side\u0026apos; (n=81) was defined as the positive class and \u0026apos;Right side\u0026apos; (n=76) as the negative class. The total number of patients was 157.\u003c/p\u003e\n\u003cp\u003eAbbreviations: PPV, Positive Predictive Value; NPV, Negative Predictive Value; MCC, Matthews Correlation Coefficient.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4. Performance of GPT-4.1 in Lateralization and Localization Across Different Epilepsy Types\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"884\" class=\"fr-table-selection-hover\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTypes of Epilepsy\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLateralization accuracy, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLocalization score, median [Q1, Q3]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFLE (n=39)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e39 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e65.0 [60.0,70.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePLE (n=15)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e15 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e60.0 [35.0,60.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTLE (n=73)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e72 (98.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e70.0 [70.0,75.5]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOLE (n=4)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e4 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e55.0 [12.5,60.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eILE (n=8)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e8 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e15.0 [2.5,45.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 272px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMultilobar Epilepsy\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;(n=18)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 236px;\"\u003e\n \u003cp\u003e16 (88.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 376px;\"\u003e\n \u003cp\u003e80.0 [50.0,90.0]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eData are presented as n (%) for lateralization accuracy and median [Q1, Q3] for localization score.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAbbreviations:\u0026nbsp;\u003c/strong\u003eFLE: Frontal Lobe Epilepsy; ILE: Insular Lobe Epilepsy; OLE: Occipital Lobe Epilepsy; TLE: Temporal Lobe Epilepsy; PLE: Parietal Lobe Epilepsy.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Drug-resistant epilepsy, Epileptogenic zone localization, Large language models, Multi-source clinical text analysis, SEEG recommendation","lastPublishedDoi":"10.21203/rs.3.rs-6853779/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6853779/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLarge language models (LLMs) show promise for biomedical text analysis but are underused in pre-surgical epileptogenic zone (EZ) localization. This study evaluated LLMs' ability to analyze multi-source clinical text (medical records, EEG, MRI, FDG-PET, MEG) in 157 drug-resistant epilepsy patients (2020\u0026ndash;2024) at Beijing Tiantan Hospital. Three LLMs (GPT-4.1, Deepseek-R1, Claude 3.7 Sonnet) performed (1) EZ laterality classification, (2) probabilistic lobar localization (top 3 most likely lobes), and (3) SEEG recommendation scoring (0-100), benchmarked against multidisciplinary team (MDT)-defined resection sites. Modality ablation and stability analyses were conducted. Results showed GPT-4.1 and Deepseek-R1 achieved 98.1% laterality accuracy, vs. 97.5% for Claude 3.7 Sonnet (p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). GPT-4.1 and Claude 3.7 Sonnet had median lobar localization scores of 70 vs. 60 for Deepseek-R1 (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Textual information source ablation study revealed MRI reports and medical records were critical for GPT-4.1\u0026rsquo;s localization. GPT-4.1\u0026rsquo;s stability analysis using a two-way mixed-effects model showed excellent consistency for localization scores: single-measure ICC\u0026thinsp;=\u0026thinsp;0.960 (95% CI: 0.948\u0026ndash;0.969, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and average-measure ICC\u0026thinsp;=\u0026thinsp;0.986 (95% CI: 0.982\u0026ndash;0.990, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Additionally, we used GPT-4.1 for the SEEG recommendation task: its SEEG scores were significantly higher in patients undergoing SEEG (90.00 [IQR: 85.00\u0026ndash;90.00] vs. 25.00 [IQR: 10.00\u0026ndash;90.00] for non-SEEG, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). LLMs accurately inferred EZ laterality/lobar localization, aligning with MDT consensus, and their SEEG stratification potential may reduce invasive monitoring. Future research should focus on fine-tuning and multimodal fusion to optimize drug-resistant epilepsy outcomes.\u003c/p\u003e","manuscriptTitle":"Can Large Language Models Aid Pre-surgical Epileptogenic Zone Localization? A Multi-source Text Analysis Performance Study in Drug-Resistant Epilepsy","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-16 12:05:20","doi":"10.21203/rs.3.rs-6853779/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a235855a-2af1-468a-9c26-848f65d4c2a6","owner":[],"postedDate":"June 16th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-07-07T11:53:51+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-16 12:05:20","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6853779","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6853779","identity":"rs-6853779","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00