Beyond Traditional Prognostics: Integrating RAG-Enhanced AtlasGPT and ChatGPT 4.0 into Aneurysmal Subarachnoid Hemorrhage Outcome Prediction

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-14

This study compared AI language models AtlasGPT and ChatGPT 4.0 against traditional scales for predicting aneurysmal subarachnoid hemorrhage outcomes, finding AtlasGPT showed high accuracy for predicting decompressive hemicraniectomy.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

This retrospective cohort study of 82 patients with aneurysmal subarachnoid hemorrhage evaluated whether the large language models AtlasGPT and ChatGPT 4.0, when given standardized prompts and 17 baseline clinical/radiologic/laboratory parameters, could predict in-hospital mortality, favorable neurological outcome (mRS ≤ 2) at discharge and at 6 months, and the need for decompressive hemicraniectomy within a week, compared with the WFNS, Fisher, and SEBES clinical scales. AtlasGPT achieved the highest diagnostic accuracy for predicting decompressive hemicraniectomy (AUC 0.80), while WFNS performed best for long-term functional outcome prediction (AUC 0.76); for discharge outcomes, both WFNS and AtlasGPT showed similar prognostic performance (AUCs ~0.74–0.75). The authors used a small single-center sample and evaluated prompts answered as yes/no outputs in March 2024 rather than prospective testing, with outcomes derived from chart/ICU data and repeated question aggregation. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Background To assess the predictive accuracy of advanced AI language models and established clinical scales in prognosticating outcomes for patients with aneurysmal subarachnoid hemorrhage (aSAH). Methods This retrospective cohort study included 82 patients suffering from aSAH. We evaluated the predictive efficacy of AtlasGPT and ChatGPT 4.0 by examining the area under the curve (AUC), sensitivity, specificity, and Youden's Index, in comparison to established clinical grading scales such as the World Federation of Neurological Surgeons (WFNS) scale, Simplified Endovascular Brain Edema Score (SEBES), and Fisher scale. This assessment focused on four endpoints: in-hospital mortality, need for decompressive hemicraniectomy, and functional outcomes at discharge and after 6-month follow-up. Results In-hospital mortality occurred in 22% of the cohort, and 34.1% required decompressive hemicraniectomy during treatment. At hospital discharge, 28% of patients exhibited a favorable outcome (mRS ≤ 2), which improved to 46.9% at the 6-month follow-up. Prognostication utilizing the WFNS grading scale for 30-day in-hospital survival revealed an AUC of 0.72 with 59.4% sensitivity and 83.3% specificity. AtlasGPT provided the highest diagnostic accuracy (AUC 0.80, 95% CI: 0.70–0.91) for predicting the need for decompressive hemicraniectomy, with 82.1% sensitivity and 77.8% specificity. Similarly, for discharge outcomes, the WFNS score and AtlasGPT demonstrated high prognostic values with AUCs of 0.74 and 0.75, respectively. Long-term functional outcome predictions were best indicated by the WFNS scale, with an AUC of 0.76. Conclusions The study demonstrates the potential of integrating AI models such as AtlasGPT with clinical scales to enhance outcome prediction in aSAH patients. While established scales like WFNS remain reliable, AI language models show promise, particularly in predicting the necessity for surgical intervention and short-term functional outcomes.
Full text 100,505 characters · extracted from preprint-html · click to expand
Beyond Traditional Prognostics: Integrating RAG-Enhanced AtlasGPT and ChatGPT 4.0 into Aneurysmal Subarachnoid Hemorrhage Outcome Prediction | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Beyond Traditional Prognostics: Integrating RAG-Enhanced AtlasGPT and ChatGPT 4.0 into Aneurysmal Subarachnoid Hemorrhage Outcome Prediction Alim Emre Basaran, Agi Güresir, Hanna Knoch, Martin Vychopen, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4621973/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background To assess the predictive accuracy of advanced AI language models and established clinical scales in prognosticating outcomes for patients with aneurysmal subarachnoid hemorrhage (aSAH). Methods This retrospective cohort study included 82 patients suffering from aSAH. We evaluated the predictive efficacy of AtlasGPT and ChatGPT 4.0 by examining the area under the curve (AUC), sensitivity, specificity, and Youden's Index, in comparison to established clinical grading scales such as the World Federation of Neurological Surgeons (WFNS) scale, Simplified Endovascular Brain Edema Score (SEBES), and Fisher scale. This assessment focused on four endpoints: in-hospital mortality, need for decompressive hemicraniectomy, and functional outcomes at discharge and after 6-month follow-up. Results In-hospital mortality occurred in 22% of the cohort, and 34.1% required decompressive hemicraniectomy during treatment. At hospital discharge, 28% of patients exhibited a favorable outcome (mRS ≤ 2), which improved to 46.9% at the 6-month follow-up. Prognostication utilizing the WFNS grading scale for 30-day in-hospital survival revealed an AUC of 0.72 with 59.4% sensitivity and 83.3% specificity. AtlasGPT provided the highest diagnostic accuracy (AUC 0.80, 95% CI: 0.70–0.91) for predicting the need for decompressive hemicraniectomy, with 82.1% sensitivity and 77.8% specificity. Similarly, for discharge outcomes, the WFNS score and AtlasGPT demonstrated high prognostic values with AUCs of 0.74 and 0.75, respectively. Long-term functional outcome predictions were best indicated by the WFNS scale, with an AUC of 0.76. Conclusions The study demonstrates the potential of integrating AI models such as AtlasGPT with clinical scales to enhance outcome prediction in aSAH patients. While established scales like WFNS remain reliable, AI language models show promise, particularly in predicting the necessity for surgical intervention and short-term functional outcomes. Aneurysmal subarachnoid hemorrhage Artificial intelligence AtlasGPT ChatGPT neurologic outcomes prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Description The study explored the use of advanced AI language models, AtlasGPT and ChatGPT 4.0, to predict outcomes for patients with aneurysmal subarachnoid hemorrhage (aSAH). It found that AtlasGPT provided the highest diagnostic accuracy for predicting the need for decompressive hemicraniectomy, outperforming traditional clinical scales, while both AI models showed promise in enhancing outcome predictions when integrated with established clinical assessment tools. Introduction Aneurysmal subarachnoid hemorrhage (aSAH) is a severe medical condition with high pre- and in-hospital mortality 1 – 3 . Clinical outcome depends mainly on initial severity of aSAH 4 – 6 . In patients with intractably elevated intracranial pressure (ICP), decompressive craniectomy (DC) may be indicated as a life-saving procedure 7 . Predicting the prognosis of aSAH patients is a challenging task. Machine learning algorithms have the potential to make prognostic predictions based on individual patient parameters 8 , 9 . A pilot study has shown that large language models or Generative Pre-trained Transformers (GPT) such as ChatGPT have the potential to predict short-term functional outcomes in acute ischemic stroke patients after thrombectomy more accurately than the existing risk scores 10 . However, it should be noted that the large language models used in the studies have not yet been explicitly developed for medical purposes. Atlas GPT was specifically developed and trained for neurosurgical questions, enabling more precise answers to complex questions and potential improved prognostic performance 11 . To our knowledge, no study has yet investigated the use of large language models in relation to aSAH. However, based on the few studies conducted so far in other medical fields, there is promising potential. Against this backdrop, we compared the prognostic accuracy of the large language models ChatGPT 4.0 and AtlasGPT based on individual aSAH patient parameters for predicting in-hospital mortality, neurological outcome, and the need for DC. Methods Study setting & participants A retrospective analysis was conducted on patients diagnosed with aSAH at the Department of Neurosurgery at Leipzig University Hospital. The study included patients treated according to corresponding guidelines between 2019 and 2022. Ethics: The retrospective study was conducted in compliance with the Declaration of Helsinki and its amendments and was approved by the local medical ethics committee (387/23). Data collection: Retrospective data was collected from patient records and ICU reports for the study. The following clinical, radiological, and laboratory data were recorded: Baseline demographic patient characteristics (Age, sex), Word Federation of Neurological Surgeons (WFNS) Score, Glasgow Coma Scale (GCS) score, pupillary reflex, neuroradiological parameters (Fisher score, intracerebral hemorrhage, midline shift, location of ruptured aneurysm, presence of initial hydrocephalus, Subarachnoid Hemorrhage Early Brain Edema Score (SEBES)) surgical treatment modalities (clipping, endovascular coiling), and laboratory values (hemoglobin, CRP, blood lactate, serum creatinine) 12 , 13 . Outcomes The study's primary outcome was in-hospital survival. The secondary outcome was favorable neurological outcome, which was based on the modified Rankin Scale (mRS) and defined as a mRS score of 0–2 14 . Tertiary outcome parameter was the need for DC. Development of a chat prompt for Atlas GPT and ChatGPT 4.0 In the creation of a standardized dialogue prompt, our methodology adopted an iterative process, as recommended by Kanjee et al 15 . An initial text was composed and subsequently refined through a process of trial and error, with the aim of eliciting specific responses from ChatGPT-4 and AtlasGPT. This initial text comprehensively described the task and context to the language model. The full standardized dialogue prompt is accessible (see supplementary file 1 ). In summary, the language model was instructed to assume the role of an “AI intensive care physician” or an “AI neurosurgeon” tasked with managing a patient who was admitted on hospital for aSAH and aneurysms were treated via Coiling or Clipping within 24 hours after the aSAH event. Additionally, the model was provided with seventeen baseline patient-, disease-, and procedure-specific parameters, chosen for their recognized value in predicting outcomes after aSAH. This standardized anonymized way was chosen because the unstructured upload of medical records to a cloud-based language model would result in significant data privacy concerns. The text was developed in a standardized way until AtlasGPT (Atlasmeditech LLC. 2024) and ChatGPT 4.0 (OpenAI, Inc., San Francisco, USA) could answer questions with a simple 'yes' or 'no'. The following baseline factors of aSAH at admission were considered in the analysis: WFNS grade, pupillary reflex, age, gender, Glasgow coma scale, Fisher scale, intracerebral hemorrhage, intraventricular hemorrhage, midline-shift, location of ruptured aneurysm, treatment modality, medical treatment with antiplatelet drugs after endovascular treatment, initial hydrocephalus treatment such as external ventricular drain, hemoglobin, CRP, blood lactate, and serum creatinine. AtlasGPT and ChatGPT 4.0 were asked four questions that required ‘yes’ or ‘no’ answers after entering the text and parameters (see supplementary file 1 ). 1)Will this patient survive to hospital discharge? Please provide a yes/no answer. 2)Will this patient experience a good neurological outcome at hospital discharge as defined by the modified ranking scale (0–2). Please provide a yes/no answer. 3) Will this patient experience a good neurological outcome at 6-months after aneurysmal subarachnoid hemorrhage e as defined by the modified ranking scale (0–2). Please provide a yes/no answer. “ 4) Will the patient have to be treated by a decompressive craniectomy within the next week? Please provide a yes/no answer. The responses were recorded in an Excel file. Each question was asked three times using both AtlasGPT and ChatGPT 4.0. A new chat was opened for each question. In the case of dichotomous answers (yes/no), the most frequent response from the repeated questions was considered. For example, if the responses were yes/yes/no, the overall response was recorded as ‘yes’. In cases where ‘no’ was the more frequent response, we selected ‘no’ as the final answer. If we did not receive a response to a question, we reopened the chat and repeated the question until we received a complete and unambiguous answer. We used only complete and unambiguous answers for statistical analysis. The chat prompts were performed between the 5th and 30th of March in 2024. Statistical analysis: Receiver operating characteristics curves (ROC) curves and area under curves (AUC) values were calculated to determine the prognostic significance of AtlasGPT, ChatGPT 4.0, WFNS grading, Fisher scale, and SEBES. Median values and interquartile ranges (IQR) of metric data are reported. The prognostic significance of those parameters were determined by sensitivity, specificity, cut-off values, positive/negative predictive value, and the Youden’s index. Statistical analyses was performed using SPSS Statistics version 29.0.2.0 (IBM, Armonk, New York). Statistical graphics were created using SPSS version 29.02.0 and were modified with corresponding tables using BioRender.com. Results Study Cohort Between January 2019, and December 2022, 120 spontaneous aSAH patients were admitted to the present institution. After the exclusion of patients with an withdraw or withhold from life sustaining therapies, patients with an angiogram-negative SAH, and 7 patients who were lost to follow-up, 82 consecutive aSAH patients were included in the present investigation. Figure 1 summarizes the process flow chart. Patient characteristics The cohort comprised predominantly middle-aged individuals with a median age of 56 years (IQR 44.8–67.3), with a slight female predominance (49/82; 59.8%). The majority of patients presented with poor-grade subarachnoid hemorrhage, reflected by WFNS gradings IV and IV in 50 patients (50/82; 61.0%). A considerable number exhibited signs of brain edema on imaging as evidenced by SEBES scores, with 19 patients (23.2%) being assigned the highest score (IV). Baseline vigilance was generally compromised, indicated by a median Glasgow Coma Scale score of 7 (IQR 3.0–14.0). Intracranial pathology was confirmed by the presence of midline shift in 24 patients (29.3%). The aneurysms were most frequently located at the anterior cerebral artery (ACA) complex (36/82; 43.9%) and the middle cerebral artery (MCA) complex (25/82; 30.5%). Most patients underwent endovascular treatment for their aneurysms (54/82; 65.9%). Further patient- and disease-specific characteristics are summarized in Table 1 . Table 1 Patient characteristics Characteristic Frequency ( n = 82) Age, Median (IQR) 56.0 (44.8–67.3) Sex Female Male 49 (59.8%) 33 (40.2%) WFNS grading 1 2 3 4 5 18 (22.0%) 9 (11.0%) 5 (6.1%) 9 (11.0%) 41 (50.0%) Fisher scale 1 2 3 4 1 (1.2%) 3 (3.7%) 62 (75.6%) 16 (19.5%) SEBES 0 1 2 3 4 16 (19.5%) 12 (14.6%) 20 (24.4%) 15 (18.3%) 19 (23.2%) GCS, Median (IQR) 7.0 (3.0–14.0) Midline-shift Present 24 (29.3%) Aneurysm location Anterior complex ACA Pericallosal artery AcoA ICA Bifurcation Ophtalmic artery AChoA PcoA MCA M1 Bifurcation M2 Posterior circulation Basilar artery Vertebral artery PICA PCA 1 (1.2%) 5 (6.1%) 30 (36.6%) 4 (4.9%) 1 (1.2%) 1 (1.2%) 3 (3.7%) 5 (6.1%) 17 (20.7%) 3 (3.7%) 6 (7.3%) 2 (2.4%) 3 (3.7%) 1 (1.2%) Arterial hypertension 54 (65.9%) Antiplatelet therapy 7 (8.5%) Anticoagulation intake 3 (3.7%) Baseline hydrocephalus (EVD placement) 45 (54.9%) Type of aneurysm treatment Clipping Endovascular 28 (34.1%) 54 (65.9%) Lactate (mmol/I), Median (IQR) 1.6 (1.0-2.5) C-reactive protein (mg/I), Median (IQR) 2.3 (1.1–4.5) Serum creatinine (µmol/I), Median (IQR) 68.0 (57.0-85.5) Hemoglobin (mmol/I), Median (IQR) 8.2 (7.4–8.7) Abbrevations: IQR = Interquartile range; SEBES = subarachnoid hemorrhage early brain edema score; EVD = external ventricular drain; ACA = anterior cerebral artery; AcoA = anterior communicating artery; ICA = internal carotid artery; AChoA = anterior choroidal artery; MCA = middle cerebral artery; PICA = posterior inferior cerebellar artery; PCA = posterior cerebral artery Outcome parameters During the hospitalization period, 22% of patients ( n = 18) expired. DC was performed in 34.1% ( n = 28) during hospital therapy. Primary decompressive hemicraniectomy was performed in 15 cases during surgical clipping (15/28; 53.6%). The residual decompressive hemicraniectomies were performed secondary due to intracerebral pressure elevations after surgery because of subdural-, epidural- and intracerebral hematomas (4/28; 14.3%), brain edema without signs of infarction (4/28; 14.3%), and radiologically determined infarction (5/28; 17.9%). At the time of hospital discharge, a favorable prognosis, denoted by a modified Rankin Scale score of ≤ 2, was observed in 28% of patients (n = 23). At the 6-month follow-up, an improvement in outcomes was noted; 46.9% of the survivors (n = 30) achieved a favorable functional status. Supplementary table S1 summarizes the outcome parameters, which were requested from the artificial intelligence language models. Prognostication In-hospital Survival For the endpoint of 30-day in-hospital survival, the World Federation of Neurological Surgeons (WFNS) grading scale manifested the superior discriminative capability with an AUC of 0.72 (95% CI: 0.60–0.84), coupled with a sensitivity of 59.4% and specificity of 83.3% with a cut off set at ≤ 4/>4. The positive predictive value and negative predictive values were 92.7% and 36.6%, respectively (see supplementary table S2 ). The WFNS scale's Youden's Index was 0.43, indicating a robust balance between sensitivity and specificity in this acute prognostic scenario at admission. Although AtlasGPT had a slightly lower AUC, it demonstrated a comparable performance in terms of specificity (83.3%) when responding that the patient will survive for at least two times of three chat prompt runs. AtlasGPT and ChatGPT 4.0 showed similar capabilities with AUCs of 0.70 (95% CI: 0.57–0.83) and 0.67 (95% CI: 0.53–0.80) respectively, while SEBES and FISHER scales demonstrated lower discriminative power with AUCs of 0.53 (95% CI: 0.38–0.68) and 0.54 (95% CI: 0.40–0.69). Figure 2 illustrates the findings. Need for decompressive hemicraniectomy during hospital therapy AtlasGPT presented the most pronounced diagnostic accuracy in anticipating the requirement for decompressive hemicraniectomy upon admission with an AUC of 0.80 (95% CI: 0.70–0.91). The model's sensitivity reached 82.1%, with specificity at 77.8%, and a Youden's Index of 0.6, when prognosticating the need for decompressive hemicraniectomy in three dialog runs in a row. The positive predictive value and negative predictive values of AtlasGPT indicating the need for decompressive hemicraniectomy were 65.7% and 89.4%, respectively (see supplementary table 2 ). AtlasGPT´s predictive power was followed by ChatGPT 4.0, which had an AUC of 0.78 (95% CI: 0.68–0.88), and the WFNS score with an AUC of 0.76 (95% CI: 0.66–0.86). SEBES and FISHER scales had the lowest AUCs, indicating less predictive value for this outcome. Figure 3 summarizes the results of this analysis. Favorable outcome (mRS ≤ 2) at discharge The WFNS score and AtlasGPT both provided high prognostic value for predicting a good functional outcome at discharge, with AUCs of 0.74 (95% CI: 0.61–0.87) and 0.75 (95% CI: 0.62–0.87), respectively. The sensitivity and specificity of AtlasGPT were 69.6% and 79.7%, respectively, with an optimal cutoff at ≥ 2/<2 yes and a Youden's Index of 0.49, suggesting good prognostic performance at the point of discharge. The positive predictive value and negative predictive values of AtlasGPT prognosticating favorable mRS at discharge were 57.1% and 87.0%, respectively (see supplementary table 1 ). ChatGPT 4.0 also performed adequately with an AUC of 0.72 (95% CI: 0.59–0.85). The SEBES and FISHER scales were less predictive to assess clinical outcome, as reflected by their AUCs of 0.53 (95% CI: 0.38–0.67) and 0.62 (95% CI: 0.48–0.76). Figure 4 displays the results of the analyses regarding functional outcome at discharge. Favorable outcome (mRS ≤ 2) at 6-months For the prediction of a favorable functional outcome at a 6-month interval, the WFNS grading scale was pre-eminent, delivering an AUC of 0.76 (95% CI: 0.64–0.88), with a sensitivity of 76.7% and specificity of 71.9%. The calculated Youden's Index was 0.49, indicating a balanced predictive performance for long-term recovery assessments. The positive predictive value and negative predictive values of WFNS grading to prognosticate favorable mRS outcome at 6-months were 71.9% and 76.7%, respectively (see supplementary table 2 ). AtlasGPT and ChatGPT 4.0 provided both comparable AUCs of 0.69 (95% CI: 0.56–0.83). SEBES and FISHER scores were less effective in predicting long-term outcomes, with AUCs of 0.58 (95% CI: 0.43–0.73) and 0.58 (95% CI: 0.44–0.72), respectively. Supplementary file 2 outlines the results. Response variability of language models concerning the prediction of endpoints During each of the three trials involving AtlasGPT and ChatGPT 4.0, we encountered occurrences of divergent outputs, classified as response variability in response to the prompts given. AtlasGPT delivered different responses for the endpoints In-hospital mortality, need for decompressive hemicraniectomy, favorable outcomes at discharge, and favorable outcomes at 6-months after aSAH in 5 (6.1%), 7 (8.5%), 3 (3.7%), and 3 (3.7%) runs, respectively. ChatGPT 4.0 likewise generated diverse outcomes for in-hospital mortality, favorable discharge outcomes, and favorable 6-month outcomes post-aSAH, each occurring in 1 (1.2%) instance. No divergent responses for the necessity of decompressive hemicraniectomy was observed using ChatGPT 4.0. Discussion The present study compared the prognostic value of the retrieval-augmented generation techniques enhanced large language model AtlasGPT based on peer-reviewed reliable neurosurgery-specific evidence, the large language model ChatGPT-4.0 with well-validated clinical and imaging scales. The prognostic performance of AtlasGPT to predict the functional outcome at discharge and the need for decompressive hemicraniectomy in aSAH based on characteristics at hospital admission demonstrated substantial roles for artificial intelligence in clinical practice (see supplementary file 3 ). Nevertheless, these results necessitate extensive discussion. Both language models showed responses variabilities with slightly more variabilities in the AtlasGPT language model. However, we ran three iterations of four binary questions for 82 patients in two language models and observed only 21 instances of response variability (21/656; 3.2%). To address this issue seriously and reduce potential bias caused response variability, we ultimately took the mean value of the responses for statistical processing. Nevertheless, it has to be noted that our investigation is focused on the most advanced ChatGPT version currently available, which requires a paid subscription (namely, the GPT-4-based model). The freely accessible and more commonly used version is based on the 3.5 model, which might produce less reliable predictions with more response variability and more temporal instability 16 . This demonstrates that artificial intelligence can be most effective when integrated with “human intelligence”, such as the expertise of a supervision by a clinician. Additionally, it highlights the necessity for users to closely monitor the application of large language models in clinical settings. Decisions to treatment limitations and prognosis are inherently challenging, demanding human attributes such as long-term professional expertise, empathy, and emotional insight. In contrast, large language models are purely machines operating on stochastic processes, lacking any form of consciousness or emotional capacity 17 . While the use of large language models in the medical field is growing, neurovascular research on their ability to predict patient outcomes remains limited. A retrospective analysis of clinical, neuroimaging, and procedure-related data involving 163 patients with acute ischemic stroke found that ChatGPT adequately predicted short-term functional mRS outcomes at 3-months after mechanical thrombectomy and was superior to existing risk scores 18 . Mortality and functional prognosis in aSAH is complex, with scales like Hunt and Hess and WFNS gauging risk based on initial assessments and traits but ignoring the effects of ending life support, which might bias outcomes 19 . Furthermore, there´s growing concern that pessimistic views on severely injured patient´s recoveries might worsen outcomes, termed a “self-fulfilling prophecy” 20 , 21 . However, it is of paramount importance to scientifically validate these language models because patients and their relatives will potentially start putting trust in what conventional AI, like ChatGPT, predicts. Despite language model designers took steps to prevent these consultations by making sure that language models are not physicians and cannot give medical advices, such protocols can be bypassed by “jailbreak“ prompts, such as feigned scenarios as demonstrated in the present study 22 . Physicians evaluating large language models should be also mindful of the “stochastic parrot” concept 23 , 24 . This principle highlights that due to its algorithmic nature, large language models lack comprehension of both the information it receives and the produced responses, merely replicating learned patterns and biases, including stereotypes and social disparities 25 . These mechanisms might account for why the large language model´s performances in survival prediction and long-term neurological outcome prediction is comparable and not superior established clinical grading systems. This observation is consistent with findings from other research indicating similar levels of effectiveness in clinical or theoretical settings 26 , 27 . However, we found a superior role of a generated dialog describing baseline patient-, disease-, and procedure-specific aSAH patient characteristics which might facilitate the identification of those patients who will potentially undergo decompressive hemicraniectomy. Brain swelling and elevated ICP are established factors worsening outcomes after aSAH 28 . Despite decompressive hemicraniectomy is a well-established procedure with high-level evidence for space-occupying stroke 29 , the currently used language models and also AtlasGPT still have to rely on retrospective data. Further insights on the role of primary decompressive hemicraniectomy in aSAH and which patient will potentially benefit from this procedure might inform the randomized controlled PICASSO trial investigating decompressive hemicraniectomy in poor-grade (WFNS IV&V) aSAH patients 30 . Limitations To our knowledge, this is the first investigation using large language models for predicting outcomes following aSAH with real-world patient data. It adopts a practical method focused on reproducibility and the integrity of data. Nonetheless, this study is not without its drawbacks. First, there is the tendency of large language models to produce response variability which have to be addressed by repetitive approaches and trusting the mean values. Additionally, ChatGPT 4.0 was not specifically developed for medical use compared to AtlasGPT. Hence, the effectiveness and accuracy of ChatGPT 4.0 in addressing clinical issues has yet to be established. Finally, the present study represents a single-center cohort, which may limit their applicability to other settings or regions, underscoring the need for further studies across various environments to confirm the universality of these results. Finally, the present results have not been compared to a blinded experienced senior physician estimating the outcomes of the individual patients with the same baseline data. Conclusions AtlasGPT demonstrated a good performance in predicting functional outcome at discharge and the need for decompressive hemicraniectomy and thus may be a helpful tool for early endpoints in risk predictions of aSAH patients. Nevertheless, large language models still need human supervision by professionals. Future research should focus on specific trained large language models and further validate various clinical endpoints. Declarations The Ethics Committee at the Medical Faculty of the University of Leipzig has no ethical or scientific objections to the study (ID: 387/23). The manuscript has been prepared in accordance with the instructions provided by the authors and has been approved by all of them. All of the requirements set out by the authors have been met and approved by all of them.The authors have no conflicts of interest to declare. All co-authors have seen and agree with the content of the manuscript and there is no financial interest to report. We certify that the submission is original work and is not under review at any other publication. Author Contribution Conceptualization: Alim Emre Basaran, Johannes Wach. Investigation: Johannes Wach, Alim Emre Basaran. Visualization: Johannes Wach, Alim Emre Basaran. Supervision: Erdem Güresir, Johannes Wach. Writing-original draft: Alim Emre Basaran, Johannes Wach. Writing-review & editing: Alim Emre Basaran, Johannes Wach, Agi Güresir, Hanna Knoch, Martin Vychopen, Erdem Güresir, Johannes Wach. Funding: No funding was received to assist with the preparation of this manuscript. Acknowledgment: The visual abstract was created using BioRender Data Availability Statement: Data available on request from the authors Institutional Review Board Statement The study was conducted in accordance with the Declaration of Helsinki. References Etminan N, Chang HS, Hackenberg K et al (2019) Worldwide Incidence of Aneurysmal Subarachnoid Hemorrhage According to Region, Time Period, Blood Pressure, and Smoking Prevalence in the Population: A Systematic Review and Meta-analysis. JAMA Neurol 76(5):588–597. 10.1001/jamaneurol.2019.0006 Korja M, Lehto H, Juvela S, Kaprio J (2016) Incidence of subarachnoid hemorrhage is decreasing together with decreasing smoking rates. Neurology. ;87(11):1118-23. 10.1212/WNL.0000000000003091.Epub 2016 Aug 12 SVIN COVID-19 Global SAH Registry Global impact of the COVID-19 pandemic on subarachnoid haemorrhage hospitalisations, aneurysm treatment and in-hospital mortality: 1-year follow-up. J Neurol Neurosurg Psychiatry 2022 Jul 28:jnnp–2022. 10.1136/jnnp-2022-329200 Roquer J, Cuadrado-Godia E, Guimaraens L (2020) Short- and long-term outcome of patients with aneurysmal subarachnoid hemorrhage. Neurology 95(13):e1819–e1829. 10.1212/WNL.0000000000010618 Hammer A, Steiner A, Ranaie G et al (2018) Impact of Comorbidities and Smoking on the Outcome in Aneurysmal Subarachnoid Hemorrhage. Sci Rep. 8: 12335. Published online 2018 Aug 17 10.1038/s41598-018-30878-9 Report of World Federation of Neurological Surgeons Committee on a Universal Subarachnoid Hemorrhage Grading Scale (1988) J Neurosurg 68(6):985–986. 10.3171/jns.1988.68.6.0985 Güresir E, Raabe A, Setzer M, Vatter H, Gerlach R, Seifert V, Beck J (2009) Decompressive hemicraniectomy in subarachnoid haemorrhage: the influence of infarction, haemorrhage and brain swelling. J Neurol Neurosurg Psychiatry 80(7):799–801. 10.1136/jnnp.2008.155630 Johnsson J, Björnsson O, Andersson P et al (2020) Artificial neural networks improve early outcome prediction and risk classification in out-of-hospital cardiac arrest patients admitted to intensive care. Crit Care 24(1):474. 10.1186/s13054-020-03103 Chung CC, Chiu WT, Huang YH et al Identifying prognostic factors and developing accurate outcome predictions for in-hospital cardiac arrest by using artificial neural networks. J Neurol Sci 2021 Jun 15:425:117445. 10.1016/j.jns.2021.117445 . Epub 2021 Apr 18. Pedro T, Sousa JM, Ronseca L et al Exploring the use of ChatGPT in predicting anterior circulation stroke functional outcomes after mechanical thrombectomy: a pilot study. J Neurointerv Surg 2024 Mar 7:jnis–2024. 10.1136/jnis-2024-021556 Hopkins SB, Carter B, Lord J et al Editorial. AtlasGPT: dawn of a new era in neurosurgery for intelligent care augmentation, operative planning, and performance. J Neurosurg 2024 Feb 27:1–410.3171/2024.2.JNS232997 Ahn SH, Savarraj JP, Pervez M, Jones W, Park J, Jeon SB, Kwon SU, Chang TR, Lee K, Kim DH, Day AL, Choi HA (2018) The Subarachnoid Hemorrhage Early Brain Edema Score Predicts Delayed Cerebral Ischemia and Clinical Outcomes. Neurosurgery 83(1):137–145. 10.1093/neuros/nyx364 Fisher CM, Kistler JP, Davis JM (1980) Relation of cerebral vasospasm to subarachnoid hemorrhage visualized by computerized tomographic scanning. Neurosurgery 6(1):1–9. 10.1227/00006123-198001000-00001 Weisscher N, Vermeulen M, Roos YB, de Haan RJ (2008) What should be defined as good outcome in stroke trials; a modified Rankin score of 0–1 or 0–2? Neurol 255(6):867–874. 10.1007/s00415-008-0796-8 Epub 2008 Mar 14 Kanjee Z, Crowe B, Rodman A (2023) Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330(1):78–80. 10.1001/jama.2023.8288 Chen L, Zaharia M, Zou J How Is ChatGPT´s behavior changing over time? arXiv:2307.09009v3. https://doi.org/10.48550/arXiv.2307.09009 Khan B, Fatima H, Qureshi A et al Drawbacks of Artificial Intelligence and Their Potential Solutions in the Healthcare sector. Biomed Mater Devices 2023 Feb 8: 1–8. 10.1007/s44174-023-00063-2 Pedro T, Sousa JM, Fonseca L et al Exploring the use of ChatGPT in predicting anterior circulation stroke functional outcomes after mechanical thrombectomy: a pilot stuy. J Neurointerv Surg 2024 Mar 7:jnis–2024. 10.1136/jnis-2024-021556 Van Heuven AW, Mees SMD, Algra A, Rinkel GJE (2008) Validation of a prognostic subarachnoid hemorrhage grading scale derived directly from the Glasgow Coma Scale. Stroke 39(4):1347–1348. 10.1161/STROKEAHA.107.498345 Epub 2008 Feb 28 Becker KJ, Baxter AB, Cohen WA et al (2001) Withdrawal of support in intracerebral hemorrhage may lead to self-fulfilling prophecies. Neurology 56(6):766–772. 10.1212/wnl.56.6.766 Hemphill JC 3rd, White DB (2009) Clinical nihilism in neuroemergencies. Emerg Med Clin North Am. ;27(1):27–37, vii-viii. 10.1016/j.emc.2008.08.009 Liu Y, Deng G, Zhengzi, Xu et al Jailbreaking ChatGPT via Prompt Engineering: An empirical Study. arXiv:2305.13860v2. Bender EM, McMillan-Major A, Gebru T, Shmitchell S On the On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Conference on Fairness, Accountability, and Transparency (FAccT ´21), March 3–10, 2021, Virtual Event, Canada. ACM, New York, NY, USA, 14 pages. http://doi.org/10.1145/3442188.3445922 Boussen S, Denis JC, Simeone P et al (2023) ChatGPT and the stochastic parrot: artificial intelligence in medical research. Br J Anaesth 131(4):e120–e121. 10.1016/j.bja.2023.06.065 Epub 2023 Jul 27 Deshpande A, Murahari V, Rajpurohit T et al Toxicity in ChatGPT: Analyzing Persona-assigned Language Models. arXiv:2304.05335V1 Kanjee Z, Crowe B, Rodman A (2023) Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA 330:78–80. 10.1001/jama.2023.8288 Gilson A, Safranek CW, Huang T et al (2023) How does ChatGPT perform on the united states medical licensing examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 9:e45312. 10.2196/45312 Heuer GG, Smith MJ, Elliott JP et al (2004) Relationship between intracranial pressure and other clinical variables in patients with aneurysmal subarachnoid hemorrhage. Neurosurg 101(3):408–416. 10.3171/jns.2004.101.3.0408 Vahedi K, Hofmeijer J, Juettler E et al (2007) Early decompressive surgery in malignant infarction of the middle cerebral artery: a pooled analysis of three randomized controlled trials. Lancet Neurol 6(3):215–222. 10.1016/S1474-4422(07)70036-4 Güresir E, Lampmann T, Brandecker S et al (2022) PrImary decompressive Craniectomy in AneurySmal Subarachnoid hemOrrhage (PICASSO) trial: study protocol for a randomized controlled trial. Trials 23(1):1027. 10.1186/s13063-022-06969-4 Additional Declarations No competing interests reported. Supplementary Files Supplementaryfile1.docx Supplementaryfile2asah.png supplementaryfile3asah.png Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4621973","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":329559434,"identity":"a3b75fa0-9007-41d2-9f30-b4ac0741a437","order_by":0,"name":"Alim Emre Basaran","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABVklEQVRIie3RsUrDQBjA8e84yC0HWU9q21eIFKqC5l7ljkCnlnbqImileJ3UNeDgK9Tl6JhyQ5fQuYNDoZA5IghCQRNTalJRcBPJnwyXu/y4DwJQVvYXI2gAyZMVAyC1eaVAAAIAlqxFgeAcQX6B4O9I+umGYPq53hxlFYg9xFdjNHG5fWei1Wn7pHpjG7VE6nEfMJ6a3vrIHRAvyBFmkFqg0JP+vHXY6OhWQ8F05CAVJYNZwviUeQMaFa75IAoLCKFZ6WgjFbpUDClDuaGOoYx5wNpOTtQzcsHrIXmpHOs3qTDKCGA7TlRKunGOOBkxaBzSZgXpQCprSygYKpib3JKf6yAlMpzJh5D2966111A0IWKeEssxfsCERaP8YLXZKFo8Tc54LSSavWq3en87i1jcNxzIcPXcW59zm3hL2Ensbmx30l8j1Zfjn0oJ/5UoKysr+4e9A0l/dr8JjEaiAAAAAElFTkSuQmCC","orcid":"","institution":"University hospital Leipzig","correspondingAuthor":true,"prefix":"","firstName":"Alim","middleName":"Emre","lastName":"Basaran","suffix":""},{"id":329559435,"identity":"e855c9ef-8f73-40d4-873c-836cfdc48e6c","order_by":1,"name":"Agi Güresir","email":"","orcid":"","institution":"University hospital Leipzig","correspondingAuthor":false,"prefix":"","firstName":"Agi","middleName":"","lastName":"Güresir","suffix":""},{"id":329559436,"identity":"3a812d05-eb27-4cfd-b3ea-a0203395477e","order_by":2,"name":"Hanna Knoch","email":"","orcid":"","institution":"University hospital Leipzig","correspondingAuthor":false,"prefix":"","firstName":"Hanna","middleName":"","lastName":"Knoch","suffix":""},{"id":329559438,"identity":"1c8861cd-7b49-4b5b-b60c-53da97bba34c","order_by":3,"name":"Martin Vychopen","email":"","orcid":"","institution":"University hospital Leipzig","correspondingAuthor":false,"prefix":"","firstName":"Martin","middleName":"","lastName":"Vychopen","suffix":""},{"id":329559440,"identity":"c3fef9e6-d724-4417-86e5-cd1f727e6a69","order_by":4,"name":"Erdem Güresir","email":"","orcid":"","institution":"University hospital Leipzig","correspondingAuthor":false,"prefix":"","firstName":"Erdem","middleName":"","lastName":"Güresir","suffix":""},{"id":329559441,"identity":"5832242f-1c8b-4d8f-96c6-d074d581b93a","order_by":5,"name":"Johannes Wach","email":"","orcid":"","institution":"University hospital Leipzig","correspondingAuthor":false,"prefix":"","firstName":"Johannes","middleName":"","lastName":"Wach","suffix":""}],"badges":[],"createdAt":"2024-06-22 12:53:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4621973/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4621973/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":60811323,"identity":"488c6089-cd0d-4824-9d66-6ca42f8c0ba3","added_by":"auto","created_at":"2024-07-22 10:54:33","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":208145,"visible":true,"origin":"","legend":"\u003cp\u003eStudy flowchart over 48 months with a total study cohort of 82 consecutive patients\u003c/p\u003e","description":"","filename":"OnlineFigure1asah.png","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/986762774c6dfdf1d4152c33.png"},{"id":60811327,"identity":"eef2f2c4-ee0f-42b5-9b34-b44c4fe952b4","added_by":"auto","created_at":"2024-07-22 10:54:34","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":497529,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of ROC curves for in-hospital survival. Abbreviations include AUC: Area under the Curve, CI: Confidence Interval, GPT: Generative Pre-Trained Transformer, mRS: Modified Rankin Scale\u003c/p\u003e","description":"","filename":"OnlineFigure2asah.png","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/b8cfa627da2f59c6d479ac31.png"},{"id":60811328,"identity":"18f9f447-640c-40bb-83bc-bfdd0f77f088","added_by":"auto","created_at":"2024-07-22 10:54:34","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":559942,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of ROC curves for need for decompressive hemicraniectomy. Abbreviations include AUC: Area under the Curve, CI: Confidence Interval, GPT: Generative Pre-Trained Transformer, mRS: Modified Rankin Scale\u003c/p\u003e","description":"","filename":"OnlineFigure3asah.png","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/4fae131781aeac901e9864ea.png"},{"id":60811329,"identity":"48a1d140-843b-4993-8290-aa659d50134b","added_by":"auto","created_at":"2024-07-22 10:54:34","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":523188,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of ROC curves for favorable functional outcome (mRS≤2) at discharge. Abbreviations include AUC: Area under the Curve, CI: Confidence Interval, GPT: Generative Pre-Trained Transformer, mRS: Modified Rankin Scale\u003c/p\u003e","description":"","filename":"OnlineFigure4asah.png","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/fafdb4072cfa79d5c9e22b93.png"},{"id":63092392,"identity":"928ca035-a4bd-401a-af28-980e72a183b8","added_by":"auto","created_at":"2024-08-23 04:41:46","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3114240,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/55956b20-acef-4f84-86d8-aaa8126d14c8.pdf"},{"id":60811324,"identity":"cf1b4cc9-8fd2-4893-aa51-ae9cd25afb3f","added_by":"auto","created_at":"2024-07-22 10:54:33","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":14818,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfile1.docx","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/cdf6bfb7cc3fcea4dd275952.docx"},{"id":60811325,"identity":"8950dfd8-f840-48f5-aac5-f0ef5768f2c1","added_by":"auto","created_at":"2024-07-22 10:54:33","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":3737269,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfile2asah.png","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/f34f94655c678de929c23d65.png"},{"id":60811326,"identity":"d59fb95e-6383-4547-b53b-26d821870188","added_by":"auto","created_at":"2024-07-22 10:54:33","extension":"png","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":3212671,"visible":true,"origin":"","legend":"","description":"","filename":"supplementaryfile3asah.png","url":"https://assets-eu.researchsquare.com/files/rs-4621973/v1/44eb55783e0a3529db8895ef.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Beyond Traditional Prognostics: Integrating RAG-Enhanced AtlasGPT and ChatGPT 4.0 into Aneurysmal Subarachnoid Hemorrhage Outcome Prediction","fulltext":[{"header":"Description","content":"\u003cp\u003eThe study explored the use of advanced AI language models, AtlasGPT and ChatGPT 4.0, to predict outcomes for patients with aneurysmal subarachnoid hemorrhage (aSAH). It found that AtlasGPT provided the highest diagnostic accuracy for predicting the need for decompressive hemicraniectomy, outperforming traditional clinical scales, while both AI models showed promise in enhancing outcome predictions when integrated with established clinical assessment tools.\u003c/p\u003e"},{"header":"Introduction","content":"\u003cp\u003eAneurysmal subarachnoid hemorrhage (aSAH) is a severe medical condition with high pre- and in-hospital mortality \u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Clinical outcome depends mainly on initial severity of aSAH \u003csup\u003e\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. In patients with intractably elevated intracranial pressure (ICP), decompressive craniectomy (DC) may be indicated as a life-saving procedure \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003ePredicting the prognosis of aSAH patients is a challenging task. Machine learning algorithms have the potential to make prognostic predictions based on individual patient parameters \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. A pilot study has shown that large language models or Generative Pre-trained Transformers (GPT) such as ChatGPT have the potential to predict short-term functional outcomes in acute ischemic stroke patients after thrombectomy more accurately than the existing risk scores \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. However, it should be noted that the large language models used in the studies have not yet been explicitly developed for medical purposes. Atlas GPT was specifically developed and trained for neurosurgical questions, enabling more precise answers to complex questions and potential improved prognostic performance \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTo our knowledge, no study has yet investigated the use of large language models in relation to aSAH. However, based on the few studies conducted so far in other medical fields, there is promising potential. Against this backdrop, we compared the prognostic accuracy of the large language models ChatGPT 4.0 and AtlasGPT based on individual aSAH patient parameters for predicting in-hospital mortality, neurological outcome, and the need for DC.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy setting \u0026amp; participants\u003c/h2\u003e \u003cp\u003eA retrospective analysis was conducted on patients diagnosed with aSAH at the Department of Neurosurgery at Leipzig University Hospital. The study included patients treated according to corresponding guidelines between 2019 and 2022.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eEthics:\u003c/h2\u003e \u003cp\u003e\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e The retrospective study was conducted in compliance with the Declaration of Helsinki and its amendments and was approved by the local medical ethics committee (387/23).\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eData collection:\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eRetrospective data was collected from patient records and ICU reports for the study. The following clinical, radiological, and laboratory data were recorded: Baseline demographic patient characteristics (Age, sex), Word Federation of Neurological Surgeons (WFNS) Score, Glasgow Coma Scale (GCS) score, pupillary reflex, neuroradiological parameters (Fisher score, intracerebral hemorrhage, midline shift, location of ruptured aneurysm, presence of initial hydrocephalus, Subarachnoid Hemorrhage Early Brain Edema Score (SEBES)) surgical treatment modalities (clipping, endovascular coiling), and laboratory values (hemoglobin, CRP, blood lactate, serum creatinine) \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eOutcomes\u003c/h2\u003e \u003cp\u003eThe study's primary outcome was in-hospital survival. The secondary outcome was favorable neurological outcome, which was based on the modified Rankin Scale (mRS) and defined as a mRS score of 0\u0026ndash;2 \u003csup\u003e14\u003c/sup\u003e. Tertiary outcome parameter was the need for DC.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eDevelopment of a chat prompt for Atlas GPT and ChatGPT 4.0\u003c/h2\u003e \u003cp\u003eIn the creation of a standardized dialogue prompt, our methodology adopted an iterative process, as recommended by Kanjee et al\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. An initial text was composed and subsequently refined through a process of trial and error, with the aim of eliciting specific responses from ChatGPT-4 and AtlasGPT. This initial text comprehensively described the task and context to the language model. The full standardized dialogue prompt is accessible (see \u003cb\u003esupplementary file 1\u003c/b\u003e). In summary, the language model was instructed to assume the role of an \u0026ldquo;AI intensive care physician\u0026rdquo; or an \u0026ldquo;AI neurosurgeon\u0026rdquo; tasked with managing a patient who was admitted on hospital for aSAH and aneurysms were treated via Coiling or Clipping within 24 hours after the aSAH event. Additionally, the model was provided with seventeen baseline patient-, disease-, and procedure-specific parameters, chosen for their recognized value in predicting outcomes after aSAH. This standardized anonymized way was chosen because the unstructured upload of medical records to a cloud-based language model would result in significant data privacy concerns. The text was developed in a standardized way until AtlasGPT (Atlasmeditech LLC. 2024) and ChatGPT 4.0 (OpenAI, Inc., San Francisco, USA) could answer questions with a simple 'yes' or 'no'. The following baseline factors of aSAH at admission were considered in the analysis: WFNS grade, pupillary reflex, age, gender, Glasgow coma scale, Fisher scale, intracerebral hemorrhage, intraventricular hemorrhage, midline-shift, location of ruptured aneurysm, treatment modality, medical treatment with antiplatelet drugs after endovascular treatment, initial hydrocephalus treatment such as external ventricular drain, hemoglobin, CRP, blood lactate, and serum creatinine. AtlasGPT and ChatGPT 4.0 were asked four questions that required \u0026lsquo;yes\u0026rsquo; or \u0026lsquo;no\u0026rsquo; answers after entering the text and parameters (see \u003cb\u003esupplementary file 1\u003c/b\u003e).\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e1)Will this patient survive to hospital discharge? Please provide a yes/no answer.\u003c/p\u003e\u003cp\u003e2)Will this patient experience a good neurological outcome at hospital discharge as defined by the modified ranking scale (0\u0026ndash;2). Please provide a yes/no answer.\u003c/p\u003e\u003cp\u003e3) Will this patient experience a good neurological outcome at 6-months after aneurysmal subarachnoid hemorrhage e as defined by the modified ranking scale (0\u0026ndash;2). Please provide a yes/no answer. \u0026ldquo;\u003c/p\u003e\u003cp\u003e4) Will the patient have to be treated by a decompressive craniectomy within the next week? Please provide a yes/no answer.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe responses were recorded in an Excel file. Each question was asked three times using both AtlasGPT and ChatGPT 4.0. A new chat was opened for each question. In the case of dichotomous answers (yes/no), the most frequent response from the repeated questions was considered. For example, if the responses were yes/yes/no, the overall response was recorded as \u0026lsquo;yes\u0026rsquo;. In cases where \u0026lsquo;no\u0026rsquo; was the more frequent response, we selected \u0026lsquo;no\u0026rsquo; as the final answer. If we did not receive a response to a question, we reopened the chat and repeated the question until we received a complete and unambiguous answer. We used only complete and unambiguous answers for statistical analysis. The chat prompts were performed between the 5th and 30th of March in 2024.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis:\u003c/h2\u003e \u003cp\u003eReceiver operating characteristics curves (ROC) curves and area under curves (AUC) values were calculated to determine the prognostic significance of AtlasGPT, ChatGPT 4.0, WFNS grading, Fisher scale, and SEBES. Median values and interquartile ranges (IQR) of metric data are reported. The prognostic significance of those parameters were determined by sensitivity, specificity, cut-off values, positive/negative predictive value, and the Youden\u0026rsquo;s index. Statistical analyses was performed using SPSS Statistics version 29.0.2.0 (IBM, Armonk, New York). Statistical graphics were created using SPSS version 29.02.0 and were modified with corresponding tables using BioRender.com.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eStudy Cohort\u003c/h2\u003e \u003cp\u003eBetween January 2019, and December 2022, 120 spontaneous aSAH patients were admitted to the present institution. After the exclusion of patients with an withdraw or withhold from life sustaining therapies, patients with an angiogram-negative SAH, and 7 patients who were lost to follow-up, 82 consecutive aSAH patients were included in the present investigation. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e summarizes the process flow chart.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003ePatient characteristics\u003c/h2\u003e \u003cp\u003eThe cohort comprised predominantly middle-aged individuals with a median age of 56 years (IQR 44.8\u0026ndash;67.3), with a slight female predominance (49/82; 59.8%). The majority of patients presented with poor-grade subarachnoid hemorrhage, reflected by WFNS gradings IV and IV in 50 patients (50/82; 61.0%). A considerable number exhibited signs of brain edema on imaging as evidenced by SEBES scores, with 19 patients (23.2%) being assigned the highest score (IV). Baseline vigilance was generally compromised, indicated by a median Glasgow Coma Scale score of 7 (IQR 3.0\u0026ndash;14.0). Intracranial pathology was confirmed by the presence of midline shift in 24 patients (29.3%). The aneurysms were most frequently located at the anterior cerebral artery (ACA) complex (36/82; 43.9%) and the middle cerebral artery (MCA) complex (25/82; 30.5%). Most patients underwent endovascular treatment for their aneurysms (54/82; 65.9%). Further patient- and disease-specific characteristics are summarized in Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePatient characteristics\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCharacteristic\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFrequency (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;82)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge, Median (IQR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e56.0 (44.8\u0026ndash;67.3)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003cp\u003eFemale\u003c/p\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e49 (59.8%)\u003c/p\u003e \u003cp\u003e33 (40.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWFNS grading\u003c/p\u003e \u003cp\u003e1\u003c/p\u003e \u003cp\u003e2\u003c/p\u003e \u003cp\u003e3\u003c/p\u003e \u003cp\u003e4\u003c/p\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e18 (22.0%)\u003c/p\u003e \u003cp\u003e9 (11.0%)\u003c/p\u003e \u003cp\u003e5 (6.1%)\u003c/p\u003e \u003cp\u003e9 (11.0%)\u003c/p\u003e \u003cp\u003e41 (50.0%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFisher scale\u003c/p\u003e \u003cp\u003e1\u003c/p\u003e \u003cp\u003e2\u003c/p\u003e \u003cp\u003e3\u003c/p\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 (1.2%)\u003c/p\u003e \u003cp\u003e3 (3.7%)\u003c/p\u003e \u003cp\u003e62 (75.6%)\u003c/p\u003e \u003cp\u003e16 (19.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSEBES\u003c/p\u003e \u003cp\u003e0\u003c/p\u003e \u003cp\u003e1\u003c/p\u003e \u003cp\u003e2\u003c/p\u003e \u003cp\u003e3\u003c/p\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16 (19.5%)\u003c/p\u003e \u003cp\u003e12 (14.6%)\u003c/p\u003e \u003cp\u003e20 (24.4%)\u003c/p\u003e \u003cp\u003e15 (18.3%)\u003c/p\u003e \u003cp\u003e19 (23.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGCS, Median (IQR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.0 (3.0\u0026ndash;14.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMidline-shift\u003c/p\u003e \u003cp\u003ePresent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e24 (29.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAneurysm location\u003c/p\u003e \u003cp\u003eAnterior complex\u003c/p\u003e \u003cp\u003eACA\u003c/p\u003e \u003cp\u003ePericallosal artery\u003c/p\u003e \u003cp\u003eAcoA\u003c/p\u003e \u003cp\u003eICA\u003c/p\u003e \u003cp\u003eBifurcation\u003c/p\u003e \u003cp\u003eOphtalmic artery\u003c/p\u003e \u003cp\u003eAChoA\u003c/p\u003e \u003cp\u003ePcoA\u003c/p\u003e \u003cp\u003eMCA\u003c/p\u003e \u003cp\u003eM1\u003c/p\u003e \u003cp\u003eBifurcation\u003c/p\u003e \u003cp\u003eM2\u003c/p\u003e \u003cp\u003ePosterior circulation\u003c/p\u003e \u003cp\u003eBasilar artery\u003c/p\u003e \u003cp\u003eVertebral artery\u003c/p\u003e \u003cp\u003ePICA\u003c/p\u003e \u003cp\u003ePCA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 (1.2%)\u003c/p\u003e \u003cp\u003e5 (6.1%)\u003c/p\u003e \u003cp\u003e30 (36.6%)\u003c/p\u003e \u003cp\u003e4 (4.9%)\u003c/p\u003e \u003cp\u003e1 (1.2%)\u003c/p\u003e \u003cp\u003e1 (1.2%)\u003c/p\u003e \u003cp\u003e3 (3.7%)\u003c/p\u003e \u003cp\u003e5 (6.1%)\u003c/p\u003e \u003cp\u003e17 (20.7%)\u003c/p\u003e \u003cp\u003e3 (3.7%)\u003c/p\u003e \u003cp\u003e6 (7.3%)\u003c/p\u003e \u003cp\u003e2 (2.4%)\u003c/p\u003e \u003cp\u003e3 (3.7%)\u003c/p\u003e \u003cp\u003e1 (1.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eArterial hypertension\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e54 (65.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAntiplatelet therapy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7 (8.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAnticoagulation intake\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3 (3.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBaseline hydrocephalus (EVD placement)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e45 (54.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eType of aneurysm treatment\u003c/p\u003e \u003cp\u003eClipping\u003c/p\u003e \u003cp\u003eEndovascular\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e28 (34.1%)\u003c/p\u003e \u003cp\u003e54 (65.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLactate (mmol/I), Median (IQR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.6 (1.0-2.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC-reactive protein (mg/I), Median (IQR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.3 (1.1\u0026ndash;4.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSerum creatinine (\u0026micro;mol/I), Median (IQR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e68.0 (57.0-85.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHemoglobin (mmol/I), Median (IQR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.2 (7.4\u0026ndash;8.7)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eAbbrevations: IQR\u0026thinsp;=\u0026thinsp;Interquartile range; SEBES\u0026thinsp;=\u0026thinsp;subarachnoid hemorrhage early brain edema score; EVD\u0026thinsp;=\u0026thinsp;external ventricular drain; ACA\u0026thinsp;=\u0026thinsp;anterior cerebral artery; AcoA\u0026thinsp;=\u0026thinsp;anterior communicating artery; ICA\u0026thinsp;=\u0026thinsp;internal carotid artery; AChoA\u0026thinsp;=\u0026thinsp;anterior choroidal artery; MCA\u0026thinsp;=\u0026thinsp;middle cerebral artery; PICA\u0026thinsp;=\u0026thinsp;posterior inferior cerebellar artery; PCA\u0026thinsp;=\u0026thinsp;posterior cerebral artery\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eOutcome parameters\u003c/h2\u003e \u003cp\u003eDuring the hospitalization period, 22% of patients (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;18) expired. DC was performed in 34.1% (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;28) during hospital therapy. Primary decompressive hemicraniectomy was performed in 15 cases during surgical clipping (15/28; 53.6%). The residual decompressive hemicraniectomies were performed secondary due to intracerebral pressure elevations after surgery because of subdural-, epidural- and intracerebral hematomas (4/28; 14.3%), brain edema without signs of infarction (4/28; 14.3%), and radiologically determined infarction (5/28; 17.9%). At the time of hospital discharge, a favorable prognosis, denoted by a modified Rankin Scale score of \u0026le;\u0026thinsp;2, was observed in 28% of patients (n\u0026thinsp;=\u0026thinsp;23). At the 6-month follow-up, an improvement in outcomes was noted; 46.9% of the survivors (n\u0026thinsp;=\u0026thinsp;30) achieved a favorable functional status. \u003cb\u003eSupplementary table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e summarizes the outcome parameters, which were requested from the artificial intelligence language models.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003ePrognostication\u003c/h2\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003eIn-hospital Survival\u003c/h2\u003e \u003cp\u003eFor the endpoint of 30-day in-hospital survival, the World Federation of Neurological Surgeons (WFNS) grading scale manifested the superior discriminative capability with an AUC of 0.72 (95% CI: 0.60\u0026ndash;0.84), coupled with a sensitivity of 59.4% and specificity of 83.3% with a cut off set at \u0026le;\u0026thinsp;4/\u0026gt;4. The positive predictive value and negative predictive values were 92.7% and 36.6%, respectively (see \u003cb\u003esupplementary table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e). The WFNS scale's Youden's Index was 0.43, indicating a robust balance between sensitivity and specificity in this acute prognostic scenario at admission. Although AtlasGPT had a slightly lower AUC, it demonstrated a comparable performance in terms of specificity (83.3%) when responding that the patient will survive for at least two times of three chat prompt runs. AtlasGPT and ChatGPT 4.0 showed similar capabilities with AUCs of 0.70 (95% CI: 0.57\u0026ndash;0.83) and 0.67 (95% CI: 0.53\u0026ndash;0.80) respectively, while SEBES and FISHER scales demonstrated lower discriminative power with AUCs of 0.53 (95% CI: 0.38\u0026ndash;0.68) and 0.54 (95% CI: 0.40\u0026ndash;0.69). Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e illustrates the findings.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eNeed for decompressive hemicraniectomy during hospital therapy\u003c/h2\u003e \u003cp\u003eAtlasGPT presented the most pronounced diagnostic accuracy in anticipating the requirement for decompressive hemicraniectomy upon admission with an AUC of 0.80 (95% CI: 0.70\u0026ndash;0.91). The model's sensitivity reached 82.1%, with specificity at 77.8%, and a Youden's Index of 0.6, when prognosticating the need for decompressive hemicraniectomy in three dialog runs in a row. The positive predictive value and negative predictive values of AtlasGPT indicating the need for decompressive hemicraniectomy were 65.7% and 89.4%, respectively (see \u003cb\u003esupplementary table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u003c/b\u003e). AtlasGPT\u0026acute;s predictive power was followed by ChatGPT 4.0, which had an AUC of 0.78 (95% CI: 0.68\u0026ndash;0.88), and the WFNS score with an AUC of 0.76 (95% CI: 0.66\u0026ndash;0.86). SEBES and FISHER scales had the lowest AUCs, indicating less predictive value for this outcome. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e summarizes the results of this analysis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eFavorable outcome (mRS\u0026thinsp;\u0026le;\u0026thinsp;2) at discharge\u003c/h2\u003e \u003cp\u003eThe WFNS score and AtlasGPT both provided high prognostic value for predicting a good functional outcome at discharge, with AUCs of 0.74 (95% CI: 0.61\u0026ndash;0.87) and 0.75 (95% CI: 0.62\u0026ndash;0.87), respectively. The sensitivity and specificity of AtlasGPT were 69.6% and 79.7%, respectively, with an optimal cutoff at \u0026ge;\u0026thinsp;2/\u0026lt;2 yes and a Youden's Index of 0.49, suggesting good prognostic performance at the point of discharge. The positive predictive value and negative predictive values of AtlasGPT prognosticating favorable mRS at discharge were 57.1% and 87.0%, respectively (see supplementary table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). ChatGPT 4.0 also performed adequately with an AUC of 0.72 (95% CI: 0.59\u0026ndash;0.85). The SEBES and FISHER scales were less predictive to assess clinical outcome, as reflected by their AUCs of 0.53 (95% CI: 0.38\u0026ndash;0.67) and 0.62 (95% CI: 0.48\u0026ndash;0.76). Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e displays the results of the analyses regarding functional outcome at discharge.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eFavorable outcome (mRS\u0026thinsp;\u0026le;\u0026thinsp;2) at 6-months\u003c/h2\u003e \u003cp\u003eFor the prediction of a favorable functional outcome at a 6-month interval, the WFNS grading scale was pre-eminent, delivering an AUC of 0.76 (95% CI: 0.64\u0026ndash;0.88), with a sensitivity of 76.7% and specificity of 71.9%. The calculated Youden's Index was 0.49, indicating a balanced predictive performance for long-term recovery assessments. The positive predictive value and negative predictive values of WFNS grading to prognosticate favorable mRS outcome at 6-months were 71.9% and 76.7%, respectively (see \u003cb\u003esupplementary table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u003c/b\u003e). AtlasGPT and ChatGPT 4.0 provided both comparable AUCs of 0.69 (95% CI: 0.56\u0026ndash;0.83). SEBES and FISHER scores were less effective in predicting long-term outcomes, with AUCs of 0.58 (95% CI: 0.43\u0026ndash;0.73) and 0.58 (95% CI: 0.44\u0026ndash;0.72), respectively. \u003cb\u003eSupplementary file 2\u003c/b\u003e outlines the results.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eResponse variability of language models concerning the prediction of endpoints\u003c/h2\u003e \u003cp\u003eDuring each of the three trials involving AtlasGPT and ChatGPT 4.0, we encountered occurrences of divergent outputs, classified as response variability in response to the prompts given. AtlasGPT delivered different responses for the endpoints In-hospital mortality, need for decompressive hemicraniectomy, favorable outcomes at discharge, and favorable outcomes at 6-months after aSAH in 5 (6.1%), 7 (8.5%), 3 (3.7%), and 3 (3.7%) runs, respectively. ChatGPT 4.0 likewise generated diverse outcomes for in-hospital mortality, favorable discharge outcomes, and favorable 6-month outcomes post-aSAH, each occurring in 1 (1.2%) instance. No divergent responses for the necessity of decompressive hemicraniectomy was observed using ChatGPT 4.0.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe present study compared the prognostic value of the retrieval-augmented generation techniques enhanced large language model AtlasGPT based on peer-reviewed reliable neurosurgery-specific evidence, the large language model ChatGPT-4.0 with well-validated clinical and imaging scales. The prognostic performance of AtlasGPT to predict the functional outcome at discharge and the need for decompressive hemicraniectomy in aSAH based on characteristics at hospital admission demonstrated substantial roles for artificial intelligence in clinical practice (see \u003cb\u003esupplementary file 3\u003c/b\u003e). Nevertheless, these results necessitate extensive discussion.\u003c/p\u003e \u003cp\u003eBoth language models showed responses variabilities with slightly more variabilities in the AtlasGPT language model. However, we ran three iterations of four binary questions for 82 patients in two language models and observed only 21 instances of response variability (21/656; 3.2%). To address this issue seriously and reduce potential bias caused response variability, we ultimately took the mean value of the responses for statistical processing. Nevertheless, it has to be noted that our investigation is focused on the most advanced ChatGPT version currently available, which requires a paid subscription (namely, the GPT-4-based model). The freely accessible and more commonly used version is based on the 3.5 model, which might produce less reliable predictions with more response variability and more temporal instability \u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. This demonstrates that artificial intelligence can be most effective when integrated with \u0026ldquo;human intelligence\u0026rdquo;, such as the expertise of a supervision by a clinician. Additionally, it highlights the necessity for users to closely monitor the application of large language models in clinical settings. Decisions to treatment limitations and prognosis are inherently challenging, demanding human attributes such as long-term professional expertise, empathy, and emotional insight. In contrast, large language models are purely machines operating on stochastic processes, lacking any form of consciousness or emotional capacity \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eWhile the use of large language models in the medical field is growing, neurovascular research on their ability to predict patient outcomes remains limited. A retrospective analysis of clinical, neuroimaging, and procedure-related data involving 163 patients with acute ischemic stroke found that ChatGPT adequately predicted short-term functional mRS outcomes at 3-months after mechanical thrombectomy and was superior to existing risk scores \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eMortality and functional prognosis in aSAH is complex, with scales like Hunt and Hess and WFNS gauging risk based on initial assessments and traits but ignoring the effects of ending life support, which might bias outcomes \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Furthermore, there\u0026acute;s growing concern that pessimistic views on severely injured patient\u0026acute;s recoveries might worsen outcomes, termed a \u0026ldquo;self-fulfilling prophecy\u0026rdquo; \u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. However, it is of paramount importance to scientifically validate these language models because patients and their relatives will potentially start putting trust in what conventional AI, like ChatGPT, predicts. Despite language model designers took steps to prevent these consultations by making sure that language models are not physicians and cannot give medical advices, such protocols can be bypassed by \u0026ldquo;jailbreak\u0026ldquo; prompts, such as feigned scenarios as demonstrated in the present study \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003ePhysicians evaluating large language models should be also mindful of the \u0026ldquo;stochastic parrot\u0026rdquo; concept \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. This principle highlights that due to its algorithmic nature, large language models lack comprehension of both the information it receives and the produced responses, merely replicating learned patterns and biases, including stereotypes and social disparities \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. These mechanisms might account for why the large language model\u0026acute;s performances in survival prediction and long-term neurological outcome prediction is comparable and not superior established clinical grading systems. This observation is consistent with findings from other research indicating similar levels of effectiveness in clinical or theoretical settings \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. However, we found a superior role of a generated dialog describing baseline patient-, disease-, and procedure-specific aSAH patient characteristics which might facilitate the identification of those patients who will potentially undergo decompressive hemicraniectomy. Brain swelling and elevated ICP are established factors worsening outcomes after aSAH \u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Despite decompressive hemicraniectomy is a well-established procedure with high-level evidence for space-occupying stroke \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, the currently used language models and also AtlasGPT still have to rely on retrospective data. Further insights on the role of primary decompressive hemicraniectomy in aSAH and which patient will potentially benefit from this procedure might inform the randomized controlled PICASSO trial investigating decompressive hemicraniectomy in poor-grade (WFNS IV\u0026amp;V) aSAH patients \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eTo our knowledge, this is the first investigation using large language models for predicting outcomes following aSAH with real-world patient data. It adopts a practical method focused on reproducibility and the integrity of data. Nonetheless, this study is not without its drawbacks.\u003c/p\u003e \u003cp\u003eFirst, there is the tendency of large language models to produce response variability which have to be addressed by repetitive approaches and trusting the mean values. Additionally, ChatGPT 4.0 was not specifically developed for medical use compared to AtlasGPT. Hence, the effectiveness and accuracy of ChatGPT 4.0 in addressing clinical issues has yet to be established. Finally, the present study represents a single-center cohort, which may limit their applicability to other settings or regions, underscoring the need for further studies across various environments to confirm the universality of these results. Finally, the present results have not been compared to a blinded experienced senior physician estimating the outcomes of the individual patients with the same baseline data.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003eAtlasGPT demonstrated a good performance in predicting functional outcome at discharge and the need for decompressive hemicraniectomy and thus may be a helpful tool for early endpoints in risk predictions of aSAH patients. Nevertheless, large language models still need human supervision by professionals. Future research should focus on specific trained large language models and further validate various clinical endpoints.\u003c/p\u003e "},{"header":"Declarations","content":" \u003cp\u003e The Ethics Committee at the Medical Faculty of the University of Leipzig has no ethical or scientific objections to the study (ID: 387/23).\u003c/p\u003e\u003cp\u003eThe manuscript has been prepared in accordance with the instructions provided by the authors and has been approved by all of them. All of the requirements set out by the authors have been met and approved by all of them.The authors have no conflicts of interest to declare. All co-authors have seen and agree with the content of the manuscript and there is no financial interest to report. We certify that the submission is original work and is not under review at any other publication.\u003c/p\u003e\n\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\n\u003cp\u003eConceptualization: Alim Emre Basaran, Johannes Wach. Investigation: Johannes Wach, Alim Emre Basaran. Visualization: Johannes Wach, Alim Emre Basaran. Supervision: Erdem G\u0026uuml;resir, Johannes Wach. Writing-original draft: Alim Emre Basaran, Johannes Wach. Writing-review \u0026amp; editing: Alim Emre Basaran, Johannes Wach, Agi G\u0026uuml;resir, Hanna Knoch, Martin Vychopen, Erdem G\u0026uuml;resir, Johannes Wach.\u0026nbsp;\u003c/p\u003e\u003ch2\u003eFunding:\u003c/h2\u003e \u003cp\u003eNo funding was received to assist with the preparation of this manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgment:\u003c/h2\u003e \u003cp\u003eThe visual abstract was created using BioRender\u003c/p\u003e\u003ch2\u003eData Availability Statement:\u003c/h2\u003e \u003cp\u003eData available on request from the authors\u003c/p\u003e\u003ch2\u003eInstitutional Review Board Statement\u003c/strong\u003e \u003cp\u003e The study was conducted in accordance with the Declaration of Helsinki.\u003c/p\u003e \u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eEtminan N, Chang HS, Hackenberg K et al (2019) Worldwide Incidence of Aneurysmal Subarachnoid Hemorrhage According to Region, Time Period, Blood Pressure, and Smoking Prevalence in the Population: A Systematic Review and Meta-analysis. JAMA Neurol 76(5):588\u0026ndash;597. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamaneurol.2019.0006\u003c/span\u003e\u003cspan address=\"10.1001/jamaneurol.2019.0006\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKorja M, Lehto H, Juvela S, Kaprio J (2016) Incidence of subarachnoid hemorrhage is decreasing together with decreasing smoking rates. Neurology. ;87(11):1118-23. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1212/WNL.0000000000003091.Epub\u003c/span\u003e\u003cspan address=\"10.1212/WNL.0000000000003091.Epub\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e 2016 Aug 12\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSVIN COVID-19 Global SAH Registry Global impact of the COVID-19 pandemic on subarachnoid haemorrhage hospitalisations, aneurysm treatment and in-hospital mortality: 1-year follow-up. J Neurol Neurosurg Psychiatry 2022 Jul 28:jnnp\u0026ndash;2022. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/jnnp-2022-329200\u003c/span\u003e\u003cspan address=\"10.1136/jnnp-2022-329200\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoquer J, Cuadrado-Godia E, Guimaraens L (2020) Short- and long-term outcome of patients with aneurysmal subarachnoid hemorrhage. Neurology 95(13):e1819\u0026ndash;e1829. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1212/WNL.0000000000010618\u003c/span\u003e\u003cspan address=\"10.1212/WNL.0000000000010618\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHammer A, Steiner A, Ranaie G et al (2018) Impact of Comorbidities and Smoking on the Outcome in Aneurysmal Subarachnoid Hemorrhage. Sci Rep. 8: 12335. Published online 2018 Aug 17 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-018-30878-9\u003c/span\u003e\u003cspan address=\"10.1038/s41598-018-30878-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReport of World Federation of Neurological Surgeons Committee on a Universal Subarachnoid Hemorrhage Grading Scale (1988) J Neurosurg 68(6):985\u0026ndash;986. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/jns.1988.68.6.0985\u003c/span\u003e\u003cspan address=\"10.3171/jns.1988.68.6.0985\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG\u0026uuml;resir E, Raabe A, Setzer M, Vatter H, Gerlach R, Seifert V, Beck J (2009) Decompressive hemicraniectomy in subarachnoid haemorrhage: the influence of infarction, haemorrhage and brain swelling. J Neurol Neurosurg Psychiatry 80(7):799\u0026ndash;801. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/jnnp.2008.155630\u003c/span\u003e\u003cspan address=\"10.1136/jnnp.2008.155630\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJohnsson J, Bj\u0026ouml;rnsson O, Andersson P et al (2020) Artificial neural networks improve early outcome prediction and risk classification in out-of-hospital cardiac arrest patients admitted to intensive care. Crit Care 24(1):474. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s13054-020-03103\u003c/span\u003e\u003cspan address=\"10.1186/s13054-020-03103\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChung CC, Chiu WT, Huang YH et al Identifying prognostic factors and developing accurate outcome predictions for in-hospital cardiac arrest by using artificial neural networks. J Neurol Sci 2021 Jun 15:425:117445. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jns.2021.117445\u003c/span\u003e\u003cspan address=\"10.1016/j.jns.2021.117445\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2021 Apr 18.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedro T, Sousa JM, Ronseca L et al Exploring the use of ChatGPT in predicting anterior circulation stroke functional outcomes after mechanical thrombectomy: a pilot study. J Neurointerv Surg 2024 Mar 7:jnis\u0026ndash;2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/jnis-2024-021556\u003c/span\u003e\u003cspan address=\"10.1136/jnis-2024-021556\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHopkins SB, Carter B, Lord J et al Editorial. AtlasGPT: dawn of a new era in neurosurgery for intelligent care augmentation, operative planning, and performance. J Neurosurg 2024 Feb 27:1\u0026ndash;410.3171/2024.2.JNS232997\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhn SH, Savarraj JP, Pervez M, Jones W, Park J, Jeon SB, Kwon SU, Chang TR, Lee K, Kim DH, Day AL, Choi HA (2018) The Subarachnoid Hemorrhage Early Brain Edema Score Predicts Delayed Cerebral Ischemia and Clinical Outcomes. Neurosurgery 83(1):137\u0026ndash;145. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/neuros/nyx364\u003c/span\u003e\u003cspan address=\"10.1093/neuros/nyx364\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFisher CM, Kistler JP, Davis JM (1980) Relation of cerebral vasospasm to subarachnoid hemorrhage visualized by computerized tomographic scanning. Neurosurgery 6(1):1\u0026ndash;9. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1227/00006123-198001000-00001\u003c/span\u003e\u003cspan address=\"10.1227/00006123-198001000-00001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeisscher N, Vermeulen M, Roos YB, de Haan RJ (2008) What should be defined as good outcome in stroke trials; a modified Rankin score of 0\u0026ndash;1 or 0\u0026ndash;2? Neurol 255(6):867\u0026ndash;874. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00415-008-0796-8\u003c/span\u003e\u003cspan address=\"10.1007/s00415-008-0796-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003eEpub 2008 Mar 14\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKanjee Z, Crowe B, Rodman A (2023) Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330(1):78\u0026ndash;80. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jama.2023.8288\u003c/span\u003e\u003cspan address=\"10.1001/jama.2023.8288\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen L, Zaharia M, Zou J How Is ChatGPT\u0026acute;s behavior changing over time? arXiv:2307.09009v3. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2307.09009\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2307.09009\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan B, Fatima H, Qureshi A et al Drawbacks of Artificial Intelligence and Their Potential Solutions in the Healthcare sector. Biomed Mater Devices 2023 Feb 8: 1\u0026ndash;8. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s44174-023-00063-2\u003c/span\u003e\u003cspan address=\"10.1007/s44174-023-00063-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedro T, Sousa JM, Fonseca L et al Exploring the use of ChatGPT in predicting anterior circulation stroke functional outcomes after mechanical thrombectomy: a pilot stuy. J Neurointerv Surg 2024 Mar 7:jnis\u0026ndash;2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/jnis-2024-021556\u003c/span\u003e\u003cspan address=\"10.1136/jnis-2024-021556\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Heuven AW, Mees SMD, Algra A, Rinkel GJE (2008) Validation of a prognostic subarachnoid hemorrhage grading scale derived directly from the Glasgow Coma Scale. Stroke 39(4):1347\u0026ndash;1348. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/STROKEAHA.107.498345\u003c/span\u003e\u003cspan address=\"10.1161/STROKEAHA.107.498345\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003eEpub 2008 Feb 28\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBecker KJ, Baxter AB, Cohen WA et al (2001) Withdrawal of support in intracerebral hemorrhage may lead to self-fulfilling prophecies. Neurology 56(6):766\u0026ndash;772. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1212/wnl.56.6.766\u003c/span\u003e\u003cspan address=\"10.1212/wnl.56.6.766\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHemphill JC 3rd, White DB (2009) Clinical nihilism in neuroemergencies. Emerg Med Clin North Am. ;27(1):27\u0026ndash;37, vii-viii. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.emc.2008.08.009\u003c/span\u003e\u003cspan address=\"10.1016/j.emc.2008.08.009\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Y, Deng G, Zhengzi, Xu et al Jailbreaking ChatGPT via Prompt Engineering: An empirical Study. arXiv:2305.13860v2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBender EM, McMillan-Major A, Gebru T, Shmitchell S On the On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Conference on Fairness, Accountability, and Transparency (FAccT \u0026acute;21), March 3\u0026ndash;10, 2021, Virtual Event, Canada. ACM, New York, NY, USA, 14 pages. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://doi.org/10.1145/3442188.3445922\u003c/span\u003e\u003cspan address=\"10.1145/3442188.3445922\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoussen S, Denis JC, Simeone P et al (2023) ChatGPT and the stochastic parrot: artificial intelligence in medical research. Br J Anaesth 131(4):e120\u0026ndash;e121. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.bja.2023.06.065\u003c/span\u003e\u003cspan address=\"10.1016/j.bja.2023.06.065\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003eEpub 2023 Jul 27\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeshpande A, Murahari V, Rajpurohit T et al Toxicity in ChatGPT: Analyzing Persona-assigned Language Models. arXiv:2304.05335V1\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKanjee Z, Crowe B, Rodman A (2023) Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA 330:78\u0026ndash;80. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jama.2023.8288\u003c/span\u003e\u003cspan address=\"10.1001/jama.2023.8288\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGilson A, Safranek CW, Huang T et al (2023) How does ChatGPT perform on the united states medical licensing examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 9:e45312. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/45312\u003c/span\u003e\u003cspan address=\"10.2196/45312\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeuer GG, Smith MJ, Elliott JP et al (2004) Relationship between intracranial pressure and other clinical variables in patients with aneurysmal subarachnoid hemorrhage. Neurosurg 101(3):408\u0026ndash;416. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/jns.2004.101.3.0408\u003c/span\u003e\u003cspan address=\"10.3171/jns.2004.101.3.0408\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVahedi K, Hofmeijer J, Juettler E et al (2007) Early decompressive surgery in malignant infarction of the middle cerebral artery: a pooled analysis of three randomized controlled trials. Lancet Neurol 6(3):215\u0026ndash;222. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S1474-4422(07)70036-4\u003c/span\u003e\u003cspan address=\"10.1016/S1474-4422(07)70036-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG\u0026uuml;resir E, Lampmann T, Brandecker S et al (2022) PrImary decompressive Craniectomy in AneurySmal Subarachnoid hemOrrhage (PICASSO) trial: study protocol for a randomized controlled trial. Trials 23(1):1027. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s13063-022-06969-4\u003c/span\u003e\u003cspan address=\"10.1186/s13063-022-06969-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Aneurysmal subarachnoid hemorrhage, Artificial intelligence, AtlasGPT, ChatGPT, neurologic outcomes, prediction","lastPublishedDoi":"10.21203/rs.3.rs-4621973/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4621973/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eTo assess the predictive accuracy of advanced AI language models and established clinical scales in prognosticating outcomes for patients with aneurysmal subarachnoid hemorrhage (aSAH).\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThis retrospective cohort study included 82 patients suffering from aSAH. We evaluated the predictive efficacy of AtlasGPT and ChatGPT 4.0 by examining the area under the curve (AUC), sensitivity, specificity, and Youden's Index, in comparison to established clinical grading scales such as the World Federation of Neurological Surgeons (WFNS) scale, Simplified Endovascular Brain Edema Score (SEBES), and Fisher scale. This assessment focused on four endpoints: in-hospital mortality, need for decompressive hemicraniectomy, and functional outcomes at discharge and after 6-month follow-up.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eIn-hospital mortality occurred in 22% of the cohort, and 34.1% required decompressive hemicraniectomy during treatment. At hospital discharge, 28% of patients exhibited a favorable outcome (mRS\u0026thinsp;\u0026le;\u0026thinsp;2), which improved to 46.9% at the 6-month follow-up. Prognostication utilizing the WFNS grading scale for 30-day in-hospital survival revealed an AUC of 0.72 with 59.4% sensitivity and 83.3% specificity. AtlasGPT provided the highest diagnostic accuracy (AUC 0.80, 95% CI: 0.70\u0026ndash;0.91) for predicting the need for decompressive hemicraniectomy, with 82.1% sensitivity and 77.8% specificity. Similarly, for discharge outcomes, the WFNS score and AtlasGPT demonstrated high prognostic values with AUCs of 0.74 and 0.75, respectively. Long-term functional outcome predictions were best indicated by the WFNS scale, with an AUC of 0.76.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThe study demonstrates the potential of integrating AI models such as AtlasGPT with clinical scales to enhance outcome prediction in aSAH patients. While established scales like WFNS remain reliable, AI language models show promise, particularly in predicting the necessity for surgical intervention and short-term functional outcomes.\u003c/p\u003e","manuscriptTitle":"Beyond Traditional Prognostics: Integrating RAG-Enhanced AtlasGPT and ChatGPT 4.0 into Aneurysmal Subarachnoid Hemorrhage Outcome Prediction","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-07-22 10:54:28","doi":"10.21203/rs.3.rs-4621973/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0c2572be-1e43-4439-96f3-ee690cb2f983","owner":[],"postedDate":"July 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-08-23T04:33:32+00:00","versionOfRecord":[],"versionCreatedAt":"2024-07-22 10:54:28","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4621973","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4621973","identity":"rs-4621973","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00