Cerebrovascular disease case identification in inpatient electronic medical record data using natural language processing

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Background: Abstracting cerebrovascular disease (CeVD) from inpatient electronic medical records (EMRs) through natural language processing (NLP) is pivotal for automated disease surveillance and improving patient outcomes. Existing methods rely on coders’ abstraction, which has time delays and under-coding issues. This study sought to develop an NLP-based method to detect CeVD using EMR clinical notes. Methods CeVD status was confirmed through a chart review on randomly selected hospitalized patients who were 18 years or older and discharged from 3 hospitals in Calgary, Alberta, Canada, between January 1 and June 30, 2015. These patients’ chart data were linked to administrative discharge abstract database (DAD) and Sunrise TM Clinical Manager (SCM) EMR database records by Personal Health Number (a unique lifetime identifier) and admission date. We trained multiple natural language processing (NLP) predictive models by combining two clinical concept extraction methods and two supervised machine learning (ML) methods: random forest and XGBoost. Using chart review as the reference standard, we compared the model performances with those of the commonly applied International Classification of Diseases (ICD-10-CA) codes, on the metrics of sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Result Of the study sample (n=3036), the prevalence of CeVD was 11.8% (n=360); the median patient age was 63; and females accounted for 50.3% (n=1528) based on chart data. Among 49 extracted clinical documents from the EMR, four document types were identified as the most influential text sources for identifying CeVD disease (“nursing transfer report,” “discharge summary,” “nursing notes,” and “inpatient consultation.”). The best performing NLP model was XGBoost, combining the Unified Medical Language System concepts extracted by cTAKES (e.g., top-ranked concepts, “Cerebrovascular accident” and “Transient ischemic attack”), and the term frequency-inverse document frequency vectorizer. Compared with ICD codes, the model achieved higher validity overall, such as sensitivity (25.0% vs 70.0%), specificity (99.3% vs 99.1%), PPV (82.6 vs. 87.8%), and NPV (90.8% vs 97.1%). Conclusion The NLP algorithm developed in this study performed better than the ICD code algorithm in detecting CeVD. The NLP models could result in an automated EMR tool for identifying CeVD cases and be applied for future studies such as surveillance, and longitudinal studies.
Full text 100,322 characters · extracted from preprint-html · click to expand
Cerebrovascular disease case identification in inpatient electronic medical record data using natural language processing | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Cerebrovascular disease case identification in inpatient electronic medical record data using natural language processing Jie Pan, Zilong Zhang, Steven Ray Peters, Shabnam Vatanpour, Robin L. Walker, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2640617/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Sep, 2023 Read the published version in Brain Informatics → Version 1 posted 7 You are reading this latest preprint version Abstract Background Abstracting cerebrovascular disease (CeVD) from inpatient electronic medical records (EMRs) through natural language processing (NLP) is pivotal for automated disease surveillance and improving patient outcomes. Existing methods rely on coders’ abstraction, which has time delays and under-coding issues. This study sought to develop an NLP-based method to detect CeVD using EMR clinical notes. Methods CeVD status was confirmed through a chart review on randomly selected hospitalized patients who were 18 years or older and discharged from 3 hospitals in Calgary, Alberta, Canada, between January 1 and June 30, 2015. These patients’ chart data were linked to administrative discharge abstract database (DAD) and Sunrise TM Clinical Manager (SCM) EMR database records by Personal Health Number (a unique lifetime identifier) and admission date. We trained multiple natural language processing (NLP) predictive models by combining two clinical concept extraction methods and two supervised machine learning (ML) methods: random forest and XGBoost. Using chart review as the reference standard, we compared the model performances with those of the commonly applied International Classification of Diseases (ICD-10-CA) codes, on the metrics of sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Result Of the study sample (n=3036), the prevalence of CeVD was 11.8% (n=360); the median patient age was 63; and females accounted for 50.3% (n=1528) based on chart data. Among 49 extracted clinical documents from the EMR, four document types were identified as the most influential text sources for identifying CeVD disease (“nursing transfer report,” “discharge summary,” “nursing notes,” and “inpatient consultation.”). The best performing NLP model was XGBoost, combining the Unified Medical Language System concepts extracted by cTAKES (e.g., top-ranked concepts, “Cerebrovascular accident” and “Transient ischemic attack”), and the term frequency-inverse document frequency vectorizer. Compared with ICD codes, the model achieved higher validity overall, such as sensitivity (25.0% vs 70.0%), specificity (99.3% vs 99.1%), PPV (82.6 vs. 87.8%), and NPV (90.8% vs 97.1%). Conclusion The NLP algorithm developed in this study performed better than the ICD code algorithm in detecting CeVD. The NLP models could result in an automated EMR tool for identifying CeVD cases and be applied for future studies such as surveillance, and longitudinal studies. Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Accurate identification of patients with cerebrovascular diseases (CeVD) is important for health services research, surveillance and monitoring, risk adjustment, and quality improvement measurement [1, 2]. The standard approach to identify conditions is coded administrative hospital data using International Classification of Disease (ICD) terminology. Although structured codes are widely available and highly standardized, some conditions, including CeVD, are under-coded. Quan et al. [3] validated the ICD algorithms against chart review and reported a sensitivity of 46.3% for detecting CeVD diseases in both ICD-9 and ICD-10-CA. To overcome the shortcomings of ICD code-based algorithms, medical chart reviews act as a gold standard for case identification. Unfortunately, chart review is time- and resource-intensive requiring health professionals familiar with specific conditions [4, 5]. Electronic medical records (EMRs) are becoming increasingly popular for collecting health information [6], and can be used to improve the accuracy of identifying conditions such as CeVD. Among the components of EMR, free text notes contain detailed descriptions and give health professionals great flexibility to report conditions and comorbidities. Natural Language Processing (NLP) is an artificial intelligence technique to analyze human languages and retrieve clinically relevant information for detecting and predicting medical conditions [7]. A recent literature conducted by our team yielded few studies using NLP on clinical notes for patients with CeVD conditions [8]. Existing studies have focused on identifying ischemic stroke [9–11] and cerebral aneurysms [12], predicting the cerebrovascular causes of ischemia [13], and detecting complications of stroke [14]. Most previous studies focus on specific conditions within CeVD and have limited access to a complete set of clinical notes from EMRs, using only admission notes or radiology reports. In this study, we explored all available types of inpatient clinical notes from an EMR to identify a broad spectrum of CeVD cases. The CeVD cases were defined by our previous ICD-10 algorithm [3]. We hypothesized that using NLP techniques on these clinical notes would better detect CeVD cases than ICD-based algorithms and existing ML algorithms with limited data source types. Methods Study Population In this retrospective cohort study, we randomly selected patients who were at least 18 years of age and discharged from three acute care facilities in Calgary, Canada, between January 1 and June 30, 2015. Obstetric admissions were excluded because they have a short length of stay and lack conditions of interest. We randomly selected one hospitalization per patient if multiple discharges occurred during the study period [15]. Six nurses reviewed charts to determine the existence of CeVD [15]. Data sources EMR: Sunrise Clinical Manager (SCM) The EMR data are from SCM, a city-wide, population level EMR system used in the three acute care hospitals in Calgary. SCM provides patient-level clinical information containing medical and nursing orders, medication records, clinical documentation, diagnostic imaging and lab results [16]. Administrative Discharge Abstract Database: DAD The inpatients’ administrative, clinical, and demographic information at the time of discharge is coded in the DAD [17]. The clinical coder records up to 25 diagnostics codes for each inpatient based on available information from patient charts. The DAD, EMR data and chart data were linked with Personal Health Number (a unique lifetime identifier), chart number (a distinctive number associated with a patient’s admission), and admission date. Phenotyping Algorithm Framework We trained, validated, and tested an EMR data-driven phenotyping algorithm using NLP techniques to detect CeVD. NLP techniques are used to process and analyze human language, and contain a wide range of tasks, including named entity recognition (NER), information extraction, and text classification [11, 16]. They were applied to analyze the free text clinical notes and derive a CeVD phenotype to detect the disease automatically. As depicted in Figure 1, the general framework consists of 1) input document selection from patients’ clinical notes, 2) model training, and 3) performance evaluation using chart review as a reference standard. Document selection and feature engineering Many types of clinical notes could be generated during the hospitalization of patients involved in this study, such as nursing transfer reports, inpatient consultations, discharge summaries, and surgical assessment and history. However, not all document types contribute equally to the detection of CeVD. Noise and redundant information can hamper the detection performance of ML models [19]. The first step is determining and selecting the appropriate document type(s) sensitive to CeVD identification. The method we used is a feedforward sequential selection method [20], to iteratively add the document type that contributes most to model performance, until the performance stops increasing or reaches a predefined criterion, as shown in Figure 2. All the documents are first converted into vectors by 1) extracting relevant medical concepts from the text and 2) turning concepts into numeric features. To examine the extraction performance, we compared two types of commonly used concept extraction methods: Bag of Words (BOW) using ScispaCy [21] and Concept Unique Identifiers (CUIs) from the Unified Medical Language System using cTAKES (see Appendix 1 for a detailed explanation) [22]. We also compared two types of feature construction methods: Term Frequency-Inverse Document Frequency (TF-IDF) and word count. The obtained vectors are fed into the ML models and validated by the model performance. To estimate better generalization of the selected document types, 5-fold cross validation was applied to the selected patients (i.e., 80% training, n= 2429 and 20% test, n=607). The model development is detailed in the following section. Model development The model outcome is a binary classification where hospitalized patients with CeVD are considered positive cases. Two supervised ML methods were trained, validated, and tested using the obtained input vectors and chart review output labels, including random forest (RF) and XGBoost [19, 20]. The two methods are known for handling datasets with high dimensionality, missing data and outliers, and providing accurate and reliable predictions, especially for NLP tasks containing thousands of concept features [21, 22]. With the different combinations among methods of concept extraction, vectorization, and ML models, we have 8 model variations, such as “BOW + TF-IDF + RF” and “CUI + TF-IDF + XGBoost.” As both methods, RF and XGBoost, use decision trees as the base models, we assigned 100 decision trees to them, respectively. These models’ performance was then estimated by 5-fold cross validation, maintaining the same proportion of positive and negative patients in each group. Performance metrics To evaluate and compare the models developed, we calculated their sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and F1 score using chart data as a reference standard. We also calculated binomial proportional confidence intervals for all the metrics. We compare the results with ICD-based CeVD identification algorithms in DAD after defining CeVD using ICD-10 codes (e.g., G45-46, I60-69, H34, see Table S1 in Appendix) [3]. The performance metrics of the developed NLP models were reported on the same level of specificity as with the ICD-based algorithm. Results Characteristics of the study cohort Among the 3036 patients, chart reviewers identified 360 patients with CeVD (see Table 2 ). Characteristics that were statistically significantly different (P < .05) between the CeVD positive cohort and negative cohort are: age, comorbidities such as atrial fibrillation, angina, hypertension, peripheral vascular disease (PVD) and obesity Table 1 Patients Characteristics. Characteristics All (percentage) Patients with CEVD (percentage) Patients without CEVD (percentage) P value N = 3036 (100%) 360(11.9%) 2676(88.1%) Demographic Median of Age (IQR) 63.0 (48.9–76.5) 77.4 (67.0-85.9) 60.9 (46.4–74.2) < 0.0001 Female 1528 (50.3%) 175 (48.6%) 1353(50.6%) 0.5 Comorbidities Atrial fibrillation 370 (12.2%) 106 (29.4%) 264 (9.9%) < 0.0001 Angina 203 (6.7%) 41 (11.4%) 162 (6.1%) 0.0002 Myocardial infarction 102 (3.4%) 18 (5.0%) 84 (3.1%) 0.06 Hypertension 1469 (48.4%) 267 (74.2%) 1202 (44.9%) < 0.0001 Peripheral vascular disease 148 (4.9%) 46 (12.8%) 102 (3.8%) < 0.0001 Obesity 736 (24.2%) 68 (18.9%) 668 (25.0%) 0.01 Alcohol abuse 230 (7.6%) 19 (5.3%) 211 (7.9%) 0.08 Smoking 605 (19.9%) 66 (18.3%) 539 (20.1%) 0.4 IQR = Interquartile range Characteristics of selected document types We collected 49 types of clinical documents of patients during hospitalization, such as nursing transfer reports, inpatient consultations, and discharge summaries. The detailed text statistics for these document types can be found in Appendix Table S2. For a better explanation, we consolidated these document types into 9 categories (see Table S3 in Appendix ). Using the feedforward sequential selection, we identified four essential document types, “nursing transfer report,” “discharge summary,” “nursing notes,” and “inpatient consultation.” These documents are sensitive and informative for CeVD detection. Table 2 shows the statistics of patients, documents, and words. At least 90% of patients (with or without CeVD) have at least 2 types of documents. These four types of documents complement each other in providing sufficient clinical information to identify CeVD. Table 2 Characteristics of extracted documents. Document type All (n = 3036) Patients with CeVD (n = 360) Patients without CeVD (n = 2676) Median number of notes per patient (IQR) 2.0 (1.0–2.0) 2.0 (1.0–2.0) 2.0 (1.0–2.0) Number of patients with at least 2 types of documents (%) 2774 (91.4) 344 (95.6) 2430 (90.8) Median word count per note (IQR) 430.0 (310.0-678.0) 434.5 (322.2–723.0) 428.0 (308.0-675.0) Detailed document types: nursing transfer report - emergency department to inpatient, discharge summary-medical; surgical assessment and history, inpatient consultations, and discharge summary. To examine how these document types contribute to CeVD detection, the top ten key concepts in each document type were analyzed, as shown in Fig. 3 . There are some common and vital concepts across four document types, such as “C0038454” (stroke-related concepts) and “C0007787” (transient ischemic attack). It is reasonable that the existence of these concepts can directly reflect the CeVD status. The remaining concepts are less overlapped and unique to each document type, such as “C0012169” (low sodium diet) in “nursing notes,” “C0004134” (ataxia) in “nursing transfer report,” “C0202691” (CAT scan of head) in “discharge summary,” and “C0001962” (ethanol) in “inpatient consultation.” This demonstrated that these document types contain essential concepts and can supplement each other to gain more comprehensive information in CeVD detection. Classification performance The top 4 trained models were shown in Table 3 . XGBoost generally outperformed the random forest method. TF-IDF performed better than term count when comparing models “CUI + word count + XGBoost” and “CUI + TF-IDF + XGBoost.” Similarly, the concept extraction method “CUI” had better performance than “Bag of Words (BOW).” Consequently, the combination of XGBoost, TF-IDF, and CUI achieved the best performance over other ML models in the metrics of sensitivity (70%), specificity (99.1%), PPV (87.8%), NPV (97.1%), F1 (77.8%), and accuracy (96.5%). We also compared the model performance with ICD-10-CA-based methods. With similar specificity (99.3% in ICD-10-CA vs 99.1% in model “CUI + TF-IDF + XGBoost”), the performance in other metrics is improved hugely by the obtained model, such as sensitivity increased from 25.0–70.0%, and F1 increased from 38.4–77.8%. Table 3 CeVD case identification with DAD and EMR. Model Sensitivity% (95% CI) Specificity% (95% CI) PPV% (95% CI) NPV% (95% CI) F1% Accuracy% (95% CI) ICD-10-CA-codes in DAD 25.0 (20.6–29.8) 99.3 (98.9–99.6) 82.6 (74.5–88.5) 90.8 (90.3–91.3) 38.4 90.5 (89.4–91.5) CUI + TF-IDF + RF 65.8 (60.7–70.7) 98.5 (98.0–99.0) 85.9 (81.5–89.3) 95.5 (94.9–96.1) 74.1 94.7 (93.8–95.4) CUI + word count + XGBoost 68.1 (63.0-72.8) 98.6 (98.1–99.0) 86.9 (82.7–90.2) 95.8 (95.2–96.4) 76.2 95.0 (94.2–95.7) CUI + TF-IDF + XGBoost* 70.00 (65.0-74.7) 99.1 (98.7–99.3) 87.8 (83.7–91.0) 97.1 (96.6–97.5) 77.8 96.5 (95.8–97.0) BOW + TF-IDF + XGBoost 59.2 (53.9–64.3) 98.7 (98.1–99.1) 85.5 (80.9–89.2) 94.7 (94.1–95.3) 69.4 94.0 (93.1–94.8) We included the four metrics of the four NLP models with changing threshold values from 0.05 to 0.95, as shown in Fig. 4 . Since the ICD algorithm is deterministic, its threshold is not changeable. The PPVs of “CUI + TF-IDF + RF,” “CUI + word count + XGBoost,” “BOW + TF-IDF + XGBoost,” and “CUI + TF-IDF + XGBoost” started to exceed the performance of ICD at thresholds 0.32, 0.25, 0.28, and 0.17 within the threshold bound, respectively. “CUI + word count + XGBoost” and “CUI + TF-IDF + XGBoost” had very similar and robust performance with the change of thresholds, whereas “CUI + TF-IDF + RF” was affected significantly. Generally, the “CUI + TF-IDF + XGBoost” algorithm achieved better and more robust performance with smaller thresholds. Discussion This paper shows that EMR textual information abstracted by NLP techniques outperforms traditional ICD codes for assessing cerebrovascular disease, and compares favourably with resource-intensive chart review using a fraction of human resources. With the prevalence of 11.8% CeVD in over 3000 records, the developed NLP model significantly improves the validity of DAD-based ICD algorithm (sensitivity: 70% vs. 25% and PPV: 88% vs. 83%). EMR data is more informative and efficient in identifying CeVD patients than conventionally used hospitalization data (i.e., DAD). First, due to the high volume of discharges, coders have limited time to code patients comprehensively, causing missing codes and low quality. Second, there is no uniform international definition of the most responsible diagnosis, which varies between the primary reason for admission and the condition with intensive resource usage [ 27 ]. When looking for conditions contributing primarily to the length of stay in hospital (a Canada-wide used definition), CeVD is likely under-coded as it can be a comorbidity causing admission. Conversely, EMRs contain many documents not usually used by medical coders. As identified in this study, four types of documents (i.e., “nursing transfer report,” “discharge summary,” “nursing notes,” and “inpatient consultation”) jointly contribute to the accurate detection of CeVD by providing more comprehensive medical information. Restricting the analysis to a specific document type, therefore, has the potential to impede detection. To abstract the knowledge from these EMRs textual data, NLP techniques are essential. Information extraction from unstructured text is known to be difficult, and contains subtasks including NER, relation extraction, and pattern extraction. The text-based classification assigns categorical labels for a text fragment by finding the patterns composed of NERs and their relationships. By comparing different combinations of NLP models, we identified the optimal model, CUI + TF-IDF + XGBoost. The TF-IDF performs better than word count because it can efficiently eliminate low-sensitive concepts in differentiating positive and negative groups. CUI is a better concept extraction method than the NER by scispaCy because cTAKES can merge similar concepts into one, such as “stroke,” “CVA,” and “brain vascular accidents” are mapped to the same CUI “C0038454”. Sine XGBoost has better capability in dealing with overfitting and allows a more general model than random forest, it shows a slightly better performance in detecting CeVD, as shown in Table 3 . The widespread use of text based EMR algorithms to supplement ICD codes and traditional chart reviews has many potential advantages for epidemiology and health outcomes research. CeVD status is frequently used as an important factor in stratifying outcomes in population health research. While some outcomes, such as ischemic stroke, have reasonable validity, other aspects of CeVD, such as carotid atherosclerosis, are likely poorly coded. This probably explains the poor sensitivity (25%) of ICD codes for CeVD in our study. We achieved 88% PPV and 70% sensitivity, an improvement over the widely adopted ICD-based algorithm. Given the amount of knowledge contained in clinical text, the algorithm is applicable to detecting many other diseases, especially conditions with under-coding issues. Text based EMR algorithms may be used to periodically re-evaluate the validity of existing ICD code-based approaches and ensure that ICD code validity is not changing over time. Limitations There are some limitations in this study. First, further examination of missing cases is needed, as 30% of cases are still missed by the proposed algorithm using EMR data. The missing cases are likely caused by variations in clinical documents and the capability of NLP models to detect them. We believe that the performance of the NLP models can be further improved by having better NER and incorporating sequential and contextual patterns among recognized concepts. Second, the data we studied is only from one city (i.e., Calgary). EMR diversities in format and content could be subject to change when larger populations and geographies are considered. The identified sensitive document types will vary accordingly. Lastly, we did not validate the algorithms in external databases. We encourage researchers to apply this method to their datasets for validation and improvement. Conclusion Compared to the widely used ICD-based algorithm, the EMR NLP model significantly improved the sensitivity and PPV while maintaining similar specificity. This algorithm could be used to enhance existing ICD databases, for health research and surveillance. Declarations Ethical Approval This study was approved by the Conjoint Health Research Ethics Board at the University of Calgary (REB19-0088). Competing interests The authors declare that they have no competing interests to disclosure. Authors’ contributions Jie Pan wrote the main manuscript text and conducted the study design and analysis. Zilong Zhang developed machine learning models and assisted with writing. Steven Ray Peters provided subject matter expertise and insights about the discussion. Shabnam Vatanpour assisted with the literature review and writing. Robin L. Walker provided subject matter expertise and assisted with the writing. Seungwon Lee and Elliot A. Martin assisted with the result analysis and writing. Hude Quan was responsible for the study design and provided the interpretation framework of experimental results. All authors reviewed the manuscript from the perspectives of soundness, completeness, and novelty. Funding This work was supported by a Canadian Institutes of Health Research Operating Project Grant (201809FDN-409926-FDN-CBBA-114817). Availability of data and materials The data sets analyzed in this study are not publicly available due to the risk of exposing idenficiable information contained within the clinical notes. Access to the data is restricted to those collaborate with the Centre for Health Informatics and Alberta Health Services. References C. P. Friedman, A. K. Wong, and D. Blumenthal, ‘Policy: Achieving a nationwide learning health system’, Sci Transl Med, vol. 2, no. 57, Nov. 2010, doi: 10.1126/SCITRANSLMED.3001456/ASSET/31DF6FBB-61EA-4899-8446-9F05C371B44A/ASSETS/GRAPHIC/257CM29-F1.JPEG . A. K. Bonkhoff and C. Grefkes, ‘Precision medicine in stroke: towards personalized outcome predictions using artificial intelligence’, Brain, vol. 145, no. 2, pp. 457–475, Apr. 2022, doi: 10.1093/BRAIN/AWAB439 . H. Quan et al., ‘Assessing validity of ICD-9-CM and ICD-10 administrative data in recording clinical conditions in a unique dually coded database’, Health Serv Res, vol. 43, no. 4, pp. 1424–1441, Aug. 2008, doi: 10.1111/j.1475-6773.2007.00822.x . W. W. Yim, M. Yetisgen, W. P. Harris, and W. K. Sharon, ‘Natural Language Processing in Oncology Review’, JAMA Oncology, vol. 2, no. 6. American Medical Association, pp. 797–804, Jun. 01, 2016. doi: 10.1001/jamaoncol.2016.0213 . A. Y. X. Yu et al., ‘Use and utility of administrative health data for stroke research and surveillance’, Stroke, vol. 47, no. 7, pp. 1946–1952, Jul. 2016, doi: 10.1161/STROKEAHA.116.012390 . C. S. Kruse, K. Kothman, K. Anerobi, and L. Abanaka, ‘Adoption Factors of the Electronic Health Record: A Systematic Review’, JMIR Med Inform 2016;4(2):e19 https://medinform.jmir.org/2016/2/e19 , vol. 4, no. 2, p. e5525, Jun. 2016, doi: 10.2196/MEDINFORM.5525 . S. Wu et al., ‘Deep learning in clinical natural language processing: a methodical review’, Journal of the American Medical Informatics Association, vol. 27, no. 3, pp. 457–470, Mar. 2020, doi: 10.1093/JAMIA/OCZ200 . S. Lee et al., ‘Electronic Medical Record–Based Case Phenotyping for the Charlson Conditions: Scoping Review’, JMIR Med Inform 2021;9(2):e23934 https://medinform.jmir.org/2021/2/e23934 , vol. 9, no. 2, p. e23934, Feb. 2021, doi: 10.2196/23934 . W. Guan et al., ‘Automated Electronic Phenotyping of Cardioembolic Stroke’, Stroke, vol. 52, no. 1, pp. 181–189, Jan. 2021, doi: 10.1161/STROKEAHA.120.030663 . R. Garg, E. Oh, A. Naidech, K. Kording, and S. Prabhakaran, ‘Automating Ischemic Stroke Subtype Classification Using Machine Learning and Natural Language Processing’, Journal of Stroke and Cerebrovascular Diseases, vol. 28, no. 7, pp. 2045–2051, Jul. 2019, doi: 10.1016/J.JSTROKECEREBROVASDIS.2019.02.004 . S. F. Sung, C. Y. Lin, and Y. H. Hu, ‘EMR-Based Phenotyping of Ischemic Stroke Using Supervised Machine Learning and Text Mining Techniques’, IEEE J Biomed Health Inform, vol. 24, no. 10, pp. 2922–2931, Oct. 2020, doi: 10.1109/JBHI.2020.2976931 . V. M. Castro et al., ‘Large-scale identification of patients with cerebral aneurysms using natural language processing’, Neurology, vol. 88, no. 2, p. 164, Jan. 2017, doi: 10.1212/WNL.0000000000003490 . S. Bacchi, L. Oakden-Rayner, T. Zerner, T. Kleinig, S. Patel, and J. Jannes, ‘Deep learning natural language processing successfully predicts the cerebrovascular cause of transient ischemic attack-like presentations’, Stroke, vol. 50, no. 3, pp. 758–760, Mar. 2019, doi: 10.1161/STROKEAHA.118.024124 . M. I. Miller et al., ‘Natural Language Processing of Radiology Reports to Detect Complications of Ischemic Stroke’, Neurocrit Care, vol. 37, no. 2, pp. 291–302, Aug. 2022, doi: 10.1007/S12028-022-01513-3/FIGURES/3 . C. A. Eastwood, D. A. Southern, S. Khair, C. Doktorchik, W. A. Ghali, and H. Quan, ‘The ICD-11 field trial: Creating a large dually coded database’, Research Square Preprint, 2021. S. Lee et al., ‘Unlocking the Potential of Electronic Health Records for Health Research’, Int J Popul Data Sci, vol. 5, no. 1, Jan. 2020, doi: 10.23889/IJPDS.V5I1.1123 . H. Quan, M. Smith, G. Bartlett-Esquilant, H. Johansen, K. Tu, and L. Lix, ‘Mining Administrative Health Databases to Advance Medical Science: Geographical Considerations and Untapped Potential in Canada’, Canadian Journal of Cardiology, vol. 28, no. 2, pp. 152–154, Mar. 2012, doi: 10.1016/j.cjca.2012.01.005 . K. P. Liao et al., ‘Development of phenotype algorithms using electronic medical records and incorporating natural language processing’, BMJ, vol. 350, Apr. 2015, doi: 10.1136/BMJ.H1885 . F. Bagherzadeh-Khiabani, A. Ramezankhani, F. Azizi, F. Hadaegh, E. W. Steyerberg, and D. Khalili, ‘A tutorial on variable selection for clinical prediction models: feature selection methods in data mining could improve the results’, J Clin Epidemiol, vol. 71, pp. 76–85, Mar. 2016, doi: 10.1016/J.JCLINEPI.2015.10.002 . G. H. John, R. Kohavi, and K. Pfleger, ‘Irrelevant Features and the Subset Selection Problem’, Machine Learning Proceedings 1994, pp. 121–129, Jan. 1994, doi: 10.1016/B978-1-55860-335-6.50023-4 . M. Neumann, D. King, I. Beltagy, and W. Ammar, ‘ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing’, BioNLP 2019 - SIGBioMed Workshop on Biomedical Natural Language Processing, Proceedings of the 18th BioNLP Workshop and Shared Task, pp. 319–327, Feb. 2019, doi: 10.18653/v1/W19-5034 . O. Bodenreider, ‘The Unified Medical Language System (UMLS): integrating biomedical terminology’, Nucleic Acids Res, vol. 32, no. suppl_1, pp. D267–D270, Jan. 2004, doi: 10.1093/NAR/GKH061 . L. Breiman, ‘Random forests’, Mach Learn, vol. 45, no. 1, pp. 5–32, Oct. 2001, doi: 10.1023/A:1010933404324/METRICS . T. Chen and C. Guestrin, ‘XGBoost: A Scalable Tree Boosting System’, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, doi: 10.1145/2939672 . Z. Xu, G. Huang, K. Q. Weinberger, and A. X. Zheng, ‘Gradient boosted feature selection’, Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 522–531, 2014, doi: 10.1145/2623330.2623635 . Y. Qi, ‘Random Forest for Bioinformatics’, Ensemble Machine Learning, pp. 307–323, 2012, doi: 10.1007/978-1-4419-9326-7_11 . H. Quan et al., ‘International variation in the definition of “main condition” in ICD-coded health data’, International Journal for Quality in Health Care, vol. 26, no. 5, pp. 511–515, Oct. 2014, doi: 10.1093/INTQHC/MZU064 . Additional Declarations No competing interests reported. Supplementary Files Appendix.docx Cite Share Download PDF Status: Published Journal Publication published 02 Sep, 2023 Read the published version in Brain Informatics → Version 1 posted Editorial decision: Major revision 16 May, 2023 Reviews received at journal 15 May, 2023 Reviewers agreed at journal 14 May, 2023 Reviewers invited by journal 06 Mar, 2023 Editor assigned by journal 02 Mar, 2023 Submission checks completed at journal 02 Mar, 2023 First submitted to journal 28 Feb, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2640617","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":180286631,"identity":"9c8a3306-9f80-4425-b07b-0138eac956d2","order_by":0,"name":"Jie Pan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIie2OOwrCQBCGRxaSZtU2hY8rbMgBxJtsmqQxELC1WBBip625RbyBMpA0OUAKEUWwjiBiJbqIltnYCe5XzAPm4x8AjeYHaUxfvSfn9VeKI0s95Y0raitkRtYknGz9JMP95jLZucLEveIxg5NlegqS3GPYSceuoB5TKJQRamCQFMDQMrgrLFAp7ZLQO/qsMEu07lIxS1UKkGaEnBWUbc6RVKgqxWAYz9GO81GIjTl3IjoKKxV7gYdjeMV+K8tWh9uVdxdmllQrAgDfC6EAw6jy/klfXn6+vAEMVIZGo9H8Hw9eGktirs6RRAAAAABJRU5ErkJggg==","orcid":"","institution":"University of Calgary","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jie","middleName":"","lastName":"Pan","suffix":""},{"id":180286634,"identity":"cfd1a41f-7fe0-491b-8b56-06fda45a0c86","order_by":1,"name":"Zilong Zhang","email":"","orcid":"","institution":"University of Calgary","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zilong","middleName":"","lastName":"Zhang","suffix":""},{"id":180286637,"identity":"f117a263-4571-4f06-afaf-ea55deb9873d","order_by":2,"name":"Steven Ray Peters","email":"","orcid":"","institution":"University of Calgary","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Steven","middleName":"Ray","lastName":"Peters","suffix":""},{"id":180286640,"identity":"b54d2106-0f8f-4fb2-9c5b-4372b972c81a","order_by":3,"name":"Shabnam Vatanpour","email":"","orcid":"","institution":"University of Calgary","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shabnam","middleName":"","lastName":"Vatanpour","suffix":""},{"id":180286644,"identity":"94c4463c-8b31-434d-8254-27b2621b3847","order_by":4,"name":"Robin L. Walker","email":"","orcid":"","institution":"Alberta Health Services","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Robin","middleName":"L.","lastName":"Walker","suffix":""},{"id":180286646,"identity":"607cd908-f174-4356-b957-35b7bbff1edd","order_by":5,"name":"Seungwon Lee","email":"","orcid":"","institution":"University of Calgary","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Seungwon","middleName":"","lastName":"Lee","suffix":""},{"id":180286650,"identity":"81acf23a-f30b-4f18-96b7-c68ac1edc95f","order_by":6,"name":"Elliot A. Martin","email":"","orcid":"","institution":"University of Calgary","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Elliot","middleName":"A.","lastName":"Martin","suffix":""},{"id":180286652,"identity":"bc3c74a4-1a37-4c00-8e61-ca83118e0c08","order_by":7,"name":"Hude Quan","email":"","orcid":"","institution":"University of Calgary","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hude","middleName":"","lastName":"Quan","suffix":""}],"badges":[],"createdAt":"2023-02-28 23:59:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2640617/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2640617/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s40708-023-00203-w","type":"published","date":"2023-09-02T15:09:14+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":33984682,"identity":"31194ddb-757d-4739-ba98-35d60a5382d9","added_by":"auto","created_at":"2023-03-08 23:52:16","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":197480,"visible":true,"origin":"","legend":"\u003cp\u003eNLP-based CeVD detection framework using EMR data. It consists of manual chart review, data preprocessing, featurization, and model training, development, and validation. FPR represents three acute care facilities in Calgary, Foothills Medical Centre, Peter Lougheed Centre and Rockyview General Hospital; NER represents Named Entity Recognition (NER), a subtask of NLP that seeks to identify named objects from free-text; CUI represents Concept Unique Identifiers which map synonyms to a unique identifier; BOW represents bag of words; TF-IDF represents term frequency and inverse document frequency.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-2640617/v1/7c488d6e867627762c409f99.png"},{"id":33984684,"identity":"65d1fbb1-13a4-4b0b-944a-5eaa4b37a20b","added_by":"auto","created_at":"2023-03-08 23:52:16","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":218822,"visible":true,"origin":"","legend":"\u003cp\u003eDocument selection and featuring process based on the developed NLP models. All the document types are from 3036 patients’ clinical notes during hospitalization.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-2640617/v1/46bce20ddec6748e9c36e027.png"},{"id":33985643,"identity":"10405f90-3ed4-4613-b4c0-3bb32754e551","added_by":"auto","created_at":"2023-03-09 00:00:16","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":298777,"visible":true,"origin":"","legend":"\u003cp\u003eTop 10 key concepts for detecting CeVD in each selected document type. The concepts were UMLS terms extracted by cTAKES. The impurity-based feature importance measured the importance of classifying CeVD.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-2640617/v1/175bb4a83188bb72194053b0.png"},{"id":33984683,"identity":"35ee39e1-a14b-4cb0-a029-84c1ba0b037c","added_by":"auto","created_at":"2023-03-08 23:52:16","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":84755,"visible":true,"origin":"","legend":"\u003cp\u003ePPV, NPV, sensitivity, and specificity of the four NLP models and ICD algorithm, with changing thresholds ranging between 0.05 to 0.95. The two dashed lines in each subfigure represent the 0.05 and 0.95 threshold bounds, respectively. TFIDF-CUI-RF represents algorithm “CUI + TF-IDF + RF”; WC-CUI-XGBoost represents algorithm “CUI + word count + XGBoost”; TFIDF-BOW-XGBoost represents algorithm “BOW +TF-IDF + XGBoost”; TFIDF-CUI-XGBoost represents algorithm “CUI + TF-IDF + XGBoost”; ICD represents the ICD-10-CA-codes in DAD algorithms, respectively.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-2640617/v1/134d90516983486d496e2e28.png"},{"id":42781934,"identity":"1f86806a-6b9d-4b24-8cf7-da80fc0bf1f8","added_by":"auto","created_at":"2023-09-07 15:15:01","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1116149,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2640617/v1/ba8832d3-b3e6-4e4f-b42d-858080b31859.pdf"},{"id":33985642,"identity":"dc2d1a8c-b9c4-45be-89d3-42c3a8b9a24f","added_by":"auto","created_at":"2023-03-09 00:00:16","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":38031,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix.docx","url":"https://assets-eu.researchsquare.com/files/rs-2640617/v1/2d85a429650b39cc80a2773d.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Cerebrovascular disease case identification in inpatient electronic medical record data using natural language processing","fulltext":[{"header":"Introduction","content":"\u003cp\u003eAccurate identification of patients with cerebrovascular diseases (CeVD) is important for health services research, surveillance and monitoring, risk adjustment, and quality improvement measurement [1, 2]. The standard approach to identify conditions is coded administrative hospital data using International Classification of Disease (ICD) terminology. Although structured codes are widely available and highly standardized,\u0026nbsp;some conditions, including CeVD, are under-coded.\u0026nbsp;Quan et al. [3] validated the ICD algorithms against chart review and reported a sensitivity of 46.3% for detecting CeVD diseases in both ICD-9 and ICD-10-CA. To overcome the shortcomings of ICD code-based algorithms, medical chart reviews act as a gold standard for case identification. Unfortunately, chart review is time- and resource-intensive requiring health professionals familiar with specific conditions [4, 5].\u003c/p\u003e\n\u003cp\u003eElectronic medical records (EMRs) are becoming increasingly popular for collecting health information\u0026nbsp;[6], and can be used to improve the accuracy of identifying conditions such as CeVD.\u0026nbsp;Among the components of EMR, free text notes contain detailed descriptions and give health professionals great flexibility to report conditions and comorbidities. Natural Language Processing (NLP) is an artificial intelligence technique to analyze human languages and retrieve clinically relevant information for detecting and predicting medical conditions [7]. A recent literature conducted by our team yielded few studies using NLP on clinical notes for patients with CeVD conditions [8]. Existing studies have focused on identifying ischemic stroke [9\u0026ndash;11] and cerebral aneurysms [12], predicting the cerebrovascular causes of ischemia [13], and detecting complications of stroke [14]. Most previous studies focus on specific conditions within CeVD and have limited access to a complete set of clinical notes from EMRs, using only admission notes or radiology reports.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn this study, we explored all available types of inpatient clinical notes from an EMR to identify a broad spectrum of CeVD cases. The CeVD cases were defined by our previous ICD-10 algorithm [3]. We hypothesized that using NLP techniques on these clinical notes would better detect CeVD cases than ICD-based algorithms and existing ML algorithms with limited data source types.\u003c/p\u003e"},{"header":"Methods","content":"\u003ch2\u003e\u003cstrong\u003eStudy Population\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eIn this retrospective cohort study, we randomly selected patients who were at least 18 years of age and discharged\u0026nbsp;from three acute care facilities in Calgary, Canada, between January 1 and June 30, 2015. Obstetric admissions were excluded because they have a short length of stay and lack conditions of interest. We randomly selected one hospitalization per patient if multiple discharges occurred during the study period [15]. Six nurses reviewed charts to determine the existence of CeVD [15].\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eData sources\u003c/strong\u003e\u003c/h2\u003e\n\u003ch3\u003eEMR: Sunrise Clinical Manager (SCM)\u003c/h3\u003e\n\u003cp\u003eThe EMR data are from SCM, a city-wide, population level EMR system used in the three acute care hospitals in Calgary. SCM provides patient-level clinical information containing medical and nursing orders, medication records, clinical documentation, diagnostic imaging and lab results [16].\u003c/p\u003e\n\u003ch3\u003eAdministrative Discharge Abstract Database: DAD\u003c/h3\u003e\n\u003cp\u003eThe inpatients\u0026rsquo; administrative, clinical, and demographic information at the time of discharge is coded in the DAD [17]. The clinical coder records up to 25 diagnostics codes for each inpatient based on available information from patient charts. The DAD, EMR data and chart data were linked with Personal Health Number (a unique lifetime identifier), chart number (a distinctive number associated with a patient\u0026rsquo;s admission), and admission date.\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003ePhenotyping Algorithm Framework\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eWe trained, validated, and tested an EMR data-driven phenotyping algorithm using NLP techniques to detect CeVD. NLP techniques are used to process and analyze human language, and contain a wide range of tasks, including named entity recognition (NER), information extraction, and text classification [11, 16]. They were applied to analyze the free text clinical notes and derive a CeVD phenotype to detect the disease automatically. As depicted in Figure 1, the general framework consists of 1) input document selection from patients\u0026rsquo; clinical notes, 2) model training, and 3) performance evaluation using chart review as a reference standard.\u003c/p\u003e\n\u003ch3\u003eDocument selection and feature engineering\u003c/h3\u003e\n\u003cp\u003eMany types of clinical notes could be generated during the hospitalization of patients involved in this study, such as nursing transfer reports, inpatient consultations, discharge summaries, and surgical assessment and history. However, not all document types contribute equally to the detection of CeVD. Noise and redundant information can hamper the detection performance of ML models [19]. The first step is determining and selecting the appropriate document type(s) sensitive to CeVD identification.\u003c/p\u003e\n\u003cp\u003eThe method we used is a feedforward sequential selection method [20], to iteratively add the document type that contributes most to model performance, until the performance stops increasing or reaches a predefined criterion, as shown in Figure 2. All the documents are first converted into vectors by 1) extracting relevant medical concepts from the text and 2) turning concepts into numeric features. To examine the extraction performance, we compared two types of commonly used concept extraction methods: Bag of Words (BOW) using ScispaCy [21] and Concept Unique Identifiers (CUIs) from the Unified Medical Language System using cTAKES (see Appendix 1 for a detailed explanation) [22]. We also compared two types of feature construction methods: Term Frequency-Inverse Document Frequency (TF-IDF) and word count. The obtained vectors are fed into the ML models and validated by the model performance. To estimate better generalization of the selected document types, 5-fold cross validation was applied to the selected patients (i.e., 80% training, n= 2429 and 20% test, n=607). The model development is detailed in the following section.\u003c/p\u003e\n\u003ch3\u003eModel development\u003c/h3\u003e\n\u003cp\u003eThe model outcome is a binary classification where hospitalized patients with CeVD are considered positive cases. Two supervised ML methods were trained, validated, and tested using the obtained input vectors and chart review output labels, including random forest (RF) and XGBoost [19, 20]. The two methods are known for handling datasets with high dimensionality, missing data and outliers, and providing accurate and reliable predictions, especially for NLP tasks containing thousands of concept features [21, 22].\u003c/p\u003e\n\u003cp\u003eWith the different combinations among methods of concept extraction, vectorization, and ML models, we have 8 model variations, such as \u0026ldquo;BOW + TF-IDF + RF\u0026rdquo; and \u0026ldquo;CUI + TF-IDF + XGBoost.\u0026rdquo; As both methods, RF and XGBoost, use decision trees as the base models, we assigned 100 decision trees to them, respectively. These models\u0026rsquo; performance was then estimated by 5-fold cross validation, maintaining the same proportion of positive and negative patients in each group.\u003c/p\u003e\n\u003ch3\u003ePerformance metrics\u003c/h3\u003e\n\u003cp\u003eTo evaluate and compare the models developed, we calculated their sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and F1 score using chart data as a reference standard. We also calculated binomial proportional confidence intervals for all the metrics.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWe compare the results with ICD-based CeVD identification algorithms in DAD after defining CeVD using ICD-10 codes (e.g., G45-46, I60-69, H34, see Table S1 in Appendix) [3]. The performance metrics of the developed NLP models were reported on the same level of specificity as with the ICD-based algorithm.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv class=\"Section2\" id=\"Sec11\"\u003e\n \u003ch2\u003eCharacteristics of the study cohort\u003c/h2\u003e\n \u003cp\u003eAmong the 3036 patients, chart reviewers identified 360 patients with CeVD (see Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). Characteristics that were statistically significantly different (P\u0026thinsp;\u0026lt;\u0026thinsp;.05) between the CeVD positive cohort and negative cohort are: age, comorbidities such as atrial fibrillation, angina, hypertension, peripheral vascular disease (PVD) and obesity\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003ctable border=\"1\" id=\"Tab1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003ePatients Characteristics.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eCharacteristics\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAll (percentage)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePatients with CEVD (percentage)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePatients without CEVD (percentage)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eP\u003c/em\u003e value\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eN =\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e3036 (100%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e360(11.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2676(88.1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"3\"\u003e\n \u003cp\u003e\u003cstrong\u003eDemographic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eMedian of Age (IQR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e63.0\u003c/p\u003e\n \u003cp\u003e(48.9\u0026ndash;76.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e77.4\u003c/p\u003e\n \u003cp\u003e(67.0-85.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e60.9\u003c/p\u003e\n \u003cp\u003e(46.4\u0026ndash;74.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eFemale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1528 (50.3%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e175 (48.6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1353(50.6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"3\"\u003e\n \u003cp\u003e\u003cstrong\u003eComorbidities\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eAtrial fibrillation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e370 (12.2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e106 (29.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e264 (9.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eAngina\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e203 (6.7%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e41 (11.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e162 (6.1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0002\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eMyocardial infarction\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e102 (3.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e18 (5.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e84 (3.1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.06\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eHypertension\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1469 (48.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e267 (74.2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1202 (44.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003ePeripheral vascular disease\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e148 (4.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e46 (12.8%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e102 (3.8%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eObesity\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e736 (24.2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e68 (18.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e668 (25.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.01\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eAlcohol abuse\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e230 (7.6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e19 (5.3%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e211 (7.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eSmoking\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e605 (19.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e66 (18.3%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e539 (20.1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.4\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ctfoot\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"7\"\u003eIQR\u0026thinsp;=\u0026thinsp;Interquartile range\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tfoot\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv class=\"Section2\" id=\"Sec12\"\u003e\n \u003ch2\u003eCharacteristics of selected document types\u003c/h2\u003e\n \u003cp\u003eWe collected 49 types of clinical documents of patients during hospitalization, such as nursing transfer reports, inpatient consultations, and discharge summaries. The detailed text statistics for these document types can be found in \u003cspan class=\"InternalRef\"\u003eAppendix\u003c/span\u003e Table S2. For a better explanation, we consolidated these document types into 9 categories (see Table S3 in \u003cspan class=\"InternalRef\"\u003eAppendix\u003c/span\u003e).\u003c/p\u003e\n \u003cp\u003eUsing the feedforward sequential selection, we identified four essential document types, \u0026ldquo;nursing transfer report,\u0026rdquo; \u0026ldquo;discharge summary,\u0026rdquo; \u0026ldquo;nursing notes,\u0026rdquo; and \u0026ldquo;inpatient consultation.\u0026rdquo; These documents are sensitive and informative for CeVD detection. Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e shows the statistics of patients, documents, and words. At least 90% of patients (with or without CeVD) have at least 2 types of documents. These four types of documents complement each other in providing sufficient clinical information to identify CeVD.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003ctable border=\"1\" id=\"Tab2\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eCharacteristics of extracted documents.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eDocument type\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAll (n\u0026thinsp;=\u0026thinsp;3036)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePatients with CeVD (n\u0026thinsp;=\u0026thinsp;360)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePatients without CeVD (n\u0026thinsp;=\u0026thinsp;2676)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMedian number of notes per patient (IQR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.0 (1.0\u0026ndash;2.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.0 (1.0\u0026ndash;2.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.0 (1.0\u0026ndash;2.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNumber of patients with at least 2 types of documents (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2774 (91.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e344 (95.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2430 (90.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMedian word count per note (IQR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e430.0\u003c/p\u003e\n \u003cp\u003e(310.0-678.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e434.5\u003c/p\u003e\n \u003cp\u003e(322.2\u0026ndash;723.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e428.0\u003c/p\u003e\n \u003cp\u003e(308.0-675.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eDetailed document types: nursing transfer report - emergency department to inpatient, discharge summary-medical; surgical assessment and history, inpatient consultations, and discharge summary.\u003c/p\u003e\n \u003cp\u003eTo examine how these document types contribute to CeVD detection, the top ten key concepts in each document type were analyzed, as shown in Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. There are some common and vital concepts across four document types, such as \u0026ldquo;C0038454\u0026rdquo; (stroke-related concepts) and \u0026ldquo;C0007787\u0026rdquo; (transient ischemic attack). It is reasonable that the existence of these concepts can directly reflect the CeVD status. The remaining concepts are less overlapped and unique to each document type, such as \u0026ldquo;C0012169\u0026rdquo; (low sodium diet) in \u0026ldquo;nursing notes,\u0026rdquo; \u0026ldquo;C0004134\u0026rdquo; (ataxia) in \u0026ldquo;nursing transfer report,\u0026rdquo; \u0026ldquo;C0202691\u0026rdquo; (CAT scan of head) in \u0026ldquo;discharge summary,\u0026rdquo; and \u0026ldquo;C0001962\u0026rdquo; (ethanol) in \u0026ldquo;inpatient consultation.\u0026rdquo; This demonstrated that these document types contain essential concepts and can supplement each other to gain more comprehensive information in CeVD detection.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv class=\"Section2\" id=\"Sec13\"\u003e\n \u003ch2\u003eClassification performance\u003c/h2\u003e\n \u003cp\u003eThe top 4 trained models were shown in Table \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. XGBoost generally outperformed the random forest method. TF-IDF performed better than term count when comparing models \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;word count\u0026thinsp;+\u0026thinsp;XGBoost\u0026rdquo; and \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost.\u0026rdquo; Similarly, the concept extraction method \u0026ldquo;CUI\u0026rdquo; had better performance than \u0026ldquo;Bag of Words (BOW).\u0026rdquo; Consequently, the combination of XGBoost, TF-IDF, and CUI achieved the best performance over other ML models in the metrics of sensitivity (70%), specificity (99.1%), PPV (87.8%), NPV (97.1%), F1 (77.8%), and accuracy (96.5%).\u003c/p\u003e\n \u003cp\u003eWe also compared the model performance with ICD-10-CA-based methods. With similar specificity (99.3% in ICD-10-CA vs 99.1% in model \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost\u0026rdquo;), the performance in other metrics is improved hugely by the obtained model, such as sensitivity increased from 25.0\u0026ndash;70.0%, and F1 increased from 38.4\u0026ndash;77.8%.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003ctable border=\"1\" id=\"Tab3\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eCeVD case identification with DAD and EMR.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eModel\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSensitivity%\u003c/p\u003e\n \u003cp\u003e(95% CI)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSpecificity%\u003c/p\u003e\n \u003cp\u003e(95% CI)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePPV%\u003c/p\u003e\n \u003cp\u003e(95% CI)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eNPV%\u003c/p\u003e\n \u003cp\u003e(95% CI)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eF1%\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAccuracy%\u003c/p\u003e\n \u003cp\u003e(95% CI)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eICD-10-CA-codes\u0026nbsp;in DAD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e25.0\u003c/p\u003e\n \u003cp\u003e(20.6\u0026ndash;29.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e99.3\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(98.9\u0026ndash;99.6)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e82.6\u003c/p\u003e\n \u003cp\u003e(74.5\u0026ndash;88.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e90.8\u003c/p\u003e\n \u003cp\u003e(90.3\u0026ndash;91.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e38.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e90.5\u003c/p\u003e\n \u003cp\u003e(89.4\u0026ndash;91.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;RF\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e65.8\u003c/p\u003e\n \u003cp\u003e(60.7\u0026ndash;70.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e98.5\u003c/p\u003e\n \u003cp\u003e(98.0\u0026ndash;99.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e85.9\u003c/p\u003e\n \u003cp\u003e(81.5\u0026ndash;89.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e95.5\u003c/p\u003e\n \u003cp\u003e(94.9\u0026ndash;96.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e74.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e94.7\u003c/p\u003e\n \u003cp\u003e(93.8\u0026ndash;95.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCUI\u0026thinsp;+\u0026thinsp;word count\u0026thinsp;+\u0026thinsp;XGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e68.1\u003c/p\u003e\n \u003cp\u003e(63.0-72.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e98.6\u003c/p\u003e\n \u003cp\u003e(98.1\u0026ndash;99.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e86.9\u003c/p\u003e\n \u003cp\u003e(82.7\u0026ndash;90.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e95.8\u003c/p\u003e\n \u003cp\u003e(95.2\u0026ndash;96.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e76.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e95.0\u003c/p\u003e\n \u003cp\u003e(94.2\u0026ndash;95.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eCUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost*\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e70.00\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(65.0-74.7)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e99.1\u003c/p\u003e\n \u003cp\u003e(98.7\u0026ndash;99.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e87.8\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(83.7\u0026ndash;91.0)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e97.1\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(96.6\u0026ndash;97.5)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e77.8\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e96.5\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(95.8\u0026ndash;97.0)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eBOW +\u0026nbsp;TF-IDF +\u003c/p\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e59.2\u003c/p\u003e\n \u003cp\u003e(53.9\u0026ndash;64.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e98.7\u003c/p\u003e\n \u003cp\u003e(98.1\u0026ndash;99.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e85.5\u003c/p\u003e\n \u003cp\u003e(80.9\u0026ndash;89.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e94.7\u003c/p\u003e\n \u003cp\u003e(94.1\u0026ndash;95.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e69.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e94.0\u003c/p\u003e\n \u003cp\u003e(93.1\u0026ndash;94.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eWe included the four metrics of the four NLP models with changing threshold values from 0.05 to 0.95, as shown in Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e. Since the ICD algorithm is deterministic, its threshold is not changeable. The PPVs of \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;RF,\u0026rdquo; \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;word count\u0026thinsp;+\u0026thinsp;XGBoost,\u0026rdquo; \u0026ldquo;BOW\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost,\u0026rdquo; and \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost\u0026rdquo; started to exceed the performance of ICD at thresholds 0.32, 0.25, 0.28, and 0.17 within the threshold bound, respectively. \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;word count\u0026thinsp;+\u0026thinsp;XGBoost\u0026rdquo; and \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost\u0026rdquo; had very similar and robust performance with the change of thresholds, whereas \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;RF\u0026rdquo; was affected significantly. Generally, the \u0026ldquo;CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost\u0026rdquo; algorithm achieved better and more robust performance with smaller thresholds.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis paper shows that EMR textual information abstracted by NLP techniques outperforms traditional ICD codes for assessing cerebrovascular disease, and compares favourably with resource-intensive chart review using a fraction of human resources. With the prevalence of 11.8% CeVD in over 3000 records, the developed NLP model significantly improves the validity of DAD-based ICD algorithm (sensitivity: 70% vs. 25% and PPV: 88% vs. 83%).\u003c/p\u003e \u003cp\u003eEMR data is more informative and efficient in identifying CeVD patients than conventionally used hospitalization data (i.e., DAD). First, due to the high volume of discharges, coders have limited time to code patients comprehensively, causing missing codes and low quality. Second, there is no uniform international definition of the most responsible diagnosis, which varies between the primary reason for admission and the condition with intensive resource usage [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. When looking for conditions contributing primarily to the length of stay in hospital (a Canada-wide used definition), CeVD is likely under-coded as it can be a comorbidity causing admission. Conversely, EMRs contain many documents not usually used by medical coders. As identified in this study, four types of documents (i.e., \u0026ldquo;nursing transfer report,\u0026rdquo; \u0026ldquo;discharge summary,\u0026rdquo; \u0026ldquo;nursing notes,\u0026rdquo; and \u0026ldquo;inpatient consultation\u0026rdquo;) jointly contribute to the accurate detection of CeVD by providing more comprehensive medical information. Restricting the analysis to a specific document type, therefore, has the potential to impede detection.\u003c/p\u003e \u003cp\u003eTo abstract the knowledge from these EMRs textual data, NLP techniques are essential. Information extraction from unstructured text is known to be difficult, and contains subtasks including NER, relation extraction, and pattern extraction. The text-based classification assigns categorical labels for a text fragment by finding the patterns composed of NERs and their relationships. By comparing different combinations of NLP models, we identified the optimal model, CUI\u0026thinsp;+\u0026thinsp;TF-IDF\u0026thinsp;+\u0026thinsp;XGBoost. The TF-IDF performs better than word count because it can efficiently eliminate low-sensitive concepts in differentiating positive and negative groups. CUI is a better concept extraction method than the NER by scispaCy because cTAKES can merge similar concepts into one, such as \u0026ldquo;stroke,\u0026rdquo; \u0026ldquo;CVA,\u0026rdquo; and \u0026ldquo;brain vascular accidents\u0026rdquo; are mapped to the same CUI \u0026ldquo;C0038454\u0026rdquo;. Sine XGBoost has better capability in dealing with overfitting and allows a more general model than random forest, it shows a slightly better performance in detecting CeVD, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eThe widespread use of text based EMR algorithms to supplement ICD codes and traditional chart reviews has many potential advantages for epidemiology and health outcomes research. CeVD status is frequently used as an important factor in stratifying outcomes in population health research. While some outcomes, such as ischemic stroke, have reasonable validity, other aspects of CeVD, such as carotid atherosclerosis, are likely poorly coded. This probably explains the poor sensitivity (25%) of ICD codes for CeVD in our study. We achieved 88% PPV and 70% sensitivity, an improvement over the widely adopted ICD-based algorithm. Given the amount of knowledge contained in clinical text, the algorithm is applicable to detecting many other diseases, especially conditions with under-coding issues. Text based EMR algorithms may be used to periodically re-evaluate the validity of existing ICD code-based approaches and ensure that ICD code validity is not changing over time.\u003c/p\u003e \u003cp\u003eLimitations\u003c/p\u003e \u003cp\u003eThere are some limitations in this study. First, further examination of missing cases is needed, as 30% of cases are still missed by the proposed algorithm using EMR data. The missing cases are likely caused by variations in clinical documents and the capability of NLP models to detect them. We believe that the performance of the NLP models can be further improved by having better NER and incorporating sequential and contextual patterns among recognized concepts. Second, the data we studied is only from one city (i.e., Calgary). EMR diversities in format and content could be subject to change when larger populations and geographies are considered. The identified sensitive document types will vary accordingly. Lastly, we did not validate the algorithms in external databases. We encourage researchers to apply this method to their datasets for validation and improvement.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eCompared to the widely used ICD-based algorithm, the EMR NLP model significantly improved the sensitivity and PPV while maintaining similar specificity. This algorithm could be used to enhance existing ICD databases, for health research and surveillance.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthical Approval\u003c/h2\u003e\n\u003cp\u003eThis study was approved by the Conjoint Health Research Ethics Board at the University of Calgary (REB19-0088).\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests to disclosure.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eAuthors\u0026rsquo; contributions\u003c/h2\u003e\n\u003cp\u003eJie Pan wrote the main manuscript text and conducted the study design and analysis. Zilong Zhang developed machine learning models and assisted with writing. Steven Ray Peters provided subject matter expertise and insights about the discussion. Shabnam Vatanpour assisted with the literature review and writing. Robin L. Walker provided subject matter expertise and assisted with the writing. Seungwon Lee and Elliot A. Martin assisted with the result analysis and writing. Hude Quan was responsible for the study design and provided the interpretation framework of experimental results. All authors reviewed the manuscript from the perspectives of soundness, completeness, and novelty.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis \u0026nbsp;work \u0026nbsp; was \u0026nbsp;supported \u0026nbsp;by \u0026nbsp;a \u0026nbsp;Canadian \u0026nbsp; Institutes \u0026nbsp;of \u0026nbsp;Health \u0026nbsp; Research Operating Project Grant (201809FDN-409926-FDN-CBBA-114817).\u003c/p\u003e\n\u003ch2\u003eAvailability of data and materials\u003c/h2\u003e\n\u003cp\u003eThe data sets analyzed in this study are not publicly available due to the risk of exposing idenficiable information contained within the clinical notes. Access to the data is restricted to those collaborate with the Centre for Health Informatics and Alberta Health Services.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eC. P. Friedman, A. K. Wong, and D. Blumenthal, \u0026lsquo;Policy: Achieving a nationwide learning health system\u0026rsquo;, Sci Transl Med, vol.\u0026nbsp;2, no. 57, Nov. 2010, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1126/SCITRANSLMED.3001456/ASSET/31DF6FBB-61EA-4899-8446-9F05C371B44A/ASSETS/GRAPHIC/257CM29-F1.JPEG\u003c/span\u003e\u003cspan address=\"10.1126/SCITRANSLMED.3001456/ASSET/31DF6FBB-61EA-4899-8446-9F05C371B44A/ASSETS/GRAPHIC/257CM29-F1.JPEG\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA. K. Bonkhoff and C. Grefkes, \u0026lsquo;Precision medicine in stroke: towards personalized outcome predictions using artificial intelligence\u0026rsquo;, Brain, vol.\u0026nbsp;145, no. 2, pp.\u0026nbsp;457\u0026ndash;475, Apr. 2022, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/BRAIN/AWAB439\u003c/span\u003e\u003cspan address=\"10.1093/BRAIN/AWAB439\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eH. Quan et al., \u0026lsquo;Assessing validity of ICD-9-CM and ICD-10 administrative data in recording clinical conditions in a unique dually coded database\u0026rsquo;, Health Serv Res, vol.\u0026nbsp;43, no. 4, pp.\u0026nbsp;1424\u0026ndash;1441, Aug. 2008, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/j.1475-6773.2007.00822.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1475-6773.2007.00822.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eW. W. Yim, M. Yetisgen, W. P. Harris, and W. K. Sharon, \u0026lsquo;Natural Language Processing in Oncology Review\u0026rsquo;, JAMA Oncology, vol.\u0026nbsp;2, no. 6. American Medical Association, pp.\u0026nbsp;797\u0026ndash;804, Jun. 01, 2016. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamaoncol.2016.0213\u003c/span\u003e\u003cspan address=\"10.1001/jamaoncol.2016.0213\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eA. Y. X. Yu et al., \u0026lsquo;Use and utility of administrative health data for stroke research and surveillance\u0026rsquo;, Stroke, vol.\u0026nbsp;47, no. 7, pp.\u0026nbsp;1946\u0026ndash;1952, Jul. 2016, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/STROKEAHA.116.012390\u003c/span\u003e\u003cspan address=\"10.1161/STROKEAHA.116.012390\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eC. S. Kruse, K. Kothman, K. Anerobi, and L. Abanaka, \u0026lsquo;Adoption Factors of the Electronic Health Record: A Systematic Review\u0026rsquo;, JMIR Med Inform 2016;4(2):e19 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://medinform.jmir.org/2016/2/e19\u003c/span\u003e\u003cspan address=\"https://medinform.jmir.org/2016/2/e19\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, vol.\u0026nbsp;4, no. 2, p. e5525, Jun. 2016, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/MEDINFORM.5525\u003c/span\u003e\u003cspan address=\"10.2196/MEDINFORM.5525\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Wu et al., \u0026lsquo;Deep learning in clinical natural language processing: a methodical review\u0026rsquo;, Journal of the American Medical Informatics Association, vol.\u0026nbsp;27, no. 3, pp.\u0026nbsp;457\u0026ndash;470, Mar. 2020, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/JAMIA/OCZ200\u003c/span\u003e\u003cspan address=\"10.1093/JAMIA/OCZ200\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Lee et al., \u0026lsquo;Electronic Medical Record\u0026ndash;Based Case Phenotyping for the Charlson Conditions: Scoping Review\u0026rsquo;, JMIR Med Inform 2021;9(2):e23934 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://medinform.jmir.org/2021/2/e23934\u003c/span\u003e\u003cspan address=\"https://medinform.jmir.org/2021/2/e23934\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, vol.\u0026nbsp;9, no. 2, p. e23934, Feb. 2021, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/23934\u003c/span\u003e\u003cspan address=\"10.2196/23934\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eW. Guan et al., \u0026lsquo;Automated Electronic Phenotyping of Cardioembolic Stroke\u0026rsquo;, Stroke, vol.\u0026nbsp;52, no. 1, pp.\u0026nbsp;181\u0026ndash;189, Jan. 2021, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/STROKEAHA.120.030663\u003c/span\u003e\u003cspan address=\"10.1161/STROKEAHA.120.030663\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eR. Garg, E. Oh, A. Naidech, K. Kording, and S. Prabhakaran, \u0026lsquo;Automating Ischemic Stroke Subtype Classification Using Machine Learning and Natural Language Processing\u0026rsquo;, Journal of Stroke and Cerebrovascular Diseases, vol.\u0026nbsp;28, no. 7, pp.\u0026nbsp;2045\u0026ndash;2051, Jul. 2019, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/J.JSTROKECEREBROVASDIS.2019.02.004\u003c/span\u003e\u003cspan address=\"10.1016/J.JSTROKECEREBROVASDIS.2019.02.004\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. F. Sung, C. Y. Lin, and Y. H. Hu, \u0026lsquo;EMR-Based Phenotyping of Ischemic Stroke Using Supervised Machine Learning and Text Mining Techniques\u0026rsquo;, IEEE J Biomed Health Inform, vol.\u0026nbsp;24, no. 10, pp.\u0026nbsp;2922\u0026ndash;2931, Oct. 2020, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/JBHI.2020.2976931\u003c/span\u003e\u003cspan address=\"10.1109/JBHI.2020.2976931\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eV. M. Castro et al., \u0026lsquo;Large-scale identification of patients with cerebral aneurysms using natural language processing\u0026rsquo;, Neurology, vol.\u0026nbsp;88, no. 2, p.\u0026nbsp;164, Jan. 2017, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1212/WNL.0000000000003490\u003c/span\u003e\u003cspan address=\"10.1212/WNL.0000000000003490\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Bacchi, L. Oakden-Rayner, T. Zerner, T. Kleinig, S. Patel, and J. Jannes, \u0026lsquo;Deep learning natural language processing successfully predicts the cerebrovascular cause of transient ischemic attack-like presentations\u0026rsquo;, Stroke, vol.\u0026nbsp;50, no. 3, pp.\u0026nbsp;758\u0026ndash;760, Mar. 2019, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/STROKEAHA.118.024124\u003c/span\u003e\u003cspan address=\"10.1161/STROKEAHA.118.024124\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. I. Miller et al., \u0026lsquo;Natural Language Processing of Radiology Reports to Detect Complications of Ischemic Stroke\u0026rsquo;, Neurocrit Care, vol.\u0026nbsp;37, no. 2, pp.\u0026nbsp;291\u0026ndash;302, Aug. 2022, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/S12028-022-01513-3/FIGURES/3\u003c/span\u003e\u003cspan address=\"10.1007/S12028-022-01513-3/FIGURES/3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eC. A. Eastwood, D. A. Southern, S. Khair, C. Doktorchik, W. A. Ghali, and H. Quan, \u0026lsquo;The ICD-11 field trial: Creating a large dually coded database\u0026rsquo;, Research Square Preprint, 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eS. Lee et al., \u0026lsquo;Unlocking the Potential of Electronic Health Records for Health Research\u0026rsquo;, Int J Popul Data Sci, vol.\u0026nbsp;5, no. 1, Jan. 2020, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.23889/IJPDS.V5I1.1123\u003c/span\u003e\u003cspan address=\"10.23889/IJPDS.V5I1.1123\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eH. Quan, M. Smith, G. Bartlett-Esquilant, H. Johansen, K. Tu, and L. Lix, \u0026lsquo;Mining Administrative Health Databases to Advance Medical Science: Geographical Considerations and Untapped Potential in Canada\u0026rsquo;, Canadian Journal of Cardiology, vol.\u0026nbsp;28, no. 2, pp.\u0026nbsp;152\u0026ndash;154, Mar. 2012, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cjca.2012.01.005\u003c/span\u003e\u003cspan address=\"10.1016/j.cjca.2012.01.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK. P. Liao et al., \u0026lsquo;Development of phenotype algorithms using electronic medical records and incorporating natural language processing\u0026rsquo;, BMJ, vol.\u0026nbsp;350, Apr. 2015, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/BMJ.H1885\u003c/span\u003e\u003cspan address=\"10.1136/BMJ.H1885\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eF. Bagherzadeh-Khiabani, A. Ramezankhani, F. Azizi, F. Hadaegh, E. W. Steyerberg, and D. Khalili, \u0026lsquo;A tutorial on variable selection for clinical prediction models: feature selection methods in data mining could improve the results\u0026rsquo;, J Clin Epidemiol, vol.\u0026nbsp;71, pp.\u0026nbsp;76\u0026ndash;85, Mar. 2016, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/J.JCLINEPI.2015.10.002\u003c/span\u003e\u003cspan address=\"10.1016/J.JCLINEPI.2015.10.002\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG. H. John, R. Kohavi, and K. Pfleger, \u0026lsquo;Irrelevant Features and the Subset Selection Problem\u0026rsquo;, Machine Learning Proceedings 1994, pp.\u0026nbsp;121\u0026ndash;129, Jan. 1994, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/B978-1-55860-335-6.50023-4\u003c/span\u003e\u003cspan address=\"10.1016/B978-1-55860-335-6.50023-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM. Neumann, D. King, I. Beltagy, and W. Ammar, \u0026lsquo;ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing\u0026rsquo;, BioNLP 2019 - SIGBioMed Workshop on Biomedical Natural Language Processing, Proceedings of the 18th BioNLP Workshop and Shared Task, pp.\u0026nbsp;319\u0026ndash;327, Feb. 2019, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.18653/v1/W19-5034\u003c/span\u003e\u003cspan address=\"10.18653/v1/W19-5034\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eO. Bodenreider, \u0026lsquo;The Unified Medical Language System (UMLS): integrating biomedical terminology\u0026rsquo;, Nucleic Acids Res, vol.\u0026nbsp;32, no. suppl_1, pp.\u0026nbsp;D267\u0026ndash;D270, Jan. 2004, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/NAR/GKH061\u003c/span\u003e\u003cspan address=\"10.1093/NAR/GKH061\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL. Breiman, \u0026lsquo;Random forests\u0026rsquo;, Mach Learn, vol.\u0026nbsp;45, no. 1, pp.\u0026nbsp;5\u0026ndash;32, Oct. 2001, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1023/A:1010933404324/METRICS\u003c/span\u003e\u003cspan address=\"10.1023/A:1010933404324/METRICS\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT. Chen and C. Guestrin, \u0026lsquo;XGBoost: A Scalable Tree Boosting System\u0026rsquo;, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2939672\u003c/span\u003e\u003cspan address=\"10.1145/2939672\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZ. Xu, G. Huang, K. Q. Weinberger, and A. X. Zheng, \u0026lsquo;Gradient boosted feature selection\u0026rsquo;, Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp.\u0026nbsp;522\u0026ndash;531, 2014, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/2623330.2623635\u003c/span\u003e\u003cspan address=\"10.1145/2623330.2623635\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eY. Qi, \u0026lsquo;Random Forest for Bioinformatics\u0026rsquo;, Ensemble Machine Learning, pp.\u0026nbsp;307\u0026ndash;323, 2012, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-1-4419-9326-7_11\u003c/span\u003e\u003cspan address=\"10.1007/978-1-4419-9326-7_11\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eH. Quan et al., \u0026lsquo;International variation in the definition of \u0026ldquo;main condition\u0026rdquo; in ICD-coded health data\u0026rsquo;, International Journal for Quality in Health Care, vol.\u0026nbsp;26, no. 5, pp.\u0026nbsp;511\u0026ndash;515, Oct. 2014, doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/INTQHC/MZU064\u003c/span\u003e\u003cspan address=\"10.1093/INTQHC/MZU064\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"brain-informatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"brai","sideBox":"Learn more about [Brain Informatics](https://braininformatics.springeropen.com/about)","snPcode":"40708","submissionUrl":"https://submission.nature.com/new-submission/40708/3","title":"Brain Informatics","twitterHandle":"@SpringerOpen","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-2640617/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2640617/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e\n\u003cp\u003eAbstracting cerebrovascular disease (CeVD) from inpatient electronic medical records (EMRs) through natural language processing (NLP) is pivotal for automated disease surveillance and improving patient outcomes. Existing methods rely on coders’ abstraction, which has time delays and under-coding issues. This study sought to develop an NLP-based method to detect CeVD using EMR clinical notes.\u003c/p\u003e\n\u003ch2\u003eMethods\u003c/h2\u003e\n\u003cp\u003eCeVD status was confirmed through a chart review on randomly selected hospitalized patients who were 18 years or older and discharged from 3 hospitals in Calgary, Alberta, Canada, between January 1 and June 30, 2015. These patients’ chart data were linked to administrative discharge abstract database (DAD) and Sunrise\u003csup\u003eTM\u003c/sup\u003e Clinical Manager (SCM) EMR database records by Personal Health Number (a unique lifetime identifier) and admission date. We trained multiple natural language processing (NLP) predictive models by combining two clinical concept extraction methods and two supervised machine learning (ML) methods: random forest and XGBoost. Using chart review as the reference standard, we compared the model performances with those of the commonly applied International Classification of Diseases (ICD-10-CA) codes, on the metrics of sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV).\u003c/p\u003e\n\u003ch2\u003eResult\u003c/h2\u003e\n\u003cp\u003eOf the study sample (n=3036), the prevalence of CeVD was 11.8% (n=360); the median patient age was 63; and females accounted for 50.3% (n=1528) based on chart data. Among 49 extracted clinical documents from the EMR, four document types were identified as the most influential text sources for identifying CeVD disease (“nursing transfer report,” “discharge summary,” “nursing notes,” and “inpatient consultation.”). The best performing NLP model was XGBoost, combining the Unified Medical Language System concepts extracted by cTAKES (e.g., top-ranked concepts, “Cerebrovascular accident” and “Transient ischemic attack”), and the term frequency-inverse document frequency vectorizer. Compared with ICD codes, the model achieved higher validity overall, such as sensitivity (25.0% vs 70.0%), specificity (99.3% vs 99.1%), PPV (82.6 vs. 87.8%), and NPV (90.8% vs 97.1%).\u003c/p\u003e\n\u003ch2\u003eConclusion\u003c/h2\u003e\n\u003cp\u003eThe NLP algorithm developed in this study performed better than the ICD code algorithm in detecting CeVD. The NLP models could result in an automated EMR tool for identifying CeVD cases and be applied for future studies such as surveillance, and longitudinal studies.\u003c/p\u003e","manuscriptTitle":"Cerebrovascular disease case identification in inpatient electronic medical record data using natural language processing","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-03-08 23:52:11","doi":"10.21203/rs.3.rs-2640617/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-05-16T07:13:41+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-05-15T04:09:20+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"527c91f2-6c94-4fdb-96c5-546fbc250047","date":"2023-05-15T01:44:54+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-03-06T05:45:39+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-03-02T14:21:57+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-03-02T14:21:56+00:00","index":"","fulltext":""},{"type":"submitted","content":"Brain Informatics","date":"2023-02-28T23:50:10+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"brain-informatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"brai","sideBox":"Learn more about [Brain Informatics](https://braininformatics.springeropen.com/about)","snPcode":"40708","submissionUrl":"https://submission.nature.com/new-submission/40708/3","title":"Brain Informatics","twitterHandle":"@SpringerOpen","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e442892d-5867-462e-9753-03c027879f79","owner":[],"postedDate":"March 8th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-09-07T15:11:58+00:00","versionOfRecord":{"articleIdentity":"rs-2640617","link":"https://doi.org/10.1186/s40708-023-00203-w","journal":{"identity":"brain-informatics","isVorOnly":false,"title":"Brain Informatics"},"publishedOn":"2023-09-02 15:09:14","publishedOnDateReadable":"September 2nd, 2023"},"versionCreatedAt":"2023-03-08 23:52:11","video":"","vorDoi":"10.1186/s40708-023-00203-w","vorDoiUrl":"https://doi.org/10.1186/s40708-023-00203-w","workflowStages":[]},"version":"v1","identity":"rs-2640617","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2640617","identity":"rs-2640617","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0