EP4719177A2 - Methods of training an algorithm to predict ischemic stroke etiology - Google Patents

Methods of training an algorithm to predict ischemic stroke etiology

Info

Publication number
EP4719177A2
EP4719177A2 EP24816455.0A EP24816455A EP4719177A2 EP 4719177 A2 EP4719177 A2 EP 4719177A2 EP 24816455 A EP24816455 A EP 24816455A EP 4719177 A2 EP4719177 A2 EP 4719177A2
Authority
EP
European Patent Office
Prior art keywords
model
stroke
refined
features
etiology
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24816455.0A
Other languages
German (de)
French (fr)
Inventor
Richa Sharma
Ho-Joon Lee
Lee SCHWAMM
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yale University
Original Assignee
Yale University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yale University filed Critical Yale University
Publication of EP4719177A2 publication Critical patent/EP4719177A2/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/10Machine learning using kernel methods, e.g. support vector machines [SVM]
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H10/00ICT specially adapted for the handling or processing of patient-related medical or healthcare data
    • G16H10/60ICT specially adapted for the handling or processing of patient-related medical or healthcare data for patient-specific data, e.g. for electronic patient records
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/20ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/70ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients

Landscapes

  • Engineering & Computer Science (AREA)
  • Medical Informatics (AREA)
  • Health & Medical Sciences (AREA)
  • Public Health (AREA)
  • Data Mining & Analysis (AREA)
  • General Health & Medical Sciences (AREA)
  • Epidemiology (AREA)
  • Primary Health Care (AREA)
  • Biomedical Technology (AREA)
  • Pathology (AREA)
  • Databases & Information Systems (AREA)
  • Software Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Physics & Mathematics (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Artificial Intelligence (AREA)
  • Investigating Or Analysing Biological Materials (AREA)
  • Measuring And Recording Apparatus For Diagnosis (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

Provided herein are methods of training and building algorithms to predict stroke etiology. The method of training an algorithm to predict stroke etiology includes processing electronic health records (EHRs) in a derivation dataset, the processing including extracting clinical variables from the EHRs by detecting concept unique identifiers (CUIs) through natural language processing (NLP), and extracting other covariates from the EHRs through regular expression, wherein the clinical variables and other covariates form a training dataset; training two or more base models with the training set to form two or more trained base models, the base models selected from the group including Random Forests (RF), XGBoost (XGB), support vector classifier (SVC), and logistic regression (LR); refining the two or more trained base models using hyperparameters to form two or more refined base models; and building an ensemble model with the two or more refined base models. Also provided herein are articles and methods of predicting ischemic stroke etiology using the trained algorithm.

Description

TITLE OF THE INVENTION Methods of Training an Algorithm to Predict Ischemic Stroke Etiology CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No.63/505,006, filed May 30, 2023, which application is incorporated herein by reference in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with government support under K23NS121634 awarded by National Institutes of Health. The government has certain rights in the invention. BACKGROUND OF THE INVENTION In the United States, there are nearly 676,000 cases of ischemic stroke per year, a quarter of whom have had a prior stroke. While the management of conventional atherosclerotic risk factors is paramount for secondary stroke prevention, therapies targeting important specific mechanisms of stroke have additionally been proven to reduce the risk of recurrent stroke from these etiologies. For example, strategies proven effective include oral anticoagulation for atrial fibrillation (hazard ratio, HR for stroke or secondary embolism with apixaban versus aspirin 0.45, 95% C.I.0.32-0.62) (Connolly et al., N Engl J Med 2011; 364:806-817), internal carotid artery revascularization for symptomatic carotid artery stenosis (17% absolute risk reduction for ipsilateral stroke in 2 years) (North American Symptomatic Carotid Endarterectomy Trial Collaborators, N Engl J Med 1991; 325:445-453), dual antiplatelet therapy for intracranial atherosclerosis, and closure of patent foramen ovale (odds ratio of ischemic stroke recurrence 0.34, 95% CI 0.21-0.90) (Ntaios et al., Stroke 2018; 49: 412–418), among others. Although current practice guidelines endorse the implementation of secondary stroke prevention therapies targeting the underlying mechanism, they are often underutilized in real- world clinical settings. For example, in one study of Veterans Affairs acute ischemic stroke (AIS) patients with symptomatic stenosis, only 34% underwent carotid revascularization. In another study, only 73% of AIS patients with atrial fibrillation-related strokes were prescribed an oral anticoagulant upon discharge. While many factors are at play when determining eligibility and timing for mechanism-specific therapies, one potential reason may be that identifying the precise etiology of a stroke is diagnostically challenging. Unlike myocardial infarction, which is almost uniformly attributable to atherosclerosis, the spectrum of causative mechanisms for ischemic stroke is much broader including atherosclerosis of the aortic arch, carotid artery, and intracranial arteries; cardiac structural and functional pathologies; small vessel lipohyalinosis; coagulopathies; and rare infectious, inflammatory, or neoplastic disorders. In addition to the broader spectrum of causative mechanisms, clinicians typically collect and synthesize a myriad of data sources during stroke investigation. Moreover, only a subset of providers who care for ischemic stroke patients have had formal subspecialty training in vascular neurology, which includes developing expertise in ischemic stroke management using complex diagnostic tools. In a study of a Medicare cohort, contrary to nearly 87% of patients with acute myocardial infarction being evaluated by cardiologists, only 75% of ischemic stroke patients were evaluated by a neurologist, and only 1 in 6 patients were evaluated by board-certified vascular neurologists (Sacchetti et al., Neurohospitalist 2020 Jul; 10(3): 181–187). Not only do these factors make diagnosing ischemic stroke etiology challenging, they also impact the ability to timely administer therapies tailored to the individual patient’s diagnosis. Thus, the risk of a recurrent stroke may remain unmitigated. Accordingly, there is a need in the art for articles and methods that improve on existing articles and methods of identifying ischemic stroke etiology. The present invention addresses this need. SUMMARY OF THE INVENTION In one aspect, provided herein is a method of training an algorithm to predict ischemic stroke etiology, the method including processing electronic health records (EHRs) in a derivation dataset, the processing including extracting clinical variables from the EHRs by detecting concept unique identifiers (CUIs) by applying natural language processing (NLP) tools and extracting other covariates from the EHRs through regular expressions, wherein the clinical variables and other covariates form a training dataset; training two or more base models with the training set given certain hyperparameters, the base models selected from the group including Random Forests (RF), XGBoost (XGB), support vector classifier (SVC), and logistic regression (LR); refining the two or more base models using a number of different combinations of hyperparameter values to form two or more refined base models; and building an ensemble model with the two or more refined base models. In some embodiments, each of the CUIs belongs to a category selected from the group including disease or syndrome; neoplastic process; sign or symptom; and combinations thereof. In some embodiments, the other covariates include one or more of age; sex; clinical information not captured by the CUIs (HEX); radiological features (RAD); cardiac features (HRT); and laboratory features (LAB). In some embodiments, the HEX includes one or more of social history; National Institutes of Health Stroke Severity scale; and vital signs. In some embodiments, the RAD includes one or more of neuroanatomical location of the ischemic stroke; presence of moderate or severe stenosis or occlusion of specific head and neck arteries; and occurrence of intracranial hemorrhage encoded as a binary variable. In some embodiments, the one or more features are selected from the group including magnetic resonance imaging of the brain, magnetic resonance angiography of the head, magnetic angiography of the neck, computed tomography of the head, computed tomography angiography of the head, computed tomography angiography of the neck, carotid ultrasound, transcranial Doppler, conventional angiography. In some embodiments, the HRT include one or more features from electrocardiography reports, transthoracic and transesophageal echocardiography reports, or combinations thereof. In some embodiments, the one or more features are selected from the group including reduced left ventricular ejection fraction, vegetation, thrombus, mass, patent foramen ovale, mitral annular calcification, wall motion abnormality, left ventricular hypertrophy, aortic atheroma, aortic stenosis, mitral regurgitation, aortic regurgitation, mitral valve prolapse, aortic root dilation, diastolic dysfunction, pericardial effusion, T wave inversion, prolonged PR interval, premature atrial contraction, premature ventricular contraction, tachycardia, bradycardia, left ventricular hypertrophy, prolonged QT interval, atrial fibrillation, and supraventricular tachycardia. In some embodiments, one or more of the HRT, LAB, or HRT and LAB features are discretized. In some embodiments, the discretized features include a cutoff between high and normal levels of ejection fraction – 40; NIHSS – 6; sodium – 136; BUN – 24; ALT – 36; AST – 36; white blood cell count – 11; triglycerides – 200; HDL – 40; LDL – 100; TSH – 4.2; PTT – 29.9; A1C – 6.5. In some embodiments, the discretized features include cutoffs between high and normal and normal and low levels of hematocrit – 35 and 46; hemoglobin for female – 11.7 and 15.5; hemoglobin for male – 13.2 and 17.1. In some embodiments, the method further includes the step of reducing feature dimensionality through principal component analysis (PCA). In some embodiments, the hyperparameters are selected through a grid search of a pre- defined hyperparameter space. In some embodiments, refining the two or more trained base models includes validating the models through a stratified cross-validation (CV) strategy of 5 splits with 20% validation sets. In some embodiments, the method further includes training and validating the two or more refined base models using a repeated multi-fold CV strategy. In some embodiments, the repeated multi-fold CV strategy includes performing 2-fold, 3-fold, 4-fold, 5-fold, and 10-fold CV with 30, 20, 15, 12, and 6 repetitions with different random seeds, respectively. In some embodiments, the ensemble model comprises a refined RF model, a refined XGB model, a refined SVC model, a refined second support vector classifier model (SVC2), a refined LR model, a MEAN model, a MEDIAN model, a MAX model, and a MIN model. In some embodiments, the ensemble model generates a TOAST prediction using a voting system among the refined RF model, the refined XGB model, the refined SVC model, the refined LR model, the refined SVC2 model, the MEAN model, the MEDIAN model, the MAX model, and the MIN model. In some embodiments, the ischemic stroke etiology is selected from the group including large artery atherosclerosis (TOAST 1); cardioembolism (TOAST 2); small vessel disease (TOAST 3); and other determined (TOAST 4). In another aspect, provided herein is an apparatus for predicting ischemic stroke etiology, the apparatus including a processor; a memory unit; and a communication interface; wherein the processor is connected to the memory unit and the communication interface; and wherein the processor and memory are configured to implement the method of any one of the previous claims. In another aspect, provided herein is a computer readable storage medium storing computer-executable instructions for performing the method according to any of the embodiments disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS FIG.1 shows a table illustrating TOAST classifications and targeted therapies. FIG.2 shows a schematic illustrating an overview of a method of training a machine learning model to predict ischemic stroke etiology. FIGS.3A-C show graphs illustrating exploratory data analysis. (A) Percentage comparison of discharge summary records with radiology-related features among the 3 cohorts. (B) Numbers of PCs for each PCA total variance cutoff for 2,027 YNHH and MGH features in the case of non-discretized features with all standardized continuous features, discretized features with the standardized age feature, and discretized features with no standardization. (C) Scatter plots of PC1 and PC2 for the three cases in (B) by class and by cohort. (D) Top features which are present in >50% of non-cryptogenic stroke records for each TOAST class and their significance by chi-squared tests. FIGS.4A-E show graphs illustrating model performance. (A) Performances and fit times of each refined model for each feature group by 5-fold CV (mean +/- SD). (B-E) AUCROC and fit times of the (B) LR, (C) SVC, (D) XGB, and (E) RF PCA-based refined models. FIGS.5A-J show graphs illustrating performance comparisons among all feature groups and the 4 ML algorithms and between the training and validation sets of 5-fold cross validation. (A) Full performance results of each refined model for each feature group by 5-fold CV (mean +/- SD). (B) Performances of training and validation sets of 5-fold CV (mean +/- SD) for the feature group, combn1d.age.sex.v1.maxinfo.pca99, by the refined model, LR. (C-F) AUCROC and Accuracy of training and validation sets of 5-fold CV (mean +/- SD) for each feature group by the refined (C) LR, (D) SVC, (E) XGB, and (F) RF models, respectively. (G-J) AUCROC of training and validation sets and fit times for each feature group by the refined LR, SVC, XGB, and RF models, respectively. FIGS.6A-F show graphs illustrating model validation by RMFCV300. (A) ROC and (B) PR curves for each refined model and each CV fold by the RMFCV300 strategy. AUCROC and AUPRC are shown for each class vs. the rest. (C-F) Distributions of multiple performance metrics for each refined model and each class (vs. the rest) as well as (weighted) averages. FIGS.7A-C show graphs illustrating RMFCV300 performances in terms of AUCROC and AUPRC for each optimized model with combn1d.age.sex.v1. (A-B) All classes are combined into a single graph. (C) Each graph is for each class. FIGS.8A-B show graphs illustrating feature importance by SHAP and statistical tests. (A) Top 10 features in terms of means of absolute SHAP values, mean(|SHAP|), across all classes for each refined model for non-PCA-based and PCA-based feature groups. (B) Top 10 features (non-PCA) in terms of SHAP values for each class for each refined model. FIG.9 shows a graph illustrating frequency distributions of the top 10 features contributing to the top 5 PCs by SHAP analysis of the 4 PCA-based refined models. FIGS.10A-B show graphs illustrating correlation between SHAP analysis and statistical tests. (A) Feature correlations between mean(|SHAP|) averaged over the 4 optimized models for each class and the statistic, D (top row), and p-value (bottom row) by Kolmogorov-Smirnov tests for each class vs. the rest. The top 10 features are shown in dark red and Pearson correlation coefficients, r, on the top left.. (B) Similar analyses to (A) by Student’s t-tests. FIG.11 shows graphs illustrating the top 10 features of misclassified samples for each class by the consensus model from RMFCV300. FIGS.12A-B show graphs illustrating prediction of cryptogenic samples and highly frequent features for each predicted class. (A) The bar graphs show a prediction distribution of all cryptogenic patients by StrokeClassifier (left) and a resultant prediction distribution of all of non-cryptogenic and cryptogenic patients (right). (B) The bar plots show class-wide frequency distributions of highly frequent features. There are 26 features which are present in >50% of those cryptogenic samples of any predicted TOAST class. The significance was tested by chi- squared tests. FIG.13 shows graphs illustrating the results of 5-fold cross-validation by stacked generalization models. DETAILED DESCRIPTION OF THE INVENTION Definitions Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are described. The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods. Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range. Detailed Description Provided herein are methods of training and building algorithms to predict stroke etiology. In some embodiments, the method is a computer-implemented machine learning method. In some embodiments, the method includes processing electronic health records (EHRs) in a derivation dataset to form a training dataset, training two or more base models with the training set to form two or more trained base models, refining the two or more trained base models using hyperparameters to form two or more refined base models, and building an ensemble model with the two or more refined base models. In some embodiments, the ensemble model predicts an underlying etiology of a cryptogenic and/or non-cryptogenic stroke according to the TOAST classification system. For example, in one embodiment, the ensemble model predicts one of the following stroke etiologies: TOAST 1- large artery atherosclerosis, TOAST 2- cardioembolism, TOAST 3- small vessel disease, TOAST 4- other determined etiology, and TOAST 5- undetermined etiology. In another embodiment, the ensemble model predicts an etiology from TOAST 1-4. The derivation dataset and an external validation dataset independently include any suitable collection of EHRs containing stroke patients, including, but not limited to, hospitalization records from hospitals and/or stroke centers. This information may be obtained from the hospitals directly (e.g., institutional Get-With-The-Guidelines (GWTG) stroke databases) or through any other suitable manner, such as, but not limited to, Medical Information Mart for Intensive Care (MIMIC-III, a publicly available, de-identified health record repository). Acute ischemic stroke patients may similarly be identified within the databases through any suitable method, such as, but not limited to, through administrative billing codes (e.g., International Classification of Diseases (ICD)). For example, in one embodiment, the derivation dataset includes semi-structured discharge summary plain ASCII text files linked with corresponding stroke hospitalizations from the GWTG databases, while the external validation dataset includes discharge summary plain ASCII text files from stroke patients identified in MIMIC-III. A primary outcome of stroke etiology may be abstracted or ascribed to each patient based upon the EHRs. After identifying and/or obtaining the derivation dataset, the processing of the EHRs includes extracting one or more clinical variables therefrom. In some embodiments, extracting clinical variables from the EHRs includes detecting concept unique identifiers (CUIs) through natural language processing (NLP). Any suitable NLP or text mining tool may be used to detect the CUIs, such as, but not limited to, MetaMap, developed by the National Library of Medicine (NLM). The CUIs include any suitable, relevant CUIs that have been associated with stroke risk. For example, suitable CUIs may include those associated with stroke risk in the literature and belonging to one of the following categories: Disease or Syndrome, Neoplastic Process, or Sign or Symptom. In some embodiments, the CUIs include one or more of those shown in Table S1 of Lee, HJ., Schwamm, L.H., Sansing, L.H. et al. StrokeClassifier: ischemic stroke etiology classification by ensemble consensus modeling using electronic health records. npj Digit. Med.7, 130 (2024) (doi.org/10.1038/s41746-024-01120-w) (hereinafter “StrokeClassifier 2024”), which is incorporated herein by reference. Additionally, the processing step includes extracting one or more other covariates from the EHRs. In some embodiments, extracting the other covariates includes using regular expression to identify the desired variables. In some embodiments, the covariates include one or more of demographic variables, clinical information not captured by the CUIs (HEX), radiological features (RAD), cardiac features (HRT), and laboratory features (LAB). Suitable demographic variables include, but are not limited to, age, sex, or a combination thereof. Suitable HEX variables include, but are not limited to, social history, National Institutes of Health Stroke Severity scale, and vital signs. Suitable RAD variables include, but are not limited to, information about the neuroanatomical location of the ischemic stroke, the presence of moderate or severe stenosis or occlusion of specific head and neck arteries, the occurrence of intracranial hemorrhage encoded as a binary variable, or combinations thereof. Suitable HRT variables include, but are not limited to, electrocardiography and/or echocardiography reports in the discharge summary. In one embodiment, for example, the other covariates include age, sex, 6 HEX variables, 40 RAD variables, 36 HRT variables, and 18 LAB variables (Table 1). Although a specific number of variables is provided as an example, as will be appreciated by those skilled in the art, the disclosure is not so limited and may include any other suitable number of variables. TABLE 1 Together, the clinical variables and the other covariates form the features of a training dataset. In some embodiments, one or more of the features are discretized to reduce measurement noise or error. For example, one or more of the features may be discretized for two levels with one cutoff (high and normal), three levels with two cutoffs (high, normal, and low), or a combination thereof. In one embodiment, the single cutoff value for certain features includes ejection fraction – 40; NIHSS – 6; sodium – 136; BUN – 24; ALT – 36; AST – 36; white blood cell count – 11; triglycerides – 200; HDL – 40; LDL – 100; TSH – 4.2; PTT – 29.9; A1C – 6.5. In another embodiment, the two cutoff values for certain features include hematocrit – 35 and 46; hemoglobin for female – 11.7 and 15.5; hemoglobin for male – 13.2 and 17.1. Additionally or alternatively, in some embodiments, missing data is imputed into the training dataset and/or feature dimensionality is reduced. The missing data may be imputed by any suitable method, such as, but not limited to, by Multivariate Imputation by Chained Equations (MICE; using predictive mean matching (pmm)), Random Forests, or a combination thereof. The feature dimensionality may be reduced by any suitable method, such as, but not limited to, principal component analysis (PCA). For example, in one embodiment, PCA is applied to all features and the top PCs are selected for each of the following 10 thresholds of the total variance: 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, and 99%. In some embodiments, the training dataset includes one or more different feature groups, with each feature group representing a different subset of the features. The different subsets include any suitable grouping of features, such as, but not limited to, CUIs, RAD, HRT, HEX, discretized HEX (HEXd), LAB, discretized LAB (LABd), portions thereof, and combinations thereof. For example, in one embodiment, the training dataset includes one or more of the following feature groups: CUIs, RAD, HRT, HEX, HEXd, LAB, LABd, RAD + HRT + HEX + LAB (i.e., excluding CUIs), CUIs + HRT + HEX + LAB (i.e., excluding RAD), CUIs + RAD + HEX + LAB (i.e., excluding HRT), CUIs + RAD + HRT + LAB (i.e., excluding HEX), CUIs + RAD + HRT + HEX (i.e., excluding LAB), CUIs + RAD + HRT + HEXd (i.e., excluding LABd), CUIs + RAD + HRT + HEX + LAB, and CUIs + RAD + HRT + HEXd + LABd. In another embodiment, the training dataset includes each of the aforementioned feature groups. Once defined, the training dataset is then used to train and/or validate two or more base models. In some embodiments, each base model is built with a different machine learning algorithm. Suitable machine learning algorithms include, but are not limited to, Random Forests (RF), XGBoost (XGB), support vector classifier (SVC), and logistic regression (LR). For example, in some embodiments, the training of the two or more base models includes training and validating a first base model using the RF machine learning algorithm, a second base model using the XGB machine learning algorithm, a third base model using the SVC machine learning algorithm, and a fourth base model using the LR machine learning algorithm. In some embodiments, the training and/or validating of the base models includes refining each base model using hyperparameters for each feature group in the training dataset. For example, in one embodiment, each model is refined with a grid search of a pre-defined hyperparameter space for each of 24 feature groups in the training datasets, i.e., a total of 96 (= 4*24) hyperparameter refinement (HPR) runs. In some embodiments, each base model is also cross-validated (CV). The CV includes any suitable CV strategy, such as, but not limited to, a stratified CV strategy of 5 splits with 20% validation sets using StratifiedShuffleSplit from the scikit-learn library in Python. In some embodiments, the randomness of the stratified CV is controlled by setting the parameter random_state. For example, in one embodiment, the randomness of the stratified CV is controlled by setting the parameter, random_state = 1701. Following the training, refining, and cross-validating, one of each of the base models including refined parameters (refined base models) is selected for inclusion in the ensemble model. The selection may be by any suitable method, such as, but not limited to, based on the maximum AUCROC (the area under the curve of the receiver operating characteristic). In some embodiments, the selected base models may be further trained and validated using a repeated multi-fold CV strategy. For example, in one embodiment, the further training and validation includes performing 2-fold, 3-fold, 4-fold, 5-fold, and 10-fold CV with 30, 20, 15, 12, and 6 repetitions with different random seeds, respectively (i.e., 60 * 5 = 300 CV experiments in total), using RepeatedStratifiedKFold from the scikit-learn library in Python. In some embodiments, the further training and validation provides less statistical bias and more robustness compared to the single 5-fold CV strategy. The selected base models are then used to build an ensemble model. In some embodiments, the ensemble model also includes one or more additional base models and/or one or more ensemble classifiers. The ensemble classifiers may be generated by any suitable method, such as, but not limited to, through determining, for each class, the mean, median, maximum, and/or minimum of predicted probabilities generated from the base models, and normalizing them across the classes as ensemble classifier models MEAN, MEDIAN, MAX, and/or MIN, respectively. In one embodiment, for example, the ensemble model includes a refined Random Forests model (RF*), a refined XGBoost model (XGB*), a refined support vector classifier model (SVC*), and a refined logistic regression model (LR*). In another embodiment, in addition to the four refined base models, the ensemble model includes a second refined support vector classifier model (SVC2). In a further embodiment, in addition to the five refined base models, the ensemble model includes the ensemble classifier models MEAN, MEDIAN, MAX, and MIN, which are the mean, median, maximum, and minimum for each class from predicted probabilities generated from the five base models, normalized across the four classes. Also provided herein are methods of predicting ischemic stroke etiology. In some embodiments, the method includes training an algorithm to predict stroke etiology according to one or more of the embodiments disclosed herein, and applying the algorithm to an EHR from a subject, the algorithm predicting an ischemic stroke etiology of the subject based upon the EHR. In some embodiments, the ischemic stroke etiology is one of large artery atherosclerosis (TOAST 1); cardioembolism (TOAST 2); small vessel disease (TOAST 3); or other determined (TOAST 4). Additionally or alternatively, the stroke may be cryptogenic and/or non-cryptogenic stroke. In some embodiments, the ensemble model generates the TOAST prediction based upon a consensus prediction of the underlying classifiers (i.e., base models and/or ensemble classifiers). For example, in one embodiment, the algorithm predicts the ischemic stroke etiology from consensus predictions among the nine classifiers (i.e., RF*, XGB*, SVC*, LR* SVC2, MEAN, MEDIAN, MAX, and MIN) as a meta-classifier or a voting system. Without wishing to be bound by theory, it is believed that the consensus prediction using the ensemble model reduces or averages out bias and/or variance to provide improved performance and/or robustness as compared to a single classifier. Additionally or alternatively, in some embodiments, the methods disclosed herein provide automated ischemic stroke etiology prediction from free, unstructured text at any stage of collection. Accordingly, in some embodiments, the methods disclosed herein predict automated ischemic stroke etiology in real- time and/or within a few minutes for a sizable block of patients, with a simple output of one of 4 potential stroke etiology classes. Furthermore, in some embodiments, the methods may link the predicted ischemic stroke etiology with the corresponding guideline-recommended secondary stroke prevention therapy. By providing the capability of classifying stroke etiology during the time of the stroke hospitalization, the methods disclosed herein can enhance quality improvement, clinical trial, epidemiologic, and health policy initiatives for which stroke etiology information is essential. Additionally or alternatively, the methods disclosed herein may be applied downstream from classification of ischemic stroke etiology, such as, for example, as a clinical decision support system. Further provided herein are a computer readable storage medium and an apparatus for classifying ischemic stroke etiology. In some embodiments, the computer readable storage medium includes any suitable computer readable storage medium storing computer-executable instructions for performing the method according to any of the embodiments disclosed herein. In some embodiments, the apparatus for classifying ischemic stroke etiology includes a processor, a memory unit, and a communication interface. In one embodiment, the processor is connected to the memory unit and the communication interface. In another embodiment, the processor and memory are configured to implement the method according to any one of embodiments disclosed herein. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, numerous equivalents to the specific procedures, embodiments, claims, and examples described herein. Such equivalents are considered to be within the scope of this invention and covered by the claims appended hereto. It is to be understood that wherever values and ranges are provided herein, all values and ranges encompassed by these values and ranges, are meant to be encompassed within the scope of the present invention. Moreover, all values that fall within these ranges, as well as the upper or lower limits of a range of values, are also contemplated by the present application. The following examples further illustrate aspects of the present invention. However, they are in no way a limitation of the teachings or disclosure of the present invention as set forth herein. EXAMPLES EXAMPLE 1 Determining the etiology of an acute ischemic stroke (AIS) is fundamental to secondary stroke prevention efforts, but can be diagnostically challenging. In this Example, an automated classification machine intelligence tool, StrokeClassifier, was trained and validated using electronic health record (EHR) text data from 2,039 non-cryptogenic AIS patients at 2 academic hospitals to predict the 4-level outcome of stroke etiology determined by agreement of at least 2 board-certified vascular neurologists’ review of the stroke hospitalization EHR. StrokeClassifier is an ensemble consensus meta-model of 9 machine learning classifiers applied to features extracted from discharge summary texts by natural language processing. StrokeClassifier was externally validated in 406 discharge summaries from the MIMIC-III dataset reviewed by a vascular neurologist to ascertain stroke etiology. Compared with stroke etiologies adjudicated by vascular neurologists, StrokeClassifier achieved the mean cross-validated accuracy of 0.74 ( ±0.01) and weighted F1 of 0.74 ( ±0.01) for the multi-class classification. In the MIMIC-III cohort, the accuracy and weighted F1 of StrokeClassifier were 0.70 and 0.71, respectively. In the case of binary classification, the two metrics range from 0.77 to 0.96. SHapley Additive exPlanation analysis elucidated that the top 5 features contributing to stroke etiology prediction were atrial fibrillation, age, middle cerebral artery occlusion, internal carotid artery occlusion, and frontal stroke location. A certainty heuristic was then designed to deem a StrokeClassifier diagnosis as confidently non-cryptogenic by the degree of consensus among the 9 classifiers and applied it to 788 cryptogenic patients. This reduced the percentage of the cryptogenic strokes from 25.2% to 7.2% of all ischemic strokes. StrokeClassifier is a validated artificial intelligence tool that rivals the performance of vascular neurologists in classifying ischemic stroke etiology for individual patients. INTRODUCTION Identifying the etiology of an ischemic stroke is a clinically challenging and consequential task. In the United States, there are nearly 676,000 cases of ischemic stroke per year, a quarter of whom have had a prior stroke. Among stroke survivors, another stroke can lead to death or further disability. The causative mechanism or etiology of an ischemic stroke can be heterogeneous including large artery atherosclerosis, cardioembolism, small vessel disease, and other rare, determined etiologies. Nearly 20-30% of ischemic stroke patients in the U.S. are considered cryptogenic with no etiology determined after evaluation. The risk of recurrent stroke after a cryptogenic stroke is heightened at 5.6% at 3 months and between 14-20% at 2 years. In one study, at 21 months, cryptogenic strokes were associated with higher risk of recurrent stroke in comparison with cardioembolic (HR 1.83, p=0.028) and non-cardioembolic stroke patients with known source (HR 2.4, p=0.046). An analysis of the NOR-FIB study demonstrated an annual risk of stroke recurrence of 7.7% versus 2.8% among individuals with cryptogenic versus non-cryptogenic strokes, respectively. In the Athens Stroke Registry, the stroke recurrence rate in patients with cryptogenic stroke was 29% over a mean of 30.5 months, significantly higher compared with all non-cardioembolic stroke subtypes. The diagnosis of ischemic stroke etiology determined by a patient’s treating clinician may partly contribute to the differential rates of stroke recurrence by etiology as each diagnosis prompts a specific secondary stroke prevention treatment plan. Evidence-based, etiology-specific treatments that are proven to reduce the risk of recurrent stroke to varying degrees include carotid revascularization for symptomatic severe carotid stenosis, anticoagulation for atrial fibrillation or left ventricular thrombus, dual antiplatelet therapy after intracranial stenosis- related stroke, and patent foramen ovale closure when it is implicated, among others (FIG.1). Despite high-level evidence supporting the efficacy of such therapies to prevent recurrent stroke, secondary stroke prevention treatments are significantly underutilized both in the U.S. and globally after an ischemic stroke. This implementation gap may underlie the observation that the majority of recurrent strokes are from the same etiology as the index stroke. Furthermore, a cryptogenic stroke diagnosis precludes the institution of any guideline-recommended therapy that targets specific stroke mechanisms and reduce the risk of recurrent stroke from culprit sources. The ability to tailor and implement secondary stroke prevention strategies fundamentally hinges on the diagnosis of the culprit mechanism of an ischemic stroke. To determine the causative mechanism of an ischemic stroke, clinicians synthesize a vast array of data including clinical history and physical examination, laboratory data, cardiac rhythm interrogation, cardiac imaging, and neuroradiologic studies. Utilization of diagnostic tools has increased with time, nevertheless, a significant proportion of patients remain cryptogenic. Diagnostic uncertainty arises due to 1) an inadequate or incomplete workup with further results pending after discharge, 2) a complete workup yielding no known stroke etiology, or 3) multiple, competing possible etiologies, resulting in a diagnosis of stroke of undetermined etiology. An exacerbating factor may be the lack of widespread neurovascular experts specifically trained to collect and examine data to ascertain stroke etiology. A study has demonstrated that compared to evaluation by a non-vascular neurologist, evaluation by a vascular neurologist was associated with a more comprehensive diagnostic investigation that may change management. There is a shortage of vascular neurologists in the United States, with only one in every 6 ischemic stroke patients treated by a board-certified vascular neurologist. In this context, there is an opportunity for an automated, artificial intelligence solution to standardize the process of diagnosing the causative mechanism of stroke. Artificial intelligence has been heavily adapted for clinical use to help determine patient eligibility for acute stroke therapies such as thrombectomy to abort a stroke, but only minimally for the purpose of stroke prevention. There have been several studies of machine learning classifiers to predict stroke etiology, however these have been limited by use of manually curated discrete features, single center samples, insufficient adjudication of stroke etiology outcomes, exclusion of patients with multiple potential etiologies, reliance on a singular model, lack of model explainability, or broad, heterogeneous categorization of stroke etiology. In this multi- center study, we aim to develop and externally validate a multi-level, automated ischemic stroke etiology classifier by applying natural language and innovative machine learning tools applied directly to semi-structured text data from the EHR compiled during the AIS hospitalization. ABBREVIATIONS AF Atrial fibrillation AIS Acute ischemic stroke AUC Area under the curve AUCROC Area under the curve of the receiver operating characteristic AUPRC Area under the precision-recall curve BA Balanced accuracy CUI Concept Unique Identifier EHR Electronic health record FDR False discovery rate FNR False negative rate FPR False positive rate HPO Hyperparameter optimization HRT Heart-related features HEX Social history, NIHSS, vital signs features KAPP Cohen’s kappa LAB Laboratory test features LR Logistic regression NIHSS National Institutes of Health Stroke Scale PR Precision-Recall PPV Positive predictive value PCA Principal component analysis RAD Radiology-related features ROC Receiver operating characteristic RF Random Forests SHAP SHapley Additive explanation SVC Support vector classifier TTE Transthoracic echocardiography UMLS Unified Medical Language System XGB XGboost METHODS Study Population and Data Sources The derivation cohort consisted of hospitalizations at 2 academic, Comprehensive Stroke Centers of Yale New Haven Hospital (YNHH) and Massachusetts General Hospital (MGH) from 2015 to 2020. Institutional Review Board approval was obtained from both YNHH and MGH. The external validation cohort was a subgroup of hospitalizations at the academic, Comprehensive Stroke Center of Beth Israel Deaconess Medical Center from 2001 to 2012. Access to this cohort’s data was obtained through the MIMIC-III (Medical Information Mart for Intensive Care) warehouse which contains records of 46,520 hospitalizations from 2001 to 2012 at Beth Israel Deaconess Medical Center. MIMIC-III is a publicly available, de-identified health record repository that was developed and approved by the Beth Israel Deaconess Medical Center and Massachusetts Institute of Technology IRBs. Acute ischemic stroke hospitalizations at YNHH and MGH were identified by each institution’s Get-with-the-guidelines stroke database. Get-With-The-Guidelines (GWTG)-Stroke database is a quality improvement initiative in which participating hospitals enter clinical and radiographic data of all patients hospitalized with an ischemic stroke diagnosis. Acute ischemic stroke patients are identified by administrative billing codes (International Classification of Diseases (ICD), 10th Revision). Data abstraction, entry, and adjudication are performed by trained study personnel. There are logic checks and form controls to minimize data entry errors. The database was queried for all ischemic stroke patients >18 years admitted from January 2015 to December 2020 at MGH and YNHH to assemble the ischemic stroke cohort. The EHR platform for both institutions is Epic (Epic Systems Corporation), the most prevalent EHR system in the United States. Stroke hospitalizations from the GWTG databases were linked with corresponding semi-structured discharge summary plain ASCII text files, resulting in a total 1,269 and 1,493 records from YNHH and MGH, respectively. The MIMIC-III dataset was queried for the ICD-9 codes of 433.X and 434.X that are associated with ischemic stroke, resulting in a total of 2,563 hospitalization records from patients ages >18 years admitted to BIDMC from 2001 to 2012. A subset of these, a convenience sample of the first consecutive 500 records, were included in this study for external validation and their discharge summary plain ASCII text files were analyzed. BIDMC utilizes its own customized, hospital-wide EHR system. A description of the study populations from the 3 institutions represented in this analysis is provided in Table 2. TABLE 2 - Description of study cohorts Outcomes The primary study outcome was stroke etiology as defined by the 5 mutually exclusive causative mechanisms of stroke per the TOAST classification system: 1- large artery atherosclerosis, 2- cardioembolism, 3- small vessel disease, 4- other determined etiology, and 5- undetermined etiology (cryptogenic). Stroke etiology was determined by the agreement of two board-certified vascular neurologists. The first vascular neurologist was the discharging treating clinician, when applicable, who documented a stroke etiology impression in the EHR. The second vascular neurologist was a co-author who reviewed the entire stroke hospitalization record and viewed the neuroimaging. When either there was disagreement about the stroke etiology between the two vascular neurologists or the discharging treating clinician was not a vascular neurologist (4% and 2% of the YNHH and MGH cohort, respectively), a third vascular neurologist at each of the two institutions reviewed the entire stroke hospitalization record and provided stroke etiology diagnosis impressions. The final stroke etiology diagnosis was the etiology ascribed by the majority. If there was no majority, the stroke etiology diagnosed by the senior-most vascular neurologist was utilized. In the external validation cohort, the co-author, R.S., reviewed the text of each discharge summary and designated a TOAST classification based on the data recorded in the text corpus. Covariates Demographic Variables- Using regular expressions, age and sex were extracted from discharge summary text. The YNHH dataset did not contain sex information in a structured format in the discharge summary, unlike the MGH data. To identify sex information from the YNHH data, a customized R code was used to search for "her” or “his" in the EHR texts to assign female or male to each EHR, respectively. The accuracy of this extraction was compared with the age and sex fields hardcoded in the corresponding institutional GWTG-stroke registry. The proxy variable of race was intentionally not included as a covariate for model training and testing because the datasets lack measures of the social environment which may be more relevant indicators of stroke etiology than ancestry alone. Clinical Variables Derived from MetaMap- Natural language processing tools were applied to the corpus of discharge summary texts to engineer clinical variables that may associate with stroke etiology. Firstly, discharge summaries were processed using the natural language processing (NLP) or text mining tool, MetaMap, developed by the National Library of Medicine (NLM) to extract terms from text and link them to standard biomedical concepts in the Unified Medical Language System (UMLS) Metathesaurus. Each discharge summary is a semi- structured text that can be processed by MetaMap to detect unique concepts or concept unique identifiers (CUIs) from the UMLS which contains over 1 million biomedical concepts in an automated manner. MetaMap was applied to the discharge summary text of each hospitalization and extracted CUIs that belong to the following 3 types or categories: “Disease or Syndrome”, “Neoplastic Process”, and “Sign or Symptom” (Table S1). The rationale for selecting MetaMap CUIs was that it was designed to retrieve medical concepts by lexical analysis and tokenization. MetaMap allows for abbreviations, acronyms, negations, and parts-of-speech tagging. It facilitates lookups in the SPECIALIST system that is supported by the UMLS Metathesaurus and Semantic Network, a repository of biomedical concepts and their interrelationships that is updated quarterly and incorporates SNOMED CT content which is routinely utilized in SNOMED CT-enabled EHR systems to enable meaning-based retrieval of information and maps to ICD-9 and ICD-10 coding systems. MetaMap also performs word sense disambiguation by which concepts are favored if semantically consistent with surrounding text. There is also flexibility in input and output data formats permissible by MetaMap. Finally, MetaMap has been rigorously tested in various biomedical research applications. Compared with other clinical entity extraction tools, MetaMap was demonstrated to have the highest recall and F1-score when tasked with identifying obesity-related symptoms. In one study, MetaMap extracted biomarker types from pathology reports with > 95% accuracy. Other Variables: By employing customized regular expressions, 4 other categories of features were curated from discharge summaries. First, clinical information not captured by CUIs was extracted, including social history (tobacco, ethanol, and illicit drug use), National Institutes of Health Stroke Severity scale, and vital signs, which we designate as 6 HEX features. Second, 40 radiologic features (RAD) were extracted from studies performed during the stroke hospitalization including information about the neuroanatomical location of the ischemic stroke, the presence of moderate or severe stenosis or occlusion of specific head and neck arteries, and the occurrence of intracranial hemorrhage encoded as a binary variable (FIG.3A). The accuracy of the automated method of radiology data extraction in a random sample of 100 selected for each variable was 98% for neuroanatomic location and 99% for vessel abnormality. Third, 36 cardiac features (HRT) were extracted from electrocardiography and echocardiography reports in the discharge summary. Finally, we extracted 18 laboratory features (LAB). All lab values were generated during the stroke hospitalization encounter. In a random sample of 5 YNHH and 5 MGH patients, the accuracy of the HRT and LAB features that were extracted was 100%. In order to reduce measurement noise or error, the continuous values of the HEX and LAB features were discretized into clinically relevant categories: Ejection fraction was dichotomized as < 40% which is defined as severely reduced versus >= 40%; NIHSS was dichotomized as <6 defining a minor stroke versus >= 6; sodium level < 136 mmol/liter which is defined as hyponatremia versus >= 136mmol/liter; BUN >= 24mg/dL versus < 24mg/dL; ALT and AST < 36U/L versus >= 36U/L; white blood cell count < 11x1000/microliter versus >= 11x1000/microliter which defines leukocytosis; hematocrit < 35% (anemia), 35-45% (normal), and >=46% (erythrocytosis); hemoglobin in females < 11.7 (anemia), 11.7-15.5 (normal), and > 15.5 (erythrocytosis); hemoglobin in males < 13.2 g/dL (anemia), 13.2-17.1 g/dL (normal), and > 17.1 g/dL (erythrocytosis); triglyceride >= 200 mg/dL which defines hypertriglyceridemia versus < 200 mg/dL; HDL mg/dL < 40 versus >= 40 mg/dL; LDL >= 100 mg/dL versus < 100 mg/dL; TSH < 4.2 micro IU/mL versus >= 4.2 micro IU/mL; PTT < 29.9 versus >= 30 seconds; hemoglobin A1c >= 6.5% which defines diabetes versus < 6.5%. The rationale for selecting these thresholds for categorization is detailed in Supplementary Information below (Section 2). The discretized feature groups are denoted by HEXd and LABd. Model performance was assessed based on each of the 5 feature groups, all of the 5 groups, or those 5 combinations excluding each group. Completeness of the investigation for stroke etiology during the hospitalization was assessed based on values available for each of these groups. Imputation of missing data A multiple imputation method, MICE (Multivariate Imputation by Chained Equations) (Azur, M. J., et al. Multiple imputation by chained equations: what is it and how does it work? 20 Int J Methods Psychiatr Res 40-49 (2011); Raghunathan, T. E., et al. A multivariate technique for multiply imputing missing values using a sequence of regression models, 27 Survey Methodology 85-95 (2001)), was deployed from the mice package in R to impute missing values in categorical and numerical features of the YNHH and MGH data using the built-in method of predictive mean matching (pmm) with the default parameters. The missing MIMIC features were also imputed using the built-in method of Random Forests (rf; with the default parameters), which was found to better deal with larger fractions of missing values than pmm or other built-in imputation methods. Dimensionality reduction of features by principal component analysis Since the number of features totaled 2,027, the relationship between dimensionality reduction of features and model training and performance was explored. Principal component analysis (PCA) was chosen to reduce the feature dimensionality because of its clear interpretation of each principal component as a linear combination of all features. PCA was applied to all features and selected the top PCs for each of the following 10 thresholds of the total variance: 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, and 99%. Validation and test datasets were transformed based on PCA of training datasets. Machine learning model development and evaluation We analyzed non-cryptogenic ischemic stroke hospitalization records of discharge summaries from the merged YNHH and MGH datasets for model training and validation. FIG.2 shows an overview of the present workflow. Records from non-cryptogenic ischemic stroke hospitalizations in the MIMIC dataset were used as the test dataset. We built models using the following 20 different feature groups individually: CUIs; RAD; HRT; HEX; HEXd; LAB; LABd; RAD + HRT + HEX + LAB; CUIs + HRT + HEX + LAB; CUIs + RAD + HEX + LAB; CUIs + RAD + HRT + LAB; CUIs + RAD + HRT + HEX; CUIs + RAD + HRT + HEXd; CUIs + RAD + HRT + HEX + LAB; and CUIs + RAD + HRT + HEXd + LABd. For the last two groups, a filtering of samples was also applied based on maximum information (MaxInfo) ≥ 4 (the number of feature categories present) and the 11 PCA-based feature groups described above. Base models were built using 4 different supervised machine learning algorithms to classify the 4-level non-cryptogenic stroke etiology outcome: logistic regression (LR), support vector classifier (SVC), Random Forests (RF), and XGBoost (XGB). Each model was refined with a grid search of a pre-defined hyperparameter space for each of 24 training datasets, i.e., a total of 96 (= 4*24) hyperparameter refinement (HPR) runs, and a stratified cross-validation (CV) strategy of 5 splits of 20% validation sets using StratifiedShuffleSplit from the scikit-learn library in Python. The randomness of the stratified CV was controlled by setting the parameter, random_state = 1701, in this work. The best models with refined parameters were selected based on the maximum AUCROC (the area under the curve of the receiver operating characteristic). Details about models and configurations for HPR are provided in the Supplementary Information below. For the 4 best models with the optimal parameters identified by the above strategy, more comprehensive training and validation was performed next using a repeated multi- fold CV strategy to reduce statistical bias and ensure robustness compared to the single 5-fold CV strategy above.2-fold, 3-fold, 4-fold, 5-fold, and 10-fold CV were performed with 30, 20, 15, 12, and 6 repetitions with different random seeds, respectively (using RepeatedStratifiedKFold from the scikit-learn library in Python), i.e., 60 * 5 = 300 CV experiments in total. This strategy is denoted as RMFCV300. Next, 4 ensemble models were built using the 4 refined models selected above as well as SVC with alternative prediction probabilities, which is termed SVC2 (See Supplementary Information) as base models. The rationale for building ensemble models is that ensemble learning has demonstrated success in improving performances over single models in reducing variance or bias. From predicted probabilities generated from these 5 base models, the mean, median, maximum, and minimum for each class were normalized across the 4 classes as 4 ensemble models: MEAN, MEDIAN, MAX, and MIN, respectively. The summary-statistic- based ensemble models are a variant of stacked generalization (Wolpert, D. H. Stacked generalization, 5 Neural Networks 241-259 (1992)) without additional training. This yielded a 9- classifier system of 5 refined base and 4 ensemble classifiers. Consensus predictions were obtained among those 9 classifiers as a meta-classifier or a voting system to reduce or average out any bias from a single classifier and improve robustness. The resulting algorithm was designated as StrokeClassifier. StrokeClassifier was additional analyzed by (1) training on the YNHH dataset and testing on the MGH and MIMIC datasets and (2) training on the MGH dataset and testing on the YNHH and MIMIC datasets for a 5-way cross-hospital validation in total. For the purpose of comparison, stacking ensemble models using each of LR, SVC, RF, and XGB were also employed as a meta model (See Supplementary Information). For model performance evaluation, the following 7 performance metrics based on weighted averages for one-vs-rest classification were used: AUCROC, area under the precision- recall curve (AUPRC or average precision), accuracy (i.e., weighted recall), balanced accuracy (i.e., macro recall or the arithmetic mean of sensitivity and specificity), precision, F1, and Cohen’s kappa. As for the qualitative interpretation of Cohen’s kappa values, the scheme by Landis and Koch (Landis, J. R. & Koch, G. G. The Measurement of Observer Agreement for Categorical Data, 33 Biometrics 159-174 (1977)) was followed: kappa < 0 as no agreement, 0 – 0.20 as slight, 0.21 – 0.40 as fair, 0.41 – 0.60 as moderate, 0.61 – 0.80 as substantial, and 0.81 – 1 as almost perfect agreement. For model interpretation and feature importance, the game-theoretic Shapley value-based SHAP (Shapley Additive exPlanations) analysis was performed using the shap package in Python (Lundberg, S. M. & Lee, S.-I. in Advances in Neural Information Processing Systems 30 (eds I. Guyon et al.) 4765-4774 (Curran Associates, Inc., 2017); Lundberg, S. M. et al. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence 2, 56-67, doi:10.1038/s42256-019-0138-9 (2020)), as done previously (Lee, H.-J. An interactome landscape of SARS-CoV-2 virus-human protein-protein interactions by protein sequence-based multi-label classifiers. bioRxiv (2021); Smith, K., Shen, F., Lee, H. J. & Chandrasekaran, S. Metabolic signatures of regulation by phosphorylation and acetylation, iScience, 103730 (2022)). TreeSHAP was used for RF and XGB and KernelSHAP for LR and SVC with a k-means background with k = 100 for computational efficiency. As an alternative approach to ascertain feature importance, classifier-agnostic Student’s t-tests and Kolmogorov- Smirnov tests were performed for one-vs-rest comparisons for each class and each feature. Exploratory analyses were performed to evaluate etiologic predictions by StrokeClassifier for cryptogenic strokes adjudicated by vascular neurologists. Various certainty heuristics defined computationally by thresholds of diagnostic confidence were examined. These diagnostic confidence thresholds were designated by the number of consensus supports provided by the 9 individual classifiers in the ensemble model for each non-cryptogenic stroke etiology. As a proof of concept, the threshold of the first quartile of frequencies of support were applied for each etiology from the external validation of the MIMIC-III cohort to predict the etiologies of cryptogenic patients (788 in total) and evaluated the distribution of predicted etiologies. Those predictions with the consensus frequencies less than the thresholds were deemed persistently cryptogenic. Etiology distributions yielded by other quartile thresholds and the means of the support frequencies were also examined. Using the first quartile thresholds, a repertoire of EHR signatures associated with each predicted TOAST class were identified for cryptogenic strokes by evaluating feature frequencies from StrokeClassifier. Finally, a longitudinal analysis of StrokeClassifier was performed by dividing the combined cohort of YNHH and MGH into a training set of 1,688 discharge summaries from 2015 to 2019 and a test set of 244 discharge summaries from 2020. StrokeClassifier was re- trained using the training set along with a stratified 5-fold CV and hyperparameter optimization as above and then the optimal model was longitudinally validated using the test set. All analyses were performed in Python and R using a macOS laptop with 2.6 GHz 6-Core Intel Core i7 and 32GB memory in the case of RF and LR and a high-performance computing cluster with 64 cores and 1GB memory per core in the case of XGB and SVC. RESULTS Study Participants The study sample consisted of 3,262 discharge summaries with AIS diagnoses (N=1,269 at YNHH from 2015-2020; N=1,493 at MGH from 2016-2019; N=500 at BIDMC from 2001- 2012). The characteristics of the 3 cohorts are presented in Table 2. The derivation cohorts of YNHH and MGH were similar with some exceptions. The YNHH cohort was significantly older (median age 71 years [IQR 59-82]) compared with the MGH cohort (median age 69 [IQR 59- 79]) (p=0.013). The median word count of the YNHH discharge summaries (1639 words [IQR 1274-2064]) was significantly lower than in the MGH discharge summaries (2058 words [IQR 1593-2554]) (p=1.21e-35). The YNHH cohort was significantly more likely than the MGH cohort to have hyperlipidemia (32.9% versus 11.5%, p=0.001) and coronary artery disease (17.8% versus 4.0%, p=0.003). The YNHH and MGH cohorts had similar distributions of stroke etiologies: large artery atherosclerosis (19.8% versus 21.0%), cardioembolism (32.9% versus 29.9%), small vessel disease (15.3% versus 10.7%), other determined etiology (8.9% versus 9.6%), and cryptogenic etiology (23.1% versus 28.8%). The degree of completeness of extracted features was comparable between YNHH and MGH with respect to UMLS CUIs (extracted from 95.7% versus 94.5%), neuroimaging features (extracted from 94.1% versus 92.0%), cardiac features (95.4% versus 93.0%), clinical history (90.3% versus 91.5%), and laboratory features (90.0% versus 92.3%). Characteristics of the combined derivation cohort were compared with those of the external validation MIMIC-III cohort. The external validation cohort was comparable in age as the combined derivation cohort. Median word count of the external validation cohort discharge summaries was significantly lower (1712 words [IQR 1160-2294], p=0.002). The external validation cohort was more likely to have heart failure (27.3% versus 12.5%, p=0.019). The distribution of stroke etiologies differed significantly between the derivation and external validation cohorts (p=0.001). Large artery atherosclerosis (8.8% versus 20.5%, p=0.031) and small vessel disease (3.6% versus 12.8%, p=0.023) were significantly less frequent in the external validation cohort while cardioembolism was significantly more frequent (51.2% versus 31.3%, p=0.028). The derivation and external validation cohorts were similar in terms of feature completeness (p=0.638-0.979) (Table 2; FIG.3A). Data Post-Processing and Principal Component Analysis Of the 2,039 samples in the YNHH and MGH cohorts, a total of 1,932 samples were included for model derivation having successfully been post-processed by MetaMap. Imputation of missing entries in categorical and numerical features was performed using MICE in the derivation cohort and Random Forests-based imputation in the external validation cohort (Table 3). The levels of missingness for the categorical and numerical features were 91.9% (76.8% to 99.9%) and 73.4% (2.3% to 99.9%) on average, respectively. Imputation of several features was failed and they were excluded subsequently. All subsequent analyses were performed on the imputed datasets. Table 3 Principal component analysis of features For the non-cryptogenic stroke derivation cohort of 1,932 patients from YNHH and MGH analyzed for model development, PCA was performed on all of 2,027 features, either discretized or not, to reduce dimensionality or noise. The top PCs were then selected for each of the 10 thresholds of the total variance (see Methods) for model development. It was found that 99% of the total variance could be explained by less than half of all features, the first principal component with about 4.5% variance discriminating between the two cohorts (FIGS.3B and 3C). Base Models with Optimized Hyperparameters and Model Performances 96 hyperparameter refinements (HPRs) were performed for the 4 supervised machine- learning algorithms of LR, SVC, RF, and XGBoost and 24 training datasets (Tables 4A-B, FIG. 4A, Tables 5A-C). Based on the AUCROC rankings in the 5-fold CV (Table 6), the best model for each of the 4 strategies is denoted as LR*, SVC*, RF*, and XGB*, respectively, hereafter. All 4 best models were built using the full features with discretization (age + sex + CUI + RAD + HRT + HEXd + LABd, denoted by (Tables 4A-B). AUCROC and mean cross-validated accuracy were 89.8% and 74.7% for LR*, 90.1% and 71.9% for SVC*, 91.3% and 74.6% for XGB*, and 90.5% and 69.1% for RF*. Similar performances were observed with PCA of the full features (denoted by except for RF* (Tables 4A-B). Fit times for XGB* with were particularly longer (>235 sec) than those for the other 3 models (Tables 4A-B). It was also observed that XGB and RF tend to overfit (FIGS.4B-E and 5A-J). Table 4A. Validation results of the refined base models for each feature group Table 4B. Validation results of the refined base models for each feature group
Table 5A Table 5B Table 5C Table 6 CUIs contributed most to model performance as measured by AUCROC, while the radiologic features ranked second. Decrease in performance was largest for each model when CUIs were excluded from the full feature group. On the other hand, excluding the LAB and HEX features tend to improve the performances. There was no performance improvement with those samples of high feature information defined by the presence of at least 4 features groups. Next, the performance of each optimized model was evaluated for the full cohort of the 1,932 samples. SVC2 model was also built and examined, which calculates alternative prediction probabilities as a different calibration approach using the optimized hyperparameters from SVC* (See Supplementary Information). The runtimes for the 5 models of LR*, SVC*, RF*, XGB*, and SVC2 were 114ms, 10.8s, 258ms, 475ms, and 10.8s, respectively, and their accuracies were 90.4%, 86.2%, 92.4%, 97.6%, and 88.1%. The numbers of samples correctly predicted by N = 1, 2, 3, 4, and 5 models (i.e., supports) are 59 (3.1%), 74 (3.8%), 92 (4.8%), 108 (5.6%), and 1574 (81.5%), respectively. In other words, 91.9% of all samples were correctly predicted by at least 3 models. The remaining 25 samples (1.3%) were incorrectly predicted by all the 5 models. The [numbers, percentages] of 1,002 MGH and 930 YNHH samples with N = 0 to 5 supports are [(13, 12), (1.3%, 1.3%)], [(32, 27), (3.2%, 2.9%)], [(31, 43), (3.1%, 4.6%)], [(44, 48), (4.4%, 5.2%)], [(57, 51), (5.7%, 5.5%)], and [(825, 749), (82.3%, 80.5%)], respectively. When those 59 samples correctly predicted by a single model (N = 1) were analyzed, RF* was found to correctly predict 49 (83.1%) samples, in particular for TOAST 1 and 2 (22 and 16 samples or 37.3% and 27.1%, respectively). Performance of Ensemble Models and Consensus Meta-Model, StrokeClassifier The 4 refined models built using full features and samples were aggregated, X along with SVC2, into 4 ensemble models with 4 pre-specified summary statistics (see Methods). The 5-fold CV performance metrics associated with these ensemble models are shown in Table 4C. Performance improvement was observed using the ensemble models by up to 0.7% on average (F1 score) in MEAN across the 7 metrics compared to the individual base models. No single ensemble model performed better than the rest in predicting each TOAST classification; there was variability among models which predicted each TOAST classification most accurately (Tables 6-8). Spearman correlation and Cohen’s kappa values among the 9 base classifiers range from 0.78 and 0.81 (between RF* and SVC2) to 0.96 and 0.97 (between MEAN and MEDIAN), respectively. This observation supported the decision to use the consensus ensemble meta-model, designated as StrokeClassifier, to harness the varying predictive capacities of the 9 classifiers while diluting the bias introduced by individual models, bolstering the robustness and generalizability of the model’s output.
Table 7
Table 7 (Continued)
Table 8
StrokeClassifier demonstrated the following performance measures on average for predicting the 4-level outcome of non-cryptogenic stroke etiology: accuracy of 0.744, balanced accuracy of 0.710, weighted F1 of 0.740, and Cohen’s kappa of 0.629 (Table 4C), indicating substantial agreement with vascular neurologist-adjudicated stroke etiology. The mean accuracy of StrokeClassifier for each specific etiology versus not as a binary outcome ranged from 0.829 for TOAST 2 to 0.913 for TOAST 4 (Table 9). Table 4C. Validation results of the ensemble/meta models using combn1d.age.sex.v1 ( X( ^1)) Table 9. Performance of the consensus model for each TOAST classification Performance Validation Using 300 Repeated Multi-Fold CV Splits Since cross-validation strategies such as the 5-fold CV used for HPR are anchored to a particular seed number, which is subjective, 300 training-validation data splits were used by repeated multi-fold CV, RMFCV300, to derive better estimates of model performance and generalization errors. RMFCV300 was performed for the 4 best models refined by the HPR, focusing on AUCROC and AUPRC model performances metrics (FIGS.6A-7C; Tables S8-S10 in StrokeClassifier 2024). While there was variability in the magnitude of model performance measures for each TOAST class among the 4 models, all 4 models performed best in predicting TOAST 3 in terms of AUCROC, while they performed best in predicting TOAST 2 in terms of AUPRC, regardless of the number of CV folds employed. For each TOAST class, the means and standard deviations of both AUCROC and AUPRC for the CV fold repetitions consistently increased with the increasing CV folds across the 4 models. Analysis of Age-Sex Strata To evaluate whether there was heterogeneity in model performances based on age, sex, and race, model performances in age-sex-race subgroups was assessed using the RMFCV300 validation sets (Tables 10-11 and S12-S14 of StrokeClassifier 2024-18). It was observed that StrokeClassifier tended to perform worse in the stratum of Male/Age>=65, in particular for predicting TOAST 3 and 4 (lowest mean F1 of 64.6% and 36.3% across all strata, respectively). The stratum of Black or African American also showed a relatively worse performance for TOAST 1 (lowest mean F1 of 63.8%). In contrast, StrokeClassifier performed better in the stratum of Female/Age<65, in particular for predicting TOAST 3 and 4 (highest mean F1 of 80.6% and 68.7% across the strata, respectively). It is noted that all mean performance values were greater than 60%, except F1 scores in TOAST 4 for the 4 strata of Male (51.4%±8.1%), Age>=65 (50.8%±10.4%), Male/Age>=65 (36.3%±16.9%), Male/Age<65 (56.1%±8.9%), White (59.9%±6.5%), Black or African American (53.4%±21.7%), and Others (57.7%±13%). Table 10. Performance of StrokeClassifier in age-sex-race strata Table 11 Feature Importance Analysis Feature importance or the contribution of features to predict TOAST classification were examined by SHAP analysis for each of the 4 refined base models. The top 10 features in terms of mean absolute SHAP values for each model are shown in FIG.8A. The top feature for all 4 models is AF. The top second feature is either frontal location of infarct noted on radiography or patient age. For PCA, the top two features are PC1 and PC3 (the second and fourth principal components, respectively; 0-indexed). The largest impact of both AF and PC1 is on TOAST 2. The top 10 features for each class for each model were also examined as shown in FIG.8B. The features that contribute the most to prediction of TOAST 1 by all models were AF, carotid occlusion, and atherosclerosis; for TOAST 2 were AF, patient age, and frontal location of infarct; for TOAST 3 were frontal location of infarct, occluded middle cerebral artery, AF, and thalamus location of infarct; and for TOAST 4, patient age, AF, and hypercoagulability or thrombophilia. For the PCA-based refined models, we examined the top 5 PCs and the top 10 most contributing features for each PC for each class (FIG.9, Table S15 See StrokeClassifier 2024). Similar important features were observed including age, sex, and NIHSS. This method identified multiple unique features contributing to stroke etiology classes. For example, the following 6 features in PC11 were unique to TOAST 2 by 3 models (SVC*, XGB*, and RF*): blood pressure (HEX), mass of body region (C0577573), Macrophage Activation Syndrome (C1096155), cyclic neutropenia (C0221023), sinus (HRT), and hemorrhagic (RAD). The following 4 features in PC10 are unique to TOAST 3 by 3 models (LR*, SVC*, and XGB*): left ventricular hypertrophy (HRT; C0149721), pericardial effusion (C0031039), and agitation (C0085631). The top features by the model-agnostic Kolmogorov-Smirnov test and Student’s t- test are largely in agreement (FIG.10 and Supplementary Information). Analysis of Misclassification Misclassified samples for each class and the top 10 features of the highest frequency among those misclassified samples were examined. Classification results by StrokeClassifier were analyzed for both training and validation from the merged RMFCV300 results. The misclassification or error rates (= 1 – accuracy, Table 12) for training were 4.5 ±0.6%, 5.3 ±0.7%, 2.5 ±0.4%, and 2.0 ±0.4% for the 4 classes, respectively, and those for validation were 16.2 ±1.4%, 16.8 ±1.7%, 9.4 ±1.2%, and 9.4 ±1.2% for the 4 classes, respectively. The top 10 most frequent features among misclassified samples for each class in each training or validation set are found to be present in >=54.8% of those samples (Table S16 See Stroke Classifier 2024). Frequencies of those top 10 features in the 300 training or validation sets for each misclassified class are shown in Table 12 and FIG.11. There are 6 features which are among the top 10 in all of the 300 training or validation sets: cerebrovascular accident, ejection fraction, body substance discharge, respiratory rate, sodium, and infantile neuroaxonal dystrophy. Table 12. Top 10 features of the highest frequency for misclassification by the consensus model Model Generalizability by 5-way Cross-Hospital and Longitudinal Validation To test the model generalizability, the 9 base models (with X were applied to the curated MIMIC discharge summaries (Tables 13A-B).3 versions of the MIMIC data were used as external validation: (1) MIMIC0 = 375 non-cryptogenic samples with 1,406 features in common with YNHH and MGH, (2) MIMIC1 = 405 non-cryptogenic samples imputed by Random Forests using MICE, and (3) MIMIC2 = 405 non-cryptogenic samples imputed by random sampling using MICE. For MIMIC1, AUCROC ranged from 0.834 to 0.860 (0.847 ±0.009), accuracy from 0.667 to 0.711 (0.691 ±0.014), and F1 from 0.587 to 0.717 (0.690 ±0.039) by the 9 base classifiers, while StrokeClassifier showed AUCROC of 0.809, AUPRC 0.719, accuracy of 0.699, F1 of 0.708, and kappa 0.467 (Table 13A). Performances in MIMIC0 and MIMIC2 or those by the PCA-based models were similar (Table 14). Overall, the performance of StrokeClassifier in the external dataset was reduced by less than 5% in comparison with the internal 5-fold CV (Table 4C). Class-wide performances of StrokeClassifier in MIMIC1 was also examined. Prediction of TOAST 1 was associated with the lowest PPV of 37.0%, the lowest kappa of 0.377, and the highest false positive rate (FPR) of 11.4%; Prediction of TOAST 2 was associated with the lowest accuracy of 78.0%, the lowest F1 of 78.2%, the highest false negative rate (FNR) of 12.3%, the highest PPV of 84.1%, and the highest kappa of 0.535; Prediction of TOAST 3 was associated with the highest accuracy of 94.1%, the highest F1 of 94.6%, the lowest FPR of 4.0%, and the lowest FNR of 2.0%; performance measures for predicting TOAST 4 were moderate (Table 13B).. Similar performances are observed for MIMIC0 and MIMIC2 (Table 15). Table 13. Model generalizability (A) Global performances (weighted averages over all classes) on MIMIC by individual models (B) Cross hospital and longitudinal class-wide performances by StrokeClassifier Table 14 Table 15 For an additional test of generalizability with X the 4 base models were trained and refined the same way as above using the MGH data of 1,002 non-cryptogenic samples and applied to the YNHH and MIMIC data for external validation (Tables 13B and 15). The 4 best models, LR*MGH, SVC*MGH, XGB*MGH, and RF*MGH, yielded mean cross-validated AUCROC of 91.0%, 90.9%, 92.3%, and 91.1%, respectively, and accuracy of 74.4%, 73.6%, 76.8%, and 68.1%, respectively. The external validation of the YNHH and MIMIC1 data by StrokeClassifier resulted in accuracy of 68.9% and 70.9%, respectively. Similarly, the models using the YNHH data of 930 non-cryptogenic samples were next tested for training and the MGH and MIMIC data for external validation (Tables 13B and 15). The 4 best models, LR*YNHH, SVC*YNHH, XGB*YNHH, and RF*YNHH, yielded mean cross-validated AUCROC of 86.8%, 86.5%, 87.6%, and 87.3%, respectively, and accuracy of 69.4%, 68.6%, 69.4%, and 60.6%, respectively. The external validation of the MGH and MIMIC1 data by StrokeClassifier resulted in accuracy of 70.3% and 66.4%, respectively. Performances in MIMIC0 and MIMIC2 were similar (Table 15). To address a longitudinal useability of StrokeClassifier, the model was re-trained and optimized with a new training set of discharge summaries from 2015 to 2019 in the combined cohort of YNHH and MGH and then longitudinally validated the optimal model using a test set from 2020. The performances are AUCROC of 86.8%, AUPRC of 71.4%, accuracy of 74.2%, F1 of 74.0%, and Cohen’s kappa of 0.64 for multi-class classification. For binary classification of each of the 4 TOAST classes, accuracy and F1 range from 83.2% to 90.6% (Table 13B). Predicting Etiologies of Cryptogenic Stroke Using StrokeClassifier The next aim was to classify a potential etiology of strokes in a cohort of adjudicated cryptogenic strokes using a variety of certainty heuristics as a proof-of-concept. In the pooled cohort of YNHH, MGH, and MIMIC1 datasets, there were a total of 788 stroke patients (285, 409, and 94, respectively), which were deemed to be cryptogenic strokes by vascular neurologists (Table 16). The heuristic that was employed in this study was built on a threshold of the first quartile (25% or moderate confidence) of the number of consensus supports among the 9 base classifiers for each TOAST classification based on the MIMIC1 external validation results: 7 supports for TOAST 1, 9 for TOAST 2, 7.2 for TOAST 3, and 7 for TOAST 4 (Tables 17A-B). If the number of supports for a particular sample was greater than or equal to the prespecified TOAST class threshold, the ischemic stroke was classified as the corresponding TOAST class. If the number of supports was less than any of the pre-specified TOAST class thresholds, the etiology was classified as persistently cryptogenic. Table 16 shows distributions of predicted TOAST classifications of cryptogenic patients for each cohort and the pooled cohort. FIG.12A also depicts the distributions of TOAST classification of the full cohort as adjudicated by vascular neurologists versus StrokeClassifier. Predictions for 46.3%, 54.5%, and 37.2% of the cryptogenic samples of YNHH, MGH, and MIMIC1 were agreed by all the 9 base classifiers, respectively. The prediction agreement by at least 8 base classifiers was observed for 69.8%, 72.6%, and 61.7% of the cryptogenic samples of YNHH, MGH, and MIMIC1, respectively. The most frequently predicted etiology was TOAST 2 for YNHH and MGH (32.6% and 37.9%, respectively) and TOAST 1 for MIMIC1 (27.7%), whereas the least frequently predicted etiology was TOAST 4 for YNHH and MGH (6.7% and 5.9%, respectively) and TOAST 3 for MIMIC1 (5.3%) (Table 16). The percentages of persistently cryptogenic samples for YNHH, MGH, and MIMIC1 were 30.9%, 27.1%, and 27.7%, respectively (Table 16). In other words, 28.6% of all cryptogenic samples (225 out of 788) were not predicted with high confidence by StrokeClassifier and remain cryptogenic. This reduced the percentage of cryptogenic patients from 25.2% to 7.2% in the full cohort of 3,125 stroke patients in YNHH, MGH, and MIMIC (FIG.12A). In contrast, when a certainty heuristic of the third quartile number of consensus supports (high confidence) was used, 9.9% of cryptogenic patients (309 cryptogenic patients of the full cohort; Tables 17A-B) remained persistently cryptogenic. Table 16. Application of StrokeClassifier to cryptogenic stroke patients Table 17A. Summary statistics of the number of consensus supports by the 9 base classifiers Table 17B. Persistently cryptogenic among the 788 cryptogenic samples by prediction confidence level cutoffs Finally, a repertoire of EHR signatures of predicted TOAST classes was generated for cryptogenic strokes (excluding the 225 persistently cryptogenic strokes) using feature frequencies from StrokeClassifier. The focus was on those features which were present in >50% of the cryptogenic stroke samples in each predicted class.26 such features (FIG.12B) were identified. Six of these 26 features were class-specific with p-value < 0.01 by chi-squared tests: hypercoagulability/thrombophilia (high-frequency for TOAST 4; p = 1.19e-15), AF (high- frequency for TOAST 2; p = 2.69e-12), basal ganglia (high-frequency for TOAST 3; p = 2.93e- 12), age >65 (low-frequency for TOAST 4; p = 1.68e-05), frontal (low-frequency for TOAST 3; p = 8.60e-05), and hypertensive disease (low-frequency for TOAST 4; p = 5.66e-03). DISCUSSION A novel, accurate, and computationally efficient automated tool, StrokeClassifier, was developed and validated to predict AIS etiology using EHR text-based data collected during the stroke hospitalization. StrokeClassifier is a meta-classifier of a majority voting ensemble built from 9 base classifiers trained using adjudicated outcomes curated from institutions with vascular neurology expertise. Standardized CUI features extracted from unstructured or semi- structured text corpora by an NLP method were particularly powerful predictors. It was found that the predictive capacity of StrokeClassifier was generalizable in 5-way external validation cohorts as well as a longitudinal analysis. This work is a promising multi-cohort and multi-class study of stroke subtype classification. The external and longitudinal validation accuracies were about 70% and 74%, respectively, for multi-class classification, while they were 77-96% for binary classification. These accuracies are higher than the minimum accuracy of 70% desired by a convenience sample of 13 international clinicians who care for stroke patients to adopt an Al stroke etiology diagnostic tool into clinical practice. By applying StrokeClassifier to a cohort of cryptogenic stroke patients to predict non-cryptogenic stroke etiologies with a certainty heuristic, the proportion of ischemic stroke patients in the full cohort with a persistently cryptogenic diagnosis was 7.2%, which was 71% lower than the rate adjudicated by vascular neurologists. It is believed that StrokeClassifier can aide stroke etiology diagnosis during the stroke hospitalization and timely administration of secondary stroke prevention therapies. It may also inform future clinical and population research investigations.
In contrast to prior attempts at machine learning classifiers for ischemic stroke TOAST classification, the present Example utilized a stepwise approach, with a goal of ultimately classifying subtypes. Cryptogenic samples were not considered during training because they were comprised of a mixture of potential etiologies. Instead, distributions of the 4 predicted non- cryptogenic etiologies were investigated for cryptogenic samples. Various certainty heuristics were then developed to predict the probability of stroke etiologies, both non-cryptogenic and persistently cryptogenic. This scalable property of StrokeClassifier is promising since the patients it is tasked to classify will not be pre-specified as cryptogenic or non-cryptogenic. Additionally, in contrast to prior stroke etiology classifiers that were trained and tested at a single center and which may not generalize to other centers in the U.S. or globally, StrokeClassifier was tested in separate hospital cohorts with various EHR systems and robustness was demonstrated. Furthermore, the present Example leveraged the UMLS conceptual framework developed by the National Library of Medicine to ensure the operability of StrokeClassifier irrespective of clinician and computer environment. For computational efficiency, PCA was utilized to capture multi-dimensional contributions of a wide array of features. StrokeClassifier was uniquely trained on adjudicated stroke etiologies upon review by at least two board-certified vascular neurologists. Since there was variability among individual optimized models in predicting each etiology, the 4 optimized models along with SVC2 were aggregated into ensemble models. The ensemble modeling (meta-model) herein includes a diversity of models with summary-statistic based ensemble models.
Several measures were taken to minimize bias. To address overfitting, investigated sub- optimal models within 1 standard deviation of the optimized models were investigated in terms of AUCROC, showing performance reduction by up to 4% across different metrics and CV folds. Additionally, in an effort to offset bias introduced by relying on a single choice of CV folds and a particular random seed, the RMFCV300 strategy analysis offers a more robust framework to assess model performance and generalization errors. Finally, SHAP analyses were performed to assess the degrees to which features contributed to stroke etiology prediction. The features contributing to the prediction of each stroke etiology were biologically plausible, lending validity to StrokeClassifier.
There are multiple potential applications of a trained, automated, accurate, and computationally efficient stroke etiology classifier. It can be implemented in health systems to perform the complex task of synthesizing the copious, semi -structured data collected during an AIS hospitalization and rapidly classifying the underlying stroke etiology in an automated manner for millions of patients. Most proximally, automated stroke etiology prediction can cue a treating clinician to consider instituting a targeted treatment by reducing diagnostic uncertainty, diagnostic errors due to human cognitive biases, oversight, and therapeutic inertia. In healthcare settings where vascular neurology expertise is sparse or unavailable, StrokeClassifier may be especially valuable. A classifier such as StrokeClassifier can be harnessed by informaticians to create nudges or progress notes indicating predicted etiologies and guideline-recommended therapies for individual patients. Stroke etiology data fields collected by manual extraction are currently incomplete in registries in the U.S. at all levels, and when populated, are often inaccurate. Stroke etiology predictions can be linked to institutional, regional, and country -wide registries to facilitate quality improvement, clinical trials, public health, and health services research efforts. Finally, it may identify patients with established stroke etiologies and risk factors which may render them eligible for clinical trials studying novel secondary stroke prevention therapies.
While StrokeClassifier was trained with the task of classifying etiology at the time of discharge, the predictive factors identified may be collected at an earlier timepoint during the hospitalization. The classifier was trained using data collected during the course of the AIS hospitalization and populated into the discharge summary, which is typically finalized at the completion of the hospital encounter. It was observed that the sources of information that contributed most to the model’s diagnostic performance individually and in the leave-one-out analysis in descending order as presented in Tables 4A-C were (1) concept unique identifiers or CUIs (AUCROC range: 0.87-0.89), (2) radiologic features of neuroanatomic location of the ischemic stroke, vessel patency, and hemorrhagic transformation (AUCROC range: 0.76-0.77), and (3) cardiac features from electrocardiographic and echocardiographic reports (AUCROC range: 0.61-0.63). While CUIs representing a baseline medical history, conventional neuroimaging such as computed tomography with angiography, and electrocardiograms are collected at the time of presentation during an acute stroke code, other data such as diagnoses accrued during the stroke hospitalization encounter, advanced neuroimaging such as magnetic resonance imaging, and cardiac imaging including echocardiography are typically obtained during later timepoints, if at all, depending on the resources and level of expertise housed within a healthcare setting.
The capacity to predict an underlying etiology of cryptogenic strokes using StrokeClassifier is promising. The predicted etiology among cryptogenic patients in the YNHH and MGH cohorts was predominantly cardioembolism, varying from 33% to 38%, followed by large artery atherosclerosis in 19% to 22%. Secondary analysis of the NAVIGATE ESUS study demonstrated that among ESUS patients, there were multiple potential etiologies including atrial cardiopathy (37%), left ventricular disease (36%), and arterial atherosclerosis (29%), with no potential etiology found in only 23% of patients and more than 1 potential etiology in 41% of patients. Given that many cryptogenic stroke patients have multiple potential sources, applying an algorithm such as StrokeClassifier can be especially fruitful because its supervised learning of features that may non-linearly associate with etiologies may be transferable. StrokeClassifier is a majority -voting consensus prediction tool from multiple base classifiers. This property is harnessed to address uncertainty that arises when a patient has multiple competing potential sources of stroke. This is represented by StrokeClassifier assigning confidence levels in terms of the degree of agreements among the base classifiers, a construct denoted herein as a certainty heuristic. When the number of individual classifiers voting for two potential etiologies is equal for a patient, the patient’s etiology is classified as cryptogenic due to uncertainty. To provide interpretability in instances when an etiology is deemed cryptogenic due to multiple potential sources, the output of StrokeClassifier can include voting results of the individual classifiers so that the user is informed about the percentage of classifiers which voted for a particular etiology (e.g. Table S20 See StrokeClassifier 2024 for the MIMIC data).
EHR signatures corresponding to the predicted etiology of cryptogenic stroke patients were derived. It begins to provide a conceptual and workflow framework for strokes traditionally deemed cryptogenic. For instance, cryptogenic patients with predicted etiology of large artery atherosclerosis by StrokeClassifier tend to be older and have frontal infarct, hypertension, and no AF. Thus, predicted stroke etiology classification of patients with these features during stroke hospitalization may prompt deeper, streamlined inquiry into this potential mechanism such as more advanced vascular imaging to assess the characteristics of a sub-stenotic carotid plaque. It may also obviate the need for broad, unnecessary testing that leads to health care expenditure. Predictions may also tip clinicians uncertain about which of multiple competing etiologies led to the stroke into a singular direction. This information and subsequent diagnostic investigation may then lead to initiation of evidence-based targeted secondary stroke prevention therapy. Finally, in an era of biomarker-based clinical studies, the potential stroke etiology signatures yielded by classifiers such as StrokeClassifier may advance research by identifying an enriched population of cryptogenic ischemic stroke patients who may benefit from specific trial interventions for secondary stroke prevention. In conclusion, present herein is StrokeClassifier, a validated diagnostic tool developed using an innovative modeling strategy which allows automated, real-time classification of stroke etiology in an accurate and computationally efficient manner with EHR text data inputs. Its immediate application may be as a clinical decision support tool to aide in the diagnosis of stroke etiology, prompting targeted secondary stroke prevention therapies in a timely manner. Furthermore, StrokeClassifier may facilitate abstraction of stroke etiology in population-based registries to aide epidemiologic, health policy, and clinical research efforts. SUPPLEMENTARY INFORMATION Supplementary Notes 1. Discretization of continuous values of the HEX and LAB features In order to reduce measurement noise or error, the continuous values of the HEX and LAB features were discretized into clinically relevant categories. Ejection fraction was dichotomized as < 40% which is defined as severely reduced versus >= 40%, NIHSS was dichotomized as <6 defining a minor stroke and >= 6, sodium level < 136 mmol/liter which is defined as hyponatremia and >= 136mmol/liter , BUN >= 24mg/dL which is the upper limit of its normal range including in the elderly and < 24mg/dL and per the clinical laboratories of Yale and MGH, ALT and AST < 36U/L versus >= 36U/L per the clinical laboratory of Yale (ucsfhealth.org/medical-tests/alanine-transaminase-(alt)-blood-test#), white blood cell count < 11x1000/microliter versus >= 11x1000/microliter which defines leukocytosis and per the clinical laboratories of Yale and MGH, hematocrit < 35% (anemia), 35-45% (normal), >=46% (erythrocytosis) per Yale and MGH clinical laboratories, hemoglobin in females < 11.7 (anemia), 11.7-15.5 (normal), and > 15.5 (erythrocytosis) per Yale’s clinical laboratory, hemoglobin in males < 13.2 g/dL (anemia), 13.2-17.1 g/dL (normal), and > 17.1 g/dL (erythrocytosis) per Yale’s clinical laboratory, triglyceride >= 200 mg/dL which defines hypertriglyceridemia and per Yale and MGH clinical laboratory versus < 200 mg/dL, HDL mg/dL < 40 versus >= 40 mg/dL, LDL >= 100 mg/dL versus < 100 mg/dL, TSH < 4.2 micro IU/mL versus >= 4.2 micro IU/mL, PTT < 29.9 versus >= 30 seconds per Yale clinical laboratory, and hemoglobin A1c >= 6.5% which defines diabetes versus < 6.5%. 2. Mathematical notations for classifier models In this work, M = 4 models (LR, SVC, XGB, RF), N = 2,700 samples, max(Ll) = 2,051 features, Q = 20 feature groups, and K = 4 TOAST classes were investigated. 3. Parameter configurations for hyperparameter optimization of the 4 base models Logistic regression LogisticRegression from the sklearn library in Python was used. The following parameter values were used for a grid search of 143 combinations with penalty = ‘elasticnet’ (elastic net, lasso, or ridge regularization), the saga solver, and 500 max iteration: C = (1e-2, 1e-1, 1e+0, 1e+1, 1e+2, 1e+3, 1e+4, 1e+5, 1e+6, 1e+7, 1e+8, 1e+9, 1e+10) and l1_ratio = (0.0 , 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0). The refined parameters are C = 0.01 and l1_ratio = 0.0. Support vector classifier SVC from the sklearn library in Python was used. The following parameter values were used for a grid search of 676 combinations with decision_function_shape = ‘ovr’ (one vs. the rest), class_weight = ‘balanced’, and 1000 max iteration: C = (1e-2, 1e-1, 1e+0, 1e+1, 1e+2, 1e+3, 1e+4, 1e+5, 1e+6, 1e+7, 1e+8, 1e+9, 1e+10), gamma = (1e-9, 1e-8, 1e-7, 1e-6, 1e-5, 1e-4, 1e-3, 1e-2, 1e-1, 1e+0, 1e+1, 1e+2, 1e+3), kernel = (linear, poly, rbf, sigmoid). Its refined parameters are C = 1.0 and gamma = 0.01 with the RBF kernel. For prediction probabilities, the default outputs are based on Platt scaling using the libsvm library. As Platt scaling is controversial, alternative prediction probabilities were also calculated using normalized decision_function scores implemented in sklearn based on the refined parameters for building downstream ensemble models and refer it to SVC2. XGBoost XGBClassifier from the xgboost library in Python was used. The hyperparameter refinement was performed by a grid search of 1,620 combinations of the following parameter values: n_estimators = (500, 1000); max_depth = (4, 5, 6); learning_rate = (0.01, 0.1, 0.3, 0.5, 1); gamma = (0.0, 5.0, 10.0); reg_lambda = (0.0, 0.5, 1.0); reg_alpha = (0.0, 0.5, 1.0); subsample = (1.0, 0.75). Its refined parameters are n_estimators = 1000, max_depth = 5, learning rate = 0.01, gamma = 0.0, reg_lambda = 0.0, reg_alpha = 0.0, and subsample = 0.75. Random Forests RandomForestClassifier from the sklearn library in Python was used. The following parameter values were used for a grid search of 48 combinations with min_samples_leaf = 2 and the saga solver: n_estimators = (200, 500, 1000); max_depth = (10, 20, 50, 100); criterion = (gini, entropy); max_features = (sqrt, log2). Its refined parameters are n_estimators = 1000, max_depth = 20, criterion = ‘gini’, and max_features = ‘sqrt’. 4. Ensemble models and consensus meta-model, StrokeClassifier 4 summary statistics-based ensemble models were first built using the 4 refined best models along with SVC2 for a feature group, F, as base models, B = { LR, SVC, SVC2, XGB, RF}. The 4 ensemble models are mean, median, maximum, and minimum normalized functions of prediction probabilities, Pb, of the base models mapping from each sample, si , i = {1,2, ... , n} to each class or label, l ∈ {1, 2, 3, ... , k}: An ensemble consensus-by-voting model, StrokeClassifier, was then built using those 5 refined and 4 ensemble models as base models, i.e., a 9-classifier meta-model as our final model: 5. Other models of stacked generalization With the 4 refined base models for the feature group of combn1d.age.sex.v1.maxinfo, several ensemble models of stacked generalization were tested. Different combinations of those 4 refined models were taken as level-0 or base models and each of LR and SVC as the level-1 or meta model.5-fold CV was performed with seed = 1701 for this purpose. There was no significant improvement over the refined base models, except for RF* (See FIG.13; “ST|...” for stacked models). 6. Feature importance analysis alternative to SHAP As an alternative method to the SHAP analysis (FIG.8) to examine feature importance in a model-agnostic way, Kolmogorov-Smirnov tests and Student’s t-tests were performed for each of 2,027 features for one class versus the rest. The correlations between |t| or D statistics (or their p-values) and means of absolute SHAP values averaged over the 4 models for the 4 classes ranged between 0.43 and 0.89 (FIG.10). Thus, the top features identified by the model-specific and model-agnostic methods for each class were largely concordant. 7. Supplementary Discussion Compared to previous studies on stroke etiology classification, StrokeClassifier performed well with respect to a variety of metrics. While Garg et al. reported Cohen’s kappa values of 0.25 using radiology reports alone and 0.57 using combined data (Garg et al., 2019), highest cross-validated kappa values of 0.400 ±0.029 using radiology features alone by XGB* and 0.632 ±0.017 using X by LR* (Tables 5A-C) were achieved herein. It is noted that Cohen’s kappa is considered to be a controversial statistic to quantify agreement between two raters and compare different studies with different categories or classes. On the other hand, compared to the c-statistic (i.e., AUCROC) of 0.85 reported by Kamel et al. for a binary classifier (Kamel et al., 2020), the mean c-statistics achieved by the 9 base models used to build StrokeClassifier were higher (0.887-0.912; Table 4C). Higher c-statistics were also achieved for each individual stroke subtype than by Turner et al. (Turner et al., 2022) and comparable c- statistics to Wang et al. (Wang et al., 2022) (FIGS.6A-B). More importantly, StrokeClassifier outperformed both in terms of predicting stroke etiologies that were in closest agreement with etiologies diagnosed by board-certified vascular neurologists upon review of the entire stroke hospitalization medical record. Furthermore, in comparison to machine/deep learning classifiers in other diseases, the current study achieved better classification performances with neither big data nor deep learning, demonstrating robust generalizability. A deep learning study on classification of low vision achieved AUCROC of 82% and AUPRC of 79%, worse than the present RMFCV300 performance with TOAST 2 of similar prevalence of about 40% (AUROC = 90.9 ±1.2%, AUPRC = 87.9 ±1.7%; Table S10 See StrokeClassifier 2024). The authors used >5,500 EHRs from a single-center cohort and a different NLP tool to extract CUIs. Another study on prediction of red blood cell transfusion needs for acute gastrointestinal bleeding by a recurrent neural network model of long short-term memory with 2,032 training samples and 62 features showed AUCROC of 81% for internal validation and 65% for external validation. Machine learning models extracting cardiovascular disease from text performed similarly with significantly more data. There was variability in the predictive capacity of StrokeClassifier for each TOAST class. For instance, the model’s accuracy or F1 for predicting cardioembolism (TOAST 2) was lowest at 83% with a high false positive rate of 10% (Table 9). The largest contributor to the cardioembolic etiology prediction was AF (FIG. 8). However, not all ischemic strokes among patients with AF are due to cardioembolism. Patients with AF share risk factors for large artery atherosclerosis and small vessel disease. Further model fine-tuning is necessary to learn the roles of other features and stroke etiology in the context of AF. StrokeClassifier" s predictive capacity for stroke etiology also varied by age and sex subgroups. The performance for predicting the large artery atherosclerosis etiology (TOAST 1) per the metrics of the F1 score and balanced accuracy was lower among females, especially those 65 years and older (Table 3). The F1 score and balanced accuracy were lower for rare causes of stroke (TOAST 4) among older patients, particularly among older males (Table 10). It is unclear what is driving these differences, but we hypothesize that these patterns are reflective of the real-world prevalence of these etiologies within each subgroup (Table 11).
The disclosures of each and every patent, patent application, and publication cited herein are hereby incorporated herein by reference in their entirety.
While this invention has been disclosed with reference to specific embodiments, it is apparent that other embodiments and variations of this invention may be devised by others skilled in the art without departing from the true spirit and scope of the invention. The appended claims are intended to be construed to include all such embodiments and equivalent variations.
REFERENCES
The entirety of the references cited herein are incorporated herein by reference.

Claims

CLAIMS What is claimed is: 1. A method of training an algorithm to predict ischemic stroke etiology, the method comprising: processing electronic health records (EHRs) in a derivation dataset, the processing including: extracting clinical variables from the EHRs by detecting concept unique identifiers (CUIs) by applying natural language processing (NLP) tools; and extracting other covariates from the EHRs through regular expressions; wherein the clinical variables and other covariates form a training dataset; training two or more base models with the training set given certain hyperparameters, the base models selected from the group consisting of: Random Forests (RF); XGBoost (XGB); support vector classifier (SVC); and logistic regression (LR); refining the two or more base models using a number of different combinations of hyperparameter values to form two or more refined base models; and building an ensemble model with the two or more refined base models.
2. The method of claim 1, wherein each of the CUIs belongs to a category selected from the group consisting of: disease or syndrome; neoplastic process; sign or symptom; and combinations thereof.
3. The method of claim 1, wherein the other covariates include one or more of: age; sex; clinical information not captured by the CUIs (HEX); radiological features (RAD); cardiac features (HRT); and laboratory features (LAB).
4. The method of claim 3, wherein the HEX includes one or more of: social history; National Institutes of Health Stroke Severity scale; and vital signs.
5. The method of claim 3, wherein the RAD includes one or more of: neuroanatomical location of the ischemic stroke; presence of moderate or severe stenosis or occlusion of specific head and neck arteries; and occurrence of intracranial hemorrhage encoded as a binary variable.
6. The method of claim 5, wherein the one or more features are selected from the group consisting of: magnetic resonance imaging of the brain, magnetic resonance angiography of the head, magnetic angiography of the neck, computed tomography of the head, computed tomography angiography of the head, computed tomography angiography of the neck, carotid ultrasound, transcranial Doppler, conventional angiography
7. The method of claim 3, wherein the HRT include one or more features from electrocardiography reports, transthoracic and transesophageal echocardiography reports, or combinations thereof.
8. The method of claim 7, wherein the one or more features are selected from the group consisting of: reduced left ventricular ejection fraction, vegetation, thrombus, mass, patent foramen ovale, mitral annular calcification, wall motion abnormality, left ventricular hypertrophy, aortic atheroma, aortic stenosis, mitral regurgitation, aortic regurgitation, mitral valve prolapse, aortic root dilation, diastolic dysfunction, pericardial effusion, T wave inversion, prolonged PR interval, premature atrial contraction, premature ventricular contraction, tachycardia, bradycardia, left ventricular hypertrophy, prolonged QT interval, atrial fibrillation, and supraventricular tachycardia.
9. The method of any one of claims 3 to 8, wherein one or more of the HRT, LAB, or HRT and LAB features are discretized.
10. The method of claim 9, wherein the discretized features include a cutoff between high and normal levels of: ejection fraction – 40; NIHSS – 6; sodium – 136; BUN – 24; ALT – 36; AST – 36; white blood cell count – 11; triglycerides – 200; HDL – 40; LDL – 100; TSH – 4.2; PTT – 29.9; A1C – 6.5.
11. The method of claim 9, wherein the discretized features include cutoffs between high and normal and normal and low levels of: hematocrit – 35 and 46; hemoglobin for female – 11.7 and 15.5; hemoglobin for male – 13.2 and 17.1.
12. The method of claim 1, further comprising the step of reducing feature dimensionality through principal component analysis (PCA).
13. The method of claim 1, wherein the hyperparameters are selected through a grid search of a pre-defined hyperparameter space.
14. The method of claim 1, wherein refining the two or more trained base models includes validating the models through a stratified cross-validation (CV) strategy of 5 splits with 20% validation sets.
15. The method of claim 1, further comprising training and validating the two or more refined base models using a repeated multi-fold CV strategy.
16. The method of claim 15, wherein the repeated multi-fold CV strategy includes performing 2-fold, 3-fold, 4-fold, 5-fold, and 10-fold CV with 30, 20, 15, 12, and 6 repetitions with different random seeds, respectively.
17. The method of claim 1, wherein the ensemble model comprises a refined RF model, a refined XGB model, a refined SVC model, a refined second support vector classifier model (SVC2), a refined LR model, a MEAN model, a MEDIAN model, a MAX model, and a MIN model.
18. The method of claim 17, wherein the ensemble model generates a TOAST prediction using a voting system among the refined RF model, the refined XGB model, the refined SVC model, the refined LR model, the refined SVC2 model, the MEAN model, the MEDIAN model, the MAX model, and the MIN model.
19. The method of claim 18, wherein the ischemic stroke etiology is selected from the group comprising: large artery atherosclerosis (TOAST 1); cardioembolism (TOAST 2); small vessel disease (TOAST 3); and other determined (TOAST 4).
20. An apparatus for predicting ischemic stroke etiology, the apparatus comprising: a processor; a memory unit; and a communication interface; wherein the processor is connected to the memory unit and the communication interface; and wherein the processor and memory are configured to implement the method of any one of the previous claims.
21. A computer readable storage medium storing computer-executable instructions for performing the method of any one of claims 1-19.
EP24816455.0A 2023-05-30 2024-05-30 Methods of training an algorithm to predict ischemic stroke etiology Pending EP4719177A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202363505006P 2023-05-30 2023-05-30
PCT/US2024/031761 WO2024249688A2 (en) 2023-05-30 2024-05-30 Methods of training an algorithm to predict ischemic stroke etiology

Publications (1)

Publication Number Publication Date
EP4719177A2 true EP4719177A2 (en) 2026-04-08

Family

ID=93658847

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24816455.0A Pending EP4719177A2 (en) 2023-05-30 2024-05-30 Methods of training an algorithm to predict ischemic stroke etiology

Country Status (2)

Country Link
EP (1) EP4719177A2 (en)
WO (1) WO2024249688A2 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3029466A1 (en) * 2014-12-03 2016-06-08 Fundació Hospital Universitari Vall d' Hebron - Institut de Recerca Methods for differentiating ischemic stroke from hemorrhagic stroke
US20170140115A1 (en) * 2017-01-24 2017-05-18 Liwei Ma Intelligent stroke risk prediction and monitoring system
WO2022251747A1 (en) * 2021-05-28 2022-12-01 Tempus Labs, Inc. Ecg-based cardiovascular disease detection systems and related methods

Also Published As

Publication number Publication date
WO2024249688A2 (en) 2024-12-05
WO2024249688A3 (en) 2025-01-23

Similar Documents

Publication Publication Date Title
Seetharam et al. Artificial intelligence in cardiovascular medicine
Newaz et al. Survival prediction of heart failure patients using machine learning techniques
Al-shanableh et al. Advanced ensemble machine learning techniques for optimizing diabetes mellitus prognostication: A detailed examination of hospital data
Ahmad et al. Clinical implications of chronic heart failure phenotypes defined by cluster analysis
Yoon et al. Application and potential of artificial intelligence in heart failure: past, present, and future
Lee et al. StrokeClassifier: ischemic stroke etiology classification by ensemble consensus modeling using electronic health records
EP3433614A1 (en) Use of clinical parameters for the prediction of sirs
Ru et al. Comparison of machine learning algorithms for predicting hospital readmissions and worsening heart failure events in patients with heart failure with reduced ejection fraction: modeling study
US11449792B2 (en) Methods and systems for generating a supplement instruction set using artificial intelligence
WO2020205140A1 (en) Methods and systems for utilizing diagnostics for informed vibrant constitutional guidance
Qadri et al. Heart failure survival prediction using novel transfer learning based probabilistic features
Nasiruddin et al. Predicting heart failure survival with machine learning: assessing my risk
Abdalrada et al. Predicting diabetes disease occurrence using logistic regression: An early detection approach
Wang et al. Enabling chronic obstructive pulmonary disease diagnosis through chest X-rays: A multi-site and multi-modality study
Sanders Jr et al. Machine learning: at the heart of failure diagnosis
Sumon et al. CardioTabNet: a novel hybrid transformer model for heart disease prediction using tabular medical data
Chandralekha et al. Clinical decision system for chronic kidney disease staging using machine learning
Penikalapati et al. Healthcare analytics by engaging machine learning
Sadiku et al. Classification model for prediction of hypertension
EP4719177A2 (en) Methods of training an algorithm to predict ischemic stroke etiology
Kumar et al. An efficient diagnosis of heart disease using optimized cross-layer Densenet121 pyramid mutual attention network
Sumathi et al. Enhancing diabetes prediction with minimal processing time using catboost: a comparative study
Bilal et al. Using Machine Learning Models for The Prediction of Coronary Arteries Disease
Osman et al. Breaking new ground in cardiovascular heart disease Diagnosis K-RFC: An integrated learning approach with K-means clustering and Random Forest classifier
Khan Biochemical biomarker–Driven deep learning framework with SHAP-based feature interpretation for diabetes classification

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251229

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR