Self-explaining Artificial Intelligence for the Classification of B cell Non-Hodgkin Lymphoma

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Artificial intelligence (AI) systems have been proposed for multiparameter flow cytometry (MFC) based immunophenotyping of leukemic Non-Hodgkin-Lymphoma (NHL). Lymphoma classification has progressively been revised due to the increasing molecular knowledge. In order to establish a ground truth for AI learning we sought to disclose the data’s structure embodiment in comparison to the histopathological classification decisions through self-organization of data using swarm intelligence. 19,493 data samples allowed for an unsupervised view based on multidimensional MFC data hereby providing information about the higher structures within the data and the inherent discrimination capabilities. Yet, rare lymphoma entities remain a particular challenge for AI training. We propose an AI system termed Flow XAI, which exhibits equal immunophenotyping performance as neural network based systems but has reduced the number of needed learning data by a factor of 100. The Flow XAI is capable of “self-reflection”, it reports a self-competence estimation for each case. Moreover, it selects and reports diagnostically relevant cell populations and expression patterns in a discernable and clear manner so that physicians can understand the rationale behind the AI’s decisions. The self-explaining AI system can therefore be used for real-world training on MCF based lymphoma immunophenotyping.
Full text 146,160 characters · extracted from preprint-html · click to expand
Self-explaining Artificial Intelligence for the Classification of B cell Non-Hodgkin Lymphoma | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Self-explaining Artificial Intelligence for the Classification of B cell Non-Hodgkin Lymphoma Michael Thrun, Joerg Hoffmann, Stefan Krause, Peter Krawitz, Quirin Stier, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6690890/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Artificial intelligence (AI) systems have been proposed for multiparameter flow cytometry (MFC) based immunophenotyping of leukemic Non-Hodgkin-Lymphoma (NHL). Lymphoma classification has progressively been revised due to the increasing molecular knowledge. In order to establish a ground truth for AI learning we sought to disclose the data’s structure embodiment in comparison to the histopathological classification decisions through self-organization of data using swarm intelligence. 19,493 data samples allowed for an unsupervised view based on multidimensional MFC data hereby providing information about the higher structures within the data and the inherent discrimination capabilities. Yet, rare lymphoma entities remain a particular challenge for AI training. We propose an AI system termed Flow XAI, which exhibits equal immunophenotyping performance as neural network based systems but has reduced the number of needed learning data by a factor of 100. The Flow XAI is capable of “self-reflection”, it reports a self-competence estimation for each case. Moreover, it selects and reports diagnostically relevant cell populations and expression patterns in a discernable and clear manner so that physicians can understand the rationale behind the AI’s decisions. The self-explaining AI system can therefore be used for real-world training on MCF based lymphoma immunophenotyping. Health sciences/Diseases/Cancer/Haematological cancer/Lymphoma/Non-hodgkin lymphoma/B-cell lymphoma Biological sciences/Cancer/Haematological cancer/Lymphoma/Non-hodgkin lymphoma/B-cell lymphoma Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Artificial intelligence (AI) has delivered tremendous changes and opportunities to the field of clinical diagnostic; particularly neuronal networks (NN) enabled a new era of image analysis techniques. Multiparameter flow cytometry (MFC) has been used for digital surface protein analysis of cell populations from peripheral blood for about four decades. Moreover, novel high parametric technologies such as mass cytometry and spectral flow cytometry provide numerous simultaneous measurements of variables. Thus, algorithms that facilitate the selection of relevant cell populations in machine learning-assisted analyses have been proposed [ 1 , 2 ]. Several AI algorithms were developed for automated diagnosis of leukemia and lymphoma [ 3 – 5 ], and a NN based approach for automated immunophenotyping of mature B-cell non-Hodgkin lymphoma (B-NHL) has recently been proposed [ 6 ]. Due to the limited availability of well-trained expert physicians automation of immunophenotyping is highly desirable for support and education of hematologists to interpret antigen expression patterns of different cell populations correctly. However, there are several obstacles related to this task. First, the definition of lymphoma classes has historically been fluid and therefore, taxation whether the given lymphoma identities are represented in the data is important. Moreover, recommendations for the optimal diagnostic panel of B-cell epitopes are still under investigation [ 7 – 9 ]. Rare lymphoma entities or the reduced availability of data samples in general impair performances of NN based algorithms hereby limiting such a diagnostic process to university hospitals or high throughput diagnostic centers. Transfer learning was proposed for automated B-NHL immunophenotyping to address the problem of large training dataset requirement and it had been claimed that the proposed AI systems deliver a performance comparable to that of human experts but functioning was strongly dependent on the amount of learning data and rare lymphoma entities were excluded [ 10 , 11 ]. Beyond performance and data availability, self-explanation is an important aspect of diagnostic AI. Many AI systems, particularly if based on NN, are “subsymbolic” [ 12 – 14 ]. Subsymbolic systems assign a diagnosis to a sample but are unable to provide reasons or explanations for their decisions [ 15 , 16 ], although they are skillful learners [ 17 ]. Therefore, with regard to patient safety, trustworthiness is an issue of current debate [ 18 ] and European Union laws require causal justification for decisions made by AI [ 19 , 20 ]. Finally, due to biological variances individual samples may differ from the common lymphoma class immunophenotype. To address this issue Matutes et al. had proposed a scoring system for chronic lymphocytic leukemia (CLL) [ 21 , 22 ]. With regard to these considerations, a self-explaining AI system FlowXAI was designed to address in particular the following three key elements: Integrating medical knowledge after the assessment of the structure embodiment in given data. Obtaining a sound performance, even if only very few samples are available in the learning phase. Providing self-explanation by reporting a self-competence estimation per sample and delivering of human-understandable explanations for decisions. Methods Characterization of the data sets The MLL9 and PUM2 datasets are entirely independent, both medically and technically. They originate from two separate diagnostic laboratories and were generated using different antibody panels. The MLL9 data set contains 18,274 samples from which 19,493 peripheral blood samples were selected for this work [ 6 ]. The panel of the Munich Leukemia Laboratory (MLL) consists of three 9-color tubes with the following antibodies: 1. CD19(APCA750) + CD45(KrOr) + FMC7(FITC) + CD10(PE) + IgM(ECD) + CD79b(PC5.5) + CD20(PC7) + CD23(APC) + CD5(PacBlue); 2. CD19(APCA750) + CD45(KrOr) + kappa(FITC) + lambda(PE) + CD38(ECD) + CD25(PC5.5) + CD11c(PC7) + CD103(APC) + CD22(Pac Blue); 3. CD19(APCA750) + CD45(KrOr) + CD8(FITC) + CD4(PE) + CD3(ECD) + CD56(APC) + HLA-DR(PacBlue). The PUM2 data set has uni-centrically been collected in Marburg [ 23 , 24 ] and employs two tubes: 1. CD45(KrOr) + kappa/CD8(FITC) + lambda/CD7(PE) + CD23(ECD) + CD79b/CD4(PC5.5) + CD5(PC7) + CD38(APC) + CD19(AP-A700) + CD20/CD3(APC-A750) + FMC7/CD2(PacBlue); 2. CD19(KrOr) + CD103(FITC) + CD43(PE) + CD25(ECD) + CD10(PC5.5) + CD200(PC7) + CD11c(AP-A700) + CD20(APC-A750) + IgM(Pac Blue). The flow cytometer measurements consist of N = 50,000 (MLL9) and N = 100,000 (PUM2) events for each tube. The Flow XAI system To ensure robust diagnostic accuracy while simultaneously addressing sample quality and computational efficiency, we developed FlowXAI — an AI system for automated, self-explaining sample classification designed to assist in the classification of B-cell non-Hodgkin lymphomas (B-NHL) using flow cytometry data. FlowXAI integrates multiple stages of computational analysis to mirror the diagnostic reasoning of human experts, incorporating structural analysis of underlying data consistency with clinical knowledge. Training data reduction by tile mining The first stage of the FlowXAI system, called Tile Mining (TM), aims to evaluate unsupervised the quality of the patient samples. TM identifies relevant structures in data to select representative samples. This reduces the requirement for large training data sets to these representative samples. TM systematically divides each bivariate marker combination into a fixed grid of uniform rectangular regions called hypercubes. All unique bivariate combinations of dd markers within one tube amount to \(\:\left(\genfrac{}{}{0pt}{}{d}{2}\right)=\frac{d(d-1)}{2}\) pairs. TM partitions each of these bivariate plots into 9 equally sized hypercubes, yielding a total of # \(\:H\left(d\right)=9*\frac{d(d-1)}{2}\) . Distinct hypercubes across the panel. For d = 11 antigens, this results in \(\:\#H\left(d=11\right)=495\) hypercubes. For each hypercube \(\:h\in\:H\left(d\right)\) , the density \(\:p\left(h\right)\) is defined as the proportion of associated events: $$\:p\left(h\right)=\sum\:_{e\:\:}^{}{f}_{e},{f}_{e}=\left\{\begin{array}{cc}1&\:e\:ϵ\:h\\\:0&\:e\notin\:h\end{array}\right.$$ Here, \(\:E\:\) denotes the set of all measured events (cells) in the sample. For each patient with a sample file \(\:i\) of tube \(\:t\) , the cell densities within all hypercubes are combined to measure strangeness \(\:\sigma\:\) as follows: $$\:{\sigma\:}_{i,t}=\sum\:_{h\in\:H\left(d\right)}^{}p\left(h\right)$$ The central limit theorem states that under very general conditions the sum of such variables attains a normal distribution \(\:{{N}_{\sigma\:}}_{i,t}(m,sd)\) . Samples are defined as atypical if their values \(\:{\sigma\:}_{i,t}\) lie outside the probability limits defined by the cumulative probability function of a robust estimated empirical normal distribution \(\:{{{N}^{{\prime\:}}}_{\sigma\:}}_{i,t}(m{\prime\:},sd{\prime\:})\) . As default, two-tailed limits of 1% are used. Any sample deemed atypical in at least one tube is excluded from further analysis, ensuring that only typical, representative samples enter the training set. Self-explanatory diagnostic AI committee The second stage consists of a committee of AI experts that self-explanatory diagnose a sample in a process similar to the human expert. Building on the curated sample set, the supervised ALPODS expert committee — an extension of the original ALPODS algorithm [ 25 ] — provides self-explanation through a Bayesian decision network. For each expert, a recursive process constructs a directed acyclic graph (DAG): at each node, a marker is selected based on the Simpson index [ 26 ], which quantifies the probability that two randomly drawn cells belong to the same category; conditional dependencies between parent and child nodes are then computed via Bayes’ theorem. Edges in the DAG represent these dependencies, and recursion terminates when further splitting yields subpopulations with homogeneous class labels or falls below a preset population size. The resulting network identifies key phenotypic subpopulations, which are then distilled into fast‐and‐frugal trees [ 27 ] by merging all decisions on the same marker into single range‐based conditions and ranking subgroups by effect size of Cohen’s d [ 28 ]. To capture the hierarchical nature of hematologic diagnoses, multiple tube- and task-specific experts are trained at three diagnostic levels and aggregated into an expert committee. At level 1 (distinguishing normal controls from B-cell lymphomas), TM first selects representative cases per class; computed ABC analysis [ 29 ] then partitions learned cell populations into high-frequency (set A) and low-frequency (sets B and C) groups, each handled by a dedicated expert. Depending on the dataset, four to six experts (one per group per tube) are trained, with inter-tube weighting reflecting their differential diagnostic value. Levels 2 and 3 further subdivide lymphoma subtypes: three experts distinguish chronic lymphocytic leukemia – like versus other lymphomas across all tubes, and a separate expert analyzes the tube most informative for hairy-cell leukemia. Each ALPODS expert independently provides a decision and an explanation for its decision that is ultimately aggregated into a final diagnostic output by a higher-order meta classifier. This the meta classifier uses the Gini impurity measure suited for categorical data to weigh and combine expert opinions. A weighting scheme reflects diagnostic significance per tube, assigning higher weights to more diagnostically relevant tubes. Details are provided in SI Fig. 1 and subsequent supplementary information. Self-assessment by degrees of trustworthiness The committee of AI experts yields additionally to the self-explanatory diagnosis a self-competency estimation as the likelihood \(\:{p}_{c}\) of a case being correctly classified as either an NC sample or a B-cell lymphoma sample ranging from zero to one. Recognizing the need for intuitive confidence measures in clinical contexts such as Likert scale [ 30 ], FlowXAI converts its internal probability estimates into a three-point self-assessment scale. For each sample, the self-competence \(\:{p}_{c}\) is binned as follows: The FlowXAI is Confident about the diagnosis of a sample if either \(\:{p}_{c}\) =1 or \(\:{p}_{c}=\) 0, Probable in the two bins of \(\:1>{p}_{c}\ge\:\:0.8\) or \(\:0{p}_{c}>0.2\) . Generalization and validation protocol To evaluate generalizability and robustness, FlowXAI was tested using two rigorous cross-validation protocol. The MLL9 a underwent 100 rounds of random class-balanced 80/20% train-test splits (MLL9F: N = 15,594 (train) / N = 3,899 (test) with and without TM's elimination of atypical samples in the training set by unsupervised tiles mining. In contrast to prior works [ 6 ], a repetitive cross-validation technique was chosen in order to prevent the underestimation of errors [ 31 , 32 ]. To show the efficiently of data reduction by TM, we selected N = 512 samples of typical training samples by TM and evaluate the performance on the test sets of all other typical and atypical samples. Next, for the PUM2 N = 200 training samples from the typical sample set were sampled 100 times to evaluated a test set of N = 317 + 121 samples. In all experiments, test samples were kept entirely independent throughout training to preserve external validity. Given the imbalance between different lymphoma subtypes and control samples, the Matthews correlation coefficient (MCC) was taken as the primary metric at Levels 2 and 3, owing to its superior handling of imbalance [ 33 ] is considered more reliable than the F1 measure [ 34 ]. The AI additionally reports overall accuracy, false positive rate, and false negative rate — especially when stratified by the three confidence categories. Evaluating the limits of performance, the set of training representatives was reduced to 256 lymphoma and 256 control samples. Visualization of the results with mirrored-density plots (MD-plots) [ 35 ] allowed assessment of probability density functions (pdfs) of quality measures like MCC. The MD-plot visualization can confirm if there is no multimodality within the values of the quality measures by revealing Gaussian distribution (magenta line) thereby allowing to report the point estimate as an average above. A non-gaussian distribution may indicate that the split between the test set and training set, i.e., random sampling, is not homogeneous. Structural embodiment of medical knowledge Assessment of the “structural embodiment” of known lymphoma classes was achieved through unsupervised swarm intelligence using the databionic swarm [ 36 ]. The Databionic swarm allows to assess the structures in data using cluster analysis [ 37 , 38 ], and, separately, provides an U-matrix visualization [ 39 ], [ 40 , 41 ]. This technique generated U-Matrix visualizations of the high-dimensional data [ 42 , 43 ], allowing intuitive interpretation of clustering behavior. The U-Matrix creates a topographic map where valleys represent similar sample properties and mountains represent structural dissimilarities. The multicore implementation of the Databionic swarm facilitated efficient computation [ 44 ], making it viable for large clinical datasets. Implementation details and access to FlowXAI We have created a live, interactive demonstration of our FlowXAI platform at https://plait.mathematik.uni-marburg.de/ . This online portal is specifically designed so that physicians, computational scientists, and readers can explore our proposed step-by-step classification workflow: Visitors can upload data (or select from example datasets) and see how FlowXAI automatically identifies relevant cell populations, performs classification, and provides self-explaining visual outputs. Each FlowXAI decision is grounded in traceable rules that mimic the logic of human experts. FlowXAI provides these understandable explanations stored in the FCS format. Upon acceptance we will publicly release the full code base for the platform. Results Bioinformatic view on lymphoma immunophenotype classes The presence or absence of groups or classes within a given disease entity influences the interpretation of high parametric expression data and machine learning capabilities [ 40 , 41 , 45 ]. Besides, the nomenclature and categories of lymphoma category definitions have progressively shifted due to the growing knowledge of molecular genetic aberrations and pathways [ 46 – 49 ] and a shift of categories is a particular challenge for AI design. Historical knowledge, for example, in the form of the WHO Lymphoma classification, may or may not be captured in the measurements of a given lymphoma immunophenotyping panel. Thus, structural embodiment refers to two aspects: whether the medical knowledge is appropriately represented in the data, and whether the structures correspond to the expert classification. We sought to obtain an unbiased bioinformatic view on lymphoma entities based on the previously published MLL9 dataset. This data set contains CLL, prolymphocytic leukemia (PL), monoclonal B-cell lymphocytosis (MBL), hairy cell lymphoma (HCL), mantle cell lymphoma (MCL), marginal zone lymphoma (MZL), lymphoplasmacytic lymphoma (LPL) and follicular lymphoma (FL). Evaluating structural embodiment using a framework based on swarm intelligence and self-organization presents a compelling methodology [ 36 , 39 ] (Fig. 1 ). The topographic map in Fig. 1 a depicts two main groups, two outlier groups and indicates individual highly atypical “outlier” samples. The clustering reveals a clear separation between presumed healthy individuals/normal controls (NC) and B-NHL (Fig. 1 b). Only 3 (out of 2998) NC were assigned into the B-NHL frame while a significant number of B-NHL samples appeared within the “NC cluster” (SI Table 1). The global perspective on structural embodiment differentiates NC from B-NHL so profoundly that lymphomas with subtle features, such as HCL and FL, remain unrecognizable. Interestingly, CLL samples appeared clearly detached from other lymphomas. Subsequently, the TM algorithm was applied in order to eliminate outlier samples and a second topographic map based on the entire set of lymphoma entities now revealed three main clusters and only a few remaining singular or doublet outliers (Fig. 1 c). Cluster matching with the given diagnoses is depicted in Fig. 1 d and reveals that AI driven immunophenotype B-NHL taxonomy does not entirely match with the lymphoma classes based on WHO definition. CLL, MBL and PL cluster together in one main category and FL, LPL and MZL constitute a second group. HCL appears as a clear distinct cluster surrounded by a chain of hills but this group also contains some MZL (SI Table 2). Interestingly, PL and MCL cluster closely in the same region, while the potentially smoldering small lymphocytic lymphomas MZL, LPL and FL appear far distant. Although lymphoma classification has been an issue of intense debate and controversies between specialists [ 50 ] different treatment strategies empirically emerged for CLL, HCL and common lymphocytic lymphomas such as MZL, LPL, FL and MCL hereby confirming that the AI based immunophenotypic clusters reflect true biologic similarities and divergences. Of note, even in a well-defined category such as CLL prediction of treatment requirement and response is highly variable and beyond B cell morphology is related to clinical parameters or T cell distribution patterns [ 51 ]. Whether the AI based taxonomy may reflect “treatment classes” or clinical outcome parameters remains to be addressed by further analyses of additional comprehensive lymphoma data sets. Taken together, our swarm based self-organization framework revealed corresponding categorization of the most common lymphomas in relation to WHO classes but MZL and LPL were not resolved clearly. Strategy for AI training on B-NHL data The observation that AI taxonomy does not entirely match to the predefined lymphoma diagnoses suggests compromised learning conditions for an automated lymphoma immunophenotyping approach. Based on the visualization of the self-organizing matrix and physician’s emphasis, the diagnostic tree was designed as a five-level stepwise diagnostic approach (Fig. 2 ). The TM algorithm realizes quality control at branch level zero (L0) as designating the quality of sample data is always part of diagnostic reports, and low data quality may impair AI learning. In level L1, NC samples are distinguished from those from B-NHL patients. Correct classification of NC samples is of particular importance because a false positive diagnosis of lymphoma “cancer” fundamentally affects treatment decisions. Once B-NHL has been categorized correctly in L1, the second diagnostic level (L2) distinguishes CLL (including PL and MBL, referred as “Mature B-cell neoplasms” in the 4th WHO Classification) and HCL from all other lymphomas (AOL). For L2 a good AI separation was achieved (Fig. 1 : cluster 1, 2 and 3). Moreover, a correct classification of these lymphoma entities is clinically prioritized: biopsy with histopathological confirmation is not needed to diagnose CLL [NCCN Guidelines Version 3.2023] and for HCL, immunophenotyping is an essential diagnostic requirement [NCCN Guidelines Version 1.2023]. To align our analysis with previous studies [ 11 , 52 ], we investigated the most common and distinct mature lymphoma types in the L3 AOL subcategory: mantle cell MCL, MZL, LPL and FL. Rare B-NHL subtypes or large cell lymphomas were classified into one category as L4 remaining B-NHL and not considered for further classification. Thus, the expectation for an AI system is to perform an initial quality check to identify (outlier) samples and to achieve diagnostic performance with appropriate stringency. The general performance of FlowXAI second module at the diagnostic levels The performance of FlowXAI was evaluated on the MLL9F dataset with N = 19,493 entire patient samples in comparison to the NN based approach [ 52 ]. The total number of patients in the malignant lymphoma categories (L2 and L3) ranged from 202 (HCL) to 4,274 (CLL). The B-NHL class sizes were highly imbalanced: B-NHL 45.3%; NCs 54.7%; CLL-like group (mature B-cell neoplasms) = {CLL, PL, MBL} 33.5%; HCL 1%; Group of all other lymphomas (AOL) = {MCL, MZL, LPL, FL, and rare B-NHLs} 11.8%; CLL 22.9%; MBL 7.8%; PL 2.8%; MCL 1.5%; MZL 5.3%; LPL 3.5%; FL 1.2%. The designation of classes and lymphoma identities are listed in SI Table 3. Omitting TM, cross-validation shows that FlowXAI is capable to differentiate NC (N = 2,133) from B-NHL (N = 1,766) with an average accuracy of 94.2 ± 0.4% in L1. FlowXAI classifies CLL-like samples, HCL samples, and AOL against NC samples with an average MCC of 85.4 ± 0.7% in L2. For the AOL categories, FlowXAI performed immunophenotyping, with an average MCC of 81.8 ± 0.7% in L3. The cross-validation results are illustrated in Fig. 3 a as a MD pots. Next, we incorporated the FlowXAI system’s self-assessment capability for each diagnosis: confident , probable , and challenging . To address the average performance of FlowXAI when these three degrees are considered, the MCC, accuracy (ACC), false positive rate (FPR), and false negative rate (FNR) were calculated for decisions on L3 (Fig. 3 b). FlowXAI categorizes 7–12% of test samples as challenging, 20–35% as probable, and 52–73% as confident (5th to 95th percentile). Interestingly, the distributions of the number of test samples per degree of trustworthiness appear bimodal (Fig. 3 c). A detailed look reveals 16 cross-validation cycles with less than 8% of the patients in the challenging and probable degree of trustworthiness and more corresponding classifications as confident . This heterogeneous pattern of the cross-validation results underscores the necessity of conducting at least 100 cross-validation trials because the results may appear superior (or inferior) by chance. The MCC of the test samples in the confident degree of trustworthiness ranged from 88–91% (median 90%). The performance of the probable degree stayed in the range of 70–80%, with a median of 77%, and the challenging degree had a large variance in performance, with a range between 47% and 68% (median 60%) shown in Fig. 3 d. The performance in the confident degree of trustworthiness follows a Gaussian distribution (magenta frame), while the other degrees appear less homogenous. This could be attributed to the high number of samples in the confident degree. Direct comparison reveals that the performance of the NN AI system is nearly equivalent to the average FlowXAI performance over all cross-validation iterations, which comprises all probable and confident degrees of trustworthiness (MCC = 83% versus 85.1%, Fig. 3 e). Of note, this performance was achieved for N = 3,488 (89%) samples, as indicated by the red vertical line. The deep learning AI was trained on the same dataset with one cross-validation test. Of note, Flow XAI performance in the confident degree was clearly superior to the NN AI based predictions. Finally, the FlowXAI performance was determined after TM's elimination of atypical samples. When atypical samples are discarded, the ACC and MCC improve slightly but not markedly for all trustworthiness degrees (Fig. 3 f). False negative rates predominantly declined in the challenging degree of trustworthiness. Contingency tables were computed for each cross-validation trial to reveal the detailed performance results for every malignant lymphoma entity (SI Table 4–7). In sum, by modeling the stepwise manual diagnostic approach, FlowXAI is capable of discriminating B-NHL patients with high accuracy and self-assessed competence, as needed for trustworthy diagnostic reports. Reduction of training data and extended validations of FlowXAI For a comprehensive lymphoma classification more categories of the 5th WHO classification such as B-lymphoblastic leukemia, large B cell lymphomas, Burkitt lymphoma and rare subtypes such as KSHV/HHV8-associated lymphomas or lymphoid proliferations and lymphomas associated with immune deficiency should be considered [ 48 ] but small class sizes limit AI learning abilities. Therefore, TM was constructed and combined with the AI system to FlowXAI as depicted in Fig. 4 a and explained in SI-Fig. 1. The TM function of outlier discrimination was validated with a set of 100 randomly drawn atypical samples by judgment of immunophenotyping experts. The manual analysis of the bivariate plots confirmed that 96 TM-selected data samples were indeed judged as altered or hardly evaluable by the human experts; i.e., they would not pass the quality control of L0. Only 4% of the atypical samples were considered usable for a trustworthy manual diagnostic workflow across all tubes as the base for a valid diagnostic report (SI Fig. 2 ). When both modules of the FlowXAI are employed, only 256 normal control samples, and 256 entire B-NHL samples equally distributed across all diagnoses are sufficient to learn the diagnostic levels. Despite the small training set, FlowXAI was able to classify 49% of the test samples as confident (probable : 30%, challenging: 21%). For the 7,684 test samples predicted by FlowXAI in the confident degree, an MCC of 87% was reached in L3 ( probable : 62%, challenging : 23%). The corresponding contingency tables present the details of the diagnostic performance (SI Table 8–10). When it is confident , FlowXAI demonstrates superior performance compared to the deep learning AI system in the automated diagnosis of six distinct lymphoma groups (MCC 87% vs. 83%, respectively), despite being trained on only 512 samples compared to 18,274 samples for the deep learning AI system. Furthermore, the performance of our TM-integrated Flow XAI algorithm to separate NC samples from B-NHL samples (L1) was assessed via a second independent benchmark algorithm, the cluster identification, characterization, and regression algorithm (CITRUS)[ 1 ]. Among the many interesting approaches for identifying cell frequencies or statistical sample properties that can be theoretically used in supervised approaches (e.g.,[ 2 , 53 – 57 ]), we chose CITRUS because it seemed to be most closely related to population description and lymphoma immunophenotyping. Moreover, it has been shown that ALPODS or CITRUS outperform other approaches [ 1 , 25 ]. The highest performance was achieved with FlowXAI and the preselection of typical training samples by TM improved the usage of CITRUS (Fig. 4 b and SI Methods). In addition to our cross-validation approach, which ensured true independence of training and test data, Flow XAI was challenged with a second impartial data set with 638 samples (PUM2) from the diagnostic center Philipps-University Marburg. TM considered 517 samples as typical and 121 B-NHL samples as atypical. For each trial, the training data consisted of 100 NC and 100 B-NHL randomly selected samples (30 CLL, 15 MCL, 10 FL, 10 MZL, 5 LPL, 10 HCL, 10 MBL, and 10 NS). 100 trials of cross-validation were performed on the remaining test sets of 317 typical samples. For L1, Flow XAI classified 57% of patient samples as confident , 31% as probable , and 12% as challenging . The accuracy of the performance measures was chosen due to the balanced group sizes. The accuracy was 99% when Flow XAI was confident , 89% when XAI estimated that the appropriate diagnosis was probable , and 69% when XAI judged a case as challenging . Even with this limited number of training data FlowXAI is capable of reporting confident results in the discrimination of B-NHL patients from healthy individuals with an accuracy of 99% in more than 50% of typical patients (Fig. 4 c/d). Atypical samples were classified as confident or probable in 80% with an accuracy of 98% and 90%, respectively (SI Fig. 6). Finally, for training purposes explainability is of paramount importance. FlowXAI selects and reports diagnostically relevant cell populations and expression patterns in a human-understandable manner as illustrated in SI Fig. 3 –5 and provides for an interactive evaluation of model-generated reports: https://plait.mathematik.uni-marburg.de/ . Discussion The FlowXAI system is designed to be a supportive tool for teaching the diagnostic process to identify B-NHL via multiparameter flow cytometry-based immunophenotyping. To ensure accurate classification, clinical evidence, including genetic and histopathological information, was carefully collected for both datasets, as previously published [ 6 , 23 ]. But lymphoma classification has historically been fluid. After the initial description of Hodgkin lymphoma in 1832 and Non-Hodgkin Lymphoma in 1925, various classifications coexisted until the National Lymphoma Study Group (ILSG) proposed the Revised European and American Lymphoma (R.E.A.L.) classification in 1994, which firstly considered immunopathology aspects. In 1997 it was replaced by an international classification of the World Health Organization (WHO), with the 5th edition published in 2022. Importantly, the term B prolymphocytic leukemia was eliminated in WHO-HAEMS5, and a PL subclass is now designated as the prolymphocytic progression of CLL [ 46 ]. For direct comparison with the consensus classification we gained an unbiased MFC based AI view of the structure embodiment of diagnoses in order to understand the AI data perspective and taxonomy and to prevent FlowXAI learning of spurious correlations, a problem that may occur with deep NN based learning [ 58 ]. We applied unsupervised machine learning by Databionic swarm [ 36 ] to identify high-dimensional structures in the data. Of note, we have previously shown that these unsupervised techniques correspond well with disease entities of divergent treatment decisions [ 36 , 40 , 41 ]. Correspondingly, the topographic map clearly shows separable normal control vs lymphoma classes and, subsequently, the distinct entities HCL and CLL-like, although some entities are not separable if the information is restricted to flow cytometry data. Whether the measurement of additional antigens leads to better resolution regarding delineation into WHO classes remains to be determined. Overall performance of our novel interactive Flow XAI was evaluated by computing 100 cross-validation trials with class-balanced 80/20 splits between training and test data [ 31 , 32 , 59 ] and was equivalent to a deep NN AI system. In the confident category, FlowXAI even surpassed the deep NN AI performance. Moreover, the head-to-head comparison of our FlowXAI model with the population describing algorithm CITRUS [ 60 ] in 100 cross-validations demonstrated superior L1 accuracy. With a second independent validation dataset, we demonstrated that the TM algorithm is efficient at calculating the distribution of strangeness values with only 100 samples. Thus, we show that immense training data reduction can be achieved using TM, and prototypes for each lymphoma category can be created with a few training samples. This unique feature paves the way for a broad applicability also for rare lymphoma subtypes. Beyond the aspects of training data reduction and structural embodiment, we focused on AI suitability for real-world teaching applications through self-explanation consisting of two features: the assessment of self-competence, which aligns with the manual diagnostic reporting of scores [ 61 ] and the ability to illustrate relevant cell populations, which allows for explanation and teaching in an interactive manner. Such direct interaction between a data-driven AI system and clinicians has been considered the best-performing approach [ 62 ]. Self-explanatory clinical decision support systems are essential for trust and reliance as discussed intensively from the informatics and social points of view [ 20 , 63 , 64 ]. AI influence in the clinical world is inevitable; hence, clinicians prepare themselves for this upcoming challenge [ 65 ]. To our knowledge, the design and deployment of a self-explaining symbolic AI system for lymphoma immunophenotyping has not yet been carried out. Presumably, further performance optimization may be achieved with consideration of AI based MFC lymphoma classification and by including suitable B-cell selection algorithms (see SI Literature evaluation for seeking benchmark algorithms). Notably, unselective data approaches may permit the discovery of novel putative target cell populations even beyond the well-known pathologic B cells following the basic principle of knowledge discovery. Taken together, FlowXAI diagnoses B-NHL from a small amount of non-standardized flow cytometry data with equal or superior performance levels compared with NN or other benchmark algorithms. The proposed AI system is self-explanatory, knowledge-based by design, and its self-assessment capabilities facilitate a transparent and interactive mode with human experts. It may therefore serve as a base for an AI-open-access lymphoma-immunophenotyping teaching platform. Abbreviations QC quality control, B-NHL = B-cell non-Hodgkin lymphoma NC normal control CLL-like CLL&MBL&PL HCL hairy cell leukemia, AOL = all other malignant lymphoma CLL chronic lymphocytic leukemia MBL monoclonal B-cell lymphocytosis PL prolymphocytic leukemia MCL mantle cell lymphoma MZL marginal zone lymphoma LPL lymphoplasmocytic lymphoma FL follicular lymphoma HCLv hairy cell leukemia variant DLBCL diffuse large B-cell lymphoma Others all remaining rare entities like Burkitt lymphoma (BL). Declarations Competing interests MCT owns the company IAP-GmbH Intelligent Analytics Projects, and IAP collaborates directly with Beckman Coulter. QS is an accepted PhD student at Philipps-University and works at IAP-GmbH. JH, CB, SK, AN, and AU declare no conflicts of interest. Author contributions JH, SK, CB, AU, and MCT: conceptualization. MCT and QS: data analysis, MCT investigation, and visualization. PK FCS data provision. MCT: writing the original manuscript draft. JH, SK, PK, QS, AN, AU, and CB: writing-review and editing. MCT and AU: supervision. Acknowledgments We thank Michael Schmidt for the analysis of impurity functions within the conventional decision tree software packages in R and MATLAB. Furthermore, we thank Max Zhao for communicating further elaborations about the overview of the sample files published in [52] and Monika Sikora for English corrections. Data availability The MLL9F dataset with N = 19,493 samples and the major part of the PUM2 data set with 638 samples have been published and are accessible [ 6 , 7 , 66 ]. The remaining files from the PUM2 dataset are accessible through https://plait.mathematik.uni-marburg.de/ . References Bruggner RV, Bodenmiller B, Dill DL, Tibshirani RJ, Nolan GP. Automated identification of stratifying signatures in cellular subpopulations. Proc Natl Acad Sci U S A. 2014;111(26):E2770–E7. O’Neill K, Jalali A, Aghaeepour N, Hoos H, Brinkman RR. Enhanced flowType/RchyOptimyx: a bioconductor pipeline for discovery in high-dimensional cytometry data. Bioinformatics. 2014;30(9):1329–30. Li J-L, Lin Y-C, Wang Y-F, Monaghan SA, Ko B-S, Lee C-C. A Chunking-for-Pooling Strategy for Cytometric Representation Learning for Automatic Hematologic Malignancy Classification. IEEE Journal of Biomedical and Health Informatics. 2022;26(9):4773–84. Kowarsch F, Weijler L, Wödlinger M, Reiter M, Maurer-Granofszky M, Schumich A, et al. Towards self-explainable transformers for cell classification in flow cytometry data. International Workshop on Interpretability of Machine Intelligence in Medical Image Computing: Springer; 2022. p. 22–32. Lewis JE, Cooper LA, Jaye DL, Pozdnyakova O. Automated deep learning-based diagnosis and molecular characterization of acute myeloid leukemia using flow cytometry. Modern Pathology. 2024;37(1):100373. Zhao M, Mallesh N, Höllein A, Schabath R, Haferlach C, Haferlach T, et al. Hematologist-Level Classification of Mature B‐Cell Neoplasm Using Deep Learning on Multiparameter Flow Cytometry Data. Cytometry Part A. 2020;97:1073–80. doi: 10.1002/cyto.a.24159 . Hoffmann J, Rother M, Kaiser U, Thrun MC, Wilhelm C, Gruen A, et al. Determination of CD43 and CD200 surface expression improves accuracy of B-cell lymphoma immunophenotyping. Cytometry B Clin Cytom. 2020;98(6):476–82. doi: 10.1002/cyto.b.21936 . Van Dongen J, Lhermitte L, Böttcher S, Almeida J, Van Der Velden V, Flores-Montero J, et al. EuroFlow antibody panels for standardized n-dimensional flow cytometric immunophenotyping of normal, reactive and malignant leukocytes. Leukemia. 2012;26(9):1908–75. Rawstron AC, Kreuzer KA, Soosapilla A, Spacek M, Stehlikova O, Gambell P, et al. Reproducible diagnosis of chronic lymphocytic leukemia by flow cytometry: an European Research Initiative on CLL (ERIC) & European Society for Clinical Cell Analysis (ESCCA) Harmonisation project. Cytometry Part B: Clinical Cytometry. 2018;94(1):121–8. Costa E, Pedreira CE, Barrena S, Lecrevisse Q, Flores J, Quijano S, et al. Automated pattern-guided principal component analysis vs expert-based immunophenotypic classification of B-cell chronic lymphoproliferative disorders: a step forward in the standardization of clinical immunophenotyping. Leukemia. 2010;24(11):1927–33. Mallesh N, Zhao M, Meintker L, Höllein A, Elsner F, Lüling H, et al. Knowledge transfer to enhance the performance of deep learning models for automated classification of B cell neoplasms. Patterns. 2021;2(10):100351. Cabitza F, Campagner A, Malgieri G, Natali C, Schneeberger D, Stoeger K, et al. Quod erat demonstrandum? Towards a typology of the concept of explanation for the design of explainable AI. Expert Syst Appl. 2023;213:118888. Thrun MC. Exploiting Distance-Based Structures in Data Using an Explainable AI for Stock Picking. Information. 2022;13(2):51. doi: 10.3390/info13020051 . Thrun MC, Ultsch A, Breuer L. Explainable AI framework for multivariate hydrochemical time series. Mach Learn Knowl Extr. 2021;3(1):170–205. doi: 10.3390/make3010009 . Thrun MC. Identification of explainable structures in data with a human-in-the-loop. Ger J Artif Intell. 2022;36:297–301. Goebel R, Chander A, Holzinger K, Lecue F, Akata Z, Stumpf S, et al. Explainable AI: the new 42? In: Holzinger A, Kieseberg P, Tjoa A, Weippl E, editors. Machine Learning and Knowledge Extraction: Second IFIP TC 5, TC 8/WG 84, 89, TC 12/WG 129 International Cross-Domain Conference, CD-MAKE 2018. Cham: Springer; 2018. p. 295–303. Lötsch J, Kringel D, Ultsch A. Explainable artificial intelligence (XAI) in biomedicine: Making AI decisions trustworthy for physicians and patients. BioMedInformatics. 2021;2(1):1–17. Holzinger A. The next frontier: AI we can really trust. Machine Learning and Principles and Practice of Knowledge Discovery in Databases: International Workshops of ECML PKDD 2021, Virtual Event, September 13–17, 2021, Proceedings, Part I: Springer; 2022. p. 427 – 40. Stöger K, Schneeberger D, Holzinger A. Medical artificial intelligence: the European legal perspective. Commun ACM. 2021;64(11):34–6. Holzinger A, Biemann C, Pattichis CS, Kell DB. What do we need to build explainable AI systems for the medical domain? arXiv preprint arXiv:171209923. 2017. Matutes E, Owusu-Ankomah K, Morilla R, Houlihan A, Que T, Catovsky D. The immunological profile of B-cell disorders and proposal of a scoring system for the diagnosis of CLL. Leukemia. 1994;8(10):1640–5. Moreau EJ, Matutes E, A’Hern RP, Morilla AM, Morilla RM, Owusu-Ankomah KA, et al. Improvement of the chronic lymphocytic leukemia scoring system with the monoclonal antibody SN8 (CD79b). American journal of clinical pathology. 1997;108(4):378–82. Hoffmann J, Rother M, Kaiser U, Thrun MC, Wilhelm C, Gruen A, et al. Determination of CD43 and CD200 surface expression improves accuracy of B-cell lymphoma immunophenotyping. Cytometry Part B: Clinical Cytometry. 2020;98(6):476–82. doi: 10.1002/cyto.b.21936 . Thrun MC, Hoffman J, Röhnert M, Von Bonin M, Oelschlägel U, Brendel C, et al. Flow Cytometry datasets consisting of peripheral blood and bone marrow samples for the evaluation of explainable artificial intelligence methods. Data In Brief. 2022;43:108382. doi: 10.1016/j.dib.2022.108382 . Ultsch A, Hoffman J, Röhnert M, Von Bonin M, Oelschlägel U, Brendel C, et al. An Explainable AI System for the Diagnosis of High Dimensional Biomedical Data. BioMedInformatics. 2024;4:197–218. doi: 10.3390/biomedinformatics4010013 . Simpson EH. Measurement of Diversity. Nature. 1949;163(4148):688-. doi: 10.1038/163688a0 . Luan S, Schooler LJ, Gigerenzer G. A signal-detection analysis of fast-and-frugal trees. Psychological review. 2011;118(2):316. Cohen J. Statistical power analysis for the behavioral sciences. New York: Academic Press; 2013. Ultsch A, Lötsch J. Computed ABC Analysis for Rational Selection of Most Informative Variables in Multivariate Data. PloS one. 2015;10(6):e0129767. doi: 10.1371/journal.pone.0129767 . Penner M, Senman L, Andoni L, Dupuis A, Anagnostou E, Kao S, et al. Concordance of Diagnosis of Autism Spectrum Disorder Made by Pediatricians vs a Multidisciplinary Specialist Team. JAMA Network Open. 2023;6(1):e2252879-e. Bishop CM. Pattern Recognition and Machine Learning. New York, NY: Springer; 2006. Goodfellow I, Bengio Y, Courville A. Deep learning. Cambridge, Massachusetts, USA: MIT press; 2016. Boughorbel S, Jarray F, El-Anbari M. Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric. PloS one. 2017;12(6):e0177678. Chicco D, Jurman G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC genomics. 2020;21(1):1–13. Thrun MC, Gehlert T, Ultsch A. Analyzing the Fine Structure of Distributions. PLoS ONE. 2020;15(10):e0238835. doi: 10.1371/journal.pone.0238835 Thrun MC, Ultsch A. Swarm Intelligence for Self-Organized Clustering. Artificial Intelligence. 2021;290:103237. doi: 10.1016/j.artint.2020.103237 . Thrun MC, Pape F, Ultsch A. Conventional Displays of Structures in Data Compared With Interactive Projection-Based Clustering (IPBC). International Journal of Data Science and Analytics. 2021;12(3):249–71. doi: 10.1007/s41060-021-00264-2 . Thrun MC, Ultsch A. Using Projection based Clustering to Find Distance and Density based Clusters in High-Dimensional Data. Journal of Classification. 2020;38(2):280–312. doi: 10.1007/s00357-020-09373-2 . Thrun MC, Ultsch A. Uncovering High-Dimensional Structures of Projections from Dimensionality Reduction Methods. MethodsX. 2020;7:101093. doi: 10.1016/j.mex.2020.101093 . Thrun MC. Distance-Based Clustering Challenges for Unbiased Benchmarking Studies. Nature Scientific Reports. 2021;11(1):18988. doi: 10.1038/s41598-021-98126-1 . Ultsch A, Lötsch J. Machine-learned cluster identification in high-dimensional data. Journal of Biomedical Informatics. 2017;66(C):95–104. Ultsch A. Maps for the visualization of high-dimensional data spaces. Workshop on Self organizing Maps (WSOM). Kyushu, Japan2003. p. 225 – 30. Thrun MC, Lerch F, Lötsch J, Ultsch A. Visualization and 3D Printing of Multivariate Data of Biomarkers. In: Skala V, editor. International Conference in Central Europe on Computer Graphics, Visualization and Computer Vision (WSCG). Plzen2016. p. 7–16. Stier Q, Thrun MC. An efficient multicore CPU implementation of the DatabionicSwarm. International Federation of Classification Societies (IFCS). San José, Costa Rica2024. p. 1–8. Thrun MC, Mack E, Neubauer A, Haferlach T, Frech M, Ultsch A, et al. A Bioinformatics View on Acute Myeloid Leukemia Surface Molecules by Combined Bayesian and ABC Analysis. Bioengineering. 2022;9(11):642. doi: bioengineering9110642. Alaggio R, Amador C, Anagnostopoulos I, Attygalle AD, De Oliveira Araujo IB, Berti E, et al. The 5th edition of the World Health Organization classification of haematolymphoid tumours: lymphoid neoplasms. Leukemia. 2022;36(7):1720–48. Swerdlow SH, Campo E, Pileri SA, Harris NL, Stein H, Siebert R, et al. The 2016 revision of the World Health Organization classification of lymphoid neoplasms. Blood. 2016;127(20):2375–90. Alaggio R, Amador C, Anagnostopoulos I, Attygalle AD, Araujo IBdO, Berti E, et al. The 5th edition of the World Health Organization classification of haematolymphoid tumours: lymphoid neoplasms. Leukemia. 2022;36(7):1720–48. Swerdlow SH, Campo E, Pileri SA, Harris NL, Stein H, Siebert R, et al. The 2016 revision of the World Health Organization classification of lymphoid neoplasms. Blood, The Journal of the American Society of Hematology. 2016;127(20):2375–90. Hiddemann W, Stein H. Die neue WHO-Klassifikation der malignen Lymphome: Endlich eine weltweit akzeptierte Einteilung. Deutsches Ärzteblatt. 1999;96. Hoffmann J, Eminovic S, Wilhelm C, Krause SW, Neubauer A, Thrun MC, et al. Prediction of clinical outcomes with explainable artificial intelligence in patients with chronic lymphocytic leukemia Current Oncology. 2023;30(2):1903–15. doi: 10.3390/curroncol30020148 . Zhao M, Mallesh N, Höllein A, Schabath R, Haferlach C, Haferlach T, et al. Hematologist-level classification of mature B‐cell neoplasm using deep learning on multiparameter flow cytometry data. Cytometry A. 2020;97:1073–80. doi: 10.1002/cyto.a.24159 . Greene E, Finak G, D'Amico LA, Bhardwaj N, Church CD, Morishima C, et al. New interpretable machine-learning method for single-cell data reveals correlates of clinical response to cancer immunotherapy. Patterns. 2021;2(12):100372. Aghaeepour N, Jalali A, O'Neill K, Chattopadhyay PK, Roederer M, Hoos HH, et al. RchyOptimyx: cellular hierarchy optimization for flow cytometry. Cytometry A. 2012;81(12):1022–30. Aghaeepour N, Finak G, FlowCAP Consortium, DREAM Consortium, Hoos H, Mosmann TR, et al. Critical assessment of automated flow cytometry data analysis techniques. Nat Methods. 2013;10(3):228–38. Aghaeepour N, Chattopadhyay P, Chikina M, Dhaene T, Van Gassen S, Kursa M, et al. A benchmark for evaluation of algorithms for identification of cellular correlates of clinical outcomes. Cytometry A. 2016;89(1):16–21. Van Gassen S, Callebaut B, Van Helden MJ, Lambrecht BN, Demeester P, Dhaene T, et al. FlowSOM: using self-organizing maps for visualization and interpretation of cytometry data. Cytometry A. 2015;87(7):636–45. Lapuschkin S, Wäldchen S, Binder A, Montavon G, Samek W, Müller KR. Unmasking Clever Hans predictors and assessing what machines really learn. Nat Commun. 2019;10(1):1096. Duda RO, Hart PE, Stork DG. Pattern Classification. Second Edition ed. A Wiley-Interscience Publication. New York, USA: John Wiley & Sons; 2001. Bruggner RV, Bodenmiller B, Dill DL, Tibshirani RJ, Nolan GP. Automated identification of stratifying signatures in cellular subpopulations. Proceedings of the National Academy of Sciences. 2014;111(26):E2770-E7. Moreau EJ, Matutes E, A’Hern RP, Morilla AM, Morilla RM, Owusu-Ankomah KA, et al. Improvement of the chronic lymphocytic leukemia scoring system with the monoclonal antibody SN8 (CD79b). Am J Clin Pathol. 1997;108(4):378–82. Jacobs M, Pradier MF, McCoy TH, Perlis RH, Doshi-Velez F, Gajos KZ. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Transl Psychiatry. 2021;11(1):108. Bussone A, Stumpf S, O'Sullivan D. The role of explanations on trust and reliance in clinical decision support systems. 2015 International Conference on Healthcare Informatics. Dallas, TX, USA: IEEE; 2015. p. 160–9. Miller T. Explanation in artificial intelligence: insights from the social sciences. Artif Intell. 2019;267:1–38. James CA, Wachter RM, Woolliscroft JO. Preparing clinicians for a clinical world influenced by artificial intelligence. JAMA. 2022;327(14):1333–4. Hoffmann J, Eminovic S, Wilhelm C, Krause SW, Neubauer A, Thrun MC, et al. Prediction of clinical outcomes with explainable artificial intelligence in patients with chronic lymphocytic leukemia. Curr Oncol. 2023;30(2):1903–15. doi: 10.3390/curroncol30020148 . Additional Declarations Yes there is potential conflict of interest. Supplementary Files 2025ThruneaPLAITSIadV4.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6690890","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":458819288,"identity":"549db2fb-3df7-4b6a-ac92-cbd4dc75e2a1","order_by":0,"name":"Michael Thrun","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABQElEQVRIie2RMUvDQBTHXwhkujgnJKSfQEg4SIaq/So5Au2STF0chAYCmayugaj5CoqLo3JwLkFXIQ6li5OgVCSgVJPUgDQV6SaS33D8edyP93gPoKXlT4LKZ4v4iwBIAqSDDRRAqgp8wxCqn/2hvwh6rTz8ptBRrUClAHz8qGwmYzYTL269w2R8Nb3be1dlH5mTCWQdK943n1+gqy0pJttwYjHNvKgI2GU6UgBZug2ZcXSf4liFAW4oCPNimHk+Q6biCjrSAJkSmWdcJLmYl4AWm1ml3JCEIevNnX8pxcZ6tTJarVyS06IL74XVYJVCSoV7Amo3Bivqx6FDzoqgeAcYyYHQLxUnUtmQB31gLHehKYbHcIec0NSYua9aT7oOmJwD3Y6U4JzLd7udxmEafD8EX91oLbh8XaOlpaXlH/IJTuRrJmwGRpAAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-9542-5543","institution":"Univ. Marburg","correspondingAuthor":true,"prefix":"","firstName":"Michael","middleName":"","lastName":"Thrun","suffix":""},{"id":458819289,"identity":"b37e7256-8d9f-483b-b94d-7a113b81e4ed","order_by":1,"name":"Joerg Hoffmann","email":"","orcid":"","institution":"Universität Marburg","correspondingAuthor":false,"prefix":"","firstName":"Joerg","middleName":"","lastName":"Hoffmann","suffix":""},{"id":458819290,"identity":"c0456e73-0bde-4b92-ad6d-96fad1a790fc","order_by":2,"name":"Stefan Krause","email":"","orcid":"https://orcid.org/0000-0002-5259-4651","institution":"Universitätsklinikum Erlangen","correspondingAuthor":false,"prefix":"","firstName":"Stefan","middleName":"","lastName":"Krause","suffix":""},{"id":458819291,"identity":"50a3084d-3c47-4a59-9742-bed9fccef213","order_by":3,"name":"Peter Krawitz","email":"","orcid":"","institution":"University Bonn","correspondingAuthor":false,"prefix":"","firstName":"Peter","middleName":"","lastName":"Krawitz","suffix":""},{"id":458819292,"identity":"01dde8c2-edab-459a-a73f-3e71039d0c28","order_by":4,"name":"Quirin Stier","email":"","orcid":"","institution":"Philipps University Marburg","correspondingAuthor":false,"prefix":"","firstName":"Quirin","middleName":"","lastName":"Stier","suffix":""},{"id":458819293,"identity":"8a46655a-f33f-45db-a91f-6fd39a054492","order_by":5,"name":"Andreas NEUBAUER","email":"","orcid":"https://orcid.org/0000-0002-7606-8760","institution":"Philipps University Marburg","correspondingAuthor":false,"prefix":"","firstName":"Andreas","middleName":"","lastName":"NEUBAUER","suffix":""},{"id":458819294,"identity":"75d42bcb-66df-49e4-9101-31e54dcb4a8d","order_by":6,"name":"Cornelia Brendel","email":"","orcid":"https://orcid.org/0000-0003-1612-9557","institution":"Philipps University Marburg","correspondingAuthor":false,"prefix":"","firstName":"Cornelia","middleName":"","lastName":"Brendel","suffix":""},{"id":458819295,"identity":"2bcaab98-e307-4536-8deb-7ba2c4cafcff","order_by":7,"name":"Alfred Ultsch","email":"","orcid":"","institution":"Philipps-University Marburg","correspondingAuthor":false,"prefix":"","firstName":"Alfred","middleName":"","lastName":"Ultsch","suffix":""}],"badges":[],"createdAt":"2025-05-18 09:35:06","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6690890/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6690890/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":83280476,"identity":"ac8c295d-ef7a-47b6-a1ae-9b6e01fbfb59","added_by":"auto","created_at":"2025-05-22 10:10:28","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":578014,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTopographic map of lymphoma based on the Databionic swarm and the generalized U-matrix technique.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe topographic map allows for the visualization of high-dimensional distances and densities between Databionic-swarm-projected points, identifying if the data exhibits structures that can be used in cluster analysis. Each point consists of the three sample files of a MLL9F sample. The colors of the projected points indicate the lymphoma identities. A classification by lymphoma identities was not utilized to generate the projection or visualization. Within the topographic map, valleys, and basins are depicted as clusters, while the watersheds of hills and mountains serve as cluster boundaries by the following color scheme. The blue colors represent lower elevations (e.g., sea level), green and brown for intermediate elevations (e.g., low hills), and various shades of white for higher elevations (e.g., snow-covered mountains). Hypsometric tints use distinct surface colors to represent elevation ranges, with contour lines integrated into a specific color scheme. The elevation ranges are mapped to high-dimensional data distances and densities of the projected points. Additionally, the visual borders in the topographic map exhibit cyclic connections with periodicity.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA\u003c/strong\u003e The MLL9 data set with a sample of 9000 samples (3000 NC, 3000 CLL-like, 3000 non CLL-like B-NHL) was employed for automated self-organization and revealed two main clusters\u003cstrong\u003e, B\u003c/strong\u003e a clear distinction between healthy and lymphoma samples\u003cstrong\u003e.\u003c/strong\u003e Without TM, many single outliers scatter across the matrix, illustrated as single dots in single valleys. \u003cstrong\u003eC\u003c/strong\u003e When TM is applied, and healthy samples are disregarded, the Databionic swarm gives three main clusters\u003cstrong\u003e, D \u003c/strong\u003ereveals they are corresponding to HCL, CLL-like lymphomas, and AOL\u003cstrong\u003e.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure1Leuk.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6690890/v1/856a2c6a6891a8091058f01b.jpg"},{"id":83280472,"identity":"456d8d78-6f48-4988-bb5b-8a194ef099e0","added_by":"auto","created_at":"2025-05-22 10:10:27","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":84648,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDiagnostic decision tree\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe diagnostic tree summarizes the four distinct levels (L0 to L3) with different performance demands. Further lymphoma categories were assigned to level four (L4) but disregarded for further evaluation; these categories are marked in gray.\u003c/p\u003e\n\u003cp\u003eAbbreviations: QC = quality control, B-NHL = B-cell non-Hodgkin lymphoma; NC = normal control; CLL-like = CLL\u0026amp;MBL\u0026amp;PL; HCL = hairy cell leukemia, AOL= all other malignant lymphoma; CLL = chronic lymphocytic leukemia; MBL = monoclonal B-cell lymphocytosis; PL = prolymphocytic leukemia; MCL = mantle cell lymphoma; MZL = marginal zone lymphoma; LPL = lymphoplasmocytic lymphoma; FL = follicular lymphoma; HCLv = hairy cell leukemia variant; DLBCL = diffuse large B-cell lymphoma; Others = all remaining rare entities like Burkitt lymphoma (BL).\u003c/p\u003e","description":"","filename":"Figure2Leuk.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6690890/v1/38fa3c617f5c6919c69c283f.jpg"},{"id":83280473,"identity":"e6fc0c7b-1617-4206-94cf-31f963a9f71d","added_by":"auto","created_at":"2025-05-22 10:10:27","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":26862,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAI cross-validation for different diagnostic levels and different degrees of trustworthiness.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA \u003c/strong\u003eThe performance of FlowXAI was measured in 100 cross-validation trials as the ACC in percent for level one (L1) and as the MCC in percent for L2 and L3. The MD-plot depicts 20% of the test samples (N=3,899 samples). The average performance for L1 is 94.0 ± 0.4% (accuracy), for L2 85.0 ± 0.7% (MCC), and for L3 82.0 ± 0.7% (MCC). The estimated PDFs of accuracy and MCC values are approximately Gaussian (magenta outline).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB\u003c/strong\u003eAverage performance of FlowXAI per degree of trustworthiness in diagnosing L3 samples. The degrees are confident, probable, and challenging. A contingency table was computed for each cross-validation trial, and all contingency tables were summed for each entry (provided in supplementary). All entities other than the NC samples were aggregated to B-NHL to calculate the quality measures ACC, FNR, and FPR.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC\u003c/strong\u003eMD-plot of the numbers and distributions of patients in the three degrees: \u003cem\u003econfident\u003c/em\u003e (gold), \u003cem\u003eprobable\u003c/em\u003e (green), and \u003cem\u003echallenging\u003c/em\u003e(blue). The estimated PDFs exhibit bimodality.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD\u003c/strong\u003eMD-plot of MCC for all three degrees of trustworthiness. The estimated PDF of MCC values for the confident degree is approximately Gaussian (magenta outline), while the probable degree is slightly skewed and has a large variance for the challenging degree.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eE\u003c/strong\u003eNumber of test samples and performance of FlowXAI on L3 samples. The y-axis depicts the MCC value in percent per trial of cross-validation, and the x-axis represents the percentage of test data samples (with N=3,899 samples being 100%) that are within this degree of trustworthiness. Each set of three points represents one cross-validation trial and adds up to 100%. The performance of the deep learning AI system (magenta) was 83% (MCC) for 12% (N=2,348) of the test samples.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eF \u003c/strong\u003eAverage performance of FlowXAI per degree of trustworthiness on L3 samples, Atypical samples were discarded by tile mining (TM).\u003c/p\u003e\n\u003cp\u003eAbbreviations: PDF: Probability density function, ACC = accuracy, MCC = Mathews correlation coefficient, FPR = false positive rate, FNR = false negative rate, Conf = confidence, Prob = probable, Chall = challenging\u003c/p\u003e","description":"","filename":"OnlineFigure3Leuk.png","url":"https://assets-eu.researchsquare.com/files/rs-6690890/v1/be5010c57c4efdfeef197d59.png"},{"id":83280474,"identity":"2f9a8d7f-6887-4319-958e-9bf874489bbb","added_by":"auto","created_at":"2025-05-22 10:10:28","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":90079,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow and validation of Flow XAI in combination with TM\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA \u003c/strong\u003eProcess flow for selecting typical samples by TM and learning ALPODS populations based on the events of patient samples. Human-in-the-loop (HIL) staining is needed for an atypical sample to determine if the sample is usable. The detailed construction of FlowXAI using a committee of ALPODS experts is illustrated in SI Fig. 1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB\u003c/strong\u003e MD-plot of Accuracies for the calculation of NC versus B-NHL (L1) for the Flow XAI and CITRUS algorithms with and without the use of TM. The calculation was restricted to tube 1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC\u003c/strong\u003e Evaluation of trustworthiness and performance in L1 by Flow XAI on an independent dataset in 100 cross-validation trials (PUM2)\u003cstrong\u003e. \u003c/strong\u003eMD-plot\u003cstrong\u003e \u003c/strong\u003efor 100 B-NHL patients and 100 NC training samples (L1), Flow XAI classified 57% of N=117 test samples as \u003cem\u003econfident\u003c/em\u003e, 31% as \u003cem\u003eprobable,\u003c/em\u003eand 12% as \u003cem\u003echallenging. \u003c/em\u003eThe magenta frame indicates a Gaussian distribution.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD\u003c/strong\u003e MD-plot of MCC values in percent for the three different levels of trustworthiness.\u003c/p\u003e\n\u003cp\u003eAbbreviations: Conf = confident; Prob = probable; Chall = challenging\u003c/p\u003e","description":"","filename":"Figure4Leuk.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6690890/v1/f834249662d859ad415f6f38.jpg"},{"id":85873870,"identity":"beb93fed-ec18-48ff-a8c7-9bce72c237f4","added_by":"auto","created_at":"2025-07-02 14:47:14","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1662178,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6690890/v1/c35ff5ba-279c-4817-a8bc-8713157fcc48.pdf"},{"id":83280882,"identity":"0e019929-23cd-4f7a-83eb-9ba0e3349b80","added_by":"auto","created_at":"2025-05-22 10:18:28","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":2245898,"visible":true,"origin":"","legend":"","description":"","filename":"2025ThruneaPLAITSIadV4.docx","url":"https://assets-eu.researchsquare.com/files/rs-6690890/v1/60c00475b256b733fe3eee24.docx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential conflict of interest.","formattedTitle":"Self-explaining Artificial Intelligence for the Classification of B cell Non-Hodgkin Lymphoma","fulltext":[{"header":"Introduction","content":"\u003cp\u003eArtificial intelligence (AI) has delivered tremendous changes and opportunities to the field of clinical diagnostic; particularly neuronal networks (NN) enabled a new era of image analysis techniques. Multiparameter flow cytometry (MFC) has been used for digital surface protein analysis of cell populations from peripheral blood for about four decades. Moreover, novel high parametric technologies such as mass cytometry and spectral flow cytometry provide numerous simultaneous measurements of variables. Thus, algorithms that facilitate the selection of relevant cell populations in machine learning-assisted analyses have been proposed [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Several AI algorithms were developed for automated diagnosis of leukemia and lymphoma [\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], and a NN based approach for automated immunophenotyping of mature B-cell non-Hodgkin lymphoma (B-NHL) has recently been proposed [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Due to the limited availability of well-trained expert physicians automation of immunophenotyping is highly desirable for support and education of hematologists to interpret antigen expression patterns of different cell populations correctly. However, there are several obstacles related to this task. First, the definition of lymphoma classes has historically been fluid and therefore, taxation whether the given lymphoma identities are represented in the data is important. Moreover, recommendations for the optimal diagnostic panel of B-cell epitopes are still under investigation [\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Rare lymphoma entities or the reduced availability of data samples in general impair performances of NN based algorithms hereby limiting such a diagnostic process to university hospitals or high throughput diagnostic centers. Transfer learning was proposed for automated B-NHL immunophenotyping to address the problem of large training dataset requirement and it had been claimed that the proposed AI systems deliver a performance comparable to that of human experts but functioning was strongly dependent on the amount of learning data and rare lymphoma entities were excluded [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Beyond performance and data availability, self-explanation is an important aspect of diagnostic AI. Many AI systems, particularly if based on NN, are \u0026ldquo;subsymbolic\u0026rdquo; [\u003cspan additionalcitationids=\"CR13\" citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Subsymbolic systems assign a diagnosis to a sample but are unable to provide reasons or explanations for their decisions [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], although they are skillful learners [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Therefore, with regard to patient safety, trustworthiness is an issue of current debate [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] and European Union laws require causal justification for decisions made by AI [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Finally, due to biological variances individual samples may differ from the common lymphoma class immunophenotype. To address this issue Matutes et al. had proposed a scoring system for chronic lymphocytic leukemia (CLL) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. With regard to these considerations, a self-explaining AI system FlowXAI was designed to address in particular the following three key elements:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eIntegrating medical knowledge after the assessment of the structure embodiment in given data.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eObtaining a sound performance, even if only very few samples are available in the learning phase.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eProviding self-explanation by reporting a self-competence estimation per sample and delivering of human-understandable explanations for decisions.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eCharacterization of the data sets\u003c/h2\u003e \u003cp\u003eThe MLL9 and PUM2 datasets are entirely independent, both medically and technically. They originate from two separate diagnostic laboratories and were generated using different antibody panels. The MLL9 data set contains 18,274 samples from which 19,493 peripheral blood samples were selected for this work [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The panel of the Munich Leukemia Laboratory (MLL) consists of three 9-color tubes with the following antibodies: 1. CD19(APCA750)\u0026thinsp;+\u0026thinsp;CD45(KrOr)\u0026thinsp;+\u0026thinsp;FMC7(FITC)\u0026thinsp;+\u0026thinsp;CD10(PE)\u0026thinsp;+\u0026thinsp;IgM(ECD)\u0026thinsp;+\u0026thinsp;CD79b(PC5.5)\u0026thinsp;+\u0026thinsp;CD20(PC7)\u0026thinsp;+\u0026thinsp;CD23(APC)\u0026thinsp;+\u0026thinsp;CD5(PacBlue); 2. CD19(APCA750)\u0026thinsp;+\u0026thinsp;CD45(KrOr)\u0026thinsp;+\u0026thinsp;kappa(FITC)\u0026thinsp;+\u0026thinsp;lambda(PE)\u0026thinsp;+\u0026thinsp;CD38(ECD)\u0026thinsp;+\u0026thinsp;CD25(PC5.5)\u0026thinsp;+\u0026thinsp;CD11c(PC7)\u0026thinsp;+\u0026thinsp;CD103(APC)\u0026thinsp;+\u0026thinsp;CD22(Pac Blue); 3. CD19(APCA750)\u0026thinsp;+\u0026thinsp;CD45(KrOr)\u0026thinsp;+\u0026thinsp;CD8(FITC)\u0026thinsp;+\u0026thinsp;CD4(PE)\u0026thinsp;+\u0026thinsp;CD3(ECD)\u0026thinsp;+\u0026thinsp;CD56(APC)\u0026thinsp;+\u0026thinsp;HLA-DR(PacBlue). The PUM2 data set has uni-centrically been collected in Marburg [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] and employs two tubes: 1. CD45(KrOr)\u0026thinsp;+\u0026thinsp;kappa/CD8(FITC)\u0026thinsp;+\u0026thinsp;lambda/CD7(PE)\u0026thinsp;+\u0026thinsp;CD23(ECD)\u0026thinsp;+\u0026thinsp;CD79b/CD4(PC5.5)\u0026thinsp;+\u0026thinsp;CD5(PC7)\u0026thinsp;+\u0026thinsp;CD38(APC)\u0026thinsp;+\u0026thinsp;CD19(AP-A700)\u0026thinsp;+\u0026thinsp;CD20/CD3(APC-A750)\u0026thinsp;+\u0026thinsp;FMC7/CD2(PacBlue); 2. CD19(KrOr)\u0026thinsp;+\u0026thinsp;CD103(FITC)\u0026thinsp;+\u0026thinsp;CD43(PE)\u0026thinsp;+\u0026thinsp;CD25(ECD)\u0026thinsp;+\u0026thinsp;CD10(PC5.5)\u0026thinsp;+\u0026thinsp;CD200(PC7)\u0026thinsp;+\u0026thinsp;CD11c(AP-A700)\u0026thinsp;+\u0026thinsp;CD20(APC-A750)\u0026thinsp;+\u0026thinsp;IgM(Pac Blue). The flow cytometer measurements consist of N\u0026thinsp;=\u0026thinsp;50,000 (MLL9) and N\u0026thinsp;=\u0026thinsp;100,000 (PUM2) events for each tube.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eThe Flow XAI system\u003c/h3\u003e\n\u003cp\u003eTo ensure robust diagnostic accuracy while simultaneously addressing sample quality and computational efficiency, we developed FlowXAI \u0026mdash; an AI system for automated, self-explaining sample classification designed to assist in the classification of B-cell non-Hodgkin lymphomas (B-NHL) using flow cytometry data. FlowXAI integrates multiple stages of computational analysis to mirror the diagnostic reasoning of human experts, incorporating structural analysis of underlying data consistency with clinical knowledge.\u003c/p\u003e\n\u003ch3\u003eTraining data reduction by tile mining\u003c/h3\u003e\n\u003cp\u003eThe first stage of the FlowXAI system, called Tile Mining (TM), aims to evaluate unsupervised the quality of the patient samples. TM identifies relevant structures in data to select representative samples. This reduces the requirement for large training data sets to these representative samples. TM systematically divides each bivariate marker combination into a fixed grid of uniform rectangular regions called hypercubes. All unique bivariate combinations of dd markers within one tube amount to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(\\genfrac{}{}{0pt}{}{d}{2}\\right)=\\frac{d(d-1)}{2}\\)\u003c/span\u003e\u003c/span\u003e pairs. TM partitions each of these bivariate plots into 9 equally sized hypercubes, yielding a total of #\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:H\\left(d\\right)=9*\\frac{d(d-1)}{2}\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eDistinct hypercubes across the panel. For d\u0026thinsp;=\u0026thinsp;11 antigens, this results in \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\#H\\left(d=11\\right)=495\\)\u003c/span\u003e\u003c/span\u003e hypercubes. For each hypercube \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:h\\in\\:H\\left(d\\right)\\)\u003c/span\u003e\u003c/span\u003e, the density \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:p\\left(h\\right)\\)\u003c/span\u003e\u003c/span\u003e is defined as the proportion of associated events:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:p\\left(h\\right)=\\sum\\:_{e\\:\\:}^{}{f}_{e},{f}_{e}=\\left\\{\\begin{array}{cc}1\u0026amp;\\:e\\:ϵ\\:h\\\\\\:0\u0026amp;\\:e\\notin\\:h\\end{array}\\right.$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eHere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:E\\:\\)\u003c/span\u003e\u003c/span\u003edenotes the set of all measured events (cells) in the sample. For each patient with a sample file \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i\\)\u003c/span\u003e\u003c/span\u003e of tube \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:t\\)\u003c/span\u003e\u003c/span\u003e, the cell densities within all hypercubes are combined to measure strangeness \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sigma\\:\\)\u003c/span\u003e\u003c/span\u003e as follows:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:{\\sigma\\:}_{i,t}=\\sum\\:_{h\\in\\:H\\left(d\\right)}^{}p\\left(h\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe central limit theorem states that under very general conditions the sum of such variables attains a normal distribution \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{{N}_{\\sigma\\:}}_{i,t}(m,sd)\\)\u003c/span\u003e\u003c/span\u003e. Samples are defined as atypical if their values \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\sigma\\:}_{i,t}\\)\u003c/span\u003e\u003c/span\u003e lie outside the probability limits defined by the cumulative probability function of a robust estimated empirical normal distribution \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{{{N}^{{\\prime\\:}}}_{\\sigma\\:}}_{i,t}(m{\\prime\\:},sd{\\prime\\:})\\)\u003c/span\u003e\u003c/span\u003e. As default, two-tailed limits of 1% are used. Any sample deemed atypical in at least one tube is excluded from further analysis, ensuring that only typical, representative samples enter the training set.\u003c/p\u003e\n\u003ch3\u003eSelf-explanatory diagnostic AI committee\u003c/h3\u003e\n\u003cp\u003eThe second stage consists of a committee of AI experts that self-explanatory diagnose a sample in a process similar to the human expert. Building on the curated sample set, the supervised ALPODS expert committee \u0026mdash; an extension of the original ALPODS algorithm [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] \u0026mdash; provides self-explanation through a Bayesian decision network. For each expert, a recursive process constructs a directed acyclic graph (DAG): at each node, a marker is selected based on the Simpson index [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], which quantifies the probability that two randomly drawn cells belong to the same category; conditional dependencies between parent and child nodes are then computed via Bayes\u0026rsquo; theorem. Edges in the DAG represent these dependencies, and recursion terminates when further splitting yields subpopulations with homogeneous class labels or falls below a preset population size. The resulting network identifies key phenotypic subpopulations, which are then distilled into fast‐and‐frugal trees [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] by merging all decisions on the same marker into single range‐based conditions and ranking subgroups by effect size of Cohen\u0026rsquo;s \u003cem\u003ed\u003c/em\u003e [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo capture the hierarchical nature of hematologic diagnoses, multiple tube- and task-specific experts are trained at three diagnostic levels and aggregated into an expert committee. At level 1 (distinguishing normal controls from B-cell lymphomas), TM first selects representative cases per class; computed ABC analysis [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] then partitions learned cell populations into high-frequency (set A) and low-frequency (sets B and C) groups, each handled by a dedicated expert. Depending on the dataset, four to six experts (one per group per tube) are trained, with inter-tube weighting reflecting their differential diagnostic value. Levels 2 and 3 further subdivide lymphoma subtypes: three experts distinguish chronic lymphocytic leukemia \u0026ndash; like versus other lymphomas across all tubes, and a separate expert analyzes the tube most informative for hairy-cell leukemia. Each ALPODS expert independently provides a decision and an explanation for its decision that is ultimately aggregated into a final diagnostic output by a higher-order meta classifier. This the meta classifier uses the Gini impurity measure suited for categorical data to weigh and combine expert opinions. A weighting scheme reflects diagnostic significance per tube, assigning higher weights to more diagnostically relevant tubes. Details are provided in SI Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and subsequent supplementary information.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eSelf-assessment by degrees of trustworthiness\u003c/h3\u003e\n\u003cp\u003eThe committee of AI experts yields additionally to the self-explanatory diagnosis a self-competency estimation as the likelihood \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{c}\\)\u003c/span\u003e\u003c/span\u003e of a case being correctly classified as either an NC sample or a B-cell lymphoma sample ranging from zero to one. Recognizing the need for intuitive confidence measures in clinical contexts such as Likert scale [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], FlowXAI converts its internal probability estimates into a three-point self-assessment scale. For each sample, the self-competence \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{c}\\)\u003c/span\u003e\u003c/span\u003e is binned as follows: The FlowXAI is \u003cem\u003eConfident\u003c/em\u003e about the diagnosis of a sample if either \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{c}\\)\u003c/span\u003e\u003c/span\u003e=1 or \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{p}_{c}=\\)\u003c/span\u003e\u003c/span\u003e0, \u003cem\u003eProbable\u003c/em\u003e in the two bins of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:1\u0026gt;{p}_{c}\\ge\\:\\:0.8\\)\u003c/span\u003e\u003c/span\u003e or \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:0\u0026lt;{p}_{c}\\le\\:0.2\\)\u003c/span\u003e\u003c/span\u003e, and \u003cem\u003eChallenging\u003c/em\u003e for the remaining bin of the central zone \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:0.8\u0026gt;{p}_{c}\u0026gt;0.2\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eGeneralization and validation protocol\u003c/h2\u003e \u003cp\u003eTo evaluate generalizability and robustness, FlowXAI was tested using two rigorous cross-validation protocol. The MLL9 a underwent 100 rounds of random class-balanced 80/20% train-test splits (MLL9F: N\u0026thinsp;=\u0026thinsp;15,594 (train) / N\u0026thinsp;=\u0026thinsp;3,899 (test) with and without TM's elimination of atypical samples in the training set by unsupervised tiles mining. In contrast to prior works [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], a repetitive cross-validation technique was chosen in order to prevent the underestimation of errors [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo show the efficiently of data reduction by TM, we selected N\u0026thinsp;=\u0026thinsp;512 samples of typical training samples by TM and evaluate the performance on the test sets of all other typical and atypical samples. Next, for the PUM2 N\u0026thinsp;=\u0026thinsp;200 training samples from the typical sample set were sampled 100 times to evaluated a test set of N\u0026thinsp;=\u0026thinsp;317\u0026thinsp;+\u0026thinsp;121 samples. In all experiments, test samples were kept entirely independent throughout training to preserve external validity.\u003c/p\u003e \u003cp\u003eGiven the imbalance between different lymphoma subtypes and control samples, the Matthews correlation coefficient (MCC) was taken as the primary metric at Levels 2 and 3, owing to its superior handling of imbalance [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e] is considered more reliable than the F1 measure [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. The AI additionally reports overall accuracy, false positive rate, and false negative rate \u0026mdash; especially when stratified by the three confidence categories. Evaluating the limits of performance, the set of training representatives was reduced to 256 lymphoma and 256 control samples. Visualization of the results with mirrored-density plots (MD-plots) [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] allowed assessment of probability density functions (pdfs) of quality measures like MCC. The MD-plot visualization can confirm if there is no multimodality within the values of the quality measures by revealing Gaussian distribution (magenta line) thereby allowing to report the point estimate as an average above. A non-gaussian distribution may indicate that the split between the test set and training set, i.e., random sampling, is not homogeneous.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eStructural embodiment of medical knowledge\u003c/h3\u003e\n\u003cp\u003eAssessment of the \u0026ldquo;structural embodiment\u0026rdquo; of known lymphoma classes was achieved through unsupervised swarm intelligence using the databionic swarm [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The Databionic swarm allows to assess the structures in data using cluster analysis [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e], and, separately, provides an U-matrix visualization [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. This technique generated U-Matrix visualizations of the high-dimensional data [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e], allowing intuitive interpretation of clustering behavior. The U-Matrix creates a topographic map where valleys represent similar sample properties and mountains represent structural dissimilarities. The multicore implementation of the Databionic swarm facilitated efficient computation [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e], making it viable for large clinical datasets.\u003c/p\u003e\n\u003ch3\u003eImplementation details and access to FlowXAI\u003c/h3\u003e\n\u003cp\u003eWe have created a live, interactive demonstration of our FlowXAI platform at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://plait.mathematik.uni-marburg.de/\u003c/span\u003e\u003cspan address=\"https://plait.mathematik.uni-marburg.de/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. This online portal is specifically designed so that physicians, computational scientists, and readers can explore our proposed step-by-step classification workflow: Visitors can upload data (or select from example datasets) and see how FlowXAI automatically identifies relevant cell populations, performs classification, and provides self-explaining visual outputs. Each FlowXAI decision is grounded in traceable rules that mimic the logic of human experts. FlowXAI provides these understandable explanations stored in the FCS format. Upon acceptance we will publicly release the full code base for the platform.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eBioinformatic view on lymphoma immunophenotype classes\u003c/h2\u003e \u003cp\u003eThe presence or absence of groups or classes within a given disease entity influences the interpretation of high parametric expression data and machine learning capabilities [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. Besides, the nomenclature and categories of lymphoma category definitions have progressively shifted due to the growing knowledge of molecular genetic aberrations and pathways [\u003cspan additionalcitationids=\"CR47 CR48\" citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e] and a shift of categories is a particular challenge for AI design. Historical knowledge, for example, in the form of the WHO Lymphoma classification, may or may not be captured in the measurements of a given lymphoma immunophenotyping panel. Thus, structural embodiment refers to two aspects: whether the medical knowledge is appropriately represented in the data, and whether the structures correspond to the expert classification. We sought to obtain an unbiased bioinformatic view on lymphoma entities based on the previously published MLL9 dataset. This data set contains CLL, prolymphocytic leukemia (PL), monoclonal B-cell lymphocytosis (MBL), hairy cell lymphoma (HCL), mantle cell lymphoma (MCL), marginal zone lymphoma (MZL), lymphoplasmacytic lymphoma (LPL) and follicular lymphoma (FL).\u003c/p\u003e \u003cp\u003eEvaluating structural embodiment using a framework based on swarm intelligence and self-organization presents a compelling methodology [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e] (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The topographic map in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea depicts two main groups, two outlier groups and indicates individual highly atypical \u0026ldquo;outlier\u0026rdquo; samples. The clustering reveals a clear separation between presumed healthy individuals/normal controls (NC) and B-NHL (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). Only 3 (out of 2998) NC were assigned into the B-NHL frame while a significant number of B-NHL samples appeared within the \u0026ldquo;NC cluster\u0026rdquo; (SI Table\u0026nbsp;1). The global perspective on structural embodiment differentiates NC from B-NHL so profoundly that lymphomas with subtle features, such as HCL and FL, remain unrecognizable. Interestingly, CLL samples appeared clearly detached from other lymphomas.\u003c/p\u003e \u003cp\u003eSubsequently, the TM algorithm was applied in order to eliminate outlier samples and a second topographic map based on the entire set of lymphoma entities now revealed three main clusters and only a few remaining singular or doublet outliers (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec). Cluster matching with the given diagnoses is depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ed and reveals that AI driven immunophenotype B-NHL taxonomy does not entirely match with the lymphoma classes based on WHO definition. CLL, MBL and PL cluster together in one main category and FL, LPL and MZL constitute a second group. HCL appears as a clear distinct cluster surrounded by a chain of hills but this group also contains some MZL (SI Table\u0026nbsp;2). Interestingly, PL and MCL cluster closely in the same region, while the potentially smoldering small lymphocytic lymphomas MZL, LPL and FL appear far distant. Although lymphoma classification has been an issue of intense debate and controversies between specialists [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e] different treatment strategies empirically emerged for CLL, HCL and common lymphocytic lymphomas such as MZL, LPL, FL and MCL hereby confirming that the AI based immunophenotypic clusters reflect true biologic similarities and divergences. Of note, even in a well-defined category such as CLL prediction of treatment requirement and response is highly variable and beyond B cell morphology is related to clinical parameters or T cell distribution patterns [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]. Whether the AI based taxonomy may reflect \u0026ldquo;treatment classes\u0026rdquo; or clinical outcome parameters remains to be addressed by further analyses of additional comprehensive lymphoma data sets.\u003c/p\u003e \u003cp\u003eTaken together, our swarm based self-organization framework revealed corresponding categorization of the most common lymphomas in relation to WHO classes but MZL and LPL were not resolved clearly.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eStrategy for AI training on B-NHL data\u003c/h2\u003e \u003cp\u003eThe observation that AI taxonomy does not entirely match to the predefined lymphoma diagnoses suggests compromised learning conditions for an automated lymphoma immunophenotyping approach. Based on the visualization of the self-organizing matrix and physician\u0026rsquo;s emphasis, the diagnostic tree was designed as a five-level stepwise diagnostic approach (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The TM algorithm realizes quality control at branch level zero (L0) as designating the quality of sample data is always part of diagnostic reports, and low data quality may impair AI learning. In level L1, NC samples are distinguished from those from B-NHL patients. Correct classification of NC samples is of particular importance because a false positive diagnosis of lymphoma \u0026ldquo;cancer\u0026rdquo; fundamentally affects treatment decisions. Once B-NHL has been categorized correctly in L1, the second diagnostic level (L2) distinguishes CLL (including PL and MBL, referred as \u0026ldquo;Mature B-cell neoplasms\u0026rdquo; in the 4th WHO Classification) and HCL from all other lymphomas (AOL). For L2 a good AI separation was achieved (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e: cluster 1, 2 and 3). Moreover, a correct classification of these lymphoma entities is clinically prioritized: biopsy with histopathological confirmation is not needed to diagnose CLL [NCCN Guidelines Version 3.2023] and for HCL, immunophenotyping is an essential diagnostic requirement [NCCN Guidelines Version 1.2023]. To align our analysis with previous studies [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e], we investigated the most common and distinct mature lymphoma types in the L3 AOL subcategory: mantle cell MCL, MZL, LPL and FL. Rare B-NHL subtypes or large cell lymphomas were classified into one category as L4 remaining B-NHL and not considered for further classification.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThus, the expectation for an AI system is to perform an initial quality check to identify (outlier) samples and to achieve diagnostic performance with appropriate stringency.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eThe general performance of FlowXAI second module at the diagnostic levels\u003c/h2\u003e \u003cp\u003eThe performance of FlowXAI was evaluated on the MLL9F dataset with N\u0026thinsp;=\u0026thinsp;19,493 entire patient samples in comparison to the NN based approach [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. The total number of patients in the malignant lymphoma categories (L2 and L3) ranged from 202 (HCL) to 4,274 (CLL). The B-NHL class sizes were highly imbalanced: B-NHL 45.3%; NCs 54.7%; CLL-like group (mature B-cell neoplasms) = {CLL, PL, MBL} 33.5%; HCL 1%; Group of all other lymphomas (AOL) = {MCL, MZL, LPL, FL, and rare B-NHLs} 11.8%; CLL 22.9%; MBL 7.8%; PL 2.8%; MCL 1.5%; MZL 5.3%; LPL 3.5%; FL 1.2%. The designation of classes and lymphoma identities are listed in SI Table\u0026nbsp;3. Omitting TM, cross-validation shows that FlowXAI is capable to differentiate NC (N\u0026thinsp;=\u0026thinsp;2,133) from B-NHL (N\u0026thinsp;=\u0026thinsp;1,766) with an average accuracy of 94.2\u0026thinsp;\u0026plusmn;\u0026thinsp;0.4% in L1. FlowXAI classifies CLL-like samples, HCL samples, and AOL against NC samples with an average MCC of 85.4\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7% in L2. For the AOL categories, FlowXAI performed immunophenotyping, with an average MCC of 81.8\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7% in L3. The cross-validation results are illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea as a MD pots. Next, we incorporated the FlowXAI system\u0026rsquo;s self-assessment capability for each diagnosis: \u003cem\u003econfident\u003c/em\u003e, \u003cem\u003eprobable\u003c/em\u003e, and \u003cem\u003echallenging\u003c/em\u003e. To address the average performance of FlowXAI when these three degrees are considered, the MCC, accuracy (ACC), false positive rate (FPR), and false negative rate (FNR) were calculated for decisions on L3 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb). FlowXAI categorizes 7\u0026ndash;12% of test samples as challenging, 20\u0026ndash;35% as probable, and 52\u0026ndash;73% as \u003cem\u003econfident\u003c/em\u003e (5th to 95th percentile). Interestingly, the distributions of the number of test samples per degree of trustworthiness appear bimodal (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). A detailed look reveals 16 cross-validation cycles with less than 8% of the patients in the challenging and probable degree of trustworthiness and more corresponding classifications as \u003cem\u003econfident\u003c/em\u003e. This heterogeneous pattern of the cross-validation results underscores the necessity of conducting at least 100 cross-validation trials because the results may appear superior (or inferior) by chance. The MCC of the test samples in the confident degree of trustworthiness ranged from 88\u0026ndash;91% (median 90%). The performance of the probable degree stayed in the range of 70\u0026ndash;80%, with a median of 77%, and the challenging degree had a large variance in performance, with a range between 47% and 68% (median 60%) shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ed. The performance in the confident degree of trustworthiness follows a Gaussian distribution (magenta frame), while the other degrees appear less homogenous. This could be attributed to the high number of samples in the confident degree. Direct comparison reveals that the performance of the NN AI system is nearly equivalent to the average FlowXAI performance over all cross-validation iterations, which comprises all \u003cem\u003eprobable\u003c/em\u003e and \u003cem\u003econfident\u003c/em\u003e degrees of trustworthiness (MCC\u0026thinsp;=\u0026thinsp;83% versus 85.1%, Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ee). Of note, this performance was achieved for N\u0026thinsp;=\u0026thinsp;3,488 (89%) samples, as indicated by the red vertical line. The deep learning AI was trained on the same dataset with one cross-validation test. Of note, Flow XAI performance in the \u003cem\u003econfident\u003c/em\u003e degree was clearly superior to the NN AI based predictions. Finally, the FlowXAI performance was determined after TM's elimination of atypical samples. When atypical samples are discarded, the ACC and MCC improve slightly but not markedly for all trustworthiness degrees (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ef). False negative rates predominantly declined in the challenging degree of trustworthiness. Contingency tables were computed for each cross-validation trial to reveal the detailed performance results for every malignant lymphoma entity (SI Table\u0026nbsp;4\u0026ndash;7). In sum, by modeling the stepwise manual diagnostic approach, FlowXAI is capable of discriminating B-NHL patients with high accuracy and self-assessed competence, as needed for trustworthy diagnostic reports.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eReduction of training data and extended validations of FlowXAI\u003c/h2\u003e \u003cp\u003eFor a comprehensive lymphoma classification more categories of the 5th WHO classification such as B-lymphoblastic leukemia, large B cell lymphomas, Burkitt lymphoma and rare subtypes such as KSHV/HHV8-associated lymphomas or lymphoid proliferations and lymphomas associated with immune deficiency should be considered [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e] but small class sizes limit AI learning abilities. Therefore, TM was constructed and combined with the AI system to FlowXAI as depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea and explained in SI-Fig.\u0026nbsp;1. The TM function of outlier discrimination was validated with a set of 100 randomly drawn atypical samples by judgment of immunophenotyping experts. The manual analysis of the bivariate plots confirmed that 96 TM-selected data samples were indeed judged as altered or hardly evaluable by the human experts; i.e., they would not pass the quality control of L0. Only 4% of the atypical samples were considered usable for a trustworthy manual diagnostic workflow across all tubes as the base for a valid diagnostic report (SI Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). When both modules of the FlowXAI are employed, only 256 normal control samples, and 256 entire B-NHL samples equally distributed across all diagnoses are sufficient to learn the diagnostic levels. Despite the small training set, FlowXAI was able to classify 49% of the test samples as \u003cem\u003econfident (probable\u003c/em\u003e: 30%, challenging: 21%). For the 7,684 test samples predicted by FlowXAI in the confident degree, an MCC of 87% was reached in L3 (\u003cem\u003eprobable\u003c/em\u003e: 62%, \u003cem\u003echallenging\u003c/em\u003e: 23%). The corresponding contingency tables present the details of the diagnostic performance (SI Table\u0026nbsp;8\u0026ndash;10). When it is \u003cem\u003econfident\u003c/em\u003e, FlowXAI demonstrates superior performance compared to the deep learning AI system in the automated diagnosis of six distinct lymphoma groups (MCC 87% vs. 83%, respectively), despite being trained on only 512 samples compared to 18,274 samples for the deep learning AI system.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFurthermore, the performance of our TM-integrated Flow XAI algorithm to separate NC samples from B-NHL samples (L1) was assessed via a second independent benchmark algorithm, the cluster identification, characterization, and regression algorithm (CITRUS)[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Among the many interesting approaches for identifying cell frequencies or statistical sample properties that can be theoretically used in supervised approaches (e.g.,[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan additionalcitationids=\"CR54 CR55 CR56\" citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]), we chose CITRUS because it seemed to be most closely related to population description and lymphoma immunophenotyping. Moreover, it has been shown that ALPODS or CITRUS outperform other approaches [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. The highest performance was achieved with FlowXAI and the preselection of typical training samples by TM improved the usage of CITRUS (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb and SI Methods). In addition to our cross-validation approach, which ensured true independence of training and test data, Flow XAI was challenged with a second impartial data set with 638 samples (PUM2) from the diagnostic center Philipps-University Marburg. TM considered 517 samples as typical and 121 B-NHL samples as atypical. For each trial, the training data consisted of 100 NC and 100 B-NHL randomly selected samples (30 CLL, 15 MCL, 10 FL, 10 MZL, 5 LPL, 10 HCL, 10 MBL, and 10 NS). 100 trials of cross-validation were performed on the remaining test sets of 317 typical samples. For L1, Flow XAI classified 57% of patient samples as \u003cem\u003econfident\u003c/em\u003e, 31% as \u003cem\u003eprobable\u003c/em\u003e, and 12% as \u003cem\u003echallenging\u003c/em\u003e. The accuracy of the performance measures was chosen due to the balanced group sizes. The accuracy was 99% when Flow XAI was \u003cem\u003econfident\u003c/em\u003e, 89% when XAI estimated that the appropriate diagnosis was \u003cem\u003eprobable\u003c/em\u003e, and 69% when XAI judged a case as \u003cem\u003echallenging\u003c/em\u003e. Even with this limited number of training data FlowXAI is capable of reporting confident results in the discrimination of B-NHL patients from healthy individuals with an accuracy of 99% in more than 50% of typical patients (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ec/d). Atypical samples were classified as \u003cem\u003econfident\u003c/em\u003e or \u003cem\u003eprobable\u003c/em\u003e in 80% with an accuracy of 98% and 90%, respectively (SI Fig.\u0026nbsp;6). Finally, for training purposes explainability is of paramount importance. FlowXAI selects and reports diagnostically relevant cell populations and expression patterns in a human-understandable manner as illustrated in SI Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e\u0026ndash;5 and provides for an interactive evaluation of model-generated reports: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://plait.mathematik.uni-marburg.de/\u003c/span\u003e\u003cspan address=\"https://plait.mathematik.uni-marburg.de/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe FlowXAI system is designed to be a supportive tool for teaching the diagnostic process to identify B-NHL via multiparameter flow cytometry-based immunophenotyping. To ensure accurate classification, clinical evidence, including genetic and histopathological information, was carefully collected for both datasets, as previously published [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. But lymphoma classification has historically been fluid. After the initial description of Hodgkin lymphoma in 1832 and Non-Hodgkin Lymphoma in 1925, various classifications coexisted until the National Lymphoma Study Group (ILSG) proposed the Revised European and American Lymphoma (R.E.A.L.) classification in 1994, which firstly considered immunopathology aspects. In 1997 it was replaced by an international classification of the World Health Organization (WHO), with the 5th edition published in 2022. Importantly, the term B prolymphocytic leukemia was eliminated in WHO-HAEMS5, and a PL subclass is now designated as the prolymphocytic progression of CLL [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. For direct comparison with the consensus classification we gained an unbiased MFC based AI view of the structure embodiment of diagnoses in order to understand the AI data perspective and taxonomy and to prevent FlowXAI learning of spurious correlations, a problem that may occur with deep NN based learning [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. We applied unsupervised machine learning by Databionic swarm [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e] to identify high-dimensional structures in the data. Of note, we have previously shown that these unsupervised techniques correspond well with disease entities of divergent treatment decisions [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Correspondingly, the topographic map clearly shows separable normal control vs lymphoma classes and, subsequently, the distinct entities HCL and CLL-like, although some entities are not separable if the information is restricted to flow cytometry data. Whether the measurement of additional antigens leads to better resolution regarding delineation into WHO classes remains to be determined. Overall performance of our novel interactive Flow XAI was evaluated by computing 100 cross-validation trials with class-balanced 80/20 splits between training and test data [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e] and was equivalent to a deep NN AI system. In the \u003cem\u003econfident\u003c/em\u003e category, FlowXAI even surpassed the deep NN AI performance. Moreover, the head-to-head comparison of our FlowXAI model with the population describing algorithm CITRUS [\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e] in 100 cross-validations demonstrated superior L1 accuracy. With a second independent validation dataset, we demonstrated that the TM algorithm is efficient at calculating the distribution of strangeness values with only 100 samples. Thus, we show that immense training data reduction can be achieved using TM, and prototypes for each lymphoma category can be created with a few training samples. This unique feature paves the way for a broad applicability also for rare lymphoma subtypes.\u003c/p\u003e \u003cp\u003eBeyond the aspects of training data reduction and structural embodiment, we focused on AI suitability for real-world teaching applications through self-explanation consisting of two features: the assessment of self-competence, which aligns with the manual diagnostic reporting of scores [\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e] and the ability to illustrate relevant cell populations, which allows for explanation and teaching in an interactive manner. Such direct interaction between a data-driven AI system and clinicians has been considered the best-performing approach [\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e]. Self-explanatory clinical decision support systems are essential for trust and reliance as discussed intensively from the informatics and social points of view [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e, \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e]. AI influence in the clinical world is inevitable; hence, clinicians prepare themselves for this upcoming challenge [\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e]. To our knowledge, the design and deployment of a self-explaining symbolic AI system for lymphoma immunophenotyping has not yet been carried out. Presumably, further performance optimization may be achieved with consideration of AI based MFC lymphoma classification and by including suitable B-cell selection algorithms (see SI Literature evaluation for seeking benchmark algorithms). Notably, unselective data approaches may permit the discovery of novel putative target cell populations even beyond the well-known pathologic B cells following the basic principle of knowledge discovery.\u003c/p\u003e \u003cp\u003eTaken together, FlowXAI diagnoses B-NHL from a small amount of non-standardized flow cytometry data with equal or superior performance levels compared with NN or other benchmark algorithms. The proposed AI system is self-explanatory, knowledge-based by design, and its self-assessment capabilities facilitate a transparent and interactive mode with human experts. It may therefore serve as a base for an AI-open-access lymphoma-immunophenotyping teaching platform.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eQC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003equality control, B-NHL\u0026thinsp;=\u0026thinsp;B-cell non-Hodgkin lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003enormal control\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCLL-like\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCLL\u0026amp;MBL\u0026amp;PL\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHCL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ehairy cell leukemia, AOL\u0026thinsp;=\u0026thinsp;all other malignant lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCLL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003echronic lymphocytic leukemia\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMBL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emonoclonal B-cell lymphocytosis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eprolymphocytic leukemia\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMCL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emantle cell lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMZL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emarginal zone lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLPL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003elymphoplasmocytic lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003efollicular lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHCLv\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ehairy cell leukemia variant\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDLBCL\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ediffuse large B-cell lymphoma\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eOthers\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eall remaining rare entities like Burkitt lymphoma (BL).\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":" \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eMCT owns the company IAP-GmbH Intelligent Analytics Projects, and IAP collaborates directly with Beckman Coulter. QS is an accepted PhD student at Philipps-University and works at IAP-GmbH. JH, CB, SK, AN, and AU declare no conflicts of interest.\u003c/p\u003e\u003ch2\u003eAuthor contributions\u003c/h2\u003e \u003cp\u003eJH, SK, CB, AU, and MCT: conceptualization. MCT and QS: data analysis, MCT investigation, and visualization. PK FCS data provision. MCT: writing the original manuscript draft. JH, SK, PK, QS, AN, AU, and CB: writing-review and editing. MCT and AU: supervision.\u003c/p\u003e\u003ch2\u003eAcknowledgments\u003c/h2\u003e \u003cp\u003eWe thank Michael Schmidt for the analysis of impurity functions within the conventional decision tree software packages in R and MATLAB. Furthermore, we thank Max Zhao for communicating further elaborations about the overview of the sample files published in [52] and Monika Sikora for English corrections.\u003c/p\u003e\u003ch2\u003eData availability\u003c/h2\u003e \u003cp\u003eThe MLL9F dataset with N\u0026thinsp;=\u0026thinsp;19,493 samples and the major part of the PUM2 data set with 638 samples have been published and are accessible [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e]. The remaining files from the PUM2 dataset are accessible through \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://plait.mathematik.uni-marburg.de/\u003c/span\u003e\u003cspan address=\"https://plait.mathematik.uni-marburg.de/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBruggner RV, Bodenmiller B, Dill DL, Tibshirani RJ, Nolan GP. Automated identification of stratifying signatures in cellular subpopulations. Proc Natl Acad Sci U S A. 2014;111(26):E2770\u0026ndash;E7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eO\u0026rsquo;Neill K, Jalali A, Aghaeepour N, Hoos H, Brinkman RR. Enhanced flowType/RchyOptimyx: a bioconductor pipeline for discovery in high-dimensional cytometry data. Bioinformatics. 2014;30(9):1329\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi J-L, Lin Y-C, Wang Y-F, Monaghan SA, Ko B-S, Lee C-C. A Chunking-for-Pooling Strategy for Cytometric Representation Learning for Automatic Hematologic Malignancy Classification. IEEE Journal of Biomedical and Health Informatics. 2022;26(9):4773\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKowarsch F, Weijler L, W\u0026ouml;dlinger M, Reiter M, Maurer-Granofszky M, Schumich A, et al. Towards self-explainable transformers for cell classification in flow cytometry data. International Workshop on Interpretability of Machine Intelligence in Medical Image Computing: Springer; 2022. p. 22\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLewis JE, Cooper LA, Jaye DL, Pozdnyakova O. Automated deep learning-based diagnosis and molecular characterization of acute myeloid leukemia using flow cytometry. Modern Pathology. 2024;37(1):100373.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao M, Mallesh N, H\u0026ouml;llein A, Schabath R, Haferlach C, Haferlach T, et al. Hematologist-Level Classification of Mature B‐Cell Neoplasm Using Deep Learning on Multiparameter Flow Cytometry Data. Cytometry Part A. 2020;97:1073\u0026ndash;80. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/cyto.a.24159\u003c/span\u003e\u003cspan address=\"10.1002/cyto.a.24159\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoffmann J, Rother M, Kaiser U, Thrun MC, Wilhelm C, Gruen A, et al. Determination of CD43 and CD200 surface expression improves accuracy of B-cell lymphoma immunophenotyping. Cytometry B Clin Cytom. 2020;98(6):476\u0026ndash;82. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/cyto.b.21936\u003c/span\u003e\u003cspan address=\"10.1002/cyto.b.21936\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Dongen J, Lhermitte L, B\u0026ouml;ttcher S, Almeida J, Van Der Velden V, Flores-Montero J, et al. EuroFlow antibody panels for standardized n-dimensional flow cytometric immunophenotyping of normal, reactive and malignant leukocytes. Leukemia. 2012;26(9):1908\u0026ndash;75.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRawstron AC, Kreuzer KA, Soosapilla A, Spacek M, Stehlikova O, Gambell P, et al. Reproducible diagnosis of chronic lymphocytic leukemia by flow cytometry: an European Research Initiative on CLL (ERIC) \u0026amp; European Society for Clinical Cell Analysis (ESCCA) Harmonisation project. Cytometry Part B: Clinical Cytometry. 2018;94(1):121\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCosta E, Pedreira CE, Barrena S, Lecrevisse Q, Flores J, Quijano S, et al. Automated pattern-guided principal component analysis vs expert-based immunophenotypic classification of B-cell chronic lymphoproliferative disorders: a step forward in the standardization of clinical immunophenotyping. Leukemia. 2010;24(11):1927\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMallesh N, Zhao M, Meintker L, H\u0026ouml;llein A, Elsner F, L\u0026uuml;ling H, et al. Knowledge transfer to enhance the performance of deep learning models for automated classification of B cell neoplasms. Patterns. 2021;2(10):100351.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCabitza F, Campagner A, Malgieri G, Natali C, Schneeberger D, Stoeger K, et al. Quod erat demonstrandum? Towards a typology of the concept of explanation for the design of explainable AI. Expert Syst Appl. 2023;213:118888.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC. Exploiting Distance-Based Structures in Data Using an Explainable AI for Stock Picking. Information. 2022;13(2):51. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/info13020051\u003c/span\u003e\u003cspan address=\"10.3390/info13020051\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Ultsch A, Breuer L. Explainable AI framework for multivariate hydrochemical time series. Mach Learn Knowl Extr. 2021;3(1):170\u0026ndash;205. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/make3010009\u003c/span\u003e\u003cspan address=\"10.3390/make3010009\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC. Identification of explainable structures in data with a human-in-the-loop. Ger J Artif Intell. 2022;36:297\u0026ndash;301.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoebel R, Chander A, Holzinger K, Lecue F, Akata Z, Stumpf S, et al. Explainable AI: the new 42? In: Holzinger A, Kieseberg P, Tjoa A, Weippl E, editors. Machine Learning and Knowledge Extraction: Second IFIP TC 5, TC 8/WG 84, 89, TC 12/WG 129 International Cross-Domain Conference, CD-MAKE 2018. Cham: Springer; 2018. p. 295\u0026ndash;303.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL\u0026ouml;tsch J, Kringel D, Ultsch A. Explainable artificial intelligence (XAI) in biomedicine: Making AI decisions trustworthy for physicians and patients. BioMedInformatics. 2021;2(1):1\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHolzinger A. The next frontier: AI we can really trust. Machine Learning and Principles and Practice of Knowledge Discovery in Databases: International Workshops of ECML PKDD 2021, Virtual Event, September 13\u0026ndash;17, 2021, Proceedings, Part I: Springer; 2022. p. 427\u0026thinsp;\u0026ndash;\u0026thinsp;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSt\u0026ouml;ger K, Schneeberger D, Holzinger A. Medical artificial intelligence: the European legal perspective. Commun ACM. 2021;64(11):34\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHolzinger A, Biemann C, Pattichis CS, Kell DB. What do we need to build explainable AI systems for the medical domain? arXiv preprint arXiv:171209923. 2017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMatutes E, Owusu-Ankomah K, Morilla R, Houlihan A, Que T, Catovsky D. The immunological profile of B-cell disorders and proposal of a scoring system for the diagnosis of CLL. Leukemia. 1994;8(10):1640\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoreau EJ, Matutes E, A\u0026rsquo;Hern RP, Morilla AM, Morilla RM, Owusu-Ankomah KA, et al. Improvement of the chronic lymphocytic leukemia scoring system with the monoclonal antibody SN8 (CD79b). American journal of clinical pathology. 1997;108(4):378\u0026ndash;82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoffmann J, Rother M, Kaiser U, Thrun MC, Wilhelm C, Gruen A, et al. Determination of CD43 and CD200 surface expression improves accuracy of B-cell lymphoma immunophenotyping. Cytometry Part B: Clinical Cytometry. 2020;98(6):476\u0026ndash;82. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/cyto.b.21936\u003c/span\u003e\u003cspan address=\"10.1002/cyto.b.21936\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Hoffman J, R\u0026ouml;hnert M, Von Bonin M, Oelschl\u0026auml;gel U, Brendel C, et al. Flow Cytometry datasets consisting of peripheral blood and bone marrow samples for the evaluation of explainable artificial intelligence methods. Data In Brief. 2022;43:108382. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.dib.2022.108382\u003c/span\u003e\u003cspan address=\"10.1016/j.dib.2022.108382\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUltsch A, Hoffman J, R\u0026ouml;hnert M, Von Bonin M, Oelschl\u0026auml;gel U, Brendel C, et al. An Explainable AI System for the Diagnosis of High Dimensional Biomedical Data. BioMedInformatics. 2024;4:197\u0026ndash;218. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/biomedinformatics4010013\u003c/span\u003e\u003cspan address=\"10.3390/biomedinformatics4010013\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSimpson EH. Measurement of Diversity. Nature. 1949;163(4148):688-. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/163688a0\u003c/span\u003e\u003cspan address=\"10.1038/163688a0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuan S, Schooler LJ, Gigerenzer G. A signal-detection analysis of fast-and-frugal trees. Psychological review. 2011;118(2):316.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCohen J. Statistical power analysis for the behavioral sciences. New York: Academic Press; 2013.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUltsch A, L\u0026ouml;tsch J. Computed ABC Analysis for Rational Selection of Most Informative Variables in Multivariate Data. PloS one. 2015;10(6):e0129767. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0129767\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0129767\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePenner M, Senman L, Andoni L, Dupuis A, Anagnostou E, Kao S, et al. Concordance of Diagnosis of Autism Spectrum Disorder Made by Pediatricians vs a Multidisciplinary Specialist Team. JAMA Network Open. 2023;6(1):e2252879-e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBishop CM. Pattern Recognition and Machine Learning. New York, NY: Springer; 2006.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoodfellow I, Bengio Y, Courville A. Deep learning. Cambridge, Massachusetts, USA: MIT press; 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoughorbel S, Jarray F, El-Anbari M. Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric. PloS one. 2017;12(6):e0177678.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChicco D, Jurman G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC genomics. 2020;21(1):1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Gehlert T, Ultsch A. Analyzing the Fine Structure of Distributions. PLoS ONE. 2020;15(10):e0238835. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0238835\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0238835\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Ultsch A. Swarm Intelligence for Self-Organized Clustering. Artificial Intelligence. 2021;290:103237. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.artint.2020.103237\u003c/span\u003e\u003cspan address=\"10.1016/j.artint.2020.103237\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Pape F, Ultsch A. Conventional Displays of Structures in Data Compared With Interactive Projection-Based Clustering (IPBC). International Journal of Data Science and Analytics. 2021;12(3):249\u0026ndash;71. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s41060-021-00264-2\u003c/span\u003e\u003cspan address=\"10.1007/s41060-021-00264-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Ultsch A. Using Projection based Clustering to Find Distance and Density based Clusters in High-Dimensional Data. Journal of Classification. 2020;38(2):280\u0026ndash;312. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00357-020-09373-2\u003c/span\u003e\u003cspan address=\"10.1007/s00357-020-09373-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Ultsch A. Uncovering High-Dimensional Structures of Projections from Dimensionality Reduction Methods. MethodsX. 2020;7:101093. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.mex.2020.101093\u003c/span\u003e\u003cspan address=\"10.1016/j.mex.2020.101093\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC. Distance-Based Clustering Challenges for Unbiased Benchmarking Studies. Nature Scientific Reports. 2021;11(1):18988. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-021-98126-1\u003c/span\u003e\u003cspan address=\"10.1038/s41598-021-98126-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUltsch A, L\u0026ouml;tsch J. Machine-learned cluster identification in high-dimensional data. Journal of Biomedical Informatics. 2017;66(C):95\u0026ndash;104.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUltsch A. Maps for the visualization of high-dimensional data spaces. Workshop on Self organizing Maps (WSOM). Kyushu, Japan2003. p. 225\u0026thinsp;\u0026ndash;\u0026thinsp;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Lerch F, L\u0026ouml;tsch J, Ultsch A. Visualization and 3D Printing of Multivariate Data of Biomarkers. In: Skala V, editor. International Conference in Central Europe on Computer Graphics, Visualization and Computer Vision (WSCG). Plzen2016. p. 7\u0026ndash;16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStier Q, Thrun MC. An efficient multicore CPU implementation of the DatabionicSwarm. International Federation of Classification Societies (IFCS). San Jos\u0026eacute;, Costa Rica2024. p. 1\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThrun MC, Mack E, Neubauer A, Haferlach T, Frech M, Ultsch A, et al. A Bioinformatics View on Acute Myeloid Leukemia Surface Molecules by Combined Bayesian and ABC Analysis. Bioengineering. 2022;9(11):642. doi: bioengineering9110642.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlaggio R, Amador C, Anagnostopoulos I, Attygalle AD, De Oliveira Araujo IB, Berti E, et al. The 5th edition of the World Health Organization classification of haematolymphoid tumours: lymphoid neoplasms. Leukemia. 2022;36(7):1720\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwerdlow SH, Campo E, Pileri SA, Harris NL, Stein H, Siebert R, et al. The 2016 revision of the World Health Organization classification of lymphoid neoplasms. Blood. 2016;127(20):2375\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlaggio R, Amador C, Anagnostopoulos I, Attygalle AD, Araujo IBdO, Berti E, et al. The 5th edition of the World Health Organization classification of haematolymphoid tumours: lymphoid neoplasms. Leukemia. 2022;36(7):1720\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwerdlow SH, Campo E, Pileri SA, Harris NL, Stein H, Siebert R, et al. The 2016 revision of the World Health Organization classification of lymphoid neoplasms. Blood, The Journal of the American Society of Hematology. 2016;127(20):2375\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHiddemann W, Stein H. Die neue WHO-Klassifikation der malignen Lymphome: Endlich eine weltweit akzeptierte Einteilung. Deutsches \u0026Auml;rzteblatt. 1999;96.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoffmann J, Eminovic S, Wilhelm C, Krause SW, Neubauer A, Thrun MC, et al. Prediction of clinical outcomes with explainable artificial intelligence in patients with chronic lymphocytic leukemia Current Oncology. 2023;30(2):1903\u0026ndash;15. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/curroncol30020148\u003c/span\u003e\u003cspan address=\"10.3390/curroncol30020148\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao M, Mallesh N, H\u0026ouml;llein A, Schabath R, Haferlach C, Haferlach T, et al. Hematologist-level classification of mature B‐cell neoplasm using deep learning on multiparameter flow cytometry data. Cytometry A. 2020;97:1073\u0026ndash;80. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/cyto.a.24159\u003c/span\u003e\u003cspan address=\"10.1002/cyto.a.24159\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGreene E, Finak G, D'Amico LA, Bhardwaj N, Church CD, Morishima C, et al. New interpretable machine-learning method for single-cell data reveals correlates of clinical response to cancer immunotherapy. Patterns. 2021;2(12):100372.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAghaeepour N, Jalali A, O'Neill K, Chattopadhyay PK, Roederer M, Hoos HH, et al. RchyOptimyx: cellular hierarchy optimization for flow cytometry. Cytometry A. 2012;81(12):1022\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAghaeepour N, Finak G, FlowCAP Consortium, DREAM Consortium, Hoos H, Mosmann TR, et al. Critical assessment of automated flow cytometry data analysis techniques. Nat Methods. 2013;10(3):228\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAghaeepour N, Chattopadhyay P, Chikina M, Dhaene T, Van Gassen S, Kursa M, et al. A benchmark for evaluation of algorithms for identification of cellular correlates of clinical outcomes. Cytometry A. 2016;89(1):16\u0026ndash;21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Gassen S, Callebaut B, Van Helden MJ, Lambrecht BN, Demeester P, Dhaene T, et al. FlowSOM: using self-organizing maps for visualization and interpretation of cytometry data. Cytometry A. 2015;87(7):636\u0026ndash;45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLapuschkin S, W\u0026auml;ldchen S, Binder A, Montavon G, Samek W, M\u0026uuml;ller KR. Unmasking Clever Hans predictors and assessing what machines really learn. Nat Commun. 2019;10(1):1096.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuda RO, Hart PE, Stork DG. Pattern Classification. Second Edition ed. A Wiley-Interscience Publication. New York, USA: John Wiley \u0026amp; Sons; 2001.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBruggner RV, Bodenmiller B, Dill DL, Tibshirani RJ, Nolan GP. Automated identification of stratifying signatures in cellular subpopulations. Proceedings of the National Academy of Sciences. 2014;111(26):E2770-E7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoreau EJ, Matutes E, A\u0026rsquo;Hern RP, Morilla AM, Morilla RM, Owusu-Ankomah KA, et al. Improvement of the chronic lymphocytic leukemia scoring system with the monoclonal antibody SN8 (CD79b). Am J Clin Pathol. 1997;108(4):378\u0026ndash;82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJacobs M, Pradier MF, McCoy TH, Perlis RH, Doshi-Velez F, Gajos KZ. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Transl Psychiatry. 2021;11(1):108.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBussone A, Stumpf S, O'Sullivan D. The role of explanations on trust and reliance in clinical decision support systems. 2015 International Conference on Healthcare Informatics. Dallas, TX, USA: IEEE; 2015. p. 160\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiller T. Explanation in artificial intelligence: insights from the social sciences. Artif Intell. 2019;267:1\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJames CA, Wachter RM, Woolliscroft JO. Preparing clinicians for a clinical world influenced by artificial intelligence. JAMA. 2022;327(14):1333\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoffmann J, Eminovic S, Wilhelm C, Krause SW, Neubauer A, Thrun MC, et al. Prediction of clinical outcomes with explainable artificial intelligence in patients with chronic lymphocytic leukemia. Curr Oncol. 2023;30(2):1903\u0026ndash;15. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/curroncol30020148\u003c/span\u003e\u003cspan address=\"10.3390/curroncol30020148\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6690890/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6690890/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eArtificial intelligence (AI) systems have been proposed for multiparameter flow cytometry (MFC) based immunophenotyping of leukemic Non-Hodgkin-Lymphoma (NHL). Lymphoma classification has progressively been revised due to the increasing molecular knowledge. In order to establish a ground truth for AI learning we sought to disclose the data\u0026rsquo;s structure embodiment in comparison to the histopathological classification decisions through self-organization of data using swarm intelligence. 19,493 data samples allowed for an unsupervised view based on multidimensional MFC data hereby providing information about the higher structures within the data and the inherent discrimination capabilities. Yet, rare lymphoma entities remain a particular challenge for AI training. We propose an AI system termed Flow XAI, which exhibits equal immunophenotyping performance as neural network based systems but has reduced the number of needed learning data by a factor of 100. The Flow XAI is capable of \u0026ldquo;self-reflection\u0026rdquo;, it reports a self-competence estimation for each case. Moreover, it selects and reports diagnostically relevant cell populations and expression patterns in a discernable and clear manner so that physicians can understand the rationale behind the AI\u0026rsquo;s decisions. The self-explaining AI system can therefore be used for real-world training on MCF based lymphoma immunophenotyping.\u003c/p\u003e","manuscriptTitle":"Self-explaining Artificial Intelligence for the Classification of B cell Non-Hodgkin Lymphoma","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-22 10:10:23","doi":"10.21203/rs.3.rs-6690890/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2e382af8-01a8-484b-b90f-d63d6f8ad058","owner":[],"postedDate":"May 22nd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":48741671,"name":"Health sciences/Diseases/Cancer/Haematological cancer/Lymphoma/Non-hodgkin lymphoma/B-cell lymphoma"},{"id":48741672,"name":"Biological sciences/Cancer/Haematological cancer/Lymphoma/Non-hodgkin lymphoma/B-cell lymphoma"}],"tags":[],"updatedAt":"2025-07-02T14:39:05+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-22 10:10:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6690890","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6690890","identity":"rs-6690890","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0