Radiomic phenotyping of tuberculosis histopathology from computed tomography scans of non-human primates

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Tuberculosis (TB) lesions display structural heterogeneity associated with disease status and progression. We explore the potential of radiomic features from computed tomography (CT) scans to classify TB lesions at a microscopic level. Treated marmosets and untreated macaques infected with Mycobacterium tuberculosis underwent PET-CT imaging after a minimum of 12 weeks. Lesions were classified post-necropsy into necrotizing granulomas, non-necrotizing granulomas, inflammation beyond defined granulomas, and fibrotic granulomas. CT scans were segmented, radiomic features were extracted, and the top ten features were used to develop machine learning (ML) models for classification based on a one-versus-all approach. Top performing models were evaluated on a separate validation dataset. 151 lesions from macaques and 149 lesions from marmosets were identified. Top features included GrayLevelNonUnformity and SmallAreaHighGrayLevelEmphasis for both datasets. Linear discriminate analysis for macaques and adaptive boosting for marmosets averaged AUC-ROCs of 0.66-0.70 for discriminating lesion types, ranging from 0.45-0.52 for necrotizing granulomas to 0.91 for fibrotic granulomas. Radiomics distinguished fibrotic from non-fibrotic TB granulomas in untreated macaques, potentially relevant for identifying disease progression risk. Classification of other lesion types was modest. Despite small, heterogeneous datasets across primate species, models performed consistently, supporting CT radiomics as a promising and trainable tool for non-invasive lesion phenotyping.
Full text 119,136 characters · extracted from preprint-html · click to expand
Radiomic phenotyping of tuberculosis histopathology from computed tomography scans of non-human primates | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Radiomic phenotyping of tuberculosis histopathology from computed tomography scans of non-human primates Dmitry Cherezov, Laura E. Via, Zakariya Khaleel, Edwin Klein, and 16 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6822856/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Tuberculosis (TB) lesions display structural heterogeneity associated with disease status and progression. We explore the potential of radiomic features from computed tomography (CT) scans to classify TB lesions at a microscopic level. Treated marmosets and untreated macaques infected with Mycobacterium tuberculosis underwent PET-CT imaging after a minimum of 12 weeks. Lesions were classified post-necropsy into necrotizing granulomas, non-necrotizing granulomas, inflammation beyond defined granulomas, and fibrotic granulomas. CT scans were segmented, radiomic features were extracted, and the top ten features were used to develop machine learning (ML) models for classification based on a one-versus-all approach. Top performing models were evaluated on a separate validation dataset. 151 lesions from macaques and 149 lesions from marmosets were identified. Top features included GrayLevelNonUnformity and SmallAreaHighGrayLevelEmphasis for both datasets. Linear discriminate analysis for macaques and adaptive boosting for marmosets averaged AUC-ROCs of 0.66-0.70 for discriminating lesion types, ranging from 0.45-0.52 for necrotizing granulomas to 0.91 for fibrotic granulomas. Radiomics distinguished fibrotic from non-fibrotic TB granulomas in untreated macaques, potentially relevant for identifying disease progression risk. Classification of other lesion types was modest. Despite small, heterogeneous datasets across primate species, models performed consistently, supporting CT radiomics as a promising and trainable tool for non-invasive lesion phenotyping. Biological sciences/Microbiology/Infectious disease diagnostics Health sciences/Diseases/Infectious diseases/Tuberculosis Biological sciences/Biological techniques/Imaging/X ray tomography Biological sciences/Computational biology and bioinformatics/Machine learning Machine learning tuberculosis radiology diagnostic imaging Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Tuberculosis (TB) is characterized by extensive heterogeneity at both the individual and lesion levels, spanning variations in clinical presentation, disease progression, and treatment response. This heterogeneity underpins differences in pharmacokinetics, bacterial metabolism and host immune response [ 1 – 4 ], yet remains poorly captured by conventional diagnostic tools such as sputum culture and molecular testing which offer limited insight into the spatial and biological complexity of in vivo disease [ 5 ]. More precise characterization of this heterogeneity could enable earlier diagnosis, individualized risk stratification, and tailored treatment approaches [ 6 , 7 ]. High resolution imaging modalities such as computed tomography (CT) provide a unique opportunity to characterize diverse lesion features in a non-invasive manner. Beyond simple visual interpretation, radiomics – a high throughput approach to extract quantitative features from radiologic images – enables systematic analysis of lesion texture, shape, and intensity patterns. Radiomic features have been applied in oncology to predict molecular and histopathologic attributes that distinguish benign and malignant lesions [ 8 ], stage and prognosticate tumors [ 9 ], and predict response to different therapies[ 10 , 11 ]. Radiomics have been explored less extensively in infectious diseases, including a paucity of applications for characterizing heterogeneous TB lesion pathology that may be clinically informative. Early lesion features on CT and chest-X-rays, such as the presence or absence of fibrosis, have been linked to TB disease progression risk over the following months to years [ 12 , 13 ]. Additionally, lesion microenvironments influence critical determinates of treatment response [ 6 ], such as bacterial physiologic states and persistence [ 14 – 16 ] and drug penetration [ 17 , 18 ]. A high-resolution radiomic framework could provide a non-invasive biomarker of lesion phenotypes to guide research studies in improving TB treatment regimens and personalized clinical diagnoses. In this study, we explored whether CT-derived radiomic features could predict TB lesion histopathology. Given the absence of large-scale, lesion-level PET/CT-histopathology mapped datasets in humans, we leveraged two distinct non-human primate cohorts—marmosets and macaques—to represent different spectra of TB pathology and to assess the robustness of radiomic classification across primate species. By training and validating machine learning classifiers separately in each dataset, we aimed to explore the feasibility and generalizability of CT radiomics as a non-invasive tool for TB lesion phenotyping. Methods This project uses two existing databases of Mycobacterium tuberculosis (Mtb) infected non-human primate 2-deoxy-2-[¹⁸F] fluorodeoxyglucose positron emission tomography/computed tomography (PET/CT) scans from common marmosets at the NIAID Tuberculosis Research Section (‘NIH’) and macaques at the Flynn lab at University of Pittsburgh (‘Pitt’). All direct interactions with animals occurred in the context of other studies [ 19 – 30 ]; no interactions with animals occurred as part of this study. All necessary permissions were obtained for the use of archived data and biological samples included in this study. Common marmosets generally exhibit increased susceptibility to Mtb infection compared to macaques, with more rapid disease progression and pronounced pathology that recapitulate heterogeneity and complexity of lesions in human TB disease [ 31 ]. On the other hand, macaques, particularly cynomolgus macaques, display a spectrum of disease outcomes from latent to active tuberculosis, closely mirroring the spectrum of human infection patterns [ 32 ]. These differences are likely due to species-specific immunological differences[ 32 , 33 ]. Taken together, these datasets comprise scans across disease stages from early to advanced, lesion types from single granulomas to coalescent and cavitary, and treated and untreated disease. The workflow for the acquisition and segmentation of scans, feature extraction, feature selection and supervised and unsupervised classification analyses are provided in Fig. 1 . Untreated macaques (University of Pittsburgh) Pitt samples were selected from an existing set of lesions from 29 unvaccinated and untreated macaques (27 cynomolgus and 2 rhesus, mostly males of Chinese origin) for 14 TB immunologic studies conducted between 2010 and 2018 [ 22 – 30 ]. The macaques were previously infected intra-bronchially with 2.7–85 colony forming units (CFU) of Mtb strain Erdman and allowed to progress for 12–86 weeks prior to necropsy. CT scans were obtained within 3–12 days prior to necropsy. Scans were performed using the NeuroLogica CereTom Clinical CT and Siemens MicroPET scanners, with imaging acquisition protocol as detailed in White A et al [ 34 ]. Images were exported as raw DICOM files. During necropsy, individual lung lesions were plated and identified using the pre-necropsy CT image. The plated lesions were used to prepare histologic sections consisting of formalin fixed tissues embedded in paraffin, sectioned, and stained with hematoxylin and eosin. In total, 314 histologic sections from the 29 animals in 14 studies mapped to CT were available and scored for histologic classification. All experiments on live vertebrates were approved by the University of Pittsburgh Institutional Animal Care and Use Committee (IACUC protocol numbers 15126588, 11110045, 11090030, 15055811, 15035407, 1105870, 14043492, 12060181, 17091258, 12080653 and 1003622). Experiments were performed and reported in accordance with the ARRIVE guidelines. Treated marmosets (NIH) The NIH samples consisted of lesions from 44 common marmosets ( Callithrix jacchus ) exposed to a low dose of aerosolized Mtb strain H37Rv. The disease was allowed to progress for 7 to 8 weeks, and antibiotic treatment was given for 4 or 8 weeks. The animals were treated for one to two months with linezolid, one of five other investigational oxazolidinones, or rifampicin, isoniazid, pyrazinamide or moxifloxacin or a combination regimen of these four drugs to prevent mortality and achieve an infection time of at least 12 weeks so that fibrotic cuffs and other signs of an immunological response could occur. Some of the animals and lesions used in this analysis were also previously analyzed for drug-specific responses to treatment, although in those analyses only features such as lesion mean HU, mean SUV, total lesion glycolysis and the change in these features between time points were explored [ 21 ]. To be eligible for inclusion in the dataset, the animals had a pre-necropsy scan within 48 hours of euthanasia, labeled histological samples preserved, and CFU/lesion determined [ 31 ]. Scans were also performed using the NeuroLogica CereTom Clinical CT and Siemens MicroPET scanners aligned using a common bed as with the Pitt scanner, with the same imaging acquisition protocol as Pitt detailed in White et al. [ 34 ] with the following exceptions: the NIH CT data was collected using 120 Kv, 5 mA/s, and the FDG PET scan began collection at 60 minutes post injection of the tracer as only two fields of view were necessary to cover the lung region of the marmoset in the Siemens Focus 220. After necropsy, there were 278 fixed hematoxylin and eosin-stained slides of classifiable TB lesions from the 44 animals studied. All live vertebrate procedures were performed in accordance with the recommendations of the Guide for the Care and Use of Laboratory Animals of the National Institutes of Health and approved by the NIAID Animal Care and Use Committee in Protocol LCIM-9 (Permit issued to NIH as A-4149-01). Both male and female marmosets ages 2 to 6 years old were included in the study. Histological classification of TB lesions All available TB lesion histologic sections were classified using a classification protocol for histologic scoring developed by a veterinary pathologist to capture TB lesion phenotypes shared across human and nonhuman primates [ 35 ] (Fig. 2 , Supplementary materials appendix S1). The histological classes include necrotizing granulomas (type 1), non-necrotizing granulomas (type 2), architecturally unstructured granulomatous inflammation outside of defined granuloma structure formation (type 3), and fibrotic granulomas (type 4). Archived fixed or digital slides were reviewed by pathologists who scored lesions according to one of four categories in the manual. Slides which contain multiple lesion types were scored according to the predominant lesion type. Slides that could not be scored due to missing lesions or unclear labeling were marked as indeterminate. Segmentation of TB lesions from CT TB lesions were manually segmented from CT scans using Amira (Thermo Fisher Scientific version 6.5) to create binary mask region of interest (ROI) for each TB lesion (See supplementary materials for protocol). Segmentation was performed by a trained individual guided by two individuals performing quality control. Individuals performing segmentation were blinded to lesion classification. DICOM files of scan images from the NIH marmosets and Pitt macaques were imported into Amira and segmentation of the ROIs were performed as follows: (i) 3D ROIs were created by wrapping or interpolating between manually segmented cross sections, (ii) ROIs were refined by removing non-lesion tissue (e.g. large vasculature/airways, mediastinal structures, chest wall), (iii) Any material with a Hounsfield (HU) lower than − 500 attributed to aerated, normal lung tissue was removed from the ROI. This segmentation process was performed for each lesion following a standard operating procedure (see Supplementary materials appendix S2). The segmentation was reviewed for quality by a separate reader to ensure protocol adherence, and inclusion of only TB lesions within the ROIs. Separation of validation set Data from both sites were randomly split into training and validation sets (D1 and D2 for Pitt, D3 and D4 for NIH). The split was performed at the primate level to prevent data leakage, ensuring that lesions from the same primate did not appear in both training and validation sets. The number of lesions and primates in each dataset (D1–D4) is detailed in Fig. 3 . Radiomic feature extraction A total of 74 first and second-order radiomic features (Supplementary table S1 ) were extracted from the lesion data using PyRadiomics, an open-source radiomic extraction python package [ 36 ]. The 74 3D features included: (i) six first-order statistics features to describe the basic distribution of voxel intensities within the identified lesion ROIs [ 36 ], (ii) 22 Haralick feature [ 37 ] from a gray-level co-occurrence matrix that captures textural patterns that could encode variation in the different presentations of lesions and (iii) 46 additional texture features from the gray-level run length[ 38 , 39 ] size zone [ 40 ]. PET images had larger voxel sizes compared to CT images (~ 10mm for PET vs. ~2mm for CT), resulting in imprecise alignment with the CT-based ROI. Consequently, feature extraction from PET was limited to SUVmax because SUVmax is less affected by the partial volume effect, where PET voxels extend beyond the smaller CT-based lesion ROIs. Lesion size exclusion criteria The two groups of primates in the study allowed us to observe biological variability in lesion representation on CT images. The lesion sizes differed significantly between the two groups. For each lesion, we identified the axial slice with the largest region of interest (ROI). The median value ROI area at this slice (TB lesion area) is 220 pixels in marmosets and 20 pixels in macaques. Since our study employed handcrafted features, one of the standard methods for quantifying image textures involves the use of a kernel with a size of 5 × 5 pixels [ 36 ]. In contrast, Gray Level Co-occurrence Matrix (GLCM) features rely on the statistical distribution of pixel intensities within the region of interest (ROI). Therefore, ROIs that were less than 5 pixels were excluded in the primary analysis due to concern of poor reliability and representation of the computed features [ 41 ]. Feature selection To enhance model generalizability and minimize the risk of overfitting, we performed radiomic feature selection as a pre-processing step. Feature selection was carried out using the Minimum Redundancy Maximum Relevance (mRMR) algorithm to ensure that the selected features had the maximum relevance to the target variable while minimizing redundancy. This approach facilitates the identification of a subset of features that contribute the most to prediction accuracy, while simultaneously reducing the complexity of the model[ 42 ]. We evaluated the performance of the resulting models for the top 5 and the top 10 features. Unsupervised analysis t-Distributed Stochastic Neighbor Embeddings (tSNE) of normalized training data using the top 10 selected features was used to visualize the topology of the N-dimensional clustering in a three-dimensional space [ 43 ]. This was accomplished by constructing probability distributions such that similar high-dimensional samples have higher probabilities, and dissimilar samples have lower probabilities. These probabilities were then mapped onto a three-dimensional space for visualization. The resulting plots demonstrate the effectiveness of the selected feature representation in capturing the structure of the data in a reduced-dimensional space. tSNE clustering was performed in R, version 4.1.1, using the Rtsne package [ 44 ]. Supervised analysis Supervised analysis was used to assess how well the model could generalize to new, unseen data and its ability to correctly identify the histologic class of a TB lesion. We used two different approaches to distinguish between multiple classes. The first approach employs a multi-class (MC) classification strategy [ 45 ], where a single model is trained to classify the input data into multiple classes. In this strategy, the classifier distinguishes quantitative features into distinct histological types, with each type encoded by a unique class label (e.g., histological type 1 is encoded as class 1, type 2 as class 2, etc.). The second approach follows the "one-versus-all" (OvA) strategy [ 46 ]. Unlike the multi-class method, this strategy involves designing an individual classification model for each histological type. For instance, one model is dedicated to classifying histological type 1—where histological type 1 is encoded as class 1 and all other types as class 0—and this process is repeated for every histological type. Here we focus on the results from the OvA approach. Results from the MC approach are presented in the supplementary materials. Two training strategies were employed: Leave-One-Out (LOO) cross-validation and standard training with a held-out validation set. The LOO method evaluates model performance and robustness by iteratively using every sample as a test case once, while the remaining data is used to train the model. This approach provides an unbiased accuracy estimate and maximizes data utility—particularly valuable for smaller datasets. Only the training datasets (D1 for Pittsburgh data and D3 for NIH) were used in LOO to maintain consistency with the validation split. LOO cross-validation was implemented as follows: In each iteration, one sample was held out as test data, while the rest trained the model. The mRMR algorithm dynamically selected the top 5 and top 10 most discriminative features per iteration, ensuring adaptive feature filtering. This process was repeated across all samples, systematically assessing which texture features best differentiated lesion types. By isolating test data in every iteration, LOO mitigates overfitting and offers a rigorous evaluation of feature generalizability. Following the Leave-One-Out (LOO) cross-validation performed on the training datasets (D1 for Pitt and D3 for NIH), we proceeded to train final diagnostic models using the complete D1 and D3 datasets. These models incorporated the most discriminative features identified during the LOO analysis. The trained models were then evaluated on the separate validation datasets (D2 for Pitt and D4 for NIH) to assess their diagnostic performance. This approach ensured that the validation was performed on entirely independent data, preventing any information leakage and providing a reliable measure of the models' generalization capability. Statistical analysis For both the supervised and unsupervised analysis methods, the following algorithms were used for model training: random forests (RF), support vector machine (SVM), adaptive boosting (Ada-Boost), and linear discriminant analysis (LDA). Results were quantified by calculating overall accuracy, weighted accuracy, and the area under the receiving operating characteristic curve (ROC-AUC) for each trained classifier, calculating these metrics on an OvA basis. The methodology described above, including LOO cross-validation and training with validation, was implemented and examined for two feature selection setups: selecting the top 5 and the top 10 best features. The top 10 features were chosen for the primary analysis because they gave the best performance for the training dataset in terms of AUC for the supervised models. Results of the analysis using the top 5 features are presented in the supplementary materials. Results Among 556 segmented lesions from 29 Pitt macaques and 488 segmented lesions from 44 NIH marmosets annotated on PET/CT, 151 lesions from 29 macaques and 149 lesions from 41 marmosets matched with associated histopathology class labels and met the lesion ROI size criteria (Fig. 3 ). Among these, 65 lesions from 15 marmosets and 27 lesions from 5 macaques were randomly set aside in an independent dataset (D2 and D4 respectively). The remaining eligible dataset was assigned to datasets D1 and D3. Pitt dataset The 150 untreated macaque lesions in the Pitt dataset used in the analysis were compromised of 76 (50.7%) type 1 lesions, 29 (19.3%) type 2 lesions, and 45 (30.0%) type 4 lesions. A single type 3 lesion representing granulomatous inflammation outside of the defined granuloma structure formation was removed from the analysis due to the paucity of type 3 lesions leading to severe class imbalance. Class imbalance in machine learning occurs when the distribution of samples across different classes is highly uneven, meaning some classes have significantly fewer examples than others [ 47 ]. This imbalance can harm model performance by causing the algorithm to become biased toward the majority class [ 48 ], leading to poor predictive accuracy for the underrepresented (but often critical) minority class, such as rare diseases in medical diagnosis or fraud cases in financial transactions. Therefore, no judgments can be made about separability of type 3 lesions in the Pitt dataset. Selected features from the training/test set are shown in Table S2 . Experiments were done with both the OvA and the MC classification approach, with similar results achieved using both approaches. Unsupervised tSNE clustering results of the training dataset are shown (Fig. 4 a). The LDA model was selected as the best-performing model for the leave-one-out approach, with an average AUC of 0.66 and a balanced accuracy of 52.76% (Fig. 4 b). Applied to the independent validation set, the LDA model yielded AUCs of 0.52 (95% CI = 0.26–0.77) for type 1 lesions, 0.69 (95% CI = 0.48–0.88) for type 2 lesions, and 0.91 (95% CI = 0.77–1.00) for type 4 lesions. NIH dataset The 149 treated marmoset lesions in the NIH dataset were comprised of 31 (20.8%) type 1 lesions, 28 (18.8%) type 2 lesions, 60 (40.3%) type 3 lesions, and 30 (20.1%) type 4 lesions. Selected features are shown in Table S2 . Like the analysis of the Pitt dataset, the OvA and MC approaches yielded comparable results. Unsupervised tSNE clustering results of the training dataset are shown (Fig. 5 a). Supervised classification results demonstrate a higher balanced accuracy but lower AUC compared to the results from the Pitt dataset (Fig. 5 b and 5 c). The Ada-boost model was selected as the best-performing model for the training set, with an average AUC of 0.70 and a balanced accuracy of 38.8% (Fig. 5 b). Applied to the validation set, the Ada-boost model yielded AUCs of 0.47 (95% CI = 0.29–0.65) for type 1 lesions, 0.58 (95% CI = 0.38–0.78) for type 2 lesions, 0.70 (95% CI = 0.56–0.83) for type 3 lesions, and 0.65 (95% CI = 0.50–0.79) for type 4 lesions. Discussion We applied radiomic feature extraction to identify shape- and texture-based characteristics that differentiate TB lesion histopathologic characteristics across two primate species, with and without treatment. Radiomic models have been widely used in cancer research and more recently in TB diagnosis, including to distinguish TB from other respiratory conditions [ 49 – 58 ] and predict drug resistance [ 59 , 60 ], but have not looked at the resolution of histopathology. To our knowledge, this is the first application of radiomics to discriminate TB lesion pathology at a microscopic level. Using these extracted features and existing experimental data from a limited cohort of untreated macaques and treated marmosets, we trained machine learning models that exhibited modest performance in classifying TB lesion histology types among a small and biologically diverse non-human primate dataset, supporting the feasibility and trainability of this approach with further refinement on existing datasets. Notably, the models distinguished fibrotic from non-fibrotic granulomas with high fidelity in untreated macaques, highlighting the potential for radiomics to resolve biologically meaningful lesion types that may relate to clinical risk. Compared to macaques, marmosets exhibited larger and more extensive lesions. Model performance metrics were generally consistent across species, feature sets, and classifier types, suggesting that the approach does not require extensive fine-tuning to yield stable results. However, performance varied by lesion type in the independent validation sets. Specifically, necrotizing granulomas (Type I lesions) were the most difficult to classify, with AUC ROC values close to 0.5 in both datasets, and were often misclassified as Type IV lesions—or as Type III lesions in the NIH dataset. In contrast, Type IV lesion predictions were more accurate in untreated macaques, with an AUC ROC of 0.9 [95% CI 0.77–1.00], than in treated marmosets, where the AUC ROC dropped to 0.47 [95% CI 0.29–0.65]. This discrepancy may reflect treatment-induced changes in lesion phenotype among marmosets, which could obscure distinctions present in untreated disease. Treated animals may also have shown a greater prevalence of mixed histopathologic features, which were categorized based on the predominant phenotype. In clinical imaging surveys of chest X-ray and PET-CT among individuals with risks for TB exposure, non-fibrotic (“active”) TB-like lesions were several times more likely to progress to active TB than fibrotic (“inactive”) lesions [ 12 , 13 ]. Therefore, the ability to distinguish fibrotic (Type IV) lesions from non-fibrotic types in untreated individuals may become a valuable biomarker for predicting and characterizing early TB before individuals develop illness or transmit to others. Notably, recent and ongoing studies are now applying CT and chest-X-ray imaging in high-risk human populations to identify early TB, offering an opportunity to test and refine this classifier while informing training of new clinical models of TB progression risk. Interestingly, SUVmax was not among the top-selected features in either dataset, indicating that classification was driven by CT-based radiomics rather than PET metabolic activity. The contribution of SUVmax may have also been limited by the considerably lower PET resolution compared to the CT-based ROI. Rather, features associated with structural organization and gray-level intensity variations were more predictive, reflecting the structural complexity of lesions. Features such as Small Area High Gray Level Emphasis were higher in inflammation beyond granulomas (Type 3; marmosets), potentially reflecting greater fragmentation and high-intensity structures with dense cellular infiltration, while fibrotic granulomas had lower values potentially reflecting more uniform collagen deposition. In both datasets, necrotizing granulomas (Type 1) and fibrotic granulomas (Type 4) also tended to have higher Contrast values, potentially reflecting sharp intensity transitions in necrotic cores and fibrotic lesions, whereas inflammation beyond granulomas (Type 3; marmosets) tended to have lower contrast, potentially due to more diffuse inflammatory involvement. These findings suggest that radiomic features, particularly those related to small-scale heterogeneity and intensity transitions, may be relevant in distinguishing TB pathology across species and treatment conditions. Early data from large autopsy studies have identified granulomas with culturable Mtb among individuals who died from non-TB causes, suggesting a prevalence and range of clinically-occult pathology that may range from self-resolution to requiring various preventive interventions [ 13 ]. The ability to phenotype lesions in-vivo would provide an opportunity to characterize the range of minimal TB disease in contacts or high-risk individuals that herald a likelihood of progression to bacteriologically positive disease. Additionally, in-vivo lesion phenotyping may also provide a useful tool to personalize treatment approaches and design host-directed or anti-microbial drug regimens effective across the range of lesion microenvironments [ 3 ]. This high-resolution imaging-based treatment response tool could apply to sputum paucibacillary and drug-resistant TB where treatment monitoring methods are limited or critical. There were several limitations of the study. Our decision not to analyze a combined NIH and Pitt dataset due to the presence of batch effects highlights an important limitation: the differences in acquisition related factors can have a substantial impact on the corresponding radiomic features, even in CT scans[ 61 , 62 ]. Therefore, models were separately generated for each dataset. Additionally, the study used existing datasets collected from different facilities, primate species, and infection/treatment protocols for different primary objectives, conditions leading to substantial heterogeneity and reduced sample sizes for each group in part because not all observed lesions in the PET/CT scans were captured in a corresponding histological section. As a result of the small sample sizes, the analysis was restricted to traditional classification methods. Deep learning approaches such as convolutional neural networks or vision transformers were not explored in this study. Acquiring larger datasets may allow for the use of deep learning models to validate our results. Future work to explore methods of normalizing the extracted features across imaging acquisition sites such as ComBat batch effect adjustment [ 63 ] are also relevant to verify these findings. Nevertheless, these limitations may be viewed as a representation of the diversity of human pathology from early to later disease stages and supports the ability of radiomic classifiers to profile TB lesions to microscopic resolutions among vastly different disease states. Finally, a substantial proportion of the macaque lesions were small (< 5 voxels in the largest slice) which did not meet the run length of several radiomic features and were thus excluded. Therefore, in very small lesions (e.g., early lesions or certain animal models), there may be a limit to what the radiomic models can reliability measure. Altogether, these findings highlight the potential of radiomics to non-invasively classify histologic TB lesion types, with consistent performance for fibrotic lesions and modest results for others. With expanding access to annotated experimental and clinical datasets, model refinement may improve classification of other lesion types and advance lesion-level phenotyping tools for clinical and research applications in TB pathogenesis, treatment strategies, and detection of early TB and TB across the disease spectrum. Declarations Data Availability All extracted radiomic feature data generated and/or analyzed during this study are included in this published article (and its Supplementary Materials files). PET/CT scans are available from the corresponding author on reasonable request. Acknowledgements Funding was provided in part by the Bill and Melinda Gates Foundation through OPP053284 (PLL), OPP1162695 and OPP1024021 (CEB), OPP1034408 and INV020435 (JF), Grand Challenges – Annual Meeting “Call to Action” (LX, JF), and in part by the Division of Intermural Research, NIAID, NIH (CEB and LEV). We thank the Comparative Medicine Branch of NIAID, NIH for clinical care of the marmosets and appreciate the technical expertise of Emmanual Dayao DVM, and Becky Sloan, and of NIAID. Author Contributions EK, LEV, HY performed lesion analysis. LEV, DMW, MS, MP, ED, HJB, AW, PM, CS, PLL collected the PET/CT studies, labeled scans, catalogued lesions, and performed necropsies. DC and VD performed the machine learning analysis under the supervision of AM. ZK performed segmentations, data cleaning, data analysis and manuscript writing. RR performed data analysis and manuscript writing. JF and CEB supervised the non-human primate studies and helped conceptualize the original study design. YLX conceptualized and supervised the study design, analyses, and manuscript writing. Competing Interests The authors declare no competing interests. References Cadena, A.M., S.M. Fortune, and J.L. Flynn, Heterogeneity in tuberculosis. Nat Rev Immunol, 2017. 17 (11): p. 691-702. Dhar, N., J. McKinney, and G. Manina, Phenotypic Heterogeneity in Mycobacterium tuberculosis. Microbiol Spectr, 2016. 4 (6). Lenaerts, A., C.E. Barry Iii, and V. Dartois, Heterogeneity in tuberculosis pathology, microenvironments and therapeutic responses. Immunological Reviews, 2015. 264 (1): p. 288-307. Lin, P.L., et al., Radiologic Responses in Cynomolgus Macaques for Assessing Tuberculosis Chemotherapy Regimens. Antimicrob Agents Chemother, 2013. 57 (9): p. 4237-4244. Chen, X. and T.Y. Hu, Strategies for advanced personalized tuberculosis diagnosis: Current technologies and clinical approaches. Precis Clin Med, 2021. 4 (1): p. 35-44. Xie, Y.L., et al., Fourteen-day PET/CT imaging to monitor drug combination activity in treated individuals with tuberculosis. Sci Transl Med, 2021. 13 (579). Walter, N.D., et al., Lung microenvironments harbor Mycobacterium tuberculosis phenotypes with distinct treatment responses. Antimicrob Agents Chemother, 2023. 67 (9): p. e0028423. Gillies, R.J., P.E. Kinahan, and H. Hricak, Radiomics: Images Are More than Pictures, They Are Data. Radiology, 2016. 278 (2): p. 563-77. Aerts, H.J.W.L., et al., Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nature Communications, 2014. 5 (1): p. 4006. Jain, P., et al., Novel Non-Invasive Radiomic Signature on CT Scans Predicts Response to Platinum-Based Chemotherapy and Is Prognostic of Overall Survival in Small Cell Lung Cancer. Front Oncol, 2021. 11 : p. 744724. Bera, K., et al., Predicting cancer outcomes with radiomics and artificial intelligence in radiology. Nat Rev Clin Oncol, 2022. 19 (2): p. 132-146. Esmail, H., et al., High resolution imaging and five-year tuberculosis contact outcomes. medRxiv, 2023. Sossen, B., et al., The natural history of untreated pulmonary tuberculosis in adults: a systematic review and meta-analysis. The Lancet Respiratory Medicine, 2023. 11 (4): p. 367-379. Lavin, R.C. and S. Tan, Spatial relationships of intra-lesion heterogeneity in Mycobacterium tuberculosis microenvironment, replication status, and drug efficacy. PLoS Pathog, 2022. 18 (3): p. e1010459. Gold, B. and C. Nathan, Targeting Phenotypically Tolerant Mycobacterium tuberculosis. Microbiol Spectr, 2017. 5 (1). Barry, C.E., et al., The spectrum of latent tuberculosis: rethinking the biology and intervention strategies. Nature Reviews Microbiology, 2009. 7 (12): p. 845-855. Prideaux, B., et al., The association between sterilizing activity and drug distribution into tuberculosis lesions. Nature Medicine, 2015. 21 (10): p. 1223-1227. Dartois, V., The path of anti-tuberculosis drugs: from blood to lesions to mycobacterial cells. Nature Reviews Microbiology, 2014. 12 (3): p. 159-167. Budak, M., et al., A systematic efficacy analysis of tuberculosis treatment with BPaL-containing regimens using a multiscale modeling approach. CPT Pharmacometrics Syst Pharmacol, 2024. 13 (4): p. 673-685. Boshoff, H.I.M., et al., Mtb-Selective 5-Aminomethyl Oxazolidinone Prodrugs: Robust Potency and Potential Liabilities. ACS Infectious Diseases, 2024. 10 (5): p. 1679-1695. Greenstein T, V.L., Moraes MP, Weiner DM, et al., PET/CT multivariate tuberculosis treatment response profiles in marmosets unify disparate preclinical biomarkers. Sci Transl Med, 2025. Maiello, P., et al., Rhesus Macaques Are More Susceptible to Progressive Tuberculosis than Cynomolgus Macaques: a Quantitative Comparison. Infect Immun, 2018. 86 (2). Ganchua, S.K.C., et al., Lymph nodes are sites of prolonged bacterial persistence during Mycobacterium tuberculosis infection in macaques. PLoS Pathog, 2018. 14 (11): p. e1007337. Lin, P.L., et al., PET CT Identifies Reactivation Risk in Cynomolgus Macaques with Latent M. tuberculosis. PLoS Pathog, 2016. 12 (7): p. e1005739. Medrano, J.M., et al., Characterizing the Spectrum of Latent Mycobacterium tuberculosis in the Cynomolgus Macaque Model: Clinical, Immunologic, and Imaging Features of Evolution. J Infect Dis, 2023. 227 (4): p. 592-601. Coleman, M.T., et al., Early Changes by (18)Fluorodeoxyglucose positron emission tomography coregistered with computed tomography predict outcome after Mycobacterium tuberculosis infection in cynomolgus macaques. Infect Immun, 2014. 82 (6): p. 2400-4. Gideon, H.P., et al., Variability in tuberculosis granuloma T cell responses exists, but a balance of pro- and anti-inflammatory cytokines is associated with sterilization. PLoS Pathog, 2015. 11 (1): p. e1004603. Rodgers, M.A., et al., Preexisting Simian Immunodeficiency Virus Infection Increases Susceptibility to Tuberculosis in Mauritian Cynomolgus Macaques. Infect Immun, 2018. 86 (12). Lin, P.L., et al., Sterilization of granulomas is common in active and latent tuberculosis despite within-host variability in bacterial killing. Nat Med, 2014. 20 (1): p. 75-9. Darrah, P.A., et al., Prevention of tuberculosis in macaques after intravenous BCG immunization. Nature, 2020. 577 (7788): p. 95-102. Via Laura, E., et al., Differential Virulence and Disease Progression following Mycobacterium tuberculosis Complex Infection of the Common Marmoset (Callithrix jacchus). Infection and Immunity, 2013. 81 (8): p. 2909-2919. Scanga, C.A. and J.L. Flynn, Modeling tuberculosis in nonhuman primates. Cold Spring Harb Perspect Med, 2014. 4 (12): p. a018564. Peña Juliet, C. and W.-Z. Ho, Non-Human Primate Models of Tuberculosis. Microbiology Spectrum, 2016. 4 (4): p. 10.1128/microbiolspec.tbtb2-0007-2016. White, A.G., et al., Analysis of 18FDG PET/CT Imaging as a Tool for Studying Mycobacterium tuberculosis Infection and Treatment in Non-human Primates. J Vis Exp, 2017(127). Leong, F.J., Dartois, V., & Dick, T. (Eds.), A Color Atlas of Comparative Pathology of Pulmonary Tuberculosis (1st ed.) . 2010: CRC Press. van Griethuysen, J.J.M., et al., Computational Radiomics System to Decode the Radiographic Phenotype. Cancer Research, 2017. 77 (21): p. e104-e107. Haralick, R.M., K. Shanmugam, and I. Dinstein, Textural Features for Image Classification. IEEE Transactions on Systems, Man, and Cybernetics, 1973. SMC-3 (6): p. 610-621. Galloway, M.M., Texture analysis using gray level run lengths. Computer Graphics and Image Processing, 1975. 4 (2): p. 172-179. Chu, A., C.M. Sehgal, and J.F. Greenleaf, Use of gray value distribution of run lengths for texture analysis. Pattern Recognition Letters, 1990. 11 (6): p. 415-419. Thibault, G., et al., Texture Indexes and Gray Level Size Zone Matrix Application to Cell Nuclei Classification , in 10th International Conference on Pattern Recognition and Information Processing . 2009. Jensen, L.J., et al. Stability of Radiomic Features across Different Region of Interest Sizes—A CT and MR Phantom Study . Tomography, 2021. 7 , 238-252 DOI: 10.3390/tomography7020022. van Timmeren, J.E., et al., Feature selection methodology for longitudinal cone-beam CT radiomics. Acta Oncol, 2017. 56 (11): p. 1537-1543. Van der Maaten, L. and G. Hinton, Visualizing data using t-SNE. Journal of machine learning research, 2008. 9 (11). Krijthe, J.H., Rtsne: T-distributed Stochastic Neighbor Embedding using Barnes-Hut Implementation . 2015. Aly, M., Survey on multiclass classification methods. Neural Netw, 2005. 19 (1-9): p. 2. Rifkin, R. and A. Klautau, In defense of one-vs-all classification. Journal of machine learning research, 2004. 5 (Jan): p. 101-141. Guo, X., et al. On the Class Imbalance Problem . in 2008 Fourth International Conference on Natural Computation . 2008. Japkowicz, N. and S. Stephen, The class imbalance problem: A systematic study. Intell. Data Anal., 2002. 6 (5): p. 429–449. Li, P., et al., A CT-based radiomics predictive nomogram to identify pulmonary tuberculosis from community-acquired pneumonia: a multicenter cohort study. Frontiers in Cellular and Infection Microbiology, 2024. 14 . Zhang, X., et al., Deep learning PET/CT-based radiomics integrates clinical data: A feasibility study to distinguish between tuberculosis nodules and lung cancer. Thorac Cancer, 2023. 14 (19): p. 1802-1811. Li, Y., et al., Machine learning-based radiomics to distinguish pulmonary nodules between lung adenocarcinoma and tuberculosis. Thoracic Cancer, 2024. 15 (6): p. 466-476. Thattaamuriyil Padmakumari, L., et al., The Role of Chest CT Radiomics in Diagnosis of Lung Cancer or Tuberculosis: A Pilot Study. Diagnostics, 2022. 12 (3): p. 739. Hu, Y., et al., Lung CT-based multi-lesion radiomic model to differentiate between nontuberculous mycobacteria and Mycobacterium tuberculosis. Med Phys, 2025. 52 (2): p. 1086-1095. Li, H.L., et al., Multimodal machine learning-based model for differentiating nontuberculous mycobacteria from mycobacterium tuberculosis. Front Public Health, 2025. 13 : p. 1470072. Li, P., et al., A CT-based radiomics predictive nomogram to identify pulmonary tuberculosis from community-acquired pneumonia: a multicenter cohort study. Front Cell Infect Microbiol, 2024. 14 : p. 1388991. Zhou, L., et al., A retrospective study differentiating nontuberculous mycobacterial pulmonary disease from pulmonary tuberculosis on computed tomography using radiomics and machine learning algorithms. Ann Med, 2024. 56 (1): p. 2401613. Yan, Q., et al., CT ‑based radiomics analysis of consolidation characteristics in differentiating pulmonary disease of non ‑tuberculous mycobacterium from pulmonary tuberculosis. Exp Ther Med, 2024. 27 (3): p. 112. Wei, S., et al., Differentiating mass-like tuberculosis from lung cancer based on radiomics and CT features. Translational Cancer Research, 2021. 10 (10): p. 4454-4463. Jiang, F., et al., A CT-based radiomics analyses for differentiating drug ‑resistant and drug-sensitive pulmonary tuberculosis. BMC Med Imaging, 2024. 24 (1): p. 307. Li, Y., et al., Radiomics analysis of lung CT for multidrug resistance prediction in active tuberculosis: a multicentre study. Eur Radiol, 2023. 33 (9): p. 6308-6317. Patrick, L., et al., Evaluating stability of histomorphometric features across scanner and staining variations: prostate cancer diagnosis from whole slide images. Journal of Medical Imaging, 2016. 3 (4): p. 047502. Cherezov, D., et al., Rank acquisition impact on radiomics estimation (AсquIRE) in chest CT imaging: A retrospective multi-site, multi-use-case study. Comput Methods Programs Biomed, 2024. 244 : p. 107990. Zhang, Y., G. Parmigiani, and W.E. Johnson, ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genomics and Bioinformatics, 2020. 2 (3): p. lqaa078. Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterialsTBradiomicmsv4.213May2025.docx SupplementaryMaterialsappendixS1IllustrationsandFurtherClarificationforRadiomicsGranulomaClassificationSystem.pdf SupplementaryMaterialsappendixS2SOPforsegmentation14May2025.docx SupplementaryMaterialsappendixS3RadiologicFeaturesDataset.csv SupplementarymaterialsappendixS4Primatelists.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6822856","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":526262084,"identity":"34eeb0a8-4681-45e6-95b4-01cf48872446","order_by":0,"name":"Dmitry Cherezov","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Dmitry","middleName":"","lastName":"Cherezov","suffix":""},{"id":526262087,"identity":"2e3498af-b000-4d97-a257-7826293bdd6a","order_by":1,"name":"Laura E. Via","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Laura","middleName":"E.","lastName":"Via","suffix":""},{"id":526262089,"identity":"eedf92ae-9d41-4f33-b15e-8a1e61006ff2","order_by":2,"name":"Zakariya Khaleel","email":"","orcid":"","institution":"Rutgers New Jersey Medical School","correspondingAuthor":false,"prefix":"","firstName":"Zakariya","middleName":"","lastName":"Khaleel","suffix":""},{"id":526262091,"identity":"72cf487a-0901-4743-9fe5-e37386c09dc8","order_by":3,"name":"Edwin Klein","email":"","orcid":"","institution":"University of Pittsburgh","correspondingAuthor":false,"prefix":"","firstName":"Edwin","middleName":"","lastName":"Klein","suffix":""},{"id":526262093,"identity":"21bb56e8-2f8b-41f6-804d-56e6fd163001","order_by":4,"name":"Robert Reiss","email":"","orcid":"","institution":"Rutgers New Jersey Medical School","correspondingAuthor":false,"prefix":"","firstName":"Robert","middleName":"","lastName":"Reiss","suffix":""},{"id":526262095,"identity":"ca68f8f5-bf36-4b2c-8a42-139adfe84988","order_by":5,"name":"Vincent Dong","email":"","orcid":"","institution":"University of Pennsylvania","correspondingAuthor":false,"prefix":"","firstName":"Vincent","middleName":"","lastName":"Dong","suffix":""},{"id":526262097,"identity":"97a4dbbc-800b-4f35-bed1-e443fc7013d3","order_by":6,"name":"H. Jacob Borish","email":"","orcid":"","institution":"University of Pittsburgh","correspondingAuthor":false,"prefix":"","firstName":"H.","middleName":"Jacob","lastName":"Borish","suffix":""},{"id":526262101,"identity":"76805ec0-e882-4434-9c9b-a532ec7be958","order_by":7,"name":"Hee-Jeong Yang","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Hee-Jeong","middleName":"","lastName":"Yang","suffix":""},{"id":526262103,"identity":"c2385d87-5cd0-43e2-afe4-c17cd6c8e0db","order_by":8,"name":"Danielle M. Weiner","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Danielle","middleName":"M.","lastName":"Weiner","suffix":""},{"id":526262105,"identity":"652832c5-1cae-41d6-af4b-05f936b2443e","order_by":9,"name":"Michelle Sutphin","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Michelle","middleName":"","lastName":"Sutphin","suffix":""},{"id":526262106,"identity":"ec22936d-4e5d-4b7a-ad38-9fb417cdf75b","order_by":10,"name":"Michaela K. Piazza","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Michaela","middleName":"K.","lastName":"Piazza","suffix":""},{"id":526262107,"identity":"dbe03e10-208a-46f7-a7a4-1b82964ed4b6","order_by":11,"name":"Emmanuel Dayao","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Emmanuel","middleName":"","lastName":"Dayao","suffix":""},{"id":526262108,"identity":"937c0677-69be-4145-941e-a3f8d8918455","order_by":12,"name":"Pauline Maiello","email":"","orcid":"","institution":"University of Pittsburgh","correspondingAuthor":false,"prefix":"","firstName":"Pauline","middleName":"","lastName":"Maiello","suffix":""},{"id":526262109,"identity":"94e32ae2-a217-4c02-9e5e-3ae9bb2072e1","order_by":13,"name":"Alexander White","email":"","orcid":"","institution":"University of Pittsburgh","correspondingAuthor":false,"prefix":"","firstName":"Alexander","middleName":"","lastName":"White","suffix":""},{"id":526262110,"identity":"c63767e5-83ff-4bbf-b890-828731b792ee","order_by":14,"name":"Charles A. Scanga","email":"","orcid":"","institution":"University of Pittsburgh","correspondingAuthor":false,"prefix":"","firstName":"Charles","middleName":"A.","lastName":"Scanga","suffix":""},{"id":526262111,"identity":"e42ad20b-bd3b-432b-8272-996950f1f14a","order_by":15,"name":"Philana Ling Lin","email":"","orcid":"","institution":"University of Pittsburgh School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Philana","middleName":"Ling","lastName":"Lin","suffix":""},{"id":526262112,"identity":"1f7b397f-313e-4101-953a-41007714f4ce","order_by":16,"name":"Clifton E. III Barry","email":"","orcid":"","institution":"National Institute of Allergy and Infectious Diseases","correspondingAuthor":false,"prefix":"","firstName":"Clifton","middleName":"E. III","lastName":"Barry","suffix":""},{"id":526262113,"identity":"f6573f35-d70a-410d-a804-ba2865e35acf","order_by":17,"name":"JoAnne Flynn","email":"","orcid":"","institution":"University of Pittsburgh","correspondingAuthor":false,"prefix":"","firstName":"JoAnne","middleName":"","lastName":"Flynn","suffix":""},{"id":526262114,"identity":"4da1a08a-cd5d-4e44-b12f-5dc390b7abee","order_by":18,"name":"Anant Madabhushi","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Anant","middleName":"","lastName":"Madabhushi","suffix":""},{"id":526262115,"identity":"bb1b07d6-2076-4449-a32e-c6e39c398686","order_by":19,"name":"Yingda Xie","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9ElEQVRIiWNgGAWjYDACdgaGAx8Y2FAFJfBqYWZgPDgDpuUAkVqYD/PAOERpMTjMY3DYdgdftHwD7zPpD79s8uUbmA/e5sGrhS3hcO4ZttwNB9jNJA72pVluOMCWbI1fC/OBw7ltQC0MbGwSB3sOGxgw8JhJ49fC2HDYEqhlfgNYy38D+Qb+bwS0AG1hBGppOADUcuDHAQOGAzxseLVIAv1ysBfksMNszBZnG5INgL4ztpyDRwvf8R7jDz/bjuXOb29jvFHxx85Avr354Y03eLQoHABTx0ARxMDA2MYAYeAD8g1gqgbK/UNA+SgYBaNgFIxIAACk90z/CCPvqAAAAABJRU5ErkJggg==","orcid":"","institution":"Rutgers New Jersey Medical School","correspondingAuthor":true,"prefix":"","firstName":"Yingda","middleName":"","lastName":"Xie","suffix":""}],"badges":[],"createdAt":"2025-06-04 17:53:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6822856/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6822856/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104340212,"identity":"c6a3dd8f-1fd2-4857-9d81-5e4f616f6f4a","added_by":"auto","created_at":"2026-03-10 16:36:23","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":613364,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow of image analysis, feature extraction and selection, and application of machine learning algorithms for radiomic prediction of TB lesion histopathology in marmosets and macaques. Performance of hierarchical classifier methods using train-test split for predicting the 4 TB lesion histopathologic classes shown in bottom right.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/e5dedc1e8a74c4946e413aae.png"},{"id":104405170,"identity":"b94e1289-58ea-4659-901e-c3624888ee97","added_by":"auto","created_at":"2026-03-11 12:21:59","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":529326,"visible":true,"origin":"","legend":"\u003cp\u003eHistology and descriptions of the four TB lesion classes: (1) Type 1: Necrotizing granulomas: These lesions demonstrate central necrosis with eosinophilic, amorphous acellular debris (caseum) +/- occasional stippled mineralization, concentrically surrounded by a mantle of epithelioid macrophages/multinucleated giant cells, and further peripherally marginated by lymphoplasmacytic inflammatory infiltrates. (2) Type 2: Non-necrotizing granulomas: Circumscribed aggregates of inflammatory cells (generally histiocytic). Demonstrate an architectural organization, although without a necrotic component and the various zones and cellular types associated therein, their architecture is generally less complex. (3) Type 3: Granulomatous inflammation outside of defined granuloma structure formation: Large area of coalescing or confluent necrotizing (caseous) granulomas or lesions in which histiocytic inflammation without frank granuloma structure extends in a direct, seemingly unimpeded fashion from alveolus-to-alveolus. This type of lesion is sometimes designated as tuberculous pneumonia. (4) Type 4: Fibrotic granulomas. Granulomatous architecture has been largely replaced by evolving dense fibrous connective tissue.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/4e70247858e2281bdede0c73.jpeg"},{"id":104340216,"identity":"3433055f-eb42-4392-8fa0-39a65d365782","added_by":"auto","created_at":"2026-03-10 16:36:23","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":269438,"visible":true,"origin":"","legend":"\u003cp\u003eStudy flowchart.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/9d9b08af6a0724e6b930d189.png"},{"id":104405542,"identity":"070f4598-900f-4e2f-98e8-726f869d9bfb","added_by":"auto","created_at":"2026-03-11 12:23:13","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":195572,"visible":true,"origin":"","legend":"\u003cp\u003eUnsupervised and supervised classification results for the Pitt dataset with (a) three-dimensional tSNE projection, (b) performance metrics for the LDA classifier based on LOO cross-validation of the training dataset D1 and (c) ROC curves and AUC values for the LDA classifier based on the validation dataset D2.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/df58ec7412913cd9a4145c23.png"},{"id":104340219,"identity":"bf390e43-a39d-4741-a0e6-0662357b9a6d","added_by":"auto","created_at":"2026-03-10 16:36:23","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":292408,"visible":true,"origin":"","legend":"\u003cp\u003eUnsupervised and supervised classification results for the NIH dataset with (a) three-dimensional tSNE projection, (b) performance metrics for the ADA-boost classifier based on LOO cross-validation of the training dataset D3 and (c) ROC curves and AUC values for the ADA-boost classifier based on the validation dataset D4.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/9ea4b5d14a1e844eb3638abf.jpeg"},{"id":104409393,"identity":"719dcc1a-d8ed-43fa-8890-0cc23e5bba35","added_by":"auto","created_at":"2026-03-11 12:44:58","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2617381,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/4e2c5814-1d8d-40a9-9ba3-c9f38a6bd0f8.pdf"},{"id":104340220,"identity":"4216b593-5fa2-43c2-8403-e89b718bb960","added_by":"auto","created_at":"2026-03-10 16:36:23","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":8533472,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialsTBradiomicmsv4.213May2025.docx","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/92f74347c5609fddabb625fe.docx"},{"id":104340221,"identity":"c82af890-7fd4-4777-97a8-c1a479c3c47c","added_by":"auto","created_at":"2026-03-10 16:36:25","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":58315670,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialsappendixS1IllustrationsandFurtherClarificationforRadiomicsGranulomaClassificationSystem.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/18b0cf3d626e32aa4754d54e.pdf"},{"id":104340213,"identity":"308ffea6-c9d0-44f1-aca6-104f4aa7768f","added_by":"auto","created_at":"2026-03-10 16:36:23","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":45713,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialsappendixS2SOPforsegmentation14May2025.docx","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/c307cc343d31c823e5f6867e.docx"},{"id":104340218,"identity":"3afb2703-9ead-4d6f-ae3f-a78091ab1976","added_by":"auto","created_at":"2026-03-10 16:36:23","extension":"csv","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":276076,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterialsappendixS3RadiologicFeaturesDataset.csv","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/5e71ede0fedbf52f3c63ad4b.csv"},{"id":104405411,"identity":"578c03f8-dc2e-473d-b337-dc532e8c36e4","added_by":"auto","created_at":"2026-03-11 12:22:49","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":16003,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementarymaterialsappendixS4Primatelists.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6822856/v1/98ed6b9a801dbb5cf3fbb0a5.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eRadiomic phenotyping of tuberculosis histopathology from computed tomography scans of non-human primates\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eTuberculosis (TB) is characterized by extensive heterogeneity at both the individual and lesion levels, spanning variations in clinical presentation, disease progression, and treatment response. This heterogeneity underpins differences in pharmacokinetics, bacterial metabolism and host immune response [\u003cspan additionalcitationids=\"CR2 CR3\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], yet remains poorly captured by conventional diagnostic tools such as sputum culture and molecular testing which offer limited insight into the spatial and biological complexity of in vivo disease [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. More precise characterization of this heterogeneity could enable earlier diagnosis, individualized risk stratification, and tailored treatment approaches [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eHigh resolution imaging modalities such as computed tomography (CT) provide a unique opportunity to characterize diverse lesion features in a non-invasive manner. Beyond simple visual interpretation, radiomics \u0026ndash; a high throughput approach to extract quantitative features from radiologic images \u0026ndash; enables systematic analysis of lesion texture, shape, and intensity patterns. Radiomic features have been applied in oncology to predict molecular and histopathologic attributes that distinguish benign and malignant lesions [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], stage and prognosticate tumors [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], and predict response to different therapies[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eRadiomics have been explored less extensively in infectious diseases, including a paucity of applications for characterizing heterogeneous TB lesion pathology that may be clinically informative. Early lesion features on CT and chest-X-rays, such as the presence or absence of fibrosis, have been linked to TB disease progression risk over the following months to years [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Additionally, lesion microenvironments influence critical determinates of treatment response [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], such as bacterial physiologic states and persistence [\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] and drug penetration [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. A high-resolution radiomic framework could provide a non-invasive biomarker of lesion phenotypes to guide research studies in improving TB treatment regimens and personalized clinical diagnoses.\u003c/p\u003e \u003cp\u003eIn this study, we explored whether CT-derived radiomic features could predict TB lesion histopathology. Given the absence of large-scale, lesion-level PET/CT-histopathology mapped datasets in humans, we leveraged two distinct non-human primate cohorts\u0026mdash;marmosets and macaques\u0026mdash;to represent different spectra of TB pathology and to assess the robustness of radiomic classification across primate species. By training and validating machine learning classifiers separately in each dataset, we aimed to explore the feasibility and generalizability of CT radiomics as a non-invasive tool for TB lesion phenotyping.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eThis project uses two existing databases of \u003cem\u003eMycobacterium tuberculosis\u003c/em\u003e (Mtb) infected non-human primate 2-deoxy-2-[\u0026sup1;⁸F] fluorodeoxyglucose positron emission tomography/computed tomography (PET/CT) scans from common marmosets at the NIAID Tuberculosis Research Section (\u0026lsquo;NIH\u0026rsquo;) and macaques at the Flynn lab at University of Pittsburgh (\u0026lsquo;Pitt\u0026rsquo;). All direct interactions with animals occurred in the context of other studies [\u003cspan additionalcitationids=\"CR20 CR21 CR22 CR23 CR24 CR25 CR26 CR27 CR28 CR29\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]; no interactions with animals occurred as part of this study. All necessary permissions were obtained for the use of archived data and biological samples included in this study.\u003c/p\u003e \u003cp\u003eCommon marmosets generally exhibit increased susceptibility to Mtb infection compared to macaques, with more rapid disease progression and pronounced pathology that recapitulate heterogeneity and complexity of lesions in human TB disease [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. On the other hand, macaques, particularly cynomolgus macaques, display a spectrum of disease outcomes from latent to active tuberculosis, closely mirroring the spectrum of human infection patterns [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. These differences are likely due to species-specific immunological differences[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Taken together, these datasets comprise scans across disease stages from early to advanced, lesion types from single granulomas to coalescent and cavitary, and treated and untreated disease. The workflow for the acquisition and segmentation of scans, feature extraction, feature selection and supervised and unsupervised classification analyses are provided in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eUntreated macaques (University of Pittsburgh)\u003c/h2\u003e \u003cp\u003ePitt samples were selected from an existing set of lesions from 29 unvaccinated and untreated macaques (27 cynomolgus and 2 rhesus, mostly males of Chinese origin) for 14 TB immunologic studies conducted between 2010 and 2018 [\u003cspan additionalcitationids=\"CR23 CR24 CR25 CR26 CR27 CR28 CR29\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. The macaques were previously infected intra-bronchially with 2.7\u0026ndash;85 colony forming units (CFU) of Mtb strain Erdman and allowed to progress for 12\u0026ndash;86 weeks prior to necropsy. CT scans were obtained within 3\u0026ndash;12 days prior to necropsy. Scans were performed using the NeuroLogica CereTom Clinical CT and Siemens MicroPET scanners, with imaging acquisition protocol as detailed in White A et al [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Images were exported as raw DICOM files. During necropsy, individual lung lesions were plated and identified using the pre-necropsy CT image. The plated lesions were used to prepare histologic sections consisting of formalin fixed tissues embedded in paraffin, sectioned, and stained with hematoxylin and eosin. In total, 314 histologic sections from the 29 animals in 14 studies mapped to CT were available and scored for histologic classification. All experiments on live vertebrates were approved by the University of Pittsburgh Institutional Animal Care and Use Committee (IACUC protocol numbers 15126588, 11110045, 11090030, 15055811, 15035407, 1105870, 14043492, 12060181, 17091258, 12080653 and 1003622). Experiments were performed and reported in accordance with the ARRIVE guidelines.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTreated marmosets (NIH)\u003c/h3\u003e\n\u003cp\u003eThe NIH samples consisted of lesions from 44 common marmosets (\u003cem\u003eCallithrix jacchus\u003c/em\u003e) exposed to a low dose of aerosolized Mtb strain H37Rv. The disease was allowed to progress for 7 to 8 weeks, and antibiotic treatment was given for 4 or 8 weeks. The animals were treated for one to two months with linezolid, one of five other investigational oxazolidinones, or rifampicin, isoniazid, pyrazinamide or moxifloxacin or a combination regimen of these four drugs to prevent mortality and achieve an infection time of at least 12 weeks so that fibrotic cuffs and other signs of an immunological response could occur. Some of the animals and lesions used in this analysis were also previously analyzed for drug-specific responses to treatment, although in those analyses only features such as lesion mean HU, mean SUV, total lesion glycolysis and the change in these features between time points were explored [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. To be eligible for inclusion in the dataset, the animals had a pre-necropsy scan within 48 hours of euthanasia, labeled histological samples preserved, and CFU/lesion determined [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Scans were also performed using the NeuroLogica CereTom Clinical CT and Siemens MicroPET scanners aligned using a common bed as with the Pitt scanner, with the same imaging acquisition protocol as Pitt detailed in White et al. [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] with the following exceptions: the NIH CT data was collected using 120 Kv, 5 mA/s, and the FDG PET scan began collection at 60 minutes post injection of the tracer as only two fields of view were necessary to cover the lung region of the marmoset in the Siemens Focus 220. After necropsy, there were 278 fixed hematoxylin and eosin-stained slides of classifiable TB lesions from the 44 animals studied. All live vertebrate procedures were performed in accordance with the recommendations of the Guide for the Care and Use of Laboratory Animals of the National Institutes of Health and approved by the NIAID Animal Care and Use Committee in Protocol LCIM-9 (Permit issued to NIH as A-4149-01). Both male and female marmosets ages 2 to 6 years old were included in the study.\u003c/p\u003e\n\u003ch3\u003eHistological classification of TB lesions\u003c/h3\u003e\n\u003cp\u003eAll available TB lesion histologic sections were classified using a classification protocol for histologic scoring developed by a veterinary pathologist to capture TB lesion phenotypes shared across human and nonhuman primates [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, Supplementary materials appendix S1). The histological classes include necrotizing granulomas (type 1), non-necrotizing granulomas (type 2), architecturally unstructured granulomatous inflammation outside of defined granuloma structure formation (type 3), and fibrotic granulomas (type 4).\u003c/p\u003e \u003cp\u003eArchived fixed or digital slides were reviewed by pathologists who scored lesions according to one of four categories in the manual. Slides which contain multiple lesion types were scored according to the predominant lesion type. Slides that could not be scored due to missing lesions or unclear labeling were marked as indeterminate.\u003c/p\u003e\n\u003ch3\u003eSegmentation of TB lesions from CT\u003c/h3\u003e\n\u003cp\u003eTB lesions were manually segmented from CT scans using Amira (Thermo Fisher Scientific version 6.5) to create binary mask region of interest (ROI) for each TB lesion (See supplementary materials for protocol). Segmentation was performed by a trained individual guided by two individuals performing quality control. Individuals performing segmentation were blinded to lesion classification. DICOM files of scan images from the NIH marmosets and Pitt macaques were imported into Amira and segmentation of the ROIs were performed as follows: (i) 3D ROIs were created by wrapping or interpolating between manually segmented cross sections, (ii) ROIs were refined by removing non-lesion tissue (e.g. large vasculature/airways, mediastinal structures, chest wall), (iii) Any material with a Hounsfield (HU) lower than \u0026minus;\u0026thinsp;500 attributed to aerated, normal lung tissue was removed from the ROI. This segmentation process was performed for each lesion following a standard operating procedure (see Supplementary materials appendix S2). The segmentation was reviewed for quality by a separate reader to ensure protocol adherence, and inclusion of only TB lesions within the ROIs.\u003c/p\u003e\n\u003ch3\u003eSeparation of validation set\u003c/h3\u003e\n\u003cp\u003eData from both sites were randomly split into training and validation sets (D1 and D2 for Pitt, D3 and D4 for NIH). The split was performed at the primate level to prevent data leakage, ensuring that lesions from the same primate did not appear in both training and validation sets. The number of lesions and primates in each dataset (D1\u0026ndash;D4) is detailed in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eRadiomic feature extraction\u003c/h2\u003e \u003cp\u003eA total of 74 first and second-order radiomic features (Supplementary table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e) were extracted from the lesion data using PyRadiomics, an open-source radiomic extraction python package [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The 74 3D features included: (i) six first-order statistics features to describe the basic distribution of voxel intensities within the identified lesion ROIs [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], (ii) 22 Haralick feature [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e] from a gray-level co-occurrence matrix that captures textural patterns that could encode variation in the different presentations of lesions and (iii) 46 additional texture features from the gray-level run length[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e] size zone [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. PET images had larger voxel sizes compared to CT images (~\u0026thinsp;10mm for PET vs. ~2mm for CT), resulting in imprecise alignment with the CT-based ROI. Consequently, feature extraction from PET was limited to SUVmax because SUVmax is less affected by the partial volume effect, where PET voxels extend beyond the smaller CT-based lesion ROIs.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eLesion size exclusion criteria\u003c/h3\u003e\n\u003cp\u003eThe two groups of primates in the study allowed us to observe biological variability in lesion representation on CT images. The lesion sizes differed significantly between the two groups. For each lesion, we identified the axial slice with the largest region of interest (ROI). The median value ROI area at this slice (TB lesion area) is 220 pixels in marmosets and 20 pixels in macaques.\u003c/p\u003e \u003cp\u003eSince our study employed handcrafted features, one of the standard methods for quantifying image textures involves the use of a kernel with a size of 5 \u0026times; 5 pixels [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. In contrast, Gray Level Co-occurrence Matrix (GLCM) features rely on the statistical distribution of pixel intensities within the region of interest (ROI). Therefore, ROIs that were less than 5 pixels were excluded in the primary analysis due to concern of poor reliability and representation of the computed features [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e].\u003c/p\u003e\n\u003ch3\u003eFeature selection\u003c/h3\u003e\n\u003cp\u003eTo enhance model generalizability and minimize the risk of overfitting, we performed radiomic feature selection as a pre-processing step. Feature selection was carried out using the Minimum Redundancy Maximum Relevance (mRMR) algorithm to ensure that the selected features had the maximum relevance to the target variable while minimizing redundancy. This approach facilitates the identification of a subset of features that contribute the most to prediction accuracy, while simultaneously reducing the complexity of the model[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. We evaluated the performance of the resulting models for the top 5 and the top 10 features.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eUnsupervised analysis\u003c/h2\u003e \u003cp\u003et-Distributed Stochastic Neighbor Embeddings (tSNE) of normalized training data using the top 10 selected features was used to visualize the topology of the N-dimensional clustering in a three-dimensional space [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. This was accomplished by constructing probability distributions such that similar high-dimensional samples have higher probabilities, and dissimilar samples have lower probabilities. These probabilities were then mapped onto a three-dimensional space for visualization. The resulting plots demonstrate the effectiveness of the selected feature representation in capturing the structure of the data in a reduced-dimensional space. tSNE clustering was performed in R, version 4.1.1, using the Rtsne package [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eSupervised analysis\u003c/h2\u003e \u003cp\u003eSupervised analysis was used to assess how well the model could generalize to new, unseen data and its ability to correctly identify the histologic class of a TB lesion. We used two different approaches to distinguish between multiple classes. The first approach employs a multi-class (MC) classification strategy [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e], where a single model is trained to classify the input data into multiple classes. In this strategy, the classifier distinguishes quantitative features into distinct histological types, with each type encoded by a unique class label (e.g., histological type 1 is encoded as class 1, type 2 as class 2, etc.).\u003c/p\u003e \u003cp\u003eThe second approach follows the \"one-versus-all\" (OvA) strategy [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. Unlike the multi-class method, this strategy involves designing an individual classification model for each histological type. For instance, one model is dedicated to classifying histological type 1\u0026mdash;where histological type 1 is encoded as class 1 and all other types as class 0\u0026mdash;and this process is repeated for every histological type. Here we focus on the results from the OvA approach. Results from the MC approach are presented in the supplementary materials.\u003c/p\u003e \u003cp\u003eTwo training strategies were employed: Leave-One-Out (LOO) cross-validation and standard training with a held-out validation set. The LOO method evaluates model performance and robustness by iteratively using every sample as a test case once, while the remaining data is used to train the model. This approach provides an unbiased accuracy estimate and maximizes data utility\u0026mdash;particularly valuable for smaller datasets. Only the training datasets (D1 for Pittsburgh data and D3 for NIH) were used in LOO to maintain consistency with the validation split. LOO cross-validation was implemented as follows:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eIn each iteration, one sample was held out as test data, while the rest trained the model.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThe mRMR algorithm dynamically selected the top 5 and top 10 most discriminative features per iteration, ensuring adaptive feature filtering.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThis process was repeated across all samples, systematically assessing which texture features best differentiated lesion types.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eBy isolating test data in every iteration, LOO mitigates overfitting and offers a rigorous evaluation of feature generalizability.\u003c/p\u003e \u003cp\u003eFollowing the Leave-One-Out (LOO) cross-validation performed on the training datasets (D1 for Pitt and D3 for NIH), we proceeded to train final diagnostic models using the complete D1 and D3 datasets. These models incorporated the most discriminative features identified during the LOO analysis. The trained models were then evaluated on the separate validation datasets (D2 for Pitt and D4 for NIH) to assess their diagnostic performance. This approach ensured that the validation was performed on entirely independent data, preventing any information leakage and providing a reliable measure of the models' generalization capability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eFor both the supervised and unsupervised analysis methods, the following algorithms were used for model training: random forests (RF), support vector machine (SVM), adaptive boosting (Ada-Boost), and linear discriminant analysis (LDA). Results were quantified by calculating overall accuracy, weighted accuracy, and the area under the receiving operating characteristic curve (ROC-AUC) for each trained classifier, calculating these metrics on an OvA basis.\u003c/p\u003e \u003cp\u003eThe methodology described above, including LOO cross-validation and training with validation, was implemented and examined for two feature selection setups: selecting the top 5 and the top 10 best features. The top 10 features were chosen for the primary analysis because they gave the best performance for the training dataset in terms of AUC for the supervised models. Results of the analysis using the top 5 features are presented in the supplementary materials.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eAmong 556 segmented lesions from 29 Pitt macaques and 488 segmented lesions from 44 NIH marmosets annotated on PET/CT, 151 lesions from 29 macaques and 149 lesions from 41 marmosets matched with associated histopathology class labels and met the lesion ROI size criteria (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Among these, 65 lesions from 15 marmosets and 27 lesions from 5 macaques were randomly set aside in an independent dataset (D2 and D4 respectively). The remaining eligible dataset was assigned to datasets D1 and D3.\u003c/p\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003ePitt dataset\u003c/h2\u003e \u003cp\u003eThe 150 untreated macaque lesions in the Pitt dataset used in the analysis were compromised of 76 (50.7%) type 1 lesions, 29 (19.3%) type 2 lesions, and 45 (30.0%) type 4 lesions. A single type 3 lesion representing granulomatous inflammation outside of the defined granuloma structure formation was removed from the analysis due to the paucity of type 3 lesions leading to severe class imbalance. Class imbalance in machine learning occurs when the distribution of samples across different classes is highly uneven, meaning some classes have significantly fewer examples than others [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. This imbalance can harm model performance by causing the algorithm to become biased toward the majority class [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e], leading to poor predictive accuracy for the underrepresented (but often critical) minority class, such as rare diseases in medical diagnosis or fraud cases in financial transactions. Therefore, no judgments can be made about separability of type 3 lesions in the Pitt dataset.\u003c/p\u003e \u003cp\u003eSelected features from the training/test set are shown in Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e. Experiments were done with both the OvA and the MC classification approach, with similar results achieved using both approaches. Unsupervised tSNE clustering results of the training dataset are shown (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). The LDA model was selected as the best-performing model for the leave-one-out approach, with an average AUC of 0.66 and a balanced accuracy of 52.76% (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003eApplied to the independent validation set, the LDA model yielded AUCs of 0.52 (95% CI\u0026thinsp;=\u0026thinsp;0.26\u0026ndash;0.77) for type 1 lesions, 0.69 (95% CI\u0026thinsp;=\u0026thinsp;0.48\u0026ndash;0.88) for type 2 lesions, and 0.91 (95% CI\u0026thinsp;=\u0026thinsp;0.77\u0026ndash;1.00) for type 4 lesions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eNIH dataset\u003c/h2\u003e \u003cp\u003eThe 149 treated marmoset lesions in the NIH dataset were comprised of 31 (20.8%) type 1 lesions, 28 (18.8%) type 2 lesions, 60 (40.3%) type 3 lesions, and 30 (20.1%) type 4 lesions. Selected features are shown in Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e. Like the analysis of the Pitt dataset, the OvA and MC approaches yielded comparable results. Unsupervised tSNE clustering results of the training dataset are shown (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea). Supervised classification results demonstrate a higher balanced accuracy but lower AUC compared to the results from the Pitt dataset (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ec). The Ada-boost model was selected as the best-performing model for the training set, with an average AUC of 0.70 and a balanced accuracy of 38.8% (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb). Applied to the validation set, the Ada-boost model yielded AUCs of 0.47 (95% CI\u0026thinsp;=\u0026thinsp;0.29\u0026ndash;0.65) for type 1 lesions, 0.58 (95% CI\u0026thinsp;=\u0026thinsp;0.38\u0026ndash;0.78) for type 2 lesions, 0.70 (95% CI\u0026thinsp;=\u0026thinsp;0.56\u0026ndash;0.83) for type 3 lesions, and 0.65 (95% CI\u0026thinsp;=\u0026thinsp;0.50\u0026ndash;0.79) for type 4 lesions.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eWe applied radiomic feature extraction to identify shape- and texture-based characteristics that differentiate TB lesion histopathologic characteristics across two primate species, with and without treatment. Radiomic models have been widely used in cancer research and more recently in TB diagnosis, including to distinguish TB from other respiratory conditions [\u003cspan class=\"CitationRef\"\u003e49\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e58\u003c/span\u003e] and predict drug resistance [\u003cspan class=\"CitationRef\"\u003e59\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e60\u003c/span\u003e], but have not looked at the resolution of histopathology. To our knowledge, this is the first application of radiomics to discriminate TB lesion pathology at a microscopic level. Using these extracted features and existing experimental data from a limited cohort of untreated macaques and treated marmosets, we trained machine learning models that exhibited modest performance in classifying TB lesion histology types among a small and biologically diverse non-human primate dataset, supporting the feasibility and trainability of this approach with further refinement on existing datasets. Notably, the models distinguished fibrotic from non-fibrotic granulomas with high fidelity in untreated macaques, highlighting the potential for radiomics to resolve biologically meaningful lesion types that may relate to clinical risk.\u003c/p\u003e\n\u003cp\u003eCompared to macaques, marmosets exhibited larger and more extensive lesions. Model performance metrics were generally consistent across species, feature sets, and classifier types, suggesting that the approach does not require extensive fine-tuning to yield stable results. However, performance varied by lesion type in the independent validation sets. Specifically, necrotizing granulomas (Type I lesions) were the most difficult to classify, with AUC ROC values close to 0.5 in both datasets, and were often misclassified as Type IV lesions\u0026mdash;or as Type III lesions in the NIH dataset. In contrast, Type IV lesion predictions were more accurate in untreated macaques, with an AUC ROC of 0.9 [95% CI 0.77\u0026ndash;1.00], than in treated marmosets, where the AUC ROC dropped to 0.47 [95% CI 0.29\u0026ndash;0.65]. This discrepancy may reflect treatment-induced changes in lesion phenotype among marmosets, which could obscure distinctions present in untreated disease. Treated animals may also have shown a greater prevalence of mixed histopathologic features, which were categorized based on the predominant phenotype. In clinical imaging surveys of chest X-ray and PET-CT among individuals with risks for TB exposure, non-fibrotic (\u0026ldquo;active\u0026rdquo;) TB-like lesions were several times more likely to progress to active TB than fibrotic (\u0026ldquo;inactive\u0026rdquo;) lesions [\u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e13\u003c/span\u003e]. Therefore, the ability to distinguish fibrotic (Type IV) lesions from non-fibrotic types in untreated individuals may become a valuable biomarker for predicting and characterizing early TB before individuals develop illness or transmit to others. Notably, recent and ongoing studies are now applying CT and chest-X-ray imaging in high-risk human populations to identify early TB, offering an opportunity to test and refine this classifier while informing training of new clinical models of TB progression risk.\u003c/p\u003e\n\u003cp\u003eInterestingly, SUVmax was not among the top-selected features in either dataset, indicating that classification was driven by CT-based radiomics rather than PET metabolic activity. The contribution of SUVmax may have also been limited by the considerably lower PET resolution compared to the CT-based ROI. Rather, features associated with structural organization and gray-level intensity variations were more predictive, reflecting the structural complexity of lesions. Features such as Small Area High Gray Level Emphasis were higher in inflammation beyond granulomas (Type 3; marmosets), potentially reflecting greater fragmentation and high-intensity structures with dense cellular infiltration, while fibrotic granulomas had lower values potentially reflecting more uniform collagen deposition. In both datasets, necrotizing granulomas (Type 1) and fibrotic granulomas (Type 4) also tended to have higher Contrast values, potentially reflecting sharp intensity transitions in necrotic cores and fibrotic lesions, whereas inflammation beyond granulomas (Type 3; marmosets) tended to have lower contrast, potentially due to more diffuse inflammatory involvement. These findings suggest that radiomic features, particularly those related to small-scale heterogeneity and intensity transitions, may be relevant in distinguishing TB pathology across species and treatment conditions.\u003c/p\u003e\n\u003cp\u003eEarly data from large autopsy studies have identified granulomas with culturable Mtb among individuals who died from non-TB causes, suggesting a prevalence and range of clinically-occult pathology that may range from self-resolution to requiring various preventive interventions [\u003cspan class=\"CitationRef\"\u003e13\u003c/span\u003e]. The ability to phenotype lesions in-vivo would provide an opportunity to characterize the range of minimal TB disease in contacts or high-risk individuals that herald a likelihood of progression to bacteriologically positive disease. Additionally, in-vivo lesion phenotyping may also provide a useful tool to personalize treatment approaches and design host-directed or anti-microbial drug regimens effective across the range of lesion microenvironments [\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e]. This high-resolution imaging-based treatment response tool could apply to sputum paucibacillary and drug-resistant TB where treatment monitoring methods are limited or critical.\u003c/p\u003e\n\u003cp\u003eThere were several limitations of the study. Our decision not to analyze a combined NIH and Pitt dataset due to the presence of batch effects highlights an important limitation: the differences in acquisition related factors can have a substantial impact on the corresponding radiomic features, even in CT scans[\u003cspan class=\"CitationRef\"\u003e61\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e62\u003c/span\u003e]. Therefore, models were separately generated for each dataset. Additionally, the study used existing datasets collected from different facilities, primate species, and infection/treatment protocols for different primary objectives, conditions leading to substantial heterogeneity and reduced sample sizes for each group in part because not all observed lesions in the PET/CT scans were captured in a corresponding histological section. As a result of the small sample sizes, the analysis was restricted to traditional classification methods. Deep learning approaches such as convolutional neural networks or vision transformers were not explored in this study. Acquiring larger datasets may allow for the use of deep learning models to validate our results. Future work to explore methods of normalizing the extracted features across imaging acquisition sites such as ComBat batch effect adjustment [\u003cspan class=\"CitationRef\"\u003e63\u003c/span\u003e] are also relevant to verify these findings. Nevertheless, these limitations may be viewed as a representation of the diversity of human pathology from early to later disease stages and supports the ability of radiomic classifiers to profile TB lesions to microscopic resolutions among vastly different disease states. Finally, a substantial proportion of the macaque lesions were small (\u0026lt;\u0026thinsp;5 voxels in the largest slice) which did not meet the run length of several radiomic features and were thus excluded. Therefore, in very small lesions (e.g., early lesions or certain animal models), there may be a limit to what the radiomic models can reliability measure.\u003c/p\u003e\n\u003cp\u003eAltogether, these findings highlight the potential of radiomics to non-invasively classify histologic TB lesion types, with consistent performance for fibrotic lesions and modest results for others. With expanding access to annotated experimental and clinical datasets, model refinement may improve classification of other lesion types and advance lesion-level phenotyping tools for clinical and research applications in TB pathogenesis, treatment strategies, and detection of early TB and TB across the disease spectrum.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll extracted radiomic feature data generated and/or analyzed during this study are included in this published article (and its Supplementary Materials files). PET/CT scans are available from the corresponding author on reasonable request.\u003cbr\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFunding was provided in part by the Bill and Melinda Gates Foundation through OPP053284 (PLL), OPP1162695 and OPP1024021 (CEB), OPP1034408 and INV020435 (JF), Grand Challenges – Annual Meeting “Call to Action” (LX, JF), and in part by the Division of Intermural Research, NIAID, NIH (CEB and LEV). We thank the Comparative Medicine Branch of NIAID, NIH for clinical care of the marmosets and appreciate the technical expertise of Emmanual Dayao DVM, and Becky Sloan, and of NIAID.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEK, LEV, HY performed lesion analysis. LEV, DMW, MS, MP, ED, HJB, AW, PM, CS, PLL collected the PET/CT studies, labeled scans, catalogued lesions, and performed necropsies. DC and VD performed the machine learning analysis under the supervision of AM. ZK performed segmentations, data cleaning, data analysis and manuscript writing. RR performed data analysis \u0026nbsp; and manuscript writing. JF and CEB supervised the non-human primate studies and helped conceptualize the original study design. \u0026nbsp;YLX conceptualized and supervised the study design, analyses, and manuscript writing.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eCadena, A.M., S.M. Fortune, and J.L. Flynn, \u003cem\u003eHeterogeneity in tuberculosis.\u003c/em\u003e Nat Rev Immunol, 2017. \u003cstrong\u003e17\u003c/strong\u003e(11): p. 691-702.\u003c/li\u003e\n \u003cli\u003eDhar, N., J. McKinney, and G. Manina, \u003cem\u003ePhenotypic Heterogeneity in Mycobacterium tuberculosis.\u003c/em\u003e Microbiol Spectr, 2016. \u003cstrong\u003e4\u003c/strong\u003e(6).\u003c/li\u003e\n \u003cli\u003eLenaerts, A., C.E. Barry Iii, and V. Dartois, \u003cem\u003eHeterogeneity in tuberculosis pathology, microenvironments and therapeutic responses.\u003c/em\u003e Immunological Reviews, 2015. \u003cstrong\u003e264\u003c/strong\u003e(1): p. 288-307.\u003c/li\u003e\n \u003cli\u003eLin, P.L., et al., \u003cem\u003eRadiologic Responses in Cynomolgus Macaques for Assessing Tuberculosis Chemotherapy Regimens.\u003c/em\u003e Antimicrob Agents Chemother, 2013. \u003cstrong\u003e57\u003c/strong\u003e(9): p. 4237-4244.\u003c/li\u003e\n \u003cli\u003eChen, X. and T.Y. Hu, \u003cem\u003eStrategies for advanced personalized tuberculosis diagnosis: Current technologies and clinical approaches.\u003c/em\u003e Precis Clin Med, 2021. \u003cstrong\u003e4\u003c/strong\u003e(1): p. 35-44.\u003c/li\u003e\n \u003cli\u003eXie, Y.L., et al., \u003cem\u003eFourteen-day PET/CT imaging to monitor drug combination activity in treated individuals with tuberculosis.\u003c/em\u003e Sci Transl Med, 2021. \u003cstrong\u003e13\u003c/strong\u003e(579).\u003c/li\u003e\n \u003cli\u003eWalter, N.D., et al., \u003cem\u003eLung microenvironments harbor Mycobacterium tuberculosis phenotypes with distinct treatment responses.\u003c/em\u003e Antimicrob Agents Chemother, 2023. \u003cstrong\u003e67\u003c/strong\u003e(9): p. e0028423.\u003c/li\u003e\n \u003cli\u003eGillies, R.J., P.E. Kinahan, and H. Hricak, \u003cem\u003eRadiomics: Images Are More than Pictures, They Are Data.\u003c/em\u003e Radiology, 2016. \u003cstrong\u003e278\u003c/strong\u003e(2): p. 563-77.\u003c/li\u003e\n \u003cli\u003eAerts, H.J.W.L., et al., \u003cem\u003eDecoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach.\u003c/em\u003e Nature Communications, 2014. \u003cstrong\u003e5\u003c/strong\u003e(1): p. 4006.\u003c/li\u003e\n \u003cli\u003eJain, P., et al., \u003cem\u003eNovel Non-Invasive Radiomic Signature on CT Scans Predicts Response to Platinum-Based Chemotherapy and Is Prognostic of Overall Survival in Small Cell Lung Cancer.\u003c/em\u003e Front Oncol, 2021. \u003cstrong\u003e11\u003c/strong\u003e: p. 744724.\u003c/li\u003e\n \u003cli\u003eBera, K., et al., \u003cem\u003ePredicting cancer outcomes with radiomics and artificial intelligence in radiology.\u003c/em\u003e Nat Rev Clin Oncol, 2022. \u003cstrong\u003e19\u003c/strong\u003e(2): p. 132-146.\u003c/li\u003e\n \u003cli\u003eEsmail, H., et al., \u003cem\u003eHigh resolution imaging and five-year tuberculosis contact outcomes.\u003c/em\u003e medRxiv, 2023.\u003c/li\u003e\n \u003cli\u003eSossen, B., et al., \u003cem\u003eThe natural history of untreated pulmonary tuberculosis in adults: a systematic review and meta-analysis.\u003c/em\u003e The Lancet Respiratory Medicine, 2023. \u003cstrong\u003e11\u003c/strong\u003e(4): p. 367-379.\u003c/li\u003e\n \u003cli\u003eLavin, R.C. and S. Tan, \u003cem\u003eSpatial relationships of intra-lesion heterogeneity in Mycobacterium tuberculosis microenvironment, replication status, and drug efficacy.\u003c/em\u003e PLoS Pathog, 2022. \u003cstrong\u003e18\u003c/strong\u003e(3): p. e1010459.\u003c/li\u003e\n \u003cli\u003eGold, B. and C. Nathan, \u003cem\u003eTargeting Phenotypically Tolerant Mycobacterium tuberculosis.\u003c/em\u003e Microbiol Spectr, 2017. \u003cstrong\u003e5\u003c/strong\u003e(1).\u003c/li\u003e\n \u003cli\u003eBarry, C.E., et al., \u003cem\u003eThe spectrum of latent tuberculosis: rethinking the biology and intervention strategies.\u003c/em\u003e Nature Reviews Microbiology, 2009. \u003cstrong\u003e7\u003c/strong\u003e(12): p. 845-855.\u003c/li\u003e\n \u003cli\u003ePrideaux, B., et al., \u003cem\u003eThe association between sterilizing activity and drug distribution into tuberculosis lesions.\u003c/em\u003e Nature Medicine, 2015. \u003cstrong\u003e21\u003c/strong\u003e(10): p. 1223-1227.\u003c/li\u003e\n \u003cli\u003eDartois, V., \u003cem\u003eThe path of anti-tuberculosis drugs: from blood to lesions to mycobacterial cells.\u003c/em\u003e Nature Reviews Microbiology, 2014. \u003cstrong\u003e12\u003c/strong\u003e(3): p. 159-167.\u003c/li\u003e\n \u003cli\u003eBudak, M., et al., \u003cem\u003eA systematic efficacy analysis of tuberculosis treatment with BPaL-containing regimens using a multiscale modeling approach.\u003c/em\u003e CPT Pharmacometrics Syst Pharmacol, 2024. \u003cstrong\u003e13\u003c/strong\u003e(4): p. 673-685.\u003c/li\u003e\n \u003cli\u003eBoshoff, H.I.M., et al., \u003cem\u003eMtb-Selective 5-Aminomethyl Oxazolidinone Prodrugs: Robust Potency and Potential Liabilities.\u003c/em\u003e ACS Infectious Diseases, 2024. \u003cstrong\u003e10\u003c/strong\u003e(5): p. 1679-1695.\u003c/li\u003e\n \u003cli\u003eGreenstein T, V.L., Moraes MP, Weiner DM, et al., \u003cem\u003ePET/CT multivariate tuberculosis treatment response profiles in marmosets unify disparate preclinical biomarkers.\u003c/em\u003e Sci Transl Med, 2025.\u003c/li\u003e\n \u003cli\u003eMaiello, P., et al., \u003cem\u003eRhesus Macaques Are More Susceptible to Progressive Tuberculosis than Cynomolgus Macaques: a Quantitative Comparison.\u003c/em\u003e Infect Immun, 2018. \u003cstrong\u003e86\u003c/strong\u003e(2).\u003c/li\u003e\n \u003cli\u003eGanchua, S.K.C., et al., \u003cem\u003eLymph nodes are sites of prolonged bacterial persistence during Mycobacterium tuberculosis infection in macaques.\u003c/em\u003e PLoS Pathog, 2018. \u003cstrong\u003e14\u003c/strong\u003e(11): p. e1007337.\u003c/li\u003e\n \u003cli\u003eLin, P.L., et al., \u003cem\u003ePET CT Identifies Reactivation Risk in Cynomolgus Macaques with Latent M. tuberculosis.\u003c/em\u003e PLoS Pathog, 2016. \u003cstrong\u003e12\u003c/strong\u003e(7): p. e1005739.\u003c/li\u003e\n \u003cli\u003eMedrano, J.M., et al., \u003cem\u003eCharacterizing the Spectrum of Latent Mycobacterium tuberculosis in the Cynomolgus Macaque Model: Clinical, Immunologic, and Imaging Features of Evolution.\u003c/em\u003e J Infect Dis, 2023. \u003cstrong\u003e227\u003c/strong\u003e(4): p. 592-601.\u003c/li\u003e\n \u003cli\u003eColeman, M.T., et al., \u003cem\u003eEarly Changes by (18)Fluorodeoxyglucose positron emission tomography coregistered with computed tomography predict outcome after Mycobacterium tuberculosis infection in cynomolgus macaques.\u003c/em\u003e Infect Immun, 2014. \u003cstrong\u003e82\u003c/strong\u003e(6): p. 2400-4.\u003c/li\u003e\n \u003cli\u003eGideon, H.P., et al., \u003cem\u003eVariability in tuberculosis granuloma T cell responses exists, but a balance of pro- and anti-inflammatory cytokines is associated with sterilization.\u003c/em\u003e PLoS Pathog, 2015. \u003cstrong\u003e11\u003c/strong\u003e(1): p. e1004603.\u003c/li\u003e\n \u003cli\u003eRodgers, M.A., et al., \u003cem\u003ePreexisting Simian Immunodeficiency Virus Infection Increases Susceptibility to Tuberculosis in Mauritian Cynomolgus Macaques.\u003c/em\u003e Infect Immun, 2018. \u003cstrong\u003e86\u003c/strong\u003e(12).\u003c/li\u003e\n \u003cli\u003eLin, P.L., et al., \u003cem\u003eSterilization of granulomas is common in active and latent tuberculosis despite within-host variability in bacterial killing.\u003c/em\u003e Nat Med, 2014. \u003cstrong\u003e20\u003c/strong\u003e(1): p. 75-9.\u003c/li\u003e\n \u003cli\u003eDarrah, P.A., et al., \u003cem\u003ePrevention of tuberculosis in macaques after intravenous BCG immunization.\u003c/em\u003e Nature, 2020. \u003cstrong\u003e577\u003c/strong\u003e(7788): p. 95-102.\u003c/li\u003e\n \u003cli\u003eVia Laura, E., et al., \u003cem\u003eDifferential Virulence and Disease Progression following Mycobacterium tuberculosis Complex Infection of the Common Marmoset (Callithrix jacchus).\u003c/em\u003e Infection and Immunity, 2013. \u003cstrong\u003e81\u003c/strong\u003e(8): p. 2909-2919.\u003c/li\u003e\n \u003cli\u003eScanga, C.A. and J.L. Flynn, \u003cem\u003eModeling tuberculosis in nonhuman primates.\u003c/em\u003e Cold Spring Harb Perspect Med, 2014. \u003cstrong\u003e4\u003c/strong\u003e(12): p. a018564.\u003c/li\u003e\n \u003cli\u003ePe\u0026ntilde;a Juliet, C. and W.-Z. Ho, \u003cem\u003eNon-Human Primate Models of Tuberculosis.\u003c/em\u003e Microbiology Spectrum, 2016. \u003cstrong\u003e4\u003c/strong\u003e(4): p. 10.1128/microbiolspec.tbtb2-0007-2016.\u003c/li\u003e\n \u003cli\u003eWhite, A.G., et al., \u003cem\u003eAnalysis of 18FDG PET/CT Imaging as a Tool for Studying Mycobacterium tuberculosis Infection and Treatment in Non-human Primates.\u003c/em\u003e J Vis Exp, 2017(127).\u003c/li\u003e\n \u003cli\u003eLeong, F.J., Dartois, V., \u0026amp; Dick, T. (Eds.), \u003cem\u003eA Color Atlas of Comparative Pathology of Pulmonary Tuberculosis (1st ed.)\u003c/em\u003e. 2010: CRC Press.\u003c/li\u003e\n \u003cli\u003evan Griethuysen, J.J.M., et al., \u003cem\u003eComputational Radiomics System to Decode the Radiographic Phenotype.\u003c/em\u003e Cancer Research, 2017. \u003cstrong\u003e77\u003c/strong\u003e(21): p. e104-e107.\u003c/li\u003e\n \u003cli\u003eHaralick, R.M., K. Shanmugam, and I. Dinstein, \u003cem\u003eTextural Features for Image Classification.\u003c/em\u003e IEEE Transactions on Systems, Man, and Cybernetics, 1973. \u003cstrong\u003eSMC-3\u003c/strong\u003e(6): p. 610-621.\u003c/li\u003e\n \u003cli\u003eGalloway, M.M., \u003cem\u003eTexture analysis using gray level run lengths.\u003c/em\u003e Computer Graphics and Image Processing, 1975. \u003cstrong\u003e4\u003c/strong\u003e(2): p. 172-179.\u003c/li\u003e\n \u003cli\u003eChu, A., C.M. Sehgal, and J.F. Greenleaf, \u003cem\u003eUse of gray value distribution of run lengths for texture analysis.\u003c/em\u003e Pattern Recognition Letters, 1990. \u003cstrong\u003e11\u003c/strong\u003e(6): p. 415-419.\u003c/li\u003e\n \u003cli\u003eThibault, G., et al., \u003cem\u003eTexture Indexes and Gray Level Size Zone Matrix Application to Cell Nuclei Classification\u003c/em\u003e, in \u003cem\u003e10th International Conference on Pattern Recognition and Information Processing\u003c/em\u003e. 2009.\u003c/li\u003e\n \u003cli\u003eJensen, L.J., et al. \u003cem\u003eStability of Radiomic Features across Different Region of Interest Sizes\u0026mdash;A CT and MR Phantom Study\u003c/em\u003e. Tomography, 2021. \u003cstrong\u003e7\u003c/strong\u003e, 238-252 DOI: 10.3390/tomography7020022.\u003c/li\u003e\n \u003cli\u003evan Timmeren, J.E., et al., \u003cem\u003eFeature selection methodology for longitudinal cone-beam CT radiomics.\u003c/em\u003e Acta Oncol, 2017. \u003cstrong\u003e56\u003c/strong\u003e(11): p. 1537-1543.\u003c/li\u003e\n \u003cli\u003eVan der Maaten, L. and G. Hinton, \u003cem\u003eVisualizing data using t-SNE.\u003c/em\u003e Journal of machine learning research, 2008. \u003cstrong\u003e9\u003c/strong\u003e(11).\u003c/li\u003e\n \u003cli\u003eKrijthe, J.H., \u003cem\u003eRtsne: T-distributed Stochastic Neighbor Embedding using Barnes-Hut Implementation\u003c/em\u003e. 2015.\u003c/li\u003e\n \u003cli\u003eAly, M., \u003cem\u003eSurvey on multiclass classification methods.\u003c/em\u003e Neural Netw, 2005. \u003cstrong\u003e19\u003c/strong\u003e(1-9): p. 2.\u003c/li\u003e\n \u003cli\u003eRifkin, R. and A. Klautau, \u003cem\u003eIn defense of one-vs-all classification.\u003c/em\u003e Journal of machine learning research, 2004. \u003cstrong\u003e5\u003c/strong\u003e(Jan): p. 101-141.\u003c/li\u003e\n \u003cli\u003eGuo, X., et al. \u003cem\u003eOn the Class Imbalance Problem\u003c/em\u003e. in \u003cem\u003e2008 Fourth International Conference on Natural Computation\u003c/em\u003e. 2008.\u003c/li\u003e\n \u003cli\u003eJapkowicz, N. and S. Stephen, \u003cem\u003eThe class imbalance problem: A systematic study.\u003c/em\u003e Intell. Data Anal., 2002. \u003cstrong\u003e6\u003c/strong\u003e(5): p. 429\u0026ndash;449.\u003c/li\u003e\n \u003cli\u003eLi, P., et al., \u003cem\u003eA CT-based radiomics predictive nomogram to identify pulmonary tuberculosis from community-acquired pneumonia: a multicenter cohort study.\u003c/em\u003e Frontiers in Cellular and Infection Microbiology, 2024. \u003cstrong\u003e14\u003c/strong\u003e.\u003c/li\u003e\n \u003cli\u003eZhang, X., et al., \u003cem\u003eDeep learning PET/CT-based radiomics integrates clinical data: A feasibility study to distinguish between tuberculosis nodules and lung cancer.\u003c/em\u003e Thorac Cancer, 2023. \u003cstrong\u003e14\u003c/strong\u003e(19): p. 1802-1811.\u003c/li\u003e\n \u003cli\u003eLi, Y., et al., \u003cem\u003eMachine learning-based radiomics to distinguish pulmonary nodules between lung adenocarcinoma and tuberculosis.\u003c/em\u003e Thoracic Cancer, 2024. \u003cstrong\u003e15\u003c/strong\u003e(6): p. 466-476.\u003c/li\u003e\n \u003cli\u003eThattaamuriyil Padmakumari, L., et al., \u003cem\u003eThe Role of Chest CT Radiomics in Diagnosis of Lung Cancer or Tuberculosis: A Pilot Study.\u003c/em\u003e Diagnostics, 2022. \u003cstrong\u003e12\u003c/strong\u003e(3): p. 739.\u003c/li\u003e\n \u003cli\u003eHu, Y., et al., \u003cem\u003eLung CT-based multi-lesion radiomic model to differentiate between nontuberculous mycobacteria and Mycobacterium tuberculosis.\u003c/em\u003e Med Phys, 2025. \u003cstrong\u003e52\u003c/strong\u003e(2): p. 1086-1095.\u003c/li\u003e\n \u003cli\u003eLi, H.L., et al., \u003cem\u003eMultimodal machine learning-based model for differentiating nontuberculous mycobacteria from mycobacterium tuberculosis.\u003c/em\u003e Front Public Health, 2025. \u003cstrong\u003e13\u003c/strong\u003e: p. 1470072.\u003c/li\u003e\n \u003cli\u003eLi, P., et al., \u003cem\u003eA CT-based radiomics predictive nomogram to identify pulmonary tuberculosis from community-acquired pneumonia: a multicenter cohort study.\u003c/em\u003e Front Cell Infect Microbiol, 2024. \u003cstrong\u003e14\u003c/strong\u003e: p. 1388991.\u003c/li\u003e\n \u003cli\u003eZhou, L., et al., \u003cem\u003eA retrospective study differentiating nontuberculous mycobacterial pulmonary disease from pulmonary tuberculosis on computed tomography using radiomics and machine learning algorithms.\u003c/em\u003e Ann Med, 2024. \u003cstrong\u003e56\u003c/strong\u003e(1): p. 2401613.\u003c/li\u003e\n \u003cli\u003eYan, Q., et al., \u003cem\u003eCT\u003c/em\u003e\u003cem\u003e‑based radiomics analysis of consolidation characteristics in differentiating pulmonary disease of non\u003c/em\u003e\u003cem\u003e‑tuberculous mycobacterium from pulmonary tuberculosis.\u003c/em\u003e Exp Ther Med, 2024. \u003cstrong\u003e27\u003c/strong\u003e(3): p. 112.\u003c/li\u003e\n \u003cli\u003eWei, S., et al., \u003cem\u003eDifferentiating mass-like tuberculosis from lung cancer based on radiomics and CT features.\u003c/em\u003e Translational Cancer Research, 2021. \u003cstrong\u003e10\u003c/strong\u003e(10): p. 4454-4463.\u003c/li\u003e\n \u003cli\u003eJiang, F., et al., \u003cem\u003eA CT-based radiomics analyses for differentiating drug\u003c/em\u003e\u003cem\u003e‑resistant and drug-sensitive pulmonary tuberculosis.\u003c/em\u003e BMC Med Imaging, 2024. \u003cstrong\u003e24\u003c/strong\u003e(1): p. 307.\u003c/li\u003e\n \u003cli\u003eLi, Y., et al., \u003cem\u003eRadiomics analysis of lung CT for multidrug resistance prediction in active tuberculosis: a multicentre study.\u003c/em\u003e Eur Radiol, 2023. \u003cstrong\u003e33\u003c/strong\u003e(9): p. 6308-6317.\u003c/li\u003e\n \u003cli\u003ePatrick, L., et al., \u003cem\u003eEvaluating stability of histomorphometric features across scanner and staining variations: prostate cancer diagnosis from whole slide images.\u003c/em\u003e Journal of Medical Imaging, 2016. \u003cstrong\u003e3\u003c/strong\u003e(4): p. 047502.\u003c/li\u003e\n \u003cli\u003eCherezov, D., et al., \u003cem\u003eRank acquisition impact on radiomics estimation (AсquIRE) in chest CT imaging: A retrospective multi-site, multi-use-case study.\u003c/em\u003e Comput Methods Programs Biomed, 2024. \u003cstrong\u003e244\u003c/strong\u003e: p. 107990.\u003c/li\u003e\n \u003cli\u003eZhang, Y., G. Parmigiani, and W.E. Johnson, \u003cem\u003eComBat-seq: batch effect adjustment for RNA-seq count data.\u003c/em\u003e NAR Genomics and Bioinformatics, 2020. \u003cstrong\u003e2\u003c/strong\u003e(3): p. lqaa078.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, tuberculosis, radiology, diagnostic imaging","lastPublishedDoi":"10.21203/rs.3.rs-6822856/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6822856/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTuberculosis (TB) lesions display structural heterogeneity associated with disease status and progression. We explore the potential of radiomic features from computed tomography (CT) scans to classify TB lesions at a microscopic level.\u003c/p\u003e\n\u003cp\u003eTreated marmosets and untreated macaques infected with \u003cem\u003eMycobacterium tuberculosis\u003c/em\u003e underwent PET-CT imaging after a minimum of 12 weeks. Lesions were classified post-necropsy into necrotizing granulomas, non-necrotizing granulomas, inflammation beyond defined granulomas, and fibrotic granulomas. CT scans were segmented, radiomic features were extracted, and the top ten features were used to develop machine learning (ML) models for classification based on a one-versus-all approach. Top performing models were evaluated on a separate validation dataset.\u003c/p\u003e\n\u003cp\u003e151 lesions from macaques and 149 lesions from marmosets were identified. Top features included GrayLevelNonUnformity and SmallAreaHighGrayLevelEmphasis for both datasets. Linear discriminate analysis for macaques and adaptive boosting for marmosets averaged AUC-ROCs of 0.66-0.70 for discriminating lesion types, ranging from 0.45-0.52 for necrotizing granulomas to 0.91 for fibrotic granulomas.\u003c/p\u003e\n\u003cp\u003eRadiomics distinguished fibrotic from non-fibrotic TB granulomas in untreated macaques, potentially relevant for identifying disease progression risk. Classification of other lesion types was modest. Despite small, heterogeneous datasets across primate species, models performed consistently, supporting CT radiomics as a promising and trainable tool for non-invasive lesion phenotyping.\u003c/p\u003e","manuscriptTitle":"Radiomic phenotyping of tuberculosis histopathology from computed tomography scans of non-human primates","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-10 16:36:12","doi":"10.21203/rs.3.rs-6822856/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8d76c2fb-a2ef-429a-80ca-77462507aa8c","owner":[],"postedDate":"March 10th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":64219920,"name":"Biological sciences/Microbiology/Infectious disease diagnostics"},{"id":64219921,"name":"Health sciences/Diseases/Infectious diseases/Tuberculosis"},{"id":64219922,"name":"Biological sciences/Biological techniques/Imaging/X ray tomography"},{"id":64219923,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"}],"tags":[],"updatedAt":"2026-03-10T16:36:13+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-10 16:36:12","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6822856","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6822856","identity":"rs-6822856","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00