The Impact of Data Treatment Strategies on the Analysis of Bacterial Metabolomics Data

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract This study investigates the impact of various data treatment strategies on the analysis of bacterial metabolomics data obtained through GC-MS. Metabolomics data was generated from different sets of Staphylococcus aureus and Pseudomonas aeruginosa samples, including clinical isolates and type strains, cultured in different media conditions. The raw data was preprocessed using GC-MS data preprocessing and metabolite annotation software, followed by manual curation. We focused on evaluating the effects of missing value imputation, data transformation, data centering, and data normalization on the dataset. Descriptive statistics and data distribution are used to assess the impact of different data treatment steps. Univariate analysis using t-tests and Mann-Whitney U tests, and multivariate analysis employing OPLS-DA, and OPLS-EP were used to assess biological variation of compared groups. The results highlight the importance of careful data quality check throughout the data treatment process to ensure accurate and reliable interpretation of metabolomics data.
Full text 135,485 characters · extracted from preprint-html · click to expand
The Impact of Data Treatment Strategies on the Analysis of Bacterial Metabolomics Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article The Impact of Data Treatment Strategies on the Analysis of Bacterial Metabolomics Data Oleksandr Ilchenko, Henrik Antti This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7433902/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This study investigates the impact of various data treatment strategies on the analysis of bacterial metabolomics data obtained through GC-MS. Metabolomics data was generated from different sets of Staphylococcus aureus and Pseudomonas aeruginosa samples, including clinical isolates and type strains, cultured in different media conditions. The raw data was preprocessed using GC-MS data preprocessing and metabolite annotation software, followed by manual curation. We focused on evaluating the effects of missing value imputation, data transformation, data centering, and data normalization on the dataset. Descriptive statistics and data distribution are used to assess the impact of different data treatment steps. Univariate analysis using t-tests and Mann-Whitney U tests, and multivariate analysis employing OPLS-DA, and OPLS-EP were used to assess biological variation of compared groups. The results highlight the importance of careful data quality check throughout the data treatment process to ensure accurate and reliable interpretation of metabolomics data. metabolomics data treatment statistical analysis microbial interactions Staphylococcus aureus Pseudomonas aeruginosa OPLS-DA OPLS-EP Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Metabolomics is the comprehensive analysis of metabolites within a biological system, has become a powerful tool for studying microbial interactions, understanding metabolic pathways, and identifying biomarkers. Proper data treatment of complex metabolomics data is crucial for ensuring the accuracy and reliability of the metabolomics data analysis. This involves a series of steps, including missing value imputation, data transformation, data centering, and data normalization, each of which can influence the output results and extract meaningful biological information. These steps are necessary to address issues such as technical variability, instrumental drift, and differences in sample preparation, which can obscure true biological variations. In addition to data treatment, quality checks are an essential part of the metabolomics workflow 1 – 3 . However, the lack of standardized protocols for statistical data treatment in metabolomics poses a significant challenge, as inappropriate preprocessing methods can introduce biased results, distort biological signals, or lead to false conclusions. The selection of optimal strategies depends on data characteristics, such as missing value patterns, noise distribution, and biological variability, yet no universal guidelines exist for all scenarios. Previous studies have demonstrated that different normalization and transformation techniques can drastically alter downstream statistical results, affecting biomarker discovery and metabolic pathway interpretation 4 , 5 . Therefore, a systematic evaluation of data treatment approaches is essential to ensure robustness and reproducibility in metabolomics studies. This study aims to evaluate the impact of different data treatment strategies on bacterial GC-MS metabolomics data. We used metabolomics data from different sets of samples with S. aureus and P. aeruginosa . The sets include clinical isolates and type strains, strains cultured in different media conditions, including different sample fractions. The effects of various data treatment steps on the dataset were assessed using descriptive statistics and type of data distribution. Univariate and multivariate analysis were used to identify data variation. In conclusion, this study highlights the importance of data treatment quality check and careful consideration of data treatment techniques to ensure accurate and reliable interpretation of metabolomics data. Materials and methods Sample Preparation and Experimental Conditions In the study, we used metabolomics datasets from different sets of Staphylococcus aureus and Pseudomonas aeruginosa samples, including clinical isolates and type strains. The bacteria isolation method, cultivation, and metabolomics analysis are described in detail elsewhere [ref]. Briefly, bacterial strains were grown in three different media conditions to assess the impact of nutrient availability: M1 (TSB 50% + 1% glucose), M2 (TSB 50% + Chelex + 1% glucose), and M3 (TSB 50% + Chelex). Chelex is a chelating material that is mixed with the medium before usage and thereafter removed. The procedure is done in order to remove iron from the culture medium, thus the label “Chelex” is here used to denote iron-free medium. Cell and supernatant fractions were analyzed separately to differentiate intracellular and extracellular metabolomic changes. The bacterial samples included both single-strain cultures and mixed co-cultures to investigate interspecies interactions. Single-strain cultures were prepared for each clinical isolate of S. aureus and P. aeruginosa and laboratory type strains of S. aureus Newman and P. aeruginosa PAO1 separately, while mixed co-cultures combined both species in the same environment (Table 1 ). To ensure reproducibility, all bacterial cultures were grown under controlled conditions, maintaining identical incubation times and temperatures. Sample collection was standardized to minimize variability, with cells and supernatant fractions separated immediately after harvesting. Samples were stored at -80°C prior to metabolomic extraction and analysis. Table 1 Sample types, their names and quantity. Sample represented by three different media (M1, M2, M3). co-culture sing sing sing sing sing sing mix mix mix isolate clinical clinical type clinical clinical type clinical clinical type genus S. aureus S. aureus S. aureus P. aeruginosa P. aeruginosa P. aeruginosa S. aureus + P. aeruginosa S. aureus + P. aeruginosa S. aureus + P. aeruginosa strain A B Newman A B PAO1 A + A B + B Newman + PAO1 acronym SA SB SN PA PB P1 SAPA SBPB SNP1 M1 a 2 2 2 2 2 2 2 2 2 b 2 2 2 2 2 2 2 2 2 M2 a 2 2 2 2 2 2 2 2 2 b 2 2 2 2 2 2 2 2 2 M3 a 2 2 2 2 2 2 2 2 2 b 2 2 2 2 2 2 2 2 2 Data Preprocessing and Metabolite Annotation Raw data was preprocessed processed using in-house GC-MS data preprocessing and a metabolite annotation software 6 , 7 , thereafter the NIST Library was used for metabolite verification. The processed datasets were refined through the manual curation method by visual assessment of peaks. The metabolites list with the mass spectral data was verified with NIST MS Search 08 software against electron ionization mass spectra (EI-MS) databases with nistAutoSearch tool ( https://github.com/CreMoProduction/nistAutoSearch ) 8 . The NIST databases contain 92599 spectra in 17 libraries. In case of metabolite duplicates, a metabolite with the lowest NIST scores was discarded. If a metabolite feature was detected in < 50% of QC samples, it was removed from the rest of the data analysis. Data treatment Effective data treatment is crucial for ensuring accuracy and reliability of the data analysis. Data treatment was done in four steps (Table 2 ) and three different data normalization methods were compared. First, missing value imputation was used to eliminate sparse data (zero values), ensuring the completeness and reliability of the dataset. Second, data transformation was applied to modify the distribution of the data into a more relevant form, achieving normality and reducing skewness. Third, data centering was performed to adjust the range of data values, making them comparable across different variables or samples. Mean centering transforms data to a range from negative to positive around the value zero. Finally, different data normalization methods (Mean normalization, Pareto normalization, Autoscale normalization) were used to remove systematic variation in the data not related to variation of biological interest, such as instrument drift or sample preparation differences. Table 2 Data treatment steps Step Type Method Equation 1 Missing value imputation Half-minimum \(\:xᵢⱼ=\frac{x̄ᵢ}{2}\) 2 Data Transformation Log 10 \(\:x̃ᵢⱼ\:=\:log₁₀\left(xᵢⱼ\right)\) 3 Data Centering Mean centering \(\:\hat xᵢⱼ\:=\:x̃ᵢⱼ\:-\:x̄̃ᵢ\) 4 Data Normalization Mean \(\:\stackrel{\sim}{x}ᵢⱼ=\frac{\hat xᵢⱼ}{x̄ᵢ}\:\:\) 4 Data Normalization Pareto \(\:x̃ᵢⱼ\:=\:\frac{(\hat xᵢⱼ\:-\:x̄ᵢ)}{\surd\:sᵢ}\) 4 Data Normalization Autoscale \(\:x̃ᵢⱼ\:=\frac{(\hat xᵢⱼ\:-\:x̄ᵢ)}{{s}_{i}}\) Pairwise Sample Analysis Pairwise comparisons were conducted to assess metabolic differences between different bacterial strains and conditions. The analyses were grouped into two main types (Table 3 ): Strain Comparisons . Clinical Isolate vs. Clinical Isolate: Investigate metabolic differences between single clinical strains under identical conditions (e.g., SA vs. SB). Clinical Isolates vs. Laboratory Strains: Compare metabolic profiles of clinical isolates against reference laboratory strains (e.g., SA & SB vs. SN). Co-culture of Clinical Isolates vs. Co-culture of Laboratory Strains: Examine metabolic differences between co-cultures of clinical strains and co-cultures of laboratory strains (e.g., SAPA & SBPB vs. SNP1). Conditional Comparisons . Assess how environmental conditions influence metabolite profiles by comparing samples across different media conditions: examine the impact of chelex presence and glucose availability by comparing metabolite profiles (M1 vs. M2, M2 vs. M3, and M1 vs. M3). Table 3 Types of pairwise comparisons. Bacterial strain/isolate Factors Media Clinical (c) Lab (l) A B L M1 SA SB SN (Strain test) How different are clinical vs lab strains? What extent? (%) glucose PA PB P1 SAPA SBPB SNP1 Chelex M2 SA SB SN glucose PA PB P1 SAPA SBPB SNP1 M3 SA SB SN Chelex PA PB P1 SAPA SBPB SNP1 (Conditional test) How conditional factors influence bacteria metabolite profiles? What extent? (%) All comparisons aimed to determine the extent of metabolic variation and the influence of external conditions on bacterial interactions. Statistical significance and biological relevance were evaluated for each comparison to provide a comprehensive overview of microbial metabolic adaptations. Statistical Analysis Statistical analysis was performed to evaluate data quality and identify significant differences for the pairwise comparisons. Interpretation of metabolic variations to determine biological relevance is presented in further analysis as it is not a scope of this paper. The following approaches were applied: Descriptive Statistics . Descriptive statistical analysis assessed data normality and variance. Key metrices included kurtosis, skewness, and the coefficient of determination (R²) for Gaussian function fit. The Kolmogorov-Smirnov test evaluated an overall dataset normality as an alternative method to R² fit to Gaussian function. Univariate Analysis . Univariate statistical tests determined significant differences between metabolite levels at α = 0.05. A one-sample Kolmogorov-Smirnov (KS) test was used to check normality for every pairwise comparison group. Based on data distribution, either independent Student’s t-test (parametric) or the Mann-Whitney U test (non-parametric) was applied. Homogeneity of variance was assessed using Levene’s test. Benjamini-Hochberg procedure controlled the false discovery rate (FDR) in multiple comparisons, ensuring reliable results across all types of statistical tests. Percentage ratios of significant variables were calculated for all pairwise comparisons. The Student’s t-test, Mann-Whitney U test, Levene’s test, and Kolmogorov-Smirnov test were computed using in-built functions in R. Multivariate Analysis . Multivariate statistical methods were employed to assess overall data structure and group separations. Principal Component Analysis (PCA) provided an unsupervised overview of sample clustering and potential outliers. Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA) was used for supervised discrimination of independent sample groups 9 , while Orthogonal Partial Least Squares Effect Projection (OPLS-EP) analyzed the effect change assuming sample dependencies 10 . The OPLS based methods comes with the benefit of allowing separation of predictive and orthogonal systematic variation which enhances interpretation of relevant systematic changes in multivariate data. Although, compared groups of samples (M1, M2, M3) are not explicitly dependent, OPLS-EP effectively account for their basic conditional similarities, presence or absence of glucose and chelex. For instance, the M1 and M2 comparison aimed to isolate the effect of chelex in M2 media, and the M2 and M3 comparison aimed to isolate the effect of glucose (Table 3 ). Leave-one-out cross-validation (LOOCV) was performed to validate the model, and the best significance of CV-ANOVA was calculated depending on the number of components to assess the best model 11 . The quality of the data was further evaluated using score plots and CV scores plots, along with loadings plots to interpret the variables contributing to the model. The goodness of prediction (Q²) and goodness of fit (R²) values calculated to determine predictive ability and model fitness. Descriptive data statistics and univariate analysis were carried out using R version 4.4.2 12 , and multivariate analysis in the SIMCA 17 software 13 . Results Raw data processing and metabolites annotation In total 117 metabolites were found in the cell fraction and 79 in supernatant fraction respectively and 23 metabolites were found in both the cells and supernatant fractions. The analysis of missing values revealed that the supernatant fraction had generally fewer missing values compared to the cell fraction. In quality control (QC) samples, missing values accounted for only 0.11% in the supernatant fraction, whereas the cell fraction exhibited 1.93% missing values. Similarly, in analytical samples, missing values were lower in the supernatant fraction (2.33%) than in the cell fraction (3.23%). Descriptive Data Statistics The descriptive data analysis of the cells and supernatant datasets revealed differences in data distribution depending on different data treatment steps ( Sup table 1 , Fig. 1 ). For the fraction cells, the raw dataset showed high kurtosis and skewness, with a low R² Gaussian fit of 0.084 and 60% significant p-values from the KS test. In contrast, log 10 transformation significantly improved the R² Gaussian fit to 0.967 while reducing kurtosis and skewness to 3.4 and 0.12 respectively, with 31% of significant KS p-values. Mean normalization resulted in an extreme kurtosis of 677 and a skewness of 9.5, with 100% significant p-values from KS test. The fraction supernatant showed similar results. Raw dataset had high kurtosis and skewness with a low R² Gaussian fit 0.025 and 55% significant p-values. Log 10 transformation improved the R² Gaussian fit to 0.947 and reduced kurtosis to 3.3 and skewness to 0.071. Mean centering transformation for supernatant showed kurtosis of 10.1, skewness of 1.8, and 85% KS significant p values. These results indicate that log 10 transformation is effective for reducing skewness, and autoscale normalization further refines the data distribution in both cell and supernatant fractions. The analysis of the relationship between standard deviation (SD) and mean intensity under different data transformation and normalization techniques for samples cells (Fig. 2 A) and supernatant (Fig. 2 B) revealed distinct effects of these methods on the datasets. The raw data showed extreme values with a strong positive correlation between mean intensity and SD, indicating heteroscedasticity. The log 10 transformed dataset revealed extreme values mainly in the fraction supernatant, while mean transformed datasets contain extreme values in both datasets. Mean normalization and autoscaling normalization introduced discreteness into the data, evident from the clustering of points into distinct horizontal and vertical bands. Pareto normalization reduced the dependence between mean intensity and SD while maintaining a more continuous distribution. The analysis of the relative log abundance (RLA) of samples across different samples under different data transformation and normalization techniques for samples cells (Fig. 3 A) and supernatant (Fig. 3 B) showed effects on datasets of applied data manipulation techniques. The raw data showed a wide range of abundances, with extreme values and high variability, especially in the fraction supernatant. That emphasizes a crucial need to transform and normalize the data. RLA plot of log 10 transformed datasets did not show implicit changes, but it significantly improved data distribution, kurtosis and skewness as seen in in Fig. 1 and Sup table 1 . In the fraction cells mean normalization showed an artifact behavior evident from an outlier sample with extreme values, while this was not observed with other normalization methods. Pareto and autoscale normalization reduced the range and brought the data closer to a more uniform variation, though some samples still display substantial variability that belongs to biological variation. However, autoscale normalization particularly introduces more uniform variability among features, potentially masking inherent biological differences while Pareto normalization keep the data mainly intact and closer to the original measurement 14 . Univariate Statistical Data Analysis The results of the t-test and Mann-Whitney U test are summarized in Fig. 4 and Sup table 2. This table presents the percentage of significant metabolite differences. It presents pairwise comparisons based on the strain test and the conditional test (Table 3 ) for the cells and supernatant fractions. Higher percentages indicate greater metabolic divergence, while lower values suggest metabolic similarities. The analysis of the t-test results comparing A and B strains under different conditions revealed a low significant difference for the cell fractions. These results indicated that clinical strains A and B may be considered as one set of samples for further analysis. The analysis of the t-test results between clinical and lab strains under different conditions also revealed relatively low significance. Only SPc vs. SPl showed significant difference with a p-value of 31%. For the comparison between conditions M1 and M3, the S condition exhibited high significance at 51% as did the P condition at 58%. The SP clinical co-culture vs SP lab co-culture had a significant difference with a p-value of 43.6%. Percentage of significant variable based on Mann Whitney U test or t-test for fraction supernatant for strain pairwise comparison and conditional comparison. Results for the strain test between A and B strains under different conditions in the supernatant fraction revealed higher significant differences compared to data from the cell fraction. Nevertheless, clinical strains A and B may be considered as one group of samples for further analysis, in order to treat the cell fraction and supernatant fractions in a similar way. When comparing clinical and lab strains, the percentage of significant p-values in supernatants varied with the lowest difference for Sc and Sl cultured under M1 media (1.28%) and the highest of 33% for Sc and Sl under M2 media, giving 20% significant p-values, in average. For conditional tests comparing media types, significant p-values in supernatants ranged from 18–94%. Specifically, M2 and M3 showed the highest percentages, with 94% for S, 89% for P, and 76% for SP, while M1 and M2 revealed the lowest difference. The results of Levene’s test are not presented in the results section since they did not show any correlation from any of the comparisons, so it is not possible to draw any conclusion from it. R 2 for the linear regression model, when assessing the relationship between Levene’s test and Mann Whitney U test or t-test ranged from 0.005 to 0.1. Multivariate Statistical Data Analysis The multivariate analysis based on OPLS-DA and OPLS-EP models focused on identifying metabolic differences within conditional comparison. It compares M1 and M2, M2 and M3 with OPLS-EP models in order to assess the effect matrix by subtracting conditional factors, such as chelex + glucose vs. glucose; chelex + glucose vs. chelex (Table 3 ). An OPLS-DA model was used for M1 and M3 conditional comparison only, as there are no shared conditional factors (glucose vs. chelex). In the fraction cells the comparison between M1 and M3 yielded the highest model predictive ability (Q2) up to 0.95 with the minimum of 0.91, indicating good predictive ability of the model in distinguishing between these two groups (Table 4 ). The comparisons between M1 and M2, M2 and M3 showed model predictive ability in average of 0.846. However, the goodness of fit (R2X) values are lower, indicating that perhaps only a small portion of the total variance was fitted in the model. In the supernatant fraction the average goodness of prediction Q2 is higher for all comparison groups and varies between 0.921 and 0.99. The low CV-ANOVA p values and high percentages of significant p values across most comparisons underscore the statistical significance of the observed metabolic differences. The high R2X values in many of the supernatant comparisons suggest that the model is explaining a larger portion of the total variance compared to the cells fraction. This implies that the metabolic differences of biological variation between the experimental groups are more pronounced in the fraction supernatant compared to the fraction cells that is proven in univariate analysis with t-test (Fig. 4 ), hence OPLS-DA and OPLS-EP performance metrics are relatively higher. Table 4 The results of OPLS-DA and OPLS-EP models comparing different media conditions (M1, M2, M3) for the cells and supernatant fractions. The table displays the number of components used in the OPLS models, number of observations (N), the R2X (variance explained by the model in the X matrix), R2Y (variance explained by the model in the Y matrix), Q2 (predictive ability of the model), p-value CV-ANOVA, and the percentage of significant p values based on univariate independent t-test. Cells Groups Compared OPLS Components N R2X R2Y Q2 p value (CV-ANOVA) t-test significant p value, % S M1 M2 EP 1 + 0 + 0 12 0.443 0.921 0.885 2.02E-05 4.27 P EP 1 + 0 + 0 12 0.434 0.916 0.849 7.76E-05 32.48 SP EP 1 + 0 + 0 12 0.319 0.925 0.845 8.84E-05 22.22 S M2 M3 EP 1 + 0 + 0 12 0.32 0.912 0.841 1.02E-04 23.93 P EP 1 + 0 + 0 12 0.288 0.918 0.835 1.23E-04 41.03 SP EP 1 + 0 + 0 12 0.335 0.898 0.822 1.80E-04 33.33 S M1 M3 DA 2 + 1 + 0 24 0.404 0.985 0.95 4.72E-12 51.28 P DA 1 + 2 + 0 24 0.624 0.984 0.939 2.02E-09 58.12 SP DA 1 + 0 + 0 24 0.302 0.936 0.91 1.07E-11 43.59 Supernatant S M1 M2 EP 1 + 0 + 0 12 0.98 0.998 0.99 8.75E-07 17.72 P EP 1 + 0 + 0 12 0.332 0.992 0.945 4.34E-05 36.71 SP EP 1 + 0 + 0 12 0.685 0.966 0.921 1.85E-04 27.85 S M2 M3 EP 1 + 0 + 0 12 0.88 0.993 0.986 1.78E-07 93.67 P EP 1 + 0 + 0 12 0.915 0.997 0.991 6.00E-07 88.61 SP EP 1 + 0 + 0 12 0.752 0.96 0.955 1.93E-07 75.95 S M1 M3 DA 2 + 1 + 0 24 0.555 0.988 0.985 8.74E-20 83.54 P DA 1 + 2 + 0 24 0.722 0.998 0.991 2.03E-16 74.68 SP DA 1 + 0 + 0 24 0.535 0.991 0.987 1.84E-20 75.95 Discussion Out of the total metabolites identified, 23 were found in both the cell and supernatant fractions. Specifically, 117 metabolites were found in the cell fraction, and 79 were found in the supernatant fraction. The Kolmogorov-Smirnov (KS) test results as well as the analysis of the R² fit to a Gaussian distribution suggest that different data transformations steps can significantly impact the outcome of the type of the data distribution, highlighting the importance of choosing appropriate preprocessing methods for statistical analysis. Mean centering is not necessary for data transformation and is primarily used to enable the comparison of two datasets (cells and supernatant). If there is no comparison within or across datasets, its effect is mainly aesthetic. A disadvantage of mean centering is that it complicates the calculation of fold changes for zero-centered data and requires the usage of adjusted fold change values, thereby introducing extra data manipulation steps. Mean normalization and autoscale normalization introduced artefacts into datasets, they were specifically influenced by the presence of outlier data points that had extreme values in the cell fraction. These outliers led to an increase in kurtosis and skewness, as well as an absolute reduction in the R 2 fit to Gaussian distributions, so the dataset lost its normal distribution. Further analysis would be required to investigate this artefact further and explain such dataset output. The Mean Standard Deviation vs log 10 Mean Intensity plot revealed discreteness, while the RLA plot showed extreme outlier values in the cell fraction when mean normalization was applied. Log 10 transformation and Pareto normalization demonstrated effective removal of outliers while preserving the patterns with the biological variation of interest. Investigation of QC samples performed to detect technical drift in the datasets revealed that there is no need to use normalization based on QCs correction – e.g. the Regression model for QCs local polynomial fits correction 15 . The analysis of the t-test results between A and B strains for the cell fraction revealed low statistical significance while a relatively high statistical significance was observed for the supernatant fraction. Nevertheless, strains A and B from both fractions could be grouped together for further analysis, to not introduce differences in the data treatment for cells and supernatants originating from the same samples. Merging them further increases the sample size, which can enhance statistical power and the reliability of the obtained results. The sample size of a compared group, depending on the pairwise comparison type, varied from 4 to 12 samples per group. The analysis of the t-test results between clinical and lab strains revealed homogeneous differences with high significance in the supernatant fraction. SP clinical co-culture vs. SP lab co-culture showed over 30% spike of statistically significant variables only in cells fraction under condition M2 (chelex with glucose). The conditional testing for co-cultures showed differences with a low statistical significance compared to monocultures, for both cell and supernatant fractions. Comparison of M1 and M3 in the cell fraction demonstrated the highest difference, while in supernatant fraction the highest difference was observed for the comparison of the conditions M2 and M3. The t-test was used in all pairwise comparisons in the analysis, instead of the Mann Whitney U test, as the data of the compared groups followed a Gaussian distribution. However, the number of samples compared were insufficient to confirm the Gaussian distribution with high precision. In other words, a larger samples size would increase the likelihood of deviation from the Gaussian distribution, as the KS test under those conditions become more sensitive to even small deviations from the hypothesized distribution. It is therefore (?) necessary to consider the magnitude of the KS test statistics, but not only the p value significance 16 . Samples followed a Gaussian distribution due to the log 10 transformation, which helped to remove skewness, as well as Pareto normalization. Mean centering made it impossible to measure the relative standard deviation (RSD) to evaluate the effectiveness of the type of data normalization. So, mean centering had a second disadvantage and was therefore not used in the analyses. However, data centering is still might be applicable, e.g. min-max scaling. Levene’s test did not show any correlation from all the comparisons, so it is not possible to draw any conclusion from it. R 2 of the linear regression model demonstrated weak correlation between Levene’s test, and t-test or Mann Whitney U test. It means that there was no correlation between homogeneity of variance and actual statistical significance in the pairwise comparison of these types of biological samples. Unequal variance does not relate to equal means. Hence, we ignored Levene’s test results. But when testing statistical significance using the ANOVA test, homogeneity of variance might be important for consideration because the ANOVA test assumes equal variances among the compared samples 17 . OPLS-DA and OPLS-EP models revealed metabolic differences between the experimental groups with more pronounced differences observed in the supernatant fraction. These results were in line with the t-test results. The results suggest that the metabolic profiles of M1 and M3 for the cell fraction, as represented by the P. aeruginosa (P) dataset, are distinctly separated and well-modeled by the OPLS-DA. The variation in R2Y and R2X values across the datasets further highlights the complexity of the metabolic profiles and the varying degrees to which the models captured the variance within and between the compared groups. Conclusions This study emphasizes the importance of data quality checks and a thorough understanding of the impact of each transformation method. Careful consideration of data treatment techniques, ensuring alignment with specific dataset characteristics (outliers, distribution assumptions, and potential artifacts), is essential for enabling accurate and robust conclusions to be drawn from metabolomics data. Based on our datasets we suggest prioritizing methods that demonstrate robustness against outliers and keep normal distribution, such as log 10 transformation and Pareto normalization, while critically evaluating the necessity of mean centering. Declarations Funding: Metabolomics data used in this study were acquired in a project funded by the Swedish Research Council (#2018–03879), Kempestiftelserna (#JCK 22–0071) and Umeå University. Author Contribution O.I. and H.A. designed and conceptualized the study. O.I. conducted the data preprocessing, treatment, statistics, coding, and wrote the original draft of the manuscript. H.A. provided supervision and critically reviewed and edited the manuscript. All authors have reviewed and approved the final manuscript. Data Availability The data supporting the findings of this study are available in the Zenodo archive: https://doi.org/10.5281/zenodo.15089833 References Martínez-Arranz, I., Mayo, R., Pérez-Cormenzana, M., Mincholé, I., Salazar, L., Alonso, C., & Mato, J. M. (2015). Enhancing Metabolomics Research through Data Mining. J Proteomics , 127 , 275–288. https://doi.org/10.1016/j.jprot.2015.01.019 Fu, J., Zhang, Y., Liu, J., Lian, X., Tang, J., Zhu, F., & Pharmacometabonomics (2021). Data Processing and Statistical Analysis. Briefings In Bioinformatics , 22 (5), bbab138. https://doi.org/10.1093/bib/bbab138 Li, B., Tang, J., Yang, Q., Cui, X., Li, S., Chen, S., Cao, Q., Xue, W., Chen, N., & Zhu, F. (2016). Performance Evaluation and Online Realization of Data-Driven Normalization Methods Used in LC/MS Based Untargeted Metabolomics Analysis. Scientific Reports , 6 (1), 38881. https://doi.org/10.1038/srep38881 Di Guida, R., Engel, J., Allwood, J. W., Weber, R. J. M., Jones, M. R., Sommer, U., Viant, M. R., Dunn, W. B., & Non-Targeted (2016). Investigation of Normalisation, Missing Value Imputation, Transformation and Scaling. Metabolomics , 12 , 93. https://doi.org/10.1007/s11306-016-1030-9 . UHPLC-MS Metabolomic Data Processing Methods: A Comparative. Armitage, E. G., Godzien, J., Alonso-Herranz, V., López-Gonzálvez, Á., & Barbas, C. (2015). Missing Value Imputation Strategies for Metabolomics Data. Electrophoresis , 36 (24), 3050–3060. https://doi.org/10.1002/elps.201500352 Jonsson, P., Johansson, A. I., Gullberg, J., Trygg, J., A, J., Grung, B., Marklund, S., Sjöström, M., Antti, H., & Moritz, T. (2005). High-Throughput Data Analysis for Detecting and Identifying Differences between Samples in GC/MS-Based Metabolomic Analyses. Analytical Chemistry , 77 (17), 5635–5642. https://doi.org/10.1021/ac050601e Jonsson, P., Gullberg, J., Nordström, A., Kusano, M., Kowalczyk, M., Sjöström, M., & Moritz, T. A. (2004). Strategy for Identifying Differences in Large Series of Metabolomic Samples Analyzed by GC/MS. Analytical Chemistry , 76 (6), 1738–1745. https://doi.org/10.1021/ac0352427 Ilchenko, O. N. I. S. T., Auto Search, R., & Tool (2025). https://doi.org/10.5281/zenodo.15095049 Trygg, J., & Wold, S. (2002). Orthogonal projections to latent structures (O-PLS). Journal Of Chemometrics , 16 (3), 119–128. https://doi.org/10.1002/cem.695 Jonsson, P., Wuolikainen, A., Thysell, E., Chorell, E., Stattin, P., Wikström, P., & Antti, H. (2015). Constrained Randomization and Multivariate Effect Projections Improve Information Extraction and Biomarker Pattern Discovery in Metabolomics Studies Involving Dependent Samples. Metabolomics Off J Metabolomic Soc , 11 (6), 1667–1678. https://doi.org/10.1007/s11306-015-0818-3 Eriksson, L., Trygg, J., & Wold, S. (2008). CV-ANOVA for Significance Testing of PLS and OPLS® Models. Journal Of Chemometrics , 22 (11–12), 594–600. https://doi.org/10.1002/cem.1187 R: The R Project for Statistical Computing . https://www.r-project.org/ (accessed 2025-03-13). SIMCA® - Multivariate Data Analysis Software . Sartorius. https://www.sartorius.com/en/products/process-analytical-technology/data-analytics-software/mvda-software/simca (accessed 2023-01-12). van den Berg, R. A., Hoefsloot, H. C., Westerhuis, J. A., Smilde, A. K., & van der Werf, M. J. (2006). Centering, Scaling, and Transformations: Improving the Biological Information Content of Metabolomics Data. Bmc Genomics , 7 , 142. https://doi.org/10.1186/1471-2164-7-142 Ranjbar, N., & Rasoul, M. (2015). Novel Preprocessing and Normalization Methods for Analysis of GC/LC-MS Data. Razali, N. M., & Wah, Y. B. Power Comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling Tests. Sawyer, S. F. (2009). Analysis of Variance: The Fundamental Concepts. The Journal Of Manual & Manipulative Therapy , 17 (2), 27. https://doi.org/10.1179/jmt.2009.17.2.27E . E-38E. Ilchenko, O. (2025). GC-MS Metabolomics Data for Staphylococcus Aureus and Pseudomonas Aeruginosa under Different Culture Conditions, https://doi.org/10.5281/zenodo.15089833 Additional Declarations No competing interests reported. Supplementary Files SupTable1.xlsx Sup table 1. Descriptive statistics of the cells and supernatant datasets under various data transformation steps, including kurtosis, skewness, Gaussian fit (R²), and KS test significance in percent, a lower percentage of p value indicates a lower probability that the data is drawn from a normal distribution. SupTable2.xlsx Sup table 2. Analysis of significant differences between strains and media using t-tests for the fraction cells and supernatant. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7433902","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":506388286,"identity":"ab751e69-7fbb-4ed8-9cf8-f3a1abe21a82","order_by":0,"name":"Oleksandr Ilchenko","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCElEQVRIiWNgGAWjYDACdgZmEMXYwMDYeABIMkgwMB9geIBPCzNCSwNUC1sCQwJxWhgYoFp4DPBq4WdmPmz4o+aO7PYG5oYDP3fY5Em293z+kNjGYM+PQ4tkM1tyMs+xZ8ZzgFYc7D2TVizNc3abBFBL4swG7FoMDvMYH2ZgO5w4A+i2w4xthxPnSeRuYwBqSTA4gEsL/+eDP/7BtfwHasl5DHaYPU4tPMwJvG1wLQcSZ0vkMIAcxrgBt1+MjXn7DhvPYAb5pS05cWbPMTOJhHMSiTNw2MLP3vxY8se3w7Iz2NsfPvjZZpc443jz4w8fymzs+XF4HwGYUbkShNSPglEwCkbBKMADAI/0YH6I46rJAAAAAElFTkSuQmCC","orcid":"","institution":"Swedish University of Agricultural Sciences","correspondingAuthor":true,"prefix":"","firstName":"Oleksandr","middleName":"","lastName":"Ilchenko","suffix":""},{"id":506388287,"identity":"b6321826-071f-4367-9f2e-7f1fc600237f","order_by":1,"name":"Henrik Antti","email":"","orcid":"","institution":"Umeå University","correspondingAuthor":false,"prefix":"","firstName":"Henrik","middleName":"","lastName":"Antti","suffix":""}],"badges":[],"createdAt":"2025-08-22 11:08:25","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7433902/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7433902/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":90340794,"identity":"470cbbf8-e255-4823-92a8-1fcb64a914de","added_by":"auto","created_at":"2025-09-01 15:09:09","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":299988,"visible":true,"origin":"","legend":"\u003cp\u003eFrequency plot. Fraction Cells (\u003cstrong\u003eA\u003c/strong\u003e) and Supernatant (\u003cstrong\u003eB\u003c/strong\u003e). Kurt – kurtosis, Skew – skewness, R\u003csup\u003e2\u003c/sup\u003e – coefficient of determination fit to a Gaussian distribution. Upper left – raw data, upper middle – log10, upper right – mean centering, bottom left – mean normalization, bottom middle – pareto normalization, bottom right – autoscale normalization. Red line – Gaussian distribution.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/4ef1803bfcd8a181382b684c.png"},{"id":90341171,"identity":"4e9a3055-14b2-4ae2-b0da-58fb5d871c6a","added_by":"auto","created_at":"2025-09-01 15:17:09","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":438284,"visible":true,"origin":"","legend":"\u003cp\u003eScatter plot. Log\u003csub\u003e10\u003c/sub\u003e Standard Deviation vs log\u003csub\u003e10\u003c/sub\u003e Mean Intensity. Fraction Cells (\u003cstrong\u003eA\u003c/strong\u003e) and Supernatant (\u003cstrong\u003eB\u003c/strong\u003e). Upper left – raw data, upper middle – log\u003csub\u003e10\u003c/sub\u003e, upper right – mean centering, bottom left – mean normalization, bottom middle – pareto normalization, bottom right – autoscale normalization\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/389e89c614cec8eb8ee828a7.png"},{"id":90340810,"identity":"18274b26-3210-4597-a337-b2c76635860b","added_by":"auto","created_at":"2025-09-01 15:09:10","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":487440,"visible":true,"origin":"","legend":"\u003cp\u003eRLA plots of samples of fraction cells (\u003cstrong\u003eA\u003c/strong\u003e) and supernatant (\u003cstrong\u003eB\u003c/strong\u003e). Upper left – raw data, upper middle – log\u003csub\u003e10\u003c/sub\u003e, upper right – mean centering, bottom left – mean normalization, bottom middle – pareto normalization, bottom right – autoscale normalization\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/e787f9661e9294594d87a964.png"},{"id":90340791,"identity":"de54ca83-60ba-471b-82f8-2af678312825","added_by":"auto","created_at":"2025-09-01 15:09:09","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":80924,"visible":true,"origin":"","legend":"\u003cp\u003eAnalysis of significant differences between strains and media using t-tests. Plots display the percentage of significant p-values for pairwise comparisons within strain tests (A vs. B, clinical vs. lab) and conditional tests (media comparisons: M1, M2, M3), for cells and supernatant samples.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/cbdc5c0f5761ae688e5f2d08.png"},{"id":104399706,"identity":"52648abe-f86f-4318-8e93-78907c30c563","added_by":"auto","created_at":"2026-03-11 12:07:19","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2086492,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/898d4335-011d-4e0c-83fb-058673db508d.pdf"},{"id":90340790,"identity":"32dfccca-a3f7-472f-bf8b-e00d33fb7469","added_by":"auto","created_at":"2025-09-01 15:09:09","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":9745,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSup table 1. \u003c/strong\u003eDescriptive statistics of the cells and supernatant datasets under various data transformation steps, including kurtosis, skewness, Gaussian fit (R²), and KS test significance in percent, a lower percentage of p value indicates a lower probability that the data is drawn from a normal distribution.\u003c/p\u003e","description":"","filename":"SupTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/2e8fd021d90b69282862b74b.xlsx"},{"id":90340805,"identity":"33f889cc-d1ee-40df-98a6-19e250c5cf05","added_by":"auto","created_at":"2025-09-01 15:09:09","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":10511,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSup table 2\u003c/strong\u003e.\u003cstrong\u003e \u003c/strong\u003eAnalysis of significant differences between strains and media using t-tests for the fraction cells and supernatant.\u003c/p\u003e","description":"","filename":"SupTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7433902/v1/d6e7d74cbfa5ae2ec8641754.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"The Impact of Data Treatment Strategies on the Analysis of Bacterial Metabolomics Data","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMetabolomics is the comprehensive analysis of metabolites within a biological system, has become a powerful tool for studying microbial interactions, understanding metabolic pathways, and identifying biomarkers. Proper data treatment of complex metabolomics data is crucial for ensuring the accuracy and reliability of the metabolomics data analysis. This involves a series of steps, including missing value imputation, data transformation, data centering, and data normalization, each of which can influence the output results and extract meaningful biological information. These steps are necessary to address issues such as technical variability, instrumental drift, and differences in sample preparation, which can obscure true biological variations. In addition to data treatment, quality checks are an essential part of the metabolomics workflow \u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eHowever, the lack of standardized protocols for statistical data treatment in metabolomics poses a significant challenge, as inappropriate preprocessing methods can introduce biased results, distort biological signals, or lead to false conclusions. The selection of optimal strategies depends on data characteristics, such as missing value patterns, noise distribution, and biological variability, yet no universal guidelines exist for all scenarios. Previous studies have demonstrated that different normalization and transformation techniques can drastically alter downstream statistical results, affecting biomarker discovery and metabolic pathway interpretation \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Therefore, a systematic evaluation of data treatment approaches is essential to ensure robustness and reproducibility in metabolomics studies.\u003c/p\u003e\u003cp\u003eThis study aims to evaluate the impact of different data treatment strategies on bacterial GC-MS metabolomics data. We used metabolomics data from different sets of samples with \u003cem\u003eS. aureus\u003c/em\u003e and \u003cem\u003eP. aeruginosa\u003c/em\u003e. The sets include clinical isolates and type strains, strains cultured in different media conditions, including different sample fractions. The effects of various data treatment steps on the dataset were assessed using descriptive statistics and type of data distribution. Univariate and multivariate analysis were used to identify data variation.\u003c/p\u003e\u003cp\u003eIn conclusion, this study highlights the importance of data treatment quality check and careful consideration of data treatment techniques to ensure accurate and reliable interpretation of metabolomics data.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eSample Preparation and Experimental Conditions\u003c/h2\u003e\u003cp\u003eIn the study, we used metabolomics datasets from different sets of \u003cem\u003eStaphylococcus aureus\u003c/em\u003e and \u003cem\u003ePseudomonas aeruginosa\u003c/em\u003e samples, including clinical isolates and type strains. The bacteria isolation method, cultivation, and metabolomics analysis are described in detail elsewhere [ref].\u003c/p\u003e\u003cp\u003eBriefly, bacterial strains were grown in three different media conditions to assess the impact of nutrient availability: M1 (TSB 50% + 1% glucose), M2 (TSB 50% + Chelex\u0026thinsp;+\u0026thinsp;1% glucose), and M3 (TSB 50% + Chelex). Chelex is a chelating material that is mixed with the medium before usage and thereafter removed. The procedure is done in order to remove iron from the culture medium, thus the label \u0026ldquo;Chelex\u0026rdquo; is here used to denote iron-free medium. Cell and supernatant fractions were analyzed separately to differentiate intracellular and extracellular metabolomic changes.\u003c/p\u003e\u003cp\u003eThe bacterial samples included both single-strain cultures and mixed co-cultures to investigate interspecies interactions. Single-strain cultures were prepared for each clinical isolate of \u003cem\u003eS. aureus\u003c/em\u003e and \u003cem\u003eP. aeruginosa\u003c/em\u003e and laboratory type strains of \u003cem\u003eS. aureus\u003c/em\u003e Newman and \u003cem\u003eP. aeruginosa\u003c/em\u003e PAO1 separately, while mixed co-cultures combined both species in the same environment (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eTo ensure reproducibility, all bacterial cultures were grown under controlled conditions, maintaining identical incubation times and temperatures. Sample collection was standardized to minimize variability, with cells and supernatant fractions separated immediately after harvesting. Samples were stored at -80\u0026deg;C prior to metabolomic extraction and analysis.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSample types, their names and quantity. Sample represented by three different media (M1, M2, M3).\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"11\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eco-culture\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003esing\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003esing\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003esing\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003esing\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003esing\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003esing\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003emix\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c10\"\u003e\u003cp\u003emix\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c11\"\u003e\u003cp\u003emix\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eisolate\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eclinical\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eclinical\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003etype\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eclinical\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eclinical\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003etype\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eclinical\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003eclinical\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003etype\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003egenus\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cem\u003eS. aureus\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cem\u003eS. aureus\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cem\u003eS. aureus\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cem\u003eP. aeruginosa\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cem\u003eP. aeruginosa\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e\u003cem\u003eP. aeruginosa\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e\u003cem\u003eS. aureus\u0026thinsp;+\u0026thinsp;P. aeruginosa\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e\u003cem\u003eS. aureus\u0026thinsp;+\u0026thinsp;P. aeruginosa\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e\u003cem\u003eS. aureus\u0026thinsp;+\u0026thinsp;P. aeruginosa\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003estrain\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNewman\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003ePAO1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eA\u0026thinsp;+\u0026thinsp;A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003eB\u0026thinsp;+\u0026thinsp;B\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003eNewman\u0026thinsp;+\u0026thinsp;PAO1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003eacronym\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSN\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003ePA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003ePB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eP1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eSAPA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003eSBPB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003eSNP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eM1\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003ea\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eb\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eM2\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003ea\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eb\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eM3\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003ea\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eb\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eData Preprocessing and Metabolite Annotation\u003c/h3\u003e\n\u003cp\u003eRaw data was preprocessed processed using in-house GC-MS data preprocessing and a metabolite annotation software \u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e, thereafter the NIST Library was used for metabolite verification. The processed datasets were refined through the manual curation method by visual assessment of peaks. The metabolites list with the mass spectral data was verified with NIST MS Search 08 software against electron ionization mass spectra (EI-MS) databases with nistAutoSearch tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/CreMoProduction/nistAutoSearch\u003c/span\u003e\u003cspan address=\"https://github.com/CreMoProduction/nistAutoSearch\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. The NIST databases contain 92599 spectra in 17 libraries. In case of metabolite duplicates, a metabolite with the lowest NIST scores was discarded. If a metabolite feature was detected in \u0026lt;\u0026thinsp;50% of QC samples, it was removed from the rest of the data analysis.\u003c/p\u003e\n\u003ch3\u003eData treatment\u003c/h3\u003e\n\u003cp\u003eEffective data treatment is crucial for ensuring accuracy and reliability of the data analysis. Data treatment was done in four steps (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) and three different data normalization methods were compared. First, missing value imputation was used to eliminate sparse data (zero values), ensuring the completeness and reliability of the dataset. Second, data transformation was applied to modify the distribution of the data into a more relevant form, achieving normality and reducing skewness. Third, data centering was performed to adjust the range of data values, making them comparable across different variables or samples. Mean centering transforms data to a range from negative to positive around the value zero. Finally, different data normalization methods (Mean normalization, Pareto normalization, Autoscale normalization) were used to remove systematic variation in the data not related to variation of biological interest, such as instrument drift or sample preparation differences.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eData treatment steps\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStep\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eType\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMethod\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEquation\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMissing value imputation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHalf-minimum\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:xᵢⱼ=\\frac{x̄ᵢ}{2}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Transformation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLog\u003csub\u003e10\u003c/sub\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x̃ᵢⱼ\\:=\\:log₁₀\\left(xᵢⱼ\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Centering\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMean centering\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\hat xᵢⱼ\\:=\\:x̃ᵢⱼ\\:-\\:x̄̃ᵢ\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Normalization\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMean\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{\\sim}{x}ᵢⱼ=\\frac{\\hat xᵢⱼ}{x̄ᵢ}\\:\\:\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Normalization\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePareto\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x̃ᵢⱼ\\:=\\:\\frac{(\\hat xᵢⱼ\\:-\\:x̄ᵢ)}{\\surd\\:sᵢ}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Normalization\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAutoscale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x̃ᵢⱼ\\:=\\frac{(\\hat xᵢⱼ\\:-\\:x̄ᵢ)}{{s}_{i}}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\n\u003ch3\u003ePairwise Sample Analysis\u003c/h3\u003e\n\u003cp\u003ePairwise comparisons were conducted to assess metabolic differences between different bacterial strains and conditions. The analyses were grouped into two main types (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e):\u003c/p\u003e\u003cp\u003e\u003cb\u003eStrain Comparisons\u003c/b\u003e. Clinical Isolate vs. Clinical Isolate: Investigate metabolic differences between single clinical strains under identical conditions (e.g., SA vs. SB). Clinical Isolates vs. Laboratory Strains: Compare metabolic profiles of clinical isolates against reference laboratory strains (e.g., SA \u0026amp; SB vs. SN). Co-culture of Clinical Isolates vs. Co-culture of Laboratory Strains: Examine metabolic differences between co-cultures of clinical strains and co-cultures of laboratory strains (e.g., SAPA \u0026amp; SBPB vs. SNP1).\u003c/p\u003e\u003cp\u003e\u003cb\u003eConditional Comparisons\u003c/b\u003e. Assess how environmental conditions influence metabolite profiles by comparing samples across different media conditions: examine the impact of chelex presence and glucose availability by comparing metabolite profiles (M1 vs. M2, M2 vs. M3, and M1 vs. M3).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTypes of pairwise comparisons.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colspan=\"3\" nameend=\"c5\" namest=\"c3\"\u003e\u003cp\u003eBacterial strain/isolate\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eFactors\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eMedia\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eClinical (c)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eLab (l)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eL\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSN\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\" morerows=\"8\" rowspan=\"9\"\u003e\u003cp\u003e\u003cb\u003e(Strain test)\u003c/b\u003e\u003c/p\u003e\u003cp\u003eHow different are clinical vs lab strains? What extent? (%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eglucose\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSAPA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSBPB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSNP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChelex\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSN\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eglucose\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSAPA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSBPB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSNP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSN\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChelex\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSAPA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSBPB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eSNP1\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colspan=\"3\" nameend=\"c5\" namest=\"c3\"\u003e\u003cp\u003e\u003cb\u003e(Conditional test)\u003c/b\u003e\u003c/p\u003e\u003cp\u003eHow conditional factors influence bacteria metabolite profiles? What extent? (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAll comparisons aimed to determine the extent of metabolic variation and the influence of external conditions on bacterial interactions. Statistical significance and biological relevance were evaluated for each comparison to provide a comprehensive overview of microbial metabolic adaptations.\u003c/p\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003eStatistical Analysis\u003c/h2\u003e\u003cp\u003eStatistical analysis was performed to evaluate data quality and identify significant differences for the pairwise comparisons. Interpretation of metabolic variations to determine biological relevance is presented in further analysis as it is not a scope of this paper. The following approaches were applied:\u003c/p\u003e\u003cp\u003e\u003cb\u003eDescriptive Statistics\u003c/b\u003e. Descriptive statistical analysis assessed data normality and variance. Key metrices included kurtosis, skewness, and the coefficient of determination (R\u0026sup2;) for Gaussian function fit. The Kolmogorov-Smirnov test evaluated an overall dataset normality as an alternative method to R\u0026sup2; fit to Gaussian function.\u003c/p\u003e\u003cp\u003e\u003cb\u003eUnivariate Analysis\u003c/b\u003e. Univariate statistical tests determined significant differences between metabolite levels at α\u0026thinsp;=\u0026thinsp;0.05. A one-sample Kolmogorov-Smirnov (KS) test was used to check normality for every pairwise comparison group. Based on data distribution, either independent Student\u0026rsquo;s t-test (parametric) or the Mann-Whitney U test (non-parametric) was applied. Homogeneity of variance was assessed using Levene\u0026rsquo;s test. Benjamini-Hochberg procedure controlled the false discovery rate (FDR) in multiple comparisons, ensuring reliable results across all types of statistical tests. Percentage ratios of significant variables were calculated for all pairwise comparisons. The Student\u0026rsquo;s t-test, Mann-Whitney U test, Levene\u0026rsquo;s test, and Kolmogorov-Smirnov test were computed using in-built functions in R.\u003c/p\u003e\u003cp\u003e\u003cb\u003eMultivariate Analysis\u003c/b\u003e. Multivariate statistical methods were employed to assess overall data structure and group separations. Principal Component Analysis (PCA) provided an unsupervised overview of sample clustering and potential outliers. Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA) was used for supervised discrimination of independent sample groups \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e, while Orthogonal Partial Least Squares Effect Projection (OPLS-EP) analyzed the effect change assuming sample dependencies \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. The OPLS based methods comes with the benefit of allowing separation of predictive and orthogonal systematic variation which enhances interpretation of relevant systematic changes in multivariate data. Although, compared groups of samples (M1, M2, M3) are not explicitly dependent, OPLS-EP effectively account for their basic conditional similarities, presence or absence of glucose and chelex. For instance, the M1 and M2 comparison aimed to isolate the effect of chelex in M2 media, and the M2 and M3 comparison aimed to isolate the effect of glucose (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Leave-one-out cross-validation (LOOCV) was performed to validate the model, and the best significance of CV-ANOVA was calculated depending on the number of components to assess the best model \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. The quality of the data was further evaluated using score plots and CV scores plots, along with loadings plots to interpret the variables contributing to the model. The goodness of prediction (Q\u0026sup2;) and goodness of fit (R\u0026sup2;) values calculated to determine predictive ability and model fitness.\u003c/p\u003e\u003cp\u003eDescriptive data statistics and univariate analysis were carried out using R version 4.4.2 \u003csup\u003e12\u003c/sup\u003e, and multivariate analysis in the SIMCA 17 software \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003eRaw data processing and metabolites annotation\u003c/h2\u003e\u003cp\u003eIn total 117 metabolites were found in the cell fraction and 79 in supernatant fraction respectively and 23 metabolites were found in both the cells and supernatant fractions. The analysis of missing values revealed that the supernatant fraction had generally fewer missing values compared to the cell fraction. In quality control (QC) samples, missing values accounted for only 0.11% in the supernatant fraction, whereas the cell fraction exhibited 1.93% missing values. Similarly, in analytical samples, missing values were lower in the supernatant fraction (2.33%) than in the cell fraction (3.23%).\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003e\u003c/h3\u003e\n\u003cdiv class=\"Heading\"\u003e\u003cb\u003eDescriptive Data Statistics\u003c/b\u003e\u003c/div\u003e\u003cp\u003eThe descriptive data analysis of the cells and supernatant datasets revealed differences in data distribution depending on different data treatment steps (\u003cb\u003eSup table 1\u003c/b\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For the fraction cells, the raw dataset showed high kurtosis and skewness, with a low R\u0026sup2; Gaussian fit of 0.084 and 60% significant p-values from the KS test. In contrast, log\u003csub\u003e10\u003c/sub\u003e transformation significantly improved the R\u0026sup2; Gaussian fit to 0.967 while reducing kurtosis and skewness to 3.4 and 0.12 respectively, with 31% of significant KS p-values. Mean normalization resulted in an extreme kurtosis of 677 and a skewness of 9.5, with 100% significant p-values from KS test. The fraction supernatant showed similar results. Raw dataset had high kurtosis and skewness with a low R\u0026sup2; Gaussian fit 0.025 and 55% significant p-values. Log\u003csub\u003e10\u003c/sub\u003e transformation improved the R\u0026sup2; Gaussian fit to 0.947 and reduced kurtosis to 3.3 and skewness to 0.071. Mean centering transformation for supernatant showed kurtosis of 10.1, skewness of 1.8, and 85% KS significant p values. These results indicate that log\u003csub\u003e10\u003c/sub\u003e transformation is effective for reducing skewness, and autoscale normalization further refines the data distribution in both cell and supernatant fractions.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe analysis of the relationship between standard deviation (SD) and mean intensity under different data transformation and normalization techniques for samples cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA) and supernatant (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB) revealed distinct effects of these methods on the datasets. The raw data showed extreme values with a strong positive correlation between mean intensity and SD, indicating heteroscedasticity. The log\u003csub\u003e10\u003c/sub\u003e transformed dataset revealed extreme values mainly in the fraction supernatant, while mean transformed datasets contain extreme values in both datasets. Mean normalization and autoscaling normalization introduced discreteness into the data, evident from the clustering of points into distinct horizontal and vertical bands. Pareto normalization reduced the dependence between mean intensity and SD while maintaining a more continuous distribution.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe analysis of the relative log abundance (RLA) of samples across different samples under different data transformation and normalization techniques for samples cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA) and supernatant (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB) showed effects on datasets of applied data manipulation techniques. The raw data showed a wide range of abundances, with extreme values and high variability, especially in the fraction supernatant. That emphasizes a crucial need to transform and normalize the data. RLA plot of log\u003csub\u003e10\u003c/sub\u003e transformed datasets did not show implicit changes, but it significantly improved data distribution, kurtosis and skewness as seen in in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cb\u003eSup table 1\u003c/b\u003e. In the fraction cells mean normalization showed an artifact behavior evident from an outlier sample with extreme values, while this was not observed with other normalization methods. Pareto and autoscale normalization reduced the range and brought the data closer to a more uniform variation, though some samples still display substantial variability that belongs to biological variation. However, autoscale normalization particularly introduces more uniform variability among features, potentially masking inherent biological differences while Pareto normalization keep the data mainly intact and closer to the original measurement \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e\u003cb\u003eUnivariate Statistical Data Analysis\u003c/b\u003e\u003c/h2\u003e\u003cp\u003eThe results of the t-test and Mann-Whitney U test are summarized in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cb\u003eSup table 2.\u003c/b\u003e This table presents the percentage of significant metabolite differences. It presents pairwise comparisons based on the strain test and the conditional test (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) for the cells and supernatant fractions. Higher percentages indicate greater metabolic divergence, while lower values suggest metabolic similarities.\u003c/p\u003e\u003cp\u003eThe analysis of the t-test results comparing A and B strains under different conditions revealed a low significant difference for the cell fractions. These results indicated that clinical strains A and B may be considered as one set of samples for further analysis. The analysis of the t-test results between clinical and lab strains under different conditions also revealed relatively low significance. Only SPc vs. SPl showed significant difference with a p-value of 31%. For the comparison between conditions M1 and M3, the S condition exhibited high significance at 51% as did the P condition at 58%. The SP clinical co-culture vs SP lab co-culture had a significant difference with a p-value of 43.6%.\u003c/p\u003e\u003cp\u003ePercentage of significant variable based on Mann Whitney U test or t-test for fraction supernatant for strain pairwise comparison and conditional comparison. Results for the strain test between A and B strains under different conditions in the supernatant fraction revealed higher significant differences compared to data from the cell fraction. Nevertheless, clinical strains A and B may be considered as one group of samples for further analysis, in order to treat the cell fraction and supernatant fractions in a similar way. When comparing clinical and lab strains, the percentage of significant p-values in supernatants varied with the lowest difference for Sc and Sl cultured under M1 media (1.28%) and the highest of 33% for Sc and Sl under M2 media, giving 20% significant p-values, in average. For conditional tests comparing media types, significant p-values in supernatants ranged from 18\u0026ndash;94%. Specifically, M2 and M3 showed the highest percentages, with 94% for S, 89% for P, and 76% for SP, while M1 and M2 revealed the lowest difference.\u003c/p\u003e\u003cp\u003eThe results of Levene\u0026rsquo;s test are not presented in the results section since they did not show any correlation from any of the comparisons, so it is not possible to draw any conclusion from it. R\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e for the linear regression model, when assessing the relationship between Levene\u0026rsquo;s test and Mann Whitney U test or t-test ranged from 0.005 to 0.1.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e\u003cb\u003eMultivariate Statistical Data Analysis\u003c/b\u003e\u003c/h2\u003e\u003cp\u003eThe multivariate analysis based on OPLS-DA and OPLS-EP models focused on identifying metabolic differences within conditional comparison. It compares M1 and M2, M2 and M3 with OPLS-EP models in order to assess the effect matrix by subtracting conditional factors, such as chelex\u0026thinsp;+\u0026thinsp;glucose vs. glucose; chelex\u0026thinsp;+\u0026thinsp;glucose vs. chelex (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). An OPLS-DA model was used for M1 and M3 conditional comparison only, as there are no shared conditional factors (glucose vs. chelex). In the fraction cells the comparison between M1 and M3 yielded the highest model predictive ability (Q2) up to 0.95 with the minimum of 0.91, indicating good predictive ability of the model in distinguishing between these two groups (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). The comparisons between M1 and M2, M2 and M3 showed model predictive ability in average of 0.846. However, the goodness of fit (R2X) values are lower, indicating that perhaps only a small portion of the total variance was fitted in the model. In the supernatant fraction the average goodness of prediction Q2 is higher for all comparison groups and varies between 0.921 and 0.99. The low CV-ANOVA p values and high percentages of significant p values across most comparisons underscore the statistical significance of the observed metabolic differences. The high R2X values in many of the supernatant comparisons suggest that the model is explaining a larger portion of the total variance compared to the cells fraction. This implies that the metabolic differences of biological variation between the experimental groups are more pronounced in the fraction supernatant compared to the fraction cells that is proven in univariate analysis with t-test (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e), hence OPLS-DA and OPLS-EP performance metrics are relatively higher.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eThe results of OPLS-DA and OPLS-EP models comparing different media conditions (M1, M2, M3) for the cells and supernatant fractions. The table displays the number of components used in the OPLS models, number of observations (N), the R2X (variance explained by the model in the X matrix), R2Y (variance explained by the model in the Y matrix), Q2 (predictive ability of the model), p-value CV-ANOVA, and the percentage of significant p values based on univariate independent t-test.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"11\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colspan=\"11\" nameend=\"c11\" namest=\"c1\"\u003e\u003cp\u003eCells\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"3\" nameend=\"c3\" namest=\"c1\"\u003e\u003cp\u003eGroups Compared\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eOPLS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eComponents\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eR2X\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eR2Y\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eQ2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003ep value (CV-ANOVA)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003et-test significant p value, %\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.443\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.921\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.885\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.02E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e4.27\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.434\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.916\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.849\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e7.76E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e32.48\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.319\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.925\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.845\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e8.84E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e22.22\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.32\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.912\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.841\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.02E-04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e23.93\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.288\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.918\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.835\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.23E-04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e41.03\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.335\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.898\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.822\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.80E-04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e33.33\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u0026thinsp;+\u0026thinsp;1\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.404\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.985\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e4.72E-12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e51.28\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;2\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.624\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.984\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.939\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.02E-09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e58.12\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.302\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.936\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.07E-11\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e43.59\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"11\" nameend=\"c11\" namest=\"c1\"\u003e\u003cp\u003eSupernatant\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.998\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.99\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e8.75E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e17.72\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.332\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.992\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.945\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e4.34E-05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e36.71\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.685\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.966\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.921\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.85E-04\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e27.85\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.88\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.993\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.986\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.78E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e93.67\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.915\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.997\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.991\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e6.00E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e88.61\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.752\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.96\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.955\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.93E-07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e75.95\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eS\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eM3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u0026thinsp;+\u0026thinsp;1\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.555\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.988\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.985\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e8.74E-20\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e83.54\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;2\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.722\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.998\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.991\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.03E-16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e74.68\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1\u0026thinsp;+\u0026thinsp;0\u0026thinsp;+\u0026thinsp;0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e24\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.535\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.991\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.987\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e1.84E-20\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e75.95\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eOut of the total metabolites identified, 23 were found in both the cell and supernatant fractions. Specifically, 117 metabolites were found in the cell fraction, and 79 were found in the supernatant fraction.\u003c/p\u003e\u003cp\u003eThe Kolmogorov-Smirnov (KS) test results as well as the analysis of the R\u0026sup2; fit to a Gaussian distribution suggest that different data transformations steps can significantly impact the outcome of the type of the data distribution, highlighting the importance of choosing appropriate preprocessing methods for statistical analysis. Mean centering is not necessary for data transformation and is primarily used to enable the comparison of two datasets (cells and supernatant). If there is no comparison within or across datasets, its effect is mainly aesthetic. A disadvantage of mean centering is that it complicates the calculation of fold changes for zero-centered data and requires the usage of adjusted fold change values, thereby introducing extra data manipulation steps.\u003c/p\u003e\u003cp\u003eMean normalization and autoscale normalization introduced artefacts into datasets, they were specifically influenced by the presence of outlier data points that had extreme values in the cell fraction. These outliers led to an increase in kurtosis and skewness, as well as an absolute reduction in the R\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e fit to Gaussian distributions, so the dataset lost its normal distribution. Further analysis would be required to investigate this artefact further and explain such dataset output. The Mean Standard Deviation vs log\u003csub\u003e10\u003c/sub\u003e Mean Intensity plot revealed discreteness, while the RLA plot showed extreme outlier values in the cell fraction when mean normalization was applied. Log\u003csub\u003e10\u003c/sub\u003e transformation and Pareto normalization demonstrated effective removal of outliers while preserving the patterns with the biological variation of interest. Investigation of QC samples performed to detect technical drift in the datasets revealed that there is no need to use normalization based on QCs correction \u0026ndash; e.g. the Regression model for QCs local polynomial fits correction \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eThe analysis of the t-test results between A and B strains for the cell fraction revealed low statistical significance while a relatively high statistical significance was observed for the supernatant fraction. Nevertheless, strains A and B from both fractions could be grouped together for further analysis, to not introduce differences in the data treatment for cells and supernatants originating from the same samples. Merging them further increases the sample size, which can enhance statistical power and the reliability of the obtained results. The sample size of a compared group, depending on the pairwise comparison type, varied from 4 to 12 samples per group.\u003c/p\u003e\u003cp\u003eThe analysis of the t-test results between clinical and lab strains revealed homogeneous differences with high significance in the supernatant fraction. SP clinical co-culture vs. SP lab co-culture showed over 30% spike of statistically significant variables only in cells fraction under condition M2 (chelex with glucose).\u003c/p\u003e\u003cp\u003eThe conditional testing for co-cultures showed differences with a low statistical significance compared to monocultures, for both cell and supernatant fractions. Comparison of M1 and M3 in the cell fraction demonstrated the highest difference, while in supernatant fraction the highest difference was observed for the comparison of the conditions M2 and M3.\u003c/p\u003e\u003cp\u003eThe t-test was used in all pairwise comparisons in the analysis, instead of the Mann Whitney U test, as the data of the compared groups followed a Gaussian distribution. However, the number of samples compared were insufficient to confirm the Gaussian distribution with high precision. In other words, a larger samples size would increase the likelihood of deviation from the Gaussian distribution, as the KS test under those conditions become more sensitive to even small deviations from the hypothesized distribution. It is therefore (?) necessary to consider the magnitude of the KS test statistics, but not only the p value significance \u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. Samples followed a Gaussian distribution due to the log\u003csub\u003e10\u003c/sub\u003e transformation, which helped to remove skewness, as well as Pareto normalization. Mean centering made it impossible to measure the relative standard deviation (RSD) to evaluate the effectiveness of the type of data normalization. So, mean centering had a second disadvantage and was therefore not used in the analyses. However, data centering is still might be applicable, e.g. min-max scaling. Levene\u0026rsquo;s test did not show any correlation from all the comparisons, so it is not possible to draw any conclusion from it. R\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e of the linear regression model demonstrated weak correlation between Levene\u0026rsquo;s test, and t-test or Mann Whitney U test. It means that there was no correlation between homogeneity of variance and actual statistical significance in the pairwise comparison of these types of biological samples. Unequal variance does not relate to equal means. Hence, we ignored Levene\u0026rsquo;s test results. But when testing statistical significance using the ANOVA test, homogeneity of variance might be important for consideration because the ANOVA test assumes equal variances among the compared samples \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eOPLS-DA and OPLS-EP models revealed metabolic differences between the experimental groups with more pronounced differences observed in the supernatant fraction. These results were in line with the t-test results. The results suggest that the metabolic profiles of M1 and M3 for the cell fraction, as represented by the \u003cem\u003eP. aeruginosa\u003c/em\u003e (P) dataset, are distinctly separated and well-modeled by the OPLS-DA. The variation in R2Y and R2X values across the datasets further highlights the complexity of the metabolic profiles and the varying degrees to which the models captured the variance within and between the compared groups.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis study emphasizes the importance of data quality checks and a thorough understanding of the impact of each transformation method. Careful consideration of data treatment techniques, ensuring alignment with specific dataset characteristics (outliers, distribution assumptions, and potential artifacts), is essential for enabling accurate and robust conclusions to be drawn from metabolomics data. Based on our datasets we suggest prioritizing methods that demonstrate robustness against outliers and keep normal distribution, such as log\u003csub\u003e10\u003c/sub\u003e transformation and Pareto normalization, while critically evaluating the necessity of mean centering.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eFunding:\u003c/h2\u003e\u003cp\u003eMetabolomics data used in this study were acquired in a project funded by the Swedish Research Council (#2018\u0026ndash;03879), Kempestiftelserna (#JCK 22\u0026ndash;0071) and Ume\u0026aring; University.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eO.I. and H.A. designed and conceptualized the study. O.I. conducted the data preprocessing, treatment, statistics, coding, and wrote the original draft of the manuscript. H.A. provided supervision and critically reviewed and edited the manuscript. All authors have reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe data supporting the findings of this study are available in the Zenodo archive: https://doi.org/10.5281/zenodo.15089833\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eMart\u0026iacute;nez-Arranz, I., Mayo, R., P\u0026eacute;rez-Cormenzana, M., Minchol\u0026eacute;, I., Salazar, L., Alonso, C., \u0026amp; Mato, J. M. (2015). Enhancing Metabolomics Research through Data Mining. \u003cem\u003eJ Proteomics\u003c/em\u003e, \u003cem\u003e127\u003c/em\u003e, 275\u0026ndash;288. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jprot.2015.01.019\u003c/span\u003e\u003cspan address=\"10.1016/j.jprot.2015.01.019\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFu, J., Zhang, Y., Liu, J., Lian, X., Tang, J., Zhu, F., \u0026amp; Pharmacometabonomics (2021). Data Processing and Statistical Analysis. \u003cem\u003eBriefings In Bioinformatics\u003c/em\u003e, \u003cem\u003e22\u003c/em\u003e(5), bbab138. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/bib/bbab138\u003c/span\u003e\u003cspan address=\"10.1093/bib/bbab138\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi, B., Tang, J., Yang, Q., Cui, X., Li, S., Chen, S., Cao, Q., Xue, W., Chen, N., \u0026amp; Zhu, F. (2016). Performance Evaluation and Online Realization of Data-Driven Normalization Methods Used in LC/MS Based Untargeted Metabolomics Analysis. \u003cem\u003eScientific Reports\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(1), 38881. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/srep38881\u003c/span\u003e\u003cspan address=\"10.1038/srep38881\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDi Guida, R., Engel, J., Allwood, J. W., Weber, R. J. M., Jones, M. R., Sommer, U., Viant, M. R., Dunn, W. B., \u0026amp; Non-Targeted (2016). Investigation of Normalisation, Missing Value Imputation, Transformation and Scaling. \u003cem\u003eMetabolomics\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e, 93. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11306-016-1030-9\u003c/span\u003e\u003cspan address=\"10.1007/s11306-016-1030-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. UHPLC-MS Metabolomic Data Processing Methods: A Comparative.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eArmitage, E. G., Godzien, J., Alonso-Herranz, V., L\u0026oacute;pez-Gonz\u0026aacute;lvez, \u0026Aacute;., \u0026amp; Barbas, C. (2015). Missing Value Imputation Strategies for Metabolomics Data. \u003cem\u003eElectrophoresis\u003c/em\u003e, \u003cem\u003e36\u003c/em\u003e(24), 3050\u0026ndash;3060. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/elps.201500352\u003c/span\u003e\u003cspan address=\"10.1002/elps.201500352\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJonsson, P., Johansson, A. I., Gullberg, J., Trygg, J., A, J., Grung, B., Marklund, S., Sj\u0026ouml;str\u0026ouml;m, M., Antti, H., \u0026amp; Moritz, T. (2005). High-Throughput Data Analysis for Detecting and Identifying Differences between Samples in GC/MS-Based Metabolomic Analyses. \u003cem\u003eAnalytical Chemistry\u003c/em\u003e, \u003cem\u003e77\u003c/em\u003e(17), 5635\u0026ndash;5642. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1021/ac050601e\u003c/span\u003e\u003cspan address=\"10.1021/ac050601e\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJonsson, P., Gullberg, J., Nordstr\u0026ouml;m, A., Kusano, M., Kowalczyk, M., Sj\u0026ouml;str\u0026ouml;m, M., \u0026amp; Moritz, T. A. (2004). Strategy for Identifying Differences in Large Series of Metabolomic Samples Analyzed by GC/MS. \u003cem\u003eAnalytical Chemistry\u003c/em\u003e, \u003cem\u003e76\u003c/em\u003e(6), 1738\u0026ndash;1745. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1021/ac0352427\u003c/span\u003e\u003cspan address=\"10.1021/ac0352427\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIlchenko, O. N. I. S. T., Auto Search, R., \u0026amp; Tool (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.15095049\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.15095049\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTrygg, J., \u0026amp; Wold, S. (2002). Orthogonal projections to latent structures (O-PLS). \u003cem\u003eJournal Of Chemometrics\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(3), 119\u0026ndash;128. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/cem.695\u003c/span\u003e\u003cspan address=\"10.1002/cem.695\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJonsson, P., Wuolikainen, A., Thysell, E., Chorell, E., Stattin, P., Wikstr\u0026ouml;m, P., \u0026amp; Antti, H. (2015). Constrained Randomization and Multivariate Effect Projections Improve Information Extraction and Biomarker Pattern Discovery in Metabolomics Studies Involving Dependent Samples. \u003cem\u003eMetabolomics Off J Metabolomic Soc\u003c/em\u003e, \u003cem\u003e11\u003c/em\u003e(6), 1667\u0026ndash;1678. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11306-015-0818-3\u003c/span\u003e\u003cspan address=\"10.1007/s11306-015-0818-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEriksson, L., Trygg, J., \u0026amp; Wold, S. (2008). CV-ANOVA for Significance Testing of PLS and OPLS\u0026reg; Models. \u003cem\u003eJournal Of Chemometrics\u003c/em\u003e, \u003cem\u003e22\u003c/em\u003e(11\u0026ndash;12), 594\u0026ndash;600. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/cem.1187\u003c/span\u003e\u003cspan address=\"10.1002/cem.1187\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e\u003cem\u003eR: The R Project for Statistical Computing\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.r-project.org/\u003c/span\u003e\u003cspan address=\"https://www.r-project.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 2025-03-13).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e\u003cem\u003eSIMCA\u0026reg; - Multivariate Data Analysis Software\u003c/em\u003e. Sartorius. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.sartorius.com/en/products/process-analytical-technology/data-analytics-software/mvda-software/simca\u003c/span\u003e\u003cspan address=\"https://www.sartorius.com/en/products/process-analytical-technology/data-analytics-software/mvda-software/simca\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 2023-01-12).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003evan den Berg, R. A., Hoefsloot, H. C., Westerhuis, J. A., Smilde, A. K., \u0026amp; van der Werf, M. J. (2006). Centering, Scaling, and Transformations: Improving the Biological Information Content of Metabolomics Data. \u003cem\u003eBmc Genomics\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e, 142. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/1471-2164-7-142\u003c/span\u003e\u003cspan address=\"10.1186/1471-2164-7-142\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRanjbar, N., \u0026amp; Rasoul, M. (2015). Novel Preprocessing and Normalization Methods for Analysis of GC/LC-MS Data.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRazali, N. M., \u0026amp; Wah, Y. B. Power Comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling Tests.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSawyer, S. F. (2009). Analysis of Variance: The Fundamental Concepts. \u003cem\u003eThe Journal Of Manual \u0026amp; Manipulative Therapy\u003c/em\u003e, \u003cem\u003e17\u003c/em\u003e(2), 27. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1179/jmt.2009.17.2.27E\u003c/span\u003e\u003cspan address=\"10.1179/jmt.2009.17.2.27E\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. E-38E.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIlchenko, O. (2025). GC-MS Metabolomics Data for Staphylococcus Aureus and Pseudomonas Aeruginosa under Different Culture Conditions, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.15089833\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.15089833\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"metabolomics, data treatment, statistical analysis, microbial interactions, Staphylococcus aureus, Pseudomonas aeruginosa, OPLS-DA, OPLS-EP","lastPublishedDoi":"10.21203/rs.3.rs-7433902/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7433902/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis study investigates the impact of various data treatment strategies on the analysis of bacterial metabolomics data obtained through GC-MS. Metabolomics data was generated from different sets of \u003cem\u003eStaphylococcus aureus\u003c/em\u003e and \u003cem\u003ePseudomonas aeruginosa\u003c/em\u003e samples, including clinical isolates and type strains, cultured in different media conditions. The raw data was preprocessed using GC-MS data preprocessing and metabolite annotation software, followed by manual curation. We focused on evaluating the effects of missing value imputation, data transformation, data centering, and data normalization on the dataset. Descriptive statistics and data distribution are used to assess the impact of different data treatment steps. Univariate analysis using t-tests and Mann-Whitney U tests, and multivariate analysis employing OPLS-DA, and OPLS-EP were used to assess biological variation of compared groups. The results highlight the importance of careful data quality check throughout the data treatment process to ensure accurate and reliable interpretation of metabolomics data.\u003c/p\u003e","manuscriptTitle":"The Impact of Data Treatment Strategies on the Analysis of Bacterial Metabolomics Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-01 15:09:04","doi":"10.21203/rs.3.rs-7433902/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ea03e8cf-63bb-45f4-bbb7-a3ed8924c6f6","owner":[],"postedDate":"September 1st, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-02-28T18:39:01+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-01 15:09:04","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7433902","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7433902","identity":"rs-7433902","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00