{"paper_id":"462d9dca-6e9d-4133-b348-f7d2f2a82e6c","body_text":"Automated Metabolite Formula Ranking Using Formula Subset Analysis for LC-MS/MS-Based Metabolomics | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Automated Metabolite Formula Ranking Using Formula Subset Analysis for LC-MS/MS-Based Metabolomics Ke-Shiuan Lynn This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7764688/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 01 Dec, 2025 Read the published version in Metabolomics → Version 1 posted 9 You are reading this latest preprint version Abstract Introduction Metabolite identification remains a bottleneck in untargeted liquid chromatography–tandem mass spectrometry (LC-MS) metabolomics studies, especially when the underlying metabolite is not found in the tandem mass spectrometry (MS/MS) databases. Objective A new approach, formula subset analysis (FSA), was developed to effectively rank the chemical formula candidates for an MS/MS spectrum. Methods This approach first computes mother-daughter relationships (MDRs) among possible formulas of fragments and the precursor under a given mass tolerance and then determines the characteristic fragments (CFs) that only present one MDR with the precursor and other fragments. Subsequently, the precursor formula candidates are ranked by the scores derived from the number of MDRs. Results A numerical study using seven large datasets totaling 30357 MS/MS spectra from 6612 metabolites consisting of C, H, O, N, S, and P showed that FSA ranked the correct chemical formula as the top-1 candidate for a metabolite in 84.50% of the cases and in the top-5 candidates in 97.20% of the cases. The average processing time for each spectrum was 0.033 seconds. Moreover, FSA does not require training data, not rely on MS/MS databases, can be applied to a wide mass range, and can be quickly expanded with more chemical elements and formulas to identify different chemical species. Conclusions FSA has not utilized structural information yet and therefore its accuracy may not be competitive with some of the state-of-the-art identification tools. However, its advantages in speed, expandability, and applicability, make it suitable for prescreening candidates in untargeted LC-MS metabolomics studies. Metabolomics Formula Ranking LC-MS/MS Mother-Daughter Relationship Figures Figure 1 Figure 2 1 Introduction Metabolomics, which analyzes metabolite level changes in biological samples, has become a popular research area in systems biology. It provides an instantaneous snapshot of the physiology of a living system including its interactions with the environment, which is relatively limited in other omics such as genomics and proteomics. Liquid chromatography–mass spectrometry (LC-MS) is a commonly used platform for metabolomics studies, mostly because of its wide range of detectable polarity and molecular mass. However, metabolite identification remains a bottleneck in LC-MS-based metabolomics mainly due to component complexity and low database integrity (Dunn et al., 2011 ; Peisl et al., 2018 ; Want et al., 2010 ). Based on the five confidence levels for compound annotation and identification determined by the Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI), an unknown metabolite is claimed identified with “level-1 confidence” if its 2D structure matches with that of an authentic standard using at least two orthogonal techniques (such as MS/MS spectrum and retention time) (Blazenovic et al., 2018 ; Sumner et al., 2007 ; Viant et al., 2017 ). However, the availability of standard metabolites is limited, while some, although available, are very expensive. Moreover, finding the correct standard is generally a trial-and-error process. It is also almost impossible to find standards for novel or modified metabolites. Owing to such difficulties, putative identification (level 2–3 confidence) has become acceptable, wherein the peaks of the acquired spectrum are matched with those of spectra in public databases. Nevertheless, even with the continuous efforts in collecting MS/MS spectra in many public databases, the number of collected metabolites is still far less than the demand (Blazenovic et al., 2018 ; Dunn et al., 2013 ; Guijas et al., 2018 ; Kind et al., 2018 ). Other obstacles to spectrum match include variations in instruments, environment conditions, and sample preparation methods. Two major approaches were adopted for MS/MS spectrum identification. One attempts to predict all possible fragmentation patterns of compounds that agree with the given precursor mass, and then determine the most probable compound according to the matches between the predicted fragments and those in the given spectrum (Kind et al., 2013 ; Liebal et al., 2020 ; Ridder et al., 2014 ; Ruttkies et al., 2016 ; Wang et al., 2021 ; Wang et al., 2014 ). As this spectrum-prediction approach does not rely on the availability of MS/MS spectra in databases or any additional equipment, it is generally more applicable to metabolite identifications than the aforementioned approach. The generation of possible fragmentation patterns demands extensive chemical knowledge, which is usually obtained from specialists or via machine learning (Liebal et al., 2020 ; Wang et al., 2021 ). Such knowledge or rules may seem reasonable when a single compound is considered. However, metabolomics concerns a mix of many compounds; thus, under complicated interactions between metabolites and environmental influences, these rules may be subject to change. This leads to differences between predicted and actual spectra. For example, according to the test result of CFM-ID 4.0, the average similarity (measured by dot product) between in silico spectra and METLIN spectra is 0.38 (Wang et al., 2021 ). Furthermore, rules generated from a machine learning approach may not generalize well to unseen compound species. A re-train process is usually needed to account for such data. Another approach is oriented in the opposite direction. The possible chemical and physical features are extracted from an MS spectrum of the compound to be identified and the corresponding chemical structure is identified from databases of known compounds via chemical feature matching (Brouard et al., 2016 ; Dührkop et al., 2019 ; Heinonen et al., 2012 ; Ludwig et al., 2020 ; Shen et al., 2014 ; Tsugawa et al., 2016 ). However, this approach usually requires intensive computation and database queries, thus being time-, memory-, and network resource-consuming. For example, according to the test result of SIRIUS, the identification time for 5518 spectra was 5 h 41 min, resulting in an average of 3.71 s per spectrum, not including its long computation time (in days) for compounds with large molecular weight (≥ 850 Dalton (Da)) (Dührkop et al., 2019 ). Similar to the previous approach, features that are determined via machine learning are highly related to the training data and may require a re-train process if an unseen chemical class or species is present. In summary, current metabolite identification/ranking tools exhibit the following seven issues: 1) identification/ranking accuracy, 2) computation time, 3) applicability to unknown compounds, 4) applicability to large metabolites, 5) batch computation capability, 6) result reproducibility, and 7) internet accessibility. Considering these issues, in this work, a simple but effective approach was designed and developed, namely formula subset analysis (FSA), to rank chemical formula candidates during identification in LC-MS/MS-based metabolomics studies. Unlike MS/MS spectrum match, which relies on the availability of MS/MS spectra in public databases, FSA only requires chemical formulas and their corresponding monoisotopic molecular weights that are publicly available in large chemical databases, such as PubChem and ChemSpider. The proposed approach employs a mass match on each of the fragment monoisotopic masses in Da for possible chemical formulas, followed by a series of searches for mother–daughter relationships (MDRs) among the fragments. Here, an MDR indicates the relationship between a pair of ions in a spectrum where the chemical compositions of one ion is a subset of those of the other ion. The formula candidate of the precursor ion that comprises the most MDRs among all fragment ion formula candidates, including itself, is considered the most probable identity for the underlying metabolite. Notably, the FSA was designed to provide a ranking for the formula candidates of the underlying compound by using a list of “centroid and deisotoped” peak m/z values in an MS/MS spectrum. Thus, isomers that share the same molecular formula but have different atomic arrangements in space are not distinguishable by the proposed algorithm. Additionally, the MS/MS spectra of a metabolite acquired from different experimental conditions (e.g., collision energy, association method, temperature, and humidity) can exhibit different fragment peaks. Such variations in the fragmentation pattern may result in differences in the MDRs. FSA was designed to sustain such variation and maintain consistency provided that sufficient and accurate fragment peaks are presented in a spectrum. The designed FSA algorithm was tested on seven experimental datasets. Its formula ranking performance including accuracy, consistency, scalability, and suitable types of spectra for FSA are discussed herein. 2. Materials and methods 2.1 Construction of the Test Datasets Seven datasets were excerpted from two public databases, Human Metabolome Database (HMDB) (Wishart et al., 2018 ) and Global Natural Products Social Molecular Networking (GNPS) (Wang et al., 2016 ) to evaluate the performance of the proposed algorithm. The HMDB database mainly contains information for small molecule metabolites found in the human body, whereas the GNPS is a repository for sharing of MS/MS spectrometry data from various species, instruments, and omics studies. These seven datasets (three from HMDB and four from GNPS) were constructed to cover metabolites in various species, instruments, and a wild range of molecular weights (0–1500 Da). The first three datasets, namely HMDB_QTOF, HMDB_ITFT, and HMDB_QFT, were composed of the experimental MS/MS spectra acquired from LC-ESI-QTOF, LC-ESI-ITFT, and LC-ESI-QFT MS, respectively, in the HMDB ( http://www.hmdb.ca/downloads , release 2018-10-19). The LC-ESI-QTOF represents “Liquid Chromatography coupled with Electrospray Ionization-Quadrupole-Time of Flight” whereas the ITFT and QFT in other two instruments stand for Ion Trap-Fourier Transform and Quadrupole Fourier Transform, respectively. Although more spectra acquired from other instruments were also available in HMDB, only those acquired from the above three instruments were used as they had more spectra per instrument with sufficient mass accuracy (the provided m/z values had at least two decimal digits). The latter four datasets were excerpted from the GNPS database https://gnps-external.ucsd.edu/gnpslibrary on 2022/06/05. Of the four sets, three were extracted from the GNPS NIH Natural Products Library (both round 1 and round 2, to be referred as GNPS-NIH-NPL) and one was extracted from the GNPS Pacific Northwest National Lab Lipids Library (referred to as GNPS-PNNL-LIPIDS). The three datasets NPL_MAXIS, NPL_QTOF, and NPL_OT extracted from GNPS-NIH-NPL were constructed for three different instruments, LC-ESI-MAXIS, LC-ESI-QTOF and LC-ESI-OT (OT for orbitrap), respectively. Although MAXIS is actually a QTOF (Maxis II HD Q-TOF Bruker ™), it generated spectra with a lower mass accuracy than other QTOF spectra probably due to an older model. The last dataset, LIPID_OT extracted from GNPS-PNNL-LIPIDS, was acquired from LC-ESI-OT. Notably, two fragmentation types, collision-induced dissociation (CID) and higher energy collisional dissociation (HCD), of MS/MS spectra were provided for each compound in this dataset. To avoid redundancy, the HCD spectra were used because they usually exhibit more extensive fragmentation and thus more structural information about the ions (Jedrychowski et al., 2011 ). In addition, for a proof-of-concept study and consistency across datasets, the following four criteria were imposed to all datasets to select qualified spectra: Compounds are combinations of only C, H, O, N, S, and P where C and H are necessary elements. Molecular monoisotopic mass is < 1500 Da. The precursor type is either [M + H] + or [M − H]−. At least three peaks are present in the centroid and deisotoped MS/MS spectrum. The first two criteria are general definitions for metabolites. Notably, the compounds in GNPS-NIH-NPL did not involve P, whereas those in GNPS-PNNL-LIPIDS did not involve S. These elements were removed from the first criterion when candidate chemical formulas were searched in the corresponding datasets. The third criterion was added to simplify the composition in a spectrum. Spectra of metabolites with adduct such as Na and K will be addressed in the discussion section. The fourth criterion was set for sufficient fragments in a spectrum for its identification. As a result, the number of qualified spectra and their corresponding number of compounds in the seven datasets are listed in Table 1 . Table 1 Number of spectra and compounds of the seven datasets after filtering using the established criteria Dataset Source Instrument Original Spectra (Compounds) Qualified Spectra (Compounds) HMDB_QTOF HMDB_ITFT HMDB_QFT LC-ESI-QTOF 3743 (995) 3029 (821) LC-ESI-ITFT 4074 (314) 2601 (242) LC-ESI-QFT 2195 (334) 1439 (221) NPL_MAXIS NPL_QTOF NPL_OT LC-ESI-MAXIS 7641 (3624) 3561 (2833) LC-ESI-QTOF 2090 (1890) 1373 (1360) LC-ESI-OT 1314 (639) 270 (239) LIPIDS_OT LC-ESI-OT (HCD) 46724 (1656) 18084 (896) To construct the MDRs among the ions in an MS/MS spectrum, the possible chemical formulas for each fragment are required. In this study, the PubChem database (Kim et al., 2021 ) ( ftp://ftp.ncbi.nlm.nih.gov/pubchem/Compound/CURRENT-Full/XML/ , downloaded on 2018/11/08) was used to find the possible formulas for each fragment. The database originally comprises 109,383,767 compounds. These compounds were first filtered by the first two criteria (except that both C and H were optional here as fragments may not include these elements) and then isomers were removed, leaving a total of 928,369 reference chemical formulas. For each of the peaks (including the precursor and its fragments) in a spectrum, the possible chemical formulas were searched from the reference formulas within a given matching tolerance. 2.2 Spectral Preprocessing 2.2.1 Denoising, Centroiding, and Deisotoping As previously mentioned, FSA was designed to perform precursor formula ranking using centroid and deisotoped peak lists converted from a MS/MS spectrum in profile mode. However, denoising is usually performed in advance to remove background noise in the spectra and to reduce the computations for centroiding and deisotoping. All these tasks can be performed using commercial software bundled with a mass spectrometer or free software such as msconvert in ProteoWizard (Chambers et al., 2012 ). In fact, most MS/MS spectral databases provide centroid instead of profile peak lists probably due to file size considerations. Nevertheless, some crucial information such as charge state, isotopic ratio, coelution status, and, sometimes, small fragments are lost due to the conversion. 2.2.2 Monoisotopic Mass Restoration The MS/MS spectra were acquired in either positive- or negative-ion mode, and the masses of the fragment peaks in the spectra are represented by m/z values in Thomson (Th). These values need to be converted (for example, via Eq. (1) as shown below) to their corresponding monoisotopic masses in Da before they can be used for mass match with the compounds in the PubChem database. monoisotopic mass = ( m⁄z ) × (charge state) − adduct. (1) Some common adduct values are 1.007276, 22.9892, and 38.9632 Da for [M + H] + , [M + Na] + , and [M + K] + in positive mode and −1.007276, 34.9694, and 78.918885 for [M−H] − , [M + Cl] − , and [M + Br] − in negative mode, respectively. A useful calculator for common adducts can be found at https://fiehnlab.ucdavis.edu/staff/kind/metabolomics/ms-adduct-calculator/ . Notably, multiple-charged ions are likely but rarely seen in metabolite MS/MS spectra. However, isotope information is lost during the deisotoping process; thus, the charge state cannot be evaluated. In this work, for a proof of principle of the designed FSA algorithm, all fragment peaks in a spectrum were considered as single-charged ions. 2.2.3 Contamination and Ambiguous Peak Removal All peaks with masses greater than the precursor mass after mass conversion were removed as they were likely attributed to adducts, multimers, coelution, false centroiding/deisotoping processes, and the presence of contaminants. Furthermore, peaks within [0, −2] Da of the precursor mass were also removed because they were usually caused by false centroiding or deisotoping. An artificial precursor peak was added back to the MS/MS spectrum. Its m/z value was obtained from the precursor ion selection scan and its height was set to the same value as that of the highest peak in the spectrum. The purpose of this manipulation was to avoid inaccurate or missing precursors due to false centroiding, false deisotoping, or high-energy collisions. Finally, all the peak heights in the MS/MS spectrum were normalized such that the largest height was set to 100. To reduce false MDRs caused by ambiguous peaks, all peaks with height less than 0.5 after the normalization were also removed. After all the preprocessing steps, spectra that did not contain any fragment peaks were assigned with the closest formula by conventional mass match. 2.3 Determination of Mass Matching Tolerance A mass matching tolerance shall be determined before the possible chemical formulas can be searched for each peak in a MS/MS spectrum. An appropriate tolerance mostly depends on the accuracy of the mass spectrometer, including instrument types (QTOF, QTOF, OT, ITFT, QFT, etc.), vendors (Agilent, Bruker, Thermo Scientific, Waters, etc.), and years of release. However, such information was generally not available in the databases. Some commonly selected tolerance in the literature ranged from 5 to 25 parts per million (ppm) (Dührkop et al., 2019 ; Ludwig et al., 2020 ; Schymanski et al., 2017 ; Tsugawa et al., 2016 ; Wang et al., 2021 ; Wang et al., 2014 ). In this study, instead of relative tolerances in ppm, absolute tolerances in Da were adopted for spectra from different instruments. The selection was based on the observations that a suitable ppm value for a precursor is sometimes too stringent for its small fragments. In the seven datasets, the mass matching tolerances for QTOF, OT, ITFT, and QFT spectra were found to be 1.3, 1.1, 0.6, and 0.5 milli-Da (mDa), respectively. These values were consistent across datasets excerpted from different databases. Notably, the MAXIS spectra exhibited a lower mass accuracy than other QTOF spectra and its mass matching tolerance was found to be 2.9 mDa instead of 1.3 mDa for other QTOF spectra. These settings will be justified later in the discussion section. The optimal tolerance is instrument dependent. It is suggested that the FSA can be conducted on some known metabolites to probe for the optimal tolerance before applying it to a real experiment. Spectra whose precursor mass matched with only one chemical formula within the matching tolerance were classified as sole candidates without being further analyzed by FSA. In addition, spectra whose precursor does not match with any formula within the matching tolerance were denoted as unidentifiable. However, spectra whose fragments did not match with any formula (thus, no MDR) were assigned with the closest formula to their corresponding precursors by conventional mass match. 2.4 Mother–Daughter Relationship and Formula Ranking Score Given a total of n peaks, including a precursor with monoisotopic mass ( m p ) and n − 1 fragments with monoisotopic masses ( m f 1 , m f 2 , …, m f ( n −1) , n > 1 and m fi < m fj for fi < fj ) in a preprocessed MS/MS spectrum, let there be d possible chemical formulas c p 1 , c p 2 , …, c p d for precursor m p within a mass tolerance ε ( ε is dependent of the instrument accuracy) and d i chemical formulas c fi 1 , c fi 2 , …, c fi di for fragment m fi , 1 ≤ i ≤ (n − 1). Let c fj t ⊃ c fi s denote the MDR between the t -th chemical formula of m fj and the s -th chemical formula of m fi , where m fj < m fi , and the chemical components of c fi s are a subset of those of c fj t . For instance, the MDR C 10 H 15 N 3 O 4 ⊃ C 5 H 7 N 3 O exists as the C, H, N, and O components in C 5 H 7 N 3 O are a subset of C 10 H 15 N 3 O 4 . The precursor formula with the highest ratio of MDRs among all fragment formulas is anticipated to be the most likely formula of the underlying compound. A numerical study will be presented to demonstrate the effectiveness of this point. Therefore, the ratio of MDRs among all fragment formulas is used as a score to rank all chemical formulas of a precursor, as shown in Eq. ( 2 ). $$\\:\\text{s}\\text{c}\\text{o}\\text{r}\\text{e}=\\frac{\\text{N}\\text{u}\\text{m}\\text{b}\\text{e}\\text{r}\\:\\text{o}\\text{f}\\:\\text{M}\\text{D}\\text{R}\\text{s}\\:\\text{a}\\text{m}\\text{o}\\text{n}\\text{g}\\:\\text{p}\\text{e}\\text{a}\\text{k}\\:\\text{f}\\text{o}\\text{r}\\text{m}\\text{u}\\text{l}\\text{a}\\text{s}}{\\left(\\genfrac{}{}{0pt}{}{n}{2}\\right)}$$ 2 where score ∈ [0, 1]. However, multiple formula candidates can receive the same score, particularly when the number of fragments is small. In such cases, the average of absolute differences between the fragment mass and those of their matched formulas is computed for each of the precursor formulas and the one with the smaller value is assigned with a better rank. 2.5 Characteristic Fragments For a MS/MS spectrum that contains many fragments ( n is large), some of which also associated with many formulas ( d i is large, 1 ≤ i ≤ n ) within the given tolerance, the number of all possible MDR combinations can be enormous. To resolve this problem, a limit (MaxComb) is set for the maximum number of MDR combinations. For each precursor formula in a spectrum, if the combination number was greater than MaxComb, the fragment with the most possible formulas was omitted when computing the combinations. The process was continued until the number of combinations was equal or less than the MaxComb value. However, the adequate value for MaxComb is unknown and affects the formula ranking result as well. A numerical study was conducted to explore the relationship between the MaxComb value and the ranking result using the score in Eq. ( 2 ). The percentages of metabolites whose correct formula exhibited the highest score (rank #1) among all the possible formulas were recorded at various MaxComb values of 2 0 , 2 1 , …, 2 8 , and 2 9 for all the seven test datasets. Surprisingly, the percentage of the correct rank#1 precursor formula increased as the MaxComb value was decreased and reached a maximum when the MaxComb value was reduced to 1, as shown in Supplementary Figure S1 . The above result for reaching the highest-ranking percentage as MaxComb = 1 reveals that any fragment exhibiting multiple MDRs with other peaks of higher mass should be removed when evaluating the score of a precursor formula. A real MS/MS spectrum of salbutamol (HMDB ID: HMDB0001937) is used as a demonstration in Fig. 1 , where the smallest fragment (mass = 164.0710) exhibits two MDRs (C 11 H 19 N 4 O 2 → C 9 H 10 NO 2 ) and (C 11 H 19 N 4 O 2 → C 7 H 8 N 4 O) with the precursor formula C 11 H 19 N 4 O 2 ; thus, the fragment should be removed (or the associated MDRs should be neglected) when evaluating the score of the formula. As a consequence, all the remaining fragments are represented by a single formula and hence exhibit unique MDRs with other peaks of higher mass. Such fragments are called characteristic fragments (CFs) in this work because the number of MDRs among them are crucial for the score, and thus the ranking, of the precursor formula. Noted that whether a fragment is a CF or not is precursor formula-dependent. For instance, in Fig. 1 , the smallest fragment is a CF for the precursor formula C 13 H 21 NO 3 but not for C 11 H 19 N 4 O 2 and C 9 H 24 N 2 O 3 P. In addition, the consistent results across datasets in S1 suggest that the capability of the CFs in ranking the correct precursor formula was dataset- and instrument-invariant. 2.6 Formula Subset Analysis With the critical terms defined, the proposed FSA method can be implemented by the following steps: Step 1. Spectral Preprocessing: Given a peak list of m/z values in an MS/MS spectrum, compute the monoisotopic mass of the precursor and those of the fragments. Perform other manipulations as mentioned in the Spectral Preprocessing section if applicable. Step 2. Formula Search: Search for possible chemical formulas for the precursor and for each of the fragments under a given mass matching tolerance. If no formula can be found for the precursor, the spectrum is considered unidentifiable and the ranking score is set to −1. Additionally, if only one formula is found for the precursor, the spectrum is considered identified and the score is computed by Eq. ( 3 ). $$\\:\\text{s}\\text{c}\\text{o}\\text{r}\\text{e}=\\text{m}\\text{a}\\text{x}\\left(0,\\:\\frac{\\text{m}\\text{a}\\text{t}\\text{c}\\text{h}\\text{i}\\text{n}\\text{g}\\:\\text{t}\\text{o}\\text{l}\\text{e}\\text{r}\\text{a}\\text{n}\\text{c}\\text{e}-\\text{m}\\text{a}\\text{s}\\text{s}\\:\\text{d}\\text{i}\\text{f}\\text{f}\\text{e}\\text{r}\\text{e}\\text{n}\\text{c}\\text{e}}{\\text{m}\\text{a}\\text{t}\\text{c}\\text{h}\\text{i}\\text{n}\\text{g}\\:\\text{t}\\text{o}\\text{l}\\text{e}\\text{r}\\text{a}\\text{n}\\text{c}\\text{e}}\\right)$$ 3 Step 3. MDR Search: Search all MDRs between chemical formulas of two peaks. If no MDR can be found among all peaks, the spectrum is noted inapplicable and the chemical formulas of the precursor are found via mass match and ranked by the scores computed by Eq. ( 3 ). Step 4. CF Determination: For each precursor formula, determine its corresponding CFs, and its score is computed by Eq. ( 4 ). $$\\:\\text{s}\\text{c}\\text{o}\\text{r}\\text{e}=\\frac{\\text{N}\\text{u}\\text{m}\\text{b}\\text{e}\\text{r}\\:\\text{o}\\text{f}\\:\\text{u}\\text{n}\\text{i}\\text{q}\\text{u}\\text{e}\\:\\text{M}\\text{D}\\text{R}\\text{s}\\:\\text{a}\\text{m}\\text{o}\\text{n}\\text{g}\\:\\text{C}\\text{F}\\:\\text{f}\\text{o}\\text{r}\\text{m}\\text{u}\\text{l}\\text{a}\\text{s}}{\\left(\\genfrac{}{}{0pt}{}{n}{2}\\right)}+1$$ 4 , where n is the total number of fragments (including the precursor) whose chemical formula can be found and the denominator denotes the number of all the 2-combinations of the n fragments. The additional number one in the formula is used to distinguish from the mass matching scores in Eq. ( 3 ), whose range is [0, 1]. Step 5. Formula Ranking: Rank the precursor formulas using their scores. If ties are encountered, assign a better rank to the formula with the smaller average of absolute mass differences between the fragments and their matched formulas. The MS/MS spectrum of salbutamol in Fig. 1 is used as an example to demonstrate the effectiveness of FSA: The spectrum contains a group of three peaks with m/z values of 218.1191, 190.1234, and 163.0637 Th in negative ion mode (the precursor m/z is not present). After spectral preprocessing, the resultant spectrum consisted of a precursor and three fragments with monoisotopic molecular masses of 239.1521, 219.1264, 191.1307, 164.0710, and 83.0739 Da (i.e., n = 4). The search for possible chemical formulas under a tolerance of 10 ppm revealed accordingly 3, 4, 4, and 3 formulas for the four peaks, respectively. Despite the 3×4×4×3 = 144 possible combinations, only 16 MDRs among the peak formulas were found. The best combination corresponding to the three precursor formulas and their associated MDRs are indicated by arrows. Although the original number of MDRs was the same for the three combinations (all 5 s), some MDRs in the latter two combinations exhibited multiple choices (dashed lines in Fig. 1 ) for the daughter fragment (hence, not a CF). After removing such daughter fragments and the associated ambiguous MDRs, the first combination, which consisted of C 13 H 21 NO 3 , C 13 H 17 NO 2 , C 12 H 17 NO, and C 9 H 10 NO 2 , exhibited the most MDRs and the highest score of 1.8333 (5/6 + 1 by Eq. ( 4 )). The precursor formula of C 13 H 21 NO 3 was indeed the correct formula for salbutamol. The flowchart of the overall formula ranking process is illustrated in Fig. 2 . First, MS/MS spectra were selected using the four criteria mentioned in the Construction of the Test Datasets section. Qualified spectra were preprocessed to convert the protonated/deprotonated peaks to their monoisotopic forms, and peaks with masses larger than that of the precursor were removed. Possible formulas of the remaining peaks in the preprocessed spectra were searched via mass match under a given tolerance. Spectra whose precursor did not display any chemical formula were considered unidentifiable, whereas those with a single formula were regarded as identified. The remaining spectra that did not present any MDR were identified by mass match. Otherwise, the spectra were processed with CF determination and then formula ranking using the scores in Eq. ( 4 ). 2.7 Performance Evaluation and Improvement The performance of FSA was analyzed on various spectra in seven datasets. The analysis results can demonstrate the formula ranking capability of FSA toward a variety of metabolite species, mass ranges, experiment settings, instrument types, and manufacturers. A comparison with a state-of-the-art formula ranking tool was also conducted. However, as mentioned earlier, different MS/MS spectra of a metabolite can be acquired under different experimental conditions, resulting in different fragmentation patterns and different formula rankings from FSA. Several factors including fragment number, variance of fragment masses, and ratio of CFs were analyzed to provide guidelines for a better ranking result. Furthermore, the expandability of FSA for involving more chemical elements such as halogens was tested and several performance improvement ideas were investigated. 3. RESULTS 3.1 Determination of Mass Matching Tolerance To determine a suitable mass matching tolerance for spectra in each dataset, various values ranging from 0.1 to 3 with a step size of 0.1 mDa were tested and the percentage of spectra in each dataset that were assigned with the correct formula was recorded, as shown in Supplementary Figure S2 . The annotated points showing matching tolerance and correct ranking percentage were mostly corresponding to the highest percentage readings, except for the three HMDB datasets. The exceptions were due to situations in which the precursor information was missing in these datasets and therefore the theoretical value was used for mass match. As a result, the correct formulas of these HMDB spectra were always available during mass match, leading to the persistent percentage increase with the decreasing tolerance. The matching tolerances of the three HMDB datasets were instead evaluated directly from their MS/MS spectra. The idea was to measure the major mass differences between the possible precursors in the MS/MS spectra and their theoretical values. To avoid false precursor peaks, only spectra with peaks close (< 3 mDa) to their theoretical precursor values were included and the tolerance was selected at 80% of the qualified spectra. The histograms and their accumulated version of the qualified spectra in the three datasets are shown in Supplementary Figure S3, from which the mass matching tolerances for the three datasets HMDB_QTOF, HMDB_ITFT, and HMDB_QFT, were 1.3, 0.6, and 0.3 mDa, respectively. However, for MS/MS spectra generated from a high-resolution instrument such as QFT, a small mass matching tolerance of 0.3 mDa usually leads to a single candidate from mass match. In fact, in the HMDB_QFT dataset, more than 77.8% of spectra had their chemical formula identified via mass match using the tolerance of 0.3 mDa, leaving a very limited room for the FSA to demonstrate its effectiveness in such spectra. Alternatively, a larger tolerance of 0.5 mDa was used in this dataset such that the percentage of mass matched spectra was considerably reduced from 77.83% to 20.64%, whereas the overall correction percentage slightly decreased from 99.86% to 98.05%, as shown in Supplementary Figure S4. In summary, the mass matching tolerances in FSA for the MAXIS, QTOF, OT, ITFT, and QFT spectra were set to 2.9, 1.3, 1.1, 0.6, and 0.5 mDa, respectively. 3.2 Determination of Effective Spectral Peaks As previously mentioned, some peaks in the MS/MS spectra are attributed to coelutions, false centroiding or deisotoping. Moreover, some unstable fragments may not exist in nature environments and cannot be found in the chemical database. Such fragments can degenerate the performance of FSA. Fortunately, these fragments usually appear in the spectra with low abundances. An experiment was conducted to elucidate the relationship between the usage of the top abundant peaks in the spectra and the corresponding formula ranking performance of FSA. Supplementary Figure S5 shows the experiment result, in which various numbers of the top abundant peaks ranging from 5 to 100 were used in FSA and the performance was represented by the percentage of spectra that were assigned with the correct formula. Although the best performance in the seven datasets occurred at different numbers of top abundant peaks (as annotated in the figure), these numbers were mostly near 30 and the performances around the top-30 abundant peaks in each dataset were similar. The experimental result suggests that, by using only the top-30 abundant peaks, the performance of FSA can be mostly improved or at least maintained in comparison with using all the fragments in the spectra. More importantly, the reduction in peak number can considerably reduce the effort in MDR evaluations, resulting in significant savings in computation time. 3.3 Performance Evaluation The FSA formula ranking process was applied to the seven test datasets using the aforementioned mass matching tolerances, and the top-30 abundant fragments in each spectrum. The performances are summarized in Table 2 , and the detailed results, including the FSA scores and ranking of the correct metabolite of the seven test datasets, are provided in Supplementary File 1. Table 2 Percentage of metabolites/spectra in the seven test datasets with correctly assigned chemical formula Datasets Spectra / Metabolites Total qualified The correct formula is the sole candidate the rank#1 in top-3 ranks in top-5 ranks HMDB_QTOF Spectra 3029 426 (14.06%) 2105 (69.49%) 2677 (88.38%) 2852 (94.16%) Metabolites 821 135 (16.44%) 670 (81.61%) 776 (94.52%) 803 (97.81%) HMDB_ITFT Spectra 2601 1006 (38.68%) 2495 (95.92%) 2596 (99.81%) 2601 (100.00%) Metabolites 242 87 (35.95%) 239 (98.76%) 242 (100.00%) 242 (100.00%) HMDB_QFT Spectra 1439 297 (20.64%) 1411 (98.05%) 1439 (100.00%) 1439 (100.00%) Metabolites 221 49 (22.17%) 221 (100.00%) 221 (100.00%) 221 (100.00%) NPL_MAXIS Spectra 3561 2 (0.06%) 2966 (83.29%) 3342 (93.85%) 3421 (96.07%) Metabolites 2833 2 (0.07%) 2484 (87.68%) 2722 (96.08%) 2765 (97.60%) NPL_QTOF Spectra 1373 48 (3.50%) 939 (68.39%) 1200 (87.40%) 1274 (92.79%) Metabolites 1360 48 (3.53%) 938 (68.97%) 1193 (87.72%) 1265 (93.01%) NPL_OT Spectra 270 13 (4.81%) 249 (92.22%) 267 (98.89%) 268 (99.26%) Metabolites 239 13 (5.44%) 221 (92.47%) 237 (99.16%) 238 (99.58%) LIPIDS_OT Spectra 18084 7491 (41.42%) 12219 (67.57%) 17781 (98.32%) 18075 (99.95%) Metabolites 896 363 (40.51%) 814 (90.85%) 891 (99.44%) 893 (99.67%) Total Spectra 30357 9283 (30.58%) 22384 (73.74%) 29302 (96.52%) 29930 (98.59%) Metabolites 6612 699 (10.57%) 5587 (84.50%) 6282 (95.01%) 6427 (97.20%) The ranking capability of the FSA was represented by three values: rank#1, in top-3 ranks, and in top-5 ranks. Rank#1 shows the number of metabolites/spectra whose correct chemical formulas were ranked first among all the possible candidates by FSA. The “in top-3 ranks” and “in top-5 ranks” mean that the correct formula of the precursor was among the top-3 and top-5 ranked formula candidates, respectively. The rank#1 demonstrates the capability of FSA to pinpoint the correct formula from all the possible candidates, whereas top-3 and top-5 ranks show its capability of reducing the number of possible candidates. Notably, to present the overall formula ranking capability of the proposed process in Table. 2 and to compare with the state-of-the-art performance, the numbers of rank#1, top-3 rank, and top-5 rank in Table 2 include those inapplicable spectra whose correct chemical formulas were determined by mass match under the dataset-specific mass matching tolerance. Moreover, because different MS/MS spectra (e.g., different instrument settings) for a compound may obtain different ranking results, the subsequent test results focused on metabolite instead of spectrum. Nevertheless, some spectra of a metabolite could be found inapplicable, whereas others were applicable. Consequently, the numbers of inapplicable and applicable metabolites do not add up to the number of total metabolites in Table 2 . The issue of which types of spectra are more likely found with the correct formula will be addressed in the Discussion section. The rank#1 data in Table 2 shows that FSA correctly assigned 84.50% of the metabolite (73.74% of the spectra) formulas in the seven test datasets spanning three different data sources of HMDB, GNPS-NIH-NPL, and GNPS-PNNL-LIPIDS. These results demonstrate that by simply evaluating the MDRs of the CFs in a MS/MS spectrum, the underlying chemical formula is highly likely (> 84%) to be identified with rank#1 using FSA. If the top-3 and top-5 ranks were allowed, the percentage would improve to 95.01% and 97.20%, respectively. Although top-3 and top-5 ranks do not actually identify the formula of the underlying compound, they can help rank/reduce the possible candidates, hence reducing the subsequent identification cost. In addition to the accuracy, the precursor formula ranking performed by FSA is fast. The overall computation time of the total 30357 MS/MS spectra in the seven datasets was approximately 16 min 33 s (0.033 s per spectrum) in the Windows 10 environment with an Intel® Core™ i9-10900F CPU. Both formula selection and MDR computation can be achieved via logical operations, making FSA a very effective algorithm for precursor formula ranking. 4. DISCUSSION 4.1 Probable Causes of Datasets with Low Formula Ranking Accuracies The formula ranking accuracies for spectra in the three datasets, HMDB_QTOF, NPL_QTOF, and LIPID_OT, were lower (< 70%) than those of the other four datasets. The causes of the low accuracy were investigated. First, the low-ranking accuracies may be attributed to the low CF ratios in the three datasets. This argument can be inferred from Supplementary Figure S6, where the ratios of informative fragments in the spectra of the seven test datasets are displayed. Recalled that the proposed FSA uses the 30 most abundant fragments in a spectrum for the precursor formula ranking. Supplementary Figure S6(a) shows the ratio of the 30 fragments in a spectrum that can be found to match with at least one chemical formula (hereafter referred to as fragments with formula), whereas Supplementary Figure S6(b) demonstrates the ratio of these fragments with formula that were CFs of the correct precursor formula. As shown in Supplementary Figure S6(a), the fragment-with-formula ratios were rather high, but the CF ratios were among the lowest in HMDB_QTOF and NPL_QTOF. The high fragment-with-formula ratios were most likely caused by low mass accuracy and thus large mass matching tolerance, whereas the low CF ratios mainly resulted from unstable fragments and false peaks. The unstable fragments denote product ions that only exist during high-energy dissociation, not in the natural environment; thus, their formulas are not in the compound database. In contrast, the false peaks are signals in a spectrum that do not represent any compound-related product ions. These peaks generally result from false deisotope/centroiding process, contamination, or coelution. Both unstable fragments and false peaks may be matched with false formulas under large tolerance that usually exhibit no MDR with the precursor, leading to a high fragment-with-formula ratio and a low CF ratio in the corresponding spectrum. Although the spectra in the NPL_MAXIS dataset were also generated from QTOF MS, the differences in collision energy, experiment settings, and instrument manufacturer can lead to different performances. However, the causes of low-ranking accuracy in the LIPID_OT dataset were different from those in the previous two QTOF datasets as the spectra therein exhibited both low fragment-with-formula and low CF ratios. Such phenomena may be due to special functional groups, various branch chain structures, and different double bond locations in the lipids. Particular fragmentation mechanisms (e.g., charge-driven and charge-remote processes) of lipid may provide clues for known and predicted fragments in a spectrum. However, such predicted fragments may not have been included in the database, leading to low fragment-with-formula ratios. However, the high molecular mass of some lipid fragments increased the chance of false formulas during mass match, which in turns reduced the number of MDRs and the CF ratios. Furthermore, in the NPL_QTOF dataset, the formula candidates of the precursor in 81 (5.9%) spectra contained no correct formula within the given mass matching tolerance (i.e., 1.3 mDa). Although such spectra also existed in the NPL_OT and LIPID_OT datasets, the respective percentages were very low (0.74% (2 spectra) and 0.05% (9 spectra, respectively). Moreover, 1141, 5, 5, 3, 218, and 1 spectra were created in the years of 2014, 2017, 2018, 2019, 2020, and 2021; 74 of the 81 problematic spectra were from 2020. The high percentage (33.94%, 74 out of 218) of spectra with large mass differences in 2020 may be due to changes in experiment conditions. However, such large number of unidentifiable spectra was attributed to the low formula assignment accuracy of FSA for this dataset. 4.2 Comparison with Other Formula Assignment Tools Several tools have been proposed for automated metabolite formula assignment/ranking for untargeted metabolomics studies [13–24, 30]. A state-of-the-art tool, SIRIUS (version 5.5.7, downloaded on 2022/10/23) was adopted to compare with FSA. The tool was selected for its top performance in category 2 (best automatic structural identification - in silico fragmentation only) of the CASMI 2016 (Schymanski et al., 2017 ) contest and the top performance in \"correct elemental formulas\" of the CASMI 2022. Notably, although ZODIAC (Ludwig et al., 2019 ) outperformed SIRIUS in the 2016 contest, it is now also implemented in SIRIUS as an option for improvement in formula ranking. The comparisons were conducted among SIRIUS alone, SIRIUS with the ZODIAC option (SIRIUS + ZODIAC), and FSA. In addition, only tools in category 2 were compared as only the MS/MS spectra were given in the contest, which was similar to the situation discussed herein. The comparisons were performed on the seven test datasets, as listed in Table 1 . The same chemical elements were specified in the comparison: CHONSP in the three HMDB datasets, CHONS in the three GNPS-NIH-NPL datasets, and CHONP in the GNPS-PNNL-LIPIDS dataset. However, as the mass matching tolerance in SIRIUS can only be assigned in ppm, various values were tested to determine the best performance of SIRIUS. The comparison result is summarized in Table 3 , where the winner for each dataset is indicated in boldfaced numbers. Table 3 A Comparison of formula assignment accuracy among SIRIUS, SIRIUS + ZODIAC, and FSA Datasets (Total Spectra) SIRIUS Alone SIRIUS + ZODIAC FSA with 30 M.A.P. HMDB_QTOF (3029) 2190 (72.30%) (14 ppm, 13 N/A) 2286 (75.47%) (14 ppm, 13 N/A) 2105 (69.49%) (1.3 mDa) HMDB_ITFT (2601) 2256 (86.74%) (6 ppm) 2318 (89.12%) (6 ppm) 2495 (95.92%) (0.6 mDa) HMDB_QFT (1439) 1372 (95.34%) (3 ppm) 1405 (97.64%) (3 ppm) 1411 (98.05%) (0.5 mDa,) NPL_MAXIS (3561) 2669 (74.95%) (20 ppm, 258 N/A) 2669 (74.95%) (20 ppm, 258 N/A) 2966 (83.29%) (2.9 mDa) NPL_QTOF (1373) 951 (69.26%) (15 ppm, 16 N/A) 1069 (77.86%) (15 ppm, 16 N/A) 939 (68.39%) (1.3 mDa) NPL_OT (270) 252 (93.33%) (7 ppm, 2 N/A) 255 (94.44%) (7 ppm, 2 N/A) 249 (92.22%) (1.1 mDa) LIPIDS_OT (18084) 11212 (62.00%) (18 ppm, 1905 N/A) 11269 (62.31%) (20 ppm, 1905 N/A) 12219 (67.57%) (1.1 mDa) Total (30357) 20902 (68.85%) 21271 (70.07%) 22384 (73.74%) M.A.P. stands for the most abundant peaks; N/A. stands for not applicable due to the 850 Da mass limit of the SIRIUS. Overall, FSA achieved the best formula assignment accuracy (73.74%) among the three tools and SIRIUS + ZODIAC (70.07%) was better than SIRIUS alone (68.85%). Individually, SIRIUS + ZODIAC was better than SIRIUS alone in all datasets and FSA attained similar or better assignment accuracies than both SIRIUS alone and SIRIUS + ZODIAC in five of the seven datasets, except for the two QTOF datasets. The inferior performance of FSA in the QTOF spectra may be due to the low CF ratios in these datasets, as demonstrated in Supplementary Figure S6(b). Although the FSA employed CFs to remove ambiguous MDRs in a rather efficient and effective manner, more sophisticated rules, such as those employed in SIRIUS and ZODIAC, which involve structural information, are needed to remove false MDRs. However, SIRIUS faces difficulties to address compounds with masses larger than 850 Da [21]. Because of this limitation, some compounds in the seven datasets were not applicable (N/A in Table 3 ) to SIRIUS, making it less suitable for large metabolites such as lipids. Furthermore, the assignment result from SIRIUS can exhibit small variations in different runs, which may be due to some statistical approximations in the algorithm. Table 3 also indicates that the selected mass matching tolerances for the SIRIUS + ZODIAC to achieve the best formula assignment accuracy in the seven datasets were 14, 6, 3, 20, 15, 7, and 20 ppm, respectively. Formula assignment accuracies versus various mass matching tolerances for SIRIUS alone and for SIRIUS + ZODIAC are provided in Supplementary Table S1 , whereas the GUI interface of SIRIUS is illustrated in Supplementary Figure S7. These tolerances selected for SIRIUS + ZODIAC are highly correlated (correlation coefficient = 0.77) with those for FSA (i.e., 1.3, 0.6, 0.5, 2.9, 1.3, 1.1, and 1.1 mDa). This phenomenon demonstrates that the matching tolerances used in FSA actually reflected the properties of the spectra in the corresponding dataset rather than were tailored for the performance of FSA. Notably, the tolerances used in SIRIUS + ZODIAC were selected to achieve the best accuracies in the seven datasets such as to compare with those of FSA. By considering consistency in instrument (as in FSA), some of the accuracies of SIRIUS and SIRIUS + ZODIAC were reduced from those in Table 3 . In addition to accuracy, computation time was also tested among the three tools. From Table 4 , FSA was the fastest among the three tools and SIRIUS alone was generally faster than SIRIUS + ZODIAC. In average, FSA was more than 30 times faster than both SIRIUS alone and SIRIUS + ZODIAC. Such a comparison may not seem fair as FSA only imposed the 30 most abundant fragments in a spectrum for the formula assignment, whereas SIRIUS + ZODIAC used all the fragments. However, even if all the fragments are used for the assignment, FSA can still achieve an averaged computation time of 0.072 s per spectrum, which is more than 13 times faster than both SIRIUS alone and SIRIUS + ZODIAC. Table 4 Computation time spent by SIRIUS, SIRIUS + ZODIAC, and FSA on the test datasets Datasets Total Spectra SIRIUS Alone SIRIUS + ZODIAC FSA with 30 M.A.P. HMDB_QTOF 3029 17 min 11 min 1.8 min (108 s) HMDB_ITFT 2601 5 min 15 min 0.8 min (48 s) HMDB_QFT 1439 2 min 3 min 0.5 min (30 s) NPL_MAXIS 3561 171 min 185 min 4.08 min (245 s) NPL_QTOF 1373 10 min 13 min 2.03 min (122 s) NPL_OT 270 1 min 1 min 0.22 min (13 s) LIPIDS_OT 18084 296 min 335 min 7.12 min (427 s) Total 30357 502 min 563 min 16.55 min (993 s) Average 1 0.99 s 1.11 s 0.033 s M.A.P. stands for the most abundant peaks. 4.3 Performance Consistency in Metabolite Mass The formula ranking capability of FSA toward different metabolite masses in the seven test datasets was analyzed. The analysis result is shown in Supplementary Figure S8, in which only the metabolites whose formulas were actually assigned by FSA are displayed. Supplementary Figure S8(a) shows the number of such metabolites in total and in each dataset, whereas Supplementary Figure S8(b) illustrates the percentage of metabolites in (a) that were assigned with the correct formulas. Some of the lines are disconnected at certain mass ranges due to lack of metabolite in those mass ranges. In addition, the metabolites were divided into different mass ranges from 0 to 1500 Da for every 100 Da. The number of metabolites instead of spectra were considered as multiple spectra can correspond to the same metabolite, each with different collision energies, particularly in the HMDB datasets. Some of such spectra may either contain few fragments (low energy) or only very low abundant fragments (high energy), resulting in few MDRs, being not suitable for FSA. As indicated in Supplementary Figure S8(b), the averaged assignment accuracy of FSA (green line) exhibits a rather stable behavior (mostly ≥ 75%) across the entire mass range. The larger variations in the high mass range 1000–1500 Da compared to the low mass range 100–1000 Da may be due to small sample size (< 40) in the datasets. From the aspect of individual dataset, FSA exhibited the best performance in the HMDB_QFT, followed by HMDB_ITFT, NPL_OT, NPL_MAXIS, LIPID_OT, HMDB_QTOF, and NPL_QTOF, which coincides with the data in Table 2 . The worse performance in the two QTOF datasets may be due to the low percentage of CFs in the spectra, as previously discussed. The performance of FSA to large mass (> 600 Da) metabolites in the two QTOF datasets (red and blue lines in Supplementary Figure S8(b)) seemed contradictory. However, the sample sizes were small in these two sets. More test data is required before deriving any conclusion. 4.4 Performance after Involving Halogen Elements As previously mentioned, FSA requires only a list of monoisotopic masses and chemical formulas of all possible compounds to perform precursor formula ranking for a MS/MS spectrum. In the previous tests, compounds consisting of only C, H, O, N, S, and P were considered. In this test, compounds involving additional halogen elements F, Cl, Br, and I were considered. As a result, the original list of 928,369 reference chemical formulas from PubChem was increased to 2,613,203 formulas. In addition, the number of qualified MS/MS spectra in the test datasets was also increased, as presented in Supplementary Table S2 . Notably, the LIPID_OT dataset was removed from this test as no halogen element was found in the dataset. The performances of FSA after including the halogen elements are summarized in Supplementary Table S3. The rank#1 data in the table shows that the proposed FSA correctly identified 76.80% of the metabolite (73.99% of the spectra) formulas in the six test datasets. These results demonstrate that after including the halogen elements in the formula list, FSA can still maintain > 76% accuracy in pinpointing (rank#1) the underlying chemical formula in the seven datasets. If the top-3 and top-5 ranks were used, the percentage would improve to 88.91% and 92.11%, respectively. In comparison with the previous test, the performance of FSA decreased less than 9% after the four halogen elements were added. Notably, FSA exhibited considerable performance reductions in HMDB_QTOF (81.61%→67.56%) and NPL_QTOF (68.59%→59.73%). This may be due to the combination effect of high energy collision and halogen elements in the spectra, producing abundant fragments that were nonexistent in the compound database. If the two QTOF datasets were not considered, the performance dropped from 89.54% (3176 of 3547 metabolites) to 85.99% (3238 of 3770 metabolites), less than 4% of reduction. In addition, the total computation time for the 14169 spectra was 30 min 37 s with an average of 0.13 s per spectrum, which was 3.94 (0.13/0.033) times longer than that of the previous test. However, the increases of 1.67 times (6→10) in element number and 2.81 times (928,369→2,613,203) in reference formula number resulted in increases of 1.67× n and 2.81 times in MDR calculation and formula candidates search, respectively, where n is the additional number of formula candidates from all fragments. With such additional computations and large overhead in passing data (more than half of the computation time), the proposed FSA was still able to process an MS/MS spectrum in a fraction of second and with a small performance decrease of < 9%, showing its capability of performing reliable formula ranking with the expansion of compound compositions. 4.5 Suitable Types of Spectra As MDRs of CFs are both critical to the FSA, the accuracy, amount, and location of the fragment peaks can all contribute toward the final score and rank of a spectrum. The manner fragments should appear in a spectrum for the FSA to provide an accurate formula ranking was also investigated. First, the relationship between the formula ranking capability of the FSA and number of fragments in a spectrum was analyzed. Supplementary Figure S9 shows boxplots illustrating whether the fragment numbers were differently distributed between spectra whose correct formulas were ranked #1 (R1) and those that were otherwise ranked (R2+) in each of the seven test datasets. At the top of each boxplot, a p -value computed by the one-sided Wilcoxon rank sum test is displayed, indicating their differences in the median values as the fragment numbers were not normally distributed ( p < 0.01 by both the Kolmogorov–Smirnov (K–S) test and the Anderson–Darling (A–D) test). The fragment number in the vertical axes is drawn in the logarithmic scale due to the large value range (102–104) among datasets. The median fragment numbers of the R1 spectra were 14, 14, 22, 219, 1276, 30, and 74, whereas those of the R2 + spectra were 13, 13.5, 13, 101, 802, 17, and 66 in the seven test datasets, respectively. Although the fragment numbers in both groups were not significantly different, the fragment numbers in the R1 spectra were higher than those in the R2 + spectra, suggesting that higher fragment numbers in a spectrum may have a higher chance for correct formula assignment by FSA. Next, the distribution of fragment peaks in a spectrum was analyzed based on the standard deviation (SD) of the absolute distances between two neighboring peaks. The precursor peak was added to the spectrum (in case it was missing due to high-energy collision) to distinguish between peak clusters with high and low mass values. The boxplots of the SDs in the rank#1 spectra versus those in the non-rank#1 ones for the seven datasets are presented in Supplementary Figure S10. The median SDs of the R1 spectra by the FSA method were 13.00, 15.03, 12.82, 3.71, 0.77, 13.23, and 22.79, whereas those of the R2 + spectra were 15.98, 22.93, 22.89, 7.45, 1.21, 55.75, and 24.76 in the seven test datasets, respectively. The one-sided Wilcoxon rank sum test showed that the SDs of the neighboring peak distances in the R1 spectra were significantly lower ( p < 0.01) than those in the R2 + spectra in all the seven datasets, suggesting that a lower SD of peak distances in a spectrum may also improve formula assignment performance. The result is understandable because fragment peaks with small inter-peak distance SDs are likely to be more evenly distributed in an MS/MS spectrum. This reflects a higher chance for more MDRs and a higher score for the underlying formula of the spectrum. Two other factors, namely the CF ratio and the difference between the top two scores, were also investigated. Supplementary Figure S11 shows the box plots of the CF ratios in the R1 and R2 + spectra of the seven datasets. The median ratios of the CF to the 30 most abundant fragments with formula were 0.66, 0.80, 0.90, 0.80, 0.66, 0.90, and 0.46 in the R1 spectra and 0.27, 0.25, 0.80, 0.43, 0.44, 0.75 and 0.27 in the R2 + spectra for the seven test datasets, respectively. The one-sided Wilcoxon rank sum test revealed that the CF ratios were significantly higher ( p < 0.01) in the R1 spectra than in the R2 + spectra, suggesting that a higher CF ratio in a spectrum may imply a higher chance for correct precursor formula assignment by FSA. In addition, the difference between the top-two FSA scores represents the effectiveness of the scoring mechanism proposed in (4). Supplementary Figure S12 shows boxplots of the differences between the R1 and the R2 + spectra of the seven datasets. The median differences of the top two FSA scores in the R1 spectra were 0.13, 0.40, 0.47, 0.19, 0.10, 0.24, and 0.0441, whereas those of the R2 + spectra were 0.04, 0.05, 0.13, 0.04, 0.03, 0.03 and 0.0444 in the seven datasets, respectively. The one-sided Wilcoxon rank sum test showed significant ( p < 0.01) differences between the R1 and the R2 + spectra in all seven datasets, suggesting that a higher score difference in a spectrum may indicate a higher chance that the proposed scoring mechanism can identify the correct precursor formula. However, the median value of the R1 group was slightly smaller than that of the R2 + group in the LIPID_OT dataset, which is opposite to what occurred with the other datasets. The exceptions for the LIPID species are probably due to their branch chain structures, resulting in alternative MDRs and reducing the CF ratios (refer to Supplementary Figure S6(b)) and scores. 4.6 Possible Improvements Several experiments were conducted to improve the FSA performance; however, limited progress was achieved. First, the length of the fragmentation path was considered in addition to the sum of MDRs. Many attempts to include additional terms in Eq. ( 4 ), to encourage a long fragmentation path (consecutive MDRs in a series of fragments), were tested. Second, possible formulas were searched for fragment peaks without a match in the PubChem database. Adducts may attach to a compound leading to a different fragment pattern in the spectrum. Some fragments in the spectrum may also be associated with the adducts; hence, their formulas may not be found in the PubChem database. Furthermore, some fragment peaks that were generated during the collision process may not be stable in the natural environment; therefore, their formulas cannot be found in PubChem. Efforts were made to develop the formulas of such fragments by adding/subtracting different known adducts and neutral losses (Ma et al., 2014 ) from fragments with known formulas. Third, combinations of different scoring strategies were also tested. As mentioned earlier, the current scoring method only computes MDRs for fragments with a single formula (i.e., CFs). Although this method achieved the best performance, fragments with multiple possible formulas may still provide clues to identify the spectra. Different weighted combinations of such multi-formula fragments were tested through a genetic algorithm approach. Although some of the above modifications did increase the formula ranking accuracy in some datasets, the improvements were limited and not consistent across all the test datasets. Although several strategies have been tested to improve the formula ranking accuracy in the test datasets, few chemical rules have been implemented to rank possible formulas and filter out unfeasible ones. Such rules can be crucial in certain compound species and in compounds with adducts. In addition, species-related rules can also be applied to improve the identification. As some compound species contain a specific functional group, fragments related to such groups can be used to reduce possible candidates. Including more chemical formula databases, such as ChemSpider, ChEBI, and SciFinder and updating the PubChem database in FSA may also improve the performance of FSA. For example, the current PubChem version (2022/11/27), which consists of 112,419,402 compounds in comparison with the employed version of 109,383,767 compounds, may help FSA increase the fragment-with-formula ratio and the CF ratio so as to analyze spectra with no or few MDRs. Furthermore, instrument manufacturers usually produce new models with better mass accuracy approximately every two years. As the FSA performance improves with increasing mass accuracy, an improved formula ranking capability of the developed FSA method is expected for spectra generated by modern instruments. Declarations Supplementary Information The online version contains supplementary material available at GitHub (https://github.com/kslynn128171/FSA.git). Author contributions KSL conceived the concept, derived the related programs, performed the analyses, and wrote the manuscript. Funding This work was supported by the National Science and Technology Council of Taiwan under Grant No. NSTC114-2113-M-030-011. Data availability The MATLAB® programs that perform FSA are freely available at GitHub (https://github.com/kslynn128171/FSA.git) along with Python programs that extract the formula database from PubChem and the seven test datasets extracted from HMDB and GNPS. Conflict of interest Ke-Shiuan Lynn declares that he has no conflict of interest. Human and animal participants This article does not contain any studies with human and/or animal participants performed by any of the authors. References Blazenovic, I., Kind, T., Ji, J. and Fiehn, O. (2018) Software Tools and Approaches for Compound Identification of LC-MS/MS Data in Metabolomics. Metabolites 8 . Brouard, C., Szafranski, M. and D'Alché-Buc, F. (2016) Input output kernel regression: supervised and semi-supervised structured output prediction with operator-valued kernels. J. Mach. Learn. Res. 17, 6105–6152. Chambers, M.C., Maclean, B., Burke, R., Amodei, D., Ruderman, D.L., Neumann, S., Gatto, L., Fischer, B., Pratt, B., Egertson, J., Hoff, K., Kessner, D., Tasman, N., Shulman, N., Frewen, B., Baker, T.A., Brusniak, M.Y., Paulse, C., Creasy, D., Flashner, L., Kani, K., Moulding, C., Seymour, S.L., Nuwaysir, L.M., Lefebvre, B., Kuhlmann, F., Roark, J., Rainer, P., Detlev, S., Hemenway, T., Huhmer, A., Langridge, J., Connolly, B., Chadick, T., Holly, K., Eckels, J., Deutsch, E.W., Moritz, R.L., Katz, J.E., Agus, D.B., MacCoss, M., Tabb, D.L. and Mallick, P. (2012) A cross-platform toolkit for mass spectrometry and proteomics. Nat Biotechnol 30, 918-20. Dührkop, K., Fleischauer, M., Ludwig, M., Aksenov, A.A., Melnik, A.V., Meusel, M., Dorrestein, P.C., Rousu, J. and Böcker, S. (2019) SIRIUS 4: a rapid tool for turning tandem mass spectra into metabolite structure information. Nature Methods 16, 299-302. Dunn, W.B., Broadhurst, D.I., Atherton, H.J., Goodacre, R. and Griffin, J.L. (2011) Systems level studies of mammalian metabolomes: the roles of mass spectrometry and nuclear magnetic resonance spectroscopy. Chem Soc Rev 40, 387-426. Dunn, W.B., Erban, A., Weber, R.J.M., Creek, D.J., Brown, M., Breitling, R., Hankemeier, T., Goodacre, R., Neumann, S., Kopka, J. and Viant, M.R. (2013) Mass appeal: metabolite identification in mass spectrometry-focused untargeted metabolomics. Metabolomics 9, 44-66. Guijas, C., Montenegro-Burke, J.R., Domingo-Almenara, X., Palermo, A., Warth, B., Hermann, G., Koellensperger, G., Huan, T., Uritboonthai, W., Aisporna, A.E., Wolan, D.W., Spilker, M.E., Benton, H.P. and Siuzdak, G. (2018) METLIN: A Technology Platform for Identifying Knowns and Unknowns. Analytical Chemistry 90, 3156-3164. Heinonen, M., Shen, H., Zamboni, N. and Rousu, J. (2012) Metabolite identification and molecular fingerprint prediction through machine learning. Bioinformatics 28, 2333-2341. Jedrychowski, M.P., Huttlin, E.L., Haas, W., Sowa, M.E., Rad, R. and Gygi, S.P. (2011) Evaluation of HCD- and CID-type fragmentation within their respective detection platforms for murine phosphoproteomics. Mol Cell Proteomics 10, M111 009910. Kim, S., Chen, J., Cheng, T., Gindulyte, A., He, J., He, S., Li, Q., Shoemaker, B.A., Thiessen, P.A., Yu, B., Zaslavsky, L., Zhang, J. and Bolton, E.E. (2021) PubChem in 2021: new data content and improved web interfaces. Nucleic Acids Res 49, D1388-D1395. Kind, T., Liu, K.-H., Lee, D.Y., DeFelice, B., Meissen, J.K. and Fiehn, O. (2013) LipidBlast in silico tandem mass spectrometry database for lipid identification. Nature methods 10, 755-758. Kind, T., Tsugawa, H., Cajka, T., Ma, Y., Lai, Z.J., Mehta, S.S., Wohlgemuth, G., Barupal, D.K., Showalter, M.R., Arita, M. and Fiehn, O. (2018) Identification of small molecules using accurate mass MS/MS search. Mass Spectrometry Reviews 37, 513-532. Liebal, U.W., Phan, A.N.T., Sudhakar, M., Raman, K. and Blank, L.M. (2020) Machine Learning Applications for Mass Spectrometry-Based Metabolomics. Metabolites 10 . Ludwig, M., Nothias, L.-F., Dührkop, K., Koester, I., Fleischauer, M., Hoffmann, M.A., Petras, D., Vargas, F., Morsy, M., Aluwihare, L., Dorrestein, P.C. and Böcker, S. (2019) ZODIAC: database-independent molecular formula annotation using Gibbs sampling reveals unknown small molecules. bioRxiv , 842740. Ludwig, M., Nothias, L.-F., Dührkop, K., Koester, I., Fleischauer, M., Hoffmann, M.A., Petras, D., Vargas, F., Morsy, M., Aluwihare, L., Dorrestein, P.C. and Böcker, S. (2020) Database-independent molecular formula annotation using Gibbs sampling through ZODIAC. Nature Machine Intelligence 2, 629-641. Ma, Y., Kind, T., Yang, D., Leon, C. and Fiehn, O. (2014) MS2Analyzer: A software for small molecule substructure annotations from accurate tandem mass spectra. Anal Chem 86, 10724-31. Peisl, B.Y.L., Schymanski, E.L. and Wilmes, P. (2018) Dark matter in host-microbiome metabolomics: Tackling the unknowns-A review. Anal Chim Acta 1037, 13-27. Ridder, L., van der Hooft, J.J.J., Verhoeven, S., de Vos, R.C.H., Vervoort, J. and Bino, R.J. (2014) In Silico Prediction and Automatic LC–MSn Annotation of Green Tea Metabolites in Urine. Analytical Chemistry 86, 4767-4774. Ruttkies, C., Schymanski, E.L., Wolf, S., Hollender, J. and Neumann, S. (2016) MetFrag relaunched: incorporating strategies beyond in silico fragmentation. Journal of Cheminformatics 8, 3. Schymanski, E.L., Ruttkies, C., Krauss, M., Brouard, C., Kind, T., Duhrkop, K., Allen, F., Vaniya, A., Verdegem, D., Bocker, S., Rousu, J., Shen, H., Tsugawa, H., Sajed, T., Fiehn, O., Ghesquiere, B. and Neumann, S. (2017) Critical Assessment of Small Molecule Identification 2016: automated methods. J Cheminform 9, 22. Shen, H., Dührkop, K., Böcker, S. and Rousu, J. (2014) Metabolite identification through multiple kernel learning on fragmentation trees. Bioinformatics 30, i157-i164. Sumner, L.W., Amberg, A., Barrett, D., Beale, M.H., Beger, R., Daykin, C.A., Fan, T.W., Fiehn, O., Goodacre, R., Griffin, J.L., Hankemeier, T., Hardy, N., Harnly, J., Higashi, R., Kopka, J., Lane, A.N., Lindon, J.C., Marriott, P., Nicholls, A.W., Reily, M.D., Thaden, J.J. and Viant, M.R. (2007) Proposed minimum reporting standards for chemical analysis Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI). Metabolomics 3, 211-221. Tsugawa, H., Kind, T., Nakabayashi, R., Yukihira, D., Tanaka, W., Cajka, T., Saito, K., Fiehn, O. and Arita, M. (2016) Hydrogen Rearrangement Rules: Computational MS/MS Fragmentation and Structure Elucidation Using MS-FINDER Software. Anal Chem 88, 7946-58. Viant, M.R., Kurland, I.J., Jones, M.R. and Dunn, W.B. (2017) How close are we to complete annotation of metabolomes? Current Opinion in Chemical Biology 36, 64-69. Wang, F., Liigand, J., Tian, S., Arndt, D., Greiner, R. and Wishart, D.S. (2021) CFM-ID 4.0: More Accurate ESI-MS/MS Spectral Prediction and Compound Identification. Anal Chem 93, 11692-11700. Wang, M., Carver, J.J., Phelan, V.V., Sanchez, L.M., Garg, N., Peng, Y., Nguyen, D.D., Watrous, J., Kapono, C.A., Luzzatto-Knaan, T., Porto, C., Bouslimani, A., Melnik, A.V., Meehan, M.J., Liu, W.T., Crusemann, M., Boudreau, P.D., Esquenazi, E., Sandoval-Calderon, M., Kersten, R.D., Pace, L.A., Quinn, R.A., Duncan, K.R., Hsu, C.C., Floros, D.J., Gavilan, R.G., Kleigrewe, K., Northen, T., Dutton, R.J., Parrot, D., Carlson, E.E., Aigle, B., Michelsen, C.F., Jelsbak, L., Sohlenkamp, C., Pevzner, P., Edlund, A., McLean, J., Piel, J., Murphy, B.T., Gerwick, L., Liaw, C.C., Yang, Y.L., Humpf, H.U., Maansson, M., Keyzers, R.A., Sims, A.C., Johnson, A.R., Sidebottom, A.M., Sedio, B.E., Klitgaard, A., Larson, C.B., P, C.A.B., Torres-Mendoza, D., Gonzalez, D.J., Silva, D.B., Marques, L.M., Demarque, D.P., Pociute, E., O'Neill, E.C., Briand, E., Helfrich, E.J.N., Granatosky, E.A., Glukhov, E., Ryffel, F., Houson, H., Mohimani, H., Kharbush, J.J., Zeng, Y., Vorholt, J.A., Kurita, K.L., Charusanti, P., McPhail, K.L., Nielsen, K.F., Vuong, L., Elfeki, M., Traxler, M.F., Engene, N., Koyama, N., Vining, O.B., Baric, R., Silva, R.R., Mascuch, S.J., Tomasi, S., Jenkins, S., Macherla, V., Hoffman, T., Agarwal, V., Williams, P.G., Dai, J., Neupane, R., Gurr, J., Rodriguez, A.M.C., Lamsa, A., Zhang, C., Dorrestein, K., Duggan, B.M., Almaliti, J., Allard, P.M., Phapale, P. et al. (2016) Sharing and community curation of mass spectrometry data with Global Natural Products Social Molecular Networking. Nat Biotechnol 34, 828-837. Wang, Y., Kora, G., Bowen, B.P. and Pan, C. (2014) MIDAS: A Database-Searching Algorithm for Metabolite Identification in Metabolomics. Analytical Chemistry 86, 9496-9503. Want, E.J., Wilson, I.D., Gika, H., Theodoridis, G., Plumb, R.S., Shockcor, J., Holmes, E. and Nicholson, J.K. (2010) Global metabolic profiling procedures for urine using UPLC-MS. Nature Protocols 5, 1005-1018. Wishart, D.S., Feunang, Y.D., Marcu, A., Guo, A.C., Liang, K., Vazquez-Fresno, R., Sajed, T., Johnson, D., Li, C., Karu, N., Sayeeda, Z., Lo, E., Assempour, N., Berjanskii, M., Singhal, S., Arndt, D., Liang, Y., Badran, H., Grant, J., Serra-Cayuela, A., Liu, Y., Mandal, R., Neveu, V., Pon, A., Knox, C., Wilson, M., Manach, C. and Scalbert, A. (2018) HMDB 4.0: the human metabolome database for 2018. Nucleic Acids Res 46, D608-D617. Additional Declarations No competing interests reported. Supplementary Files SupplementaryFile1.xlsx SupplementaryFigureandTable.doc Cite Share Download PDF Status: Published Journal Publication published 01 Dec, 2025 Read the published version in Metabolomics → Version 1 posted Editorial decision: Revision requested 07 Oct, 2025 Reviews received at journal 07 Oct, 2025 Reviewers agreed at journal 06 Oct, 2025 Reviews received at journal 05 Oct, 2025 Reviewers agreed at journal 05 Oct, 2025 Reviewers invited by journal 05 Oct, 2025 Editor assigned by journal 03 Oct, 2025 Submission checks completed at journal 03 Oct, 2025 First submitted to journal 02 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-7764688\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":false,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":526160614,\"identity\":\"5a62f614-5a29-4b73-9dfa-b7d9b85c82b9\",\"order_by\":0,\"name\":\"Ke-Shiuan Lynn\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+UlEQVRIie3RsWoCQRCA4TkWzmbw2hOCeYUBy0h8EJsNgmluwdLyQNDmHiCVvkJsrOdYSHWJbSCNEEiVQrFJIcE54YoUrqYT3L8alvmYYgF8vgssrI3Y7qgNoFQKQyzftJvU8eWBcdAXEqRQVIQdpBknLca1lbEkcAYJoa85pmU3mgjhGwtRLSG1GbuIZSb6ME+2JGihkX0T5C4SjDVrIakQtRZC74mQhYMoJDnzZmbVlc5JEiLlKbF5rgjFpwiGWtZ6Zi6EGR8xLr4G/Pp7nNzOPu0WdvdmurT5Ks/umtGkN18Ni+PkTxxkcPhMPhOU/fxj1+fz+a6mPS0mW0ia9aiiAAAAAElFTkSuQmCC\",\"orcid\":\"\",\"institution\":\"Fu Jen Catholic University\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Ke-Shiuan\",\"middleName\":\"\",\"lastName\":\"Lynn\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2025-10-02 08:08:11\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-7764688/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-7764688/v1\",\"draftVersion\":[],\"editorialEvents\":[{\"content\":\"https://doi.org/10.1007/s11306-025-02379-0\",\"type\":\"published\",\"date\":\"2025-12-01T15:57:04+00:00\"}],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":93729883,\"identity\":\"e53ccb2a-f99e-4a60-b91b-aaa7a2bc84ab\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:10\",\"extension\":\"docx\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":264479,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"AutomatedMetaboliteFormulaRankingUsingFormulaSubsetAnalysisforCharacteristicFragmentsinLiquidChromatographyTandemMassSpectrometryLynnsubmit.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/c31790420b5c44bcd9c0635e.docx\"},{\"id\":93731819,\"identity\":\"d3e13bd2-bd56-453c-af07-2a092425533f\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:24:10\",\"extension\":\"json\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":3383,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"bf74faf60a534885a832d0df36f3d980.json\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/f0ad2fded723d93bf595d485.json\"},{\"id\":93729877,\"identity\":\"48d1a397-81a9-4b90-813e-5c37309b3b97\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:10\",\"extension\":\"doc\",\"order_by\":2,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":1674240,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"SupplementaryFigureandTable.doc\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/5f9bb8112fe9f5883e7a7edd.doc\"},{\"id\":93729887,\"identity\":\"c4141364-bdfc-4874-a524-c3382d45b32a\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:11\",\"extension\":\"xlsx\",\"order_by\":3,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":3673276,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"SupplementaryFile1.xlsx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/9e9aa1f390fb6b86a5af37f1.xlsx\"},{\"id\":93729884,\"identity\":\"e8be85ba-75c7-43fb-8fb6-dcd5b3a0bc0a\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:10\",\"extension\":\"xml\",\"order_by\":4,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":174544,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"bf74faf60a534885a832d0df36f3d9801enriched.xml\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/aded67d795730b60d989380a.xml\"},{\"id\":93731816,\"identity\":\"906fe75a-df88-4ea2-8fc8-364c71fd4c93\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:24:10\",\"extension\":\"jpeg\",\"order_by\":5,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":49332,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"floatimage1.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/bfe3686d61651bc29db0180e.jpeg\"},{\"id\":93729878,\"identity\":\"b776d55f-ae5c-4c42-87ec-bd3fe8d1ec1b\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:10\",\"extension\":\"jpeg\",\"order_by\":6,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":99880,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"floatimage2.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/7db53070fcd1ce58de2ee793.jpeg\"},{\"id\":93729886,\"identity\":\"0beb133a-e0e1-4281-8ce0-4252e38e1a2f\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:11\",\"extension\":\"xml\",\"order_by\":9,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":174219,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"bf74faf60a534885a832d0df36f3d9801structuring.xml\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/2c2f2703732bda10a8b98470.xml\"},{\"id\":93729885,\"identity\":\"50153c27-8c5c-48cb-accd-a418b68d49ad\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:10\",\"extension\":\"html\",\"order_by\":10,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":181357,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"earlyproof.html\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/fb1626d6bf3a223edb4d45a0.html\"},{\"id\":93729876,\"identity\":\"70819ede-821e-4d2c-a1db-b777d9bb6dbe\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:10\",\"extension\":\"png\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":34758,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003ePossible chemical formulas of the ion peaks in a salbutamol MS/MS spectrum (HMDB ID: HMDB0001937) and MDRs among the peaks. The numbers in the row of MDRs represent all the MDRs for the corresponding precursor formula, whereas those in the parentheses represent the MDRs to CFs. The numbers in parentheses are used as numerators in (2) to rank the precursor formulas.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Onlinefloatimage1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/bda753bd7628b7258d26bea5.png\"},{\"id\":93732527,\"identity\":\"ddc5a566-b058-48d0-ad86-bba939f40274\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:32:10\",\"extension\":\"png\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":62549,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eFlowchart of the proposed formula ranking process.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Onlinefloatimage2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/6d8ae5cc8116c3259163397f.png\"},{\"id\":97723848,\"identity\":\"9df199d7-de6a-4523-9c3e-6e89944d13d4\",\"added_by\":\"auto\",\"created_at\":\"2025-12-08 16:08:44\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":1735577,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/99ab230d-606d-4ab0-946e-fa58b642098a.pdf\"},{\"id\":93729889,\"identity\":\"7e442b5b-4880-4563-a0ea-dd054218e463\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:16:11\",\"extension\":\"xlsx\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":3673276,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"SupplementaryFile1.xlsx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/e29e693aa09a740ba0c7e0ea.xlsx\"},{\"id\":93731817,\"identity\":\"aeed47d1-72ea-4c66-92fc-821a1d1cf774\",\"added_by\":\"auto\",\"created_at\":\"2025-10-17 02:24:10\",\"extension\":\"doc\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":1674240,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"SupplementaryFigureandTable.doc\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7764688/v1/f28d3a7d4165ab95e65909cf.doc\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"Automated Metabolite Formula Ranking Using Formula Subset Analysis for LC-MS/MS-Based Metabolomics\",\"fulltext\":[{\"header\":\"1 Introduction\",\"content\":\"\\u003cp\\u003eMetabolomics, which analyzes metabolite level changes in biological samples, has become a popular research area in systems biology. It provides an instantaneous snapshot of the physiology of a living system including its interactions with the environment, which is relatively limited in other omics such as genomics and proteomics. Liquid chromatography\\u0026ndash;mass spectrometry (LC-MS) is a commonly used platform for metabolomics studies, mostly because of its wide range of detectable polarity and molecular mass. However, metabolite identification remains a bottleneck in LC-MS-based metabolomics mainly due to component complexity and low database integrity (Dunn et al., \\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e2011\\u003c/span\\u003e; Peisl et al., \\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e2018\\u003c/span\\u003e; Want et al., \\u003cspan citationid=\\\"CR28\\\" class=\\\"CitationRef\\\"\\u003e2010\\u003c/span\\u003e).\\u003c/p\\u003e\\u003cp\\u003eBased on the five confidence levels for compound annotation and identification determined by the Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI), an unknown metabolite is claimed identified with \\u0026ldquo;level-1 confidence\\u0026rdquo; if its 2D structure matches with that of an authentic standard using at least two orthogonal techniques (such as MS/MS spectrum and retention time) (Blazenovic et al., \\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e2018\\u003c/span\\u003e; Sumner et al., \\u003cspan citationid=\\\"CR22\\\" class=\\\"CitationRef\\\"\\u003e2007\\u003c/span\\u003e; Viant et al., \\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e2017\\u003c/span\\u003e). However, the availability of standard metabolites is limited, while some, although available, are very expensive. Moreover, finding the correct standard is generally a trial-and-error process. It is also almost impossible to find standards for novel or modified metabolites. Owing to such difficulties, putative identification (level 2\\u0026ndash;3 confidence) has become acceptable, wherein the peaks of the acquired spectrum are matched with those of spectra in public databases. Nevertheless, even with the continuous efforts in collecting MS/MS spectra in many public databases, the number of collected metabolites is still far less than the demand (Blazenovic et al., \\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e2018\\u003c/span\\u003e; Dunn et al., \\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e2013\\u003c/span\\u003e; Guijas et al., \\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e2018\\u003c/span\\u003e; Kind et al., \\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e2018\\u003c/span\\u003e). Other obstacles to spectrum match include variations in instruments, environment conditions, and sample preparation methods.\\u003c/p\\u003e\\u003cp\\u003eTwo major approaches were adopted for MS/MS spectrum identification. One attempts to predict all possible fragmentation patterns of compounds that agree with the given precursor mass, and then determine the most probable compound according to the matches between the predicted fragments and those in the given spectrum (Kind et al., \\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e2013\\u003c/span\\u003e; Liebal et al., \\u003cspan citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e2020\\u003c/span\\u003e; Ridder et al., \\u003cspan citationid=\\\"CR18\\\" class=\\\"CitationRef\\\"\\u003e2014\\u003c/span\\u003e; Ruttkies et al., \\u003cspan citationid=\\\"CR19\\\" class=\\\"CitationRef\\\"\\u003e2016\\u003c/span\\u003e; Wang et al., \\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e2021\\u003c/span\\u003e; Wang et al., \\u003cspan citationid=\\\"CR27\\\" class=\\\"CitationRef\\\"\\u003e2014\\u003c/span\\u003e). As this spectrum-prediction approach does not rely on the availability of MS/MS spectra in databases or any additional equipment, it is generally more applicable to metabolite identifications than the aforementioned approach. The generation of possible fragmentation patterns demands extensive chemical knowledge, which is usually obtained from specialists or via machine learning (Liebal et al., \\u003cspan citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e2020\\u003c/span\\u003e; Wang et al., \\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e2021\\u003c/span\\u003e). Such knowledge or rules may seem reasonable when a single compound is considered. However, metabolomics concerns a mix of many compounds; thus, under complicated interactions between metabolites and environmental influences, these rules may be subject to change. This leads to differences between predicted and actual spectra. For example, according to the test result of CFM-ID 4.0, the average similarity (measured by dot product) between \\u003cem\\u003ein silico\\u003c/em\\u003e spectra and METLIN spectra is 0.38 (Wang et al., \\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e2021\\u003c/span\\u003e). Furthermore, rules generated from a machine learning approach may not generalize well to unseen compound species. A re-train process is usually needed to account for such data.\\u003c/p\\u003e\\u003cp\\u003eAnother approach is oriented in the opposite direction. The possible chemical and physical features are extracted from an MS spectrum of the compound to be identified and the corresponding chemical structure is identified from databases of known compounds via chemical feature matching (Brouard et al., \\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2016\\u003c/span\\u003e; D\\u0026uuml;hrkop et al., \\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e2019\\u003c/span\\u003e; Heinonen et al., \\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e2012\\u003c/span\\u003e; Ludwig et al., \\u003cspan citationid=\\\"CR15\\\" class=\\\"CitationRef\\\"\\u003e2020\\u003c/span\\u003e; Shen et al., \\u003cspan citationid=\\\"CR21\\\" class=\\\"CitationRef\\\"\\u003e2014\\u003c/span\\u003e; Tsugawa et al., \\u003cspan citationid=\\\"CR23\\\" class=\\\"CitationRef\\\"\\u003e2016\\u003c/span\\u003e). However, this approach usually requires intensive computation and database queries, thus being time-, memory-, and network resource-consuming. For example, according to the test result of SIRIUS, the identification time for 5518 spectra was 5 h 41 min, resulting in an average of 3.71 s per spectrum, not including its long computation time (in days) for compounds with large molecular weight (\\u0026ge; 850 Dalton (Da)) (D\\u0026uuml;hrkop et al., \\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e2019\\u003c/span\\u003e). Similar to the previous approach, features that are determined via machine learning are highly related to the training data and may require a re-train process if an unseen chemical class or species is present. In summary, current metabolite identification/ranking tools exhibit the following seven issues: 1) identification/ranking accuracy, 2) computation time, 3) applicability to unknown compounds, 4) applicability to large metabolites, 5) batch computation capability, 6) result reproducibility, and 7) internet accessibility.\\u003c/p\\u003e\\u003cp\\u003eConsidering these issues, in this work, a simple but effective approach was designed and developed, namely formula subset analysis (FSA), to rank chemical formula candidates during identification in LC-MS/MS-based metabolomics studies. Unlike MS/MS spectrum match, which relies on the availability of MS/MS spectra in public databases, FSA only requires chemical formulas and their corresponding monoisotopic molecular weights that are publicly available in large chemical databases, such as PubChem and ChemSpider. The proposed approach employs a mass match on each of the fragment monoisotopic masses in Da for possible chemical formulas, followed by a series of searches for mother\\u0026ndash;daughter relationships (MDRs) among the fragments. Here, an MDR indicates the relationship between a pair of ions in a spectrum where the chemical compositions of one ion is a subset of those of the other ion. The formula candidate of the precursor ion that comprises the most MDRs among all fragment ion formula candidates, including itself, is considered the most probable identity for the underlying metabolite. Notably, the FSA was designed to provide a ranking for the formula candidates of the underlying compound by using a list of \\u0026ldquo;centroid and deisotoped\\u0026rdquo; peak \\u003cem\\u003em/z\\u003c/em\\u003e values in an MS/MS spectrum. Thus, isomers that share the same molecular formula but have different atomic arrangements in space are not distinguishable by the proposed algorithm. Additionally, the MS/MS spectra of a metabolite acquired from different experimental conditions (e.g., collision energy, association method, temperature, and humidity) can exhibit different fragment peaks. Such variations in the fragmentation pattern may result in differences in the MDRs. FSA was designed to sustain such variation and maintain consistency provided that sufficient and accurate fragment peaks are presented in a spectrum. The designed FSA algorithm was tested on seven experimental datasets. Its formula ranking performance including accuracy, consistency, scalability, and suitable types of spectra for FSA are discussed herein.\\u003c/p\\u003e\"},{\"header\":\"2. Materials and methods\",\"content\":\"\\u003cdiv id=\\\"Sec3\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.1 Construction of the Test Datasets\\u003c/h2\\u003e\\u003cp\\u003eSeven datasets were excerpted from two public databases, Human Metabolome Database (HMDB) (Wishart et al., \\u003cspan citationid=\\\"CR29\\\" class=\\\"CitationRef\\\"\\u003e2018\\u003c/span\\u003e) and Global Natural Products Social Molecular Networking (GNPS) (Wang et al., \\u003cspan citationid=\\\"CR26\\\" class=\\\"CitationRef\\\"\\u003e2016\\u003c/span\\u003e) to evaluate the performance of the proposed algorithm. The HMDB database mainly contains information for small molecule metabolites found in the human body, whereas the GNPS is a repository for sharing of MS/MS spectrometry data from various species, instruments, and omics studies. These seven datasets (three from HMDB and four from GNPS) were constructed to cover metabolites in various species, instruments, and a wild range of molecular weights (0\\u0026ndash;1500 Da).\\u003c/p\\u003e\\u003cp\\u003eThe first three datasets, namely HMDB_QTOF, HMDB_ITFT, and HMDB_QFT, were composed of the experimental MS/MS spectra acquired from LC-ESI-QTOF, LC-ESI-ITFT, and LC-ESI-QFT MS, respectively, in the HMDB (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttp://www.hmdb.ca/downloads\\u003c/span\\u003e\\u003cspan address=\\\"http://www.hmdb.ca/downloads\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e, release 2018-10-19). The LC-ESI-QTOF represents \\u0026ldquo;Liquid Chromatography coupled with Electrospray Ionization-Quadrupole-Time of Flight\\u0026rdquo; whereas the ITFT and QFT in other two instruments stand for Ion Trap-Fourier Transform and Quadrupole Fourier Transform, respectively. Although more spectra acquired from other instruments were also available in HMDB, only those acquired from the above three instruments were used as they had more spectra per instrument with sufficient mass accuracy (the provided \\u003cem\\u003em/z\\u003c/em\\u003e values had at least two decimal digits).\\u003c/p\\u003e\\u003cp\\u003eThe latter four datasets were excerpted from the GNPS database \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://gnps-external.ucsd.edu/gnpslibrary\\u003c/span\\u003e\\u003cspan address=\\\"https://gnps-external.ucsd.edu/gnpslibrary\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e on 2022/06/05. Of the four sets, three were extracted from the GNPS NIH Natural Products Library (both round 1 and round 2, to be referred as GNPS-NIH-NPL) and one was extracted from the GNPS Pacific Northwest National Lab Lipids Library (referred to as GNPS-PNNL-LIPIDS). The three datasets NPL_MAXIS, NPL_QTOF, and NPL_OT extracted from GNPS-NIH-NPL were constructed for three different instruments, LC-ESI-MAXIS, LC-ESI-QTOF and LC-ESI-OT (OT for orbitrap), respectively. Although MAXIS is actually a QTOF (Maxis II HD Q-TOF Bruker \\u0026trade;), it generated spectra with a lower mass accuracy than other QTOF spectra probably due to an older model. The last dataset, LIPID_OT extracted from GNPS-PNNL-LIPIDS, was acquired from LC-ESI-OT. Notably, two fragmentation types, collision-induced dissociation (CID) and higher energy collisional dissociation (HCD), of MS/MS spectra were provided for each compound in this dataset. To avoid redundancy, the HCD spectra were used because they usually exhibit more extensive fragmentation and thus more structural information about the ions (Jedrychowski et al., \\u003cspan citationid=\\\"CR9\\\" class=\\\"CitationRef\\\"\\u003e2011\\u003c/span\\u003e). In addition, for a proof-of-concept study and consistency across datasets, the following four criteria were imposed to all datasets to select qualified spectra:\\u003c/p\\u003e\\u003cp\\u003e\\u003col\\u003e\\u003cspan\\u003e\\u003cli\\u003e\\u003cp\\u003eCompounds are combinations of only C, H, O, N, S, and P where C and H are necessary elements.\\u003c/p\\u003e\\u003c/li\\u003e\\u003c/span\\u003e\\u003cspan\\u003e\\u003cli\\u003e\\u003cp\\u003eMolecular monoisotopic mass is \\u0026lt;\\u0026thinsp;1500 Da.\\u003c/p\\u003e\\u003c/li\\u003e\\u003c/span\\u003e\\u003cspan\\u003e\\u003cli\\u003e\\u003cp\\u003eThe precursor type is either [M\\u0026thinsp;+\\u0026thinsp;H]\\u0026thinsp;+\\u0026thinsp;or [M\\u0026thinsp;\\u0026minus;\\u0026thinsp;H]\\u0026minus;.\\u003c/p\\u003e\\u003c/li\\u003e\\u003c/span\\u003e\\u003cspan\\u003e\\u003cli\\u003e\\u003cp\\u003eAt least three peaks are present in the centroid and deisotoped MS/MS spectrum.\\u003c/p\\u003e\\u003c/li\\u003e\\u003c/span\\u003e\\u003c/ol\\u003e\\u003c/p\\u003e\\u003cp\\u003eThe first two criteria are general definitions for metabolites. Notably, the compounds in GNPS-NIH-NPL did not involve P, whereas those in GNPS-PNNL-LIPIDS did not involve S. These elements were removed from the first criterion when candidate chemical formulas were searched in the corresponding datasets. The third criterion was added to simplify the composition in a spectrum. Spectra of metabolites with adduct such as Na and K will be addressed in the discussion section. The fourth criterion was set for sufficient fragments in a spectrum for its identification. As a result, the number of qualified spectra and their corresponding number of compounds in the seven datasets are listed in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eNumber of spectra and compounds of the seven datasets after filtering using the established criteria\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"4\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eDataset\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eSource Instrument\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eOriginal\\u003c/p\\u003e\\u003cp\\u003eSpectra (Compounds)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eQualified\\u003c/p\\u003e\\u003cp\\u003eSpectra (Compounds)\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"2\\\" rowspan=\\\"3\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_ITFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-QTOF\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e3743 (995)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e3029 (821)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-ITFT\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e4074 (314)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e2601 (242)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-QFT\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e2195 (334)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e1439 (221)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"2\\\" rowspan=\\\"3\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_MAXIS\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-MAXIS\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e7641 (3624)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e3561 (2833)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-QTOF\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e2090 (1890)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e1373 (1360)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-OT\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e1314 (639)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e270 (239)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eLIPIDS_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eLC-ESI-OT (HCD)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e46724 (1656)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e18084 (896)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eTo construct the MDRs among the ions in an MS/MS spectrum, the possible chemical formulas for each fragment are required. In this study, the PubChem database (Kim et al., \\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e2021\\u003c/span\\u003e) (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003eftp://ftp.ncbi.nlm.nih.gov/pubchem/Compound/CURRENT-Full/XML/\\u003c/span\\u003e\\u003cspan address=\\\"http://ftp://ftp.ncbi.nlm.nih.gov/pubchem/Compound/CURRENT-Full/XML/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e, downloaded on 2018/11/08) was used to find the possible formulas for each fragment. The database originally comprises 109,383,767 compounds. These compounds were first filtered by the first two criteria (except that both C and H were optional here as fragments may not include these elements) and then isomers were removed, leaving a total of 928,369 reference chemical formulas. For each of the peaks (including the precursor and its fragments) in a spectrum, the possible chemical formulas were searched from the reference formulas within a given matching tolerance.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec4\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.2 Spectral Preprocessing\\u003c/h2\\u003e\\u003cdiv id=\\\"Sec5\\\" class=\\\"Section3\\\"\\u003e\\u003ch2\\u003e2.2.1 Denoising, Centroiding, and Deisotoping\\u003c/h2\\u003e\\u003cp\\u003eAs previously mentioned, FSA was designed to perform precursor formula ranking using centroid and deisotoped peak lists converted from a MS/MS spectrum in profile mode. However, denoising is usually performed in advance to remove background noise in the spectra and to reduce the computations for centroiding and deisotoping. All these tasks can be performed using commercial software bundled with a mass spectrometer or free software such as msconvert in ProteoWizard (Chambers et al., \\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e2012\\u003c/span\\u003e). In fact, most MS/MS spectral databases provide centroid instead of profile peak lists probably due to file size considerations. Nevertheless, some crucial information such as charge state, isotopic ratio, coelution status, and, sometimes, small fragments are lost due to the conversion.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec6\\\" class=\\\"Section3\\\"\\u003e\\u003ch2\\u003e2.2.2 Monoisotopic Mass Restoration\\u003c/h2\\u003e\\u003cp\\u003eThe MS/MS spectra were acquired in either positive- or negative-ion mode, and the masses of the fragment peaks in the spectra are represented by \\u003cem\\u003em/z\\u003c/em\\u003e values in Thomson (Th). These values need to be converted (for example, via Eq.\\u0026nbsp;(1) as shown below) to their corresponding monoisotopic masses in Da before they can be used for mass match with the compounds in the PubChem database.\\u003c/p\\u003e\\u003cp\\u003emonoisotopic mass = (\\u003cem\\u003em\\u0026frasl;z\\u003c/em\\u003e) \\u0026times; (charge state) \\u0026minus; adduct. (1)\\u003c/p\\u003e\\u003cp\\u003eSome common adduct values are 1.007276, 22.9892, and 38.9632 Da for [M\\u0026thinsp;+\\u0026thinsp;H]\\u003csup\\u003e+\\u003c/sup\\u003e, [M\\u0026thinsp;+\\u0026thinsp;Na]\\u003csup\\u003e+\\u003c/sup\\u003e, and [M\\u0026thinsp;+\\u0026thinsp;K]\\u0026thinsp;+\\u0026thinsp;in positive mode and \\u0026minus;1.007276, 34.9694, and 78.918885 for [M\\u0026minus;H]\\u003csup\\u003e\\u0026minus;\\u003c/sup\\u003e, [M\\u0026thinsp;+\\u0026thinsp;Cl]\\u003csup\\u003e\\u0026minus;\\u003c/sup\\u003e, and [M\\u0026thinsp;+\\u0026thinsp;Br]\\u003csup\\u003e\\u0026minus;\\u003c/sup\\u003e in negative mode, respectively. A useful calculator for common adducts can be found at \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://fiehnlab.ucdavis.edu/staff/kind/metabolomics/ms-adduct-calculator/\\u003c/span\\u003e\\u003cspan address=\\\"https://fiehnlab.ucdavis.edu/staff/kind/metabolomics/ms-adduct-calculator/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e. Notably, multiple-charged ions are likely but rarely seen in metabolite MS/MS spectra. However, isotope information is lost during the deisotoping process; thus, the charge state cannot be evaluated. In this work, for a proof of principle of the designed FSA algorithm, all fragment peaks in a spectrum were considered as single-charged ions.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec7\\\" class=\\\"Section3\\\"\\u003e\\u003ch2\\u003e2.2.3 Contamination and Ambiguous Peak Removal\\u003c/h2\\u003e\\u003cp\\u003eAll peaks with masses greater than the precursor mass after mass conversion were removed as they were likely attributed to adducts, multimers, coelution, false centroiding/deisotoping processes, and the presence of contaminants. Furthermore, peaks within [0, \\u0026minus;2] Da of the precursor mass were also removed because they were usually caused by false centroiding or deisotoping. An artificial precursor peak was added back to the MS/MS spectrum. Its \\u003cem\\u003em/z\\u003c/em\\u003e value was obtained from the precursor ion selection scan and its height was set to the same value as that of the highest peak in the spectrum. The purpose of this manipulation was to avoid inaccurate or missing precursors due to false centroiding, false deisotoping, or high-energy collisions. Finally, all the peak heights in the MS/MS spectrum were normalized such that the largest height was set to 100. To reduce false MDRs caused by ambiguous peaks, all peaks with height less than 0.5 after the normalization were also removed. After all the preprocessing steps, spectra that did not contain any fragment peaks were assigned with the closest formula by conventional mass match.\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec8\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.3 Determination of Mass Matching Tolerance\\u003c/h2\\u003e\\u003cp\\u003eA mass matching tolerance shall be determined before the possible chemical formulas can be searched for each peak in a MS/MS spectrum. An appropriate tolerance mostly depends on the accuracy of the mass spectrometer, including instrument types (QTOF, QTOF, OT, ITFT, QFT, etc.), vendors (Agilent, Bruker, Thermo Scientific, Waters, etc.), and years of release. However, such information was generally not available in the databases. Some commonly selected tolerance in the literature ranged from 5 to 25 parts per million (ppm) (D\\u0026uuml;hrkop et al., \\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e2019\\u003c/span\\u003e; Ludwig et al., \\u003cspan citationid=\\\"CR15\\\" class=\\\"CitationRef\\\"\\u003e2020\\u003c/span\\u003e; Schymanski et al., \\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e2017\\u003c/span\\u003e; Tsugawa et al., \\u003cspan citationid=\\\"CR23\\\" class=\\\"CitationRef\\\"\\u003e2016\\u003c/span\\u003e; Wang et al., \\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e2021\\u003c/span\\u003e; Wang et al., \\u003cspan citationid=\\\"CR27\\\" class=\\\"CitationRef\\\"\\u003e2014\\u003c/span\\u003e).\\u003c/p\\u003e\\u003cp\\u003eIn this study, instead of relative tolerances in ppm, absolute tolerances in Da were adopted for spectra from different instruments. The selection was based on the observations that a suitable ppm value for a precursor is sometimes too stringent for its small fragments. In the seven datasets, the mass matching tolerances for QTOF, OT, ITFT, and QFT spectra were found to be 1.3, 1.1, 0.6, and 0.5 milli-Da (mDa), respectively. These values were consistent across datasets excerpted from different databases. Notably, the MAXIS spectra exhibited a lower mass accuracy than other QTOF spectra and its mass matching tolerance was found to be 2.9 mDa instead of 1.3 mDa for other QTOF spectra. These settings will be justified later in the discussion section. The optimal tolerance is instrument dependent. It is suggested that the FSA can be conducted on some known metabolites to probe for the optimal tolerance before applying it to a real experiment.\\u003c/p\\u003e\\u003cp\\u003eSpectra whose precursor mass matched with only one chemical formula within the matching tolerance were classified as sole candidates without being further analyzed by FSA. In addition, spectra whose precursor does not match with any formula within the matching tolerance were denoted as unidentifiable. However, spectra whose fragments did not match with any formula (thus, no MDR) were assigned with the closest formula to their corresponding precursors by conventional mass match.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec9\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.4 Mother\\u0026ndash;Daughter Relationship and Formula Ranking Score\\u003c/h2\\u003e\\u003cp\\u003eGiven a total of \\u003cem\\u003en\\u003c/em\\u003e peaks, including a precursor with monoisotopic mass (\\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ep\\u003c/em\\u003e\\u003c/sub\\u003e) and n \\u0026minus; 1 fragments with monoisotopic masses (\\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ef\\u003c/em\\u003e1\\u003c/sub\\u003e, \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ef\\u003c/em\\u003e2\\u003c/sub\\u003e, \\u0026hellip;, \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ef\\u003c/em\\u003e(\\u003cem\\u003en\\u003c/em\\u003e\\u0026minus;1)\\u003c/sub\\u003e, \\u003cem\\u003en\\u003c/em\\u003e\\u0026thinsp;\\u0026gt;\\u0026thinsp;1 and \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sub\\u003e \\u0026lt; \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efj\\u003c/em\\u003e\\u003c/sub\\u003e for \\u003cem\\u003efi\\u003c/em\\u003e\\u0026thinsp;\\u0026lt;\\u0026thinsp;\\u003cem\\u003efj\\u003c/em\\u003e) in a preprocessed MS/MS spectrum, let there be \\u003cem\\u003ed\\u003c/em\\u003e possible chemical formulas \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003ep\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e1\\u003c/sub\\u003e, \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003ep\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e2\\u003c/sub\\u003e, \\u0026hellip;, \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003ep\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e\\u003cem\\u003ed\\u003c/em\\u003e\\u003c/sub\\u003e for precursor \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ep\\u003c/em\\u003e\\u003c/sub\\u003e within a mass tolerance \\u003cem\\u003eε\\u003c/em\\u003e (\\u003cem\\u003eε\\u003c/em\\u003e is dependent of the instrument accuracy) and \\u003cem\\u003ed\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ei\\u003c/em\\u003e\\u003c/sub\\u003e chemical formulas \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e1\\u003c/sub\\u003e, \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e2\\u003c/sub\\u003e, \\u0026hellip;, \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e\\u003cem\\u003edi\\u003c/em\\u003e\\u003c/sub\\u003e for fragment \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sub\\u003e, 1 \\u0026le; i \\u0026le; (n \\u0026minus; 1). Let \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efj\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e\\u003cem\\u003et\\u003c/em\\u003e\\u003c/sub\\u003e \\u0026sup; \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e\\u003cem\\u003es\\u003c/em\\u003e\\u003c/sub\\u003e denote the MDR between the \\u003cem\\u003et\\u003c/em\\u003e-th chemical formula of \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efj\\u003c/em\\u003e\\u003c/sub\\u003e and the \\u003cem\\u003es\\u003c/em\\u003e-th chemical formula of \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sub\\u003e, where \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efj\\u003c/em\\u003e\\u003c/sub\\u003e \\u0026lt; \\u003cem\\u003em\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sub\\u003e, and the chemical components of \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efi\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e\\u003cem\\u003es\\u003c/em\\u003e\\u003c/sub\\u003e are a subset of those of \\u003cem\\u003ec\\u003c/em\\u003e\\u003csup\\u003e\\u003cem\\u003efj\\u003c/em\\u003e\\u003c/sup\\u003e\\u003csub\\u003e\\u003cem\\u003et\\u003c/em\\u003e\\u003c/sub\\u003e. For instance, the MDR C\\u003csub\\u003e10\\u003c/sub\\u003eH\\u003csub\\u003e15\\u003c/sub\\u003eN\\u003csub\\u003e3\\u003c/sub\\u003eO\\u003csub\\u003e4\\u003c/sub\\u003e \\u0026sup; C\\u003csub\\u003e5\\u003c/sub\\u003eH\\u003csub\\u003e7\\u003c/sub\\u003eN\\u003csub\\u003e3\\u003c/sub\\u003eO exists as the C, H, N, and O components in C\\u003csub\\u003e5\\u003c/sub\\u003eH\\u003csub\\u003e7\\u003c/sub\\u003eN\\u003csub\\u003e3\\u003c/sub\\u003eO are a subset of C\\u003csub\\u003e10\\u003c/sub\\u003eH\\u003csub\\u003e15\\u003c/sub\\u003eN\\u003csub\\u003e3\\u003c/sub\\u003eO\\u003csub\\u003e4\\u003c/sub\\u003e.\\u003c/p\\u003e\\u003cp\\u003eThe precursor formula with the highest ratio of MDRs among all fragment formulas is anticipated to be the most likely formula of the underlying compound. A numerical study will be presented to demonstrate the effectiveness of this point. Therefore, the ratio of MDRs among all fragment formulas is used as a score to rank all chemical formulas of a precursor, as shown in Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ1\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e).\\u003cdiv id=\\\"Equ1\\\" class=\\\"Equation\\\"\\u003e\\u003cdiv format=\\\"TEX\\\" class=\\\"mathdisplay\\\" id=\\\"FileID_Equ1\\\" name=\\\"EquationSource\\\"\\u003e\\n$$\\\\:\\\\text{s}\\\\text{c}\\\\text{o}\\\\text{r}\\\\text{e}=\\\\frac{\\\\text{N}\\\\text{u}\\\\text{m}\\\\text{b}\\\\text{e}\\\\text{r}\\\\:\\\\text{o}\\\\text{f}\\\\:\\\\text{M}\\\\text{D}\\\\text{R}\\\\text{s}\\\\:\\\\text{a}\\\\text{m}\\\\text{o}\\\\text{n}\\\\text{g}\\\\:\\\\text{p}\\\\text{e}\\\\text{a}\\\\text{k}\\\\:\\\\text{f}\\\\text{o}\\\\text{r}\\\\text{m}\\\\text{u}\\\\text{l}\\\\text{a}\\\\text{s}}{\\\\left(\\\\genfrac{}{}{0pt}{}{n}{2}\\\\right)}$$\\u003c/div\\u003e\\u003cdiv class=\\\"EquationNumber\\\"\\u003e2\\u003c/div\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003ewhere score \\u0026isin; [0, 1]. However, multiple formula candidates can receive the same score, particularly when the number of fragments is small. In such cases, the average of absolute differences between the fragment mass and those of their matched formulas is computed for each of the precursor formulas and the one with the smaller value is assigned with a better rank.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec10\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.5 Characteristic Fragments\\u003c/h2\\u003e\\u003cp\\u003eFor a MS/MS spectrum that contains many fragments (\\u003cem\\u003en\\u003c/em\\u003e is large), some of which also associated with many formulas (\\u003cem\\u003ed\\u003c/em\\u003e\\u003csub\\u003e\\u003cem\\u003ei\\u003c/em\\u003e\\u003c/sub\\u003e is large, 1 \\u0026le; \\u003cem\\u003ei\\u003c/em\\u003e \\u0026le; \\u003cem\\u003en\\u003c/em\\u003e) within the given tolerance, the number of all possible MDR combinations can be enormous. To resolve this problem, a limit (MaxComb) is set for the maximum number of MDR combinations. For each precursor formula in a spectrum, if the combination number was greater than MaxComb, the fragment with the most possible formulas was omitted when computing the combinations. The process was continued until the number of combinations was equal or less than the MaxComb value.\\u003c/p\\u003e\\u003cp\\u003eHowever, the adequate value for MaxComb is unknown and affects the formula ranking result as well. A numerical study was conducted to explore the relationship between the MaxComb value and the ranking result using the score in Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ1\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e). The percentages of metabolites whose correct formula exhibited the highest score (rank #1) among all the possible formulas were recorded at various MaxComb values of 2\\u003csup\\u003e0\\u003c/sup\\u003e, 2\\u003csup\\u003e1\\u003c/sup\\u003e, \\u0026hellip;, 2\\u003csup\\u003e8\\u003c/sup\\u003e, and 2\\u003csup\\u003e9\\u003c/sup\\u003e for all the seven test datasets. Surprisingly, the percentage of the correct rank#1 precursor formula increased as the MaxComb value was decreased and reached a maximum when the MaxComb value was reduced to 1, as shown in Supplementary Figure \\u003cspan refid=\\\"MOESM1\\\" class=\\\"InternalRef\\\"\\u003eS1\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003eThe above result for reaching the highest-ranking percentage as MaxComb\\u0026thinsp;=\\u0026thinsp;1 reveals that any fragment exhibiting multiple MDRs with other peaks of higher mass should be removed when evaluating the score of a precursor formula. A real MS/MS spectrum of salbutamol (HMDB ID: HMDB0001937) is used as a demonstration in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e, where the smallest fragment (mass\\u0026thinsp;=\\u0026thinsp;164.0710) exhibits two MDRs (C\\u003csub\\u003e11\\u003c/sub\\u003eH\\u003csub\\u003e19\\u003c/sub\\u003eN\\u003csub\\u003e4\\u003c/sub\\u003eO\\u003csub\\u003e2\\u003c/sub\\u003e \\u0026rarr; C\\u003csub\\u003e9\\u003c/sub\\u003eH\\u003csub\\u003e10\\u003c/sub\\u003eNO\\u003csub\\u003e2\\u003c/sub\\u003e) and (C\\u003csub\\u003e11\\u003c/sub\\u003eH\\u003csub\\u003e19\\u003c/sub\\u003eN\\u003csub\\u003e4\\u003c/sub\\u003eO\\u003csub\\u003e2\\u003c/sub\\u003e \\u0026rarr; C\\u003csub\\u003e7\\u003c/sub\\u003eH\\u003csub\\u003e8\\u003c/sub\\u003eN\\u003csub\\u003e4\\u003c/sub\\u003eO) with the precursor formula C\\u003csub\\u003e11\\u003c/sub\\u003eH\\u003csub\\u003e19\\u003c/sub\\u003eN\\u003csub\\u003e4\\u003c/sub\\u003eO\\u003csub\\u003e2\\u003c/sub\\u003e; thus, the fragment should be removed (or the associated MDRs should be neglected) when evaluating the score of the formula. As a consequence, all the remaining fragments are represented by a single formula and hence exhibit unique MDRs with other peaks of higher mass. Such fragments are called characteristic fragments (CFs) in this work because the number of MDRs among them are crucial for the score, and thus the ranking, of the precursor formula. Noted that whether a fragment is a CF or not is precursor formula-dependent. For instance, in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e, the smallest fragment is a CF for the precursor formula C\\u003csub\\u003e13\\u003c/sub\\u003eH\\u003csub\\u003e21\\u003c/sub\\u003eNO\\u003csub\\u003e3\\u003c/sub\\u003e but not for C\\u003csub\\u003e11\\u003c/sub\\u003eH\\u003csub\\u003e19\\u003c/sub\\u003eN\\u003csub\\u003e4\\u003c/sub\\u003eO\\u003csub\\u003e2\\u003c/sub\\u003e and C\\u003csub\\u003e9\\u003c/sub\\u003eH\\u003csub\\u003e24\\u003c/sub\\u003eN\\u003csub\\u003e2\\u003c/sub\\u003eO\\u003csub\\u003e3\\u003c/sub\\u003eP. In addition, the consistent results across datasets in S1 suggest that the capability of the CFs in ranking the correct precursor formula was dataset- and instrument-invariant.\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec11\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.6 Formula Subset Analysis\\u003c/h2\\u003e\\u003cp\\u003eWith the critical terms defined, the proposed FSA method can be implemented by the following steps:\\u003c/p\\u003e\\u003cp\\u003eStep 1. Spectral Preprocessing: Given a peak list of \\u003cem\\u003em/z\\u003c/em\\u003e values in an MS/MS spectrum, compute the monoisotopic mass of the precursor and those of the fragments. Perform other manipulations as mentioned in the \\u003cem\\u003eSpectral Preprocessing\\u003c/em\\u003e section if applicable.\\u003c/p\\u003e\\u003cp\\u003eStep 2. Formula Search: Search for possible chemical formulas for the precursor and for each of the fragments under a given mass matching tolerance. If no formula can be found for the precursor, the spectrum is considered unidentifiable and the ranking score is set to \\u0026minus;1. Additionally, if only one formula is found for the precursor, the spectrum is considered identified and the score is computed by Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ2\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e).\\u003cdiv id=\\\"Equ2\\\" class=\\\"Equation\\\"\\u003e\\u003cdiv format=\\\"TEX\\\" class=\\\"mathdisplay\\\" id=\\\"FileID_Equ2\\\" name=\\\"EquationSource\\\"\\u003e\\n$$\\\\:\\\\text{s}\\\\text{c}\\\\text{o}\\\\text{r}\\\\text{e}=\\\\text{m}\\\\text{a}\\\\text{x}\\\\left(0,\\\\:\\\\frac{\\\\text{m}\\\\text{a}\\\\text{t}\\\\text{c}\\\\text{h}\\\\text{i}\\\\text{n}\\\\text{g}\\\\:\\\\text{t}\\\\text{o}\\\\text{l}\\\\text{e}\\\\text{r}\\\\text{a}\\\\text{n}\\\\text{c}\\\\text{e}-\\\\text{m}\\\\text{a}\\\\text{s}\\\\text{s}\\\\:\\\\text{d}\\\\text{i}\\\\text{f}\\\\text{f}\\\\text{e}\\\\text{r}\\\\text{e}\\\\text{n}\\\\text{c}\\\\text{e}}{\\\\text{m}\\\\text{a}\\\\text{t}\\\\text{c}\\\\text{h}\\\\text{i}\\\\text{n}\\\\text{g}\\\\:\\\\text{t}\\\\text{o}\\\\text{l}\\\\text{e}\\\\text{r}\\\\text{a}\\\\text{n}\\\\text{c}\\\\text{e}}\\\\right)$$\\u003c/div\\u003e\\u003cdiv class=\\\"EquationNumber\\\"\\u003e3\\u003c/div\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eStep 3. MDR Search: Search all MDRs between chemical formulas of two peaks. If no MDR can be found among all peaks, the spectrum is noted inapplicable and the chemical formulas of the precursor are found via mass match and ranked by the scores computed by Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ2\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e).\\u003c/p\\u003e\\u003cp\\u003eStep 4. CF Determination: For each precursor formula, determine its corresponding CFs, and its score is computed by Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ3\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e).\\u003cdiv id=\\\"Equ3\\\" class=\\\"Equation\\\"\\u003e\\u003cdiv format=\\\"TEX\\\" class=\\\"mathdisplay\\\" id=\\\"FileID_Equ3\\\" name=\\\"EquationSource\\\"\\u003e\\n$$\\\\:\\\\text{s}\\\\text{c}\\\\text{o}\\\\text{r}\\\\text{e}=\\\\frac{\\\\text{N}\\\\text{u}\\\\text{m}\\\\text{b}\\\\text{e}\\\\text{r}\\\\:\\\\text{o}\\\\text{f}\\\\:\\\\text{u}\\\\text{n}\\\\text{i}\\\\text{q}\\\\text{u}\\\\text{e}\\\\:\\\\text{M}\\\\text{D}\\\\text{R}\\\\text{s}\\\\:\\\\text{a}\\\\text{m}\\\\text{o}\\\\text{n}\\\\text{g}\\\\:\\\\text{C}\\\\text{F}\\\\:\\\\text{f}\\\\text{o}\\\\text{r}\\\\text{m}\\\\text{u}\\\\text{l}\\\\text{a}\\\\text{s}}{\\\\left(\\\\genfrac{}{}{0pt}{}{n}{2}\\\\right)}+1$$\\u003c/div\\u003e\\u003cdiv class=\\\"EquationNumber\\\"\\u003e4\\u003c/div\\u003e\\u003c/div\\u003e,\\u003c/p\\u003e\\u003cp\\u003ewhere \\u003cem\\u003en\\u003c/em\\u003e is the total number of fragments (including the precursor) whose chemical formula can be found and the denominator denotes the number of all the 2-combinations of the \\u003cem\\u003en\\u003c/em\\u003e fragments. The additional number one in the formula is used to distinguish from the mass matching scores in Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ2\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e), whose range is [0, 1].\\u003c/p\\u003e\\u003cp\\u003eStep 5. Formula Ranking: Rank the precursor formulas using their scores. If ties are encountered, assign a better rank to the formula with the smaller average of absolute mass differences between the fragments and their matched formulas.\\u003c/p\\u003e\\u003cp\\u003eThe MS/MS spectrum of salbutamol in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e is used as an example to demonstrate the effectiveness of FSA: The spectrum contains a group of three peaks with \\u003cem\\u003em/z\\u003c/em\\u003e values of 218.1191, 190.1234, and 163.0637 Th in negative ion mode (the precursor \\u003cem\\u003em/z\\u003c/em\\u003e is not present). After spectral preprocessing, the resultant spectrum consisted of a precursor and three fragments with monoisotopic molecular masses of 239.1521, 219.1264, 191.1307, 164.0710, and 83.0739 Da (i.e., \\u003cem\\u003en\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;4). The search for possible chemical formulas under a tolerance of 10 ppm revealed accordingly 3, 4, 4, and 3 formulas for the four peaks, respectively. Despite the 3\\u0026times;4\\u0026times;4\\u0026times;3\\u0026thinsp;=\\u0026thinsp;144 possible combinations, only 16 MDRs among the peak formulas were found. The best combination corresponding to the three precursor formulas and their associated MDRs are indicated by arrows. Although the original number of MDRs was the same for the three combinations (all 5 s), some MDRs in the latter two combinations exhibited multiple choices (dashed lines in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e) for the daughter fragment (hence, not a CF). After removing such daughter fragments and the associated ambiguous MDRs, the first combination, which consisted of C\\u003csub\\u003e13\\u003c/sub\\u003eH\\u003csub\\u003e21\\u003c/sub\\u003eNO\\u003csub\\u003e3\\u003c/sub\\u003e, C\\u003csub\\u003e13\\u003c/sub\\u003eH\\u003csub\\u003e17\\u003c/sub\\u003eNO\\u003csub\\u003e2\\u003c/sub\\u003e, C\\u003csub\\u003e12\\u003c/sub\\u003eH\\u003csub\\u003e17\\u003c/sub\\u003eNO, and C\\u003csub\\u003e9\\u003c/sub\\u003eH\\u003csub\\u003e10\\u003c/sub\\u003eNO\\u003csub\\u003e2\\u003c/sub\\u003e, exhibited the most MDRs and the highest score of 1.8333 (5/6\\u0026thinsp;+\\u0026thinsp;1 by Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ3\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e)). The precursor formula of C\\u003csub\\u003e13\\u003c/sub\\u003eH\\u003csub\\u003e21\\u003c/sub\\u003eNO\\u003csub\\u003e3\\u003c/sub\\u003e was indeed the correct formula for salbutamol.\\u003c/p\\u003e\\u003cp\\u003eThe flowchart of the overall formula ranking process is illustrated in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e. First, MS/MS spectra were selected using the four criteria mentioned in the \\u003cem\\u003eConstruction of the Test Datasets\\u003c/em\\u003e section. Qualified spectra were preprocessed to convert the protonated/deprotonated peaks to their monoisotopic forms, and peaks with masses larger than that of the precursor were removed. Possible formulas of the remaining peaks in the preprocessed spectra were searched via mass match under a given tolerance. Spectra whose precursor did not display any chemical formula were considered unidentifiable, whereas those with a single formula were regarded as identified. The remaining spectra that did not present any MDR were identified by mass match. Otherwise, the spectra were processed with CF determination and then formula ranking using the scores in Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ3\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e).\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec12\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e2.7 Performance Evaluation and Improvement\\u003c/h2\\u003e\\u003cp\\u003eThe performance of FSA was analyzed on various spectra in seven datasets. The analysis results can demonstrate the formula ranking capability of FSA toward a variety of metabolite species, mass ranges, experiment settings, instrument types, and manufacturers. A comparison with a state-of-the-art formula ranking tool was also conducted. However, as mentioned earlier, different MS/MS spectra of a metabolite can be acquired under different experimental conditions, resulting in different fragmentation patterns and different formula rankings from FSA. Several factors including fragment number, variance of fragment masses, and ratio of CFs were analyzed to provide guidelines for a better ranking result. Furthermore, the expandability of FSA for involving more chemical elements such as halogens was tested and several performance improvement ideas were investigated.\\u003c/p\\u003e\\u003c/div\\u003e\"},{\"header\":\"3. RESULTS\",\"content\":\"\\u003cdiv id=\\\"Sec14\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e3.1 Determination of Mass Matching Tolerance\\u003c/h2\\u003e\\u003cp\\u003eTo determine a suitable mass matching tolerance for spectra in each dataset, various values ranging from 0.1 to 3 with a step size of 0.1 mDa were tested and the percentage of spectra in each dataset that were assigned with the correct formula was recorded, as shown in Supplementary Figure \\u003cspan refid=\\\"MOESM2\\\" class=\\\"InternalRef\\\"\\u003eS2\\u003c/span\\u003e. The annotated points showing matching tolerance and correct ranking percentage were mostly corresponding to the highest percentage readings, except for the three HMDB datasets. The exceptions were due to situations in which the precursor information was missing in these datasets and therefore the theoretical value was used for mass match. As a result, the correct formulas of these HMDB spectra were always available during mass match, leading to the persistent percentage increase with the decreasing tolerance.\\u003c/p\\u003e\\u003cp\\u003eThe matching tolerances of the three HMDB datasets were instead evaluated directly from their MS/MS spectra. The idea was to measure the major mass differences between the possible precursors in the MS/MS spectra and their theoretical values. To avoid false precursor peaks, only spectra with peaks close (\\u0026lt;\\u0026thinsp;3 mDa) to their theoretical precursor values were included and the tolerance was selected at 80% of the qualified spectra. The histograms and their accumulated version of the qualified spectra in the three datasets are shown in Supplementary Figure S3, from which the mass matching tolerances for the three datasets HMDB_QTOF, HMDB_ITFT, and HMDB_QFT, were 1.3, 0.6, and 0.3 mDa, respectively.\\u003c/p\\u003e\\u003cp\\u003eHowever, for MS/MS spectra generated from a high-resolution instrument such as QFT, a small mass matching tolerance of 0.3 mDa usually leads to a single candidate from mass match. In fact, in the HMDB_QFT dataset, more than 77.8% of spectra had their chemical formula identified via mass match using the tolerance of 0.3 mDa, leaving a very limited room for the FSA to demonstrate its effectiveness in such spectra. Alternatively, a larger tolerance of 0.5 mDa was used in this dataset such that the percentage of mass matched spectra was considerably reduced from 77.83% to 20.64%, whereas the overall correction percentage slightly decreased from 99.86% to 98.05%, as shown in Supplementary Figure S4. In summary, the mass matching tolerances in FSA for the MAXIS, QTOF, OT, ITFT, and QFT spectra were set to 2.9, 1.3, 1.1, 0.6, and 0.5 mDa, respectively.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec15\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e3.2 Determination of Effective Spectral Peaks\\u003c/h2\\u003e\\u003cp\\u003eAs previously mentioned, some peaks in the MS/MS spectra are attributed to coelutions, false centroiding or deisotoping. Moreover, some unstable fragments may not exist in nature environments and cannot be found in the chemical database. Such fragments can degenerate the performance of FSA. Fortunately, these fragments usually appear in the spectra with low abundances. An experiment was conducted to elucidate the relationship between the usage of the top abundant peaks in the spectra and the corresponding formula ranking performance of FSA. Supplementary Figure S5 shows the experiment result, in which various numbers of the top abundant peaks ranging from 5 to 100 were used in FSA and the performance was represented by the percentage of spectra that were assigned with the correct formula. Although the best performance in the seven datasets occurred at different numbers of top abundant peaks (as annotated in the figure), these numbers were mostly near 30 and the performances around the top-30 abundant peaks in each dataset were similar. The experimental result suggests that, by using only the top-30 abundant peaks, the performance of FSA can be mostly improved or at least maintained in comparison with using all the fragments in the spectra. More importantly, the reduction in peak number can considerably reduce the effort in MDR evaluations, resulting in significant savings in computation time.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec16\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e3.3 Performance Evaluation\\u003c/h2\\u003e\\u003cp\\u003eThe FSA formula ranking process was applied to the seven test datasets using the aforementioned mass matching tolerances, and the top-30 abundant fragments in each spectrum. The performances are summarized in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e, and the detailed results, including the FSA scores and ranking of the correct metabolite of the seven test datasets, are provided in Supplementary File 1.\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003ePercentage of metabolites/spectra in the seven test datasets with correctly assigned chemical formula\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"17\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c7\\\" colnum=\\\"7\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c8\\\" colnum=\\\"8\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c9\\\" colnum=\\\"9\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c10\\\" colnum=\\\"10\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c11\\\" colnum=\\\"11\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c12\\\" colnum=\\\"12\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c13\\\" colnum=\\\"13\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c14\\\" colnum=\\\"14\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c15\\\" colnum=\\\"15\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c16\\\" colnum=\\\"16\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c17\\\" colnum=\\\"17\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003eDatasets\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003eSpectra / Metabolites\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003eTotal qualified\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"11\\\" nameend=\\\"c17\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003eThe correct formula is\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003ethe sole candidate\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c11\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003ethe rank#1\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c14\\\" namest=\\\"c12\\\"\\u003e\\u003cp\\u003ein top-3 ranks\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c17\\\" namest=\\\"c15\\\"\\u003e\\u003cp\\u003ein top-5 ranks\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e3029\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e426 (14.06%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2105 (69.49%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e2677 (88.38%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e2852 (94.16%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e821\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e135 (16.44%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e670 (81.61%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e776 (94.52%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e803 (97.81%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_ITFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e2601\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e1006 (38.68%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2495 (95.92%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e2596 (99.81%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e2601 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e242\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e87 (35.95%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e239 (98.76%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e242 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e242 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e1439\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e297 (20.64%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e1411 (98.05%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e1439 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e1439 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e221\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e49 (22.17%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e221 (100.00%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e221 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e221 (100.00%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_MAXIS\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e3561\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e2 (0.06%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2966 (83.29%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e3342 (93.85%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e3421 (96.07%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e2833\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e2 (0.07%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2484 (87.68%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e2722 (96.08%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e2765 (97.60%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" morerows=\\\"1\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e1373\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e48 (3.50%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e939 (68.39%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e1200 (87.40%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e1274 (92.79%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e1360\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e48 (3.53%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e938 (68.97%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e1193 (87.72%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e1265 (93.01%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c2\\\" namest=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c4\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c6\\\" namest=\\\"c5\\\"\\u003e\\u003cp\\u003e270\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c8\\\" namest=\\\"c7\\\"\\u003e\\u003cp\\u003e13 (4.81%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e249 (92.22%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c13\\\" namest=\\\"c11\\\"\\u003e\\u003cp\\u003e267 (98.89%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c16\\\" namest=\\\"c14\\\"\\u003e\\u003cp\\u003e268 (99.26%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"1\\\" nameend=\\\"c17\\\" namest=\\\"c17\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c3\\\" namest=\\\"c2\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c5\\\" namest=\\\"c4\\\"\\u003e\\u003cp\\u003e239\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c7\\\" namest=\\\"c6\\\"\\u003e\\u003cp\\u003e13 (5.44%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c9\\\" namest=\\\"c8\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e221 (92.47%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c12\\\" namest=\\\"c10\\\"\\u003e\\u003cp\\u003e237 (99.16%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c15\\\" namest=\\\"c13\\\"\\u003e\\u003cp\\u003e238 (99.58%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c17\\\" namest=\\\"c16\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eLIPIDS_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c3\\\" namest=\\\"c2\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c5\\\" namest=\\\"c4\\\"\\u003e\\u003cp\\u003e18084\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c7\\\" namest=\\\"c6\\\"\\u003e\\u003cp\\u003e7491 (41.42%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c9\\\" namest=\\\"c8\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e12219 (67.57%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c12\\\" namest=\\\"c10\\\"\\u003e\\u003cp\\u003e17781 (98.32%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c15\\\" namest=\\\"c13\\\"\\u003e\\u003cp\\u003e18075 (99.95%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c17\\\" namest=\\\"c16\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c3\\\" namest=\\\"c2\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c5\\\" namest=\\\"c4\\\"\\u003e\\u003cp\\u003e896\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c7\\\" namest=\\\"c6\\\"\\u003e\\u003cp\\u003e363 (40.51%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c9\\\" namest=\\\"c8\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e814 (90.85%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c12\\\" namest=\\\"c10\\\"\\u003e\\u003cp\\u003e891 (99.44%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c15\\\" namest=\\\"c13\\\"\\u003e\\u003cp\\u003e893 (99.67%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c17\\\" namest=\\\"c16\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eTotal\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c3\\\" namest=\\\"c2\\\"\\u003e\\u003cp\\u003eSpectra\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c5\\\" namest=\\\"c4\\\"\\u003e\\u003cp\\u003e30357\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c7\\\" namest=\\\"c6\\\"\\u003e\\u003cp\\u003e9283 (30.58%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c9\\\" namest=\\\"c8\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e22384 (73.74%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c12\\\" namest=\\\"c10\\\"\\u003e\\u003cp\\u003e29302 (96.52%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c15\\\" namest=\\\"c13\\\"\\u003e\\u003cp\\u003e29930 (98.59%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c17\\\" namest=\\\"c16\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c3\\\" namest=\\\"c2\\\"\\u003e\\u003cp\\u003eMetabolites\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c5\\\" namest=\\\"c4\\\"\\u003e\\u003cp\\u003e6612\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c7\\\" namest=\\\"c6\\\"\\u003e\\u003cp\\u003e699 (10.57%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c9\\\" namest=\\\"c8\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e5587 (84.50%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c12\\\" namest=\\\"c10\\\"\\u003e\\u003cp\\u003e6282 (95.01%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"3\\\" nameend=\\\"c15\\\" namest=\\\"c13\\\"\\u003e\\u003cp\\u003e6427 (97.20%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c17\\\" namest=\\\"c16\\\"\\u003e\\u0026nbsp;\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eThe ranking capability of the FSA was represented by three values: rank#1, in top-3 ranks, and in top-5 ranks. Rank#1 shows the number of metabolites/spectra whose correct chemical formulas were ranked first among all the possible candidates by FSA. The \\u0026ldquo;in top-3 ranks\\u0026rdquo; and \\u0026ldquo;in top-5 ranks\\u0026rdquo; mean that the correct formula of the precursor was among the top-3 and top-5 ranked formula candidates, respectively. The rank#1 demonstrates the capability of FSA to pinpoint the correct formula from all the possible candidates, whereas top-3 and top-5 ranks show its capability of reducing the number of possible candidates. Notably, to present the overall formula ranking capability of the proposed process in Table. 2 and to compare with the state-of-the-art performance, the numbers of rank#1, top-3 rank, and top-5 rank in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e include those inapplicable spectra whose correct chemical formulas were determined by mass match under the dataset-specific mass matching tolerance.\\u003c/p\\u003e\\u003cp\\u003eMoreover, because different MS/MS spectra (e.g., different instrument settings) for a compound may obtain different ranking results, the subsequent test results focused on metabolite instead of spectrum. Nevertheless, some spectra of a metabolite could be found inapplicable, whereas others were applicable. Consequently, the numbers of inapplicable and applicable metabolites do not add up to the number of total metabolites in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e. The issue of which types of spectra are more likely found with the correct formula will be addressed in the Discussion section.\\u003c/p\\u003e\\u003cp\\u003eThe rank#1 data in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e shows that FSA correctly assigned 84.50% of the metabolite (73.74% of the spectra) formulas in the seven test datasets spanning three different data sources of HMDB, GNPS-NIH-NPL, and GNPS-PNNL-LIPIDS. These results demonstrate that by simply evaluating the MDRs of the CFs in a MS/MS spectrum, the underlying chemical formula is highly likely (\\u0026gt;\\u0026thinsp;84%) to be identified with rank#1 using FSA. If the top-3 and top-5 ranks were allowed, the percentage would improve to 95.01% and 97.20%, respectively. Although top-3 and top-5 ranks do not actually identify the formula of the underlying compound, they can help rank/reduce the possible candidates, hence reducing the subsequent identification cost. In addition to the accuracy, the precursor formula ranking performed by FSA is fast. The overall computation time of the total 30357 MS/MS spectra in the seven datasets was approximately 16 min 33 s (0.033 s per spectrum) in the Windows 10 environment with an Intel\\u0026reg; Core\\u0026trade; i9-10900F CPU. Both formula selection and MDR computation can be achieved via logical operations, making FSA a very effective algorithm for precursor formula ranking.\\u003c/p\\u003e\\u003c/div\\u003e\"},{\"header\":\"4. DISCUSSION\",\"content\":\"\\u003cdiv id=\\\"Sec18\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e4.1 Probable Causes of Datasets with Low Formula Ranking Accuracies\\u003c/h2\\u003e\\u003cp\\u003eThe formula ranking accuracies for spectra in the three datasets, HMDB_QTOF, NPL_QTOF, and LIPID_OT, were lower (\\u0026lt;\\u0026thinsp;70%) than those of the other four datasets. The causes of the low accuracy were investigated. First, the low-ranking accuracies may be attributed to the low CF ratios in the three datasets. This argument can be inferred from Supplementary Figure S6, where the ratios of informative fragments in the spectra of the seven test datasets are displayed. Recalled that the proposed FSA uses the 30 most abundant fragments in a spectrum for the precursor formula ranking. Supplementary Figure S6(a) shows the ratio of the 30 fragments in a spectrum that can be found to match with at least one chemical formula (hereafter referred to as fragments with formula), whereas Supplementary Figure S6(b) demonstrates the ratio of these fragments with formula that were CFs of the correct precursor formula. As shown in Supplementary Figure S6(a), the fragment-with-formula ratios were rather high, but the CF ratios were among the lowest in HMDB_QTOF and NPL_QTOF. The high fragment-with-formula ratios were most likely caused by low mass accuracy and thus large mass matching tolerance, whereas the low CF ratios mainly resulted from unstable fragments and false peaks. The unstable fragments denote product ions that only exist during high-energy dissociation, not in the natural environment; thus, their formulas are not in the compound database. In contrast, the false peaks are signals in a spectrum that do not represent any compound-related product ions. These peaks generally result from false deisotope/centroiding process, contamination, or coelution. Both unstable fragments and false peaks may be matched with false formulas under large tolerance that usually exhibit no MDR with the precursor, leading to a high fragment-with-formula ratio and a low CF ratio in the corresponding spectrum. Although the spectra in the NPL_MAXIS dataset were also generated from QTOF MS, the differences in collision energy, experiment settings, and instrument manufacturer can lead to different performances.\\u003c/p\\u003e\\u003cp\\u003eHowever, the causes of low-ranking accuracy in the LIPID_OT dataset were different from those in the previous two QTOF datasets as the spectra therein exhibited both low fragment-with-formula and low CF ratios. Such phenomena may be due to special functional groups, various branch chain structures, and different double bond locations in the lipids. Particular fragmentation mechanisms (e.g., charge-driven and charge-remote processes) of lipid may provide clues for known and predicted fragments in a spectrum. However, such predicted fragments may not have been included in the database, leading to low fragment-with-formula ratios. However, the high molecular mass of some lipid fragments increased the chance of false formulas during mass match, which in turns reduced the number of MDRs and the CF ratios.\\u003c/p\\u003e\\u003cp\\u003eFurthermore, in the NPL_QTOF dataset, the formula candidates of the precursor in 81 (5.9%) spectra contained no correct formula within the given mass matching tolerance (i.e., 1.3 mDa). Although such spectra also existed in the NPL_OT and LIPID_OT datasets, the respective percentages were very low (0.74% (2 spectra) and 0.05% (9 spectra, respectively). Moreover, 1141, 5, 5, 3, 218, and 1 spectra were created in the years of 2014, 2017, 2018, 2019, 2020, and 2021; 74 of the 81 problematic spectra were from 2020. The high percentage (33.94%, 74 out of 218) of spectra with large mass differences in 2020 may be due to changes in experiment conditions. However, such large number of unidentifiable spectra was attributed to the low formula assignment accuracy of FSA for this dataset.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec19\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e4.2 Comparison with Other Formula Assignment Tools\\u003c/h2\\u003e\\u003cp\\u003eSeveral tools have been proposed for automated metabolite formula assignment/ranking for untargeted metabolomics studies [13\\u0026ndash;24, 30]. A state-of-the-art tool, SIRIUS (version 5.5.7, downloaded on 2022/10/23) was adopted to compare with FSA. The tool was selected for its top performance in category 2 (best automatic structural identification - \\u003cem\\u003ein silico\\u003c/em\\u003e fragmentation only) of the CASMI 2016 (Schymanski et al., \\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e2017\\u003c/span\\u003e) contest and the top performance in \\\"correct elemental formulas\\\" of the CASMI 2022. Notably, although ZODIAC (Ludwig et al., \\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e2019\\u003c/span\\u003e) outperformed SIRIUS in the 2016 contest, it is now also implemented in SIRIUS as an option for improvement in formula ranking. The comparisons were conducted among SIRIUS alone, SIRIUS with the ZODIAC option (SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC), and FSA. In addition, only tools in category 2 were compared as only the MS/MS spectra were given in the contest, which was similar to the situation discussed herein.\\u003c/p\\u003e\\u003cp\\u003eThe comparisons were performed on the seven test datasets, as listed in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e. The same chemical elements were specified in the comparison: CHONSP in the three HMDB datasets, CHONS in the three GNPS-NIH-NPL datasets, and CHONP in the GNPS-PNNL-LIPIDS dataset. However, as the mass matching tolerance in SIRIUS can only be assigned in ppm, various values were tested to determine the best performance of SIRIUS. The comparison result is summarized in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e, where the winner for each dataset is indicated in boldfaced numbers.\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab3\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 3\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eA Comparison of formula assignment accuracy among SIRIUS, SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC, and FSA\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"4\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eDatasets\\u003c/p\\u003e\\u003cp\\u003e(Total Spectra)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eSIRIUS Alone\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eSIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eFSA with 30 M.A.P.\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(3029)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e2190 (72.30%)\\u003c/p\\u003e\\u003cp\\u003e(14 ppm, 13 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2286 (75.47%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(14 ppm, 13 N/A)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e2105 (69.49%)\\u003c/p\\u003e\\u003cp\\u003e(1.3 mDa)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_ITFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(2601)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e2256 (86.74%)\\u003c/p\\u003e\\u003cp\\u003e(6 ppm)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e2318 (89.12%)\\u003c/p\\u003e\\u003cp\\u003e(6 ppm)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2495 (95.92%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(0.6 mDa)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(1439)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1372 (95.34%)\\u003c/p\\u003e\\u003cp\\u003e(3 ppm)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e1405 (97.64%)\\u003c/p\\u003e\\u003cp\\u003e(3 ppm)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e1411 (98.05%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(0.5 mDa,)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_MAXIS\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(3561)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e2669 (74.95%)\\u003c/p\\u003e\\u003cp\\u003e(20 ppm, 258 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e2669 (74.95%)\\u003c/p\\u003e\\u003cp\\u003e(20 ppm, 258 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2966 (83.29%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(2.9 mDa)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(1373)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e951 (69.26%)\\u003c/p\\u003e\\u003cp\\u003e(15 ppm, 16 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e1069 (77.86%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(15 ppm, 16 N/A)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e939 (68.39%)\\u003c/p\\u003e\\u003cp\\u003e(1.3 mDa)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(270)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e252 (93.33%)\\u003c/p\\u003e\\u003cp\\u003e(7 ppm, 2 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e255 (94.44%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(7 ppm, 2 N/A)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e249 (92.22%)\\u003c/p\\u003e\\u003cp\\u003e(1.1 mDa)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eLIPIDS_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(18084)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e11212 (62.00%)\\u003c/p\\u003e\\u003cp\\u003e(18 ppm, 1905 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e11269 (62.31%)\\u003c/p\\u003e\\u003cp\\u003e(20 ppm, 1905 N/A)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e12219 (67.57%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cb\\u003e(1.1 mDa)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eTotal (30357)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e20902 (68.85%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e21271 (70.07%)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e22384 (73.74%)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cem\\u003eM.A.P. stands for the most abundant peaks; N/A. stands for not applicable due to the 850 Da mass limit of the SIRIUS.\\u003c/em\\u003e\\u003c/p\\u003e\\u003cp\\u003eOverall, FSA achieved the best formula assignment accuracy (73.74%) among the three tools and SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC (70.07%) was better than SIRIUS alone (68.85%). Individually, SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC was better than SIRIUS alone in all datasets and FSA attained similar or better assignment accuracies than both SIRIUS alone and SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC in five of the seven datasets, except for the two QTOF datasets. The inferior performance of FSA in the QTOF spectra may be due to the low CF ratios in these datasets, as demonstrated in Supplementary Figure S6(b). Although the FSA employed CFs to remove ambiguous MDRs in a rather efficient and effective manner, more sophisticated rules, such as those employed in SIRIUS and ZODIAC, which involve structural information, are needed to remove false MDRs. However, SIRIUS faces difficulties to address compounds with masses larger than 850 Da [21]. Because of this limitation, some compounds in the seven datasets were not applicable (N/A in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e) to SIRIUS, making it less suitable for large metabolites such as lipids. Furthermore, the assignment result from SIRIUS can exhibit small variations in different runs, which may be due to some statistical approximations in the algorithm.\\u003c/p\\u003e\\u003cp\\u003eTable\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e also indicates that the selected mass matching tolerances for the SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC to achieve the best formula assignment accuracy in the seven datasets were 14, 6, 3, 20, 15, 7, and 20 ppm, respectively. Formula assignment accuracies versus various mass matching tolerances for SIRIUS alone and for SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC are provided in Supplementary Table \\u003cspan refid=\\\"MOESM1\\\" class=\\\"InternalRef\\\"\\u003eS1\\u003c/span\\u003e, whereas the GUI interface of SIRIUS is illustrated in Supplementary Figure S7. These tolerances selected for SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC are highly correlated (correlation coefficient\\u0026thinsp;=\\u0026thinsp;0.77) with those for FSA (i.e., 1.3, 0.6, 0.5, 2.9, 1.3, 1.1, and 1.1 mDa). This phenomenon demonstrates that the matching tolerances used in FSA actually reflected the properties of the spectra in the corresponding dataset rather than were tailored for the performance of FSA. Notably, the tolerances used in SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC were selected to achieve the best accuracies in the seven datasets such as to compare with those of FSA. By considering consistency in instrument (as in FSA), some of the accuracies of SIRIUS and SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC were reduced from those in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003eIn addition to accuracy, computation time was also tested among the three tools. From Table\\u0026nbsp;\\u003cspan refid=\\\"Tab4\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e, FSA was the fastest among the three tools and SIRIUS alone was generally faster than SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC. In average, FSA was more than 30 times faster than both SIRIUS alone and SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC. Such a comparison may not seem fair as FSA only imposed the 30 most abundant fragments in a spectrum for the formula assignment, whereas SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC used all the fragments. However, even if all the fragments are used for the assignment, FSA can still achieve an averaged computation time of 0.072 s per spectrum, which is more than 13 times faster than both SIRIUS alone and SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC.\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab4\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 4\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eComputation time spent by SIRIUS, SIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC, and FSA on the test datasets\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"5\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eDatasets\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eTotal Spectra\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eSIRIUS Alone\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eSIRIUS\\u0026thinsp;+\\u0026thinsp;ZODIAC\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eFSA with 30 M.A.P.\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e3029\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e17 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e11 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e1.8 min (108 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_ITFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e2601\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e5 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e15 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.8 min (48 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eHMDB_QFT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1439\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e2 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e3 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.5 min (30 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_MAXIS\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e3561\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e171 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e185 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e4.08 min (245 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_QTOF\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1373\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e10 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e13 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e2.03 min (122 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPL_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e270\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e1 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e1 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.22 min (13 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eLIPIDS_OT\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e18084\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e296 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e335 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e7.12 min (427 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eTotal\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e30357\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e502 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e563 min\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e16.55 min (993 s)\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eAverage\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.99 s\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e1.11 s\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.033 s\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cem\\u003eM.A.P. stands for the most abundant peaks.\\u003c/em\\u003e\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec20\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e4.3 Performance Consistency in Metabolite Mass\\u003c/h2\\u003e\\u003cp\\u003eThe formula ranking capability of FSA toward different metabolite masses in the seven test datasets was analyzed. The analysis result is shown in Supplementary Figure S8, in which only the metabolites whose formulas were actually assigned by FSA are displayed. Supplementary Figure S8(a) shows the number of such metabolites in total and in each dataset, whereas Supplementary Figure S8(b) illustrates the percentage of metabolites in (a) that were assigned with the correct formulas. Some of the lines are disconnected at certain mass ranges due to lack of metabolite in those mass ranges. In addition, the metabolites were divided into different mass ranges from 0 to 1500 Da for every 100 Da. The number of metabolites instead of spectra were considered as multiple spectra can correspond to the same metabolite, each with different collision energies, particularly in the HMDB datasets. Some of such spectra may either contain few fragments (low energy) or only very low abundant fragments (high energy), resulting in few MDRs, being not suitable for FSA.\\u003c/p\\u003e\\u003cp\\u003eAs indicated in Supplementary Figure S8(b), the averaged assignment accuracy of FSA (green line) exhibits a rather stable behavior (mostly \\u0026ge; 75%) across the entire mass range. The larger variations in the high mass range 1000\\u0026ndash;1500 Da compared to the low mass range 100\\u0026ndash;1000 Da may be due to small sample size (\\u0026lt;\\u0026thinsp;40) in the datasets. From the aspect of individual dataset, FSA exhibited the best performance in the HMDB_QFT, followed by HMDB_ITFT, NPL_OT, NPL_MAXIS, LIPID_OT, HMDB_QTOF, and NPL_QTOF, which coincides with the data in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e. The worse performance in the two QTOF datasets may be due to the low percentage of CFs in the spectra, as previously discussed. The performance of FSA to large mass (\\u0026gt;\\u0026thinsp;600 Da) metabolites in the two QTOF datasets (red and blue lines in Supplementary Figure S8(b)) seemed contradictory. However, the sample sizes were small in these two sets. More test data is required before deriving any conclusion.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec21\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e4.4 Performance after Involving Halogen Elements\\u003c/h2\\u003e\\u003cp\\u003eAs previously mentioned, FSA requires only a list of monoisotopic masses and chemical formulas of all possible compounds to perform precursor formula ranking for a MS/MS spectrum. In the previous tests, compounds consisting of only C, H, O, N, S, and P were considered. In this test, compounds involving additional halogen elements F, Cl, Br, and I were considered. As a result, the original list of 928,369 reference chemical formulas from PubChem was increased to 2,613,203 formulas. In addition, the number of qualified MS/MS spectra in the test datasets was also increased, as presented in Supplementary Table \\u003cspan refid=\\\"MOESM2\\\" class=\\\"InternalRef\\\"\\u003eS2\\u003c/span\\u003e. Notably, the LIPID_OT dataset was removed from this test as no halogen element was found in the dataset.\\u003c/p\\u003e\\u003cp\\u003eThe performances of FSA after including the halogen elements are summarized in Supplementary Table S3. The rank#1 data in the table shows that the proposed FSA correctly identified 76.80% of the metabolite (73.99% of the spectra) formulas in the six test datasets. These results demonstrate that after including the halogen elements in the formula list, FSA can still maintain\\u0026thinsp;\\u0026gt;\\u0026thinsp;76% accuracy in pinpointing (rank#1) the underlying chemical formula in the seven datasets. If the top-3 and top-5 ranks were used, the percentage would improve to 88.91% and 92.11%, respectively. In comparison with the previous test, the performance of FSA decreased less than 9% after the four halogen elements were added. Notably, FSA exhibited considerable performance reductions in HMDB_QTOF (81.61%\\u0026rarr;67.56%) and NPL_QTOF (68.59%\\u0026rarr;59.73%). This may be due to the combination effect of high energy collision and halogen elements in the spectra, producing abundant fragments that were nonexistent in the compound database. If the two QTOF datasets were not considered, the performance dropped from 89.54% (3176 of 3547 metabolites) to 85.99% (3238 of 3770 metabolites), less than 4% of reduction.\\u003c/p\\u003e\\u003cp\\u003eIn addition, the total computation time for the 14169 spectra was 30 min 37 s with an average of 0.13 s per spectrum, which was 3.94 (0.13/0.033) times longer than that of the previous test. However, the increases of 1.67 times (6\\u0026rarr;10) in element number and 2.81 times (928,369\\u0026rarr;2,613,203) in reference formula number resulted in increases of 1.67\\u0026times;\\u003cem\\u003en\\u003c/em\\u003e and 2.81 times in MDR calculation and formula candidates search, respectively, where \\u003cem\\u003en\\u003c/em\\u003e is the additional number of formula candidates from all fragments. With such additional computations and large overhead in passing data (more than half of the computation time), the proposed FSA was still able to process an MS/MS spectrum in a fraction of second and with a small performance decrease of \\u0026lt;\\u0026thinsp;9%, showing its capability of performing reliable formula ranking with the expansion of compound compositions.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec22\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e4.5 Suitable Types of Spectra\\u003c/h2\\u003e\\u003cp\\u003eAs MDRs of CFs are both critical to the FSA, the accuracy, amount, and location of the fragment peaks can all contribute toward the final score and rank of a spectrum. The manner fragments should appear in a spectrum for the FSA to provide an accurate formula ranking was also investigated.\\u003c/p\\u003e\\u003cp\\u003eFirst, the relationship between the formula ranking capability of the FSA and number of fragments in a spectrum was analyzed. Supplementary Figure S9 shows boxplots illustrating whether the fragment numbers were differently distributed between spectra whose correct formulas were ranked #1 (R1) and those that were otherwise ranked (R2+) in each of the seven test datasets. At the top of each boxplot, a \\u003cem\\u003ep\\u003c/em\\u003e-value computed by the one-sided Wilcoxon rank sum test is displayed, indicating their differences in the median values as the fragment numbers were not normally distributed (\\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.01 by both the Kolmogorov\\u0026ndash;Smirnov (K\\u0026ndash;S) test and the Anderson\\u0026ndash;Darling (A\\u0026ndash;D) test). The fragment number in the vertical axes is drawn in the logarithmic scale due to the large value range (102\\u0026ndash;104) among datasets. The median fragment numbers of the R1 spectra were 14, 14, 22, 219, 1276, 30, and 74, whereas those of the R2\\u0026thinsp;+\\u0026thinsp;spectra were 13, 13.5, 13, 101, 802, 17, and 66 in the seven test datasets, respectively. Although the fragment numbers in both groups were not significantly different, the fragment numbers in the R1 spectra were higher than those in the R2\\u0026thinsp;+\\u0026thinsp;spectra, suggesting that higher fragment numbers in a spectrum may have a higher chance for correct formula assignment by FSA.\\u003c/p\\u003e\\u003cp\\u003eNext, the distribution of fragment peaks in a spectrum was analyzed based on the standard deviation (SD) of the absolute distances between two neighboring peaks. The precursor peak was added to the spectrum (in case it was missing due to high-energy collision) to distinguish between peak clusters with high and low mass values. The boxplots of the SDs in the rank#1 spectra versus those in the non-rank#1 ones for the seven datasets are presented in Supplementary Figure S10. The median SDs of the R1 spectra by the FSA method were 13.00, 15.03, 12.82, 3.71, 0.77, 13.23, and 22.79, whereas those of the R2\\u0026thinsp;+\\u0026thinsp;spectra were 15.98, 22.93, 22.89, 7.45, 1.21, 55.75, and 24.76 in the seven test datasets, respectively. The one-sided Wilcoxon rank sum test showed that the SDs of the neighboring peak distances in the R1 spectra were significantly lower (\\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.01) than those in the R2\\u0026thinsp;+\\u0026thinsp;spectra in all the seven datasets, suggesting that a lower SD of peak distances in a spectrum may also improve formula assignment performance. The result is understandable because fragment peaks with small inter-peak distance SDs are likely to be more evenly distributed in an MS/MS spectrum. This reflects a higher chance for more MDRs and a higher score for the underlying formula of the spectrum.\\u003c/p\\u003e\\u003cp\\u003eTwo other factors, namely the CF ratio and the difference between the top two scores, were also investigated. Supplementary Figure S11 shows the box plots of the CF ratios in the R1 and R2\\u0026thinsp;+\\u0026thinsp;spectra of the seven datasets. The median ratios of the CF to the 30 most abundant fragments with formula were 0.66, 0.80, 0.90, 0.80, 0.66, 0.90, and 0.46 in the R1 spectra and 0.27, 0.25, 0.80, 0.43, 0.44, 0.75 and 0.27 in the R2\\u0026thinsp;+\\u0026thinsp;spectra for the seven test datasets, respectively. The one-sided Wilcoxon rank sum test revealed that the CF ratios were significantly higher (\\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.01) in the R1 spectra than in the R2\\u0026thinsp;+\\u0026thinsp;spectra, suggesting that a higher CF ratio in a spectrum may imply a higher chance for correct precursor formula assignment by FSA.\\u003c/p\\u003e\\u003cp\\u003eIn addition, the difference between the top-two FSA scores represents the effectiveness of the scoring mechanism proposed in (4). Supplementary Figure S12 shows boxplots of the differences between the R1 and the R2\\u0026thinsp;+\\u0026thinsp;spectra of the seven datasets. The median differences of the top two FSA scores in the R1 spectra were 0.13, 0.40, 0.47, 0.19, 0.10, 0.24, and 0.0441, whereas those of the R2\\u0026thinsp;+\\u0026thinsp;spectra were 0.04, 0.05, 0.13, 0.04, 0.03, 0.03 and 0.0444 in the seven datasets, respectively. The one-sided Wilcoxon rank sum test showed significant (\\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.01) differences between the R1 and the R2\\u0026thinsp;+\\u0026thinsp;spectra in all seven datasets, suggesting that a higher score difference in a spectrum may indicate a higher chance that the proposed scoring mechanism can identify the correct precursor formula. However, the median value of the R1 group was slightly smaller than that of the R2\\u0026thinsp;+\\u0026thinsp;group in the LIPID_OT dataset, which is opposite to what occurred with the other datasets. The exceptions for the LIPID species are probably due to their branch chain structures, resulting in alternative MDRs and reducing the CF ratios (refer to Supplementary Figure S6(b)) and scores.\\u003c/p\\u003e\\u003c/div\\u003e\\u003cdiv id=\\\"Sec23\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e4.6 Possible Improvements\\u003c/h2\\u003e\\u003cp\\u003eSeveral experiments were conducted to improve the FSA performance; however, limited progress was achieved. First, the length of the fragmentation path was considered in addition to the sum of MDRs. Many attempts to include additional terms in Eq.\\u0026nbsp;(\\u003cspan refid=\\\"Equ3\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e), to encourage a long fragmentation path (consecutive MDRs in a series of fragments), were tested. Second, possible formulas were searched for fragment peaks without a match in the PubChem database. Adducts may attach to a compound leading to a different fragment pattern in the spectrum. Some fragments in the spectrum may also be associated with the adducts; hence, their formulas may not be found in the PubChem database. Furthermore, some fragment peaks that were generated during the collision process may not be stable in the natural environment; therefore, their formulas cannot be found in PubChem. Efforts were made to develop the formulas of such fragments by adding/subtracting different known adducts and neutral losses (Ma et al., \\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e2014\\u003c/span\\u003e) from fragments with known formulas. Third, combinations of different scoring strategies were also tested. As mentioned earlier, the current scoring method only computes MDRs for fragments with a single formula (i.e., CFs). Although this method achieved the best performance, fragments with multiple possible formulas may still provide clues to identify the spectra. Different weighted combinations of such multi-formula fragments were tested through a genetic algorithm approach. Although some of the above modifications did increase the formula ranking accuracy in some datasets, the improvements were limited and not consistent across all the test datasets.\\u003c/p\\u003e\\u003cp\\u003eAlthough several strategies have been tested to improve the formula ranking accuracy in the test datasets, few chemical rules have been implemented to rank possible formulas and filter out unfeasible ones. Such rules can be crucial in certain compound species and in compounds with adducts. In addition, species-related rules can also be applied to improve the identification. As some compound species contain a specific functional group, fragments related to such groups can be used to reduce possible candidates. Including more chemical formula databases, such as ChemSpider, ChEBI, and SciFinder and updating the PubChem database in FSA may also improve the performance of FSA. For example, the current PubChem version (2022/11/27), which consists of 112,419,402 compounds in comparison with the employed version of 109,383,767 compounds, may help FSA increase the fragment-with-formula ratio and the CF ratio so as to analyze spectra with no or few MDRs. Furthermore, instrument manufacturers usually produce new models with better mass accuracy approximately every two years. As the FSA performance improves with increasing mass accuracy, an improved formula ranking capability of the developed FSA method is expected for spectra generated by modern instruments.\\u003c/p\\u003e\\u003c/div\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cp\\u003e\\u003cstrong\\u003eSupplementary Information\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThe online version contains supplementary material available at GitHub (https://github.com/kslynn128171/FSA.git).\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eAuthor contributions\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eKSL conceived the concept, derived the related programs, performed the analyses, and wrote the manuscript.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eFunding\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThis work was supported by the National Science and Technology Council of Taiwan under Grant No. NSTC114-2113-M-030-011.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eData availability\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThe MATLAB\\u0026reg; programs that perform FSA are freely available at GitHub (https://github.com/kslynn128171/FSA.git) along with Python programs that extract the formula database from PubChem and the seven test datasets extracted from HMDB and GNPS.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eConflict of interest\\u003c/strong\\u003e Ke-Shiuan Lynn declares that he has no conflict of interest.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eHuman and animal participants\\u003c/strong\\u003e This article does not contain any studies with human and/or animal participants performed by any of the authors.\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003cp\\u003eBlazenovic, I., Kind, T., Ji, J. and Fiehn, O. (2018) Software Tools and Approaches for Compound Identification of LC-MS/MS Data in Metabolomics. \\u003cem\\u003eMetabolites\\u003c/em\\u003e \\u003cstrong\\u003e8\\u003c/strong\\u003e.\\u003c/p\\u003e\\n\\u003cp\\u003eBrouard, C., Szafranski, M. and D'Alché-Buc, F. (2016) Input output kernel regression: supervised and semi-supervised structured output prediction with operator-valued kernels. \\u003cem\\u003eJ. Mach. Learn. Res.\\u003c/em\\u003e \\u003cstrong\\u003e17,\\u003c/strong\\u003e 6105–6152.\\u003c/p\\u003e\\n\\u003cp\\u003eChambers, M.C., Maclean, B., Burke, R., Amodei, D., Ruderman, D.L., Neumann, S., Gatto, L., Fischer, B., Pratt, B., Egertson, J., Hoff, K., Kessner, D., Tasman, N., Shulman, N., Frewen, B., Baker, T.A., Brusniak, M.Y., Paulse, C., Creasy, D., Flashner, L., Kani, K., Moulding, C., Seymour, S.L., Nuwaysir, L.M., Lefebvre, B., Kuhlmann, F., Roark, J., Rainer, P., Detlev, S., Hemenway, T., Huhmer, A., Langridge, J., Connolly, B., Chadick, T., Holly, K., Eckels, J., Deutsch, E.W., Moritz, R.L., Katz, J.E., Agus, D.B., MacCoss, M., Tabb, D.L. and Mallick, P. (2012) A cross-platform toolkit for mass spectrometry and proteomics. \\u003cem\\u003eNat Biotechnol\\u003c/em\\u003e \\u003cstrong\\u003e30,\\u003c/strong\\u003e 918-20.\\u003c/p\\u003e\\n\\u003cp\\u003eDührkop, K., Fleischauer, M., Ludwig, M., Aksenov, A.A., Melnik, A.V., Meusel, M., Dorrestein, P.C., Rousu, J. and Böcker, S. (2019) SIRIUS 4: a rapid tool for turning tandem mass spectra into metabolite structure information. \\u003cem\\u003eNature Methods\\u003c/em\\u003e \\u003cstrong\\u003e16,\\u003c/strong\\u003e 299-302.\\u003c/p\\u003e\\n\\u003cp\\u003eDunn, W.B., Broadhurst, D.I., Atherton, H.J., Goodacre, R. and Griffin, J.L. (2011) Systems level studies of mammalian metabolomes: the roles of mass spectrometry and nuclear magnetic resonance spectroscopy. \\u003cem\\u003eChem Soc Rev\\u003c/em\\u003e \\u003cstrong\\u003e40,\\u003c/strong\\u003e 387-426.\\u003c/p\\u003e\\n\\u003cp\\u003eDunn, W.B., Erban, A., Weber, R.J.M., Creek, D.J., Brown, M., Breitling, R., Hankemeier, T., Goodacre, R., Neumann, S., Kopka, J. and Viant, M.R. (2013) Mass appeal: metabolite identification in mass spectrometry-focused untargeted metabolomics. \\u003cem\\u003eMetabolomics\\u003c/em\\u003e \\u003cstrong\\u003e9,\\u003c/strong\\u003e 44-66.\\u003c/p\\u003e\\n\\u003cp\\u003eGuijas, C., Montenegro-Burke, J.R., Domingo-Almenara, X., Palermo, A., Warth, B., Hermann, G., Koellensperger, G., Huan, T., Uritboonthai, W., Aisporna, A.E., Wolan, D.W., Spilker, M.E., Benton, H.P. and Siuzdak, G. (2018) METLIN: A Technology Platform for Identifying Knowns and Unknowns. \\u003cem\\u003eAnalytical Chemistry\\u003c/em\\u003e \\u003cstrong\\u003e90,\\u003c/strong\\u003e 3156-3164.\\u003c/p\\u003e\\n\\u003cp\\u003eHeinonen, M., Shen, H., Zamboni, N. and Rousu, J. (2012) Metabolite identification and molecular fingerprint prediction through machine learning. \\u003cem\\u003eBioinformatics\\u003c/em\\u003e \\u003cstrong\\u003e28,\\u003c/strong\\u003e 2333-2341.\\u003c/p\\u003e\\n\\u003cp\\u003eJedrychowski, M.P., Huttlin, E.L., Haas, W., Sowa, M.E., Rad, R. and Gygi, S.P. (2011) Evaluation of HCD- and CID-type fragmentation within their respective detection platforms for murine phosphoproteomics. \\u003cem\\u003eMol Cell Proteomics\\u003c/em\\u003e \\u003cstrong\\u003e10,\\u003c/strong\\u003e M111 009910.\\u003c/p\\u003e\\n\\u003cp\\u003eKim, S., Chen, J., Cheng, T., Gindulyte, A., He, J., He, S., Li, Q., Shoemaker, B.A., Thiessen, P.A., Yu, B., Zaslavsky, L., Zhang, J. and Bolton, E.E. (2021) PubChem in 2021: new data content and improved web interfaces. \\u003cem\\u003eNucleic Acids Res\\u003c/em\\u003e \\u003cstrong\\u003e49,\\u003c/strong\\u003e D1388-D1395.\\u003c/p\\u003e\\n\\u003cp\\u003eKind, T., Liu, K.-H., Lee, D.Y., DeFelice, B., Meissen, J.K. and Fiehn, O. (2013) LipidBlast in silico tandem mass spectrometry database for lipid identification. \\u003cem\\u003eNature methods\\u003c/em\\u003e \\u003cstrong\\u003e10,\\u003c/strong\\u003e 755-758.\\u003c/p\\u003e\\n\\u003cp\\u003eKind, T., Tsugawa, H., Cajka, T., Ma, Y., Lai, Z.J., Mehta, S.S., Wohlgemuth, G., Barupal, D.K., Showalter, M.R., Arita, M. and Fiehn, O. (2018) Identification of small molecules using accurate mass MS/MS search. \\u003cem\\u003eMass Spectrometry Reviews\\u003c/em\\u003e \\u003cstrong\\u003e37,\\u003c/strong\\u003e 513-532.\\u003c/p\\u003e\\n\\u003cp\\u003eLiebal, U.W., Phan, A.N.T., Sudhakar, M., Raman, K. and Blank, L.M. (2020) Machine Learning Applications for Mass Spectrometry-Based Metabolomics. \\u003cem\\u003eMetabolites\\u003c/em\\u003e \\u003cstrong\\u003e10\\u003c/strong\\u003e.\\u003c/p\\u003e\\n\\u003cp\\u003eLudwig, M., Nothias, L.-F., Dührkop, K., Koester, I., Fleischauer, M., Hoffmann, M.A., Petras, D., Vargas, F., Morsy, M., Aluwihare, L., Dorrestein, P.C. and Böcker, S. (2019) ZODIAC: database-independent molecular formula annotation using Gibbs sampling reveals unknown small molecules. \\u003cem\\u003ebioRxiv\\u003c/em\\u003e\\u003cstrong\\u003e,\\u003c/strong\\u003e 842740.\\u003c/p\\u003e\\n\\u003cp\\u003eLudwig, M., Nothias, L.-F., Dührkop, K., Koester, I., Fleischauer, M., Hoffmann, M.A., Petras, D., Vargas, F., Morsy, M., Aluwihare, L., Dorrestein, P.C. and Böcker, S. (2020) Database-independent molecular formula annotation using Gibbs sampling through ZODIAC. \\u003cem\\u003eNature Machine Intelligence\\u003c/em\\u003e \\u003cstrong\\u003e2,\\u003c/strong\\u003e 629-641.\\u003c/p\\u003e\\n\\u003cp\\u003eMa, Y., Kind, T., Yang, D., Leon, C. and Fiehn, O. (2014) MS2Analyzer: A software for small molecule substructure annotations from accurate tandem mass spectra. \\u003cem\\u003eAnal Chem\\u003c/em\\u003e \\u003cstrong\\u003e86,\\u003c/strong\\u003e 10724-31.\\u003c/p\\u003e\\n\\u003cp\\u003ePeisl, B.Y.L., Schymanski, E.L. and Wilmes, P. (2018) Dark matter in host-microbiome metabolomics: Tackling the unknowns-A review. \\u003cem\\u003eAnal Chim Acta\\u003c/em\\u003e \\u003cstrong\\u003e1037,\\u003c/strong\\u003e 13-27.\\u003c/p\\u003e\\n\\u003cp\\u003eRidder, L., van der Hooft, J.J.J., Verhoeven, S., de Vos, R.C.H., Vervoort, J. and Bino, R.J. (2014) In Silico Prediction and Automatic LC–MSn Annotation of Green Tea Metabolites in Urine. \\u003cem\\u003eAnalytical Chemistry\\u003c/em\\u003e \\u003cstrong\\u003e86,\\u003c/strong\\u003e 4767-4774.\\u003c/p\\u003e\\n\\u003cp\\u003eRuttkies, C., Schymanski, E.L., Wolf, S., Hollender, J. and Neumann, S. (2016) MetFrag relaunched: incorporating strategies beyond in silico fragmentation. \\u003cem\\u003eJournal of Cheminformatics\\u003c/em\\u003e \\u003cstrong\\u003e8,\\u003c/strong\\u003e 3.\\u003c/p\\u003e\\n\\u003cp\\u003eSchymanski, E.L., Ruttkies, C., Krauss, M., Brouard, C., Kind, T., Duhrkop, K., Allen, F., Vaniya, A., Verdegem, D., Bocker, S., Rousu, J., Shen, H., Tsugawa, H., Sajed, T., Fiehn, O., Ghesquiere, B. and Neumann, S. (2017) Critical Assessment of Small Molecule Identification 2016: automated methods. \\u003cem\\u003eJ Cheminform\\u003c/em\\u003e \\u003cstrong\\u003e9,\\u003c/strong\\u003e 22.\\u003c/p\\u003e\\n\\u003cp\\u003eShen, H., Dührkop, K., Böcker, S. and Rousu, J. (2014) Metabolite identification through multiple kernel learning on fragmentation trees. \\u003cem\\u003eBioinformatics\\u003c/em\\u003e \\u003cstrong\\u003e30,\\u003c/strong\\u003e i157-i164.\\u003c/p\\u003e\\n\\u003cp\\u003eSumner, L.W., Amberg, A., Barrett, D., Beale, M.H., Beger, R., Daykin, C.A., Fan, T.W., Fiehn, O., Goodacre, R., Griffin, J.L., Hankemeier, T., Hardy, N., Harnly, J., Higashi, R., Kopka, J., Lane, A.N., Lindon, J.C., Marriott, P., Nicholls, A.W., Reily, M.D., Thaden, J.J. and Viant, M.R. (2007) Proposed minimum reporting standards for chemical analysis Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI). \\u003cem\\u003eMetabolomics\\u003c/em\\u003e \\u003cstrong\\u003e3,\\u003c/strong\\u003e 211-221.\\u003c/p\\u003e\\n\\u003cp\\u003eTsugawa, H., Kind, T., Nakabayashi, R., Yukihira, D., Tanaka, W., Cajka, T., Saito, K., Fiehn, O. and Arita, M. (2016) Hydrogen Rearrangement Rules: Computational MS/MS Fragmentation and Structure Elucidation Using MS-FINDER Software. \\u003cem\\u003eAnal Chem\\u003c/em\\u003e \\u003cstrong\\u003e88,\\u003c/strong\\u003e 7946-58.\\u003c/p\\u003e\\n\\u003cp\\u003eViant, M.R., Kurland, I.J., Jones, M.R. and Dunn, W.B. (2017) How close are we to complete annotation of metabolomes? \\u003cem\\u003eCurrent Opinion in Chemical Biology\\u003c/em\\u003e \\u003cstrong\\u003e36,\\u003c/strong\\u003e 64-69.\\u003c/p\\u003e\\n\\u003cp\\u003eWang, F., Liigand, J., Tian, S., Arndt, D., Greiner, R. and Wishart, D.S. (2021) CFM-ID 4.0: More Accurate ESI-MS/MS Spectral Prediction and Compound Identification. \\u003cem\\u003eAnal Chem\\u003c/em\\u003e \\u003cstrong\\u003e93,\\u003c/strong\\u003e 11692-11700.\\u003c/p\\u003e\\n\\u003cp\\u003eWang, M., Carver, J.J., Phelan, V.V., Sanchez, L.M., Garg, N., Peng, Y., Nguyen, D.D., Watrous, J., Kapono, C.A., Luzzatto-Knaan, T., Porto, C., Bouslimani, A., Melnik, A.V., Meehan, M.J., Liu, W.T., Crusemann, M., Boudreau, P.D., Esquenazi, E., Sandoval-Calderon, M., Kersten, R.D., Pace, L.A., Quinn, R.A., Duncan, K.R., Hsu, C.C., Floros, D.J., Gavilan, R.G., Kleigrewe, K., Northen, T., Dutton, R.J., Parrot, D., Carlson, E.E., Aigle, B., Michelsen, C.F., Jelsbak, L., Sohlenkamp, C., Pevzner, P., Edlund, A., McLean, J., Piel, J., Murphy, B.T., Gerwick, L., Liaw, C.C., Yang, Y.L., Humpf, H.U., Maansson, M., Keyzers, R.A., Sims, A.C., Johnson, A.R., Sidebottom, A.M., Sedio, B.E., Klitgaard, A., Larson, C.B., P, C.A.B., Torres-Mendoza, D., Gonzalez, D.J., Silva, D.B., Marques, L.M., Demarque, D.P., Pociute, E., O'Neill, E.C., Briand, E., Helfrich, E.J.N., Granatosky, E.A., Glukhov, E., Ryffel, F., Houson, H., Mohimani, H., Kharbush, J.J., Zeng, Y., Vorholt, J.A., Kurita, K.L., Charusanti, P., McPhail, K.L., Nielsen, K.F., Vuong, L., Elfeki, M., Traxler, M.F., Engene, N., Koyama, N., Vining, O.B., Baric, R., Silva, R.R., Mascuch, S.J., Tomasi, S., Jenkins, S., Macherla, V., Hoffman, T., Agarwal, V., Williams, P.G., Dai, J., Neupane, R., Gurr, J., Rodriguez, A.M.C., Lamsa, A., Zhang, C., Dorrestein, K., Duggan, B.M., Almaliti, J., Allard, P.M., Phapale, P.\\u003cem\\u003e et al.\\u003c/em\\u003e (2016) Sharing and community curation of mass spectrometry data with Global Natural Products Social Molecular Networking. \\u003cem\\u003eNat Biotechnol\\u003c/em\\u003e \\u003cstrong\\u003e34,\\u003c/strong\\u003e 828-837.\\u003c/p\\u003e\\n\\u003cp\\u003eWang, Y., Kora, G., Bowen, B.P. and Pan, C. (2014) MIDAS: A Database-Searching Algorithm for Metabolite Identification in Metabolomics. \\u003cem\\u003eAnalytical Chemistry\\u003c/em\\u003e \\u003cstrong\\u003e86,\\u003c/strong\\u003e 9496-9503.\\u003c/p\\u003e\\n\\u003cp\\u003eWant, E.J., Wilson, I.D., Gika, H., Theodoridis, G., Plumb, R.S., Shockcor, J., Holmes, E. and Nicholson, J.K. (2010) Global metabolic profiling procedures for urine using UPLC-MS. \\u003cem\\u003eNature Protocols\\u003c/em\\u003e \\u003cstrong\\u003e5,\\u003c/strong\\u003e 1005-1018.\\u003c/p\\u003e\\n\\u003cp\\u003eWishart, D.S., Feunang, Y.D., Marcu, A., Guo, A.C., Liang, K., Vazquez-Fresno, R., Sajed, T., Johnson, D., Li, C., Karu, N., Sayeeda, Z., Lo, E., Assempour, N., Berjanskii, M., Singhal, S., Arndt, D., Liang, Y., Badran, H., Grant, J., Serra-Cayuela, A., Liu, Y., Mandal, R., Neveu, V., Pon, A., Knox, C., Wilson, M., Manach, C. and Scalbert, A. (2018) HMDB 4.0: the human metabolome database for 2018. \\u003cem\\u003eNucleic Acids Res\\u003c/em\\u003e \\u003cstrong\\u003e46,\\u003c/strong\\u003e D608-D617.\\u003c/p\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":false,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":true,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"metabolomics\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"mebo\",\"sideBox\":\"Learn more about [Metabolomics](http://link.springer.com/journal/11306)\",\"snPcode\":\"11306\",\"submissionUrl\":\"https://submission.nature.com/new-submission/11306/3\",\"title\":\"Metabolomics\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"em\",\"reportingPortfolio\":\"Springer Hybrid\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":false},\"keywords\":\"Metabolomics, Formula Ranking, LC-MS/MS, Mother-Daughter Relationship\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-7764688/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-7764688/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003ch2\\u003eIntroduction\\u003c/h2\\u003e\\u003cp\\u003eMetabolite identification remains a bottleneck in untargeted liquid chromatography\\u0026ndash;tandem mass spectrometry (LC-MS) metabolomics studies, especially when the underlying metabolite is not found in the tandem mass spectrometry (MS/MS) databases.\\u003c/p\\u003e\\u003ch2\\u003eObjective\\u003c/h2\\u003e\\u003cp\\u003eA new approach, formula subset analysis (FSA), was developed to effectively rank the chemical formula candidates for an MS/MS spectrum.\\u003c/p\\u003e\\u003ch2\\u003eMethods\\u003c/h2\\u003e\\u003cp\\u003eThis approach first computes mother-daughter relationships (MDRs) among possible formulas of fragments and the precursor under a given mass tolerance and then determines the characteristic fragments (CFs) that only present one MDR with the precursor and other fragments. Subsequently, the precursor formula candidates are ranked by the scores derived from the number of MDRs.\\u003c/p\\u003e\\u003ch2\\u003eResults\\u003c/h2\\u003e\\u003cp\\u003eA numerical study using seven large datasets totaling 30357 MS/MS spectra from 6612 metabolites consisting of C, H, O, N, S, and P showed that FSA ranked the correct chemical formula as the top-1 candidate for a metabolite in 84.50% of the cases and in the top-5 candidates in 97.20% of the cases. The average processing time for each spectrum was 0.033 seconds. Moreover, FSA does not require training data, not rely on MS/MS databases, can be applied to a wide mass range, and can be quickly expanded with more chemical elements and formulas to identify different chemical species.\\u003c/p\\u003e\\u003ch2\\u003eConclusions\\u003c/h2\\u003e\\u003cp\\u003eFSA has not utilized structural information yet and therefore its accuracy may not be competitive with some of the state-of-the-art identification tools. However, its advantages in speed, expandability, and applicability, make it suitable for prescreening candidates in untargeted LC-MS metabolomics studies.\\u003c/p\\u003e\",\"manuscriptTitle\":\"Automated Metabolite Formula Ranking Using Formula Subset Analysis for LC-MS/MS-Based Metabolomics\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2025-10-17 02:16:05\",\"doi\":\"10.21203/rs.3.rs-7764688/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0},{\"type\":\"decision\",\"content\":\"Revision requested\",\"date\":\"2025-10-07T20:09:04+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2025-10-07T20:05:53+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"114252158092834379665117232643066656850\",\"date\":\"2025-10-06T15:15:36+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"editorInvitedReview\",\"content\":\"\",\"date\":\"2025-10-05T18:34:52+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewerAgreed\",\"content\":\"101752038381754263745401315668905786379\",\"date\":\"2025-10-05T16:37:13+00:00\",\"index\":\"hide\",\"fulltext\":\"\"},{\"type\":\"reviewersInvited\",\"content\":\"\",\"date\":\"2025-10-05T09:49:38+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorAssigned\",\"content\":\"\",\"date\":\"2025-10-03T13:56:23+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"checksComplete\",\"content\":\"\",\"date\":\"2025-10-03T13:56:17+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"submitted\",\"content\":\"Metabolomics\",\"date\":\"2025-10-02T07:52:30+00:00\",\"index\":\"\",\"fulltext\":\"\"}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"metabolomics\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"mebo\",\"sideBox\":\"Learn more about [Metabolomics](http://link.springer.com/journal/11306)\",\"snPcode\":\"11306\",\"submissionUrl\":\"https://submission.nature.com/new-submission/11306/3\",\"title\":\"Metabolomics\",\"twitterHandle\":\"\",\"acdcEnabled\":true,\"dfaEnabled\":true,\"editorialSystem\":\"em\",\"reportingPortfolio\":\"Springer Hybrid\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":false}}],\"origin\":\"\",\"ownerIdentity\":\"d92e9012-1f76-49d0-ae1d-058004d1a7b9\",\"owner\":[],\"postedDate\":\"October 17th, 2025\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"published-in-journal\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2025-12-08T16:01:21+00:00\",\"versionOfRecord\":{\"articleIdentity\":\"rs-7764688\",\"link\":\"https://doi.org/10.1007/s11306-025-02379-0\",\"journal\":{\"identity\":\"metabolomics\",\"isVorOnly\":false,\"title\":\"Metabolomics\"},\"publishedOn\":\"2025-12-01 15:57:04\",\"publishedOnDateReadable\":\"December 1st, 2025\"},\"versionCreatedAt\":\"2025-10-17 02:16:05\",\"video\":\"\",\"vorDoi\":\"10.1007/s11306-025-02379-0\",\"vorDoiUrl\":\"https://doi.org/10.1007/s11306-025-02379-0\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-7764688\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-7764688\",\"identity\":\"rs-7764688\",\"version\":[\"v1\"]},\"buildId\":\"8U1c8b4HqxoKbykW_rLl7\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}