{"paper_id":"282425b1-5098-41de-b220-4afe16b03af7","body_text":"1 \n \nIn-source fragmentation in mass spectrometry-based proteomics: prevalence, 1 \nimpact, and strategies for mitigation 2 \nThorben Schramm 1, Ludovic Gillet 1, Viviane Reber 1, Natalie de Souza 1,2,3, Matthias Gstaiger 1, 3 \nPaola Picotti1,* 4 \n1Institute of Molecular Systems Biology, Department of Biology, ETH Zürich, Otto-Stern-Weg 3, 5 \n8093 Zürich, Switzerland 6 \n2Department of Quantitative Biomedicine, Universität Zürich, Winterthurerstrasse 190, 8057 7 \nZürich, Switzerland 8 \n3Institute of Molecular Health Sciences, Department of Biology, ETH Zürich, Otto-Stern-Weg 7, 9 \n8093 Zürich, Switzerland 10 \n*Correspondence: ppicotti@ethz.ch 11 \n 12 \nAbstract 13 \nPeptide-level analyses are becoming increasingly popular in mass spectrometry-based proteomics and 14 \nare being applied, for example, in immunopeptidomics, structural proteomics, and analyses of post-15 \ntranslational modifications. In such analyses, peptides that are not biologically meaningful but instead 16 \narise as artifacts prior to mass spectrometry analysis pose the risk of data misinterpretation. Here, we 17 \ndescribe an approach based on retention time analysis and precise chromatographic peak matching to 18 \nidentify peptides generated by in-source fragmentation (ISF), which occurs between chromatographic 19 \nseparation of peptide mixtures and the first mass filter of a tandem mass spectrometer (MS). To 20 \nunderstand the prevalence and properties of ISF, we generated 13 proteomics datasets and analyzed 21 \nthem along with additional 25 previously published datasets spanning a broad range of sample types, 22 \nMS, and proteomics approaches including classical bottom-up proteomics, immunopeptidomics, 23 \nstructural proteomics, and phosphoproteomics. We found that, in typical trypsin-digested samples on 24 \naverage 1 % of fully-tryptic peptides and 22 % of semi-tryptic peptides originated from ISF. However, 25 \nwe observed large variations between datasets, and in-source fragments exceeded, in some cases, a 26 \nthird of the total peptide identifications. The extent of ISF was dependent on the peptide sequence, the 27 \ninstrument, method parameters, and sample complexity. Although ISF did not impair relative 28 \nquantification across samples, it generated peptides that could be misinterpreted qualitatively, inflated 29 \npeptide identifications, and comprised up to 37 percent of peptides shorter than 9 amino acids in 30 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n2 \n \nimmunopeptidomics datasets. We propose that, for peptide-centric applications, our open-source ISF 31 \ndetection approach be used to re-annotate peptides generated by ISF and remove them to avoid 32 \nmisinterpretation of data. ISF is an increasing concern with improving mass spectrometers, as they 33 \nenable detection of an ever-increasing number of m/z features, including low abundance features like 34 \nISF products. Our work thus addresses a growing issue in proteomics and presents solutions to mitigate 35 \nthe impact of in-source fragment peptides. In the future, improved feature detection algorithms may 36 \nenable elucidation of new ISF patterns affecting side chains that have been missed so far, which could 37 \ncontribute to explaining the vast space of as-yet unannotated proteomics data. 38 \n 39 \nKeywords 40 \nproteomics / in-source fragmentation / mass spectrometry / retention time / 41 \nimmunopeptidomics / phosphoproteomics 42 \n 43 \nIntroduction 44 \nMass spectrometry (MS) has emerged as a standard tool for the analysis of complex metabolite1,2 45 \nand protein 3,4 samples. While MS has been broadly applied in scientific and industrial research 46 \nfor decades, an increasing number of MS methods are now applied in clinical diagnostics as well5. 47 \nIn MS, analytes such as metabolites and peptides are ionized prior to measurement 6, with 48 \nelectrospray ionization (ESI)7 being one of the most frequently used ionization methods in current  49 \napproaches. During ESI, a liquid phase containing analytes is sprayed through an injection needle 50 \nand subjected to high temperatures (25 – 500 °C) 8 and electric potentials (500 - 4500 V) 6. This 51 \nleads to the ionization of the analytes and evaporation of the liquid phase. Although ESI is often 52 \ndescribed as a soft or gentle ionization technique 6,9–11, analytes can be chemically modified or 53 \npartially fragmented during ESI11–13. Such in-source fragmentation (ISF) is typically considered an 54 \nartifact because it creates new molecular species that can be misinterpreted as biologically 55 \nrelevant. 56 \nIn the metabolomics field, ISF has been extensively studied 10–18, and depending on the study, 5 57 \nup to 70 % of the data have been reported to originate from ISF10,11,13,15–18. Several bioinformatic 58 \ntools and strategies have been developed to annotate in-source fragments in metabolomics 59 \ndata10,11,15,19. If analytes are separated by chromatography before MS measurements, in-source 60 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n3 \n \nfragments can readily be identified because they share the same retention time and 61 \nchromatographic elution profiles as their parental molecule 11,19. In cases where retention times 62 \nare not available, for example in high-throughput metabolomics approaches that rely on flow 63 \ninjection analyses20, network approaches can identify in-source fragments13,15. 64 \nIn contrast to the metabolomics field, ISF has been much less studied in the context of 65 \nproteomics. A possible reason for this is that the primary goal of classical bottom-up proteomics 66 \nhas been the differential abundance analysis of proteins, which was dominated by data 67 \ndependent acquisition (DDA) MS methods and by peptide search engines that mostly focus on 68 \nidentifying fully-tryptic peptides. In this context, the probability of detecting fully-tryptic peptides 69 \nas artifacts arising from ISF is low because for such a mis-annotation to occur, a longer parental 70 \npeptide arising from missed trypsin cleavage events would have to fragment into a fully-tryptic 71 \npeptide of sufficient intensity to be selected for MS2 measurements by DDA. In protein 72 \nabundance analysis, potential effects of ISF on the accuracy of peptide quantification would be 73 \ndiminished during aggregation of peptide quantities to protein quantities. Therefore, ISF has not 74 \nbeen considered of major concern so far. In support of this, a previous study identified ISF as a 75 \nmajor source of half-tryptic peptides in DDA MS, but showed that it had a low impact for complex 76 \nsamples measured, as in-source fragments represented only 1 - 3 % of the analyzed DDA data 21. 77 \nHowever, data independent acquisition (DIA) MS methods have gained much in popularity in 78 \nrecent years 22,23. Together with new mass spectrometers with drastically increased sensitivity, 79 \nthey enable the detection of an ever-increasing number of peptides, and it is unclear whether 80 \nsuch approaches are more prone to identifying in-source fragments than previous methods. 81 \nFurther, it is known that instrument parameters have an impact on ISF, but the extent of these 82 \neffects and the relative impact of different parameters is unclear. To complicate this matter, the 83 \nnumber of proteomics applications that rely on peptide-centric analyses to quantify half- or non-84 \ntryptic peptides, or post-translational modifications (PTMs), is steadily increasing 24. In these 85 \nanalyses, the quantitative impact of ISF is expected to be more pronounced as peptide quantities 86 \nwould not be aggregated to protein quantities. In addition, qualitative effects of ISF on peptide 87 \nidentification could substantially compromise data interpretation. For instance, in mass 88 \nspectrometry-based immunopeptidomics or structural proteomics approaches using limited 89 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n4 \n \nproteolysis25,26, biological findings rely on the precise characterization of every peptide from a 90 \ngiven protein, including those with semi- or non-tryptic termini. It remains unclear how ISF 91 \nimpacts peptide-centric analyses, quantitatively and qualitatively. 92 \nWe have therefore revisited the prevalence, characteristics, and impact of ISF in the context of 93 \nproteomics focusing on DIA datasets acquired with state-of-the-art mass spectrometers. We 94 \nsystematically evaluated a range of instrument parameters and sample types, with particular 95 \nemphasis on assessing the implications of ISF for peptide-centric proteomics approaches. 96 \nFurther, we provide practical solutions on how to mitigate ISF in proteomics. 97 \n 98 \nResults 99 \nDetecting in-source fragment peptides in proteomics data 100 \nTo analyze the prevalence and impact of ISF products, we first sought to develop a bioinformatic 101 \npipeline for their detection and quantitative analysis. Peptides are usually separated by liquid 102 \nchromatography prior to electrospray ionization and mass spectrometry in bottom-up 103 \nproteomics (Supplementary Fig. 1). Since ISF occurs after chromatographic peptide separation 104 \nand prior to the first mass filter, in-source fragments and their respective parent peptides have 105 \nthe same retention time profiles despite being of different molecular mass. Fragment peptides 106 \nalso share their amino acid sequence with their parents. We therefore developed an algorithm 107 \nto detect peptide ISF by identifying peptides with shared sequences that co-elute, or, more 108 \naccurately for in-source fragments, have the same retention time (RT). To illustrate our approach, 109 \nwe consider a theoretical example with 8 different peptides that share a sequence (Fig 1.a), which 110 \nyield 28 unique peptide pairs. For each pair, the difference between the apex retention times 111 \n(∆RT) was calculated by subtracting the RT of the shorter (or equally sized) peptide (peptide 1) 112 \nfrom the RT of the longer peptide (peptide 2) (Fig. 1.b). Peptides that resulted from ISF, which 113 \nwere peptides I, II, VI, and VII in our example, should have small or negligible ∆RTs if paired with 114 \ntheir respective ISF parent peptides. We thus identified these peptides by defining a ∆RT cutoff 115 \nand inspecting which peptide pairs have a ∆RT that fell below it. 116 \nAs RT distributions, chromatographic peak shapes, and resolution can vary between instrumental 117 \nsetups or runs, we determined the “co-elution” ∆RT cutoff for each sample individually. To obtain 118 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n5 \n \nthis cutoff, we analyzed the data at the peptide precursors level, as throughout this study if not 119 \nstated otherwise, and used the ∆RT distribution of peptides that differ only in charge and not in 120 \nmass, which were peptides IV and V in our example. Because these peptides must have the same 121 \nRT, they provide ground truth for the natural variation of ∆RT values of co-eluting peptides (Fig. 122 \n1.c). As false-discovery-rates are commonly set at 1 % for peptide searches, a small fraction of 123 \npeptide pairs will contain a falsely annotated peptide resulting in a continuous uniform 124 \ndistribution of ∆RT values corresponding to these false positives. The ∆RT cutoff for co-eluting 125 \npeptides is defined as the ∆RT at which the normal distribution of true-positives intersects with 126 \nthe continuous uniform distribution of the false-positives. In the next step, peptide pairs were 127 \nclassified as pairs with ISF evidence, if their ∆RT was below the ∆RT cutoff, or as pairs without ISF 128 \nevidence (Fig. 1.d). In the last step of our ISF detection, peptide pairs were uncoupled and, in 129 \ncase of ISF evidence, shorter, type 1 peptides annotated as fragment peptides, and longer, type 130 \n2 peptides as parent peptides (Fig. 1.e). If there was no ISF evidence, peptides were referred to 131 \nas non-ISF peptides. 132 \nIn case of a parent peptide yielding more than one fragment peptide, the intermediary fragments 133 \nwill still be annotated as fragments but not as parent peptides, implying that parent and fragment 134 \npeptide annotations are mutually exclusive. Multiple in-source fragments of a parent peptide 135 \nalso enable the construction of ISF networks, in which each peptide is a node and each edge a 136 \npeptide pairing with a ∆RT. These fragmentation networks enable the identification of in-source 137 \nfragments that share the same parent and can be used to define new peptide groups, for which 138 \nall individual peptide quantities of each group can be aggregated to single peptide group 139 \nquantities. We further identify C-terminal ISF if the N-terminal part of the parent is detected as 140 \na fragment, and vice versa for N-terminal ISF. We denote the first amino acid at the C-terminal 141 \nside of the ISF site F1, and the first amino acid at the N-terminal side F1’, analogous to the 142 \nnomenclature of proteolytic cleavage sites (P1, P1’, etc.). While the detection of F1 and F1’ 143 \nimplies an ISF that alters the amino acid sequence of the parent peptide, it is in principle possible 144 \nthat an ISF does not change the amino acid sequence but affects a side chain or PTM. As some 145 \nPTMs like oxidation of methionine are included in most peptide searches by default, we also 146 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n6 \n \nconsider PTM in-source fragmentations (PTM ISF), defined as an ISF event that detaches the PTM 147 \nwithout altering the amino acid sequence. 148 \nWe tested our ISF detection algorithm on DIA data from a Saccharomyces cerevisiae 288c sample 149 \nspiked with 8 proteins, which we used as a reference sample throughout this study. We chose 150 \nthis sample as a reference because we suspected that peptides with high concentrations were 151 \nmore likely to yield measurable in-source fragments, and this sample yielded such peptides (from 152 \nthe spiked in proteins) in a controlled fashion without compromising on overall sample 153 \ncomplexity. Our model identified a ∆RT cutoff of ± 0.129 min for co-eluting peptides with a recall 154 \nof 97.1 % for the true positive peptide pairs, which were the peptides of equal mass but different 155 \ncharge states (Supplementary Fig. 2a). This ∆RT cutoff was more conservative than the median 156 \nfull width at half maximum (FWHM), which was 0.19 min and is often used to determine peak 157 \nseparation. Next, we used the ∆RT cutoff to determine ISF peptides and found that 24.4 % of all 158 \npeptide pairs that differed in molecular mass fell within the cutoff (Supplementary Fig. 2b). This 159 \nresult provides the first evidence that in-source fragmentation is prevalent in DIA proteomics 160 \ndata. 161 \nIn the ∆RT distribution, we further observed a broad peak of peptide pairs around -10 min. We 162 \nfound that this was due to peptides carrying an oxidation group as PTM (Fisher’s exact test 163 \nchecking for enrichment between -13 and -5 min, right-tailed, P-value < 1e-200). We also 164 \nobserved a long tail of the ∆RT distribution for ∆RT > 0, which was expected because peptide 165 \npairs that were not due to ISF typically have a longer type 2 peptide that is often stronger retained 166 \non the chromatographic column than the shorter type 1 peptide. 167 \nTo our knowledge, there are no published, open access software tools to detect ISF in DIA 168 \nproteomics data to which we could compare our ISF detection algorithm. However, we compared 169 \nit to another, commercially available and unpublished algorithm that was developed in parallel 170 \nto our study as a feature for Spectronaut 20 (Biognosys). For our yeast reference sample, 97.7 % 171 \nof all ISF annotations across n = 4 technical replicates (4716 out of 4828) were identical, showing 172 \nbroad agreement between both algorithms (Supplementary Fig. 3) and providing a cross-173 \nvalidation. 174 \nIn-source fragmentation is abundant 175 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n7 \n \nNext, to determine how abundant peptide ISF is in data from typical trypsin-digested bottom-up 176 \nproteomics samples measured by DIA, we analyzed 16 published datasets, acquired with 9 177 \ndifferent mass spectrometers models, from different research groups (Supplementary Tab. 1) 27–178 \n42. These datasets comprised 317 samples from a broad range of sources, including Escherichia 179 \ncoli, Saccharomyces cerevisiae , Arabidopsis thaliana , Mus musculus , and Homo sapiens, and 180 \ncomprising both laboratory cultures and human biopsies. The number of detected protein groups 181 \nranged from 3 to 6,180. In-source fragments made up 3.4 ± 3.2 % of the total peptide precursor 182 \nidentifications (mean and standard deviation, n = 317 samples) across the datasets, and 4.8 ± 4.2 183 \n% of the total peptide intensity (Fig. 1a). However, we observed a large variation in the extent of 184 \nISF, ranging from 0.2 ± 0.1 % by intensity (PXD062917, n = 9 samples; 0.5 ± 0.5 % by 185 \nidentifications) to 15.5 ± 0.3 % by intensity (PXD069457, n = 6; 13.1 ± 0.1 % by identifications). 186 \nIncreasing sample complexity generally yielded lower shares of in-source fragments within a 187 \ndataset (% by identifications and by intensities), although not every sample with a low complexity 188 \nhad high levels of ISF. To better understand how sample complexity affects ISF, we prepared the 189 \nfollowing 7 trypsin-digested samples of ascending complexity and compared the prevalence of 190 \nISFs: (1) mix of 11 peptides, (2) mix of 8 proteins, (3) E. coli  BW25113 lysate, (4) affinity-191 \npurification MS (APMS) from a H. sapiens kinase pull-down (5) S. cerevisiae S288c lysate, (6) H. 192 \nsapiens HEK-293 lysate, and (7) the tribrid proteome sample obtained by combining samples 3, 193 \n5, and 6. We detected ISF peptides in all samples (Fig. 2.b). The tribrid proteome sample (n = 4 194 \ntechnical replicates) had the lowest level of ISF, with fragments representing 4.01 ± 0.06 % of 195 \ntotal peptide intensity (2.591 ± 0.005 % of identifications, mean and standard deviation), while 196 \nthe mix of 8 proteins had the highest value (24.1 ± 1.3 % by intensity; 31.0 ± 0.5 % by 197 \nidentifications). Across all samples, both the intensities as well as the numbers of fragment 198 \npeptide identifications increased with decreasing sample complexity. The absolute number of 199 \nparent peptides was on par with the number of fragments, except for the low complexity samples 200 \n(peptide and protein mixes), which had more fragments than parents (Fig. 2.c). Although we 201 \nobserved fewer parent peptides than non-ISF peptides, the intensities of the parent peptides 202 \nwere in the same range as the non-ISF peptides (Fig. 2.c). This also meant that at least 32 % ( H. 203 \nsapiens) and up to 95 % (mix of peptides) of the total peptide intensity could be associated with 204 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n8 \n \nISF, combining the intensities of fragments and parents (Fig. 2.d). This was consistent with our 205 \nobservations in the published datasets, in which on average 37 ± 22 % of the total peptide 206 \nintensity could be associated with ISF, with large variation between individual datasets. Across 207 \nall datasets (published and our new data reflecting a range of complexity), 82 % of the in-source 208 \nfragments were semi-tryptic peptides (Fig. 2.e), and in-source fragments represented 0.2 – 83.3 209 \n% of semi-tryptic peptides, depending on the sample (Fig. 2.f). 210 \nTaken together, these results showed that ISF varies strongly between datasets, depends on 211 \noverall sample complexity, and can account for a large share of the mass spectrometry signal. ISF 212 \nis especially prevalent in low complexity samples, highlighting the importance of ISF detection in 213 \nthose cases. 214 \n 215 \nCharacterization of ISF peptides 216 \nTo learn general properties of the observed ISF peptides, we analyzed combined data with 217 \n2,407,805 identified peptides from all the (published and newly generated) datasets described 218 \nabove. 30.3 % of the peptides shared their sequence with another peptide and formed around 219 \n13.7 million peptide pairs (Fig. 3.a). 1,222,589 of the peptide pairs (8.9 %) showed evidence of 220 \nISF, of which 67.9 % had sequence-altering ISF, and 26.6 % were fragment-to-fragment pairs.  221 \nMost fragment peptides originated from N-terminal ISF of fully-tryptic peptides and had a C-222 \nterminal arginine or lysine (Supplementary Fig. 4.a). Further, non-ISF peptides had mostly lower 223 \nintensities than in-source fragments and parents, with parent peptides having higher intensities 224 \nthan fragments (Fig. 3.b and Supplementary Fig. 4.b). This indicated that the detected in-source 225 \nfragments and parents either derived from abundant proteins, were especially amenable to 226 \nmeasurement by MS, or both. As expected, in-source fragment peptides often had a lower 227 \nnumber of charges than their respective parent peptides and were shorter than the parent and 228 \nnon-ISF peptides, as 47.3 % of the ISF peptide pairs showed losses in charge and 92.5 % showed 229 \nlosses of amino acids due to ISF (Supplementary Fig. 4.c, d, e, and f). 230 \nTo understand if ISF events had a bias for certain amino acid bonds, we analyzed how often an 231 \namino acid occurred at F1’ or F1 of an ISF site. Valine followed by leucine, isoleucine, and alanine 232 \nwere most abundant at F1’, while proline was most abundant at F1, followed by glycine and 233 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n9 \n \nalanine (Fig. 3.c). An enrichment analysis of peptides bonds at ISF sites, using all possible peptide 234 \nbonds of detected peptides as null distribution (two-tailed Fisher’s exact test), showed that 56 235 \nout of 80 possible peptide bonds with short, hydrophobic amino acids (valine, isoleucine, leucine, 236 \nand alanine) at F1’ showed a significant enrichment (Benjamini-Hochberg 43 adjusted P-values < 237 \ne-20), and 17 out of 20 bonds with proline at F1 were enriched at ISF sites (Fig. 3.d). The 238 \naspartate-proline bond is known to be susceptible to non-enzymatic cleavage under acidic 239 \nconditions44. Here, we did not observe an enrichment of the aspartate-proline bond at 240 \nfragmentation sites but rather a depletion. However, we detected enrichment of the aspartate-241 \nproline bond at non-tryptic termini of semi-tryptic non-ISF peptides (Supplementary Fig. 5). 242 \nAs they made up over a quarter of the ISF peptide pairs, fragment-to-fragment pairs indicated 243 \nthat multiple ISF events of peptides are abundant. We thus constructed ISF networks with 244 \nsequence-altering ISFs to better understand this effect and its prevalence but also to assign ISF 245 \npeptides to new peptide groups. Merging information from peptides with multiple charges, we 246 \nobserved that 72.4 % of the ISF networks had two unique peptide sequences and 72.5 % a one-247 \nto-one ratio between in-source fragment and parent peptides (Fig. 3.e), implying that the 248 \nmajority of ISFs were from a single ISF parent to a single fragment peptide. However, the 249 \nremaining 27.6 % of the ISF networks contained more than two unique peptide sequences, with 250 \nthe largest ISF network having 23 unique peptide sequences (Supplementary Fig. 6). 27.3 % of 251 \nthe ISF networks also had an imbalance between the number of in-source fragments and parent 252 \npeptides (Fig. 3.e). Detecting both N- and C-terminal fragments from a single ISF event appeared 253 \nto be challenging as we only observed 9 cases across the 23 datasets, in which the combined 254 \nsequences of two in-source fragments matched the parent sequence. 255 \nThese results show that peptide ISF mostly occurs at specific peptide bonds prone to fragmenting 256 \nin the gas phase and often creates a complex pattern of peptide products and m/z features, which 257 \ncan be deciphered by network analysis. 258 \n 259 \nIn-source fragmentation of PTMs 260 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n10 \n \nIf ISF only removes the PTM group of a peptide, the resulting fragment peptide can appear as an 261 \nunmodified peptide, which could be mistaken for an authentic identification of the unmodified 262 \npeptide. We analyzed how often this effect occurs for several PTMs. 263 \nMethionine oxidation and N-terminal acetylation are commonly observed PTMs in proteomics 264 \ndata; we included them as variable modifications in the peptide searches of our combined 265 \ndataset and studied if they undergo ISF. We detected that ISF of these PTMs comprised 5.5 % of 266 \nthe total ISF events (Fig. 3.a) and that the intensity differences between in-source fragments and 267 \nparents were much smaller for PTM ISF than for sequence-altering ISF (Supplementary Fig. 4.b). 268 \nMost of the cases involved methionine oxidation, and only a few N-terminal acetylation (0.4 % of 269 \nall PTM ISF) (Fig. 4.a). While there was a large variation in the number of methionine oxidation 270 \nISF between the datasets, we did not observe a correlation with the number of peptides carrying 271 \na methionine oxidation (Fig. 4.b). This indicated that the amount of PTM ISF was dependent on 272 \nother factors, e.g. instrument parameters. 273 \nIt is known that phosphorylated metabolites like ATP tend to lose phospho-groups due to in-274 \nsource fragmentation 11, and we thus assessed the prevalence of ISF events affecting 275 \nphosphopeptides. We first directly re-searched our S. cerevisiae  S288c dataset for 276 \nphosphopeptides. Second, we analyzed S. cerevisiae  S288c digests after phosphopeptide 277 \nenrichment by titanium dioxide chromatography 45, and third, we included two additional, 278 \npublished datasets that used various phosphopeptide enrichment strategies (PXD014525 46 and 279 \nPXD04448247). As expected, the number of phosphopeptides were much higher in the three 280 \nphospho-enriched datasets, where they made up 68 - 91 % of all identified peptides as compared 281 \nto 0.6 % in the standard proteomics dataset (Fig. 4.c). Similarly, the number of ISF events leading 282 \nto loss of a phospho-group was higher in the phosphoproteomics datasets (0.17 - 0.63 % of in-283 \nsource fragments by identifications, 0.46 – 1.06 % by intensity, mean values, number of samples 284 \nn = 4 (phospho-enriched yeast dataset), n = 18 (PXD014525), and n = 10 (PXD044482)) than in 285 \nthe standard proteomics dataset (0.0048 ± 4e-5 % by identifications, 0.0037 ± 1e-4 % by intensity, 286 \nmean and standard deviation, n = 4 technical replicates) (Fig. 4.d). In the two published datasets, 287 \nwe observed a bias of phospho-group ISFs towards phosphorylated tyrosines (Y) indicating that 288 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n11 \n \nthis amino acid is more prone to phospho-group ISF than phosphorylated serines (S) and 289 \nthreonines (T) (Supplementary Fig. 7). 290 \nIn conclusion, PTM ISF occurs at a detectable level in proteomics data and varies in a dataset and 291 \nPTM-dependent manner, although it was less prevalent than ISF of regular peptide bonds. 292 \n 293 \nPeptide ISF is instrument and parameter dependent 294 \nThe ion funnel design and its radiofrequency (RF) are known to have an impact on the ISF of 295 \npeptides48, the temperature of the ion transfer capillary and the RF affect lipid ISF 49, and 296 \nincreasing accumulation times during trapped ion mobility spectrometry (TIMS) time-of-flight 297 \nmass spectrometry (TOF) also increases ISF50. Based on this information, we expected ISF to vary 298 \nbetween different mass spectrometers. To assess the extent of such variation, we measured a 299 \ntrypsin-digested yeast reference sample mixed with 8 pure proteins on three Thermo Scientific 300 \nmass spectrometers, an Orbitrap Exploris 480, an Orbitrap Astral, and a Q Exactive Plus Hybrid 301 \nQuadrupole-Orbitrap (QE+), each using LC and DIA methods based on previously established 302 \nworkflows51,52. With 1.07 ± 0.02 and 1.51 ± 0.02 % (mean and standard deviation, n = 4 technical 303 \nreplicates), the QE+ and the Astral mass spectrometers yielded a much lower fraction of peptide 304 \nidentifications corresponding to in-source fragments than the Exploris 480 (5.68 ± 0.03 %) 305 \n(Supplementary Fig. 8). This was also reflected in the total intensity of in-source fragments (QE+: 306 \n2.2 ± 0.3 %, Exploris 480: 8.8 ± 0.2 %, Astral: 3.70 ± 0.05 %). Thus, ISF was indeed highly dependent 307 \non the instrument. 308 \nAs the RF and the ITC temperature were shown to affect ISF48,49, we subsequently quantified the 309 \nimpact of these two parameters as well as the electrospray voltage on peptide ISF using the 310 \nOrbitrap Exploris 480. We used the yeast reference sample again and measured it at a spray 311 \nvoltage fixed at 2500 V with seven variations of the funnel radio frequency and six variations of 312 \nthe ITC temperature. We also measured the same sample with a relative funnel radio frequency 313 \nfixed at 50 % with six variations each of the other two parameters (ITC temperature and spray 314 \nvoltage), which amounted to a total of 72 unique sets of parameters tested. 315 \nThe chosen parameters resulted in a broad range of ISF, from low (0.2 % of peptide precursor 316 \nidentifications across all samples, 301 fragments at 100 °C, 2500 V, 10 % RF, n = 1 replicate) to 317 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n12 \n \nhigh (4.9 %, 6566 fragments at 350 °C, 2500 V, 60 % RF) (Fig. 5.a and Supplementary Fig. 9.a). 318 \nIncreasing the funnel RF increased both the relative intensity and the number of in-source 319 \nfragments (Fig. 5.b). Simultaneously, it only increased the overall signal up to a RF of 40 %, 320 \nindicating that higher funnel RF values are disadvantageous. The effect of the ITC temperature 321 \non ISF was dependent on the funnel RF and led to increases in ISF (in signal and identifications) 322 \nat high RF values but had little effect on ISF at low RF values. In contrast to both other parameters, 323 \nthe electrospray voltage had little effect on overall signals, number of peptide identifications, 324 \nand ISF, except for the very low setting of 1500 V that decreased overall signal drastically. 325 \nThe number of precursor identifications is often used as a quality control measure during data 326 \nacquisition. Because instrument parameters increase the number of in-source fragments while 327 \ndecreasing the number of identifications of genuine peptides with no evidence of fragmentation, 328 \nnot accounting for ISF will lead to an overestimation of identifications and can thus lead to 329 \nchoosing suboptimal conditions (Fig. 5.a). 330 \nIn conclusion, these results show that ISF was strongly dependent on the instrument, funnel RF, 331 \nand ITC temperature, indicating that proteomics experiments can be optimized to avoid high ISF 332 \nrates. Since ISF can lead to an overestimation of the total peptide identifications, minimizing it 333 \nshould generally be part of method optimization efforts alongside the optimization of other 334 \nparameters like total peptide identifications and signal intensities. 335 \n 336 \nISF does not impair relative quantification 337 \nTo test if ISF has an impact on relative quantification of peptide precursors, peptide groups, and 338 \nprotein groups, we conducted an experiment in which we mixed, at equal volumes, a yeast 339 \nsample at a fixed concentration with 8 samples each with a different concentration of E. coli 340 \npeptides (Fig. 6.a). Peptide precursor fold changes were calculated relative to the sample with 341 \nthe highest E. coli proteome concentration. 342 \nFor all samples, the yeast proteome remained stable around a fold change of 1, whereas the E. 343 \ncoli peptide precursors followed the dilutions (Fig. 6.b). At lower concentrations of the E. coli 344 \nproteome, the peptide precursor fold changes did not match the theoretical values as accurately 345 \nas at higher concentrations, which was probably due to ion suppression and poor peptide 346 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n13 \n \ndetectability. However, comparing the data between non-ISF, fragment, and parent peptide 347 \nprecursors revealed that the non-ISF peptide precursors had broader fold change distributions 348 \nthan the other peptide precursors (Fig. 6.b and Supplementary Fig. 10) and that coefficients of 349 \nvariation (n = 3 technical replicates) of non-ISF peptide precursors were often larger than those 350 \nof in-source fragments and parents (Fig. 6.c, left plot, P < e-200 for all pairwise comparisons, 351 \nWilcoxon rank sum test, two-tailed). Because the dilution factors were known, we could calculate 352 \na relative error between the measured fold changes of the precursor peptide quantities and the 353 \nexpected, theoretical fold changes. These relative errors showed that the non-ISF precursor 354 \npeptides recovered the expected, theoretical values better than the in-source fragments by only 355 \na small margin (0.015 difference in the median relative error); interestingly, the ISF parents 356 \nrecovered the theoretical values the best, which is probably due to their generally good 357 \nmeasurability and high signals (Fig. 6.c, right plot, P < e-13 for all pairwise comparisons, Wilcoxon 358 \nrank sum test, two-tailed). These results show that ISF has little effect on relative quantification 359 \nof peptide precursors. 360 \nSo far, our analyses have been at the peptide precursor level but most applications in proteomics 361 \nrely on peptide or protein level analyses, in which data of peptide precursors that only differ in 362 \ncharge or even PTMs are combined to improve quantification and reduce redundancy. Since ISF 363 \nonly had a minor impact on relative peptide precursor quantification, we tested if assigning in-364 \nsource fragments to their parental peptide groups or protein groups provide an advantage, e.g. 365 \nby reducing errors, or if in-source fragments should simply be filtered out. 366 \nFirst, we used network analysis to group in-source fragments with their respective ISF parents, 367 \nand calculated peptide quantities by summing all peptide precursors intensities of each peptide 368 \ngroup. Subsequently, we compared the peptide groups of ISF parents with and without added in-369 \nsource fragments as well as all the peptide groups unaffected by ISF with each other. The CVs of 370 \npeptide groups that contain in-source fragments were smaller than the same peptide groups 371 \nwithout in-source fragments, although the difference was minor (Fig. 6.d, P = 0.013, Wilcoxon 372 \nrank sum test, two-tailed). The absolute relative errors showed that parent peptides (without 373 \nmerging with in-source fragments) much better followed the theoretical fold changes than non-374 \nISF peptides (Fig. 6.d). However, these data also revealed that merging in-source fragment data 375 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n14 \n \nwith their parental peptide groups worsens quantification ( P = 2.6e-14, Wilcoxon rank sum test, 376 \ntwo-tailed) but again the effect is small. 377 \nSecond, we calculated protein-level quantities by summing up all peptide precursors that belong 378 \nto a certain protein, once with in-source fragments and once without them. Comparing the CVs 379 \nand relative errors revealed that in-source fragments had close to no effect on the protein-level 380 \nquantities (Fig. 6.e, P > 0.4, Wilcoxon rank sum test, two-tailed). 381 \nTaken together, these results show that ISF has little effect on relative quantification, that in-382 \nsource fragments and parent peptides are among those peptides that are usable for relative 383 \nquantification, and that ISF parents were the best peptides for quantification. These findings are 384 \nlikely because in-source peptides have typically higher intensities than non-ISF peptides. 385 \nRegrouping in-source fragments with their parent peptides or including them in protein groups 386 \nalso had little effect, with a small tendency to reduce CVs of the new peptide groups but at the 387 \ncost of distorting relative quantities across samples. In sum, while in-source fragments present 388 \nno major challenge for quantitative analyses in general, they are only a proxy of their parent 389 \npeptides, with slightly worse quantitative performance (CVs and relative errors), and we 390 \ntherefore recommend filtering them out. 391 \n 392 \nImpact of ISF on immunopeptidomics 393 \nSince ISF of tryptic peptides produces mostly peptides with semi-tryptic sequences 394 \n(Supplementary Fig. 4), we wondered how ISF would impact peptide-centric proteomic 395 \napproaches relying on relaxed trypsin specificity searches, such as immunopeptidomics. During 396 \nthe immune response, peptides produced by proteolytic in vivo-processes are presented to killer 397 \nT cells as antigens25,53. Identifying those antigens is crucial for a better understanding of immune 398 \nresponses but also for the development of new medical treatments. Misidentification of an in-399 \nsource fragment as true biological antigen can thus have costly consequences. The peptide 400 \nantigens show narrow length distributions, typically between 9 and 14 amino acids, depending 401 \non the peptide class, and they do not show specific terminal amino acids 54. Therefore, 402 \nimmunopeptidomics requires unspecific searches for peptide identification, which causes an 403 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n15 \n \nexponential increase of the search space and increases the risk of confusing in-source fragments 404 \nwith biologically relevant peptides. 405 \nHere, we used four datasets from the literature (PXD022950, PXD051490, PXD054417, and 406 \nPXD05880)55–58 to assess the impact of ISF on immunopeptidomics. Across the datasets, we 407 \nobserved large variation in the number of in-source fragments (Fig. 7.a), and similarly to our 408 \nprevious results, most ISFs occurred at the N-terminus (Fig. 7.b). As expected, the length 409 \ndistribution of peptides from immunopeptidomes was much narrower than in our data of trypsin-410 \ndigested proteomes, but in-source fragments were again typically 1 to 2 amino acids shorter than 411 \nparent peptides (Fig. 7.c). In the datasets PXD058880 and PXD022950, in-source fragments made 412 \nup 16.6 and 36.8 % of all peptides below 9 amino acids (2.6 and 2.3 % by intensity) but only 0.07 413 \nand 0.92 % of longer peptides (0.04 and 0.8 % by intensity), respectively. The other datasets that 414 \nwe analyzed (PXD054417 and PXD05149) had much less ISF; the fragment identifications were 415 \nagain higher for peptides below 9 amino acids (0.53 and 0.90 % by identifications, 0.003 and 0.17 416 \n% by intensity) than for longer peptides (0.11 and 0.07 % by identifications, 0.06 and 0.01 % by 417 \nintensity). 418 \nThese results show that, in immunopeptidomics based on HLA-I peptides, ISF is likely to impact 419 \npeptides shorter than 9 amino acids. Therefore, a simple strategy to mitigate this would be to 420 \nexclude peptides shorter than 9 amino acids before further analyses. Alternatively or in addition, 421 \napplication of our RT-based detection algorithm to immunopeptidomics data will enable 422 \nidentifying in-source fragments of any length that contaminate the pool of biologically relevant 423 \npeptides. 424 \n 425 \nIn-source fragmentation in limited proteolysis data 426 \nLimited proteolysis coupled with mass spectrometry (LiP-MS) is another peptide-centric 427 \nproteomic approach, which enables the global analysis of protein structural changes in proteome 428 \nextracts26. It relies on proteases with broad specificity that cleave proteins for a brief period of 429 \ntime such that structural properties govern cleavage events. These primary cleavages can 430 \nproduce large protein fragments that may not be directly amenable for MS and are thus further 431 \ndigested by trypsin under denaturing conditions in a second step. LiP-MS therefore intrinsically 432 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n16 \n \nfeatures many more semi-tryptic peptides than other proteomic datasets. As most in-source 433 \nfragments appear as semi-tryptic peptides based on our analyses, we asked how many of the 434 \nidentified semi-tryptic peptides in a typical LiP experiment were due to ISF and therefore could 435 \nmistakenly be identified as structurally relevant peptides. 436 \nWe used three published LiP datasets to study this (PXD022297, PXD033826, and PXD034606)59–437 \n61 and found that 0.6 % to 2.0 % of all identified peptides were due to ISF (1.3 to 3.6 % of the total 438 \nintensity) (Fig. 7.d). However, the in-source fragments made up only 0.9 to 3.7 % of the semi-439 \ntryptic peptides (2.1 to 5.9 % by intensity), showing that most semi-tryptic peptides were indeed 440 \ntrue-positives. In these LiP datasets, only 7.9 to 18.1 % of the in-source fragments appeared were 441 \nnon-tryptic (Fig. 7.f), corresponding to 0.08 – 0.12 % of all identifications (0.07 – 0.14 % by 442 \nintensity). 443 \nThese results show that ISF had only little impact on LiP datasets as the vast majority of semi-444 \ntryptic peptides were indeed caused by limited proteolysis and not by ISF. While the inherently 445 \nhigh amounts of semi-tryptic peptides in LiP data did yield a small number of in-source fragments, 446 \ndetection of these peptides requires unspecific searches, which are typically not performed for 447 \nstandard LiP datasets. Overall, our data indicates that ISF is unlikely to be an issue for LiP-MS 448 \nanalyses. 449 \n 450 \nCurrent peptide-centric DIA analyses underestimate the extent of ISF 451 \nOur analyses so far focused on DIA datasets because DIA offers higher confidence than DDA in 452 \nassessing co-elution of peptides and therefore in detecting ISF. DIA data however come with an 453 \nimportant caveat: most peptide-centric DIA algorithms assume that a peptide elutes as a single 454 \npeak at a specific retention time. However, if ISF occurs, it is possible that a peptide has more 455 \nthan one peak because it could stem from ISF in addition to a biological source that gives rise to 456 \na genuine peptide. Therefore, current search engines will either miss the identification of a 457 \ngenuine peptide or underrepresent ISF in cases in which a peptide has multiple peaks 458 \n(Supplementary Fig. 11.a and b). In contrast to current DIA analysis approaches, DDA analysis 459 \ndoes not make any assumption regarding the elution of peptides but reports all RTs at which a 460 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n17 \n \nsequence is identified, although RT determination can be imprecise due to dynamic peak 461 \nexclusion settings (Supplementary Fig. 11.C). 462 \nTo understand the extent to which current “single-peak” peptide-centric DIA algorithms 463 \nunderestimate ISF, we analyzed DIA datasets using an approach that relies on modified DDA 464 \nlibraries62 (also see Methods part M.5). In brief, peptides that were identified in multiple DDA 465 \nwindows (each two or more minutes apart) were assigned to different bins of spectral assays 466 \nusing unique identifiers. Searches of DIA data with such modified spectral libraries would thus 467 \nretain assays of the same peptide that are separated by RT and allow analysis of multiple peaks. 468 \nWe used this strategy to generate a modified spectral library from the DDA files acquired for the 469 \nLiP dataset PXD022297 59, which contains many genuine semi-tryptic peptides due to limited 470 \nproteolysis. The library had 44,046 peptide precursor assays and included 2,961 (6.7 %) assays of 471 \nprecursors that were present in more than one RT window. Next, we used this library to search 472 \nthe DIA data from the same dataset and compared the library-based peptide peak identifications 473 \nto those obtained by a “direct” peptide search of the DIA data, which did not rely on a spectral 474 \nlibrary from DDA. This revealed that the library-free DIA extraction identified 1,505 (57 %) 475 \nprecursors of the 2,961 assays with multiple peaks from our modified spectral library at the 476 \nexpected RT, misidentified 940 assays at a different RT, and completely missed 211 assays. Using 477 \nthe ISF parent and fragment annotations of the library-free DIA extraction showed that the 478 \n(direct) peptide-centric DIA analysis underestimated the number of ISF peptides by 218 and the 479 \nnumber of genuine peptide sequences by 933. 480 \nUsing a spectral library for DIA extraction introduces the usual biases of DDA measurements, such 481 \nas undersampling low abundant signals or underrepresenting singly charged precursors. 482 \nTherefore, our results on the multiple peak extractions still present an underestimation of ISF 483 \nand the true number of cases in which a peptide sequence occurs multiple times in a gradient. 484 \nHowever, this analysis shows that current peptide-centric DIA analysis tools that only account for 485 \na single RT per peptide sequence systematically underestimate both the extent of ISF as well as 486 \nthe number of genuine peptides, which are not created by ISF. These results thus encourage the 487 \ndevelopment of peptide-centric algorithms that account for multiple peaks. 488 \n 489 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n18 \n \nDiscussion 490 \nIt is known that ISF affects peptides and can be a source of peptides with non-tryptic termini in 491 \nbottom-up proteomics data21. The proteomics field has evolved in past years, with an increasing 492 \nuse of peptide-centric proteomics approaches like structural proteomics and 493 \nimmunopeptidomics that rely on semi- and non-tryptic peptides. Further, DIA methods that 494 \nenable detection of low intensity m/z features, as opposed to DDA methods that likely miss many 495 \nlow intensity fragments, have found mainstream adaptation, and the development of the new 496 \nmass spectrometers with vastly increased sensitivity has introduced new instrument geometries 497 \nand designs that could affect the prevalence of ISF. It is therefore timely to conduct an 498 \nassessment of ISF in the context of DIA proteomics, with a particular focus on peptide-based 499 \napproaches. 500 \nWe described an approach to re-annotate peptides based on RT patterns. The novelty of our 501 \napproach is to use intra-sample information from peptides with multiple charge states to 502 \nestimate natural distributions of ∆RT values and, based on these, determine ∆RT cutoffs for ISF 503 \ndetection. This approach can therefore account for differences in LC performance between runs, 504 \nbatches, or HPLCs. In the future, our ISF detection could even be further improved by using 505 \ndynamic ∆RT cutoffs that depend on the RT of the peptide peaks. This could address dynamically 506 \nvarying ∆RT value distributions in non-linear gradients. 507 \nWe observed that ISF can affect a large share of a proteomics dataset acquired by DIA and 508 \naccount for more than 50 % of the detected semi-tryptic peptides. We also found that the 509 \nnumbers of detected in-source fragments anticorrelate with sample complexity; this has 510 \nimplications for MS approaches such as crosslinking MS, hydrogen/deuterium exchange mass 511 \nspectrometry (HDX-MS), affinity purification coupled to MS, or (co-)fractionation MS that often 512 \nrely on the analysis of low complexity samples63–66. These observations match well with previous 513 \nreports based on DDA data21. Further, our analysis of the intensity distributions showed that the 514 \nobserved ISF parents often had very high intensities resulting also in fragments with higher 515 \nintensities than of non-ISF peptides. This observation could be the result of a bias of the search 516 \nengine against in-source fragments with low intensities. 517 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n19 \n \nGenerally, we observed a large variation in the amount of detected ISF among different datasets, 518 \nwhich can be explained by different samples complexities, instruments, and chosen parameters, 519 \nall of which we have shown here to have an impact on the prevalence of ISF. The combinatorial 520 \neffect of all these factors makes it difficult to estimate whether ISF could be a problem for a 521 \nparticular sample or dataset and, considering that ISF can in some cases account for more than 522 \n30 % of the peptide identifications, we suggest that ISF products should be detected and 523 \nannotated as a routine data analysis step. 524 \nThe current peptide-centric DIA analysis tools assume that a peptide sequence occurs only once 525 \nacross a chromatogram. Using a RT-based splitting strategy to generate spectral libraries from 526 \nDDA data, we were able to analyze peptides occurring multiple times across the LC gradient in 527 \nDIA data. Our analysis showed that both the extent of ISF and the number genuine, non-ISF 528 \npeptides were underestimated depending on the peak selection. Underestimating ISF might not 529 \npresent a problem for typical applications but, whenever an in-source fragment is identified, it 530 \ncan come at the cost of losing an identification of a genuine peptide that shares its sequence with 531 \nthe in-source fragment. With current peptide search strategies for DIA data, ISF can thus mask 532 \nthe identification of genuine peptides and lead to loss of information. Furthermore, for peptides 533 \nwith high signals, there is a reasonable chance that an in-source fragment can be observed in the 534 \ndata. This effect could be used in future peptide search engines to improve peptide identification 535 \nand estimation of false-discovery rates while increasing overall annotations of proteomics data. 536 \nIn immunopeptidomics, peptides are not generated by in vitro digestion and are often identified 537 \nby unspecific peptide searches. Therefore, immunopeptidomics data could be much more prone 538 \nto interference by ISF, which often generates peptides with non-tryptic termini. Our results show 539 \nthat in immunopeptidomics data with HLA-I peptides, most in-source fragments are shorter than 540 \n9 amino acids and can account for more than a third of all peptides of such short length. This 541 \nraises the question of how many of the immunopeptides shorter than 9 amino acids are truly 542 \nauthentic peptides because current approaches still underestimate how many of the detected 543 \npeptides are in-source fragments. We recommend using our ISF detection to minimize the risk of 544 \nfalsely identifying peptides. Alternatively, or in addition, immunopeptides smaller than 9 amino 545 \nacids can be filtered out to strongly reduce the impact of ISF on the data. 546 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n20 \n \nAs in-source fragments have the same RT as their parent peptides, the RT of the in-source 547 \nfragment does not correctly reflect the peptide separation by chromatography and is different 548 \nfrom the RT of a non-ISF peptide with the same amino acid sequence. Peptide ISF can therefore 549 \nhave an impact on the RT realignment during peptide searches. Depending on the MS approach, 550 \nthe strength of this effect is likely to vary, and will be more impactful in cases where a few 551 \nreference peptides are chosen for realignment, like analyses by single reaction monitoring (SRM), 552 \nor when the prevalence of ISF is particularly high. 553 \nWe observed little PTM ISF for phosphorylation, oxidation of methionine, and N-terminal 554 \nacetylation, but we cannot exclude that other PTMs such as glycosylation, which is known to be 555 \nprone to ISF 67, might be affected to a much larger extent. Some PTMs might even be so labile 556 \nthat we only detect the in-source fragments, which would require exceptionally mild ionization 557 \nand ion routing approaches to study them. Sequence-altering ISF of parent peptides that carry a 558 \nPTM could also explain some of the detected semi-tryptic peptides in typical trypsin-digested 559 \nbottom-up proteomics samples, which have not been associated with an ISF parent so far, 560 \nbecause many peptides with PTMs are very challenging to detect in non-enriched datasets by 561 \ncurrent search engines. Other cases of ISF that have not been addressed so far are those in which 562 \nfragmentation occurs not at peptide bonds (between C1 and N2) or at PTM-peptide bonds but 563 \nrather at side chains. Detecting these cases presents a major challenge but potentially could 564 \nexplain many of the detected but uncharacterized m/z features in proteomics data and help 565 \nseparate these from more biologically meaningful uncharacterized portions of the proteome 68. 566 \nMoreover, in-source chemical reactions that add to the mass of a molecule have been observed 567 \nfor metabolites before at a high prevalence 13 and, in principal, could also affect peptides. Taken 568 \ntogether, this poses the question: how many of the m/z features in a proteomics dataset that are 569 \nunannotated can be explained by ISF or in-source reactions? Answering this question will result 570 \nin new peptide groups without annotation, for which the mode of fragmentation or reaction 571 \ncould provide insights into the identity of the parent peptides. 572 \nIn summary, in-source fragments inflate the number of identifications and, especially in low 573 \ncomplexity samples, can make up most of the signal and lead to data misinterpretation. Further, 574 \nsince ISF varies substantially dependent on numerous parameters, predicting the impact of these 575 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n21 \n \nfragments a priori in a given analysis is challenging. Based on our findings in this study, we thus 576 \nconclude that ISF should be identified by default in current proteomics approaches and the in-577 \nsource fragments filtered out bioinformatically. 578 \n  579 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n22 \n \nFigures 580 \n 581 \nFig. 1. Illustration of the ISF detection algorithm 582 \na, Graph showing hypothetical chromatograms of different peptide precursors that share amino acid sequences and 583 \nthus form peptide pairs. The axes display the retention time, the m/z, and the signal intensity as indicated. The 584 \npeptides form three coeluting groups. 585 \nb, Heatmap illustrating all pairwise combinations of peptides from the example in A and their respective ∆RTs. The 586 \n∆RT is calculated by subtracting the apex RT of the shorter peptide from that of the longer peptide. Peptide pairs 587 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n23 \n \nthat do not show co-elution and have ∆RTs larger than a cutoff are blue. Peptide pairs with evidence of ISF are red. 588 \nPairs of peptides that only differ in charge are yellow. 589 \nc, Diagram illustrating a theoretical ∆RT distribution of peptides pairs that only differ in charge. An observed ∆RT 590 \ndistribution (blue line) is the sum of two sub-distributions: a normal distribution of true-positive peptide pairs (blue 591 \nshade) centered around zero and a uniform distribution of false-positive peptide pairs (red shade), of which at least 592 \none peptide per pair is falsely annotated. The ∆RT cutoff is the ∆RT at which the value of the normal distribution 593 \nequals the uniform distribution. 594 \nd, Graph showing a theoretical ∆RT distribution of all peptide pairs that are not part of C. The ∆RT cutoff is used to 595 \nclassify peptide pairs into those with ISF evidence, if they fall within the cutoff, and those without ISF evidence. 596 \ne, Illustration of peptide annotation: peptide pairs with ISF evidence are deconvoluted, and peptide precursors 597 \nassigned as in-source fragments or parent peptides. Peptides that show ISF often form complex fragmentation 598 \nnetworks, in which edges indicate ISF evidence (labelled with the respective ∆RT) and nodes are peptides. 599 \nFragmentation networks can be used to create new peptide groups. Amino acid sequence-altering ISF can occur N-600 \nterminally, C-terminally, or at both ends (omitted in the illustration). ISF of exclusively PTM groups are examples of 601 \namino acid sequence-conserving ISF. 602 \n  603 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n24 \n \n 604 \nFig. 2. Evidence for abundant ISF across proteomics data of different complexity  605 \na, Dot plot showing the share of in-source fragments (% by intensity) per number of detected protein groups 606 \n(minimum of 5 unique peptide sequences per group) in 317 samples across 16 published proteomics datasets (as 607 \nindicated by color and symbol shape). 608 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n25 \n \nb, Dot plot showing the share of in-source fragments (dots: % by intensity, squares: % by identifications) per number 609 \nof identified protein groups (minimum of 5 unique peptide sequences per group) across 7 datasets (as indicated by 610 \ncolors) that ranged from a mix of 11 peptides up to a mix of three whole proteomes. Dots and squares are mean 611 \nvalues (n = 4), vertical lines indicate standard deviations. 612 \nc, Bar graphs showing the number of identifications and the absolute intensity of in-source fragments, ISF parents, 613 \nand peptides without ISF evidence for the datasets shown in B. Bars indicate mean values (n = 4), vertical lines 614 \nstandard deviations. 615 \nd, Bar plot showing the percentage of total intensity covered by in-source fragments and parents combined. 616 \ne, Pie chart showing the share (%) of semi- and fully-tryptic peptides among in-source fragments across all the 617 \ndatasets from A. 618 \nf, Bar plot depicting the average share of semi-tryptic in-source fragments (%) among all detected semi-tryptic 619 \npeptides per dataset from A. Blue bars are intensity-based representations (yellow bars: number of identifications-620 \nbased). Bars indicate mean values (n = 4), horizontal lines the standard deviation. 621 \n  622 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n26 \n \n 623 \nFig. 3. ISF peptide characterization 624 \na, Three pie charts illustrating: (1) the fractions of peptides with and without peptide pairing (in %), (2) the number 625 \nof peptide pairs with and without evidence of ISF (%), and (3) the number of peptide pairs with ISF evidence that 626 \nhave a sequence-altering ISF, are fragment-to-fragment pairs, and feature PTM ISF (%). The data used for this figure 627 \nstems from 23 datasets (also see Fig. 2) that were merged. 628 \nb, Graph showing the z-scored (modified) log 2-intensity distributions (in %) of in-source fragment (red) and parent 629 \n(yellow) peptides as well as peptides without ISF evidence (blue). Box whisker plots above the distributions indicate 630 \nthe median values and interquartile ranges. Modified z-scores are calculated with median values and median 631 \nabsolute deviations. 632 \nc, Dot plot showing the abundance of amino acids at the F1’ position at ISF sites over occurrences at F1. 633 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n27 \n \nd, Heatmap showing the adjusted P-values from an enrichment analysis (Fisher’s exact test, two-tailed, P-value 634 \nadjustment by Benjamini-Hochberg procedure 43) of all peptide bonds observed at IFS sites. Blue squares indicate a 635 \nnegative odds ratio, yellow a positive ratio. 636 \ne, Histograms showing the number of ISF networks (in %) with the indicated number of unique peptides per network 637 \n(left) or with the indicated in-source fragment to parent ratio (right). 638 \n  639 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n28 \n \n 640 \nFig. 4. PTM in-source fragmentation 641 \na, Bar plot showing the number of peptide pairs with a PTM ISF (as labelled) among all pairs with ISF evidence from 642 \n23 combined datasets (also see Fig. 3.). 643 \nb, Dot plot showing the number of ISF of methionine oxidation (%) per number of precursors with a single 644 \nmethionine oxidation (%). Colors and symbols indicate 23 different datasets. 645 \nc, Bar plot showing the percentages of phosphopeptides (red) and other peptides (blue) in four different datasets. 646 \nThe two S. cerevisiae  datasets were acquired from the same sample source: one as a standard yeast proteomics 647 \ndataset and the other as phosphoproteomics dataset (including phosphopeptide enrichment). The two literature 648 \ndatasets (PXD014525 and PXD044482) relied on phosphopeptide enrichment. 649 \nd, Bar graph displaying the number of in-source fragments (% by identifications) in the four datasets of C. 650 \n  651 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n29 \n \n 652 \nFig. 5. Impact of the funnel RF, electrospray voltage, and ITC temperature on peptide ISF.  653 \na, Bar plot displaying the number of detected peptide precursors that are non-ISF and parent peptides (blue bars) 654 \nand that are in-source fragments (red bars) (% of all identifications across all samples) in a yeast reference sample 655 \nacquired at different funnel radiofrequencies (%), electrospray voltages (V), and ITC temperatures (°C) (n = 1, except 656 \n5 pairs of technical replicates as indicated). The samples are sorted from left to right in ascending order of increasing 657 \nISF. 658 \nb, Dot plots showing the relative intensity (%) (upper left ), the total intensity (lower left), the relative number of 659 \nidentifications (%) (upper right ), and absolute number of identifications (lower right ) of in-source fragments in the 660 \nsamples from A. 661 \n  662 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n30 \n \n 663 \nFig. 6. Impact of ISF on quantification  664 \na, Whole proteome digests of S. cerevisiae S288c and E. coli BW25113 were mixed at different ratios as indicated as 665 \ndilution factors below the blue bars. The yeast proteome concentration was kept constant. The  E. coli  proteome 666 \nconcentration was varied, and volumes compensated with water.  667 \nb, Dot plots showing the mean log 2-precursor fold changes (relative to the sample with the undiluted E. coli  668 \nproteome) per mean log2-precursor intensity (n = 3 technical replicates). Red dots indicate in-source fragments, blue 669 \ndots ISF parents, and grey dots non-ISF peptides. 670 \nc, Box whisker plot showing the median coefficient of variation (CV, %, n = 3 technical replicates) and the median 671 \nrelative error between theoretical and measured fold changes (n = 3 technical replicates) of all peptide precursors 672 \nfrom the samples in B. Data of in-source fragments is indicated in red (parents: blue, non-ISF precursors: grey). Boxes 673 \ndepict interquartile ranges, and whiskers extend to 1.5-fold of the interquartile ranges. 674 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n31 \n \nd, Box whisker plot showing the median coefficient of variation (CV, %, n = 3 technical replicates) and the median 675 \nrelative error between theoretical and measured fold changes (n = 3 technical replicates) of all peptide groups from 676 \nthe samples in B. Peptide groups of all ISF parents without (blue) and with (yellow) merging with their respective in-677 \nsource fragments are compared with all peptide groups without ISF evidence (green). Boxes depict interquartile 678 \nranges, and whiskers extend to 1.5-fold of the interquartile ranges. P-values were calculated using the two-tailed 679 \nWilcoxon rank sum test. Peptide quantities were calculated by summing up precursor intensities. 680 \ne, Same as D but depicting protein groups that contain (yellow) or omit (green) in-source fragments. Protein 681 \nquantities were calculated by summing up precursor intensities. 682 \n  683 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n32 \n \n 684 \nFig. 7. ISF in immunopeptidomics and limited proteolysis datasets  685 \na, Bar plot showing the average percentage of in-source fragments in immunopeptidomics datasets (PXD058880: n 686 \n= 12, PXD054417: n = 12, PXD022950: n = 6, PXD051490: n = 12), by peptide intensities and numbers of 687 \nidentifications. 688 \nb, Bar plot showing the number of precursor pairs with evidence of ISF (per dataset in A) (%) that feature N-terminal 689 \nISF, C-terminal ISF, or ISF at both termini. 690 \nc, Dot plots displaying the number of peptides (%, normalized per group) of different lengths (number of amino 691 \nacids) in each dataset of A. Non-ISF peptides are indicated in blue (ISF parents: yellow, in-source fragments: red). 692 \nd, Bar graph showing the mean fraction of in-source fragments within each indicated LiP dataset (yellow bars: % by 693 \nintensity, blue bars: % by the number of identifications) (n = 3). Data results from trypsin unspecific peptide searches. 694 \ne, Bar plot showing the percentage of in-source fragments among semi-tryptic peptides (green bars: % by intensity, 695 \nviolet bars: % by identifications) in the indicated LiP datasets. 696 \nf, Bar plot showing the percentage of in-source fragments by their apparent trypticity in the indicated LiP datasets. 697 \nRed bars indicate in-source fragments that appear as non-tryptic peptides (yellow bars: semi-tryptic in-source 698 \nfragments, blue bars: fully-tryptic in-source fragments). 699 \n  700 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n33 \n \nMethods 701 \nM.1 Sample preparation for MS-based proteomics 702 \nM.1.1 Saccharomyces cerevisiae S288c whole-proteome sample 703 \nS. cerevisiae S288c (ATCC #204508) was streaked from cryo stock onto an YPD agar plate and 704 \nincubated at room temperature. After single colonies became visible, the plate was stored at 4 705 \n°C. 25 mL synthetic defined (SD) medium containing 20 g/L D-glucose (Merck, Sigma-Aldrich 706 \n#G7021), 5 g/L ammonium sulfate (Fluka #09982), and yeast base (Merck, Millipore #Y1251) was 707 \ninoculated from a single colony and incubate in 100 mL Erlenmeyer flasks at 30 °C under shaking 708 \nof 200 rpm for circa 20 h. 40 mL SD medium (2 w/v-% glucose) in 500 mL Erlenmeyer flasks was 709 \ninoculated from the previous 25 mL-culture to a start optical density (OD at 600 nm) of 0.004 and 710 \nincubated at 30 °C under shaking at 200 rpm. 300 mL SD medium (2 w/v-% glucose) in 2 L 711 \nErlenmeyer flasks were inoculated from the previous 40 mL-culture to a start OD of 0.08 and 712 \nincubated at 30 °C under shaking at 200 rpm. The OD was measured regularly to calculate growth 713 \nrates and exclude growth defects. At a final OD of circa 0.6, samples for proteomics were 714 \ncollected using a method that is analogous to a phosphoproteomics sampling approach 45 and 715 \nrelies on quenching enzyme activity by trichloroacetic acid (TCA): 275 mL of culture were 716 \ntransferred to 500 mL harvesting tubes on ice. 18.3 mL 4 °C-cold 100 w/v-% TCA (final 717 \nconcentration 6.2 % TCA; Merck, Sigma-Aldrich, #91228) was added. The cell suspension was 718 \nincubated for 10 min on ice and, subsequently, centrifuged for 5 min at 4 °C and 3428 g. The 719 \nsupernatant was discarded, and the cell pellet resuspended in 40 mL 4 °C-cold acetone (Merck, 720 \nSupelco, #1.00014). The suspension was transferred to 50 mL reaction tubes and centrifuged for 721 \n5 min at 4 °C and 3,428 g. The supernatant was discarded, the cell pellet resuspended in 40 mL 4 722 \n°C-cold acetone, and the suspension centrifuged for 5 min at 4 °C and 3,428 g. The supernatant 723 \nwas discarded, and the cell pellet resuspended in 5 mL 4 °C-cold acetone. The cell suspensions of 724 \n4 independent biological replicates (= 4 different colonies) were pooled and distributed to 1.5 mL 725 \nscrew cap-reaction tubes. After 5 min of centrifugation at 3,400 g and 4 °C, the supernatant was 726 \ndiscarded, and the remaining cell pellets flash frozen in liquid nitrogen and stored at -80 °C. Cell 727 \npellets were resuspended in 500 µL lysis buffer, which was 8 M urea (Merck, Empure Essential, 728 \n#1.08486.1000) 100 mM ammonium bicarbonate (Merck, Sigma-Aldrich, #A6141) at pH 7.8. After 729 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n34 \n \nadding an equal volume of glass beads (0.5 mm diameter, Merck, Sigma-Aldrich # G8772), cells 730 \nwere lysed at 4 °C using 8 cycles of bead beating at 6.5 m/s for 30 s with 200 s pause. After bead 731 \nremoval, lysates were centrifuged for 15 min at 17,000 g and 4 °C, and all supernatants pooled 732 \nfor further processing. After tryptic digest and peptide clean-up, the reference yeast sample was 733 \nobtained by mixing 800 µL of the S. cerevisiae S288c sample with 200 µL of the mix of 8 proteins 734 \nsample. 735 \n 736 \nM.1.2 Escherichia coli BW25113 whole-proteome sample 737 \nE. coli BW25113 (DSMZ, #27469) was streaked from cryo stock onto a LB agar plate and incubated 738 \nat 37 °C overnight. The plate was stored at 4 °C. 25 mL M9 minimal medium was inoculated from 739 \na single colony and incubate in 100 mL Erlenmeyer flasks at 37 °C under shaking of 200 rpm 740 \novernight. M9 minimal medium contained: 5 g/L D-glucose (Merck, Sigma-Aldrich, #G7021), 6 g/L 741 \nNa2HPO4 (Carl Roth, #P030.2), 3 g/L KH 2PO4 (Carl Roth, #3904.1), 0.5 g/L NaCl (Carl Roth, 742 \n#9265.1), 1.5 g/L (NH 4)2SO4 (Merck, Sigma-Aldrich, #A3920), 1.8 mg/L ZnSO 4x 7 H 2O (Fluka, 743 \n#96500), 1.2 mg/L CuCl2 x 2 H2O (Fluka, #61174), 1.2 mg/mL MnSO4 x H2O (Merck, Sigma-Aldrich, 744 \n#M7631), 1.8 mg/L CoCl2 x 6 H2O (Riedel de Haën, #12914), 2.8 mM thiamine-HCl (Merck, Sigma-745 \nAldrich, #T4625), 1mM MgSO 4 x 7 H 2O (Merck, Sigma-Aldrich, #63138), 100 µM CaCl 2 x 2 H 2O 746 \n(Merck, Sigma-Aldrich, #C8106), and 100 µM FeCl 3 x 6 H2O (Merck, Sigma-Aldrich, #31232). 300 747 \nmL M9 minimal medium (0.5 w/v-% glucose) in 2 L Erlenmeyer flasks were inoculated from the 748 \nprevious 25 mL-culture to a start OD of circa 0.004 and incubated at 37 °C under shaking at 200 749 \nrpm. The OD was measured regularly to calculate growth rates and exclude growth defects. At a 750 \nfinal OD of circa 0.6, samples for proteomics were collected: 275 mL of culture were transferred 751 \nto 500 mL harvesting tubes on ice. 18.3 mL 4 °C-cold 100 w/v-% TCA was added. The cell 752 \nsuspension was incubated for 10 min on ice and, subsequently, centrifuged for 5 min at 4 °C and 753 \n3,428 g. The supernatant was discarded, and the cell pellet resuspended in 40 mL 4 °C-cold 754 \nacetone. The suspension was transferred to 50 mL reaction tubes and centrifuged for 5 min at 4 755 \n°C and 3,428 g. The supernatant was discarded, the cell pellet resuspended in 40 mL 4 °C-cold 756 \nacetone, and the suspension centrifuged for 5 min at 4 °C and 3,428 g. The supernatant was 757 \ndiscarded, and the cell pellet resuspended in 5 mL 4 °C-cold acetone (Merck, Supelco, #1.00014). 758 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n35 \n \nThe cell suspensions of 4 independent biological replicates (= 4 different colonies) were pooled 759 \nand distributed to 1.5 mL screw cap-reaction tubes. After 5 min of centrifugation at 3,400 g and 760 \n4 °C, the supernatant was discarded and remaining cell pellets flash frozen in liquid nitrogen. 761 \nSamples were stored at -80 °C. Cells were lysed like S. cerevisiae S288c samples. 762 \n 763 \nM.1.3 Homo sapiens HEK-293 whole-proteome sample 764 \nH. sapiens HEK-293 cells (ATCC, #CRL-1573) were thawed from cryo stock and transferred to a 765 \ncell culture-grade, sterile, 10 cm-diameter petri dish. 10 mL Dulbecco’s Modified Eagle Medium 766 \n(DMEM, Thermo Fisher Scientific, Gibco #41965039) containing 10 v/v-% heat inactivated foetal 767 \ncalf serum (FCS, BioConcept, #2-01F16-I) and 1 v/v-% penicillin-streptomycin solution (P/S, 768 \n10,000 U/mL, Thermo Fisher Scientific, Gibco #15140122) were added, and cells incubated for 3 769 \ndays at 37 °C, 5 %CO 2, and a relative humidity of 85 %. The supernatant was removed, 10 mL 770 \nphosphate buffered saline solution (PBS, pH 7.4, Thermo Fisher Scientific, Gibco #10010015) 771 \nadded to wash surface adherent cells. The supernatant was removed and 1 mL 0.25 % Trypsin-772 \nEDTA solution (Thermo Fisher Scientific, Gibco #25200056) added. 24 mL of fresh medium were 773 \nadded to a 15 cm-diameter petri dish, trypsin-suspended cells transferred to the plate with fresh 774 \nmedium, and cells incubated for 4 days. The cell culture was tested negative for mycoplasma 775 \ncontamination by the MycoGenie Rapid Mycoplasma Detection Kit (AssayGenie, #MORV0011-776 \n50). The supernatant was removed and cells washed with 25 mL PBS. 3 mL 0.25 % Trypsin-EDTA 777 \nsolution added, and 5 new plates each with 23 mL fresh medium prepared. 8 mL medium were 778 \nadded to the trypsinated culture, each 2 mL of the cell suspension added to a new plate, and cells 779 \nincubated for 2 days. Proteomics samples were taken by transferring the plates onto ice, remove 780 \nthe supernatant, and, for each plate, adding 10 mL 4 °C-cold PBS containing 1 mM EDTA (Merck, 781 \nSigma-Aldrich, #03677). The cell suspension was transferred to 15 mL reaction tubes and 782 \ncentrifuged for 4 min at 300 g and 4 °C. The supernatant was removed, and cells resuspended in 783 \n1 mL of the 4 °C-cold PBS-EDTA solution. The cell suspensions from 4 plates were pooled, 784 \ndistributed to 2 mL reaction tubes, and centrifuged for 4 min at 300 g and 4 °C. The supernatant 785 \nwas discarded, and cell pellets flash frozen in liquid nitrogen and stored at -80 °C. Cell pellets 786 \nwere resuspended in 500 µL lysis buffer and lysed by vortexing twice for 30 s with a pause of at 787 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n36 \n \nleast 30 s on ice. Lysates were centrifuged for 15 min at 17.000 g and 4 °C. Supernatants were 788 \nused for further processing. 789 \n 790 \nM.1.4 Mix of 8 proteins 791 \nSeparate solutions of bovine catalase (Uniprot ID: P00432, Merck, Sigma-Aldrich, #C40), rabbit 792 \ncreatine kinase (P00563, Merck, Sigma-Aldrich, #C3755), rabbit fructose-bisphosphate aldolase 793 \nA (P00883, Merck, Sigma-Aldrich, #A2714), bovine lactoferrin (P24627, Merck, Sigma-Aldrich, 794 \n#L9507), chicken ovotransferrin (P02789, Fluka, #27695), rabbit pyruvat kinase (P11974, Merck, 795 \nSigma-Aldrich, #P9136), bovine serotransferrin (Q29443, Merck, Sigma-Aldrich, #T1408), and 796 \nbovine serum albumin (P02769, Merck, Sigma-Aldrich, #A7638) were prepared at 1 µg protein/µL 797 \nin 8 M urea 100 mM ammonium bicarbonate. After tryptic digestion and sample clean-up, 798 \npeptides from the individual proteins were pooled together yielding the mix of 8 proteins sample. 799 \n 800 \nM.1.4 Tryptic digest and sample clean-up 801 \nProtein concentrations in cell lysates were measured using a  bicinchoninic acid-based assay 802 \n(Thermo Scientific, Pierce BCA Protein Assay Kit, #23225). At a protein concentration of 1 µg/µL, 803 \ndisulfide bonds were first reduced by incubation with 5 mM tris(2-carboxyethyl)phosphin -804 \nhydrochlorid (TCEP, Merck, Sigma-Aldrich, #C4706) for 30 min at 37 °C under 200 rpm of shaking 805 \nand second alkylated by incubation with 12 mM iodoacetamide (IAA, Merck, Sigma-Aldrich, 806 \n#I1149) for 15 min in the dark at 37 °C under 200 rpm of shaking. Samples were diluted with 100 807 \nmM ammonium bicarbonate (Merck, Sigma-Aldrich, #A6141) to a final urea concentration of 1 808 \nM. Sequencing-grade trypsin (Promega, #V5111) was added at a protease to protein ratio of 809 \n1:100, and samples incubated overnight at 37 °C. The digestion was stopped with formic acid (FA, 810 \nMerck, Sigma-Aldrich, #33015) at a final concentration of 2 %. Samples were desalted using C18 811 \ncolumns (Waters Sep-Pak Vac 3cc, 500 ng). After washing the columns with 2.5 mL methanol, 812 \ntwice with 2.5 mL Buffer B (50 % 0.1 % FA), and thrice with 2.5 mL Buffer A (0.1 % FA), samples 813 \nwere loaded onto the columns, and columns washed thrice with 2.5 mL Buffer A. Peptides were 814 \neluted with 2 mL Buffer B, dried by vacuum centrifugation at 40 °C, and resuspended in Buffer A 815 \nat a concentration of ca. 1 µg/µL. 816 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n37 \n \n 817 \nM.1.5 Affinity purification samples 818 \nA H. sapiens HEK-293 cell line expressing the Strep-HA-tagged human kinase SRPK3 (Uniprot ID: 819 \nQ9UPE1) under an inducible promoter was obtained from Varjosalo et al.,69 and cells seeded at 820 \na density of 10 7 in 25 mL DMEM supplemented with 10 % FCS and 1 % P/S in a 15 cm dish and 821 \nincubated overnight at 37 °C. Protein expression was induced with 1.3 μg/mL doxycycline for 24 822 \nhours. Per sample, cells from two 15 cm dishes were harvested by scraping off in 4 °C-cold PBS. 823 \nCells were washed once with PBS, and cell pellets flash frozen in liquid nitrogen. Pellets were 824 \nlysed by resolubilizing in 2 mL HNN buffer (50 mM HEPES pH 8.0, 150 mM NaCl, 50 mM NaF) 825 \nsupplemented with 0.5 % IGEPAL, 400 nM Na 3VO4, 1 mM phenylmethylsulfonyl fluoride, 0.2 % 826 \nprotease inhibitor cocktail (Merck, Sigma-Adrich, #P8340) and 0.5 μL/mL benzonase (Merck, 827 \nSigma-Aldrich, #1.01695). The lysate was incubated on ice for 10 min and centrifuged at 18,000 828 \ng for 20 min at 4 ˚C. The supernatant was added to 80 μL StrepTatcin Sepharose resin (IBA 829 \nLifesciences, #2-1201) and incubated on a rotary shaker for 60 min at 4 ˚C. The beads were loaded 830 \nonto a 1 μm glass filter column and washed 3 times with 1 mL HNN buffer supplemented with 831 \n0.5 % IGEPAL and 400 nM Na 3VO4, 3 times with 1 mL HNN buffer, and 3 times with 100 mM 832 \nammonium bicarbonate. Samples were eluted twice by incubation with 30 μL of 0.5 mM biotin 833 \nin 100 mM ammonium bicarbonate for 15 min. To the samples, 60 μL 8 M urea in 100 mM 834 \nammonium bicarbonate for a final concentration of 4 M urea. Samples were reduced with 5 mM 835 \nTCEP for 40 min at 37 °C and 200 rpm and alkylated with 40 mM IAA for 30 min at 30 °C in the 836 \ndark, at 200 rpm. Samples were diluted with 100 mM ammonium bicarbonate to a urea 837 \nconcentration of 1 M. To each sample, 1 μg Lys-C and 1 μg trypsin were added, and samples were 838 \nincubated overnight at 37 °C and 200 rpm. The digestion was stopped with 2 % formic acid. 839 \nSamples were desalted using HNFR S18V desalting plates (Nest Group). The resin was activated 840 \nusing 200 μL methanol and washed 3 times with 200 μL Buffer B. The resin was equilibrated 3 841 \ntimes with 200 μL Buffer A. Samples were loaded and washed 3 times with 200 μL Buffer A. 842 \nSamples were eluted twice with 100 μL Buffer B and dried at 40 °C in a vacuum centrifuge. 843 \n 844 \nM.1.6 Mix of 11 peptides sample 845 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n38 \n \nWe used a commercially available mix of 11 peptides (Biognosys iRT Kit, Bruker Daltonics, 846 \n#1816351). 847 \n 848 \nM.1.7 Phosphoproteomics samples 849 \nPhosphopeptides of S. cerevisiae S288c samples were enriched based on a protocol described in 850 \nBodenmiller et al. 45. Between 0.5 and 1 mg of desalted peptides were incubated in a rotary 851 \nshaker for 1 h with 1.25 mg of TiO 2 resin (GL Sciences, Japan), preequilibrated twice with 500 µL 852 \nof methanol, and twice with 500 µL of a solution saturated with phthalic acid (Merck, Sigma-853 \nAldrich, #402915). Peptides bound to the TiO2 resin were washed twice with 500 µL phthalic acid 854 \nsolution, twice with 80 % acetonitrile 0.1% formic acid solution, and finally twice with 0.1% formic 855 \nacid. The phosphopeptides were eluted from the beads twice with 150 µl of 0.3 M ammonium 856 \nhydroxide at pH 10.5 and immediately acidified with 50 µL 5% trifluoroacetic acid to reach about 857 \npH 2.0. The enriched phosphopeptides were desalted on microspin columns (The Nest Group, 858 \nUSA), dried using a vacuum centrifuge, and resolubilized in 10 µl of 0.1 % formic acid. 859 \n 860 \nM.2 LC-MS 861 \nIf not stated otherwise, an EASY-nLC 1200 (Thermo Scientific) coupled to an Exploris 480 (Thermo 862 \nScientific) was used to measure proteomics samples in DIA mode. A sample volume 863 \ncorresponding to 1 µg of peptides was injected. A 40 cm x 0.75 µm (inner diameter) column (New 864 \nObjective, 10 µm tip, PicoFrit, #PF360-75-10-N-5) packed with 3 μm C18 beads (Dr. Maisch, 120 865 \nÅ pore size, 300 m² surface area, #Reprosil-Pur 120) was used for separating peptides over a 866 \nlinear 120 min gradient. Buffer A was 0.1 % FA, and Buffer B 50 % ACN 0.1 % FA. The gradient 867 \nstarted at 3 % Buffer B and ended at 30 % Buffer B. The flow rate was 300 nL/min. Peptides were 868 \nmeasured in positive ionization mode. The electrospray voltage was 2500 V, the ion transfer 869 \ncapillary temperature was 275 °C, and the funnel radiofrequency 50 % if not indicated else. To 870 \nreduce contamination of the MS, the electrospray voltage was set to 1500 V during the initial 8 871 \nmin of the method and to 0 V during column washes. The MS1 full scan range was from 350 to 872 \n1150 m/z at a resolution of 120,000. with an automatic gain control (AGC) target of 200 % and a 873 \nmaximum injection time of 264 ms. Precursors were fragmented in the higher-energy collisional 874 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n39 \n \ndissociation (HCD) cell at 30 % relative collision energy. For DIA, 41 variable-size MS2 scan 875 \nwindows were used between 350 and 1,150 m/z, each at a resolution of 30,000, an AGC target 876 \nof 200 %, a maximum injection time of 66 ms, and a window overlap of 1 Da. 877 \nFor measurements with an Q Exactive Plus (Thermo Scientific), which was also coupled to an 878 \nEASY-nLC 1000 (Thermo Scientific), the same parameters were used as for the Exploris 480, if 879 \npossible. MS1 full scans differed by the resolution (70,000), the AGC target (3e6), and the 880 \nmaximum injection time (120 ms). DIA MS2 measurements differed by the resolution (35,000), 881 \nAGC target (3e6), and the maximum injection time (60 ms). 882 \nFor measurements with an Astral (Thermo Scientific) coupled to a Vanquish Neo LC system 883 \n(Thermo Scientific), the flow rate was 400 nL/min, and a 41.8 min gradient was used. Starting 884 \nwith Buffer B at 3 %, buffer B was changed in linear steps to 32 % at 25.5 min, 45 % at 30.5 min, 885 \n95 % at 32.5 min, where it was kept constant for 8 min, and 3 % at 41.8 min. MS1 full scans were 886 \nacquired between 350 and 1400 m/z with a resolution of 240,000. The AGC target was 5e6. For 887 \nDIA, 524 fixed-size (2 Da) MS2 scan windows between 350 and 1400 m/z were used with a 888 \nmaximum injection time of 3 ms, a relative HCD collision energy of 27 %, the RF lens at 40 %, and 889 \nan AGC target of 5e4. 890 \n 891 \nM.3 Peptide searches and data processing 892 \nSpectronaut 20 (Biognosys, Schlieren, Switzerland) was used to search DIA data for peptides 893 \nusing “directDIA”. The data was searched against different proteome databases appropriate to 894 \neach sample and related FASTA-files were obtained from UniProt 70. The enzyme specificity was 895 \nset to semi-specific with cleavage rules for trypsin (or if applicable also for LysC). Peptide lengths 896 \nbetween 7 and 52 amino acids, two missed cleavages, and a maximum of 5 variable modifications 897 \nwere permitted. The immunopeptidomics and LiP datasets were searched for peptides with 898 \nunspecific cleavage sites and 5 to 16 amino acids in length. N-Acetylation at protein N-termini 899 \nand oxidation of methionine were set as variable modifications, and carbamidomethylation of 900 \ncysteine as a fixed modification. Searches for phosphopeptides included phosphorylation of 901 \nserine, threonine, and tyrosine. Data from the parameter test (also see Fig. 4) was searched 902 \ntwice: (1) all samples were searched individually to obtain an accurate representation of the 903 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n40 \n \nnumber of peptide precursor identifications, and (2) all samples were searched together to 904 \nobtain accurate quantitative data (intensities). Supplementary Table 2 provides an overview of 905 \nwhich samples of the literature datasets (also see Supplementary Tab. 1) were used in this study 906 \nand included in peptide searches. 907 \nData from peptide searches were analyzed with MATLAB (R2023b, 23.2.0.2409890). Venn 908 \ndiagrams were created using the MATLAB function “venn” 71. The data was analyzed at the 909 \npeptide precursor level if not specified else. Precursors with identification Q-values greater than 910 \n10-3 and precursors without quantities or “FGMS2RawQuantity” smaller than 2 were filtered out, 911 \nexcept for the algorithm comparison (also see Supplementary Fig. 3) for which no filter was 912 \napplied. 913 \n 914 \nM.4 ISF detection 915 \nISF were detected in each sample separately. First, all peptide precursors were paired up against 916 \neach other, for which one precursor’s amino acid sequence contained the sequence from the 917 \nother precursor. Then, looping over all precursors and their pairings from small to large (length), 918 \n∆RTs were calculated by subtracting the peak apex retention times of the smaller peptide from 919 \nthe larger peptide. Subsequently, ∆RT cutoffs were determined by fitting the model function 920 \n(EQ1) to the ∆RT distribution of precursor pairs that only differ in charge. The function consisted 921 \nof two components. One of which was a normal distribution with mean value µ and standard 922 \ndeviation σ that was scaled with parameter a. The other component was the uniform distribution 923 \nk0. ∆RT cutoff was the ∆RT, for which the absolute difference between the normal distribution 924 \nand k0 was smallest. 925 \n𝑦(𝑥) =  𝑎 ∙\n௘\nష (౮  షഋ )మ\nమ഑ మ\nఙ √ଶగ + 𝑘଴   (EQ1) 926 \nThe distribution of measured ∆RTs was estimated ten times with different bin sizes and, 927 \naccordingly, the model function fitted to each estimation using least square optimization. The 928 \nfinal ∆RT cutoff was the median across the ten ∆RT cutoffs from the ten individual fits. If the 929 \nnumber of datapoints of false-positive precursor pairs that only differ in charge was insufficient 930 \nto model the uniform distribution, the ∆RT cutoffs could not be accurately determined and 931 \ntypically became very large. In such cases, in which our model-based approach yields a ∆RT cutoff 932 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n41 \n \nlarger than 0.15 min, the full width at half maximum (FWHM) of the normal distribution was 933 \nused. In extreme cases, in which the data was even insufficient to model the normal distribution 934 \nof true positives, a ∆RT cutoff of 0.1 min was taken. The lower boundary of the ∆RT cutoff was 935 \nthe smallest ∆RT value of the respective sample. 936 \n 937 \nM.5 Multiple peak extraction from DIA data 938 \nTo analyze peptide precursors with multiple peaks in DIA data, we followed a previously 939 \npublished strategy 62. If the same peptide sequence is identified in DDA beyond a certain time 940 \nspan, which is typically longer than the average elution peak width (e.g. 2 min, which we used 941 \nhere), the peptide-spectrum matches were binned into separate spectral libraries. Subsequently, 942 \nfor each of the peptides, a local average retention time and consensus MS/MS fragmentation 943 \nspectrum were calculated. To force the DIA analysis software to extract multiple identical peptide 944 \nsequences at different retention times, we “anonymized” the peptides with a set of characters 945 \nthat cannot be properly parsed by the software (e.g. adding numbers or special characters). 946 \nWhen faced with such non-parsable sequences, the search tool defers back to the only other 947 \ninformation available to extract the DIA data: (1) the m/z and charge state of the precursor, (2) 948 \nthe m/z, charge states and relative intensities of the fragments, and (3) the peptide retention 949 \ntime. This strategy enabled us to estimate how often a peptide sequence could be identified 950 \nmultiple times in DIA file. 951 \n 952 \nData and code availability 953 \nProteomics data will be made available on PRIDE and relevant MATLAB code will be provided on 954 \nGitHub and Zenodo upon publication. The source data of figures will be provided as 955 \nsupplementary information upon publication. 956 \n 957 \nAcknowledgments 958 \nWe thank Oliver Bernhardt, Tejas Gandhi, Roland Bruderer, and Monika Pepelnjak for 959 \ndiscussions. We thank Biognosys for early access to the ISF detection feature of Spectronaut 20. 960 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n42 \n \nP. P. was funded by the Promedica Stiftung (2025-0022/M) and by an ETH Zurich Grant (25-2 ETH-961 \n027). 962 \n 963 \nAuthor contributions 964 \nThorben Schramm:  Conceptualization, Methodology, Software, Validation, Formal analysis, 965 \nInvestigation, Resources, Data Curation, Writing - Original Draft, Writing - Review & Editing, 966 \nVisualization, Project administration 967 \nLudovic Gillet:  Conceptualization, Methodology, Software, Validation, Formal analysis, 968 \nInvestigation, Resources, Writing - Original Draft, Writing - Review & Editing 969 \nViviane Reber: Resources 970 \nNatalie de Souza: Writing - Review & Editing 971 \nMatthias Gstaiger: Supervision 972 \nPaola Picotti: Conceptualization, Writing - Original Draft, Writing - Review & Editing, Supervision, 973 \nProject administration, Funding acquisition 974 \n 975 \nDeclaration of interests 976 \nP. P. is a scientific advisor for the company Biognosys AG (Schlieren, Switzerland) and an inventor 977 \nof a patent licensed by Biognosys AG that covers the LiP-MS method. The remaining authors 978 \ndeclare no competing interests. 979 \n 980 \nReferences 981 \n1. Alseekh, S. et al. Mass spectrometry-based metabolomics: a guide for annotation, quantification 982 \nand best reporting practices. Nat. Methods 18, 747–756 (2021). 983 \n2. Collins, S. L., Koo, I., Peters, J. M., Smith, P. B. & Patterson, A. D. Current Challenges and Recent 984 \nDevelopments in Mass Spectrometry–Based Metabolomics. Annu. Rev. Anal. Chem. 14, 467–487 985 \n(2021). 986 \n3. Guo, T., Steen, J. A. & Mann, M. Mass-spectrometry-based proteomics: from single cells to clinical 987 \napplications. Nature 638, 901–911 (2025). 988 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n43 \n \n4. Aebersold, R. & Mann, M. Mass-spectrometric exploration of proteome structure and function. 989 \nNature 537, 347–355 (2016). 990 \n5. Thomas, S. N., French, D., Jannetto, P. J., Rappold, B. A. & Clarke, W. A. Liquid chromatography–991 \ntandem mass spectrometry for clinical diagnostics. Nat. Rev. Methods Primer 2, 96 (2022). 992 \n6. Glish, G. L. & Vachet, R. W. The basics of mass spectrometry in the twenty-first century. Nat. Rev. 993 \nDrug Discov. 2, 140–150 (2003). 994 \n7. Yamashita, M. & Fenn, J. B. Electrospray ion source. Another variation on the free-jet theme. J. Phys. 995 \nChem. 88, 4451–4459 (1984). 996 \n8. Prabhu, G. R. D., Williams, E. R., Wilm, M. & Urban, P. L. Mass spectrometry using electrospray 997 \nionization. Nat. Rev. Methods Primer 3, 23 (2023). 998 \n9. Daniel, J. M., Friess, S. D., Rajagopalan, S., Wendt, S. & Zenobi, R. Quantitative determination of 999 \nnoncovalent binding interactions using soft ionization mass spectrometry. Int. J. Mass Spectrom. 1000 \n216, 1–27 (2002). 1001 \n10. Guo, J., Shen, S., Xing, S., Yu, H. & Huan, T. ISFrag: De Novo Recognition of In-Source Fragments for 1002 \nLiquid Chromatography–Mass Spectrometry Data. Anal. Chem. 93, 10243–10250 (2021). 1003 \n11. Xu, Y.-F., Lu, W. & Rabinowitz, J. D. Avoiding Misannotation of In-Source Fragmentation Products as 1004 \nCellular Metabolites in Liquid Chromatography–Mass Spectrometry-Based Metabolomics. Anal. 1005 \nChem. 87, 2273–2281 (2015). 1006 \n12. Domingo-Almenara, X. et al. Autonomous METLIN-Guided In-source Fragment Annotation for 1007 \nUntargeted Metabolomics. Anal. Chem. 91, 3246–3253 (2019). 1008 \n13. Farke, N., Schramm, T., Verhülsdonk, A., Rapp, J. & Link, H. Systematic analysis of in-source 1009 \nmodifications of primary metabolites during flow-injection time-of-flight mass spectrometry. Anal. 1010 \nBiochem. 664, 115036 (2023). 1011 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n44 \n \n14. Chen, L. et al. Widespread occurrence of in-source fragmentation in the analysis of natural 1012 \ncompounds by liquid chromatography–electrospray ionization mass spectrometry. Rapid Commun. 1013 \nMass Spectrom. 37, e9519 (2023). 1014 \n15. Schmid, R. et al. Ion identity molecular networking for mass spectrometry-based metabolomics in 1015 \nthe GNPS environment. Nat. Commun. 12, 3832 (2021). 1016 \n16. Lu, W. et al. Improved Annotation of Untargeted Metabolomics Data through Buffer Modifications 1017 \nThat Shift Adduct Mass and Intensity. Anal. Chem. 92, 11573–11581 (2020). 1018 \n17. Giera, M., Aisporna, A., Uritboonthai, W. & Siuzdak, G. The hidden impact of in-source 1019 \nfragmentation in metabolic and chemical mass spectrometry data interpretation. Nat. Metab. 6, 1020 \n1647–1648 (2024). 1021 \n18. El Abiead, Y. et al. Discovery of metabolites prevails amid in-source fragmentation. Nat. Metab. 7, 1022 \n435–437 (2025). 1023 \n19. Senan, O. et al. CliqueMS: a computational tool for annotating in-source metabolite ions from LC-1024 \nMS untargeted metabolomics data based on a coelution similarity network. Bioinformatics 35, 1025 \n4089–4097 (2019). 1026 \n20. Fuhrer, T., Heer, D., Begemann, B. & Zamboni, N. High-Throughput, Accurate Mass Metabolome 1027 \nProfiling of Cellular Extracts by Flow Injection–Time-of-Flight Mass Spectrometry. Anal. Chem. 83, 1028 \n7074–7080 (2011). 1029 \n21. Kim, J.-S., Monroe, M. E., Camp, D. G. I., Smith, R. D. & Qian, W.-J. In-Source Fragmentation and the 1030 \nSources of Partially Tryptic Peptides in Shotgun Proteomics. J. Proteome Res. 12, 910–916 (2013). 1031 \n22. Gillet, L. C. et al. Targeted Data Extraction of the MS/MS Spectra Generated by Data-independent 1032 \nAcquisition: A New Concept for Consistent and Accurate Proteome Analysis*. Mol. Cell. Proteomics 1033 \n11, O111.016717 (2012). 1034 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n45 \n \n23. Gillet, L. C., Leitner, A. & Aebersold, R. Mass Spectrometry Applied to Bottom-Up Proteomics: 1035 \nEntering the High-Throughput Era for Hypothesis Testing. Annu. Rev. Anal. Chem. 9, 449–472 (2016). 1036 \n24. Ting, Y. S. et al. Peptide-Centric Proteome Analysis: An Alternative Strategy for the Analysis of 1037 \nTandem Mass Spectrometry Data*. Mol. Cell. Proteomics 14, 2301–2307 (2015). 1038 \n25. Chong, C., Coukos, G. & Bassani-Sternberg, M. Identification of tumor antigens with 1039 \nimmunopeptidomics. Nat. Biotechnol. 40, 175–188 (2022). 1040 \n26. Feng, Y. et al. Global analysis of protein structural changes in complex proteomes. Nat. Biotechnol. 1041 \n32, 1036–1044 (2014). 1042 \n27. Kalxdorf, M., Müller, T., Stegle, O. & Krijgsveld, J. IceR improves proteome coverage and data 1043 \ncompleteness in global and single-cell proteomics. Nat. Commun. 12, 4787 (2021). 1044 \n28. Midha, M. K. et al. A comprehensive spectral assay library to quantify the Escherichia coli proteome 1045 \nby DIA/SWATH-MS. Sci. Data 7, 389 (2020). 1046 \n29. Jayavelu, A. K. et al. The proteogenomic subtypes of acute myeloid leukemia. Cancer Cell 40, 301-1047 \n317.e12 (2022). 1048 \n30. Salovska, B. et al. Peroxiredoxin 6 protects irradiated cells from oxidative stress and shapes their 1049 \nsenescence-associated cytokine landscape. Redox Biol. 49, 102212 (2022). 1050 \n31. Gunter, H. M. et al. A universal molecular control for DNA, mRNA and protein expression. Nat. 1051 \nCommun. 15, 2480 (2024). 1052 \n32. Bradley, D. et al. The fitness cost of spurious phosphorylation. EMBO J. 43, 4720–4751 (2024). 1053 \n33. Guzman, U. H. et al. Ultra-fast label-free quantification and comprehensive proteome coverage with 1054 \nnarrow-window data-independent acquisition. Nat. Biotechnol. 42, 1855–1866 (2024). 1055 \n34. Kattelus, R. et al. Phenotypic profiling of human induced regulatory T cells at early differentiation: 1056 \ninsights into distinct immunosuppressive potential. Cell. Mol. Life Sci. 81, 399 (2024). 1057 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n46 \n \n35. Zare, A. et al. Axonal tau reduction ameliorates tau and amyloid pathology in a mouse model of 1058 \nAlzheimer’s disease. Transl. Neurodegener. 14, 39 (2025). 1059 \n36. Zhang, H. et al. Heterochromatome wide analyses reveal MBD2 as a phase separation scaffold for 1060 \nheterochromatin compartmentalization and composition. Nucleic Acids Res. 53, gkaf1380 (2025). 1061 \n37. Romero-Pérez, P. S. et al. Protein surface chemistry encodes an adaptive tolerance to desiccation. 1062 \nCell Syst. 16, 101407 (2025). 1063 \n38. Arroyo-Gomez, J. et al. Functional landscape of ubiquitin linkages couples K29-linked ubiquitylation 1064 \nto epigenome integrity. EMBO J. 44, 6944–6978 (2025). 1065 \n39. Botella, J. et al. Sprint interval exercise disrupts mitochondrial ultrastructure driving a unique 1066 \nmitochondrial stress response and remodelling in men. Nat. Commun. 17, 71 (2025). 1067 \n40. Pereyra, G. et al. SFRP1 upregulation causes hippocampal synaptic dysfunction and memory 1068 \nimpairment. Cell Rep. 44, 115535 (2025). 1069 \n41. Gharibi, B. et al. Post-gastrulation amnioids as a stem cell-derived model of human extra-embryonic 1070 \ndevelopment. Cell 188, 3757-3774.e20 (2025). 1071 \n42. Su, J. et al. Polymerization-mediated SRFR1 condensation in upper lateral root cap cells regulates 1072 \nroot growth. Plant Cell 38, koaf292 (2026). 1073 \n43. Benjamini, Y. & Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful 1074 \nApproach to Multiple Testing. J. R. Stat. Soc. Ser. B Methodol. 57, 289–300 (1995). 1075 \n44. Hinterholzer, A. et al. Detecting aspartate isomerization and backbone cleavage after aspartate in 1076 \nintact proteins by NMR spectroscopy. J. Biomol. NMR 75, 71–82 (2021). 1077 \n45. Bodenmiller, B. et al. Phosphoproteomic Analysis Reveals Interconnected System-Wide Responses 1078 \nto Perturbations of Kinases and Phosphatases in Yeast. Sci. Signal. 3, rs4–rs4 (2010). 1079 \n46. Bekker-Jensen, D. B. et al. Rapid and site-specific deep phosphoproteome profiling by data-1080 \nindependent acquisition without the need for spectral libraries. Nat. Commun. 11, 787 (2020). 1081 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n47 \n \n47. Wang, Y. et al. GABAA receptor π forms channels that stimulate ERK through a G-protein-dependent 1082 \npathway. Mol. Cell 85, 166-176.e5 (2025). 1083 \n48. He, Y. et al. Evaluation of the Orbitrap Ascend Tribrid Mass Spectrometer for Shotgun Proteomics. 1084 \nAnal. Chem. 95, 10655–10663 (2023). 1085 \n49. Criscuolo, A., Zeller, M. & Fedorova, M. Evaluation of Lipid In-Source Fragmentation on Different 1086 \nOrbitrap-based Mass Spectrometers. J. Am. Soc. Mass Spectrom. 31, 463–466 (2020). 1087 \n50. Yu, F. et al. Fast Quantitative Analysis of timsTOF PASEF Data with MSFragger and IonQuant. Mol. 1088 \nCell. Proteomics 19, 1575–1585 (2020). 1089 \n51. Reber, V. et al. Paradoxical non-catalytic kinase functions are driven by inhibitor-induced 1090 \ndisplacement of autoinhibitory domains. 2025.11.06.687012 Preprint at 1091 \nhttps://doi.org/10.1101/2025.11.06.687012 (2026). 1092 \n52. Elsässer, F. et al. Limited proteolysis-coupled mass spectrometry captures proteome-wide protein 1093 \nstructural alterations and biomolecular condensation in living cells. Mol. Syst. Biol. (2026) 1094 \ndoi:10.1038/s44320-025-00182-6. 1095 \n53. Abelin, J. G. & Cox, A. L. Innovations Toward Immunopeptidomics. Mol. Cell. Proteomics 23, 100823 1096 \n(2024). 1097 \n54. Purcell, A. W., Ramarathinam, S. H. & Ternette, N. Mass spectrometry–based identification of MHC-1098 \nbound peptides for immunopeptidomics. Nat. Protoc. 14, 1687–1707 (2019). 1099 \n55. Pak, H. et al. Sensitive Immunopeptidomics by Leveraging Available Large-Scale Multi-HLA Spectral 1100 \nLibraries, Data-Independent Acquisition, and MS/MS Prediction. Mol. Cell. Proteomics 20, 100080 1101 \n(2021). 1102 \n56. Kessler, A. L. et al. HLA I immunopeptidome of synthetic long peptide pulsed human dendritic cells 1103 \nfor therapeutic vaccine design. Npj Vaccines 10, 12 (2025). 1104 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n48 \n \n57. Dorvash, M., Illing, P. T., Croft, N. P., Ramarathinam, S. H. & Purcell, A. W. Deep Exploration of the 1105 \nImmunopeptidome of a Pancreatic Cancer Cell Line: Implications for Clinical Immunopeptidomics 1106 \nand Immunotherapy. Mol. Cell. Proteomics 24, 101030 (2025). 1107 \n58. Tanuwidjaya, E. et al. SAPrIm 2.0: a semi-automated protocol for mid-throughput soluble HLA 1108 \nimmunopeptidomics. Front. Immunol. 16, (2025). 1109 \n59. Cappelletti, V. et al. Dynamic 3D proteomes reveal protein functional alterations at high resolution 1110 \nin situ. Cell 184, 545-559.e22 (2021). 1111 \n60. Mehta, V. et al. Structure of Mycobacterium tuberculosis Cya, an evolutionary ancestor of the 1112 \nmammalian membrane adenylyl cyclases. eLife 11, e77032 (2022). 1113 \n61. Li, K. et al. A peptide-centric local stability assay enables proteome-scale identification of the 1114 \nprotein targets and binding regions of diverse ligands. Nat. Methods 22, 278–282 (2025). 1115 \n62. Schubert, O. T. et al. Building high-quality assay libraries for targeted analysis of SWATH MS data. 1116 \nNat. Protoc. 10, 426–441 (2015). 1117 \n63. Dunham, W. H., Mullin, M. & Gingras, A.-C. Affinity-purification coupled to mass spectrometry: Basic 1118 \nprinciples and strategies. PROTEOMICS 12, 1576–1590 (2012). 1119 \n64. O’Reilly, F. J. & Rappsilber, J. Cross-linking mass spectrometry: methods and applications in 1120 \nstructural, molecular and systems biology. Nat. Struct. Mol. Biol. 25, 1000–1008 (2018). 1121 \n65. Goel, R. K., Bithi, N. & Emili, A. Trends in co-fractionation mass spectrometry: A new gold-standard 1122 \nin global protein interaction network discovery. Curr. Opin. Struct. Biol. 88, 102880 (2024). 1123 \n66. Malta, C. F. et al. Pushing the limits of hydrogen/deuterium exchange mass spectrometry to study 1124 \nprotein:fragment low affinity interactions. Commun. Chem. 8, 405 (2025). 1125 \n67. Harvey, D. J. Analysis of Protein Glycosylation by Mass Spectrometry. in Analysis of Protein Post-1126 \nTranslational Modifications by Mass Spectrometry 89–159 (John Wiley & Sons, Ltd, 2016). 1127 \ndoi:10.1002/9781119250906.ch3. 1128 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint \n\n49 \n \n68. Cardon, T., Fournier, I. & Salzet, M. Chasing the Ghost Proteome in the Dark Matter. Mol. Cell. 1129 \nProteomics 24, 101076 (2025). 1130 \n69. Varjosalo, M. et al. The Protein Interaction Landscape of the Human CMGC Kinase Group. Cell Rep. 1131 \n3, 1306–1320 (2013). 1132 \n70. The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Res. 1133 \n53, D609–D617 (2025). 1134 \n71. Gamble, D. venn (https://ch.mathworks.com/matlabcentral/fileexchange/22282-venn), MATLAB 1135 \nCentral File Exchange. (2026). 1136 \n 1137 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted March 30, 2026. ; https://doi.org/10.64898/2026.03.27.714398doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}