DeepLC can predict retention times for peptides that carry as-yet unseen modifications | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article DeepLC can predict retention times for peptides that carry as-yet unseen modifications Lennart Martens, Robbin Bouwmeester, Ralf Gabriels, Niels Hulstaert, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-275246/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The inclusion of peptide retention time prediction promises to remove peptide identification ambiguity in complex LC-MS identification workflows. However, due to the way peptides are encoded in current prediction models, accurate retention times cannot be predicted for modified peptides. This is especially problematic for fledgling open modification searches, which will benefit from accurate retention time prediction for modified peptides to reduce identification ambiguity. We here therefore present DeepLC, a novel deep learning peptide retention time predictor utilizing a new peptide encoding based on atomic composition that allows the retention time of (previously unseen) modified peptides to be predicted accurately. We show that DeepLC performs similarly to current state-of-the-art approaches for unmodified peptides, and, more importantly, accurately predicts retention times for modifications not seen during training. Moreover, we show that DeepLC’s ability to predict retention times for any modification enables potentially incorrect identifications to be flagged in an open modification search of CD8-positive T-cell proteome data. DeepLC is available under the permissive Apache 2.0 open source license and comes with a user-friendly graphical user interface, as well as a Python package on PyPI, Bioconda, and BioContainers for effortless workflow integration. Bioinformatics Computational Biology Artificial Intelligence and Machine Learning Liquid Chromatography (LC) Mass Spectrometry (MS) proteomics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Full Text Additional Declarations There is NO Competing Interest. Supplementary Files DeepLCmainsupplementalresubmissionfinal.pdf Supplemental information for: DeepLC can predict retention times for peptides that carry as-yet unseen modifications Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-275246","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":16062529,"identity":"aeaba72d-b9e8-4237-bcf2-ffb887118b2f","order_by":0,"name":"Lennart Martens","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABLElEQVRIiWNgGAWjYBACPnQBOQZmHhibB10SDNgglAGIYGwAEsZwLTzEaklsYCCkhb332IOfe/4w6Lafff7g44669A3HeQ8w8/yyybNn4D34AJsWnnPphj3PDBjMzqQbNs48czh3w2G+BGbevrRiHga+ZANsWiRyzCR4DgC1HEhjbOZtO5A7s5nH/Ddvz+HEHgYeMwlsWuTfmEn+AWk5/4yx+W9bXbpkM48BM2/Pf5AW8x9YbeExkwbbcgNoC2MbcwI/M1ALz48DYFuwep8nx0xa5oAxj9mNZ4wze9sOG/Yz8yUwzm1ITuw5zJeMzWH87GfMJN8ckJMzO5/G8OFnW508G//ZAwxv/tgltrf3HvyANZghAC0KGNuABDMe9VjAH9KUj4JRMApGwbAGAM1SWiunG9aZAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-4277-658X","institution":"UGent and VIB","correspondingAuthor":true,"prefix":"","firstName":"Lennart","middleName":"","lastName":"Martens","suffix":""},{"id":16062530,"identity":"2b811a9a-7ee0-453c-82ea-75f4436c7da8","order_by":1,"name":"Robbin Bouwmeester","email":"","orcid":"https://orcid.org/0000-0001-6807-7029","institution":"VIB-UGent","correspondingAuthor":false,"prefix":"","firstName":"Robbin","middleName":"","lastName":"Bouwmeester","suffix":""},{"id":16062531,"identity":"25ff1968-dbcd-4bae-a883-a86f93b2a110","order_by":2,"name":"Ralf Gabriels","email":"","orcid":"https://orcid.org/0000-0002-1679-1711","institution":"Ghent University","correspondingAuthor":false,"prefix":"","firstName":"Ralf","middleName":"","lastName":"Gabriels","suffix":""},{"id":16062532,"identity":"20be9016-31fc-49a6-8507-2be685a41b8d","order_by":3,"name":"Niels Hulstaert","email":"","orcid":"","institution":"Flanders Institute for Biotechnology","correspondingAuthor":false,"prefix":"","firstName":"Niels","middleName":"","lastName":"Hulstaert","suffix":""},{"id":16062533,"identity":"713cd40c-98c8-40a6-8d9b-34b56f7ce2f2","order_by":4,"name":"Sven Degroeve","email":"","orcid":"","institution":"Flanders Institute for Biotechnology","correspondingAuthor":false,"prefix":"","firstName":"Sven","middleName":"","lastName":"Degroeve","suffix":""}],"badges":[],"createdAt":"2021-02-24 18:46:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-275246/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-275246/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":6905774,"identity":"91655081-ec4b-4d84-8c4c-8dcf1754e850","added_by":"auto","created_at":"2021-03-12 23:32:25","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":87610,"visible":true,"origin":"","legend":"Scatter plot of predicted against observed on three of the largest data sets; SWATH Library, HeLa HF, and DIA HF.","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/fe26484c71aaece0d9cb1946.jpg"},{"id":6905475,"identity":"dbe3149e-ff3f-45ba-8bef-303405b8c725","added_by":"auto","created_at":"2021-03-12 23:29:25","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":73557,"visible":true,"origin":"","legend":"For the twenty data sets the number of training peptides (y-axis) is plotted against the relative MAE (left), relative Δ𝑡95% (middle) and Pearson correlation (right).","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/e324187b803d4617efba8bea.jpg"},{"id":6905780,"identity":"5af52f03-3445-4047-91dd-705897115da1","added_by":"auto","created_at":"2021-03-12 23:32:26","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":86762,"visible":true,"origin":"","legend":"Learning curves for each of the three selected data sets. Prediction performances (R and MAE) for models trained on different training set sizes (x-axis) are computed for a fixed test set.","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/603b17fdfea2ce8dfced5726.jpg"},{"id":6906245,"identity":"b54ba731-c45b-42fa-8b88-53028964a578","added_by":"auto","created_at":"2021-03-12 23:35:25","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":58840,"visible":true,"origin":"","legend":"The modification that was excluded for training is shown on the horizontal axis, and the vertical axis shows the retention time error (experimental - predicted) when the modification was either not encoded (red) or encoded during the predictions (blue).","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/7a47dd6ef865d0e8cb23fef5.jpg"},{"id":6905776,"identity":"da17b0b7-d00d-4747-8f95-4343a2886057","added_by":"auto","created_at":"2021-03-12 23:32:25","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":77993,"visible":true,"origin":"","legend":"Each amino acid that was excluded for training is shown as a circle, where the size of the circle and color indicates the remaining training peptides and chemical property, respectively. The amino acid is either encoded as glycine (vertical axis) or as its own atomic composition (horizontal axis) and its position depicts the MAE for all amino acid containing peptides. This means that everything above the diagonal line is predicted with a higher accuracy when the amino acid is encoded as itself, while the reverse is true if it is below the diagonal line.","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/716b9b405e72c1b1a0bc744b.jpg"},{"id":6905778,"identity":"7f25eb7a-29f6-435d-88d4-361233071909","added_by":"auto","created_at":"2021-03-12 23:32:25","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":143019,"visible":true,"origin":"","legend":"Predicted retention time analysis for open modification results of a CD8-positive T-cell data set. (A) Predicted and observed retention times split by q-value. Panel (B) shows the error distribution split by q-value and any points that are higher than 1.5 times the interquartile range plus the relevant quartile range are excluded from the plot. Panel (C) only shows PSMs with q-values \u003c 0.01. Colors indicate four different subsets of detected modifications. (1) PSMs carrying the top ten most abundant modifications (2) PSMs carrying the ten modifications with the largest absolute mean error (3) PSMs carrying modifications that are not expected to occur in the sample; and (4) PSMs carrying the top ten mutations with the largest absolute mean error. Panel (D) contains PSMs that were identified to contain a dethiomethyl modification and predictions for the same peptides where the dethiomethyl is replaced with an oxidation.","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/40bf9eb901b77d68904ae20f.jpg"},{"id":6905479,"identity":"5edd0f9b-c98b-4918-bfe1-c95adc7cca2b","added_by":"auto","created_at":"2021-03-12 23:29:25","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":98636,"visible":true,"origin":"","legend":"Visualization of DeepLC’s convolutional architecture with the four individual paths named: One-hot encoding, Global features, Diamino acids composition, and Amino acids composition. These individual paths are concatenated in the Combination of representations path.","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/2990c9a8102088399270bf78.jpg"},{"id":13606621,"identity":"41b529ea-f80f-4b0a-a82b-d6e8aa439750","added_by":"auto","created_at":"2021-09-17 06:08:49","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":10957913,"visible":true,"origin":"","legend":"Article File","description":"","filename":"DeepLCmainmanuscriptresubmissionfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1_covered.pdf"},{"id":6906411,"identity":"a7d4ba28-5f0a-42c8-8484-b2f9c849fce8","added_by":"auto","created_at":"2021-03-12 23:38:29","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5649404,"visible":true,"origin":"","legend":"Article File","description":"","filename":"DeepLCmainmanuscriptresubmissionfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1_stamped.pdf"},{"id":6906246,"identity":"3eb96797-9606-4b96-8489-ef976ef14534","added_by":"auto","created_at":"2021-03-12 23:35:25","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":14363816,"visible":true,"origin":"","legend":"Supplemental information for: DeepLC can predict retention times for peptides that carry as-yet unseen modifications","description":"","filename":"DeepLCmainsupplementalresubmissionfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/e35e183f277bdbd3a3edfe71.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"DeepLC can predict retention times for peptides that carry as-yet unseen modifications","fulltext":[{"header":"Full Text","content":"\u003cp\u003eThis preprint is available for \u003ca href='/article/rs-275246/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Liquid Chromatography (LC), Mass Spectrometry (MS), proteomics","lastPublishedDoi":"10.21203/rs.3.rs-275246/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-275246/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The inclusion of peptide retention time prediction promises to remove peptide identification ambiguity in complex LC-MS identification workflows. However, due to the way peptides are encoded in current prediction models, accurate retention times cannot be predicted for modified peptides. This is especially problematic for fledgling open modification searches, which will benefit from accurate retention time prediction for modified peptides to reduce identification ambiguity. We here therefore present DeepLC, a novel deep learning peptide retention time predictor utilizing a new peptide encoding based on atomic composition that allows the retention time of (previously unseen) modified peptides to be predicted accurately. We show that DeepLC performs similarly to current state-of-the-art approaches for unmodified peptides, and, more importantly, accurately predicts retention times for modifications not seen during training. Moreover, we show that DeepLC’s ability to predict retention times for any modification enables potentially incorrect identifications to be flagged in an open modification search of CD8-positive T-cell proteome data. DeepLC is available under the permissive Apache 2.0 open source license and comes with a user-friendly graphical user interface, as well as a Python package on PyPI, Bioconda, and BioContainers for effortless workflow integration.","manuscriptTitle":"DeepLC can predict retention times for peptides that carry as-yet unseen modifications","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-03-12 23:29:23","doi":"10.21203/rs.3.rs-275246/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4db58cad-0862-4b79-84a3-56d58ce5e507","owner":[],"postedDate":"March 12th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":2941835,"name":"Bioinformatics"},{"id":2941836,"name":"Computational Biology"},{"id":2941837,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2021-04-28T20:41:15+00:00","versionOfRecord":[],"versionCreatedAt":"2021-03-12 23:29:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-275246","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-275246","identity":"rs-275246","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.