DeepLC can predict retention times for peptides that carry as-yet unseen modifications

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-14

DeepLC, a novel deep learning peptide retention time predictor, accurately predicts retention times for modified peptides, including those with unseen modifications, by using an atomic composition-based encoding.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

DeepLC is a deep learning peptide retention time predictor designed to accurately forecast LC-MS retention times for modified peptides whose modifications were not present in the model’s training data, using an atomic composition–based peptide encoding to overcome limitations of prior retention-time encodings. The study reports performance comparable to state-of-the-art methods for unmodified peptides and improved accuracy for “unseen” modifications, and it demonstrates a use case in which potentially incorrect identifications in an open modification search of CD8-positive T-cell proteomics can be flagged using retention time prediction errors. A stated caveat is that the preprint has not been peer reviewed. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract The inclusion of peptide retention time prediction promises to remove peptide identification ambiguity in complex LC-MS identification workflows. However, due to the way peptides are encoded in current prediction models, accurate retention times cannot be predicted for modified peptides. This is especially problematic for fledgling open modification searches, which will benefit from accurate retention time prediction for modified peptides to reduce identification ambiguity. We here therefore present DeepLC, a novel deep learning peptide retention time predictor utilizing a new peptide encoding based on atomic composition that allows the retention time of (previously unseen) modified peptides to be predicted accurately. We show that DeepLC performs similarly to current state-of-the-art approaches for unmodified peptides, and, more importantly, accurately predicts retention times for modifications not seen during training. Moreover, we show that DeepLC’s ability to predict retention times for any modification enables potentially incorrect identifications to be flagged in an open modification search of CD8-positive T-cell proteome data. DeepLC is available under the permissive Apache 2.0 open source license and comes with a user-friendly graphical user interface, as well as a Python package on PyPI, Bioconda, and BioContainers for effortless workflow integration.
Full text 17,991 characters · extracted from preprint-html · click to expand
DeepLC can predict retention times for peptides that carry as-yet unseen modifications | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article DeepLC can predict retention times for peptides that carry as-yet unseen modifications Lennart Martens, Robbin Bouwmeester, Ralf Gabriels, Niels Hulstaert, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-275246/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The inclusion of peptide retention time prediction promises to remove peptide identification ambiguity in complex LC-MS identification workflows. However, due to the way peptides are encoded in current prediction models, accurate retention times cannot be predicted for modified peptides. This is especially problematic for fledgling open modification searches, which will benefit from accurate retention time prediction for modified peptides to reduce identification ambiguity. We here therefore present DeepLC, a novel deep learning peptide retention time predictor utilizing a new peptide encoding based on atomic composition that allows the retention time of (previously unseen) modified peptides to be predicted accurately. We show that DeepLC performs similarly to current state-of-the-art approaches for unmodified peptides, and, more importantly, accurately predicts retention times for modifications not seen during training. Moreover, we show that DeepLC’s ability to predict retention times for any modification enables potentially incorrect identifications to be flagged in an open modification search of CD8-positive T-cell proteome data. DeepLC is available under the permissive Apache 2.0 open source license and comes with a user-friendly graphical user interface, as well as a Python package on PyPI, Bioconda, and BioContainers for effortless workflow integration. Bioinformatics Computational Biology Artificial Intelligence and Machine Learning Liquid Chromatography (LC) Mass Spectrometry (MS) proteomics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Full Text Additional Declarations There is NO Competing Interest. Supplementary Files DeepLCmainsupplementalresubmissionfinal.pdf Supplemental information for: DeepLC can predict retention times for peptides that carry as-yet unseen modifications Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-275246","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":16062529,"identity":"aeaba72d-b9e8-4237-bcf2-ffb887118b2f","order_by":0,"name":"Lennart Martens","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABLElEQVRIiWNgGAWjYBACPnQBOQZmHhibB10SDNgglAGIYGwAEsZwLTzEaklsYCCkhb332IOfe/4w6Lafff7g44669A3HeQ8w8/yyybNn4D34AJsWnnPphj3PDBjMzqQbNs48czh3w2G+BGbevrRiHga+ZANsWiRyzCR4DgC1HEhjbOZtO5A7s5nH/Ddvz+HEHgYeMwlsWuTfmEn+AWk5/4yx+W9bXbpkM48BM2/Pf5AW8x9YbeExkwbbcgNoC2MbcwI/M1ALz48DYFuwep8nx0xa5oAxj9mNZ4wze9sOG/Yz8yUwzm1ITuw5zJeMzWH87GfMJN8ckJMzO5/G8OFnW508G//ZAwxv/tgltrf3HvyANZghAC0KGNuABDMe9VjAH9KUj4JRMApGwbAGAM1SWiunG9aZAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-4277-658X","institution":"UGent and VIB","correspondingAuthor":true,"prefix":"","firstName":"Lennart","middleName":"","lastName":"Martens","suffix":""},{"id":16062530,"identity":"2b811a9a-7ee0-453c-82ea-75f4436c7da8","order_by":1,"name":"Robbin Bouwmeester","email":"","orcid":"https://orcid.org/0000-0001-6807-7029","institution":"VIB-UGent","correspondingAuthor":false,"prefix":"","firstName":"Robbin","middleName":"","lastName":"Bouwmeester","suffix":""},{"id":16062531,"identity":"25ff1968-dbcd-4bae-a883-a86f93b2a110","order_by":2,"name":"Ralf Gabriels","email":"","orcid":"https://orcid.org/0000-0002-1679-1711","institution":"Ghent University","correspondingAuthor":false,"prefix":"","firstName":"Ralf","middleName":"","lastName":"Gabriels","suffix":""},{"id":16062532,"identity":"20be9016-31fc-49a6-8507-2be685a41b8d","order_by":3,"name":"Niels Hulstaert","email":"","orcid":"","institution":"Flanders Institute for Biotechnology","correspondingAuthor":false,"prefix":"","firstName":"Niels","middleName":"","lastName":"Hulstaert","suffix":""},{"id":16062533,"identity":"713cd40c-98c8-40a6-8d9b-34b56f7ce2f2","order_by":4,"name":"Sven Degroeve","email":"","orcid":"","institution":"Flanders Institute for Biotechnology","correspondingAuthor":false,"prefix":"","firstName":"Sven","middleName":"","lastName":"Degroeve","suffix":""}],"badges":[],"createdAt":"2021-02-24 18:46:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-275246/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-275246/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":6905774,"identity":"91655081-ec4b-4d84-8c4c-8dcf1754e850","added_by":"auto","created_at":"2021-03-12 23:32:25","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":87610,"visible":true,"origin":"","legend":"Scatter plot of predicted against observed on three of the largest data sets; SWATH Library, HeLa HF, and DIA HF.","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/fe26484c71aaece0d9cb1946.jpg"},{"id":6905475,"identity":"dbe3149e-ff3f-45ba-8bef-303405b8c725","added_by":"auto","created_at":"2021-03-12 23:29:25","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":73557,"visible":true,"origin":"","legend":"For the twenty data sets the number of training peptides (y-axis) is plotted against the relative MAE (left), relative Δ𝑡95% (middle) and Pearson correlation (right).","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/e324187b803d4617efba8bea.jpg"},{"id":6905780,"identity":"5af52f03-3445-4047-91dd-705897115da1","added_by":"auto","created_at":"2021-03-12 23:32:26","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":86762,"visible":true,"origin":"","legend":"Learning curves for each of the three selected data sets. Prediction performances (R and MAE) for models trained on different training set sizes (x-axis) are computed for a fixed test set.","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/603b17fdfea2ce8dfced5726.jpg"},{"id":6906245,"identity":"b54ba731-c45b-42fa-8b88-53028964a578","added_by":"auto","created_at":"2021-03-12 23:35:25","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":58840,"visible":true,"origin":"","legend":"The modification that was excluded for training is shown on the horizontal axis, and the vertical axis shows the retention time error (experimental - predicted) when the modification was either not encoded (red) or encoded during the predictions (blue).","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/7a47dd6ef865d0e8cb23fef5.jpg"},{"id":6905776,"identity":"da17b0b7-d00d-4747-8f95-4343a2886057","added_by":"auto","created_at":"2021-03-12 23:32:25","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":77993,"visible":true,"origin":"","legend":"Each amino acid that was excluded for training is shown as a circle, where the size of the circle and color indicates the remaining training peptides and chemical property, respectively. The amino acid is either encoded as glycine (vertical axis) or as its own atomic composition (horizontal axis) and its position depicts the MAE for all amino acid containing peptides. This means that everything above the diagonal line is predicted with a higher accuracy when the amino acid is encoded as itself, while the reverse is true if it is below the diagonal line.","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/716b9b405e72c1b1a0bc744b.jpg"},{"id":6905778,"identity":"7f25eb7a-29f6-435d-88d4-361233071909","added_by":"auto","created_at":"2021-03-12 23:32:25","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":143019,"visible":true,"origin":"","legend":"Predicted retention time analysis for open modification results of a CD8-positive T-cell data set. (A) Predicted and observed retention times split by q-value. Panel (B) shows the error distribution split by q-value and any points that are higher than 1.5 times the interquartile range plus the relevant quartile range are excluded from the plot. Panel (C) only shows PSMs with q-values \u003c 0.01. Colors indicate four different subsets of detected modifications. (1) PSMs carrying the top ten most abundant modifications (2) PSMs carrying the ten modifications with the largest absolute mean error (3) PSMs carrying modifications that are not expected to occur in the sample; and (4) PSMs carrying the top ten mutations with the largest absolute mean error. Panel (D) contains PSMs that were identified to contain a dethiomethyl modification and predictions for the same peptides where the dethiomethyl is replaced with an oxidation.","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/40bf9eb901b77d68904ae20f.jpg"},{"id":6905479,"identity":"5edd0f9b-c98b-4918-bfe1-c95adc7cca2b","added_by":"auto","created_at":"2021-03-12 23:29:25","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":98636,"visible":true,"origin":"","legend":"Visualization of DeepLC’s convolutional architecture with the four individual paths named: One-hot encoding, Global features, Diamino acids composition, and Amino acids composition. These individual paths are concatenated in the Combination of representations path.","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/2990c9a8102088399270bf78.jpg"},{"id":13606621,"identity":"41b529ea-f80f-4b0a-a82b-d6e8aa439750","added_by":"auto","created_at":"2021-09-17 06:08:49","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":10957913,"visible":true,"origin":"","legend":"Article File","description":"","filename":"DeepLCmainmanuscriptresubmissionfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1_covered.pdf"},{"id":6906411,"identity":"a7d4ba28-5f0a-42c8-8484-b2f9c849fce8","added_by":"auto","created_at":"2021-03-12 23:38:29","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5649404,"visible":true,"origin":"","legend":"Article File","description":"","filename":"DeepLCmainmanuscriptresubmissionfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1_stamped.pdf"},{"id":6906246,"identity":"3eb96797-9606-4b96-8489-ef976ef14534","added_by":"auto","created_at":"2021-03-12 23:35:25","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":14363816,"visible":true,"origin":"","legend":"Supplemental information for: DeepLC can predict retention times for peptides that carry as-yet unseen modifications","description":"","filename":"DeepLCmainsupplementalresubmissionfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-275246/v1/e35e183f277bdbd3a3edfe71.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"DeepLC can predict retention times for peptides that carry as-yet unseen modifications","fulltext":[{"header":"Full Text","content":"\u003cp\u003eThis preprint is available for \u003ca href='/article/rs-275246/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Liquid Chromatography (LC), Mass Spectrometry (MS), proteomics","lastPublishedDoi":"10.21203/rs.3.rs-275246/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-275246/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The inclusion of peptide retention time prediction promises to remove peptide identification ambiguity in complex LC-MS identification workflows. However, due to the way peptides are encoded in current prediction models, accurate retention times cannot be predicted for modified peptides. This is especially problematic for fledgling open modification searches, which will benefit from accurate retention time prediction for modified peptides to reduce identification ambiguity. We here therefore present DeepLC, a novel deep learning peptide retention time predictor utilizing a new peptide encoding based on atomic composition that allows the retention time of (previously unseen) modified peptides to be predicted accurately. We show that DeepLC performs similarly to current state-of-the-art approaches for unmodified peptides, and, more importantly, accurately predicts retention times for modifications not seen during training. Moreover, we show that DeepLC’s ability to predict retention times for any modification enables potentially incorrect identifications to be flagged in an open modification search of CD8-positive T-cell proteome data. DeepLC is available under the permissive Apache 2.0 open source license and comes with a user-friendly graphical user interface, as well as a Python package on PyPI, Bioconda, and BioContainers for effortless workflow integration.","manuscriptTitle":"DeepLC can predict retention times for peptides that carry as-yet unseen modifications","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-03-12 23:29:23","doi":"10.21203/rs.3.rs-275246/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4db58cad-0862-4b79-84a3-56d58ce5e507","owner":[],"postedDate":"March 12th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":2941835,"name":"Bioinformatics"},{"id":2941836,"name":"Computational Biology"},{"id":2941837,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2021-04-28T20:41:15+00:00","versionOfRecord":[],"versionCreatedAt":"2021-03-12 23:29:23","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-275246","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-275246","identity":"rs-275246","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00