Impact of Variant-level Batch Effects on Identification of Genetic Risk Factors in Large Sequencing Studies

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background: Genetic studies have shifted to sequencing-based rare variants discovery after decades of success in identifying common disease variants by Genome-Wide Association Studies using Single Nucleotide Polymorphism chips. Sequencing-based studies require large sample sizes for statistical power but often inadvertently introduce batch effects because samples are typically collected, processed, and sequenced at multiple centers. Conventionally, batch effects are first detected and visualized using Principal Components Analysis and then controlled by including batch covariates in the disease association models. For sequencing-based genetic studies, because all variants included in the association analyses have passed quality control measures, this conventional approach treats every variant as equal and ignores the substantial differences still remaining in variant qualities and characteristics such as genotype quality scores, alternative allele fractions (fraction of reads supporting alternative allele at a variant position) and sequencing depths. Results: In the Alzheimer’s Disease Sequencing Project (ADSP) exome dataset of 9,904 cases and controls, we discovered hidden variant-level differences between sample batches of three sequencing centers and two exome capture kits. Although sequencing centers were included as a covariate in our association models, we observed differences at the variant level in genotype quality and alternative allele fraction between samples processed by different exome capture kits that significantly impacted both the confidence of variant detection and the identification of disease-associated variants. Furthermore, we found that the association signals of a subset of top disease risk variants came exclusively from samples processed by one exome capture kit that was more effective at capturing the alternative alleles compared to the other kit. Conclusions: Our findings highlight the importance of additional variant-level quality control for large sequencing-based genetic studies. More importantly, we demonstrate that automatically filtering out variants with batch differences may lead to false negatives if the batch discordance came largely from quality differences and if the variants from one batch had better quality scores.
Full text 23,969 characters · extracted from preprint-html · click to expand
Impact of Variant-level Batch Effects on Identification of Genetic Risk Factors in Large Sequencing Studies | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Impact of Variant-level Batch Effects on Identification of Genetic Risk Factors in Large Sequencing Studies Daniel P Wickland, Yingxue Ren, Jason P Sinnwell, Joseph S Reddy, and 11 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-44710/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Genetic studies have shifted to sequencing-based rare variants discovery after decades of success in identifying common disease variants by Genome-Wide Association Studies using Single Nucleotide Polymorphism chips. Sequencing-based studies require large sample sizes for statistical power but often inadvertently introduce batch effects because samples are typically collected, processed, and sequenced at multiple centers. Conventionally, batch effects are first detected and visualized using Principal Components Analysis and then controlled by including batch covariates in the disease association models. For sequencing-based genetic studies, because all variants included in the association analyses have passed quality control measures, this conventional approach treats every variant as equal and ignores the substantial differences still remaining in variant qualities and characteristics such as genotype quality scores, alternative allele fractions (fraction of reads supporting alternative allele at a variant position) and sequencing depths. Results: In the Alzheimer’s Disease Sequencing Project (ADSP) exome dataset of 9,904 cases and controls, we discovered hidden variant-level differences between sample batches of three sequencing centers and two exome capture kits. Although sequencing centers were included as a covariate in our association models, we observed differences at the variant level in genotype quality and alternative allele fraction between samples processed by different exome capture kits that significantly impacted both the confidence of variant detection and the identification of disease-associated variants. Furthermore, we found that the association signals of a subset of top disease risk variants came exclusively from samples processed by one exome capture kit that was more effective at capturing the alternative alleles compared to the other kit. Conclusions: Our findings highlight the importance of additional variant-level quality control for large sequencing-based genetic studies. More importantly, we demonstrate that automatically filtering out variants with batch differences may lead to false negatives if the batch discordance came largely from quality differences and if the variants from one batch had better quality scores. Medical Genetics Batch effects Exome capture Alternative allele fraction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Full Text Due to technical limitations, full-text HTML conversion of this manuscript could not be completed. However, the manuscript can be downloaded and accessed as a PDF. Supplementary Files figS4.png figS1.png figS2.png figS3.png Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-44710","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":1380515,"identity":"fd8dba0f-703c-4173-b42e-b44b222ad65e","order_by":0,"name":"Daniel P Wickland","email":"","orcid":"","institution":"Mayo Clinic Department of Health Sciences Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"P","lastName":"Wickland","suffix":""},{"id":1380516,"identity":"a26225fc-c56d-4d60-b8fa-fa72d421148c","order_by":1,"name":"Yingxue Ren","email":"","orcid":"","institution":"Mayo Clinic Department of Health Sciences Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yingxue","middleName":"","lastName":"Ren","suffix":""},{"id":1380517,"identity":"74f2d5b7-e8ff-4502-a298-f6c23cd0f32c","order_by":2,"name":"Jason P Sinnwell","email":"","orcid":"","institution":"Mayo Clinic Department of Health Sciences Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jason","middleName":"P","lastName":"Sinnwell","suffix":""},{"id":1380518,"identity":"82d17ba3-19d6-4b91-b3ed-749c6a3ee1c0","order_by":3,"name":"Joseph S Reddy","email":"","orcid":"","institution":"Mayo Clinic Department of Health Sciences Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Joseph","middleName":"S","lastName":"Reddy","suffix":""},{"id":1380519,"identity":"20ec9d34-9829-49aa-b1d2-9a12c938e710","order_by":4,"name":"Cyril Pottier","email":"","orcid":"","institution":"Mayo Clinic's Campus in Florida","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Cyril","middleName":"","lastName":"Pottier","suffix":""},{"id":1380520,"identity":"4defd2c3-12cf-4791-a44d-af008ddee756","order_by":5,"name":"Vivekananda Sarangi","email":"","orcid":"","institution":"Mayo Clinic Department of Health Sciences Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Vivekananda","middleName":"","lastName":"Sarangi","suffix":""},{"id":1380521,"identity":"5a6fb616-957b-46b9-87e3-086bdc94c7bb","order_by":6,"name":"Minerva M Carrasquillo","email":"","orcid":"","institution":"Mayo Clinic's Campus in Florida","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Minerva","middleName":"M","lastName":"Carrasquillo","suffix":""},{"id":1380522,"identity":"18194082-426d-4d85-ac84-73870270f15a","order_by":7,"name":"Owen A Ross","email":"","orcid":"","institution":"Mayo Clinic's Campus in Florida","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Owen","middleName":"A","lastName":"Ross","suffix":""},{"id":1380523,"identity":"07bb35ac-e8e0-40cb-9d2b-f4710453b7fa","order_by":8,"name":"Steven G Younkin","email":"","orcid":"","institution":"Mayo Clinic's Campus in Florida","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Steven","middleName":"G","lastName":"Younkin","suffix":""},{"id":1380524,"identity":"91990be5-55b8-414d-ac31-0ecd6ffad206","order_by":9,"name":"Nilüfer Ertekin-Taner","email":"","orcid":"","institution":"Mayo Clinic's Campus in Florida","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nilüfer","middleName":"","lastName":"Ertekin-Taner","suffix":""},{"id":1380525,"identity":"72eb8d2c-3316-451b-b8a9-021a674ad9d5","order_by":10,"name":"Rosa Rademakers","email":"","orcid":"","institution":"Mayo Clinic's Campus in Florida","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rosa","middleName":"","lastName":"Rademakers","suffix":""},{"id":1380526,"identity":"4c615031-4042-4f22-ac11-b725f837d993","order_by":11,"name":"Matthew E Hudson","email":"","orcid":"","institution":"University of Illinois at Urbana-Champaign","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"E","lastName":"Hudson","suffix":""},{"id":1380527,"identity":"a200beb5-6bbd-4311-9aeb-ab5cacdddfbb","order_by":12,"name":"Liudmila Sergeevna Mainzer","email":"","orcid":"","institution":"University of Illinois at Urbana-Champaign","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Liudmila","middleName":"Sergeevna","lastName":"Mainzer","suffix":""},{"id":1380528,"identity":"3ca76458-7822-4aef-9338-a1813bdf7ae7","order_by":13,"name":"Joanna M Biernacka","email":"","orcid":"","institution":"Mayo Clinic Department of Health Sciences Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Joanna","middleName":"M","lastName":"Biernacka","suffix":""},{"id":1380529,"identity":"9eb591d8-7137-458c-98be-6c301f2ea807","order_by":14,"name":"Yan Asmann","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2klEQVRIiWNgGAWjYHAD5gNQRgLRWthgSonXwmNAnBb+/tOJnwt3MCT2z+75uuFnmw0DP3uOAV4tEjdyN0vPPMOQOOPO2W03e9vSGCR73uDXwnCDd4M0bxtDYsON3G03eNsOMxjcIGCL/Pmzm3+DtMy/kfPs5l+gFntCWgwO5G4D27LhRg7bbbAtEgS0GALdY83bJmG88Uaa2W2Zc2k8EmeeFeDVIgd0GNBwG9l5N5Kf3XxTZiPH3568Aa8WKJBwbABRjGwMPMQoBwN7CPWHaA2jYBSMglEwggAAhV5L3SUtzwEAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-8896-2647","institution":"Mayo Clinic","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Yan","middleName":"","lastName":"Asmann","suffix":""}],"badges":[],"createdAt":"2020-07-17 10:17:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-44710/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-44710/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":1911147,"identity":"264eb105-d2cc-4f46-8552-2d37528aff85","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":216636,"visible":true,"origin":"","legend":" Principal Component (PC) eigenvector plots using genotypes of a pruned set of 16,187 high quality common variants for 9,904 ADSP individuals. Each data point represents a single individual. Clustering of samples for a particular variable signifies genotypic similarity between individuals for the trait\n represented by that color. (A) PCs of the genotypes. (B) PCs color coded based on sub-population. (C) PCs\ncolor coded based on center. (D) PCs color coded based on Gender. (E) PCs color coded based on AD\nphenotype. As expected, clustering is apparent only by sub-population. ","description":"","filename":"fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig1.png"},{"id":1911148,"identity":"5dea73b1-4966-4a36-8843-a88e531761ab","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":43983,"visible":true,"origin":"","legend":"Covariates included in each model for association tests. ","description":"","filename":"fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig2.png"},{"id":1911149,"identity":"083ad714-e9cd-401f-8e09-705a23afba79","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":176650,"visible":true,"origin":"","legend":"Sequencing center specific association p-values of SNPs that reached exome-wide significance\n (denoted by the dashed horizontal lines) in the full-dataset analysis. (A) Seven SNPs in TOMM40 and APOE. (B) Twenty-nine novel SNPs \n","description":"","filename":"fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig3.png"},{"id":1911150,"identity":"f39e0754-9cdc-45a0-8e0e-aa96a6da3c4e","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":277800,"visible":true,"origin":"","legend":"PC eigenvector plots of genotypes at 29 exome-wide significant SNPs. Each data point represents a single individual. Clustering of samples for a particular variable indicates genotypic similarity between\nindividuals for the trait represented by that color. (A) PCs color coded based on sub-population. (B) PCs color\ncoded based on gender. (C) PCs color coded based on capture kit. The NimbleGen-captured samples cluster\ntightly together, indicating their genotypic similarity that is distinct from the Illumina-captured samples. \n","description":"","filename":"fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig4.png"},{"id":1911151,"identity":"51638068-c3a8-4210-b17b-2babc776b168","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":78365,"visible":true,"origin":"","legend":" Density plots of variant quality parameters between two exome capture kits. Mean values were\ncomputed across all samples for each variant. The solid lines show the distributions of all 166,947 variants\nused in the association analyses, and the scattered dots represent the 29 novel SNPs. ","description":"","filename":"fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig5.png"},{"id":1911152,"identity":"028bf3e8-1c4b-47f1-b0ac-19b1721ff76c","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":164897,"visible":true,"origin":"","legend":"Minor Allele Frequency (MAF) of 29 exome-wide significant SNPs in AD control exomes processed by two capture kits and in the ExAC Non-Finnish European (NFE) population. ","description":"","filename":"fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig6.png"},{"id":1911153,"identity":"2e7a27f5-4f1b-478b-9139-cd68a541d10b","added_by":"auto","created_at":"2020-08-13 14:10:40","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":247221,"visible":true,"origin":"","legend":" PC eigenvector plots of genotypes at variants lying in different sections of quality-metric ratio\n distributions. Each data point represents a single individual, color coded according to capture kit. (A) PCs of\n variants in either 5% tail. (B) PCs of variants in right 5% tail. (C) PCs of variants in left 5% tail. (D) Variants in\n middle 90% of distributions. Variants in the tails, in particular the left 5% tail (better quality in NimbleGen kit),\nshow clear separation by capture kit in both cases and controls. \n","description":"","filename":"fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/fig7.png"},{"id":13522803,"identity":"c0cd9cf8-dc5a-4cad-8922-59a4c48ccff0","added_by":"auto","created_at":"2021-09-17 00:37:29","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6430637,"visible":true,"origin":"","legend":"","description":"","filename":"All.pdf","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1_covered.pdf"},{"id":1911159,"identity":"5b6dad3d-f286-465f-81d2-2c1f94c32752","added_by":"auto","created_at":"2020-08-13 14:10:44","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6549704,"visible":true,"origin":"","legend":"","description":"","filename":"All.pdf","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1_stamped.pdf"},{"id":1911154,"identity":"6b265341-b8b4-4e9b-b731-eb1b1be5496f","added_by":"auto","created_at":"2020-08-13 14:10:41","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6498423,"visible":true,"origin":"","legend":"","description":"","filename":"All.pdf","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/All.pdf"},{"id":1911155,"identity":"7815ceba-ba11-4393-8d2a-65ab65974620","added_by":"auto","created_at":"2020-08-13 14:10:41","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":46295,"visible":true,"origin":"","legend":"","description":"","filename":"figS4.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/figS4.png"},{"id":1911156,"identity":"bc318522-4cf9-407e-bd59-3cba07798bc7","added_by":"auto","created_at":"2020-08-13 14:10:41","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":75009,"visible":true,"origin":"","legend":"","description":"","filename":"figS1.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/figS1.png"},{"id":1911157,"identity":"43d63d84-e436-4b3c-b59c-e85ec5750edc","added_by":"auto","created_at":"2020-08-13 14:10:41","extension":"png","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":35411,"visible":true,"origin":"","legend":"","description":"","filename":"figS2.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/figS2.png"},{"id":1911158,"identity":"0874cea3-ee96-4cb0-9d62-87d9d3bab0a9","added_by":"auto","created_at":"2020-08-13 14:10:41","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":70274,"visible":true,"origin":"","legend":"","description":"","filename":"figS3.png","url":"https://assets-eu.researchsquare.com/files/rs-44710/v1/figS3.png"}],"financialInterests":"","formattedTitle":"\u003cp\u003eImpact of Variant-level Batch Effects on Identification of Genetic Risk Factors in Large Sequencing Studies\u003c/p\u003e","fulltext":[{"header":"Full Text","content":"\u003cp\u003eDue to technical limitations, full-text HTML conversion of this \nmanuscript could not be completed. However, \nthe manuscript can be downloaded and accessed as a PDF.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Batch effects, Exome capture, Alternative allele fraction ","lastPublishedDoi":"10.21203/rs.3.rs-44710/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-44710/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Background: Genetic studies have shifted to sequencing-based rare variants discovery after decades of success in identifying common disease variants by Genome-Wide Association Studies using Single Nucleotide Polymorphism chips. Sequencing-based studies require large sample sizes for statistical power but often inadvertently introduce batch effects because samples are typically collected, processed, and sequenced at multiple centers. Conventionally, batch effects are first detected and visualized using Principal Components Analysis and then controlled by including batch covariates in the disease association models. For sequencing-based genetic studies, because all variants included in the association analyses have passed quality control measures, this conventional approach treats every variant as equal and ignores the substantial differences still remaining in variant qualities and characteristics such as genotype quality scores, alternative allele fractions (fraction of reads supporting alternative allele at a variant position) and sequencing depths.\n\nResults: In the Alzheimer’s Disease Sequencing Project (ADSP) exome dataset of 9,904 cases and controls, we discovered hidden variant-level differences between sample batches of three sequencing centers and two exome capture kits. Although sequencing centers were included as a covariate in our association models, we observed differences at the variant level in genotype quality and alternative allele fraction between samples processed by different exome capture kits that significantly impacted both the confidence of variant detection and the identification of disease-associated variants. Furthermore, we found that the association signals of a subset of top disease risk variants came exclusively from samples processed by one exome capture kit that was more effective at capturing the alternative alleles compared to the other kit.\n\nConclusions: Our findings highlight the importance of additional variant-level quality control for large sequencing-based genetic studies. More importantly, we demonstrate that automatically filtering out variants with batch differences may lead to false negatives if the batch discordance came largely from quality differences and if the variants from one batch had better quality scores.","manuscriptTitle":"Impact of Variant-level Batch Effects on Identification of Genetic Risk Factors in Large Sequencing Studies","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-08-13 14:09:19","doi":"10.21203/rs.3.rs-44710/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b18ac9cb-c9b9-4c30-a96c-ed045a563ab9","owner":[],"postedDate":"August 13th, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":291038,"name":"Medical Genetics"}],"tags":[],"updatedAt":"2020-10-01T17:50:13+00:00","versionOfRecord":[],"versionCreatedAt":"2020-08-13 14:09:19","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-44710","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-44710","identity":"rs-44710","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00