A new hybrid record linkage process to render epidemiological databases interoperable: application to the GEMO and GENEPSO studies involving BRCA1 and BRCA2 mutation carriers | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article A new hybrid record linkage process to render epidemiological databases interoperable: application to the GEMO and GENEPSO studies involving BRCA1 and BRCA2 mutation carriers YUE JIAO, Fabienne Lesueur, Chloé-Agathe Azencott, Maïté Laurent, and 12 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-64751/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 29 Jul, 2021 Read the published version in BMC Medical Research Methodology → Version 1 posted 12 You are reading this latest preprint version Abstract Background Linking independent sources of data related to same individuals enable innovative epidemiological and health studies but requires a robust record linkage approach. We describe a hybrid record linkage process to link databases from two independent ongoing national studies, GEMO (Genetic Modifiers of BRCA1 and BRCA2 ), which focuses on the identification of genetic factors modifying cancer risk of BRCA1 and BRCA2 mutation carriers, and GENEPSO (prospective cohort of BRCAx mutation carriers), which focuses on environmental and lifestyle risk factors. Methods To identify the maximum individuals participating in the two studies but may not be registered by a common number, we combined Probabilistic Record Linkage (PRL) and supervised Machine Learning (ML). This combined linkage was named “PRL+ML”. We built the ML model using a first version of the two databases as a training dataset on which matching status was assigned by PRL followed manual review. Results The Random Forest (RF) algorithm showed a highest sensitivity (0.985) among six widely used ML algorithms: RF, Bagged trees, AdaBoost, Support Vector Machine, Neural Network. Therefore, RF was selected to build the ML model since our goal was to identify the maximum of true matches. Our combined linkage PRL+ML showed a higher sensitivity (range 0.988-0.992) than either PRL (range 0.916-0.991) or ML (0.981) alone. It identified 2,068 individuals participating in both GEMO (6,375 participants) and GENEPSO (4,925 participants). Conclusions Our hybrid linkage process represents an efficient tool for linking GEMO and GENEPSO. It may be generalizable to other epidemiological studies involving other databases and registries. Molecular Genetics Record linkage hybrid process probabilistic linkage supervised machine learning human-in-the-loop Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Full Text Supplementary Files SupplementaryData20200807.docx Cite Share Download PDF Status: Published Journal Publication published 29 Jul, 2021 Read the published version in BMC Medical Research Methodology → Version 1 posted Editorial decision: Major revision 05 Jan, 2021 Review # 2 received at journal 27 Dec, 2020 Review # 3 received at journal 17 Dec, 2020 Reviewer # 3 agreed at journal 26 Nov, 2020 Reviewer # 2 agreed at journal 23 Nov, 2020 Review # 1 received at journal 20 Oct, 2020 Reviewers invited by journal 29 Sep, 2020 Reviewer # 1 agreed at journal 29 Sep, 2020 Editor assigned by journal 10 Sep, 2020 First submitted to journal 09 Sep, 2020 Submission checks completed at journal 09 Sep, 2020 Editor invited by journal 09 Sep, 2020 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-64751","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":2289162,"identity":"72b73a05-058a-40bb-adf4-e9d184659e60","order_by":0,"name":"YUE JIAO","email":"","orcid":"https://orcid.org/0000-0001-9872-6034","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"YUE","middleName":"","lastName":"JIAO","suffix":""},{"id":2289163,"identity":"81d9a6eb-838a-4903-b40e-a3113410f6fe","order_by":1,"name":"Fabienne Lesueur","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Fabienne","middleName":"","lastName":"Lesueur","suffix":""},{"id":2289164,"identity":"d44e0a3b-cffe-4046-926d-d878144b5675","order_by":2,"name":"Chloé-Agathe Azencott","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Chloé-Agathe","middleName":"","lastName":"Azencott","suffix":""},{"id":2289165,"identity":"4c965cd1-fb7d-4a86-9d7c-7cc2f12de53d","order_by":3,"name":"Maïté Laurent","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Maïté","middleName":"","lastName":"Laurent","suffix":""},{"id":2289166,"identity":"80290892-2fd9-4894-b7c7-f1cc35f7141b","order_by":4,"name":"Noura Mebirouk","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Noura","middleName":"","lastName":"Mebirouk","suffix":""},{"id":2289167,"identity":"cb116f0b-8c8c-4fd2-a63b-6839c7825697","order_by":5,"name":"Lilian Laborde","email":"","orcid":"","institution":"Institut Paoli-Calmettes","correspondingAuthor":false,"prefix":"","firstName":"Lilian","middleName":"","lastName":"Laborde","suffix":""},{"id":2289168,"identity":"b81565da-d764-4061-879e-2d08b8ee443f","order_by":6,"name":"Juana Beauvallet","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Juana","middleName":"","lastName":"Beauvallet","suffix":""},{"id":2289169,"identity":"ca786d3a-020a-435f-b524-c82b289f15d4","order_by":7,"name":"Marie-Gabrielle Dondon","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Marie-Gabrielle","middleName":"","lastName":"Dondon","suffix":""},{"id":2289170,"identity":"67a62c87-8048-42fe-acf0-500fa3fdc365","order_by":8,"name":"Séverine Eon-Marchais","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Séverine","middleName":"","lastName":"Eon-Marchais","suffix":""},{"id":2289171,"identity":"45579e5a-4832-46f1-aa3c-adf91d84b571","order_by":9,"name":"Anthony Laugé","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Anthony","middleName":"","lastName":"Laugé","suffix":""},{"id":2289172,"identity":"fbd207f2-99f8-4259-b517-5db66d8919c1","order_by":10,"name":"GEMO Study Collaborators","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"GEMO","middleName":"Study","lastName":"Collaborators","suffix":""},{"id":2289173,"identity":"ece55ea5-15d2-49b3-9ec6-dd0044783fa9","order_by":11,"name":"GENEPSO Study Collaborators","email":"","orcid":"","institution":"Institut Paoli-Calmettes","correspondingAuthor":false,"prefix":"","firstName":"GENEPSO","middleName":"Study","lastName":"Collaborators","suffix":""},{"id":2289174,"identity":"2e9698a4-3297-47ae-95ac-2c710e83de78","order_by":12,"name":"Catherine Noguès","email":"","orcid":"","institution":"Institut Paoli-Calmettes","correspondingAuthor":false,"prefix":"","firstName":"Catherine","middleName":"","lastName":"Noguès","suffix":""},{"id":2289175,"identity":"5ef5b8a3-3682-4192-9bc1-5c62335cc68a","order_by":13,"name":"Nadine Andrieu","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Nadine","middleName":"","lastName":"Andrieu","suffix":""},{"id":2289176,"identity":"fa9694ce-e10c-4c68-81d2-f6af6291e80c","order_by":14,"name":"Dominique Stoppa-Lyonnet","email":"","orcid":"","institution":"Institut Curie","correspondingAuthor":false,"prefix":"","firstName":"Dominique","middleName":"","lastName":"Stoppa-Lyonnet","suffix":""},{"id":2289177,"identity":"a3cb4fe8-cced-44f6-baed-3a1f8f2ed7e7","order_by":15,"name":"Sandrine M. Caputo","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAUlEQVRIie3PMUvDQBjG8QcEs5x0PRGar/CGg0ogH+aKcFkSdbNbb9Kl0DWDH8LJOeEGl4prRAdFyJShk3TIoElbi9CLq8P9OXLHwY/3Arhc/zBfAxy0uxi2n3yz743yX0RCbIiwE7TkJ4mxXp96iPdYvEwuX4fwTPVxtIrieZa+50vQhfVfZudn4YIqAaZOBZMqzcqYigwUapvJk9GxJjPWHKMTJk16xxUMQ0M2gae6I1PNvc+WxNSSBmQn5XqKBGfdFNkR9BAqaxF+k+CaJVfBrVJBtqhQzMhO/HkSPOvG+APv4f6tjiJ/cKMOlqtJz8O2HXZrN/1PsFUul8vl2tMXp+hQNzbHqeQAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-5338-9388","institution":"Institut Curie","correspondingAuthor":true,"prefix":"","firstName":"Sandrine","middleName":"M.","lastName":"Caputo","suffix":""}],"badges":[],"createdAt":"2020-08-24 10:29:48","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-64751/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-64751/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12874-021-01299-6","type":"published","date":"2021-07-29T15:02:19+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":2433106,"identity":"50f65633-6f72-40d9-b06e-a5086bdcce25","added_by":"auto","created_at":"2020-09-16 14:19:22","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":33560,"visible":true,"origin":"","legend":"Elaboration of hybrid record linkage process and main steps. (a) Assignment the matching status by PRL followed manual review. We manual reviewed all the pairs whose scores exceed the threshold and also the pairs whose scores below and near the threshold, in order to minimize the FP and FN in PRL. (b) Selection of the supervised machine learning algorithm. (c) Selection of the combined linkage method PRL+RF. (d) Building of RF model on initial databases. (e) Application of PRL+RF combined linkage method to classify the updated data.","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/Fig1.png"},{"id":2433107,"identity":"faf7497e-bc46-40ea-9a19-778999aa245e","added_by":"auto","created_at":"2020-09-16 14:19:22","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":123066,"visible":true,"origin":"","legend":"Score distribution of 15,653,232 record pairs in dataset 1. (a) Whole score distribution. (b) Zoom on the distribution for the highest scores.","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/Fig2.png"},{"id":2433108,"identity":"a2371de6-1b49-438a-806b-20b0faa4b9e2","added_by":"auto","created_at":"2020-09-16 14:19:22","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":33442,"visible":true,"origin":"","legend":"Performance of three linkage methods: PRL (Probabilistic Record Linkage), RF (Random Forest) and PRL+RF. PRL has thresholds varying from 0.6 to 0.8. (a) Comparison of their sensitivities. (b) Comparison of their precisions.","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/Fig3.png"},{"id":2433109,"identity":"7e72511d-70ec-4e48-9d4e-dcb3581bbf0e","added_by":"auto","created_at":"2020-09-16 14:19:22","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":28749,"visible":true,"origin":"","legend":"Comparison of candidate matches predicted by the RF and PRL models for the updated databases. (a) Before manual review, RF model and PRL predicted 819 and 1,268 new candidate matches, respectively; 772 candidate matches were found by both approaches. (b) After manual review, PRL+RF identified 738 true matches, in which 727 true matches were identified by PRL and 715 true matches were identified by RF. 704 true matches were identified by both approaches. 23 true matches were identified only by PRL, and 11 true matches were identified only by the RF model.","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/Fig4.png"},{"id":2433110,"identity":"460a8fd6-b327-41ec-980d-3362a35ecf14","added_by":"auto","created_at":"2020-09-16 14:19:23","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":24270,"visible":true,"origin":"","legend":"General overview of the hybrid record linkage process. (a) Probabilistic record linkage (PRL) followed by a stage of manual review is first applied to build a machine learning (ML) model. (b) The PRL+ML combined linkage is then used to classify the updated datasets (Record pair comparison from Database X’ and Database Y’). The ML model obtained in (a) is used (dotted arrow) for the prediction in (b).","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/Fig5.png"},{"id":13530816,"identity":"ac80dd4e-dab7-432f-9948-c3881b8242ed","added_by":"auto","created_at":"2021-09-17 01:10:19","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":797471,"visible":true,"origin":"","legend":"","description":"","filename":"CANSOP20200807mainnofig.pdf","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1_covered.pdf"},{"id":2433113,"identity":"a6a551cc-f9cd-4370-968d-2c9313b256d9","added_by":"auto","created_at":"2020-09-16 14:19:26","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":873456,"visible":true,"origin":"","legend":"","description":"","filename":"CANSOP20200807mainnofig.pdf","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1_stamped.pdf"},{"id":2433111,"identity":"e50383da-c693-4581-a795-8b3cbdd560b2","added_by":"auto","created_at":"2020-09-16 14:19:23","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":798820,"visible":true,"origin":"","legend":"","description":"","filename":"CANSOP20200807mainnofig.pdf","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/CANSOP20200807mainnofig.pdf"},{"id":2433112,"identity":"59a585ba-9012-4672-b0bd-c03c7aca9284","added_by":"auto","created_at":"2020-09-16 14:19:23","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":57239,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryData20200807.docx","url":"https://assets-eu.researchsquare.com/files/rs-64751/v1/SupplementaryData20200807.docx"}],"financialInterests":"","formattedTitle":"A new hybrid record linkage process to render epidemiological databases interoperable: application to the GEMO and GENEPSO studies involving BRCA1 and BRCA2 mutation carriers","fulltext":[{"header":"Full Text","content":"\u003cp\u003eThis preprint is available for \u003ca href='/article/rs-64751/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Record linkage, hybrid process, probabilistic linkage, supervised machine learning, human-in-the-loop","lastPublishedDoi":"10.21203/rs.3.rs-64751/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-64751/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBackground\u003c/p\u003e\u003cp\u003eLinking independent sources of data related to same individuals enable innovative epidemiological and health studies but requires a robust record linkage approach. We describe a hybrid record linkage process to link databases from two independent ongoing national studies, GEMO (Genetic Modifiers of \u003cem\u003eBRCA1\u003c/em\u003e and \u003cem\u003eBRCA2\u003c/em\u003e), which focuses on the identification of genetic factors modifying cancer risk of \u003cem\u003eBRCA1\u003c/em\u003e and\u003cem\u003e BRCA2\u003c/em\u003e mutation carriers, and GENEPSO (prospective cohort of \u003cem\u003eBRCAx \u003c/em\u003emutation carriers), which focuses on environmental and lifestyle risk factors.\u003c/p\u003e\u003cp\u003eMethods\u003c/p\u003e\u003cp\u003eTo identify the maximum individuals participating in the two studies but may not be registered by a common number, we combined Probabilistic Record Linkage (PRL) and supervised Machine Learning (ML). This combined linkage was named “PRL+ML”. We built the ML model using a first version of the two databases as a training dataset on which matching status was assigned by PRL followed manual review. \u003c/p\u003e\u003cp\u003eResults\u003c/p\u003e\u003cp\u003eThe Random Forest (RF) algorithm showed a highest sensitivity (0.985) among six widely used ML algorithms: RF, Bagged trees, AdaBoost, Support Vector Machine, Neural Network.\u003c/p\u003e\u003cp\u003eTherefore, RF was selected to build the ML model since our goal was to identify the maximum of true matches. Our combined linkage PRL+ML showed a higher sensitivity (range 0.988-0.992) than either PRL (range 0.916-0.991) or ML (0.981) alone. It identified 2,068 individuals participating in both GEMO (6,375 participants) and GENEPSO (4,925 participants).\u003c/p\u003e\u003cp\u003eConclusions\u003c/p\u003e\u003cp\u003eOur hybrid linkage process represents an efficient tool for linking GEMO and GENEPSO. It may be generalizable to other epidemiological studies involving other databases and registries.\u003c/p\u003e","manuscriptTitle":"A new hybrid record linkage process to render epidemiological databases interoperable: application to the GEMO and GENEPSO studies involving BRCA1 and BRCA2 mutation carriers","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-09-16 14:19:21","doi":"10.21203/rs.3.rs-64751/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2021-01-06T00:00:00+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2020-12-28T00:00:00+00:00","index":2,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nJiao et al proposed a new hybrid record linkage process with an initial probabilistic linkage with manual review generating labelled data and a machine learning algorithm based on the labelled data. This hybrid process was found to identify more matches than either algorithm alone. The hybrid record linkage is an interesting idea which could be quite useful in many applications. However, I have a major concern about the data sets used in the linkage process. The two databases GEMO and GENEPSO were independent databases set up at different time by different coordinating centers and investigators. Family number and individual number were used in the matching process - but if the studies are independent, wouldn't assignment of family numbers and subject numbers independent as well and hence can't be used for matching? In addition, participants' names and addresses were not available. With only BRCA status, date of birth, gender, and consultation center, I am not sure participants can be reliably identified. Other comments are as follows:\n1. The process of the hybrid matching is not very clear. Specifically, figure 1 showing the process is quite confusing. It was said that the labelled data (all record pairs identified by blocking?) were split into training and testing data, while the 5-fold cross validation was performed using the training data to tune the parameters of the ML algorithms. But in Table 2 when the ML algorithms were tested, mean and SD of sensitivity and precision were reported: it was unclear how the averages were calculated. It was also mentioned that 5-fold cross-validation was performed for the hybrid process. What about the 40% testing data?\n2. The ML algorithms were applied to the labeled data that were identified by blocking. How many pairs were there and how many were matches? It is essential that there were sufficient number of record pairs so that ML algorithms can be used so please show the sample size. If these were split into training and testing data, please show the sample size in each data.\n3. In the process, there were two rounds of manual review that appeared to serve different purposes. The first round happened after PRL with the purpose to establish the training data. It was said that all pairs classified as matches by PRL plus some pairs with score around the threshold were manually reviewed. Please show the exact number of pairs being reviewed, as well as how \"around the threshold\" was determined. The second round happened after linkage process was completed, with all record pairs classified as matches by linkage algorithms manually reviewed. The purpose of the second round of review appeared to be for evaluation of performance of the linkage process. If this was correct, please clarify as it was quite confusing.\n4. Please also clarify that PRL+ML identifies record pairs as matches if they were identified by either PRL or ML.\n5. The threshold of 0.6 selected by the authors for PRL+ML seemed to give very similar performance as PRL alone. That was why RF only identified 47 additional pairs (figure 4). If a higher threshold is used, then many more pairs will be missed by PRL and picked up by ML. So the statement \"PRL+ML combined method …can improve the linkage by identifying as many true matches as possible without paying more additional manual review than PRL\" (page 14 line 51) is not true. The precision is 727/1268=0.5733 for PRL and 738/1315=0.5612 for PRL+ML. They were very close and both were very poor due to the low threshold of 0.6. It appeared a higher threshold would make more sense with a slightly reduced sensitivity and much better precision.\n6. The two databases used in the linkage process were both quite small. In many applications, very large databases are involved, resulting in a large number of matches. It would be impossible to manually review all these record pairs plus some more pairs with score around the threshold. Please discuss this limitation and provide insights on how the proposed hybrid process can be modified.\n\n\n* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons upon publication of the manuscript. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **No**\n* Declaration of competing interests: **I declare that I have no competing interests**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please do not publish my name with my report. (default)**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **No**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **No**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"editorInvitedReview","content":"","date":"2020-12-18T00:00:00+00:00","index":3,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nThank you for the opportunity to review the manuscript titled \"A new hybrid record linkage process to render epidemiological databases interoperable: application to the GEMO and GENEPSO studies involving BRCA1 and BRCA2 mutation carriers\". In this manuscript, the authors introduce a hybrid record linkage method that combines a probabilistic record linkage and a machine-learning-based method. The linkage result of probabilistic methods, together with manual reviews was used as the gold standard to train the ML method. The results show that the hybrid method outperformed the standalone methods. Record linkage research is essential to enhance the quality and completeness of biomedical data. Constant efforts to improve the performance RL methods are important to improve the accuracy of linkage results. Therefore, this work is highly relevant. The study was rigorously designed. The writing and diagrams are straightforward. Here are my comments\n\n- It is not clear to me the role of the manual review step in this process. Is it part of the linkage methods, or is it a step to generate the gold standard for performance evaluation? It is suggested in the writing that it is part of the method, but I'm not sure.\n- The size of the datasets used to develop these methods is very small. Granted that the findings of this study may be applicable to other scenarios in which linkage datasets are small, I have serious doubt about the scalability of these. For example, PRL without blocking variable is certain not to work with large datasets.\n- The description of the linkage methods is very inadequate. More details on how the PRL, ML, manual review, and data manipulation methods are needed for the paper to be helpful. Data Imputation in record linkage must be done carefully, especially on variables with categorical values. As a reader, I didn't learn much from this process at all. Statements like \"The likelihood score threshold of PRL was designated low enough\" or \"We then manually reviewed all pairs whose scores exceed the threshold and the pairs whose scores below and near the threshold\" are too generic to be useful.\n- Deterministic methods can be much more than exact matching. The current definition of deterministic linkage equates to simplistic exact linkages, which is too narrow.\n- In the Data section, please describe the shared linkage variables in greater detail for international audiences. For example, what is an \"individual number\"? Is it a national identifier in France?\n- For audiences who are not a genome expert, like myself, the authors need to explain why picking BRAC1 and BRAC2 as blocking variables is appropriate in this case.\nOther comments:\n- Change the term \"common number\" to \"shared identifier\".\n- While this is a record linkage-focused paper, the contribution of RL methods should be mentioned earlier in the Background section.\n- Consider changing the term sensitivity to recall as precision and recall usually go together.\n* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons upon publication of the manuscript. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **Yes**\n* Declaration of competing interests: **I declare that I have no competing interests**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please do not publish my name with my report. (default)**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **No**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **Yes**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"reviewerAgreed","content":"","date":"2020-11-27T00:00:00+00:00","index":3,"fulltext":""},{"type":"reviewerAgreed","content":"","date":"2020-11-24T00:00:00+00:00","index":2,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2020-10-20T12:00:00+00:00","index":1,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nThis is a fascinating work that tackles an important problem - how we can accurately link records across multiple data sets. Setting aside the ethical question of whether such an effort conflicts with the principles of informed consent patients theoretically provided when agreeing to participate in these registries, the methodology proposed by the authors allows for the completion of multiple analyses that would not be possible without the linkage. There is no doubt that this is an impactful proof of concept.\n\nThat said, I have some issues with the paper and cannot recommend it in its current form. These points are divided into two categories - larger/conceptual issues and lower-level grammar/usage quibbles.\n\n-It is relatively unclear to the reader until midway through the paper how the authors obtained the gold standard of true matches used to generate the confusion matrix and assess performance. Do participants in GEMO and GENEPSO provide a singular identifier that allowed the researchers to assess the performance of their linkage methods? What is this identifier? How reliable is it? On line 54 of page 3 we learn that the individuals may not be registered by a common number. Then, on page 7, line 46, we learn that both data sets contain the participant's variant, gender, center number, family number, subject identifier, and birth date. However, it is still unclear whether the SUJID is persistent across both data sets and serves as an authoritative source for calculating performance. If it is, the linkage method is unnecessary and the advantages proposed in the introduction - e.g. predicting response to treatment according to mutation status and variant type - are already available to researchers without use of the PRL-ML hybrid technique. Presumably given the page 3 line 54 reference it is not - but if it is not, it is still unclear to the reader how performance was calculated! This is made even worse by the SUJID in the example in Table 1.\n\n-The use of Jaro-Winkler for calculating similarity of HGVS notation seems questionable to me. Why allow for fuzziness in the variant notation, where a difference of one character can yield a completely different variant with a different effect, but insist on exact binary similarity for other variables potentially subject to data entry error, such as birth day? Especially puzzling given how many different ways there are in HGVS to annotate the same variant.\n\n-Again, for the ML component of the methods, it's unclear what was used as the gold standard for the test data set - was it the objective truth (e.g. matching SUJID) or are we assuming that the PRL-derived matches are the gold (silver?) standard?\n\n-The authors emphasize the strengths and weaknesses of each component of their process and highlight that different use cases may favor differing levels of sensitivity/specificity - this is commendable\n\nMinor comments:\n\n-Errors in grammar and usage throughout. Consider a copy editor. Some examples follow:\n\n-Second sentence of first paragraph of \"Background\" is a runon\n-Page 4 line 36: \"will allow to better understand the cancer risk for people...\" Allow for whom?\n-Page 4 line 49: sentence starting \"Record linkage is the process...\" is grammatically incorrect and confusing.\n-Page 5 line 1: \"computational expensive\" should read \"computationally expensive\"\n-Page 5 line 5: \"blocking that splits the datasets into blocks\" - as opposed to what other kind of blocking? This definition is tautological.\n-Page 5 line 15: sentence starting \"The records comparison involves...\" is a runon and confusing\n-Many more throughout - please consider a proofreader\n\n-On page 5, final paragraph, you state that if data have more than 5% missing or incorrect in any matching variable, deterministic record linkage will produce a large number of FP. Wouldn't this be just as likely - or more likely - to produce FNs? If 5% of the single unique matching variable is missing this by definition can't create more FPs, only FNs!\n\n-\n* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons upon publication of the manuscript. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **No**\n* Declaration of competing interests: **I declare that I have no competing interests.**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please do not publish my name with my report. (default)**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **Yes**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **No**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"reviewersInvited","content":"","date":"2020-09-29T12:00:00+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2020-09-29T12:00:00+00:00","index":1,"fulltext":""},{"type":"editorAssigned","content":"","date":"2020-09-10T12:00:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"","date":"2020-09-09T12:00:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2020-09-09T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2020-09-09T12:00:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8a0c337d-3f38-4ba1-bb51-4d73a2f8ba3d","owner":[],"postedDate":"September 16th, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":525479,"name":"Molecular Genetics"}],"tags":[],"updatedAt":"2021-08-22T15:13:05+00:00","versionOfRecord":{"articleIdentity":"rs-64751","link":"https://doi.org/10.1186/s12874-021-01299-6","journal":{"identity":"bmc-medical-research-methodology","isVorOnly":false,"title":"BMC Medical Research Methodology"},"publishedOn":"2021-07-29 15:02:19","publishedOnDateReadable":"July 29th, 2021"},"versionCreatedAt":"2020-09-16 14:19:21","video":"","vorDoi":"10.1186/s12874-021-01299-6","vorDoiUrl":"https://doi.org/10.1186/s12874-021-01299-6","workflowStages":[]},"version":"v1","identity":"rs-64751","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-64751","identity":"rs-64751","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.