{"paper_id":"07712b98-a17e-44de-aa0b-f59c13676b80","body_text":"1 \n \nMaternal and Infant Research Electronic Data Analysis (MIREDA): A protocol for creating a common data model for \nfederated analysis of UK birth cohorts and the life course.  \n \nMJ Seaborne1 4 5, HE Jones1 5, N Cockburn14, S Durbaba7 13, TC Giles10, A González-Izquierdo14, A Hough9, D Mason9, A Mendez-Villalon10,  \nC. Sanchez-Soriano6, C. Orton16, D Ford16, P Quinlan10 11, K Nirantharakumar*14 15, L. Poston*7 13, RM Reynolds*6, G Santorelli*9, S Brophy1 2 3 5 8 \n*Senior authors with equal contribution \nAffiliations: \n1 National Centre for Population Health and Wellbeing Research, Swansea University Medical School. \n2 Health Data Research UK. \n3 Administrative Data Research Wales, Swansea University Medical School, Swansea, Wales. \n4 Data Lab, National Centre for Population Health and Wellbeing Research, Swansea University. \n5 Born in Wales, National Centre for Population Health and Wellbeing Research, Swansea University. \n6 Born in Scotland, Centre for Cardiovascular Science, University of Edinburgh, Edinburgh. \n7 eLIXIR, Born in South London, King’s College London, Strand, London. \n8 Health and Care Research Wales, Cardiff, Wales. \n9 Born in Bradford, Bradford Institute for Health Research, Bradford Royal Infirmary, Bradford. \n10 Digital Research Service, University of Nottingham, Nottingham. \n11 School of Medicine, University of Nottingham, Nottingham. \n12 Department of Women and Children’s Health, King’s College London, London. \n13 King’s College London, School of Life Course & Population Sciences, London. \n14 Institute of Applied Health Research, University of Birmingham, Birmingham. \n15 Mum-PreDiCT, UK. \n16 Health Data Science, Swansea University, Swansea, Wales. \n \n \nAbstract: \nIntroduction \nBirth cohorts are valuable resources for studying early life, the determinants of health, disease, and development. They are \nessential for studying life course. Electronic cohorts are live, dynamic longitudinal cohorts using anonymised, routinely collected \ndata. There is no selection bias through direct recruitment, but they are limited to health and administrative system data and may \nlack contextual information. \nThe MIREDA (Maternal and Infant Research Electronic Data Analysis) partnership creates a UK-wide birth cohort by aligning \nexisting electronic birth cohorts to have the same structure, content, and vocabularies, enabling UK-wide federated analyses. \nObjectives \n1) Create a core dynamic, live UK-wide electronic birth cohort with approximately 100,000 new births per year using a common \ndata model (CDM). \n2) Provide data linkage and automation for long-term follow up of births from MuM-PreDiCT and the ‘Born in’ initiatives of Bradford, \nWales, Scotland, and South London for comparable analyses. \nMethods \nWe will establish core data content and collate linkable data. Use a suite of extraction, transformation, and load (ETL) tools will be \nused to transform the data for each birth cohort into the CDM. Transformed datasets will remain within each cohort’s trusted \nresearch environment (TRE). Metadata will be uploaded for the public to the Health Data Research (HDRUK) Innovation Gateway. \nWe will develop a single online data access request for researchers. A cohort profile will be developed for researchers to reference \nthe resource. \nEthics \nEach cohort has approval from their TRE through compliance with their project application processes and information governance.  \nDissemination \nWe will engage with researchers in the field to promote our resource through partnership networking, publication, research \ncollaborations, conferences, social media, and marketing communications strategies. \n \nKeywords   \nBirth Cohort, Life Course Perspective, Data Science, Data Curation, Routinely Collected Health Data, Electronic Health Records, \nUnified Medical Language System.\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \nNOTE: This preprint reports new research that has not been certified by peer review and should not be used to guide clinical practice.\n\n2 \n \nIntroduction \nRapid socio-economic changes including the recent increased cost of living in the UK are contributing to \nwidening inequalities which will impact population health (1). This is coupled with the increased burden to the \nNHS, social care services and education disruption after the pandemic (2). These changes are likely to adversely \naffect maternal, infant and child health outcomes as there are strong negative associations with deprivation and \npoor diet (3), unhealthy maternal BMI (4,5), shorter intervals between pregnancies (6), pregnancy in adolescence \n(7), negative health behaviours (8), and conditions such as anaemia (9). These alone can be responsible for \ncausal sequences which may result in deteriorating maternal, infant and child health and wellbeing, and can be \ndamaging to development and future health and prosperity of offspring through their life course (10). There is a \nclear need for continued monitoring and evaluation of interventions for preventative strategies that reduce the \nrisks of poor health and wellbeing in these population groups throughout their life course. \nAlthough new technologies and treatments continue to improve population health, there is a greater need for \ninformed preventative action that reduces the burden on healthcare resources (11). Measures to optimise \nmaternal, infant and child health during the perinatal period and beyond will inevitably confer substantial benefits \nin the prevention of ill-health.  \nTraditionally, data from birth cohorts have been valuable resources for studying early life and the determinants of \nhealth, disease and development (12). They have played an essential role in profiling the life course of a child \nfrom conception through to all other stages of development. \nLarge birth cohorts have historically contributed to life course approaches to disease prevention and public \nhealth, they are often through direct recruitment, and are usually specific to a country, region, or other measure \nof locality. Directly recruited cohorts can limit representation of the target population because of selective \ninclusion. They are often limited by the period for which the cohort is defined and may not represent those born \noutside the cohort recruitment window (12–14).  \nLife course studies from birth are invariably complex as there are multiple and often interacting factors to \nconsider which may influence outcomes. As such, the type and content of data collected is extensive and may \ninclude maternal characteristics (age, parity, gravida, ethnicity, deprivation scores, medical history, etc), \npregnancy complications - both physical and mental health, birth outcomes (gestational age, birth weight, level of \ncare required, etc), and other longitudinal data which relates to social care, developmental milestones, \neducation, household composition, etc. \nWe are living in a new era with increasing access to vast amounts of health data and with growing analytical \nexpertise, to interpret and disseminate meaningful insights to the health of our nations. This provides an exciting \nnew capacity to further our research into the life course. Much work was done during the COVID-19 pandemic to \nbring together large amounts of health data to monitor its effects, but significant issues emerged when comparing \nroutinely collected data across different nations and regions within the UK. These were compounded by a lack of \nstandardisation of clinical coding systems, terminology, definitions, formats, difference in type and/or quantity of \ndata collected, and governance relating to access to data (15). These differences were a barrier to accessibility \nof easily comparable data. \nThe evolution of data science has led to the development of accessible routinely collected anonymised data for \nlarge populations in trusted research environments (TREs) that can be linked to many different data sources. \nThey provide rich, detailed information and remove obstacles such as the need to recruit, and so provide a \nmeans for which a cohort representative of that population can grow over time and the data remains relevant in \nthe future. \nIn the UK, there is currently no central repository to access birth data for all four nations which is suitably \nanonymised and capable of linking to the plethora of data sources that help define the life course of each birth. \nNational core datasets exist, but they are used for more general statistical output and without the granularity of \ndata available through data linkage. There are several electronic birth cohorts which use different approaches to \ncompiling routinely collected data for analysing population level data and the health life course. Each has similar \ndata but with additional linkable datasets which lend them different strengths (16–20). They have arisen, in \nparallel, with improved relative ease of access, the removal of some research barriers and the need for analyses \nwithin specific populations which cannot necessarily be extrapolated from other, dissimilar populations. \nHarmonisation of these cohorts is a goal that aligns with the FAIR Guiding Principles for scientific data \nmanagement and stewardship (21). The principles aim to establish comparability and consistency of analyses \nthrough a widely accepted standard. Achieving this would result in a network of cohorts gathering uniform core \ndata, augmented with additional datasets that could be linked to extend the breadth and depth of research \npossibilities. This harmonisation effort would not only improve the findability and accessibility of birth data across \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n3 \n \nthe four nations but also enhance the interoperability and reusability of the data for diverse research needs, thus \nfulfilling the FAIR criteria and significantly advancing the field of life course health research. \nUsing the established standards of the Observational Medical Outcomes Partnership (OMOP) Common Data \nModel (CDM) could help achieve all this. This approach is used to standardise both the structure and content \n(terminologies, vocabularies, and coding schemes) of observational data and has been used successfully in \nother health data analyses to produce reliable evidence and enable comparisons. The OMOP CDM contains 37 \nstandardised tables relating to: clinical data and vocabularies. It also contains additional domains relating to \nhealth systems, health economics, derived elements, and metadata (16). \nThe Mother and Infant Research Electronic Data Analysis (MIREDA) partnership will facilitate a platform in which \nUK wide, dynamic live birth cohorts will be collated into common core datasets using the OMOP CDM. The core \ndata resource will initially contain the harmonised data for 1,000,000 live births and will grow by approximately \n100,000 births per year, with long-term follow up using routine data. It will contain each respective cohort’s \nlinkable data from public health, neonatal health, imaging, primary and secondary care and will be a central \naccess point for longitudinal birth cohort federated analyses; federated analyses with the purpose of improving \ninfant and maternal health across the UK by comparing differences in policies and practises.  \nThe cohort will be facilitated by Health Data Research UK (HDR-UK) and access will be governed by their TRE \nprinciples and best practices (22). Accredited researchers wishing to undertake research with the harmonised \nOMOP cohorts will apply for access using a single, centrally approved HDR-UK application. \nThe metadata for the MIREDA resource will be prepared and made accessible from the Health Data Research \n(HDRUK) Innovation Gateway using a standardised request form. \nThe anonymised birth cohorts included are: Born in Wales (BiW) (17), Born in Scotland (BiS) (20), Born in \nBradford (BiBBS and BiB4ALL) (23–25), the early Life data Cross-Linkage in Research (eLIXIR) partnership \n(Born in South London) (18), and Multimorbidity in Pregnancy: Determinant, Consequences, Clusters and \nTrajectories (MuM-PreDiCT) (19). \nAdditional linkable data, where available for each cohort will be provided and used to enrich datasets beyond \ntheir core content. \n \nThis phase of the project will run for one year from June 2023. \n \nMethods \nStudy design \nThe MIREDA partnership will bring together standardised, harmonised birth cohorts relating to mother, baby, and \nchild in a federated common data framework. This will include structured definitions, descriptions and \ndocumentation of the data using the OMOP common data model (CDM), a comprehensive data dictionary, a \ndata quality plan, and data governance policy. The data will be available for creating a sustainable and reactive \nhealth system to inform public health decisions and policy. \nThe project unites the cohorts from BiW, eLIXIR (BiSL), BiBBS/BiB4ALL, BiS, and MuM-PreDiCT, encompassing \ndata from all four UK nations. The initial combined cohort will represent approximately 350,000 live births from \nthe first four cohorts. The share from each of these cohorts is apportioned as 82%, 13%, 4%, and 1% \nrespectively. MuM-PreDiCT will represent approximately eight million births across the UK. \nWe estimate that the combined cohorts will amass an additional 100,000 births per year. \nAll live births on or after January 1st, 2014, will be identified from routine data for BiW and represent the earliest \nbirths in the collection. Others will contribute birth data from the start of their respective cohorts. \nA suite of extraction, transformation and load (ETL) tools (WhiteRabbit (26), Convenient and Reusable Rapid \nOMOP Transformer (CaRROT) Mapper (27), and CaRROT-CDM (28) will be used to restructure concepts into \nthe OMOP CDM of standard tables, fields and vocabularies. \nData storage \nEach electronic birth cohort with person-level data including outcomes and exposures will remain inside their own \ntrusted research environment (TRE) and will be subject to all the established rules of their TREs. \nCore data: Cohort data alignment \nThe MIREDA partnership will collaborate to review the data accessible to each birth cohort and establish a \nconsensus for their common data. This will be used to form a mutually agreed core dataset to be implemented \nfor each birth cohort (\n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n4 \n \nTable 1). \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n5 \n \nTable 1. Core data for birth cohort dataset modules. \nDataset modules Data Additional information \nMaternal socio-\ndemographics \nWeek/Date of birth   \nEthnicity   \nDeprivation Quintile By locality at start of pregnancy \nRural/urban By locality at start of pregnancy \nEducation   \nQualifications   \nOccupation   \nBaby socio-\ndemographics \nWeek/Date of birth   \nSex   \nEthnicity   \nDeprivation quintile By locality at birth and school age \nRural/urban residence By locality at birth and school age \nDisability codes   \nPregnancy Date of antenatal visit   \nMaternal age At booking visit/conception \nMaternal smoking status   \nMaternal Height   \nMaternal Weight   \nMaternal BMI   \nExisting diagnosis codes Comorbidities existing at booking visit \nNew event type New events since previous visit e.g. diagnosis, \nprocedure, medicines, etc. \nNew event code  \nNew event date  \nNew event source Event sources e.g. GP/hospital. \nData of expected Delivery   \nFolic acid supplements indicator   \nDomestic violence indicator   \nWHOOLEY score Or other depression screening indicators. \nGravidity   \nParity   \nPrevious C-Section indicator   \nDate of previous C-section   \nPrevious preterm birth indicator   \nDate of previous preterm birth   \nGestational age At visit. \n \nDataset modules Data Additional information \nPost-birth \n(maternal) \nPostnatal depression indicator   \nEvent type Events within 10 weeks of birth date; may \ninclude: diagnosis, procedure, medication, etc. \nEvent code  \nEvent date  \nEvent source Event sources e.g. GP/hospital. \nBirth Date of delivery   \nGestational age At birth. \nBirth weight   \nEvent type Events from birth to 28 days; may include: \ndiagnosis, procedure, medication, etc. \nEvent code  \nEvent date  \nEvent source Event sources e.g. GP/hospital. \nMode of delivery E.g. c-section, forceps, vaginal, etc. \nInfancy Event type As above but for events from 29 days to school \nage \nEvent code  \nEvent date  \nEvent source  \nNeonatal admission indicator   \nNeonatal event indicator   \nBreastfeeding indicator Breastfeeding, mixed or formula indicator \nBreastfeeding duration   \nSchool age Event type As above but for events from school start age to \n12 years. \nEvent code  \nEvent date  \nEvent source  \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n6 \n \nDetails of linkable data for each cohort will also be provided in later work Table 2. \n \nTable 2. Linkable datasets. \n Born in \nWales \nElixir: Born in \nSouth London  Born in Scotland \nBorn in Bradford  \nCPRD  \nBiBBS BiB4All \nCore data \nPregnancy, birth, \ninfancy. Yes Yes Yes Yes Yes Yes \nRoutine data \nHealth Visitor data Yes Will have Will have Yes Yes - \nEducation data Yes Yes Children too young but \nwill have in the future  Children too \nyoung at present - \nGP Yes Yes Will have Yes Yes Yes \nHospital admission Yes Yes Will have Yes Yes Yes \nA&E Yes  Will have (HES) Will have Yes Yes - \nCongenital \nabnormalities Yes Yes Will have Yes Yes  \nNational community \nchild health data \n(vaccination, breast \nfeeding) \nYes  \nHealth visitor \ndata for \nbreastfeeding. \nNIV data for \nvaccination - \napplied for \nWill have Yes Yes  \nSurvey with family \nPregnancy  Yes No  Yes Yes - \n18 months Yes No     \n3-5 years Yes No     \nPrimary school Yes No     \nBiological samples       \nBloods - Yes Yes Yes Yes  \nOther datasets  \nCensus 2011 Yes Could apply for \nthis    - \nCensus 2021 Yes As above    - \nLooked after \nchildren  Yes As above     \nChild protection Yes  As above     \nChild in receipt of \ncare and support Yes As above     \n \nThe agreed core datasets will be created within each cohort. They will include all live births from 2014 or earliest \nyear available after 2014. \nWhere available, core data will relate to both mother and baby with socio-demographic data such as date of \nbirth, sex of baby, maternal education/qualifications, ethnicity, deprivation level, relationship status, and \ngeographical location type (urban or rural). Other data will be divided into the following modules: pregnancy, \nmaternal ten weeks post-birth, birth, infancy, and school age (Table 1). \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n7 \n \nData extraction and transformation \nWhen data wrangling is complete and the core datasets are assembled, each site will use WhiteRabbit (24,25) \nsoftware to scan them and produce scan reports. The reports will be anonymised and contain detailed \ninformation about tables, fields, and values that appear in fields. The anonymous summary reports will be \nextracted from the TREs for mapping and to create transformation rules. \nThe transformation rules will be imported into each TRE as a .json file and will be used by the locally installed \nCaRROT-CDM tool to transform the core datasets into OMOP CDM.  \n \nData access \nData access will be granted by central request via the Health Data Research (HDRUK) Innovation Gateway. \nHere  the MIREDA collection will be available. It will detail the metadata for each cohort and their linkable \ndatasets. We will adapt the current HDRUK data access request (DAR) form to satisfy each TRE’s requirements \nand create a single access request form for all MIREDA cohorts. No person-level data will leave any of the TREs. \nInstead, access will be used to create anonymised, aggregated tables and results of statistical analyses within \neach TRE. To do so, common R scripts will be developed and sent from a central hub to each TRE (step 1), \nwhere they will be implemented locally to produce the output required and will be inspected by local analysts \n(step 2) before transferring to the central hub (step 3) for synthesis and further analysis (step 4), Figure 1. \n \n \nFigure 1. Federated analysis infrastructure. The numbers in the figure describe the following steps: \n1. Analysis R script sent to central data hub (TRE) with specifications for the data model and analyses. \n2. These are relayed to each birth cohort’s TRE. \n3. R scripts are run on each TRE’s own system to create anonymised data tables and statistics. \n4. These are sent back to the central hub for analysis. \n5. Central hub compiles and analyses the data. \n \nDiscussion \nThis protocol outlines the construction of a federated UK longitudinal birth cohort, leveraging routine data from \nUK healthcare providers, structured in a multi-modular format encompassing various stages from pregnancy to \nearly childhood. The adoption of the OMOP CDM for this initiative is pivotal, ensuring that the FAIR principles are \nembedded within the data management framework. By defining core variables and incorporating supplementary \ndata from diverse sources, the protocol facilitates the enhancement and contextualisation of the core datasets. \nUtilising the OMOP CDM allows for the standardisation of data, enabling uniform analyses across cohorts and \nfostering the use of a single, centralised coding strategy across different Trusted Research Environments \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n8 \n \n(TREs). Consequently, this standardised approach permits the execution of a single analysis script across \nmultiple cohorts, obviating the necessity for distinct scripts for each dataset. \nThe harmonised datasets created under this common data model will empower network studies, allowing for \nintegrated research across different centres while ensuring that TREs maintain control over their original data. \nThis is crucial for upholding the security protocols and governance standards of each TRE. Moreover, the central \nmanagement of access requests via the HDR-UK gateway ensures a consistent application process that meets \nthe specific requirements of each TRE site. This centralised system aligns with the FAIR principles by providing a \ntransparent, efficient, and universally applicable approach to data access, thereby advancing the integrity and \nutility of research conducted on these valuable cohorts. \n \nStrengths and limitations of this study \nIt is a strength that this methodology uses an established standard for data harmonisation and standardisation \nthrough use of the OMOP common data model to create comparable core datasets for research. \nThe process will establish a comprehensive, UK-wide database which consolidates clinical data from maternity, \nneonatal, child health and education records. It will also be enriched by quantitative and qualitative results from \nsurveys conducted by Born in Wales and Born in Bradford providing context and insight into detail not \nrepresented by routinely collected data. \nThe limitation of using routine and administrative data is that data may be missing or prone to errors. The data \ncould lack context and there is a potential for loss of data pertaining to individuals who relocate outside of a \ncohort’s remit during pregnancy or after the child’s birth. \n \nFuture directions \nAs the OMOP CDM has been used as an accepted standard internationally, future development will include \nworking with other cohorts, and other countries for international research comparisons and insights. MIREDA \nintends to expand the OMOP standard to other non-healthcare data e.g. releated social and community data, \nand to apply a pseudo-OMOP style solution to inclusion. Differences between cohorts’ foci will be exploited for \nextrapolation to others where the same data or depth of data is not available. MIREDA aims to use pooled cohort \ndata to address rare conditions. We seek to holistically prevent and reduce risk factors that adversely affect child \nhealth and wellbeing. We will work with HDR-UK to develop training and education materials, and to build a \nlibrary of commonly used coding scripts for rapid output analysis. \n \n \nConclusion \nOMOP CDM is used internationally as an accepted standard for healthcare data and thus has a well-developed \ntraining and support infrastructure, so it is admirably suited to standardising data for collaborative use and to \nbuild stronger networks with improved insight into specific research topics. It allows for rapid, repeatable \ncomparison of data between different regions, countries, and other localities. As there are many software tools \nand a network of experts available to support such work via the Observational Health Data Sciences and \nInformatics (OHDSI) group it is a relatively easy methodology to adopt and assimilate into existing TRE \nenvironments so that existing policies, procedures and standards can be maintained. \n \nAcknowledgements \nThis work was supported by an MRC Partnership Grant [MR/X02055X/1], MatCHNet pump-priming \n[U20005/302873] and an MRC Programme Grant [MR/X009742/1]. \n \nStatement of conflicts of interest \nThere are no known competing interests relating to any members of the MIREDA partnership and those involved \nin this study. \n \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n9 \n \nEthics Statement \nAccess to data is granted according to the information governance requirements of each TRE. The Data \nProtection Act 2018 is not applicable to anonymised data and the OMOP CDM will be anonymised and provide \naggregated data and statistics only. Each TRE has ethical approval for its operation and use, thus no additional \nethical approval was required beyond the standard project approval by official channels. \n \nDissemination: \nThe MIREDA partnership’s activities will actively engage, communicate, and disseminate this new data resource \namong researchers in maternal and infant health, the target research community. We will promote the resources, \ntools, collaborative expertise and infrastructure through publication, research collaborations, conferences, social \nmedia/marketing communications strategies. \nMIREDA will present at meetings and events held by relevant health, and health data, research organisations. \nWorkshops at targeted conferences will promote the linked data to encourage applications for grants and \nresearch collaborations.  \nFace-to-face workshops will present research with MIREDA data resources. Recordings will be uploaded as \nonline webinars. Key groups invited will include policymakers, NHS healthcare professionals from maternity, and \nindustry stakeholders.  \nMOOCs (massive open online courses) will be developed to improve skills and expertise in maternal and infant \nlinked data analysis. Regular blogs and podcasts will discuss data use. \nWe will work with international partners to raise awareness of MIREDA for research collaborations. \nMIREDA information will be hosted on the websites of MIREDA partners, and resources will be promoted through \nADR Digital Insights. \nTraining will be provided in association with HDRUK training group, ADR training and NIHR to develop capacity \nand expertise. Pump priming funding will enable early career researchers to utilise the data. \n \nPatient and public involvement \nPatients and/or the public were involved in the design, or conduct, or reporting, or dissemination plans of this \nresearch. We follow the Co-production of Research and Strategy (CORDS) standard operating procedure and \nthe UK standards for PPI involvement and National Institute for Health Research (NIHR) guidance from \nINVOLVE (29). Records of PPI activity are maintained using the Public Involvement in Research Impact Toolkit \n(PIRIT) (30). Each cohort and co-connect works with its own PPI groups under different names/guises to \ncontribute to each step of this process. \n \nData availability statement \nData will be available upon reasonable request through the Health Data Research (HDRUK) Innovation \nGateway. \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n10  \n Word count = 4,927 \n \n         \nReferences \n1. Broadbent P, Thomson R, Kopasker D, Mccartney G, Meier P, Richiardi M, et al. The public health \nimplications of the cost-of-living crisis: outlining mechanisms and modelling consequences. Lancet Reg \nHealth - Eur. 2023;27:100585.  \n2. Fears R, Gillett W, Haines A, Norton M, Meulen V ter. Post-pandemic recovery: use of scientific advice to \nachieve social equity, planetary health, and economic benefits. Lancet Planet Health. 2020;4:e383–4.  \n3. Haggarty P, Campbell DM, Duthie S, Andrews K, Hoad G, Piyathilake C, et al. Diet and deprivation in \npregnancy. Br J Nutr. 2009 Nov;102(10):1487–97.  \n4. Black RE, Allen LH, Bhutta ZA, Caulfield LE, Onis M de, Ezzati M, et al. Maternal and child undernutrition: \nglobal and regional exposures and health consequences. Lancet Lond Engl. 2008;371(9608):243–60.  \n5. Nurul-Farehah S, Rohana AJ. Maternal obesity and its determinants: A neglected issue? Malays Fam \nPhysician Off J Acad Fam Physicians Malays. 2020;15(2):34.  \n6. Schummers L, Hutcheon JA, Hernandez-Diaz S, Williams PL, Hacker MR, Vanderweele TJ, et al. \nAssociation of Short Interpregnancy Interval With Pregnancy Outcomes According to Maternal Age. JAMA \nIntern Med. 2018 Dec;178(12):1661–70.  \n7. Chen XK, Wen SW, Fleming N, Demissie K, Rhoads GG, Walker M. Teenage pregnancy and adverse birth \noutcomes: a large population based retrospective cohort study. Int J Epidemiol. 2007 Apr;36(2):368–73.  \n8. Smoking and pregnancy patient information leaflet | RCOG [Internet]. Available from: \nhttps://www.rcog.org.uk/for-the-public/browse-our-patient-information/smoking-and-pregnancy-patient-\ninformation-leaflet/ \n9. Nair M, Churchill D, Robinson S, Nelson-Piercy C, Stanworth SJ, Knight M. Association between maternal \nhaemoglobin and stillbirth: a cohort study among a multi-ethnic population in England. Br J Haematol. 2017 \nDec;179(5):829–37.  \n10. Mikkelsen B, Williams J, Rakovac I, Wickramasinghe K, Hennis A, Shin HR, et al. Life course approach to \nprevention and control of non-communicable diseases. BMJ [Internet]. 2019 Jan;364. Available from: \nhttps://www.bmj.com/content/364/bmj.l257 https://www.bmj.com/content/364/bmj.l257.abstract \n11. Hochlaf D, Quilter-Pinner H, Kibasi T. THE CASE FOR A NEW APPROACH TO PUBLIC HEALTH AND \nPREVENTION The progressive policy think tank. 2019; Available from: www.ippr.org \n12. O’connor M, Spry E, Patton G, Moreno-Betancur M, Arnup S, Downes M, et al. Better together: Advancing \nlife course research through multi-cohort analytic approaches. Adv Life Course Res [Internet]. 2022;53. \nAvailable from: http://creativecommons.org/licenses/by/4.0/ \n13. Downs JM, Ford T, Stewart R, Epstein S, Shetty H, Little R, et al. An approach to linking education, social \ncare and electronic health records for children and young people in South London: a linkage study of child \nand adolescent mental health service data. Available from: http://bmjopen.bmj.com/ \n14. Overy C, Reynolds LA, Tansey EM, Group H of BR. History of the Avon longitudinal study of parents and \nchildren (ALSPAC), c. 1980-2000 : the transcript of a Witness Seminar held by the History of Modern \nBiomedicine Research Group, Queen Mary, University of London, on 24 May 2011. 2012;122.  \n15. Data Standardization – OHDSI [Internet]. Available from: https://www.ohdsi.org/data-standardization/ \n16. Wright J, Small N, Raynor P, Tuffnell D, Bhopal R, Cameron N, et al. Cohort Profile: The Born in Bradford \nmulti-ethnic family cohort study. Int J Epidemiol. 2013;42:978–91.  \n17. Jones H, Seaborne MJ, Kennedy NL, James M, Dredge S, Bandyopadhyay A, et al. Cohort profile: Born in \nWales—a birth cohort with maternity, parental and child data linkage for life course research in Wales, UK. \nBMJ Open. 2024 Jan;14(1):e076711.  \n18. Carson LE, Azmi B, Jewell A, Taylor CL, Flynn A, Gill C, et al. Cohort profile: the eLIXIR Partnership-a \nmaternity-child data linkage for life course research in South London, UK. BMJ Open. 2020;10:39583.  \n19. Lee SI, Eastwood KA, Moss N, Azcoaga-Lorenzo A, Subramanian A, Anand A, et al. Protocol for the \ndevelopment of a core outcome set for studies of pregnant women with pre-existing multimorbidity. BMJ \nOpen. 2021 Oct;11(10):e044919.  \n20. Born in Scotland | The University of Edinburgh [Internet]. Available from: \nhttps://www.ed.ac.uk/cardiovascular-science/born-in-scotland \n21. Wilkinson MD, Dumontier M, Aalbersberg IjJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding \nPrinciples for scientific data management and stewardship. Sci Data. 2016 Mar 15;3(1):160018.  \n22. Alliance UHDR, NHSX. Building Trusted Research Environments - Principles and Best Practices; Towards \nTRE ecosystems. 2021 Dec; Available from: https://zenodo.org/record/5767586 \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint \n\n11  \n Word count = 4,927 \n \n         \n23. Dickerson J, Bridges S, Willan K, Kelly B, Moss RH, Lister J, et al. Born in Bradford’s Better Start (BiBBS) \ninterventional birth cohort study: Interim cohort profile. Wellcome Open Res. 2022 Oct;7:244.  \n24. BiB4ALL [Internet]. Born In Bradford. [cited 2024 Feb 13]. Available from: \nhttps://borninbradford.nhs.uk/what-we-do/cohort-studies/born-in-bradford/ \n25. Cohort studies: BiB4ALL [Internet]. 2024. Available from: https://borninbradford.nhs.uk/what-we-do/cohort-\nstudies/born-in-bradford/ \n26. WhiteRabbit for ETL design – OHDSI [Internet]. Available from: https://www.ohdsi.org/analytic-\ntools/whiterabbit-for-etl-design/ \n27. About - CaRROT-Mapper [Internet]. Available from: https://hdruk.github.io/CaRROT-Docs/CaRROT-\nMapper/about/ \n28. About - CaRROT-CDM [Internet]. Available from: https://hdruk.github.io/CaRROT-Docs/CaRROT-\nCDM/About/ \n29. INVOLVE. Guidance on co-producing a research project [Internet]. INVOLVE; 2018. p. 1–20. Available from: \nhttps://www.invo.org.uk/posttypepublication/guidance-on-co-producing-a-research-project/ \n30. Public Involvement in Research Impact Toolkit (PIRIT) - Marie Curie Research Centre - Cardiff University \n[Internet]. Available from: https://www.cardiff.ac.uk/marie-curie-research-centre/patient-and-public-\ninvolvement/public-involvement-in-research-impact-toolkit-pirit \n \n \nAbbreviations \nMIREDA  Maternal and Infant Research Electronic Data Analysis \nOMOP   Observational Medical Outcomes Partnership \nCDM   Common Data Model \nTRE   Trusted Research Environment \nBiW   Born in Wales \neLIXIR   early Life data Cross-Linkage in Research \nBiSL   Born in South London \nBiB4All  Local name for BaBi Bradford \nBaBi Bradford Born and Bred in Bradford \nBiBBS   Born in Bradford Birth Study \nBiS   Born in Scotland \nMuM-PreDiCT  Multimorbidity in Pregnancy: Determinant, Consequences, Clusters and Trajectories \nETL   Extraction, Transformation, and Load \nHDR-UK   Health Data Research UK \nGP    General Practitioner \nCPRD   Clinical Practice Research Datalink \nHES   Hospital Episode Statistics \nNIV   National Immunisation Vaccine \nDAR   Data Access Request \nMOOC   Massive Open Online Courses \nWHO   World Health Organization \nNIHR   National Institute for Health Research \nOHDSI   Observational Health Data Sciences and Informatics \nPPI   Patient and Public Involvement \nFAIR   Findable, Accessible, Interoperable, Reusable \n \n \nFull URLs contained within manuscript \nHealth Data Research Innovation Gateway   https://www.healthdatagateway.org/ \nObservational Health Data Sciences and Informatics (OHDSI) https://www.ohdsi.org/ \n . CC-BY 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted April 10, 2024. ; https://doi.org/10.1101/2024.04.08.24305489doi: medRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}