Harmonising definitions of multiple long term conditions for inflammation research: a co-production approach.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by claude@2026-08, 2026-08-09

This paper developed harmonized codelists for 60 multiple long-term conditions, incorporating patient input and clinical consensus, to support inflammation-focused epidemiological research.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

OBJECTIVE: Inflammation is implicated in many chronic diseases, but its role in driving multiple long-term conditions (MLTCs) remains unclear. The InflAIM programme aims to explore this potential link. Existing MLTC codelists are valuable but have not been designed for inflammation-focused research and often lack harmonisation across coding systems. This work aimed to develop a transparent, co-produced MLTC framework with harmonised codelists for epidemiological research, focusing on inflammatory conditions within the InflAIM programme. RESULTS: Using a systematic, co-produced approach, we combined literature review, clinical consensus, and patient input. From 8 sources, 363 candidate conditions were identified, refined to 60 validated MLTCs. Patients informed terminology and diagnostic experiences. Codelists were drawn from CALIBER (43), HDR UK Phenotype Library (4), MULTIPLY (8), with 7 developed de novo. Harmonisation covered Read v2, SNOMED CT, and Medcodeids. Quality assurance excluded 16.5% of codes and resolved 92 overlaps.We produced harmonised codelists for 60 MLTCs, comprising 7,426 validated MedCodeids. Clinical consensus identified 17 conditions with significant inflammatory components. Additional lists were developed for autoimmune and infectious diseases to support inflammation-focused analyses. This framework extends existing resources to meet the needs of inflammation-focused research, offering a transparent, harmonised, and patient-informed foundation for the InflAIM programme.
Full text 20,359 characters · extracted from pmc-nxml · 5 sections · click to expand

Methods

We employed a two-part methodology: (1) co-produced condition selection using an iterative, consensus-based process, and (2) the development of harmonised, interoperable codelists across UK EHR systems. We conducted a pragmatic literature review to identify key sources informing disease inclusion criteria for defining MLTCs. Starting with the seminal paper by Barnett et al. [ 2 ], we conducted forward citation searching, supplemented by keyword searches across PubMed, Google Scholar, Scopus, and Web of Science. An overview of outcomes of the search strategy is shown in Fig.  1 . Fig. 1 Flowchart of inclusion/exclusion*. Flow diagram showing the process of identifying and refining multiple long-term conditions (MLTCs). Candidate conditions were sourced from Barnett et al. [ 2 ], McRae et al. [ 4 ], Ho et al. [ 3 ], Cooper et al. [ 10 ], Lyall et al. [ 7 ], and the GEMINI codelist (2023). Following clinical review ( n  = 363), conditions were aggregated to ICD-10 subchapter level ( n  = 307), filtered for inclusion if listed in ≥ 2 sources and validated through clinical review ( n  = 103), reviewed by the PPIE core group ( n  = 62), and subjected to final coding refinements. This process provided a final list of 60 MLTCs. (n) = number of long-term conditions reported Flowchart of inclusion/exclusion*. Flow diagram showing the process of identifying and refining multiple long-term conditions (MLTCs). Candidate conditions were sourced from Barnett et al. [ 2 ], McRae et al. [ 4 ], Ho et al. [ 3 ], Cooper et al. [ 10 ], Lyall et al. [ 7 ], and the GEMINI codelist (2023). Following clinical review ( n  = 363), conditions were aggregated to ICD-10 subchapter level ( n  = 307), filtered for inclusion if listed in ≥ 2 sources and validated through clinical review ( n  = 103), reviewed by the PPIE core group ( n  = 62), and subjected to final coding refinements. This process provided a final list of 60 MLTCs. (n) = number of long-term conditions reported This process identified eight key sources: Barnett et al. [ 2 ], Ho et al. [ 3 ] MacRae et al. [ 4 ], GEMINI study [ 5 ], Dodds et al. [ 6 ], Lyall et al. [ 7 ], Head et al. [ 8 ], and Beaney et al. [ 9 ]. We systematically compiled all conditions identified across these sources. Conditions appearing in at least two of the eight sources comprised our provisional list of 363 conditions. A panel of three clinical experts—representing primary and secondary care, public health and psychiatry—reviewed the candidate conditions over several rounds, resolving disagreements by consensus. The final resulting refined list of MLTCs were categorised as chronic physical, mental, or infectious diseases. Defining “chronic” was a key issue: published definitions range from disease durations of greater than 3 to 12 months [ 9 , 11 ]. Rather than applying a strict threshold, the panel used clinical judgement instead of a strict cut-off. Conditions were excluded if they were acute, hereditary, congenital, childhood or pregnancy related. Long COVID was also excluded because it is coded as a broad, multi-system entity that cannot be positioned within a single-condition framework. Our PPIE group comprised 5–7 individuals with lived experience of two long terms conditions or of caring for someone with them. Members’ conditions included rheumatoid arthritis (RA), depression and diabetes among others. Participants were all aged over 40 with 55% female and 45% male. The group including people from different ethnic, socioeconomic and seldom-heard communities. Participants attended three separate group sessions involving whiteboard discussions on MLTCs, reviewing the clinical condition list and mind-mapping to integrate lived and clinical perspectives. These activities revealed differences in language and symptom framing, ensuring the findings reflected both patient experience and clinical understanding. Following the initial review stages, we re-examined the literature and coding sources to ensure completeness and identify any overlooked conditions, incorporating findings from the PPIE process. The clinical panel then undertook a final review to refine the list, resolve outstanding disagreements and reach consensus. The provisional conditions were then aggregated to ICD-10 subchapter level to align with established epidemiological groupings and ensure consistency across coding systems (Read v2, SNOMED CT, ICD-10). Following condition selection, we developed comprehensive phenotyping algorithms for the condition set across three coding systems relevant to UK primary care research: Systematized Nomenclature of Medicine – Clinical Terms (SNOMED CT, UK Clinical Extension & International Edition), Read Code Version 2, and CPRD (Clinical Practice Research Datalink) Aurum Medcodeids which were mapped to the relevant ICD-10 subchapters. These systems reflect the range of EHRs and research datasets included in the InflAIM programme. We used a hierarchical approach to develop phenotyping algorithms, prioritising sources with strong methodological rigor, validation and research use. This ensured consistent quality and semantic interoperability. Our primary source was the CALIBER (CArdiovascular disease research using LInked Bespoke studies and Electronic health Records) algorithms [ 12 ], chosen for their validated definitions and reproducibility across multiple studies. As CALIBER was developed using the CPRD GOLD database, we adapted codelists for CPRD Aurum using the mapping approach from Head et al. [ 8 ]. Secondary sources included the HDR UK Phenotype Library, recognised for expert clinical validation, and the MULTIPLY (MULTI-national consortia to study Polygenic risk scores by Large-scale health sYstems Initiative), which provides consensus-based phenotypes with transparent methods and primary care relevance [ 12 ]. For conditions lacking published definitions, we implemented a systematic semi-automated development process. This began with ICD-10 concept expansion using the WHO API (World Health Organization, Application Programming Interface) (2019), followed by cross-mapping using NHS TRUD (Technology Reference data Update Distribution) files and the CPRD Aurum Medical Dictionary. Prioritising sensitivity in the initial stage, we applied automated filtering across five sequential categories: procedure-related, administrative and non-diagnostic, acute event, congenital/hereditary, and finally “not otherwise specified” or “unspecified” code exclusions. Where a code appeared across multiple condition lists, we used clinical adjudication to assign it appropriately. Supplementary analyses followed two steps. First, the clinical panel identified a subset of the primary condition list with significant inflammatory pathophysiology. Second, for autoimmune and infectious diseases, we systematically sourced published codelists (e.g. Lyall et al. [ 7 ], Watson et al. [ 13 ], Janbek et al. [ 14 ]), harmonised across coding systems using the CPRD Aurum Medical Dictionary and the same methods as for primary codelists. Final lists were clinically curated and de-duplicated to exclude overlaps with the primary MLTC set.

Results

The literature review identified 363 candidate conditions across eight key sources. Through a five-stage co-production process—including clinical expert review, patient and public involvement, and final adjudication. This was reduced to 60 conditions (Table  1 ; Fig.  2 ), each agreed by unanimous clinical consensus as relevant to MLTCs for the InflAIM programme. Table 1 MLTC conditions, ICD-10 chapter and inflammatory indicator ICD-10 MLTC condition A15-A19 Tuberculosis* A69 Other spirochaetal infections* B18-B19 Chronic viral hepatitis* B20-B24 HIV disease C00-C80, C97, D37-D44, D47, D48 Malignant neoplasms C81-C96 Malignant neoplasms, stated or presumed to be primary, of lymphoid, haematopoietic and related tissue D00-D09 In situ neoplasms D50 Iron deficiency anaemia E00-E07 Disorders of thyroid gland E10-E14 Diabetes E20 Hypoparathyroidism E21 Hyperparathyroidism E27 Addison’s disease F00-F03 Dementia F10-F19 Drug or alcohol misuse F20-F29 Schizophrenia F30-F39 Mood disorders F40-F48 Neurotic, stress-related and somatoform disorders F50 Eating disorders G20-G21 Parkinson disease G35, H46 Multiple sclerosis (and optic neuritis)* G40, F80.3 Epilepsy G43 Migraine G47 Sleep disorders G61-G64 Polyneuropathies and other disorders of the peripheral nervous system H25 Cataract H33 Retinal detachments and breaks H35 Other retinal disorders (Macular degeneration) H40-H42 Glaucoma H53-H54 Visual disturbances and blindness H81.0 Meniere disease H90-H91 Conductive and sensorineural hearing loss I05-I09 Chronic rheumatic heart diseases* I10-I15 Hypertensive diseases I20-I25 Ischaemic heart diseases I30-I52 Other forms of heart disease I60-I69 Cerebrovascular Disease I70-I79 Disease of arteries, arterioles and capillaries I80-I89 Diseases of veins, lymphatic vessels and lymph nodes, note elsewhere classified J30-J32 Rhinitis and Sinusitis* J40-J47 Chronic lower respiratory diseases* J84 Other interstitial pulmonary diseases* K21, K27, K29 Diseases of oesophagus, stomach and duodenum K50-K52 Noninfective enteritis and colitis* K57 Diverticular disease of the intestine* K58 Irritable Bowel Syndrome K70-K77 Diseases of the liver K80-K87 Disorders of gallbladder, biliary tract and pancreas L20-L30 Dermatitis and eczema* L40 Psoriasis* M05-M14 Inflammatory polyarthropathies* M15-M19 Arthrosis M30-M36 Systemic connective tissue disorders* M45 Ankylosing spondylitis* M47, M51 Spinal degenerative disease M80-M85 Disorders of bone density and structure N18-N19 Chronic kidney disease N40-N42 Disorders of the prostate N70-N77 Inflammatory diseases of the female pelvic organs* N80 Endometriosis* *Inflammatory condition MLTC conditions, ICD-10 chapter and inflammatory indicator *Inflammatory condition Fig. 2 Workflow for the development and validation of 60 harmonised phenotyping algorithms. * Workflow for development and validation of harmonised phenotyping algorithms. The diagram illustrates the hierarchical sourcing strategy, integration of published and de novo algorithms, and the multi-stage filtering and adjudication processes that produced the final validated codelists across three UK primary care coding systems. Counts shown in each block indicate the number of unique codes (concepts/terms) present at that stage within the respective coding system Workflow for the development and validation of 60 harmonised phenotyping algorithms. * Workflow for development and validation of harmonised phenotyping algorithms. The diagram illustrates the hierarchical sourcing strategy, integration of published and de novo algorithms, and the multi-stage filtering and adjudication processes that produced the final validated codelists across three UK primary care coding systems. Counts shown in each block indicate the number of unique codes (concepts/terms) present at that stage within the respective coding system A major methodological step was aggregating conditions to ICD-10 subchapters: 307 of the 363 were consolidated into broader groupings to enable harmonised cross-mapping, though this reduced granularity. After clinical and PPIE review of 103 subchapter-level groupings, the final set of 60 long-term conditions was selected. Conditions were classed as inflammatory when they involved active immune-driven tissue inflammation; grouped conditions were labelled inflammatory when key subtypes met this criterion. Autoimmune and infectious diseases were handled separately. Insights from the PPIE group informed several aspects of condition selection. Members highlighted differences between clinical diagnoses and lived experience—especially for mental health and fluctuating conditions—prompting closer scrutiny of coding. Members described gout as chronic but episodic, and inflammatory bowel disease (IBD) as marked by debilitating flares, emphasising that such fluctuations often go uncaptured in clinical records. This input supported retention of conditions that may be inconsistently coded but consistently described as impactful. Participants further stressed that conditions such as, RA, Multiple Sclerosis and systemic lupus erythematosus affect multiple organ systems, prompting scrutiny of how conditions were grouped for analysis. The final list included conditions from a wide range of clinical domains: cardiovascular, metabolic, endocrine, musculoskeletal, respiratory, neurological, mental health, gastrointestinal, renal, dermatological, haematological, oncological, and infectious diseases. Of the final MLTC list, 17 conditions were identified by clinical consensus as having significant inflammatory pathophysiology, including RA, inflammatory bowel disease and asthma (see Table  1 , inflammatory conditions indicated by *). To further support InflAIM’s focus on inflammation, we identified a comprehensive list of autoimmune and infectious disease conditions through separate literature and clinical reviews for targeted phenotyping. Phenotyping algorithms were successfully developed for all 60 conditions using a hierarchical approach to sourcing, based on methodological rigour, validation status, and research adoption. CALIBER provided codelists for 43 conditions (71.6%), forming the majority of the final set. HDR UK Phenotype Library contributed 4 (6.6%), including ischaemic heart disease and tuberculosis. MULTIPLY Initiative provided 8 (13.3%), including systemic connective tissue disorders. Codelists for seven conditions (11.6%) were developed de novo through a semi-automated process, for conditions such as hypoparathyroidism and in situ neoplasms. Of the conditions sourced above, two (3.2%) required hybrid approaches combining multiple sources. CALIBER provided codelists for 43 conditions (71.6%), forming the majority of the final set. HDR UK Phenotype Library contributed 4 (6.6%), including ischaemic heart disease and tuberculosis. MULTIPLY Initiative provided 8 (13.3%), including systemic connective tissue disorders. Codelists for seven conditions (11.6%) were developed de novo through a semi-automated process, for conditions such as hypoparathyroidism and in situ neoplasms. Of the conditions sourced above, two (3.2%) required hybrid approaches combining multiple sources. The initial compilation included 8,898 Medcodeids across three UK coding systems. Following automated and manual filtering to improve diagnostic precision, 1,472 codes (16.5%) were excluded, resulting in 7,426 validated codes. These were harmonised across SNOMED CT, Read Version 2, and CPRD Aurum Medcodeids (Fig.  2 ). Automated quality checks flagged 92 SNOMED CT concepts and 160 Read/Medcodeids that mapped to multiple conditions. These were adjudicated by clinical experts, with codes either reassigned to a single most appropriate condition or retained across multiple lists where justified by clinical overlap (Additional file, Table 1). The taxonomic structure of the clinical codelists for 60 LTCs is visualised as a radial dendrogram in the Additional file, Fig. 1. Supplementary codelists were developed for 228 autoimmune and 2,107 infectious diseases, based on a systematic sourcing and harmonisation process. The autoimmune codelist included 489 Read codes and 341 SNOMED CT codes. The infectious disease list comprised 2,107 Read codes and 1,390 SNOMED CT codes (see Additional files, Tables 2 and 3).

Discussion

This study presents a systematically developed MLTC framework, tailored to support inflammation-focused research. Through a five-stage co-production process, we integrated evidence from the literature, clinical consensus, and public involvement to define a robust condition set and harmonised codelists. This addresses limitations in prior frameworks, which often lack transparency, patient input, or interchangeability across coding systems. Our approach was shaped by the aims of the InflAIM programme, exploring inflammatory mechanisms in multimorbidity. This influenced condition selection—prioritising those with inflammatory pathophysiology—and the decision to develop supplementary autoimmune and infectious disease codelists. Choices around granularity were guided by biological relevance; for example, grouping spinal degenerative conditions while treating inflammatory arthropathies as distinct. Patient and public contributors played a central role, prompting closer scrutiny of mental health and fluctuating conditions and highlighting mismatches between clinical codes and lived experience. This strengthened the framework’s alignment with both diagnostic logic and real-world complexity. Codelists were developed using a hierarchical strategy, prioritising validated algorithms across three UK coding systems. Where definitions were absent, we derived new algorithms using a semi-automated pipeline with clinical adjudication and automated checks to improve specificity and resolve overlaps. ICD-10 alignment and mapping tables enable consistent and reproducible use across datasets such as CPRD, UK Biobank, ELSA (English Longitudinal Study of Ageing) and international resources (e.g. Global Burden of Disease, WHO Global Health Observatory, OECD, Organisation for Economic Co-operation and Development). Future work will include validating use of the framework by linking it to key epidemiological outcomes including morbidity, mortality and age-specific disease burden. Our literature review focused on key sources, not exhaustive coverage. Clinical review involved judgement calls on chronicity and condition grouping. Aggregation of conditions to ICD-10 sub-chapter level will inevitably have resulted in some loss of granularity for certain conditions. Despite rigorous harmonisation, some coding inconsistencies may persist. The PPIE process was qualitative and influential, but its specific impact is not easily measured. Overall, this work delivers a flexible, transparent and interoperable MLTC identification framework with broad applicability. By explicitly incorporating inflammation as a core mechanism and combining clinical evidence with patient perspectives, it enhances the validity and relevance of MLTC research.

Introduction

The UK’s ageing population is living longer with complex health needs: two-thirds of adults aged 65 or older have at least two long-term conditions (MLTCs) [ 1 – 4 ], compared with one-third globally [ 1 ]. Policy efforts, including NICE (National Institute for Health and Care Excellence) guidelines [ 5 ] and NIHR’s (National Institute for Health and Care Research) strategic focus on MLTCs [ 6 ], are limited by the lack of standardised definitions. Existing frameworks vary in condition selection, coding, and disease groupings [ 7 – 9 ], limiting comparability across studies. The InflAIM programme examines whether inflammation drives MLTC development, especially in socioeconomically disadvantaged groups. Large-scale EHR data are crucial but pose methodological challenges: most MLTC codelists do not separate inflammatory from non-inflammatory conditions—e.g., Osteoarthritis and Rheumatoid Arthritis (RA) are often combined—reducing their value for inflammation-focused research. Differences in coding systems hinder harmonisation, and few frameworks incorporate lived experience, despite evidence that patients describe symptoms differently from clinical labels [ 2 , 3 ]. Examining this with large-scale electronic health record (EHR) data is important but methodologically challenging. To address these challenges, we developed a co-produced MLTC framework with three key innovations: A targeted focus on identifying and prioritising inflammatory conditions within MLTC definitions; The development of harmonised codelists across three major UK primary care coding systems; and. A methodology that integrates patient experience with clinical and coding expertise. A targeted focus on identifying and prioritising inflammatory conditions within MLTC definitions; The development of harmonised codelists across three major UK primary care coding systems; and. A methodology that integrates patient experience with clinical and coding expertise.

Supplementary Material

Supplementary material 1. Supplementary material 1. Supplementary material 2. Supplementary material 2. Supplementary material 3. Supplementary material 3. Supplementary material 4. Supplementary material 4.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-09T06:10:49.860119+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-ND-4.0