Standardizing data exchange for clinical research protocols and case report forms: An assessment of the suitability of the Clinical Data Interchange Standards Consortium (CDISC) Operational Data Model (ODM).

OA: closed
⚙ AI-generated deep summary by qwen3.7-flash, 2026-08-26 · read from full text ⓘ

This study evaluates the suitability of the Clinical Data Interchange Standards Consortium Operational Data Model for standardizing the exchange of clinical research protocols and case report forms. Using a single NIH protocol as a case study, the authors analyze the complete lifecycle from drafting to data sharing to determine if a single format can support all stages of clinical trial management. The paper highlights that while existing standards cover specific elements like registration or case report forms, no current standard effectively transports the full protocol document in a structured sense across the entire study duration. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Efficient communication of a clinical study protocol and case report forms during all stages of a human clinical study is important for many stakeholders. An electronic and structured study representation format that can be used throughout the whole study life-span can improve such communication and potentially lower total study costs. The most relevant standard for representing clinical study data, applicable to unregulated as well as regulated studies, is the Operational Data Model (ODM) in development since 1999 by the Clinical Data Interchange Standards Consortium (CDISC). ODM's initial objective was exchange of case report forms data but it is increasingly utilized in other contexts. An ODM extension called Study Design Model, introduced in 2011, provides additional protocol representation elements. Using a case study approach, we evaluated ODM's ability to capture all necessary protocol elements during a complete clinical study lifecycle in the Intramural Research Program of the National Institutes of Health. ODM offers the advantage of a single format for institutions that deal with hundreds or thousands of concurrent clinical studies and maintain a data warehouse for these studies. For each study stage, we present a list of gaps in the ODM standard and identify necessary vendor or institutional extensions that can compensate for such gaps. The current version of ODM (1.3.2) has only partial support for study protocol and study registration data mainly because it is outside the original development goal. ODM provides comprehensive support for representation of case report forms (in both the design stage and with patient level data). Inclusion of requirements of observational, non-regulated or investigator-initiated studies (outside Food and Drug Administration (FDA) regulation) can further improve future revisions of the standard.
Full text 41,881 characters · extracted from pmc-nxml · 6 sections · click to expand

Section 2

A clinical study (or protocol ) goes through several stages, including protocol drafting by a research team, protocol submission and approval by one or more IRBs, study registration within a clinical trial registry, study recruitment pre-screening ( pre-screen recruitment questions ), actual patient recruitment at the study site ( in-person recruitment questions ), and collection of study data, perhaps using electronic Case Report Forms (eCRF) within a Clinical Trial Data Management System (CTDMS) or an electronic health record (EHR). Ideally, protocols should be represented in a format that supports common protocol data elements, such as study title, locations or enrollment goal, as well as stage-specific data and metadata elements, such as information required for IRB approval, sharing study details and study eCRFs with all study sites (including clinical research organizations or sponsors), and submission of final data to a statistician, regulator or data sharing platform. A single format for all these tasks, from study drafting (prior to enrolling the first patient) through study completion and follow-up (after the last patient's data have been collected) would be preferable. We define a study protocol as a detailed document that is typically 10–80 pages long and includes the study schedule, detailed description of all study events as well as other elements defined by the Good Clinical Practice guideline (E6) [ 3 ] from the International Conference on Harmonization. This guideline, which was created with input from US as well as EU authorities, standardizes numerous protocol sections, such as, withdrawal criteria, blinding or adverse event reporting. Within the protocol, we also distinguish a short set of study metadata elements that we refer to as study registration information (such as title, principal investigator, research sites, study design, or enrollment goal). These elements are typically required by trial registries or internal study administration systems. Table 1 provides an overview of different study documentation components, as well as relevant policies for each component. We consider design of case report forms to be an important attachment to the study protocol and an integral part of good study documentation. Study protocols are also needed when study results are obtained and communicated, and hence we also include in Table 1 study results data elements. These include summary data that are required by some trial registries (and by US law) and individual patient level data, that are needed for submission to a regulator or to some data sharing platforms (eg, TrialShare from National Institutes of Allergy and Infectious Diseases (NIAID) Immune Tolerance Network) [ 4 ]. The resulting representation will need to be computable, for example, an Extensible Markup Language (XML) file that can be consumed both by information systems and by humans (after transformation into formats such as Hypertext Markup Language (HTML) or Portable Document Format (PDF). Existing standards cover to some extent the representation of study registration information and CRFs; however, there are no standards capable of transporting (in a structured sense) the full protocol document. The dominant standard development organization (SDO) for creating clinical research informatics standard is the Clinical Data Interchange Standards Consortium (CDISC), which was established in 1999. CDISC standards development efforts are organized into different workgroups, with the Protocol Representation Group (formed in 2002) working to create such a standardized format. CDISC intends to use this or a similar format to standardize submission of protocol data to clinical trial registries. For example, currently it is not possible to use a single format to submit clinical study data to USA's clinical trial registry ( ClinicalTrials.gov ) and EU's registry run by European Medicines Agency (EMA). Protocol representation formats have been the subject of several prior studies. The most relevant standard is the CDISC Operational Data Model. Prior studies using ODM focused mostly on case report forms, rather than strictly on protocol representation. Bickel et al. developed an i2b2-based tool that can import the CDISC ODM formatted data into an i2b2-based data warehouse [ 5 ]. Dugas at al. developed a Web-based platform [ 6 ] and an R package [ 7 ] supporting exchange of empty eCRFs in ODM format. Karam et al. from World Health Organization developed an ODM extension focused on clinical trial registration [ 8 ]. ODM was one of the first standards developed by CDISC and it was meant from the start to be a foundation standard with building blocks for capturing a range of clinical study data. ODM was initially created in 1999 with updates to version 1.3 in 2005 and a small update to version 1.3.2 in 2013. It was designed to meet an initial set of requirements for transporting case report data between research collaborators (e.g. sponsors, clinical research organization, electronic data capture vendors and others) including adherence to US regulatory requirements (21CFR11) for data provenance and electronic signatures. ODM has pre-defined XML elements that can represent protocol-level metadata ( Study element), single case report form data ( FormDef element), form sections ( ItemGroupDef element), individual CRF questions ( ItemDef element) and allowable values for questions ( CodeListDef element). The standard is essentially an XML schema with an accompanying schema usage specification document. In addition to its function as a stand-alone standard, it is also a foundation for several other standards: Define-XML, Dataset-XML [ 9 ] and Study Design Model (SDM). All of these standards are based on the foundational syntax defined by ODM and are, in an XML-schema sense, ODM XML schema extensions. Comprehensive protocol representation was outside the initial focus of ODM and only later extensions, such as SDM, focused on this issue. To separate the various forms of eCRFs and recruitment stages, Weng at al. defined the concept of pre-screening [ 10 ] and analyzed whether EHR records can be used to pre-screen potential research subjects. Richesson at al. looked overall at clinical research informatics standards [ 11 ] and identified gaps that exist [ 12 ]. In Europe, the Electronic Health Records for Clinical Research (EHR4CR) project is attempting to integrate collection of clinical study data within EHRs and is using ODM as the main standard. In the US, the Office of the National Coordinator for Health Information Technology is working to improve standards adoption by US EHR manufacturers through the Structured Data Capture Group of its Standards & Interoperability Initiative. The ClinicalTrials.gov team developed an XML schema that is used to register trials [ 13 ]. Finally, the NIH framework for defining and coordinating clinical research Common Data Elements (CDE) aims to enable data collection harmonization across studies [ 14 ]. Besides the CDISC ODM, there are other relevant CDISC standards and models, such as the Study Data Tabulation Model (SDTM) and Define-XML that are regulatory submission standards. However, the majority of prior protocol data exchange efforts use the more generic CDISC ODM standard. A high level CDISC conceptual model called BRIDG is another candidate protocol representation framework. BRIDG was first introduced in 2006, seven years after the creation of ODM. CDISC strives to align all current and future standards to this common domain model that specifies basic elements such as investigator, subject, study or intervention. BRIDG is expected to be a final International Standard Organization (ISO) standard by April 2015 (it passed draft international standard ballot in January 2015). Most recently, standard integration efforts are most apparent within the CDISC Shared Health and Clinical Research Electronic Library (SHARE) [ 15 ]. SHARE was initially made available to CDISC Platinum members and it will be made available to others pending an appropriate business model for sustainability. BRIDG was also embraced by Health Level 7 and International Organization for Standardization (ISO), FDA and National Cancer Institute. There are no published reports or roadmaps that comment on how well the latest version of ODM is harmonized with the central BRIDG model. While BRIDG is potentially a much more comprehensive model, our choice to use ODM was driven by the fact that BRIDG, as a Unified Modeling Language (UML) model, does not offer a well-defined data storage format. In other words, it is not possible to use several Unified Modeling Language tools to model a single clinical trial using BRIDG, and be certain to produce identical files (e.g., NCT001234.BRIDG). In this article, we use a case study approach in which we analyze the complete protocol lifecycle of a single protocol in the National Institutes of Health (NIH) Intermural Research Program (IRP). The intramural research program has currently over 2300 active protocols (as of August 2014). On average, 231 new protocols are initiated every year. Since 1953, the Intramural Research Program has registered a total of 8017 completed studies and maintains a repository [ 16 ] of data collected in those studies. Approximately half of the studies are observational natural history studies (based on an analysis of 1155 intramural studies initiated during the last 5 years). To better illustrate the protocol representation and communication challenges, we provide below a brief description of the protocol lifecycle in the Intramural Research Program. There are several large intramural IT systems, and each needs to represent slightly different clinical study metadata or various types of case report forms (see also Fig. 1 and Table 2 ). ProtoType is a Web-based protocol authoring tool used by some intramural investigators during protocol drafting. Finalized protocols are then passed to the IRBs which use electronic eIRB systems to track each submitted and approved study. At NIH, there are 10 IRBs and they use one of the two kinds of eIRB systems (referred to as PTMS and iRIS; see Appendix A for a full overview). In our analysis, we focus on the home-grown PTMS systems because of greater control over its features. If a study is approved by the IRB, the NIH Office of Protocol Services imports study data into a repository of active study protocols (called ProTrak). Among other tasks, the ProTrak system and the ProTrak team registers the study in the ClinicalTrials.gov registry (as required by US law for certain types of interventional studies). Approved and registered studies may initiate patient recruitment, which is facilitated (for studies that choose to do so) by the NIH Patient Recruitment and Public Liaison (PRPL) Office. The Recruitment Team maintains a database of research volunteers and tracks every study phone screening session within the Recruitment Volunteer System (RVS) [ 17 ]. RVS enrollment criteria are specified in the IRB protocol as phone screening CRFs . If a patient meets pre-screening phone criteria, his contact information is passed on to the study research coordinator who typically arranges in-person screening visit at the research clinic. Screening visit data are driven by screening CRFs and protocol event schedules (such as screening laboratory tests procedures). Protocol-specific eCRF data are recorded within a CTDMS (eg, Pelvic Pain Screening Questionnaire). Within the Intramural Research Program, a total of 6 different CTDMSs are used. See Appendix A for a full list. The most widely used is Clinical Trial Database (CTDB). Other relevant study data (eg, blood sample laboratory results or EKG) are recorded within an EHR system. The main intramural EHR system is currently Allscripts Sunrise Clinical Manager. Data from the EHR and some CTDMSs are integrated into a data repository called the Biomedical Translational Research Information System (BTRIS) [ 16 ]. For researchers conducting secondary analyses executed within BTRIS, protocol metadata of terminated protocols are needed for proper data interpretation. For example, data from ‘Natural History’ studies are usually longitudinal and come from multiple visits while data from a ‘Phase 1 Clinical Trial’ study type may be limited to a single visit. Table 2 shows an overview of various systems that need to represent protocol data and what representation standards they currently support. The last column indicates whether the CDISC ODM standard contains all necessary fields to be able to replace the current standard and whether ODM is currently used. A graphical overview of data flow is provided in Fig. 1 . We set out to investigate the best standard and mechanism that would support moving protocol data and metadata across these stages/systems.

Intro

There are increasing pressures to lower the cost of conducting human clinical studies. One way to achieve this is to streamline the communication of clinical study protocol information to study sites and other stakeholders, such as trial registries or institutional review boards (IRBs). Because completed clinical studies represent a significant past investment and the need to re-analyze the data is common, many institutions that are consolidating data from clinical studies into larger repositories will benefit from such streamlining as well. A meta-analysis [ 1 ] reported that between 9% to 49% of randomized control trials report on outcomes that were not declared in a trial registry. This indicates that post hoc analyses can be quite frequent. Public pressure for comprehensive sharing of clinical study data will likewise benefit from improved exchange of study data and metadata [ 2 ]. The US National Institutes of Health (NIH) shares all of the above motivations for standardizing protocol information. We use an example of one NIH protocol to examine the issues of standardization and explore the suitability of the Clinical Data Interchange Standards Consortium (CDISC) Operational Data Model (ODM) standard to facilitate exchange of clinical protocol data. In contrast to prior studies, we evaluate the use of a single format that could cover the complete study life-cycle from study inception to termination and sharing of study results data.

Methods

We demonstrate the challenges of protocol representation using a case study approach. We selected a single protocol (“The Safety and Effectiveness of Surgery with or without Raloxifene for the Treatment of Pelvic Pain Caused by Endometriosis”) and modeled its data representation in various stages of its lifecycle. Throughout this article, we use the word “study” or “trial” to refer to the raloxifene trial and the word “analysis” to refer to our representation format analysis. For each study stage, we quantified, where possible, the number of study elements that could be represented directly in ODM. For some study stages, where we found that ODM provides limited support and coverage, we investigated and described which other CDISC standards could be used instead. The majority of the protocols of completed intramural research studies are available as PDF documents which are image scans of the paper protocols. We identified the raloxifene file and converted it into a textual format. We used the WmHelp XMLPad editor (version 3.0.2.1; available as freeware from xmlpad-mobile.com ) to create the study ODM files. We pre-loaded the ODM XML schema (version 1.3.2) to enable XML schema validation while editing the study's XML file. Because a second CDISC standard called Study Design Model (SDM) offers improved support for some study metadata, we have also added the SDM XML in addition to the ODM schema. This is a preferred approach recommended by CDISC, since SDM was released as an extension to the ODM schema. In addition to the XML editor, we also used a Web-based wizard tool [ 18 ] for creating CDISC study outline files. To facilitate accurate characterization of the study lifecycle, we delineate several study stages below and provide additional stage-specific methodological, institutional or informatics context. The goal of the study drafting stage is to generate study registration information and the full protocol that describes the steps and procedures of the protocol. These data are needed to either support a study funding decision process or to communicate the study to the larger research team or external study sites. The Intramural Research Program standardized data relating to the registration of the study itself (as opposed to registration of research subjects) by instituting a common form (Initial Review Application; formerly referred to as form 1195). The full form is available at http://dx.doi.org/10.6084/m9.figshare.1096216 . It contains, for example, fields for protocol title, subject accrual characteristics, protocol type and phase, principal and associate investigators, and related FDA identifiers, such as Investigational New Drug [IND] or Investigational Device Exemption [IDE]). Because study registration data are required by many systems, most investigators provide them in structured form either in the ProtoType system or in an eIRB system. Neither ProtoType nor PTMS support export of these data in a structured form. ProtoType allows export of the data in Microsoft Word format and PTMS supports generation of a PDF file that recreates the 1195 form. Fig. 2 shows structured data entry of form 1195 subject accrual characteristics data within ProtoType. The study drafting stage also involves creation of the detailed protocol. The majority of protocols in the NIH Intramural Research Program are authored in word processing software (e.g., Microsoft Word) and communicated in Office Open XML format (.docx) or PDF format. PTMS, in the current version, does not support authoring the full protocol. The ProtoType system, on the other hand, enables principal investigators to author protocols using a Web-based system. In contrast to using word processing software on a local PC, ProtoType provides team features such as protocol commenting and approval by collaborating scientists. Similar to word processing software, it offers the ability to add images and references to any part of the protocol. By requiring entry in some protocol sections, it encourages data completeness. Because registration of most interventional clinical studies is required by US law, we consider submission to a trial registry in a separate registration stage. In the Intramural Research Program, the ProTrak system and team facilitate submission of data to the ClinicalTrials.gov registry. Data entered into ProTrak and submitted to the registry can be re-used to extend the protocol metadata created in the study drafting stage. The study initiation stage starts upon IRB approval. In order to collect patient level data, one or more CRFs must be defined. In a multi-site trail, each site may be using a different electronic data capture system and the ability to import and export CRFs is an important function of a protocol representation format and can save time spent on duplicate entries of the same CRFs into multiple system. It is important to note that communication of empty CRFs is free of any patient privacy considerations because only empty forms are being transmitted. In our analysis of this study stage we focus on empty CRFs representations in this stage and consider CRFs with patient level data in the next stage. The study termination phase begins when last patient data are collected. During the study termination phase, a data capture format for study results representation is needed, in addition to the protocol representation (see Table 1 ). Export of collected data for analysis is the most important step in the termination phase. A data export, however, may also occur during the study execution phase to support, for example, interim study monitoring by the Data Safety Monitoring Board (DSMB) (e.g., premature termination of the study due to lack of efficacy or other reasons). A standardized export of study data is helpful to statisticians, who deal with multiple studies, or to data repositories responsible for long-term storage of clinical study data.

Results

The selected protocol “The Safety and Effectiveness of Surgery With or Without Raloxifene for the Treatment of Pelvic Pain Caused by Endometriosis” is a description of the trial registered at ClinicalTrial.gov under identification number NCT00001848 ( http://clinicaltrials.gov/show/NCT00001848 ). The trial examined whether 6 months of raloxifene was effective in the treatment of chronic pelvic pain in women with endometriosis. Women with chronic pelvic pain underwent laparoscopy and were randomly allocated to raloxifene or placebo. A second laparoscopy was performed at 2 years, or earlier, if pain returned. Detailed study results for raloxifene are available in the primary trial result publication [ 19 ] with additional publications describing secondary analyses of migraine prevalence [ 20 ] and the relationship of pain location to biopsy-proven lesions [ 21 ]. The study data collection was completed in 2004. We obtained the permission of the study principal investigator to use trial's empty case report forms and parts of the protocol for our analysis. The raloxifene trial had 25 forms defined with a total number of 686 questions. The most used forms were ‘Quality of life questionnaire’, ‘Pelvic Pain Questionnaire’ and ‘Sexual Function Questionnaire’. See Table 3 for complete overview. The trial had 10 study intervals defined (e.g., Screening, Baseline, 3 Month, 6 Month, 9 Month, 24 Month) with many forms assigned for repeated collection at multiple intervals. The Supplemental File S1 contains all protocol metadata available within ODM that are typically specified during early study drafting. Most of those are elements from the NIH form 1195. The elements directly supported by ODM were: title, study ID, and study description (also referred to as study précis). Using the SDM extension, we were also able to capture study inclusion and exclusion criteria (as a free text block, without separating individual criteria). Ideally, text-based criteria would be augmented with annotations that could facilitate computerized determination of eligibility [ 22 , 23 ]. Such a feature would currently be only achievable via an additional ODM extension or a revision of the existing SDM extension. ODM lacks designated XML elements to represent brief title, scientific keywords, study type (observational, interventional), interventional trial phase (0,1,1-2,2,3,4), inclusion of healthy volunteers, use of ionizing radiation (yes/no flag), or capturing of associated identifiers for Investigational New Drug (IND) application or Investigational Device Exemption (IDE) that are associated with a study. The study drafting data also include the names of the principal investigator, the contact for scientific queries, and the collaborators (such as lead associate investigator, associate investigators and medical advisory investigator). While these additional contacts can be represented in ODM using the “User” element with an optional UserType attribute, ODM does not mandate the use of any commonly used user types (e.g. principal investigator). The ODM User element is predominantly meant for management of data access for clinical study coordinators at various study sites during the study execution stage. Possible changes to the ODM model suggested in the charter of the newly formed CDISC Clinical Trials Registration group (CTR) may improve ODM support for this study phase [ 24 ]. Moreover, such changes should not only consider minimum regulatory data elements but also elements that would facilitate protocol-level data exchange between research systems, such as eIRB systems. We converted the raloxifene study protocol PDF into a structured HTML5 (Hypertext Markup Language) document using optical character recognition in combination with manual editing and review. Since ODM provides almost no support for representing the full protocol (except for brief study description field), the most optimal way to model these data is via an institutional ODM extension. The Raloxifene protocol sections are shown in Fig. 3 . However, limited capture of protocol data is possible using other CDISC models and tools. The CDISC Protocol Representation (PR) Group [ 25 ] developed between 2002 and 2010 a set of more than 300 standardized protocol concepts (such as study objective, phase, design, study population age criteria) that were modeled in the Unified Modeling Language (UML). CDISC formally released this set in 2010 as the Protocol Representation Model v1.0 (PRM). The CDISC PR group included representatives from academic medical centers and pharmaceutical companies, WHO, EMEA and NIH. The resulting PRM model was harmonized with the CDISC common BRIDG model. To illustrate the content of PRM, Appendix A includes an overview of all PRM sections and a subset of elements within each section with links to the corresponding part of the IHC Good Clinical Practice E6 guideline. Due to limited adoption of UML-based tools by principal investigators and study managers [ 26 ], in 2011 CDISC released a subset of the 30 most-used basic protocol concepts [ 27 ] (referred to as the protocol outline ) and provided a Microsoft Word template [ 28 ] that enables protocol writers to enter semi-structured content using a standard and familiar word processor software. The purpose of the protocol outline, also referred to as “Study Synopsis” or “Study Concept Sheet”, is to provide an outline of the key concepts which define the study design prior to commencing authorship of the full protocol and development of individual study CRFs. Supplemental file S2 demonstrates the protocol outline data for the raloxifene trial that uses the provided template. In addition to the template, CDISC provides a link to a Web-based protocol authoring tool [ 18 ] that can generate a protocol outline file in PDF and CSV format. A protocol outline subset includes elements that directly populate the study metadata required for FDA submission (elements of the Trial Summary domain in the Study Data Tabulation Model). Supplemental files 3a, 3b and 3c show the PDF file and two comma separated value files (representing the CDISC Study Data Tabulation Model format) for the raloxifene trial that were created with the Web-based tool. ODM provides very limited support for representing study registration data. Only 7 of the 20 elements defined in the WHO minimum trial metadata [ 8 , 29 , 30 ] can be directly represented in ODM These are: Scientific Title, Countries of Recruitment (via the ODM Location element), Key Inclusion and Exclusion Criteria, and via the user and user type elements: Source of Monetary or Material Support, Primary Sponsor, Secondary Sponsor, Contact for Public Queries and Contact for Scientific Queries). Appendix A shows a complete list of all 20 elements together with their mapping to ClinicalTrials.gov fields. In some cases, the ClinicalTrials.gov registry uses multiple fields to capture a single WHO element. The newly formed CDISC Clinical Trials Registration group (CTR) is working toward an ODM extension that addresses the remaining 13 elements (e.g., target sample size) marked in Appendix A as currently not covered by ODM. An initial attempt to model study registration data was undertaken in 2011 as a joint project of CDISC and HL7. The resulting UML model [ 31 ] is the basis of a new project and working group (Clinical Trials Registry and Results) that is developing a standard format to reconcile all metadata required by ClinicalTrials.gov and EudraCT. The targeted format, scheduled to be published in 2014, will be a BRIDG-compatible, ODM-based XML schema that can be used to electronically exchange registry information between a study sponsor and a registry organization. Until this format is available, an alternative approach to such an exchange is the inclusion of all ClinicalTrial.gov elements using the extension mechanism of ODM. Supplemental file S4 shows an example of implementing this extension-based approach. Because ClinicalTrials.gov registry collects relatively detailed study metadata (either at initial study registration or later deposition of basic summary results), it may be beneficial to import data entered into ClinicalTrials.gov back into local research systems. We have created an eXtensible Stylesheet Language transformation (XSLT) that takes study NCT identifier as input and produces an ODM file with mappable ClinicalTrials.gov data. The XSLT file is available at http://dx.doi.org/10.6084/m9.figshare.1096216 . In contrast to the previous two stages, ODM support for the representation of CRFs and their patient level data is very comprehensive since those were key design requirements during its creation. ODM has a thorough set of XML elements to model CRFs. A single ODM file can represent one or multiple empty CRFs. In ODM, forms are divided into sections ( ItemGroup elements), which consist of individual questions ( Item elements). A form question may reference a pre-defined value set ( CodeList element). For example a value set consisting of choices ‘All of the time’(1), ‘Most of the time’ (2), ‘Some of the time’ (3), ‘A little of the time’ (4), and ‘None of the time’ (5) can be linked to several questions that use the same value set. A CodeList may use integers (shown above), characters or strings as identifiers for individual code list choices. With such a value set definition in place, several questions, such as ‘Did you feel full of life?’ or ‘Have you been a very nervous person?’ can both reference such pre-defined value set. If an ODM file formally defines clinical study intervals ( StudyEvent element; for example, screening or 3 month follow up visit), a form definition may be linked to zero, one or multiple study intervals. While value set standardization within a study is beneficial, of greater importance is standardization across studies, which is potentially achievable with efforts such as National Library of Medicine's Value Set Authority Center [ 32 ] or Common Terminology Services (CTS 2) [ 33 ]. ODM includes an ExternalCodeList element with an optional URL attribute that allows linking to external resources. If such a link is used, however, the standard currently does not specify requirements for the format or functionality of the URL link. A binding guidance (issued by FDA) for regulatory submissions specifies NIH/National Cancer Institute Enterprise Vocabulary Services as a terminology resource. With respect to ODM and case report form, another CDISC initiative called Clinical Data Acquisition Standards Harmonization (CDASH) [ 34 ] aims to provide standardized forms and form elements, and it is released as an ODM XML file. It was first released in 2008 and a new version is expected in 2015. ODM also includes support for branching logic (eg, dynamic forms that display certain questions based on prior form entries or prior patient data); however, ODM does not prescribe use of any specific expression language, such as XPath or JavaScript. This flexible, but incomplete approach limits the form portability across various electronic data capture systems. Supplemental file S5 shows an example ODM file that represents a single form from the raloxifene trial. It was produced by a newly implemented export feature within the intramural electronic data capture system (CTDB). Even though ODM can represent many CRF features relatively well, some vendors of electronic data capture systems (eg, OpenClinica, Formedix, Medidata Solutions or Oracle) and institutions (NIH Intramural Research Program) developed extensions to handle advanced form features. Examples of features covered by extensions are: (1) form fonts or colors (eg, color coding or rich text formatting of certain items in a value set or word emphasis in a question); (2) presence of images on forms (eg, recording exact location of pain using a diagram); (3) addition of explanatory notes linked to particular form questions; or (4) complex rules for checking user-entered values. The current version of ODM focuses mostly on CRF data representation relevant for final submission of study data to a regulator. This keeps the standard relatively simple but requires using extensions which limit portability across systems. Future revisions of the standard may include additional requirements for non-regulated studies or requirements for simple export of empty, but rich formatted forms. Our provided example focuses on a single form; however, ODM also supports representation of multiple forms and proper linkage of individual forms to relevant study events. Table 4 lists several ODM use scenarios that are possible. The six rows of the table show a range of possible ODM data exchange scenarios that range from ranging from a mere protocol description (without any CRFs; rows 1–2) to various forms of CRF exchange (rows 3–5) or a complete study archival scenario (row 6). The columns of the table indicate utilization of different high-level groups of ODM elements utilized in these scenarios. Similarly to the study initiation stage, ODM provides good support for export of patient level data collected via CRFs. Study data in ODM can be in a snapshot form , which is a copy of the data at a given time, or in a more detailed, transactional form that includes full audit trails (prior values for each data point and how it changed over time). Audit trails can, for example, fully accommodate double entry CRFs that offer greater data accuracy. ODM with patient level data can be used to export study data for final analysis ( Fig. 1 ), and for sharing with other researchers and regulators. Prior to 2014, ODM-based data had to be yet further converted into SAS-based XPORT format (.xpt) as specified by the CDISC STDM standard. However, since May 2014, the FDA is exploring possible acceptance of ODM-based submissions (using CDISC DataSet-XML standard which is essentially a subset of ODM) in a pilot project. The ODM format can also be used to transmit study data to interested researchers, as mandated by pioneering journals (such as the British Medical Journal, Annals of Internal Medicine or PLoS Medicine [ 35 ]) or other initiatives, such as the Yoda Project [ 36 ]. ODM can also be used for sharing de-identified patient-level level data, based on the January 2014 statement by EMA to make data available to independent researchers that wish to reanalyze the data or conduct new analyses after the marketing authorization process has completed. Similar development toward data sharing may be expected from FDA based on June 2013 request for comments (Availability of Masked and De-identified Non-Summary Safety and Efficacy Data [ 37 ]. Supplemental file S5 provides an example of patient level data collected within the Surgical Findings CRF (2 made-up patients).

Discussion

The ODM standard is the best baseline standard for representing clinical research study protocol. It natively supports local extensions for data elements not covered by the canonical standard. For data repositories, dealing with a single format for study protocols and study results decreases the development time required to import studies into the repository or to exchange data between systems. ODM constructs for value sets, form sections and study events may also provide guidance for developers of electronic data capture system and contribute to better data structure compatibility across studies. The current ODM version 1.3.2 has limited support for capturing full protocol and study registration data but supports case report form representation well (in the design phase and with patient-level data). These limitations exist mainly because this was not the original use case for creating the standard. Recent initiatives within CDISC, however, try to increase the coverage either via additional ODM extensions or a new version of ODM. To further increase the adoption of ODM as a data exchange standard among research systems, focus on satisfying minimum data exchange needs in addition to data elements required for a regulatory submission is critical. Inclusion of requirements of observational, non-regulated or investigator-initiated studies (outside FDA context) can further improve future versions of the standards. For example, addition of more user roles, study type and phase, capturing full protocol sections, support for rich-text formatting and more restricted syntax for CRF data validation. One of the authors (VH) participates within CDISC CTR&R, PR and XML groups to facilitate such a change. For long-term data storage, an important feature of ODM is the annotation of CRF questions with coded concepts from any terminology (eg, SNOMED or LOINC) or data element definition scheme (eg, Common Data Elements) [ 14 ]. Such annotation can facilitate discovery of relevant data by users of large data repositories. A minor limitation of this annotation feature is that annotations are linked to the whole questions and cannot reference particular parts of the questions. For example, a form question “Are you pregnant or breast feeding?” would be annotated as a whole with two coded concepts as opposed to a more elaborate structure, such as “Are you [pregnant (SNOMED:77386006)] or [breast feeding (SNOMED:69840006)]?”. Another important feature, somewhat outside the scope of protocol representation, is actual use of ODM by healthcare institutions and EHR vendors. The ability to present a research form at the right time inside an EHR system is a feature that has been implemented using ODM [ 38 ]. El Fadly at al. described a RE-USE project that uses ODM metadata message to create an empty research form in the EHR, and the use of an ODM mediator to exchange the form data [ 39 ]. ODM is also the standard of choice used in the Retrieve Form for Data-capture (RFD) profile maintained by the Integrating Healthcare Enterprise (IHE). A related IHE standard, called Retrieve Process for Execution, aims to facilitate automation of protocol related processes specified by the protocol study design. CDISC published that GE, Cerner, Allscripts, Epic, Greenway, Tiani Spirit and eMDs vendors implemented the Retrieve Form for Data capture capabilities; however, there are few or none published reports on its use in recent research trials. The US meaningful use regulations that advocated for greater use of healthcare IT did not include any mandate of EHR systems to interoperate with research systems or provide research capabilities. From an institutional perspective, the benefits of using ODM depend on the overall volume of studies, proportion of studies eventually submitted to FDA and enterprise electronic research systems used at a given site. During study drafting and registration, the overall volume of study data is very limited and repeated entering of the data into various systems is feasible. The size of the protocol data increases greatly in the study initiation phase (empty CRFs specifications) and the motivation for standardization in this phase is greater. Standardization benefits are most apparent for multi-site studies using multiple electronic data capture systems (non-centralized) and for studies later submitting data to the FDA or data sharing platforms. Proper study representation has the greatest value when data are re-analyzed by investigators that have no or limited access to the principal investigator of the original study (eg, the original PI has left the institution). Use of ODM from the study inception greatly simplifies export of study data for submission to the FDA. This approach is widely advocated for and referred to as CDISC-end-to-end approach [ 40 ]. Our analysis of protocol representation using ODM is limited by the fact that we used only a single interventional study and the context was only a single research institution. In selecting the raloxifene trial as a case study, we first identified all terminated interventional trials that contributed CRF data into our institutional data repository and ordered them in descending order by total count of research forms defined. We picked the most complex study for which we were able to secure the permission of the study principal investigator to use his/her study for an informatics analysis. By using a set of several studies with different study designs instead of a single study, we would most likely arrive at a more comprehensive ODM evaluation. In terms of future work, we have explored and hope to continue to explore archiving of intramural research studies using ODM (including, for example, screening protocols that are not subject to mandatory registry submission) and the potential use of ODM in submitting data to trial data sharing platforms. The first author (VH) also participates within the CDISC XML team in several initiatives relevant to protocol representation.

Conclusions

ODM supports a growing number of protocol metadata elements and it offers advantages of a single format for institutions that have to deal with hundreds or thousands of concurrent clinical studies and maintain a repository with data from all completed trials. For each study stage we presented a list of gaps in the standard and necessary vendor or institutional extensions that can compensate for such gaps together with an institution-centric and life-time view of the trial. Protocol representation was outside of ODM's initial focus of exchange of case report form data; however, existing and emerging extensions and use cases will most likely increase support for protocol-level metadata within ODM.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-10-04T09:26:46.659050+00:00