Enhancing reusability of veterinary epidemiological data by creating contextual metadata

preprint OA: closed CC-BY-4.0

Abstract

Abstract Despite the global adoption of the FAIR guiding principles, significant barriers remain in their practical implementation due to the lack of domain-specific standards for "rich metadata". While generic metadata schemes provide basic information, they often fail to capture the complexity of data quality and the nuances of specialized scientific fields. This gap is particularly evident in veterinary epidemiology, a field characterized by complex, multi-scale data spanning various domains where analytical frameworks trend to be prioritized over raw data description, leading to inconsistent reporting and low data reusability. This paper introduces the first community-developed rich metadata guidelines specifically designed for veterinary epidemiology datasets, accompanied by practical templates and real-world examples. By integrating existing standards with specific guidance on domain-specific attributes and data quality, these guidelines provide a comprehensive, single-source framework accessible to researchers across academic, public, and private sectors. They aim to support researchers and stakeholders in enhancing the reusability of animal health data.
Full text 241,213 characters · extracted from preprint-html · click to expand
Enhancing reusability of veterinary epidemiological data by creating contextual metadata | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Enhancing reusability of veterinary epidemiological data by creating contextual metadata Celine Faverjon, Camille Delavenne, Servane Bareille, Angus Cameron, and 18 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9050779/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Despite the global adoption of the FAIR guiding principles, significant barriers remain in their practical implementation due to the lack of domain-specific standards for "rich metadata". While generic metadata schemes provide basic information, they often fail to capture the complexity of data quality and the nuances of specialized scientific fields. This gap is particularly evident in veterinary epidemiology, a field characterized by complex, multi-scale data spanning various domains where analytical frameworks trend to be prioritized over raw data description, leading to inconsistent reporting and low data reusability. This paper introduces the first community-developed rich metadata guidelines specifically designed for veterinary epidemiology datasets, accompanied by practical templates and real-world examples. By integrating existing standards with specific guidance on domain-specific attributes and data quality, these guidelines provide a comprehensive, single-source framework accessible to researchers across academic, public, and private sectors. They aim to support researchers and stakeholders in enhancing the reusability of animal health data. Metadata veterinary epidemiology FAIR guidelines 1. Introduction Over the last decade, a growing number of governments, funding agencies and individuals have promoted data reusability to maximize research investments’ efficiency, enhance reproducibility and transparency, and foster collaboration and innovation. Several tools and guidelines have been developed to support these efforts. The most widely adopted are the FAIR guiding principles (i.e., Findable, Accessible, Interoperable, Reusable) which are intended to, “ act as a guideline for those wishing to enhance the reusability of their data ” 1 , and which have become the reference in many organizations for good research data management practices. The FAIR guiding principles highlight the importance of metadata for supporting data reusability and insist on the importance of having ‘rich metadata’ (i.e., Guiding principles ‘ F2: Data are described with rich metadata’ and ‘R1: meta(data) are richly described with a plurality of accurate and relevant attributes’ ). However, the guiding principles stop short of providing detailed information on what metadata or rich metadata should look like and even pushing the responsibility of deciding which metadata elements are most relevant to each individual scientific community 2 . To help scientists and other stakeholders, several metadata schemes have been developed. These schemes often cover, in some detail, general information relating to digital objects (e.g., title, unique identifier, creators, etc.) useful to meet several of the FAIR guiding principles. The concept of rich metadata is often simply covered by a section entitled, “ Description of the data ”, where the users are invited to provide additional information about the data that did not fit in any of the generic categories they have previously defined (See for examples 3 – 5 ) leaving plenty of room for interpretation. In addition, the existing generic metadata standards usually do not include consideration of data quality, even though having access to information related to dataset flaws has been identified as a critical issue for data (re)usability and interoperability 6 . Having generic specifications makes sense when the purpose of the metadata scheme is to cover a broad range of datasets. However, it also comes with challenges for users when they try to describe their own datasets and associated limitations. The absence of existing specific guidance on dataset flaws has been, for example, identified as a major barrier to the reporting of data quality 7 . Indeed, it can be daunting to identify what information is required or useful for describing a dataset and its quality, often pushing authors to report it inconsistently, or to simply overlook this task 8 . This significantly impairs data reusability as the information needed by the person trying to reuse the data is rarely available or sufficiently understandable, especially when datasets are complex. The difficulty is that what makes sense in terms of data description and data quality changes, depending on the nature of the data involved and the specific domain they are related to. To meet this challenge, more and more communities are developing their own standards to ensure that the data they generate contain the minimum information required to ensure that they can be easily interpreted and therefore reused. For example there are minimum information standards relating to microarray experiments (MIAME) 9 , arthropod abundance 10 , or more recently crops data in agriculture 11 ; but many are still pending (see, for example, a call to action for food composition data 12 and the numerous initiatives in public health such as Sinaci et al. (2020) and Gregory et al. (2023)). In the field of animal health, and in particular in the veterinary epidemiology, the lack of fit-for-purpose standards has been identified as one of the reasons that could explain the overall low level of FAIR-ness of the data used in this domain and likely contributes to limiting data reuse 15 , 16 . Animal health is of key importance not only to ensure sustainable access to safe animal-sourced foods, but also to protect public health as highlighted by the emergence of the ‘One Health’ concept 17 . One of its specificities is that data in veterinary epidemiology are often complex, multi-scale, and span diverse research domains, such as genetics, population dynamics, economics and biology. Indeed, if epidemiology can be defined as the study of disease in populations and of factors that determine its occurrence, veterinary epidemiology goes beyond and includes the study of other health-related factors such as livestock productivity, economics and biosecurity. Veterinary epidemiology is also used in a wide range of species and production systems and operates at the interface of ecology when wild animal populations or vector-borne diseases are being studied. Therefore, the purposes of the data used in veterinary epidemiology and their associated degree of complexity vary greatly. This diversity comes with additional challenges when trying to define standards to support data reuse as these standards must, in this context, meet two competing goals: being consistent and flexible. Work conducted by Kloeze et al. (2012) and Schwantes et al. (2025) has tried to partially fill this gap by defining minimum data standards for laboratory data and wildlife disease research and surveillance, respectively. However, these standards focus on standard data formats (and not metadata) to be used for a very specific kind of veterinary epidemiology data. Guidelines have also been developed to strengthen the reporting of epidemiological studies (see for examples 20 , 21 ).These frameworks focus primarily on describing the analytical framework and its bias, offering only a secondary role to the data. Indeed, the data are assessed exclusively within the context of the study that uses them. Therefore, there remain significant gaps in terms of mechanisms to produce comprehensive enough information to facilitate the broad reuse of the wide array of veterinary epidemiological data. The purpose of this work is to propose comprehensive and practical instructions on how to create domain-relevant rich metadata in veterinary epidemiology to support data reuse in this field. In these guidelines, we mostly focus on elements that are specific to this domain, but, for the sake of completeness, we also briefly cover generic aspects of the metadata, by referring to existing standards. Special attention is also given to the description of the dataset quality, as it has often been found to be poorly documented in the field of animal health 15 , 16 , 22 . These instructions intend to provide guidance to data owners or data holders working with these data, whether they are researchers creating data as part of research studies, or public or private stakeholders involved in animal health and production and generating data on a routine basis. 2. Method Following the work done by Crystal-Ornelas et al. (2022), we initiated a community-centric model that engaged domain scientists to develop formats for common veterinary epidemiology data types. The objective was to create formatting guidelines and templates that would gather the minimum but sufficient metadata necessary for data interpretation and reuse. The guidelines and associated templates were developed following four rounds of reviews by members of the community of veterinary epidemiologists. In total, 25 individuals representing 13 institutions provided inputs at various stages of the development process. This process is detailed below. The team leaders first reviewed existing metadata guidelines and used them as a basis to develop a first draft of the guidelines. The first version was developed as a simple text document, which was initially shared and discussed within a small group of researchers. A second version was then developed and discussed with members of the community during a workshop focusing on Data Stewardship and organized as part of the annual conference of the Society for Veterinary Epidemiology and Preventive Medicine in Berlin in March 2025. This second round of comments led to a third version of the guidelines and to the development of metadata templates. These templates were developed using a tabular strategy to make it easier for others to review and document the metadata, and check errors 24 . The objective was also to avoid a potential lack of engagement of contributors because of the eventual lack of knowledge on machine-readable data formats, as it could be an issue in this community 16 . The tabular templates were proposed in .csv and .xlsx formats, which are data formats often used by veterinary epidemiologists. This third version of the guidelines and associated templates were shared with the members of the core group and one member of the advisory board of a large Horizon 2020 European-commissioned research project named DECIDE (i.e., Data-driven control and prioritization of non-EU-regulated contagious animal diseases), which gathered 19 partners from the public and private sectors across 11 European countries. A fourth version was developed after including their comments and was then shared with the rest of the consortium of the DECIDE project and with interested members of the veterinary epidemiology community identified amongst the participants of the workshop organized in Berlin in March 2025. The contributors piloted the guidelines and associated templates within their research groups using different datasets to ensure they were practical and useful for scientists who collect and reuse data. This last round of reviews gathered only relatively minor feedback and led to the version of the guidelines and their associated templates presented in this paper. The tabular strategy used to develop the templates had the added benefit of being easily extendable later to other metadata formats (e.g., Resource Description Framework (RDF)). Once the guidelines and associated templates were finalized, real life examples were also developed from the DECIDE project to better illustrate how the templates could be filled for two different datasets. The first dataset was related to Polish poultry production data, which had been aggregated and presented a number of data quality issues 25 . The second dataset was related to Norwegian salmonids production data 26 . These two examples are used in the following sections to demonstrate the use of the guidelines. 3. Results 3.1. Data covered by the guidelines These guidelines do not cover all types of data used in veterinary epidemiology. Specific data, such as those used in statistics or molecular epidemiology research, the branch of epidemiology using genetic sequences of pathogens or hosts to describe disease patterns are not considered. Indeed, these specific domains have already made significant advances in terms of data sharing, where standards exist and are already in widespread use (for example, in molecular epidemiology 15 , 27 , 28 and in statistics, with the SDMX initiative 29 ). The focus of these guidelines is on nonspecific data used in veterinary epidemiology, such as data covering animal populations, disease prevalence, antimicrobial usage, presence or absence of risk factors, economic data, production performance, sensor data or animal movements. The datasets described using these guidelines can be made publicly available or simply preserved in private repositories. Indeed, creating good metadata should be considered independently from the action of sharing the data with external people. Whatever the data sharing status, the data should be thoroughly described with rich metadata to facilitate their re-use by the members of the organization owning the data, and by others when applicable. It is also important to highlight the fact that metadata describes a specific digital object. This means that, if only parts of the data maintained in private repositories are made publicly available, and sensitive information, for example, is removed before data sharing, then these data should be considered a new digital object with slightly different characteristics than the original data. Therefore, they should be associated with different metadata than the original data. In practice, their metadata will be largely similar, but will slightly vary in terms of, for example, update frequency, and data anonymization. 3.2. How and when to use these guidelines These guidelines should be consulted before creating metadata to have a good understanding of the information that would be required and how it could be structured. The best time to create metadata is during or before the dataset creation or collection process. Steps should be taken to facilitate data reuse from the very beginning to facilitate the task and save resources. Indeed, even if it is always possible to create metadata a posteriori , it will be a lot more time-consuming than if considered from the very beginning and integrated into the workflow. The metadata created using the guidelines should be stored together with the data they describe in an appropriate repository. Multiplying the repositories where the data and their metadata are deposited is of interest to increase data longevity and maximize data discoverability. The appropriate repositories to be used depend on the nature of the data, and the rules set by each organization owning the data. However, whatever workflow is decided, it should take into account the evolving services provided by the repository, so that all data versions over time are properly linked. In these guidelines, we are proposing structuring the metadata around three components, which are further described in the following sections: General information – this section refers to basic information associated with any digital object and includes data access rules. High-level description – this section intends to provide a basic understanding of what the dataset is about, defining what should be included in the description of a dataset to make it meaningful for others. Specific content description – this section intends to provide a specific content description of the data for someone else to be able to actually (re)use it in an appropriate way 30 . It focuses on detailed information, such as data models, data dictionaries, and data quality (e.g., completeness and accuracy). Note that these guidelines present the items that the metadata cover but also propose metadata templates based on files in .csv and .xlsx format (see appendices). Real life examples of datasets described using the templates are also available in 25 and 26 . 3.3. Section 1: General information For the sake of completeness, our guidelines include the key elements usually reported in the existing metadata schemes, based on the scheme proposed by the DataCite Metadata Working Group (2024). Some of the items highlighted in this section of our guidelines may be found in slightly different versions or under different names, depending on the platform where the dataset is published and/or the generic metadata scheme used as a reference. The list of generic information to be reported about any dataset is presented below. Such information (as well as some elements from Section 2) is generally embedded with the data in formal data repositories and online platforms. However, an example of a template to be filled in order to provide structured general information about a dataset is provided in Appendix 1 in case the dataset is not stored in such repository. Title : A name or title by which a resource is known 4 . A way to identify or refer to the dataset is needed to be able to describe it. It can take the form of a plain language name, which may include the organisation creating the dataset, the purpose or content, and, in the case of a multi-part collection, an identifier of the specific part (e.g., the year or version). Unique Identifier : the title or name of the dataset should be complemented with a unique identifier in particular if the dataset is published in a data catalogue or a data repository. The unique identifier can be a number (e.g., a digital object identifier 1 (DOI®) that is commonly used to reference journal articles) or a unique text identifier based on the name such as used in some data catalogues, like OpenDataSoft 2 . Creators : The main researchers involved in producing the data, and/or the authors of the publication, in priority order, including contact details and when relevant their contribution to the data. For instruments collecting data this is the manufacturer or developer of the instrument 4 . Publisher : The name of the entities that hold, archive, publish, print, distribute, or release the resource 4 . Contributors : The institution or person responsible for managing, distributing, or otherwise contributing to the development of the resource 4 . Publication date : Date when this version of the dataset was made available. Funding reference : Information about financial support (funding) for the resource being registered 4 . Size : Size of the dataset (e.g., bytes, pages, inches) or duration (extent, e.g., hours, minutes, days) of a resource 4 . Version : the version number of the resource or dataset. Language : This is the (human) language or languages that the dataset includes. This does not refer to the encoding system used in the data (see “Format” below). While many datasets may be highly quantitative (containing mostly numbers), interpretation of these values requires data dictionaries with text descriptions. Some datasets may also include free text fields recorded in one or more languages. Keywords or subject : Subject, keyword, classification code, or key phrase describing the resource 4 . Format : Technical format of the resource 4 . This refers to the data format and encoding system used to store the data. It could be a simple XML or CSV file, but other formats could be used, such as OWL/RDF, JSON, BSON, PDF, SQLite, or DB files. Special attention should be given to the rights associated with the dataset. The rights of the datasets aim at reporting the rules for who can access and reuse the data, for what purposes and under what conditions. It should clearly define the roles and responsibilities, as well as the data reuse policy associated with the dataset. More specifically, it should include the following items: Data owner : Data owner means a legal person who, in accordance with applicable law, has the right to grant access to or to share certain personal or non-personal data under its control and is concerned with risk and appropriate access to data. For example, in a surveillance scheme, a data collector can be a laboratory, which stores the surveillance results in a governmental database according to legal framework, and where the data owner is the government. Other key stakeholders : In certain cases, other key stakeholders involved in data access should be reported. For example, the data processor (i.e., a third party that processes personal data on behalf of a data controller) and/or data collector (i.e., natural or legal person responsible for creating or capturing data) could be also mentioned there if deemed necessary. License or agreement : If there is any, the license or data sharing agreement associated with the dataset should be mentioned. If the data are open access, it should be also mentioned there. Access procedures : This item should describe who makes the decision, what is the process, what are the conditions placed on the data reuse, and how the conditions are communicated and enforced. Information about embargo, classified information or presence of sensitive personal data should also be included if applicable. If the data are classified according to certain security standards, the item should also include a reference to the classification model and what class has been assigned to the data according to that model. 3.4. Section 2: General data description The purpose of the high-level overview section is to enable somebody interested in data reuse to determine if the dataset is relevant and potentially usable. It provides basic information to understand what the data are about, focusing on items meaningful for data in animal health. An example of a template to be filled in order to provide a structured high-level description of a dataset used in veterinary epidemiology is provided in Appendix 2. The template includes examples of possible answers, which are highlighted in yellow within the document. 3.4.1. Context Domain : the general domain that the data relate to. It may sometimes overlap with the “keywords or subject” item reported in the previous section, depending on the platform used to publish the data. Examples : A broiler production company may have a production database, and the domain might be ‘ Commercial poultry production’ A veterinary diagnostic laboratory might have data in the domain of ‘ Livestock diagnostic testing’ A government department may maintain a spatial database of aquaculture production leases with the domain of ‘ Registration of aquaculture production sites ’ Original purpose : This is the purpose for which the dataset was initially created. Specifying the purpose of a dataset is useful to understand why and how the data were collected and therefore their potential values and biases. When a dataset has been created based on the integration or analysis of multiple other data sources, it is likely that the purpose of the integrated dataset is different from the purpose of the source datasets. Examples : Diagnostic, research and regulatory specimen test results generated by a diagnostic laboratory in Poland. WOAH WAHIS’s primary purpose may be to provide information on national disease status to prevent international spread of livestock diseases and facilitate trade in animals and animal products. However, it is based on reports from national animal disease surveillance databases, which may have different original purposes such as disease control and surveillance. Scope : This is a definition of what types of data are eligible to be in the dataset, and what types are not. It is related to “What” is included or excluded from the data. In this section, any relevant fact about the data could be also reported. For example, for longitudinal data, reporting important events, which occurred during the time covered by the data and that may have impacted the data could be of interest. Examples : Production data from all the Scottish production sites. Please note that the year YYYY was an unusually severe season of jellyfish, which has likely affected the fish mortality rate and maybe other variables that year. Organic farms and conventional farms with less than 10 pigs are not included in the dataset. Related documents : Any digital object that might be connected to the data and could be useful for data reuse could be also mentioned in this section. For example, scientific papers describing or using the data could be cited there. Note that some of these objects could be also used as reference in other sections of the metadata. 3.4.2. Data resolution This section aims to describe the data dimensionality but focusing only on data in low-dimensional spaces (high number of variables per record), which remain very common in veterinary epidemiology. For data in high-dimensional spaces, the stakeholders would be invited to consider the addition of other attributes to properly describe their data to others. Subject unit identifies the subject unit from which individual data items are drawn. This defines what a row is, or in a non-tabular dataset, what defines a unique entry. This definition is essential to properly use the dataset, count records and understand and identify the duplicates. For flat datasets related to epidemiology, this would refer to the epidemiological unit of interest in the dataset such as an individual animal, herd, or region. If there are multiple subject units (e.g., relational database) connected by a hierarchical structure, it should be carefully highlighted and described there. Detailed information on each subject unit can be also provided directly within the data model (see below). Examples : A dataset related to poultry production may be structured at the flock level with each flock being associated with a specific farm . A dataset focusing on dairy production could have data relating to individuals , which belong to specific groups of animals, which are themselves in different farms. A dataset on antimicrobial usage may contain data extracted from veterinary clinics, but the data have been aggregated by region. In that case, the subject unit would be “ all the veterinary clinics in a given region ”. Spatial scope is the geographical area covered by the data. It may be a country, an administrative subdivision (state, province, county etc.) or indirectly defined (e.g., the area of activity of the production company and all its contracted farms) 31 . Spatial resolution identifies the spatial unit from which individual data items are drawn. This item may also be thought of as the level of precision used to identify a location 31 . Sometimes individual data items are not linked to any spatial information. In that case, it is important to make clear to future users that there is no spatial information available in the data. Examples : A dataset on wildlife density may include one estimate for each 5km-by-5km grid cell. The shapefile of reference for the grid is available in XXXX. Georeferenced farm locations may be recorded in latitude and longitude, (degrees-minutes-seconds format) to the closest whole second, or alternatively, in decimal degrees to the third decimal place. A dataset on antimicrobial usage may contain data extracted from veterinary clinics, but the data have been aggregated by region. In that case, the spatial unit would be “region defined as NUTS 2 according to the NUTS classification released in 2021 https://ec.europa.eu/eurostat/fr/web/nuts/history ”. A dataset on antimicrobial usage may contain the data collected by veterinary clinics, but there is no information on the spatial localization of these establishments. In that case the spatial resolution of this dataset could be stated as “No spatial information available in the data”. Period covered is indicated by the start date and the end date. For an operational data set that continues to capture data, there is no (currently known) end date 31 , but an update frequency should be reported in the next point. Temporal resolution refers to the precision used to measure when events occurred. Commonly, an event is recorded, giving a temporal resolution of one day. Some systems may use a timestamp to record when an event occurred (for example, when data were submitted) with a resolution of milliseconds 31 . Regulatory databases often summarize data into reporting periods, and therefore have figures aggregated by week, month, quarter or year. Update frequency and mechanisms - For operational, continuously growing datasets, data streams and any data receiving regular updates, this item would describe how and when new data are added to the dataset. This may be continuous (with no specified interval) when data are added to the dataset as soon as they are created, or be based on a regular reporting or data processing cycle (daily, weekly, monthly, quarterly or annually (in that case, specifying the period of the year is encouraged)) 31 . If any backwards corrections to the previous versions of the data are made during the update, these should be also noted and clearly described in the data lineage section. Examples : Continuous update: The automatic feeder is recording data every time a cow visits it. Regular reporting or data processing: New data from production sites are automatically extracted once per month. Data from the previous month are also re-extracted at that time as these data could have been modified or corrected by the data collectors. Data of the previous month are therefore overwritten every month with the most up-to-date data. 3.4.3. Quantity of data Number of unique entries – Number of data points or unique entries present in the dataset ideally for each hierarchical level (e.g., number of farms, number of flocks, number of isolates per antibiotic testing and/or the number of raw antibiotic test results). For fixed dataset, the number of records in each component (e.g., sheet or table) should be documented. For continuously growing operational datasets, indicate the number of records present at a specified date instead. Number of new records added or modified per update - For operational, continuously growing datasets and data streams, it is useful to provide an estimate of the number of new records or new inserts added or modified per update. This may be exactly known (e.g., one new submission by each of the 27 EU member states per month ), or estimated (e.g., about 200 new disease outbreak reports per month ). Note that changes in the number of variables, rather than in the number of records, should be mentioned in the data lineage section. 3.4.4. Data lineage Data lineage describes the journey data have taken to arrive at the current data set and includes tracking how it moves between formats, software and location, and what processing or changes occur at each step (Mitchell et al. 2022; Dublin Core™ Metadata Initiative, n.d.). Not only does it explain how the data were captured, but this type of description also provides important information on how to use and interpret the data, as well as any potential quality issues. (e.g., if several data collectors are involved, maybe their data could be handled separately). For a primary dataset, this may be relatively simple but it can be more complex for secondary datasets that have been integrated or derived from the primary sources. The description of data lineage can be communicated in writing and/or by a flow diagram. For secondary datasets, a description of data lineage may be more complex. In this case, data flow diagrams may be particularly useful (see example of data lineage provided with the example dataset from Śmiałek et al. (2025)). In both cases, common considerations to note in the data lineage include: Data collection agent and method - The data collection and transcription method should be specified to ensure that future data users understand the potential biases associated with the data. Indeed, data collection and transcription can be automated (e.g., with sensors) or performed manually. When done manually through human observation, it may be linked to data entry bias and/or inconsistencies (e.g., keying errors such as typographical errors, transposed numbers, and misspelt names). In other cases, such as laboratory test results, knowing the precise laboratory method used is essential to understand the data. Similarly, when the collection agent is a sensor, knowing the actual product and links to the technical description is useful to understand potential bias in the data. Note that this section intends to provide high-level information on how the data were collected. When relevant, more specific details could be provided in the specific content description for each variable. Data processing steps - The number and nature of processing steps that data undergoes may also impact usability and quality and should be detailed as thoroughly as possible. For example, data cleaning and strategies for handling missing or incorrect data (such as replacing it with interpolated values), duplicated entries, and/or data anonymisation can alter the original data in ways that may not be useful for certain reuse purposes. Similarly, if outliers have been removed, the process used to identify the outliers and reasons for exclusion should be documented. When organizations or people are involved in specific data collection or processing steps, attaching their function to the tasks they perform can be useful to allow data users to access additional information in case the information provided in the metadata is not sufficient. For long term preservation of the data, providing the computer code used for the data processing is the best way to share with others how the data processing was performed. In some cases, it may be impossible or impractical to provide full details of original data collection systems. For example, the WOAH WAHIS dataset is based on national animal disease surveillance systems from all member countries. There is a large variety and complexity within the national systems, and in many cases, the details of the internal data lineage may not have been published or analyzed. Even if the process is impossible or impractical to describe, it does not mean that this item should be just left blank. Explicitly informing future data users that the data lineage is particularly complex and cannot be fully described is crucial, as it may impact data quality. Example Farm-level production data may go through the following steps : o Barn staff make regular observations of the animals in each room and routine management activities (feeding, water consumption, environmental conditions, mortality, treatments, etc.) o These observations are recorded on a whiteboard in the barn daily. o The observations are transcribed onto a paper form once a week and transmitted to the office. o The figures are then aggregated across rooms within a barn and entered into an Excel spreadsheet by office staff. o A separate sheet is used for each barn and each production cycle. In this example, the data lineage could be described as follows : o Collection agent and method - All information is collected via human observation, except for the room temperature and humidity which are measured via the appropriate devices. The data are transcribed twice (from whiteboard to form once a day, and from form to spreadsheet once a week). o Processing – Data are aggregated across rooms within a barn and entered into an Excel spreadsheet by office staff. No further data cleaning is done after that. 3.5. Section 3: Specific content description While the high-level overview helps to determine if the data are relevant and potentially useful for a specific purpose, a detailed description is required to reuse the data for the desired purpose. Despite its technical nature and the effort required for its development, this metadata component remains critical. Indeed, comprehensive knowledge of a dataset's content and quality is crucial for its effective usage. An example of a template to be filled is provided in Appendix 3. It is partially based on the data model developed by the European Food Safety Authority (EFSA) as part of the SIGMA project to collect animal population data 33 . However, we modified this data model to make it more generic and therefore suitable for different types of datasets and to add considerations related to data quality. The template includes examples of possible answers, which are highlighted in yellow throughout the document. 3.5.1. Data model The data model proposed in this framework is made of three elements, which are further described below. Components and their relationships Clearly identifying any distinct components of a dataset is the first necessary step to document its content and structure. Each component or table of the dataset should be named so that it can be described afterwards. For example, if the dataset is made of multiple spreadsheets or multiple pages within a spreadsheet file, each of them should have a specific name. Examples of components of a dataset include: multiple spreadsheets or multiple pages within a spreadsheet file tables within a relational database separate Word or PDF documents When there are multiple tables, the relationship between them should be documented, including how they are linked (e.g., presence of unique identifiers). A more complex relational database will require a more comprehensive description of the relationships between the separate tables, but it will consist of the same essential information. For each table, identify which other tables it is related to, and which identifiers are present in both tables to allow them to be linked. Entity-relationship diagrams (ERDs) are often used to capture this information. A complex integrated database may have dozens or hundreds of tables, making an ERD hard to visualise. In this case, it is often more useful to present a series of partial ERDs that include only the relevant tables for a particular module. Example A farm may maintain separate spreadsheets for : o daily observations (feed consumption, mortality, environmental conditions) o production cycle figures (start date, end date, number of animals, average weight at sale, sale price). These sheets are related as they both refer to information about the animals in the barns during a production cycle. The link is established by two identifiers that appear in both spreadsheets : o The production cycle identifier (e.g., the 2nd cycle starting in 2023: 2023-02) o The barn identifier (e.g., barn 4). Using these identifiers, it is possible to link the data between the two spreadsheets to calculate, for example, the average feed conversion ratio for each barn over the cycle. Data elements A data element is defined as a unit of data for which the definition, identification, representation (term used to represent it), and permissible values are specified using a set of attributes 34 . Structured data are often organized in rows and columns, but other approaches may be used (such as key/value pairs grouped into objects or arrays, as used in formats such as JSON or XML). In any case, metadata help to understand and use these data elements (columns of a table, or keys of a structured object). A common approach is to use a data dictionary , which is often organized as a table listing the data elements in each component. The information required includes: Name – This is the name used in the data, be it the fieldname in a database, the column header in a spreadsheet or CSV file, or the key in a key/value pair. For example, “ dateOfSampleSubmission” or “ totalAnimals” . Using a consistent and appropriate naming convention where phrases are written without spaces or punctuation and with capitalized words (e.g., camelCase or PascalCase) is strongly recommended to facilitate readability of the data. Label - This is how the Name used in the data could be translated in plain language. For example, the label associated with the name “ dateOfSampleSubmission” would be “ date of sample submission”. Similarly, the label associated with the name “ totalAnimals” would be “ total number of animals”. Definition – This defines the element identified by a Name and a Label. The definition should be as specific as possible, indicating any relevant reference to be mentioned. For example, in the case of results of a diagnostic test, it may include the exact name of the test used and its associated known sensitivity and specificity. When dates, times, timestamps and durations are used, the standards used to report them should be specified such as the time zone if relevant. Using ISO 8601 standard is usually recommended to report these data (e.g., October 31st, 2025 at 6 p.m. should be represented as 2025-10-31 18:00:00.000), but others could be used. In case the data contains “NA” or other values used to indicate the presence of missing values, the meaning of these special values should be also clearly specified there. In general, it is important to clearly distinguish between unknown and zero values. Examples: dateOfSampleSubmission : the date upon which a sample submission form was completed by the person submitting the sample to the laboratory. The date has a length of ten digits and is presented in the following format: DDMMYYYY (D = day, M=month, Y=year). totalAnimals : the total number of live animals of the specified species, of any age on the specified premises at the date of the visit. Missing values are indicated using − 1 testResult : results obtained with the diagnostic test X (sensitivity: 80%, specificity: 92%) for further information about the diagnostic test, follow link X Type – this indicates the data type of the element. The available data types and terminology vary slightly according to the software and format used, but the main types (and variants) are: Number (integer, float, decimal) Text (string, character, variable-length character) Date (date-time) Logical (Boolean, true/false) Enumerated (a closed list of categorical values nominal or ordinal) Other special field types. Some software may have dedicated field types for specific types of data, such as geographical coordinates, or a universally unique identifier (UUID). When represented in other formats, these values may be represented in one of the main formats listed above (such as text for a UUID, or a pair of float values for coordinates). Additional information is needed for categorical values. Indeed, the meaning of each category should be properly described in a vocabulary (sometimes also called a glossary or data catalogue ). The vocabulary can also be linked to an external reference when a standardized vocabulary is used. A standardized vocabulary or thesaurus is a vocabulary that has been established as a standard within a specific domain or community. When using standardized vocabulary, metadata should include a reference to the existing standard, rather than trying to reproduce it. Indeed, several ontologies or standard vocabularies have already been developed in the animal health sector, including FAANG for the animal genome 35 , FoodOn for the food chain 36 , ICAR the global standard for livestock data 37 , ATCvet code for the classification of veterinary medicines 38 , VetSCT the veterinary extension to SNOMED CT for veterinary health records 39 , AGROVOC for agricultural concepts 40 ; ATOL, EOL and AHOL for livestock traits, environment and health 41 ; AHSO on animal health surveillance 42 ; and DECIDE’s ontology initially developed for aquaculture and later extended to LivestockHealthOntology (LHO) for livestock diseases 43 . None of the existing standards is perfect because they only cover certain concepts and their quality varies (e.g., inconsistent Identifiers or Uniform Resource Identifiers (IDs/URIs), limited mapping to other vocabularies, sparse documentation, weak maintenance). However, reusing or extending them when required, instead of creating new standards, support data interoperability and improve long-term maintenance and comparison across studies. If a new vocabulary / thesaurus must be developed, it should ideally be done using existing standards to ensure interoperability. The key standards of the domain are the ISO 25964, the international standard for thesauri and interoperability with other vocabularies 44 and the Simple Knowledge Organisation System (SKOS), a common data model for sharing controlled vocabularies via the web in a machine-readable format. Examples : An element called “language”, could be associated with the label “the language in which the respondent answered the questionnaire”, and defined as “language of the respondent based on the ISO 3166-1 two-letter language code - https://www.iso.org/iso-3166-country-codes.html “ A spatial unit called “NUTS” could be associated with the label “Nomenclature of territorial units for statistics” and defined as “NUTS as defined in the classification published by Eurostat in 2021 - https://ec.europa.eu/eurostat/web/gisco/geodata/statistical-units/territorial-units-statistics “ History Metadata should include a change log documenting the date an alteration was made and any significant changes in the data collection, processing, storage, structure, dictionary, or vocabularies that may have an impact on the interpretation or analysis of the data. For example, a change in the data dictionary may result in a field having a different definition before and after the change date, which might need to be considered for the analysis. Any information on the impact of that change for future data analysis should be made as explicit as possible. Examples : Before the change date, the indices in column XXX started at 1, while after the indices start at 0. Old data have not been recoded meaning that indices 1 before the change date should be considered equivalent to indices 0 in the data collected after the change date. After the change date, the category “Equidae” from the catalogue XXX was divided into two categories: “horse” and “donkey”. Data collected by the farmer association Y were added to the dataset after the change date. 3.5.2. Data quality Data quality is an important consideration when deciding whether existing data are suitable for reuse for a specific purpose 45 . The first complexity arises from the fact that the definition of “good data quality” depends on the nature of the data being considered and its intended use. Providing an exhaustive description of data quality in the case of generic data reuse is likely out of reach. The purpose of this framework is therefore not to provide a thorough description of what is meant by good data quality, but rather to ensure that the metadata offers at least basic information related to data quality, allowing users to understand the key biases of the data. The second issue comes with the wide variations reported in the list of criteria to be used for data quality assessment and in their definitions, making standardized evaluations of data quality very challenging 46 – 49 . For example, the term “validity” can be used to represent concepts such as “accuracy” or “completeness”, which can also be seen as two distinct concepts. To overcome these issues, our framework focuses solely on generic quality criteria, independent of the task to be performed with the data, and on information not already covered in other sections of the framework. For example, the concept of “Timeliness” was not considered, as it is very specific to the task to be accomplished with the data. Moreover, when “timeliness” is understood as the period covered by the data, the concept is already covered under the item “Data resolution > period covered” of this framework. When understood as the general update frequency of the dataset, it would be covered by the item “Data resolution > Update frequency and mechanisms”. When understood as the time between an event and its actual reporting (e.g., time lag between sampling and analysis by a laboratory), it would be often more appropriate to report it under the description of the corresponding data element (i.e., “Data model > Data element > Definition”). In any case, it does not need to be repeated here. In this framework, data quality is described based on four key items, which are further described below: data completeness, data accuracy, reliability, and quality assurance measures. These categories are not intended to support quantitative scoring, but to structure qualitative reporting. However, if additional activities have been performed to assess other data quality criteria, documenting them in the metadata would be considered of added value and should be encouraged. Such additional information could avoid future users from repeating data quality assessments which have already been performed. Moreover, if the quality of the data has changed at a certain time point in the dataset, it would be also important to highlight it in the description of the data quality. Appendix 3 proposes a template to document data quality information in a standardized way. Completeness Data completeness can be defined as the extent to which data are of sufficient breadth, depth and scope for the task at hand 50 . When the task at hand is unknown, completeness can be defined as a measure of the proportion of expected data that is actually present. It can be evaluated in two ways: Coverage – measures the completeness at the record level . It can also be seen as the representativeness of the data. For example, if there are known to be x units of interest in the population, but the dataset only has data on y units, the coverage is y/x . This may be easy to calculate if reliable data on the population exists (e.g., all the farms from a given region were sampled). However, the true size of the population of interest may not be known. In that case, the coverage may either be unknown or only estimated especially if it changes over time due to the data collection process. Missing data items – measure the completeness at the column or field level . For these items, the proportion of records with missing values should be calculated. Examples: Registered aquaculture sites may be audited periodically. Coverage - an audit dataset may contain audit results on only 20% of the true total number of sites before 01/02/2025, but it reached 70% after that date due to addition of a new data source. Missing data items – within that audit dataset, information about the person in charge of data collection might be missing in 20% of the records. Accuracy Data accuracy or correctness can be defined as the extent to which data are correct and/or certified by a third party (Wang and Strong 1996). It can be also understood as the inverse of the error rate, which provides an estimate of the proportion of the data that contains errors, including invalid or incorrect values. For the purpose of this framework, we restrict the definition of data correctness to invalid and incorrect data that were not removed during the initial data cleaning phase. Invalid values are data that are not consistent with the others in terms of format or content. They are due to data entry errors or system glitches. For example, a capital letter ‘O’ used instead of the digit zero in a field intended to store numbers would be considered as an invalid value. A text value that is not a valid member of the defined code list or vocabulary in a categorical field would also be considered invalid. Well-designed systems with fields proposing pre-typed entries for selection and quality assurance systems (documented in the business rules ) should contain no invalid data. Systems that allow free-text entry of categorical values are, however, prone to typographical, capitalization and punctuation inconsistencies, requiring careful data cleaning before analysis starts. In any case, the proportion of invalid present in the shared data should be documented. If invalid data has been corrected before data sharing, the correction process should be documented in the data history section. Incorrect data are a different type of error where the data may have a valid value, but those values do not reflect the true state. Incorrect data may be due to various issues, including data capture and data entry errors resulting from typographical mistakes or selection of the incorrect categorical value. Identifying incorrect data is more difficult because it often appears to be correct. Quantifying them implies comparing the available data with a known point of truth or biological value, which may not always be easily available. Therefore, in practice, incorrect data can often be just defined as “ unknown values ”. Examples: Registered aquaculture sites may be audited periodically. Invalid data – the dataset contains no invalid data. However, be aware that the data contains three free-text fields to document the evolution of disease events. The content of these fields cannot be validated. Incorrect data – The dataset contains data from Belgium; however, 2% of the latitude and longitudes coordinate pairs reported are located in Poland. In addition, 1% of the cattle appear to have an age largely above the life expectancy of a cow (> 50 years) indicating probable incorrect data. Reliability Reliability provides an estimate of the confidence we can put in the results. In this framework, it covers the concepts of data consistency and data credibility . Data consistency represents the extent to which data are compatible or coherent with similar data (Wang and Strong 1996). This refers to internal and external discrepancies present in the data. Internal discrepancy means that the data present in the dataset are not coherent, but the truth is unknown. For example, partial duplicated entries might be present in the data, but there is no way to identify which entry should be used for analysis. External discrepancy focuses on the comparison between the data present in the dataset and those available elsewhere. In general, external data discrepancy may be difficult to assess in a standardized way, especially for users who lack in-depth knowledge of the data context. However, if such information is available, it should be reported. Data credibility refers to the objective and subjective components of the trustworthiness of data. In that case, biases are suspected but the truth is unknown. Indeed, a data value may be valid, and it may reflect what the generator of the data intended, but it still may not reflect the truth. For example, for human input into a database, some analysis and judgement bias may be involved. Examples : Internal consistency : 10% of the mortality data present internal consistency issues: the percentage of mortality does not align with the category of mortality the flock has been assigned to (e.g., Poultry flock mortality was sometimes reported as 90%, but was at the same time labelled as “very low” in another column). This is an internal discrepancy between these two pieces of information and we are unable to know which one is true. 200.000 fish were reported dead at the end of the month, yet no fish were reported to be present at that location. External consistency : Some latitude and longitude of the cases reported in the data do not match with the official description of the outbreak location. Similarly, some of the symptoms reported in the data are not aligned with the symptoms expected in this production type. In particular, drop in milk production has been reported by 1.2% of the beef farms present in the data. The total volume of slaughtered fish reported by salmon farmers in [Region/Area] for [YYYY–YYYY] is [A] tonnes. For the same period, the volume reported to the Norwegian Tax Administration is [B] tonnes. This discrepancy ([A–B] tonnes; [(A–B)/A × 100]%) indicates an external inconsistency between data sources. Moreover, the dataset contains values that are inconsistent with regulations. For X% of locations, fish are reported present within one month of the end of the previous production cycle, whereas the regulations require a minimum two-month fallowing period between cycles. Credibility: A field that provides clinical diagnosis may be considered relatively reliable when a veterinarian provides the data, but less so when it is provided by a farmer without veterinary training. Similarly, a diagnostic test performed in difficult field conditions could be considered less credible than a test conducted in an accredited laboratory. Quality assurance measures This final section describes any measures in place to ensure the quality of the data. Examples include one-off (for static datasets) or periodic (for operational datasets) data assessment and cleaning, or elements of system design, such as: Business rules or built-in automated data validation steps, which may have been applied to the data. Business rules are a set of rules related to data quality that are applied to the data. They are important to document how to reuse the data properly. For example, a business rule might be that all values above a certain threshold are considered abnormal and are replaced by a constant or set to missing. Whenever possible, business rules should be expressed and shared in a machine-actionable format (e.g., script), or at the minimum with a strict formalism using mathematical, statistical or conditional explicit formulas to ensure they are univocal. Feedback systems which provide the data generator with an opportunity to validate the data contained in the dataset. For example, within online survey systems, the mechanisms put in place to validate the data and ensure internal consistency could be reported here. 4. Discussion This paper presents the first community-developed rich metadata guidelines tailored to veterinary epidemiology datasets with ready-to-use templates and real-life examples. Thanks to the large number of stakeholders involved in the process, the guidelines should be easily understood by the members of the community (i.e. researchers in university but also scientists working in public or private organizations), and target the key information needed to support internal and external data reusability in this domain. They provide a single, comprehensive source of information, so researchers do not need to consult the metadata literature, which is often complex and overly technical for newcomers to the field. If having good metadata increases the value of the collected data, creating metadata primarily benefits the data-owning organization by ensuring institutional memory, making internal data reuse easier and more cost effective, and therefore maximizing the value they get from the data they collect. Therefore, creating rich metadata, even if only in a private repository, should be driven by the direct benefits they provide to the data owner, and not by the willingness to share data with external stakeholders. In the modern era where data are becoming big and complex and essential for running all kinds of operations, relying solely on people’s memory or availability to be able to work with data are a critical risk that public and private organizations should not be taking anymore. Having strong organizational incentives to support the creation of metadata is therefore critical to help individuals engage in the process even if they may not have a personal benefit for doing so. For example, in the public sector, including Data Management Plans (DMPs) requirements in grant contracts as done by numerous funding bodies is a first step, but, this should be associated with a specific allocation of resources to not strip them of their meaning and truly achieve data sustainability 51 . At a smaller scale, it would be also worth requiring metadata creation from students as part of their first projects during their studies and internships to make metadata creation a habit right from the start. This would be part of the development of their data literacy, which must be good enough to allow them later to handle several issues. Developing different career advancement criteria in academia, for example based on a dataset citation index, may also help researchers to dedicate time to this activity. In the private sector, the role of data as a valuable resource to better manage production processes, identify new opportunities or leverage business partnerships is not new 52 , 53 . We can only encourage stakeholders in the animal production business to invest more in good data management practices to be able to fully embrace this opportunity. In some special cases, researchers may want to reuse non-scholarly data for conducting research, but the data owner is not willing to invest in the creation of metadata. In that context, we believe the responsibility of generating the best possible metadata is with the researchers, who are intending to generate and share scientific knowledge based on this data. It is indeed part of their ethical responsibility as scientists to carry out and document their work in a transparent and accurate manner 54 . In this situation, it could be interesting to investigate whether sharing back and cross validating with data owners the metadata created during the research process could be considered as an added benefit, and whether this practice could help improve the accessibility of non-scholarly data for scientists. Data sharing is not a prerequisite to the creation of metadata, and metadata can be published even when data sharing is restricted. Data sharing is indeed often complex especially in the agriculture domain because of privacy and confidentiality issues (see for examples 55 , 56 ). On the other hand, metadata sharing is easier as it is not meant to contain any sensitive information and it still provides a high degree of “FAIRness” even in the absence of FAIR publication of the data itself by facilitating data discovery and including clear rules regarding the process for accessing the data 1 . The work conducted by Fraga-González et al. (2025) presents a number of recommendations on how and where to share metadata that can be applied to the case of data in veterinary epidemiology. We emphasize their recommendations of a two-stage approach using Git to continuously catalog data and documentation, and a well-established research repository like Zenodo ( https://zenodo.org/ ) to describe the Git repository with the catalog and point at its URL. These guidelines intend to facilitate the process of creating rich metadata in veterinary epidemiology, but it is worth noting that it will be always more challenging to use them post-hoc, i.e., after data production, than concurrently with the data production 51 . We therefore cannot stress enough the importance of DMPs to make sure the creation of metadata is embedded in the workflow and is not an extra task raised only at the end of a project. Moreover, it should be clear that these guidelines will not necessarily help improving metadata quality scores as defined by some frameworks such as the one developed by the official portal for European data 57 . Indeed, such frameworks assess the quality of metadata only via a selected list of FAIR guiding principles focusing on concepts that are easy to evaluate in a standardized and domain-agnostic manner. They do not consider the “richness” of metadata and how they can be effectively used to support data reuse, as this task is necessarily domain dependent and a lot more difficult to achieve. The development of domain-specific guidelines such as those proposed in this paper is a step forward to be also able to assess this aspect of metadata quality, but more work would be needed to develop a formal evaluation process. During the development of these guidelines, the researchers often navigated between two competing goals: providing detailed enough guidance while remaining flexible to adapt to a large range of datasets. Some authors would have preferred to go into even more detailed reporting, while others found the proposed list of items too long and overwhelming and may have preferred to prioritize some items. The version presented in this paper was deemed to provide a suitable balance between both objectives. It is important to note that these guidelines intend to be a first version that may evolve over time to better meet the needs of the professionals working in veterinary epidemiology, and stakeholders should feel free to adapt them to their specific or evolving needs. Indeed, additional items may need to be added to make specific kinds of data understandable by others and prevent potential misuse of the data. In general, when it comes to rich metadata, there is never too much information available so adding more should never be considered as an issue. However, to do so, we strongly recommend continuing using a structured approach and building on existing standards to ensure any new information added can be easily retrieved from the metadata by any stakeholders. Moreover, some parts of the guidelines could be extended to be handled in a more standardized manner. For example, the guidance provided in this study to report data quality is a first step but could be further refined to better support good data quality check practices. At this stage, several definitions were kept broad on purpose to cover a large diversity of situations, which may lead to different interpretations and therefore different assessments. For example, assessing internal data consistency in practice is likely to remain challenging in the absence of more detailed information on how to proceed. To allow for such evolution and ensure the sustainability of these guidelines, this version is available in a version control platform ( https://github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines ), which is associated to a long-term data repository ( https://doi.org/10.5281/zenodo.18889286 ). New versions developed should ensure they refer to these first versions and that differences are clearly highlighted to help stakeholders navigate potential multiple versions of the guidelines and associated templates. For inexperienced professionals in our community, the guidelines may appear complex and lengthy at first. However, a more thorough examination will demonstrate that it would be difficult to reuse a dataset without having access to the basic information listed in the guidelines. During the development process, several rounds of discussions and reviews have led to retaining only the critical elements and removing anything that was agreed to be unnecessary. Some compromises had to be made between the authors to lead to these choices. It is expected that similar disagreements will also be raised by other professionals in our community. However, we expect that the potential barriers to the adoption of these guidelines are less related to their content, but rather to the true or perceived lack of resources and incentives to fill the templates. These concerns echo with what has been reported elsewhere when considering barriers to data sharing and the adoption of good data management practices in related domains 8 , 58 – 61 . First, it is important to highlight the fact that none of the information gathered through these guidelines is really new in the sense that anyone who is working with the data should be aware of it. Therefore, the information already exists “somewhere”, the only effort required is to document it once in a standardized and centralized way. Moreover, not every type of data would require the same amount of work to generate metadata. For example, the development of sensor data in modern agricultural systems partially solved the question of metadata creation for these data which usually come with robust data schemas and quality descriptions. However, information related to the context of use of these technologies remains rarely available even if technologies are being developed to fill this gap 62 . In that context, only an ad-hoc collection of contextual metadata would be necessary to ensure proper data reuse as other information should be already made available by the manufacturer of the tool. Then, regarding the lack of resources to generate metadata, there is likely a biased perception that current practices are not costing resources, which has been shown to be largely false at least in the academic world 63 , 64 . As said by Mons (2020): “ if data are treated properly, researchers will have significantly more time to do research ”. There is therefore a need to change the narrative from “the time lost by creating metadata” to “the time saved by creating metadata” in this community. New technologies currently being developed can help address some of the barriers identified to the creation of rich metadata such as the lack of resources. In particular, the integration of artificial intelligence (AI) into metadata management processes has the potential to address many of the challenges faced by traditional systems, particularly in the areas of metadata generation, reuse, and lifecycle management 66 . These technologies can be particularly useful when metadata needs to be created post-hoc and many documents describing the data already exist and can be used as reference to generate rich metadata. Indeed, AI models can only be of use when the information already exists and is accessible in a digital format somewhere. Moreover, if AI models are effective at pointing out statistical anomalies, they are still struggling to identify issues of data quality which are domain-dependent (e.g., differencing a normal from an abnormal value is difficult without a deep contextual understanding of the data). Therefore, our guidelines are unlikely to become obsolete because of AI technologies, and they would rather be used as a source of information by these tools. AI may also help address the fact that rich metadata can probably only be unified up to a certain point. Indeed, our guidelines propose several items to be filled, but the content of each of them remain mostly textual and is likely to vary greatly among datasets depending for example on their type or structure. AI tools could help turn these partially structured data into fully structured data that can be more easily queried while not impairing human usability of the metadata collection templates. If AI provides powerful tools to save resources and support metadata creation, the users always need to keep in mind that such tools may also lack consistency, validation and may compromise data privacy and security 66 . They should therefore be used with great caution. We chose to propose the templates associated with these guidelines in .csv and .xlsx files formats to facilitate their adoption by a community not yet very engaged in the process of making their data reusable. However, these formats are not satisfactory from a machine-readable perspective nor for long-term or large-scale data sharing. They should be considered a first but important step towards greater data FAIRness in the domain of veterinary epidemiology. Indeed, looking at the FAIR continuum proposed by García-Closas et al. (2023), our guidelines mostly focus on steps 1 to 3 (i.e., Metadata exist in a human and basic nonproprietary machine readable format), providing a stepping stone for our community to reach step 4 (i.e., “linked data framework”) in the future. The future development of these guidelines may therefore propose more advanced formats to better support machine-readability such as Resource Description Framework (RDF), JSON-LD or ontology-based systems, which can help automate metadata creation and make them more consistent and reusable. Such a stepwise process has proven successful, for example by the Cattle Barometer, a data-driven decision support tool which was initially developed based on direct data sharing using spreadsheets and CSV files received by email and manually uploaded and cleaned 68 . Later versions transitioned to an ontology-driven pipeline mapping records to the Livestock Health Ontology for creating interoperable and fully automated dashboards 69 , 70 . The workflow is documented in the DECIDE Cattle barometer tutorial 71 . Work on more advanced format for the guidelines has already been started and the first steps are available on various platforms (see for example GitHub 72 ). Our guidelines are a first step in the transition to more advanced data management practices, but to go further, more awareness campaigns and fit-for-purpose training will be needed to increase the overall skills of the community on these questions. Increasing collaboration and relationships with computer scientists may also help to facilitate this transition. Declarations Funding statement Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the granting authority. Neither the European Union nor the granting authority can be held responsible for them. This work was specifically supported by the European Union’s Horizon 2020 research and innovation programme under the grant agreement No 101000494 (Data-driven control and prioritisation of non-EU-regulated contagious animal diseases [DECIDE]) and by the European Partnership on Animal Health and Welfare (EUPAHW) under the grant agreement No 101136346. Author Contribution C.F. conceived the study, developed the theoretical framework and performed the experiments. C.D. aided in the development of the framework and implementation of the experiments. C.F. supervised the study. K.R.D., V.H.S.O. and M.N.O created the salmon case study. C.D. created the poultry case study. G.v.S. acquired the funding and did the project administration. All authors took part of the development of the guidelines and associated templates and provided critical feedback and helped shape the research, analysis and manuscript. Acknowledgement All the people in the different research teams involved tested the templates on their data. Special thank you to Guita Niang from INRAE, Robert Johansson from the Swedish Veterinary Agency. Data Availability All the templates associated with the paper are available on a version control platform [https://github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines](https:/github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines) , which is associated to a long-term data repository ( [https://doi.org/10.5281/zenodo.18889286](https:/doi.org/10.5281/zenodo.18889286) ). References Wilkinson, M. D. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data 3, 160018 (2016). https://doi.org/10.1038/sdata.2016.18 Jacobsen, A. et al. FAIR Principles: Interpretations and Implementation Considerations. Data Intell. 10–29 (2020) doi: 10.1162/dint_r_00024 . Dublin Core ™ Metadata Initiative. Metadata Terms. DataCite Metadata Working Group. DataCite Metadata Schema Documentation for the Publication and Citation of Research Data and Other Research Outputs. (2024) doi: https://doi.org/10.14454/mzv1-5b55 . Data Catalog Vocabulary (DCAT). (2024). OECD. Responding to Societal Challenges with Data: Access, Sharing, Stewardship and Control . vol. 342 https://doi.org/10.1787/2182ce9f-en . (2023). Callahan, T., Barnard, J., Helmkamp, L., Maertens, J. & Kahn, M. Reporting Data Quality Assessment Results: Identifying Individual and Organizational Barriers and Solutions. eGEMs 16 (2017) doi: 10.5334/egems.214 . Tenopir, C. et al. Data sharing, management, use, and reuse: Practices and perceptions of scientists worldwide. PLOS ONE e0229003 (2020) doi: 10.1371/journal.pone.0229003 . Brazma, A. et al. Minimum information about a microarray experiment (MIAME)-toward standards for microarray data. Nat. Genet. 29, 365–371 (2001). DOI: 10.1038/ng1201-365 Rund, S. S. C. et al. MIReAD, a minimum information standard for reporting arthropod abundance data. Sci. Data 6, 40 (2019). https://doi.org/10.1038/s41597-019-0042-5 Tenorio, F. A. M. et al. Filling the agronomic data gap through a minimum data collection approach. Field Crops Res. 308, 109278 (2024). https://doi.org/10.1016/j.fcr.2024.109278 Blumberg, K. L., McKillop, K., Pehrsson, P. R. & Fukagawa, N. K. Call to Action: A Need for Community-Driven Minimum Information Standards for Food Composition Data. Am. J. Clin. Nutr. https://doi.org/10.1016/j.ajcnut.2025.06.027 (2025) doi:10.1016/j.ajcnut.2025.06.027. Sinaci, A. A. et al. From Raw Data to FAIR Data: The FAIRification Workflow for Health Research. Methods Inf. Med. 59, e21–e32 (2020). DOI: 10.1055/s-0040-1713684 Gregory, A. et al. WorldFAIR Project (D7.1) Population Health Data Implementation Guide. https://doi.org/10.5281/zenodo.7887385 (2023) doi:10.5281/zenodo.7887385. Meyer, A., Faverjon, C., Hostens, M., Stegeman, A. & Cameron, A. Systematic review of the status of veterinary epidemiological research in two species regarding the FAIR guiding principles. BMC Vet. Res. 270 (2021) doi: 10.1186/s12917-021-02971-1 . Delavenne, C., Cameron, A., van Schaik, G., Frössling, J. & Faverjon, C. Reusability challenges of livestock production data to improve animal health. Sci. Data 12, (2025) https://doi.org/10.1038/s41597-025-04785-4 Zinsstag, J., Schelling, E., Waltner-Toews, D. & Tanner, M. From “one medicine” to “one health” and systemic approaches to health and well-being. Prev. Vet. Med. 101, 148–156 (2011) doi: 10.1016/j.prevetmed.2010.07.003 Kloeze, H. et al. A minimum data set of animal health laboratory data to allow for collation and analysis across jurisdictions for the purpose of surveillance. Transbound. Emerg. Dis. 59, 264–268 (2012) doi: 10.1111/j.1865-1682.2011.01264.x Schwantes, C. J. et al. A minimum data standard for wildlife disease research and surveillance. Sci. Data 12, 1054 (2025) doi: 10.1038/s41597-025-05332-x O’Connor, A. M. et al. Explanation and Elaboration Document for the STROBE-Vet Statement: Strengthening the Reporting of Observational Studies in Epidemiology – Veterinary Extension. Zoonoses Public Health 63, 662–698 (2016), doi: 10.1111/jvim.14592 Plate, K., Knüppel, S., Hornbacher, A., Greiner, M. & Müller-Graf, C. Development of a tool for rapid assessment of the evidence of human observational epidemiological studies with focus on risk of bias in the context of public health and risk assessment. EFSA Support. Publ. 21, 8925E (2024) https://doi.org/10.2903/sp.efsa. 2024.EN-8925Digital Object Identifier (DOI) McKechnie, I., Raymond, K. & Stacey, D. Identifying Inconsistencies in Data Quality Between FAOSTAT, WOAH, UN Agriculture Census, and National Data | Data Science Journal. https://doi.org/10.5334/dsj-2024-044 (2024) doi:10.5334/dsj-2024-044. Crystal-Ornelas, R. et al. Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats. Sci. Data 9, 700 (2022). https://doi.org/10.1038/s41597-022-01606-w Lortie, C. J., Vargas Poulsen, C., Brun, J. & Kui, L. Tabular strategies for metadata in ecology, evolution, and the environmental sciences. Ecol. Evol. 12, e9245 (2022). https://doi.org/10.1002/ece3. 9245Digital Object Identifier (DOI) Śmiałek, M., Delavenne, C. & Faverjon, C. Production and health performance of a selection of polish poultry broiler flocks (Ross308) between 2018 and 2023. Zenodo https://doi.org/10.5281/zenodo.17406208 (2025). Norwegian Veterinary Institute, Oliveira, V. H. S., Osnes, M. N. & Dean, K. R. Salmonid mortality and losses in Norwegian aquaculture. https://doi.org/10.5281/zenodo.17791861 (2025). Byrd, J. B., Greene, A. C., Prasad, D. V., Jiang, X. & Greene, C. S. Responsible, practical genomic data sharing that accelerates research. Nat. Rev. Genet. 21, 615–629 (2020). https://doi.org/10.1038/s41576-020-0257-5 Cernava, T. et al. Metadata harmonization–Standards are the key for a better usage of omics data for integrative microbiome analysis. Environ. Microbiome 17, 33 (2022). doi: 10.1186/s40793-022-00425-1 SDMX community. https://sdmx.org/ . Fairchild, G. et al. Epidemiological Data Challenges: Planning for a More Robust Future Through Data Standards. Front. Public Health https://doi.org/10.3389/fpubh.2018.00336 (2018) doi:10.3389/fpubh.2018.00336. Zeginis, D., Kalampokis, E., Palma, R., Atkinson, R. & Tarabanis, K. Statistical Challenges Towards a Semantic Meta-Model for Data Integration and Exploitation in Precision Agriculture and Livestock Farming. 7th Int. Workshop Semantic Stat. SemStats 2019 Co-Located 18th Int. Semantic Web Conf. ISWC2019 https://doi.org/10.3233/SW-233156 (2019) doi:10.3233/SW-233156. Mitchell, S. N. et al. FAIR data pipeline: provenance-driven data management for traceable scientific workflows. Philos. Trans. R. Soc. Math. Phys. Eng. Sci. 380, 20210300 (2022). doi: 10.1098/rsta.2021.0300 Aminalragia-Giamini, R. et al. SIGMA - Population Data Model. https://doi.org/10.5281/zenodo.7520891 (2023) doi:10.5281/zenodo.7520891. Catalogue - ORIONKnowledgeHub. https://aginfra.d4science.org:443/web/orionknowledgehub/catalogue?p_p_auth=Rc853tgT&p_p_id=49&p_p_lifecycle=1&p_p_state=normal&p_p_mode=view&_49_struts_action=%2Fmy_sites%2Fview&_49_groupId=92976941&_49_privateLayout=false. Harrison, P. W. et al. FAANG, establishing metadata standards, validation and best practices for the farmed and companion animal community. Anim. Genet. 520–526 (2018) doi: 10.1111/age.12736 . Dooley, D. M. et al. FoodOn: a harmonized food ontology to increase global food traceability, quality control and data integration. Npj Sci. Food 23 (2018) doi: 10.1038/s41538-018-0032-6 . ICAR. Global standard for livestock data. Norwegian Institute of ublic Health. ATCvet code. (2026). College of Veterinary Medicine - Virginia-Maryland. VetSCT. (2025). FAO. AGROVOC Multilingual Thesaurus. (2026). INRAE. Ontologies for Livestock. Dórea, F. C. et al. Drivers for the development of an Animal Health Surveillance Ontology (AHSO). Prev. Vet. Med. 39–48 (2019) doi: 10.1016/j.prevetmed.2019.03.002 . Noor, S. & Hostens, M. Livestock Health Ontology. (2025). ISO. ISO 25964 – the international standard for thesauri and interoperability with other vocabularies. Oliveira, P., Rodrigues, F. & Rangel Henriques, P. A Formal Definition of Data Quality Problems. Proceedings of the 2005 International Conference on Information Quality, ICIQ 2005 (2005). Chen, H., Hailey, D., Wang, N. & Yu, P. A Review of Data Quality Assessment Methods for Public Health Information Systems. Int. J. Environ. Res. Public. Health 11, 5170–5207 (2014). https://doi.org/10.3390/ijerph110505170 Weiskopf, N. G. & Weng, C. Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research. J. Am. Med. Inform. Assoc. 20, 144–151 (2013). DOI: 10.1136/amiajnl-2011-000681 Li, L., Peng, T. & Kennedy, J. A Rule Based Taxonomy of Dirty Data. GSTF Int. J. Comput. (2011). Birkegård, A. C. et al. Building the foundation for veterinary register-based epidemiology: A systematic approach to data quality assessment and validation. Zoonoses Public Health 65, 936–946 (2018). https://doi.org/10.1111/zph.12513 Wang, R. Y. & and Strong, D. M. Beyond Accuracy: What Data Quality Means to Data Consumers. J. Manag. Inf. Syst. 12, 5–33 (1996). https://www.jstor.org/stable/40398176 Fraga-González, G. et al. Affording reusable data: recommendations for researchers from a data-intensive project. Sci. Data 12, 258 (2025). https://doi.org/10.1038/s41597-025-04565-0 Grimaldi, D., Fernandez, V. & Carrasco, C. Exploring data conditions to improve business performance. J. Oper. Res. Soc. 72, 1087–1098 (2021). https://doi.org/10.1080/01605682.2019.1590136 Popovič, A., Hackney, R., Tassabehji, R. & Castelli, M. The impact of big data analytics on firms’ high value business performance. Inf. Syst. Front. 20, 209–222 (2018). https://doi.org/10.1007/s10796-016-9720-4 ALLEA - All European Academies. The European Code of Conduct for Research Integrity . (ALLEA - All European Academies, DE, 2023). Jouanjean, M.-A., Casalini, F., Wiseman, L. & Gray, E. Issues around data governance in the digital transformation of agriculture : The farmers’ perspective. OECD Food Agric. Fish. Pap. https://doi.org/10.1787/53ecf2ab-en (2020) doi:10.1787/53ecf2ab-en. Gavai, A. K. et al. Agricultural data privacy: Emerging platforms & strategies. Food Humanity 4, 100542 (2025). https://doi.org/10.1016/j.foohum.2025.100542 European Data Portal. Metadata Quality Assurance. https://data.europa.eu/mqa/methodology?locale=en (2025). Houtkoop, B. L. et al. Data sharing in psychology: A survey on barriers and preconditions. Adv. Methods Pract. Psychol. Sci. 70–85 (2018) doi: 10.1177/2515245917751886 . Hughes, L. D. et al. Addressing barriers in FAIR data practices for biomedical data. Sci. Data 98 (2023) doi: 10.1038/s41597-023-01969-8 . Perrier, L., Blondal, E. & MacDonald, H. The views, perspectives, and experiences of academic researchers with data sharing and reuse: A meta-synthesis. PloS One (2020). https://doi.org/10.1371/journal.pone.0229182 Rainey, L., Lutomski, J. E. & Broeders, M. J. FAIR data sharing: An international perspective on why medical researchers are lagging behind. Big Data Soc. https://doi.org/10.1177/20539517231171052 (2023) doi:10.1177/20539517231171052. Basir, Md. S., Zhang, Y., Buckmaster, D., Raturi, A. & Krogmeier, J. V. Meta Ag: An automatic agricultural contextual metadata collection app. Smart Agric. Technol. 12, 101073 (2025). https://doi.org/10.1016/j.atech.2025.101073 Badolato, A.-M. Cost-Benefit analysis for FAIR research data – Policy Recommendations. Ouvrir la Science https://www.ouvrirlascience.fr/cost-benefit-analysis-for-fair-research-data-policy-recommendations-2 (2019). European Commission. Cost-Benefit Analysis for FAIR Research Data: Policy Recommendations. https://data.europa.eu/doi/10.2777/706548 (2018). Mons, B. Invest 5% of research funds in ensuring data are reusable. Nature 578, (2020). doi: https://doi.org/10.1038/d41586-020-00505-7 Yang, W., Fu, R., Amin, M. B. & Kang, B. The Impact of Modern AI in Metadata Management. Hum.-Centric Intell. Syst. 5, 323–350 (2025). https://doi.org/10.1007/s44230-025-00106-5 García-Closas, M. et al. Moving Toward Findable, Accessible, Interoperable, Reusable Practices in Epidemiologic Research. Am. J. Epidemiol. 192, 995–1005 (2023). DOI: 10.1093/aje/kwad040 Bokma, J. et al. European veterinary barometer for Bovine Respiratory Diseases : a tool showing diagnostic test results and geolocation of respiratory tract samples from cattle. in European Buiatrics Congress and ECBHM Jubilee Symposium 2023 179–181 (2023). Noor, S. et al. Advancing precision livestock farming through ontology-driven interoperable health data management : extending the Livestock Health Ontology (LHO) for enhanced disease surveillance. in 11th European Conference on Precision Livestock Farming, Proceedings 571–580 (The Organising Committee of the 11th European Conference on Precision Livestock Farming (ECPLF), 2024). Noor, S., Bokma, J., Pardon, B., van Schaik, G. & Hostens, M. Agri semantics : developments to improve data interoperability to support farm information management and decision support systems in agriculture. in Smart farms : improving data-driven decision making in agriculture 75–96 (Burleigh Dodds Science, 2024). doi: 10.19103/as.2023.0132.05 . Bokma, J., Noor, S., Pardon, B. & Hostens, M. Cattle barometer. (2023). VEMO: veterinary epidemiology metadata ontology. (2026). Footnotes https://www.doi.org/ https://public.opendatasoft.com/explore/?sort=modified Additional Declarations No competing interests reported. Supplementary Files Appendix1GeneralInformation.xlsx Appendix2GeneralDataDescription.xlsx Appendix3SpecificDataDescription.xlsx Appendix1GeneralInformation.csv Appendix2GeneralDataDescription.csv Appendix3SpecificDataDescriptionDataQuality.csv Appendix3SpecificDataDescriptionOverview.csv Appendix3SpecificDataDescriptionExampleCatalogueSAMPNT.csv Appendix3SpecificDataDescriptionExampleCatalogueSPECIES.csv Appendix3SpecificDataDescriptionExampleTablelaboratorytest.csv Appendix3SpecificDataDescriptionExampleTablepopulation.csv Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9050779","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":602987506,"identity":"735141a9-cfdc-4bee-8288-465bf5b44ec1","order_by":0,"name":"Celine Faverjon","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJ0lEQVRIie3RMWuDQBTA8Xcc2EXr+qQQv4Li0A4t/SoVIVka6JghFEU4l9BZp36FTmmzJRzE5aBrwQ4thc5K9tJTNKUYO3e4/3Dqkx/eIYBK9S+j4VquZnOPAKP9C+1oiJCGWGFLvGbUEDpMYE9kftiNYICYSRTxmxmgmQbb8mz2Onk84flb9bS0jymQsrruERSbkKcCEF/GQYric7q6G/tRJgqXUaBWtuwRB/2QGwxuQQgPLManD0J3Y4MVRBKNGgeI/S7JF6Atnnc1mTgtuRwkSCQJAZ18QWty1RF/iKCQG9O3iG7OPPnE3dVC87OMFQGjJD50FjPhfKfPz3HE6QfFGbdPdbouK1Zc3Cfxpqz6pPtYs9L64vxM2//zV6T8TVQqlUrV9A3vmGphUq8S8wAAAABJRU5ErkJggg==","orcid":"","institution":"EpiMundi","correspondingAuthor":true,"prefix":"","firstName":"Celine","middleName":"","lastName":"Faverjon","suffix":""},{"id":602987507,"identity":"6d431f8a-fe24-459a-93df-2498c416d43e","order_by":1,"name":"Camille Delavenne","email":"","orcid":"","institution":"EpiMundi","correspondingAuthor":false,"prefix":"","firstName":"Camille","middleName":"","lastName":"Delavenne","suffix":""},{"id":602987508,"identity":"fa59e8d2-0d50-4725-9df2-b983aa62c8bd","order_by":2,"name":"Servane Bareille","email":"","orcid":"","institution":"University of Lyon","correspondingAuthor":false,"prefix":"","firstName":"Servane","middleName":"","lastName":"Bareille","suffix":""},{"id":602987509,"identity":"085cb6c5-41ad-4e77-ab8c-3f034bf3330c","order_by":3,"name":"Angus Cameron","email":"","orcid":"","institution":"EpiMundi","correspondingAuthor":false,"prefix":"","firstName":"Angus","middleName":"","lastName":"Cameron","suffix":""},{"id":602987510,"identity":"07caba72-29d7-4486-83da-3901964e83d9","order_by":4,"name":"Luis P. Carmo","email":"","orcid":"","institution":"Norwegian Veterinary Institute","correspondingAuthor":false,"prefix":"","firstName":"Luis","middleName":"P.","lastName":"Carmo","suffix":""},{"id":602987511,"identity":"ace2bdf2-2681-4fa2-9c6c-11415d743854","order_by":5,"name":"Katharine R. Dean","email":"","orcid":"","institution":"Norwegian Veterinary Institute","correspondingAuthor":false,"prefix":"","firstName":"Katharine","middleName":"R.","lastName":"Dean","suffix":""},{"id":602987512,"identity":"7bdd297e-8ed4-4b82-83e9-3e7b3d604d05","order_by":6,"name":"Marco De Nardi","email":"","orcid":"","institution":"University of Bologna","correspondingAuthor":false,"prefix":"","firstName":"Marco","middleName":"","lastName":"De Nardi","suffix":""},{"id":602987513,"identity":"07cda575-11d0-4f36-85cc-3ae3b9b963ec","order_by":7,"name":"Pauline Ezanno","email":"","orcid":"","institution":"National Research Institute for Agriculture, Food and Environment","correspondingAuthor":false,"prefix":"","firstName":"Pauline","middleName":"","lastName":"Ezanno","suffix":""},{"id":602987514,"identity":"2162ea4d-389a-4776-be83-51094a504181","order_by":8,"name":"Jenny Frössling","email":"","orcid":"","institution":"Swedish Veterinary Agency","correspondingAuthor":false,"prefix":"","firstName":"Jenny","middleName":"","lastName":"Frössling","suffix":""},{"id":602987515,"identity":"dfa451cd-3244-44b1-a6c1-efb225dc446a","order_by":9,"name":"Wiktor Gustafsson","email":"","orcid":"","institution":"Swedish Veterinary Agency","correspondingAuthor":false,"prefix":"","firstName":"Wiktor","middleName":"","lastName":"Gustafsson","suffix":""},{"id":602987516,"identity":"c0261bb3-7e97-45a1-afe9-eccf18316d52","order_by":10,"name":"Miel Hostens","email":"","orcid":"","institution":"Ghent University","correspondingAuthor":false,"prefix":"","firstName":"Miel","middleName":"","lastName":"Hostens","suffix":""},{"id":602987517,"identity":"650d9420-80ac-44f0-9250-258626080dfc","order_by":11,"name":"Anne Meyer","email":"","orcid":"","institution":"Episystemic","correspondingAuthor":false,"prefix":"","firstName":"Anne","middleName":"","lastName":"Meyer","suffix":""},{"id":602987518,"identity":"4e9536cd-a37b-495a-8452-639489e46c04","order_by":12,"name":"Saba Noor","email":"","orcid":"","institution":"Ghent University","correspondingAuthor":false,"prefix":"","firstName":"Saba","middleName":"","lastName":"Noor","suffix":""},{"id":602987519,"identity":"55ececc2-5809-485b-98ae-44b38576487d","order_by":13,"name":"Victor H.S. Oliveira","email":"","orcid":"","institution":"Norwegian Veterinary Institute","correspondingAuthor":false,"prefix":"","firstName":"Victor","middleName":"H.S.","lastName":"Oliveira","suffix":""},{"id":602987520,"identity":"9e083c98-8cb5-4127-81b3-88dd903afb07","order_by":14,"name":"Magnus N. Osnes","email":"","orcid":"","institution":"Norwegian Veterinary Institute","correspondingAuthor":false,"prefix":"","firstName":"Magnus","middleName":"N.","lastName":"Osnes","suffix":""},{"id":602987521,"identity":"22bc7b68-7ca3-4d5a-a6c4-cb7c14b30131","order_by":15,"name":"Sébastien Picault","email":"","orcid":"","institution":"National Research Institute for Agriculture, Food and Environment","correspondingAuthor":false,"prefix":"","firstName":"Sébastien","middleName":"","lastName":"Picault","suffix":""},{"id":602987522,"identity":"9239c2ce-e050-4fe8-b85a-ab610c0d5554","order_by":16,"name":"Crawford W. Revie","email":"","orcid":"","institution":"Food and Agriculture Organization of the United Nations","correspondingAuthor":false,"prefix":"","firstName":"Crawford","middleName":"W.","lastName":"Revie","suffix":""},{"id":602987523,"identity":"dff0ce77-bc1f-4a3a-8518-d74300c92b16","order_by":17,"name":"Javier Sanchez","email":"","orcid":"","institution":"University of Prince Edward Island","correspondingAuthor":false,"prefix":"","firstName":"Javier","middleName":"","lastName":"Sanchez","suffix":""},{"id":602987524,"identity":"ffe26782-c4e0-4d7b-a14b-7303a3eaed7c","order_by":18,"name":"Inge Santman-Berends","email":"","orcid":"","institution":"Royal GD (Netherlands)","correspondingAuthor":false,"prefix":"","firstName":"Inge","middleName":"","lastName":"Santman-Berends","suffix":""},{"id":602987525,"identity":"a73b9486-b3c7-4d02-8a45-e75f1b219812","order_by":19,"name":"Gema Vidal","email":"","orcid":"","institution":"Swedish Veterinary Agency","correspondingAuthor":false,"prefix":"","firstName":"Gema","middleName":"","lastName":"Vidal","suffix":""},{"id":602987526,"identity":"f428af31-828b-43d5-98d1-a087ed6ed909","order_by":20,"name":"Clazien J. de Vos","email":"","orcid":"","institution":"Wageningen University \u0026 Research","correspondingAuthor":false,"prefix":"","firstName":"Clazien","middleName":"J.","lastName":"de Vos","suffix":""},{"id":602987527,"identity":"28efcc59-d0f5-411a-af0d-8246306da327","order_by":21,"name":"Gerdien van Schaik","email":"","orcid":"","institution":"Utrecht University","correspondingAuthor":false,"prefix":"","firstName":"Gerdien","middleName":"van","lastName":"Schaik","suffix":""}],"badges":[],"createdAt":"2026-03-06 12:54:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9050779/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9050779/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106959192,"identity":"41f7638c-5841-4a8f-b2f3-17c10c57bce7","added_by":"auto","created_at":"2026-04-15 08:54:23","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1196322,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/54720eeb-1ea5-4855-9656-7f88b43f8705.pdf"},{"id":104780082,"identity":"7bf7e0b1-e2c4-4ef5-9688-b722ba317654","added_by":"auto","created_at":"2026-03-17 07:50:17","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":13316,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix1GeneralInformation.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/33839600d2612c96822503a9.xlsx"},{"id":104423902,"identity":"ef5877f8-ecb1-4392-980e-ef2350c2e901","added_by":"auto","created_at":"2026-03-11 14:22:35","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":15408,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix2GeneralDataDescription.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/a3380a169dfcd44e7744ac0e.xlsx"},{"id":104780086,"identity":"a80e5381-89cb-4a3c-a5ea-ac40fbd69d3a","added_by":"auto","created_at":"2026-03-17 07:50:20","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":24680,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescription.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/796bd76860a482205ea0475b.xlsx"},{"id":104780161,"identity":"728b7d31-b571-4b1b-b390-df931d963922","added_by":"auto","created_at":"2026-03-17 07:51:08","extension":"csv","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":2455,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix1GeneralInformation.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/2b53414feb531e22d867a0e6.csv"},{"id":104423906,"identity":"75dde8fd-cff6-44d7-9206-5de8611a528b","added_by":"auto","created_at":"2026-03-11 14:22:35","extension":"csv","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":4326,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix2GeneralDataDescription.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/8df8f43e3f7e5388c101df50.csv"},{"id":104423907,"identity":"b0d92b33-2187-4a94-933c-c34471b5a83f","added_by":"auto","created_at":"2026-03-11 14:22:35","extension":"csv","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":3286,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescriptionDataQuality.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/8b5cd098ad3883c011f66e1b.csv"},{"id":104423911,"identity":"f979d3ff-4d81-4e63-9e72-f492e5fbfcc3","added_by":"auto","created_at":"2026-03-11 14:22:35","extension":"csv","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":1538,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescriptionOverview.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/6df930cf2503130e69eb75e1.csv"},{"id":104423910,"identity":"2bf3f4ea-ecdf-49c1-9618-d512c3d61529","added_by":"auto","created_at":"2026-03-11 14:22:35","extension":"csv","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":3331,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescriptionExampleCatalogueSAMPNT.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/0048ef505a0c860f037508db.csv"},{"id":104779991,"identity":"2ecc4a59-4055-4c9b-906a-8d4f655740e7","added_by":"auto","created_at":"2026-03-17 07:48:55","extension":"csv","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":663,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescriptionExampleCatalogueSPECIES.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/b8fbe8520de03b6786d2dccc.csv"},{"id":104423908,"identity":"df442932-a7be-4d2c-850f-75fafa47241e","added_by":"auto","created_at":"2026-03-11 14:22:35","extension":"csv","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":774,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescriptionExampleTablelaboratorytest.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/0043e1e6e8b917c0408b0e73.csv"},{"id":104779907,"identity":"8c9f600f-167e-4cf6-b5a9-78790bffad7c","added_by":"auto","created_at":"2026-03-17 07:47:42","extension":"csv","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":2252,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3SpecificDataDescriptionExampleTablepopulation.csv","url":"https://assets-eu.researchsquare.com/files/rs-9050779/v1/2e199f07bce0820f321df216.csv"}],"financialInterests":"No competing interests reported.","formattedTitle":"Enhancing reusability of veterinary epidemiological data by creating contextual metadata","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eOver the last decade, a growing number of governments, funding agencies and individuals have promoted data reusability to maximize research investments\u0026rsquo; efficiency, enhance reproducibility and transparency, and foster collaboration and innovation. Several tools and guidelines have been developed to support these efforts. The most widely adopted are the FAIR guiding principles (i.e., Findable, Accessible, Interoperable, Reusable) which are intended to, \u0026ldquo;\u003cem\u003eact as a guideline for those wishing to enhance the reusability of their data\u003c/em\u003e\u0026rdquo; \u003csup\u003e1\u003c/sup\u003e, and which have become the reference in many organizations for good research data management practices. The FAIR guiding principles highlight the importance of metadata for supporting data reusability and insist on the importance of having \u0026lsquo;rich metadata\u0026rsquo; (i.e., Guiding principles \u0026lsquo;\u003cem\u003eF2: Data are described with rich metadata\u0026rsquo;\u003c/em\u003e and \u003cem\u003e\u0026lsquo;R1: meta(data) are richly described with a plurality of accurate and relevant attributes\u0026rsquo;\u003c/em\u003e). However, the guiding principles stop short of providing detailed information on what metadata or rich metadata should look like and even pushing the responsibility of deciding which metadata elements are most relevant to each individual scientific community \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTo help scientists and other stakeholders, several metadata schemes have been developed. These schemes often cover, in some detail, general information relating to digital objects (e.g., title, unique identifier, creators, etc.) useful to meet several of the FAIR guiding principles. The concept of rich metadata is often simply covered by a section entitled, \u0026ldquo;\u003cem\u003eDescription of the data\u003c/em\u003e\u0026rdquo;, where the users are invited to provide additional information about the data that did not fit in any of the generic categories they have previously defined (See for examples \u003csup\u003e\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e) leaving plenty of room for interpretation. In addition, the existing generic metadata standards usually do not include consideration of data quality, even though having access to information related to dataset flaws has been identified as a critical issue for data (re)usability and interoperability \u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eHaving generic specifications makes sense when the purpose of the metadata scheme is to cover a broad range of datasets. However, it also comes with challenges for users when they try to describe their own datasets and associated limitations. The absence of existing specific guidance on dataset flaws has been, for example, identified as a major barrier to the reporting of data quality \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Indeed, it can be daunting to identify what information is required or useful for describing a dataset and its quality, often pushing authors to report it inconsistently, or to simply overlook this task \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. This significantly impairs data reusability as the information needed by the person trying to reuse the data is rarely available or sufficiently understandable, especially when datasets are complex. The difficulty is that what makes sense in terms of data description and data quality changes, depending on the nature of the data involved and the specific domain they are related to. To meet this challenge, more and more communities are developing their own standards to ensure that the data they generate contain the minimum information required to ensure that they can be easily interpreted and therefore reused. For example there are minimum information standards relating to microarray experiments (MIAME) \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e, arthropod abundance \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e, or more recently crops data in agriculture \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e; but many are still pending (see, for example, a call to action for food composition data \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e and the numerous initiatives in public health such as Sinaci et al. (2020) and Gregory et al. (2023)).\u003c/p\u003e \u003cp\u003eIn the field of animal health, and in particular in the veterinary epidemiology, the lack of fit-for-purpose standards has been identified as one of the reasons that could explain the overall low level of FAIR-ness of the data used in this domain and likely contributes to limiting data reuse \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. Animal health is of key importance not only to ensure sustainable access to safe animal-sourced foods, but also to protect public health as highlighted by the emergence of the \u0026lsquo;One Health\u0026rsquo; concept \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. One of its specificities is that data in veterinary epidemiology are often complex, multi-scale, and span diverse research domains, such as genetics, population dynamics, economics and biology. Indeed, if epidemiology can be defined as the study of disease in populations and of factors that determine its occurrence, veterinary epidemiology goes beyond and includes the study of other health-related factors such as livestock productivity, economics and biosecurity. Veterinary epidemiology is also used in a wide range of species and production systems and operates at the interface of ecology when wild animal populations or vector-borne diseases are being studied. Therefore, the purposes of the data used in veterinary epidemiology and their associated degree of complexity vary greatly. This diversity comes with additional challenges when trying to define standards to support data reuse as these standards must, in this context, meet two competing goals: being consistent and flexible. Work conducted by Kloeze et al. (2012) and Schwantes et al. (2025) has tried to partially fill this gap by defining minimum data standards for laboratory data and wildlife disease research and surveillance, respectively. However, these standards focus on standard data formats (and not metadata) to be used for a very specific kind of veterinary epidemiology data. Guidelines have also been developed to strengthen the reporting of epidemiological studies (see for examples \u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e).These frameworks focus primarily on describing the analytical framework and its bias, offering only a secondary role to the data. Indeed, the data are assessed exclusively within the context of the study that uses them. Therefore, there remain significant gaps in terms of mechanisms to produce comprehensive enough information to facilitate the broad reuse of the wide array of veterinary epidemiological data.\u003c/p\u003e \u003cp\u003eThe purpose of this work is to propose comprehensive and practical instructions on how to create domain-relevant rich metadata in veterinary epidemiology to support data reuse in this field. In these guidelines, we mostly focus on elements that are specific to this domain, but, for the sake of completeness, we also briefly cover generic aspects of the metadata, by referring to existing standards. Special attention is also given to the description of the dataset quality, as it has often been found to be poorly documented in the field of animal health \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. These instructions intend to provide guidance to data owners or data holders working with these data, whether they are researchers creating data as part of research studies, or public or private stakeholders involved in animal health and production and generating data on a routine basis.\u003c/p\u003e"},{"header":"2. Method","content":"\u003cp\u003eFollowing the work done by Crystal-Ornelas et al. (2022), we initiated a community-centric model that engaged domain scientists to develop formats for common veterinary epidemiology data types. The objective was to create formatting guidelines and templates that would gather the minimum but sufficient metadata necessary for data interpretation and reuse.\u003c/p\u003e \u003cp\u003eThe guidelines and associated templates were developed following four rounds of reviews by members of the community of veterinary epidemiologists. In total, 25 individuals representing 13 institutions provided inputs at various stages of the development process. This process is detailed below.\u003c/p\u003e \u003cp\u003eThe team leaders first reviewed existing metadata guidelines and used them as a basis to develop a first draft of the guidelines. The first version was developed as a simple text document, which was initially shared and discussed within a small group of researchers. A second version was then developed and discussed with members of the community during a workshop focusing on Data Stewardship and organized as part of the annual conference of the Society for Veterinary Epidemiology and Preventive Medicine in Berlin in March 2025.\u003c/p\u003e \u003cp\u003eThis second round of comments led to a third version of the guidelines and to the development of metadata templates. These templates were developed using a tabular strategy to make it easier for others to review and document the metadata, and check errors \u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. The objective was also to avoid a potential lack of engagement of contributors because of the eventual lack of knowledge on machine-readable data formats, as it could be an issue in this community \u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. The tabular templates were proposed in .csv and .xlsx formats, which are data formats often used by veterinary epidemiologists. This third version of the guidelines and associated templates were shared with the members of the core group and one member of the advisory board of a large Horizon 2020 European-commissioned research project named DECIDE (i.e., Data-driven control and prioritization of non-EU-regulated contagious animal diseases), which gathered 19 partners from the public and private sectors across 11 European countries.\u003c/p\u003e \u003cp\u003eA fourth version was developed after including their comments and was then shared with the rest of the consortium of the DECIDE project and with interested members of the veterinary epidemiology community identified amongst the participants of the workshop organized in Berlin in March 2025. The contributors piloted the guidelines and associated templates within their research groups using different datasets to ensure they were practical and useful for scientists who collect and reuse data. This last round of reviews gathered only relatively minor feedback and led to the version of the guidelines and their associated templates presented in this paper. The tabular strategy used to develop the templates had the added benefit of being easily extendable later to other metadata formats (e.g., Resource Description Framework (RDF)).\u003c/p\u003e \u003cp\u003eOnce the guidelines and associated templates were finalized, real life examples were also developed from the DECIDE project to better illustrate how the templates could be filled for two different datasets. The first dataset was related to Polish poultry production data, which had been aggregated and presented a number of data quality issues \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. The second dataset was related to Norwegian salmonids production data \u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. These two examples are used in the following sections to demonstrate the use of the guidelines.\u003c/p\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Data covered by the guidelines\u003c/h2\u003e \u003cp\u003eThese guidelines do not cover all types of data used in veterinary epidemiology. Specific data, such as those used in statistics or molecular epidemiology research, the branch of epidemiology using genetic sequences of pathogens or hosts to describe disease patterns are not considered. Indeed, these specific domains have already made significant advances in terms of data sharing, where standards exist and are already in widespread use (for example, in molecular epidemiology \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e and in statistics, with the SDMX initiative \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e). The focus of these guidelines is on nonspecific data used in veterinary epidemiology, such as data covering animal populations, disease prevalence, antimicrobial usage, presence or absence of risk factors, economic data, production performance, sensor data or animal movements.\u003c/p\u003e \u003cp\u003eThe datasets described using these guidelines can be made publicly available or simply preserved in private repositories. Indeed, creating good metadata should be considered independently from the action of sharing the data with external people. Whatever the data sharing status, the data should be thoroughly described with rich metadata to facilitate their re-use by the members of the organization owning the data, and by others when applicable.\u003c/p\u003e \u003cp\u003eIt is also important to highlight the fact that metadata describes a specific digital object. This means that, if only parts of the data maintained in private repositories are made publicly available, and sensitive information, for example, is removed before data sharing, then these data should be considered a new digital object with slightly different characteristics than the original data. Therefore, they should be associated with different metadata than the original data. In practice, their metadata will be largely similar, but will slightly vary in terms of, for example, update frequency, and data anonymization.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2. How and when to use these guidelines\u003c/h2\u003e \u003cp\u003eThese guidelines should be consulted before creating metadata to have a good understanding of the information that would be required and how it could be structured. The best time to create metadata is during or before the dataset creation or collection process. Steps should be taken to facilitate data reuse from the very beginning to facilitate the task and save resources. Indeed, even if it is always possible to create metadata \u003cem\u003ea posteriori\u003c/em\u003e, it will be a lot more time-consuming than if considered from the very beginning and integrated into the workflow.\u003c/p\u003e \u003cp\u003eThe metadata created using the guidelines should be stored together with the data they describe in an appropriate repository. Multiplying the repositories where the data and their metadata are deposited is of interest to increase data longevity and maximize data discoverability. The appropriate repositories to be used depend on the nature of the data, and the rules set by each organization owning the data. However, whatever workflow is decided, it should take into account the evolving services provided by the repository, so that all data versions over time are properly linked.\u003c/p\u003e \u003cp\u003eIn these guidelines, we are proposing structuring the metadata around three components, which are further described in the following sections:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eGeneral information\u003c/em\u003e \u0026ndash; this section refers to basic information associated with any digital object and includes data access rules.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eHigh-level description\u003c/em\u003e \u0026ndash; this section intends to provide a basic understanding of what the dataset is about, defining what should be included in the description of a dataset to make it meaningful for others.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eSpecific content description \u0026ndash;\u003c/em\u003e this section intends to provide a specific content description of the data for someone else to be able to actually (re)use it in an appropriate way \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. It focuses on detailed information, such as data models, data dictionaries, and data quality (e.g., completeness and accuracy).\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eNote that these guidelines present the items that the metadata cover but also propose metadata templates based on files in .csv and .xlsx format (see appendices). Real life examples of datasets described using the templates are also available in \u003csup\u003e25\u003c/sup\u003e and \u003csup\u003e26\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Section 1: General information\u003c/h2\u003e \u003cp\u003eFor the sake of completeness, our guidelines include the key elements usually reported in the existing metadata schemes, based on the scheme proposed by the DataCite Metadata Working Group (2024). Some of the items highlighted in this section of our guidelines may be found in slightly different versions or under different names, depending on the platform where the dataset is published and/or the generic metadata scheme used as a reference. The list of generic information to be reported about any dataset is presented below. Such information (as well as some elements from Section 2) is generally embedded with the data in formal data repositories and online platforms. However, an example of a template to be filled in order to provide structured general information about a dataset is provided in Appendix 1 in case the dataset is not stored in such repository.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eTitle\u003c/b\u003e: A name or title by which a resource is known \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. A way to identify or refer to the dataset is needed to be able to describe it. It can take the form of a plain language name, which may include the organisation creating the dataset, the purpose or content, and, in the case of a multi-part collection, an identifier of the specific part (e.g., the year or version).\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eUnique Identifier\u003c/b\u003e: the title or name of the dataset should be complemented with a unique identifier in particular if the dataset is published in a data catalogue or a data repository. The unique identifier can be a number (e.g., a digital object identifier\u003csup\u003e1\u003c/sup\u003e (DOI\u0026reg;) that is commonly used to reference journal articles) or a unique text identifier based on the name such as used in some data catalogues, like OpenDataSoft\u003csup\u003e2\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eCreators\u003c/b\u003e: The main researchers involved in producing the data, and/or the authors of the publication, in priority order, including contact details and when relevant their contribution to the data. For instruments collecting data this is the manufacturer or developer of the instrument \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003ePublisher\u003c/b\u003e: The name of the entities that hold, archive, publish, print, distribute, or release the resource \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eContributors\u003c/b\u003e: The institution or person responsible for managing, distributing, or otherwise contributing to the development of the resource \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003ePublication date\u003c/b\u003e: Date when this version of the dataset was made available.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eFunding reference\u003c/b\u003e: Information about financial support (funding) for the resource being registered \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSize\u003c/b\u003e: Size of the dataset (e.g., bytes, pages, inches) or duration (extent, e.g., hours, minutes, days) of a resource \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eVersion\u003c/b\u003e: the version number of the resource or dataset.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eLanguage\u003c/b\u003e: This is the (human) language or languages that the dataset includes. This does not refer to the encoding system used in the data (see \u0026ldquo;Format\u0026rdquo; below). While many datasets may be highly quantitative (containing mostly numbers), interpretation of these values requires data dictionaries with text descriptions. Some datasets may also include free text fields recorded in one or more languages.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eKeywords or subject\u003c/b\u003e: Subject, keyword, classification code, or key phrase describing the resource \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eFormat\u003c/b\u003e: Technical format of the resource \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. This refers to the data format and encoding system used to store the data. It could be a simple XML or CSV file, but other formats could be used, such as OWL/RDF, JSON, BSON, PDF, SQLite, or DB files.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eSpecial attention should be given to the rights associated with the dataset. The rights of the datasets aim at reporting the rules for who can access and reuse the data, for what purposes and under what conditions. It should clearly define the roles and responsibilities, as well as the data reuse policy associated with the dataset. More specifically, it should include the following items:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData owner\u003c/b\u003e: Data owner means a legal person who, in accordance with applicable law, has the right to grant access to or to share certain personal or non-personal data under its control and is concerned with risk and appropriate access to data. For example, in a surveillance scheme, a data collector can be a laboratory, which stores the surveillance results in a governmental database according to legal framework, and where the data owner is the government.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eOther key stakeholders\u003c/b\u003e: In certain cases, other key stakeholders involved in data access should be reported. For example, the data processor (i.e., a third party that processes personal data on behalf of a data controller) and/or data collector (i.e., natural or legal person responsible for creating or capturing data) could be also mentioned there if deemed necessary.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eLicense or agreement\u003c/b\u003e: If there is any, the license or data sharing agreement associated with the dataset should be mentioned. If the data are open access, it should be also mentioned there.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eAccess procedures\u003c/b\u003e: This item should describe who makes the decision, what is the process, what are the conditions placed on the data reuse, and how the conditions are communicated and enforced. Information about embargo, classified information or presence of sensitive personal data should also be included if applicable. If the data are classified according to certain security standards, the item should also include a reference to the classification model and what class has been assigned to the data according to that model.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.4. Section 2: General data description\u003c/h2\u003e \u003cp\u003eThe purpose of the high-level overview section is to enable somebody interested in data reuse to determine if the dataset is relevant and potentially usable. It provides basic information to understand what the data are about, focusing on items meaningful for data in animal health.\u003c/p\u003e \u003cp\u003eAn example of a template to be filled in order to provide a structured high-level description of a dataset used in veterinary epidemiology is provided in Appendix 2. The template includes examples of possible answers, which are highlighted in yellow within the document.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e3.4.1. Context\u003c/h2\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eDomain\u003c/b\u003e: the general domain that the data relate to. It may sometimes overlap with the \u0026ldquo;keywords or subject\u0026rdquo; item reported in the previous section, depending on the platform used to publish the data.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA broiler production company may have a production database, and the domain might be \u0026lsquo;\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003eCommercial poultry production\u0026rsquo;\u003c/span\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA veterinary diagnostic laboratory might have data in the domain of \u0026lsquo;\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003eLivestock diagnostic testing\u0026rsquo;\u003c/span\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA government department may maintain a spatial database of aquaculture production leases with the domain of \u0026lsquo;\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003eRegistration of aquaculture production sites\u003c/span\u003e \u003cem\u003e\u0026rsquo;\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eOriginal purpose\u003c/b\u003e: This is the purpose for which the dataset was initially created. Specifying the purpose of a dataset is useful to understand why and how the data were collected and therefore their potential values and biases. When a dataset has been created based on the integration or analysis of multiple other data sources, it is likely that the purpose of the integrated dataset is different from the purpose of the source datasets.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eDiagnostic, research and regulatory specimen test results generated by a diagnostic laboratory in Poland.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eWOAH WAHIS\u0026rsquo;s primary purpose may be to provide information on national disease status to prevent international spread of livestock diseases and facilitate trade in animals and animal products. However, it is based on reports from national animal disease surveillance databases, which may have different original purposes such as disease control and surveillance.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eScope\u003c/b\u003e: This is a definition of what types of data are eligible to be in the dataset, and what types are not. It is related to \u0026ldquo;What\u0026rdquo; is included or excluded from the data. In this section, any relevant fact about the data could be also reported. For example, for longitudinal data, reporting important events, which occurred during the time covered by the data and that may have impacted the data could be of interest.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eProduction data from all the Scottish production sites. Please note that the year YYYY was an unusually severe season of jellyfish, which has likely affected the fish mortality rate and maybe other variables that year.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eOrganic farms and conventional farms with less than 10 pigs are not included in the dataset.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eRelated documents\u003c/b\u003e: Any digital object that might be connected to the data and could be useful for data reuse could be also mentioned in this section. For example, scientific papers describing or using the data could be cited there. Note that some of these objects could be also used as reference in other sections of the metadata.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.4.2. Data resolution\u003c/h2\u003e \u003cp\u003eThis section aims to describe the data dimensionality but focusing only on data in low-dimensional spaces (high number of variables per record), which remain very common in veterinary epidemiology. For data in high-dimensional spaces, the stakeholders would be invited to consider the addition of other attributes to properly describe their data to others.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSubject unit\u003c/b\u003e identifies the subject unit from which individual data items are drawn. This defines what a row is, or in a non-tabular dataset, what defines a unique entry. This definition is essential to properly use the dataset, count records and understand and identify the duplicates. For flat datasets related to epidemiology, this would refer to the epidemiological unit of interest in the dataset such as an individual animal, herd, or region. If there are multiple subject units (e.g., relational database) connected by a hierarchical structure, it should be carefully highlighted and described there. Detailed information on each subject unit can be also provided directly within the data model (see below).\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA dataset related to poultry production may be structured at the\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003eflock level\u003c/span\u003e \u003cem\u003ewith each flock being associated with a specific\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003efarm\u003c/span\u003e. \u003cem\u003eA dataset focusing on dairy production could have data relating to\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003eindividuals\u003c/span\u003e, \u003cem\u003ewhich belong to specific\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003egroups\u003c/span\u003e \u003cem\u003eof animals, which are themselves in different\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003efarms.\u003c/span\u003e\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA dataset on antimicrobial usage may contain data extracted from veterinary clinics, but the data have been aggregated by region. In that case, the subject unit would be \u0026ldquo;\u003c/em\u003e \u003cspan type=\"ItalicUnderline\" class=\"ItalicUnderline\" name=\"Emphasis\"\u003eall the veterinary clinics in a given region\u003c/span\u003e \u003cem\u003e\u0026rdquo;.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSpatial scope\u003c/b\u003e is the geographical area covered by the data. It may be a country, an administrative subdivision (state, province, county etc.) or indirectly defined (e.g., the area of activity of the production company and all its contracted farms) \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSpatial resolution\u003c/b\u003e identifies the spatial unit from which individual data items are drawn. This item may also be thought of as the level of precision used to identify a location \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Sometimes individual data items are not linked to any spatial information. In that case, it is important to make clear to future users that there is no spatial information available in the data.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA dataset on wildlife density may include one estimate for each 5km-by-5km grid cell. The shapefile of reference for the grid is available in XXXX.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eGeoreferenced farm locations may be recorded in latitude and longitude, (degrees-minutes-seconds format) to the closest whole second, or alternatively, in decimal degrees to the third decimal place.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA dataset on antimicrobial usage may contain data extracted from veterinary clinics, but the data have been aggregated by region. In that case, the spatial unit would be \u0026ldquo;region defined as NUTS 2 according to the NUTS classification released in 2021\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ec.europa.eu/eurostat/fr/web/nuts/history\u003c/span\u003e\u003cspan address=\"https://ec.europa.eu/eurostat/fr/web/nuts/history\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cem\u003e\u0026rdquo;.\u003c/em\u003e\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA dataset on antimicrobial usage may contain the data collected by veterinary clinics, but there is no information on the spatial localization of these establishments. In that case the spatial resolution of this dataset could be stated as \u0026ldquo;No spatial information available in the data\u0026rdquo;.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003ePeriod covered\u003c/b\u003e is indicated by the start date and the end date. For an operational data set that continues to capture data, there is no (currently known) end date \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e, but an update frequency should be reported in the next point.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eTemporal resolution\u003c/b\u003e refers to the precision used to measure when events occurred. Commonly, an event is recorded, giving a temporal resolution of one day. Some systems may use a timestamp to record when an event occurred (for example, when data were submitted) with a resolution of milliseconds \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Regulatory databases often summarize data into reporting periods, and therefore have figures aggregated by week, month, quarter or year.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eUpdate frequency and mechanisms -\u003c/b\u003e For operational, continuously growing datasets, data streams and any data receiving regular updates, this item would describe how and when new data are added to the dataset. This may be continuous (with no specified interval) when data are added to the dataset as soon as they are created, or be based on a regular reporting or data processing cycle (daily, weekly, monthly, quarterly or annually (in that case, specifying the period of the year is encouraged)) \u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. If any backwards corrections to the previous versions of the data are made during the update, these should be also noted and clearly described in the data lineage section.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eContinuous update: The automatic feeder is recording data every time a cow visits it.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eRegular reporting or data processing: New data from production sites are automatically extracted once per month. Data from the previous month are also re-extracted at that time as these data could have been modified or corrected by the data collectors. Data of the previous month are therefore overwritten every month with the most up-to-date data.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e3.4.3. Quantity of data\u003c/h2\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eNumber of unique entries\u003c/b\u003e \u0026ndash; Number of data points or unique entries present in the dataset ideally for each hierarchical level (e.g., number of farms, number of flocks, number of isolates per antibiotic testing and/or the number of raw antibiotic test results). For fixed dataset, the number of records in each component (e.g., sheet or table) should be documented. For continuously growing operational datasets, indicate the number of records present at a specified date instead.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eNumber of new records added or modified per update -\u003c/b\u003e For operational, continuously growing datasets and data streams, it is useful to provide an estimate of the number of new records or new inserts added or modified per update. This may be exactly known (e.g., \u003cem\u003eone new submission by each of the 27 EU member states per month\u003c/em\u003e), or estimated (e.g., \u003cem\u003eabout 200 new disease outbreak reports per month\u003c/em\u003e). Note that changes in the number of variables, rather than in the number of records, should be mentioned in the data lineage section.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e3.4.4. Data lineage\u003c/h2\u003e \u003cp\u003eData lineage describes the journey data have taken to arrive at the current data set and includes tracking how it moves between formats, software and location, and what processing or changes occur at each step (Mitchell et al. 2022; Dublin Core\u0026trade; Metadata Initiative, n.d.). Not only does it explain how the data were captured, but this type of description also provides important information on how to use and interpret the data, as well as any potential quality issues. (e.g., if several data collectors are involved, maybe their data could be handled separately). For a primary dataset, this may be relatively simple but it can be more complex for secondary datasets that have been integrated or derived from the primary sources.\u003c/p\u003e \u003cp\u003eThe description of data lineage can be communicated in writing and/or by a flow diagram. For secondary datasets, a description of data lineage may be more complex. In this case, data flow diagrams may be particularly useful (see example of data lineage provided with the example dataset from Śmiałek et al. (2025)). In both cases, common considerations to note in the data lineage include:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData collection agent and method\u003c/b\u003e \u003cem\u003e-\u003c/em\u003e The data collection and transcription method should be specified to ensure that future data users understand the potential biases associated with the data. Indeed, data collection and transcription can be automated (e.g., with sensors) or performed manually. When done manually through human observation, it may be linked to data entry bias and/or inconsistencies (e.g., keying errors such as typographical errors, transposed numbers, and misspelt names). In other cases, such as laboratory test results, knowing the precise laboratory method used is essential to understand the data. Similarly, when the collection agent is a sensor, knowing the actual product and links to the technical description is useful to understand potential bias in the data. Note that this section intends to provide high-level information on how the data were collected. When relevant, more specific details could be provided in the specific content description for each variable.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData processing steps\u003c/b\u003e - The number and nature of processing steps that data undergoes may also impact usability and quality and should be detailed as thoroughly as possible. For example, data cleaning and strategies for handling missing or incorrect data (such as replacing it with interpolated values), duplicated entries, and/or data anonymisation can alter the original data in ways that may not be useful for certain reuse purposes. Similarly, if outliers have been removed, the process used to identify the outliers and reasons for exclusion should be documented.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWhen organizations or people are involved in specific data collection or processing steps, attaching their function to the tasks they perform can be useful to allow data users to access additional information in case the information provided in the metadata is not sufficient. For long term preservation of the data, providing the computer code used for the data processing is the best way to share with others how the data processing was performed.\u003c/p\u003e \u003cp\u003eIn some cases, it may be impossible or impractical to provide full details of original data collection systems. For example, the WOAH WAHIS dataset is based on national animal disease surveillance systems from all member countries. There is a large variety and complexity within the national systems, and in many cases, the details of the internal data lineage may not have been published or analyzed. Even if the process is impossible or impractical to describe, it does not mean that this item should be just left blank. Explicitly informing future data users that the data lineage is particularly complex and cannot be fully described is crucial, as it may impact data quality.\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eExample\u003c/strong\u003e \u003cp\u003e \u003cem\u003eFarm-level production data may go through the following steps\u003c/em\u003e:\u003c/p\u003e \u003cp\u003eo \u003cem\u003eBarn staff make regular observations of the animals in each room and routine management activities (feeding, water consumption, environmental conditions, mortality, treatments, etc.)\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eThese observations are recorded on a whiteboard in the barn daily.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eThe observations are transcribed onto a paper form once a week and transmitted to the office.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eThe figures are then aggregated across rooms within a barn and entered into an Excel spreadsheet by office staff.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eA separate sheet is used for each barn and each production cycle.\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cem\u003eIn this example, the data lineage could be described as follows\u003c/em\u003e:\u003c/p\u003e \u003cp\u003eo \u003cem\u003eCollection agent and method - All information is collected via human observation, except for the room temperature and humidity which are measured via the appropriate devices. The data are transcribed twice (from whiteboard to form once a day, and from form to spreadsheet once a week).\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eProcessing \u0026ndash; Data are aggregated across rooms within a barn and entered into an Excel spreadsheet by office staff. No further data cleaning is done after that.\u003c/em\u003e\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.5. Section 3: Specific content description\u003c/h2\u003e \u003cp\u003eWhile the high-level overview helps to determine if the data are relevant and potentially useful for a specific purpose, a detailed description is required to reuse the data for the desired purpose. Despite its technical nature and the effort required for its development, this metadata component remains critical. Indeed, comprehensive knowledge of a dataset's content and quality is crucial for its effective usage.\u003c/p\u003e \u003cp\u003eAn example of a template to be filled is provided in Appendix 3. It is partially based on the data model developed by the European Food Safety Authority (EFSA) as part of the SIGMA project to collect animal population data \u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. However, we modified this data model to make it more generic and therefore suitable for different types of datasets and to add considerations related to data quality. The template includes examples of possible answers, which are highlighted in yellow throughout the document.\u003c/p\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e3.5.1. Data model\u003c/h2\u003e \u003cp\u003eThe data model proposed in this framework is made of three elements, which are further described below.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eComponents and their relationships\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eClearly identifying any distinct components of a dataset is the first necessary step to document its content and structure. Each component or table of the dataset should be named so that it can be described afterwards. For example, if the dataset is made of multiple spreadsheets or multiple pages within a spreadsheet file, each of them should have a specific name.\u003c/p\u003e \u003cp\u003eExamples of components of a dataset include:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003emultiple spreadsheets or multiple pages within a spreadsheet file\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003etables within a relational database\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eseparate Word or PDF documents\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWhen there are multiple tables, the relationship between them should be documented, including how they are linked (e.g., presence of unique identifiers). A more complex relational database will require a more comprehensive description of the relationships between the separate tables, but it will consist of the same essential information. For each table, identify which other tables it is related to, and which identifiers are present in both tables to allow them to be linked. Entity-relationship diagrams (ERDs) are often used to capture this information. A complex integrated database may have dozens or hundreds of tables, making an ERD hard to visualise. In this case, it is often more useful to present a series of partial ERDs that include only the relevant tables for a particular module.\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eExample\u003c/strong\u003e \u003cp\u003e \u003cem\u003eA farm may maintain separate spreadsheets for\u003c/em\u003e:\u003c/p\u003e \u003cp\u003eo \u003cem\u003edaily observations (feed consumption, mortality, environmental conditions)\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eproduction cycle figures (start date, end date, number of animals, average weight at sale, sale price).\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cem\u003eThese sheets are related as they both refer to information about the animals in the barns during a production cycle. The link is established by two identifiers that appear in both spreadsheets\u003c/em\u003e:\u003c/p\u003e \u003cp\u003eo \u003cem\u003eThe production cycle identifier (e.g., the 2nd cycle starting in 2023: 2023-02)\u003c/em\u003e\u003c/p\u003e \u003cp\u003eo \u003cem\u003eThe barn identifier (e.g., barn 4).\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cem\u003eUsing these identifiers, it is possible to link the data between the two spreadsheets to calculate, for example, the average feed conversion ratio for each barn over the cycle.\u003c/em\u003e \u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData elements\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eA data element is defined as a unit of data for which the definition, identification, representation (term used to represent it), and permissible values are specified using a set of attributes \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. Structured data are often organized in rows and columns, but other approaches may be used (such as key/value pairs grouped into objects or arrays, as used in formats such as JSON or XML). In any case, metadata help to understand and use these data elements (columns of a table, or keys of a structured object).\u003c/p\u003e \u003cp\u003eA common approach is to use a \u003cem\u003edata dictionary\u003c/em\u003e, which is often organized as a table listing the data elements in each component. The information required includes:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eName\u003c/em\u003e \u0026ndash; This is the name used in the data, be it the fieldname in a database, the column header in a spreadsheet or CSV file, or the key in a key/value pair. For example, \u0026ldquo;\u003cem\u003edateOfSampleSubmission\u0026rdquo;\u003c/em\u003e or \u0026ldquo;\u003cem\u003etotalAnimals\u0026rdquo;\u003c/em\u003e. Using a consistent and appropriate naming convention where phrases are written without spaces or punctuation and with capitalized words (e.g., camelCase or PascalCase) is strongly recommended to facilitate readability of the data.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eLabel\u003c/em\u003e \u003cb\u003e-\u003c/b\u003e This is how the Name used in the data could be translated in plain language. For example, the label associated with the name \u0026ldquo;\u003cem\u003edateOfSampleSubmission\u0026rdquo;\u003c/em\u003e would be \u0026ldquo;\u003cem\u003edate of sample submission\u0026rdquo;.\u003c/em\u003e Similarly, the label associated with the name \u0026ldquo;\u003cem\u003etotalAnimals\u0026rdquo;\u003c/em\u003e would be \u0026ldquo;\u003cem\u003etotal number of animals\u0026rdquo;.\u003c/em\u003e\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eDefinition\u003c/em\u003e \u0026ndash; This defines the element identified by a Name and a Label. The definition should be as specific as possible, indicating any relevant reference to be mentioned. For example, in the case of results of a diagnostic test, it may include the exact name of the test used and its associated known sensitivity and specificity. When dates, times, timestamps and durations are used, the standards used to report them should be specified such as the time zone if relevant. Using ISO 8601 standard is usually recommended to report these data (e.g., October 31st, 2025 at 6 p.m. should be represented as 2025-10-31 18:00:00.000), but others could be used. In case the data contains \u0026ldquo;NA\u0026rdquo; or other values used to indicate the presence of missing values, the meaning of these special values should be also clearly specified there. In general, it is important to clearly distinguish between unknown and zero values.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eExamples:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003edateOfSampleSubmission\u003c/em\u003e: the date upon which a sample submission form was completed by the person submitting the sample to the laboratory. The date has a length of ten digits and is presented in the following format: DDMMYYYY (D\u0026thinsp;=\u0026thinsp;day, M=month, Y=year).\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003etotalAnimals\u003c/em\u003e: the total number of live animals of the specified species, of any age on the specified premises at the date of the visit. Missing values are indicated using\u0026thinsp;\u0026minus;\u0026thinsp;1\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003etestResult\u003c/em\u003e: results obtained with the diagnostic test X (sensitivity: 80%, specificity: 92%) for further information about the diagnostic test, follow link X\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eType\u003c/em\u003e \u0026ndash; this indicates the data type of the element. The available data types and terminology vary slightly according to the software and format used, but the main types (and variants) are:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eNumber (integer, float, decimal)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eText (string, character, variable-length character)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDate (date-time)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eLogical (Boolean, true/false)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eEnumerated (a closed list of categorical values nominal or ordinal)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eOther special field types. Some software may have dedicated field types for specific types of data, such as geographical coordinates, or a universally unique identifier (UUID). When represented in other formats, these values may be represented in one of the main formats listed above (such as text for a UUID, or a pair of float values for coordinates).\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eAdditional information is needed for categorical values. Indeed, the meaning of each category should be properly described in a \u003cem\u003evocabulary\u003c/em\u003e (sometimes also called a \u003cem\u003eglossary\u003c/em\u003e or \u003cem\u003edata catalogue\u003c/em\u003e). The vocabulary can also be linked to an external reference when a standardized vocabulary is used. A \u003cem\u003estandardized vocabulary\u003c/em\u003e or \u003cem\u003ethesaurus\u003c/em\u003e is a vocabulary that has been established as a standard within a specific domain or community. When using standardized vocabulary, metadata should include a reference to the existing standard, rather than trying to reproduce it. Indeed, several ontologies or standard vocabularies have already been developed in the animal health sector, including FAANG for the animal genome \u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e, FoodOn for the food chain \u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e, ICAR the global standard for livestock data \u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e, ATCvet code for the classification of veterinary medicines \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e, VetSCT the veterinary extension to SNOMED CT for veterinary health records \u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e, AGROVOC for agricultural concepts \u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e; ATOL, EOL and AHOL for livestock traits, environment and health \u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e; AHSO on animal health surveillance \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e; and DECIDE\u0026rsquo;s ontology initially developed for aquaculture and later extended to LivestockHealthOntology (LHO) for livestock diseases \u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. None of the existing standards is perfect because they only cover certain concepts and their quality varies (e.g., inconsistent Identifiers or Uniform Resource Identifiers (IDs/URIs), limited mapping to other vocabularies, sparse documentation, weak maintenance). However, reusing or extending them when required, instead of creating new standards, support data interoperability and improve long-term maintenance and comparison across studies. If a new vocabulary / thesaurus must be developed, it should ideally be done using existing standards to ensure interoperability. The key standards of the domain are the ISO 25964, the international standard for thesauri and interoperability with other vocabularies \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e and the Simple Knowledge Organisation System (SKOS), a common data model for sharing controlled vocabularies via the web in a machine-readable format.\u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eAn element called \u0026ldquo;language\u0026rdquo;, could be associated with the label \u0026ldquo;the language in which the respondent answered the questionnaire\u0026rdquo;, and defined as \u0026ldquo;language of the respondent based on the ISO 3166-1 two-letter language code -\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.iso.org/iso-3166-country-codes.html\u003c/span\u003e\u003cspan address=\"https://www.iso.org/iso-3166-country-codes.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e \u003cem\u003e\u0026ldquo;\u003c/em\u003e\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eA spatial unit called \u0026ldquo;NUTS\u0026rdquo; could be associated with the label \u0026ldquo;Nomenclature of territorial units for statistics\u0026rdquo; and defined as \u0026ldquo;NUTS as defined in the classification published by Eurostat in 2021 -\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ec.europa.eu/eurostat/web/gisco/geodata/statistical-units/territorial-units-statistics\u003c/span\u003e\u003cspan address=\"https://ec.europa.eu/eurostat/web/gisco/geodata/statistical-units/territorial-units-statistics\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e \u003cem\u003e\u0026ldquo;\u003c/em\u003e\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eHistory\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eMetadata should include a change log documenting the date an alteration was made and any significant changes in the data collection, processing, storage, structure, dictionary, or vocabularies that may have an impact on the interpretation or analysis of the data. For example, a change in the data dictionary may result in a field having a different definition before and after the change date, which might need to be considered for the analysis. Any information on the impact of that change for future data analysis should be made as explicit as possible.\u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eBefore the change date, the indices in column XXX started at 1, while after the indices start at 0. Old data have not been recoded meaning that indices 1 before the change date should be considered equivalent to indices 0 in the data collected after the change date.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eAfter the change date, the category \u0026ldquo;Equidae\u0026rdquo; from the catalogue XXX was divided into two categories: \u0026ldquo;horse\u0026rdquo; and \u0026ldquo;donkey\u0026rdquo;.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eData collected by the farmer association Y were added to the dataset after the change date.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003e3.5.2. Data quality\u003c/h2\u003e \u003cp\u003eData quality is an important consideration when deciding whether existing data are suitable for reuse for a specific purpose \u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. The first complexity arises from the fact that the definition of \u0026ldquo;good data quality\u0026rdquo; depends on the nature of the data being considered and its intended use. Providing an exhaustive description of data quality in the case of generic data reuse is likely out of reach. The purpose of this framework is therefore not to provide a thorough description of what is meant by good data quality, but rather to ensure that the metadata offers at least basic information related to data quality, allowing users to understand the key biases of the data.\u003c/p\u003e \u003cp\u003eThe second issue comes with the wide variations reported in the list of criteria to be used for data quality assessment and in their definitions, making standardized evaluations of data quality very challenging \u003csup\u003e\u003cspan additionalcitationids=\"CR47 CR48\" citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. For example, the term \u0026ldquo;validity\u0026rdquo; can be used to represent concepts such as \u0026ldquo;accuracy\u0026rdquo; or \u0026ldquo;completeness\u0026rdquo;, which can also be seen as two distinct concepts. To overcome these issues, our framework focuses solely on generic quality criteria, independent of the task to be performed with the data, and on information not already covered in other sections of the framework. For example, the concept of \u0026ldquo;Timeliness\u0026rdquo; was not considered, as it is very specific to the task to be accomplished with the data. Moreover, when \u0026ldquo;timeliness\u0026rdquo; is understood as the period covered by the data, the concept is already covered under the item \u0026ldquo;Data resolution\u0026thinsp;\u0026gt;\u0026thinsp;period covered\u0026rdquo; of this framework. When understood as the general update frequency of the dataset, it would be covered by the item \u0026ldquo;Data resolution\u0026thinsp;\u0026gt;\u0026thinsp;Update frequency and mechanisms\u0026rdquo;. When understood as the time between an event and its actual reporting (e.g., time lag between sampling and analysis by a laboratory), it would be often more appropriate to report it under the description of the corresponding data element (i.e., \u0026ldquo;Data model\u0026thinsp;\u0026gt;\u0026thinsp;Data element\u0026thinsp;\u0026gt;\u0026thinsp;Definition\u0026rdquo;). In any case, it does not need to be repeated here.\u003c/p\u003e \u003cp\u003eIn this framework, data quality is described based on four key items, which are further described below: data completeness, data accuracy, reliability, and quality assurance measures. These categories are not intended to support quantitative scoring, but to structure qualitative reporting. However, if additional activities have been performed to assess other data quality criteria, documenting them in the metadata would be considered of added value and should be encouraged. Such additional information could avoid future users from repeating data quality assessments which have already been performed. Moreover, if the quality of the data has changed at a certain time point in the dataset, it would be also important to highlight it in the description of the data quality. Appendix 3 proposes a template to document data quality information in a standardized way.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eCompleteness\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eData completeness can be defined as the extent to which data are of sufficient breadth, depth and scope for the task at hand \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. When the task at hand is unknown, completeness can be defined as a measure of the proportion of expected data that is actually present. It can be evaluated in two ways:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eCoverage\u003c/em\u003e \u0026ndash; measures the completeness at the \u003cem\u003erecord level\u003c/em\u003e. It can also be seen as the representativeness of the data. For example, if there are known to be \u003cem\u003ex\u003c/em\u003e units of interest in the population, but the dataset only has data on \u003cem\u003ey\u003c/em\u003e units, the coverage is \u003cem\u003ey/x\u003c/em\u003e. This may be easy to calculate if reliable data on the population exists (e.g., all the farms from a given region were sampled). However, the true size of the population of interest may not be known. In that case, the coverage may either be unknown or only estimated especially if it changes over time due to the data collection process.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eMissing data items \u0026ndash;\u003c/em\u003e measure the completeness at the \u003cem\u003ecolumn or field level\u003c/em\u003e. For these items, the proportion of records with missing values should be calculated.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eExamples: Registered aquaculture sites may be audited periodically.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eCoverage\u003c/em\u003e - an audit dataset may contain audit results on only 20% of the true total number of sites before 01/02/2025, but it reached 70% after that date due to addition of a new data source.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eMissing data items\u003c/em\u003e \u0026ndash; within that audit dataset, information about the person in charge of data collection might be missing in 20% of the records.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eAccuracy\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eData accuracy or correctness can be defined as the extent to which data are correct and/or certified by a third party (Wang and Strong 1996). It can be also understood as the inverse of the error rate, which provides an estimate of the proportion of the data that contains errors, including invalid or incorrect values. For the purpose of this framework, we restrict the definition of data correctness to invalid and incorrect data that were not removed during the initial data cleaning phase.\u003c/p\u003e \u003cp\u003e \u003cem\u003eInvalid values\u003c/em\u003e are data that are not consistent with the others in terms of format or content. They are due to data entry errors or system glitches. For example, a capital letter \u0026lsquo;O\u0026rsquo; used instead of the digit zero in a field intended to store numbers would be considered as an invalid value. A text value that is not a valid member of the defined code list or vocabulary in a categorical field would also be considered invalid. Well-designed systems with fields proposing pre-typed entries for selection and quality assurance systems (documented in the \u003cem\u003ebusiness rules\u003c/em\u003e) should contain no \u003cem\u003einvalid data.\u003c/em\u003e Systems that allow free-text entry of categorical values are, however, prone to typographical, capitalization and punctuation inconsistencies, requiring careful data cleaning before analysis starts. In any case, the proportion of invalid present in the shared data should be documented. If invalid data has been corrected before data sharing, the correction process should be documented in the data history section.\u003c/p\u003e \u003cp\u003e \u003cem\u003eIncorrect data\u003c/em\u003e are a different type of error where the data may have a valid value, but those values do not reflect the true state. Incorrect data may be due to various issues, including data capture and data entry errors resulting from typographical mistakes or selection of the incorrect categorical value. Identifying \u003cem\u003eincorrect data\u003c/em\u003e is more difficult because it often appears to be correct. Quantifying them implies comparing the available data with a known point of truth or biological value, which may not always be easily available. Therefore, in practice, incorrect data can often be just defined as \u0026ldquo;\u003cem\u003eunknown values\u003c/em\u003e\u0026rdquo;.\u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples: Registered aquaculture sites may be audited periodically.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eInvalid data \u0026ndash; the dataset contains no invalid data. However, be aware that the data contains three free-text fields to document the evolution of disease events. The content of these fields cannot be validated.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eIncorrect data \u0026ndash; The dataset contains data from Belgium; however, 2% of the latitude and longitudes coordinate pairs reported are located in Poland. In addition, 1% of the cattle appear to have an age largely above the life expectancy of a cow (\u0026gt;\u0026thinsp;50 years) indicating probable incorrect data.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eReliability\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eReliability provides an estimate of the confidence we can put in the results. In this framework, it covers the concepts of \u003cem\u003edata consistency\u003c/em\u003e and \u003cem\u003edata credibility\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003cem\u003eData consistency\u003c/em\u003e represents the extent to which data are compatible or coherent with similar data (Wang and Strong 1996). This refers to internal and external discrepancies present in the data.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eInternal discrepancy\u003c/em\u003e means that the data present in the dataset are not coherent, but the truth is unknown. For example, partial duplicated entries might be present in the data, but there is no way to identify which entry should be used for analysis.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eExternal discrepancy\u003c/em\u003e focuses on the comparison between the data present in the dataset and those available elsewhere. In general, external data discrepancy may be difficult to assess in a standardized way, especially for users who lack in-depth knowledge of the data context. However, if such information is available, it should be reported.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eData credibility\u003c/em\u003e refers to the objective and subjective components of the trustworthiness of data. In that case, biases are suspected but the truth is unknown. Indeed, a data value may be valid, and it may reflect what the generator of the data intended, but it still may not reflect the truth. For example, for human input into a database, some analysis and judgement bias may be involved.\u003c/p\u003e \u003cp\u003e \u003cem\u003eExamples\u003c/em\u003e:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eInternal consistency\u003c/em\u003e:\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003e10% of the mortality data present internal consistency issues: the percentage of mortality does not align with the category of mortality the flock has been assigned to (e.g., Poultry flock mortality was sometimes reported as 90%, but was at the same time labelled as \u0026ldquo;very low\u0026rdquo; in another column). This is an internal discrepancy between these two pieces of information and we are unable to know which one is true.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003e200.000 fish were reported dead at the end of the month, yet no fish were reported to be present at that location.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eExternal consistency\u003c/em\u003e:\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eSome latitude and longitude of the cases reported in the data do not match with the official description of the outbreak location. Similarly, some of the symptoms reported in the data are not aligned with the symptoms expected in this production type. In particular, drop in milk production has been reported by 1.2% of the beef farms present in the data.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eThe total volume of slaughtered fish reported by salmon farmers in [Region/Area] for [YYYY\u0026ndash;YYYY] is [A] tonnes. For the same period, the volume reported to the Norwegian Tax Administration is [B] tonnes. This discrepancy ([A\u0026ndash;B] tonnes; [(A\u0026ndash;B)/A \u0026times; 100]%) indicates an external inconsistency between data sources. Moreover, the dataset contains values that are inconsistent with regulations. For X% of locations, fish are reported present within one month of the end of the previous production cycle, whereas the regulations require a minimum two-month fallowing period between cycles.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eCredibility: A field that provides clinical diagnosis may be considered relatively reliable when a veterinarian provides the data, but less so when it is provided by a farmer without veterinary training. Similarly, a diagnostic test performed in difficult field conditions could be considered less credible than a test conducted in an accredited laboratory.\u003c/em\u003e \u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eQuality assurance measures\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThis final section describes any measures in place to ensure the quality of the data. Examples include one-off (for static datasets) or periodic (for operational datasets) data assessment and cleaning, or elements of system design, such as:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eBusiness rules or built-in automated data validation steps, which may have been applied to the data. Business rules are a set of rules related to data quality that are applied to the data. They are important to document how to reuse the data properly. For example, a business rule might be that all values above a certain threshold are considered abnormal and are replaced by a constant or set to missing. Whenever possible, business rules should be expressed and shared in a machine-actionable format (e.g., script), or at the minimum with a strict formalism using mathematical, statistical or conditional explicit formulas to ensure they are univocal.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFeedback systems which provide the data generator with an opportunity to validate the data contained in the dataset. For example, within online survey systems, the mechanisms put in place to validate the data and ensure internal consistency could be reported here.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis paper presents the first community-developed rich metadata guidelines tailored to veterinary epidemiology datasets with ready-to-use templates and real-life examples. Thanks to the large number of stakeholders involved in the process, the guidelines should be easily understood by the members of the community (i.e. researchers in university but also scientists working in public or private organizations), and target the key information needed to support internal and external data reusability in this domain. They provide a single, comprehensive source of information, so researchers do not need to consult the metadata literature, which is often complex and overly technical for newcomers to the field.\u003c/p\u003e \u003cp\u003eIf having good metadata increases the value of the collected data, creating metadata primarily benefits the data-owning organization by ensuring institutional memory, making internal data reuse easier and more cost effective, and therefore maximizing the value they get from the data they collect. Therefore, creating rich metadata, even if only in a private repository, should be driven by the direct benefits they provide to the data owner, and not by the willingness to share data with external stakeholders. In the modern era where data are becoming big and complex and essential for running all kinds of operations, relying solely on people\u0026rsquo;s memory or availability to be able to work with data are a critical risk that public and private organizations should not be taking anymore. Having strong organizational incentives to support the creation of metadata is therefore critical to help individuals engage in the process even if they may not have a personal benefit for doing so. For example, in the public sector, including Data Management Plans (DMPs) requirements in grant contracts as done by numerous funding bodies is a first step, but, this should be associated with a specific allocation of resources to not strip them of their meaning and truly achieve data sustainability \u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. At a smaller scale, it would be also worth requiring metadata creation from students as part of their first projects during their studies and internships to make metadata creation a habit right from the start. This would be part of the development of their data literacy, which must be good enough to allow them later to handle several issues. Developing different career advancement criteria in academia, for example based on a dataset citation index, may also help researchers to dedicate time to this activity. In the private sector, the role of data as a valuable resource to better manage production processes, identify new opportunities or leverage business partnerships is not new \u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e,\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e. We can only encourage stakeholders in the animal production business to invest more in good data management practices to be able to fully embrace this opportunity. In some special cases, researchers may want to reuse non-scholarly data for conducting research, but the data owner is not willing to invest in the creation of metadata. In that context, we believe the responsibility of generating the best possible metadata is with the researchers, who are intending to generate and share scientific knowledge based on this data. It is indeed part of their ethical responsibility as scientists to carry out and document their work in a transparent and accurate manner \u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e. In this situation, it could be interesting to investigate whether sharing back and cross validating with data owners the metadata created during the research process could be considered as an added benefit, and whether this practice could help improve the accessibility of non-scholarly data for scientists.\u003c/p\u003e \u003cp\u003eData sharing is not a prerequisite to the creation of metadata, and metadata can be published even when data sharing is restricted. Data sharing is indeed often complex especially in the agriculture domain because of privacy and confidentiality issues (see for examples \u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e,\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e). On the other hand, metadata sharing is easier as it is not meant to contain any sensitive information and it still provides a high degree of \u0026ldquo;FAIRness\u0026rdquo; even in the absence of FAIR publication of the data itself by facilitating data discovery and including clear rules regarding the process for accessing the data \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. The work conducted by Fraga-Gonz\u0026aacute;lez et al. (2025) presents a number of recommendations on how and where to share metadata that can be applied to the case of data in veterinary epidemiology. We emphasize their recommendations of a two-stage approach using Git to continuously catalog data and documentation, and a well-established research repository like Zenodo (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://zenodo.org/\u003c/span\u003e\u003cspan address=\"https://zenodo.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to describe the Git repository with the catalog and point at its URL.\u003c/p\u003e \u003cp\u003eThese guidelines intend to facilitate the process of creating rich metadata in veterinary epidemiology, but it is worth noting that it will be always more challenging to use them post-hoc, i.e., after data production, than concurrently with the data production \u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. We therefore cannot stress enough the importance of DMPs to make sure the creation of metadata is embedded in the workflow and is not an extra task raised only at the end of a project. Moreover, it should be clear that these guidelines will not necessarily help improving metadata quality scores as defined by some frameworks such as the one developed by the official portal for European data \u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e. Indeed, such frameworks assess the quality of metadata only via a selected list of FAIR guiding principles focusing on concepts that are easy to evaluate in a standardized and domain-agnostic manner. They do not consider the \u0026ldquo;richness\u0026rdquo; of metadata and how they can be effectively used to support data reuse, as this task is necessarily domain dependent and a lot more difficult to achieve. The development of domain-specific guidelines such as those proposed in this paper is a step forward to be also able to assess this aspect of metadata quality, but more work would be needed to develop a formal evaluation process.\u003c/p\u003e \u003cp\u003eDuring the development of these guidelines, the researchers often navigated between two competing goals: providing detailed enough guidance while remaining flexible to adapt to a large range of datasets. Some authors would have preferred to go into even more detailed reporting, while others found the proposed list of items too long and overwhelming and may have preferred to prioritize some items. The version presented in this paper was deemed to provide a suitable balance between both objectives. It is important to note that these guidelines intend to be a first version that may evolve over time to better meet the needs of the professionals working in veterinary epidemiology, and stakeholders should feel free to adapt them to their specific or evolving needs. Indeed, additional items may need to be added to make specific kinds of data understandable by others and prevent potential misuse of the data. In general, when it comes to rich metadata, there is never too much information available so adding more should never be considered as an issue. However, to do so, we strongly recommend continuing using a structured approach and building on existing standards to ensure any new information added can be easily retrieved from the metadata by any stakeholders. Moreover, some parts of the guidelines could be extended to be handled in a more standardized manner. For example, the guidance provided in this study to report data quality is a first step but could be further refined to better support good data quality check practices. At this stage, several definitions were kept broad on purpose to cover a large diversity of situations, which may lead to different interpretations and therefore different assessments. For example, assessing internal data consistency in practice is likely to remain challenging in the absence of more detailed information on how to proceed. To allow for such evolution and ensure the sustainability of these guidelines, this version is available in a version control platform (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines\u003c/span\u003e\u003cspan address=\"https://github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e ), which is associated to a long-term data repository (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.18889286\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.18889286\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). New versions developed should ensure they refer to these first versions and that differences are clearly highlighted to help stakeholders navigate potential multiple versions of the guidelines and associated templates.\u003c/p\u003e \u003cp\u003eFor inexperienced professionals in our community, the guidelines may appear complex and lengthy at first. However, a more thorough examination will demonstrate that it would be difficult to reuse a dataset without having access to the basic information listed in the guidelines. During the development process, several rounds of discussions and reviews have led to retaining only the critical elements and removing anything that was agreed to be unnecessary. Some compromises had to be made between the authors to lead to these choices. It is expected that similar disagreements will also be raised by other professionals in our community. However, we expect that the potential barriers to the adoption of these guidelines are less related to their content, but rather to the true or perceived lack of resources and incentives to fill the templates. These concerns echo with what has been reported elsewhere when considering barriers to data sharing and the adoption of good data management practices in related domains \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan additionalcitationids=\"CR59 CR60\" citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e. First, it is important to highlight the fact that none of the information gathered through these guidelines is really new in the sense that anyone who is working with the data should be aware of it. Therefore, the information already exists \u0026ldquo;somewhere\u0026rdquo;, the only effort required is to document it once in a standardized and centralized way. Moreover, not every type of data would require the same amount of work to generate metadata. For example, the development of sensor data in modern agricultural systems partially solved the question of metadata creation for these data which usually come with robust data schemas and quality descriptions. However, information related to the context of use of these technologies remains rarely available even if technologies are being developed to fill this gap \u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e. In that context, only an ad-hoc collection of contextual metadata would be necessary to ensure proper data reuse as other information should be already made available by the manufacturer of the tool. Then, regarding the lack of resources to generate metadata, there is likely a biased perception that current practices are not costing resources, which has been shown to be largely false at least in the academic world \u003csup\u003e\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e,\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e. As said by Mons (2020): \u0026ldquo;\u003cem\u003eif data are treated properly, researchers will have significantly more time to do research\u003c/em\u003e\u0026rdquo;. There is therefore a need to change the narrative from \u0026ldquo;the time lost by creating metadata\u0026rdquo; to \u0026ldquo;the time saved by creating metadata\u0026rdquo; in this community.\u003c/p\u003e \u003cp\u003eNew technologies currently being developed can help address some of the barriers identified to the creation of rich metadata such as the lack of resources. In particular, the integration of artificial intelligence (AI) into metadata management processes has the potential to address many of the challenges faced by traditional systems, particularly in the areas of metadata generation, reuse, and lifecycle management \u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e. These technologies can be particularly useful when metadata needs to be created post-hoc and many documents describing the data already exist and can be used as reference to generate rich metadata. Indeed, AI models can only be of use when the information already exists and is accessible in a digital format somewhere. Moreover, if AI models are effective at pointing out statistical anomalies, they are still struggling to identify issues of data quality which are domain-dependent (e.g., differencing a normal from an abnormal value is difficult without a deep contextual understanding of the data). Therefore, our guidelines are unlikely to become obsolete because of AI technologies, and they would rather be used as a source of information by these tools. AI may also help address the fact that rich metadata can probably only be unified up to a certain point. Indeed, our guidelines propose several items to be filled, but the content of each of them remain mostly textual and is likely to vary greatly among datasets depending for example on their type or structure. AI tools could help turn these partially structured data into fully structured data that can be more easily queried while not impairing human usability of the metadata collection templates. If AI provides powerful tools to save resources and support metadata creation, the users always need to keep in mind that such tools may also lack consistency, validation and may compromise data privacy and security \u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e. They should therefore be used with great caution.\u003c/p\u003e \u003cp\u003eWe chose to propose the templates associated with these guidelines in .csv and .xlsx files formats to facilitate their adoption by a community not yet very engaged in the process of making their data reusable. However, these formats are not satisfactory from a machine-readable perspective nor for long-term or large-scale data sharing. They should be considered a first but important step towards greater data FAIRness in the domain of veterinary epidemiology. Indeed, looking at the FAIR continuum proposed by Garc\u0026iacute;a-Closas et al. (2023), our guidelines mostly focus on steps 1 to 3 (i.e., Metadata exist in a human and basic nonproprietary machine readable format), providing a stepping stone for our community to reach step 4 (i.e., \u0026ldquo;linked data framework\u0026rdquo;) in the future. The future development of these guidelines may therefore propose more advanced formats to better support machine-readability such as Resource Description Framework (RDF), JSON-LD or ontology-based systems, which can help automate metadata creation and make them more consistent and reusable. Such a stepwise process has proven successful, for example by the Cattle Barometer, a data-driven decision support tool which was initially developed based on direct data sharing using spreadsheets and CSV files received by email and manually uploaded and cleaned \u003csup\u003e\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e\u003c/sup\u003e. Later versions transitioned to an ontology-driven pipeline mapping records to the Livestock Health Ontology for creating interoperable and fully automated dashboards \u003csup\u003e\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e,\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e. The workflow is documented in the DECIDE \u003cem\u003eCattle barometer tutorial\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e\u003c/sup\u003e. Work on more advanced format for the guidelines has already been started and the first steps are available on various platforms (see for example GitHub \u003csup\u003e\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e). Our guidelines are a first step in the transition to more advanced data management practices, but to go further, more awareness campaigns and fit-for-purpose training will be needed to increase the overall skills of the community on these questions. Increasing collaboration and relationships with computer scientists may also help to facilitate this transition.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eFunding statement\u003c/h2\u003e \u003cp\u003eFunded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the granting authority. Neither the European Union nor the granting authority can be held responsible for them. This work was specifically supported by the European Union\u0026rsquo;s Horizon 2020 research and innovation programme under the grant agreement No 101000494 (Data-driven control and prioritisation of non-EU-regulated contagious animal diseases [DECIDE]) and by the European Partnership on Animal Health and Welfare (EUPAHW) under the grant agreement No 101136346.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eC.F. conceived the study, developed the theoretical framework and performed the experiments. C.D. aided in the development of the framework and implementation of the experiments. C.F. supervised the study. K.R.D., V.H.S.O. and M.N.O created the salmon case study. C.D. created the poultry case study. G.v.S. acquired the funding and did the project administration. All authors took part of the development of the guidelines and associated templates and provided critical feedback and helped shape the research, analysis and manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eAll the people in the different research teams involved tested the templates on their data. Special thank you to Guita Niang from INRAE, Robert Johansson from the Swedish Veterinary Agency.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eAll the templates associated with the paper are available on a version control platform [https://github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines](https:/github.com/EpiMundi/faverjon-2026-vet-epi-contextual-metadata-guidelines) , which is associated to a long-term data repository ( [https://doi.org/10.5281/zenodo.18889286](https:/doi.org/10.5281/zenodo.18889286) ).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWilkinson, M. D. \u003cem\u003eet al.\u003c/em\u003e The FAIR Guiding Principles for scientific data management and stewardship. \u003cem\u003eSci. Data\u003c/em\u003e 3, 160018 (2016). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/sdata.2016.18\u003c/span\u003e\u003cspan address=\"10.1038/sdata.2016.18\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJacobsen, A. \u003cem\u003eet al.\u003c/em\u003e FAIR Principles: Interpretations and Implementation Considerations. \u003cem\u003eData Intell.\u003c/em\u003e 10\u0026ndash;29 (2020) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1162/dint_r_00024\u003c/span\u003e\u003cspan address=\"10.1162/dint_r_00024\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDublin Core\u003csup\u003e\u0026trade;\u003c/sup\u003e Metadata Initiative. Metadata Terms.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDataCite Metadata Working Group. DataCite Metadata Schema Documentation for the Publication and Citation of Research Data and Other Research Outputs. (2024) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.14454/mzv1-5b55\u003c/span\u003e\u003cspan address=\"10.14454/mzv1-5b55\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eData Catalog Vocabulary (DCAT). (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOECD. \u003cem\u003eResponding to Societal Challenges with Data: Access, Sharing, Stewardship and Control\u003c/em\u003e. vol. 342 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1787/2182ce9f-en\u003c/span\u003e\u003cspan address=\"10.1787/2182ce9f-en\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCallahan, T., Barnard, J., Helmkamp, L., Maertens, J. \u0026amp; Kahn, M. Reporting Data Quality Assessment Results: Identifying Individual and Organizational Barriers and Solutions. \u003cem\u003eeGEMs\u003c/em\u003e 16 (2017) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5334/egems.214\u003c/span\u003e\u003cspan address=\"10.5334/egems.214\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTenopir, C. \u003cem\u003eet al.\u003c/em\u003e Data sharing, management, use, and reuse: Practices and perceptions of scientists worldwide. \u003cem\u003ePLOS ONE\u003c/em\u003e e0229003 (2020) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0229003\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0229003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrazma, A. \u003cem\u003eet al.\u003c/em\u003e Minimum information about a microarray experiment (MIAME)-toward standards for microarray data. \u003cem\u003eNat. Genet.\u003c/em\u003e 29, 365\u0026ndash;371 (2001). DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/ng1201-365\u003c/span\u003e\u003cspan address=\"10.1038/ng1201-365\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRund, S. S. C. \u003cem\u003eet al.\u003c/em\u003e MIReAD, a minimum information standard for reporting arthropod abundance data. \u003cem\u003eSci. Data\u003c/em\u003e 6, 40 (2019). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41597-019-0042-5\u003c/span\u003e\u003cspan address=\"10.1038/s41597-019-0042-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTenorio, F. A. M. \u003cem\u003eet al.\u003c/em\u003e Filling the agronomic data gap through a minimum data collection approach. \u003cem\u003eField Crops Res.\u003c/em\u003e 308, 109278 (2024). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.fcr.2024.109278\u003c/span\u003e\u003cspan address=\"10.1016/j.fcr.2024.109278\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlumberg, K. L., McKillop, K., Pehrsson, P. R. \u0026amp; Fukagawa, N. K. Call to Action: A Need for Community-Driven Minimum Information Standards for Food Composition Data. \u003cem\u003eAm. J. Clin. Nutr.\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ajcnut.2025.06.027\u003c/span\u003e\u003cspan address=\"10.1016/j.ajcnut.2025.06.027\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025) doi:10.1016/j.ajcnut.2025.06.027.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinaci, A. A. \u003cem\u003eet al.\u003c/em\u003e From Raw Data to FAIR Data: The FAIRification Workflow for Health Research. \u003cem\u003eMethods Inf. Med.\u003c/em\u003e 59, e21\u0026ndash;e32 (2020). DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1055/s-0040-1713684\u003c/span\u003e\u003cspan address=\"10.1055/s-0040-1713684\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGregory, A. \u003cem\u003eet al.\u003c/em\u003e WorldFAIR Project (D7.1) Population Health Data Implementation Guide. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.7887385\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.7887385\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023) doi:10.5281/zenodo.7887385.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeyer, A., Faverjon, C., Hostens, M., Stegeman, A. \u0026amp; Cameron, A. Systematic review of the status of veterinary epidemiological research in two species regarding the FAIR guiding principles. \u003cem\u003eBMC Vet. Res.\u003c/em\u003e 270 (2021) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12917-021-02971-1\u003c/span\u003e\u003cspan address=\"10.1186/s12917-021-02971-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDelavenne, C., Cameron, A., van Schaik, G., Fr\u0026ouml;ssling, J. \u0026amp; Faverjon, C. Reusability challenges of livestock production data to improve animal health. \u003cem\u003eSci. Data\u003c/em\u003e 12, (2025) \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41597-025-04785-4\u003c/span\u003e\u003cspan address=\"10.1038/s41597-025-04785-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZinsstag, J., Schelling, E., Waltner-Toews, D. \u0026amp; Tanner, M. From \u0026ldquo;one medicine\u0026rdquo; to \u0026ldquo;one health\u0026rdquo; and systemic approaches to health and well-being. \u003cem\u003ePrev. Vet. Med.\u003c/em\u003e 101, 148\u0026ndash;156 (2011) doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.prevetmed.2010.07.003\u003c/span\u003e\u003cspan address=\"10.1016/j.prevetmed.2010.07.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKloeze, H. \u003cem\u003eet al.\u003c/em\u003e A minimum data set of animal health laboratory data to allow for collation and analysis across jurisdictions for the purpose of surveillance. \u003cem\u003eTransbound. Emerg. Dis.\u003c/em\u003e 59, 264\u0026ndash;268 (2012) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/j.1865-1682.2011.01264.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1865-1682.2011.01264.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchwantes, C. J. \u003cem\u003eet al.\u003c/em\u003e A minimum data standard for wildlife disease research and surveillance. \u003cem\u003eSci. Data\u003c/em\u003e 12, 1054 (2025) doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41597-025-05332-x\u003c/span\u003e\u003cspan address=\"10.1038/s41597-025-05332-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eO\u0026rsquo;Connor, A. M. \u003cem\u003eet al.\u003c/em\u003e Explanation and Elaboration Document for the STROBE-Vet Statement: Strengthening the Reporting of Observational Studies in Epidemiology \u0026ndash; Veterinary Extension. \u003cem\u003eZoonoses Public Health\u003c/em\u003e 63, 662\u0026ndash;698 (2016), doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/jvim.14592\u003c/span\u003e\u003cspan address=\"10.1111/jvim.14592\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePlate, K., Kn\u0026uuml;ppel, S., Hornbacher, A., Greiner, M. \u0026amp; M\u0026uuml;ller-Graf, C. Development of a tool for rapid assessment of the evidence of human observational epidemiological studies with focus on risk of bias in the context of public health and risk assessment. \u003cem\u003eEFSA Support. Publ.\u003c/em\u003e 21, 8925E (2024) \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2903/sp.efsa.\u003c/span\u003e\u003cspan address=\"10.2903/sp.efsa.\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e2024.EN-8925Digital Object Identifier (DOI)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcKechnie, I., Raymond, K. \u0026amp; Stacey, D. Identifying Inconsistencies in Data Quality Between FAOSTAT, WOAH, UN Agriculture Census, and National Data | Data Science Journal. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5334/dsj-2024-044\u003c/span\u003e\u003cspan address=\"10.5334/dsj-2024-044\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024) doi:10.5334/dsj-2024-044.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCrystal-Ornelas, R. \u003cem\u003eet al.\u003c/em\u003e Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats. \u003cem\u003eSci. Data\u003c/em\u003e 9, 700 (2022). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41597-022-01606-w\u003c/span\u003e\u003cspan address=\"10.1038/s41597-022-01606-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLortie, C. J., Vargas Poulsen, C., Brun, J. \u0026amp; Kui, L. Tabular strategies for metadata in ecology, evolution, and the environmental sciences. \u003cem\u003eEcol. Evol.\u003c/em\u003e 12, e9245 (2022). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/ece3.\u003c/span\u003e\u003cspan address=\"10.1002/ece3.\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e9245Digital Object Identifier (DOI)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eŚmiałek, M., Delavenne, C. \u0026amp; Faverjon, C. Production and health performance of a selection of polish poultry broiler flocks (Ross308) between 2018 and 2023. Zenodo \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.17406208\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.17406208\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNorwegian Veterinary Institute, Oliveira, V. H. S., Osnes, M. N. \u0026amp; Dean, K. R. Salmonid mortality and losses in Norwegian aquaculture. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.17791861\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.17791861\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eByrd, J. B., Greene, A. C., Prasad, D. V., Jiang, X. \u0026amp; Greene, C. S. Responsible, practical genomic data sharing that accelerates research. \u003cem\u003eNat. Rev. Genet.\u003c/em\u003e 21, 615\u0026ndash;629 (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41576-020-0257-5\u003c/span\u003e\u003cspan address=\"10.1038/s41576-020-0257-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCernava, T. \u003cem\u003eet al.\u003c/em\u003e Metadata harmonization\u0026ndash;Standards are the key for a better usage of omics data for integrative microbiome analysis. \u003cem\u003eEnviron. Microbiome\u003c/em\u003e 17, 33 (2022). doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s40793-022-00425-1\u003c/span\u003e\u003cspan address=\"10.1186/s40793-022-00425-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSDMX community. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://sdmx.org/\u003c/span\u003e\u003cspan address=\"https://sdmx.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFairchild, G. \u003cem\u003eet al.\u003c/em\u003e Epidemiological Data Challenges: Planning for a More Robust Future Through Data Standards. \u003cem\u003eFront. Public Health\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpubh.2018.00336\u003c/span\u003e\u003cspan address=\"10.3389/fpubh.2018.00336\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018) doi:10.3389/fpubh.2018.00336.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeginis, D., Kalampokis, E., Palma, R., Atkinson, R. \u0026amp; Tarabanis, K. Statistical Challenges Towards a Semantic Meta-Model for Data Integration and Exploitation in Precision Agriculture and Livestock Farming. \u003cem\u003e7th Int. Workshop Semantic Stat. SemStats\u003c/em\u003e2019 \u003cem\u003eCo-Located 18th Int. Semantic Web Conf. ISWC2019\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3233/SW-233156\u003c/span\u003e\u003cspan address=\"10.3233/SW-233156\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019) doi:10.3233/SW-233156.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMitchell, S. N. \u003cem\u003eet al.\u003c/em\u003e FAIR data pipeline: provenance-driven data management for traceable scientific workflows. \u003cem\u003ePhilos. Trans. R. Soc. Math. Phys. Eng. Sci.\u003c/em\u003e 380, 20210300 (2022). doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1098/rsta.2021.0300\u003c/span\u003e\u003cspan address=\"10.1098/rsta.2021.0300\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAminalragia-Giamini, R. \u003cem\u003eet al.\u003c/em\u003e SIGMA - Population Data Model. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.7520891\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.7520891\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023) doi:10.5281/zenodo.7520891.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCatalogue - ORIONKnowledgeHub. https://aginfra.d4science.org:443/web/orionknowledgehub/catalogue?p_p_auth=Rc853tgT\u0026amp;p_p_id=49\u0026amp;p_p_lifecycle=1\u0026amp;p_p_state=normal\u0026amp;p_p_mode=view\u0026amp;_49_struts_action=%2Fmy_sites%2Fview\u0026amp;_49_groupId=92976941\u0026amp;_49_privateLayout=false.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarrison, P. W. \u003cem\u003eet al.\u003c/em\u003e FAANG, establishing metadata standards, validation and best practices for the farmed and companion animal community. \u003cem\u003eAnim. Genet.\u003c/em\u003e 520\u0026ndash;526 (2018) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/age.12736\u003c/span\u003e\u003cspan address=\"10.1111/age.12736\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDooley, D. M. \u003cem\u003eet al.\u003c/em\u003e FoodOn: a harmonized food ontology to increase global food traceability, quality control and data integration. \u003cem\u003eNpj Sci. Food\u003c/em\u003e 23 (2018) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41538-018-0032-6\u003c/span\u003e\u003cspan address=\"10.1038/s41538-018-0032-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eICAR. Global standard for livestock data.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNorwegian Institute of ublic Health. ATCvet code. (2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollege of Veterinary Medicine - Virginia-Maryland. VetSCT. (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFAO. AGROVOC Multilingual Thesaurus. (2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eINRAE. Ontologies for Livestock.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD\u0026oacute;rea, F. C. \u003cem\u003eet al.\u003c/em\u003e Drivers for the development of an Animal Health Surveillance Ontology (AHSO). \u003cem\u003ePrev. Vet. Med.\u003c/em\u003e 39\u0026ndash;48 (2019) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.prevetmed.2019.03.002\u003c/span\u003e\u003cspan address=\"10.1016/j.prevetmed.2019.03.002\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoor, S. \u0026amp; Hostens, M. Livestock Health Ontology. (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eISO. ISO 25964 \u0026ndash; the international standard for thesauri and interoperability with other vocabularies.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOliveira, P., Rodrigues, F. \u0026amp; Rangel Henriques, P. \u003cem\u003eA Formal Definition of Data Quality Problems. Proceedings of the\u003c/em\u003e 2005 \u003cem\u003eInternational Conference on Information Quality, ICIQ 2005\u003c/em\u003e (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, H., Hailey, D., Wang, N. \u0026amp; Yu, P. A Review of Data Quality Assessment Methods for Public Health Information Systems. \u003cem\u003eInt. J. Environ. Res. Public. Health\u003c/em\u003e 11, 5170\u0026ndash;5207 (2014). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/ijerph110505170\u003c/span\u003e\u003cspan address=\"10.3390/ijerph110505170\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeiskopf, N. G. \u0026amp; Weng, C. Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research. \u003cem\u003eJ. Am. Med. Inform. Assoc.\u003c/em\u003e 20, 144\u0026ndash;151 (2013). DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/amiajnl-2011-000681\u003c/span\u003e\u003cspan address=\"10.1136/amiajnl-2011-000681\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, L., Peng, T. \u0026amp; Kennedy, J. A Rule Based Taxonomy of Dirty Data. \u003cem\u003eGSTF Int. J. Comput.\u003c/em\u003e (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBirkeg\u0026aring;rd, A. C. \u003cem\u003eet al.\u003c/em\u003e Building the foundation for veterinary register-based epidemiology: A systematic approach to data quality assessment and validation. \u003cem\u003eZoonoses Public Health\u003c/em\u003e 65, 936\u0026ndash;946 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/zph.12513\u003c/span\u003e\u003cspan address=\"10.1111/zph.12513\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, R. Y. \u0026amp; and Strong, D. M. Beyond Accuracy: What Data Quality Means to Data Consumers. \u003cem\u003eJ. Manag. Inf. Syst.\u003c/em\u003e 12, 5\u0026ndash;33 (1996). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.jstor.org/stable/40398176\u003c/span\u003e\u003cspan address=\"https://www.jstor.org/stable/40398176\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFraga-Gonz\u0026aacute;lez, G. \u003cem\u003eet al.\u003c/em\u003e Affording reusable data: recommendations for researchers from a data-intensive project. \u003cem\u003eSci. Data\u003c/em\u003e 12, 258 (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41597-025-04565-0\u003c/span\u003e\u003cspan address=\"10.1038/s41597-025-04565-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrimaldi, D., Fernandez, V. \u0026amp; Carrasco, C. Exploring data conditions to improve business performance. \u003cem\u003eJ. Oper. Res. Soc.\u003c/em\u003e 72, 1087\u0026ndash;1098 (2021). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/01605682.2019.1590136\u003c/span\u003e\u003cspan address=\"10.1080/01605682.2019.1590136\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePopovič, A., Hackney, R., Tassabehji, R. \u0026amp; Castelli, M. The impact of big data analytics on firms\u0026rsquo; high value business performance. \u003cem\u003eInf. Syst. Front.\u003c/em\u003e 20, 209\u0026ndash;222 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10796-016-9720-4\u003c/span\u003e\u003cspan address=\"10.1007/s10796-016-9720-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eALLEA - All European Academies. \u003cem\u003eThe European Code of Conduct for Research Integrity\u003c/em\u003e. (ALLEA - All European Academies, DE, 2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJouanjean, M.-A., Casalini, F., Wiseman, L. \u0026amp; Gray, E. Issues around data governance in the digital transformation of agriculture : The farmers\u0026rsquo; perspective. \u003cem\u003eOECD Food Agric. Fish. Pap.\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1787/53ecf2ab-en\u003c/span\u003e\u003cspan address=\"10.1787/53ecf2ab-en\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020) doi:10.1787/53ecf2ab-en.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGavai, A. K. \u003cem\u003eet al.\u003c/em\u003e Agricultural data privacy: Emerging platforms \u0026amp; strategies. \u003cem\u003eFood Humanity\u003c/em\u003e 4, 100542 (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.foohum.2025.100542\u003c/span\u003e\u003cspan address=\"10.1016/j.foohum.2025.100542\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEuropean Data Portal. Metadata Quality Assurance. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://data.europa.eu/mqa/methodology?locale=en\u003c/span\u003e\u003cspan address=\"https://data.europa.eu/mqa/methodology?locale=en\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoutkoop, B. L. \u003cem\u003eet al.\u003c/em\u003e Data sharing in psychology: A survey on barriers and preconditions. \u003cem\u003eAdv. Methods Pract. Psychol. Sci.\u003c/em\u003e 70\u0026ndash;85 (2018) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1177/2515245917751886\u003c/span\u003e\u003cspan address=\"10.1177/2515245917751886\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHughes, L. D. \u003cem\u003eet al.\u003c/em\u003e Addressing barriers in FAIR data practices for biomedical data. \u003cem\u003eSci. Data\u003c/em\u003e 98 (2023) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41597-023-01969-8\u003c/span\u003e\u003cspan address=\"10.1038/s41597-023-01969-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerrier, L., Blondal, E. \u0026amp; MacDonald, H. The views, perspectives, and experiences of academic researchers with data sharing and reuse: A meta-synthesis. \u003cem\u003ePloS One\u003c/em\u003e (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.pone.0229182\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0229182\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRainey, L., Lutomski, J. E. \u0026amp; Broeders, M. J. FAIR data sharing: An international perspective on why medical researchers are lagging behind. \u003cem\u003eBig Data Soc.\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/20539517231171052\u003c/span\u003e\u003cspan address=\"10.1177/20539517231171052\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023) doi:10.1177/20539517231171052.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBasir, Md. S., Zhang, Y., Buckmaster, D., Raturi, A. \u0026amp; Krogmeier, J. V. Meta Ag: An automatic agricultural contextual metadata collection app. \u003cem\u003eSmart Agric. Technol.\u003c/em\u003e 12, 101073 (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.atech.2025.101073\u003c/span\u003e\u003cspan address=\"10.1016/j.atech.2025.101073\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBadolato, A.-M. Cost-Benefit analysis for FAIR research data \u0026ndash; Policy Recommendations. \u003cem\u003eOuvrir la Science\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ouvrirlascience.fr/cost-benefit-analysis-for-fair-research-data-policy-recommendations-2\u003c/span\u003e\u003cspan address=\"https://www.ouvrirlascience.fr/cost-benefit-analysis-for-fair-research-data-policy-recommendations-2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEuropean Commission. \u003cem\u003eCost-Benefit Analysis for FAIR Research Data: Policy Recommendations.\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://data.europa.eu/doi/10.2777/706548\u003c/span\u003e\u003cspan address=\"https://data.europa.eu/doi/10.2777/706548\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMons, B. Invest 5% of research funds in ensuring data are reusable. \u003cem\u003eNature\u003c/em\u003e 578, (2020). doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/d41586-020-00505-7\u003c/span\u003e\u003cspan address=\"10.1038/d41586-020-00505-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, W., Fu, R., Amin, M. B. \u0026amp; Kang, B. The Impact of Modern AI in Metadata Management. \u003cem\u003eHum.-Centric Intell. Syst.\u003c/em\u003e 5, 323\u0026ndash;350 (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s44230-025-00106-5\u003c/span\u003e\u003cspan address=\"10.1007/s44230-025-00106-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarc\u0026iacute;a-Closas, M. \u003cem\u003eet al.\u003c/em\u003e Moving Toward Findable, Accessible, Interoperable, Reusable Practices in Epidemiologic Research. \u003cem\u003eAm. J. Epidemiol.\u003c/em\u003e 192, 995\u0026ndash;1005 (2023). DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/aje/kwad040\u003c/span\u003e\u003cspan address=\"10.1093/aje/kwad040\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBokma, J. \u003cem\u003eet al.\u003c/em\u003e European veterinary barometer for Bovine Respiratory Diseases : a tool showing diagnostic test results and geolocation of respiratory tract samples from cattle. in \u003cem\u003eEuropean Buiatrics Congress and ECBHM Jubilee Symposium 2023\u003c/em\u003e 179\u0026ndash;181 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoor, S. \u003cem\u003eet al.\u003c/em\u003e Advancing precision livestock farming through ontology-driven interoperable health data management : extending the Livestock Health Ontology (LHO) for enhanced disease surveillance. in \u003cem\u003e11th European Conference on Precision Livestock Farming, Proceedings\u003c/em\u003e 571\u0026ndash;580 (The Organising Committee of the 11th European Conference on Precision Livestock Farming (ECPLF), 2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoor, S., Bokma, J., Pardon, B., van Schaik, G. \u0026amp; Hostens, M. Agri semantics : developments to improve data interoperability to support farm information management and decision support systems in agriculture. in \u003cem\u003eSmart farms : improving data-driven decision making in agriculture\u003c/em\u003e 75\u0026ndash;96 (Burleigh Dodds Science, 2024). doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.19103/as.2023.0132.05\u003c/span\u003e\u003cspan address=\"10.19103/as.2023.0132.05\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBokma, J., Noor, S., Pardon, B. \u0026amp; Hostens, M. Cattle barometer. (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVEMO: veterinary epidemiology metadata ontology. (2026).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Footnotes","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.doi.org/\u003c/span\u003e\u003cspan address=\"https://www.doi.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://public.opendatasoft.com/explore/?sort=modified\u003c/span\u003e\u003cspan address=\"https://public.opendatasoft.com/explore/?sort=modified\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Metadata, veterinary epidemiology, FAIR, guidelines","lastPublishedDoi":"10.21203/rs.3.rs-9050779/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9050779/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDespite the global adoption of the FAIR guiding principles, significant barriers remain in their practical implementation due to the lack of domain-specific standards for \"rich metadata\". While generic metadata schemes provide basic information, they often fail to capture the complexity of data quality and the nuances of specialized scientific fields. This gap is particularly evident in veterinary epidemiology, a field characterized by complex, multi-scale data spanning various domains where analytical frameworks trend to be prioritized over raw data description, leading to inconsistent reporting and low data reusability. This paper introduces the first community-developed rich metadata guidelines specifically designed for veterinary epidemiology datasets, accompanied by practical templates and real-world examples. By integrating existing standards with specific guidance on domain-specific attributes and data quality, these guidelines provide a comprehensive, single-source framework accessible to researchers across academic, public, and private sectors. They aim to support researchers and stakeholders in enhancing the reusability of animal health data.\u003c/p\u003e","manuscriptTitle":"Enhancing reusability of veterinary epidemiological data by creating contextual metadata","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-11 14:22:30","doi":"10.21203/rs.3.rs-9050779/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a8127a36-f781-40fb-bc98-95b23b17b9c5","owner":[],"postedDate":"March 11th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"284972432806050106161828364779252665850","date":"2026-05-01T01:12:17+00:00","index":35,"fulltext":""},{"type":"reviewerAgreed","content":"171918485699456953187594668947590955667","date":"2026-04-30T12:33:00+00:00","index":34,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-07T14:09:16+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-11 14:22:30","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9050779","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9050779","identity":"rs-9050779","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0