FAIR in practice: minimum metadata schema for bioinformatics analytics by machines | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article FAIR in practice: minimum metadata schema for bioinformatics analytics by machines Daphne Wijnbergen, Núria Queralt-Rosinach, Valérie Barbié, Emma Verkinderen, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6396831/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Background One pillar of FAIR principles adoption is Reusability by machines to enable for example more efficient data analytics in fields such as Bioinformatics. However, it is not clear to what extent current metadata exposed by datasets and tools in common repositories enable this. In practice, metadata often lacks in machine actionability due to incomplete standardised metadata and lack of ontological descriptions. Results In this work, we identified minimal metadata that is needed to improve the machine actionability of bioinformatics tools and proposed a schema to address current limitations. The schema consists of metadata properties for the identification, selection, validation, and execution of tools. We also aligned this metadata to the metadata of datasets, in order to improve their integration for analytics by machines. Conclusions The identified minimal metadata improves the machine actionability of tools and data, and can be incorporated into platforms for tool and data sharing, and in FAIR infrastructures. FAIR metadata machine-actionability bioinformatics Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction The reuse of data is crucial to maximize the value gained from existing data resources, which are often the result of significant investments in terms of money and time. For example, in biology, new discoveries can be made using existing datasets by integrating them with other (new or existing) datasets, or by applying new tools and hypotheses. To facilitate the reuse of data, the FAIR principles were coined by Wilkinson et al. [ 1 ]. These principles show how the findability, accessibility, interoperability and reusability of data can be improved through, for example, metadata provision and formal knowledge representation. In practice, the FAIR principles have been applied in various forms to enable human and machine data analysis in heterogeneous data types, including, heterogeneous genomic data [ 2 ], (multi-)omics data [ 3 ], and rare disease patient registries [ 4 ]. The need for adhering to the FAIR principles was especially highlighted during the COVID pandemic, when the reuse of data was hindered by limitations in data sharing (e.g., unstable URLs and limited description of the data) during a crisis in which efficient reuse and integration of data was crucial [ 5 ]. In contrast to many data sharing standards, the FAIR principles have been developed with a particular focus on the machine actionability of data (e.g. by recommending formal knowledge description) which is essential for enabling larger scaled analysis by allowing computational agents to take over tasks from analysts [ 1 ]. The machine actionability of data will become increasingly more important for computational data-driven research especially as AI advances. In order to scalably reuse data in practice, it becomes a necessity to perform data analysis on a FAIR ecosystem so that datasets from heterogeneous sources can be found simultaneously without consulting every source individually. This is especially important for rare diseases for which data is scarce. However, being able to link datasets from different sources into a FAIR knowledge graph is a huge challenge. In the Swiss Personalized Health Network (SPHN) for example, a national FAIR data schema for health data ( https://www.biomedit.ch/rdf/sphn-schema/sphn ) and a corresponding toolstack is developed to ensure hospitals deliver interoperable data that can be easily used in research projects [ 6 ]. Similarly, in the European Joint Programme on Rare Diseases (EJP RD), a FAIR infrastructure was developed in order to allow the discovery and privacy preserving accessibility of various resources such as patient registries, biobanks and datasets in an interoperable manner within its Virtual Platform [ 7 , 8 ]. Another example, currently undergoing development, is the FAIR Data Train [ 9 ]. This is a FAIR federated ecosystem consisting most importantly of FAIR Data Stations containing datasets and FAIR Data Trains containing algorithms, which both require metadata to empower computational agents to perform analysis. Despite the introduction of the FAIR principles, current metadata standards and practices in data repositories, as for example in omics data repositories, remain a bottleneck for analytics by humans, and even more so for machines [ 10 – 12 ]. One reason for this bottleneck is that current metadata standards and practices are often incomplete, and thus are missing elements that would greatly benefit the reuse not only of datasets but also of computational tools in bioinformatics. Additionally, most metadata are focused on human readability, and as a result are not accessible for machines. For example, a lot of information only exists in textual descriptions of data instead of using ontological definitions. Like datasets, tools in bioinformatics can also benefit from reuse. For this reason, the FAIR principles are increasingly being applied to software. For example, the FAIR4RS initiative has created an adaptation of the FAIR principles for analytical tools and pipelines to improve their reproducibility and reusability [ 13 , 14 ]. In this adaptation the principles have been modified to be applicable to software. De Visser et al. have followed up on the FAIR4RS initiative by recommending practical applications to increase the FAIRness of workflows [ 15 ]. For example, uploading a workflow in Common Workflow Language [ 16 ] to WorkflubHub [ 17 ] with sufficient metadata will make it more reusable [ 18 , 19 ]. Despite these efforts, much software is still contained in repositories with insufficient metadata for machine actionable analysis and hampered machine readability of the metadata. For example, memory requirements are not reported in repositories such as bio.tools [ 20 ], CRAN [ 21 ] and BioConductor [ 22 ] though this is essential for understanding whether a system can run the tool. Similarly, software maturity is not reported in Schema.org, CRAN and BioConductor, despite its usefulness in finding tools that are established and thus likely to be reliable. Finally, despite much focus on the FAIRness and metadata of data, and to a lesser degree on tools, currently popular metadata standards for FAIR data and tools have not been investigated in a connected manner. This connected view is important, since for machine based analytics, it is necessary that both tools and data are FAIR. The metadata should ideally also be aligned to facilitate integrated use of data and tools. In this paper, we propose a metadata schema that addresses limitations of current metadata for tools and data and enables machine actionability of computational bioinformatics tools and data by connecting their metadata. To achieve this goal, we review current metadata standards and identify what metadata for tools is minimally necessary for the purpose of machine actionability. We also identify and propose which metadata is needed for interactions between tools and data to enable FAIR based data analytics. Methods Requirements and strategy To steer our minimal metadata schemas (one for the machine actionability of tools and one for the alignment of this schema to the metadata of datasets) towards a clearly established direction, we defined several goals: The metadata should facilitate the reuse of data and tools The metadata should facilitate the connection between data and tools The metadata should follow the FAIR principles and consequently be usable by machines The metadata should be valuable for the biomedical science and bioinformatics domains Our strategy followed a community-driven approach for decision-making that consisted of forming a group of stakeholders that are involved in the ELIXIR Rare Disease Community. These stakeholders represent various domains within bioinformatics and healthcare data such as genomics, transcriptomics, health infrastructure and FAIR. Within this group, we organized regular calls in order to jointly investigate, identify and then decide what the minimal metadata schemas should consist of, starting from selected tools widely used in the RD community. Approach Our approach consists of several steps (Fig. 1 ), which are described in more detail in the following sections. First, we selected a list of tools that are used in the Rare Disease Community by conducting a survey and querying bio.tools for tools tagged with “Rare Disease”. For a selection of popular tools from this list, we identified metadata that is available for these tools in common repositories and platforms for tools. We then aligned these lists to each other to compare them and to identify which metadata is necessary for machine actionability. Finally, we made an alignment between metadata for data, and metadata for tools to increase their interoperability. On the resulting minimal metadata, we executed several queries to compare machine actionability of resources with and without FAIRification. Identification of minimal metadata for tools To gain insight into the current state of metadata for bioinformatics tools and to build upon this, we identified which metadata is currently published in repositories. For this purpose, we selected a sample of bioinformatics tools used in the rare disease community. We sent out an email survey to members of the ELIXIR Rare Diseases Community [ 23 ] and the EJP RD [ 24 ], asking which tools they used in their research and the function of each tool. Additionally, we identified tools used in rare disease research by querying bio.tools for tools tagged with “Rare Disease”. We made a selection from these lists based on the popularity of the tools, which contained the following tools: LIMMA [ 25 ], ORVAL [ 26 ], orsum [ 27 ], Variomes [ 28 ], RD-Connect GPAP [ 29 ], MOGAMUN [ 30 ], BridgeDB [ 31 ], VEP [ 32 ], PathVisio [ 33 ], mixomics [ 34 , 35 ], ODAMnet [ 36 ] and Robust Rank Aggregation [ 37 ]. For these tools, we identified the available metadata in one or more platforms that facilitate the sharing of tools, specifically bio.tools [ 20 ], CRAN.R [ 21 ], Bioconductor [ 22 ] and the EJP RD Virtual Platform [ 7 , 8 ]. Metadata was retrieved manually using the available web interfaces and APIs during a time window from October 9 to October 27 in 2023. In order to further investigate the retrieved metadata, we created a spreadsheet with the metadata properties on the rows, and different tools and repositories on the columns. For each combination of property and tool/repository, we filled in the metadata property or metadata value. In order to initiate the alignment of metadata properties for different software repositories, we used crosswalk [ 38 ] from the CodeMeta Project [ 29 ] as a base since this project is devoted to providing a table to connect the diverse metadata. Metadata not included in crosswalk was added at the bottom. Next, using the spreadsheet, based on each platform and tools metadata, we identified for each property whether it is necessary for analytics by machines. Additionally, we identified which FAIR principle is satisfied by each metadata property in the set of minimal metadata. Alignment of metadata for data and tools To create an alignment of metadata for data and tools, we jointly investigated the minimal tool metadata that we identified previously, and voted on whether this metadata property is also applicable to datasets. For the purposes of this study study, we defined two degrees of alignment. For the first degree, it is simply a metadata property that is applicable to both tools and data. For the second degree, it is a metadata property that helps a computational agent know how the tool and dataset interact. For example, by specifying the file and input types, or by specifying storage requirements. Every metadata property that satisfied one of these definitions, was added to the aligned metadata for data and tools. Evaluation of using metadata with and without FAIRification We compared the performance of FAIR metadata to non-FAIR metadata to assess their machine actionability. In order to do so, we created three different queries (overview shown in Fig. 2 ). In the first query, we searched for datasets about the disease “Inclusion Body Myositis” in both the Gene Expression Omnibus (GEO), and in FAIR Data Points [ 40 ] in December 2023. In GEO ( https://www.ncbi.nlm.nih.gov/geo/ ) we used the search term “Inclusion Body Myositis”, while in the FAIR Data Points ( https://index.vp.ejprarediseases.org/ ) , we searched for datasets with an ontology term for “Inclusion Body Myositis” (DOID:3429 from the Human Disease Ontology [ 41 ]). We calculated the number of true and false positives from the results that we retrieved by querying both repositories. In the second query, we performed a query that matches datasets to tools by using ontology terms from a shared ontology. Specifically, we used terms from the EDAM ontology [ 42 ], which is an ontology for bioscientific data analysis and data management, and includes concepts such as data types, formats and operations. In this query, we queried an Inclusion Body Myositis dataset for its ontological description by DCAT themes [ 43 ]. We then searched for tools that have one of these themes as an input. Finally, in the third query, to investigate the use of information embedded in ontologies, we performed a query that finds datasets for a group of related disorders by using the semantic information that is embedded in the Human Disease Ontology [ 41 ]. Specifically, we searched for datasets for all diseases in the group of “myositis” (DOID:633). All queries are available at https://github.com/LUMC-BioSemantics/Tool-and-data-queries . Results Recommended minimal metadata of tools for machine actionability We identified 32 tools used in rare disease research in the survey, and 248 tools tagged with rare disease in bio.tools. For 12 tools that were selected based on their popularity, we identified their metadata in several platforms. Based on this metadata, a selection of minimal metadata for machine actionable analysis was created, as shown in Fig. 3 . Since the resulting metadata list was relatively long, we divided the metadata into four groups representing four subsequent steps in machine actionable analysis. First, to be identified, a tool needs an identifier and a name. Second, to select tools that are applicable for a dataset, metadata such as the application category, permissions and license are needed. Third, for the validation of the tool, multiple quality indicators, such as the maturity, development status, and use in different communities are valuable. Finally, for the execution of the tools, metadata covering some computational aspects are necessary. This includes metadata such as hardware and software requirements, but also metadata that is needed to know how the tool should be used. For each metadata property, we indicate the relevant FAIR principles. An alignment of the metadata properties to the source repositories is shown in Table 1 and Additional file 1, and can be used to see where each property fits in the source metadata model. Table 1 Snapshot of mappings for the identified minimal metadata in several repositories. The full table is shown in Additional file 1. Property CodeMeta/Schema.org Bio.tools EJP RD/DCAT CRAN Bioconductor Identifier schema:Thing/identifier tool::biotoolsID dct:identifier Package - Name schema:Thing/name tool::name dct:title Title Title Description schema:Thing/description tool::description dct:description Description Description Keywords schema:CreativeWork/keywords tool::topic (term, uri - both EDAM) dct:keyword - - Accessibility - tool::accessibility dct:accessRights - - Recommended common metadata to connect tools and datasets Starting from the minimal metadata that was identified for tools, we have aligned the metadata for tools and datasets. An overview of the aligned metadata properties is shown in Fig. 4 . This alignment resulted in a schema that contains metadata that is relevant for both datasets and tools, and can be used to establish connections between the two for data analysis. One important property for alignment is the data input and output type, which can be described for example with EDAM ontology terms. Here, a tool might be described as requiring FASTQ files ( http://edamontology.org/format_1930 ) as input, while a dataset can be described as being in the FASTQ format. These metadata properties can then be used to match the tool to the dataset. Another example is the license, since it is necessary to have the rights to use both the data and the tool, when performing an analysis. For example, a computational agent may check the licenses for both a dataset and a tool to understand under what conditions they can be accessed, used and shared by the organisation. Finally, the storage requirement for both data and tools needs to be added together to get the total amount of storage that is required to perform an analysis. Machine actionability of tools and data with and without FAIRification We performed queries ( https://github.com/LUMC-BioSemantics/Tool-and-data-queries ) for three tool and data identification scenarios. When querying GEO for the term “Inclusion Body Myositis”, we found 11 true positives and 4 false positives. The false positives were retrieved because the name “Inclusion Body Myositis” or “IBM” is mentioned in the description as free text either because the study is related to Inclusion Body Myositis, or because the statistical software SPSS from the company “IBM” has been used in the study. In contrast, when querying FAIR Data Points for the ontology term “Inclusion Body Myositis” (DOID:3429), we found 4 true positives and 0 false positives (in this case, there are less results since adoption of the FDP is lower). In addition, findability time for GEO was 5 minutes and 9 seconds due to needed manual curation, while for FAIR metadata this was 13 seconds. For the second query, we queried one of the identified transcriptomics datasets for Inclusion Body Myositis in a FAIR Data Point Index for its data format. Here we found the “expression data” (EDAM:2603) EDAM term. Using this EDAM term, we then queried FAIR Data Points for tools that have this EDAM term as their input. As a result, out of 4 tools that were indexed in the FAIR Data Point Index, we found the DESEQ2 tool, which performs differential gene expression testing, and is applicable on this dataset. In the third query, we first queried the ORDO ontology [ 44 ] for diseases that are a subclass of “myositis”. Using this list of diseases, we subsequently queried FAIR Data Points for datasets that are described with at least one of these disease terms. We found 4 datasets for Inclusion Body Myositis out of the 21 total dataset that were connected to this FAIR Data Point index. Discussion In this work, we identified important gaps in current metadata standards for machine actionable bioinformatics analysis. Our main findings are a lack of machine readability, and a limited alignment of metadata for tools and datasets. As a solution, we propose minimal metadata for tools, and an alignment of metadata for tools and datasets. We identified a minimal set of metadata properties that are minimally necessary for machine actionable analysis. By drawing from existing repositories, we include properties that were only included in one of some of these repositories. This includes memory requirements and software maturity, which are important for machine actionability. Additionally, we aligned the tool metadata to tool metadata which provided several useful alignments. For example, describing inputs and outputs of tools, and format of data, allows computer agents to know which tool works for which data. Another example is storage requirements, where the storage requirement of a tool and dataset has to be summed to get the total required storage. This alignment can improve the interoperability between tools and datasets for machine actionable analysis. Queries on our metadata schemas showed that the machine actionability aspect of FAIR has several important benefits in research. For example, we were able to find datasets with more precision by incorporating ontology terms into the search instead of plain text. This highlights the importance of the use of ontologies to avoid ambiguity for terms and acronyms that have multiple meanings. Another query showed that ontologies also increased interoperability between datasets and tools, by allowing them to be described with identical or mapped ontology terms, ensuring the machine understands that the resources annotated with these terms have the same meaning without being affected by differences in spelling or ambiguity of natural language. Finally, we found that the use of FAIR and ontologies also enabled machines to use knowledge embedded in ontologies, for example, to learn which diseases belong to a disease group. This makes queries more powerful, and thus improves machine actionability. There are several limitations to note in this work. The scope of our work is focused on biomedical sciences and bioinformatics and although our work could be applicable to other domains, existing data sharing models and standards differ between fields. Further, the comparison between analysis with and without FAIRification has limited quantification, which limits the strength of its conclusions. With a higher adoption of FAIR comparable to platforms that are not FAIRified, a more comprehensive comparison could have been made. In future work, the minimal metadata that we identified can serve as a reference for data and software sharing platforms to improve the machine actionability of their metadata and improve the alignment between tools and datasets. Our minimal metadata can also be implemented in FAIR ecosystems such as the SPHN and EJPRD. The minimal metadata would especially benefit federated infrastructure such as the FAIR Data Train [ 9 ], which requires a high degree of machine actionable communication between FAIR Data Stations (containing datasets) and FAIR Data Trains (containing tools). Our minimal metadata schema can serve as a starting point in implementing this communication, for example, by describing the memory requirements of a train, so that sufficient memory is allocated by a FAIR Data Station when running a tool. The minimal metadata that we identified can also be expanded further in order to achieve an increasingly higher degree of machine actionability for tools and datasets. Efforts could for example focus on improving the interaction between tools and datasets by enabling more granular file format descriptions of datasets and of tool inputs and outputs thereby connecting the two in a way that is precise enough for computer agents to work increasingly more independently. Conclusion In this work, we propose two metadata schemas to address current limitations of metadata. One is focused on minimal metadata for tools, and the second one is focused on alignment between metadata of data and tools. Through queries, we showed that our metadata schema performs well in terms of precision, aligning data with tools, and allowing for more complex and powerful queries. These schemas can be used in platforms for bioinformatics tools and in FAIR infrastructures, but can also apply beyond the Rare Disease domain, and serve other areas. Declarations Ethics approval and consent to participate Not applicable Consent for publication Not applicable Availability of data and materials The queries that support our findings are available on GitHub, https://github.com/LUMC-BioSemantics/Tool-and-data-queries. Competing interests The authors declare that they have no competing interests Funding This work was funded by ELIXIR, the research infrastructure for life science data. This initiative has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement N°825575 (The European Joint Programme Rare Diseases, EJP RD). This project is supported by the Innovative Health Initiative Joint Undertaking (IHI JU) under grant agreement No 101132943. The JU receives support from the European Union’s Horizon Europe research and innovation programme, COCIR, EFPIA, EuropaBio, MedTech Europe, Vaccines Europe (and the FSHD Society and TREAT-NMD Services Ltd as contributing partners participating in the project). The project will also receive funding from UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee [grant numbers 10105921, 10103989, and 10083579]. NQR is partially funded as contributing partners of SYNTHIA and ERDERA projects. SYNTHIA is supported by the Innovative Health Initiative Joint Undertaking (IHI JU) under grant agreement No 101172872. The JU receives support from the EU Horizon Europe programme and COCIR, EFPIA, Europa Bio, MedTech Europe, and Vaccines Europe and DNV. ERDERA has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement N°101156595. Authors' contributions DW: Methodology, Software, Investigation, Writing - Original draft. NQR: Funding acquisition, Conceptualization, Methodology, Writing - Review & Editing. VB: Investigation, Writing - Review & Editing. EV: Software,Investigation, Writing - Review & Editing. NB: Investigation, Writing - Review & Editing. AJ: Funding acquisition, Writing - Review & Editing. PACtH: Writing - Review & Editing. CC: Writing - Review & Editing. MR: Writing - Review & Editing. EM: Funding acquisition, Writing - Review & Editing, Methodology, Supervision. Acknowledgements We thank the members of the ELIXIR Rare Disease Implementation Study for the collaboration and fruitful discussions. We would also like to thank all the members of the ELIXIR Rare Disease Community and EJP RD for contributing to our survey and for their valuable input that was foundational for this study. References Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. Comment: The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:1–9. van der Horst E, Kaliyaperumal R, Tatum Z, Thompson M, Schultes E, Mina E et al. Finding novel associations across domains using linked data: a case study on genetic variants disrupting transcription start sites. Liao X, Niehues A, de Visser C, Huang J, Doornbos C, Ederveen THA et al. FAIR Data Cube, a FAIR data infrastructure for integrated multi-omics data analysis. 2023;:2023.04.23.23289000. dos Santos Vieira B, Bernabé CH, Zhang S, Abaza H, Benis N, Cámara A, et al. Towards FAIRification of sensitive and fragmented rare disease patient data: challenges and solutions in European reference network registries. Orphanet J Rare Dis. 2022;17:436. Herre B, Rodés-Guirao L, Mathieu E, Ritchie H, Giattino C, Hasell J, et al. Best practices for government agencies to publish data: lessons from COVID-19. Lancet Public Health. 2024;9:e407–10. Touré V, Krauss P, Gnodtke K, Buchhorn J, Unni D, Horki P, et al. FAIRification of health-related data using semantic web technologies in the Swiss Personalized Health Network. Sci Data. 2023;10:127. Kaliyaperumal R, Wilkinson MD, Moreno PA, Benis N, Cornet R, dos Santos Vieira B, et al. Semantic modelling of common data elements for rare disease registries, and a prototype workflow for their deployment over registry data. J Biomed Semant. 2022;13:9. ejp-rd-vp/resource-metadata-schema. 2024. da Silva Santos LB, Burger K, Kaliyaperumal R. FAIR Data Train specifications. 2023. https://specs.fairdatatrain.org/ . Accessed 19 Feb 2025. Wang Z, Lachmann A, Ma’ayan A. Mining data and metadata from the gene expression omnibus. Biophys Rev. 2019;11:103–10. Lim N, Tesar S, Belmadani M, Poirier-Morency G, Mancarci BO, Sicherman J, et al. Curation of over 10 000 transcriptomic studies to enable data reuse. Database. 2021;2021:baab006. Carbon S, Champieux R, McMurry JA, Winfree L, Wyatt LR, Haendel MA. An analysis and metric of reusable data licensing practices for biomedical resources. PLoS ONE. 2019;14:e0213090. Barker M, Chue Hong NP, Katz DS, Lamprecht A-L, Martinez-Ortiz C, Psomopoulos F, et al. Introducing the FAIR Principles for research software. Sci Data. 2022;9:622. Chue Hong NP, Katz DS, Barker M, Lamprecht A-L, Martinez C, Psomopoulos FE et al. FAIR Principles for Research Software (FAIR4RS Principles). 2022. https://doi.org/10.15497/RDA00068 de Visser C, Johansson LF, Kulkarni P, Mei H, Neerincx P, van der Velde KJ, et al. Ten quick tips for building FAIR workflows. PLOS Comput Biol. 2023;19:e1011369. Crusoe MR, Abeln S, Iosup A, Amstutz P, Chilton J, Tijanić N, et al. Methods Included: Standardizing Computational Reuse and Portability with the Common Workflow Language. Commun ACM. 2022;65:54–63. Gustafsson OJR, Wilkinson SR, Bacall F, Pireddu L, Soiland-Reyes S, Leo S et al. WorkflowHub: a registry for computational workflows. 2024. Wijnbergen D, Johari M, Ozisik O, ‘t Hoen PAC, Ehrhart F, Baudot A, et al. Multi-omics analysis in inclusion body myositis identifies mir-16 responsible for HLA overexpression. Orphanet J Rare Dis. 2025;20:27. Wijnbergen D, Johari M, Udd B, Roos M, Mina E. Inclusion Body Myositis Active Subnetwork Identification Workflow. 2024. Ison J, Ienasescu H, Chmura P, Rydza E, Ménager H, Kalaš M, et al. The bio.tools registry of software tools and data resources for the life sciences. Genome Biol. 2019;20:164. R Core Team. R: A Language and Environment for Statistical Computing. Gentleman RC, Carey VJ, Bates DM, Bolstad B, Dettling M, Dudoit S, et al. Bioconductor: open software development for computational biology and bioinformatics. Genome Biol. 2004;5:R80. Rare Diseases Community | ELIXIR. https://elixir-europe.org/communities/rare-diseases . Accessed 20 Feb 2025. EJP, RD – European Joint Programme on Rare Diseases. https://www.ejprarediseases.org/ . Accessed 20 Feb 2025. Ritchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43:e47. Renaux A, Papadimitriou S, Versbraegen N, Nachtegael C, Boutry S, Nowé A, et al. ORVAL: a novel platform for the prediction and exploration of disease-causing oligogenic variant combinations. Nucleic Acids Res. 2019;47:W93–8. Ozisik O, Térézol M, Baudot A. orsum: a Python package for filtering and comparing enrichment analyses using a simple principle. BMC Bioinformatics. 2022;23:293. Pasche E, Mottaz A, Caucheteur D, Gobeill J, Michel P-A, Ruch P. Variomes: a high recall search engine to support the curation of genomic variants. Bioinformatics. 2022;38:2595–601. Laurie S, Piscia D, Matalonga L, Corvó A, Fernández-Callejo M, Garcia-Linares C, et al. The RD-Connect Genome-Phenome Analysis Platform: Accelerating diagnosis, research, and gene discovery for rare diseases. Hum Mutat. 2022;43:717–33. Novoa-del-Toro EM, Mezura-Montes E, Vignes M, Térézol M, Magdinier F, Tichit L, et al. A multi-objective genetic algorithm to find active modules in multiplex biological networks. PLOS Comput Biol. 2021;17:e1009263. van Iersel MP, Pico AR, Kelder T, Gao J, Ho I, Hanspers K, et al. The BridgeDb framework: standardized access to gene, protein and metabolite identifier mapping services. BMC Bioinformatics. 2010;11:5. McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:1–14. Kutmon M, van Iersel MP, Bohler A, Kelder T, Nunes N, Pico AR, et al. PathVisio 3: An Extendable Pathway Analysis Toolbox. PLOS Comput Biol. 2015;11:e1004085. Lê Cao K-A, González I, Déjean S. integrOmics: an R package to unravel relationships between two omics datasets. Bioinformatics. 2009;25:2855–6. Lê Cao K-A, Rohart F, Gonzalez I, Dejean S. mixOmics: Omics Data Integration Project. R package version 6.1.1. 2016. Térézol M, Baudot A, Ozisik O, ODAMNet. A Python package to identify molecular relationships between chemicals and rare diseases using overlap, active module and random walk approaches. SoftwareX. 2024;26:101701. Kolde R, Laur S, Adler P, Vilo J. Robust rank aggregation for gene list integration and meta-analysis. Bioinformatics. 2012;28:573–80. codemeta/crosswalk.csv at master · codemeta/codemeta. https://github.com/codemeta/codemeta/blob/master/crosswalk.csv . Accessed 21 Jan 2025. The CodeMeta Project. https://codemeta.github.io/ . Accessed 21 Jan 2025. da Silva Santos LOB, Burger K, Kaliyaperumal R, Wilkinson MD. FAIR Data Point: A FAIR-Oriented Approach for Metadata Publication. Data Intell. 2023;5:163–83. Schriml LM, Munro JB, Schor M, Olley D, McCracken C, Felix V, et al. The Human Disease Ontology 2022 update. Nucleic Acids Res. 2022;50:D1255–61. Black M, Lamothe L, Eldakroury H, Kierkegaard M, Priya A, Machinda A et al. EDAM: the bioscientific data analysis ontology (update 2021). F1000Research. 2022;11. Albertoni R, Browning D, Cox S, Gonzalez Beltran A, Perego A, Winstanley P et al. Data Catalog Vocabulary (DCAT) - Version 2. 2020. https://www.w3.org/TR/vocab-dcat-2/ . Accessed 10 Feb 2023. Vasant D, Chanas L, Malone J, Hanauer M, Olry A, Jupp S et al. ORDO: An Ontology Connecting Rare Disease, Epidemiology and Genetic Data. 2014. Additional Declarations No competing interests reported. Supplementary Files additionalfile1.xls Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 31 Aug, 2025 Reviews received at journal 02 Jun, 2025 Reviews received at journal 28 May, 2025 Reviewers agreed at journal 25 May, 2025 Reviewers agreed at journal 23 May, 2025 Reviewers agreed at journal 23 May, 2025 Reviewers invited by journal 02 May, 2025 Editor assigned by journal 11 Apr, 2025 Submission checks completed at journal 11 Apr, 2025 First submitted to journal 07 Apr, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6396831","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":450969906,"identity":"80c6c6df-630d-46e7-a518-d47d62a48900","order_by":0,"name":"Daphne Wijnbergen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1UlEQVRIiWNgGAWjYDCCA0AkASJB4EMBmGI8QLQWxhkGcEH8WuAkMw8xWvgO8BgesNxxR868f/GxxzYGh6P5G3gP4NUieYDH4IDkmWfGMjeepRvnGBzOnXGALwGvFoMDvBsOSLYdTpwhccZMGqRlAwPQEGK01M+QOP9N2oIULQkS/D1s0gzEaJE8zP8BqOWZ4QwJNjPJHoP03BmHCWjhO96W/Fmy7Y68BP/hZxI/Kqxz+9t7DB/g08LADEQSIIZEAkKEIGD8ACL58bpmFIyCUTAKRjIAAJG4UhtUeewYAAAAAElFTkSuQmCC","orcid":"","institution":"Leiden University Medical Center","correspondingAuthor":true,"prefix":"","firstName":"Daphne","middleName":"","lastName":"Wijnbergen","suffix":""},{"id":450969910,"identity":"82ad3c4a-9115-49a2-ac0c-1788c649e978","order_by":1,"name":"Núria Queralt-Rosinach","email":"","orcid":"","institution":"Leiden University Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Núria","middleName":"","lastName":"Queralt-Rosinach","suffix":""},{"id":450969914,"identity":"2f1f3854-a647-4fe0-9378-48290e11f5d7","order_by":2,"name":"Valérie Barbié","email":"","orcid":"","institution":"Swiss Institute of Bioinformatics","correspondingAuthor":false,"prefix":"","firstName":"Valérie","middleName":"","lastName":"Barbié","suffix":""},{"id":450969916,"identity":"ba07bcdb-eacf-4a67-a953-918a6b8d2a2e","order_by":3,"name":"Emma Verkinderen","email":"","orcid":"","institution":"Université Libre de Bruxelles","correspondingAuthor":false,"prefix":"","firstName":"Emma","middleName":"","lastName":"Verkinderen","suffix":""},{"id":450969919,"identity":"ed135907-add2-485d-abe3-4a7cac4fefd3","order_by":4,"name":"Nirupama Benis","email":"","orcid":"","institution":"Amsterdam UMC location University of Amsterdam","correspondingAuthor":false,"prefix":"","firstName":"Nirupama","middleName":"","lastName":"Benis","suffix":""},{"id":450969921,"identity":"0be608d2-1d7e-426a-9245-06e7a4e8cc53","order_by":5,"name":"Annika Jacobsen","email":"","orcid":"","institution":"Leiden University Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Annika","middleName":"","lastName":"Jacobsen","suffix":""},{"id":450969923,"identity":"8859109b-119d-4738-83f6-f24bbc1d7333","order_by":6,"name":"Peter A.C. ’t Hoen","email":"","orcid":"","institution":"Radboud University Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Peter","middleName":"A.C. ’t","lastName":"Hoen","suffix":""},{"id":450969925,"identity":"02a38d69-b9e8-4567-84b0-8623172d3b2d","order_by":7,"name":"Claudio Carta","email":"","orcid":"","institution":"Istituto Superiore di Sanità","correspondingAuthor":false,"prefix":"","firstName":"Claudio","middleName":"","lastName":"Carta","suffix":""},{"id":450969926,"identity":"e8fa5a49-ecd4-44da-bfc3-c0aac668abff","order_by":8,"name":"Marco Roos","email":"","orcid":"","institution":"Leiden University Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Marco","middleName":"","lastName":"Roos","suffix":""},{"id":450969927,"identity":"7d96c970-bc84-4f23-9c58-0054616eebd5","order_by":9,"name":"Eleni Mina","email":"","orcid":"","institution":"Leiden University Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Eleni","middleName":"","lastName":"Mina","suffix":""}],"badges":[],"createdAt":"2025-04-07 18:38:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6396831/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6396831/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":82213099,"identity":"910e82cf-a445-48b9-8e4f-fa85602da98f","added_by":"auto","created_at":"2025-05-07 20:36:41","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":649222,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow for the identification of the minimal tool metadata, and alignment to dataset metadata. Sources of data and knowledge are shown on the left side of the figure, while the results are displayed on the right side of the figure. The processes by which the results were produced from the inputs are shown in the middle.\u003c/p\u003e","description":"","filename":"figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-6396831/v1/b2bef7a552a6b7c2566a2c73.png"},{"id":82213101,"identity":"d8b24355-fef5-4d38-9197-993921f89002","added_by":"auto","created_at":"2025-05-07 20:36:41","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":443655,"visible":true,"origin":"","legend":"\u003cp\u003ethree queries we executed to investigate the impact of FAIRifications on finding tools and datasets.\u003c/p\u003e","description":"","filename":"figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-6396831/v1/fb06fdc2d06b059a00b4684b.png"},{"id":82213320,"identity":"7d19f7ff-8b67-4faf-9b15-6e28fc585972","added_by":"auto","created_at":"2025-05-07 20:44:41","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":746095,"visible":true,"origin":"","legend":"\u003cp\u003eThe identified minimal metadata for tools divided in four sequential categories: identification, selection, validation and execution.\u003c/p\u003e","description":"","filename":"figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-6396831/v1/d08ce785f835a745613cfd45.png"},{"id":82213100,"identity":"ec3aada6-d2dd-48c7-b1dc-f09289a2589e","added_by":"auto","created_at":"2025-05-07 20:36:41","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":224381,"visible":true,"origin":"","legend":"\u003cp\u003eAligned metadata for both tools and data. These metadata properties can be matched between tools and data, and thereby allow interaction between tools and data to perform bioinformatics analysis.\u003c/p\u003e","description":"","filename":"figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-6396831/v1/d5f66e02c27d1ebae229a916.png"},{"id":82213680,"identity":"3571e066-b49c-420d-bc87-b8373127dcd7","added_by":"auto","created_at":"2025-05-07 21:00:42","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2623069,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6396831/v1/3346f871-85de-4d2f-a29e-bd93c5653899.pdf"},{"id":82213321,"identity":"a3f48762-29a9-48f1-9ed0-bd755ad30cd7","added_by":"auto","created_at":"2025-05-07 20:44:41","extension":"xls","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":27136,"visible":true,"origin":"","legend":"","description":"","filename":"additionalfile1.xls","url":"https://assets-eu.researchsquare.com/files/rs-6396831/v1/63078788414c590b3bb1aee2.xls"}],"financialInterests":"No competing interests reported.","formattedTitle":"FAIR in practice: minimum metadata schema for bioinformatics analytics by machines","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe reuse of data is crucial to maximize the value gained from existing data resources, which are often the result of significant investments in terms of money and time. For example, in biology, new discoveries can be made using existing datasets by integrating them with other (new or existing) datasets, or by applying new tools and hypotheses. To facilitate the reuse of data, the FAIR principles were coined by Wilkinson et al. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. These principles show how the findability, accessibility, interoperability and reusability of data can be improved through, for example, metadata provision and formal knowledge representation. In practice, the FAIR principles have been applied in various forms to enable human and machine data analysis in heterogeneous data types, including, heterogeneous genomic data [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], (multi-)omics data [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], and rare disease patient registries [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The need for adhering to the FAIR principles was especially highlighted during the COVID pandemic, when the reuse of data was hindered by limitations in data sharing (e.g., unstable URLs and limited description of the data) during a crisis in which efficient reuse and integration of data was crucial [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In contrast to many data sharing standards, the FAIR principles have been developed with a particular focus on the machine actionability of data (e.g. by recommending formal knowledge description) which is essential for enabling larger scaled analysis by allowing computational agents to take over tasks from analysts [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The machine actionability of data will become increasingly more important for computational data-driven research especially as AI advances.\u003c/p\u003e \u003cp\u003eIn order to scalably reuse data in practice, it becomes a necessity to perform data analysis on a FAIR ecosystem so that datasets from heterogeneous sources can be found simultaneously without consulting every source individually. This is especially important for rare diseases for which data is scarce. However, being able to link datasets from different sources into a FAIR knowledge graph is a huge challenge. In the Swiss Personalized Health Network (SPHN) for example, a national FAIR data schema for health data (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.biomedit.ch/rdf/sphn-schema/sphn\u003c/span\u003e\u003cspan address=\"https://www.biomedit.ch/rdf/sphn-schema/sphn\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e and a corresponding toolstack is developed to ensure hospitals deliver interoperable data that can be easily used in research projects [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Similarly, in the European Joint Programme on Rare Diseases (EJP RD), a FAIR infrastructure was developed in order to allow the discovery and privacy preserving accessibility of various resources such as patient registries, biobanks and datasets in an interoperable manner within its Virtual Platform [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Another example, currently undergoing development, is the FAIR Data Train [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. This is a FAIR federated ecosystem consisting most importantly of FAIR Data Stations containing datasets and FAIR Data Trains containing algorithms, which both require metadata to empower computational agents to perform analysis.\u003c/p\u003e \u003cp\u003eDespite the introduction of the FAIR principles, current metadata standards and practices in data repositories, as for example in omics data repositories, remain a bottleneck for analytics by humans, and even more so for machines [\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. One reason for this bottleneck is that current metadata standards and practices are often incomplete, and thus are missing elements that would greatly benefit the reuse not only of datasets but also of computational tools in bioinformatics. Additionally, most metadata are focused on human readability, and as a result are not accessible for machines. For example, a lot of information only exists in textual descriptions of data instead of using ontological definitions.\u003c/p\u003e \u003cp\u003eLike datasets, tools in bioinformatics can also benefit from reuse. For this reason, the FAIR principles are increasingly being applied to software. For example, the FAIR4RS initiative has created an adaptation of the FAIR principles for analytical tools and pipelines to improve their reproducibility and reusability [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. In this adaptation the principles have been modified to be applicable to software. De Visser et al. have followed up on the FAIR4RS initiative by recommending practical applications to increase the FAIRness of workflows [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. For example, uploading a workflow in Common Workflow Language [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] to WorkflubHub [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] with sufficient metadata will make it more reusable [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Despite these efforts, much software is still contained in repositories with insufficient metadata for machine actionable analysis and hampered machine readability of the metadata. For example, memory requirements are not reported in repositories such as bio.tools [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], CRAN [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] and BioConductor [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] though this is essential for understanding whether a system can run the tool. Similarly, software maturity is not reported in Schema.org, CRAN and BioConductor, despite its usefulness in finding tools that are established and thus likely to be reliable.\u003c/p\u003e \u003cp\u003eFinally, despite much focus on the FAIRness and metadata of data, and to a lesser degree on tools, currently popular metadata standards for FAIR data and tools have not been investigated in a connected manner. This connected view is important, since for machine based analytics, it is necessary that both tools and data are FAIR. The metadata should ideally also be aligned to facilitate integrated use of data and tools. In this paper, we propose a metadata schema that addresses limitations of current metadata for tools and data and enables machine actionability of computational bioinformatics tools and data by connecting their metadata. To achieve this goal, we review current metadata standards and identify what metadata for tools is minimally necessary for the purpose of machine actionability. We also identify and propose which metadata is needed for interactions between tools and data to enable FAIR based data analytics.\u003c/p\u003e"},{"header":"Methods","content":" \u003cp\u003eRequirements and strategy\u003c/p\u003e \u003cp\u003eTo steer our minimal metadata schemas (one for the machine actionability of tools and one for the alignment of this schema to the metadata of datasets) towards a clearly established direction, we defined several goals:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eThe metadata should facilitate the reuse of data and tools\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe metadata should facilitate the connection between data and tools\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe metadata should follow the FAIR principles and consequently be usable by machines\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe metadata should be valuable for the biomedical science and bioinformatics domains\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eOur strategy followed a community-driven approach for decision-making that consisted of forming a group of stakeholders that are involved in the ELIXIR Rare Disease Community. These stakeholders represent various domains within bioinformatics and healthcare data such as genomics, transcriptomics, health infrastructure and FAIR. Within this group, we organized regular calls in order to jointly investigate, identify and then decide what the minimal metadata schemas should consist of, starting from selected tools widely used in the RD community.\u003c/p\u003e \u003cp\u003eApproach\u003c/p\u003e \u003cp\u003eOur approach consists of several steps (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), which are described in more detail in the following sections. First, we selected a list of tools that are used in the Rare Disease Community by conducting a survey and querying bio.tools for tools tagged with \u0026ldquo;Rare Disease\u0026rdquo;. For a selection of popular tools from this list, we identified metadata that is available for these tools in common repositories and platforms for tools. We then aligned these lists to each other to compare them and to identify which metadata is necessary for machine actionability. Finally, we made an alignment between metadata for data, and metadata for tools to increase their interoperability. On the resulting minimal metadata, we executed several queries to compare machine actionability of resources with and without FAIRification.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIdentification of minimal metadata for tools\u003c/p\u003e \u003cp\u003eTo gain insight into the current state of metadata for bioinformatics tools and to build upon this, we identified which metadata is currently published in repositories. For this purpose, we selected a sample of bioinformatics tools used in the rare disease community. We sent out an email survey to members of the ELIXIR Rare Diseases Community [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] and the EJP RD [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], asking which tools they used in their research and the function of each tool. Additionally, we identified tools used in rare disease research by querying bio.tools for tools tagged with \u0026ldquo;Rare Disease\u0026rdquo;. We made a selection from these lists based on the popularity of the tools, which contained the following tools: LIMMA [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], ORVAL [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], orsum [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], Variomes [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], RD-Connect GPAP [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], MOGAMUN [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], BridgeDB [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e], VEP [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], PathVisio [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], mixomics [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], ODAMnet [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e] and Robust Rank Aggregation [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. For these tools, we identified the available metadata in one or more platforms that facilitate the sharing of tools, specifically bio.tools [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], CRAN.R [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], Bioconductor [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] and the EJP RD Virtual Platform [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Metadata was retrieved manually using the available web interfaces and APIs during a time window from October 9 to October 27 in 2023.\u003c/p\u003e \u003cp\u003eIn order to further investigate the retrieved metadata, we created a spreadsheet with the metadata properties on the rows, and different tools and repositories on the columns. For each combination of property and tool/repository, we filled in the metadata property or metadata value. In order to initiate the alignment of metadata properties for different software repositories, we used crosswalk [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] from the CodeMeta Project [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] as a base since this project is devoted to providing a table to connect the diverse metadata. Metadata not included in crosswalk was added at the bottom. Next, using the spreadsheet, based on each platform and tools metadata, we identified for each property whether it is necessary for analytics by machines. Additionally, we identified which FAIR principle is satisfied by each metadata property in the set of minimal metadata.\u003c/p\u003e \u003cp\u003eAlignment of metadata for data and tools\u003c/p\u003e \u003cp\u003eTo create an alignment of metadata for data and tools, we jointly investigated the minimal tool metadata that we identified previously, and voted on whether this metadata property is also applicable to datasets. For the purposes of this study study, we defined two degrees of alignment. For the first degree, it is simply a metadata property that is applicable to both tools and data. For the second degree, it is a metadata property that helps a computational agent know how the tool and dataset interact. For example, by specifying the file and input types, or by specifying storage requirements. Every metadata property that satisfied one of these definitions, was added to the aligned metadata for data and tools.\u003c/p\u003e \u003cp\u003eEvaluation of using metadata with and without FAIRification\u003c/p\u003e \u003cp\u003eWe compared the performance of FAIR metadata to non-FAIR metadata to assess their machine actionability. In order to do so, we created three different queries (overview shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the first query, we searched for datasets about the disease \u0026ldquo;Inclusion Body Myositis\u0026rdquo; in both the Gene Expression Omnibus (GEO), and in FAIR Data Points [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e] in December 2023. In GEO (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e we used the search term \u0026ldquo;Inclusion Body Myositis\u0026rdquo;, while in the FAIR Data Points (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://index.vp.ejprarediseases.org/\u003c/span\u003e\u003cspan address=\"https://index.vp.ejprarediseases.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e, we searched for datasets with an ontology term for \u0026ldquo;Inclusion Body Myositis\u0026rdquo; (DOID:3429 from the Human Disease Ontology [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]). We calculated the number of true and false positives from the results that we retrieved by querying both repositories.\u003c/p\u003e \u003cp\u003eIn the second query, we performed a query that matches datasets to tools by using ontology terms from a shared ontology. Specifically, we used terms from the EDAM ontology [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e], which is an ontology for bioscientific data analysis and data management, and includes concepts such as data types, formats and operations. In this query, we queried an Inclusion Body Myositis dataset for its ontological description by DCAT themes [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. We then searched for tools that have one of these themes as an input.\u003c/p\u003e \u003cp\u003eFinally, in the third query, to investigate the use of information embedded in ontologies, we performed a query that finds datasets for a group of related disorders by using the semantic information that is embedded in the Human Disease Ontology [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Specifically, we searched for datasets for all diseases in the group of \u0026ldquo;myositis\u0026rdquo; (DOID:633). All queries are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/LUMC-BioSemantics/Tool-and-data-queries\u003c/span\u003e\u003cspan address=\"https://github.com/LUMC-BioSemantics/Tool-and-data-queries\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eRecommended minimal metadata of tools for machine actionability\u003c/p\u003e \u003cp\u003eWe identified 32 tools used in rare disease research in the survey, and 248 tools tagged with rare disease in bio.tools. For 12 tools that were selected based on their popularity, we identified their metadata in several platforms. Based on this metadata, a selection of minimal metadata for machine actionable analysis was created, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Since the resulting metadata list was relatively long, we divided the metadata into four groups representing four subsequent steps in machine actionable analysis. First, to be identified, a tool needs an identifier and a name. Second, to select tools that are applicable for a dataset, metadata such as the application category, permissions and license are needed. Third, for the validation of the tool, multiple quality indicators, such as the maturity, development status, and use in different communities are valuable. Finally, for the execution of the tools, metadata covering some computational aspects are necessary. This includes metadata such as hardware and software requirements, but also metadata that is needed to know how the tool should be used. For each metadata property, we indicate the relevant FAIR principles. An alignment of the metadata properties to the source repositories is shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Additional file 1, and can be used to see where each property fits in the source metadata model.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSnapshot of mappings for the identified minimal metadata in several repositories. The full table is shown in Additional file 1.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProperty\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCodeMeta/Schema.org\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBio.tools\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEJP RD/DCAT\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCRAN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eBioconductor\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eIdentifier\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eschema:Thing/identifier\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etool::biotoolsID\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edct:identifier\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePackage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eName\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eschema:Thing/name\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etool::name\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edct:title\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTitle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTitle\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eDescription\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eschema:Thing/description\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etool::description\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edct:description\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eKeywords\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eschema:CreativeWork/keywords\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etool::topic (term, uri - both EDAM)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edct:keyword\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAccessibility\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etool::accessibility\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edct:accessRights\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eRecommended common metadata to connect tools and datasets\u003c/p\u003e \u003cp\u003eStarting from the minimal metadata that was identified for tools, we have aligned the metadata for tools and datasets. An overview of the aligned metadata properties is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. This alignment resulted in a schema that contains metadata that is relevant for both datasets and tools, and can be used to establish connections between the two for data analysis. One important property for alignment is the data input and output type, which can be described for example with EDAM ontology terms. Here, a tool might be described as requiring FASTQ files (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://edamontology.org/format_1930\u003c/span\u003e\u003cspan address=\"http://edamontology.org/format_1930\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) as input, while a dataset can be described as being in the FASTQ format. These metadata properties can then be used to match the tool to the dataset. Another example is the license, since it is necessary to have the rights to use both the data and the tool, when performing an analysis. For example, a computational agent may check the licenses for both a dataset and a tool to understand under what conditions they can be accessed, used and shared by the organisation. Finally, the storage requirement for both data and tools needs to be added together to get the total amount of storage that is required to perform an analysis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eMachine actionability of tools and data with and without FAIRification\u003c/p\u003e \u003cp\u003eWe performed queries (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/LUMC-BioSemantics/Tool-and-data-queries\u003c/span\u003e\u003cspan address=\"https://github.com/LUMC-BioSemantics/Tool-and-data-queries\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e for three tool and data identification scenarios. When querying GEO for the term \u0026ldquo;Inclusion Body Myositis\u0026rdquo;, we found 11 true positives and 4 false positives. The false positives were retrieved because the name \u0026ldquo;Inclusion Body Myositis\u0026rdquo; or \u0026ldquo;IBM\u0026rdquo; is mentioned in the description as free text either because the study is related to Inclusion Body Myositis, or because the statistical software SPSS from the company \u0026ldquo;IBM\u0026rdquo; has been used in the study. In contrast, when querying FAIR Data Points for the ontology term \u0026ldquo;Inclusion Body Myositis\u0026rdquo; (DOID:3429), we found 4 true positives and 0 false positives (in this case, there are less results since adoption of the FDP is lower). In addition, findability time for GEO was 5 minutes and 9 seconds due to needed manual curation, while for FAIR metadata this was 13 seconds.\u003c/p\u003e \u003cp\u003eFor the second query, we queried one of the identified transcriptomics datasets for Inclusion Body Myositis in a FAIR Data Point Index for its data format. Here we found the \u0026ldquo;expression data\u0026rdquo; (EDAM:2603) EDAM term. Using this EDAM term, we then queried FAIR Data Points for tools that have this EDAM term as their input. As a result, out of 4 tools that were indexed in the FAIR Data Point Index, we found the DESEQ2 tool, which performs differential gene expression testing, and is applicable on this dataset.\u003c/p\u003e \u003cp\u003eIn the third query, we first queried the ORDO ontology [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e] for diseases that are a subclass of \u0026ldquo;myositis\u0026rdquo;. Using this list of diseases, we subsequently queried FAIR Data Points for datasets that are described with at least one of these disease terms. We found 4 datasets for Inclusion Body Myositis out of the 21 total dataset that were connected to this FAIR Data Point index.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this work, we identified important gaps in current metadata standards for machine actionable bioinformatics analysis. Our main findings are a lack of machine readability, and a limited alignment of metadata for tools and datasets. As a solution, we propose minimal metadata for tools, and an alignment of metadata for tools and datasets.\u003c/p\u003e \u003cp\u003eWe identified a minimal set of metadata properties that are minimally necessary for machine actionable analysis. By drawing from existing repositories, we include properties that were only included in one of some of these repositories. This includes memory requirements and software maturity, which are important for machine actionability. Additionally, we aligned the tool metadata to tool metadata which provided several useful alignments. For example, describing inputs and outputs of tools, and format of data, allows computer agents to know which tool works for which data. Another example is storage requirements, where the storage requirement of a tool and dataset has to be summed to get the total required storage. This alignment can improve the interoperability between tools and datasets for machine actionable analysis.\u003c/p\u003e \u003cp\u003eQueries on our metadata schemas showed that the machine actionability aspect of FAIR has several important benefits in research. For example, we were able to find datasets with more precision by incorporating ontology terms into the search instead of plain text. This highlights the importance of the use of ontologies to avoid ambiguity for terms and acronyms that have multiple meanings. Another query showed that ontologies also increased interoperability between datasets and tools, by allowing them to be described with identical or mapped ontology terms, ensuring the machine understands that the resources annotated with these terms have the same meaning without being affected by differences in spelling or ambiguity of natural language. Finally, we found that the use of FAIR and ontologies also enabled machines to use knowledge embedded in ontologies, for example, to learn which diseases belong to a disease group. This makes queries more powerful, and thus improves machine actionability.\u003c/p\u003e \u003cp\u003eThere are several limitations to note in this work. The scope of our work is focused on biomedical sciences and bioinformatics and although our work could be applicable to other domains, existing data sharing models and standards differ between fields. Further, the comparison between analysis with and without FAIRification has limited quantification, which limits the strength of its conclusions. With a higher adoption of FAIR comparable to platforms that are not FAIRified, a more comprehensive comparison could have been made.\u003c/p\u003e \u003cp\u003eIn future work, the minimal metadata that we identified can serve as a reference for data and software sharing platforms to improve the machine actionability of their metadata and improve the alignment between tools and datasets. Our minimal metadata can also be implemented in FAIR ecosystems such as the SPHN and EJPRD. The minimal metadata would especially benefit federated infrastructure such as the FAIR Data Train [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], which requires a high degree of machine actionable communication between FAIR Data Stations (containing datasets) and FAIR Data Trains (containing tools). Our minimal metadata schema can serve as a starting point in implementing this communication, for example, by describing the memory requirements of a train, so that sufficient memory is allocated by a FAIR Data Station when running a tool. The minimal metadata that we identified can also be expanded further in order to achieve an increasingly higher degree of machine actionability for tools and datasets. Efforts could for example focus on improving the interaction between tools and datasets by enabling more granular file format descriptions of datasets and of tool inputs and outputs thereby connecting the two in a way that is precise enough for computer agents to work increasingly more independently.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn this work, we propose two metadata schemas to address current limitations of metadata. One is focused on minimal metadata for tools, and the second one is focused on alignment between metadata of data and tools. Through queries, we showed that our metadata schema performs well in terms of precision, aligning data with tools, and allowing for more complex and powerful queries. These schemas can be used in platforms for bioinformatics tools and in FAIR infrastructures, but can also apply beyond the Rare Disease domain, and serve other areas.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics approval and consent to participate\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003eConsent for publication\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003eAvailability of data and materials\u003c/h2\u003e\n\u003cp\u003eThe queries that support our findings are available on GitHub, https://github.com/LUMC-BioSemantics/Tool-and-data-queries.\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis work was funded by ELIXIR, the research infrastructure for life science data. This initiative has received funding from the European Union\u0026rsquo;s Horizon 2020 research and innovation programme under grant agreement N\u0026deg;825575 (The European Joint Programme Rare Diseases, EJP RD). This project is supported by the Innovative Health Initiative Joint Undertaking (IHI JU) under grant agreement No 101132943. The JU receives support from the European Union\u0026rsquo;s Horizon Europe research and innovation programme, COCIR, EFPIA, EuropaBio, MedTech Europe, Vaccines Europe (and the FSHD Society and TREAT-NMD Services Ltd as contributing partners participating in the project). The project will also receive funding from UK Research and Innovation (UKRI) under the UK government\u0026rsquo;s Horizon Europe funding guarantee [grant numbers 10105921, 10103989, and 10083579]. NQR is partially funded as contributing partners of SYNTHIA and ERDERA projects. SYNTHIA is supported by the Innovative Health Initiative Joint Undertaking (IHI JU) under grant agreement No 101172872. The JU receives support from the EU Horizon Europe programme and COCIR, EFPIA, Europa Bio, MedTech Europe, and Vaccines Europe and DNV. ERDERA has received funding from the European Union\u0026rsquo;s Horizon Europe research and innovation programme under grant agreement N\u0026deg;101156595.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eAuthors\u0026apos; contributions\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eDW:\u003c/strong\u003e Methodology, Software, Investigation, Writing - Original draft. \u003cstrong\u003eNQR:\u003c/strong\u003e Funding acquisition, Conceptualization, Methodology, Writing - Review \u0026amp; Editing. \u003cstrong\u003eVB:\u003c/strong\u003e Investigation, Writing - Review \u0026amp; Editing. \u003cstrong\u003eEV:\u0026nbsp;\u003c/strong\u003eSoftware,Investigation, Writing - Review \u0026amp; Editing.\u003cstrong\u003e\u0026nbsp;NB:\u003c/strong\u003e Investigation, Writing - Review \u0026amp; Editing. \u003cstrong\u003eAJ:\u0026nbsp;\u003c/strong\u003eFunding acquisition, Writing - Review \u0026amp; Editing.\u003cstrong\u003e\u0026nbsp;PACtH:\u0026nbsp;\u003c/strong\u003eWriting - Review \u0026amp; Editing. \u003cstrong\u003eCC:\u0026nbsp;\u003c/strong\u003eWriting - Review \u0026amp; Editing.\u003cstrong\u003e\u0026nbsp;MR:\u0026nbsp;\u003c/strong\u003eWriting - Review \u0026amp; Editing. \u003cstrong\u003eEM:\u0026nbsp;\u003c/strong\u003eFunding acquisition, Writing - Review \u0026amp; Editing, Methodology, Supervision.\u003c/p\u003e\n\u003ch2\u003eAcknowledgements\u003c/h2\u003e\n\u003cp\u003eWe thank the members of the ELIXIR Rare Disease Implementation Study for the collaboration and fruitful discussions. We would also like to thank all the members of the ELIXIR Rare Disease Community and EJP RD for contributing to our survey and for their valuable input that was foundational for this study.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. Comment: The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:1\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan der Horst E, Kaliyaperumal R, Tatum Z, Thompson M, Schultes E, Mina E et al. Finding novel associations across domains using linked data: a case study on genetic variants disrupting transcription start sites.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao X, Niehues A, de Visser C, Huang J, Doornbos C, Ederveen THA et al. FAIR Data Cube, a FAIR data infrastructure for integrated multi-omics data analysis. 2023;:2023.04.23.23289000.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003edos Santos Vieira B, Bernab\u0026eacute; CH, Zhang S, Abaza H, Benis N, C\u0026aacute;mara A, et al. Towards FAIRification of sensitive and fragmented rare disease patient data: challenges and solutions in European reference network registries. Orphanet J Rare Dis. 2022;17:436.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHerre B, Rod\u0026eacute;s-Guirao L, Mathieu E, Ritchie H, Giattino C, Hasell J, et al. Best practices for government agencies to publish data: lessons from COVID-19. Lancet Public Health. 2024;9:e407\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTour\u0026eacute; V, Krauss P, Gnodtke K, Buchhorn J, Unni D, Horki P, et al. FAIRification of health-related data using semantic web technologies in the Swiss Personalized Health Network. Sci Data. 2023;10:127.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaliyaperumal R, Wilkinson MD, Moreno PA, Benis N, Cornet R, dos Santos Vieira B, et al. Semantic modelling of common data elements for rare disease registries, and a prototype workflow for their deployment over registry data. J Biomed Semant. 2022;13:9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eejp-rd-vp/resource-metadata-schema. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eda Silva Santos LB, Burger K, Kaliyaperumal R. FAIR Data Train specifications. 2023. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://specs.fairdatatrain.org/\u003c/span\u003e\u003cspan address=\"https://specs.fairdatatrain.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 19 Feb 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Z, Lachmann A, Ma\u0026rsquo;ayan A. Mining data and metadata from the gene expression omnibus. Biophys Rev. 2019;11:103\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim N, Tesar S, Belmadani M, Poirier-Morency G, Mancarci BO, Sicherman J, et al. Curation of over 10 000 transcriptomic studies to enable data reuse. Database. 2021;2021:baab006.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarbon S, Champieux R, McMurry JA, Winfree L, Wyatt LR, Haendel MA. An analysis and metric of reusable data licensing practices for biomedical resources. PLoS ONE. 2019;14:e0213090.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarker M, Chue Hong NP, Katz DS, Lamprecht A-L, Martinez-Ortiz C, Psomopoulos F, et al. Introducing the FAIR Principles for research software. Sci Data. 2022;9:622.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChue Hong NP, Katz DS, Barker M, Lamprecht A-L, Martinez C, Psomopoulos FE et al. FAIR Principles for Research Software (FAIR4RS Principles). 2022. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.15497/RDA00068\u003c/span\u003e\u003cspan address=\"10.15497/RDA00068\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ede Visser C, Johansson LF, Kulkarni P, Mei H, Neerincx P, van der Velde KJ, et al. Ten quick tips for building FAIR workflows. PLOS Comput Biol. 2023;19:e1011369.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCrusoe MR, Abeln S, Iosup A, Amstutz P, Chilton J, Tijanić N, et al. Methods Included: Standardizing Computational Reuse and Portability with the Common Workflow Language. Commun ACM. 2022;65:54\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGustafsson OJR, Wilkinson SR, Bacall F, Pireddu L, Soiland-Reyes S, Leo S et al. WorkflowHub: a registry for computational workflows. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWijnbergen D, Johari M, Ozisik O, \u0026lsquo;t Hoen PAC, Ehrhart F, Baudot A, et al. Multi-omics analysis in inclusion body myositis identifies mir-16 responsible for HLA overexpression. Orphanet J Rare Dis. 2025;20:27.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWijnbergen D, Johari M, Udd B, Roos M, Mina E. Inclusion Body Myositis Active Subnetwork Identification Workflow. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIson J, Ienasescu H, Chmura P, Rydza E, M\u0026eacute;nager H, Kalaš M, et al. The bio.tools registry of software tools and data resources for the life sciences. Genome Biol. 2019;20:164.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eR Core Team. R: A Language and Environment for Statistical Computing.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGentleman RC, Carey VJ, Bates DM, Bolstad B, Dettling M, Dudoit S, et al. Bioconductor: open software development for computational biology and bioinformatics. Genome Biol. 2004;5:R80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRare Diseases Community | ELIXIR. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://elixir-europe.org/communities/rare-diseases\u003c/span\u003e\u003cspan address=\"https://elixir-europe.org/communities/rare-diseases\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 20 Feb 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEJP, RD \u0026ndash; European Joint Programme on Rare Diseases. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ejprarediseases.org/\u003c/span\u003e\u003cspan address=\"https://www.ejprarediseases.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 20 Feb 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRitchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43:e47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRenaux A, Papadimitriou S, Versbraegen N, Nachtegael C, Boutry S, Now\u0026eacute; A, et al. ORVAL: a novel platform for the prediction and exploration of disease-causing oligogenic variant combinations. Nucleic Acids Res. 2019;47:W93\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOzisik O, T\u0026eacute;r\u0026eacute;zol M, Baudot A. orsum: a Python package for filtering and comparing enrichment analyses using a simple principle. BMC Bioinformatics. 2022;23:293.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePasche E, Mottaz A, Caucheteur D, Gobeill J, Michel P-A, Ruch P. Variomes: a high recall search engine to support the curation of genomic variants. Bioinformatics. 2022;38:2595\u0026ndash;601.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaurie S, Piscia D, Matalonga L, Corv\u0026oacute; A, Fern\u0026aacute;ndez-Callejo M, Garcia-Linares C, et al. The RD-Connect Genome-Phenome Analysis Platform: Accelerating diagnosis, research, and gene discovery for rare diseases. Hum Mutat. 2022;43:717\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNovoa-del-Toro EM, Mezura-Montes E, Vignes M, T\u0026eacute;r\u0026eacute;zol M, Magdinier F, Tichit L, et al. A multi-objective genetic algorithm to find active modules in multiplex biological networks. PLOS Comput Biol. 2021;17:e1009263.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Iersel MP, Pico AR, Kelder T, Gao J, Ho I, Hanspers K, et al. The BridgeDb framework: standardized access to gene, protein and metabolite identifier mapping services. BMC Bioinformatics. 2010;11:5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:1\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKutmon M, van Iersel MP, Bohler A, Kelder T, Nunes N, Pico AR, et al. PathVisio 3: An Extendable Pathway Analysis Toolbox. PLOS Comput Biol. 2015;11:e1004085.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL\u0026ecirc; Cao K-A, Gonz\u0026aacute;lez I, D\u0026eacute;jean S. integrOmics: an R package to unravel relationships between two omics datasets. Bioinformatics. 2009;25:2855\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL\u0026ecirc; Cao K-A, Rohart F, Gonzalez I, Dejean S. mixOmics: Omics Data Integration Project. R package version 6.1.1. 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eT\u0026eacute;r\u0026eacute;zol M, Baudot A, Ozisik O, ODAMNet. A Python package to identify molecular relationships between chemicals and rare diseases using overlap, active module and random walk approaches. SoftwareX. 2024;26:101701.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKolde R, Laur S, Adler P, Vilo J. Robust rank aggregation for gene list integration and meta-analysis. Bioinformatics. 2012;28:573\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ecodemeta/crosswalk.csv at master \u0026middot; codemeta/codemeta. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/codemeta/codemeta/blob/master/crosswalk.csv\u003c/span\u003e\u003cspan address=\"https://github.com/codemeta/codemeta/blob/master/crosswalk.csv\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 21 Jan 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThe CodeMeta Project. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://codemeta.github.io/\u003c/span\u003e\u003cspan address=\"https://codemeta.github.io/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 21 Jan 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eda Silva Santos LOB, Burger K, Kaliyaperumal R, Wilkinson MD. FAIR Data Point: A FAIR-Oriented Approach for Metadata Publication. Data Intell. 2023;5:163\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchriml LM, Munro JB, Schor M, Olley D, McCracken C, Felix V, et al. The Human Disease Ontology 2022 update. Nucleic Acids Res. 2022;50:D1255\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlack M, Lamothe L, Eldakroury H, Kierkegaard M, Priya A, Machinda A et al. EDAM: the bioscientific data analysis ontology (update 2021). F1000Research. 2022;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlbertoni R, Browning D, Cox S, Gonzalez Beltran A, Perego A, Winstanley P et al. Data Catalog Vocabulary (DCAT) - Version 2. 2020. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.w3.org/TR/vocab-dcat-2/\u003c/span\u003e\u003cspan address=\"https://www.w3.org/TR/vocab-dcat-2/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 10 Feb 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVasant D, Chanas L, Malone J, Hanauer M, Olry A, Jupp S et al. ORDO: An Ontology Connecting Rare Disease, Epidemiology and Genetic Data. 2014.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"journal-of-biomedical-semantics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jbsm","sideBox":"Learn more about [Journal of Biomedical Semantics](http://jbiomedsem.biomedcentral.com/)","snPcode":"13326","submissionUrl":"https://submission.nature.com/new-submission/13326/3","title":"Journal of Biomedical Semantics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"FAIR, metadata, machine-actionability, bioinformatics","lastPublishedDoi":"10.21203/rs.3.rs-6396831/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6396831/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eOne pillar of FAIR principles adoption is Reusability by machines to enable for example more efficient data analytics in fields such as Bioinformatics. However, it is not clear to what extent current metadata exposed by datasets and tools in common repositories enable this. In practice, metadata often lacks in machine actionability due to incomplete standardised metadata and lack of ontological descriptions.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eIn this work, we identified minimal metadata that is needed to improve the machine actionability of bioinformatics tools and proposed a schema to address current limitations. The schema consists of metadata properties for the identification, selection, validation, and execution of tools. We also aligned this metadata to the metadata of datasets, in order to improve their integration for analytics by machines.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThe identified minimal metadata improves the machine actionability of tools and data, and can be incorporated into platforms for tool and data sharing, and in FAIR infrastructures.\u003c/p\u003e","manuscriptTitle":"FAIR in practice: minimum metadata schema for bioinformatics analytics by machines","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-07 20:36:36","doi":"10.21203/rs.3.rs-6396831/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-08-31T20:04:33+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-02T20:04:08+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-05-28T22:27:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"159754748488426417693563706939345257693","date":"2025-05-25T20:01:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"155650745073686013090162255413124149029","date":"2025-05-23T14:33:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"180564414346001509792163342975503030224","date":"2025-05-23T09:32:39+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-05-02T08:21:30+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-04-11T13:17:34+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-04-11T13:14:50+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Biomedical Semantics","date":"2025-04-07T18:33:40+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"journal-of-biomedical-semantics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jbsm","sideBox":"Learn more about [Journal of Biomedical Semantics](http://jbiomedsem.biomedcentral.com/)","snPcode":"13326","submissionUrl":"https://submission.nature.com/new-submission/13326/3","title":"Journal of Biomedical Semantics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c806de82-4146-407b-bde1-a9148060027d","owner":[],"postedDate":"May 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-29T03:54:09+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-07 20:36:36","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6396831","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6396831","identity":"rs-6396831","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.