Milestone MS32 The design and prototype of a workflow integrating Wikidata into validation and linking

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This work introduces a workflow integrating Wikidata to automate and improve the linking of collector name strings to persistent identifiers, enhancing efficiency and data interoperability.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This paper describes the design and development of a prototype workflow that integrates Wikidata into validation and linking processes, aiming to improve how structured data are checked and connected. The work is presented as a preprint focused on the software/workflow architecture rather than on biomedical experimentation, using a high-level description of system components and how Wikidata is incorporated into validation/linking steps. A key caveat is that the contribution is a prototype workflow, with emphasis on design and implementation details rather than providing validated performance results from large-scale deployments. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

In this task, the aim is to develop a workflow that should facilitate the linking process of collector name strings to PIDs for those collectors. Such a workflow should help scale up the number of links being made, make the process more efficient and should take advantage as much as possible of existing work and infrastructures, so as not to reinvent the wheel. As such, the work can be roughly split into a few subtasks:- Make existing linking workflows more easily implementable in other contexts and by other infrastructures. This includes finding ways for such workflows to produce links that can easily be published, i.e. in a standardised format compatible with existing infrastructure. The suitability of different infrastructures for making established links available should also be assessed.- Establish, document and improve the comprehensiveness, findability and interoperability of the content in PID-minting resources, in particular Wikidata as it can be edited openly.- Refine the decision making process of establishing links, by implementing and improving the methods that can be used to validate potential links. In this document, the focus lies on linking people. We will propose a workflow to 'roundtrip' links established through the Bionomia platform back to the collections holding the attributed specimens, as well as making them available for use by other BiCIKL infrastructures. We will also refine existing automated linking workflows and pilot the new functionalities on the (botanical) collections of the task partners. These refinements will be influenced by an assessment of the current state of Wikidata, investigated through shape expressions constructed from commonly used queries and from Wikidata records which have been linked in previous efforts such as the Botany Pilot, Bionomia and published specimen data to GBIF.
Full text 853 characters · extracted from oa-doi-fallback · click to expand
Preprint ARPHA Preprints https://doi.org/10.3897/arphapreprints.e114920 (31 Oct 2023) https://doi.org/10.3897/arphapreprints.e114920 (31 Oct 2023) Submitted to Research Ideas and Outcomes Other versions: - Preprint InfoPreprint Info - CiteCite - MetricsMetrics - CommentComment - RelatedRelated - CitedCited ARPHA Preprints doi: 10.3897/arphapreprints.e114920 First posted 31 Oct 2023 Authors Mathias Dillen - Corresponding author Meise Botanic Garden, Meise, Belgium Botanical Garden and Botanical Museum, Berlin, Germany Conflict of interest The authors have declared that no competing interests exist. This is an open access preprint distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00