Automated Identification of Contextually Relevant Biomedical Entities with Grounded LLMs

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

This study investigates the effectiveness of different large language models (LLMs) for automated biomedical entity annotation in research articles with a focus on contextualized and grounded results. A 4-step generative workflow iteratively generates and refines entity candidates by considering a metadata schema for context and agentic tool use for validation with the PubTator 3 data base. The precision of this flow was assessed with a random effects meta-analysis after face-to-face interviews with authors of six papers from the Collaborative Research Center (CRC) 1453 “NephGen”. With an overall precision of 91.3%, the selected models provide qualitatively valuable annotations, with models GPT 4.1, GPT-4o Mini, and Gemini 2.0 Flash showing the highest precision. While GPT 4.1 and Gemini 2.0 Flash excelled in the total number of correct annotations, GPT-4o Mini and Gemini 2.0 Flash were fastest and most cost-effective. Large variations in annotation count and the conflation of publication and dataset-specific annotations highlight that human review (“human-in-the-loop”) is still important. The results further highlight the trade-offs between precision, total number of correct annotations, cost, and speed. While quality is paramount in collaborative research settings, cost-effectiveness could be more critical in public implementations.
Full text 3,523 characters · extracted from oa-doi-fallback · click to expand
Abstract This study investigates the effectiveness of different large language models (LLMs) for automated biomedical entity annotation in research articles with a focus on contextualized and grounded results. A 4-step generative workflow iteratively generates and refines entity candidates by considering a metadata schema for context and agentic tool use for validation with the PubTator 3 data base. The precision of this flow was assessed with a random effects meta-analysis after face-to-face interviews with authors of six papers from the Collaborative Research Center (CRC) 1453 “NephGen”. With an overall precision of 91.3%, the selected models provide qualitatively valuable annotations, with models GPT 4.1, GPT-4o Mini, and Gemini 2.0 Flash showing the highest precision. While GPT 4.1 and Gemini 2.0 Flash excelled in the total number of correct annotations, GPT-4o Mini and Gemini 2.0 Flash were fastest and most cost-effective. Large variations in annotation count and the conflation of publication and dataset-specific annotations highlight that human review (“human-in-the-loop”) is still important. The results further highlight the trade-offs between precision, total number of correct annotations, cost, and speed. While quality is paramount in collaborative research settings, cost-effectiveness could be more critical in public implementations. Competing Interest Statement The authors have declared no competing interest. Funding Statement This work was supported by the DFG (German Research Foundation) through the following projects: Project-ID 441891347 SFB 1479 (CG, HB), Project-ID 499552394 SFB 1597 (HB), Project-ID 491676693 TRR 359 (GB, HB), Project-ID 256073931 SFB 1160 (FE, GB, HB), Project-ID 514483642 TRR 384 (FE, GB, CG, HB), Project-ID 259373024 TRR 167 (HB), and Project-ID 431984000 SFB 1453 (HB) and EXC 2189 (HB). Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Footnotes This work was supported by the DFG (German Research Foundation) through the following projects: Project-ID 441891347 – SFB 1479 (CG, HB), Project-ID 499552394 – SFB 1597 (HB), Project-ID 491676693 – TRR 359 (GB, HB), Project-ID 256073931 – SFB 1160 (FE, GB, HB), Project-ID 514483642 – TRR 384 (FE, GB, CG, HB), Project-ID 259373024 – TRR 167 (HB), and Project-ID 431984000 – SFB 1453 (HB) and EXC 2189 (HB). Data Availability All data produced in the present work are contained in the manuscript

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00