dgiLIT: A Method for Prioritization and AI Curation of Drug-Gene Interactions

preprint OA: closed
Full text JSON View at publisher

Abstract

ABSTRACT IMPORTANCE The Drug-Gene Interaction Database (DGIdb) has a long history of driving hypothesis generation for biomedical research through the careful curation of drug-gene interaction data from primary and secondary sources with supporting literature. Recent advances in large-language model (LLM) and artificial intelligence (AI) technologies have enabled new paradigms for knowledge extraction and biocuration. The accelerating growth of biomedical literature presents a significant challenge for maintaining up-to-date interaction data. With more than 38 million citations indexed in PubMed alone, new strategies must evolve to identify and incorporate new interaction data into DGIdb. OBJECTIVE Identify new cost-effective AI curation strategies for incorporating new drug-gene interactions into DGIdb. METHODS We present a methodology that leverages deterministic natural language processing techniques, existing harmonization frameworks, and AI-assisted curation to systematically narrow the literature space and identify new drug-gene interactions from published studies for inclusion in DGIdb. RESULTS We demonstrate the use of lemmatization to prioritize a set of 100 abstracts containing high amounts of interaction words for downstream AI curation. From our set of abstracts, we were then able to identify 137 drug-gene interactions via an AI curation task, with 121 (88.3%) of these interactions being completely novel to DGIdb. A human expert evaluator reviewed this interaction set and was able to validate 134 of 137 (97.8%) interactions as being valid based on the text provided. CONCLUSION Taken together, our results highlight a promising, cost-effective method of ingesting new interactions into DGIdb.
Full text 1,808 characters · extracted from oa-doi-fallback · 5 sections · click to expand

Abstract

IMPORTANCE The Drug-Gene Interaction Database (DGIdb) has a long history of driving hypothesis generation for biomedical research through the careful curation of drug-gene interaction data from primary and secondary sources with supporting literature. Recent advances in large-language model (LLM) and artificial intelligence (AI) technologies have enabled new paradigms for knowledge extraction and biocuration. The accelerating growth of biomedical literature presents a significant challenge for maintaining up-to-date interaction data. With more than 38 million citations indexed in PubMed alone, new strategies must evolve to identify and incorporate new interaction data into DGIdb.

Objective

Identify new cost-effective AI curation strategies for incorporating new drug-gene interactions into DGIdb.

Methods

We present a methodology that leverages deterministic natural language processing techniques, existing harmonization frameworks, and AI-assisted curation to systematically narrow the literature space and identify new drug-gene interactions from published studies for inclusion in DGIdb.

Results

We demonstrate the use of lemmatization to prioritize a set of 100 abstracts containing high amounts of interaction words for downstream AI curation. From our set of abstracts, we were then able to identify 137 drug-gene interactions via an AI curation task, with 121 (88.3%) of these interactions being completely novel to DGIdb. A human expert evaluator reviewed this interaction set and was able to validate 134 of 137 (97.8%) interactions as being valid based on the text provided.

Conclusion

Taken together, our results highlight a promising, cost-effective method of ingesting new interactions into DGIdb. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00