Biotext: Exploiting Biological-Text Format for Text Mining

preprint OA: closed
📄 Open PDF View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This paper presents BIOTEXT, a package of strategies for converting natural language text into biological-like data, enabling text mining using bioinformatics tools.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

ABSTRACT The large amount of existing textual data justifies the development of new text mining tools. Bioinformatics tools can be brought to Text Mining, increasing the arsenal of resources. Here, we present BIOTEXT, a package of strategies for converting natural language text into biological-like information data, providing a general protocol with standardized functions, allowing to share, encode and decode textual data for amino acid and DNA. The package was used to encode the arbitrary information present in the headings of the biological sequences found in a BLAST survey. The protocol implemented in this study consists of 12 steps, which can be easily executed and/ or changed by the user, depending on the study area. BIOTEXT empowers users to perform text mining using bioinformatics tools. BIOTEXT is freely available at https://pypi.org/project/BIOTEXT/ (Python package) and https://sourceforge.net/projects/BIOTEXTtools/files/AMINOcode_GUI/ (Standalone tool).

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-06-04T02:00:05.705006+00:00