Graph Based Imbalanced Multi-Text Classification: A Study on Low Resource Language

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

A pre-labeled dataset is required for any machine learning or deep learning tasks in Natural Language Processing (NLP). Certain languages lack adequate resources, hence impeding research efforts, even for seemingly straightforward natural language processing (NLP) problems. This work introduces a novel methodology for enhancing data in the context of Natural Language Processing (NLP) tasks, specifically focusing on Part-of-Speech (POS) tagging and Named Entity Recognition (NER). The representation of a limited quantity of data takes the form of a directed graph, which is utilized in conjunction with the Random Walk method to produce sentence instances for a given word. A Long Short Term Memory (LSTM) neural network is employed to create a hybrid Part-of-Speech (POS) and Named Entity Recognition (NER) model. This model is designed to annotate words within sentence instances. The empirical findings demonstrate that the model applied to the Manipuri dataset achieves exceptional performance, as evidenced by an F-score metric of 1 for both POS and NER classification.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0