Research on Automatic Labelling of Imbalanced Texts of Customer Complaints Based on Text Enhancement and Layer-By-Layer Semantic Matching
preprint
OA: closed
Abstract
[Purpose/meaning] Due to its potential impact on business efficiency, automated customer complaint labeling and classification are of great importance for management decisions and business-level applications. The majority of the current research on automated labeling uses large and well-balanced datasets. However, customer complaints' labels are hierarchical in structure, with many labels at the lowest hierarchy level. Relying on lower-level labels leads to small and imbalanced samples, thus rendering the current automatic labeling practices not applicable to customer complaints. [Methodology/process] This article proposes an automatic labeling model incorporating the BERT and Word2Vec methods. The model is validated on electric utility customer complaints data. Within the model, the BERT method serves to obtain shallow-level text tags. Further, text enhancement is used to mitigate the problem of uneven samples that emerges when the number of labels is large. Finally, the Word2Vec model is utilized for the deep-level text analysis. [Findings/conclusions] The experiments demonstrate the proposed model's efficiency in automating customer complaint labeling. Consequently, the proposed model supports the enterprises in improving their service quality while simultaneously reducing labor costs.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00