Enhancing Sentiment Analysis with Term Sentiment Entropy: Capturing Nuanced Sentiment in Text Classification

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

Sentiment analysis benefits from representations that highlight polarity-bearing terms while suppressing sentiment-ambivalent ones. This paper introduces Term Sentiment Entropy (TSE), a supervised, information-theoretic global factor for sparse text representation. TSE quantifies how selectively a term associates with sentiment labels in the training fold, and it is composed with TF-IDF to up-weight terms that are distributionally concentrated within a class and down-weight those that are diffuse across classes. We evaluate the approach on four public datasets spanning product reviews, social media, and long-form movie reviews under a fixed protocol with Naïve Bayes, Random Forest, and a linear Support Vector Classifier. Results reported as Accuracy, Macro-Precision, Macro-Recall, and Macro-F1 show that TF-IDF plus TSE often matches or improves performance on short and noisy texts such as Amazon cell-phone reviews and two Twitter corpora, while achieving near-ceiling parity with strong baselines on IMDb. The method is lightweight, reproducible, and compatible with conventional preprocessing and feature-selection pipelines because it requires only label statistics from the training data and no external lexicons. We also discuss limitations related to label quality and class imbalance, and we outline imbalance-aware and learned variants of TSE as natural extensions.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-05T02:00:03.366016+00:00
License: CC-BY-4.0