Bug Report Classification with Ensemble Learning for Closed-Source Software

preprint OA: closed
View at publisher

Abstract

This paper introduces a set of datasets grouped under the name Turkish Software Bug Reports (TSBR), which comprises commercial software bug reports from a closed-source project. We investigate and report the statistical properties and classification difficulty of the TSBR datasets. We employ various methods from the text classification literature to apply several classification tasks related to software development on the TSBR datasets. The methods we employ include traditional machine learning (ML) methods such as k-nearest neighbors (KNN) and random forest (RF); sequential deep learning (DL) models such as gated recurrent unit (GRU) and convolutional neural network (CNN); transformer-based language models; and ensembles of the employed models. Our work is among the first efforts in automated bug report classification literature that uses ensembles of DL models.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00