Benchmarking of Machine Learning Models to Assist the Prognosis of Tuberculosis

Maicon Herverton Lino Ferreira da Silva Barros; Geovanne Oliveira Alves; Lubnnia Morais Florêncio Souza; Élisson da Silva Rocha; João Fausto Lorenzato de Oliveira; Theo Lynn; Vanderson Sampaio; Patricia Takako Endo

doi:10.20944/preprints202103.0284.v2

Submitted:

09 April 2021

Posted:

12 April 2021

You are already at the latest version

Abstract

Tuberculosis (TB) is an airborne infectious disease caused by organisms in the Mycobacterium tuberculosis (Mtb) complex. In many low and middle-income countries, TB remains a major cause of morbidity and mortality. Once a patient has been diagnosed with TB, it is critical that healthcare workers make the most appropriate treatment decision given the individual conditions of the patient and the likely course of the disease based on medical experience. Depending on the prognosis, delayed or inappropriate treatment can result in unsatisfactory results including the exacerbation of clinical symptoms, poor quality of life, and increased risk of death. This work benchmarks machine learning models to aid TB prognosis using a Brazilian health database of confirmed cases and deaths related to TB in the State of Amazonas. The goal is to predict the probability of death by TB thus aiding the prognosis of TB and associated treatment decision making process. In its original form, the data set comprised 36,228 records and 130 fields but suffered from missing, incomplete, or incorrect data. Following data cleaning and preprocessing, a revised data set was generated comprising 24,015 records and 38 fields, including 22,876 reported cured TB patients and 1,139 deaths by TB. To explore how the data imbalance impacts model performance, two controlled experiments were designed using (1) imbalanced and (2) balanced data sets. The best result is achieved by the Gradient Boosting (GB) model using the balanced data set to predict TB-mortality, and the ensemble model composed by the Random Forest (RF), GB and Multi-layer Perceptron (MLP) models is the best model to predict the cure class.

Keywords:

machine learning

;

benchmarking

;

tuberculosis

;

prognosis

Subject:

Computer Science and Mathematics - Algebra and Number Theory

Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.

Benchmarking of Machine Learning Models to Assist the Prognosis of Tuberculosis

Abstract

Keywords:

Subject:

MDPI Initiatives

Important Links

Subscribe