Clinical text classification research trends: Systematic literature review and open issues.
Saved in:
| Title: | Clinical text classification research trends: Systematic literature review and open issues. |
|---|---|
| Authors: | Mujtaba, Ghulam1,2 mujtaba@siswa.um.edu.my, Shuib, Liyana1 liyanashuib@um.edu.my, Idris, Norisma3 norisma@um.edu.my, Hoo, Wai Lam1 wlhoo@um.edu.my, Raj, Ram Gopal3 ramdr@um.edu.my, Khowaja, Kamran4, Shaikh, Khairunisa5 khairunisashaikh@siswa.um.edu.my, Nweke, Henry Friday1,6 henrynweke@siswa.um.edu.my |
| Source: | Expert Systems with Applications. Feb2019, Vol. 116, p494-520. 27p. |
| Subjects: | Clinical medicine research, Literature classification, Machine learning, Rule-based programming, Key performance indicators (Management) |
| Abstract: | Highlights • To review free-text clinical text classification approaches from six aspects. • In selected studies, mostly content-based and concept-based features were used. • The datasets used in selected studies were categorized into four distinct types. • Selected studies used either supervised machine learning or rule-based approaches. • Ten open research challenges are presented in clinical text classification domain. Abstract The pervasive use of electronic health databases has increased the accessibility of free-text clinical reports for supplementary use. Several text classification approaches, such as supervised machine learning (SML) or rule-based approaches, have been utilized to obtain beneficial information from free-text clinical reports. In recent years, many researchers have worked in the clinical text classification field and published their results in academic journals. However, to the best of our knowledge, no comprehensive systematic literature review (SLR) has recapitulated the existing primary studies on clinical text classification in the last five years. Thus, the current study aims to present SLR of academic articles on clinical text classification published from January 2013 to January 2018. Accordingly, we intend to maximize the procedural decision analysis in six aspects, namely, types of clinical reports, data sets and their characteristics, pre-processing and sampling techniques, feature engineering, machine learning algorithms, and performance metrics. To achieve our objective, 72 primary studies from 8 bibliographic databases were systematically selected and rigorously reviewed from the perspective of the six aspects. This review identified nine types of clinical reports, four types of data sets (i.e., homogeneous–homogenous, homogenous–heterogeneous, heterogeneous–homogenous, and heterogeneous–heterogeneous), two sampling techniques (i.e., over-sampling and under-sampling), and nine pre-processing techniques. Moreover, this review determined bag of words, bag of phrases, and bag of concepts features when represented by either term frequency or term frequency with inverse document frequency, thereby showing improved classification results. SML-based or rule-based approaches were generally employed to classify the clinical reports. To measure the performance of these classification approaches, we used precision, recall, F-measure, accuracy, AUC, and specificity in binary class problems. In multi-class problems, we primarily used micro or macro-averaging precision, recall, or F-measure. Lastly, open research issues and challenges are presented for future scholars who are interested in clinical text classification. This SLR will definitely be a beneficial resource for researchers engaged in clinical text classification. [ABSTRACT FROM AUTHOR] |
| Copyright of Expert Systems with Applications is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| Abstract: | Highlights • To review free-text clinical text classification approaches from six aspects. • In selected studies, mostly content-based and concept-based features were used. • The datasets used in selected studies were categorized into four distinct types. • Selected studies used either supervised machine learning or rule-based approaches. • Ten open research challenges are presented in clinical text classification domain. Abstract The pervasive use of electronic health databases has increased the accessibility of free-text clinical reports for supplementary use. Several text classification approaches, such as supervised machine learning (SML) or rule-based approaches, have been utilized to obtain beneficial information from free-text clinical reports. In recent years, many researchers have worked in the clinical text classification field and published their results in academic journals. However, to the best of our knowledge, no comprehensive systematic literature review (SLR) has recapitulated the existing primary studies on clinical text classification in the last five years. Thus, the current study aims to present SLR of academic articles on clinical text classification published from January 2013 to January 2018. Accordingly, we intend to maximize the procedural decision analysis in six aspects, namely, types of clinical reports, data sets and their characteristics, pre-processing and sampling techniques, feature engineering, machine learning algorithms, and performance metrics. To achieve our objective, 72 primary studies from 8 bibliographic databases were systematically selected and rigorously reviewed from the perspective of the six aspects. This review identified nine types of clinical reports, four types of data sets (i.e., homogeneous–homogenous, homogenous–heterogeneous, heterogeneous–homogenous, and heterogeneous–heterogeneous), two sampling techniques (i.e., over-sampling and under-sampling), and nine pre-processing techniques. Moreover, this review determined bag of words, bag of phrases, and bag of concepts features when represented by either term frequency or term frequency with inverse document frequency, thereby showing improved classification results. SML-based or rule-based approaches were generally employed to classify the clinical reports. To measure the performance of these classification approaches, we used precision, recall, F-measure, accuracy, AUC, and specificity in binary class problems. In multi-class problems, we primarily used micro or macro-averaging precision, recall, or F-measure. Lastly, open research issues and challenges are presented for future scholars who are interested in clinical text classification. This SLR will definitely be a beneficial resource for researchers engaged in clinical text classification. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 09574174 |
| DOI: | 10.1016/j.eswa.2018.09.034 |