Discriminating between Empirical Studies and Nonempirical Works Using Automated Text Classification

Saved in:
Bibliographic Details
Title: Discriminating between Empirical Studies and Nonempirical Works Using Automated Text Classification
Language: English
Authors: Langlois, Alexis (ORCID 0000-0002-9280-2320), Nie, Jian-Yun, Thomas, James, Hong, Quan Nha, Pluye, Pierre
Source: Research Synthesis Methods. Dec 2018 9(4):587-601.
Availability: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
Peer Reviewed: Y
Page Count: 15
Publication Date: 2018
Document Type: Journal Articles
Reports - Research
Descriptors: Mixed Methods Research, Databases, Information Retrieval, Search Strategies, Documentation, Accuracy, Metadata, Comparative Analysis, Research Reports, Classification, Vocabulary, Reference Materials, Medical Research
DOI: 10.1002/jrsm.1317
ISSN: 1759-2879
Abstract: Objective: Identify the most performant automated text classification method (eg, algorithm) for differentiating empirical studies from nonempirical works in order to facilitate systematic mixed studies reviews. Methods: The algorithms were trained and validated with 8050 database records, which had previously been manually categorized as empirical or nonempirical. A Boolean mixed filter developed for filtering MEDLINE records (title, abstract, keywords, and full texts) was used as a baseline. The set of features (eg, characteristics from the data) included observable terms and concepts extracted from a metathesaurus. The efficiency of the approaches was measured using sensitivity, precision, specificity, and accuracy. Results: The decision trees algorithm demonstrated the highest performance, surpassing the accuracy of the Boolean mixed filter by 30%. The use of full texts did not result in significant gains compared with title, abstract, keywords, and records. Results also showed that mixing concepts with observable terms can improve the classification. Significance: Screening of records, identified in bibliographic databases, for relevant studies to include in systematic reviews can be accelerated with automated text classification.
Abstractor: As Provided
Entry Date: 2020
Accession Number: EJ1255688
Database: ERIC
Full text is not displayed to guests.
Description
Abstract:Objective: Identify the most performant automated text classification method (eg, algorithm) for differentiating empirical studies from nonempirical works in order to facilitate systematic mixed studies reviews. Methods: The algorithms were trained and validated with 8050 database records, which had previously been manually categorized as empirical or nonempirical. A Boolean mixed filter developed for filtering MEDLINE records (title, abstract, keywords, and full texts) was used as a baseline. The set of features (eg, characteristics from the data) included observable terms and concepts extracted from a metathesaurus. The efficiency of the approaches was measured using sensitivity, precision, specificity, and accuracy. Results: The decision trees algorithm demonstrated the highest performance, surpassing the accuracy of the Boolean mixed filter by 30%. The use of full texts did not result in significant gains compared with title, abstract, keywords, and records. Results also showed that mixing concepts with observable terms can improve the classification. Significance: Screening of records, identified in bibliographic databases, for relevant studies to include in systematic reviews can be accelerated with automated text classification.
ISSN:1759-2879
DOI:10.1002/jrsm.1317