Discriminating between Empirical Studies and Nonempirical Works Using Automated Text Classification
Saved in:
| Title: | Discriminating between Empirical Studies and Nonempirical Works Using Automated Text Classification |
|---|---|
| Language: | English |
| Authors: | Langlois, Alexis (ORCID |
| Source: | Research Synthesis Methods. Dec 2018 9(4):587-601. |
| Availability: | Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA |
| Peer Reviewed: | Y |
| Page Count: | 15 |
| Publication Date: | 2018 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Mixed Methods Research, Databases, Information Retrieval, Search Strategies, Documentation, Accuracy, Metadata, Comparative Analysis, Research Reports, Classification, Vocabulary, Reference Materials, Medical Research |
| DOI: | 10.1002/jrsm.1317 |
| ISSN: | 1759-2879 |
| Abstract: | Objective: Identify the most performant automated text classification method (eg, algorithm) for differentiating empirical studies from nonempirical works in order to facilitate systematic mixed studies reviews. Methods: The algorithms were trained and validated with 8050 database records, which had previously been manually categorized as empirical or nonempirical. A Boolean mixed filter developed for filtering MEDLINE records (title, abstract, keywords, and full texts) was used as a baseline. The set of features (eg, characteristics from the data) included observable terms and concepts extracted from a metathesaurus. The efficiency of the approaches was measured using sensitivity, precision, specificity, and accuracy. Results: The decision trees algorithm demonstrated the highest performance, surpassing the accuracy of the Boolean mixed filter by 30%. The use of full texts did not result in significant gains compared with title, abstract, keywords, and records. Results also showed that mixing concepts with observable terms can improve the classification. Significance: Screening of records, identified in bibliographic databases, for relevant studies to include in systematic reviews can be accelerated with automated text classification. |
| Abstractor: | As Provided |
| Entry Date: | 2020 |
| Accession Number: | EJ1255688 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHs68DeR7CIeEZpxzpuWQxUAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDAQvNzCzGWEDYxz3iQIBEICBm10Ma0qmzi3g4rYT5RdFEH0662XZxp9WuUK1Gr3yCLfKs4x5AKE1l5eLtjSOvj2ywtcaAxDxyPTy_Lq2WeBTobJ1Y2-B7_jLzJa__P7rCjy-spWVrB8BPVT716c7Po-pFcbKisT5WrSoSjhAyuc1zKrDHHqqp99VWeskEtgk4VmNaz7ACPL7CmjgA5RnQnyWOHFznFo7qT97tb8S Text: Availability: 1 Value: <anid>AN0133441849;[bdct]01dec.18;2018Dec10.05:59;v2.2.500</anid> <title id="AN0133441849-1">Discriminating between empirical studies and nonempirical works using automated text classification </title> <p>Objective: Identify the most performant automated text classification method (eg, algorithm) for differentiating empirical studies from nonempirical works in order to facilitate systematic mixed studies reviews. Methods: The algorithms were trained and validated with 8050 database records, which had previously been manually categorized as empirical or nonempirical. A Boolean mixed filter developed for filtering MEDLINE records (title, abstract, keywords, and full texts) was used as a baseline. The set of features (eg, characteristics from the data) included observable terms and concepts extracted from a metathesaurus. The efficiency of the approaches was measured using sensitivity, precision, specificity, and accuracy. Results: The decision trees algorithm demonstrated the highest performance, surpassing the accuracy of the Boolean mixed filter by 30%. The use of full texts did not result in significant gains compared with title, abstract, keywords, and records. Results also showed that mixing concepts with observable terms can improve the classification. Significance: Screening of records, identified in bibliographic databases, for relevant studies to include in systematic reviews can be accelerated with automated text classification.</p> <p>Keywords: automated text classification; decision tree; health care; research method; support vector machine; systematic review</p> <p>Researchers,  policymakers, and practitioners are increasingly interested in literature reviews can be used to justify, design, and interpret results of primary studies. Their growing popularity is mainly due to the increasing interest in evidence‐informed decision‐making and the need to have rigourous methods to identify and synthesize research. To synthesize research results, preference is given to systematic reviews since they use reproducible methods and are reported in a transparent manner.[<reflink idref="bib1" id="ref1">1</reflink>] Systematic reviews are considered epistemologically, methodologically, and practically relevant since they synthesize the best available evidence for a specific question. Moreover, they are increasing in popularity; the growth of the annual number of published systematic reviews largely exceeds that of other types of publications at least since 2010.[<reflink idref="bib2" id="ref2">2</reflink>]</p> <p>Over the past decade, mixed studies reviews have emerged as a new type of systematic review. They apply mixed methods approaches to critically analyse, synthesize, and integrate the findings of empirical studies.[<reflink idref="bib3" id="ref3">3</reflink>], [<reflink idref="bib4" id="ref4">4</reflink>], [<reflink idref="bib5" id="ref5">5</reflink>] Moreover, given they combine empirical evidence from qualitative, quantitative, and mixed methods studies, these reviews can provide a rich understanding of complex phenomena. Although empirical research has a clear definition (based directly on observation, experiment, or simulation, rather than on reasoning or theory alone),[<reflink idref="bib6" id="ref6">6</reflink>], [<reflink idref="bib7" id="ref7">7</reflink>] because mixed studies reviews include all types of empirical research designs,[<reflink idref="bib3" id="ref8">3</reflink>] search strategies often yield a high number of records to screen (sometimes more than 10 000). In fact, many of these records are totally irrelevant. The high yield means that the screening process is time consuming. Unlike reviews of randomized controlled trials, for example, because systematic mixed studies reviews include all types of design, no term referring to study design can be used to capture them. Empirical research is not referred to as empirical in articles, but rather by study design.</p> <p>In addition to that, it is estimated that approximately 1.4 million articles are written every year in scientific journals.[<reflink idref="bib8" id="ref9">8</reflink>] Estimates also indicate that the entire systematic review process typically takes about 12  months,[<reflink idref="bib9" id="ref10">9</reflink>] which may include 1 or 2  months for manual screening of records. This time scale can be problematic for researchers limited in resources. As a result, a high number of irrelevant entries must be filtered. One common practice in systematic reviews is to use highly sensitive search filters to narrow the search for relevant records.</p> <p>The filters (or classifiers) have been developed for a very specific purpose, specific study type design (eg, randomized controlled trials[<reflink idref="bib10" id="ref11">10</reflink>]) or discipline (eg, primary care[<reflink idref="bib11" id="ref12">11</reflink>]). Traditional search strategies in bibliographic databases generally have high sensitivity (ie, recall in computer science) and specificity for randomized controlled trials  but are limited for other types of research study designs.[<reflink idref="bib12" id="ref13">12</reflink>] Since mixed studies reviews are interested in several types of designs, these filters cannot be used. Also, several nonempirical works such as opinion letters, commentaries, editorials, reviews, and errata form a group of irrelevant records that are difficult to identify using traditional search filters because they often follow a research paper format (introduction, method, results, and discussion).</p> <p>El Sherif et al[<reflink idref="bib13" id="ref14">13</reflink>] proposed a mixed filter based on Boolean expressions to facilitate the identification of empirical studies for systematic mixed studies reviews. This Boolean filter covers quantitative, qualitative, and mixed methods studies  and includes keywords and subject headings for identifying empirical studies and excluding nonempirical works. This filter has shown high sensitivity (89.5%), but its precision and specificity are just over 50%.</p> <p>The task of identifying empirical studies can be cast as a text classification problem since it can be resolved with two classes: relevant (empirical) or irrelevant (nonempirical). Automated text classification is "the activity of labelling natural language texts with thematic categories from a predefined set of data."[<reflink idref="bib14" id="ref15">14</reflink>] Also, automated text classification algorithms have the potential to provide users with a confidence or likelihood scale for each prediction. Automated text classification approaches are promising avenues for reducing the burden of screening of thousands of irrelevant records often captured in bibliographic data base searches for systematic reviews. In medical topic‐specific searches, it was shown that these methods may reduce screening time by half without any loss of relevant records.[<reflink idref="bib15" id="ref16">15</reflink>] Studies about the effectiveness of automated text classification for screening papers in systematic reviews are increasingly being published.[<reflink idref="bib16" id="ref17">16</reflink>], [<reflink idref="bib17" id="ref18">17</reflink>], [<reflink idref="bib18" id="ref19">18</reflink>], [<reflink idref="bib19" id="ref20">19</reflink>]</p> <p>Extant research mainly focusses on topic‐specific algorithm training where algorithms are conditioned to measure the relationship between a research question and a study. Little is known about the automated identification of potential relevant studies for systematic mixed studies reviews. Indeed, no research has been done to evaluate the performance of automated text classification for reviews based on research methods. Therefore, the objective of this study was to identify the most performant algorithm to distinguish empirical studies from nonempirical works, thereby facilitating the search and filtering of qualitative, quantitative, and mixed methods studies. The objectives of this study cover the following points:</p> <p></p> <ulist> <item> Identify the relevant characteristics (ie, features) of both classes of document (ie, empirical and nonempirical).</item> <p></p> <item> Compare the most popular text classification methods with the Boolean "mixed filter."</item> <p></p> <item> Design a fitted model based on the most efficient algorithm and features.</item> </ulist> <hd id="AN0133441849-2">METHODS</hd> <p></p> <hd id="AN0133441849-3">Text collection</hd> <p>The text collection is a training set of preclassified records that are used to test the algorithms. This text collection consisted of sets of titles, abstracts, and full texts.</p> <hd id="AN0133441849-4">Titles and abstracts</hd> <p>In order to train and test the different algorithms, we used several collections. The first contains the 5516 entries extracted from seven journals (covering three areas: medical informatics, public health, and primary care) assembled by the developers of the Boolean mixed filter for evaluating its performance.[<reflink idref="bib13" id="ref21">13</reflink>] Second, we reused screened records and results from previous systematic reviews.[<reflink idref="bib20" id="ref22">20</reflink>], [<reflink idref="bib21" id="ref23">21</reflink>], [<reflink idref="bib22" id="ref24">22</reflink>], [<reflink idref="bib23" id="ref25">23</reflink>], [<reflink idref="bib24" id="ref26">24</reflink>], [<reflink idref="bib25" id="ref27">25</reflink>], [<reflink idref="bib26" id="ref28">26</reflink>], [<reflink idref="bib27" id="ref29">27</reflink>] These reviews cover a broad range of topics, from electronic prescription usage and participatory research to dementia and online health care. In total, approximately 10  000 records were gathered. After removing entries without abstract or full text, 8050 were included in the final collection. The relevant entries (ie, empirical) were labelled "1," and the irrelevant entries (ie, nonempirical) were labelled "0." Only titles and abstracts were considered in our initial experimentations. Subsequently, full texts were used for performance comparison. Table 1 shows the final collection distribution.</p> <p>Summary of the collection</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Subcollection&lt;/th&gt;&lt;th align="left"&gt;Empirical&lt;/th&gt;&lt;th align="left"&gt;Nonempirical&lt;/th&gt;&lt;th align="left"&gt;Total&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;El Sherif et al&lt;xref ref-type="bibr" rid="bibr13"&gt;13&lt;/xref&gt;&lt;/td&gt;&lt;td align="left"&gt;2207&lt;/td&gt;&lt;td align="left"&gt;3309&lt;/td&gt;&lt;td align="left"&gt;5516&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Khanassov et al (&lt;xref ref-type="bibr" rid="bibr24"&gt;24&lt;/xref&gt;, &lt;xref ref-type="bibr" rid="bibr25"&gt;25&lt;/xref&gt;, &lt;xref ref-type="bibr" rid="bibr26"&gt;26&lt;/xref&gt;)&lt;/td&gt;&lt;td align="left"&gt;459&lt;/td&gt;&lt;td align="left"&gt;214&lt;/td&gt;&lt;td align="left"&gt;673&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Gagnon et al&lt;xref ref-type="bibr" rid="bibr20"&gt;20&lt;/xref&gt;&lt;/td&gt;&lt;td align="left"&gt;33&lt;/td&gt;&lt;td align="left"&gt;39&lt;/td&gt;&lt;td align="left"&gt;72&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Jagosh et al(&lt;xref ref-type="bibr" rid="bibr22"&gt;22&lt;/xref&gt;, &lt;xref ref-type="bibr" rid="bibr23"&gt;23&lt;/xref&gt;) and Macaulay et al&lt;xref ref-type="bibr" rid="bibr27"&gt;27&lt;/xref&gt;&lt;/td&gt;&lt;td align="left"&gt;613&lt;/td&gt;&lt;td align="left"&gt;670&lt;/td&gt;&lt;td align="left"&gt;1283&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Granikov et al&lt;xref ref-type="bibr" rid="bibr21"&gt;21&lt;/xref&gt;&lt;/td&gt;&lt;td align="left"&gt;306&lt;/td&gt;&lt;td align="left"&gt;200&lt;/td&gt;&lt;td align="left"&gt;506&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&lt;/td&gt;&lt;td align="left"&gt;3618&lt;/td&gt;&lt;td align="left"&gt;4432&lt;/td&gt;&lt;td align="left"&gt;8050&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0133441849-5">Full texts</hd> <p>Researchers can obtain full texts automatically from reference management software, provided their institution has access. Thus, we also measured the benefits of incorporating full texts to the classification task. It should be noted that this evaluation is experimental since the availability of such content depends on database subscriptions.</p> <p>Full texts (PDF) were automatically using <emph>EndNote</emph> or retrieved manually via <emph>Google Scholar</emph>. In order to convert PDF files into usable text files, we used Tika, a content analysis toolkit developed for different document formats. It should be noted that this conversion can be fully automated using the Tika application program interface.</p> <hd id="AN0133441849-6">Datasets</hd> <p>To train the automatic classifiers (algorithms) and, thus, adjust the parameters of their mathematical functions described below, the final collection had to be separated into three datasets: a training set, a validation set, and a test set. The classifiers were tested on the same entries as the Boolean mixed filter using a fourfold cross validation. Therefore, each distinct fold contained 1136 entries for testing, 1000 entries for validation (ie, optimization), and 5914 entries for training. Entries were selected randomly while keeping the same category ratio (ie, empirical/nonempirical) between folds.</p> <hd id="AN0133441849-7">Baseline</hd> <p>The algorithms were compared with the Boolean mixed filter[<reflink idref="bib13" id="ref30">13</reflink>] as it is the only approach to distinguish empirical studies from nonempirical works. Developed by librarians and researchers with expertise in systematic mixed studies reviews, this filter consists of a combination of subject headings and keywords associated with randomized controlled trials, nonrandomized and descriptive quantitative studies, and qualitative and mixed methods studies and has been implemented for MEDLINE, an online bibliographic database. With a search engine like the one provided by MEDLINE, it is possible to build complex queries using the Boolean operators AND (ie, all keywords included), OR (ie, any keywords included), and NOT (ie, keywords not included). As such, the filter includes the expression "NOT (letter OR comment OR editorial OR newspaper article).pt." to exclude possible irrelevant publication types (".pt."). Terms associated with relevant methodologies like "case‐control," "focus group," and "grounded theory" are combined with the operator OR and searched for in titles and abstracts. To maintain flexibility, some keywords are truncated with the operator "*," allowing the search engine to look for a portion of the words like "random*," "control*," and "evaluation stud*." The Boolean filter and its toolkit are available online.  </p> <hd id="AN0133441849-8">Text characteristics</hd> <p>Automatic text classification relies on features (ie, characteristics or properties) extracted from the texts. The features we used are terms and concepts as outlined below.</p> <hd id="AN0133441849-9">Terms</hd> <p>Terms are stemmed words that we generated as follows. The words composing the abstracts and titles were used to create the initial representation of each record. Terms were determined as follows. First, common words such as "of" and "from" were removed from the documents. Words were then stemmed using the Porter algorithm.[<reflink idref="bib28" id="ref31">28</reflink>] The latter is commonly used in natural language processing to standardize singular and plural forms as well as inflected words. An internal representation of a document was then created using the extracted terms as well as their weighting. An example of internal document representation is a vector in the space formed by all the terms. Numerous indexation methods can be used for this.[<reflink idref="bib29" id="ref32">29</reflink>] TF‐IDF is the most common method for term weighting and it balances the local representativeness of a term within a document and the global discrimination of the term in the whole dataset. It should be noted that this is the technique mostly commonly used in text classification.[<reflink idref="bib30" id="ref33">30</reflink>] The values can be calculated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mfenced separators="" open="|" close="|"&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mfrac&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;msub&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>f</emph><subs><emph>t</emph>,<emph>d</emph></subs> is the frequency of term <emph>t</emph> in document <emph>d</emph>, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mfenced separators="" open="|" close="|"&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> is the length of document <emph>d</emph>, <emph>N</emph> is the total number of documents and <emph>n</emph><subs><emph>t</emph></subs> is the number of documents containing term <emph>t</emph>.</p> <hd id="AN0133441849-10">Feature selection approaches</hd> <p>Not all of the selected terms may be useful for the task of classification. Thus, to eliminate irrelevant terms and decrease computational load, features were filtered using a feature selection approach.[<reflink idref="bib31" id="ref34">31</reflink>] We compared three different feature selection methods: information gain, <emph>χ</emph><sups>2</sups> statistic test, and document frequency. Information gain can be translated as the difference between the portion of irrelevant entries considering all features and the portion of irrelevant entries given a specific feature:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>H</emph>(<emph>E</emph>) is the portion of irrelevant entries in the collection <emph>E</emph> and <emph>H</emph>(<emph>E</emph>|<emph>t</emph>) is the portion of irrelevant entries in <emph>E</emph> given a feature <emph>t</emph>.</p> <p>The <emph>χ</emph><sups>2</sups> statistic test method measures the dependency between a term and its category (empirical or nonempirical):</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msup&gt;&lt;mi&gt;&amp;#967;&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>A</emph> is the number of times term <emph>t</emph> and category <emph>c</emph> co‐occur, <emph>B</emph> is the number of times <emph>t</emph> occurs without <emph>c</emph>, <emph>C</emph> is the number of times <emph>c</emph> occurs without <emph>t</emph>, <emph>D</emph> is the number of times neither <emph>c</emph> nor <emph>t</emph> occurs, and <emph>N</emph> is the number of documents.</p> <p>The document frequency method measures the number of times a term <emph>t</emph> occurs in a document (ie, the text representing a record).</p> <p>Based on these three calculations, the features obtaining the highest values are selected and used in the classification algorithms. Using our text collection, information gain and <emph>χ</emph><sups>2</sups> statistic test generated zero values for terms excluded from the top 8000. As a result, the amount of terms selected for each measure was set to 8000, accordingly.</p> <hd id="AN0133441849-11">Concepts</hd> <p>Many concepts in the Boolean mixed filter[<reflink idref="bib13" id="ref35">13</reflink>] are compound words and cannot be captured by single terms. Using a metathesaurus is a simple way to consider complex and potentially important concepts in the indexation process. To this end, we used the Unified Medical Language System (UMLS) that provides a set of possible expressions for each concept, and relationships between concepts.[<reflink idref="bib32" id="ref36">32</reflink>] The selection process used a custom script divided in two parts: concepts in UMLS metathesaurus were stemmed and then searched for in the documents. For this task, all the concept identifiers (CUI) listed by UMLS were considered, and their associated names were added in the new set of features. The concept identifiers are located in a rich release format (RRF) file provided with the metathesaurus. In total, 2101 relevant concepts were extracted from the dictionary and added to the vectors.</p> <hd id="AN0133441849-12">Algorithms</hd> <p>Multiple studies have compared traditional text classification approaches for various problems.[<reflink idref="bib14" id="ref37">14</reflink>], [<reflink idref="bib33" id="ref38">33</reflink>], [<reflink idref="bib34" id="ref39">34</reflink>] Below, we describe these approaches that are strong options for easily exploiting machine learning algorithms for automatic text classification.</p> <hd id="AN0133441849-13">K‐nearest neighbours (kNNs)</hd> <p>K‐nearest neighbour predicts the category of a test document using the most common category of the surrounding documents (ie, nearest neighbours) in the feature space. K‐nearest neighbour is one of the best‐known statistical approaches for supervised text classification.[<reflink idref="bib35" id="ref40">35</reflink>] Among a set of training documents, the algorithm tries to identify the <emph>k</emph> closest entries from a test entry <emph>x</emph>. The majority category of the <emph>k</emph> entries is then used to classify <emph>x</emph> following a proximity weighting formula. For a test document <emph>x</emph> and a distinct training entry <emph>v</emph>, we used the Euclidian distance to represent the similarity (ie, proximity) of both entries:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;msqrt&gt;&lt;mrow&gt;&lt;mstyle displaystyle="true"&gt;&lt;munderover&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;/munderover&gt;&lt;/mstyle&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;v&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/msqrt&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>x</emph><subs><emph>i</emph></subs> and <emph>v</emph><subs><emph>i</emph></subs> are the <emph>i</emph>th features of weight vectors <emph>x</emph> and <emph>v</emph>, respectively.</p> <p>The <emph>k</emph> documents with the highest <emph>s</emph><emph>i</emph><emph>m</emph>(<emph>x</emph>,<emph>v</emph>) values were selected to represent the category of <emph>x</emph>. The estimated probabilities of <emph>x</emph> being empirical or nonempirical were calculated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"&gt;P0(x)=1&amp;#8721;i=1ksim(x,Vi)&amp;#8721;i=1kg(x,Vi,0)P1(x)=1&amp;#8721;i=1ksim(x,Vi)&amp;#8721;i=1kg(x,Vi,1),&lt;/math&gt; </ephtml> </p> <p>where <emph>V</emph><subs><emph>i</emph></subs> represents the <emph>i</emph>th nearest neighbour and <emph>P</emph><subs>0</subs>(<emph>x</emph>) and <emph>P</emph><subs>1</subs>(<emph>x</emph>) represent the likelihood of negative and positive categories, respectively.</p> <p>Weighting function <emph>g</emph> can be formulated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;V&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced separators="" open="{" close=""&gt;&lt;mrow&gt;&lt;mspace width="1em" /&gt;1sim(x,v)ifyi=c,0else&lt;mspace width="1em" /&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>y</emph><subs><emph>i</emph></subs> is the category of document <emph>V</emph><subs><emph>i</emph></subs>. The final category was based on the maximum between <emph>P</emph><subs>0</subs>(<emph>x</emph>) and <emph>P</emph><subs>1</subs>(<emph>x</emph>).</p> <hd id="AN0133441849-14">Naive Bayes</hd> <p>Naive Bayes classifiers are commonly used for automated text classification. Despite the fact that Naive Bayes approaches ignore all dependencies between features, they are still competitive with high‐capacity algorithms.[<reflink idref="bib36" id="ref41">36</reflink>] Because of this strong assumption, Naive Bayes may identify the winning category with disproportionate probabilities in some cases. Hence, the approach may provide inaccurate estimations but can still be efficient in providing the correct predictions with a large enough dataset. The typical assumption is that continuous data or features (ie, quantitative data that can be measured) are distributed according to a normal distribution. Two estimators were used for both categories of documents. Training of the classifiers for a document <emph>x</emph> of dimension <emph>m</emph> was calculated with the following conditional probability:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mover accent="true"&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mstyle displaystyle="true"&gt;&lt;munderover&gt;&lt;mo&gt;&amp;#8719;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;/munderover&gt;&lt;/mstyle&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>As stated above, the probability of observing component <emph>x</emph><subs><emph>i</emph></subs> with category <emph>c</emph> is modelled as a normal distribution. The final model follows Bayes' formula and choose the category with the highest probability:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>P</emph>(<emph>c</emph>) is the prior likelihood of category <emph>c</emph>.</p> <hd id="AN0133441849-15">Support vector machine (SVM)</hd> <p>Support vector machine can be considered as a representation of entries as points in space, where the greatest possible distance between entries from opposite categories is sought. It is one of the most popular approaches for binary classification. Based on risk minimization, the objective of the algorithm is to find the optimal hyperplane <emph>w</emph><sups><emph>T</emph></sups><emph>x</emph> + <emph>b</emph> that separate two predefined categories. To address non‐linearity, soft margins and higher dimension projections may be considered. We used the LibSVM implementation with a linear kernel to generate our classifier.[<reflink idref="bib37" id="ref42">37</reflink>]</p> <hd id="AN0133441849-16">Decision trees</hd> <p>Decision trees combine a set of approaches based mainly on rules.[<reflink idref="bib38" id="ref43">38</reflink>] They are especially useful for text classification problems since their predictions are easily interpretable. Many versions are exploitable and can be differentiated by their underlying algorithms and pruning techniques. The most common variants for this category of approaches are ID3 and its successor C4.5.[<reflink idref="bib39" id="ref44">39</reflink>] We used the latter along with its reduced error pruning (REP) method.</p> <p>C4.5 tries to minimize the entropy (ie, portion of irrelevant entries) of a group of documents by splitting them into two different subsets using a rule generated by discretization. The latter process aims to summarize the behaviour of the features using conditional operators such as &gt;, &lt;, ≤, or ≥. Let <emph>E</emph> be the initial training set and let <emph>E</emph><subs>1</subs> and <emph>E</emph><subs>2</subs> be the two subsets resulting from the separation of <emph>E</emph> using a split based on a feature. Using the entropy of these three sets, the best possible separation rule is defined as the one that provides the highest information gain. This can be calculated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"&gt;H(E)=&amp;#8722;&amp;#8721;cE1,clog2E1,c&amp;#8722;&amp;#8721;cE2,clog2E2,cIG(E)=&amp;#8722;&amp;#8721;cEclog2Ec)&amp;#8722;H(E),&lt;/math&gt; </ephtml> </p> <p>where <emph>E</emph><subs><emph>i</emph>,<emph>c</emph></subs> and <emph>E</emph><subs><emph>c</emph></subs> are the proportion of documents belonging to category <emph>c</emph> in <emph>E</emph><subs><emph>i</emph></subs> and <emph>E</emph>, respectively.</p> <p>This process is recursively applied until entropy cannot be further minimized. Pruning is then used to eliminate unnecessary splits based on the predictions of left out documents (ie, randomly and automatically selected from the training sets before the pruning process).</p> <hd id="AN0133441849-17">Method refinement</hd> <p>To improve the classification results of the approaches mentioned above, we used additional techniques: bagging and booting, feature combination, linear interpolation, and titles. It is important to mention that these techniques do not represent additional distinctive algorithms but can be seen as different ways to enhance the performance of the approaches already presented.</p> <hd id="AN0133441849-18">Bagging and boosting</hd> <p>The previous algorithms can be combined and seen as a series of prediction votes (ie, voting techniques). It has been demonstrated that voting techniques have the potential to increase the stability and capacity of traditional algorithms for automated text classification.([<reflink idref="bib40" id="ref45">40</reflink>], [<reflink idref="bib41" id="ref46">41</reflink>]) Comparisons have shown appreciable gain of precision using diversified datasets. Since a vote simply corresponds to the aggregation of predictions provided by a group of classifiers, voting techniques can be applied without additional complexity. They can be seen as meta‐algorithms since they rely on the predictions of first‐level algorithms. For the most part, aggregation represents the average of the predictions generated by high‐capacity classifiers. This representation is also referred to as bagging (ie, bootstrap aggregating). It is also possible to aggregate the predictions of multiple low capacity classifiers or weak learners (ie, boosting). The following formulas describe these two approaches.</p> <p>Assuming an arbitrary training set <emph>E</emph> separated into <emph>j</emph> subsets randomly generated with replacement. For each subset <emph>E</emph><subs><emph>i</emph></subs>, a traditional classifier <emph>H</emph><subs><emph>i</emph></subs> can be trained. In order to aggregate the predictions for a test document <emph>x</emph>, the following formula was used:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/mfrac&gt;&lt;mstyle displaystyle="true"&gt;&lt;munderover&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/munderover&gt;&lt;/mstyle&gt;&lt;msub&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>H</emph><subs><emph>i</emph></subs>(<emph>x</emph>) is the prediction of the classifier <emph>H</emph><subs><emph>i</emph></subs> given <emph>x</emph>.</p> <p>As for the boosting approach, a first weak learner <emph>H</emph><subs><emph>i</emph></subs> is trained on dataset <emph>E</emph>. Prediction results are then memorized in a vector. Subsequently, a second weak learner <emph>H</emph><subs><emph>i</emph> + 1</subs> is trained on <emph>E</emph> while making sure misclassified entries from <emph>H</emph><subs><emph>i</emph></subs> are better categorized. A total of <emph>m</emph> weak learners are trained iteratively following the same operation. The importance of each learner <emph>H</emph> is determined by a coefficient <emph>α</emph> that is based on the error rate of the learner. The error rate often represents the sum of the errors generated by the weak learner. Hence, a learner producing fewer errors will have a greater <emph>α</emph> value. Similar to the bagging approach, weak learners are then combined to determine the category of a document:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mstyle displaystyle="true"&gt;&lt;munderover&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/munderover&gt;&lt;/mstyle&gt;&lt;msub&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>α</emph><subs><emph>i</emph></subs> ≥ 0.</p> <p>The Adaboost.M1 algorithm was used to represent this approach.[<reflink idref="bib42" id="ref47">42</reflink>]</p> <hd id="AN0133441849-19">Feature combination</hd> <p>Quantitative research methods rely largely on statistical explanations. Thus, numerical terms represent an important part of the entries implicated in the classification process. For instance, numbers may be observed in the form of percentages, <emph>P</emph> values or quantities. Because the variation of number values should not influence the predictions of the classifiers, a separate <emph>Numbers</emph> feature of documents was generated by merging these particular features. The feature was weighted as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msup&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mtext&gt;&amp;#8943;&lt;/mtext&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mfenced separators="" open="|" close="|"&gt;&lt;mi&gt;Q&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mrow&gt;&lt;mfenced separators="" open="|" close="|"&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;munder&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mi&gt;Q&lt;/mi&gt;&lt;/mrow&gt;&lt;/munder&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>f</emph><subs><emph>n</emph>,<emph>d</emph></subs> is the frequency of a numeric expression <emph>n</emph> in document <emph>d</emph>, <emph>Q</emph> is the group of numeric expressions in document <emph>d</emph>, and |<emph>d</emph>| is the length of the document.</p> <p>Mathematical and statistical symbols are commonly observed in documents containing quantitative research methods. In addition to percentages (%), a large number of texts contains variables (eg, <emph>σ</emph>, <emph>α</emph>, <emph>β</emph>, and <emph>μ</emph>), operators (eg, +, =, ±, &lt;, and &gt;), and fractions or calculus symbols (eg, <sups>2</sups>, <sups>3</sups>, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msqrt&gt;&lt;mrow /&gt;&lt;/msqrt&gt;&lt;/math&gt; </ephtml> , <ephtml> &lt;math display="inline" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mfrac&gt;&lt;/math&gt; </ephtml> , and  <ephtml> &lt;math display="inline" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mn&gt;4&lt;/mn&gt;&lt;/mfrac&gt;&lt;/math&gt; </ephtml> ). Their occurrences in a document provide additional clues regarding its category. Thus, an additional <emph>Maths</emph> feature was created and weighted in the same way as the number feature.</p> <p>Unified Medical Language System provides concept associations such as synonyms. As such, by merging terms based on these relations, features may gain in homogeneity. Thus, we generated additional <emph>S</emph><emph>y</emph><emph>n</emph><emph>o</emph><emph>n</emph><emph>y</emph><emph>m</emph><subs><emph>k</emph></subs> features combining the weights (ie, frequencies) of concepts and terms appearing in an observed group of synonym <emph>k</emph>. It is important to mention that number and symbol combinations presented above could have a bigger impact on quantitative methods.</p> <p>Merging of features was done separately for the three methods. Afterward, an additional evaluation was performed using a mix of all combinations (ie, <emph>Numbers</emph>, <emph>Maths</emph>, and <emph>Synonyms</emph>).</p> <hd id="AN0133441849-20">Linear interpolation</hd> <p>The different text characteristics described above (ie, terms and concepts) can be combined during the classification process. Yet, the significance of both types of feature can also be measured in order to grant a greater degree of importance to a specific group of terms or concepts. Smoothing techniques are often used for such evaluation and are particularly popular for natural language models.[<reflink idref="bib43" id="ref48">43</reflink>] Linear interpolation (ie, Jelinek‐Mercer's method) is a common approach that allows the combination of two different classification models. Specifically, the approach uses a coefficient <emph>λ</emph> that controls the influence of two separate groups of characteristics (<emph>θ</emph><subs><emph>A</emph></subs> and <emph>θ</emph><subs><emph>B</emph></subs>):</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>Smoothing is particularly useful for classifying the model based on decision trees (M1). For upper nodes, decision trees are inclined to favour terms that are unrelated to the problem when separating training data (eg, terms associated with a journal rather than a research method). This phenomenon may affect the generalization of the two categories. Therefore, weights associated with this kind of feature should be penalized. Using the 8000 terms and 2000 concepts previously calculated, let <ephtml> &lt;math display="inline" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msup&gt;&lt;/math&gt; </ephtml> be the weight vector of terms not included in UMLS and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;/msup&gt;&lt;/math&gt; </ephtml> be the weight vector of matching concepts for document <emph>x</emph>. Predictions based on linear interpolation and decision trees can be translated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mstyle displaystyle="true"&gt;&lt;munderover&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/munderover&gt;&lt;/mstyle&gt;&lt;msub&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mstyle displaystyle="true"&gt;&lt;munderover&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/munderover&gt;&lt;/mstyle&gt;&lt;msub&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;/math&gt; </ephtml> </p> <p>where <emph>P</emph><subs><emph>i</emph></subs>(<emph>x</emph>) represents the probability distribution of document <emph>x</emph> generated by the <emph>i</emph>th tree.</p> <p>Linear interpolation was tested with λ‐values set to 0, 0.25, 0.5, 0.75, and 1.</p> <hd id="AN0133441849-21">Titles</hd> <p>Examining terms in the article titles provides important indications regarding the methodology used. To date, in the description of text characteristics, document representations do not differentiate terms from the abstracts and titles. Although term frequency for titles is meaningless, presence and absence indications may be valuable. These features can be represented as simple binary values. Let title(x) be the title of document x. New features α<subs>i</subs> ∈ {0,1} can be generated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msub&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced separators="" open="{" close=""&gt;&lt;mrow&gt;&lt;mspace width="1em" /&gt;1ifti&amp;#8712;title(x)0else,&lt;mspace width="1em" /&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/math&gt; </ephtml> </p> <p>where t<subs>i</subs> is the ith term observable in the titles.</p> <p>By reconsidering the model presented in , α components can be merged with vectors T and C. Since concepts are considerably less frequent in titles than regular terms, vectors T were chosen to carry the new features:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"&gt;xtitle=(x1,x2,&amp;#8943;,xa,&amp;#945;1,&amp;#945;2,&amp;#8943;,&amp;#945;l)Pr(xtitle|T,C)=&amp;#955;&amp;#8721;i=1kPi(xtitle|T)+(1&amp;#8722;&amp;#955;)&amp;#8721;i=1kPi(x|C),&lt;/math&gt; </ephtml> </p> <p>where l is the total number of terms observable in titles.</p> <p>Terms composing the titles were also evaluated separately in order to measure their capacity to describe the nature of a study.</p> <hd id="AN0133441849-22">Implementation</hd> <p>The approaches were implemented using Weka,  an application program interface that provides a collection of several machine learning algorithms. The features were extracted using custom scripts developed in programming language Python. The entries were indexed (ie, term weighting) in this same language.</p> <p>Once the best method was selected, a more user‐friendly and convenient tool was programmed for researchers. The source code (Java and Python) as well as our original datasets are openly accessible at the same location and can be tested on projects and additional entries, thus, improved. New entries will also be made available over time along with the tool. Otherwise, please do not hesitate to contact the authors for an access to the data.</p> <hd id="AN0133441849-23">Evaluation</hd> <p>Algorithms were directly compared with the Boolean mixed filter (labelled "baseline"). Since sensitivity, precision, specificity, and accuracy were used to evaluate the filter and considered for the new automatic text classifiers (algorithms). The four indices were calculated as follows:</p> <p> <ephtml> &lt;math display="block" altimg="urn:x-wiley:jrsm:media:jrsm1317:jrsm1317-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"&gt;Sensitivity=TPTP+FNPrecision=TPTP+FPSpecificity=TNTN+FPAccuracy=TP+TNTP+TN+FP+FN,&lt;/math&gt; </ephtml> </p> <p>where TP= number of true positives, TN= number of true negatives, FP= number of false positives, and FN= number of false negatives.</p> <hd id="AN0133441849-24">RESULTS</hd> <p></p> <hd id="AN0133441849-25">Algorithms</hd> <p>A total of 8000 terms exclusively chosen by information gain were kept. Table 2 shows the performance of the six most efficient automatic text classification approaches tested. Note for first‐level classifiers that were not improved by bagging and boosting techniques, the associated results are not included in Table 2. The additional method refinement techniques were evaluated separately (see Section 2). Bagging was tested with decision trees, Naive Bayes, kNN, and SVM. Boosting was tested with decision trees, Naive Bayes, and kNN. Most classifiers tended to perform better with nonempirical documents. The decision trees with bagging (M1) approach performed well for empirical entries (&gt;0.8) and increased the accuracy by 31.7% compared with the baseline. Support vector machine (M3) outperformed kNN and Naïve Bayes as well. These results informed the subsequent evaluations that were performed using the two best families of algorithm, that is, the decision trees (with bagging) and SVM.</p> <p>Algorithm comparison</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Algorithm&lt;/th&gt;&lt;th align="center" /&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy, %&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Bagging&amp;#8208;decision trees&lt;/td&gt;&lt;td align="center"&gt;(M1)&lt;/td&gt;&lt;td align="left"&gt;0.805&lt;/td&gt;&lt;td align="left"&gt;0.853&lt;/td&gt;&lt;td align="left"&gt;0.899&lt;/td&gt;&lt;td align="left"&gt;88.35&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Boosting&amp;#8208;decision trees&lt;/td&gt;&lt;td align="center"&gt;(M2)&lt;/td&gt;&lt;td align="left"&gt;0.776&lt;/td&gt;&lt;td align="left"&gt;0.852&lt;/td&gt;&lt;td align="left"&gt;0.879&lt;/td&gt;&lt;td align="left"&gt;87.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SVM&lt;/td&gt;&lt;td align="center"&gt;(M3)&lt;/td&gt;&lt;td align="left"&gt;0.778&lt;/td&gt;&lt;td align="left"&gt;0.825&lt;/td&gt;&lt;td align="left"&gt;0.884&lt;/td&gt;&lt;td align="left"&gt;86.42&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Decision trees&lt;/td&gt;&lt;td align="center"&gt;(M4)&lt;/td&gt;&lt;td align="left"&gt;0.763&lt;/td&gt;&lt;td align="left"&gt;0.789&lt;/td&gt;&lt;td align="left"&gt;0.878&lt;/td&gt;&lt;td align="left"&gt;84.81&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;kNN&lt;/td&gt;&lt;td align="center"&gt;(M5)&lt;/td&gt;&lt;td align="left"&gt;0.591&lt;/td&gt;&lt;td align="left"&gt;0.365&lt;/td&gt;&lt;td align="left"&gt;0.85&lt;/td&gt;&lt;td align="left"&gt;68.81&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Naive Bayes&lt;/td&gt;&lt;td align="center"&gt;(M6)&lt;/td&gt;&lt;td align="left"&gt;0.5&lt;/td&gt;&lt;td align="left"&gt;0.981&lt;/td&gt;&lt;td align="left"&gt;0.515&lt;/td&gt;&lt;td align="left"&gt;66.9&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Boolean mixed filter&lt;/td&gt;&lt;td align="center"&gt;(Baseline)&lt;/td&gt;&lt;td align="left"&gt;0.604&lt;/td&gt;&lt;td align="left"&gt;0.895&lt;/td&gt;&lt;td align="left"&gt;0.545&lt;/td&gt;&lt;td align="left"&gt;56.6&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0133441849-26">Concepts</hd> <p>Figure 1A,B shows the progression of accuracy for decision trees with bagging (M1) and SVM (M3) when concepts provided by the metathesaurus are added to the weight vectors. A maximum gain of 0.2% can be observed for decision trees. When 2000 concepts are considered in the classification process, precision, sensitivity, and specificity of decision trees with bagging increase by 0.38%, 0.1%, and 0.23%, respectively. As for SVM, accuracy gained 0.5% at 2000 additional concepts. At the same level, precision increased by 1%, sensitivity by 0.2%, and specificity by 0.6%. Table 3 gives an overview of the new performances for both algorithms.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01dec18/jrsm1317-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1317-fig-0001.jpg" title="Accuracy of decision trees with bagging (M1) and support vector machine (SVM) (M3) using concepts" /> </p> <p></p> <p>Performances of decision trees with bagging (M1) and SVM (M3) using 2000 concepts</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Algorithm&lt;/th&gt;&lt;th align="center" /&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy, %&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Bagging&amp;#8208;decision trees&lt;/td&gt;&lt;td align="center"&gt;(M1)&lt;/td&gt;&lt;td align="left"&gt;0.809&lt;/td&gt;&lt;td align="left"&gt;0.854&lt;/td&gt;&lt;td align="left"&gt;0.9&lt;/td&gt;&lt;td align="left"&gt;88.53&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SVM&lt;/td&gt;&lt;td align="center"&gt;(M3)&lt;/td&gt;&lt;td align="left"&gt;0.788&lt;/td&gt;&lt;td align="left"&gt;0.827&lt;/td&gt;&lt;td align="left"&gt;0.89&lt;/td&gt;&lt;td align="left"&gt;86.92&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>We experimented with different numbers of concepts as features. Figure 1 shows how accuracy changes according to the number of concepts with M1 and M3. Our results indicate the ideal number of concepts for M1 and M3 is, respectively, 2000 and 1200. </p> <p>To further assess the influence of concepts, we examined the top 50 features (including terms) selected using information gain. Our results indicate that 67% of these features are concepts included in UMLS. This shows that concepts are extensively used by the classification algorithms. The fact that the addition of concepts did not increase performance measures by large margins can be explained by the overlap between terms and concepts: most of these concepts would have been covered by terms if concepts are not used.</p> <p>Figure 2 shows a list of lemmatized concepts from the initial group of 2000 frequently selected by decision trees. Results indicate that these are multi‐word concepts (which are more precise than single words or terms).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01dec18/jrsm1317-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1317-fig-0002.jpg" title="Twenty concepts selected by decision trees with bagging (M1)" /> </p> <p></p> <hd id="AN0133441849-29">Method refinement</hd> <p></p> <hd id="AN0133441849-30">Feature combination</hd> <p>Features were combined following the methods presented in the previous section. Table 4 provides an overview of how the different combinations, using decision trees with bagging and SVM, performed.</p> <p>Performances of decision trees with bagging (M1) and SVM (M3) with feature combination</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Type&lt;/th&gt;&lt;th align="left"&gt;Classifier&lt;/th&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy, %&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;italic&gt;N&lt;/italic&gt;&lt;italic&gt;u&lt;/italic&gt;&lt;italic&gt;m&lt;/italic&gt;&lt;italic&gt;b&lt;/italic&gt;&lt;italic&gt;e&lt;/italic&gt;&lt;italic&gt;r&lt;/italic&gt;&lt;italic&gt;s&lt;/italic&gt;&lt;/td&gt;&lt;td align="left"&gt;M1&lt;/td&gt;&lt;td align="left"&gt;0.807&lt;/td&gt;&lt;td align="left"&gt;0.849&lt;/td&gt;&lt;td align="left"&gt;0.901&lt;/td&gt;&lt;td align="left"&gt;88.35 (&amp;#8201;&amp;#8722;&amp;#8201;0.18)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;M3&lt;/td&gt;&lt;td align="left"&gt;0.776&lt;/td&gt;&lt;td align="left"&gt;0.833&lt;/td&gt;&lt;td align="left"&gt;0.882&lt;/td&gt;&lt;td align="left"&gt;86.56 (&amp;#8201;&amp;#8722;&amp;#8201;0.36)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;italic&gt;S&lt;/italic&gt;&lt;italic&gt;y&lt;/italic&gt;&lt;italic&gt;m&lt;/italic&gt;&lt;italic&gt;b&lt;/italic&gt;&lt;italic&gt;o&lt;/italic&gt;&lt;italic&gt;l&lt;/italic&gt;&lt;italic&gt;s&lt;/italic&gt;&lt;/td&gt;&lt;td align="left"&gt;M1&lt;/td&gt;&lt;td align="left"&gt;0.785&lt;/td&gt;&lt;td align="left"&gt;0.856&lt;/td&gt;&lt;td align="left"&gt;0.885&lt;/td&gt;&lt;td align="left"&gt;87.55 (&amp;#8201;&amp;#8722;&amp;#8201;0.98)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;M3&lt;/td&gt;&lt;td align="left"&gt;0.788&lt;/td&gt;&lt;td align="left"&gt;0.828&lt;/td&gt;&lt;td align="left"&gt;0.89&lt;/td&gt;&lt;td align="left"&gt;86.94 (&amp;#8201;+&amp;#8201;0.02)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;italic&gt;S&lt;/italic&gt;&lt;italic&gt;y&lt;/italic&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;italic&gt;o&lt;/italic&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;italic&gt;y&lt;/italic&gt;&lt;italic&gt;m&lt;/italic&gt;&lt;italic&gt;s&lt;/italic&gt;&lt;/td&gt;&lt;td align="left"&gt;M1&lt;/td&gt;&lt;td align="left"&gt;0.811&lt;/td&gt;&lt;td align="left"&gt;0.834&lt;/td&gt;&lt;td align="left"&gt;0.905&lt;/td&gt;&lt;td align="left"&gt;88.12 (&amp;#8201;&amp;#8722;&amp;#8201;0.41)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;M3&lt;/td&gt;&lt;td align="left"&gt;0.791&lt;/td&gt;&lt;td align="left"&gt;0.826&lt;/td&gt;&lt;td align="left"&gt;0.892&lt;/td&gt;&lt;td align="left"&gt;87.01 (&amp;#8201;+&amp;#8201;0.09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;All&lt;/td&gt;&lt;td align="left"&gt;M1&lt;/td&gt;&lt;td align="left"&gt;0.788&lt;/td&gt;&lt;td align="left"&gt;0.836&lt;/td&gt;&lt;td align="left"&gt;0.889&lt;/td&gt;&lt;td align="left"&gt;87.19 (&amp;#8201;&amp;#8722;&amp;#8201;1.34)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;M3&lt;/td&gt;&lt;td align="left"&gt;0.777&lt;/td&gt;&lt;td align="left"&gt;0.831&lt;/td&gt;&lt;td align="left"&gt;0.882&lt;/td&gt;&lt;td align="left"&gt;86.51 (&amp;#8201;&amp;#8722;&amp;#8201;0.41)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Combining synonyms and symbols slightly increased SVM accuracy (+0.09% and +0.02%, respectively). However, most combinations negatively affected decision tree final predictions.</p> <hd id="AN0133441849-31">Linear interpolation</hd> <p>Results are shown in table 5. When <emph>λ</emph> = 0.75, feature smoothing increased accuracy and precision by 0.16% and 0.5%, respectively. Although concepts from the thesaurus provide substantial support to predictions, regular terms still have a greater impact on the model. Compared with the approach that combined concepts and terms, the smoothing approach is slightly more effective.</p> <p>Performances of the model based on interpolation</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;italic&gt;&amp;#955;&lt;/italic&gt;&amp;#8208;value&lt;/th&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy, %&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;No smoothing&lt;/td&gt;&lt;td align="left"&gt;0.809&lt;/td&gt;&lt;td align="left"&gt;0.854&lt;/td&gt;&lt;td align="left"&gt;0.9&lt;/td&gt;&lt;td align="left"&gt;88.53&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0.803&lt;/td&gt;&lt;td align="left"&gt;0.848&lt;/td&gt;&lt;td align="left"&gt;0.898&lt;/td&gt;&lt;td align="left"&gt;88.14 (&amp;#8201;&amp;#8722;&amp;#8201;0.39)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.25&lt;/td&gt;&lt;td align="left"&gt;0.812&lt;/td&gt;&lt;td align="left"&gt;0.851&lt;/td&gt;&lt;td align="left"&gt;0.903&lt;/td&gt;&lt;td align="left"&gt;88.6 (&amp;#8201;+&amp;#8201;0.07)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.5&lt;/td&gt;&lt;td align="left"&gt;0.813&lt;/td&gt;&lt;td align="left"&gt;0.851&lt;/td&gt;&lt;td align="left"&gt;0.904&lt;/td&gt;&lt;td align="left"&gt;88.66 (&amp;#8201;+&amp;#8201;0.13)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.75&lt;/td&gt;&lt;td align="left"&gt;0.814&lt;/td&gt;&lt;td align="left"&gt;0.852&lt;/td&gt;&lt;td align="left"&gt;0.904&lt;/td&gt;&lt;td align="left"&gt;88.69 (&amp;#8201;+&amp;#8201;0.16)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;0.805&lt;/td&gt;&lt;td align="left"&gt;0.858&lt;/td&gt;&lt;td align="left"&gt;0.898&lt;/td&gt;&lt;td align="left"&gt;88.35 (&amp;#8201;&amp;#8722;&amp;#8201;0.18)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0133441849-32">Titles</hd> <p>Table 6 lists some examples of predominant design indications that can be observed in titles and associated with a specific category.</p> <p>Design indications in titles</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Indication&lt;/th&gt;&lt;th align="left"&gt;Frequency&lt;/th&gt;&lt;th align="left"&gt;Most Likely Category&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Review&lt;/td&gt;&lt;td align="left"&gt;403&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Comment&lt;/td&gt;&lt;td align="left"&gt;359&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Analysis&lt;/td&gt;&lt;td align="left"&gt;265&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Controlled trial&lt;/td&gt;&lt;td align="left"&gt;214&lt;/td&gt;&lt;td align="left"&gt;Empirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Systematic review&lt;/td&gt;&lt;td align="left"&gt;196&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Qualitative&lt;/td&gt;&lt;td align="left"&gt;132&lt;/td&gt;&lt;td align="left"&gt;Empirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Cohort profile&lt;/td&gt;&lt;td align="left"&gt;119&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Erratum/corrigendum&lt;/td&gt;&lt;td align="left"&gt;95&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Response&lt;/td&gt;&lt;td align="left"&gt;87&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Cohort study&lt;/td&gt;&lt;td align="left"&gt;79&lt;/td&gt;&lt;td align="left"&gt;Empirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Meta&amp;#8208;analysis&lt;/td&gt;&lt;td align="left"&gt;76&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Case study&lt;/td&gt;&lt;td align="left"&gt;46&lt;/td&gt;&lt;td align="left"&gt;Empirical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Opinion/editorial&lt;/td&gt;&lt;td align="left"&gt;24&lt;/td&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 7 summarizes the scores of the new model in comparison with decision trees and bagging without smoothing for abstracts. The best results were obtained when <emph>λ</emph> is perfectly balanced (0.5). Accuracy increased by 0.07% as opposed to the previous smoothed models, and by 0.23% as opposed to decision trees without smoothing. In total, 137 occurrences of features associated with the titles are exploited by decision trees to split the training set. However, the contribution of these new features is arguable.</p> <p>Performances of the model based on interpolation with title features</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;italic&gt;&amp;#955;&lt;/italic&gt;&amp;#8208;value&lt;/th&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy, %&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;No smoothing&lt;/td&gt;&lt;td align="left"&gt;0.809&lt;/td&gt;&lt;td align="left"&gt;0.854&lt;/td&gt;&lt;td align="left"&gt;0.9&lt;/td&gt;&lt;td align="left"&gt;88.53&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0.803&lt;/td&gt;&lt;td align="left"&gt;0.848&lt;/td&gt;&lt;td align="left"&gt;0.898&lt;/td&gt;&lt;td align="left"&gt;88.14 (&amp;#8201;&amp;#8722;&amp;#8201;0.39)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.25&lt;/td&gt;&lt;td align="left"&gt;0.814&lt;/td&gt;&lt;td align="left"&gt;0.85&lt;/td&gt;&lt;td align="left"&gt;0.905&lt;/td&gt;&lt;td align="left"&gt;88.64 (&amp;#8201;+&amp;#8201;0.11)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.5&lt;/td&gt;&lt;td align="left"&gt;0.817&lt;/td&gt;&lt;td align="left"&gt;0.85&lt;/td&gt;&lt;td align="left"&gt;0.906&lt;/td&gt;&lt;td align="left"&gt;88.76 (&amp;#8201;+&amp;#8201;0.23)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.75&lt;/td&gt;&lt;td align="left"&gt;0.812&lt;/td&gt;&lt;td align="left"&gt;0.847&lt;/td&gt;&lt;td align="left"&gt;0.903&lt;/td&gt;&lt;td align="left"&gt;88.48 (&amp;#8201;&amp;#8722;&amp;#8201;0.05)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;0.808&lt;/td&gt;&lt;td align="left"&gt;0.853&lt;/td&gt;&lt;td align="left"&gt;0.901&lt;/td&gt;&lt;td align="left"&gt;88.48 (&amp;#8201;&amp;#8722;&amp;#8201;0.05)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0133441849-33">Full texts</hd> <p>Based on previous results, we used the classifier based on decision trees with bagging to evaluate its performance using full texts exclusively.</p> <p>Table 8 shows the gains from abstract to full text representations when concepts are added to the vectors and the three combination approaches are applied. Full text classification is positively influenced by the new features in every case, with the exception of synonyms. Combining the numbers has the greatest positive impact on predictions, with a precision increase of 0.6%. Concepts have a sensitivity gain of 1.23% compared with 0.1% for abstracts.</p> <p>Gain (%) provided by additional features for full texts compared with abstracts</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Type&lt;/th&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="center"&gt;Concepts&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Abstracts&lt;/td&gt;&lt;td align="left"&gt;+0.4&lt;/td&gt;&lt;td align="left"&gt;+0.1&lt;/td&gt;&lt;td align="left"&gt;+0.2&lt;/td&gt;&lt;td align="left"&gt;+0.3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Full texts&lt;/td&gt;&lt;td align="left"&gt;+0.3&lt;/td&gt;&lt;td align="left"&gt;+1.23&lt;/td&gt;&lt;td align="left"&gt;+0.1&lt;/td&gt;&lt;td align="left"&gt;+0.45&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="center"&gt;Number combination&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Abstracts&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.2&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.5&lt;/td&gt;&lt;td align="left"&gt;+0.1&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Full texts&lt;/td&gt;&lt;td align="left"&gt;+0.6&lt;/td&gt;&lt;td align="left"&gt;+0&lt;/td&gt;&lt;td align="left"&gt;+0.3&lt;/td&gt;&lt;td align="left"&gt;+0.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="center"&gt;Symbol combination&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Abstracts&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;2.4&lt;/td&gt;&lt;td align="left"&gt;+0.2&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.5&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.98&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Full texts&lt;/td&gt;&lt;td align="left"&gt;+0.6&lt;/td&gt;&lt;td align="left"&gt;+0.1&lt;/td&gt;&lt;td align="left"&gt;+0.3&lt;/td&gt;&lt;td align="left"&gt;+0.22&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="center"&gt;Synonym combination&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Abstracts&lt;/td&gt;&lt;td align="left"&gt;+0.2&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;2&lt;/td&gt;&lt;td align="left"&gt;+0.5&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.41&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Full texts&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.9&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.2&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.6&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.44&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="center"&gt;All combinations&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Abstracts&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;2.1&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;1.8&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;1.1&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;1.34&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Full texts&lt;/td&gt;&lt;td align="left"&gt;+0&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.5&lt;/td&gt;&lt;td align="left"&gt;+0&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.2&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Table 9 shows the overall performance of the interpolation model (<emph>λ</emph> = 0.5) for both empirical and nonempirical entries on abstracts and full texts. Although full texts include more detail than abstracts, the final scores for both types of content is similar. When feature combination is active, the most discriminating terms/concepts reported by decision trees (M1) are both involved in full texts and abstracts, which explains the similar results.</p> <p>Overall performances of the final model</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Category&lt;/th&gt;&lt;th align="left"&gt;Precision&lt;/th&gt;&lt;th align="left"&gt;Sensitivity&lt;/th&gt;&lt;th align="left"&gt;Specificity&lt;/th&gt;&lt;th align="left"&gt;Accuracy, %&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="center"&gt;Abstracts (best &lt;italic&gt;&amp;#955;&lt;/italic&gt;&amp;#8201;=&amp;#8201;0.75)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Empirical&lt;/td&gt;&lt;td align="left"&gt;0.814&lt;/td&gt;&lt;td align="left"&gt;0.852&lt;/td&gt;&lt;td align="left"&gt;0.904&lt;/td&gt;&lt;td align="left"&gt;...&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;td align="left"&gt;0.925&lt;/td&gt;&lt;td align="left"&gt;0.904&lt;/td&gt;&lt;td align="left"&gt;0.852&lt;/td&gt;&lt;td align="left"&gt;...&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Average&lt;/td&gt;&lt;td align="left"&gt;0.87&lt;/td&gt;&lt;td align="left"&gt;0.878&lt;/td&gt;&lt;td align="left"&gt;0.878&lt;/td&gt;&lt;td align="left"&gt;88.7&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="center"&gt;Full texts (best &lt;italic&gt;&amp;#955;&lt;/italic&gt;&amp;#8201;=&amp;#8201;0.5)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Empirical&lt;/td&gt;&lt;td align="left"&gt;0.863&lt;/td&gt;&lt;td align="left"&gt;0.854&lt;/td&gt;&lt;td align="left"&gt;0.933&lt;/td&gt;&lt;td align="left"&gt;...&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Nonempirical&lt;/td&gt;&lt;td align="left"&gt;0.928&lt;/td&gt;&lt;td align="left"&gt;0.933&lt;/td&gt;&lt;td align="left"&gt;0.854&lt;/td&gt;&lt;td align="left"&gt;...&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Average&lt;/td&gt;&lt;td align="left"&gt;0.896&lt;/td&gt;&lt;td align="left"&gt;0.894&lt;/td&gt;&lt;td align="left"&gt;0.894&lt;/td&gt;&lt;td align="left"&gt;90.71&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0133441849-34">DISCUSSION</hd> <p>The general observation on the classification algorithms shows that decision trees and bagging perform best, followed by SVM. These three algorithms are clearly better than Naive Bayes and kNN algorithms tested in this study. Moreover, they performed better than the manual Boolean filter (baseline) suggesting they can be used in pace of this filter. An important advantage of automatic classifiers is that they can be trained automatically. Our experiments show that words (terms) are the basic useful features that one can extract and select from abstracts and full texts. Additional features based on numerical and mathematical expressions, as well as concepts, can provide small, but limited, improvements (especially when full texts are used).</p> <p>Prediction errors generated by decision trees and SVM (M1 and M3) occur with various research methods. Therefore, it is not possible to propose a general solution to improve the classifiers. Additionally, some publication types are often mentioned in both empirical and nonempirical records. For example, "action research" occurs in 246 abstracts of nonempirical works and 384 abstracts of empirical studies. This issue is not uncommon. Our results indicated that predictions for randomized controlled trials are influenced by ambiguous terms like "trial" (475 negative abstracts and 241 positive abstracts). However, most of the prediction errors made by the decision trees with linear interpolation model share common characteristics regarding false positives. Numerous entries labelled as negative and containing empirical research method keywords were incorrectly identified by the algorithm. Meta‐analysis and reviews are directly linked to this problem. In our study, it was not unusual to observe co‐occurrences of concepts related to opposite classes such as "review" and "controlled trial" (325 abstracts), "review" and "cohort study" (176 abstracts), "meta‐analysis" and "controlled trial" (133 abstracts), as well as "meta‐analysis" and "case‐control" (46 abstracts).</p> <p>False negatives were less common given that letters, editorials, commentaries, and errata were usually correctly identified by both decision trees and SVM. In fact, precision for the negative class was considerably higher (+92%). However, there are a few similarities among false negatives for all the classification methods. More than half of these abstracts did not follow a typical structure with keywords such as "objective," "results," and "conclusion". Also, short abstracts with vague descriptions were often rejected by all the algorithms we tested. Finally, negative concepts are sometimes included in empirical studies. For instance, we observed "review" 293 times in false negatives, "systematic" 121 times, and "analysis" 192 times.</p> <p>A benefit of using automated text classification methods, other than SVM, for categorizing empirical studies is their ability to provide a confidence score along with the predictions. Even though Naive Bayes and kNN provided irregular distributions for correct and incorrect predictions, decision trees resulted in a relatively coherent model for confidence scores. Regarding feature interpolation, the average disparity between the actual and the predicted classes was 19.21% with a median of 18.3%. In practice, for librarians requiring a reliable confidence scale, these results may be acceptable. To illustrate, a user who chooses to set the confidence threshold of the algorithm to 33% is able to get a greater sensitivity without undue interference on the precision.</p> <p>Regarding the three methods of feature combination on abstracts, poor overall performance was observed when used on abstracts only. These results can be explained in four ways: insufficient detail in abstracts, ambiguous connotations of numerical terms, uneven distribution of mathematical symbols within categories, and limited coverage of synonyms. An important aspect that is difficult to capture by number combinations is the variation of meanings associated with particular features. Occurrences of term "2" in different sequences such as "2 years old," "type 2 diabetes," "<emph>p</emph> = 2," and "2 patients" do not hold the same information. Since documents are very short, this phenomenon may have a negative impact on classification. Furthermore, in this study, merging mathematical and statistical symbols in a single feature did not lead to noticeable improvements in performance. Upon further examination, our data show that symbols have low occurrence frequency per document. In fact, the median of symbol occurrences in empirical documents is close to 1 and nearly 0 for nonempirical articles.</p> <p>A similar problem can be observed for the approach based on synonym combination. Specifically, a group of synonyms contains only eight concepts, on average, with a low frequency per document. As a result, the scope of each group is considerably reduced. The use of hypernym relations (ie, broader concepts) proposed by UMLS, for instance, may address this problem. These relations are particularly popular for document and query expansion.[<reflink idref="bib44" id="ref49">44</reflink>] Despite the potential impact of synonym combinations on the classification of all three types of research methods (ie, quantitative, qualitative, and mixed), we were wary of the fact that number and symbol combinations could result in a bias towards quantitative and mixed methods. Nevertheless, the proposed automated text classification for systematic mixed studies reviews is promising as it suggests researchers can use supervised machine learning for screening records. In comparison with manually screening titles and abstracts, combining this method specific automatic text classification method with topic‐specific automated text classification could potentially save hours of work by, for the most part, reducing the number of irrelevant records to manually screen. Future work could test this. Provided that reviewers can retrieve full‐text publications in an automatic, systematic, and reliable manner, the proposed algorithms may represent an important innovation and transform systematic review processes.</p> <p>Given the absence of universal access to full‐text publications, a combination of abstracts and full texts can be used as training data to enhance the predictions of M1 and M3. Figure 3 presents a possible scenario, illustrating the performance progression according to the ratio of full texts to abstracts in the collection. There is a high correlation between the general performance of our algorithm based on decision trees/bagging and the variation of the full‐text ratio. However, sensitivity appears to be negatively affected by the mixture. There is also a decrease of almost every performance index when full‐text ratio is relatively small. Alternatively, two distinct classifiers could be used: one for abstracts and one for full texts. In such a scenario, abstracts would need to be automatically differentiated from full texts prior to classification.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01dec18/jrsm1317-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1317-fig-0003.jpg" title="Performances of decision trees with bagging (M1) mixing abstracts and full texts" /> </p> <p></p> <p>Assuming an almost complete availability of full content provided by <emph>Google Scholar</emph>, reviewers would still need an automated tool to extract full texts from the pages listed by the search engine. Such a tool may require a web crawler[<reflink idref="bib45" id="ref50">45</reflink>] and a complete evaluation on a generic data collection. It is important to note that we have not proposed a tool for this type of operation.</p> <p>The proposed automated text classification (M1) performs very well for excluding nonempirical works (high precision is important for negative class), that is negative sampling. This suggests potential applications for future systematic and nonsystematic reviews. First, in systematic mixed studies reviews, high sensitivity is key. Reviewers seek the entire population of studies (exhaustive search in a comprehensive set of bibliographic databases and in the grey literature) to answer specific questions (qualitative and/or quantitative). For example, "In population P, what is the effectiveness of the intervention I (compared to intervention C) regarding the Outcome O?" and "what are the views and the life‐experience of end‐users and their relatives with regard to the planning, implementation, evaluation and sustainability of intervention I?" Thus, researchers could consider using M1 as an initial screening/filtering procedure to exclude irrelevant documents with high precision. Two independent researchers could then proceed with manual screening to select relevant studies to include in the review. To ensure no relevant studies are lost with the initial automatic text classification step, the process could be completed with citation tracking.[<reflink idref="bib46" id="ref51">46</reflink>]</p> <p>Second, M1 can be of interest in nonsystematic reviews, where an exhaustive search is not required. For example, for theses and dissertations, graduate students do not conduct exhaustive searches of all relevant publications could save time by combining the Boolean mixed filter (high sensitivity) with M1 (high specificity) to obtain a large (good enough) sample of studies. Likewise, sensitivity is not an issue in nonsystematic scoping reviews[<reflink idref="bib47" id="ref52">47</reflink>] where reviewers seek a sample of the population of studies to address a large‐scope (broad) question. Thus, researchers could consider combining the Boolean mixed filter (high sensitivity) to rule in a large sample of publications and M1 (high specificity) to rule out nonempirical work.</p> <p>Automated text classification can be easy to use. For example, if made available online, reviewers could export their records from reference management software. The tool would classify records into empirical and nonempirical sets of records. The two sets of records could then be imported to the reference management software. We have built a complete M1 tool (including a user interface) for categorizing records saved in a spreadsheet. The Method Development component of the Quebec‐SPOR SUPPORT Unit is currently building and testing a website to disseminate the Automated Text Classification of Empirical Records (ATCER) and a user guide. The user guide will include the abovementioned recommendations for using the algorithm and its limitations. To access the website, please go to https://atcer.iro.umontreal.ca.</p> <hd id="AN0133441849-36">LIMITATION</hd> <p>One major aspect to consider regarding the results of this study is the limited amount of data used for training the algorithms. Knowing that <emph>PubMed</emph> alone includes at least 420  000 randomized controlled trials and almost 800  000 clinical studies, further tests should be performed to ensure the performance results of the algorithm we report herein are not influenced by our limited data distribution. However, for such tests, the mass extraction of training data from bibliographic databases should be supervised to ensure valid labelling (ie, empirical vs  nonempirical).</p> <p>Some drawbacks to using decision trees should be noted. The risk of overfitting is high, even with the use of pruning techniques. This occasionally applies when an algorithm is overtrained on a collection that does not represent the full population. For instance, commonly occurring research questions/disciplines and methods can greatly influence the categorization. In addition, decision trees are relatively unstable. In other words, small modifications applied to the training set can lead to very different predictions. Because of these difficulties, feature selection and training must be based on balanced collections with diversified methodologies.</p> <hd id="AN0133441849-37">CONCLUSION</hd> <p>Automated text classification of empirical studies (vs  nonempirical works) is a promising option to use when conducting nonsystematic literature reviews, but further testing is required to verify its performance for systematic reviews. We propose a supervised machine learning algorithm that can facilitate the identification of empirical studies in bibliographic databases (ie, the search for qualitative, quantitative, and mixed methods evidence) for systematic reviews. This can be used as an alternative or a complement to the Boolean mixed filter. Our results suggest that decision trees can surpass the accuracy of manual queries by at least 30% without influencing sensitivity. More importantly, the presented models obtained very high precision scores (+92%) for nonempirical works and could be used for removing entries rather than selecting studies.</p> <p>The use of separate features for concepts (extracted from a metathesaurus) and terms in titles moderately increased the performance of our methods. Varying the weights between terms and concepts provided gains as well, especially for precision (+0.5%) when the two groups of features had similar importance. In addition, the combination of features representing numbers, symbols, and synonyms was evaluated  but did not enhance results sufficiently to be considered helpful for abstracts. Finally, the use of abstracts in the classification was compared with the use of full texts. Results showed very small gains for specificity and accuracy (≈2%) and noticeable gains for precision (≈5%) when full texts were employed.</p> <p>It is important to specify that the nature of a relevant entry may slightly differ according to reviewers' perspectives and chosen topics. Hence, generic training should be followed by adjustment processes based on users' preferences. For example, the proposed classifiers can be improved online when new examples are provided during their utilization. Active learning approaches, which are commonly used to rectify classifier behaviours for automated text classification,[<reflink idref="bib48" id="ref53">48</reflink>], [<reflink idref="bib49" id="ref54">49</reflink>] can also be used. Further research is needed to evaluate the proposed models using a much larger collection, to compare our results with unsupervised machine learning, and to classify empirical records in accordance with the main study designs to facilitate syntheses (ie, qualitative research, quantitative descriptive, nonrandomized studies, randomized trials, and mixed methods research).</p> <hd id="AN0133441849-38">ACKNOWLEDGEMENTS</hd> <p>We would like to thank Reem El Sherif, Genevieve Gore, and Vera Granikov for providing the data used by the Boolean mixed filter and for their valuable input. We would also like to thank Drs  Isabelle Vedel and Marie‐Pierre Gagnon for supplying additional records used in our collection. This study was supported by the Quebec SPOR SUPPORT Unit (<ulink href="http://unitesoutiensrapqc.ca/english/">http://unitesoutiensrapqc.ca/english/</ulink>).</p> <hd id="AN0133441849-39">CONFLICT OF INTEREST</hd> <p>The author reported no conflict of interest.</p> <ref id="AN0133441849-40"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> https://tika.apache.org.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> <ulink href="http://www.ncbi.nlm.nih.gov/books/NBK3827/table/pubmedhelp.T.stopwords">www.ncbi.nlm.nih.gov/books/NBK3827/table/pubmedhelp.T.stopwords</ulink>.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref3" type="bt">3</bibl> <bibtext> <ulink href="http://toolkit4mixedstudiesreviews.pbworks.com">http://toolkit4mixedstudiesreviews.pbworks.com</ulink>.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref4" type="bt">4</bibl> <bibtext> <ulink href="http://www.cs.waikato.ac.nz/ml/weka/">http://www.cs.waikato.ac.nz/ml/weka/</ulink>.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref5" type="bt">5</bibl> <bibtext> https://atcer.iro.umontreal.ca.</bibtext> </blist> </ref> <ref id="AN0133441849-41"> <title> REFERENCES </title> <blist> <bibtext> Pluye P, Hong QN, Bush P, Vedel I. Opening‐up the definition of systematic literature review: the plurality of worldviews, methodologies and methods for reviews and syntheses. J Clin Epidemiol. 2016 ; 73 : 2 ‐ 5.</bibtext> </blist> <blist> <bibtext> Ioannidis J. The mass production of redundant, misleading, and conflicted systematic reviews and meta‐analyses. Milbank Q. 2016 ; 94 (3): 485 ‐ 514.</bibtext> </blist> <blist> <bibtext> Pluye P, Hong QN. Combining the power of stories and the power of numbers: mixed methods research and mixed studies reviews. Public Health. 2014 ; 35 (1): 29 ‐ 45.</bibtext> </blist> <blist> <bibtext> Heyvaert M, Hannes K, Onghena P. Using Mixed Methods Research Synthesis for Literature Reviews: The Mixed Methods Research Synthesis Approach. Los Angeles : SAGE Publications ; 2016.</bibtext> </blist> <blist> <bibtext> Souto RQ, Khanassov V, Hong QN, Bush P, Vedel I, Pluye P. Systematic mixed studies reviews: updating results on the reliability and efficiency of the mixed methods appraisal tool. Int J Nurs Stud. 2015 ; 52 (1): 500 ‐ 501.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref6" type="bt">6</bibl> <bibtext> Porta M, Greenland S, Hernán M, dos Santos Silva I, Last J. A Dictionary of Epidemiology. New York : Oxford University Press ; 2014.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref7" type="bt">7</bibl> <bibtext> Abbott A. The causal devolution. Sociol Methods Res. 1998 ; 27 (2): 148 ‐ 181.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref9" type="bt">8</bibl> <bibtext> Björk BC, Roos A, Lauri M. Scientific journal publishing: yearly volume and open access availability. Inf Res. 2009 ; 14 (1): 391.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref10" type="bt">9</bibl> <bibtext> Ganann R, Ciliska D, Thomas H. Expediting systematic reviews: methods and implications of rapid reviews. Implement Sci. 2010 ; 5 (1): 56.</bibtext> </blist> <blist> <bibtext> McKibbon KA, Wilczynski NL, Haynes RB. Developing optimal search strategies for retrieving qualitative studies in PsycINFO. Eval Health Prof. 2006 ; 29 (4): 440 ‐ 454.</bibtext> </blist> <blist> <bibtext> Gill PJ, Roberts NW, Wang KY, Heneghan C. Development of a search filter for identifying studies completed in primary care. Fam Pract. 2014 ; 31 (6): 739 ‐ 745.</bibtext> </blist> <blist> <bibtext> Lefebvre C, Manheimer E, Glanville J. Chapter 6: Searching for studies. Cochrane Handbook for Systematic Reviews of Interventions. Chichester (UK) : John Wiley &amp; Sons ; 2008 : 95 ‐ 150.</bibtext> </blist> <blist> <bibtext> El Sherif R, Pluye P, Gore G, Granikov V, Hong QN. Performance of a mixed filter to identify relevant studies for mixed studies reviews. J Med Libr Assoc. 2016 ; 104 (1): 47.</bibtext> </blist> <blist> <bibtext> Sebastiani F. Machine learning in automated text categorization. ACM Comput Surv. 2002 ; 34 (1): 1 ‐ 47.</bibtext> </blist> <blist> <bibtext> O'Mara‐Eves A, Thomas J, McNaught J, Miwa M, Ananiadou S. Using text mining for study identification in systematic reviews: a systematic review of current approaches. Syst Rev. 2015 ; 4 (1): 5.</bibtext> </blist> <blist> <bibtext> Shemilt I, Simon A, Hollands GJ, et al. Pinpointing needles in giant haystacks: use of text mining to reduce impractical screening workload in extremely large scoping reviews. Res Syn Meth. 2014 ; 5 (1): 31 ‐ 49.</bibtext> </blist> <blist> <bibtext> Howard BE, Phillips J, Miller K, et al. Swift‐review: a text‐mining workbench for systematic review. Syst Rev. 2016 ; 5 (1): 1.</bibtext> </blist> <blist> <bibtext> Yuanhan M, Kontonatsios G, Ananiadou S. Supporting systematic reviews using LDA‐based document representations. Syst Rev. 2015 ; 4 (1): 1.</bibtext> </blist> <blist> <bibtext> Thomas J, O'Mara‐Eves A, McNaught J, Ananiadou S. The potential of text mining to reduce screening workload in systematic reviews: a retrospective evaluation. Better Knowledge for Better Health. Abstracts of the 21st Cochrane Colloquium; 2013.</bibtext> </blist> <blist> <bibtext> Gagnon MP, Nsangou ÉR, Payne‐Gagnon J, Grenier S, Sicotte C. Barriers and facilitators to implementing electronic prescription: a systematic review of user groups' perceptions. J Am Med Inform Assoc. 2014 ; 21 (3): 535 ‐ 541.</bibtext> </blist> <blist> <bibtext> Granikov V, El Sherif R, Pluye P. Patient information aid: promoting the right to know, evaluate, and share consumer health information found on the internet. J Consum Health Internet. 2015 ; 19 (3‐4): 233 ‐ 240.</bibtext> </blist> <blist> <bibtext> Jagosh J, Macaulay AC, Pluye P, et al. Uncovering the benefits of participatory research: implications of a realist review for health research and practice. Milbank Q. 2012 ; 90 (2): 311 ‐ 346.</bibtext> </blist> <blist> <bibtext> Jagosh J, Pluye P, Macaulay AC, et al. Assessing the outcomes of participatory research: protocol for identifying, selecting, appraising and synthesizing the literature for realist review. Implement Sci. 2011 ; 6 (1): 1.</bibtext> </blist> <blist> <bibtext> Khanassov V, Vedel I, Pluye P. Barriers to implementation of case management for patients with dementia: a systematic mixed studies review. Ann Fam Med. 2014 ; 12 (5): 456 ‐ 465.</bibtext> </blist> <blist> <bibtext> Khanassov V, Vedel I, Pluye P. Case management for dementia in primary health care: a systematic mixed studies review. J Clin Interv Aging. 2014 ; 9 : 915 ‐ 928.</bibtext> </blist> <blist> <bibtext> Khanassov V, Vedel I, Pluye P. Dementia in canadian primary health care: the potential role of case management. Health Sci Inquiry. 2014 ; 5 (1): 74 ‐ 76.</bibtext> </blist> <blist> <bibtext> Macaulay AC, Jagosh J, Seller R, et al. Assessing the benefits of participatory research: a rationale for a realist review. Glob Health Promot. 2011 ; 18 (2): 45 ‐ 48.</bibtext> </blist> <blist> <bibtext> Porter MF. An algorithm for suffix stripping. Program. 1980 ; 14 (3): 130 ‐ 137.</bibtext> </blist> <blist> <bibtext> Salton G, Buckley C. Term‐weighting approaches in automatic text retrieval. Inf Process Manag. 1988 ; 24 (5): 513 ‐ 523.</bibtext> </blist> <blist> <bibtext> Lan M, Tan CL, Low HB, Sung SY. A comprehensive comparative study on term weighting schemes for text categorization with support vector machines. In: Special Interest Tracks and Posters of the 14th International Conference on World Wide Web. New York : ACM Press ; 2005 : 1032 ‐ 1033.</bibtext> </blist> <blist> <bibtext> Yang Y, Pedersen JO. A comparative study on feature selection in text categorization. ICML. 1997 ; 97 : 412 ‐ 420.</bibtext> </blist> <blist> <bibtext> O Bodenreider. The unified medical language system what is it and how to use it? Tutorial at Medinfo; 2007.</bibtext> </blist> <blist> <bibtext> Li YH, Jain AK. Classification of text documents. Comput J. 1998 ; 41 (8): 537 ‐ 546.</bibtext> </blist> <blist> <bibtext> Yang Y, Liu X. A re‐examination of text categorization methods. In: Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval; 1999 ; Berkeley, California USA : 42 ‐ 49.</bibtext> </blist> <blist> <bibtext> Aha DW, Kibler D, Albert MK. Instance‐based learning algorithms. Mach Learn. 1991 ; 6 (1): 37 ‐ 66.</bibtext> </blist> <blist> <bibtext> John GH, Langley P. Estimating continuous distributions in Bayesian classifiers. In: Proceedings of the Eleventh Conference on Uncertainty in Artificial Intelligence. San Francisco : Morgan Kaufmann Publishers Inc. ; 1995 : 338 ‐ 345.</bibtext> </blist> <blist> <bibtext> Chang CC, Lin CJ. LIBSVM: a library for support vector machines. ACM Trans Intell Syst Technol. 2011 ; 2 : 1 ‐ 27. Software available at <ulink href="http://www.csie.ntu.edu.tw/cjlin/libsvm">http://www.csie.ntu.edu.tw/cjlin/libsvm</ulink></bibtext> </blist> <blist> <bibtext> Mohan V. Decision trees: a comparison of various algorithms for building Decision Trees ; 2013.</bibtext> </blist> <blist> <bibtext> Quinlan JR. C4. 5: Programs for Machine Learning. San Francisco : Elsevier ; 2014.</bibtext> </blist> <blist> <bibtext> Breiman L. Bagging predictors. Mach Learn. 1996 ; 24 (2): 123 ‐ 140.</bibtext> </blist> <blist> <bibtext> Bauer E, Kohavi R. An empirical comparison of voting classification algorithms: bagging, boosting, and variants. Mach Learn. 1999 ; 36 (1‐2): 105 ‐ 139.</bibtext> </blist> <blist> <bibtext> Freund Y, Schapire RE. Experiments with a new boosting algorithm. ICML. 1996 ; 96 : 148 ‐ 156.</bibtext> </blist> <blist> <bibtext> Zhai C, Lafferty J. A study of smoothing methods for language models applied to information retrieval. ACM Trans Intell Syst Technol. 2004 ; 22 (2): 179 ‐ 214.</bibtext> </blist> <blist> <bibtext> Tao T, Wang X, Mei Q, Zhai C. Language model information retrieval with document expansion. In: Proceedings of the Main Conference on Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics. Morristown, NJ, USA : Association for Computational Linguistics ; 2006 : 407 ‐ 414.</bibtext> </blist> <blist> <bibtext> Shkapenyuk V, Suel T. Design and implementation of a high‐performance distributed web crawler. Data engineering. Proceedings. 18th International Conference on IEEE. San Jose California : IEEE CS Press ; 2002 : 357 ‐ 368.</bibtext> </blist> <blist> <bibtext> Kloda LA. Use Google Scholar, Scopus and Web of Science for comprehensive citation tracking. Evid Based Libr Inf Pract. 2007 ; 2 (3): 87 ‐ 90.</bibtext> </blist> <blist> <bibtext> Tricco AC, Lillie E, Zarin W, et al. A scoping review on the conduct and reporting of scoping reviews. BMC Med Res Methodol. 2016 ; 16 (1): 1.</bibtext> </blist> <blist> <bibtext> Schohn G, Cohn D. Less is more: active learning with support vector machines. In: ICML. Pittsburgh, Pennsylvania USA ; 2000 : 839 ‐ 846.</bibtext> </blist> <blist> <bibtext> Tong S, Koller D. Support vector machine active learning with applications to text classification. J Mach Learn Res. 2001 ; 2 : 45 ‐ 66.</bibtext> </blist> </ref> <aug> <p>By Alexis Langlois; Jian‐Yun Nie; James Thomas; Quan Nha Hong and Pierre Pluye</p> </aug> <nolink nlid="nl1" bibid="bib10" firstref="ref11"></nolink> <nolink nlid="nl2" bibid="bib11" firstref="ref12"></nolink> <nolink nlid="nl3" bibid="bib12" firstref="ref13"></nolink> <nolink nlid="nl4" bibid="bib13" firstref="ref14"></nolink> <nolink nlid="nl5" bibid="bib14" firstref="ref15"></nolink> <nolink nlid="nl6" bibid="bib15" firstref="ref16"></nolink> <nolink nlid="nl7" bibid="bib16" firstref="ref17"></nolink> <nolink nlid="nl8" bibid="bib17" firstref="ref18"></nolink> <nolink nlid="nl9" bibid="bib18" firstref="ref19"></nolink> <nolink nlid="nl10" bibid="bib19" firstref="ref20"></nolink> <nolink nlid="nl11" bibid="bib20" firstref="ref22"></nolink> <nolink nlid="nl12" bibid="bib21" firstref="ref23"></nolink> <nolink nlid="nl13" bibid="bib22" firstref="ref24"></nolink> <nolink nlid="nl14" bibid="bib23" firstref="ref25"></nolink> <nolink nlid="nl15" bibid="bib24" firstref="ref26"></nolink> <nolink nlid="nl16" bibid="bib25" firstref="ref27"></nolink> <nolink nlid="nl17" bibid="bib26" firstref="ref28"></nolink> <nolink nlid="nl18" bibid="bib27" firstref="ref29"></nolink> <nolink nlid="nl19" bibid="bib28" firstref="ref31"></nolink> <nolink nlid="nl20" bibid="bib29" firstref="ref32"></nolink> <nolink nlid="nl21" bibid="bib30" firstref="ref33"></nolink> <nolink nlid="nl22" bibid="bib31" firstref="ref34"></nolink> <nolink nlid="nl23" bibid="bib32" firstref="ref36"></nolink> <nolink nlid="nl24" bibid="bib33" firstref="ref38"></nolink> <nolink nlid="nl25" bibid="bib34" firstref="ref39"></nolink> <nolink nlid="nl26" bibid="bib35" firstref="ref40"></nolink> <nolink nlid="nl27" bibid="bib36" firstref="ref41"></nolink> <nolink nlid="nl28" bibid="bib37" firstref="ref42"></nolink> <nolink nlid="nl29" bibid="bib38" firstref="ref43"></nolink> <nolink nlid="nl30" bibid="bib39" firstref="ref44"></nolink> <nolink nlid="nl31" bibid="bib40" firstref="ref45"></nolink> <nolink nlid="nl32" bibid="bib41" firstref="ref46"></nolink> <nolink nlid="nl33" bibid="bib42" firstref="ref47"></nolink> <nolink nlid="nl34" bibid="bib43" firstref="ref48"></nolink> <nolink nlid="nl35" bibid="bib44" firstref="ref49"></nolink> <nolink nlid="nl36" bibid="bib45" firstref="ref50"></nolink> <nolink nlid="nl37" bibid="bib46" firstref="ref51"></nolink> <nolink nlid="nl38" bibid="bib47" firstref="ref52"></nolink> <nolink nlid="nl39" bibid="bib48" firstref="ref53"></nolink> <nolink nlid="nl40" bibid="bib49" firstref="ref54"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1255688 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Discriminating between Empirical Studies and Nonempirical Works Using Automated Text Classification – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Langlois%2C+Alexis%22">Langlois, Alexis</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-9280-2320">0000-0002-9280-2320</externalLink>)<br /><searchLink fieldCode="AR" term="%22Nie%2C+Jian-Yun%22">Nie, Jian-Yun</searchLink><br /><searchLink fieldCode="AR" term="%22Thomas%2C+James%22">Thomas, James</searchLink><br /><searchLink fieldCode="AR" term="%22Hong%2C+Quan+Nha%22">Hong, Quan Nha</searchLink><br /><searchLink fieldCode="AR" term="%22Pluye%2C+Pierre%22">Pluye, Pierre</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Research+Synthesis+Methods%22"><i>Research Synthesis Methods</i></searchLink>. Dec 2018 9(4):587-601. – Name: Avail Label: Availability Group: Avail Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 15 – Name: DatePubCY Label: Publication Date Group: Date Data: 2018 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Mixed+Methods+Research%22">Mixed Methods Research</searchLink><br /><searchLink fieldCode="DE" term="%22Databases%22">Databases</searchLink><br /><searchLink fieldCode="DE" term="%22Information+Retrieval%22">Information Retrieval</searchLink><br /><searchLink fieldCode="DE" term="%22Search+Strategies%22">Search Strategies</searchLink><br /><searchLink fieldCode="DE" term="%22Documentation%22">Documentation</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Metadata%22">Metadata</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Reports%22">Research Reports</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Vocabulary%22">Vocabulary</searchLink><br /><searchLink fieldCode="DE" term="%22Reference+Materials%22">Reference Materials</searchLink><br /><searchLink fieldCode="DE" term="%22Medical+Research%22">Medical Research</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1002/jrsm.1317 – Name: ISSN Label: ISSN Group: ISSN Data: 1759-2879 – Name: Abstract Label: Abstract Group: Ab Data: Objective: Identify the most performant automated text classification method (eg, algorithm) for differentiating empirical studies from nonempirical works in order to facilitate systematic mixed studies reviews. Methods: The algorithms were trained and validated with 8050 database records, which had previously been manually categorized as empirical or nonempirical. A Boolean mixed filter developed for filtering MEDLINE records (title, abstract, keywords, and full texts) was used as a baseline. The set of features (eg, characteristics from the data) included observable terms and concepts extracted from a metathesaurus. The efficiency of the approaches was measured using sensitivity, precision, specificity, and accuracy. Results: The decision trees algorithm demonstrated the highest performance, surpassing the accuracy of the Boolean mixed filter by 30%. The use of full texts did not result in significant gains compared with title, abstract, keywords, and records. Results also showed that mixing concepts with observable terms can improve the classification. Significance: Screening of records, identified in bibliographic databases, for relevant studies to include in systematic reviews can be accelerated with automated text classification. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2020 – Name: AN Label: Accession Number Group: ID Data: EJ1255688 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1255688 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/jrsm.1317 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 15 StartPage: 587 Subjects: – SubjectFull: Mixed Methods Research Type: general – SubjectFull: Databases Type: general – SubjectFull: Information Retrieval Type: general – SubjectFull: Search Strategies Type: general – SubjectFull: Documentation Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Metadata Type: general – SubjectFull: Comparative Analysis Type: general – SubjectFull: Research Reports Type: general – SubjectFull: Classification Type: general – SubjectFull: Vocabulary Type: general – SubjectFull: Reference Materials Type: general – SubjectFull: Medical Research Type: general Titles: – TitleFull: Discriminating between Empirical Studies and Nonempirical Works Using Automated Text Classification Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Langlois, Alexis – PersonEntity: Name: NameFull: Nie, Jian-Yun – PersonEntity: Name: NameFull: Thomas, James – PersonEntity: Name: NameFull: Hong, Quan Nha – PersonEntity: Name: NameFull: Pluye, Pierre IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 12 Type: published Y: 2018 Identifiers: – Type: issn-print Value: 1759-2879 Numbering: – Type: volume Value: 9 – Type: issue Value: 4 Titles: – TitleFull: Research Synthesis Methods Type: main |
| ResultId | 1 |