Potential Technologies Review: A Hybrid Information Retrieval Framework to Accelerate Demand-Pull Innovation in Biomedical Engineering

Saved in:
Bibliographic Details
Title: Potential Technologies Review: A Hybrid Information Retrieval Framework to Accelerate Demand-Pull Innovation in Biomedical Engineering
Language: English
Authors: Schmitz, Tom (ORCID 0000-0002-5937-765X), Bukowski, Mark (ORCID 0000-0003-4563-1159), Koschmieder, Steffen (ORCID 0000-0002-1011-8171), Schmitz-Rode, Thomas (ORCID 0000-0002-1181-2165), Farkas, Robert (ORCID 0000-0002-8199-6764)
Source: Research Synthesis Methods. Sep 2019 10(3):420-439.
Availability: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
Peer Reviewed: Y
Page Count: 20
Publication Date: 2019
Document Type: Journal Articles
Reports - Descriptive
Descriptors: Innovation, Biomedicine, Medical Research, Meta Analysis, Risk, Guidelines, Sampling, Comparative Analysis, Information Retrieval, Research Reports, Data Analysis, Outcomes of Treatment
DOI: 10.1002/jrsm.1350
ISSN: 1759-2879
Abstract: Launching biomedical innovations based on clinical demands instead of translating basic research findings to practice reduces the risk that the results will not fit the clinical routine. To realize this type of innovation, a meta-analysis of the body of research is necessary to reveal demand-matching concepts. However, both the data deluge and the narrow time constraints for innovation make it impossible to perform such reviews manually. Thus, this paper proposes a specifically adapted "Potential Technologies Review" approach focusing on automated text mining and information retrieval techniques. The novel framework combines features from both systematic and scoping reviews. It aims at high coverage and reproducibility while mapping technologies--even with a fuzzy initial scope. To achieve these goals for search and triage, a set of closely interrelated methods has been developed: (a) automated query optimization, (b) screening prioritization, and (c) recall estimation. To determine appropriate parameters, a variety of published literature corpora were used and compared with an evaluation on a real-world dataset. Our results show that it is feasible to automate the identification of relevant works using this newly introduced framework. It achieved a workload reduction of up to 91% "Work-saved-over Sampling (WSS)" with a 76% overall recall compared with manually screening search results. Reducing the workload is a prerequisite for a rapid Potential Technologies Review when conducting demand-pull innovations. Moreover, it facilitates the updating and closer monitoring of latest findings. Studying the robustness of the framework and expanding it to patent documents are future tasks.
Abstractor: As Provided
Entry Date: 2020
Accession Number: EJ1255363
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFVlNAxhKW1EP7Yg5WLEyJYAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDEMgyOAyC_sD4atLpQIBEICBmwxb-g1S1zoNhp5EJyPahTe_VKE1Z-CJfebru2AcXosbY2dfj_URWSuXcld1Rkbt47zWR7h7F6AN0CrNTRuMvMkM459ZFIk89zbVaXsrURtI5PLzb8BCikbEq9dN-PtsbTKQpGlcz-IYmqpyyyd1fS-Fqz_RZMUQWVNBOuHFOtXda8nlb8laT8jejjwDluA1rd3nzS6pkqeOgd7k
Text:
  Availability: 1
  Value: <anid>AN0138519222;[bdct]01sep.19;2019Sep11.02:56;v2.2.500</anid> <title id="AN0138519222-1">Potential Technologies Review: A hybrid information retrieval framework to accelerate demand‐pull innovation in biomedical engineering </title> <p>Launching biomedical innovations based on clinical demands instead of translating basic research findings to practice reduces the risk that the results will not fit the clinical routine. To realize this type of innovation, a meta‐analysis of the body of research is necessary to reveal demand‐matching concepts. However, both the data deluge and the narrow time constraints for innovation make it impossible to perform such reviews manually. Thus, this paper proposes a specifically adapted "Potential Technologies Review" approach focusing on automated text mining and information retrieval techniques. The novel framework combines features from both systematic and scoping reviews. It aims at high coverage and reproducibility while mapping technologies—even with a fuzzy initial scope. To achieve these goals for search and triage, a set of closely interrelated methods has been developed: (a) automated query optimization, (b) screening prioritization, and (c) recall estimation. To determine appropriate parameters, a variety of published literature corpora were used and compared with an evaluation on a real‐world dataset. Our results show that it is feasible to automate the identification of relevant works using this newly introduced framework. It achieved a workload reduction of up to 91% "Work‐saved‐over Sampling (WSS)" with a 76% overall recall compared with manually screening search results. Reducing the workload is a prerequisite for a rapid Potential Technologies Review when conducting demand‐pull innovations. Moreover, it facilitates the updating and closer monitoring of latest findings. Studying the robustness of the framework and expanding it to patent documents are future tasks.</p> <p>Keywords: cross‐validation; information retrieval; text mining; vector space model; WSS score</p> <hd id="AN0138519222-2">Highlights</hd> <p></p> <ulist> <item> The idea presented in this work is to systematically and efficiently map existing knowledge by automated information retrieval, which then can be used to support the development of an innovative biomedical device matching a given clinical demand. In this way, the innovation process can eventually be accelerated.</item> <p></p> <item> The proposed text mining system aim to fit specifically Potential Technologies Review, which combines features of scoping and—to a smaller extent—systematic review methods into a novel hybrid approach. The developed text mining technologies improve search quality and reduce the manual screening workload. To our knowledge, this is the first approach based on a (systematic) literature review that is applicable during an innovation process because it can be performed rapidly and can easily be updated during translational research and subsequent R&D.</item> <p></p> <item> While this work was influenced by prior research on systematic and scoping reviews and associated text mining methods, our findings may also contribute to research on these topics. In particular, the novel automatic query optimization algorithm addresses the challenge of performing (scoping) reviews from an initially fuzzy scope.</item> </ulist> <hd id="AN0138519222-3">BACKGROUND</hd> <p></p> <hd id="AN0138519222-4">Review as a means for translational research</hd> <p>In biomedical engineering, the innovation process is often understood as translating findings from "bench to bedside."[<reflink idref="bib1" id="ref1">1</reflink>] It starts from a certain finding, leads to a technological innovation, and is finally applied to one or more clinical tasks. This so‐called <emph>technology‐push</emph> approach can be summarized as "new technology seeks clinical application(s)."</p> <p>In contrast, to shorten development times and improve clinical applicability, the <emph>demand‐pull</emph> approach focuses on clinical demand instead of technology as a stimulus for innovation.[[<reflink idref="bib2" id="ref2">2</reflink>]] This can be summarized as "clinical demand seeks appropriate technological solution(s)." Starting the innovation process in the clinical reality, the demand‐driven approach is open to any technology capable of contributing to the solution of the clinical problem in question.</p> <p>The challenge now, however, is to find suitable technological solutions for a given clinical demand. The typical and most common approach is to ask affiliated experts. Although this approach may be successful, appropriate experts might not easily be available. In contrast, scientific and/or technological publications contain the desired information as well. They capture knowledge at a global scope and are readily accessible.</p> <p>Thus, one obvious solution is to perform a conventional literature review to extract the existing knowledge. But the results will largely depend on the reviewer and the selected search technique(s).[<reflink idref="bib5" id="ref3">5</reflink>] To overcome this risk of bias, a systematic review of literature should rather be carried out. This well‐standardized meta‐analysis addresses the problem of limited reproducibility and coverage in conventional literature reviews[[<reflink idref="bib6" id="ref4">6</reflink>]] and represents an acknowledged method for providing evidence‐based answers to clinical questions.</p> <p>However, the method shows essentially three drawbacks regarding the purpose of identifying technological concepts: First, it aggregates evidence rather than mapping research activity. Second, it requires clearly defined search boundaries. Third, it aims at a 100% recall knowing that the necessary processing time increases remarkably.</p> <p>Considering the first two issues leads to the so‐called "scoping review,"[<reflink idref="bib9" id="ref5">9</reflink>] an alternative to a systematic review, which is often conducted in advance of a systematic review to provide an overview of the topic. Compared with systematic reviews, a scoping review aims to explore the evidence and to refine the boundaries of the review question.[[<reflink idref="bib10" id="ref6">10</reflink>]] This objective fits well to the demand driven search for technological solutions, which is by its nature initially fuzzy.</p> <p>The strengths of scoping reviews seem to make it the first choice to master the task of identifying relevant technological knowledge. However, due to its iterative design it features severe drawbacks such as offering limited reproducibility that only results in coverage information and yielding even larger amounts of datasets to be screened than a systematic review.</p> <p>Hence, facing the large datasets the 100% recall requirement affects both procedures, by enlarging the workload of the reviewer and the timeline to finish. While the aim of aggregating all relevant information (100% recall) concerning a certain topic may be justified when researching a sound basis of evidence for important clinical questions or guidelines, it is overwhelming for cases involving fast innovation processes such as in biomedical engineering. The workload associated with a systematic review is very large and caused primarily by the process of identifying relevant publications.[[<reflink idref="bib12" id="ref7">12</reflink>]]</p> <p>For example, in one of the most cited systematic reviews, approximately 6000 records were screened, but only 495 of those were eventually incorporated into the review.[<reflink idref="bib15" id="ref8">15</reflink>] In other cases, the inclusion rate drops to even lower values, such as 51/1333 for the "Proton Pump Inhibitor" study[<reflink idref="bib16" id="ref9">16</reflink>] or 765/48 638 for the "Transgenerational Inheritance of Health Effects" corpus.[<reflink idref="bib17" id="ref10">17</reflink>] Although for scoping reviews the 100% recall requirement is usually relaxed, the above ratio often reaches comparable levels[<reflink idref="bib18" id="ref11">18</reflink>]: 169/2214 or 32/607 for a review on secondary findings after sequencing the whole genome.[<reflink idref="bib19" id="ref12">19</reflink>]</p> <p>Overall, the scoping review seems to suit better than the systematic review to identify technological solutions from literature. However, the workload associated with the review must be reduced, because time is crucial in translational research.[[<reflink idref="bib20" id="ref13">20</reflink>]]</p> <hd id="AN0138519222-5">Text mining accelerating the review process</hd> <p>The last decade in particular has seen considerable research efforts to accelerate the review procedure or parts thereof through automation. Various information retrieval technologies are implemented to facilitate, to limit, or to take over the manual work of the reviewer during the search and especially the screening stage.</p> <p>In a pioneering study, Cohen et al[<reflink idref="bib16" id="ref14">16</reflink>] examine the performance of a machine learning based classifier to select relevant citations when updating 15 different manually conduced systematic drug reviews. Titles and abstracts, medical subject headings (MeSH), and the publication type serve as feature base to feed the classifier, which performs basically well in 5 × 2 cross‐validation indicated by recall, precision, and F1‐measure. Further research aims to improve the performance by using different classifiers such as support vector machine (SVM)[[<reflink idref="bib23" id="ref15">23</reflink>]] or naïve Bayes,[<reflink idref="bib25" id="ref16">25</reflink>] by enlarging the feature space (using, eg, UMLS) and/or through integrating user feed‐back as active learning.[[<reflink idref="bib12" id="ref17">12</reflink>], [<reflink idref="bib26" id="ref18">26</reflink>]] Apart from improving performance, active learning shall also overcome the main drawback of conventional automated classification (AC), namely, the need for a sufficient amount of previously labeled training data, which limits the applicability to updating reviews or replacing the required second reviewer.</p> <p>O'Mara‐Eves et al[<reflink idref="bib27" id="ref19">27</reflink>] overviewed the advancements of automating the screening stage in systematic reviews. They state that in contrast to other procedures AC strives to reduce the number of citations for manual screening saving time and resources. It was again Cohen,[<reflink idref="bib16" id="ref20">16</reflink>] who was the first to introduce the Work‐saved‐over Sampling (WSS) measure indicating the proportion of citations that the reviewer does not have to assess manually in the triage process. Although some authors use different measures such as "(screening) burden,"[[<reflink idref="bib12" id="ref21">12</reflink>], [<reflink idref="bib28" id="ref22">28</reflink>]] WSS is probably the most widely used and can therefore be regarded as a benchmark for comparison. This is particularly supported by Cohens step to make his test collections publicly available.[<reflink idref="bib16" id="ref23">16</reflink>]</p> <p>Different from AC, a second group of procedures focuses on prioritizing the search results for screening.[<reflink idref="bib27" id="ref24">27</reflink>] The relevance ranking techniques are considered to be more flexible and effective than classification.[<reflink idref="bib29" id="ref25">29</reflink>] In a recent study, Howard et al[<reflink idref="bib17" id="ref26">17</reflink>] remarkably demonstrated that prioritization clearly outperforms automatic classification by doubling the WSS at 95% recall, averaged over all 15 datasets (48.8%) compared with the results from Cohen,[<reflink idref="bib16" id="ref27">16</reflink>] which are achieved with AC.</p> <p>In addition, prioritization may enable an improved workflow design, eg, by assigning the citations to different reviewers according to their proficiency[<reflink idref="bib30" id="ref28">30</reflink>] or starting the synthesis stage even before the screening has been fully completed. Finally, using a cut‐off threshold to terminate the screening would lower the number of articles for manual screening, but at the same time, it will increase the risk of bias by impairing the recall, known as the "hasty generalization" problem.</p> <p>This risk is becoming even more important when trying to narrow the search before screening and not to admit so many irrelevant citations. On the other hand, the potential gain might be remarkably large considering that usually less than a tenth of the search results are finally included in the review. Consequently, automatic term recognition (ATR)[<reflink idref="bib31" id="ref29">31</reflink>] is conducted to identify concepts of a corpus represented by terms in a weighted order either to refine the search strategy or to prioritize the screening. In addition, user feedback can be implemented to control the term extraction by an intuitive interface[<reflink idref="bib32" id="ref30">32</reflink>] or by providing an initial corpus of a small number of manually assessed relevant citations.[<reflink idref="bib33" id="ref31">33</reflink>]</p> <p>Although many text mining technologies are analyzed to accelerate review, almost all studies focus explicitly on systematic rather than scoping reviews. Among the few exceptions, eg, Stansfield et al work on clustering topics following the screening phase,[<reflink idref="bib34" id="ref32">34</reflink>] the study of Shemilt et al[<reflink idref="bib10" id="ref33">10</reflink>] clearly stands out for many reasons. In order to reduce the manual workload of screening more than 800 000 citations per scoping review to a feasible level, ATR, AC and the so‐called Reviewer Terms (RTs) are combined to form a complex, multistage procedure. In this approach, the RT, ie, reviewer dependent lists of indicative terms to predict inclusion or exclusion, represents an "active learning" contribution and are complementing the true automation tools (AC, ATR).</p> <p>Although the implementation achieves remarkable workload reductions of up to 90%, the authors admit that this is due to the low recall (38%‐85%). Moreover, to evaluate the performance and to train the classifier, a large set of 46 000 records per review is manually screened in advance. Thus, they conclude that text mining offers a lot of potential for scoping reviews, but substantial efforts are required to develop a convincing method from the current beginnings.</p> <hd id="AN0138519222-6">Objectives</hd> <p>Considering the aim of systematically identifying technological concepts fitting to a clinical demand from literature as an onset for innovation, no appropriate procedure is currently available. The defining characteristics of systematic reviews (100% recall, evidence, timeline) are not applicable to this task, but there is a substantial body of research on accelerating this established procedure while maintaining the named properties. On the other hand, the characteristics of scoping reviews fit much better to technology identification, but here, the research efforts on text mining based automation are still in their infancy. Even so‐called evidence maps cannot close the existing gap, as they neither aim to identify technologies[<reflink idref="bib35" id="ref34">35</reflink>] nor the application of information retrieval for this type of reviews has been investigated so far.[<reflink idref="bib36" id="ref35">36</reflink>]</p> <p>Thus, we propose a novel review approach that incorporates features of both systematic and scoping review techniques, which we call "Potential Technologies Review." Because the Potential Technologies Review process is specifically designed for biomedical <emph>demand‐pull</emph> innovations and their narrow timetables, we build a novel framework from text mining and information retrieval methods to reduce the workload and accelerate the review process: the query is automatically optimized on the basis of an initial corpus of relevant literature, the records in the screening process are prioritized by relevance ranking, and the current recall is monitored on the basis of cross‐validation.</p> <p>The remainder of this paper is structured as follows: In Section , we describe the process of a Potential Technologies Review, and we begin the explanation of the methods in Section with the PubMed‐based automatic query optimization and prioritization process. This is followed by the evaluation measures and procedures. The results chapter opens with the parameter selection for automatic query optimization (see Section) and continues to the prioritization performance, which depends on the implemented ranking method in conjunction with the optimized query. It is evaluated using existing data of systematic and scoping reviews. Finally, the feasibility of the Potential Technologies Review approach is demonstrated using real‐world data in Section. Discussion and conclusion will finalize the paper.</p> <hd id="AN0138519222-7">POTENTIAL TECHNOLOGIES REVIEW (PTR) OUTLINE</hd> <p></p> <hd id="AN0138519222-8">Defining features</hd> <p>Apart from the content‐related orientation, the critical features of a Potential Technologies Review are coverage, a narrow timeline and high reproducibility. Besides comprehensive information on existing research for a given topic, reviewer‐independent results must be achieved, and therefore, an appropriate design is required. As far as possible, the design should adhere to the two established systematic review procedures, which support the goals and requirements differently. Table  shows a comparison among the features of a systematic review, a scoping review, and the proposed Potential Technologies Review.</p> <p>Comparison of defining features for different types of review positioning the Potential Technologies Review between systematic and scoping review</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Feature</th><th align="left">Systematic review</th><th align="left">Potential Technologies Review</th><th align="left">Scoping review</th></tr></thead><tbody valign="top"><tr><td>Goal</td><td>Aggregate evidence</td><td>Explore promising technological concepts</td><td>Map/explore evidence</td></tr><tr><td>Boundaries</td><td>Clearly defined at the outset</td><td>Fuzzy at the outset, refined during review</td><td>Unclear at the outset; refined during review</td></tr><tr><td>Coverage</td><td>Every eligible study</td><td>High recall, but small processing time</td><td>Representative/key studies</td></tr><tr><td>Reproducibility</td><td>High (linear design)</td><td>High (despite iterative design)</td><td>Limited (iterative design)</td></tr><tr><td>Topic</td><td>Clinical</td><td>Technological</td><td>Clinical</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>. The features of aim, boundaries, and coverage are described by Shemilt et al.[<reflink idref="bib10" id="ref36">10</reflink>] The iterative nature of scoping reviews is lined out by Arksey and O'Malley.[<reflink idref="bib9" id="ref37">9</reflink>]</p> <p>Potential Technologies Review is placed right between systematic and scoping reviews. The "Boundaries" and "Goal" feature show a clear tendency towards scoping reviews because of the initially fuzzy relevance criteria to explore evidence. In contrast, high "Reproducibility" and "Coverage" of the proposed method are closer to systematic review, even if the 100% recall is not required.</p> <hd id="AN0138519222-9">Text mining‐based procedure</hd> <p>Considering the complete review process (see Figure ), the text mining‐based procedure affects only the first two stages of a systematic review, namely, "Framing questions for a review" and "Identifying relevant works."[<reflink idref="bib37" id="ref38">37</reflink>] The subsequent steps are appropriate only for clinical studies. So, in a "Potential Technologies Review," those steps are replaced.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0001.jpg" title="1 Comparison of full procedures for systematic review vs Potential Technologies Review. The five steps of a systematic review a given according to Khan et al[37] [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>In this paper, however, we will focus exclusively on the first two steps to finally identify relevant literature. Different from SR, the PTR starts with a preliminary search. It aims to support the question framing stage.[[<reflink idref="bib10" id="ref39">10</reflink>], [<reflink idref="bib34" id="ref40">34</reflink>]] In some extended applications, the comprehensive exploratory search is conducted as scoping review prior to a systematic review. However, in the Potential Technologies Review, the preliminary search phase only takes up the prior knowledge of the reviewer and external experts to quickly create a basic set of definitely relevant citations. This initial corpus forms the starting material for the subsequent automation steps (see Figure ).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0002.jpg" title="2 Process overview of the first two stages (question framing and search, screening) of the Potential Technologies Review (PTR). Manual tasks on the left side (rounded boxes), automated steps on the right (rectangles) [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>In the preliminary search phase, the initial corpus is manually compiled on the basis of the clinical demand using different search techniques such as unsystematic keyword searches in databases, citation tracking ("Snowballing") or through expert knowledge. The relevance criteria and search strategy are iteratively refined as knowledge on the topic of interest grows during this phase. Such search techniques offer limited reproducibility but are known to be efficient.[<reflink idref="bib14" id="ref41">14</reflink>] The subsequent automated expansion of the search by reusing the initial corpus is necessary to achieve the required recall and reproducibility. The prioritizing combined with a recall threshold reduces the necessary workload for a closing manual screening, which results in a final corpus of included studies. These citations form the basis for the remaining synthetizing steps of a Potential Technologies Review, as indicated in Figure .</p> <p>The boundaries of relevance are fuzzy at the beginning of a Potential Technologies Review. Therefore, formulating a database search query may be a difficult task. Existing research supporting query formulation relies heavily on expert opinion, which adds additional biases and may be time‐consuming and costly.[<reflink idref="bib38" id="ref42">38</reflink>] Thus, an automated query optimization process to formulate a search statement by analyzing the initial corpus will be introduced as a first step.</p> <p>Having applied the query, the next step involves distinguishing the most relevant findings from less relevant ones because the goal of a Potential Technologies Review is to explore existing knowledge rather than to aggregate research findings. Therefore, relevance ranking is used for prioritization during this screening process by applying information retrieval technologies.</p> <p>Finally, a termination point at which the screening process should stop must be determined to achieve the necessary reduction of the screening workload. This step involves monitoring the search performance by estimating the recall according to the initial corpus and fixing a suitable threshold. Shemilt's approach provides a feasible recall estimation, but it involves substantial additional workload.[<reflink idref="bib10" id="ref43">10</reflink>] Thus, we present a novel method that limits the workload through a system of strongly related automation steps.</p> <hd id="AN0138519222-12">METHODS</hd> <p></p> <hd id="AN0138519222-13">Automatic query optimization</hd> <p>As a starting point for the query optimization, a text corpus is required containing essential technological concepts. We assume that this initial corpus contains most of the concepts that are relevant for a certain research question.</p> <p>By recognizing terms representing the content of the corpus, we are following the <emph>bag‐of‐words</emph> approach to assemble, evaluate, and refine appropriate search queries, which are applied on PubMed database. More precisely, a set of titles and abstracts of relevant publications is the input into the algorithm, which outputs a set of suitable queries, along with their recall and precision scores based on the initial corpus.</p> <hd id="AN0138519222-14">Algorithm outline</hd> <p>The query optimization algorithm is depicted in Figure .</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0003.jpg" title="3 Query optimization algorithm processing fixed set of initial corpus, keyword‐count and AND‐combinations. It provides a list of Boolean queries with the corresponding recall of the initial corpus" /> </p> <p></p> <p>In addition to the initial corpus, the following two parameters must be provided to control the procedure.</p> <p></p> <ulist> <item> <emph>keyword‐count</emph> (kc) marks the absolute number of keywords selected to represent the initial corpus. When the keyword count parameter is small, only the most common keywords from the corpus are selected. When the keyword count parameter is larger, the algorithm also selects keywords that represent only a subgroup of records.</item> <p></p> <item> <emph>AND‐combinations</emph> (ac) controls the maximum number of keywords selected that are combined using Boolean AND. Assigning larger values for this parameter cause the resulting query to be more specific.</item> </ulist> <p>The algorithm consists of five steps:</p> <p></p> <ulist> <item> SELECT_TERMS: Terms in the initial corpus are analyzed for their document frequency (df) and sorted in descending df order. Other techniques for term recognition such as <emph>Termine</emph>[<reflink idref="bib39" id="ref44">39</reflink>] or <emph>Rapid Automatic Keyword Extraction</emph> (RAKE)[<reflink idref="bib40" id="ref45">40</reflink>] exist, but they focus on selecting terms per document. In contrast, the df analyzes the term frequency for a corpus.[<reflink idref="bib41" id="ref46">41</reflink>] The algorithm selects a number of keywords equal to the <emph>keyword‐count</emph> parameter from the top of the resulting ranked list. If the rank of the keyword with the lowest rank appears multiple times, the algorithm selects all keywords with the same rank.</item> <p></p> <item> CREATE_AND_QUERIES: The selected terms are combined using a Boolean AND operator, and the maximum number of combined terms is controlled by the <emph>AND‐combinations</emph> parameter. When the <emph>AND‐combinations</emph> parameter is set to 2, a maximum of two terms are combined by AND (eg, term A AND term B). Variants containing less terms are also considered (eg, term A). Obviously, a setting of <emph>AND‐combinations</emph> = 1 means that no AND is used.</item> <p></p> <item> PUBMED_EVALUATION: The queries are executed in PubMed and the F1‐measure is calculated on the basis of the initial corpus (see Section). The F1‐measure is particularly suitable for ranking appropriate queries forming a compromise between recall and precision. Subsequently, the gain is calculated for each query. Gain is defined as the number of publications identified from the initial corpus compared with previous AND combinations that rank higher in the query list based on the F1‐measure.</item> <p></p> <item> CREATE_OR_QUERIES: Because finding a query using only a single AND combination of keywords that achieves a high recall score is unlikely, it is typically necessary to assemble cumulative queries by combining high‐performing AND combinations with a Boolean OR. High‐performing queries are those AND combinations that attain both a high F1‐measure and a gain greater than 0. Requiring a gain greater than 0 ensures that no redundant queries are selected.</item> <p></p> <item> PUBMED_EVALUATION: Those cumulative queries are tested on PubMed and recall, precision and F1‐measure are again calculated according to the initial corpus. The cumulative queries are returned in ascending order of recall.</item> </ulist> <p>The algorithm was implemented in Python as follows. The title and abstract of each publication in the initial corpus are concatenated. The concatenated string is preprocessed as follows: (a) stop word removal by using an English stop word list,[<reflink idref="bib42" id="ref47">42</reflink>] (b) tokenizing, and (c) stemming by using the English <emph>Snowball</emph> stemmer from the <emph>Natural Language Toolkit</emph> (NLTK, <ulink href="http://www.nltk.org/),">http://www.nltk.org/),</ulink> which is a revised version of the <emph>Porter stemmer</emph>.[<reflink idref="bib43" id="ref48">43</reflink>] N‐grams are not used in this analysis.</p> <p>The document frequency (the number of documents in which a certain word appears) is calculated by using the vectorizer <emph>CountVectorizer</emph> from the <emph>scikit‐learn</emph> project.[<reflink idref="bib44" id="ref49">44</reflink>]</p> <p>To create the term combinations, the <emph>combinations</emph> function from the Python <emph>itertools</emph> library is used. The resulting queries are evaluated against PubMed using the <emph>E‐Utilities ESearch</emph> (https://<ulink href="http://www.ncbi.nlm.nih.gov/books/NBK25499">www.ncbi.nlm.nih.gov/books/NBK25499</ulink>) interface.</p> <hd id="AN0138519222-16">Parameter optimization</hd> <p>The performance of the automatic query optimization depends on the values of both the <emph>keyword‐count</emph> (<emph>kc</emph>) and <emph>AND‐combination</emph> (<emph>ac</emph>) parameter, which have to be set at procedure start.</p> <p>To identify the optimal parameter values, a grid search with <emph>kc</emph> ∈ (<reflink idref="bib5" id="ref50">5</reflink>,<reflink idref="bib10" id="ref51">10</reflink>,<reflink idref="bib15" id="ref52">15</reflink>), <emph>ac</emph> ∈ (<reflink idref="bib1" id="ref53">1</reflink>,<reflink idref="bib2" id="ref54">2</reflink>,<reflink idref="bib3" id="ref55">3</reflink>) using 5 × 2 cross‐validation was conducted to estimate the generalization performance. Consequently, the target value is given by the difference in testing recall on a random versus a systematic corpus. The rationale behind this is that an optimal query construction algorithm should be able to assemble a query with a high generalization performance for a corpus with a common underlying concept but not for a random corpus, because, in theory, it is highly unlikely that two subsets of a random dataset would have common underlying concepts.</p> <hd id="AN0138519222-17">Screening prioritization</hd> <p>To calculate relevance ranking of query results, the vector space model (VSM) with cosine similarity[<reflink idref="bib45" id="ref56">45</reflink>] is widely used.</p> <p>A search statement of a query (q) and title/abstract of a retrieved document (d) are understood as vectors of terms according to the <emph>bag‐of‐words</emph> concept. The similarity between q und d is given by the angle between the above vectors. Small angles indicate a close similarity and therefore a top position in the corresponding ranking.</p> <p>In this study, we used the implementation from the <emph>Lucene</emph>‐based <emph>Apache Solr</emph> (https://lucene.apache.org/solr/) platform.</p> <p>Thus, the prioritization is performed by using the automatically generated search statement as input for the VSM‐based ranking of the retrieved documents. No additional training data are necessary.</p> <hd id="AN0138519222-18">Evaluation</hd> <p></p> <hd id="AN0138519222-19">Datasets</hd> <p>In research on text mining in systematic reviews, publicly available corpora from existing systematic reviews[<reflink idref="bib16" id="ref57">16</reflink>] are regularly used to evaluate the performance. Howard et al summarized the available corpora.[<reflink idref="bib17" id="ref58">17</reflink>] For this study, the following datasets (see Table ) cover systematic corpora, scoping review corpora and our real‐world "Innovative Laboratory Diagnostics" (ILD) corpus. They have been chosen because of their varying properties considering size, topic, include ratio, review type, and prior results of possible workload reduction. In contrast to the systematic corpora, the datasets of the scoping review corpora are not readily available. Therefore, we reconstructed these datasets (see Supporting Information).</p> <p>Review corpora used for performance evaluation</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Dataset</th><th align="left" /><th align="left">Records (PubMed)</th><th align="left">Review</th></tr><tr><th align="left" /><th align="left">From search</th><th align="left">Included</th><th align="left" /></tr></thead><tbody valign="top"><tr><td>Transgenerational Inheritance of Health Effects<xref ref-type="bibr" rid="bibr17">17</xref></td><td>TIHE</td><td>48 638</td><td>765</td><td>(1.6%)</td><td>Systematic</td></tr><tr><td>Attention Deficit Hyperactivity Disorder<xref ref-type="bibr" rid="bibr16">16</xref></td><td>ADHD</td><td>851</td><td>20</td><td>(2.4%)</td><td>Systematic</td></tr><tr><td>Atypical Antipsychotics<xref ref-type="bibr" rid="bibr16">16</xref></td><td>AA</td><td>1120</td><td>146</td><td>(13.0%)</td><td>Systematic</td></tr><tr><td>Oral Hypoglycemics<xref ref-type="bibr" rid="bibr16">16</xref></td><td>OH</td><td>503</td><td>136</td><td>(27.0%)</td><td>Systematic</td></tr><tr><td>Aquatic Cycling<xref ref-type="bibr" rid="bibr46">46</xref></td><td>ACY</td><td>128</td><td>20</td><td>(15.6%)</td><td>Scoping</td></tr><tr><td>Patient Age in Cancer Care<xref ref-type="bibr" rid="bibr47">47</xref></td><td>PACC</td><td>844</td><td>298</td><td>(35.3%)</td><td>Scoping</td></tr><tr><td>Innovative Laboratory Diagnostics</td><td>ILD</td><td>Not fixed<xref ref-type="fn" rid="tfn2" /></td><td>60</td><td>‐</td><td>Potential technologies</td></tr></tbody></table> </ephtml> </p> <p>2 Different from the literature‐based datasets, the ILD corpus is derived from a preliminary search in the initial phase of the review (real‐world data).</p> <p>As the goal of this paper is to analyze the feasibility of a Potential Technologies Review, a real‐world dataset has been elaborated from a case study. The clinical demand is specified by the Department of Hematology, Oncology, Hemostaseology, and Stem Cell Transplant at RWTH Aachen University. The corresponding project is referred to as "Innovative Laboratory Diagnostics" (ILD). Similar to Shemilt et al,[<reflink idref="bib10" id="ref59">10</reflink>] we iteratively assemble an initial corpus of relevant citations by performing a preliminary search based on expert knowledge of the above clinicians and biomedical engineers. Provisional eligibility criteria are used to filter the search results, before the manual screening was performed. This leads to the initial ILD corpus with finally 60 relevant PubMed records, which are used as input to the automatic query optimization procedure (see Section).</p> <p>Finally, a random corpus of PubMed records (RPM) is used as a baseline for the performance assessment. Sixty records are assembled by querying random PubMed IDs between 20 000 000 and 28 000 000. The 60‐record size is chosen because it equals the size of the ILD real‐world corpus used in the feasibility section of this paper.</p> <hd id="AN0138519222-20">Performance measures</hd> <p>The comparison between the expected extent of relevant results (derived from previous knowledge) and the actual amount of relevant information found through text mining constitutes the basis for three fundamental key indicators that are also common in systematic/scoping reviews:</p> <p>Recall (r) is the fraction of retrieved relevant information over the total amount of relevant information; Precision (p) is the fraction of relevant information in all retrieved information, and F1‐measure (F1) is the harmonic mean of the aforementioned recall and precision. Further details on these basic measures are provided, eg, by O'Mara‐Eves et al[<reflink idref="bib27" id="ref60">27</reflink>] or Sammut and Webb.[<reflink idref="bib48" id="ref61">48</reflink>]</p> <p>In a multistage text mining procedure, where a query optimization is followed by prioritization, the second step is applied on the result set of the initial step, with already reduced recall. Thus, the <emph>overall recall</emph>, which refers to the whole procedure, is given by the product of the proportional recall values (r<subs>s</subs>) of all consecutive process steps (s). An <emph>overall test recall</emph> refers to the <emph>overall recall</emph> based on the test set.</p> <p> <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1350:jrsm1350-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><mtext mathvariant="italic">Overall Recall</mtext><mo>=</mo><munderover><mo>∏</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>s</mi></munderover><msub><mi mathvariant="normal">r</mi><mi>i</mi></msub></math> </ephtml> </p> <p>To evaluate the performance of a text mining implementation the WSS score was introduced by Cohen et al.[<reflink idref="bib16" id="ref62">16</reflink>] The WSS score reflects the percentage of the workload saved compared with randomly sampling the dataset. In other words, the WSS describes the proportion of studies that a reviewer no longer needs to read to achieve a given recall. The WSS at 95% recall is defined as follows:</p> <p> <ephtml> <math display="block" overflow="scroll" altimg="urn:x-wiley:17592879:media:jrsm1350:jrsm1350-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">WSS</mi><mo>@</mo><mn>95</mn><mo>=</mo><mfrac><mrow><mi mathvariant="italic">TN</mi><mo>+</mo><mi mathvariant="italic">FN</mi></mrow><mi>N</mi></mfrac><mo>−</mo><mn>0.05</mn><mo>,</mo></math> </ephtml> </p> <p>where TN and FN denotes the numbers of true negatives and false negatives identified according to the screening decision, respectively, and N is the total number of records in the dataset. A linear interpolation is performed using the closest values lower and higher than the recall of interest. This procedure returns comparable scores at certain levels (such as 95%) and is therefore capable to assess the performance of a simple prioritization or a combined text mining framework consisting of prioritization and query atomization, as illustrated in Figure .</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0004.jpg" title="4 Illustrated derivation of the work‐saved‐over‐sampling (WSS) value under different text mining conditions at 95% recall. Two examples of ranked corpora are displayed: On the right side, the WSS@95q refers to the 95% recall of the reduced results based on the query with a query depended recall reduction. The recall reduction is assumed to be of the same value in both examples. The recall of WSS@95q corresponds to the overall recall of 80% [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <hd id="AN0138519222-22">5 × 2 cross‐validation</hd> <p>The applied procedure is a 5 × 2 cross‐validation. Cross‐validation is normally used to evaluate classifiers; however, it has been applied in a related publication on automatic classification in systematic reviews, and "[it] is thought to give better estimates of actual performance than the ten‐way cross‐validation [...]."[<reflink idref="bib16" id="ref63">16</reflink>] In 5 × 2 cross‐validation, the dataset is randomly split into two samples of equal size for which two cross‐validations are performed: the first half is used to build the query (<emph>training set</emph>) and the second half is used to evaluate the performance with unseen data (<emph>test set</emph>). Then, the second half is used to build the query, and the first half is used to evaluate the performance. This process is repeated five times, and the average of the results is calculated.</p> <hd id="AN0138519222-23">Test procedure</hd> <p>Query optimization and prioritization are evaluated on literature‐based corpora (see Table ) applying 5 × 2 cross‐validation. The query optimization algorithm uses half of the included studies (the training set) with constant parameter values (<emph>keyword‐count</emph>, <emph>AND‐combinations</emph>) and is executed on the complete PubMed database. Then, the query for a <emph>training recall</emph> level of 90% based on the training set is selected and applied to the test set, which contains all the excluded studies and the other half of the included studies.</p> <p>The results are ranked by relevance score according to the query. Analogous to the procedure used in previous publications on automatic classification (see Howard et al[<reflink idref="bib17" id="ref64">17</reflink>]), a cutoff relevance score is selected such that the corpus of studies ("potentially relevant" studies) with relevance scores above the cutoff contains 95% of the studies included in the test set that appear in the result list. This enables the calculation of WSS@95.</p> <p>As illustrated in Figure , the evaluation procedure for automatic classification differs from evaluating the combination of automatic query optimization and relevance ranking: (a) When evaluating automatic query optimization only, the included studies are used, and the training is performed on the complete PubMed dataset, while automatic classification needs both positive and negative training data (included and excluded studies). (b) The <emph>overall recall</emph> for the test set achieved by the query built by automatic query optimization limits the maximum achievable recall in the relevance ranking step. Consequently, a recall of 95% (<emph>test recall</emph>) in the relevance ranking step corresponds eventually to a lower <emph>overall test recall</emph>, based on the entire <emph>test set</emph> (see Figure ).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0005.jpg" title="5 Applied procedure for evaluating the prioritization performance of the automatic query optimization and relevance ranking (right side) procedures in comparison with the procedures used in existing work on automatic classification (left side)" /> </p> <p></p> <p>Due to these differences, the comparability of the results on automatic classification and the combination of automatic query optimization and relevance ranking is limited. Nevertheless, this approach establishes the best possible connection to existing work.</p> <hd id="AN0138519222-25">RESULTS</hd> <p></p> <hd id="AN0138519222-26">Optimized queries</hd> <p></p> <hd id="AN0138519222-27">Choosing a parameter setting</hd> <p>The greatest difference between the Random PubMed sample and the systematic review corpora AA, OH, and TIHE was achieved using a parameter set of <emph>kc</emph> = 15 and <emph>ac</emph> = 3 considering the test set (see Table ).</p> <p>Effect of the parameter variation of keyword‐count (kc) and AND‐combinations (ac) on the mean recall for training (R Tr) and test (R Te) set over 5 × 2 cross‐validation for the different corpora</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">kc, ac</th><th align="left">Random</th><th align="left">AA</th><th align="left">OH</th><th align="left">TIHE</th><th align="left">ADHD</th></tr><tr><th align="left">RTr</th><th align="left">RTe</th><th align="left">RTr</th><th align="left">RTe</th><th align="left">RTr</th><th align="left">RTe</th><th align="left">RTr</th><th align="left">RTe</th><th align="left">RTr</th><th align="left">RTe</th></tr></thead><tbody valign="top"><tr><td>5, 1</td><td>0.66</td><td>0.49</td><td>0.94 (0.29)</td><td>0.94 (0.44)</td><td>0.93 (0.28)</td><td>0.93 (0.43)</td><td>0.85 (0.19)</td><td>0.84 (0.35)</td><td>0.97 (0.31)</td><td>0.97 (0.47)</td></tr><tr><td>5, 2</td><td>0.55</td><td>0.34</td><td>0.88 (0.33)</td><td>0.85 (0.51)</td><td>0.86 (0.31)</td><td>0.84 (0.50)</td><td>0.71 (0.16)</td><td>0.70 (0.35)</td><td>0.97 (0.42)</td><td>0.95 (0.61)</td></tr><tr><td>5, 3</td><td>0.47</td><td>0.28</td><td>0.77 (0.29)</td><td>0.74 (0.46)</td><td>0.80 (0.33)</td><td>0.77 (0.50)</td><td>0.60 (0.13)</td><td>0.59 (0.32)</td><td>0.78 (0.30)</td><td>0.69 (0.41)</td></tr><tr><td>10, 1</td><td>0.75</td><td>0.55</td><td>0.92 (0.17)</td><td>0.91 (0.36)</td><td>0.93 (0.18)</td><td>0.93 (0.38)</td><td>0.79 (0.05)</td><td>0.79 (0.24)</td><td>0.97 (0.00)</td><td>0.97 (0.08)</td></tr><tr><td>10, 2</td><td>0.63</td><td>0.30</td><td>0.84 (0.21)</td><td>0.80 (0.50)</td><td>0.87 (0.24)</td><td>0.84 (0.54)</td><td>0.70 (0.07)</td><td>0.69 (0.39)</td><td>0.97 (0.34)</td><td>0.95 <bold>(0.65)</bold></td></tr><tr><td>10, 3</td><td>0.55</td><td>0.20</td><td>0.79 (0.24)</td><td>0.74 (0.54)</td><td>0.81 (0.26)</td><td>0.76 (0.56)</td><td>0.65 (0.10)</td><td>0.64 (0.44)</td><td>0.74 (0.20)</td><td>0.63 (0.43)</td></tr><tr><td>15, 1</td><td>0.75</td><td>0.55</td><td>0.94 (0.19)</td><td>0.95 (0.40)</td><td>0.93 (0.17)</td><td>0.93 (0.38)</td><td>0.79 (0.04)</td><td>0.79 (0.24)</td><td>0.94 (0.19)</td><td>0.90 (0.35)</td></tr><tr><td>15, 2</td><td>0.63</td><td>0.25</td><td>0.82 (0.19)</td><td>0.82 (0.57)</td><td>0.85 (0.22)</td><td>0.82 (0.58)</td><td>0.69 (0.06)</td><td>0.68 (0.44)</td><td>0.92 (0.29)</td><td>0.9 <bold>(0.65)</bold></td></tr><tr><td>15, 3</td><td>0.57</td><td>0.13</td><td>0.76 (0.19)</td><td>0.71 <bold>(0.58)</bold></td><td>0.79 (0.22)</td><td>0.72 <bold>(0.59)</bold></td><td>0.66 (0.09)</td><td>0.64 <bold>(0.51)</bold></td><td>0.74 (0.17)</td><td>0.54 (0.40)</td></tr></tbody></table> </ephtml> </p> <p>3 <emph>Note</emph>. The respective difference values to the random corpus are given in parentheses (peak values in bold).</p> <p>Only the ADHD corpus reaches the maximum difference already at (<reflink idref="bib15" id="ref65">15</reflink>, 2), which is probably due to the small number of just 10 records in the test set. Therefore, this result may not be reliable. Since all other corpora consistently top at a parameter set of (<reflink idref="bib15" id="ref66">15</reflink>, 3), those parameter values are used for all subsequent evaluations.</p> <hd id="AN0138519222-28">Choosing a query</hd> <p>The query optimization algorithm returns a list of possible queries that yield varying recalls, but one query must be selected for the evaluation steps in Section. The following two query features are considered when making this decision.</p> <p>(a) The number of search results retrieved by a query is an important metric because—in particular for a systematic review—all the search results must be reviewed. Therefore, this number should be low. (b) The recall based on the training set should be high because high coverage of the initial corpus is a general requirement for a Potential Technologies Review.</p> <p>The twin goals of high recall with minimal search results are generally contradictory. This trade‐off also applies to the query optimization algorithm and the applied systematic review corpora; queries yielding high recall generally return large result sets (see Figure ).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0006.jpg" title="6 The mean training recall of the queries and the corresponding mean number of PubMed search results based on the 5 × 2 cross‐validation. Data are shown for the random corpus, the systematic review corpora OH, AA, ADHD and the scoping review corpora PACC, ACY" /> </p> <p></p> <p>A trade‐off must be found between a relatively small result set and a high <emph>training recall</emph>. In most cases, the greatest increase of search results is between 90% and 100% recall, on average more than 14 times greater search results (see Figure ). Therefore, the recall threshold for training is set to 90%, which is also in line with the general idea of the Potential Technologies Review.</p> <p>For further application and evaluation, we used the 90% threshold to select the query from the set of generated queries. For the evaluation in Section , the first query is taken whose <emph>training recall</emph> is closest to 90% (equal or higher).</p> <hd id="AN0138519222-30">Query results on test sets</hd> <p>Figure  shows exemplary results for the AA corpus and two parameter settings. The recall and the number of results (as indicated by "results") can vary substantially for the two different parameter settings. For example, the smallest result depicted in Figure B1 is 0.1 × 10<sups>6</sups>, while it is 0.7 × 10<sups>4</sups> in Figure B2. Obviously, the results largely depend on the parameter setting; therefore, parameter optimization is necessary.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0007.jpg" title="7 Process steps of an automatic query optimization shown by exemplary results for the AA corpus with two different parameter settings. Image (A) shows the identified terms and their document frequency (df) according to the initial corpus. In (B1) and (B2), keyword‐count is set to 10, while in (B1) the AND‐combinations parameter is set to 1, and in (B2), it is set to 2. Images (B1) and (B2) show the results and evaluation of CREATE_AND_QUERIES and CREATE_OR_QUERIES, respectively." /> </p> <p></p> <p>Only for the parameters of <emph>kc</emph> = 15 and <emph>ac</emph> = 3, for the AA Corpus in total 6613 queries are automatically generated and their PubMed results evaluated (the number of queries is based on the binomial coefficient for 1, 2, 3 AND‐combinations for all different 5 × 2 cross‐validations). Among these, 10 queries are optimized for the 90% recall threshold on the training set providing different PubMed results and recalls on the test set (see Table ).</p> <p>Exemplary results from automatic query optimization for AA corpus with 146 included studies (for training and testing 73 documents each)</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Id</th><th align="left">PubMed Results</th><th align="left">Overall Recall (Test), %</th></tr></thead><tbody valign="top"><tr><td>AA1</td><td align="char" char=".">283 662</td><td align="char" char=".">90.4</td></tr><tr><td>AA2</td><td align="char" char=".">123 458</td><td align="char" char=".">84.9</td></tr><tr><td>AA3</td><td align="char" char=".">27 464</td><td align="char" char=".">89.0</td></tr><tr><td>AA4</td><td align="char" char=".">21 338</td><td align="char" char=".">80.8</td></tr><tr><td>AA5</td><td align="char" char=".">21 315</td><td align="char" char=".">83.6</td></tr><tr><td>AA6</td><td align="char" char=".">20 905</td><td align="char" char=".">79.5</td></tr><tr><td>AA7</td><td align="char" char=".">20 884</td><td align="char" char=".">84.9</td></tr><tr><td>AA8</td><td align="char" char=".">20 491</td><td align="char" char=".">79.5</td></tr><tr><td>AA9</td><td align="char" char=".">18 950</td><td align="char" char=".">84.9</td></tr><tr><td>AA10</td><td align="char" char=".">18 040</td><td align="char" char=".">80.8</td></tr></tbody></table> </ephtml> </p> <p>4 <emph>Note</emph>. Based on 5 × 2 cross‐validation 10 queries are optimized for 90% recall on the training set (66 out of 73). The different queries lead to varying size of PubMed results that cover different recalls on the test set.</p> <p>In the following, the query of AA9 (see Table ) is shown as an example that provides the best balance (sum of ranks) between few PubMed results and a high <emph>overall test recall</emph> compared with the other queries:</p> <p> <emph>((efficac*[All Fields]) AND (score*[All Fields]) AND (antipsychot*[All Fields])) OR ((score*[All Fields]) AND (antipsychot*[All Fields]) AND (random*[All Fields])) OR ((schizophrenia[All Fields]) AND (efficac*[All Fields]) AND (score*[All Fields])) OR ((trial*[All Fields]) AND (score*[All Fields]) AND (antipsychot*[All Fields])) OR ((efficac*[All Fields]) AND (antipsychot*[All Fields]) AND (random*[All Fields])) OR ((week*[All Fields]) AND (antipsychot*[All Fields]) AND (random*[All Fields])) OR ((improv*[All Fields]) AND (score*[All Fields]) AND (antipsychot*[All Fields])) OR ((improv*[All Fields]) AND (antipsychot*[All Fields]) AND (random*[All Fields])) OR ((schizophrenia[All Fields]) AND (week*[All Fields]) AND (random*[All Fields])) OR ((improv*[All Fields]) AND (efficac*[All Fields]) AND (antipsychot*[All Fields])) OR ((schizophrenia[All Fields]) AND (antipsychot*[All Fields]) AND (random*[All Fields])) OR ((compar*[All Fields]) AND (improv*[All Fields]) AND (antipsychot*[All Fields])) OR ((clinic*[All Fields]) AND (antipsychot*[All Fields]) AND (random*[All Fields])) OR ((signific*[All Fields]) AND (trial*[All Fields]) AND (antipsychot*[All Fields])) OR ((clinic*[All Fields]) AND (week*[All Fields]) AND (antipsychot*[All Fields])) OR ((compar*[All Fields]) AND (trial*[All Fields]) AND (antipsychot*[All Fields])) OR ((trial*[All Fields]) AND (symptom*[All Fields]) AND (antipsychot*[All Fields])) OR ((signific*[All Fields]) AND (compar*[All Fields]) AND (antipsychot*[All Fields]))</emph> </p> <hd id="AN0138519222-32">Comparative performance analysis</hd> <p>The aim of this section is to evaluate the prioritization performance of the applied relevance ranking algorithm in conjunction with the query optimization algorithm. To facilitate contextualization, prioritization performance is compared with existing results of studies on automatic classification[<reflink idref="bib17" id="ref67">17</reflink>] although they were mainly dedicated explicitly to systematic reviews.</p> <p>In Table , the WSS scores, which are achieved by the combination of automatic query optimization and relevance ranking, are compared with existing results for automatic classification. Because the <emph>overall recall</emph> is generally lower in this evaluation and is an important difference from previously reported WSS scores, the <emph>overall recall</emph> is shown along with the WSS score. The WSS scores achieved by the method presented in this work are lower than the reported WSS scores for all corpora.</p> <p>Comparison of prioritization performance indicated by WSS@95 between the proposed method (Potential Technologies Review [PTR] using automatic query optimization and relevance ranking) and maximum values reported from previous studies</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Corpus</th><th align="left">Cohen (2006)<xref ref-type="bibr" rid="bibr16">16</xref></th><th align="left">Howard (2016)<xref ref-type="bibr" rid="bibr17">17</xref></th><th align="left">PTR<xref ref-type="fn" rid="tfn5" /></th></tr><tr><th align="left" /><th align="left">WSS@95</th><th align="left">WSS@95</th><th align="left">WSS@95</th><th align="left">(Overall Test Recall, %)</th></tr></thead><tbody valign="top"><tr><td>Systematic reviews</td></tr><tr><td>TIHE</td><td>N/A</td><td>71.4</td><td>9.2</td><td>(84.3%)</td></tr><tr><td>ADHD</td><td>68.0</td><td>79.3</td><td>15.8</td><td>(80.8%)</td></tr><tr><td>AA</td><td>14.1</td><td>25.1</td><td>12.5</td><td>(78.6%)</td></tr><tr><td>OH</td><td>9.0</td><td>11.7</td><td>4.1</td><td>(80.1%)</td></tr><tr><td>Scoping reviews</td></tr><tr><td>ACY</td><td>N/A</td><td>N/A</td><td>17.3</td><td>(60.8%)</td></tr><tr><td>PACC</td><td>N/A</td><td>N/A</td><td>1.8</td><td>(87.3%)</td></tr></tbody></table> </ephtml> </p> <p>5 The WSS@95 score and <emph>overall test recall</emph> are averaged based on data from the applied 5 × 2 cross‐validation.</p> <p>Compared with existing work on the use of text mining for supporting systematic reviews, a Potential Technologies Review is not necessarily tied to achieving a certain recall level. Therefore, we investigated how the algorithm performs at different <emph>overall recall</emph> levels on the test data.</p> <p>Figure  depicts the achieved WSS score for the corpora at different <emph>overall recall</emph> levels on the test set. It shows how the prioritization performance (given by WSS score) depends on the desired amount of included studies (test set) identified as potentially relevant. The peaks of the graphs reside at different recall levels. Especially for the TIHE and AA corpora, higher WSS scores can be achieved by choosing an overall <emph>test recall</emph> lower than 95%.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0008.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0008.jpg" title="8 Dependency of average WSS score on the specific overall test recall for the systematic review corpora AA, ADHD, OH, TIHE and scoping review corpora ACY, PACC based on 5 × 2 cross‐validation with 90% training recall" /> </p> <p></p> <p>Another restriction that can reduce the prioritization performance of the proposed method is its application to relatively small systematic review corpora (excluding the TIHE corpus). Therefore, we performed a PubMed‐wide example evaluation for the AA corpus to more realistically model the application of the proposed method. The AA corpus was chosen because it is a midsize corpus that has no special features. Consequently, these findings should be generalizable to other corpora.</p> <p>Figure  shows exemplarily with the AA corpus that the achievable WSS score for all <emph>overall test recall</emph> levels is substantially higher for a PubMed‐wide search compared with the search within the AA corpus data (see Figure ): the maximum WSS score of 53% is reached at an <emph>overall recall</emph> of 80% (7180 documents examined) and is essentially higher than the maximum WSS score of 22% at 65% <emph>overall recall</emph> level (319 documents examined) for the search within the AA corpus data.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0009.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0009.jpg" title="9 Dependency of average WSS score on the selected overall test recall for the systematic review corpus AA and a PubMed‐wide search and prioritization based on 5 × 2 cross‐validation with 90% training recall" /> </p> <p></p> <hd id="AN0138519222-35">Case study: Potential Technologies Review on ILD</hd> <p>The methods and parameters described above are now used for query optimization and screening prioritization in a real‐world set‐up of a Potential Technologies Review on "Innovative Laboratory Diagnostics."</p> <p>An important difference from evaluations using existing systematic review corpora is that the query trained by the query optimization algorithm is applied to the entire PubMed database of approximately 27 million records (see https://<ulink href="http://www.nlm.nih.gov/bsd/licensee/2017%5fstats/2017%5fLO.html">www.nlm.nih.gov/bsd/licensee/2017%5fstats/2017%5fLO.html</ulink>).</p> <p>The 5 × 2 cross‐validation (see Section) is also applied in this section to select the best performing query and enable a simple method for estimating the <emph>overall test recall</emph>. The rationale for estimating the recall based on cross‐validation is as follows. The test set mimics a set of unknown but relevant records. Consequently, the recall according to the test set is an estimation of the recall for unknown relevant records.</p> <p>Table  shows that the WSS substantially depends on the training subset. The WSS@95 score varies between 17.4% and 91.1%. The <emph>overall test recall</emph>, based on the full test set and not just to the records that appear in the result list, ranges from 60.2% to 82.3%.</p> <p>Work‐saved‐over sampling (WSS) score regarding test recall of 95% for each ILD data subset with the corresponding overall test recall</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Subset</th><th align="left">WSS@95</th><th align="left">Overall Test Recall</th></tr></thead><tbody valign="top"><tr><td>ILD1a</td><td align="char" char=".">64.0</td><td align="char" char="."><bold>82.3</bold></td></tr><tr><td>ILD1b</td><td align="char" char="."><bold>17.3</bold></td><td align="char" char="."><bold>60.2</bold></td></tr><tr><td>ILD2a</td><td align="char" char="."><bold>91.1</bold></td><td align="char" char=".">76.0</td></tr><tr><td>ILD2b</td><td align="char" char=".">32.4</td><td align="char" char=".">69.7</td></tr><tr><td>ILD3a</td><td align="char" char=".">74.2</td><td align="char" char=".">72.8</td></tr><tr><td>ILD3b</td><td align="char" char=".">88.3</td><td align="char" char=".">69.7</td></tr><tr><td>ILD4a</td><td align="char" char=".">80.5</td><td align="char" char=".">63.3</td></tr><tr><td>ILD4b</td><td align="char" char=".">67.5</td><td align="char" char=".">69.7</td></tr><tr><td>ILD5a</td><td align="char" char=".">62.6</td><td align="char" char=".">79.2</td></tr><tr><td>ILD5b</td><td align="char" char=".">73.5</td><td align="char" char=".">66.5</td></tr></tbody></table> </ephtml> </p> <p>6 <emph>Note</emph>. In each column, the highest and lowest value is shown in bold letters. ILD2A yields the highest workload reduction.</p> <p>While the WSS score quantifies the ability of a query to reduce the workload, the <emph>overall test recall</emph> measures the associated coverage. The results in Table  indicate that the WSS score and <emph>overall test recall</emph> are not negatively correlated: the subset yielding the lowest WSS score (ILD1B) also yields the lowest <emph>overall test recall</emph>, while the subset with the highest WSS score (ILD2A) yields the third highest <emph>overall test recall</emph>. In other words, a high workload reduction does not imply low coverage. Therefore, it is safe to choose the query from the subset yielding the highest WSS score. In this case, the ILD2a subset, which is chosen for further evaluation.</p> <p>This query optimization was performed in September 2017. We downloaded the PubMed dataset according to the query based on the ILD2A subset in November 2017. This enabled the text mining and information retrieval methods to include the latest research activities.</p> <p>Compared with random sampling, a substantial amount of screening workload can be saved by ranking the search results according to the optimized query (see Figure ). The <emph>test recall</emph> is more than 90 percentage points higher compared with random sampling after examining approximately only 4.0% of the documents (approximately 552 documents). The <emph>test recall</emph> at this point corresponds to an <emph>overall test recall</emph> of 76%. This <emph>overall test recall</emph> serves as an estimation of the recall of the Potential Technologies Review.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0010.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0010.jpg" title="10 The solid line shows the documents examined for the specific test recall for the ILD2a subset after prioritization. The interpolated WSS@95 is visualized and the corresponding overall test recall is shown. The dashed line shows the expected test recall when performing screening based on random sampling" /> </p> <p></p> <p>Based on those results the top‐ranked 552 documents were manually screened. This portion was chosen because those documents should be the most relevant and provide the greatest benefit over random sampling. Moreover, a corpus of 552 documents can be screened even with limited resources. The expected recall of 76% is reasonable for the ILD project.</p> <p>Figure  shows the positions of relevant documents within the corpus of 552 documents. According to the last column ("Total"), most of the relevant documents were ranked in the top 60 documents. In contrast, almost no relevant documents were ranked between 221 and 460, and only 11 relevant documents were found beyond rank 460. The third column ("New") show that those relevant records beyond rank 460 are mostly new relevant records that did not belong to the initial corpus.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01sep19/jrsm1350-fig-0011.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1350-fig-0011.jpg" title="11 A heat map of the manual screening results. Each cell shows the absolute number of relevant documents found at a certain ranking level. The first two columns contain the distribution of the documents of the initial corpus, the third column the newly identified relevant studies after final manual screening and the last column the total amount of relevant documents. Additionally, the first following document belonging to the test set after the screened 552 documents is shown, which is used for calculating the interpolated overall test recall of 76%" /> </p> <p></p> <p>Summing up those results, the majority of the documents belonging to the initial corpus were successfully prioritized. Most new relevant documents were also prioritized in a manner comparable with those of the initial corpus. Among those were at least one highly relevant publication providing new aspects to concepts that are contained in the initial corpus. Furthermore, for some new relevant documents, the prioritization shows a second peak below rank 460. Among those at least two publications contain completely new relevant concepts beyond those of the initial corpus.</p> <hd id="AN0138519222-38">DISCUSSION</hd> <p>We proposed a novel hybrid review approach that incorporates features of both systematic and scoping review techniques to identify suitable technological solutions for a clinical demand from scientific literature—the "Potential Technologies Review." Its defining features contain high recall (but not 100%) and reproducibility despite the fuzzy boundaries at the onset. All this must be accomplished within a particularly tight timeframe to meet the major goal of translational research. Therefore, we developed and evaluated a novel information retrieval framework to accelerate the usually most burdensome review stages, search and triage of relevant citations. Its feasibility was demonstrated on a real‐world dataset derived from a case study in oncology.</p> <p>We integrated two related text mining applications: (a) A complex query is built automatically based on an initial corpus of relevant documents. (b) The manual screening workload is substantially reduced by means of query‐based relevance ranking. All analyses are conducted as 5 × 2 cross‐validations to assess the generalization performance of the different steps.</p> <p>The query optimization algorithm is based on ATR utilizing the documents frequency. This term recognition method was selected, because it is a straightforward approach that considers the importance of a term within a certain corpus in contrast to other ATR methods such as Termine[<reflink idref="bib39" id="ref68">39</reflink>] or RAKE[<reflink idref="bib40" id="ref69">40</reflink>] that base on term frequency. Moreover, the automated query optimization is extended to include Boolean operators (AND, OR) and thus formulates a structured and comprehensive search statement. The numerous possible versions of the query applied in PubMed are finally optimized using the <emph>training recall</emph> as target function.</p> <p>Although automated query formulation has been researched in the past, mainly within the computer science domain,[[<reflink idref="bib49" id="ref70">49</reflink>]] it is—to the best of our knowledge—now the first implementation to support a review procedure. While current approaches use additional metadata[<reflink idref="bib51" id="ref71">51</reflink>] or graph databases[<reflink idref="bib52" id="ref72">52</reflink>] to improve the query formulation or to assist the searcher, our tool is limited to terms and operators, which keeps the resulting query still being understandable/human‐readable. Furthermore, no user interaction is required. The cross‐validation yields for all corpora far more than 90% recall considering the testing dataset. That means, that the tool once trained is able to identify 90% of the unknown relevant data—a remarkable generalization performance considering the initially fuzzy scope. Moreover, the results are user‐independent and reproducible for a given initial corpus.</p> <p>For the reduction of the screening workload, the results vary depending on the evaluation scenario. While the prioritization performance of the relevance ranking in conjunction with query optimization on literature‐based data of systematic or scoping reviews is inferior to existing methods,[<reflink idref="bib17" id="ref73">17</reflink>] high WSS scores are observed when applied to a real‐world dataset.</p> <p>When interpreting the findings on literature based corpora, the following has to be considered:</p> <p></p> <ulist> <item> The comparability of the calculated WSS score to those reported by Howard et al[<reflink idref="bib17" id="ref74">17</reflink>] is somewhat limited due to a procedural difference. While the optimized queries are generated using PubMed, the final prioritization is performed on the provided result datasets using Solr. Since the query and its execution is inevitable when applying the ranking to the documents, the <emph>overall recall</emph> is generally reduced (see Section).</item> <p></p> <item> The mismatch between the query derived from PubMed‐based optimization applied to a distinct result set, which is not used for generating the query, probably lowers the effect of the chosen ranking procedure. The test of the AA corpus supports this interpretation: An optimized AA query submitted to the full PubMed dataset instead of the corpus shows a substantial performance improvement from 12.5% to 53% WSS. Therefore, the query‐based relevance ranking might be ill‐posed for a prioritization of systematic review corpora that were preselected on the basis of a different (manually formulated) query.</item> </ulist> <p>In the case study on ILD corpus, the text mining procedures performed very well. The maximum achieved WSS score (91%) is superior to previously reported WSS scores for the application of text mining within systematic reviews and is comparable with the reported screening workload reduction for large scoping reviews.[<reflink idref="bib10" id="ref75">10</reflink>] Consequently, in the feasibility study, the absolute workload for the closing manual screening was reduced to only 552 documents.</p> <p>Moreover, calculating the <emph>overall recall</emph> as an estimation of the coverage of the Potential Technologies Review is achieved by performing cross‐validation. The test set, which is not used for optimizing the query, mimics unknown but relevant records. Therefore, it can be used to estimate the coverage achieved during screening. In the feasibility study, this aspect was not fully investigated because it would be necessary to screen the entire result set. The analysis of the first 552 documents, which have been screened manually, shows a difference between the distribution of documents belonging to the initial corpus and the distribution of 39 "new relevant" records.</p> <p>One possible explanation is that the applied query achieves the highest WSS score in the cross‐validation; therefore, it might be "overfitted" to the initial corpus with nearly all documents placed at top ranking positions. On the other hand, the distribution of the new relevant records shows two peaks, one among the top ranking classes and a second at the end of the ordered list. The content of this last group of relevant citations obviously differs considerably from the concepts represented by the initial corpus. The manual screening of the full‐text revealed, that two out of the 10 low‐ranked new relevant records dealt with completely novel concepts compared with the initial corpus. Thus, our procedure seems to be remarkably robust against the perhaps most dangerous pitfall of a text mining based process, the <emph>hasty generalization</emph>.[<reflink idref="bib12" id="ref76">12</reflink>] This outcome is the result of complementary effects of several individual steps such as the by cross‐validation proven generalization performance of the automated query optimization, the avoidance of excessive recall values or the combination of different text mining tools[<reflink idref="bib10" id="ref77">10</reflink>] complemented with reviewer dependent manual stages.</p> <p>The case study shows the advanced applicability of the proposed procedure especially under real conditions. Compared with machine learning‐based ACs,[[<reflink idref="bib16" id="ref78">16</reflink>], [<reflink idref="bib24" id="ref79">24</reflink>], [<reflink idref="bib53" id="ref80">53</reflink>]] the novel procedure does not require a large amount of already labeled training data and is therefore capable to perform a complete PTR from the scratch.</p> <p>For the further development of the procedure, certain aspects should be investigated to improve the approach or to extend the applicability.</p> <p>Generally, hasty generalization is a major problem. As described above, a carefully performed preliminary search is an important remedy. An interesting contribution to this stage of the search could be the approach used in the ASSERT project,[<reflink idref="bib32" id="ref81">32</reflink>] which offers an iterative procedure that requests user feedback to expand a query by similarity and supports the review of search results by document clustering methods.</p> <p>In this study, we applied a relevance ranking algorithm that uses the optimized query to prioritize records for the screening process. Different approaches could be evaluated and treated as either additional or substitute methods. Existing work in support of text mining within systematic reviews focuses on automatic classification.[[<reflink idref="bib12" id="ref82">12</reflink>], [<reflink idref="bib16" id="ref83">16</reflink>], [<reflink idref="bib55" id="ref84">55</reflink>]] However, because automatic classification requires both positive and negative training data, it might not be a good substitute for relevance ranking, but it could possibly support the screening process as some preliminary tests using a SVM classifier on the screening results of the feasibility study indicates. This would be in line with the findings of Shemilt et al.[<reflink idref="bib10" id="ref85">10</reflink>] Other text mining technologies such as topic modeling have also been applied[[<reflink idref="bib53" id="ref86">53</reflink>], [<reflink idref="bib57" id="ref87">57</reflink>]] and should also be investigated.</p> <p>In addition to further analysis of additional methodical enhancements, the procedure should be broadened to other data sources. While this work focused on scientific publications in the PubMed database, applying the framework to other document types and databases such as patents might provide up‐to‐date insights into the current research and development landscape of the topic in question.</p> <p>Finally, the automated query optimization needs very much computation power, eg, it requires approximately 13 hours for the ILD corpus, because the optimization algorithm uses the PubMed E‐Utilities interface. By applying the algorithm to a local PubMed database as suggested by Döring et al,[<reflink idref="bib58" id="ref88">58</reflink>] the procedure could be substantially accelerated.</p> <hd id="AN0138519222-39">CONCLUSION</hd> <p>We were able to show that a novel specific system of related text mining and information retrieval tools is capable to automatically support searching and screening of the existing scientific literature for relevant technologies in translational research. Starting from a relatively small, manually created corpus of relevant work matching the given clinical demand, a complex query is automatically generated from terms including Boolean operators to expand the search and then the recall‐optimized version is selected in an iterative comparison with PubMed. This procedure does not only go far beyond the ATR approaches, eg, RAKE[<reflink idref="bib40" id="ref89">40</reflink>] or Termine[<reflink idref="bib33" id="ref90">33</reflink>] used so far in systematic reviews, but in combination with the query‐based prioritization indicates that it does not tend towards "hasty generalization," but on the contrary also includes new aspects at the edges of the topic area. This has to be investigated in more detail in the future.</p> <p>With the presented text mining and information retrieval framework, not only the first two phases (search and triage) of a "Potential Technologies Review" can be handled quickly (due to reduced workload) and efficiently (reproducible, high recall). The automation procedure could also prove its worth in scoping reviews or even systematic reviews under particularly tight timelines, because it could not only be used for updates or the replacement of the required second reviewer, but also right from the start—a field for future analyses.</p> <p>In any case, the text mining framework allows to find relevant technologies within the procedure of a PTR in just a few weeks with clear reproducibility and thus makes an enabling contribution to the accelerated development of demand driven innovations in biomedical technology.</p> <hd id="AN0138519222-40">ACKNOWLEDGMENT</hd> <p>This research was supported by Klaus Tschira Stiftung gGmbH, Heidelberg, Germany.</p> <hd id="AN0138519222-41">CONFLICT OF INTEREST</hd> <p>None</p> <hd id="AN0138519222-42">DATA AVAILABILITY STATEMENT</hd> <p>Data of the different corpora are available either from published studies (systematic reviews) or from in the supporting files (reconstructed scoping reviews). The data of ILD corpus are available on reasonable request from the corresponding author. The data are not publicly available due to privacy or commercial restrictions.</p> <p>GRAPH: Appendix S1 Supporting information</p> <p>GRAPH: Data S2 Supporting information</p> <ref id="AN0138519222-43"> <title> REFERENCES </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> Kaitin KI. Translational research and the evolving landscape for biomedical innovation. J Invest Med. 2012 ; 60 (7): 995 ‐ 998. https://doi.org/10.2310/JIM.0b013e318268694f</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> Baas J, Brauer U, Dannhorn DR, et al. Innovation in der Medizintechnik: Nationaler Strategieprozess http://www.strategieprozess‐medizintechnik.de/sites/default/files/Schlussbericht%5fNSIM.pdf. Accessed April 25, 2016.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref55" type="bt">3</bibl> <bibtext> McCarthy AD, Sproson L, Wells O, Tindale W. Unmet needs: relevance to medical technology innovation. J Med Eng Technol. 2014 ; 39 (7): 382 ‐ 387. https://doi.org/10.3109/03091902.2015.1088093</bibtext> </blist> <blist> <bibl id="bib4" type="bt">4</bibl> <bibtext> Godin B, Lane JP. Pushes and pulls: hi(S)tory of the demand pull model of innovation. Sci Technol Hum Values. 2013 ; 38 (5): 621 ‐ 654. https://doi.org/10.1177/0162243912473163</bibtext> </blist> <blist> <bibl id="bib5" idref="ref3" type="bt">5</bibl> <bibtext> Booth A. Unpacking your literature search toolbox: on search styles and tactics. Health Info Libr J. 2008 ; 25 (4): 313 ‐ 317. https://doi.org/10.1111/j.1471‐1842.2008.00825.x</bibtext> </blist> <blist> <bibl id="bib6" idref="ref4" type="bt">6</bibl> <bibtext> Egger M, Smith GD, Altman DG. Systematic Reviews in Health Care. London, UK : BMJ Publishing Group ; 2001.</bibtext> </blist> <blist> <bibl id="bib7" type="bt">7</bibl> <bibtext> Cooper H. The Handbook of Research Synthesis and Meta‐Analysis. New York : Russell Sage Foundation ; 2009 <ulink href="http://gbv.eblib.com/patron/FullRecord.aspx?p=4416794">http://gbv.eblib.com/patron/FullRecord.aspx?p=4416794</ulink>.</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Higgins JP, Green S. Cochrane handbook for systematic reviews of interventions: version 5.1.0. <ulink href="http://handbook.cochrane.org/">http://handbook.cochrane.org/</ulink>. Updated März 2011. Accessed August 13, 2015.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref5" type="bt">9</bibl> <bibtext> Arksey H, O'Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Methodol. 2005 ; 8 (1): 19 ‐ 32. https://doi.org/10.1080/1364557032000119616</bibtext> </blist> <blist> <bibtext> Shemilt I, Simon A, Hollands GJ, et al. Pinpointing needles in giant haystacks: use of text mining to reduce impractical screening workload in extremely large scoping reviews. Res Synth Methods. 2014 ; 5 (1): 31 ‐ 49. https://doi.org/10.1002/jrsm.1093</bibtext> </blist> <blist> <bibtext> Pham MT, Rajić A, Greig JD, Sargeant JM, Papadopoulos A, McEwen SA. A scoping review of scoping reviews: advancing the approach and enhancing the consistency. Res Synth Methods. 2014 ; 5 (4): 371 ‐ 385. https://doi.org/10.1002/jrsm.1123</bibtext> </blist> <blist> <bibtext> Wallace BC, Trikalinos TA, Lau J, Brodley C, Schmid CH. Semi‐automated screening of biomedical citations for systematic reviews. BMC Bioinf. 2010 ; 11 (1): 55. https://doi.org/10.1186/1471‐2105‐11‐55.</bibtext> </blist> <blist> <bibtext> Allen IE, Olkin I. Estimating time to conduct a meta‐analysis from number of citations retrieved. Jama. 1999 ; 282 (7): 634 ‐ 635.</bibtext> </blist> <blist> <bibtext> Greenhalgh T, Peacock R. Effectiveness and efficiency of search methods in systematic reviews of complex evidence: audit of primary sources. BMJ. 2005 ; 331 (7524): 1064 ‐ 1065. https://doi.org/10.1136/bmj.38636.593461.68</bibtext> </blist> <blist> <bibtext> Greenhalgh T, Robert G, Macfarlane F, Bate P, Kyriakidou O. Diffusion of innovations in service organizations: systematic review and recommendations. Milbank Q. 2004 ; 82 (4): 581 ‐ 629. https://doi.org/10.1111/j.0887‐378X.2004.00325.x</bibtext> </blist> <blist> <bibtext> Cohen AM, Hersh WR, Peterson K, Yen P‐Y. Reducing workload in systematic review preparation using automated citation classification. J Am Med Inform Assoc. 2006 ; 13 (2): 206 ‐ 219. https://doi.org/10.1197/jamia.M1929</bibtext> </blist> <blist> <bibtext> Howard BE, Phillips J, Miller K, et al. SWIFT‐review: a text‐mining workbench for systematic review. Syst Rev. 2016 ; 5 (1): 87. https://doi.org/10.1186/s13643‐016‐0263‐z.</bibtext> </blist> <blist> <bibtext> Hurlimann T, Peña‐Rosas JP, Saxena A, Zamora G, Godard B. Correction: ethical issues in the development and implementation of nutrition‐related public health policies and interventions: a scoping review. PLoS ONE. 2018 ; 13 (2): e0192356. https://doi.org/10.1371/journal.pone.0192356</bibtext> </blist> <blist> <bibtext> Douglas MP, Ladabaum U, Pletcher MJ, Marshall DA, Phillips KA. Economic evidence on identifying clinically actionable findings with whole‐genome sequencing: a scoping review. Genet Med. 2016 ; 18 (2): 111 ‐ 116. https://doi.org/10.1038/gim.2015.69</bibtext> </blist> <blist> <bibtext> Morris ZS, Wooding S, Grant J. The answer is 17 years, what is the question: understanding time lags in translational research. J R Soc Med. 2011 ; 104 (12): 510 ‐ 520. https://doi.org/10.1258/jrsm.2011.110180</bibtext> </blist> <blist> <bibtext> Chalmers I, Bracken MB, Djulbegovic B, et al. How to increase value and reduce waste when research priorities are set. Lancet. 2014 ; 383 (9912): 156 ‐ 165. https://doi.org/10.1016/S0140‐6736(13)62229‐1</bibtext> </blist> <blist> <bibtext> Farkas R, Puiu AA, Hamadeh N, Bukowski M, Schmitz‐Rode T. Empirical assessment of the time course of innovation in biomedical engineering: first results of a comparative approach. Curr Dir Biomed Eng. 2016 ; 2 (1): 394. https://doi.org/10.1515/cdbme‐2016‐0132</bibtext> </blist> <blist> <bibtext> Kim S, Choi J. An SVM‐based high‐quality article classifier for systematic reviews. J Biomed Inform. 2014 ; 47 : 153 ‐ 159. https://doi.org/10.1016/j.jbi.2013.10.005</bibtext> </blist> <blist> <bibtext> Shekelle PG, Shetty K, Newberry S, Maglione M, Motala A. Machine learning versus standard techniques for updating searches for systematic reviews: a diagnostic accuracy study. Ann Intern Med. 2017 ; 167 (3): 213 ‐ 215. https://doi.org/10.7326/L17‐0124</bibtext> </blist> <blist> <bibtext> Frunza O, Inkpen D, Matwin S. Building systematic reviews using automatic text classification techniques. In: Proceedings of the 23 rd International Conference on Computational Linguistics: Posters. Stroudsburg, PA, USA : Association for Computational Linguistics ; 2010 : 303 ‐ 311 COLING'10.</bibtext> </blist> <blist> <bibtext> Liu J, Timsina P, El‐Gayar O. A comparative analysis of semi‐supervised learning: The case of article selection for medical systematic reviews. Inf Syst Front. 2018 ; 20 (2, SI): 195 ‐ 207. https://doi.org/10.1007/s10796‐016‐9724‐0</bibtext> </blist> <blist> <bibtext> O'Mara‐Eves A, Thomas J, McNaught J, Miwa M, Ananiadou S. Using text mining for study identification in systematic reviews: a systematic review of current approaches. Syst Rev. 2015 ; 4 (1): 5. https://doi.org/10.1186/2046‐4053‐4‐5.</bibtext> </blist> <blist> <bibtext> Bekhuis T, Tseytlin E, Mitchell KJ, Demner‐Fushman D. Feature engineering and a proposed decision‐support system for systematic reviewers of medical evidence. PLoS ONE. 2014 ; 9 (1): e86277. https://doi.org/10.1371/journal.pone.0086277</bibtext> </blist> <blist> <bibtext> Bui DDA, Jonnalagadda S, Del Fiol G. Automatically finding relevant citations for clinical guideline development. J Biomed Inform. 2015 ; 57 : 436 ‐ 445. https://doi.org/10.1016/j.jbi.2015.09.003</bibtext> </blist> <blist> <bibtext> Cohen AM, Ambert K, McDonagh M. Cross‐topic learning for work prioritization in systematic review creation and update. J Am Med Inform Assoc. 2009 ; 16 (5): 690 ‐ 704. https://doi.org/10.1197/jamia.M3162</bibtext> </blist> <blist> <bibtext> Thomas J, McNaught J, Ananiadou S. Applications of text mining within systematic reviews. Res Synth Methods. 2011 ; 2 (1): 1 ‐ 14. https://doi.org/10.1002/jrsm.27</bibtext> </blist> <blist> <bibtext> Ananiadou S, Rea B, Okazaki N, Procter R, Thomas J. Supporting systematic reviews using text mining. Soc Sci Comput Rev. 2009 ; 27 (4): 509 ‐ 523. https://doi.org/10.1177/0894439309332293</bibtext> </blist> <blist> <bibtext> Stansfield C, O'Mara‐Eves A, Thomas J. Text mining for search term development in systematic reviewing: a discussion of some methods and challenges. Res Synth Methods. 2017 ; 8 (3): 355 ‐ 365. https://doi.org/10.1002/jrsm.1250</bibtext> </blist> <blist> <bibtext> Stansfield C, Thomas J, Kavanagh J. Clustering' documents automatically to support scoping reviews of research: a case study. Res Synth Methods. 2013 ; 4 (3): 230 ‐ 241. https://doi.org/10.1002/jrsm.1082</bibtext> </blist> <blist> <bibtext> Miake‐Lye IM, Hempel S, Shanman R, Shekelle PG. What is an evidence map? A systematic review of published evidence maps and their definitions, methods, and products. Syst Rev. 2016 ; 5 (1): 28. https://doi.org/10.1186/s13643‐016‐0204‐x.</bibtext> </blist> <blist> <bibtext> Marshall Z, Welch V, Thomas J, et al. Documenting research with transgender and gender diverse people: protocol for an evidence map and thematic analysis. Syst Rev. 2017 ; 6 (1): 35. https://doi.org/10.1186/s13643‐017‐0427‐5</bibtext> </blist> <blist> <bibtext> Khan KS, Kunz R, Kleijnen J, Antes G. Five steps to conducting a systematic review. J R Soc Med. 2003 ; 96 (3): 118 ‐ 121.</bibtext> </blist> <blist> <bibtext> Porter AL, Youtie J, Shapira P, Schoeneck DJ. Refining search terms for nanotechnology. J Nanopart Res. 2008 ; 10 (5): 715 ‐ 728. https://doi.org/10.1007/s11051‐007‐9266‐y</bibtext> </blist> <blist> <bibtext> Frantzi K, Ananiadou S, Mima H. Automatic recognition of multi‐word terms: the C‐value/NC‐value method. Int J Digit Libr. 2000 ; 3 (2): 115 ‐ 130. https://doi.org/10.1007/s007999900023</bibtext> </blist> <blist> <bibtext> Rose S, Engel D, Cramer N, Cowley W. Automatic keyword extraction from individual documents. In: Berry MW, Kogan J, eds. Text Mining. Chichester, UK : John Wiley & Sons, Ltd ; 2010 : 1 ‐ 20.</bibtext> </blist> <blist> <bibtext> Manning CD, Raghavan P, Schütze H. Introduction to Information Retrieval. Reprinted. Cambridge : Cambridge Univ. Press ; 2009 http://reference‐tree.com/book/an‐introduction‐to‐information‐retrieval?utm_source=gbv&utm_medium=referral&utm_campaign=collaboration.</bibtext> </blist> <blist> <bibtext> Fox C. A stop list for general text. SIGIR Forum. 1989 ; 24 (1–2): 19 ‐ 21. https://doi.org/10.1145/378881.378888</bibtext> </blist> <blist> <bibtext> Porter MF. An algorithm for suffix stripping. Dent Prog. 1980 ; 14 (3): 130 ‐ 137. https://doi.org/10.1108/eb046814</bibtext> </blist> <blist> <bibtext> Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit‐learn: machine learning in Python. J Mach Learn Res. 2011 ; 12 : 2825 ‐ 2830.</bibtext> </blist> <blist> <bibtext> Salton G, Wong A, Yang CS. A vector space model for automatic indexing. Commun ACM. 1975 ; 18 (11): 613 ‐ 620. https://doi.org/10.1145/361219.361220</bibtext> </blist> <blist> <bibtext> Rewald S, Mesters I, Lenssen AF, et al. Aquatic cycling‐what do we know? A scoping review on head‐out aquatic cycling. PLoS ONE. 2017 ; 12 (5): e0177704. https://doi.org/10.1371/journal.pone.0177704</bibtext> </blist> <blist> <bibtext> Tranvag EJ, Norheim OF, Ottersen T. Clinical decision making in cancer care: a review of current and future roles of patient age. BMC Cancer. 2018 ; 18. https://doi.org/10.1186/s12885‐018‐4456‐9 (1): 546.</bibtext> </blist> <blist> <bibtext> Sammut C, Webb GI (Eds). Encyclopedia of Machine Learning: With 78 Tables. New York, NY : Springer ; 2011 Springer reference. https://doi.org/10.1007/978‐0‐387‐30164‐8.</bibtext> </blist> <blist> <bibtext> Oakes MP, Taylor MJ. Automated assistance in the formulation of search statements for bibliographic databases. Inf Process Manag. 1998 ; 34 (6): 645 ‐ 668. https://doi.org/10.1016/S0306‐4573(98)00029‐6</bibtext> </blist> <blist> <bibtext> Jansen BJ, McNeese MD. Evaluating the effectiveness of and patterns of interactions with automated searching assistance. J Am Soc Inf Sci. 2005 ; 56 (14): 1480 ‐ 1503. https://doi.org/10.1002/asi.20242</bibtext> </blist> <blist> <bibtext> Glocker K, Knurr A, Dieter J, et al. Optimizing a query by transformation and expansion. Stud Health Technol Inform. 2017 ; 243 : 197 ‐ 201.</bibtext> </blist> <blist> <bibtext> Bhowmick SS, Chua HE, Choi B, Dyreson C. VISUAL: simulation of visual subgraph query formulation to enable automated performance benchmarking. IEEE Trans Knowl Data Eng. 2017 ; 29 (8): 1765 ‐ 1778. https://doi.org/10.1109/TKDE.2017.2690392</bibtext> </blist> <blist> <bibtext> Miwa M, Thomas J, O'Mara‐Eves A, Ananiadou S. Reducing systematic review workload through certainty‐based screening. J Biomed Inform. 2014 ; 51 : 242 ‐ 253. https://doi.org/10.1016/j.jbi.2014.06.005</bibtext> </blist> <blist> <bibtext> Wallace BC, Small K, Brodley CE, et al. Toward modernizing the systematic review pipeline in genetics: efficient updating via data mining. Genet Med. 2012 ; 14 (7): 663 ‐ 669. https://doi.org/10.1038/gim.2012.7.</bibtext> </blist> <blist> <bibtext> Aphinyanaphongs Y, Tsamardinos I, Statnikov A, Hardin D, Aliferis CF. Text categorization models for high‐quality article retrieval in internal medicine. J Am Med Inform Assoc. 2005 ; 12 (2): 207 ‐ 216. https://doi.org/10.1197/jamia.M1641</bibtext> </blist> <blist> <bibtext> García Adeva JJ, Pikatza Atxa JM, Ubeda Carrillo M, Ansuategi Zengotitabengoa E. Automatic text classification to support systematic reviews in medicine. Expert Syst App. 2014 ; 41 (4): 1498 ‐ 1508. https://doi.org/10.1016/j.eswa.2013.08.047</bibtext> </blist> <blist> <bibtext> Li D, Wang Z, Shen F, Murad MH, Liu H. Towards a multi‐level framework for supporting systematic review—a pilot study. In: Zheng, Huiru. IEEE International Conference on Bioinformatics and Biomedicine (BIBM) Piscataway, NJ ; 2014 : 43 ‐ 50.</bibtext> </blist> <blist> <bibtext> Döring K, Grüning BA, Telukunta KK, Thomas P, Günther S. PubMedPortable: a framework for supporting the development of text mining applications. PLoS ONE. 2016 ; 11 (10): e0163794. https://doi.org/10.1371/journal.pone.0163794</bibtext> </blist> </ref> <aug> <p>By Tom Schmitz; Mark Bukowski; Steffen Koschmieder; Thomas Schmitz‐Rode and Robert Farkas</p> <p>Reported by Author; Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib10" firstref="ref6"></nolink> <nolink nlid="nl2" bibid="bib12" firstref="ref7"></nolink> <nolink nlid="nl3" bibid="bib15" firstref="ref8"></nolink> <nolink nlid="nl4" bibid="bib16" firstref="ref9"></nolink> <nolink nlid="nl5" bibid="bib17" firstref="ref10"></nolink> <nolink nlid="nl6" bibid="bib18" firstref="ref11"></nolink> <nolink nlid="nl7" bibid="bib19" firstref="ref12"></nolink> <nolink nlid="nl8" bibid="bib20" firstref="ref13"></nolink> <nolink nlid="nl9" bibid="bib23" firstref="ref15"></nolink> <nolink nlid="nl10" bibid="bib25" firstref="ref16"></nolink> <nolink nlid="nl11" bibid="bib26" firstref="ref18"></nolink> <nolink nlid="nl12" bibid="bib27" firstref="ref19"></nolink> <nolink nlid="nl13" bibid="bib28" firstref="ref22"></nolink> <nolink nlid="nl14" bibid="bib29" firstref="ref25"></nolink> <nolink nlid="nl15" bibid="bib30" firstref="ref28"></nolink> <nolink nlid="nl16" bibid="bib31" firstref="ref29"></nolink> <nolink nlid="nl17" bibid="bib32" firstref="ref30"></nolink> <nolink nlid="nl18" bibid="bib33" firstref="ref31"></nolink> <nolink nlid="nl19" bibid="bib34" firstref="ref32"></nolink> <nolink nlid="nl20" bibid="bib35" firstref="ref34"></nolink> <nolink nlid="nl21" bibid="bib36" firstref="ref35"></nolink> <nolink nlid="nl22" bibid="bib37" firstref="ref38"></nolink> <nolink nlid="nl23" bibid="bib14" firstref="ref41"></nolink> <nolink nlid="nl24" bibid="bib38" firstref="ref42"></nolink> <nolink nlid="nl25" bibid="bib39" firstref="ref44"></nolink> <nolink nlid="nl26" bibid="bib40" firstref="ref45"></nolink> <nolink nlid="nl27" bibid="bib41" firstref="ref46"></nolink> <nolink nlid="nl28" bibid="bib42" firstref="ref47"></nolink> <nolink nlid="nl29" bibid="bib43" firstref="ref48"></nolink> <nolink nlid="nl30" bibid="bib44" firstref="ref49"></nolink> <nolink nlid="nl31" bibid="bib45" firstref="ref56"></nolink> <nolink nlid="nl32" bibid="bib48" firstref="ref61"></nolink> <nolink nlid="nl33" bibid="bib49" firstref="ref70"></nolink> <nolink nlid="nl34" bibid="bib51" firstref="ref71"></nolink> <nolink nlid="nl35" bibid="bib52" firstref="ref72"></nolink> <nolink nlid="nl36" bibid="bib24" firstref="ref79"></nolink> <nolink nlid="nl37" bibid="bib53" firstref="ref80"></nolink> <nolink nlid="nl38" bibid="bib55" firstref="ref84"></nolink> <nolink nlid="nl39" bibid="bib57" firstref="ref87"></nolink> <nolink nlid="nl40" bibid="bib58" firstref="ref88"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1255363
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Potential Technologies Review: A Hybrid Information Retrieval Framework to Accelerate Demand-Pull Innovation in Biomedical Engineering
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Schmitz%2C+Tom%22">Schmitz, Tom</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-5937-765X">0000-0002-5937-765X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Bukowski%2C+Mark%22">Bukowski, Mark</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4563-1159">0000-0003-4563-1159</externalLink>)<br /><searchLink fieldCode="AR" term="%22Koschmieder%2C+Steffen%22">Koschmieder, Steffen</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1011-8171">0000-0002-1011-8171</externalLink>)<br /><searchLink fieldCode="AR" term="%22Schmitz-Rode%2C+Thomas%22">Schmitz-Rode, Thomas</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1181-2165">0000-0002-1181-2165</externalLink>)<br /><searchLink fieldCode="AR" term="%22Farkas%2C+Robert%22">Farkas, Robert</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8199-6764">0000-0002-8199-6764</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Research+Synthesis+Methods%22"><i>Research Synthesis Methods</i></searchLink>. Sep 2019 10(3):420-439.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 20
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2019
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Descriptive
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Innovation%22">Innovation</searchLink><br /><searchLink fieldCode="DE" term="%22Biomedicine%22">Biomedicine</searchLink><br /><searchLink fieldCode="DE" term="%22Medical+Research%22">Medical Research</searchLink><br /><searchLink fieldCode="DE" term="%22Meta+Analysis%22">Meta Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Risk%22">Risk</searchLink><br /><searchLink fieldCode="DE" term="%22Guidelines%22">Guidelines</searchLink><br /><searchLink fieldCode="DE" term="%22Sampling%22">Sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Information+Retrieval%22">Information Retrieval</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Reports%22">Research Reports</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Outcomes+of+Treatment%22">Outcomes of Treatment</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jrsm.1350
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1759-2879
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Launching biomedical innovations based on clinical demands instead of translating basic research findings to practice reduces the risk that the results will not fit the clinical routine. To realize this type of innovation, a meta-analysis of the body of research is necessary to reveal demand-matching concepts. However, both the data deluge and the narrow time constraints for innovation make it impossible to perform such reviews manually. Thus, this paper proposes a specifically adapted "Potential Technologies Review" approach focusing on automated text mining and information retrieval techniques. The novel framework combines features from both systematic and scoping reviews. It aims at high coverage and reproducibility while mapping technologies--even with a fuzzy initial scope. To achieve these goals for search and triage, a set of closely interrelated methods has been developed: (a) automated query optimization, (b) screening prioritization, and (c) recall estimation. To determine appropriate parameters, a variety of published literature corpora were used and compared with an evaluation on a real-world dataset. Our results show that it is feasible to automate the identification of relevant works using this newly introduced framework. It achieved a workload reduction of up to 91% "Work-saved-over Sampling (WSS)" with a 76% overall recall compared with manually screening search results. Reducing the workload is a prerequisite for a rapid Potential Technologies Review when conducting demand-pull innovations. Moreover, it facilitates the updating and closer monitoring of latest findings. Studying the robustness of the framework and expanding it to patent documents are future tasks.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2020
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1255363
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1255363
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jrsm.1350
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 420
    Subjects:
      – SubjectFull: Innovation
        Type: general
      – SubjectFull: Biomedicine
        Type: general
      – SubjectFull: Medical Research
        Type: general
      – SubjectFull: Meta Analysis
        Type: general
      – SubjectFull: Risk
        Type: general
      – SubjectFull: Guidelines
        Type: general
      – SubjectFull: Sampling
        Type: general
      – SubjectFull: Comparative Analysis
        Type: general
      – SubjectFull: Information Retrieval
        Type: general
      – SubjectFull: Research Reports
        Type: general
      – SubjectFull: Data Analysis
        Type: general
      – SubjectFull: Outcomes of Treatment
        Type: general
    Titles:
      – TitleFull: Potential Technologies Review: A Hybrid Information Retrieval Framework to Accelerate Demand-Pull Innovation in Biomedical Engineering
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Schmitz, Tom
      – PersonEntity:
          Name:
            NameFull: Bukowski, Mark
      – PersonEntity:
          Name:
            NameFull: Koschmieder, Steffen
      – PersonEntity:
          Name:
            NameFull: Schmitz-Rode, Thomas
      – PersonEntity:
          Name:
            NameFull: Farkas, Robert
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Type: published
              Y: 2019
          Identifiers:
            – Type: issn-print
              Value: 1759-2879
          Numbering:
            – Type: volume
              Value: 10
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Research Synthesis Methods
              Type: main
ResultId 1