Application of Text Mining Techniques on Scholarly Research Articles: Methods and Tools

Saved in:
Bibliographic Details
Title: Application of Text Mining Techniques on Scholarly Research Articles: Methods and Tools
Language: English
Authors: Thakur, Khusbu (ORCID 0000-0002-9842-3452), Kumar, Vinit (ORCID 0000-0001-8306-2087)
Source: New Review of Academic Librarianship. 2022 28(3):279-302.
Availability: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
Peer Reviewed: Y
Page Count: 24
Publication Date: 2022
Document Type: Journal Articles
Reports - Research
Descriptors: Information Retrieval, Data Analysis, Research Methodology, Trend Analysis, Sample Size, Information Technology, Educational Research, Algorithms, Models, Computer Software, Programming Languages
DOI: 10.1080/13614533.2021.1918190
ISSN: 1361-4533
1740-7834
Abstract: A vast amount of published scholarly literature is generated every day. Today, it is one of the biggest challenges for organisations to extract knowledge embedded in published scholarly literature for business and research applications. Application of text mining is gaining popularity among researchers and applications are growing exponentially in different research areas. This study investigates the variety of text mining tools, techniques, sample sizes, domains and sections of the documents preferred by the text mining researchers through a systematic and structured literature review of conceptual and empirical studies. The significant findings depict that LDA and R package is the most extensively used tool and technique among the authors, most of the researchers prefer the sample size of 1000 articles for analysis, literature belonging to the domain of ICT, and related disciplines are frequently analysed in the text mining studies and abstracts constitute the corpus of the majority of text mining studies.
Abstractor: As Provided
Entry Date: 2023
Accession Number: EJ1366037
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHXykc2D8sepvrc2vXe_G68AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDBfJVcJJ6pzWhaMMuQIBEICBmwwfLkfwtt2REK8_BRkisrJAA6YnANuUfpgr9baSRjEpArjF78r2fgAw0PpVKhLtRgvkQtGMg5Fx-DioLavgJ_N8a-KpjL3x7Bjr_11fTmlLqA6X6r7appxe-i_MFyc1Pz9Lmn-upf4oHFBN-_PDHrOuPDMsx630gwBb0XrDRd_gREnjNdmic9ri9ZpcUtJXWEb1OYFE6M1gVkMU
Text:
  Availability: 1
  Value: <anid>AN0158753239;rfw15apr.22;2022Aug30.05:37;v2.2.500</anid> <title id="AN0158753239-1">Application of Text Mining Techniques on Scholarly Research Articles: Methods and Tools </title> <p>A vast amount of published scholarly literature is generated every day. Today, it is one of the biggest challenges for organisations to extract knowledge embedded in published scholarly literature for business and research applications. Application of text mining is gaining popularity among researchers and applications are growing exponentially in different research areas. This study investigates the variety of text mining tools, techniques, sample sizes, domains and sections of the documents preferred by the text mining researchers through a systematic and structured literature review of conceptual and empirical studies. The significant findings depict that LDA and R package is the most extensively used tool and technique among the authors, most of the researchers prefer the sample size of 1000 articles for analysis, literature belonging to the domain of ICT, and related disciplines are frequently analysed in the text mining studies and abstracts constitute the corpus of the majority of text mining studies.</p> <p>Keywords: knowledge discovery; Latent Dirichlet Allocation; research trends analysis; text mining; topic modelling</p> <hd id="AN0158753239-2">Introduction</hd> <p>The amount of electronic textual information available is consistently increasing (Dang & Ahmad, [<reflink idref="bib23" id="ref1">23</reflink>]). Even the scientific literature is growing and archived for the upcoming generations in digital formats (Nie & Sun, [<reflink idref="bib79" id="ref2">79</reflink>]). Along with this, efforts are going on to automatically mine this growing literature to sieve out the meaningful information out of textual information by applying text mining techniques. These techniques are deployed to extract, analyse, and find meaningful information from a set of documents (Sumathy & Chidambaram, [<reflink idref="bib98" id="ref3">98</reflink>]). Some additional applications of text mining are text summarisation, information retrieval, information clustering, text categorisation, information extraction, language identification, authorship relation identification, identifying phrase structures and detecting key phrases, extracting "entities" such as molecules, compounds, gene-sequences, names, dates, abbreviations and locating acronyms (Witten, Don, Dewsnip, & Tablan, [<reflink idref="bib108" id="ref4">108</reflink>]).</p> <p>The text mining techniques apply statistical analysis to discover hidden information from a large corpus of text, enabling us to find patterns, trends, duplication of data, and improved conceptual understanding of the unstructured text. In recent developments, text mining techniques are applied to identify prominent topics, research trends (Jeon et al., [<reflink idref="bib41" id="ref5">41</reflink>]) and even identifying the dynamic user demands(Gaikwad, Chaugule, & Patil, [<reflink idref="bib34" id="ref6">34</reflink>]). Text and data mining in the future aspires to discover new scientific knowledge by automatically analysing available literature.</p> <p>Application of text mining is gaining popularity among researchers and applications are growing exponentially in different research areas. In recent years, there have been growing applications of text mining in automatic analysis of scholarly literature. These applications aim to present an automatic summary of relevant scholarly literature to readers, eventually saving the effort of manual reading and sieving out relevant knowledge from scholarly literature. This paper highlights a careful review of studies conducted in the last ten years with applied text mining techniques on the published scholarly articles. The findings of the present study provide indications about the preferred text mining tools and techniques, the major domain areas in which studies have been performed and the sample size preferred by the researchers while mining published scholarly literature.</p> <hd id="AN0158753239-3">Applications of text mining</hd> <p>Text mining is the process of transforming unstructured text data into machine-processable structured form to discover hidden patterns, also known as a knowledge discovery database from the text (KDT), it deals with the machine learning supported analysis of the textual data. Textual data is extracted from semi-structured and unstructured datasets such as emails, full-text documents, HTML files (Yehia, Ibrahim, & Abulkhair, [<reflink idref="bib109" id="ref7">109</reflink>]). Several text mining techniques are applied to analyse the pattern in a large text corpus. Some of the applications are - document classification (text classification, document standardisation), information retrieval (keyword search/querying and indexing), document clustering (phrase clustering), natural language processing (spelling correction, lemmatisation, grammatical parsing and word sense variant), information extraction(relationship extraction/link analysis), and web link analysis (Talib, Kashif, Ayesha, & Fatima, [<reflink idref="bib101" id="ref8">101</reflink>]). This section discusses some of the applications of text mining.</p> <hd id="AN0158753239-4">Information retrieval (IR)</hd> <p>Information Retrieval is the systematic process of retrieving relevant information from a collection of information resources (Onwuchekwa & Jegede, [<reflink idref="bib81" id="ref9">81</reflink>]). Text mining has applications in the development of information retrieval systems to provide faster results based on extracted knowledge. For a collection involving unstructured documents, the text mining techniques help extract information and build structured knowledge bases leading to higher precision retrieval. Similarly, user querying behaviour analysis and trend analysis of user keywords can also be done using text mining algorithms (Talib et al., [<reflink idref="bib101" id="ref10">101</reflink>]).</p> <hd id="AN0158753239-5">Text clustering (TC)</hd> <p>Text clustering(TC) is an "unsupervised technique" applied to classify the text documents and extract exciting patterns. It has various applications of text clustering techniques such as hierarchical, distribution, centroid, k-mean (Dang & Ahmad, [<reflink idref="bib23" id="ref11">23</reflink>]). It analyses the text documents with the help of the machine learning algorithms and Natural Language Processing (NLP) techniques to understand and categorise unstructured textual data into clusters. Text clustering helps identify meaningful patterns, includes extraction of topics, retrieves the documents, and filters the information in quick and accurate sequence (Allahyari et al., [<reflink idref="bib3" id="ref12">3</reflink>]).</p> <hd id="AN0158753239-6">Text summarization (TS)</hd> <p>Text summarisation automatically reduces the length, writing style and syntax of given text based on neural networks and fuzzy clustering techniques. The principal goal of text summarisation is to shorten the original textual data into a tiny version covering all the relevant information. It automatically selects the number of indicative sentences, passages, paragraphs from the original documents and generates the text summary. This text summarisation is an old process, but nowadays, it is an emerging trend in biomedical, education, email management and blogs. In Natural Language Processing, text summarisation is a dominant area of research to solve mining challenges.</p> <hd id="AN0158753239-7">Information extraction (IE)</hd> <p>Information Extraction (IE) is a machine learning process which performs automatic extraction of the relevant 'core keywords' information from a collection of documents. The documents may be text, image, video, web, emails and structured or unstructured formats. This technique is used to identify named entities, contact information, remove unwanted data, and organise, manage, and analyse the relevant data. Information extraction sieves out quality and accurate information from a large dataset of unstructured text (Vidhya & Aghila, [<reflink idref="bib103" id="ref13">103</reflink>]).</p> <hd id="AN0158753239-8">Text mining algorithms</hd> <p>Text mining algorithms are the statistical representatives of machine learning. It is applied to volumes of polygonal textual data that have been transformed from natural language to structured, numerical form (Miner et al., [<reflink idref="bib74" id="ref14">74</reflink>]). All these techniques have common goals to discover hidden information, trends, patterns to help in decision making. Before discussing the different text mining techniques deployed by text mining researchers to mine scholarly research output, a brief explanation about some of the popular text mining algorithms are outlined below:</p> <hd id="AN0158753239-9">Latent Dirichlet allocation (LDA)</hd> <p>LDA is an unsupervised learning technique, which is used for topic modelling. Topic modelling is a statistical process to identify topics from text corpora. It synthesises the correlations between the topics and words present in the collection of documents. Latent Dirichlet allocation (LDA) is a probabilistic model for identifying latent semantic topics from extensive text corpora collection. It provides the probability of each document to a set of topics based on a three-level hierarchical Bayesian model. Using LDA, the underlying topics of a document can automatically be identified based on the content of documents. It is crucial to distinguish LDA from a simple Dirichlet-multinomial clustering model. A classical clustering model would involve a two-level model in which a Dirichlet is sampled once for a corpus, a multinomial clustering variable is selected once for each document in the corpus, and a set of words are selected for the document conditional on the cluster variable. On the other hand, LDA involves three levels, and notably, the topic node is sampled repeatedly within the document. Under this model, documents can be associated with multiple topics (Blei, Ng, & Jordan, [<reflink idref="bib9" id="ref15">9</reflink>]). LDA has various extended applications in other fields such as NLP, speech recognition, spam filtering, web mining, and video analysis (Diane & Lawrence, [<reflink idref="bib26" id="ref16">26</reflink>]).</p> <hd id="AN0158753239-10">Naive Bayes classifier(NBC)</hd> <p>Naive Bayes Classifier is the probabilistic text mining algorithm. This algorithm is generally classified as a binary classification problem. It is one of the best techniques for forecast analysis and gives accurate results. NBC is a text classification technique, having various applications such as finding relevant documents, classification, categorisation, age/gender identification, language detection and trend analysis (Zhang, [<reflink idref="bib113" id="ref17">113</reflink>]). However, it is primarily used for text classification and categorisation (Behera & Kumar, [<reflink idref="bib7" id="ref18">7</reflink>]).</p> <hd id="AN0158753239-11">Support vector machines (SVM)</hd> <p>SVM is a "supervised machine learning algorithm" used for classification and regression analysis such as detecting spam, sentiment analysis, and document classification into categories like news, emails, articles, and web pages(Brereton & Lloyd, [<reflink idref="bib10" id="ref19">10</reflink>]). This algorithm draws a line which is known as a "hyperplane".This algorithm's principal aim is to divide the dataset accurately and find out the "maximum marginal hyperplane" (Brown & Lewis, [<reflink idref="bib11" id="ref20">11</reflink>]). SVM models are also applied to classify unlabelled dataset such as text and image classification, handwriting recognition, biometric authentication detection analysis.</p> <hd id="AN0158753239-12">Association rules</hd> <p>Association Rules is based on machine learning methods and applied to find out the interesting pattern and correlation between a large number of variables in a dataset. The widespread applications of association rules are basket analysis, cross-marketing clustering, classification, catalogue and design (Hahsler & Karpienko, [<reflink idref="bib38" id="ref21">38</reflink>]). Using this text mining algorithm, one can discover the interesting patterns, gain in-depth knowledge about the selected data in an extensive collection of datasets.</p> <hd id="AN0158753239-13">K-Nearest neighbour(KNN)</hd> <p>The k-Nearest Neighbour is a "supervised classification algorithm" which is used for text categorisation. It is deployed to classify new documents among the dataset, k-Nearest Neighbour is a straightforward algorithm and non-parametric process which learns by itself from past experiences. In addition, this algorithm classifies new points based on the concepts of similarity and the value of k-operation items of the data. The main goal of the KNN is searching for similar documents among a large dataset of items (Li, Yu, & Lu, [<reflink idref="bib69" id="ref22">69</reflink>]).</p> <hd id="AN0158753239-14">Neural networks</hd> <p>A Neural Network is an artificial neural network inspired by the functioning of the biological human brain, in general theory, it works like a network of neurons. A neural network has various nodes which are connected with each other and has three layers: "input layer", "intermediate layer" and"outer layer". The input layer receives the signals from the outer layer as well as an intermediate layer comprising the neurons and gives output. The neural network technique classifies the text automatically. These artificial networks are specially used for predictive modelling and statistical analysis where they are trained through datasets. This neural network model has solved several problems in areas such as technology, statistics, and economics, which records the data at a time and learns by comparing their classification with the authentic sources (Prasanna & Rao, [<reflink idref="bib83" id="ref23">83</reflink>]).</p> <hd id="AN0158753239-15">Decision trees</hd> <p>Decision Trees techniques is a machine learning algorithm typically used to predict the value of the target variable by learning the rules of decision trees from the training data. This predictive modelling has applications in statistics, text mining as well as in machine learning. Decision trees are also known as a method of classification (Song & Ying, [<reflink idref="bib95" id="ref24">95</reflink>]).</p> <hd id="AN0158753239-16">Text mining tools</hd> <p>The text mining techniques mentioned in the previous section are deployed in a research problem using tools that implement these techniques and provide input mechanisms, data processing, and visualisation of results. In this section, a brief detail about some of the popular tools researchers use for text mining are discussed.</p> <hd id="AN0158753239-17">R-package</hd> <p>The R-package is one of the most popular open source software for text analysis. It contains a set of library packages to perform NLP operations, sentiment analysis, topic modelling, determining word frequencies, and forming word clouds. It provides packages such as NLTK, with features for tokenization, lemmatisation, stemming, preprocessing of text, and implementations for deploying, LDA, association rules, classification algorithms and other algorithms. This software demonstrates quantitative results visualised into statistical graphics of textual data (Feinerer, [<reflink idref="bib31" id="ref25">31</reflink>]).</p> <hd id="AN0158753239-18">SAS enterprise</hd> <p>Statistical Analysis System(SAS) is a proprietary software for data processing and statistical analysis. It is a programming software that can process a high volume of data for performing text analysis. It was primarily developed for data retrieval, organising data and performing text analytics and predictive analytics. It is the industry-leading tool for business analytics which helps in statistical analysis (Jotsov & Iliev, [<reflink idref="bib47" id="ref26">47</reflink>]).</p> <hd id="AN0158753239-19">Python</hd> <p>Python is one of the fastest-growing programming languages that contain inbuilt and third-party libraries and packages for performing mining of text. Python has an NLTK package containing numbers NLP methods like sentence and word tokenisation, part of speech tagging, chunking and classification. Natural Language Toolkit is a suite of libraries and programs for statistical natural language processing for python programming languages. Many organisations and companies use this tool for multiple studies such as Topic modelling, NLP, web developing, software making, artificial intelligence, business analytics, market analysis and scientific computing. The packages like Gensim and Spacy are deployed for advanced text mining applications.</p> <hd id="AN0158753239-20">RapidMiner</hd> <p>RapidMiner is a data science tool primarily deployed by enterprises for machine learning and data mining tasks. RapidMiner provides a "drag and drop" user interface to design the workflow of the analytics process. It can connect with datasets stored in relational databases, or spreadsheets or statistical package-specific file formats. RapidMiner uses XML to describe the process of knowledge discovery with many learning algorithms from WEKA. This software helps in data mining tasks to solve business problems and predict hot trends (Kalra & Aggarwal, [<reflink idref="bib48" id="ref27">48</reflink>]) and text mining applications like sentiment analysis, fraud detection, and advertisement profiling.</p> <hd id="AN0158753239-21">Mallet</hd> <p>MALLET is an open-source software written in Java and has several machine learning applications to text mining such as text extraction, classification, clustering, topic modelling, NLP and other tasks. It implements a variety of algorithms like Naïve Bayes, Maximum Entropy, and Decision Trees (Lamba & Madhusudhan, [<reflink idref="bib61" id="ref28">61</reflink>]).</p> <hd id="AN0158753239-22">KH coder</hd> <p>KH Coder is free software which is used for text analysis and other computational linguistics operations. It provides various kinds of search and statistical analysis function back-end tools such as snowball stemmer with ability to connect with MYSQL and R. It also has capability to analyse documents from different languages like - Japanese, English, French, German, Italian, Portuguese and Spanish. It provides functions to build KWIC, Collocation, co-occurrence network, cluster analysis and multidimensional scaling (Zheng, Liang, Huang, & Liu, [<reflink idref="bib71" id="ref29">71</reflink>]).</p> <p>Among the available variety of text mining algorithms and tools the text mining researchers face a tough challenge in selecting appropriate tool and text mining technique while analysing a collection of text to tasks such as text summarisation, information retrieval, information clustering, text categorisation, information extraction, language identification, authorship relation identification, identifying phrase structures and detecting key phrases, extracting "entities" such as molecules, compounds, gene-sequences, names, dates, abbreviations and locating acronyms. The present study aims to investigate the popular text mining techniques and tools used by researchers while mining scholarly research output.</p> <hd id="AN0158753239-23">Research questions</hd> <p>The present study investigates the following research questions:</p> <p></p> <ulist> <item> <emph>RQ1</emph>: What are the tools and techniques of text mining frequently deployed by text mining researchers in mining scholarly literature?</item> <p></p> <item> <emph>RQ2</emph>: What are the prominent domain areas in which text mining researchers have majorly applied text mining techniques?</item> <p></p> <item> <emph>RQ3</emph>: What is the most preferred sample size selected by text mining researchers when applying text mining techniques to scholarly literature?</item> <p></p> <item> <emph>RQ4</emph>: Which are the notable countries where text mining researchers have published studies attempting text mining of scholarly articles?</item> <p></p> <item> <emph>RQ5</emph>: Which sections of the scholarly articles have been selected by the authors to form the corpus while text mining the scholarly literature?</item> </ulist> <hd id="AN0158753239-24">Methods</hd> <p>The data for this study has been collected from scholarly databases such as Scopus, Google Scholar, IEEE, Europe PMC, ACM, Korean Society for Internet Information using the terms, ("text mining" OR "text analysis" OR "topic modelling" OR "knowledge discovery in databases") AND ("scholarly articles" OR "research papers" OR "research articles") during March 2019 to May 2020. A total number of 426 research articles were retrieved from Scopus and other scholarly databases, which were critically analysed by screening the title, abstract, methodology, tools and techniques used by the authors in their papers. The data collected was objectively reviewed by the researcher, keeping in mind the factors such as scope, objectives and nature of studies. After the critical review of articles the relevant studies which focussed on the text mining of scholarly output were considered for further detailed review and analysis. A total of 146 research articles were excluded from the study as they were theoretical in nature and were largely based on identifying social media trends, comparing text mining tools, analysing news articles and covering the general overview of text mining applications, tools and techniques. Similarly, the articles covering the challenges of text mining were also excluded from the final set of selected articles.</p> <p>After the preliminary review of collected literature, 85 articles were found relevant to the scope of the study that have applied text mining in the scholarly output(articles and research papers) which were further considered for analysis. Fig. 1, explains the exclusion and inclusion criteria using the PRISMA flow model.</p> <p>Graph: Figure 1. Exclusion and inclusion of studies for this review.</p> <hd id="AN0158753239-25">Analysis</hd> <p>The qualitative and quantitative analysis of selected articles was performed in five phases; In the initial phase, the collected articles were assessed on the basis of the text mining techniques reported in the research paper. In the second phase, the selected articles were analysed based on the text mining tool deployed by the authors to implement the chosen text mining techniques. In the third phase of the review, the selected articles were segregated based on the number of documents taken as samples by the authors of the selected papers. In the fourth phase, the subject domains of the selected studies were established from which the sample documents were picked by the authors to conduct text mining. In the fifth step, the country of origin of the primary author of the chosen studies was established based on their affiliation. In the final phase, the sections of the scholarly literature selected by the authors of the selected studies to extract textual data for the text analysis were identified and recorded.</p> <hd id="AN0158753239-26">Results</hd> <p>The first research question of the study was to examine the prominent text mining techniques used by different scholars in their respective studies on text mining of scholarly literature.</p> <p>Table 1 represents the major techniques used by various authors in their respective research studies related to text mining of scholarly literature. LDA was found to be used by most (29.5%) scholars in their text analysis work, while 17% of the researchers used the k-means clustering algorithm followed by the NLP technique by 3.4%, the Association method by 5.1% and the SVM technique by 4.25%. The results indicate that the LDA is mostly deployed by the researchers whereas the Support Vector Machine algorithm is least frequently used for text mining of scholarly research output.</p> <p>Table 1. Represents the use of the text mining technique by the authors in their research work.</p> <p> <ephtml> <table><thead><tr><td>Techniques</td><td>Number of studies (%)</td><td>Sources</td></tr></thead><tbody valign="top"><tr><td>Latent Dirichlet Allocation (LDA)</td><td char=".">35 (29.75%)</td><td>(Dayeen, Sharma, & Derrible, <xref ref-type="bibr" rid="bibr25">2020</xref>; García et al., <xref ref-type="bibr" rid="bibr35">2020</xref>; Kang, Kim, & Kang, <xref ref-type="bibr" rid="bibr49">2019</xref>; Yoon & Suh, <xref ref-type="bibr" rid="bibr111">2019</xref>; Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr61">2018a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr62">2018b</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr63">2019a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr64">2019b</xref>; Westergaard, Staerfeldt, Tønsberg, Jensen, & Brunak, <xref ref-type="bibr" rid="bibr106">2018</xref>; Kirilenko & Stepchenkova, <xref ref-type="bibr" rid="bibr54">2018</xref>; Jia & Wu, <xref ref-type="bibr" rid="bibr43">2018</xref>; Joo, Choi, & Choi, <xref ref-type="bibr" rid="bibr46">2018</xref>; Syed, Borit, & Spruit, <xref ref-type="bibr" rid="bibr100">2018</xref>; Jeon et al., <xref ref-type="bibr" rid="bibr41">2018</xref>; Lim & Maglio, <xref ref-type="bibr" rid="bibr70">2018</xref>; Liu, Tang, Dong, Yao, & Zhou, <xref ref-type="bibr" rid="bibr71">2016</xref>; Kurata et al., <xref ref-type="bibr" rid="bibr58">2018</xref>; Zou, <xref ref-type="bibr" rid="bibr115">2018</xref>; Cortez, Moro, Rita, King, & Hall, <xref ref-type="bibr" rid="bibr22">2018</xref>; Ding, Li, & Fan, <xref ref-type="bibr" rid="bibr27">2018</xref>; Sun & Yin, <xref ref-type="bibr" rid="bibr99">2017</xref>; Moro, Alturas, Esmerado, & Costa, <xref ref-type="bibr" rid="bibr76">2017</xref>; Nie & Sun, <xref ref-type="bibr" rid="bibr79">2017</xref>; Kim, Kwahk, & Yoon, <xref ref-type="bibr" rid="bibr51">2017a</xref>; Cho, Bae, & Woo, <xref ref-type="bibr" rid="bibr17">2017</xref>; Wang et al., <xref ref-type="bibr" rid="bibr104">2016</xref>; Das, Sun, & Dutta, <xref ref-type="bibr" rid="bibr24">2016</xref>; Wang, Bowers, & Fikis, <xref ref-type="bibr" rid="bibr105">2017</xref>; Jeong & Lee, <xref ref-type="bibr" rid="bibr42">2016</xref>; Chen, Wang, & Lu, <xref ref-type="bibr" rid="bibr13">2016</xref>; Moro, Cortez, & Rita, <xref ref-type="bibr" rid="bibr77">2015</xref>; Do & Skłodowski, <xref ref-type="bibr" rid="bibr28">2014</xref>; Sugimoto, Li, Russell, Finlay, & Ding, <xref ref-type="bibr" rid="bibr96">2011</xref>; Yu & Ku, <xref ref-type="bibr" rid="bibr112">2010</xref>)</td></tr><tr><td>k-Means clustering</td><td char=".">20 (17%)</td><td>(Gülkesen & Haux, <xref ref-type="bibr" rid="bibr36">2019</xref>; Ferreira‐Mello, André, Pinheiro, Costa, & Romero, <xref ref-type="bibr" rid="bibr33">2019</xref>; Kang et al., <xref ref-type="bibr" rid="bibr49">2019</xref>; Abu-Shanab & Harb, <xref ref-type="bibr" rid="bibr1">2019</xref>; Bhanot, Singh, Sharma, Jain, & Jain, <xref ref-type="bibr" rid="bibr8">2019</xref>; Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr61">2018a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr64">2019b</xref>; Salloum, Al-Emran, Monem, & Shaalan, <xref ref-type="bibr" rid="bibr85">2018</xref>; Jitngernmadan & Boonmee, <xref ref-type="bibr" rid="bibr45">2018</xref>; Syed et al., <xref ref-type="bibr" rid="bibr100">2018</xref>; Gurung & Wagh, <xref ref-type="bibr" rid="bibr37">2017</xref>; Konstantinidis, Billis, Wharrad, & Bamidis, <xref ref-type="bibr" rid="bibr56">2017</xref>; Aich, Sain, Park, Choi, & Kim, <xref ref-type="bibr" rid="bibr2">2017</xref>; Wang et al., <xref ref-type="bibr" rid="bibr104">2016</xref>; Zheng et al., <xref ref-type="bibr" rid="bibr114">2016</xref>; White, Guldiken, Hemphill, He, & Sharifi Khoobdeh, <xref ref-type="bibr" rid="bibr107">2016</xref>; Cho & Kim <xref ref-type="bibr" rid="bibr18">2012</xref>; Hung, <xref ref-type="bibr" rid="bibr39">2012</xref>; Hung & Zhang, <xref ref-type="bibr" rid="bibr40">2012</xref>; Lee et al. <xref ref-type="bibr" rid="bibr67">2010</xref>)</td></tr><tr><td>Association rules</td><td char=".">6 (5.1%)</td><td>(Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr61">2018a</xref>; Salloum et al., <xref ref-type="bibr" rid="bibr85">2018</xref>; Jeon et al., <xref ref-type="bibr" rid="bibr41">2018</xref>; Jitngernmadan & Boonmee, <xref ref-type="bibr" rid="bibr45">2018</xref>; Fang, Yang, Gao, & Li, <xref ref-type="bibr" rid="bibr30">2018</xref>; Kim, Ohk, & Moon, <xref ref-type="bibr" rid="bibr52">2017b</xref>)</td></tr><tr><td>Support Vector Machine</td><td char=".">5(4.25%)</td><td>(Ferreira‐Mello et al., <xref ref-type="bibr" rid="bibr33">2019</xref>; Sharma, Kumar, & Chand, <xref ref-type="bibr" rid="bibr87">2018</xref>; Sulova, Todoranova, Penchev, & Nacheva, <xref ref-type="bibr" rid="bibr97">2017</xref>; O'Mara-Eves, Thomas, McNaught, Miwa, & Ananiadou, <xref ref-type="bibr" rid="bibr80">2015</xref>;Cheng, Yang, & Zhang, <xref ref-type="bibr" rid="bibr15">2015</xref>)</td></tr><tr><td>Natural Language Processing</td><td char=".">4(3.4%)</td><td>(Gülkesen & Haux, <xref ref-type="bibr" rid="bibr36">2019</xref>; Ferreira‐Mello et al., <xref ref-type="bibr" rid="bibr33">2019</xref>; Kim et al., <xref ref-type="bibr" rid="bibr50">2019</xref>; Syed et al., <xref ref-type="bibr" rid="bibr100">2018</xref>)</td></tr></tbody></table> </ephtml> </p> <hd id="AN0158753239-27">Text mining tools</hd> <p>Another objective of this study was to examine the variety of tools used to implement text mining techniques by researchers applying text mining techniques to achieve their research objectives.</p> <p>Table 2 indicates that R-Packages were used by 15.3% of selected studies followed by Python by 7.6%, RapidMiner by 5.1% and MALLET by 4.25%, SAS Enterprise by 3.4% of the selected studies.</p> <p>Table 2. Represents the tools used by the authors in their research work.</p> <p> <ephtml> <table><thead><tr><td>Tools</td><td>Number of studies (%)</td><td>Sources</td></tr></thead><tbody valign="top"><tr><td>R-Packages</td><td char=".">18 (15.3%)</td><td>(Mandujano, <xref ref-type="bibr" rid="bibr73">2019</xref>; Koch, <xref ref-type="bibr" rid="bibr55">2019</xref>; Bhanot et al., <xref ref-type="bibr" rid="bibr8">2019</xref>; Jeon et al., <xref ref-type="bibr" rid="bibr41">2018</xref>; Cortez et al., <xref ref-type="bibr" rid="bibr22">2018</xref>; Amado, Cortez, Rita, & Moro, <xref ref-type="bibr" rid="bibr4">2018</xref>; Ding et al., <xref ref-type="bibr" rid="bibr27">2018</xref>; Fang et al., <xref ref-type="bibr" rid="bibr30">2018</xref>; Kim et al., <xref ref-type="bibr" rid="bibr52">2017b</xref>; Shankaranarayanan & Blake <xref ref-type="bibr" rid="bibr86">2017</xref>; Moro et al., <xref ref-type="bibr" rid="bibr76">2017</xref>; Carnerud, <xref ref-type="bibr" rid="bibr12">2017</xref>; Cho et al., <xref ref-type="bibr" rid="bibr17">2017</xref>; Aich et al., <xref ref-type="bibr" rid="bibr2">2017</xref>; Shinde, Oza, & Kamat, <xref ref-type="bibr" rid="bibr90">2017</xref>; Das et al., <xref ref-type="bibr" rid="bibr24">2016</xref>; Wang et al., <xref ref-type="bibr" rid="bibr105">2017</xref>; Moro et al., <xref ref-type="bibr" rid="bibr77">2015</xref>)</td></tr><tr><td>SAS Enterprise</td><td char=".">4 (3.4%)</td><td>(Kim et al., <xref ref-type="bibr" rid="bibr51">2017a</xref>; Kim & Delen, <xref ref-type="bibr" rid="bibr53">2018</xref>; Hung, <xref ref-type="bibr" rid="bibr39">2012</xref>; Hung & Zhang, <xref ref-type="bibr" rid="bibr40">2012</xref>)</td></tr><tr><td>Python</td><td char=".">9 (7.65%)</td><td>(Dayeen et al., <xref ref-type="bibr" rid="bibr25">2020</xref>; Comeau, Wei, Islamaj Doğan, & Lu, <xref ref-type="bibr" rid="bibr21">2019</xref>; Abu-Shanab & Harb, <xref ref-type="bibr" rid="bibr1">2019</xref>; Bhanot et al., <xref ref-type="bibr" rid="bibr8">2019</xref>; Lim & Maglio, <xref ref-type="bibr" rid="bibr70">2018</xref>; Kurata et al., <xref ref-type="bibr" rid="bibr58">2018</xref>; Sharma et al., <xref ref-type="bibr" rid="bibr87">2018</xref>; Syed et al., <xref ref-type="bibr" rid="bibr100">2018</xref>; Cho et al., <xref ref-type="bibr" rid="bibr17">2017</xref>)</td></tr><tr><td>RapidMiner</td><td char=".">6 (5.1%)</td><td>(Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr63">2019a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr64">2019b</xref>, Lamba, <xref ref-type="bibr" rid="bibr60">2019</xref>; Salloum et al., <xref ref-type="bibr" rid="bibr85">2018</xref>; Jitngernmadan & Boonmee, <xref ref-type="bibr" rid="bibr45">2018</xref>; Sulova et al., <xref ref-type="bibr" rid="bibr97">2017</xref>)</td></tr><tr><td>MALLET</td><td char=".">5 (4.25%)</td><td>(Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr61">2018a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr64">2019b</xref>; Sun & Yin, <xref ref-type="bibr" rid="bibr99">2017</xref>; Wang et al., <xref ref-type="bibr" rid="bibr104">2016</xref>; Do & Skłodowski, <xref ref-type="bibr" rid="bibr28">2014</xref>)</td></tr><tr><td>KHcoder3</td><td char=".">2 (1.7%)</td><td>(Park & Park, <xref ref-type="bibr" rid="bibr82">2019</xref>; Zheng et al., <xref ref-type="bibr" rid="bibr114">2016</xref>)</td></tr></tbody></table> </ephtml> </p> <p>It can be said that the R-packages are deployed by most of the authors in their research work, whereas the SAS Enterprise and KHcoder3 are deployed by very few researchers for mining the scholarly literature.</p> <hd id="AN0158753239-28">Sample size</hd> <p>The text mining analysis results depend on the sample size of the documents chosen for the analysis. A sample of inappropriate sample size will yield results that cannot be generalised. Thus it is significant to analyse the sample size chosen by different researchers mining scholarly research outputs. Table 3 represents the number of documents analysed by authors of the selected studies, the table indicates that over 28.90% of the selected studies have sample sizes of less than or equal to 1000, whereas 11.05% of studies have a sample size of more than 1000 but less than 2000. The studies having sample size above 2000 is found to be 31.45%. Also, it was found that one study had sample sizes of 15 million full-text scientific articles.</p> <p>Table 3. Represents the number of documents analysed by the authors</p> <p> <ephtml> <table><thead><tr><td>No. of Documents</td><td>No. of studies(%)</td></tr></thead><tbody valign="top"><tr><td>1-1000</td><td char=".">34 (28.9%)</td></tr><tr><td>1001-2000</td><td char=".">13 (11.05%)</td></tr><tr><td>2001- Above</td><td char=".">37 (31.45%)</td></tr></tbody></table> </ephtml> </p> <p>It can be said that most of the studies in the selected group of articles found to be having sample sizes below 1000, which is considered a small sample size for text analysis. Although a good share of studies have sample size above 2000 which depicts that most of the researchers are testing with smaller sample sizes while few of them are applying text mining on larger datasets too. Another point that must be noted is that the scholarly research articles apart from the title and abstract sections are unstructured and have low uniformity thus requiring a great level of effort from the side of the researcher to pre-process them and create datasets. This may be the reason for choosing smaller sample sizes by most of the researchers.</p> <hd id="AN0158753239-29">Preferred scholarly domains</hd> <p>Another research question was to study the prominent domain areas in which researchers have tried to understand patterns of research trends, hot topics, and entity analysis applying text mining techniques. Table 4 indicates that 19.5% of selected studies belong to scholarly outputs of Information Communication and Technology domain followed by Biomedical sciences with 14.4% of the studies, while studies with samples taken from the Management domain have a share of 10.2%, and the domain of the Environmental sciences has 9.3% share of studies. Similarly, 8.5% of the studies selected in this study fell under the Library and Information Science domain.</p> <p>Table 4. Represents the research domains selected by the authors.</p> <p> <ephtml> <table><thead><tr><td>Domain</td><td>Number of studies(%)</td><td>Sources</td></tr></thead><tbody valign="top"><tr><td>Information Communication Technology</td><td char=".">23 (19.55%)</td><td>(García et al., <xref ref-type="bibr" rid="bibr35">2020</xref>; Shin & Suh, <xref ref-type="bibr" rid="bibr89">2019</xref>; Kim et al., <xref ref-type="bibr" rid="bibr50">2019</xref>; Chen, Xu, Jin, & Wanatowski, <xref ref-type="bibr" rid="bibr14">2019</xref>; Ferreira‐Mello et al., <xref ref-type="bibr" rid="bibr33">2019</xref>; Chen et al., <xref ref-type="bibr" rid="bibr14">2019</xref>; Dreisbach, Koleck, Bourne, & Bakken, <xref ref-type="bibr" rid="bibr29">2019</xref>; Abu-Shanab & Harb, <xref ref-type="bibr" rid="bibr1">2019</xref>; Usai, Pironti, Mital, & Mejri, <xref ref-type="bibr" rid="bibr102">2018</xref>; Jindal & Shweta, <xref ref-type="bibr" rid="bibr44">2018</xref>; Sharma et al., <xref ref-type="bibr" rid="bibr87">2018</xref>; Liu et al., <xref ref-type="bibr" rid="bibr71">2016</xref>; Salloum et al., <xref ref-type="bibr" rid="bibr85">2018</xref>; Lim & Maglio, <xref ref-type="bibr" rid="bibr70">2018</xref>; Cortez et al., <xref ref-type="bibr" rid="bibr22">2018</xref>; Gurung & Wagh, <xref ref-type="bibr" rid="bibr37">2017</xref>; Chen et al., <xref ref-type="bibr" rid="bibr13">2016</xref>; Ammarukleart & Kim, <xref ref-type="bibr" rid="bibr5">2017</xref>; Zheng et al., <xref ref-type="bibr" rid="bibr114">2016</xref>; Moro et al., <xref ref-type="bibr" rid="bibr77">2015</xref>; Hung, <xref ref-type="bibr" rid="bibr39">2012</xref>; Hung & Zhang, <xref ref-type="bibr" rid="bibr40">2012</xref>; Lee et al., <xref ref-type="bibr" rid="bibr67">2010</xref>)</td></tr><tr><td>Biomedical Sciences</td><td char=".">17 (14.45%)</td><td>(Shin & Suh, <xref ref-type="bibr" rid="bibr89">2019</xref>; Gülkesen & Haux, <xref ref-type="bibr" rid="bibr36">2019</xref>; Comeau et al., <xref ref-type="bibr" rid="bibr21">2019</xref>; Smink et al., <xref ref-type="bibr" rid="bibr91">2019</xref>; Park & Park, <xref ref-type="bibr" rid="bibr82">2019</xref>; Krishnamurthy & Balasubramanium, <xref ref-type="bibr" rid="bibr57">2019</xref>; Son & Kang, <xref ref-type="bibr" rid="bibr92">2019</xref>; Yoon & Suh, <xref ref-type="bibr" rid="bibr111">2019</xref>; Zou, <xref ref-type="bibr" rid="bibr115">2018</xref>; Kim & Delen, <xref ref-type="bibr" rid="bibr53">2018</xref>; Konstantinidis et al., <xref ref-type="bibr" rid="bibr56">2017</xref>; Lam et al., <xref ref-type="bibr" rid="bibr59">2016</xref>; Wang et al., <xref ref-type="bibr" rid="bibr104">2016</xref>; Liu et al., <xref ref-type="bibr" rid="bibr71">2016</xref>; Yoo, Shin, Yoo, & Shin, <xref ref-type="bibr" rid="bibr16">2015</xref>; Choi & Lee, <xref ref-type="bibr" rid="bibr16">2015</xref>; Song & Kim, <xref ref-type="bibr" rid="bibr94">2013</xref>)</td></tr><tr><td>Management</td><td char=".">12 (10.2%).</td><td>(Choo, Park, Kim, & Seo, <xref ref-type="bibr" rid="bibr20">2019</xref>; Joo et al., <xref ref-type="bibr" rid="bibr46">2018</xref>; Kirilenko & Stepchenkova, <xref ref-type="bibr" rid="bibr54">2018</xref>; Jitngernmadan & Boonmee, <xref ref-type="bibr" rid="bibr45">2018</xref>; Fang et al., <xref ref-type="bibr" rid="bibr30">2018</xref>; Amado et al., <xref ref-type="bibr" rid="bibr4">2018</xref>; Cho et al., <xref ref-type="bibr" rid="bibr17">2017</xref>; Sulova et al., <xref ref-type="bibr" rid="bibr97">2017</xref>; White et al., <xref ref-type="bibr" rid="bibr107">2016</xref>; Nagarkar & Kumbhar, <xref ref-type="bibr" rid="bibr78">2015</xref>; Ravikumar, Agrahari, & Singh, <xref ref-type="bibr" rid="bibr84">2015</xref>; Song, Park, Jung, & Song, <xref ref-type="bibr" rid="bibr93">2013</xref>).</td></tr><tr><td>Environmental sciences</td><td char=".">11(9.35%)</td><td>(Dayeen et al., <xref ref-type="bibr" rid="bibr25">2020</xref>; Maddi, Sapinho, & Baudoin, <xref ref-type="bibr" rid="bibr72">2019</xref>; Mandujano, <xref ref-type="bibr" rid="bibr73">2019</xref>; Koch, <xref ref-type="bibr" rid="bibr55">2019</xref>; Ding et al., <xref ref-type="bibr" rid="bibr27">2018</xref>; Jeon et al., <xref ref-type="bibr" rid="bibr41">2018</xref>; Syed et al., <xref ref-type="bibr" rid="bibr100">2018</xref>; Lee & Lee, <xref ref-type="bibr" rid="bibr68">2018</xref>; Jia & Wu, <xref ref-type="bibr" rid="bibr43">2018</xref>; Lee, Song, & Lee, <xref ref-type="bibr" rid="bibr66">2013</xref>; Shianghau & Jiannjong, <xref ref-type="bibr" rid="bibr88">2010</xref>)</td></tr><tr><td>Library and Information Science</td><td char=".">10(8.5%)</td><td>(Miyata et al., <xref ref-type="bibr" rid="bibr75">2020</xref>; Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr65">2020</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr63">2019a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr64">2019b</xref>, Lamba, <xref ref-type="bibr" rid="bibr60">2019</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr61">2018a</xref>, Lamba & Madhusudhan, <xref ref-type="bibr" rid="bibr62">2018b</xref>; Ferran-Ferrer, Guallar, Abadal, & Server, <xref ref-type="bibr" rid="bibr32">2017</xref>; Cho, Lim, & Hur, <xref ref-type="bibr" rid="bibr19">2014</xref>; Sugimoto et al., <xref ref-type="bibr" rid="bibr96">2011</xref>)</td></tr></tbody></table> </ephtml> </p> <p>The analysis indicates that researchers have preferred scholarly output from the ICT domain to analyse data actively in different fields such as business, machines and education. Which can be explained by the fact that the researchers of the ICT domain are traditionally more versed with the text mining techniques and have chosen their own disciplines for studying the research trends and hot topics. Another reason could be that having domain knowledge helps in interpreting the results of the analysis, thus mining documents from one's own discipline will have better interpretations. This is also true for other findings too, in the case of biomedical sciences and Management.</p> <hd id="AN0158753239-30">Notable countries</hd> <p>The affiliation of the first author of each selected study was identified and presented in Table 5 in order to determine the notable countries where text mining researchers have published studies with the goal of applying text mining to scholarly output.</p> <p>Table 5. Represents the countries published research articles on TM.</p> <p> <ephtml> <table><thead><tr><td>S. No.</td><td>Countries</td><td>No of studies (%)</td></tr></thead><tbody valign="top"><tr><td>1</td><td>USA</td><td char=".">21(17.85%)</td></tr><tr><td>2</td><td>Korea</td><td char=".">11(9.35%)</td></tr><tr><td>3</td><td>India</td><td char=".">7(5.95%)</td></tr><tr><td>4</td><td>U. K.</td><td char=".">6(5.1%)</td></tr><tr><td>6</td><td>China</td><td char=".">5(4.25%)</td></tr></tbody></table> </ephtml> </p> <p>Table 5 reveals that the USA has the largest number of publications among selected research articles with a 17.8% share of selected studies followed by Korea 9.3%, India 5.95%, U. K. 5.1% and China 4.25%. China had the lowest number of publications of research articles on the application of text mining techniques to scholarly literature.</p> <hd id="AN0158753239-31">Text corpus</hd> <p>It was interesting to know, which parts of the scholarly literature constitute the corpus for the studies mining the scholarly research output.</p> <p>Table 6 indicates that most of the studies(20.4%) preferred the abstract section of the research papers to form their corpus followed by the inclusion of the full text of the research articles(17.8%), while 10.2% of studies included the title, abstract and keywords and 8.5% of studies conducted text mining on the corpus containing only the title and abstract.</p> <p>Table 6. Represents the corpus of research articles used by the authors in their articles.</p> <p> <ephtml> <table><thead><tr><td>No.</td><td>Corpus</td><td>No. of Studies (%)</td></tr></thead><tbody valign="top"><tr><td>1</td><td>Abstracts only</td><td>24(20.4%)</td></tr><tr><td>2</td><td>Full Texts only</td><td>21(17.85%)</td></tr><tr><td>2</td><td>Title and Abstract</td><td>10(8.5%)</td></tr><tr><td>4</td><td>Title, Abstract and Keywords</td><td>12(10.2%)</td></tr></tbody></table> </ephtml> </p> <hd id="AN0158753239-32">Discussion</hd> <p>Text mining tools and techniques have promising applications in the summarisation, finding research trends, topic modelling and document classification. The recent development in the field of distance reading is one of the examples of the application of text mining. The results of the study indicate the highly popular tools, text mining techniques, prefered sample size and the domains where text mining is applied in scholarly literature. The study found that the LDA technique is preferred by the majority of the authors in their studies. This may be attributed to LDA's ability to perform topic modelling based on three-level hierarchical Bayesian models to discover the hidden information in textual data and thereby provide reliable results. This algorithm has multiple features available and also maintains the flexibility to generate the particular concepts, themes or topics but does not provide the full sense of the textual data (Mandujano, [<reflink idref="bib73" id="ref30">73</reflink>]). This finding is in line with Asmussen and Møller ([<reflink idref="bib6" id="ref31">6</reflink>]) that also reported that the LDA is an algorithm of choice for topic modelling irrespective of the sample size.</p> <p>Similarly, in the case of tools, the R-packages are deployed by the majority of the authors in their research work. This may be due to the recent surge in the awareness of R programming language and free availability of several packages implementing major text mining algorithms. In terms of the sample size, the study found that half of the selected studies have sample sizes of less than or equal to 1000 documents. Since the results of text mining highly depend on the sample size of the documents, the results of the studies must be taken into consideration considering the sample sizes. The selection of a small sample size shows that most of the studies are experimental in nature wherein the authors are testing the algorithms on smaller manageable sample size as in small sample sizes, it is trivial to extract the structured text using preprocessing pipelines. Although a small number of authors had conducted text mining on larger sample sizes too. While analysing the prefered domains, it is found that studies have been done in a variety of domains but in this paper, we found top five domains which are highly prefered in scholarly text mining literature such as Library and Information Science, Biomedical Sciences, Information Communication Technology, Environmental Sciences and Management. The results indicate that the majority of the text mining researchers have preferred scholarly output from ICT domain, which can be explained by the fact that the researchers of ICT domain are traditionally more versed with the text mining techniques and have chosen their own disciplines for studying the research trends and hot topics. Another reason could be that having domain knowledge helps in interpreting the results of the analysis, thus mining documents from one's own discipline will have better interpretations. This is also true for other findings too, in the case of biomedical sciences and Management.</p> <p>Further the results of the study depict that most of the researchers belong to the United States of America and still the text mining of English language documents are preferred by the text mining researchers. This may be due to unavailability of language-specific data sets, OCR algorithms and lower efficiency of text mining algorithms to mine corpus of other languages. It is also found that most of the studies preferred the abstract section of the research papers to form their corpus followed by the inclusion of the full text of the research articles to conduct text mining. This may be explained by the fact that the abstract level data is easily available from the scholarly databases whereas it requires a lot of effort from the part of the researcher to extract the full-text and process it further to form a corpus. It requires to write a lot of pipelines for extracting full-text from articles published from different journals having a variety of formats.</p> <p>The selected studies also highlighted the benefits of the application of text mining on scholarly literature such as improving the quality of research, highlighting the hidden information, saving time and cost, developing new services, ideas, and creating new knowledge concepts. Some of the studies selected in this study cited scholarly literature on text mining related to risks and barriers which includes, availability of multiple datasets, variety of tools and techniques for data analysis that creates confusion and requires a lot of effort in deciding appropriate techniques for their study.</p> <p>In addition, we also found some common challenges such as lack of skills, limited data sets in some domains, legal insecurity, technical problems, improper infrastructure and support for conducting text mining on large corpus size. Although this study provides some characteristics of studies published on text mining applications in the scholarly literature domain, it has some limitations too. This systematic review is based on selected studies on the scholarly text mining literature published in the last ten years. A more thorough systematic review including a larger set of articles will yield better results. Including articles from the last twenty years would also help in understanding the trends such as switching of preferences to text mining tools and technologies over the longer period.</p> <hd id="AN0158753239-33">Conclusion</hd> <p>This study would help to understand the popular text mining tools and techniques among the variety of tools and techniques used by text mining researchers to mine scholarly literature. It requires a lot of effort from the part of researchers to track and read this growing amount of literature. The availability of text mining techniques and tools have provided an opportunity to automatically analyse large sets of documents. This paper is of significance to the scholarly text mining researchers in determining the sample size, selecting which section of the literature to prefer for analysis and popular tools in their future research projects. Text mining as a technique has a greater significance to explore the unknown structures of research and also helps in determining the latest research trends.</p> <ref id="AN0158753239-34"> <title> References </title> <blist> <bibl id="bib1" type="bt">1</bibl> <bibtext> Abu-Shanab, E., & Harb, Y. (2019). E-government research insights: Text mining analysis. Electronic Commerce Research and Applications, 38, 100892. doi: 10.1016/j.elerap.2019.100892</bibtext> </blist> <blist> <bibl id="bib2" type="bt">2</bibl> <bibtext> Aich, S., Sain, M., Park, J., Choi, K. W., & Kim, H. C. (2017). A text mining approach to identify the relationship between gait-Parkinson's disease (PD) from PD based research articles. In 2017 International conference on inventive computing and informatics (ICICI), 481 – 485. IEEE. doi: 10.1109/ICICI.2017.8365398</bibtext> </blist> <blist> <bibl id="bib3" idref="ref12" type="bt">3</bibl> <bibtext> Allahyari, M., Pouriyeh, S., Assefi, M., Safaei, S., Trippe, E. D., Gutierrez, J.B., & Kochut, K. (2017). A Brief Survey of Text mining: Classification, clustering and extraction techniques. ArXiv:1707. 02919 [Cs]. <ulink href="http://arxiv.org/abs/1707.02919">http://arxiv.org/abs/1707.02919</ulink>.</bibtext> </blist> <blist> <bibl id="bib4" type="bt">4</bibl> <bibtext> Amado, A., Cortez, P., Rita, P., & Moro, S. (2018). Research trends on Big Data in Marketing: A text mining and topic modeling based literature analysis. European Research on Management and Business Economics, 24 (1), 1 – 7. doi: 10.1016/j.iedeen.2017.06.002</bibtext> </blist> <blist> <bibl id="bib5" type="bt">5</bibl> <bibtext> Ammarukleart, S., & Kim, J. (2017). Institutional repository research 2005-2015: A trend analysis using bibliometrics and text mining. Digital Library Perspectives, 33 (3), 264 – 278. doi: 10.1108/DLP-07-2016-0027</bibtext> </blist> <blist> <bibl id="bib6" idref="ref31" type="bt">6</bibl> <bibtext> Asmussen, C. B., & Møller, C. (2019). Smart literature review: A practical topic modelling approach to exploratory literature review. Journal of Big Data, 6 (1), 93. doi: 10.1186/s40537-019-0255-7</bibtext> </blist> <blist> <bibl id="bib7" idref="ref18" type="bt">7</bibl> <bibtext> Behera, S., & Kumar, N. V. (2015). Filtering of unstructured text. International Journal of Engineering Research and Development, 12 (11), e2278-067X.</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Bhanot, N., Singh, H., Sharma, D., Jain, H., & Jain, S. (2019). Python vs. R: A Text Mining Approach for Analyzing the Research Trends in Scopus Database. arXiv Preprint arXiv, 1911, 08271.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref15" type="bt">9</bibl> <bibtext> Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. Journal of Machine Learning Research, 3 (Jan), 993 – 1022.</bibtext> </blist> <blist> <bibtext> Brereton, R. G., & Lloyd, G. R. (2010). Support vector machines for classification and regression. The Analyst, 135 (2), 230 – 267. doi: 10.1039/B918972F</bibtext> </blist> <blist> <bibtext> Brown, M., & Lewis, H. (1999). Support vector machines and linear spectral unmixing for remote sensing. In S. Singh (Ed.). International Conference on Advances in Pattern Recognition (pp. 395 – 404). London : Springer.</bibtext> </blist> <blist> <bibtext> Carnerud, D. (2017). Exploring research on quality and reliability management through text mining methodology. International Journal of Quality & Reliability Management, 34 (7), 975 – 1014. doi: 10.1108/IJQRM-03-2015-0033</bibtext> </blist> <blist> <bibtext> Chen, J., Wang, T. T., & Lu, Q. (2016). THC-DAT: A document analysis tool based on topic hierarchy and context information. Library Hi-Tech, 34, 64 – 86.</bibtext> </blist> <blist> <bibtext> Chen, W., Xu, Y., Jin, R., & Wanatowski, D. (2019). Text mining–based review of articles published in the Journal of Professional Issues in Engineering Education and Practice. Journal of Professional Issues in Engineering Education and Practice, 145 (4), 06019002. doi: 10.1061/(ASCE)EI.1943-5541</bibtext> </blist> <blist> <bibtext> Cheng, F., Yang, K., & Zhang, L. (2015). A structural SVM based approach for binary classification under class imbalance. Mathematical Problems in Engineering, 2015, Article ID 269856. doi: 10.1155/2015/269856.</bibtext> </blist> <blist> <bibtext> Choi, W., & Lee, H. (2015). A text mining approach for identifying herb-chemical relationships from biomedical a rticles. Proceedings of the ACM Ninth International Workshop on Data and Text Mining in Biomedical Informatics - DTMBIO' Vol. 15, pp. 25–25. doi: 10.1145/2811163.2811178</bibtext> </blist> <blist> <bibtext> Cho, K.-W., Bae, S.-K., & Woo, Y.-W. (2017). Analysis on topic trends and topic modeling of KSHSM Journal Papers using Text Mining. The Korean Journal of Health Service Management, 11 (4), 213 – 224. doi: 10.12811/kshsm.2017.11.4.213</bibtext> </blist> <blist> <bibtext> Cho, S. G., & Kim, S. B. (2012). Identification of research patterns and trends through text mining. International Journal of Information and Educational Technology, 2 (3), 233 – 235. doi: 10.7763/IJIET.2012.V2.117.</bibtext> </blist> <blist> <bibtext> Cho, G. H., Lim, S. Y., & Hur, S. (2014). An analysis of the research methodologies and techniques in the industrial engineering using text mining. Journal of Korean Institute of Industrial Engineers, 40 (1), 52 – 59. doi: 10.7232/JKIIE.2014.40.1.052</bibtext> </blist> <blist> <bibtext> Choo, S., Park, H., Kim, T., & Seo, J. (2019). Analysis of trends in Korean BIM research and technologies using text mining. Applied Sciences, 9 (20), 4424. doi: 10.3390/app9204424</bibtext> </blist> <blist> <bibtext> Comeau, D. C., Wei, C. H., Islamaj Doğan, R., & Lu, Z. (2019). PMC text mining subset in BioC: About three million full-text articles and growing. Bioinformatics, 35 (18), 3533 – 3535. doi: 10.1093/bioinformatics/btz070</bibtext> </blist> <blist> <bibtext> Cortez, P., Moro, S., Rita, P., King, D., & Hall, J. (2018). Insights from a text mining survey on Expert Systems research from 2000 to 2016. Expert Systems, 35 (3), e12280. doi: 10.1111/exsy.12280</bibtext> </blist> <blist> <bibtext> Dang, D. S., & Ahmad, P. H. (2015). A review of text mining techniques associated with various application areas. International Journal of Science and Research (IJSR), 4 (2), 2461–2466.</bibtext> </blist> <blist> <bibtext> Das, S., Sun, X., & Dutta, A. (2016). Text mining and topic modeling of compendiums of papers from transportation research board annual meetings. Transportation Research Record: Journal of the Transportation Research Board, 2552 (1), 48 – 56. doi: 10.3141/2552-07</bibtext> </blist> <blist> <bibtext> Dayeen, F. R., Sharma, A. S., & Derrible, S. (2020). A text mining analysis of the climate change literature in industrial ecology. Journal of Industrial Ecology, 24 (2), 276 – 284. doi: 10.1111/jiec.12998</bibtext> </blist> <blist> <bibtext> Diane, H., & Lawrence, S. (2009). Latent Dirichlet allocation for text, images, and music (p. 1). San Diego : Department of Computer Science University of California.</bibtext> </blist> <blist> <bibtext> Ding, Z., Li, Z., & Fan, C. (2018). Building energy savings: Analysis of research trends based on text mining. Automation in Construction, 96, 398 – 410. doi: 10.1016/j.autcon.2018.10.008</bibtext> </blist> <blist> <bibtext> Do, Y., & Skłodowski, J. (2014). Research topics and trends over the past decade (2001-2013) of Baltic Coleopterology using text mining methods. Baltic Journal of Coleopterology, 14 (1), 1–6.</bibtext> </blist> <blist> <bibtext> Dreisbach, C., Koleck, T. A., Bourne, P. E., & Bakken, S. (2019). A systematic review of natural language processing and text mining of symptoms from electronic patient-authored text data. International Journal of Medical Informatics, 125, 37 – 46. doi: 10.1016/j.ijmedinf.2019.02.008</bibtext> </blist> <blist> <bibtext> Fang, D., Yang, H., Gao, B., & Li, X. (2018). Discovering research topics from library electronic references using latent Dirichlet allocation. Library Hi Tech, 36 (3), 400 – 410. doi: 10.1108/LHT-06-2017-0132</bibtext> </blist> <blist> <bibtext> Feinerer, I. (2013). Introduction to the tm Package Text Mining in R. Accessible en ligne: <ulink href="http://cran.r-project.org/web/packages/tm/vignettes/tm">http://cran.r-project.org/web/packages/tm/vignettes/tm</ulink>.</bibtext> </blist> <blist> <bibtext> Ferran-Ferrer, N., Guallar, J., Abadal, E., & Server, A. (2017). Research methods and techniques in Spanish library and information science journals (2012-2014). Information Research, 22 (1), paper 741.</bibtext> </blist> <blist> <bibtext> Ferreira‐Mello, R., André, M., Pinheiro, A., Costa, E., & Romero, C. (2019). Text mining in education. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9 (6), 1332. doi: 10.1002/widm</bibtext> </blist> <blist> <bibtext> Gaikwad, S. V., Chaugule, A., & Patil, P. (2014). Text mining methods and techniques. International Journal of Computer Applications, 85 (17), 42 – 45. doi: 10.5120/14937-3507</bibtext> </blist> <blist> <bibtext> García, D., Massucci, F. A., Mosca, A., Ràfols, I., Rodríguez, A., & Vassena, R. (2020). Mapping research in assisted reproduction worldwide. Reproductive BioMedicine Online, 40 (1), 71 – 81. doi: 10.1016/j.rbmo.2019.10.013</bibtext> </blist> <blist> <bibtext> Gülkesen, K. H., & Haux, R. (2019). Research subjects and research trends in medical informatics. Methods of Information in Medicine, 58 (S 01), e1 – e13. doi: 10.1055/s-0039-1681107</bibtext> </blist> <blist> <bibtext> Gurung, P., & Wagh, R. (2017). A study on topic identification using K means clustering algorithm: Big vs. small documents. Advances in Computational Sciences and Technology, 10, 221 – 233.</bibtext> </blist> <blist> <bibtext> Hahsler, M., & Karpienko, R. (2017). Visualizing association rules in hierarchical groups. Journal of Business Economics, 87 (3), 317 – 335. doi: 10.1007/s11573-016-0822-8</bibtext> </blist> <blist> <bibtext> Hung, J. (2012). Trends of e-learning research from 2000 to 2008: Use of text mining and bibliometrics: Research trends of e-learning. British Journal of Educational Technology, 43 (1), 5 – 16. doi: 10.1111/j.1467-8535.2010.01144.x</bibtext> </blist> <blist> <bibtext> Hung, J.-L., & Zhang, K. (2012). Examining mobile learning trends 2003–2008: A categorical meta-trend analysis using text mining techniques. Journal of Computing in Higher Education, 24 (1), 1 – 17. doi: 10.1007/s12528-011-9044-9</bibtext> </blist> <blist> <bibtext> Jeon, H. J., Kim, D. Y., Han, K. J., Han, D.W., Son, S. W., & Lee, C. M. (2018). An analysis of indoor environment research trends in Korea using topic modeling: Case study on abstracts from the journal of the Korean society for indoor environment. Journal of Odor and Indoor Environment, 17 (4), 322 – 329. doi: 10.15250/joie.2018.17.4.322</bibtext> </blist> <blist> <bibtext> Jeong, B., & Lee, H. (2016). Research Topics in Industrial Engineering 2001 ∼ 2015. Journal of Korean Institute of Industrial Engineers, 42 (6), 421 – 431. doi: 10.7232/JKIIE.2016.42.6.421</bibtext> </blist> <blist> <bibtext> Jia, S. S., & Wu, B. (2018). Incorporating LDA based text mining method to explore new energy vehicles in China. IEEE Access, 6, 64596 – 64602. doi: 10.1109/ACCESS.2018.2877716</bibtext> </blist> <blist> <bibtext> Jindal, R., & Shweta. (2018). A modified knowledge discovery process in the text documents. International Journal of Innovative Computing Information and Control, 14 (3), 817 – 832.</bibtext> </blist> <blist> <bibtext> Jitngernmadan, P., & Boonmee, P. (2018, October). Research trends of online marketing in social media r esearch. In 2018 International Conference on Information Technology (InCIT), IEEE, pp. 1–6. doi: 10.23919/INCIT.2018.8584887</bibtext> </blist> <blist> <bibtext> Joo, S., Choi, I., & Choi, N. (2018). Topic analysis of the research domain in knowledge organization: A latent dirichlet allocation approach. Knowledge Organization, 45 (2), 170 – 183. doi: 10.5771/0943-7444-2018-2-170</bibtext> </blist> <blist> <bibtext> Jotsov, V. S., & Iliev, E. (2015). Applications of advanced analytics methods in SAS enterprise miner. In P. Angelov, K. T. Atanassov, L. Doukovska, M. Hadjiski, V. Jotsov, J. Kacprzyk, N. Kasabov, S. Sotirov, E. Szmidt, & S. Zadrożny (Eds.), Intelligent Systems' 2014 (pp. 413 – 429). Cham : Springer.</bibtext> </blist> <blist> <bibtext> Kalra, V., & Aggarwal, R. (2017, December). Importance of text data preprocessing and implementation in RapidMiner. In Proceedings of the First International Conference on Information, ICITKM, New Delhi, pp. 71–75).</bibtext> </blist> <blist> <bibtext> Kang, H. J., Kim, C., & Kang, K. (2019). Analysis of the trends in biochemical research using Latent Dirichlet Allocation (LDA). Processes, 7 (6), 379. doi: 10.3390/pr7060379</bibtext> </blist> <blist> <bibtext> Kim, K. H., Han, Y. J., Lee, S., Cho, S. W., & Lee, C. (2019). Text mining for patent analysis to forecast emerging technologies in wireless power transfer. Sustainability, 11 (22), 6240. doi: 10.3390/su11226240.</bibtext> </blist> <blist> <bibtext> Kim, C. S., Kwahk, K. Y., & Yoon, H. J. (2017a). An analysis of research trends in tourism studies: Applying topic modeling and time series regression analysis. The Journal of Tourism & Leisure Research, 29 (12), 25 – 39.</bibtext> </blist> <blist> <bibtext> Kim, M. J., Ohk, K., & Moon, C. S. (2017b). Trend analysis by using text mining of Journal Articles Regarding Consumer Policy. New Physics: Sae Mulli, 67 (5), 555 – 561. doi: 10.3938/NPSM.67.555</bibtext> </blist> <blist> <bibtext> Kim, Y.-M., & Delen, D. (2018). Medical informatics research trend analysis: A text mining approach. Health Informatics Journal, 24 (4), 432 – 452. doi: 10.1177/1460458216678443</bibtext> </blist> <blist> <bibtext> Kirilenko, A. P., & Stepchenkova, S. (2018). Tourism research from its inception to present day: Subject area, geography, and gender distributions. PLoS ONE, 13 (11), e0206820. doi: 10.1371/journal.pone.0206820</bibtext> </blist> <blist> <bibtext> Koch, N. M. (2019). The phylogenomic revolution and its conceptual innovations: A text mining approach. Organisms Diversity & Evolution, 19 (2), 99 – 103.</bibtext> </blist> <blist> <bibtext> Konstantinidis, S. T., Billis, A., Wharrad, H., & Bamidis, P. D. (2017). Internet of things in health trends through bibliometrics and text mining. Informatics for Health: Connected Citizen-Led Wellness and Population Health, 235, 73.</bibtext> </blist> <blist> <bibtext> Krishnamurthy, R., & Balasubramanium, R. K. (2019). Using text mining to identify trends in oropharyngeal dysphagia research: A proof of concept. Communication Sciences & Disorders, 24 (1), 234 – 243. doi: 10.12963/csd.19579</bibtext> </blist> <blist> <bibtext> Kurata, K., Miyata, Y., Ishita, E., Yamamoto, M., Yang, F., & Iwase, A. (2018). Analyzing library and information science full‐text articles using a topic modeling approach. Proceedings of the Association for Information Science and Technology, 55 (1), 847 – 848. doi: 10.1002/pra2.2018.14505501143</bibtext> </blist> <blist> <bibtext> Lam, C., Lai, F.-C., Wang, C.-H., Lai, M.-H., Hsu, N., & Chung, M.-H. (2016). Text mining of journal articles for sleep disorder terminologies. PLOS ONE, 11 (5), e0156031. doi: 10.1371/journal.pone.0156031</bibtext> </blist> <blist> <bibtext> Lamba, M. (2019). Text Analysis of ETDs in ProQuest Dissertations and Theses (PQDT) Global Database (2016–2018). "International Conference on Digital Landscape (ICDL) 2019: Digital Transformation for an Agile Environment"; November 6–8, 2019; New Delhi.</bibtext> </blist> <blist> <bibtext> Lamba, M., & Madhusudhan, M. (2018a). Metadata tagging of library and information science theses: Shodhganga (2013-2017). ETD 2018 Taiwan Beyond the Boundaries of Rims and Oceans: Globalizing Knowledge with ETDs.</bibtext> </blist> <blist> <bibtext> Lamba, M., & Madhusudhan, M. (2018b). Application of topic mining and prediction modeling tools for library and information science journals. In M. R. Murali Prasad, A. Munigal, R. Nalik, M. Madhusudhan, G. Surender Rao (Eds.), Library Practices in Digital Era (pp. 395 – 401). Hyderabad : BS Publications.</bibtext> </blist> <blist> <bibtext> Lamba, M., & Madhusudhan, M. (2019a). Author-Topic Modeling of DESIDOC Journal of Library and Information Technology. (2008-2017), India.). Library Philosophy and Practice (e-journal), 2593.</bibtext> </blist> <blist> <bibtext> Lamba, M., & Madhusudhan, M. (2019b). Mapping of topics in DESIDOC Journal of Library and Information Technology, India. Scientometrics, 120 (2), 477 – 505. doi: 10.1007/s11192-019-03137-5</bibtext> </blist> <blist> <bibtext> Lamba, M., & Madhusudhan, M. (2020). Mapping of ETDs in ProQuest dissertations and theses (PQDT) global database (2014-2018). Cadernos BAD, 1, 169 – 182.</bibtext> </blist> <blist> <bibtext> Lee, H. S., Song, H. G., & Lee, H. S. (2013). Classification of photovoltaic research papers by using text-mining techniques. Applied Mechanics and Materials, 284-287, 3362 – 3369. doi: 10.4028/<ulink href="http://www.scientific.net/AMM.284-287.3362">www.scientific.net/AMM.284-287.3362</ulink></bibtext> </blist> <blist> <bibtext> Lee, J. Y., Kim, H. & Kim, P. J. (2010). Domain analysis with text mining: Analysis of digital library research trends using profiling methods. Journal of Information Science, 36 (2), 144 – 161. doi: 10.1177/0165551509353251</bibtext> </blist> <blist> <bibtext> Lee, J., & Lee, M. J. (2018, July). Measuring contribution of spatial information to environmental research using text mining techniques. In IGARSS 2018-2018 IEEE International Geoscience and Remote Sensing Symposium, IEEE, pp. 5289 – 5291.</bibtext> </blist> <blist> <bibtext> Li, B., Yu, S. & Lu, Q. (2003). An improved k-nearest neighbor algorithm for text categorization. arXiv preprint cs/0306099</bibtext> </blist> <blist> <bibtext> Lim, C., & Maglio, P. P. (2018). Data-driven understanding of smart service systems through text mining. Service Science, 10 (2), 154 – 180. doi: 10.1287/serv.2018.0208</bibtext> </blist> <blist> <bibtext> Liu, L., Tang, L., Dong, W., Yao, S., & Zhou, W. (2016). An overview of topic modeling and its current applications in bioinformatics. SpringerPlus, 5 (1), 1608. doi: 10.1186/s40064-016-3252-8</bibtext> </blist> <blist> <bibtext> Maddi, A., Sapinho, D., & Baudoin, L. (2019). Mapping an emerging research subject: Case of microbiota concept. In 17th International Conference on Scientometrics & Informetrics, ISSI2019, with a Special STI Indicators Conference Track, 2-5 September 2019, Sapienza University of Rome, Rome (Italy), pp. 1232 – 1243.</bibtext> </blist> <blist> <bibtext> Mandujano, S. (2019). Analysis and trends of photo-trapping in Mexico: Text mining in R. Therya, 10 (1), 25 – 32. doi: 10.12933/therya-19-666</bibtext> </blist> <blist> <bibtext> Miner, G., Elder, I. V., J., Fast, A., Hill, T., Nisbet, R., & Delen, D. (2012). Practical text mining and statistical analysis for non-structured text data applications. Cambridge : Academic Press.</bibtext> </blist> <blist> <bibtext> Miyata, Y., Ishita, E., Yang, F., Yamamoto, M., Iwase, A., & Kurata, K. (2020). Knowledge structure transition in library and information science: Topic modeling and visualization. Scientometrics, 125 (1), 665 – 687. doi: 10.1007/s11192-020-03657-5</bibtext> </blist> <blist> <bibtext> Moro, S., Alturas, B., Esmerado, J., & Costa, C. J. (2017). Research trends in CISTI's unveiled through text mining. 2017 12th Iberian Conference on Information Systems and Technologies (CISTI), 1 – 5. doi: 10.23919/CISTI.2017.7975765</bibtext> </blist> <blist> <bibtext> Moro, S., Cortez, P., & Rita, P. (2015). Business intelligence in banking: A literature analysis from 2002 to 2013 using text mining and latent Dirichlet allocation. Expert Systems with Applications, 42 (3), 1314 – 1324. doi: 10.1016/j.eswa.2014.09.024</bibtext> </blist> <blist> <bibtext> Nagarkar, S. P., & Kumbhar, R. (2015). Text mining: An analysis of research published under the subject category 'Information Science Library Science' in Web of Science Database during 1999-2013. Library Review, 64 (3), 248 – 262. doi: 10.1108/LR-08-2014-0091</bibtext> </blist> <blist> <bibtext> Nie, B., & Sun, S. (2017). Using text mining techniques to identify research trends: A case study of design research. Applied Sciences, 7 (4), 401. doi: 10.3390/app7040401</bibtext> </blist> <blist> <bibtext> O'Mara-Eves, A., Thomas, J., McNaught, J., Miwa, M., & Ananiadou, S. (2015). Using text mining for study identification in systematic reviews: A systematic review of current approaches. Systematic Reviews, 4 (1), 5. doi: 10.1186/2046-4053-4-5</bibtext> </blist> <blist> <bibtext> Onwuchekwa, E. O., & Jegede, O. R. (2011). Information retrieval methods in libraries and information centers. African Research Review, 5 (6), 108 – 120. doi: 10.4314/afrrev.v5i6.10</bibtext> </blist> <blist> <bibtext> Park, H., & Park, M. S. (2019). Capturing the trend of mHealth research using text mining. mHealth, 5, 48 – 48. 2019. 09. 06 doi: 10.21037/mhealth</bibtext> </blist> <blist> <bibtext> Prasanna, P. L., & Rao, D. R. (2019). A text mining research based on topic modeling using Latent Dritchlent allocation. International Journal of Recent Technology and Engineering, 7 (5), 308 – 317.</bibtext> </blist> <blist> <bibtext> Ravikumar, S., Agrahari, A., & Singh, S. N. (2015). Mapping the intellectual structure of scientometrics: A co-word analysis of the journal Scientometrics (2005–2010. Scientometrics, 102 (1), 929 – 955. doi: 10.1007/s11192-014-1402-8</bibtext> </blist> <blist> <bibtext> Salloum, S. A., Al-Emran, M., Monem, A. A., & Shaalan, K. (2018). Using text mining techniques for extracting information from research articles. In K. Shaalan, A. E. Hassanien, M. F. Tolba (Eds.), Intelligent natural language processing: Trends and applications (pp. 373 – 397). Cham : Springer.</bibtext> </blist> <blist> <bibtext> Shankaranarayanan, G., & Blake, R. (2017). From content to context: The evolution and growth of data quality research. Journal of Data and Information Quality, 8 (2), 1 – 28. doi: 10.1145/2996198</bibtext> </blist> <blist> <bibtext> Sharma, D., Kumar, B., & Chand, S. (2018). October). Trend analysis in machine learning research using text mining. In 2018 International Conference on Advances in Computing, Communication Control and Networking (ICACCCN), IEEE, pp. 136–141. doi: 10.1109/ICACCCN.2018.8748686</bibtext> </blist> <blist> <bibtext> Shianghau, W., & Jiannjong, G. (2010, October). The trend of green supply chain management research (2000–2010): A text mining analysis. 2010 8th International Conference on Supply Chain Management and Information, IEEE, pp. 1 – 6.</bibtext> </blist> <blist> <bibtext> Shin, S. Y., & Suh, C. K. (2019, October). Discovering platform government research trends using topic modeling. International Conference on Model and Data Engineering (pp. 67 – 82). Springer, Cham.</bibtext> </blist> <blist> <bibtext> Shinde, P. P., Oza, K. S., & Kamat, R. K. (2017, February). Big data predictive analysis: Using R analytical tool. 2017 International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), IEEE, pp. 839 – 842.</bibtext> </blist> <blist> <bibtext> Smink, W., Sools, A. M., van der Zwaan, J. M., Wiegersma, S., Veldkamp, B. P., & Westerhof, G. J. (2019). Towards text mining therapeutic change: A systematic review of text-based methods for therapeutic change process research. PLoS ONE, 14 (12), e0225703. doi: 10.1371/journal.pone.0225703</bibtext> </blist> <blist> <bibtext> Son, Y., & Kang, H. S. (2019). A Text Mining Analysis of HPV Vaccination Research Trends. Child Health Nursing Research, 25 (4), 458 – 467. doi: 10.4094/chnr.2019.25.4.458</bibtext> </blist> <blist> <bibtext> Song, H. J., Park, K. S., Jung, H. E., & Song, M. (2013). Trend Analysis of Korean Economy in the Economic Literature by text mining techniques. In Proceedings of the Korean Society for Information Management Conference (pp. 47 – 50). Korean Society for Information Management.</bibtext> </blist> <blist> <bibtext> Song, M., & Kim, S. Y. (2013). Detecting the knowledge structure of bioinformatics by mining full-text collections. Scientometrics, 96 (1), 183 – 201. doi: 10.1007/s11192-012-0900-9</bibtext> </blist> <blist> <bibtext> Song, Y. Y., & Ying, L. U. (2015). Decision tree methods: Applications for classification and prediction. Shanghai Archives of Psychiatry, 27 (2), 130.</bibtext> </blist> <blist> <bibtext> Sugimoto, C. R., Li, D., Russell, T. G., Finlay, S. C., & Ding, Y. (2011). The shifting sands of disciplinary development: Analyzing North American Library and Information Science dissertations using latent Dirichlet allocation. Journal of the American Society for Information Science and Technology, 62 (1), 185 – 204. doi: 10.1002/asi.21435</bibtext> </blist> <blist> <bibtext> Sulova, S., Todoranova, L., Penchev, B., & Nacheva, R. (2017). Using text mining to classify research papers. International Multidisciplinary Scientific GeoConference Surveying Geology and Mining Ecology Management, SGEM, Vol. 17(21), pp. 647 – 654.</bibtext> </blist> <blist> <bibtext> Sumathy, K. L., & Chidambaram, M. (2013). Text mining: Concepts, applications, tools and issues-an overview. International Journal of Computer Applications, 80 (4), 29 – 32. doi: 10.5120/13851-1685</bibtext> </blist> <blist> <bibtext> Sun, L., & Yin, Y. (2017). Discovering themes and trends in transportation research using topic modeling. Transportation Research Part C: Emerging Technologies, 77, 49 – 66. doi: 10.1016/j.trc.2017.01.013</bibtext> </blist> <blist> <bibtext> Syed, S., Borit, M., & Spruit, M. (2018). Narrow lenses for capturing the complexity of fisheries: A topic analysis of fisheries science from 1990 to 2016. Fish and Fisheries, 19 (4), 643 – 661. doi: 10.1111/faf.12280</bibtext> </blist> <blist> <bibtext> Talib, R., Kashif, M., Ayesha, S., & Fatima, F. (2016). Text mining: Techniques, applications and issues. International Journal of Advanced Computer Science and Applications, 7 (11), 414 – 418. doi: 10.14569/IJACSA.2016.071153</bibtext> </blist> <blist> <bibtext> Usai, A., Pironti, M., Mital, M., & Mejri, C. A. (2018). Knowledge discovery out of text data: A systematic review via text mining. Journal of Knowledge Management, 22 (7), 1471 – 1488. doi: 10.1108/JKM-11-2017-0517</bibtext> </blist> <blist> <bibtext> Vidhya, K. A., & Aghila G. (2010). Text mining process, techniques and tools: An overview. International Journal of Information Technology and Knowledge Management, 2 (2), 613 – 622.</bibtext> </blist> <blist> <bibtext> Wang, S. H., Ding, Y., Zhao, W., Huang, Y. H., Perkins, R., Zou, W., & Chen, J. J. (2016). Text mining for identifying topics in the literatures about adolescent substance use and depression. BMC Public Health, 16 (1), 279. doi: 10.1186/s12889-016-2932-1</bibtext> </blist> <blist> <bibtext> Wang, Y., Bowers, A. J., & Fikis, D. J. (2017). Automated text data mining analysis of five decades of educational leadership research literature: Probabilistic topic modeling of EAQ articles from 1965 to 2014. Educational Administration Quarterly, 53 (2), 289 – 323. doi: 10.1177/0013161X16660585</bibtext> </blist> <blist> <bibtext> Westergaard, D., Staerfeldt, H. H., Tønsberg, C., Jensen, L. J., & Brunak, S. (2018). A comprehensive and quantitative comparison of text-mining in 15 million full-text articles versus their corresponding abstracts. PLoS Computational Biology, 14 (2), e1005962. doi: 10.1371/journal.pcbi.1005962</bibtext> </blist> <blist> <bibtext> White, G. O., Guldiken, O., Hemphill, T. A., He, W., & Sharifi Khoobdeh, M. (2016). Trends in international strategic management research from 2000 to 2013: TexMining and bibliometric analysis. Management International Review, 56 (1), 35 – 65. doi: 10.1007/s11575-015-0260-9</bibtext> </blist> <blist> <bibtext> Witten, I. H., Don, K. J., Dewsnip, M., & Tablan, V. (2004). Text mining in a digital library. International Journal on Digital Libraries, 4 (1), 56 – 59. doi: 10.1007/s00799-003-0066-4</bibtext> </blist> <blist> <bibtext> Yehia, A. M., Ibrahim, L. F., & Abulkhair, M. F. (2016). Text mining and knowledge discovery from big data: Challenges and promise. International Journal of Computer Science Issues (IJCSI), 13 (3), 54.</bibtext> </blist> <blist> <bibtext> Yoo, H. H., Shin, S., Yoo, H. H., & Shin, S. (2015). Trends of research articles in the Korean Journal of Medical Education by social network analysis. Korean Journal of Medical Education, 27 (4), 247 – 254. doi: 10.3946/kjme.2015.27.4.247</bibtext> </blist> <blist> <bibtext> Yoon, J. E., & Suh, C. J. (2019). Research trend analysis by using text-mining techniques on the convergence studies of ai and healthcare technologies. Journal of Information Technology Services, 18 (2), 123 – 141.</bibtext> </blist> <blist> <bibtext> Yu, B., & Ku, M.-C. (2010). Collecting legacy corpora from social science research for text mining evaluation: Collecting legacy corpora from social science research for text mining evaluation. Proceedings of the American Society for Information Science and Technology, 47 (1), 1 – 2. doi: 10.1002/meet.14504701368</bibtext> </blist> <blist> <bibtext> Zhang, H. (2005). Exploring conditions for the optimality of naïve bayes. International Journal of Pattern Recognition and Artificial Intelligence, 19 (02), 183 – 198. doi: 10.1142/S0218001405003983</bibtext> </blist> <blist> <bibtext> Zheng, P., Liang, X., Huang, G., & Liu, X. (2016). Mapping the field of communication technology research in Asia: Content analysis and text mining of SSCI journal articles 1995–2014. Asian Journal of Communication, 26 (6), 511 – 531. doi: 10.1080/01292986.2016.1231210</bibtext> </blist> <blist> <bibtext> Zou, C. (2018). Analyzing research trends on drug safety using topic modeling. Expert Opinion on Drug Safety, 17 (6), 629 – 636. doi: 10.1080/14740338.2018.1458838</bibtext> </blist> </ref> <aug> <p>By Khusbu Thakur and Vinit Kumar</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib23" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib79" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib98" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib108" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib41" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib34" firstref="ref6"></nolink> <nolink nlid="nl7" bibid="bib109" firstref="ref7"></nolink> <nolink nlid="nl8" bibid="bib101" firstref="ref8"></nolink> <nolink nlid="nl9" bibid="bib81" firstref="ref9"></nolink> <nolink nlid="nl10" bibid="bib103" firstref="ref13"></nolink> <nolink nlid="nl11" bibid="bib74" firstref="ref14"></nolink> <nolink nlid="nl12" bibid="bib26" firstref="ref16"></nolink> <nolink nlid="nl13" bibid="bib113" firstref="ref17"></nolink> <nolink nlid="nl14" bibid="bib10" firstref="ref19"></nolink> <nolink nlid="nl15" bibid="bib11" firstref="ref20"></nolink> <nolink nlid="nl16" bibid="bib38" firstref="ref21"></nolink> <nolink nlid="nl17" bibid="bib69" firstref="ref22"></nolink> <nolink nlid="nl18" bibid="bib83" firstref="ref23"></nolink> <nolink nlid="nl19" bibid="bib95" firstref="ref24"></nolink> <nolink nlid="nl20" bibid="bib31" firstref="ref25"></nolink> <nolink nlid="nl21" bibid="bib47" firstref="ref26"></nolink> <nolink nlid="nl22" bibid="bib48" firstref="ref27"></nolink> <nolink nlid="nl23" bibid="bib61" firstref="ref28"></nolink> <nolink nlid="nl24" bibid="bib71" firstref="ref29"></nolink> <nolink nlid="nl25" bibid="bib73" firstref="ref30"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1366037
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Application of Text Mining Techniques on Scholarly Research Articles: Methods and Tools
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Thakur%2C+Khusbu%22">Thakur, Khusbu</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-9842-3452">0000-0002-9842-3452</externalLink>)<br /><searchLink fieldCode="AR" term="%22Kumar%2C+Vinit%22">Kumar, Vinit</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-8306-2087">0000-0001-8306-2087</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22New+Review+of+Academic+Librarianship%22"><i>New Review of Academic Librarianship</i></searchLink>. 2022 28(3):279-302.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 24
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Information+Retrieval%22">Information Retrieval</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Methodology%22">Research Methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Trend+Analysis%22">Trend Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Sample+Size%22">Sample Size</searchLink><br /><searchLink fieldCode="DE" term="%22Information+Technology%22">Information Technology</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Research%22">Educational Research</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Software%22">Computer Software</searchLink><br /><searchLink fieldCode="DE" term="%22Programming+Languages%22">Programming Languages</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1080/13614533.2021.1918190
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1361-4533<br />1740-7834
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: A vast amount of published scholarly literature is generated every day. Today, it is one of the biggest challenges for organisations to extract knowledge embedded in published scholarly literature for business and research applications. Application of text mining is gaining popularity among researchers and applications are growing exponentially in different research areas. This study investigates the variety of text mining tools, techniques, sample sizes, domains and sections of the documents preferred by the text mining researchers through a systematic and structured literature review of conceptual and empirical studies. The significant findings depict that LDA and R package is the most extensively used tool and technique among the authors, most of the researchers prefer the sample size of 1000 articles for analysis, literature belonging to the domain of ICT, and related disciplines are frequently analysed in the text mining studies and abstracts constitute the corpus of the majority of text mining studies.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2023
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1366037
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1366037
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/13614533.2021.1918190
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 24
        StartPage: 279
    Subjects:
      – SubjectFull: Information Retrieval
        Type: general
      – SubjectFull: Data Analysis
        Type: general
      – SubjectFull: Research Methodology
        Type: general
      – SubjectFull: Trend Analysis
        Type: general
      – SubjectFull: Sample Size
        Type: general
      – SubjectFull: Information Technology
        Type: general
      – SubjectFull: Educational Research
        Type: general
      – SubjectFull: Algorithms
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Computer Software
        Type: general
      – SubjectFull: Programming Languages
        Type: general
    Titles:
      – TitleFull: Application of Text Mining Techniques on Scholarly Research Articles: Methods and Tools
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Thakur, Khusbu
      – PersonEntity:
          Name:
            NameFull: Kumar, Vinit
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 1361-4533
            – Type: issn-electronic
              Value: 1740-7834
          Numbering:
            – Type: volume
              Value: 28
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: New Review of Academic Librarianship
              Type: main
ResultId 1