<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Comparative Review of Text Mining &amp; Related Technologies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roland Vasili</string-name>
          <email>rvasili@uogj.edu.al</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Endri Xhina</string-name>
          <email>endri.xhina@fshn.edu.al</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Souliotis</string-name>
          <email>s1778881@sms.ed.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilia Ninka</string-name>
          <email>ilia.ninka@fshn.edu.al</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Library &amp;</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Informatics, Faculty of Natural Sciences, University of Tirana</institution>
          ,
          <addr-line>1001 Tirana</addr-line>
          ,
          <country country="AL">Albania</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Informatics, University of Edinburgh</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dept. of Mathematics, Informatics &amp; Physics, Faculty of Natural Sciences, University of Gjirokastra</institution>
          ,
          <addr-line>6001 Gjirokastra</addr-line>
          ,
          <country country="AL">Albania</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Information</institution>
          ,
          <addr-line>Science</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data Mining Text Mining Figure 1: Multidisciplinary Nature of Text Mining (Composition of Fig.1 in [Tal16] and Fig. 4.1 in [Dea14])</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Text mining has become an established
discipline in both research and business
intelligence. It refers commonly to the method
of extracting interesting information and
knowledge from unstructured text. Society's
future will be closely connected to handling
large amount of data. Information may be
available in various ways, either freely on the
Web or on social networks. Text mining is a
multi-disciplinary field in view of Data
Mining, Computational Linguistics, Artificial
Intelligence and Machine Learning, Statistics,
Databases, Library and Information Sciences,
and actually the new field of Big Data. Some
of these disciplines will be compared based
on the goals, data, algorithms, techniques and
the tools they use, as well as the their outcome.
All these subjects are similar, which is based
on two fundamental facts: (1) all of them
develop methods and procedures to process
data, and (2) any data processing algorithm or
procedure may belong to some or even all.
The differences are in their perspectives. This
difference in perspectives does not affect the
procedures but it does affect the choice of
them and, even more so, interpretation of
concepts and results.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>The Text Mining field covers a wide research area and
its methods can be applied in different contexts and for
several purposes, depending on the needs of the
specific task and the availability of data and expertise.
To this aim, rather than being an exhaustive list of
techniques and research directions, the Figure 1 shows
that the text mining is a composite discipline that
overlaps several branches of science. In the Figure 1 of
[Tal16] Knowledge Data Discovery field from Fig. 4.1
in [Dea14] is added, that will help us understand that
Statistics
AI &amp; Machine</p>
      <p>Learning
* Document
Classification
* Information
Extraction
Databases
* Document</p>
      <p>Clustering
* Information</p>
      <p>Retrieval</p>
      <sec id="sec-2-1">
        <title>M*Winienbg</title>
        <p>KD
D
* Natural
Language</p>
        <p>Processing
* Concept
Extraction</p>
        <p>Computational
Linguistics
the goal of all of these disciplines is knowledge
discovery, and in this base they will be compared.</p>
        <p>So, the goal of this paper is to outline the Text
Mining landscape which in contrast to encompassed
technologies like Data Mining, Natural Language
Processing, Information Retrieval, Information
Extraction, Artificial Intelligence and Machine
Learning., it tries to depict the scale and potential
scientific interaction with classic scientific areas, such
as Statistics.</p>
        <sec id="sec-2-1-1">
          <title>1.1 Definition of Text Mining</title>
          <p>Text mining (TM), also called Intelligent Text
Analysis, Text Data Mining or Knowledge-Discovery
in Text (KDT), is mainly used to define the procedure
of extracting interesting and non-trivial data and
knowledge from unstructured text ([Gök15]). There are
many more definitions of text mining like the definition
of the Oxford English Dictionary: "as the process or
practice of examining large collections of written
resources in order to generate new information,
typically using specialized computer software". It
widely covers a large set of related topics and
algorithms for analyzing text, spanning various
communities, including information retrieval, natural
language processing, data mining, machine learning,
many application domains web and biomedical
sciences.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>1.2 Text Mining Process</title>
          <p>The text mining process (TM) can basically be
summarized in three (3) steps below ([Kar05]):
Documents
Collection
</p>
          <p>Document Collection</p>
          <p>Attribute
Selection</p>
          <p>Text Mining
Techniques /
Pattern Discovery</p>
          <p>Interpretation
/ Evaluation
Text</p>
          <p>Text
Preprocessing</p>
          <p>Text Transformation
(Attribute Generation)
The steps of TM process regarding the output of results
are shown in the Figure 2.</p>
          <p>A TM system receives a collection of documents as
input and then pre-processes each document by
checking its format and set of characters. Next, these
pre-processed documents go through the text analysis
phase, by repeating the techniques until the required
information is extracted. Figure 3 shows three
techniques of text analysis, but other techniques,
Documents
Collection</p>
          <p>Retrieve &amp;
preprocess
documents</p>
          <p>Analyze Text</p>
          <p>Information</p>
          <p>Extraction
Clustering Summarization</p>
          <p>Information
Management</p>
          <p>System</p>
          <p>Knowledge
/ Wisdom
depending on the goal and the corporation, may also be
used. Information derived from the extraction can be
accessed by an information management system,
producing valuable knowledge for the user of this
system. Figure 4 analyzes the processing steps that a
typical TM System follows.</p>
          <p>Retrieve &amp;
Preprocess Documents</p>
          <p>Feature
Selection
Feature
Generation</p>
          <p>Information
Retrieval</p>
          <p>Information</p>
          <p>Extraction</p>
          <p>TM
Techniques
Feature
Selection
Feature
Generation</p>
          <p>Information
Management</p>
          <p>System
(Knowledge)
Text Summarization</p>
          <p>Topic Discovery</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>1.2.1 Document-Text Collection</title>
          <p>The basic element of TM is the collection of documents
of any text form. The number of texts in such
collections may range from thousands to several
millions.</p>
          <p>Text collection can be static or dynamic. At the
static approach, the original textbook total remains
unchanged while at the dynamic, the textbook over time
is classified into new or gets updated. Extremely large
collections and high-rate changing text collections are
considered challenges and constitute the main object of
Text Mining Systems. A peculiar example of a large
dynamic collection of texts, used by millions around
the world, is Pub Med (US National Library of
Medicine 2018)1. It is an internet resource, which
includes literature references related to biomedical and
health sciences. It is worth pointing out that it includes
over 25 million research reports in the biomedical field
in which they are added, roughly 35,000 with 40,000
new items each month. In addition to that, unstructured
data and free text are usually most of the data we
encounter and this includes over 40 million articles in
Wikipedia, 4.5 billion Web pages, about 500 million
tweets a day, and over 1.5 trillion queries on Google in
a year.</p>
          <p>Therefore, to initiate the TM process, the user has to
choose the desired collection of texts on which the
procedure will be based on, and the variety of texts that
will constitute the source of the data.</p>
          <p>The following process involves the TM System,
which has the ability (with the help of
knowledgediscovery algorithms) to quickly and efficiently identify
the patterns among a large number of natural texts.
But the realization of this requires the existence of
elaborate text collections. For this reason, the most
important TM process is the pre-processing phase of
the texts under examination, and then, the successful
implementation of the knowledge-discovery algorithms.</p>
        </sec>
        <sec id="sec-2-1-4">
          <title>1.2.2 Text Preprocessing</title>
          <p>Though this is considered to be the preliminary step to
be conducted, before actually applying Text Mining
algorithms/methods, it is a very important process. This
routine itself is divided into a number of sub-methods
which again have optional algorithms with their own set
of advantages and disadvantages.</p>
          <p>Most of the TM approaches are based on the idea
that a text document can be described by the set of
words contained in it i.e. bag-of-words representation.
The preprocessing itself is made up of a sequence of
steps ([Gup09]) (Figure 5). The steps are:
 Text Structure Removal.
 Tokenization.
 Stopwords Removal.
 Filtering (Removing terms based on their
length).
1 https://www.ncbi.nlm.nih.gov/pubmed/: Accessed 1-4-2018




</p>
          <p>Files
(Html, Pdf)</p>
          <p>Filtering (Removing terms based on their
frequency).</p>
          <p>Part Of Speech Tagging (Syntactical and
Semantical Analysis)
Stemming.</p>
          <p>N- grams.</p>
          <p>Term weighting.</p>
          <p>Text +
Structure</p>
          <p>Structure Identification</p>
          <p>Structure Removal</p>
          <p>Text
Total</p>
          <p>Names
Stemming</p>
          <p>POS Tagging</p>
          <p>Tokenization</p>
          <p>Tokens
Stopwords
Removal</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2 Text Mining and Data Mining</title>
      <p>Data Mining (DM) is a subfield of computer science
which combines many techniques from statistics, data
science, database theory and machine learning.</p>
      <p>DM is simply the process of gathering information
from huge databases that was previously
incomprehensible and unknown and then using that
information to make relevant business decisions. More
simply, data mining is a set of various methods that are
used in the process of knowledge discovery for
distinguishing the relationships and patterns that were
previously unknown. The final goal is the description
of existing database data as well as forecasting and
clarification of new data. We can therefore define data
mining as a combination of various other fields like
artificial intelligence, data room virtual base
management, pattern recognition, visualization of data,
machine learning, statistical studies and so on. The
primary goal of data mining is to extract information
from various sets of data in an attempt to transform it in
proper and meaningful structures for eventual use. It
mainly includes procedures and tools of extracting
patterns from the data set and relates exclusively to
structured data. But in recent years, interest has also
shifted to unstructured data (e.g. texts, images,
paperwork, web pages, etc.) with the result of
knowledge discovery from text (Text Mining). This
shift is very important since most of the data nowadays
are in unstructured textual form ([Gri08]). For example,
a text file contains few structured elements such as
author, title, date of creation etc. But it also contains
large segments of unstructured text such as its summary
and its contents. This requires both sophisticated
linguistic and statistical techniques able to analyze
unstructured text formats and techniques that combine
each document with actionable metadata.</p>
      <p>TM is an intense cognitive process through which
the user interacts with a collection of texts using a set
of analysis tools ([Seh04]). Similarly, as well as DM,
TM aims at extracting useful information from data
sources through recognition (identification) and
examination of interesting patterns. Meanwhile, in the
case of TM data sources are text collections, interesting
motives are searched in unstructured textual data
([Nah02]).</p>
      <p>Given the above definition, it is argued that the TM
has its roots in the area of Knowledge Discovery (KD).
Moreover, this is also used for the DM definition
reference. Consequently, TM is similar to DM, mainly
because in both cases, knowledge detection is based on
processes of data preprocessing and pattern searching
algorithms. However, this similarity may lead to
overseeing their differences. Thus, the goal of the
majority of the studies in those two areas, is to identify
and analyze these differences.</p>
    </sec>
    <sec id="sec-4">
      <title>3 Text Mining vs. Data Mining</title>
      <p>The method of Knowledge Discovery from Data or
Data Mining, namely finding useful patterns between
data, is a very good solution for collecting and storing a
huge volume of data. Though the scope of its
implementation is extensive it is not a developing
technology.</p>
      <p>Instead, the knowledge discovery of textual data or
Text Mining is a new method in the field of Knowledge
Discovery, which is feasible because the information to
be extracted refers to text.</p>
      <p>The knowledge discovery from text resembles a lot
to the classical method of knowledge discovery from
data, since both are based on knowledge management.
But, [Fra92] and [Raj97] concluded that the difference
between these two domains is the type of data they use
for Knowledge Discovery (KD). Thus, while DM uses
data extraction techniques over structured data, TM
does the same thing but for unstructured or
semistructured data, which is often referred to as textual
data ([Gup09]).</p>
      <p>Knowledge discovery from data is implemented in
the databases where the data is structured and described
by a unique structure where each instance of a problem
is determined by a specific and fixed set of features
([Kan09]).</p>
      <p>Instead, in the case of knowledge discovery from
text, the data is semi-structured or unstructured and
cannot be described by any set of fixed features
([Liu11]). For this reason, the method tries to bring the
text in the appropriate form for the direct application of
its computing applications.</p>
      <p>In the case of knowledge discovery from texts, there
are two approaches regarding the representation of the
text. In the first approach, the presence of a feature
(word) in a text is taken into consideration. Thus, when
a new instance of the problem occurs, what is
controlled is the presence of instances of the features
(words) in different classes of the problem. The class in
which most words are present is the desired class.</p>
      <p>In the second approach, for each feature we hold the
frequency of its appearance in a text. Thus a new
instance class derives from the frequency of the
presence of text words in different classes of the
problem. The class in which the most displayed and the
most frequent word of the text is the desired class.</p>
      <p>In addition to the data type, [Dör99] separated these
fields from the complexity of the steps that followed for
knowledge discovery. The general steps followed by
DM are:
(1) identifying the data collection,
(2) preparation and features selection and
(3) distribution analysis.</p>
      <p>Even though TM does not deviate from these steps, the
selection of features is different, since it is not practical
to be responsible for the examination of the features
and decide which of them should be used.</p>
      <p>The other point where they differ is in distribution
analysis, where multi-dimensional vectors are to be
treated. This implies that there must be special versions
and implementations of DM algorithms. However,
these differences do not prevent [Hea99] from
declaring that TM is an extension of DM. It is not clear
to what extent this statement may be true as there are no
studies that agree or disagree with it. Some basic
differences between TM and DM are also presented in
[Ber09] work, which are seen in Table 1. In Table 2 we
show some additional features :</p>
      <p>However, [Fan06] considers Text Mining as an
interdisciplinary field based on other disciplines, such
as Data Mining, Information Retrieval, Computational
Statistics, Computer Science and Linguistics.</p>
    </sec>
    <sec id="sec-5">
      <title>4 Text Mining vs. NLP</title>
      <p>Natural language processing (NLP) is a subfield of
computer science (CS), artificial intelligence (AI), and
linguistics concerned with the interactions between
computers and human (natural) languages. As such,
NLP is related to the area of human–computer
interaction. Many challenges in NLP involve natural
language understanding ([Nav18]), that is, enabling
computers to derive meaning from human or natural
language input, and others involve natural language
generation.</p>
      <p>TM refers to a subset of data mining concerned with
discovering knowledge from various sources;
especially, unstructured texts, which are still considered
the greatest easily accessible source of knowledge. In
TM, the main problem arises when trying to extract
explicit and implicit ideas and semantic links among
different ideas using NLP methods. The objective is to
obtain a full understanding of vast amounts of text data.
Many of the text mining algorithms extensively make
use of NLP techniques, such as part of speech tagging
(POS tagging), syntactic parsing and other types of
linguistic analysis ([Kao07]). TM is greatly connected
to NLP, but it is also related to processes in statistics,
machine learning, information extraction, information
management etc. During its procedure of finding out
hidden secrets, TM has a very important part in
upcoming applications of NLP field, like Text
Understanding ([Sal18]).</p>
      <p>Text Mining deals with the text itself, while NLP
deals with the underlying/latent metadata.</p>
      <p>Answering questions like - frequency counts of words,
length of the sentence, presence/absence of certain
words etc. is actually text mining.</p>
      <p>NLP on the other hand allows you to answer
questions like; - What is the sentiment? - What are the
keywords? (using POS tagging &amp; parsers) - What
category of content it falls under? - Which are the
entities in the sentence? etc.</p>
      <p>Text mining is the process of mining text in the
context of data mining, when we consider as data just
text. Mining is about extracting useful information from
the available data. Information could be patterns in text
or matching structures but the semantics in the text is
not considered. The goal is not about making the
system understand what does the text convey, rather
about providing information to the user based on a
certain step by step process.</p>
      <p>Natural language is what humans use for
communication. Processing such data is NLP where the
data could be speech or text. Thus, the main goal is
understanding what is the semantic meaning conveyed
in it. Therefore, we can understand why we care about
grammatical part of speeches and the lexical relations
among them.</p>
      <p>Speech recognition systems could be a part of NLP,
but it has nothing to do with TM. It may seem like NLP
is a more general, significant concept, because it uses
TM, however, it's actually the other way round. TM
uses NLP, because it makes sense to mine the data
when we understand the data semantically.</p>
      <p>Table 3 shows the top 5 Comparison between Text
Mining vs. Natural Language Processing:</p>
      <p>Table 3 : TM vs. NLP Comparison2</p>
      <sec id="sec-5-1">
        <title>Text mining NLP</title>
        <p>2
https://www.educba.com/important-text-mining-vs-naturallanguage-processing/: Accessed 5-9-2018</p>
        <p>To conclude2, both TM and NLP try to extract
information from unstructured data. TM is concentrated
on text documents and mostly depends on a statistical
and probabilistic model to derive a representation of
documents. NLP tries to get semantic meaning from all
means of human natural communication like text,
speech or even an image. NLP has potential to
revolutionize the way humans interact with machines
e.g. AWS Echo and Google Home.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5 Text Mining vs. Web Search</title>
      <p>Text Mining is different from the concept referred as
Web Search. In addition to the differences between TM
and DM, explained above, [Gup09] tries to establish
boundaries between TM and web search.</p>
      <p>The main part of web search is the web engine. A
web engine has three main parts: (1) Crawler: Gathers
the contents of all web pages (using a program called a
crawler or spider), (2) Indexer: Organizes the contents
of the pages in a way that allows efficient retrieval
(indexing), and (3) Ranker: Takes in a query,
determines which pages match, and shows the results
(ranking and display of results).</p>
      <p>The difference from TM lies in the fact that internet
users are searching for something that exists, which has
been found and was previously written by a person,
while TM is aimed at detecting previously unknown
information ([Gök15]). So, the problem is to separate
the material that is not related to your needs and keep
the essentials in order to find the information you need.</p>
    </sec>
    <sec id="sec-7">
      <title>6 Text Mining vs. Information Retrieval</title>
      <p>Information retrieval ([Rij79]) is used to search
documents or information in documents. Generally, it is
a subject of information science and computer science.
Its main uses are for access to books and journals from
universities and public libraries and the most notable
application is as web search engines. With the great
improvement of the web, a huge amount of information
is available online for the daily user. A user will try to
retrieve relevant information from web search engines
with a question or a query. Information retrieval helps
the process to return a set of documents that meets the
requirements of the user’s query.</p>
      <p>The concept of information retrieval is really old.
The first time that someone mentioned in a paper the
ability of a computer to retrieve relevant pieces of
information was in 1945 in the article “As We May
Think” by Vannevar Bush [Sin01]. Since then, many
other techniques have been shown until the last two
decades in which web search engines have boosted the
need of a large-scale information retrieval system.
There are different mathematical models for the
information retrieval. Common models are
settheoretic, algebraic and probabilistic models.
Settheoretic models represent documents as sets of words
or phrases. Algebraic models convert documents and
words in vectors, matrices and tuples. Probabilistic
models treat the information retrieval as a probabilistic
inference.</p>
      <p>It is important to differentiate between ΤΜ and
Information Retrieval (IR). We can say that TM
represents subsequent evolution (transformation) of IR.</p>
      <p>In retrieving information, the search is conducted
only for texts that already contain the answers to
questions rather than search for new knowledge
([Hea99] &amp; [Seh04]). In general, IR’s goal is to extract
all documents that are closer to the answer of a
question. Thus, it is the activity of obtaining
information resources (usually documents) relevant to
an information need from a collection of information
resources ([Fal95], [Man08]). Searches can be based
either on metadata or on full-text indexing. Therefore,
IR mostly focuses on facilitating information access
rather than analyzing information and finding hidden
patterns, which is the main purpose of text mining. IR
does not care a lot about processing or transforming
text, whereas text mining can be considered as going
beyond information access to further aid users to
analyze and understand information and ease the
decision making. Table 4 illustrates some of the
differences between TM and IR:</p>
      <p>Table 4 : Differences between TM and IR</p>
      <p>The most important distinction between TM and IR
is the output of each process. In the IR process the
result consists of documents, some of which may be
clustered, ordered or scored but at the end to get the
information we have to read the documents. In contrast
the results of TM process can be features, patterns,
connections, profiles or trends, and to find the
information we need, we don't necessary have to read
the documents.</p>
    </sec>
    <sec id="sec-8">
      <title>7 Text Mining and Statistics</title>
      <sec id="sec-8-1">
        <title>7.1 What is Statistics?</title>
        <p>Statistics consists of a set of mathematical methods
related to the collection, organization and analyzation
of data. These techniques (and more) are used so as to
extract some useful outcomes depending on our needs,
while all the potential techniques used are categorized
in two main categories the descriptive and the
inferential.</p>
        <p>In the descriptive statistics the initial data are used
only for processing reasons and producing some useful
conclusions based on them. However, no potential
forecasts are made based on this data and no results are
really inferred other than some simple outcomes only
for the current data. These predictions are actually part
of the second big category, the inferential statistics,
where useful estimations are made for future events
based on the current data.</p>
      </sec>
      <sec id="sec-8-2">
        <title>7.2 Statistics: The Science of Learning from Data</title>
        <p>Statistics is another broad subject which deals with
the study of data, that is widely applied and plays a
very important role in all areas of science. Statistics
provides the methodology for making conclusions from
data. It gives different methods to gather data, analyze
them and interpret results and is widely used by
scientists, researchers, and mathematicians in solving
problems.</p>
        <p>Though statistics provides the methods for data
collection and analysis, it helps to obtain information
from numerical and categorical data. Categorical data
refers to unique data, e.g. blood group of a person,
marital status, etc.</p>
        <p>Statistics is highly significant in data related studies
because it helps in,
 Deciding the type of data required to address a
given problem
 Organizing and summarizing data
 Analysis to be done to draw conclusions from
data
 Assessing the effectiveness of results and to
evaluate uncertainties
The methods provided by statistics include,
 Design for planning and conducting research
 Descriptions which implies exploring and
summarizing data
 Making predictions and inference using the
phenomena represented by data.</p>
        <p>So, Statistics is essentially a part of the process of
TM. It is the science of learning from data. Also, it
provides tools and techniques for dealing with large
amounts of data. Statistics includes a number of
processes, like:
 The planning behind data collections
 Data management
 Drawing inferences from numerical data facts</p>
      </sec>
      <sec id="sec-8-3">
        <title>7.3 Text Mining vs. Statistics</title>
        <p>Scientific literature suffers from lack of articles on
comparisons such as TM and Statistics, even on DM
and Statistics. So, since TM is a subfield of DM, we
will base our comparison to DM and will check if it is
valid for TM then will present it in the comparative
Table 53 bellow.</p>
        <p>In practice, comparing Statistics means comparing
what is defined in terms of a set of tools, namely those
being taught in graduate programs, i.e. Probability
theory, Real anlysis, Measure Theory, Asymptotics,
Decision theory, Markov chains, Martingales, Ergotic
theory, etc. The field of Statistics seems to be defined
as the set of problems that can be successfully
addressed with these and related topics ([Fri98]).</p>
        <p>For this reason, our comparison will not be a
thorough one, based on multiple literature resources,
but a little more simplified. Yet this analysis will still
based on some scientific criteria. Our sources will be
multiple web sources, but mainly three scientific
articles: [Fri98] by Jerome H. Friedman of Stanford
University, that explains the connection between
Statistics and DM, [Sap00] by G. Saporta, that focuses
on how DM could be used in official statistics and
[Has14] by Hassani, Saporta, and Silva, that presents a
thorough review of published work to date on the
application of data mining in official statistics, and on
identification of the techniques that have been
explored.
3
https://www.educba.com/data-mining-vs-statistics/:
Accessed 5-9-2018
Suitable for large data sets
Suitable for smaller data sets</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>8 Conclusion</title>
      <p>In this article we attempted to briefly describe the
differences of Text Mining with other related
disciplines, while making a concise presentation.</p>
      <p>In summary, it is noted that TM and all these
sciences (even statistics) may seem indistinguishable
due to its close connection. It is clear, however, that
statistics is actually a tool or method for all these
sciences, while most of them spread over a wide
domain where a statistical method is an essential
component. Text Mining has developed recently with
big data and will continue to grow in the following
years as data growth seems to be never-ending. This
also applies to the other disciplines, which means that
the data driving the algorithms, methods and decisions
need to be high-quality. Nonetheless, all disciplinary
fields described briefly in this review, cover the major
areas of working with data and problems on various
areas related to this data. The emerging picture reveals
a blend of theory and practice that reflects each
discipline rather than a unified system. Hopefully, a
productive merging of TM approaches through
increased cross-disciplinary research can develop and
advance not only TM but all these fields. The rate of
change in the text mining field is so rapid that the
information is likely to be measurably different in the
following years.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Ber09]
          <string-name>
            <given-names>C.</given-names>
            <surname>Berkouwer</surname>
          </string-name>
          . Master Thesis:
          <article-title>The Reflection of Foresight in Defense Policy Making : A Comparative Study of the United Kingdom and the United States</article-title>
          ,
          <year>March</year>
          2009
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Dea14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Big Data, Data Mining, and Machine Learning: Value Creation for Business Leaders</article-title>
          and Practitioners: pp
          <fpage>56</fpage>
          . John Wiley and Sons, Inc.,
          <year>2014</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Dör99]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dörre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gerstl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Seiffert</surname>
          </string-name>
          . Finding Text Mining:
          <article-title>Nuggets in Mountains of Textual Data</article-title>
          .
          <source>KDD '99 Proceedings of the Fifth ACM SIGKDD Intern. Conference on K. Discovery and Data Mining</source>
          :
          <fpage>398</fpage>
          -
          <lpage>401</lpage>
          ,
          <year>August 1999</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Fan06]
          <string-name>
            <given-names>W.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wallace</surname>
          </string-name>
          , S. Rich, &amp;
          <string-name>
            <surname>Z.</surname>
          </string-name>
          <article-title>Zhang Tapping the Power of Text Mining</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>49</volume>
          (
          <issue>9</issue>
          ):
          <fpage>76</fpage>
          -
          <lpage>82</lpage>
          ,
          <year>September 2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Fal95]
          <string-name>
            <given-names>C.</given-names>
            <surname>Faloutsos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W</given-names>
            <surname>Oard</surname>
          </string-name>
          .
          <article-title>A survey of information retrieval and filtering methods</article-title>
          .
          <source>Technical Report</source>
          . University of Maryland at College Park, MD, USA 1995
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Fra92]
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Frawley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Piatetsky-Shapiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Matheus</surname>
          </string-name>
          .
          <article-title>Knowledge Discovery in Databases : An Overview</article-title>
          .
          <source>AI Magazine</source>
          ,
          <volume>13</volume>
          (
          <issue>3</issue>
          ):
          <fpage>57</fpage>
          -
          <lpage>70</lpage>
          ,
          <year>September 1992</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Fri98]
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          .
          <source>Data Mining and Statistics: What's the Connection? Computing Science and Statistics</source>
          Vol.
          <volume>29</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3</fpage>
          -9, Ed. D. Scott 1998
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Gri08]
          <string-name>
            <given-names>S.</given-names>
            <surname>Grimes</surname>
          </string-name>
          .
          <article-title>Unstructured data and the 80 percent rule</article-title>
          .
          <source>Clarabridge Bridgepoints newsletter 23</source>
          , column, “Experts Corner: Seth Grimes.”,
          <year>August 2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Gök15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gök</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Waterworth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shapira</surname>
          </string-name>
          .
          <article-title>Use of web mining in studying innovation</article-title>
          .
          <source>Scientometrics</source>
          <volume>102</volume>
          (
          <issue>1</issue>
          ):
          <fpage>653</fpage>
          -
          <lpage>671</lpage>
          , Jan. 2015
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Gup09]
          <string-name>
            <given-names>V.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lehal</surname>
          </string-name>
          .
          <article-title>A survey of text mining techniques and applications</article-title>
          .
          <source>Journal of Emerging Technologies in Web Intelligence</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>60</fpage>
          -
          <lpage>76</lpage>
          ,
          <string-name>
            <surname>Academy</surname>
            <given-names>Publisher</given-names>
          </string-name>
          ,
          <year>August 2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Has14]
          <string-name>
            <given-names>H.</given-names>
            <surname>Hassani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Saporta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Silva</surname>
          </string-name>
          .
          <article-title>Data Mining and Official Statistics: The Past, the Present and the Future</article-title>
          .
          <source>Big Data</source>
          Vol.
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>34</fpage>
          -
          <lpage>43</lpage>
          ,
          <year>March</year>
          2014
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Hea99]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <article-title>Untangling Text Data Mining</article-title>
          .
          <source>Proceedings of ACL '99: the 37th Annual Meeting of the Association for computational Linguistics</source>
          , University of Maryland,
          <year>June 1999</year>
          (invited paper)
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Kan09]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Baumgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>McCrohon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ananiadou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. B.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tsujii</surname>
          </string-name>
          .
          <article-title>Data mining: concept and techniques</article-title>
          .
          <source>Oxford Journal of Bioinformatics</source>
          , Volume
          <volume>25</volume>
          , Issue 15:
          <fpage>1997</fpage>
          -
          <lpage>1998</lpage>
          ,
          <article-title>August 2009</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Kao07]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Poteet</surname>
          </string-name>
          .
          <source>Natural language processing and text mining</source>
          . Springer,
          <year>2007</year>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Kar05]
          <string-name>
            <given-names>H.</given-names>
            <surname>Karanikas</surname>
          </string-name>
          , Th.
          <source>Mavroudakis. Text Mining Software Survey. RANLP Text Mining Workshop No</source>
          <volume>1</volume>
          :
          <fpage>39</fpage>
          -
          <lpage>48</lpage>
          ,
          <year>September 2005</year>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Lia12]
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. Y.</given-names>
            <surname>Hsiao. Data Mining</surname>
          </string-name>
          <string-name>
            <surname>Techniques</surname>
          </string-name>
          &amp;
          <article-title>Applications - A Decade Review from 2000 to 2011</article-title>
          .
          <article-title>Expert Systems with Applications</article-title>
          , Vol.
          <volume>39</volume>
          (
          <issue>12</issue>
          ):
          <fpage>11303</fpage>
          -
          <lpage>11311</lpage>
          ,
          <string-name>
            <given-names>Elsevier</given-names>
            <surname>Ltd</surname>
          </string-name>
          .,
          <source>September 2012</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Liu11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>X. Lu.</surname>
          </string-name>
          <article-title>Survey on text clustering algorithm</article-title>
          .
          <source>Proceedings of 2nd International IEEE Conference on Software Engineering and Services Science (ICSESS)</source>
          ,
          <year>China</year>
          ,
          <fpage>901</fpage>
          -
          <lpage>904</lpage>
          ,
          <year>2011</year>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>[Man08] C. D. Manning</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schütze</surname>
          </string-name>
          . Introduction to Information Retrieval. Cambridge University Press, New York 2008
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Nah02]
          <string-name>
            <given-names>U. Y.</given-names>
            <surname>Nahm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Mooney</surname>
          </string-name>
          .
          <article-title>Text Mining with Information Extraction</article-title>
          .
          <source>Technical Report SS02-06</source>
          , Department of Computer Sciences, University of Texas, March 2002
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [Nav18]
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Navigli</surname>
          </string-name>
          . Natural Language
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Artificial</given-names>
            <surname>Intelligence</surname>
          </string-name>
          :
          <fpage>5697</fpage>
          -
          <lpage>5702</lpage>
          , Early
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Career</surname>
          </string-name>
          ,
          <article-title>July 2018</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>[Raj97] M. Rajman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Besançon</surname>
          </string-name>
          .
          <article-title>Text mining: Natural language techniques and text mining applications. Data Mining and Reverse Engineering: Searching for Semantics: IFIP TC2 WG2.6 IFIP 7th Conference on Database Semantics (DS-7</article-title>
          ):
          <fpage>50</fpage>
          -
          <lpage>66</lpage>
          ,
          <year>January 1997</year>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Rij79]
          <string-name>
            <given-names>C. J. Van Rijsbergen. Information</given-names>
            <surname>Retrieval</surname>
          </string-name>
          , London: Butterworths, 2nd edition, November 1979
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Sal18]
          <string-name>
            <given-names>S.A.</given-names>
            <surname>Salloum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.Q.</given-names>
            <surname>AlHamad</surname>
          </string-name>
          , M.
          <string-name>
            <surname>Al-Emran</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Shaalan</surname>
          </string-name>
          .
          <article-title>A Survey of Arabic Text Mining</article-title>
          .
          <source>Studies in Computational Intelligence</source>
          , vol
          <volume>740</volume>
          :
          <fpage>417</fpage>
          -
          <lpage>431</lpage>
          , Springer, Cham, January 2018
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [Sap00]
          <string-name>
            <given-names>G.</given-names>
            <surname>Saporta</surname>
          </string-name>
          .
          <source>Data Mining and Official Statistics. Quinta Conferenza Nationale di Statistica</source>
          , ISTAT, Roma, November 2000
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Seh04]
          <string-name>
            <given-names>A.K.</given-names>
            <surname>Sehgal</surname>
          </string-name>
          . Text Mining:
          <article-title>The Search for Novelty in Text</article-title>
          .
          <source>Ph.D. Comprehensive Examination Report</source>
          , Dept. of Computer Science, The University of Iowa, April 2004
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [Sin01]
          <string-name>
            <given-names>A.</given-names>
            <surname>Singhal</surname>
          </string-name>
          .
          <article-title>Modern Information Retrieval: A Brief Overview</article-title>
          .
          <source>Bulletin of the IEEE Computer Society Technical Committee on Data Engineering</source>
          <volume>24</volume>
          (
          <issue>4</issue>
          ):
          <fpage>35</fpage>
          -
          <lpage>43</lpage>
          ,
          <year>December 2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [Tal16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Talib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kashif</surname>
          </string-name>
          , Sh.
          <string-name>
            <surname>Ayesha</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Fatima</surname>
          </string-name>
          . Text Mining: Techniques, Applications and Issues.
          <source>International Journal of Advanced Computer Science &amp; Applications</source>
          Vol.
          <volume>7</volume>
          (
          <issue>11</issue>
          ):
          <fpage>414</fpage>
          -
          <lpage>418</lpage>
          , November 2016
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>