<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Lviv, Ukraine, November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Building a Feature Taxonomy of the Terms Extracted from a Text Collection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Svitlana Moiseyenko</string-name>
          <email>svitlana.moiseyenko@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Vasileyko</string-name>
          <email>vasileyko.alex@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vadim Ermolayev</string-name>
          <email>vadim@ermolayev.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Zaporizhzhia National University</institution>
          ,
          <addr-line>Zhukovskogo st. 66, Zaporizhzhia</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>1</volume>
      <fpage>5</fpage>
      <lpage>16</lpage>
      <abstract>
        <p>This position paper presents an approach for feature grouping and taxonomic relationship extraction with the further objective to build a feature taxonomy of a learned ontology. The approach needs to be developed as a part of the OntoElect methodology for domain ontologies refinement. The paper contributes a review of the related work in taxonomic relationships extraction from natural language texts. Within this review, the research gaps and remaining challenges are analyzed. The paper proceeds with outlining the envisioned solution. It presents the approach to this solution starting with the research questions, followed by the initial research hypotheses to be tested. Consequently, the plan of research is presented, including the potential research problems, the rationale to use and re-use existing components, and evaluation plan. Finally, the proposed solution, and the project, are placed in the broader context of the overall OntoElect workflow.</p>
      </abstract>
      <kwd-group>
        <kwd>feature taxonomy</kwd>
        <kwd>relation extraction</kwd>
        <kwd>subsumption</kwd>
        <kwd>meronymy</kwd>
        <kwd>ontology engineering</kwd>
        <kwd>OntoElect</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper presents an approach to building a taxonomy of the features, indicating the
elicited requirements, within the OntoElect methodology [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. OntoElect is an ontology
refinement methodology, which consists of the following phases: (i) Feature
Elicitation; (ii) Conceptualization and Formalization; (iii) Evaluation. Building the taxonomy
of the developed or refined ontology is a major step within the Conceptualization and
Formalization phase. It is composed of two steps: (i) Feature Grouping; and (ii)
Relationship Extraction.
      </p>
      <p>In this paper, we present our vision of an advanced approach for extracting specific
types of relationships, namely subsumption and meronymy. These relationships are to
be further used for building a feature taxonomy for the ontology under development or
refinement. The approach is advanced as it combines the highlights of the existing
approaches that are known from the related work. Due to this combination, we envision
that the precision of the result will be higher than that of the State-of-the-Art.</p>
      <p>
        The vision of the approach is presented as an M.Sci project proposal, as it: (i) fits as
a part of the developed technology for the Feature Grouping phase of OntoElect; and
(ii) may further be extended to extract other kinds of relationships to formalize the
requirements for an ontology under development [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>The reminder of the paper is structured as follows. Section 2 reviews the related
work on ontology learning from text and taxonomy extraction. Based on this review it
outlines the research gaps. Sect. 3 presents our motivation to narrow the outlined
research gaps by combining the advantages of the State-of-the-Art approaches with the
knowledge we have already made available in the Requirements Elicitation phase of
OntoElect. Research questions and initial research hypotheses are formulated. Sect. 4
presents our informed vision of the approach to developing the proposed solution and
the plan of the project. Sect. 5 puts the proposed work and the envisioned solution in
the broader context of the OntoElect project. Finally, we offer our conclusive remarks
and a perspective view on the relevant future work in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>Various taxonomy extraction approaches have been studied in the related work. These
approaches, reviewed below, used different combinations of knowledge sources for
extracting taxonomic relationships. Based on the combinations of these sources, the
related work could be grouped into pattern-based, linguistic, statistical, graph-based,
terminology-based, and dynamic clustering types of methods.</p>
      <p>
        The pioneering works on pattern-based approaches to taxonomy extraction were [
        <xref ref-type="bibr" rid="ref3 ref4">3,
4</xref>
        ]. However, the patterns, proposed initially, were too specific to cover all the required
true positive cases that can be met in natural language texts. Therefore, further research
in the pattern-based direction looked into devising more flexible and generalizing
patterns, following and extending [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ]. Some authors explored the opposite direction
and developed more specific “doubly-anchored” patterns [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Some results have been reported in combining linguistic and statistical approaches
to taxonomic relationships extraction. It was reported that such a combination of the
appropriate techniques allows increasing the precision of results [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10, 11</xref>
        ]. Some other
works exploited negative information within a linguistic approach [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. A notable
cluster of research followed a hybrid approach and combined NLP and pattern-based
techniques, resulting in the proposal of lexico-syntactic patterns (LSP). LSP were used for
linguistic and statistical matching in retrieving concepts and their relationships [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        Some authors have noted that the use of plain texts for retrieving relationships does
not yield required quality. As an additional source of quality information, they proposed
to exploit an existing lexical resource, WordNet [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] or other Web resources for
validating the relationships extracted from text [
        <xref ref-type="bibr" rid="ref14 ref6">6, 14</xref>
        ]. Further, graph-based techniques in
NLP were used to parse textual knowledge sources and evaluate the obtained results
using Word-Net [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        A cluster of contributions adding to the quality of relationship extraction used
additional knowledge provided by the terms extracted from texts, or text collections /
corpora [
        <xref ref-type="bibr" rid="ref1 ref15 ref16 ref17 ref18 ref6">15, 16, 6, 1, 17, 18</xref>
        ]. Some of the works in this category used post-processing to
improve recall and precision. This post-processing was done by applying: (i)
statisticsbased cuts [
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ]; (ii) domain specificity / significance scores [
        <xref ref-type="bibr" rid="ref1 ref17 ref21">1, 17, 21</xref>
        ].
      </p>
      <p>
        In addition to the reviewed typology of techniques, it needs to be mentioned that,
broadly, the methods for taxonomy extraction could be classified as (i) unsupervised;
(ii) semi-supervised; and (iii) supervised. Unsupervised methods are domain-neutral as
these techniques do not require a domain-specific training set or other bits of expertize
(e.g. specific rules) that help improve the quality of extraction. For instance, the feature
grouping technique [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is domain neutral as it is based on the partial matching of
candidate term strings using a selection of string similarity measures. Supervised methods
are mainly based on the use of domain-specific machine learning models. Topical
representatives of this category are Word2Vec [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], GloVe [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and SensEmbed [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
Importantly, these supervised methods exploit terms embeddings in the text which are
important source of knowledge about relationships and the contexts of terms.
Semisupervised methods were considered in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] based on a handful of labeling and model
learning techniques to extract categories (concepts) and relations from web pages. A
semi-supervised algorithm was developed [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], that uses a root concept and recursive
surface pattern to learn automatically from the Web relation pairs.
      </p>
      <p>
        Finally, an early stage work using dynamic clustering for retrieving taxonomic
relations (a supervised approach) was proposed in [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. In its idea, this work resembles our
background work on feature grouping [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The difference is that [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] groups terms and
devises taxonomic relationships based on candidate strings matching.
      </p>
      <p>Source-wise, all the reviewed approaches exploit only a part of the available
information to improve the quality of taxonomy extraction. The proposal of this position
paper is to develop a method that exploits all of the available types of knowledge:
terminology with significance scores; LSP of term definitions; term embeddings in the
text; available commonsense lexical databases (like WordNet).
3</p>
    </sec>
    <sec id="sec-3">
      <title>Motivation, Background, Research Questions, and Hypotheses</title>
      <p>The motive to develop the hybrid approach envisioned in this paper is the extraction of
a feature taxonomy, in an unsupervised way, with higher quality than achieved
currently by the State-of-the-Art techniques (Section 2). Our premise is that using all the
available sources of information about potential taxonomic relationships will result in
filtering out false positives while keeping the true positives not negatively affected.
This will increase the balanced F-measure of our extraction – so the quality will be
higher. A pictorial representation of the envisioned approach is given in Fig. 1.</p>
      <p>
        Our approach exploits the background knowledge of the OntoElect project1 in
extracting a saturated [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] bag of terms from the document collection using the
descending citation frequency ordering of the documents in the collection to balance
terminol1
https://www.researchgate.net/project/OntoElect-a-Methodology-for-Domain-Ontology-Refinement
ogy drift in time and the difference in the terminological impact of individual
documents [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Term extraction provides a flat ordered list of the terms with their
significance values. In OntoElect, these terms are further termed as features identifying the
requirements for the refined ontology. This flat list of features could further be
transformed in a set of hierarchical groupings [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] hinting about existing subsumption or
meronymy relationships between some of the terms. These background-based steps are
numbered as (1) and (2) in Fig. 1.
      </p>
      <p>Order in Descending
Citation Frequency
(CF)</p>
      <p>1
Extracted Terms
Term Score
… …</p>
      <p>2
Extract
and
Group 3
terms</p>
      <p>Def:
…
Def:</p>
      <p>Texts
…</p>
      <p>CF=…
Improve using
Term Definitions</p>
      <p>Refine using
WordNet
4</p>
      <p>Definition
Embeddings</p>
      <p>(LSP)
WordNet Fragment
According to our background knowledge, a draft taxonomy built of the groupings
(step 2) is of poor quality. The reason for that is that only the information about
approximate string matches is used for grouping. Hence, the refinement is required. In this
context, our first research question is: (RQ1) Are there more bits of relevant
information in the text collection used for term extraction at step (2) that may help refine the
feature taxonomy built using grouping?</p>
      <p>
        Our corresponding research hypothesis (step (3) in Fig. 1) is:
(H1) The feature taxonomy may be improved by:
(i) Extracting term definition embeddings from the documents containing
the involved terms and having the highest citation frequency (CF), following
a domain-neutral rule-based approach based on LSP [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
(ii) Doing Part of Speech (PoS) tagging of these embeddings
(iii) Extracting potential taxonomic relationships from the PoS taggings
One more research question relevant regarding the quality aspect is (RQ2): Are all
the relationships, extracted at step (3), taxonomic relationships; has the semantics been
properly extracted? We propose to answer RQ2 by validating the extracted
relationships using the general-purpose (commonsense) knowledge provided by WordNet.
This is reflected as step (4) in Fig. 1. Suppose that an extracted taxonomic relationship
is supported by an explicit representation of it in WordNet, regarding a similar term or
its superclass. Then, such a support could be regarded as a commonsense evidence of
the correctness of this relationship. Otherwise, if a relationship is not supported, the
possibility of it to be a false positive grows higher. A weak aspect in this approach is
that WordNet is not supposed to contain and represent professional terminology, in a
subject domain. Hence, there is a risk that a relationship of a valid feature does not find
its support. As a remedy, the significance score of the respective term could be
evaluated. The higher the confidence score, the higher is the probability that the feature, and
the relationship, are true positives. In this context, our research hypothesis is:
(H2) The quality of the taxonomic relationships in a feature taxonomy could be
validated, in a balanced way, by:
(i) Seeking for the support of this relationship by the evidence provided by
Word
      </p>
      <p>Net
(ii) Balancing the lack of the coverage of professional terminology in WordNet by
high significance scores of respective terms retrieved at step (2)
4</p>
    </sec>
    <sec id="sec-4">
      <title>Approach to Solution and Research Plan</title>
      <p>The outline of the proposed solution for improving the quality of taxonomic
relationships extraction from texts has been given in Sect. 3. In this section, we discuss, in more
detail, the research problems that have to be solved on the approach to the working
solution. This discussion is structured using the chosen methodology and thus
represents our proposed research plan.</p>
      <p>
        Methodologically, the proposed project follows the pattern of the Scientific Method
[
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] and approaches its objectives iteratively. Every iteration starts with the formulation
or refinement of a research hypothesis. Our initial hypotheses have been formulated in
Sect. 3 as H1 and H2. Within the subsequent phase of an iteration, the instruments for
testing the hypothesis are materialized, e.g. software is implemented, adopted, or
adapted, and dataset(s) prepared. Finally, the hypothesis is tested in an evaluation
experiment and the results are assessed regarding their proximity to the project objective.
      </p>
      <p>In the remainder of this section, we outline our vision of the first iteration following
the abovementioned pattern of “hypothesize – materialize – evaluate”.
4.1</p>
      <sec id="sec-4-1">
        <title>Potential Research Problems in the Processing Pipeline</title>
        <p>Until now, we presented the envisioned approach to feature taxonomy extraction
without looking into the technical details. In this subsection, we point out the parts of the
processing pipeline, which may be problematic to develop without elaborating these
technical details and, possibly, refining or revising our research hypotheses H1 and H2.</p>
        <p>
          Term groupings and false positive relationships (H1). Our prior work on terms
grouping [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] revealed that a pair of term candidate strings might have high similarity
(0.85-0.95 out of 1.00) but carry similar (partial positives – PP) or fully different (partial
negatives - PN) semantics. PP pairs are exactly the cases in which taxonomic
relationships may be sought. However, PNs are the false positive pairs. We admitted in [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]
that there was no reliable way to filter out PNs as the similarity value intervals overlap
substantially for PPs and PNs. This was perhaps one of the major reasons for term
grouping to yield low quality in taxonomic relationships extraction. To improve this
quality, the analysis of the composition of the candidate strings might be helpful. If for
example, the PP term strings are “time interval” and “unbounded time interval” then
the second string in the pair differs from the first by the added word token “unbounded”.
Hence, the addition of a word token might indicate that a taxonomic relationship exists
in this pair. Furthermore, the added word token is an adjective. Therefore, the
recognition of PPs could be improved if the information about the parts of speech is exploited.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Patterns for reliably locating term definitions in text (H1). We devise the ra</title>
        <p>tionale for using term definitions for extracting the relationships (and properties) of
these terms from the analysis of the deliberation patterns used in defining things by
humans. Indeed, as known from Cognitive Science, humans define things:
• Either top-down, by: (i) relating the defined thing to the known thing that is more
abstract or general; and (ii) specifying the additional properties for the defined thing.
For example, “an unbounded time interval is a time interval, whose endings are not
fixed”.
• Or bottom-up, by collecting the things that are less abstract and finding their
common properties. For example, “an unbounded time interval, an open time interval, a
convex time interval are all time intervals as these are not instant in time, i.e. have
duration”.</p>
        <p>Our premise, in the context of H1, is that the parts of text that define terms could be
distinguished from the rest of the text based on their structure. Hence, the patterns of
this structure for all the ways of defining things have to be developed and tested. This
might be not that straightforwardly easy as outlined above. One possible reason is that,
even for a fixed way of defining things, there might be different styles of formulations
in a natural language. For example, for a top-down definition, the following two
definitions are correct but follow different stylistic patterns: (i) “an unbounded time interval
is a time interval, whose endings are not fixed”; (ii) “a time interval, whose endings are
not fixed is an unbounded time interval”.</p>
        <p>A proper size of a term embedding in a definition text fragment (H1). The
definitions of terms could be given in one sentence, but could also span across several
consecutive sentences. Hence, an open question in the context of locating a term definition
embedding is how broad, in sentences, is has to be. Currently, we do not envision any
formal way to specify the threshold. It is planned that the valid samples of term
definitions need to be collected and analyzed. Furthermore, different thresholds need to be
tested in the evaluation experiments on “gold standard” datasets (Sect. 4.3).</p>
        <p>Weighting WordNet support and term significance scores (H2). Intuitively, it
might be straightforward that if a candidate string, which is not a stop word, frequently
appears in the text then it might be a term used in the text. This is reflected by a high
significance score of such a string. Consequently, if there are indications of a taxonomic
relationship between two strings with high significance scores then our confidence in
the existence of this relationship is ought to be high. It is excellent if such a relationship
gets explicit support in a lexical resource, like WordNet. However, it is not always the
case as terminologies belonging to specialist domains may not be included in these
lexical resources. Therefore, an open question in this context is what sort of evidence
prevails – high text colocation confidence or no explicit support in WordNet. In the
proposed approach, it is planned to find the proper balance between these two kinds of
evidence by applying linear weighting. The weights will be empirically determined in
the experiments on the “gold standard” datasets.
4.2</p>
      </sec>
      <sec id="sec-4-3">
        <title>Materialization of the Required Components</title>
        <p>To implement the project pipeline, several research and development tasks have to be
planned regarding algorithmic and software implementation.</p>
        <p>
          Term grouping software, to be used at step (2) is available as our background as a
proof-of-concept implementation [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. One of the shortcomings of this implementation
is that the run times are too big for realistic full text document collections. The plan is
to optimize this software by applying an efficient multiple string matching algorithm
[
          <xref ref-type="bibr" rid="ref33">33</xref>
          ].
        </p>
        <p>For step 3, it is planned to develop the lexico-syntactic patterns based on the analysis
of the grammatically parsed text fragment definitions. For text parsing and PoS tagging,
Stanford CoreNLP2 will be used as the core API. It allows automatically elaborating
the structure of sentences in terms of phrases, parts of speech, and syntactic
dependencies. Hence, we expect to receive initial indicators for detecting the definitions of
features in a domain-neutral way. Based on these taggings and the analysis of their types
at a sentence level, the rules will be developed for extracting taxonomic relations from
the feature definition embeddings.</p>
        <p>For manipulating the elements in the WordNet database, it is planned to use the
LxicaDB WordNet HTTP API 3, which allows querying WordNet via http. In addition, an
algorithm for measuring the impact of the external support, balanced with the allocation
confidence, needs to be developed for the instrumental pipeline to support evaluation
experiments.
4.3</p>
      </sec>
      <sec id="sec-4-4">
        <title>Evaluation Methodology and Plan</title>
        <p>The goal of the evaluation phase is to test the hypotheses and find out how close the
results are to the expected state of affairs. For evaluation at each iteration, in addition
to the materialized solution (Sect 4.2), evaluation objectives, methodology, plan, data,
and execution environment have to be specified. In this section, we outline the
components, which will presumably remain unchanged in iterations, like methodology.
We also specify the rest for the initial hypotheses H1 and H2 (Sect. 3).</p>
        <p>Evaluation Objectives. The evaluation objectives are formulated in a way to check
if the tested hypotheses hold true. As for H1, the objective (EO1) is to measure the
quality of the extraction of taxonomic relationships using the steps 2, 3, and 4 of the
processing pipeline (Fig. 1) and find out if the measured quality increases from step to
step. To test H2 the objective is twofold. The first sub-objective (EO2.1) is finding out
if there is support for the extracted relationships in WordNet and measure the
proportion of the supported versus unsupported taxonomic relationships. This will presumably
2 https://stanfordnlp.github.io/CoreNLP/
3 http://www.lexicadb.com/lxserver/wn_http.html
allow verifying the utility of WordNet as a source of information where such support
could be sought. Secondly, and if the support found in WordNet is sufficient, the
subobjective (EO2.2) is to find out the proportion of the unsupported true positives, having
high significance scores. Achieving this sub-objective may allow finding proper
weighting for reaching the balance mentioned in H2.</p>
        <p>
          Evaluation Methodology. The outlined evaluation objectives provide the rationale
for choosing a proper methodology for experimental evaluation. EO1 suggests that it
would be rational to organize the experiments using the pattern of an ablation study4 in
which the quality measurements after steps 4, 3, and 2 are compared. Both EO1 and
EO2 are based on quality measurement. One of the mainstream approaches for
measuring the quality of extraction is the use of a measure based on a combination of recall
and precision. One of the most often used measures of this sort in information retrieval
is balanced F-measure [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]. The problem with it is that true/false positives and
negatives have to be reliably identified and counted in order to get the trusted value. This is
done using an appropriate “gold standard” dataset. In the context of the presented
project, in a “gold standard” dataset all valid terms/features, taxonomic relationships
among the features, and a feature taxonomy have to be manually extracted by a domain
expert. The result of this manual extraction is regarded as the ground truth, which is
used to measure the quality of automated extraction. Furthermore, as the goal of the
proposed research is developing a domain-neutral (unsupervised) solution, we need
several “gold standard” datasets for evaluation.
        </p>
        <p>
          Evaluation Data. Based on the evaluation methodology, we are seeking for several
collections of professional documents in different subject domains. Each of these
collections has to be processed by domain experts to be used as a “gold standard” dataset.
The experts have to manually extract terms, taxonomic relationships, and build the
feature taxonomy. This work, if done from scratch, is too laborious for being feasibly
undertaken in an individual (Master) project. Therefore, an exploratory search for
available manually pre-processed datasets have to be undertaken at the initial phase of the
project. Currently, the following collections could be pointed to as the ones for which
the required pre-processing has been done at least in part:
• The TIME paper collection and Syndicated Ontology of Time (SOT) [
          <xref ref-type="bibr" rid="ref1 ref29">29, 1</xref>
          ]. SOT
is available as our background result in the OntoElect project. It contains the feature
taxonomy manually built [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] based on the terms extracted from the TIME document
collection5. Hence, the taxonomy and taxonomic relationships are the available parts
of the “gold standard” for TIME and could be used as one of the “gold standard”
datasets in our experimental evaluation.
• Gold standard taxonomies of the SemEval initiative, taxonomy extraction
evaluation task [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ]. SemVal offers several gold standard taxonomies in different domains
together with the sets of extracted terms. One shortcoming of these data is that the
4 Ablation study is the experimental procedure for testing deep neural networks in which certain
parts of the network, e.g. layers, are gradually removed to understand the influence of each
individual part on the overall result.
5 TIME collection in plain texts:
http://dx.doi.org/10.17632/knb8fgyr8n.1#folder-d1e5f2b6c51e-4572-b10d-0e2ebccead02
source document collections are not provided. Secondly, the significance scores of
the extracted terms are not provided. Hence, these datasets could be used for
crossevaluation at the final iterations of our research workflow.
        </p>
        <p>Evaluation Plan. The plan of performing evaluation experiments is
straightforwardly inferred from Fig. 1. In each experiment for each individual dataset: (i) the
processing step (2, 3, then 4) is executed; (ii) in each step the quality of taxonomy
extraction is measured (balanced F-measure) by comparing the result of extraction to the gold
standard result.</p>
        <p>Execution Environment. All the planned computations are sufficiently lightweight
to be run on a conventional laptop computer in affordable time.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>The Work in a Broader Context</title>
      <p>
        The proposed research is targeted to fill in the research gap in the Conceptualization
Phase of the OntoElect approach for domain ontology refinement. The goal of
OntolElect is to provide the methods and instrumental support for the refinement of a domain
ontology in an arbitrary subject domain. The idea of the approach is that the
requirements for an ontology are extracted, from a complete document collection describing
the subject domain, as terms (features) in the Requirements Elicitation phase. In the
Conceptualization phase, these requirements are formalized as ontological fragments.
For that, different features, representing concepts, are grouped and the feature
taxonomy is built. Finally, the features, representing properties are added to the nodes in the
feature taxonomy. Hence, the formalized requirements are formed ass ontological
contexts [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The workflow is pictured in Fig. 2. The part of this workflow that covers the
scope of the proposed work is given in the inner rounded rectangle with solid border.
As it is seen from the sequence of tasks, the solution that will be developed implements
the instrumental support for the steps of Feature Grouping and Categorization and
Building the Feature Taxonomy.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusive Remarks and Outlook</title>
      <p>In this position paper, we outlined our vision of a focused M.Sci proposal, which is
topically aligned with our plans in the OntoElect project. We analyzed the related work
in the field of extracting taxonomic relationships from a natural language text. Based
on this analysis, we outlined the research gaps in the form of the open research
questions and formulated the relevant research hypotheses that emerged from our vision of
the potential solution.</p>
      <p>The idea of this proposal is to combine the use of more relevant sources of
information about taxonomic relationships, than used in the State-of-the-Art solutions to
date. Hence, the proposed approach is hybrid. We proposed to organize the envisioned
solution as a four-step processing pipeline in a way to add more extraction quality at
each consecutive step. We also presented our plan for performing this research that
follows the pattern “hypothesize – materialize – evaluate” within each iteration. In the
plan, we included our rationale for choosing the methodology, outlined potential
research problems related to the steps of the proposed pipeline, and presented the way to
perform experimental evaluation in the project.</p>
      <p>Our planned future work related to the proposed project is extending the solution to
the extraction of all types of relationships for incorporating properties into the
ontological representations of requirements in the subsequent steps of the Conceptualization
phase workflow.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ermolayev</surname>
          </string-name>
          , V.:
          <article-title>OntoElecting requirements for domain ontologies. The case of time domain</article-title>
          .
          <source>EMISA Int J of Conceptual Modeling</source>
          <volume>13</volume>
          (
          <article-title>Sp</article-title>
          . Issue),
          <fpage>86</fpage>
          -
          <lpage>109</lpage>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .18417/emisa.si.
          <source>hcm.9</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Moiseyenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Conceptualizing and formalizing requirements for ontology engineering</article-title>
          . In: Antoniou,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Zholtkevych</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . (eds.)
          <source>PhD Symposium at ICTERI</source>
          <year>2018</year>
          ,
          <article-title>CEUR-WS</article-title>
          , vol.
          <volume>2122</volume>
          , pp.
          <fpage>35</fpage>
          -
          <lpage>44</lpage>
          (
          <year>2018</year>
          ). online
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Automatic acquisition of hyponyms from large text corpora</article-title>
          .
          <source>In: 14th Conf on Computational Linguistics</source>
          , pp.
          <fpage>539</fpage>
          -
          <lpage>545</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kozareva</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hov</surname>
          </string-name>
          , E.:
          <article-title>A semi-supervised method to learn and construct taxonomies using the web</article-title>
          .
          <source>In: Proc. 2010 Conf on Empirical Methods in Natural Language Processing, EMNLP 2010</source>
          , pp.
          <fpage>1110</fpage>
          -
          <lpage>1118</lpage>
          , MIT, Massachusetts, USA (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ritter</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>What is this, anyway: automatic hypernym discovery</article-title>
          .
          <source>In: AAAI 2009 Spring Symposium on Learning by Reading and Learning to Read</source>
          , pp.
          <fpage>88</fpage>
          -
          <lpage>93</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Tuan</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Taxonomy construction using syntactic contextual evidence</article-title>
          .
          <source>In: 2014 Conf on Empirical Methods in Natural Language Processing, EMNLP 2014</source>
          , pp.
          <fpage>810</fpage>
          -
          <lpage>819</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Snow</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
          </string-name>
          . A.:
          <article-title>Learning syntactic patterns for automatic hypernym discovery</article-title>
          .
          <source>In: 17th Annual Conf on Neural Information Processing Systems</source>
          , pp.
          <fpage>1297</fpage>
          -
          <lpage>1304</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Betteridge</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hruschka</surname>
          </string-name>
          Jr., E.R., Mitchell, T.M.:
          <article-title>Coupled semisupervised learning for information extraction</article-title>
          .
          <source>In: 3d Int Conf on Web Search and Web Data Mining</source>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>110</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Diederich</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balke</surname>
          </string-name>
          , W.-T.:
          <article-title>The semantic growbag algorithm: automatically deriving categorization systems</article-title>
          .
          <source>In: Research and Advanced Technology for Digital Libraries, 11th European Conf, ECDL</source>
          <year>2007</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cafarella</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Downey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kok</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popescu</surname>
            <given-names>A.-M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaked</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yates</surname>
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Web-scale information extraction in knowitall (preliminary results)</article-title>
          .
          <source>In: 13th Int Conf on World Wide Web</source>
          , pp.
          <fpage>100</fpage>
          -
          <lpage>110</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>K.O.</given-names>
          </string-name>
          :
          <article-title>Probase: a probabilistic taxonomy for text understanding</article-title>
          .
          <source>In: ACM SIGMOD Int Conf on Management of Data</source>
          , pp.
          <fpage>481</fpage>
          -
          <lpage>492</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Chinese hypernym-hyponym extraction from user generated categories</article-title>
          .
          <source>In: 26th Int Conf on Computational Linguistics</source>
          , pp.
          <fpage>1350</fpage>
          -
          <lpage>1361</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Bennamoun</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Ontology learning from text: a look back and into the future</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>44</volume>
          (
          <issue>4</issue>
          ),
          <volume>20</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          :
          <fpage>36</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tuan</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Incorporating trustiness and collective synonym/contrastive evidence into taxonomy construction</article-title>
          .
          <source>In: 2015 Conf on Empirical Methods in Natural Language Processing, EMNLP 2015</source>
          , pp.
          <fpage>1013</fpage>
          -
          <lpage>1022</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Yang</surname>
          </string-name>
          , H.:
          <article-title>Constructing task-specific taxonomies for document collection browsing</article-title>
          .
          <source>In: 2012 Joint Conf on Empirical Methods in Natural Language Processing and Computational Natural Language Learning</source>
          , pp.
          <fpage>1278</fpage>
          -
          <lpage>1289</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ming</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chua</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Topic hierarchy construction for the organization of multi-source user generated contents</article-title>
          .
          <source>In: 36th Int ACM SIGIR Conf on Research and Development in Information Retrieval</source>
          , pp.
          <fpage>233</fpage>
          -
          <lpage>242</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Tatarintseva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , Keller,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Matzke</surname>
          </string-name>
          , W.-E.:
          <article-title>Quantifying ontology fitness in OntoElect using saturation- and vote-based metrics</article-title>
          . In: Ermolayev,
          <string-name>
            <surname>V.</surname>
          </string-name>
          , et al. (eds.)
          <source>Revised Selected Papers of ICTERI</source>
          <year>2013</year>
          ,
          <article-title>CCIS</article-title>
          , vol.
          <volume>412</volume>
          , pp.
          <fpage>136</fpage>
          -
          <lpage>162</lpage>
          (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -03998-
          <issue>5</issue>
          _
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kosa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaves-Fraga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keberle</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birukou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Similar terms grouping yields faster terminological saturation</article-title>
          . In: Ermolayev,
          <string-name>
            <surname>V.</surname>
          </string-name>
          et al. (eds.)
          <article-title>ICTERI 2018</article-title>
          .
          <article-title>Revised Selected Papers</article-title>
          .
          <source>CCIS</source>
          , vol.
          <volume>1007</volume>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>70</lpage>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -13929-
          <issue>2</issue>
          _
          <fpage>3</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velardi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Learning domain ontologies from document warehouses and dedicated web sites</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>30</volume>
          (
          <issue>2</issue>
          ),
          <fpage>151</fpage>
          -
          <lpage>179</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. de Knijff, J.,
          <string-name>
            <surname>Frasincar</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogenboom</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Domain taxonomy learning from text: the subsumption method versus hierarchical clustering</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          <volume>83</volume>
          ,
          <fpage>54</fpage>
          -
          <lpage>69</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Alfarone</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
          </string-name>
          , J.:
          <article-title>Unsupervised learning of an IS-A taxonomy from a limited domainspecific corpus</article-title>
          .
          <source>In: 24th Int Joint Conf on Artificial Intelligence</source>
          , pp.
          <fpage>1434</fpage>
          -
          <lpage>1441</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C. (ed.)
          <article-title>: WordNet. An Electronic Lexical Database</article-title>
          . MIT Press, Cambridge, MA (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: 27th Annual Conf on Neural Information Processing Systems</source>
          , pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Glove: global vectors for word representation</article-title>
          .
          <source>In: 2014 Conf on Empirical Methods in Natural Language Processing, EMNLP 2014</source>
          , pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Iacobacci</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pilehvar</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.:
          <article-title>Sensembed: learning sense embeddings for word and relational similarity</article-title>
          .
          <source>In: 53d Annual Meeting of the Association for Computational Linguistics and 7th Int Joint Conf on Natural Language Processing of the Asian Federation of Natural Language Processing</source>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>105</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Yamane</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takatani</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yamada</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miwa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sasaki</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Distributional hypernym generation by jointly learning clusters and projections</article-title>
          .
          <source>In: 26th Int Conf on Computational Linguistics</source>
          , pp.
          <fpage>1871</fpage>
          -
          <lpage>1879</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Kosa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chugunenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuschenko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badenes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birukou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Semantic saturation in retrospective text document collections</article-title>
          . In: Mallet,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Zholtkevych</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . (eds.)
          <article-title>ICTERI 2017 PhD Symposium</article-title>
          , CEUR-WS, vol.
          <year>1851</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          (
          <year>2017</year>
          ) online
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Kosa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaves-Fraga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumenko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuschenko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moiseyenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dobrovolskyi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasileyko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badenes-Olmedo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Birukou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The influence of the order of adding documents to datasets on terminological saturation</article-title>
          .
          <source>Technical Report TS-RTDC-TR-2018-2-v2</source>
          , Dept. of Computer Science, Zaporizhzhia National University, Ukraine (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batsakis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keberle</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tatarintseva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antoniou</surname>
          </string-name>
          , G.:
          <article-title>Ontologies of time: review and trends</article-title>
          .
          <source>Int J of Computer Science and Applications</source>
          <volume>11</volume>
          (
          <issue>3</issue>
          ),
          <fpage>57</fpage>
          -
          <lpage>115</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Bordea</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faralli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.: SemEval-2015 task 17:
          <article-title>taxonomy extraction evaluation (TExEval)</article-title>
          .
          <source>In: 9th Int. W-shop on Semantic Evaluation, SemEval</source>
          <year>2015</year>
          , pp.
          <fpage>902</fpage>
          -
          <lpage>910</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schütze</surname>
          </string-name>
          , H.: Introduction to Information Retrieval, Cambridge University Press (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Dodig-Crnkovic</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Scientific methods in Computer Science</article-title>
          . In: Conf. for the Promotion of Research in IT at New Universities and at University Colleges in Sweden (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Chowdhury</surname>
            ,
            <given-names>F. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farrell</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>An efficient approach for super and nested term indexing and retrieval</article-title>
          .
          <source>arXiv preprint arXiv:1905.09761v1 [cs.DS]</source>
          , 23 May (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>