<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Revieval of Subject Analysis: A Knowledge-based Approach facilitating Semantic Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sebastian Furth</string-name>
          <email>sebastian.furth@denkbares.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Volker Belli</string-name>
          <email>volker.belli@denkbares.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joachim Baumeister</string-name>
          <email>joachim.baumeister@denkbares.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Würzburg, Institute of Computer Science</institution>
          ,
          <addr-line>Am Hubland, 97074 Würzburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>denkbares GmbH</institution>
          ,
          <addr-line>Friedrich-Bergius-Ring 15, 97076 Würzburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Semantic Search emerged as the new system paradigm in enterprise information systems. However, usually only small amounts of textual enterprise data is semantically prepared for such systems. The manual semantification of these resources typically is a time-consuming process. The automatic semantification requires deep knowledge in Natural Language Processing. Therefore, in this paper we present a novel approach that makes the underlying Subject Indexing task rather a Knowledge Engineering than a Natural Language Processing task. The approach is based on a simple but powerful and intuitive probabilistic model that allows for the easy integration of expert knowledge.</p>
      </abstract>
      <kwd-group>
        <kwd>Subject Indexing</kwd>
        <kwd>Document Classification</kwd>
        <kwd>Semantic Search</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Historically, Subject Analysis and Subject Indexing [
        <xref ref-type="bibr" rid="ref1 ref10">10,1</xref>
        ] had been a rather manual
task, where librarians or catalogers tried to index large corpora of documents
according to a given set of controlled subjects. A more technical but prominent
example for large scale Subject Indexing is the web catalog from the early Yahoo
times, where websites had been indexed with certain topics. Regardless of which
medium was used, catalogers typically tried to determine the overall content
of a work in order to identify key terms/concepts that summarize the primary
subject(s) of the work. An indexing step enabled in-depth access to parts of the
work (chapters, articles, etc.). Therefore, the item was conceptually analyzed
(what is it about?) and subsequently tagged and cataloged with subjects from a
controlled vocabulary.
      </p>
      <p>
        Nowadays, Semantic Search [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] applications belong to the state of the art
in Information Retrieval. In contrast to traditional search engines ontologies
are used to connect multi-modal content with semantic concepts, which can
then be exploited during the retrieval to improve search results. Therefore,
users of Semantic Search applications typically formulate their search queries as
semantic concepts. Then a retrieval algorithm might expand the query considering
ontological information. Finally, a look-up method maps the concepts to actual
search results using an index from concepts to information resources.
      </p>
      <p>With the growing amount of information manually maintaining catalogs or
indices became almost impossible. However, catalogs and indices are typically
built for a specific problem domain. For many domains formal knowledge in form
of ontologies exists and comprises decent amounts of terminology and relational
information. Thus, we describe a novel approach for automatic Subject Analysis
that allows for the easy integration of formal domain knowledge. We have built an
intuitive probabilistic model that makes Subject Analysis not a Natural Language
Processing but rather a Knowledge Engineering task. Therefore, the approach
allows that domain experts can control the analysis by expressing their knowledge
about relations between concept classes and the importance of certain document
structures.</p>
      <p>The remainder of the paper is structured as follows: Section 2 formally defines
the Subject Analysis problem and discusses related work. In Section 3 we present
our Knowledge-based Subject Analysis approach. Section 4 describes experiences
made with our approach in industrial scenarios. We conclude with a discussion
of our approach in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem Description</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>Controlled Vocabularies</title>
        <p>
          The fundamental requirement for Subject Analysis and the subsequent Subject
Indexing is the existence of a controlled vocabulary. Historically, a controlled
vocabulary defined the way how concepts were expressed, provided access to
preferred terms and contained a term’s relationships to broader, narrower and
related terms. Nowadays, such information is typically modeled by standardized
ontologies [
          <xref ref-type="bibr" rid="ref12 ref13 ref9">9,13,12</xref>
          ], where terms are embedded in complex networks of concepts
covering broad fields of the underlying problem domain. Typical examples are
ontologies powering semantic enterprise information systems. In such systems
users interact using concepts that are company-wide known and valid. An
increasing amount of companies maintain corresponding ontologies as they are
the key element for the interconnection of enterprise systems and data [18]. If
such ontologies do not exist, the construction is usually very reasonable under
cost-benefit considerations, as they support not only semantic information
systems but are also a vehicle for the introduction of more elaborate services like
Semantic Autocompletions [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] or Semantic Assistants. In this paper, we formally
define a controlled vocabulary as follows:
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Definition 1 (Controlled Vocabulary). A controlled vocabulary is an ontol</title>
        <p>ogy O = (T, C, P ) that contains a set of terms T that are connected to a set of
concepts C. Concepts c ∈ C are connected to other concepts using properties
p ∈ P .
User Manual</p>
        <p>Repair</p>
        <p>Manual
Spare Parts</p>
        <p>Document: Repair Manual</p>
        <p>DoDcouIcmnufeomnUetnntit /</p>
        <p>Segment 1</p>
        <p>Information Unit: Segment 1
The necessary components
for the transmission control
such as the gear selector
switch, the electric ..</p>
        <p>
          Token
necessary
Term
gear selector switch
Assuming that such an ontology/controlled vocabulary exists the task is to
examine the subject-rich portions of the item being cataloged to identify key
words and concepts. Therefore existing textual resources must be partitioned to
sets of reasonable Information Units [
          <xref ref-type="bibr" rid="ref3 ref4 ref6">6,3,17,4</xref>
          ] (see Figure 1). Then the task can
be defined as follows:
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Definition 2 (Subject Indexing from information units). For each Infor</title>
        <p>mation Unit i ∈ I find a set of concepts Ci ⊆ C from an ontology O that describe
the topic of the corresponding text best.</p>
        <p>An information unit i has an associated bag of term matches Mi, i.e. a list of
terms from a domain ontology/controlled vocabulary that occur in a particular
information unit. Given the bag of term matches the task can be specialized as
follows:</p>
      </sec>
      <sec id="sec-2-4">
        <title>Definition 3 (Subject Indexing from bags of words). Given a bag of term</title>
        <p>matches Mi determine the underlying topics in the form of a set of concepts
Ci ⊆ C from an Ontology O.</p>
        <p>
          The availability of formalized domain knowledge is usually a valuable support
factor for tasks that cover certain aspects of a problem domain [
          <xref ref-type="bibr" rid="ref14">14,15,16</xref>
          ]. We
claim that this is also true for Subject Indexing where the selection of topics can
profit from formalized background knowledge. Thus, the integration of domain
knowledge in the annotation mechanism becomes a critical success factor and
the task can be further refined as follows:
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Definition 4 (Subject Indexing with background knowledge). Given a</title>
        <p>
          bag of term matches Mi determine the underlying topic in the form of a set of
concepts Ci ⊆ C from an Ontology O considering the domain knowledge contained
in Ontology O.
2.3
Topic Analysis is a relatively wide field of research and is strongly influenced
by Document Classification and Document Clustering approaches. Notable
approaches exist in particular among latent methods, i.e. topics are not expressed
in form of explicit concepts but as a set of key terms. Prominent examples are
Latent Dirichlet Allocation (LDA)[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and Latent Semantic Analysis (LSA)[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
Regarding the deduction of explicit topics Explicit Semantic Analysis [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is a
well-known approach.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Probabilistic Subject Analysis</title>
      <p>In the following, we will first present a basic probabilistic model that is based
mainly on weighted semantic relations between terms and concepts. The model
can be tailored to integrate expert knowledge for a certain domain specific
controlled vocabulary. The basic model will be extended in order to also consider
document characteristics, like important document structures (e.g. headlines) or
formatting information (e.g. bold text).
3.1</p>
      <sec id="sec-3-1">
        <title>Basic Probabilistic Model</title>
        <p>The basic probabilistic model is founded on observable text/term matches,
relations between these terms and potential topics (concepts) and a strong
independence assumption between all features. The model connects the features
as follows:
1. Starting from a text match match ∈ Mi in an information unit i ∈ I the
model derives potentially corresponding terms t ∈ T .
2. The model optionally weights the term t ∈ T with respect to the covering
document structure of the corresponding text match match ∈ Mi, e.g. term
occurrences in headlines might be more important.
3. Given a term t ∈ T the model looks for concepts c ∈ C that can be described
with this term, i.e. which concepts have this term as label and how specific is
this label.
4. The concepts c ∈ C derived from the model on basis of the text/term
match m ∈ Mi might have relations to topic concepts topic ∈ Ci with
Ci ⊆ C. The model exploits ontological information for the derivation of
topic concepts topic ∈ Ci from observed (term) concepts c ∈ C resulting in a
topic probability for a text/term match.
5. The derived topic probabilities for each text match match ∈ Mi get
aggregated in order to compute the overall topic probabilities for an information
unit i ∈ I.</p>
        <p>Given a bag of term matches Mi for an information unit i ∈ I, we realized
steps (1) to (3) by computing the topic probabilities for each text/term match
match ∈ Mi:</p>
        <p>P(topic | match) = α ∗ P (topic | c) ∗ P (c | t) ∗ P (t | match).</p>
        <p>Therefore, we consider the confidence of a term match P (t | match), i.e. the
probability of a certain term t given a textual match match. Additionally we take
the specificity P (c | t) of a term t for a certain concept c into account. Unique
labels have the maximum specificity of 1.0. The relevance of the concept in focus
c for a topic concept topic is P (topic | c). This relevance gets computed on basis
of ontological information between both concepts. The relevance is maximum if
both considered concepts are equal (identity). Finally, we use the constant prior
α to express the linguistic uncertainty that a certain topic is not meant given
a certain term match. This avoids that one perfect term match pretends other
topics to get more important, i.e. it regulates how many related term matches are
necessary in order to outperform one perfectly matching term. We then compute
the topic probabilities for an information unit i ∈ I (step 4) on basis of the topic
probabilities for each term match match ∈ Mi:</p>
        <p>P(topic) = 1 −</p>
        <p>Mi
Y (1 − P (topic | match)).
match
(2)
(3)</p>
        <p>The result is a set of topics topic ∈ Ci with associated probabilities that
express how well a certain topic fits to the terminology observed in the information
unit. This computation assumes independence between the term matches in Mi
according to Bayes’ Theorem. The independence assumption might not perfectly
reflect reality but is a suficient approximation in this application scenario.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Extended Probabilistic Model</title>
        <p>The basic probabilistic model can be extended, such that it also considers
distinctive document characteristics as valuable background knowledge. In many
specialized publications like technical documents or textbooks document
structures indicate the underlying topic or support at least the discrimination of
multiple topic candidates. Typical examples are headlines or formatted text
(italics, bold, underlined).</p>
        <p>The basic probabilistic model uses a constant prior α that expresses the
linguistic uncertainty that a topic is not meant by a certain term match. We
extend the basic model, such that the prior is not constant but depends on the
document structure where the term match was observed. Therefore, document
structures get weighted according to their importance for the deduction of a topic
for an information unit. Assuming that for each document structure a weight w
exists (default 1.0) the value for the prior α is computed as follows:
αadaptive = 1 − (1 − αconstant)w.</p>
        <p>This procedure also allows to discriminate document structures that are
inappropriate for the topic deduction, e.g. references/links to other documents.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Knowledge Representations and Derivation of Probabilities</title>
        <p>The preceding sections introduced a simple but powerful and intuitive probabilistic
model for Subject Analysis. However, the primary target remains that the Subject
Analysis of large document corpora becomes rather a Knowledge Engineering
than a Natural Language Processing task. Therefore, the proposed probabilistic
model allows for the easy adaptation to characteristics of a domain specific
controlled vocabulary and the corresponding corpus of documents that shall be
subject indexed. The following section describe the knowledge-based adaptation,
i.e. the definition of basic conditions for the derivation of probabilities.
Term Confidence P (t | match) The term confidence P (t | match) expresses
how certain a text match is actually a term occurrence. The computed confidence
depends on the quality of the text match. A perfect match, i.e. the text match
match is equal to the term t results in the maximum confidence of 1.0. The usage
of fuzzy string matching techniques like order independent matching, stemming
etc. might lower the confidence of term matches. Therefore, implementations
of the presented probabilistic approach should allow for the configuration of
diferent fuzziness levels and adjust the confidence accordingly.</p>
        <p>Term Specificity P (c | t) Given a term the model must derive all concepts
c ∈ O that can be described by this term. The model must also express how
specific a term is for a concept P (c | t), i.e. handle ambiguous terms like “apple”
which can be the name of a company or a fruit. In the context of technical
documents, we might encounter terms like “nut”, “engine” or “screw” that are
very ambiguous and thus unspecific. Therefore, the specificity of a term must
be distributed over all potential concepts. In the simplest case the specificity
can be distributed equally over all concepts. Unambiguous terms always have a
specificity of 1.0. However, experts’ knowledge might be used to prefer certain
concepts. This might be useful if some concepts of an ontology are not applicable,
e.g. because components they represent are not included in certain machines.
Concept Relevance P (topic | c) Then, given a concept the model must be
able to determine how relevant it is for certain topics P (topic | c). The procedure
is always the same and is explained by the example of technical documents.
In technical documents the occurrence/observation of a concept describing a
component might be relevant for a couple of concept topics: (1) machine functions
relying on this component, (2) parent components or (3) the component itself.</p>
        <p>In general, we assume that the relevance of a concept for a topic decreases the
larger the distance between both concepts is in the underlying ontology. However,
experts’ might know that in certain situations (documents) the occurrence of a
concept is much more indicative for specific topics than for others. For example
in operator manuals component terms might also indicate functions while they
typically do not in repair manuals because usually an operator wants to "operate
a function", whereas a technician usually wants to "repair a component".</p>
        <p>For the calculation of the concept relevance distances between concepts and
topic concepts are extracted/queried from the ontology. Expert knowledge can be
used to weight these distances according to the properties p ∈ P involved. This
way background knowledge regarding the relevance of certain concepts under
certain circumstances can be included in the model. Finally the weighted distances
between the concept in focus c and the topic concept topic get transformed to
a probability. We propose the usage of a normalized sigmoid function to avoid
overestimation of the distance. The parameters β and γ can be used to control
the sigmoid function and thus the overall importance of the concept relevance:
P(topic | c) =</p>
        <p>1 + e(−β)∗γ
1 + e(distance−β)∗γ
(4)
Linguistic Uncertainty α In the basic probabilistic model the parameter α is
constant. In the extended model the parameter α can be adjusted, such that it can
prefer or discriminate term occurrences in certain document structures. Therefore,
domain experts can define weights w for certain document structures (default
1.0). Values for w greater than 1.0 prefer, values smaller than 1.0 discriminate
terms in certain structures respectively. During the computation of the value for
the adaptive linguistic uncertainty αadaptive an implementation has to consider
the value accordingly.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Extended Example</title>
      <p>An exhaustive and thorough evaluation of the presented approach is subject to
future work. However, we have already applied the probabilistic model in an
ongoing industrial semantification project with promising results. In the following
we briefly describe the key aspects of the case study.
4.1</p>
      <sec id="sec-4-1">
        <title>The data set</title>
        <p>In the case study the task is to semantify a given corpus of technical
documents provided in PDF format. The corpus comprises several thousand pages of
technical information, spreaded over diferent documents like operator manuals,
functional descriptions or repair and maintenance instructions. The
semantification partitions the PDF files to reasonable segments (information units). Then,
each information unit is subject indexed with respect to an existing ontology.</p>
        <p>The ontology contains information about the hierarchical structure of
components in the corresponding machine as well as functional connections between
components (see Figure 2 for a simplified visualization). Labels are attached to
all concepts.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Parametrization of the Probabilistic Model</title>
        <p>The probabilistic model has been parametrized to incorporate existing domain
knowledge. Therefore, we used the tailoring possibilities described in Section 3.3
as follows:</p>
        <p>Label
subComponentOf
skos:altLabel</p>
        <p>Concept
rdfs:subClassOf</p>
        <p>rdfs:subClassOf
Component
refers</p>
        <p>Function
– Term Confidence P (t | match): We allowed order independent lookups
without decreasing term confidences. Matches that had only been possible
due to stemming have been discriminated.
– Term Specificity P (c | t): We have distributed the specificity equally over
all concepts, i.e. if a term is attached to two concepts, the specificity of the
term is 0.5 for both concepts.
– Concept Relevance P (topic | c): For operating manuals we slightly
preferred the refer property, in descriptive manuals the subComponentOf
property respectively.
– Linguistic Uncertainty αadaptive: We defined weights w greater than 1.0
for headlines and captions, i.e. prefered term matches occurring in the heading
of sections and the descriptions of images.</p>
        <p>A formal evaluation has not yet been performed. However, experts reviewed
the derived topics and confirmed a noticable improvement over a previous
implementation based on Explicit Semantic Analysis.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper we presented a novel approach for automatic Subject Indexing, i.e.
the indexing of information units with respect to a controlled vocabulary/ontology.
The presented approach is based on a simple but powerful and intuitive
probabilistic model. We claim that this approach does not require training data but
facilitates the easy incorporation of experts’ domain knowledge and thus is highly
adaptive. The adaptiveness through experts’ knowledge makes automatic Subject
Indexing rather a Knowledge Engineering than a Natural Language Processing
task. Thus, large scale semantification of enterprise corpora becomes possible.
The approach has not yet been evaluated thoroughly. However, its application in
industrial case studies yielded promising results.</p>
      <p>Besides an exhaustive evaluation future directions include the addition of
learning methods. Therefore, we consider incorporating latent approaches as
preprocessors to adjust concept relevances based on term frequencies in the
underlying corpus. Additionally, we plan to investigate whether simulated annealing
can be used to learn weights w for document structures. We also plan to
consider background knowledge about (hierarchical) connections between document
structures in the model, e.g. the consideration of neighbour or parent segments’
topics.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The work described in this paper is supported by the Bundesministerium für
Wirtschaft und Energie (BMWi) under the grant ZIM ZF4172701 "APOSTL
Accessible Performant Ontology Supported Text Learning".
15. Padma, T., Balasubramanie, P.: Knowledge based decision support system to assist
work-related risk analysis in musculoskeletal disorder. Knowledge-Based Systems
22(1), 72–78 (2009)
16. Puppe, F., Buscher, G., Atzmueller, M., Huettig, M., Buscher, H.P.: Clinical
experiences with a knowledge-based system in sonography (sonoconsult). In: Workshop
on Current Aspects of Knowledge Management in Medicine (KMM05), Proceedings
3rd Conference Professional Knowledge Management - Experiences and Visions,
Kaiserslautern, Germany (2005)
17. Reynar, J.C.: Statistical models for topic segmentation. In: Dale, R., Church,
K.W. (eds.) ACL. Association of Computer Linguistics (1999),
http://dblp.unitrier.de/db/conf/acl/acl1999.html#Reynar99
18. Stephens, S.: The enterprise semantic web. In: The Semantic Web, pp. 17–37.</p>
      <p>Springer (2007)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Albrechtsen</surname>
          </string-name>
          , H.:
          <article-title>Subject analysis and indexing: from automated indexing to domain analysis</article-title>
          .
          <source>Indexer</source>
          <volume>18</volume>
          ,
          <fpage>219</fpage>
          -
          <lpage>219</lpage>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Latent Dirichlet Allocation</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>3</volume>
          ,
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Borkar</surname>
            ,
            <given-names>V.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deshmukh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarawagi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Automatic segmentation of text into structured records</article-title>
          . In: Mehrotra,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sellis</surname>
          </string-name>
          , T.K. (eds.) SIGMOD Conference. pp.
          <fpage>175</fpage>
          -
          <lpage>186</lpage>
          . ACM (
          <year>2001</year>
          ), http://dblp.uni-trier.de/db/conf/sigmod/sigmod2001.html# BorkarDS01
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Choi</surname>
          </string-name>
          , F.Y.Y.:
          <article-title>Advances in domain independent linear text segmentation</article-title>
          .
          <source>In: ANLP</source>
          . pp.
          <fpage>26</fpage>
          -
          <lpage>33</lpage>
          (
          <year>2000</year>
          ), http://dblp.uni-trier.de/db/conf/anlp/anlp2000.html#Choi00
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Deerwester</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furnas</surname>
            ,
            <given-names>G.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harshman</surname>
          </string-name>
          , R.:
          <article-title>Indexing by latent semantic analysis</article-title>
          .
          <source>Journal of the American society for information science 41</source>
          (
          <issue>6</issue>
          ),
          <fpage>391</fpage>
          -
          <lpage>407</lpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Furth</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumeister</surname>
          </string-name>
          , J.:
          <source>Semantification of Large Corpora of Technical Documentation. IGI Global</source>
          (
          <year>2016</year>
          ), http://www.igi
          <article-title>-global.com/book/enterprise-big-dataengineering-analytics/145468</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gabrilovich</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markovitch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Computing semantic relatedness using Wikipediabased explicit semantic analysis</article-title>
          .
          <source>In: Proceedings of the 20th international joint conference on artificial intelligence</source>
          . vol.
          <volume>6</volume>
          , p.
          <volume>12</volume>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Guha</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCool</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Semantic search</article-title>
          .
          <source>In: Proceedings of the 12th international conference on World Wide Web</source>
          . pp.
          <fpage>700</fpage>
          -
          <lpage>709</lpage>
          . ACM (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krötzsch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parsia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel-Schneider</surname>
            ,
            <given-names>P.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rudolph</surname>
          </string-name>
          , S. (eds.)
          <source>: OWL 2 Web Ontology Language: Primer. W3C Recommendation (27 October</source>
          <year>2009</year>
          ), available at http://www.w3.org/TR/owl2-primer/
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hutchins</surname>
            ,
            <given-names>W.J.:</given-names>
          </string-name>
          <article-title>The concept of 'aboutness' in subject indexing</article-title>
          .
          <source>In: Aslib Proceedings</source>
          . vol.
          <volume>30</volume>
          , pp.
          <fpage>172</fpage>
          -
          <lpage>181</lpage>
          .
          <source>MCB UP Ltd</source>
          (
          <year>1978</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hyvönen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mäkelä</surname>
          </string-name>
          , E.:
          <article-title>Semantic autocompletion</article-title>
          .
          <source>In: The Semantic Web-ASWC</source>
          <year>2006</year>
          , pp.
          <fpage>739</fpage>
          -
          <lpage>751</lpage>
          . Springer (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Klyne</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>J.J.: Resource</given-names>
          </string-name>
          <string-name>
            <surname>Description</surname>
          </string-name>
          <article-title>Framework (RDF): Concepts and Abstract Syntax (</article-title>
          <year>Feb 2004</year>
          ), http://www.w3.org/TR/rdf-concepts/
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Miles</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bechhofer</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>SKOS Simple Knowledge Organization System Reference</article-title>
          .
          <source>W3C Recommendation 18 August</source>
          <year>2009</year>
          . (
          <year>2009</year>
          ), http://www.w3. org/TR/2009/REC-skos-reference-
          <volume>20090818</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Milne</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nicol</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trave-Massuyès</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quevedo</surname>
            ,
            <given-names>J.: TIGER</given-names>
          </string-name>
          :
          <article-title>Knowledge based gas turbine condition monitoring</article-title>
          .
          <source>AI Communications</source>
          <volume>9</volume>
          (
          <issue>3</issue>
          ),
          <fpage>92</fpage>
          -
          <lpage>108</lpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>