<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Projects and System
Demonstrations
$ rafael@dlsi.ua.es (R. Muñoz); ygutierrez@dlsi.ua.es
(Y. Gutiérrez); montoyo@dlsi.ua.es (A. Montoyo)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>T2Know: An Advance Scientific-Tecnical Text Analysis Platform for Trend and Knowledge Extraction Using NLP Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rafael Muñoz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yoan Gutiérrez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrés Montoyo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Alicante</institution>
          ,
          <addr-line>Spain. Crta. San Vicentte del Raspeig s/n, Alicante</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The project T2Know presents the use of natural language processing technologies for the creation of a semantic platform of scientific documents via knowledge graphs. This knowledge graph will link relevant parts of each document with those of other documents in such a way that trend analysis and recommendations can be achieved. The goals addressed within the scope of this project include entity recognizers development, profile definition and documents linkage through the use of transformers technologies. As a result, the relevant parts of the documents to be extracted are related not only to the title and afiliation of the authors, but also to article topics such as references, which are also considered relevant parts of the scientific article.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Semantics</kwd>
        <kwd>semantic document profile</kwd>
        <kwd>entity recognition</kwd>
        <kwd>language models</kwd>
        <kwd>trasnformers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>volumes of textual data in the form of digital documents
of a scientific-technical nature. This facilitates systematic
Health research organizations have always been an ex- analysis, environment assessment and tje proposition of
ceptional environment for identifying specific needs and plausible future scenarios. This project significantly
supgenerating new ideas that lead to innovative processes, ports the transition from a traditional model of medicine,
products and services that improve health outcomes and known as reactive or curative medicine, in which patients
the sustainability of health systems. Additionally, these go to the doctor and are treated, to a medicine that is
organizations also provide key information on the en- not satisfied with curing, but seeks to prevent, improve
vironment and market trends in their scientific areas of the quality of life, adapt to the individual, predict the
interest. evolution of the disease and put the patient at the center.</p>
      <p>Nevertheless, despite having technological surveil- In this process of change towards a new paradigm,
lance (in many cases) or competitive intelligence systems as in the case of 5P medicine — which stands for more
that enable them to configure customized alerts or the preventive, participatory, personalized, predictive and
option of performing specialized information retrieval population-based medicine—, HLTs, and tools such as
searches, they still lack solutions that support the sys- the one proposed in this project, play a fundamental
tematic analysis of the large volume of information to role owing to the great potential ofered by health data
retrieve the desired information. More specifically, ac- collection and the analysis of large collections of health
cording to industry figures, medical professionals can data.
spend up to 20 percent 1 of their working day conducting A convenient RDI planning, adapted to both the
orinformation searches to support their daily activities. ganization and its environment, will guarantee the
opti</p>
      <p>In this context, in order to provide greater value, both mization and investment in technological developments
to society and to the people who make up the structure that meet society’s demands. Thus, this would ensure
of health research institutes, we propose the implemen- efectively translating these aspects into clinical practice
tation of a platform for the advanced analysis of large through the productive ecosystem. Trend identification
in research that consequently creates new markets will
lfourish the development of companies which respond
to these new business niches. This will translate into a
significant improvement in patients’ quality of life and
in the generation of wealth and well-being in society</p>
      <sec id="sec-1-1">
        <title>1.1. Project objectives</title>
        <p>The main objective of this project is the research and
development of T2KNOW, an advanced text analysis
platform based on Natural Language Processing (NLP)
technologies, is the extraction and representation of
semantic profiles of digital entities and the identification of
research trends from the automatic analysis of
scientifictechnical documents. Starting from this general objective,
the project has the following specific objectives:
• To design and develop a flexible, scalable and
robust technological architecture for the
management and processing of large volumes of
unstructured data (text) as a necessary basis for advanced
analysis.
• To research and develop advanced text
analysis algorithms, using PLN techniques, that allow
knowledge extraction and semantic exploration
of content for the detection of research trends.
• To develop data visualization technologies to
discover and graphically represent the evolution of
research lines, topics and emerging technologies
that allow the identification of research trends.
• To design and execute a pilot test to validate
the technologies developed in a key area such
as healthcare, with the creation of specific corpus
of scientific publications.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. State of the art</title>
      <p>The process of knowledge discovery from natural
language can be seen as a flow composed of several stages:
from the initial text to a final relevant semantic
knowledge representation in the form of an ontology. The
ifrst step consists in the manual or semi-automatic
construction of annotated linguistic resources. This requires
choosing or defining an annotation scheme that is
conducive to the domain of interest. From an annotated
corpus, machine learning algorithms are trained to apply
the same annotation scheme to large volumes of text. 3. Human language technologies
Subsequently, all automatically discovered entities and
relationships are grouped into a semantic graph. At this The T2Know project is a consortium involving several
point, it is possible to perform post-processing tasks to entities such as ISABIAL (Institute of Health and
Biomedeliminate redundancies, combine similar entities, or de- ical Research of Alicante), the company DIFUSION S.L.
tect inconsistencies. Finally, a unified semantic structure and the University of Alicante. In addition, stakholders
is obtained, which can be presented in the ontology for- such as AIMPLAS (Instituto Tecnológico del Plástico) and
mat, where the relevant knowledge that was implicit in ITI (Centro Tecnológico de Investigación, Desarrollo e
the original text is represented. Innovación TIC) were invited to be part of the team, since</p>
      <p>
        To extract relevant knowledge from natural language they will contribute with expertise in areas related to the
text, PLN techniques have been introduced in systems health field such as plastics or technology. Depending on
such as ISODLE [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. the role played by each entity, these appear in diferent
      </p>
      <p>
        The use of natural language features can be used to modules, as shown in Figure 1.
build rule-based systems, such as the proposed OntoLT The focuses on three use cases or application areas
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which extracts concepts and relationships through a such as health, plastics and technology. The focuses
mapping from linguistic classes to ontology classes. An on three use cases or application areas such as health,
alternative approach is the use statistical or probabilis- plastics and technology. For these, taxonomies have been
tic models, exemplified by systems such as LEILA [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or identified, such as the one in the table 1 for health.
Text2Onto [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Another example is KnowItAll [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which
introduces a pointwise mutual information (PMI) metric
to select relevant instances. Once instances of entities
and relationships are extracted from the text, a natural
question is whether more abstract knowledge can be
inferred from these examples. Systems that address this
problem often use unsupervised techniques to attempt to
discover inherent structures. Two relevant examples of
this approach are OntoGain [6] and ASIUM [7], which
attempt to automatically construct a hierarchy of concepts
using clustering techniques. Even though most of the
aforementioned systems generally focus on one iteration
of the extraction process, more recent approaches, such
as NELL [8], attempt to continuously learn from a stream
of web data and increase over time both the quantity and
quality of the knowledge discovered.
      </p>
      <p>As we have seen, in the systems mentioned above, in
order to extract knowledge from textual sources, it is
necessary to contemplate the use and development of
techniques for the semantic representation of knowledge,
its storage and computational processing, and its metrics
for the evaluation of its quality.</p>
      <sec id="sec-2-1">
        <title>3.1. Document profile representation</title>
        <p>This task is responsible for semantically representing
documents based on a series of characteristics previously
extracted from them. This representation is governed
by the scheme defined in the figure ??, which will allow
to characterize the documents and in turn the entities
and elements included in them, so that, with the use of
semantic exploration techniques, it will be possible to
recover not only documents but also their metadata as
digital entities, with their respective profiles.</p>
        <p>For instance, a given document may have a title,
authors, an abstract, diferent topics and entities and also
citations to other documents. In turn, authors may have an
associated email, afiliation to one or several institutions,
which may have an associated country. As can be seen,
there are several characteristics that, when analyzed at a
deeper level of detail, reveal the benefit of representing
everything that can be characterized, since it enriches
the quality of the metadata and thus ofers greater
opportunities and points of view to explore the documents.
Therefore, once documents are semantically represented,
it will be possible to query not only documents, but also
metadata under multiple non-conventional criteria such
as, for example, Countries, institutions or authors most
• Huntington
– Epigenetics
– Biomarker
– Transcription
– Next Generation Sequencing
– Etc.
• Multiple Sclerosis
• Alzheimer
– Myelin
– Trained immunity
– Olygodendrocyte
– Epigenetics
– Ageing
– Inflammation
– Etc.
– Extracellular vesicles
– Exosomes
– Etc.</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. Development of technologies for semantic information extraction from documents</title>
        <sec id="sec-2-2-1">
          <title>Semantic information extraction is a task in which, by relying mainly on a set of PLN sensors, pieces of infor</title>
          <p>semi-supervised way by a human and using semantic
techniques to discern what would be the appropriate link
between profiles, either of documents or of any digital
entity involved (ie. authors, institutions, countries, etc.).
mation with semantic connotations are extracted from
the textual content. These sensors must be able to
identify, classify and extract all the necessary information to
create a knowledge base. In this task, NERC domain and
text classification sensors are developed, allowing the
application of information and terminology extraction
processes. In particular, the following phases or stages
will be followed:
• Terminology extraction. Term extraction based
on statistical and linguistic algorithms.
• Domain entity recognition. Development of
entity recognizers specialized in each type of
concept.
• Information extraction through PLN tools, e.g.,
sensors of attributes defined in the document
proifle such as date, author, language, summary,
topics, entities, among others.</p>
          <p>During this task, processes will be developed to link the
document profiles through semantic relationships, as
shown in the figure 3.3. That is, once the documents
are semantically represented, it is necessary to link them
to each other, and in this way they will serve as a
connection point between other digital entities such as, for
example, authors, subjects, etc. Not only will the
documents serve as a link between digital entities, but also
the common characteristics found between documents
already semantically represented.</p>
          <p>This linking of profile characteristics enables the
inferring and discovering of new information otherwise
harder to identify at first sight. For instance, authors or
institutions that coincide on a particular topic or have
written about a particular technology.</p>
          <p>Profile linking presents the problem of semantic
ambiguity [9] and ontology mapping [10] when it is proposed
to automatically link profiles without running the risk
of making mistakes. This issue can be dealt with in a</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>To be able to identify, at a general level, diferent cate</title>
          <p>gories to which documents may belong, and to be able
to classify within them the diferent types of domain 3.4. Data analytics and trends
entities, it is necessary to reuse and develop diferent
sensors based on natural language processing technologies. In this task, tools will be developed to detect the greatest
Therefore, machine learning models will be developed, amount of statistical information related to the document
adapted and reused for this purpose. profiles by temporal fractions. This will allow the trend</p>
          <p>With these sensors it will be possible to classify docu- identification and topic evolution related to research in
ments according to diferent topics, as well as to identify the sectors involved in the project. For this purpose,
auand categorize diferent types of entities present in the tomatic learning techniques will be resorted to, such as
contents, in order to guarantee an advanced exploration Time Series, as well as the potential ofered by their
visuof the knowledge involved. To make possible the develop- alization. This visualization allows a two-fold result: (1)
ment of these sensors based in part on machine learning, the generation of new knowledge, and (2) the
presentawe will start from the corpus annotated by domain ex- tion of the temporal and statistical evolution of the pieces
perts that will be used to train the PLN models. of information involved in the identified ontology.
The processes involved in the ontology life cycle,
3.3. Profile linking namely, the ontology creation, management, analysis
and reuse, entail workflows formed by several
activities. These have been defined taking into consideration
the main methodologies for the development of
ontology models. Additionally, it is also necessary that these
activities are supported by mechanisms and tools that
allow their eficient development. These mechanisms
are mainly robust visualization techniques and provided
with an interaction that allows the user, through its
capacity, to develop abstraction, conception,
understanding, representation and learning of knowledge. One
of the most important aspects to take into account in
this task is semantic exploration and recommendation.</p>
          <p>Given the existence of a semantic database, it is
necessary to develop mechanisms for the exploration of the
semantic network, supported by SPARQL(https://skos.
um.es/TR/rdf-sparql-query/) or Cypher (https://neo4j.
com/developer/cypher/) queries, of document profiles
and other digital entities. These mechanisms will allow
to retrieve not only documents through metadata filters,
but also to make aggregate queries to discover statistical 13th international conference on World Wide Web,
trends, and to make recommendations of profiles (e.g., ACM, 2004, pp. 100–110.
documents, authors, institutions, topics, named entities, [6] E. Drymonas, K. Zervanou, E. G. Petrakis,
Unsuetc.) through the semantic links that interconnect the pervised ontology acquisition from plain texts: the
network. OntoGain system, in: International Conference on
Application of Natural Language to Information
Systems, Springer, 2010, pp. 277–287.
4. Conclusions [7] D. Faure, T. Poibeau, First experiments of using
semantic knowledge learned by asium for
informaCurrently, the project is in an initial stage focused on the tion extraction task using intex, in: Proceedings of
capturing of both technical and functional requirements. the ECAI workshop on Ontology Learning, 2000.
In addition, the KPIs, of key importance for identification, [8] T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar,
have been identified and the system development and B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M.
Gardevaluation will begin shortly. ner, B. Kisiel, et al., Never-ending learning,
Communications of the ACM 61 (2018) 103–115.</p>
          <p>Acknowledgments [9] Y. Gutiérrez, S. Vázquez, A. Montoyo, Spreading
semantic information by word sense disambiguation,
This project is funded by the Valencian Agency for Knowledge-Based Systems 132 (2017) 47–61.
Innovation through the project INNEST/2022/24, par- [10] Y. Gutierrez, D. Tomas, I. Moreno, Developing an
tially funded by the Generalitat Valenciana (Conselle- ontology schema for enriching and linking digital
ria d’Educació, Investigació, Cultura i Esport) through media assets, Future Generation Computer Systems
the following projects NL4DISMIS: TLHs for an Equal 101 (2019) 381–397.
and Accessible Inclusive Society (CIPROM/2021/021) and
T2Know: Platform for advanced analysis of
scientifictechnical texts to extract trends and knowledge through
NLP techniques.(Innest/2022/24). Moreover, it was
backed by the work of two COST Actions: CA19134
“Distributed Knowledge Graphs” and CA19142 -
“Leading Platform for European Citizens, Industries, Academia,
and Policymakers in Media Accessibility”.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <article-title>Web-based ontology learning with isolde</article-title>
          ,
          <source>in: Proc. of the Workshop on Web Content Mining with Human Language at the International Semantic Web Conference</source>
          ,
          <string-name>
            <surname>Athens</surname>
            <given-names>GA</given-names>
          </string-name>
          , USA, volume
          <volume>11</volume>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sintek</surname>
          </string-name>
          ,
          <source>Ontolt version 1</source>
          .
          <article-title>0: Middleware for ontology extraction from text</article-title>
          ,
          <source>in: Proc. of the Demo Session at the International Semantic Web Conference</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Suchanek</surname>
          </string-name>
          , G. Ifrim, G. Weikum,
          <article-title>Leila: Learning to extract information by linguistic analysis</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Ontology Learning and Population: Bridging the Gap between Text and Knowledge</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>18</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Völker</surname>
          </string-name>
          , text2onto,
          <source>in: International Conference on Application of Natural Language to Information Systems</source>
          , Springer,
          <year>2005</year>
          , pp.
          <fpage>227</fpage>
          -
          <lpage>238</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O.</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cafarella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Downey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kok</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.- M. Popescu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Shaked</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>D. S.</given-names>
          </string-name>
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Yates</surname>
          </string-name>
          ,
          <article-title>Web-scale information extraction in knowitall:(preliminary results)</article-title>
          ,
          <source>in: Proceedings of the</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>