<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Isaac Alpizar-Chacon[</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Interlingua: Linking Textbooks Across Di erent Languages</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Utrecht University</institution>
          ,
          <addr-line>Utrecht</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>0000</year>
      </pub-date>
      <volume>0002</volume>
      <abstract>
        <p>Increasing numbers of students enrol in formal and informal courses taught in a foreign language. Studying a course from an unfamiliar university/program is di cult enough, but the di culties multiply when the transition to new course requirements is exacerbated by the necessity to learn course material in a foreign language. This paper describes Interlingua a platform where students can study textbooks in a foreign language supported by on-demand access to relevant reading material in their mother tongue. Interlingua automatically recognises important terminology within textbooks content, extracts structural models of textbooks and links sections and subsections across textbooks in di erent languages covering the same academic subject. The interface and architecture of Interlingua as well as the technologies underlying the platform are described.</p>
      </abstract>
      <kwd-group>
        <kwd>Linking textbooks Modelling textbooks Terminology extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        "Inter-lingual" students
Two parallel trends exist in the current EU education, independent one from
another, originating from di erent conditions, yet leading to a shared outcome.
From the socio-economic perspective, EU promotes ever-increasing mobility,
especially when it comes to younger population. The Bologna process [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the Youth
on Move initiative [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and students exchange programs like Erasmus contribute
to the vision of a joint European education ecosystem, where students from all
EU countries freely and actively engage in educational programs and individual
courses across borders and cultures.
      </p>
      <p>From the pedagogical (and technological) perspective, new forms of learning
have emerged supported by information and communication technologies. They
facilitate free and easy access to learning materials and promote more central
and active role of a learner. Initiatives like Open Educational Resources (OER)
and phenomena like Massive Open Online Course (MOOC) shape the new
educational reality where students have more choice and exibility in terms of which
textbook to read, which course to take, and which skill to acquire. As a result
they become less dependent on the actual institution issuing a degree.</p>
      <p>These two trends reinforce each other and jointly contribute to a
muchdesired outcome of more scalable, sustainable and a ordable education.
However, they also lead to a potential problematic situation that is occurring more
often as more international students enrol in formal university courses and/or
free MOOCs taught in a foreign language. Studying a course from an
unfamiliar university/program is challenging enough. It might be taught on a new
(more abstract or intensive) level. It might require a student to have prerequisite
knowledge and skills that s/he has not acquired yet. Yet for foreign students,
this transition to new course requirements is aggravated by the necessity to learn
material in a foreign language that they did not use when taking the prerequisite
courses. A foreign student inevitably faces a certain language barrier ampli ed
by the mismatch in the background knowledge and terminology. Unfortunately,
the current tradition of resolving these di culties is hardly e cient - teachers
report that non-native students are allowed to bring dictionaries to regular classes
and exams. As a result, we have an educational system that promotes student
mobility on the level of policies, but does not su ciently support it in on the
individual level.</p>
      <p>An e ective remedy to this problem is the provision of international students
with multilingual access to instructional material, where educational resources
in a language of a course are accompanied by resources in their native
language. There are two principle ways to achieve this: translation of the original
resource and linking between two corresponding "inter-lingual" resources. The
translation-based approach will not help solve the problem: manual translation
is not scalable, and machine-based translation is not yet capable of producing
results of adequate quality in a narrow academic domain. In this paper, we present
a solution based on automated semantic linking of related educational resources
across languages with the main focus on textbooks1. Probability theory and
statistics has been chosen as the target domain. One reason for it is the
popularity of this subject in teaching programs of many technical and social science
degrees. Another reason is that this subject uses a lot of speci c terminology
both new and borrowed from prerequisite parts of mathematics.
1.2</p>
      <sec id="sec-2-1">
        <title>Textbooks modelling and linking</title>
        <p>Textbooks are often considered non-structured information resources for the
purpose of text analysis and information extraction. Yet, textbooks are created,
structured and formatted by their authors (who are presumably domain experts)
with a primary purpose to explain the knowledge in the domain to a novice. A
textbook author uses his/her own understanding of the domain when structuring
and formatting the content of the textbook to facilitate this explanation. As a
result, the formatting and structural elements of a textbook (headers, table of</p>
        <sec id="sec-2-1-1">
          <title>1 This work was supported by the INTERREG-IVA-GR program (grant 138 GR</title>
          <p>
            DeLux 32274)
content (ToC), index) not only shape the organisation of the textbook, but also
re ect the organisation of the domain as the author sees it. When extracted they
can be formally represented as a semantic model of the textbook itself and the
domain it teaches. We have experimented with automated extraction of such
models from multiple textbooks in di erent domains (probability theory and
statistics and information retrieval) and di erent representation formats (PDF
and ePUB) [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ].
          </p>
          <p>
            A textbook model can support a variety of tasks including inference and
reasoning about knowledge taught bu the textbook, tracking and modelling
of reader's progress with the textbook, enhanced interaction through
semantic search and adaptive navigation through the textbook. In the case of multiple
textbooks in the same domain, such models can be mapped between each other
providing a semantic bridge for linking related sections and fragments across
textbooks. An external reference model can be used to facilitate the mapping
and help resolving possible terminology con icts. In [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], we report the results of
several experiments on mapping automatically extracted textbook models
between each other and to DBPedia [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]. An alternative to DBPedia can be another
global knowledge graph available in the Open Linked Data Cloud, such as
Freebase [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ], Wikidata [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ], YAGO [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ] or any other encyclopedic data set. Another
alternative can be a domain-oriented ontology or a thesaurus. In either case, to
support linking across textbooks in di erent languages, a reference model has to
specify terminology in these languages. In this project, we have used two such
multilingual models: DBPedia (which extracts textual labels and descriptions of
entities in multiple languages from corresponding resources of Wikipedia) and
the multilingual glossary of statistical terms maintained by the International
Statistical Institute (ISI) - we will refer to it as the ISI glossary from now on.
          </p>
          <p>The rest of this paper is structured as follows. Section 2 gives an overview
of related work in the eld of textbook model extraction and linking. Section
3 presents the details of the proposed technology and the architecture of
Interlingua - the Web-based platform that students can use to access textbooks in
a foreign language while getting recommendations of related reading resources
in their mother tongue. Section 4 presents the interface of Interlingua and the
ways it can support students. Section 5 concludes the paper with a discussion
and a description of future work.
2</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Related work</title>
      <p>
        Creation of knowledge models from textbooks has been explored in a limited
way. Larrn~aga et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] used natural language processing techniques, heuristic
reasoning, and ontologies to semi-automatically construct a representation of the
knowledge to be learned from electronic textbooks. Their approach uses the
document outline to create a tree-like internal representation, and extract the main
domain topics and the pedagogical relationships among them. Sosnovsky et al.
experimented with harvesting topic-based models from HTML textbooks based
on the structure of their headings; the resulting models have been
automatically mapped into a reference ontology to facilitate more ne-grained inference
and adaptation [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Wang et al. [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] have extracted concept hierarchies from
textbooks using Wikipedia. Textbooks sections were matched against Wikipedia
articles and, as a results, annotated with corresponding Wikipedia entities. The
hierarchy was reconstructed based on the hierarchy of chapters, sections and
subsections. Olney et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] used a combination of natural language processing
techniques and a manually extracted index to generate concept maps from a
biology textbook.
      </p>
      <p>
        Several projects experimented with semantic linking of relevant textbooks
written in the same language. Guerra et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] used a probabilistic topic
modelling approach to extract topic models from textbook and use them to
link sections and subsections across multiple textbooks. Later, Meng et. al [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
explored di erent content modelling approaches for textbook linking: a
termbased approach (each word is considered as a knowledge component), LDA
(latent topics representing knowledge components), and a concept-based approach
(author-assigned keywords in scienti c publications representing knowledge
components). Also, an ensemble of the three approaches was used. The semantic and
the combined approaches achieved valuable linking performance.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Interlingua technology</title>
      <p>Architecturally, Interlingua consists of two large components. The o ine
component performs the tasks of textbook modelling and linking, while the online
component supports students' interaction with the content of linked textbooks.
This section describes the details of the o ine component; the student interface
is presented in the next section.</p>
      <p>
        The overall process of adding a new textbook into the Interlingua content
repository is depicted by 1. The only action performed manually is the upload
of textbook les. A teacher decides which textbooks and in which languages
should be available for the students of the course. After a textbook le is
submitted, the textbook model is generated: the textbook is divided into sections
and sub-sections, its ToC and index are extracted, each page is identi ed, every
di erently formatted fragment is recognised and provided with a semantic label
if a corresponding rule exists (we keep working on expanding the rule base).
Index plays a special role in this process, as it provides a glossary of manually
selected terms that the author of the textbook deemed meaningful. Interlingua
extracts index terms and pages referenced by the terms and uses them as the
semantic anchors to link pages and sections of the textbook to the concepts of
the reference ontology and through them to other textbooks available in the
content repository. Finally, the self-assessment component uses the information
from DBPedia to generate multiple choice questions related to the index terms
explained by the current section of the textbook. Once processed and stored this
way, the textbook becomes accessible through the student interface. The rest of
the section describes individual phases of this process; however, we refer readers
to our two other papers for a detailed description of textbook modelling [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and
linking [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
The Textbook modelling component is in charge of processing new textbooks
and creating their internal representations. The process of extraction of
semantic models from PDF textbooks consists of four steps: extraction of raw content,
construction of the style library, identi cation of logical elements, and
construction of the model. Each step and its tasks are depicted in Fig. 2. Each of these
four main steps starts after the previous and relies on its results. We focus
mainly on PDF as the most common representation format; however, the overall
approach is also applicable to other formats that are more explicit and coherent
in their structural speci cations than PDF.
      </p>
      <p>Raw content extraction. In this step, the PDF le of the textbook is parsed
to extract the text (the characters), geometrical (X- and Y- coordinates on the
pages) information and formatting (type, size, and features of fonts) styles of the
document. The parsing of the PDF is done with the Apache PDFBox library2.</p>
      <sec id="sec-4-1">
        <title>2 https://pdfbox.apache.org/</title>
        <p>The parser reads the PDF content stream where the elements are sorted from
left to right and top to bottom according to their position on a page. It merges
smaller objects into bigger using a bottom-up approach. First, individual
characters are processed. Then, characters are grouped into words, words into lines,
and lines into pages. At each stage, the extractor compares the proximity of an
object's coordinates with the coordinates' of its preceding neighbour to form a
bigger object. Besides, the content itself, the parser also saves the styling
information if every object, such as font size, family, properties (boldness, colour,
etc.) as well as margins.</p>
        <p>Style library construction. Once all text fragments are extracted and their
styles are identi ed, the library of all the formatting styles used in the textbook
is constructed. This library is used in the next step to recognise the di erent
structural elements constituting the textbook. The main style of the textbook
content is identi ed, as well as styles for important text fragments, headings
of sections and subsections, etc. Geometric features of the text play a role in
this process as well, as they allow to combine text into paragraphs, sections and
chapters, allow to compute spacing and indentations, allow to recognise columns
and captions, etc.</p>
        <p>
          Logical element identi cation In this step, some auxiliary texts such as
headers, footers and page numbers are recognised and separated from the main
content. In addition, important structural elements of the textbook are
identied. The ToC section is recognised and parsed to construct an outline of the
textbook. Each content section and sub-section is identi ed, associated with
its header, its ToC entry, and its page interval. Finally, the textbook index is
processed. This section plays a special role, as it provides the main source of
information for subsequent textbook linking. A good textbook index is not just
a collection of words, but, essentially, a reference model produced by a domain
expert according to a prede ned set of rules. Every publisher guides the process
of index creation stipulating index length and style, suggesting what can be good
and bad candidates for index terms, advising on how to maintain consistency
when creating hierarchical indices, etc. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Each index entry is provided with
one or more links to the pages within a textbook. And these pages do not
simply mention the entries, but provide meaningful references by either introducing
corresponding terms or elaborating them.
        </p>
        <p>Model construction Finally, with all the extracted information a semantic
knowledge model of the textbook is constructed. The model allows to quickly
obtained di erent knowledge views of the textbooks to provide rich
information to the users, for example: hierarchical structure of sections, all index terms
related to one section, all pages where an index term is explained, etc. The
model is represented using SKOS3, which is a W3C standard based on RDF
and OWL. It provides a model for expressing the basic structure and content of
concept schemes such as thesauri, classi cation schemes, subject heading lists,
taxonomies and other similar types of controlled vocabulary.
3.2</p>
        <sec id="sec-4-1-1">
          <title>Textbook enrichment</title>
          <p>This component takes the model of a textbook and links it to two external
reference models: the ISI Glossary and DBpedia. The ISI Glossary4 was created
by The International Statistical Institute. It contains more than 3500
statistical terms combined into synsets and translated into 31 di erent languages. The
glossary has been parsed and represented as a reference ontology. As a result,
Interlingua can unambiguously recognise that, for example, in English term
"zscore" is the same as "standard score". What is even more importantly, we can
also identify all the ways this term is called in all 31 languages of the ISI
glossary. DBpedia is a machine-understandable source of knowledge created based on
automatically-extracted structured information from Wikipedia, which contains
4.58 million resources and is available in 125 languages. It plays two important
roles in Interlingua. First, it provides and alternative domain-independent
multilingual reference model. Although, ISI glossary has a much better granularity
and coverage statistical terms, DBPedia ensures that the overall approach of
Interlingua is applicable to other domains (and other languages) where a
dedicated high-quality multilingual thesauri cannot be obtained. The second role of
DBPedia is to provide additional structural and annotation information, such as</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3 https://www.w3.org/TR/skos-reference/</title>
      </sec>
      <sec id="sec-4-3">
        <title>4 http://isi.cbs.nl/glossary/</title>
        <p>
          dbpedia-owl:abstract for term explanations, dcterms:subject for elds of study
where the term is coming from and dbpedia-owl:wikiPageWikiLink for di erent
terms, which are mentioned in the same context, etc. The enrichment process is
shown in Fig.3. More information about the enrichment algorithm is available
in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
        </p>
        <sec id="sec-4-3-1">
          <title>Linking of the ISI glossary to DBpedia. In the beginning terms of the ISI</title>
          <p>Glossary have been (automatically) linked to its correspondent entities in
DBpedia. The linking algorithm rst constructs a core set of terms from the glossary
for which a matching DBpedia resource was unambiguously discovered. Then, it
recursively uses textual content associated with recognised DBPedia entities as
the context information to facilitate disambiguation of newly discovered
candidate matches. The process repeats itself until no new terms can be match with
su cient certainty. This only needs to be done once, before the enrichment of
individual textbook models.</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>Linking of the textbook model to the ISI Glossary. At this step, each</title>
          <p>index term extracted from the textbook is compared against each term of the
ISI Glossary, and when the similarity is over 90 %, the index term is linked to
the ISI Glossary term.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>Enrichment with semantic information. Finally, for each index term with</title>
          <p>a linked ISI Glossary term that has a DBpedia resource, semantic information
is extracted to enrich the model of the textbook. The following information is
obtained from the linked DBpedia resources: abstracts, links to the
corresponding Wikipedia pages of the resources, DBpedia categories of the resources, and
direct relations among the linked index terms using the links that exist in the
Wikipedia pages of the resources. At the end of this approach, the textbook's
model is updated with the gathered information.
3.3</p>
        </sec>
        <sec id="sec-4-3-4">
          <title>Textbook linking</title>
          <p>The linking among the di erent parts of textbooks is done based on the ISI
glossary as a reference model to build a Vector Space Model (VSM). The VSM
consists of several documents (each section) per textbook in the rows, and each
entry of the ISI glossary as the columns. The weight for each term of each
document is calculated from the strength of the association between the term
and the document. The VSM is improved pruning the terms in the glossary
that are not found in the corpora. After the VSM has been created, it is split
into matrices, one for each language, and a similarity matrix is created for each
language pair by rst multiplying the two appropriate parts of the VSM together
and then normalising the resulting matrix. These matrices contain the similarity
value for any pair of sections of textbooks from di erent languages. Finally, each
section of each textbook is linked to the ve most similar documents in each
target textbook to create the links among the textbooks in di erent languages.
3.4</p>
        </sec>
        <sec id="sec-4-3-5">
          <title>Assessment generation</title>
          <p>The assessment engine generates Multiple Choice Questions (MCQ) for the
learners to examine their understanding of the terminologies related to the
section that they are currently reading. The basic MCQ asks the learners to select
the correct translation of one of the introduced concepts in the current section
into their native language. Each question has only one right answer and several
distractors. To generate the questions the assessment engine uses the ISI
glossary to get the labels in di erent languages for the same concept. Distractors
are chosen using the relations among concepts obtained from the enrichment of
the model to select reasonably di cult distractors. In other words, the semantic
distance between the correct answer and a distarctor in the model should not
be very long and the lengths of both textual labels should be comparable.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4 Interlingua interface</title>
      <p>
        When a student accesses the Interlingua client, the entry page shows a list of
available textbooks in the target language of study. A student can indicate her
language of study and the mother tongue. Currently, Intelingua uses the
following textbooks: [
        <xref ref-type="bibr" rid="ref23 ref8">8, 23</xref>
        ] in English, [
        <xref ref-type="bibr" rid="ref13 ref20">13, 20</xref>
        ] in French, [
        <xref ref-type="bibr" rid="ref11 ref9">9, 11</xref>
        ] in German, [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] in
Spanish, and [
        <xref ref-type="bibr" rid="ref15 ref7">7, 15</xref>
        ] in Dutch.
      </p>
      <p>A student selects one of the available textbooks to open the textbook
navigation page (see Fig. 4). It consists of the outline panel on the left side and the
content panel with multiple tabs on the right side of the window.</p>
      <p>In the outline panel, a student can brows through sections and sub-sections
of the textbook. She can click on a magnifying glass icon annotating each section
title to load the respective section into the content panel. The arrow icon is used
to display a list of related readings of the selected sub-section in the mother
tongue of the student (see Fig. 5).</p>
      <p>The student can browse through a subsection content in a similar way to any
standard PDF viewer application. Index terms are recognised and highlighted in
the content to indicate that additional interaction with them is possible. When a
student clicks in a highlighted word, an action menu appears with three available
options: translate, explain, and assess (see Fig. 6). The translate action opens a
pop-up window that presents the translation of the term into the mother tongue
of the student, and the de nition of the term extracted from DBpedia (if this
index term has been found in DBpEdia). The explain action opens a pop-up
window that gives access to the sections of the textbooks in the mother tongue
of the user in which the term is explained. The assess action generates and
displays in a pop-up window an question about the translation of the term into
the mother tongue of the user.
As mention before, in the navigation page a student also can open related
readings for the current section in her mother tongue (see Fig. 4). A pop-up
window suggests a list of related readings (see Fig. 5). if a student selects any
of the suggested links, a corresponding sub-section will be loaded in a new tab.
The student can switch between the reading tabs that now contain the related
content in two di erent languages.</p>
      <p>Finally, the user can click on the 'Assessment' tab in the content panel to
generate an assessment composed of several MCQ related to the content of the
currently browsed section (the questions are similar to the one at the bottom
pop-up in Fig 6).</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion and future work</title>
      <p>The paper has presented the motivation, the approach and the implementation
of the Interlingua system that provides foreign students with on-demand access
to related textbook material in di erent languages.</p>
      <p>There might be several possible concerns both practical and pragmatic about
the scale and applicability of this approach. One question that needs to be
addressed is the availability of textbooks and the copyrights issues. In our
experience, university libraries can supply enough PDF-based textbooks on a variety of
subjects. From the point of copyright protection, if a system provides enhanced
access to these books but only to the students of the university holding
necessary subscriptions, then publishers do not have a reason to object. In the worst
case scenario, many good-quality textbooks are freely available online nowadays
in open repositories such as Openstax5/Connections6, Open Textbook Library7,
OER-Commons8, etc.</p>
      <p>Another concern is how well the automated textbook modelling technology
will cope with di erent textbook formatting. Naturally, formats can be quite
di erent between di erent textbooks, yet the authors and the publishers make
every e ort to ensure they are consistent within textbooks (otherwise they would
cause unnecessary confusion) and we believe, that the formatting patterns of
textbooks do follow a set of logical rules. When implementing the rule-based
component of the model extraction algorithm we try to capture these patterns.
We have observed that the variability is quite manageable and as we have been
processing more and more textbooks, we update our set of rules progressively
less.</p>
      <p>Finally, we realise that none of the extracted models is guaranteed to provide
high-quality representation of a domain. These models can potentially su er from
several drawbacks: (1) Subjectivity: they can contain terms are only marginally
related to an objective picture of the domain semantics; (2) Coverage: they can
miss important terms if a textbook does not cover (enough) a particular part
of a domain; (3) Granularity: the terms in an index can be too detailed, too
broad, or inconsistently alternate in their granularity; (4) Lack of semantics:
most indices are at lists of terms with no relations between them. To combat
these potential problems, glossaries extracted from individual textbooks can be
integrated between each other and to external models: global as DBPedia, or
domain-focused as ISI glossary. We plan to further investigate these and other
issues. For example, an important direction for future work is to evaluate the
e ectiveness of the system in a user study with real students from a target group.</p>
      <p>From a broader perspective, Interlingua is an interesting example of a
service that can be built on top of linked semantic models extracted from related
textbooks. However, a corpus of semantically linked high-quality educational</p>
      <sec id="sec-6-1">
        <title>5 https://openstax.org</title>
      </sec>
      <sec id="sec-6-2">
        <title>6 https://cnx.org/</title>
      </sec>
      <sec id="sec-6-3">
        <title>7 https://open.umn.edu/opentextbooks</title>
      </sec>
      <sec id="sec-6-4">
        <title>8 https://www.oercommons.org/</title>
        <p>content can be used to implement a range of di erent services including
adaptive navigation through or recommendation of textbook content, enrichment of
textbook content with external (interactive) educational resources, or extraction
of di erent types of learning objects from the textbook themselves.</p>
        <p>Finally, it is interesting to explore thew applicability of the the approach
towards automated modelling and linking of textbook that underlies
Interlingua in other, less formal domains (e.g., medicine) or domains with con icting
viewpoints (e.g. history).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alpizar-Chacon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erensoy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sosnovsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Order out of chaos: Construction of knowledge models from pdf textbooks</article-title>
          .
          <source>In: Proceedings of the 10th International Conference on Knowledge Capture</source>
          (Submitted).
          <source>K-CAP '19</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Alpizar-Chacon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sosnovsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Expanding the web of knowledge: one textbook at a time</article-title>
          .
          <source>In: Proceedings of the 30th ACM Hypertext conference (Submitted)</source>
          .
          <source>HT '19</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ament</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Indexing: a nuts-and-bolts guide for technical writers</article-title>
          .
          <source>William Andrew</source>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ives</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Dbpedia: A nucleus for a web of open data</article-title>
          .
          <source>In: The semantic web</source>
          , pp.
          <volume>722</volume>
          {
          <fpage>735</fpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bollacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paritosh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , J.:
          <article-title>Freebase: a collaboratively created graph database for structuring human knowledge</article-title>
          .
          <source>In: Proceedings of the 2008 ACM SIGMOD international conference on Management of data</source>
          . pp.
          <volume>1247</volume>
          {
          <fpage>1250</fpage>
          .
          <string-name>
            <surname>AcM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Bologna Declaration:
          <article-title>Towards the european higher european area</article-title>
          (
          <year>June 1999</year>
          ), https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=LEGISSUM:
          <fpage>c11088</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Caenepeel</surname>
          </string-name>
          , S., de Groen, P.:
          <article-title>Waarschijnlijkheidsrekening en statistiek</article-title>
          .
          <source>Vrije Universiteit Brussel</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dekking</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kraaikamp</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.H.</given-names>
            ,
            <surname>Meester</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.E.</surname>
          </string-name>
          :
          <article-title>A modern introduction to probability and statistics: understanding why and how</article-title>
          . Springer (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Eckstein</surname>
            ,
            <given-names>P.P.</given-names>
          </string-name>
          : Repetitorium Statistik. Springer Gabler (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>European</surname>
          </string-name>
          <article-title>Commission: Youth on the move - an initiative to unleash the potential of young people to achieve smart, sustainable and inclusive growth in the european union (</article-title>
          <year>2010</year>
          ), http://europa.eu/youthonthemove/docs/communication/youth
          <article-title>-onthe-move EN</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fahrmeir</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          : Statistik: der Weg zur Datenanalyse. Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Guerra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sosnovsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brusilovsky</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>When one textbook is not enough: Linking multiple textbooks using probabilistic topic models</article-title>
          . In:
          <string-name>
            <surname>Hernandez-Leo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ley</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klamma</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harrer</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . (eds.)
          <article-title>Scaling up Learning for Sustained Impact</article-title>
          . pp.
          <volume>125</volume>
          {
          <fpage>138</fpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Jourdain</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Probabilites et statistique (</article-title>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. Larran~aga,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Conde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Calvo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Elorriaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.A.</given-names>
            ,
            <surname>Arruarte</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Automatic generation of the domain module from electronic textbooks: Method and validation</article-title>
          .
          <source>IEEE Trans. on Knowl. and Data Eng</source>
          .
          <volume>26</volume>
          (
          <issue>1</issue>
          ),
          <volume>69</volume>
          {82 (Jan
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Marchant</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Statistiek</surname>
            <given-names>I. Universiteit</given-names>
          </string-name>
          <string-name>
            <surname>Gent</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Meng</surname>
            , R., Han,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brusilovsky</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Knowledgebased content linking for online textbooks</article-title>
          .
          <source>In: 2016 IEEE/WIC/ACM International Conference on Web Intelligence (WI)</source>
          . pp.
          <volume>18</volume>
          {
          <issue>25</issue>
          (Oct
          <year>2016</year>
          ). https://doi.org/10.1109/WI.
          <year>2016</year>
          .0014
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Olney</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cade</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Generating concept map exercises from textbooks</article-title>
          .
          <source>In: Proceedings of the Sixth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          . pp.
          <volume>111</volume>
          {
          <issue>119</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Sosnovsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsiao</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brusilovsky</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Adaptation in the wild: ontology-based personalization of open-corpus learning material</article-title>
          .
          <source>In: European Conference on Technology Enhanced Learning</source>
          . pp.
          <volume>425</volume>
          {
          <fpage>431</fpage>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago: a core of semantic knowledge</article-title>
          .
          <source>In: Proceedings of the 16th international conference on World Wide Web</source>
          . pp.
          <volume>697</volume>
          {
          <fpage>706</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Velenik</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Probabilites et statistique</article-title>
          . Universit de Genve (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Vrandecic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Krotzsch, M.:
          <article-title>Wikidata: a free collaborative knowledge base (</article-title>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Walpole</surname>
          </string-name>
          , R.E.:
          <article-title>Probabilidad y estadistica para ingenieros</article-title>
          .
          <source>Pearson</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Walpole</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>R.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Probability statistics for engineers scientists</article-title>
          .
          <source>Prentice Hall</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pursel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brautigam</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saul</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>Concept hierarchy extraction from textbooks</article-title>
          .
          <source>In: Proceedings of the 2015 ACM Symposium on Document Engineering</source>
          . pp.
          <volume>147</volume>
          {
          <fpage>156</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>