<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Construction of a Corpus for the Evaluation of Textual Case-based Reasoning Architectures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreas Korger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joachim Baumeister</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Angesagt GmbH</institution>
          ,
          <addr-line>Dettelbachergasse 2, D-97070 Wu ̈rzburg</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Wu ̈rzburg</institution>
          ,
          <addr-line>Am Hubland, D-97074 Wu ̈rzburg</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>denkbares GmbH</institution>
          ,
          <addr-line>Friedrich-Bergius-Ring 15, D-97076 Wu ̈rzburg</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Regulatory documents denote an interesting application domain for case-based knowledge management. These documents enumerate situations with conditions, that are often dangerous for human and environment and they give advice, rules, and instructions for prevention or handling. That type of documents is eminent in many domains and provides valuable experience knowledge which makes it a remarkable application and research domain for (textual) case-based reasoning. In this paper, an initial case-based representation of regulatory documents is introduced. We report on the construction of an open corpus of regulatory documents in the domain of nuclear safety regulations.</p>
      </abstract>
      <kwd-group>
        <kwd>Case-based Reasoning</kwd>
        <kwd>Experience Management</kwd>
        <kwd>Knowledge Management</kwd>
        <kwd>Textual Case-based Reasoning</kwd>
        <kwd>Corpus Annotation</kwd>
        <kwd>Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Case-based knowledge management approaches seem promising to handle
regulatory documents. In general, regulatory documents are published by authorities
covering the handling of situations of a specific domain. A document enumerates
noteworthy situations with conditions, that are often dangerous for human and
environment. Based on a detected situation the document gives advice, rules,
and instructions for preventing or handling a particular situation.</p>
      <p>Examples of regulatory documents are compliance documents of large
companies, safety documents for conventions and festivals, and legislative agreements
between parties.</p>
      <p>This type of document is eminent in many areas and provides valuable
experience knowledge. Furthermore, a single document often covers experience
knowledge from different domains thus requiring the collaboration of many domain
experts. For instance, the compliance document of a company will require
prevention and handling rules defined by experts from the law and human resources
departments. A safety document of a festival will cover regulatory situations
defined by experts from the fire department and the police.</p>
      <p>Today, the use of regulatory documents faces two essential challenges:
• Creation: Documents are commonly created manually and collaboratively
by domain experts in a knowledge intensive process. The reuse of existing
knowledge especially found in existing regulatory documents, is mostly not
present.
• Retrieval: The documents are usually available in plain text and therefore
the retrieval of suitable information for a given situation is complex and time
intensive.</p>
      <p>We see regulatory documents as an interesting application and research
domain for (textual) case-based reasoning and we formulate the following
hypotheses:
1. Case-based reasoning can provide a natural representation for covering
experience knowledge.
2. Case-based reasoning can handle incomplete input for the formulation and
retrieval of experience knowledge.
3. Case-based reasoning is very suitable to integrate background knowledge
into the process provided, e.g., by ontologies.</p>
      <p>To tackle these hypotheses, we started to build an open corpus of regulatory
documents that can be used to evaluate interesting research questions:
• Case-oriented representation of regulatory documents
• Textual CBR in the domain of regulatory documents
• Retrieval of regulatory experience knowledge
• Reuse and generation of new experience knowledge
• Case-based review critique of regulatory documents
• (Textual) quality assessment of regulatory documents</p>
      <p>To the knowledge of the authors there exists currently no open corpus of
regulatory documents, that can be used to work on the research questions stated
above.</p>
      <p>
        In this paper, we introduce an initial case-based representation of regulatory
documents and we report on the construction of an open corpus of regulatory
documents in the domain of nuclear safety regulations. We invite volunteers to
join this construction process. Preceding work in the field of knowledge
management and case-based reasoning laying fundamentals was presented by Korger
and Baumeister [
        <xref ref-type="bibr" rid="ref22 ref24">22, 24</xref>
        ].
      </p>
      <p>The paper is organized as follows: In Section 2 we introduce the PIRI
structure representing regulatory experience knowledge as a case-based interpretation
and we sketch how case-based reasoning is used to work with regulatory
documents. The construction of the corpus of nuclear safety regulations is described
in Section 3, where we introduce the intention and the statistics of the corpus.
We show the methods used for the construction of the corpus and we explain
how to obtain the corpus for own research. In Section 4 we discuss a number
of use cases applicable for the introduced corpus. The paper is concluded with
related and future work in Section 5.</p>
      <p>Case-Based Representation of Regulatory Documents
A regulatory document lists a collection of incidents of interest for a given life
or work context. For each incident of interest, the document usually describes
measures for prevention and measures for handling an occurring incident. To
know, which incidents are actually of interest in a certain scenario and which
measures are effectively applicable, is the result of experience collected in the
past. A key to access this tacit knowledge is to identify parts of the document
corresponding to incidents and measures, as well as to find a metric to make them
comparable within a certain context. These considerations will be reflected in
the case structure presented in the following.
2.1</p>
      <p>The Domain of Regulatory Documents
A regulatory document consists of a sequence of text passages describing specific
aspects of the domain. We call interesting passages within a document
information units.</p>
      <p>ONTOLOGY
Domain Knowledge
Incidents</p>
      <p>Measures</p>
      <p>Context</p>
      <p>INCIDENT-ORIENTED-CASE (PIRI-CASES)
= Context+Incident+relevant Measures</p>
      <p>Incident</p>
      <p>Measures
PREVENT M1</p>
      <p>M2 I M5
M3 M6</p>
      <p>Measures</p>
      <p>M4 REACT</p>
      <p>Context</p>
      <p>INFORMATION-UNIT-CASES
= Semantic Concept + Textual Passages</p>
      <p>DOCUMENTS
Tacit Knowledge</p>
      <p>Text</p>
      <p>Definition 1 (Information Unit and Corpus). An information unit u is
defined as a text passage within a document, i.e., u = (d, f, t), where d is a
document URI and f, t are offset information describing the extract of the
document. For a corpus RD, the universal set of all information units is defined as</p>
      <p>For semantic interpretation of the text passage, such information units need
to be annotated by metadata, explicitly describing the content.</p>
      <p>Definition 2 (Ontology). An ontology O = (E, R) contains metadata
relevant for the considered domain. Here, relevant semantic concepts represented as
entities e ∈ E, that are connected by relations r ∈ R.</p>
      <p>Typical examples for entities within a regulatory domain are fire, smoking
prohibition, explosion, and evacuation. Common relations in this context are
partOf and requires. A symbolic depiction of the ontology, the documents, and
the case considerations can be seen in Figure 1.</p>
      <p>When defining regulatory documents in a case-based representation, we
distinguish two design alternatives:
• Information unit cases: A case is represented by a text passage (information
unit) describing an incident or measure. The case also contains annotations
explicitly describing the incident/measure using elements of the ontology. A
case base then contains a collection of information units as cases and
incident/measure annotations are attributes representing the particular cases.
• Incident oriented cases: One distinct case is represented by one incident
together with all measures mentioned in a document for preventing and
handling the incident. The case base then contains a collection of incidents with
measures as attributes representing the particular cases. It is worth noticing,
that incident oriented cases are constructed by aggregating information unit
cases.</p>
      <p>In the following, we describe both representation alternatives in more detail
and we motivate their use by retrieval and reuse examples.
2.2</p>
      <p>
        Information Unit Case Structure
To build a bridge between semantic concepts and free text an information unit
case combines a distinct information unit with a corresponding metadata
annotation. Having in mind that the overall goal is to reuse documents, it is convenient
to consider textual passages as solutions to structured problem descriptions [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
Subsequently, it is assumed that the textual passage is reusable for similar
problems.
      </p>
      <p>Definition 3 (Information Unit Case). An information unit case cu is
defined as follows: cu = (a, u), where u ∈ IU RD is an information unit from corpus
RD and a ⊆ O is describing metadata from an ontology O. We call the case
base of information unit cases CBU = {cu1 , ..., cun } the collection of all cases
cui , ∀i ∈ {1, .., n} that are extracted and annotated from available regulatory
documents.</p>
      <p>A case cu = (pu, su) representing one information unit is defined by a set
of named entities and relations between them as problem description pu and
textual passage contained in a document indicating the fulfillment as solution su
to the problem. These textual passages are referred as fulfilling textual features.
A textual passage may be just one word up to some sentences. Depending on
the use case scenario the description and the solution of the case might switch.
For instance given a textual passage as problem description the metadata is the
solution. An exemplary information unit case is:</p>
      <p>
        cx = (piri:manualFireFighting, ”find a fire extinguisher and put out the fire”)
Similarity of Information Units To define similarity functions for
information unit cases, we exploit the taxonomic interrelation of semantic concepts.
Both incidents and measures can be well classified into a taxonomy, building
the base for the similarity assessment and the adaptation. The similarities are
calculated via the taxonomic order of its elements. Each element of the
hierarchy is assigned with a likelihood symbolizing the similarity of its sub-elements.
The similarity of the leaf elements is set to 1 and to 0 for the root element.
The similarity increases with depth d of the element according to for instance
simd = 1 − 1/2d, [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. With this atomic case structure, basic retrieval and reuse
is possible. Picking up the previous example one could state the question, what
to do, if there is no fire extinguisher in reach. This can be solved e.g. by the
retrieval of measures more special than manual fire fighting, like manual fire
fighting with clothes. A retrieved case would be for instance:
      </p>
      <p>cx = (piri:manualFireFightingWithClothes, ”take off your jacket and use it
to put out the fire by throwing it onto the flame”)
2.3</p>
      <p>Incident-Oriented Case Structure (PIRI)
A reoccurring pattern sharpens the focus on the essence of the regulatory
documents. The pattern selected for this purpose bases on the assumption that a
regulatory document delivers a core message. In terms of safety, the most
important content are the incidents it mentions as well as the measures to prevent
them or to react to their consequences. For a given context and relevant
incident induced by the context the according measures are ordered by importance
and classified into preventive and reactive measures. Other patterns might be
considered for other tasks in an analogous way.</p>
      <p>Contextual
Elements</p>
      <p>C
C
1
2
C
3</p>
      <p>Influence</p>
      <p>Preventive
Measures</p>
      <p>P1
P2
P3</p>
      <p>Incident
0,9</p>
      <p>0,95
0,85</p>
      <p>
        We call the instances of these information patterns PIRI-snippets
(PreventiveIncident-Reactive-Interrelation) [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The presented model reduces the
complexity of the real world for facilitation of assessment. Typically there is a cascade
of measures that are executed in a specific order. For instance in the case of fire,
first evacuate all people, then close the doors and windows. The PIRI-pattern
in its graphical representation can be seen in Figure 2. The corresponding case
structure is given in the following definition:
Definition 4 (The Incident-Oriented PIRI-Case). A PIRI-case is defined
as follows: cP IRI = (PI , I, RI , CI ), where I ∈ CBU is a case describing a specific
incident, PI ⊆ CBU denotes a set of cases describing preventive measures for I,
RI ⊆ CBU denotes a set of cases describing reactive measures for I, CI ⊆ CBU
denotes a set of cases describing the context of the incident I. We call the case
base of PIRI-cases CBP IRI ⊆ P(CBU ), where P(CBU ) is the set of all subsets
from CBU .
      </p>
      <p>A case describing one PIRI-snippet is the subset of the case base CBU
containing only cases of information units that are related to an incident I. In this
case definition the importance weights are neglected. We see, that PIRI-cases are
an aggregated form of information unit cases, and more sophisticated retrieval
and reuse scenarios are possible. An exemplary PIRI-case for the incident fire
looks as follows.</p>
      <p>cx = ([piri:fireAlarmSystem, piri:fireDrill], piri:fireIncident,
[piri:callFireDepartment, piri:manualFireFighting, piri:evacuate],
piri:doc id123 fire fighting in power plants)
(1)
(2)
Similarity of PIRI Cases To retrieve similar cases for instance similar
PIRIsnippets, the case base is searched for similar problem descriptions pi to the
query q1. A query is made up by a set of named entities, relations between them
and textual passages. With an aggregation function a global similarity measure
is composed by weighting the previously described local similarity functions of
information unit cases with the parameters (ωP , ωI , ωR) and summed up as
follows:</p>
      <p>SimPIRI(ck, cl) =
ωP SimP (Pk, Pl) + ωI SimI (ik, il) + ωRSimR(Rk, Rl)
ωP + ωI + ωR
If we want to compare two PIRI-snippets, then it is desirable to consider the
context. For this reason we define the following extended similarity measure
under the context C:</p>
      <p>
        SimPIRI+C(ck, cl) =
ω1SimPIRI(ck, cl) + ω2SimC (Contk, Contl)
ω1 + ω2
A similarity measure SimC of two documents for context comparison can be
obtained from the background information provided in the documents as free
text. The available documents most often contain brief background information
and summaries of document objective, scope, and structure. This information
can be exploited for context assessment. These passages can be compared to
approximate the documents similarity using textual similarity measures [
        <xref ref-type="bibr" rid="ref18 ref25">18, 25</xref>
        ].
The specific case-based modeling of the context is in scope of future work.
2.4
      </p>
      <p>
        Corpus of Nuclear Safety Regulations
For the initial construction of the corpus we annotated 143 documents summing
up to about 17.500 pages of nuclear safety regulations as described before. The
International Atomic Energy Agency (IAEA) granted their permission to be the
source for all these documents [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Yet there are many more sources of documents
in the same domain as partially summarized in Table 1.
The corpus described in this work can be rebuilt following some basic steps.
In the future (CRC of this paper), we will provide a ready-to-use download. All
resources are available via GitHub [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Tools to aid in building and working on
the corpus by oneself are explained in the source file or readme files. The
corpus consists of three major parts. Scripts written in Java provide functionality,
ontologies contain the data, a case base provides case-based similarity
assessment, retrieval, and adaptation capacity. The ontology was implemented using
the semantic wiki KnowWE [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. For the case-based implementation we made
use of the framework myCBR [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. For the textual structuring of the regulatory
documents an ontology was implemented [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Core components make use of the
SKOS ontology (Simple Knowledge Organization System) [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] and the PROV
ontology [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] as upper ontologies.
      </p>
      <p>
        To access documents without annotation the present annotation
information can be used as seeds for a case-based bootstrapping strategy to mine new
cases [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. This semi-supervised approach for entity and relation extraction
supports the user in an active learning scenario. The algorithm compares the
unknown corpus B with all n-grams retrieved from the already annotated corpus
A. It uses different similarity measures for n-grams of different sizes. To
compare large text passages a text based similarity measure like tf-idf or sentence
embeddings is used, smaller information units are compared using case-based
similarities. The result is the set of automatically extracted annotations most
similar (to a certain threshold) to manually verified information units existent
in the corpus A. The user reviews the generated annotations, adjusts them if
necessary, and adds them to the case base as manually verified annotations.
      </p>
      <p>else
Algorithm 1: Algorithm for semi-supervised case-based bootstrapping.</p>
      <p>Data: Annotated corpus A, not annotated corpus B
Result: Set of new annotations for corpus B with high similarity
gram max=length of largest annotated n-gram in A;
v=similarity switch (e.g. 10 words);
Query ontological entity labels to the corpus B;
Construct new ontology with retrieved machine annotations;
Get all annotations from the corpus A in gram-size-order;
while i &lt; gram max do
if i &lt; v then
simgram = simcase−based
simgram = simtext−based
Query i-grams to machine annotations of B with simgram;</p>
      <p>Request user review for most similar information units;
3</p>
    </sec>
    <sec id="sec-2">
      <title>Use Cases</title>
      <p>We exemplify the previous approach by use case scenarios contained in the corpus
to verify the hypothesizes stated in the introducing section. In the first part of
this section the semantic retrieval is illustrated. Afterwards, we demonstrate how
textual passages can be adapted using the hierarchical relation of entities. In the
following, we describe the usefulness of the PIRI-approach for the generation of
new document sketches.
3.1</p>
      <p>Semantic Search and Retrieval of Documents
A semantically annotated corpus allows for semantic search of documents fitting
to a given problem description. Therefore, the user query is translated into a
structured query to the case base. For instance, the question “Which measures
prevent a fire?” can be analyzed in a first step retrieving the contained entities
Measure, Fire, and the relation isPreventiveMeasureFor. The most special
relational triple according to their hierarchical interrelation, fulfilling the elements
extracted out of the user query, is &lt;Measure isPreventiveMeasureFor Fire&gt;. All
ontological elements containing entities of the hierarchy, that are equal or more
special, should be respected and thus retrieved. For instance:
• piri:fireWatch piri:isPreventiveMeasureFor piri:fireIncident
• piri:fireBarriers piri:isPreventiveMeasureFor piri:fireSpreadingIncident
• piri:smokeDetector piri:isPreventiveMeasureFor piri:smolderingIncident</p>
      <p>Following this retrieval, the user can research the particular text passages,
that were annotated with these concepts.
3.2</p>
      <p>Adaptation of Textual Passages
Text snippets can be adapted by specification or generalization of entities
according to the ontological classification hierarchy. Additionally the corpus aims
to support the mining of transformation knowledge. For instance, to find a
process how measures suitable for certain scenarios are automatically adaptable to
new and unknown incidents. An example for adaptational capacity using the
hierarchical structuring of incidents and measures bases on the following textual
snippet.</p>
      <p>
        Example 1. “In general, the fire containment approach is preferred, since it
emphasizes passive protection and thus the protection of safety systems does not
depend on the operation of a fixed fire extinguishing system.” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>It contains the entities passive protection, safety systems, and fixed fire
extinguishing system. The entity passive protection has a related entity active
protection. The entity fixed fire extinguishing system has the related entity mobile
fire extinguishing system. Those two are related in an appropriate manner. We
can check the correct adaptation by simply querying for cases, where a mobile
fire fighting system is used for active protection.
3.3</p>
      <p>
        Generation of Document Plots
Usually, documents are created by first sketching a document plot. Here,
graphbased and case-based data management complement to each other. We assume
that a document about maintenance of fire safety systems has to be written,
which also means to create a new case for this scenario [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. In our corpus there
exist two documents, one for fire safety in general and one for maintenance of
power plants. In a first step, the ontological models associated to the documents
are extracted and united using a convenient unification strategy [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The new
model then contains a relevant set of entities and relations with corresponding
information units. In a next step, the PIRI-snippets for the new document are
generated automatically. Measures that target the same incident are
accumulated into one PIRI-snippet. Afterwards, the document plot is presented to the
user. The user now adapts and revises the generated plot. If needed, further
user support is available with a case-based query for similar information units.
For each component of the maintenance requirements special documents may
give deeper advice than available in the two merged documents. We selected an
interesting textual passage for fire safety containing advice for maintenance in
the following example.
      </p>
      <p>
        Example 2. “The inspection, maintenance and testing programme should cover
the following fire protection measures:
—fire barrier closures such as fire doors and fire dampers;
—fire detection and alarm systems, including flammable gas detectors;
—emergency lighting systems;
—water based fire extinguishing systems;
—a water supply system including a water source and distribution pipe;
—gaseous and dry powder fire extinguishing systems;” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
      </p>
      <p>For instance, the maintenance of the water supply system is in the focus of
the fire safety document. The corpus contains a document subjected to the
radioactive contamination of water which provides helpful suggestions for the safe
maintenance of the water supply system. In this manner, noteworthy measures
and incidents can be added to complete the generated document plot.
3.4</p>
      <p>Results
The use cases gave a first outline how case-based reasoning can provide a natural
representation for covering experience knowledge. They showed that especially
the concept of similarity-based retrieval is suitable to handle incomplete input
for the formulation and retrieval of experience knowledge. What exceeded the
capacity of this work was to show how background knowledge could be integrated
into the process in a case-based manner. Nevertheless, it was suggested how
strategies basing on the similarity of free text can be intuitively integrated into
a structured case-based architecture.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>This paper introduced an approach for constructing a corpus exploiting
graphbased, case-based, and textual information in the domain of nuclear safety
regulations. It showed aspects of research work necessary in this field. A proposal for
the case-based structuring of a collection of similar documents and text snippets,
respectively, was made. Finally selected use cases showed how the corpus can be
used practically and will be beneficial for further research work in the domain.
4.1</p>
      <p>
        Related Work
An approach for information retrieval using nuclear safety ontologies was
presented by Ogure et. al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Bouchet and Eichenbaum-Voline [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] presented
some early work on case-based search in experience feedback reports from
nuclear power plants. Ahmad et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] pick out the issue that frequent changes in
daily life induce an increasing amount and heterogeneity of safety-related data.
Grabmair et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] presented first results of a feasibility experiment to annotate
documents on sub-sentence level with the goal of ranked document retrieval in
a certain medical law domain.
      </p>
      <p>
        The idea to use patterns for relation extraction dates back to Hearst [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and
was evolved by many following authors. A work by Krug et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] using
augmented rule-based relation extraction combining active learning with supervised
strategies has inspired the presented way of corpus construction. Fundamentals
for the use of case-based bootstrapping in document indexing was presented early
by Bru¨ninghaus and Ashley [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ]. These approaches are well known but came
into focus again meeting state-of-the-art computational and natural language
processing capacity.
      </p>
      <p>
        In this work a simplified data structure is used to assess the content of an
entire document. A similar problem statement was handled by Caro-Martinez et
al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] to use case-based strategies in an environment of incomplete information
to find explanatory examples in recommender systems. An approach for the
semi-automated proof of correctness of case solving strategies in the domain of
German law was presented by Beck et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This inspired aspects of this work,
such as it showed capacities of a rule-based approach in a similar use case.
For future work the annotation will be extended to more documents of the
domain provided by other authorities and countries. We expect benefits from
this extension, such as the chance to build a multilingual textual case-based
corpus [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Basic applications will be evaluated more thoroughly on a broader
data basis. The goal is to exploit the tacit corpus knowledge to develop more
sophisticated applications. Incorporating more NLP technologies is expected to
further facilitate text generation in this and similar knowledge intensive domains.
Other researches are welcome to use the presented corpus for their work and to
contribute to extensions of it.
      </p>
      <p>
        Acknowledgments
We wish to thank the International Atomic Energy Agency (IAEA) for their
support and the consent to use their publications as a base for this work [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. Autorit´e de Suˆret´e Nucl´eaire: https://www.asn.fr</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. Bundesamt fu¨r Strahlenschutz: https://www.bfs.de</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Canadian</given-names>
            <surname>Nuclear</surname>
          </string-name>
          Safety Commission: https://www.cnsc-ccsn.gc.ca
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. Consejo de Seguridad Nuclear: https://www.csn.es</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>5. International Atomic Energy Agency: https://www.iaea.org</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>6. Ispettorato Nazionale per la Sicurezza Nucleare: https://www.isinucleare.it</mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. U.K. Office for Nuclear Regulation: https://http://www.onr.org.uk</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. U.S. Nuclear Regulatory Commission: https://www.nrc.gov</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <article-title>Fire safety in the operation of nuclear power plants: safety guide</article-title>
          .
          <source>International Atomic Energy Agency</source>
          , Vienna (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kostov</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lalis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kremen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Ontological foundation of hazards and risks in stamp</article-title>
          .
          <source>Semantic Web Journal</source>
          (to appear) (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bach</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Althoff</surname>
          </string-name>
          , K.D.:
          <article-title>Developing case-based reasoning applications using myCBR3</article-title>
          . In: Agudo,
          <string-name>
            <given-names>B.D.</given-names>
            ,
            <surname>Watson</surname>
          </string-name>
          , I. (eds.)
          <source>Case-Based Reasoning Research and Development</source>
          . pp.
          <fpage>17</fpage>
          -
          <lpage>31</lpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Baumeister</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutelshoefer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puppe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>KnowWE: A semantic wiki for knowledge engineering</article-title>
          .
          <source>Applied Intelligence</source>
          <volume>35</volume>
          (
          <issue>3</issue>
          ),
          <fpage>323</fpage>
          -
          <lpage>344</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ifland</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kronbach</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puppe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schenke</surname>
          </string-name>
          , R.:
          <article-title>Semi-automatische Korrektur von juristischen Lo¨sungsskizzen</article-title>
          .
          <source>Zeitschrift fu¨r Didaktik der Rechtswissenschaft</source>
          <volume>3</volume>
          ,
          <fpage>242</fpage>
          -
          <lpage>251</lpage>
          (01
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Bergmann</surname>
          </string-name>
          , R.:
          <source>Experience Management</source>
          . Springer, Berlin, Heidelberg (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Bouchet</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichenbaum-Voline</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Case-based reasoning techniques applied to operation experience feedback in nuclear power plants</article-title>
          . In: Smith,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Faltings</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (eds.)
          <source>Advances in Case-Based Reasoning</source>
          . pp.
          <fpage>497</fpage>
          -
          <lpage>511</lpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Bru¨ninghaus,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Ashley</surname>
          </string-name>
          , K.D.:
          <article-title>Bootstrapping case base development with annotated case summaries</article-title>
          . In: Althoff,
          <string-name>
            <given-names>K.D.</given-names>
            ,
            <surname>Bergmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Branting</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (eds.)
          <source>CaseBased Reasoning Research and Development</source>
          . pp.
          <fpage>59</fpage>
          -
          <lpage>73</lpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. Bru¨ninghaus,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Ashley</surname>
          </string-name>
          , K.D.:
          <article-title>Reasoning with textual cases</article-title>
          . In: Mun˜oz- A´vila, H.,
          <string-name>
            <surname>Ricci</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          . (eds.)
          <source>Case-Based Reasoning Research and Development</source>
          . pp.
          <fpage>137</fpage>
          -
          <lpage>151</lpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Caro-Martinez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Recio-Garcia</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez-Diaz</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>An algorithm independent case-based explanation approach for recommender systems using interaction graphs</article-title>
          . In: Bach,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Marling</surname>
          </string-name>
          , C. (eds.)
          <source>Case-Based Reasoning Research and Development</source>
          . pp.
          <fpage>17</fpage>
          -
          <lpage>32</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Grabmair</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashley</surname>
            ,
            <given-names>K.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sureshkumar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nyberg</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>V.R.</given-names>
          </string-name>
          :
          <article-title>Introducing luima: An experiment in legal conceptual retrieval of vaccine injury decisions using a uima type system and tools</article-title>
          .
          <source>In: Proceedings of the 15th International Conference on Artificial Intelligence and Law</source>
          . p.
          <fpage>69</fpage>
          -
          <lpage>78</lpage>
          . ICAIL '
          <volume>15</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Automatic acquisition of hyponyms from large text corpora</article-title>
          .
          <source>In: COLING 1992 Volume 2: The 15th International Conference on Computational Linguistics</source>
          . pp.
          <fpage>539</fpage>
          -
          <lpage>545</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Korger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumeister</surname>
          </string-name>
          , J.: GitHub: https://github.com/regdoc/piri
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Korger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumeister</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The SECCO ontology for the retrieval and generation of security concepts</article-title>
          .
          <source>In: Cox</source>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Funk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Begum</surname>
          </string-name>
          , S. (eds.)
          <source>ICCBR. Lecture Notes in Computer Science</source>
          , vol.
          <volume>11156</volume>
          , pp.
          <fpage>186</fpage>
          -
          <lpage>201</lpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Korger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumeister</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Case-based retrieval and adaptation of regulatory documents and their context</article-title>
          .
          <source>In: LWDA Berlin</source>
          . pp.
          <fpage>292</fpage>
          -
          <lpage>303</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Korger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumeister</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Case-based generation of regulatory documents and their semantic relatedness</article-title>
          . In: Arei,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Kapoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Bhatia</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <source>Future of Information and Communication Conference San Francisco. Advances in Information and Communication</source>
          , vol.
          <volume>1130</volume>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>110</lpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Krug</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reger</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jannidis</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weimer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Madara´sz, N.,
          <string-name>
            <surname>Puppe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Overcoming data sparsity for relation detection in german novels</article-title>
          .
          <source>In: DHd 2017 Digital Humanities: multimedial &amp; multimodal</source>
          , Montreal. pp.
          <fpage>490</fpage>
          -
          <lpage>493</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Provenance: An Introduction to PROV</article-title>
          .
          <source>Synthesis Lectures on the Semantic Web: Theory and Technology</source>
          , Morgan and Claypool (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Ogure</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furuta</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Information retrieval using ontology for sharing knowledge on safety</article-title>
          . In: Spitzer,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Schmocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            ,
            <surname>Dang</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.N.</surname>
          </string-name>
          <article-title>(eds.) Probabilistic Safety Assessment and Management</article-title>
          . pp.
          <fpage>531</fpage>
          -
          <lpage>536</lpage>
          . Springer London, London (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <article-title>W3C: SKOS Simple Knowledge Organization System Reference</article-title>
          : http://www.w3.org/TR/skos-reference (
          <year>August 2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>