<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Dataset for Natural Language Requirements Processing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessio Ferrari</string-name>
          <email>alessio.ferrari@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgio O. Spagnolo</string-name>
          <email>spagnolo@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefania Gnesi</string-name>
          <email>stefania.gnesi@isti.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Information Science and Technologies (ISTI) of the Italian National Research Council (CNR)</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>[Context and motivation] The current breakthrough of natural language processing (NLP) techniques can provide the requirements engineering (RE) community with powerful tools that can help addressing speci c tasks of natural language (NL) requirements analysis, such as traceability, ambiguity detection and requirements classi cation, to name a few. [Question/problem] However, modern NLP techniques are mainly statistical, and need large NL requirements datasets, to support appropriate training, test and validation of the techniques. The RE community has experimented with NLP since long time, but datasets were often proprietary, or limited to few software projects for which requirements were publicly available. Hence, replication of the experiments and generalization have always been an issue. [Principal idea/results] Our near future commitment is to provide a publicly available NL requirements dataset. [Contribution] To this end, we are collecting requirements documents from the Web, and we are representing them in a common XML format. In this paper, we present the current version of the dataset, together with our agenda concerning formatting, extension, and annotation of the dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        As well known, requirements are normally expressed with the most human of
the communication codes, which is natural language (NL) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In recent years,
natural language processing (NLP) technologies have seen a rapid growth, and
our ability of addressing common NLP tasks, such as concept categorisation,
synonyms detection, semantic relatedness evaluation, etc., have radically
improved [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Therefore, we would like the capabilities of modern NLP technologies
to be shared also by the requirements engineering (RE) community. However,
recent techniques are mainly machine learning methods, which are statistical
in nature [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], and require large datasets to properly work. Hence, extensive
requirements datasets are needed to e ectively exploit these technologies in RE [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Several works were performed in RE, which used real-world NL requirements
to address speci c tasks. In particular, works were performed on functional and
non-functional requirements categorization [
        <xref ref-type="bibr" rid="ref13 ref2">2, 13</xref>
        ], traceability [
        <xref ref-type="bibr" rid="ref16 ref4 ref9">4, 9, 16</xref>
        ],
detection of equivalent requirements [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], ambiguity detection [
        <xref ref-type="bibr" rid="ref10 ref17 ref18 ref8">8, 10, 17, 18</xref>
        ] and model
Copyright 2017 for this paper by its authors. Copying permitted for private
and academic purposes.
synthesis [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Notwithstanding the value of these works, the majority of them
share one or both of these weaknesses: (a) experiments are hard to reproduce;
(b) results cannot be considered general. To our understanding, the only dataset
that was used by more than one work (e.g., by Gervasi and Zowghi [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and by
Sultanov and Hayes [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]) is the NASA CM-1 dataset, which concerns a
scienti c instrument to be carried on board a satellite. However, since the dataset
is focused on a speci c project, one needs to experiment with other datasets to
expect generality from the results. Tjong and Berry [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] use a dataset of seven
publicly available industrial requirements documents. Formats vary from .pdf to
.doc, and, to replicate the experiments, one need to have some uniform format
{ e.g., plain text, XML { to be sure that pre-processing of the les did not alter
the original data. Some e orts on the direction of having a common dataset for
NLP in RE are ongoing within the Trace-Lab project [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this case, the focus
is on the speci c task of requirements traceability. Overall, to our knowledge,
a large dataset of NL requirements documents from di erent sources, di erent
domains, for di erent tasks, and in a uniform format is not available yet.
      </p>
      <p>To address this benchmark gap, and at the same time be able to exploit the
capability of modern NLP techniques, our commitment is to de ne a publicly
available NL requirements dataset. To this end, we have currently retrieved 79
requirements documents from the Web. The documents cover multiple domains,
have di erent degrees of abstraction, and range from product standards, to
international project deliverables, to university projects. We also de ned a general
XML schema le (XSD) to represent these di erent documents in a uniform
format. Our short-term goal will be mapping the original requirements documents
to this common format, and share the resulting XML les. Our long-term goal,
which requires the contribution of the RE community, will be annotating the
dataset for the di erent tasks that are relevant in RE.</p>
      <p>The remainder of the paper is structured as follows. In Sect. 2, we list the
requirements documents that we retrieved from the Web. In Sect. 3, we discuss
our agenda, and the speci c challenges that we expect to have to face in our
current e ort for the RE community. Finally, Sect. 4 provides nal remarks.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Publicly Available Requirements Documents</title>
      <p>The rst stage of our work concerned the identi cation of publicly available
requirements documents from the Web. To this end, we queried Google with
the OR-linked keywords Requirements Documents, Requirements Speci cation,
System Speci cation, Software Speci cation, SRS, and we selected those links
that pointed to requirements documents. Our search led to the identi cation of
79 documents. The whole dataset can be downloaded from our Web-site 1. We
inspected each document, and labelled it according to the following main elds,
plus additional ones, which provide some rst-stage qualitative information.
{ Doc Name: an alphanumerical ID that identi es the document.
1 http://fmt.isti.cnr.it/nlreqdataset/
{ Pages: a number indicating the number of pages of the document.
{ Level: a letter indicating the degree of abstraction of the requirements. Can
be H = high-level requirements, or L = low-level requirements. The
judgment was subjectively given according to the following rationale. If further
re nement of the document was required before the system could be
implemented, we labelled the document with H. If the content of the document
was ready for implementation, we labelled it with L.
{ Structure: a letter, or combinations of letters, indicating how requirements
are structurally expressed. Can be: S = structured : if the requirements are
expressed in a structured format, as, e.g., use-cases; U = unstructured : if
requirements are expressed as unstructured NL descriptions; O = one
statement : if each requirement is expressed in a single NL statement. If mixed
ways of expressing requirements were used { e.g., if in the same document,
we found both structured requirements (S) and unstructured ones (U) {, we
combined the letters with the + operator (i.e., S + U).
{ Source: a letter indicating if the source of the requirements is a University
(U), or an Public/Private Organization (I). Documents tagged with U
normally include case studies, or excercises. Documents tagged with I include
industrial strength requirements.</p>
      <p>A complete table that summarizes all the requirements documents, and all
the elds, is available from our Web-site. Here, we show some statistics on the
di erent elds. These statistics are not meant to be a formal evaluation of the
generality and balance of the dataset, but are oriented to give a avor of what
can be found in our repository at this stage of our project.</p>
      <p>Pages (Fig. 1a) We have a maximum of 288 pages, a minimum of 7 pages, an
average of 47 pages, with a quite high standard deviation of 45 pages. This indicates
a strong variability of the dataset in terms of length. More accurate indicators
of the documents' length (e.g., number of requirements) will be provided when
all the documents will be formatted in XML.
300
250
200
150
100
50
0</p>
      <p>U + O
O + S 1%
3%
U + O
38%</p>
      <p>U +3O%+ S
U + S
11%
14U%</p>
      <p>15S%
15O%
(a) Pages</p>
      <p>(b) Structure</p>
      <p>Structure (Fig. 1b) The majority of the documents include a combination of
unstructured content and requirements expressed in one sentence (U + O, 38%).
Document with uniform formats { i.e., U, S, or O { are equally distributed, with
about 15% of the documents for each class. Less represented are the other
composite classes. However, the dataset appears already quite general and balanced
for what concerns the structure of the requirements.</p>
      <p>
        Level and Source Concerning the level of the requirements { not shown in the
pictures { we have a dominance of high-level (H) requirements, with 71% of
the documents classi ed with H, and 29% of them classi ed with L. Overall,
more low-level requirements documents shall be added to the dataset to increase
the balance. Concerning the source of the requirements, we have 62% of the
documents coming from Public/Private Organizations (I), and 38% from
Universities (U). Additional industrial requirements are needed, since each company
has its speci c jargon [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and, although the dataset includes documents from
companies, it does not cover all the potential writing styles.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Agenda and Challenges</title>
      <p>
        The work that we are sharing in this paper is at its early stages. Here, we discuss
our agenda, and the related challenges that we expect to face in the near future.
1. Extracting the Text A rst step towards a dataset in a uniform XML
format2 is the extraction of the text from the documents. Tools for text
extraction from .doc, .pdf and other formats are available3. However, the
extraction is never fully clean, and some manual post-processing is required,
to extract XML meta-data from the text (e.g., requirements ID), and to deal
with other conversion issues. Therefore, we expect to combine automated
text extraction techniques with manual work, to have a high-quality dataset
in which only clean and informative text and meta-data are included.
2. Annotating the Dataset Even though we would already have all our
requirements documents in a clean and uniform format, this would not be
su cient. Indeed, for each speci c RE task, manual annotations have to
be provided for the requirements, in order to use the documents as
training, test and validation sets, for supervised machine-learning algorithms, or
as gold standards (i.e., benchmarks) for unsupervised algorithms [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. For
example, if one has to address a classi cation task, in which requirements
are classi ed based on their functional topic, we have to go through each
requirement, and manually associate a topic to it (e.g., train braking, man
machine interface, for a document of the railway domain). In this way, we
can train a supervised classi er on a sub-set of the data, and evaluate its
performance on the remaining sub-set. Similarly, we can train a clustering
2 The generic XSD, together with requirements document examples in XML, is
available through our Web-site.
3 See, for example, textract: https://goo.gl/38NF7Z
algorithm, and check its ability of identifying clusters of topics against the
gold standard of topic classes (i.e., the manually annotated requirements). Of
course, work-around solutions can be found, using contextual information {
as e.g., title of the paragraphs { for the speci c task of functional topic
classi cation. On the other hand, for other tasks, as, e.g., ambiguity, we do not
see other options rather than manually identifying ambiguous requirements.
The annotation work is expected to involve the whole RE community
interested in NLP. Indeed, requirements annotation requires a relevant e ort, and
domain-speci c knowledge is needed to evaluate the requirements. A single
human annotator is often not acceptable, and one have to compare the
annotations of di erent subjects, and compute their inter-annotator agreement
(e.g., through Cohen's kappa [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), as common in NLP. However, we expect
that researchers interested in a speci c task will annotate our dataset for
their task, also de ning shared annotation schemes to be reused by the
community. Concerning the format to be used for annotations, we recommend
to use GATE4, which is already used within the RE community [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
3. Updating and Extending the Dataset From our preliminary analysis,
we have already seen that our dataset is partially unbalanced for what
concerns, e.g., the level of the requirements. Therefore, we expect to surf the
Web to nd additional requirements documents, as well as to include
documents described within the RE literature. On the other hand, we encourage
researchers and companies to share their documents with us, and contribute
to our challenge. For a task such as traceability, the issue of dataset extension
is more tricky. Indeed, to identify requirements traces between requirements
at di erent levels of abstraction, as performed, e.g., in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we need high-level
and low-level requirements belonging to the same project. At this stage, our
dataset includes only few documents belonging to the same project.
4. API De nition To provide researchers with a easy-to-use dataset, we also
need to develop appropriate APIs (Application Program Interfaces) to access
the XML les, and extract both text and meta-data. Luckily, given the XSD
le, technologies like JAXB (Java Architecture for XML Binding)5, can
automatically create Java classes directly from the XSD. In this way, one
can easily access XSD-compliant XML les. Our work is therefore reduced
to the de nition of more high-level APIs, to, e.g., compute statistics from
the dataset, which can work on top of JAXB.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>NLP consists of two fundamental ingredients: algorithms and data. The NLP
community can provide the RE community with advanced algorithms for text
processing. However, we cannot use these powerful tools, unless we take the
burden of providing the data. This paper presents a rst step towards the de nition
of a dataset for natural language requirements processing. To perform the next
4 https://gate.ac.uk
5 http://www.oracle.com/technetwork/articles/javase/index-140168.html
steps, we need the contribution of the whole RE community interested in NLP,
especially for what concerns the annotation of the requirements for speci c tasks,
and the de nition of reusable annotation schemes.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arora</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabetzadeh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Briand</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zimmer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Automated checking of conformance to requirements templates using natural language processing</article-title>
          .
          <source>IEEE TSE</source>
          <volume>41</volume>
          (
          <issue>10</issue>
          ),
          <volume>944</volume>
          {
          <fpage>968</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Casamayor</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Godoy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campo</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Functional grouping of natural language requirements for assistance in architectural software design</article-title>
          .
          <source>KBS</source>
          <volume>30</volume>
          ,
          <issue>78</issue>
          {
          <fpage>86</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cleland-Huang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Czauderna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dekhtyar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gotel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keenan</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leach</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maletic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poshyvanyk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , et al.:
          <article-title>Grand challenges, benchmarks, and tracelab: developing infrastructure for the software traceability research community</article-title>
          .
          <source>In: TEFSE'11</source>
          . pp.
          <volume>17</volume>
          {
          <fpage>23</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cleland-Huang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Czauderna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibiec</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Emenecker</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A machine learning approach for tracing regulatory codes to product speci c requirements</article-title>
          .
          <source>In: ICSE (1)</source>
          . pp.
          <volume>155</volume>
          {
          <fpage>164</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit</article-title>
          .
          <source>Psychological bulletin 70(4)</source>
          ,
          <volume>213</volume>
          (
          <year>1968</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Falessi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cantone</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Canfora</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Empirical principles and an industrial case study in retrieving equivalent requirements via natural language processing techniques</article-title>
          .
          <source>IEEE TSE 39(1)</source>
          ,
          <volume>18</volume>
          {
          <fpage>44</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ferrari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Orletta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esuli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gervasi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gnesi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <source>Natural Language Requirements Processing: a 4D Vision</source>
          . IEEE Software (to appear) (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ferrari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipari</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gnesi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spagnolo</surname>
            ,
            <given-names>G.O.</given-names>
          </string-name>
          :
          <article-title>Pragmatic ambiguity detection in natural language requirements</article-title>
          .
          <source>In: AIRE'14</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gervasi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zowghi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Supporting traceability through a nity mining</article-title>
          .
          <source>In: RE'14</source>
          . pp.
          <volume>143</volume>
          {
          <fpage>152</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gleich</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Creighton</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kof</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Ambiguity detection: Towards a tool explaining ambiguity sources</article-title>
          .
          <source>In: REFSQ'10</source>
          , pp.
          <volume>218</volume>
          {
          <fpage>232</fpage>
          . Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Goth</surname>
          </string-name>
          , G.:
          <article-title>Deep or shallow, NLP is breaking out</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>59</volume>
          (
          <issue>3</issue>
          ),
          <volume>13</volume>
          {
          <fpage>16</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kassab</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neill</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laplante</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>State of practice in requirements engineering: contemporary data</article-title>
          .
          <source>Innovations in Systems and Software Engineering</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <volume>235</volume>
          {
          <fpage>241</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Knauss</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>(semi-) automatic categorization of natural language requirements</article-title>
          .
          <source>In: REFSQ</source>
          , pp.
          <volume>39</volume>
          {
          <fpage>54</fpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          , Schutze, H.:
          <article-title>Foundations of statistical natural language processing</article-title>
          . MIT Press (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Robeer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucassen</surname>
            , G., van der Werf,
            <given-names>J.M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalpiaz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brinkkemper</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Automated extraction of conceptual models from user stories via nlp</article-title>
          .
          <source>In: RE'16</source>
          . pp.
          <volume>196</volume>
          {
          <fpage>205</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Sultanov</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          :
          <article-title>Application of reinforcement learning to requirements engineering: requirements tracing</article-title>
          .
          <source>In: RE'13</source>
          . pp.
          <volume>52</volume>
          {
          <fpage>61</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Tjong</surname>
            ,
            <given-names>S.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berry</surname>
            ,
            <given-names>D.M.:</given-names>
          </string-name>
          <article-title>The design of sreea prototype potential ambiguity nder for requirements speci cations and lessons learned</article-title>
          .
          <source>In: REFSQ'13</source>
          , pp.
          <volume>80</volume>
          {
          <fpage>95</fpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Roeck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gervasi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nuseibeh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Analysing anaphoric ambiguity in natural language requirements</article-title>
          .
          <source>REJ</source>
          <volume>16</volume>
          (
          <issue>3</issue>
          ),
          <volume>163</volume>
          {
          <fpage>189</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>