<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How to Deal with Inaccurate Service Requirements? Insights in Our Current Approach and New Ideas</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Frederik S. Baumer</string-name>
          <email>frederik.baeumer@upb.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michaela Geierhos</string-name>
          <email>michaela.geierhos@upb.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Paderborn</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The main idea in On-The-Fly Computing is to automatically compose existing software services according to the wishes of end-users. However, since user requirements are often ambiguous, vague and incomplete, the selection and composition of suitable software services is a challenging task. In this report paper, we present our current approach to improve requirement descriptions before they are used for software composition. This procedure is fully automated, but also has limitations, for example, if necessary information is missing. In addition, and in response to the limitations, we provide insights into our above-mentioned current work that combines the existing optimization approach with a chatbot solution.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Motivation</title>
      <p>Copyright c 2018 by the paper's authors. Copying permitted for private and academic purposes.</p>
      <p>1See https://sfb901.upb.de for more information about OTF Computing.</p>
      <p>Requirements</p>
      <p>Software
End user</p>
      <p>Service
market
Composition
+</p>
      <p>Software services</p>
      <p>Service 1</p>
      <p>...</p>
      <p>Service 2</p>
      <p>Service 3
not always expedient and, above all, not always what the user originally wanted. For this reason, we think that
minimal communication between the compensation system and the user is required, even if this means slowing
down the creation of software compositions. We also believe that this can be realized through chatbot technology.
In this case, the challenge is to keep the conversation with the user so slim and meaningful that it will not be
annoying for end-users. Furthermore, untrained users must have all the information needed for a decision (e.g.,
description of the problem and examples). This is part of our current and future research activity (see Section 4).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Project and Team Overview</title>
      <p>The development of techniques and processes for the automatic on-the- y con guration and provision of
individual IT services is the goal of the Collaborative Research Centre 901 \On-The-Fly Computing" at the
University of Paderborn, Germany. In particular, subproject B1 \Parameterized Service Speci cations" deals
with the e cient processing of di erent types of requirements speci cations, which enable the successful search,
composition and analysis of services. In the very rst phase (until June, 2015), the focus of the project was on
developing a speci cation language for services in order to automatically and e ciently process them. In the
second phase (until June, 2019), user-friendly speci cations are introduced. These are intuitively understandable
speci cations and are used primarily by end-users and domain experts. However, these requirements descriptions
have errors that need to be corrected using Natural Language Processing (NLP). This is the task of the NLP
team in subproject B1, which consists of Prof. Dr. Michaela Geierhos and Dr. Frederik Baumer. Prof. Geierhos
is a full professor for Digital Humanities at the University of Paderborn, Germany and Dr. Baumer works as a
postdoctoral researcher at the same chair.</p>
      <p>Michaela Geierhos. Before Prof. Dr. Geierhos became a professor for Digital Humanities, she was an
assistant professor for Semantic Information Processing at the University of Paderborn and worked as a postdoc
at the Ludwig-Maximilians-University in Munich. After completing her studies in computational linguistics,
computer science and phonetics, she worked as a research associate at the Center for Information and Speech
Processing from 2006 to 2012. In 2010 she obtained a doctoral degree in computer linguistics on \BiographIE
Classi cation and Extraction of Career-Speci c Information".</p>
      <p>Frederik Baumer. Dr. Frederik Baumer has been working as a research associate at the Semantic
Information Processing group in Paderborn since completing his master's degree in \Management Information
Systems" with a focus on Semantic Information Processing. Since July 2013, he has been involved in various
research projects at the Heinz Nixdorf Institute { initially as a master's student and then as doctoral student
{ and is an expert for search engine technology and information extraction. While participating in the
Collaborative Research Center 901 \On-the-Fly Computing", he has been working on his doctorate since
October 2014. Dr. Baumer successfully defended his PhD thesis in July 2017 on \Indicator-based Detection
and Compensation of Inaccurate and Incompletely Described Software Requirements".</p>
    </sec>
    <sec id="sec-3">
      <title>Past Research on NLP for Requirement Engineering</title>
      <p>At the beginning of our work, we dealt in particular with the question of the characteristics of NL requirements
descriptions [GSB15]. It quickly became clear that a major challenge is o -topic information, incompleteness and
ambiguity { problems that have already been subject of research in this area for a long time. For the automatic
composition of software services, it is particularly crucial because in OTF Computing, it is better assumed that
high-quality, formal speci cations are provided as input, which are qualitatively better than NL requirements.
There was, at this point, no automated procedure for checking and compensating requirements.</p>
      <p>As a rst challenge, we understood how to separate relevant information from irrelevant and how to identify
the core components of requirements [BDG17, DG16]. After some attempts with rule-based approaches, we
decided to go for a machine-learning approach because the text quality of user-generated requirements is very
poor and we achieved only a bad recall through using rule-based approaches [DG16]. We then focused on
the language de cits mentioned before. They are di cult to recognize and to compensate for end-users and
sometimes cause great damage in the composition process [Bau17]. However, each language de cit has di erent
characteristics and must be recognized and compensated di erently. In the case of incompleteness, we focused
on the predicate-argument structure in requirements descriptions and tried to identify missing arguments and
thus incompleteness [BG16]. For this procedure, extensive resources are necessary, which did not exist so far
and were created by ourselves. Since there are hardly any real-life requirements to nd, we have instead used
software descriptions and feature requests from app stores (e.g., download.com) and developer forums (e.g.,
sourceforge.com) [BDG17, DG16].</p>
      <p>There is a lot of work done on ambiguity in requirement descriptions. We distinguish between
lexical, syntactic and referential ambiguity. All forms require di erent detection and compensation methods
[Bau17, GB17, GSB15]. Even more di cult to detect is vagueness. Here we have developed a rule-based
approach that can recognize some forms of vagueness, but not all [GB17]. Since the compensation of vagueness
is more di cult because of their di erent characteristics compared with ambiguities, we introduce a dialogue
with the user (cf. Section 4). Based on this research and the realization that linguistic de cits are so di erent,
we focused on a strategic approach for the optimization of NL: The goal was to develop a parameterized model
(so-called strategy con guration) that automatically selects the best tting strategy to compensate de ciencies
in software speci cations [Bau17]. In order to enable this, suitable methods for the recognition and resolution
of lexical, syntactic, and referential ambiguity and for the completion of NL requirements were rst identi ed in
literature and then implemented. This is the starting point of the indicator-based strategies (cf. Figure 2).</p>
      <p>Requirements
Description</p>
      <p>Indicators</p>
      <p>Strategy index
Indicator-based</p>
      <p>Con guration
Preprocessing and</p>
      <p>compensation</p>
      <p>These strategies enable CORDULA to execute and monitor individual methods, to compare results and to
correct them if necessary as well as to use synergy e ects. For this purpose, it was necessary to de ne
contextsensitive indicators to determine the compensation need, which can trigger individual strategies. Employing these
strategies, CORDULA is able to apply existing, highly heterogeneous processing and compensation techniques
on requirements descriptions. It is only possible by means of the linguistic indicators to capture and optimize
the individual text quality in a data-driven manner, as required, because CORDULA's compensation pipeline is
con gured on the basis of the detected shortcomings. In addition, the individual results of the components are
merged and structured as a consistent compensation result.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Research Plan on NLP for Requirement Engineering</title>
      <p>As mentioned before, language de cits in NL requirements can only be compensated automatically in parts. A
complete compensation without any questions to the users is in many cases not possible, especially if vagueness
and incompleteness are present. For this reason, we are currently working to extend CORDULA with a chat
functionality. The users can submit their requirements to CORDULA via chat (left part of Figure 3) and receive
an overview of the processed requirements as return (right part of Figure 3). If CORDULA is able to detect but
not to compensate a problem, the system automatically asks the user to provide some helpful information or to
make a decision (e.g., annotate (a part of) a sentence as an o -topic, as shown in Figure 3).</p>
      <p>First attempts with existing chatbot frameworks such as Google's api.ai or IBM Watson Conversations worked
well and brought up interesting ndings. They pointed out that we need more exibility and control in the
conversation and less free communication (which the framework aims for). The dialogue should not be conducted
completely by the chatbot, but should address a speci c problem and provide targeted information for
CORDULA. For this reason, in our current vision we use our own chat solution, which can ask prede ned questions
to the user about prede ned problems. In this way, we expect to get further relevant user input for a speci c
problem just-in-time. One criticism of this bidirectional chat approach remains, however: chatting slows down
the entire OTF service composition process. Nevertheless, an improvement in the overall result can be expected,
since the requirements will be more concrete and complete. This trade o needs to be discussed in future work.
Acknowledgements
This work was partially supported by the German Research Foundation (DFG) within the Collaborative Research
Center \On-The-Fly Computing" (SFB 901).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Bau17]
          <article-title>Frederik Simon Baumer. Indikatorbasierte Erkennung und Kompensation von ungenauen und unvollstandig beschriebenen Softwareanforderungen</article-title>
          .
          <source>Phd thesis</source>
          , University of Paderborn, Paderborn, Germany,
          <year>July 2017</year>
          . ISBN:
          <fpage>978</fpage>
          -3-
          <fpage>942647</fpage>
          -91-5.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [BDG17]
          <article-title>Frederik Simon Baumer</article-title>
          , Markus Dollmann, and
          <string-name>
            <given-names>Michaela</given-names>
            <surname>Geierhos</surname>
          </string-name>
          .
          <article-title>Studying software descriptions in sourceforge and app stores for a better understanding of real-life requirements</article-title>
          . In Federica Sarro, Emad Shihab, Meiyappan Nagappan, Marie Christin Platenius, and Daniel Kaimann, editors,
          <source>Proceedings of the 2nd ACM SIGSOFT International Workshop on App Market Analytics</source>
          , pages
          <volume>19</volume>
          {
          <fpage>25</fpage>
          , Paderborn, Germany,
          <year>September 2017</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [BG16]
          <article-title>[DG16] [GB16] [GB17] Frederik Simon Baumer and Michaela Geierhos. Running out of words: How similar user stories can help to elaborate individual natural language requirement descriptions</article-title>
          .
          <source>In Giedre Dregvaite and Robertas Damasevicius</source>
          , editors,
          <source>Proceedings of the 22nd Conference on Information and Software Technologies (ICIST</source>
          <year>2016</year>
          ), volume
          <volume>639</volume>
          of Communications in Computer and Information Science,
          <article-title>chapter Running Out of Words: How Similar User Stories Can Help to Elaborate Individual Natural Language Requirement Descriptions</article-title>
          , pages
          <volume>549</volume>
          {
          <fpage>558</fpage>
          . Springer International Publishing, Druskininkai, Lithuania,
          <year>October 2016</year>
          . ISBN:
          <fpage>978</fpage>
          -3-
          <fpage>319</fpage>
          -46254-7.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Markus</given-names>
            <surname>Dollmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Michaela</given-names>
            <surname>Geierhos</surname>
          </string-name>
          .
          <article-title>On- and o -topic classi cation and semantic annotation of user-generated software requirements</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <year>1807</year>
          {
          <year>1816</year>
          , Austin, TX, USA,
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          November
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Michaela</given-names>
            <surname>Geierhos</surname>
          </string-name>
          and
          <article-title>Frederik Simon Baumer. How to complete customer requirements using concept expansion for requirement re nement</article-title>
          .
          <source>In Proceedings of the 21st International Conference on Applications of Natural Language to Information Systems (NLDB</source>
          <year>2016</year>
          ), volume
          <volume>9612</volume>
          , pages
          <fpage>37</fpage>
          {
          <fpage>47</fpage>
          ,
          <string-name>
            <surname>Salford</surname>
          </string-name>
          , UK,
          <fpage>22</fpage>
          -
          <lpage>24</lpage>
          June 2016. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Michaela</given-names>
            <surname>Geierhos</surname>
          </string-name>
          and
          <article-title>Frederik Simon Baeumer</article-title>
          .
          <article-title>Guesswork? resolving vagueness in user-generated software requirements</article-title>
          . In Henning Christiansen,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dolores Jimenez</surname>
          </string-name>
          <string-name>
            <surname>Lopez</surname>
          </string-name>
          , Roussanka Loukanova, and Lawrence S. Moss, editors,
          <source>Partiality and Underspeci cation in Information, Languages, and Knowledge, Partiality and Underspeci cation in Information, Languages, and Knowledge</source>
          , chapter
          <volume>3</volume>
          , pages
          <fpage>65</fpage>
          {
          <fpage>108</fpage>
          . Cambridge Scholars Publishing,
          <year>September 2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [GSB15]
          <article-title>Michaela Geierhos, Sabine Schulze, and Frederik Simon Baumer. What did you mean? facing the challenges of user-generated software requirements</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Agents and Arti cial Intelligence</source>
          , pages
          <fpage>277</fpage>
          {
          <fpage>283</fpage>
          ,
          <fpage>10</fpage>
          -
          <lpage>12</lpage>
          January
          <year>2015</year>
          . Lisbon, Portugal. ISBN:
          <fpage>978</fpage>
          -
          <lpage>989</lpage>
          -758-073-4.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>