<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the EVALITA 2016 Question Answering for Frequently Asked Questions (QA4FAQ) Task</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>ADAPT Centre</institution>
          ,
          <addr-line>Dublin</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science, University of Bari Aldo Moro</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Francesco Lovecchio</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>QuestionCube S.r.l</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This paper describes the first edition of the Question Answering for Frequently Asked Questions (QA4FAQ) task at the EVALITA 2016 campaign. The task concerns the retrieval of relevant frequently asked questions, given a user query. The main objective of the task is the evaluation of both question answering and information retrieval systems in this particular setting in which the document collection is composed of FAQs. The data used for the task are collected in a real scenario by AQP Risponde, a semantic retrieval engine used by Acquedotto Pugliese (AQP, the Organization for the management of the public water in the South of Italy) for supporting their customer care. The system is developed by QuestionCube, an Italian startup company which designs Question Answering tools.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Questo lavoro descrive la prima
edizione del Question Answering for
Frequently Asked Questions (QA4FAQ) task
proposto durante la campagna di
valutazione EVALITA 2016. Il task consiste
nel recuperare le domande piu` frequenti
rilevanti rispetto ad una domanda posta
dall’utente. L’obiettivo principale del task
e` la valutazione di sistemi di question
answering e di recupero dell’informazione in
un contesto applicativo reale, utilizzando i
dati provenienti da AQP Risponde, un
motore di ricerca semantico usato da
Acquedotto Pugliese (AQP, l’ente per la gestione
dell’acqua pubblica nel Sud Italia). Il
sistema e` sviluppato da QuestionCube, una
startup italiana che progetta soluzioni di
Question Answering.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Motivation</title>
      <p>Searching within the Frequently Asked Questions
(FAQ) page of a web site is a critical task:
customers might feel overloaded by many irrelevant
questions and become frustrated due to the
difficulty in finding the FAQ suitable for their
problems. Perhaps they are right there, but just worded
in a different way than they know.</p>
      <p>The proposed task consists in retrieving a list of
relevant FAQs and corresponding answers related
to the query issued by the user.</p>
      <p>Acquedotto Pugliese (AQP) developed a
semantic retrieval engine for FAQs, called AQP
Risponde1, based on Question Answering (QA)
techniques. The system allows customers to ask
their own questions, and retrieves a list of
relevant FAQs and corresponding answers.
Furthermore, customers can select one FAQ among those
retrieved by the system and can provide their
feedback about the perceived accuracy of the answer.</p>
      <p>AQP Risponde poses relevant research
challenges concerning both the usage of the Italian
language in a deep QA architecture, and the variety
of language expressions adopted by customers to
formulate the same information need.</p>
      <p>
        The proposed task is strongly related to the
one recently organized at Semeval 2015 and 2016
about Answer Selection in Community Question
Answering
        <xref ref-type="bibr" rid="ref5">(Nakov et al., 2015)</xref>
        . This task helps
to automate the process of finding good answers
to new questions in a community-created
discussion forum (e.g., by retrieving similar questions in
1http://aqprisponde.aqp.it/ask.php
the forum and by identifying the posts in the
answer threads of similar questions that answer the
original one as well). Moreover, the QA-FAQ has
some common points with the Textual Similarity
task
        <xref ref-type="bibr" rid="ref1">(Agirre et al., 2015)</xref>
        that received an
increasing amount of attention in recent years.
      </p>
      <p>The paper is organized as follows: Section 2
describes the task, while Section 3 provides details
about competing systems. Results of the task are
discussed in Section 4.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Task Description: Dataset, Evaluation</title>
    </sec>
    <sec id="sec-4">
      <title>Protocol and Measures</title>
      <p>The task concerns the retrieval of relevant
frequently asked questions, given a user query. For
defining an evaluation protocol, we need a set of
FAQs, a set of user questions and a set of relevance
judgments for each question. In order to collect
these data, we exploit an application called AQP
Risponde, developed by QuestionCube for the
Acquedotto Pugliese. AQP Risponde provides a
back-end that allows to analyze both the query log
and the customers’ feedback to discover, for
instance, new emerging problems that need to be
encoded as FAQ. AQP Risponde is provided as web
and mobile application for Android2 and iOS3 and
is currently running in the Acquedotto Pugliese
customer care. AQP received about 25,000
questions and collected about 2,500 user feedback. We
rely on these data to build the dataset for the task.
In particular, we provide:
a knowledge base of 406 FAQs. Each FAQ is
composed of a question, an answer, and a set
of tags;
a set of 1,132 user queries. The queries
are collected by analyzing the AQP Risponde
system log. From the initial set of queries, we
removed queries that contains personal data;
a set of 1,406 pairs &lt; query; relevantf aq &gt;
that are exploited to evaluate the
contestants. We build these pairs by analyzing the
user feedback provided by real users of AQP
Risponde. We manually check the user
feedback in order to remove noisy or false
feedback. The check was performed by two
experts of the AQP customer support.</p>
      <p>2https://play.google.com/store/apps/
details?id=com.questioncube.aqprisponde&amp;
hl=it</p>
      <p>3https://itunes.apple.com/it/app/
aqp-risponde/id1006106860</p>
      <p>We provided a little sample set for the system
development and a test set for the evaluation. We
did not provide a set of training data: AQP is
interested in the development of unsupervised systems
because AQP Risponde must be able to achieve
good performance without any user feedback.
Following, an example of FAQ is reported:
Question “Come posso telefonare al numero
verde da un cellulare?” How can I call the
toll-free number by a mobile phone?
Answer “E` possibile chiamare il Contact Center
AQP per segnalare un guasto o per un pronto
intervento telefonando gratuitamente anche
da cellulare al numero verde 800.735.735.
Mentre per chiamare il Contact Center AQP
per servizi commerciali 800.085.853 da un
cellulare e dall’estero e` necessario comporre
il numero +39.080.5723498 (il costo della
chiamata e` secondo il piano tariffario del
chiamante).” You can call the AQP Contact
Center to report a fault or an emergency call
without charge by the phone toll-free number
800 735 735...</p>
      <p>Tags canali, numero verde, cellulare</p>
      <p>For example, the previous FAQ is relevant for
the query: “Si puo` telefonare da cellulare al
numero verde?” Is it possible to call the toll-free
number by a mobile phone?</p>
      <p>Moreover, we provided a simple baseline based
on a classical information retrieval model.
2.1</p>
      <p>Data Format
FAQs are provided in both XML and CSV format
using “;” as separator. The file is encoded in
UTF8 format. Each FAQ is described by the following
fields:
id a number that uniquely identifies the FAQ
question the question text of the current FAQ
answer the answer text of the current FAQ
tag a set of tags separated by “,”</p>
      <p>Test data are provided as a text file composed by
two strings separated by the TAB character. The
first string is the user query id, while the second
string is the text of the user query. For example:
“1 Come posso telefonare al numero verde da un
cellulare?” and “2 Come si effettua l’autolettura
del contatore?”.
The baseline is built by using Apache Lucene (ver.
4.10.4)4. During the indexing for each FAQ, a
document with four fields (id, question, answer,
tag) is created. For searching, a query for each
question is built taking into account all the
question terms. Each field is boosted according to the
following score question=4, answer=2 and tag=1.
For both indexing and search the ItalianAnalyzer
is adopted. The top 25 documents for each query
are provided as result set. The baseline is freely
available on GitHub5 and it was released to
participants after the evaluation period.
The participants must provide results in a text file.
For each query in the test data, the participants can
provide 25 answers at the most, ranked according
by their systems. Each line in the file must contain
three values separated by the TAB character: &lt;
queryid &gt;&lt; f aqid &gt;&lt; score &gt;.</p>
      <p>
        Systems are ranked according to the
accuracy@1 (c@1). We compute the precision of the
system by taking into account only the first
correct answer. This metric is used for the final
ranking of systems. In particular, we take into account
also the number of unanswered questions,
following the guidelines of the CLEF ResPubliQA Task
        <xref ref-type="bibr" rid="ref6">(Pen˜as et al., 2009)</xref>
        . The formulation of c@1 is:
n1 (nR + nU nnR )
(1)
where nR is the number of questions correctly
answered, nU is the number of questions
unanswered, and n is the total number of questions.
      </p>
      <p>The system should not provide result for a
particular question when it is not confident about the
correctness of its answer. The goal is to reduce the
amount of incorrect responses, keeping the
number of correct ones, by leaving some questions
unanswered. Systems should ensure that only the
portion of wrong answers is reduced, maintaining
as high as possible the number of correct answers.
Otherwise, the reduction in the number of correct
answers is punished by the evaluation measure for
both the answered and unanswered questions.</p>
      <sec id="sec-4-1">
        <title>4http://lucene.apache.org/ 5https://github.com/swapUniba/qa4faq</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Systems</title>
      <p>
        Thirteen teams registered in the task, but only
three of them actually submitted the results for the
evaluation. A short description of each system
follows:
chiLab4It - The system described in
        <xref ref-type="bibr" rid="ref7 ref8">(Pipitone et
al., 2016a)</xref>
        is based on the cognitive model
proposed in
        <xref ref-type="bibr" rid="ref7 ref8">(Pipitone et al., 2016b)</xref>
        . When a
support text is provided for finding the
correct answer, QuASIt is able to use this text
to find the required information. ChiLab4It
is an adaptation of this model to the context
of FAQs, in this case the FAQ is exploited
as support text: the most relevant FAQ will
be the one whose text will best fit the user’s
question. The authors define three
similarity measures for each field of the FAQ:
question, answer and tags. Moreover, an
expansion step by exploiting synonyms is applied
to the query. The expansion module is based
on Wiktionary.
fbk4faq - In
        <xref ref-type="bibr" rid="ref4">(Fonseca et al., 2016)</xref>
        , the authors
proposed a system based on vector
representations for each query, question and answer.
Query and answer are ranked according to the
cosine distance to the query. Vectors are built
by exploring the word embeddings generated
by
        <xref ref-type="bibr" rid="ref3">(Dinu et al., 2014)</xref>
        , and combined in a way
to give more weight to more relevant words.
NLP-NITMZ the system proposed by
        <xref ref-type="bibr" rid="ref2">(Bhardwaj et al., 2016)</xref>
        is based on a classical
VSM model implemented in Apache Nutch6.
Moreover, the authors add a combinatorial
searching technique that produces a set of
queries by several combinations of all the
keywords occurring in the user query. A
custom stop word list was developed for the task,
which is freely available7.
      </p>
      <p>It is important to underline that all the systems
adopt different strategies, while only one system
(chiLab4It) is based on a typical question answer
module. We provide a more detailed analysis
about this aspect in Section 4.
Results of the evaluation in terms of c@1 are
reported in Table 1. The best performance is
obtained by the chilab4it team, that is the only one
able to outperform the baseline. Moreover, the
chilab4it team is the only one that exploits
question answering techniques: the good performance
obtained by this team proves the effectiveness of
question answering in the FAQ domain. All the
other participants had results under the baseline.
Another interesting outcome is that the baseline
exploiting a simple VSM model achieved
remarkable results.</p>
      <p>
        A deep analysis of results is reported in
        <xref ref-type="bibr" rid="ref4">(Fonseca et al., 2016)</xref>
        , where the authors have built
a custom development set by paraphrasing
original questions or generating a new question (based
on original FAQ answer), without considering the
original FAQ question. The interesting result is
that their system outperformed the baseline on the
development set. The authors underline that the
development set is completely different from the
test set which contains sometime short queries and
more realistic user’s requests. This is an
interesting point of view since one of the main challenge
of our task concerns the variety of language
expressions adopted by customers to formulate the
information need. Moreover, in their report the
authors provide some examples in which the FAQ
reported in the gold standard is less relevant than
the FAQ reported by their system, or in some cases
the system returns a correct answer that is not
annotated in the gold standard. Regarding the first
point, we want to point out that our relevance
judgments are computed according to the users’
feedback and reflect their concept of relevance8.
      </p>
      <sec id="sec-5-1">
        <title>6https://nutch.apache.org</title>
        <p>7https://github.com/SRvSaha/
QA4FAQ-EVALITA-16/blob/master/italian_
stopwords.txt
8Relevance is subjective.</p>
        <p>We tried to mitigate issues related to relevance
judgments by manually checking users’ feedback.
However, this manual annotation process might
have introduced some noise, which is common to
all participants.</p>
        <p>Regarding missing correct answers in the gold
standard: this is a typical issue in the retrieval
evaluation, since it is impossible to assess all the FAQ
for each test query. Generally, this issue can be
solved by creating a pool of results for each query.
Such pool is built by exploiting the output of
several systems. In this first edition of the task, we
cannot rely on previous evaluations on the same
set of data, therefore we chose to exploit users’
feedback. In the next editions of the task, we can
rely on previous results of participants to build that
pool of results.</p>
        <p>Finally, in Table 2 we report some
information retrieval metrics for each system9. In
particular, we compute Mean Average Precision (MAP),
Geometrical-Mean Average Precision (GMAP),
Mean Reciprocal Rank (MRR), Recall after five
(R@5) and ten (R@10) retrieved documents.
Finally we report the success 1 that is equal to c@1,
but without taking into account answered queries.
We can notice that on retrieval metrics the
baseline is the best approach. This was quite expected
since an information retrieval model tries to
optimize retrieval performance. Conversely, the best
approach according to success 1 is the chilab4it
system based on question answering, since it tries
to retrieve a correct answer in the first position.
This result suggests that the most suitable
strategy in this context is to adopt a question
answering model, rather than to adapt an information
retrieval approach. Another interesting outcome
concerns the system NLP-NITMZ.1, which obtains
an encouraging success 1, compared to the c@1.
This behavior is ascribable to the fact that the
system does not adopt a strategy that provides an
answer for all queries.
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>For the first time for the Italian language, we
propose a question answering task for frequently
asked questions. Given a user query, the
participants must provide a list of FAQs ranked by
relevance according to the user need. The collection
9Metrics are computed by the latest version of
the trec eval tool: http://trec.nist.gov/trec_
eval/
baseline
NLP-NITMZ.1
NLP-NITMZ.2
0.5190
of FAQs was built by exploiting a real
application developed by QuestionCube for Acquedotto
Pugliese. The relevance judgments for the
evaluation are built by taking into account the user
feedback.</p>
      <p>Results of the evaluation demonstrated that only
the system based on question answering
techniques is able to outperform the baseline, while
all the other participants reported results under the
baseline. Some issues pointed out by participants
suggest exploring a pool of results for building
more accurate judgments. We plan to implement
this approach in future editions of the task.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work is supported by the project
“Multilingual Entity Liking” funded by the Apulia Region
under the program FutureInResearch.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Eneko</given-names>
            <surname>Agirre</surname>
          </string-name>
          , Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalara,
          <string-name>
            <given-names>Rada</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          , et al.
          <year>2015</year>
          .
          <article-title>Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability</article-title>
          .
          <source>In Proceedings of the 9th international workshop on semantic evaluation (SemEval</source>
          <year>2015</year>
          ), pages
          <fpage>252</fpage>
          -
          <lpage>263</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Divyanshu</given-names>
            <surname>Bhardwaj</surname>
          </string-name>
          , Partha Pakray, Jereemi Bentham, Saurav Saha, and
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Question Answering System for Frequently Asked Questions</article-title>
          . In Pierpaolo Basile, Anna Corazza, Franco Cutugno, Simonetta Montemagni, Malvina Nissim, Viviana Patti, Giovanni Semeraro, and Rachele Sprugnoli, editors,
          <source>Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>Final Workshop (EVALITA</source>
          <year>2016</year>
          ).
          <article-title>Associazione Italiana di Linguistica Computazionale (AILC).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Georgiana</given-names>
            <surname>Dinu</surname>
          </string-name>
          , Angeliki Lazaridou, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Improving zero-shot learning by mitigating the hubness problem</article-title>
          .
          <source>arXiv preprint arXiv:1412</source>
          .
          <fpage>6568</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Erick R. Fonseca</surname>
          </string-name>
          , Simone Magnolini, Anna Feltracco,
          <string-name>
            <surname>Mohammed R. H. Qwaider</surname>
            , and
            <given-names>Bernardo</given-names>
          </string-name>
          <string-name>
            <surname>Magnini</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Tweaking Word Embeddings for FAQ Ranking</article-title>
          . In Pierpaolo Basile, Anna Corazza, Franco Cutugno, Simonetta Montemagni, Malvina Nissim, Viviana Patti, Giovanni Semeraro, and Rachele Sprugnoli, editors,
          <source>Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>Final Workshop (EVALITA</source>
          <year>2016</year>
          ).
          <article-title>Associazione Italiana di Linguistica Computazionale (AILC).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Preslav</given-names>
            <surname>Nakov</surname>
          </string-name>
          , Lluıs Marquez, Walid Magdy, Alessandro Moschitti, James Glass, and
          <string-name>
            <given-names>Bilal</given-names>
            <surname>Randeree</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Semeval-2015 task 3: Answer selection in community question answering</article-title>
          .
          <source>SemEval-2015</source>
          , page 269.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Anselmo</given-names>
            <surname>Pen</surname>
          </string-name>
          ˜as, Pamela Forner, Richard Sutcliffe,
          <string-name>
            <surname>A</surname>
          </string-name>
          ´ lvaro Rodrigo, Corina Fora˘scu, In˜aki Alegria, Danilo Giampiccolo, Nicolas Moreau, and
          <string-name>
            <given-names>Petya</given-names>
            <surname>Osenova</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Overview of ResPubliQA 2009: question answering evaluation over European legislation</article-title>
          .
          <source>In Workshop of the Cross-Language Evaluation Forum for European Languages</source>
          , pages
          <fpage>174</fpage>
          -
          <lpage>196</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Arianna</given-names>
            <surname>Pipitone</surname>
          </string-name>
          , Giuseppe Tirone, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Pirrone</surname>
          </string-name>
          .
          <year>2016a</year>
          .
          <article-title>ChiLab4It System in the QA4FAQ Competition</article-title>
          . In Pierpaolo Basile, Anna Corazza, Franco Cutugno, Simonetta Montemagni, Malvina Nissim, Viviana Patti, Giovanni Semeraro, and Rachele Sprugnoli, editors,
          <source>Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>Final Workshop (EVALITA</source>
          <year>2016</year>
          ).
          <article-title>Associazione Italiana di Linguistica Computazionale (AILC).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Arianna</given-names>
            <surname>Pipitone</surname>
          </string-name>
          , Giuseppe Tirone, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Pirrone</surname>
          </string-name>
          .
          <year>2016b</year>
          .
          <article-title>QuASIt: a Cognitive Inspired Approach to Question Answering System for the Italian Language</article-title>
          .
          <source>In Proceedings of the 15th International Conference on the Italian Association for Artificial Intelligence</source>
          <year>2016</year>
          . aAcademia University Press.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>