<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Approach to the Design and Evaluation of an Enterprise Search Application</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Zilio</string-name>
          <email>daniel.zilio@unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maristella Agosti</string-name>
          <email>maristella.agosti@unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Turato</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Cultural Heritage, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Research and Development, SIAV</institution>
          ,
          <addr-line>Rubano, Padua</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper reports on an experience of designing and implementing an enterprise search application. One of the motivations for the work is the lack of previous studies that deal with the problem of evaluating an enterprise search application. The reason for this lack depends on the building a signi cant test collection, because each company may be interested in dealing with di erent types of topic and data format possibly with atypical content formats.</p>
      </abstract>
      <kwd-group>
        <kwd>enterprise search</kwd>
        <kwd>enterprise search application</kwd>
        <kwd>document formats</kwd>
        <kwd>enterprise search evaluation</kwd>
        <kwd>test collection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Enterprise search is the term commonly used to de ne information retrieval in a
business setting, o ering users a way to search for informative content generated
within their companies. By comparing it with web search and desktop search, we
can identify some distinctive features: the content is stored in di erent
information systems (e.g. data sources like customer relationship management software
(CRM), enterprise resource planning software (ERP), intranets), di erent le
formats are used (e.g. pdf, docx, xls), including both structured (e.g. database)
and unstructured content (e.g. scanned documents), and di erent access levels
give users di erent rights.</p>
      <p>
        The purpose of enterprise search is to enable users to e ectively nd the
information they need to perform their tasks, while at the same time requiring
a minimal e ort for the users and low costs to be sustained by the company in
terms of ine ciencies [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Many of the problems of enterprise search have been
addressed in previous experiences, as, for example, those reported in [5{7]. The
notion of relevance, for example, can be di erent from that used in web search
where there are usually many documents relevant to a query, and the ranking
?? Siav S.p.A. has supported this work by providing the data and expertise in document
management needed to accomplish the study.
tends to favor the most popular ones. Instead the typical enterprise search query
has few correct answers. Dealing with enterprise information content has also
some bene ts to be exploited: the content is produced for dissemination purpose,
unlike web content which is usually written to attract people. Moreover, we can
obtain more contextual information on the queries: since the users are inside a
company, we can obtain detailed pro le information (e.g. role, experience, skills,
team) or very precise location inside the company's building. Nonetheless, more
e ective algorithms that leverage those kind of bene ts have to be developed.
      </p>
      <p>One of the motivations for the work reported here is the lack of previous
works that deal with the problem of evaluating an enterprise search application.
The reason for this lack may depend on the di culties in building a signi cant
test collection: each company can be interested in dealing with very di erent
types of topics and data format possibly with atypical content formats (e.g. in
the case of the study reported in this paper CRM events were also dealt with),
and the company may only be able to allocate limited resources to this type of
activity.</p>
      <p>In the present work we present a real case study of the design, development
and evaluation of an application of enterprise search within a company, and draw
attention to the issues that had to be overcome to complete the task. The paper
also details the technological stack employed to build and test the enterprise
search application.</p>
      <p>In Figure 1 we sketch the di erent steps to design and implement the
enterprise application together with its evaluation. The gure can also be used as a
sort of outline of the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>Related</title>
    </sec>
    <sec id="sec-3">
      <title>Works</title>
      <p>The bibliographic research we conducted showed a lack of work on assessment
and evaluation methodologies of enterprise search. This may be due to the fact
that the importance of conducting an assessment to contribute to the
improvement of the quality o ered by information retrieval services has not yet been
perceived, especially in the industrial sector. We must also consider that most
of the research in this eld is being carried out within companies, so we cannot
exclude the possibility that many companies simply prefer not to disclose results
and information considered to be strategic.</p>
      <p>
        The results we present are based on [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] which reports on the study of
designing and implementing a software prototype of an enterprise search application.
The software prototype has been made interoperable with the Archi ow4
document management application that is a software product of SIAV. The work
has also been inspired by the example of using a real experimental collection to
analyze the behavior of a system illustrated in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and by the work to design a
model of a graph-based enterprise search engine provided in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>The Design of an Enterprise Search Application</title>
      <p>The objective of this work was to produce a prototype able to perform the
effective retrieval of documents that constitute the di erent collections present
within the SIAV internal installation of the document management platform
Archi ow. The system is designed to integrate the collections of documents with
other content from sources other than Archi ow itself, in order to provide the
end user with a spectrum of documents related to a speci c subject of interest.
An example which clari es this aim is a search that requires the documents
related to a particular customer: in addition to providing these documents, while
maintaining e cacy as a metric, the system will provide additional related
contents, like the latest tickets opened by the customer (available in the helpdesk
management system), recently opened business opportunities with the customer
derived from CRM or news on the web concerning the customer. This allows us
to anticipate possible information needs, thus obtaining an advantage in terms
of e ciency, in accordance with a business intelligence perspective.</p>
      <p>The reference users are the professionals within SIAV. The company made
available a set of data and documents from di erent information sources,
including those managed by the di erent systems available within the company;
in this way we were able to use these information sources to build a collection
to evaluate the system.</p>
      <p>
        The platform adopted is Apache Solr [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which has been chosen after a
benchmark process that saw Apache Solr emerge as a suitable tool for industrial
applications.
      </p>
      <sec id="sec-4-1">
        <title>4 https://www.siav.com/software-solutions/archiflow/</title>
        <p>3.1</p>
        <sec id="sec-4-1-1">
          <title>The Di erent Data Sources and Their Formats</title>
          <p>The di erent data sources that have been considered are:
{ Documents: The documents are a subset of the document base of SIAV, that
is, a part of the documents that are managed by the local installation of
Archi ow at SIAV headquarters. The documents are of di erent types (such
as invoices, emails, project sheets, etc.) and of di erent formats
(including DOC, DOCX, TXT, PDF, TIF, MSG, EML, RTF, XLS, and XML). The metadata of
each document are stored in an associated card from which the metadata can
be exported and extracted in CSV format. For this category of data sources
each document, together with the metadata related to it, is considered a
basic unit of information.
{ Events: SIAV is equipped with an issue tracking system that collects the
tickets regarding the technical problems reported by their customers. The
company manages the ticket elements through a CRM, which allows the
coordination of the activities of the sta working at the helpdesk and the
technical department. In this way all the actions performed by the sta
involved are recorded in the system. Those actions are named events. It is
possible to export the events from the CRM, and each event is enriched with
a set of descriptive attributes. A total of 70,475 events were considered and
studied; those events were collected over a period of about ve years.
{ Web Pages: During the work only a focused portion of the Web was
examined; since the study on this portion of the Web is preliminary, the related
results are not reported here.</p>
          <p>It is important to note that objects within di erent sources may actually
contain related information; e.g. a document can be related to a particular project
and also a ticket item can be linked to it. One of the requirements of the system
was to allow the user to exploit this information. For all considered data sources,
the reference text language is Italian.
3.2</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Text extraction</title>
          <p>Since the documents were in di erent formats, an ingestion utility had to be
developed and used to generate a new version of each document, in a format
usable by Solr, for the necessary preprocessing and indexing procedures.</p>
          <p>The rst part of the ingestion work was the extraction of the textual contents
from the documents: for each document le the text was extracted and saved in
a simple text le, so as to allow the subsequent loading and veri cation through
Solr of the quality of the extraction procedures. This processing step has been
supported by di erent technical tools.</p>
          <p>The documents of the used collection were:
{ communications (11,974)
{ bids and tenders (59,763)
{ project reports (4,962)
for a total of 76,699 indexed documents.</p>
          <p>The permanent memory space necessary for the storage of the preprocessed
documents is 2.13 GB while the permanent memory used to store the original
documents was more than 60 GB. For each original document, the nal document
with a valid format to be imported into Solr was created by adding the metadata
of the original document to its extracted textual content. To manage and index
all the documents (textual and metadata documents) through the search engine
the library package SolrNet5 has been used.</p>
          <p>Information on document number and permanent memory size needed to
store the indexed documents are reported in Table 1 and 2.</p>
          <p>Document Sources
Total Archi ow Documents 76,699</p>
          <p>
            Total Events 70,475
Total Document Number 147,174
Evaluation in IR is a well-established practice that has been studied for many
years [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]. In accordance with this practice, an IR system or an IR application
is evaluated using a test collection C constructed speci cally for a given task.
Therefore, it is necessary to select a set D of documents relevant to the task, a set
of topics T that re ect the real information needs of users in relation to the task
and de ne the relevance judgements RJ , so that each document is considered
relevant or not relevant to a given topic.
          </p>
          <p>The test collection to be used for the evaluation has to be built before starting
the evaluation experiments and it is composed of the triple: C = D; T; RJ .</p>
          <p>In this case the possibility of using an existing test collection was rejected
because there is no test collection that is suitable to the type of tasks of
interest. Therefore the test collection needed to be built. To build the test collection
di erent alternatives were studied and the following operational strategy,
exploiting a pooling methodology, was eventually decided. Of the three categories
of documents { documents from Archi ow, events and web pages { it was decided
to use documents and events that are peculiar to the company; web pages were
not considered strategic for the design purpose. With the support of a thorough
analysis of the tasks of interest conducted with the company management, ve
topics of interest for each document category were de ned, and each topic was
translated by a domain expert into a query. The pooling was done using Solr,</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>5 https://github.com/mausch/SolrNet</title>
        <p>Terrier6 and two Desktop Search named Copernic7 and Windows Search8. For
each of these search engines, the following actions were conducted: indexing of
the document base; execution of queries to produce the run; calculation of the
number of documents to be judged on the basis of di erent depths of cut.</p>
        <p>
          The achieved results [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] con rmed the choice of using Solr, as the results
are comparable to those obtained with Terrier, while the two desktop search
systems proved to be ine cient, as they returned not relevant results. It has to
be noted, however, that while the obtained results querying the event base were
relevant, those relating to documents originating from Archi ow were not, and
this could be the consequence of having de ned only ve topics, a number that
has to be considered insu cient for this documentary base.
        </p>
        <p>The de nition of the document base to be used and the subsequent evaluation
of the results raised some issues: the formulation of the topics was really
challenging; the process of de ning the queries required the systematic cooperation
between domain user experts and IR experts; the calculation and veri cation
of all the calculated metrics was supported by MATTERS9, a software toolkit
written in Matlab for information retrieval evaluation; many activities requiring
great time and e ort were needed to obtain the pool.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>Before this project SIAV used to rely on third party products to provide its
customers with advanced search features, because the base search module in
SIAV Archi ow provided only metadata based search.</p>
      <p>The purpose of this work for SIAV was twofold: (1) develop a more exible
retrieval system and (2) acquire the necessary knowledge to put the enterprise
search application into production, maintain it and in the future be able to
enrich it with new desired features, in order to cope with new search needs
of its customers. These goals were positively met, but there is still the need
of an evaluation platform to continue improving the developed system. While
enhancing the search system, the developer will in fact need a tool to check the
e ectiveness of new versions. The methodology presented in this work can be a
valuable starting point for the development of an automated procedure.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>We thank the anonymous referees for their fruitful comments and careful reading,
which improved this article.</p>
      <sec id="sec-6-1">
        <title>6 http://terrier.org/</title>
      </sec>
      <sec id="sec-6-2">
        <title>7 https://www.copernic.com/</title>
      </sec>
      <sec id="sec-6-3">
        <title>8 https://msdn.microsoft.com/en-us/library/windows/desktop/ff628790.aspx</title>
      </sec>
      <sec id="sec-6-4">
        <title>9 http://matters.dei.unipd.it/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Feldman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The high cost of not nding information</article-title>
          .
          <source>IDC White Paper (April</source>
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Feldman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The information advantage: Information access in tomorrow's enterprise. IDC, Adapted from Hidden Costs of Information Work: A Progress Report by Susan Feldman, IDC 217936, and from Worldwide Search and Discovery Software 2009{2013 Forecast Update and 2008 Vendor Shares by Susan Feldman</article-title>
          , IDC
          <volume>219883</volume>
          (
          <year>October 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Grainger</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potter</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          : Solr in Action. Manning Publications, USA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Information Retrieval Evaluation (1st ed</article-title>
          .). Morgan and Claypool Publishers (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hawking</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Challenges in enterprise search</article-title>
          .
          <source>In: Proceedings of the Australasian Database Conference ADC2004</source>
          . pp.
          <volume>15</volume>
          {
          <issue>26</issue>
          (
          <year>January 2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hawking</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Enterprise search</article-title>
          . In: Baeza-Yates,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Ribeiro-Neto</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (eds.) Modern Information Retrieval, 2nd Ed., pp.
          <volume>641</volume>
          {
          <fpage>684</fpage>
          .
          <string-name>
            <surname>Pearson</surname>
            <given-names>Educational</given-names>
          </string-name>
          , UK (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mao</surname>
          </string-name>
          , J.:
          <article-title>Enterprise search: Tough stu</article-title>
          .
          <source>Queue</source>
          <volume>2</volume>
          (
          <issue>2</issue>
          ),
          <volume>36</volume>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rowlands</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hawking</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sankaranarayana</surname>
          </string-name>
          , R.:
          <article-title>Workload sampling for enterprise search evaluation</article-title>
          .
          <source>In: Proc. of the 30th Annual Int. ACM SIGIR Conf. on Research and Development in IR (SIGIR '07)</source>
          . pp.
          <volume>887</volume>
          {
          <fpage>888</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Tosato</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Exploiting ERP systems in enterprise search</article-title>
          .
          <source>In: Proc. of the 7th Italian Information Retrieval Workshop</source>
          , Venice, Italy (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Zilio</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Progettazione e realizzazione di un sistema di enterprise search</article-title>
          . Master thesis in computer engineering, University of Padua, Italy (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>