<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Applied Distributed Information Retrieval in Enterprise Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erwin Gunadi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Till Plumbaum</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sahin Albayrak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Technische Universität Berlin 10587 Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Distributed enterprise search as a special case of Distributed Information Retrieval (DIR) is characterized by the need to query multiple repositories in enterprise environments. However, in DIR research there is a lack of publicly available real-world datasets for evaluation purposes. As a result, there is a gap between insights gained from simulated environments and real-world investigations on distributed enterprise search. In this paper, we outline three fundamental issues based on our investigations of a large real-world distributed enterprise search system. We found that (1) the utilization of security features in enterprise repositories, (2) the adaptation of resource description and resource selection for enterprise model, and (3) repository grouping are fundamental real-world issues. We hypothesize that a better understanding of these issues will contribute to improve the distributed enterprise search and better support complex search tasks in enterprise environments. Based on our experience gained from a real-world system, we also outline needed steps to cope with these issues.</p>
      </abstract>
      <kwd-group>
        <kwd>distributed information retrieval</kwd>
        <kwd>enterprise search</kwd>
        <kwd>result aggregation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3.4 [Systems and Software]: Distributed Systems</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>Enterprise Search is an area of information retrieval which
speci cally addresses the information needs of enterprise
users. Enterprise is de ned as an organizational entity with
an exclusive memberships of its users. A typical enterprise
environment consists of multiple layers of access rights and
di erent dedicated data repositories such as web servers, le
servers, wikis, etc. Key di erences to Web search are as
follows: (1) Heterogeneous document types such as web pages:
Copyright c 2015 for the individual papers by the papers’ authors.
Copying permitted for private and academic purposes. This volume is published
and copyrighted by its editors.</p>
      <p>
        ECIR Supporting Complex Search Task Workshop ’15 Vienna, Austria
Published on CEUR-WS: http://ceur-ws.org/Vol-1338/.
wiki, pdfs, emails, word documents etc. (2) Multiple
document repositories: Documents are normally not held in a
single le server or system. (3) Access restriction: hierarchies
and roles rules for every document, and (4) Managed data
generation process: Di ering from web documents, which
are created by individual entity, each enterprise de nes their
own document creation and update policy, which e ectively
valid for all of its members [
        <xref ref-type="bibr" rid="ref13 ref5 ref9">9, 5, 13</xref>
        ].
      </p>
      <p>
        To address the above mentioned issues (1)-(4), the use
of Distributed Information Retrieval (DIR) has been
proposed [
        <xref ref-type="bibr" rid="ref13 ref4 ref5">4, 5, 13</xref>
        ]. DIR is a concept of managing di erent
resources through a broker. A broker mediates between users
and di erent sources, or repositories, to collect and combine
search results. To accomplish this task three sub-problems
need to be addressed: resource description, resource
selection and result merging [
        <xref ref-type="bibr" rid="ref11 ref2">2, 11</xref>
        ]. With exception of the TREC
Federated Web Search track 1 most of the DIR research have
been based on synthetic test collections [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Due to the lack
of appropriate datasets and the proprietary nature of
enterprise data many improvements achieved in DIR are not
directly suitable for distributed enterprise search [
        <xref ref-type="bibr" rid="ref5 ref6">6, 5</xref>
        ]. These
factors prevent a further adoption of DIR improvements in
distributed enterprise search and creates a gap between these
two areas.
      </p>
      <p>This paper has two main contributions. First, we bridge
the gap between DIR and the enterprise context, by
outlining three fundamental problems that have been widely
ignored in the eld of DIR research, but are mandatory to
be considered in the enterprise context. Second, we propose
ways to cope with these problems. The presented problems
are (1) Utilization of security features in enterprise
repositories, (2) Adaptation of resource description and resource
selection for enterprise model, and (3) Repositories grouping.
All our ndings base on real-world experiences and insights
we gain from the operation of a distributed enterprise search
system at TU Berlin and the city's administration of Berlin
with about 50.000 employees.</p>
      <p>In the next Section we outline related works which
discuss the integration of DIR in enterprise environment. In
Section 3 we describe the architecture of our agent-based
distributed enterprise search platform and its deployment.
Section 4 details the open problems arise out from our
experience. In Section 5 we conclude our paper.
2.</p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORKS</title>
      <p>
        Various works have proposed DIR as a possible paradigm
to implement enterprise search [
        <xref ref-type="bibr" rid="ref12 ref13 ref5">5, 13, 12</xref>
        ]. The main concept
1https://sites.google.com/site/trecfedweb/
of DIR is the usage of multiple resources, or repositories,
using a broker-concept in order to satisfy a users' information
need. This concept ts the main characteristic of an
enterprise environment where normally multiple data repositories
for di erent needs exists, such as web servers, le servers,
wikis, etc. [
        <xref ref-type="bibr" rid="ref13 ref5 ref9">9, 5, 13</xref>
        ].
      </p>
      <p>
        Works investigating how to secure enterprise search
system are presented in [
        <xref ref-type="bibr" rid="ref1 ref14">1, 14</xref>
        ]. Bailey et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] propose the
implementation of document level security for enterprise search
systems.They evaluated how on-the- y security checks for
each document during search result list building a ect the
search processing time. Zhou et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] propose the usage
of ontology-based user pro les to secure the search process.
The ontology models the information search service and user
role information, which can be maintained for di erent
departments in an enterprise. It is still an open question about
how security restrictions a ects search result quality in a
distributed enterprise search. Current work focuses more on
performance issues than quality.
      </p>
      <p>
        Regarding the DIR algorithms Li et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] evaluated the
performance of various result merging algorithms on
multiple enterprise repositories which are unique to each other.
Li et al. argue that repositories in an enterprise context are
not identical to each other and each of these repositories
may have di erent size, document types, intended audience
and administration control. These characteristics need to be
explored in the context of DIR in enterprise search.
      </p>
    </sec>
    <sec id="sec-4">
      <title>DISTRIBUTED ENTERPRISE SEARCH</title>
    </sec>
    <sec id="sec-5">
      <title>IN THE REAL-WORLD</title>
      <p>
        The ndings we present in this paper based on research
cooperation with the service provider of the administration of
Berlin, where we have deployed an agent-based distributed
enterprise search system [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The system is currently used
as the standard search platform for about 50.000 employees.
The structure of Berlins' network confronted us with some
challenges. Even though the whole city can be regarded as
a closed organization with state o cials as its employees,
each of the city's districts maintains its own data
management policy. This means a city district is a private network
with access to all main data repositories such as city's own
intranet, but without access rights to data repositories from
other city districts. As opposed to the classic DIR setting
these sub-networks cannot be served by a single main
broker. It requires that each network area has its' own private
broker and an extra broker installed in the main intranet.
Another use case for an additional broker is the user
desktop. To comply with the user privacy policy from the city's
administration, local desktop les should not be externally
accessible. Because of such restrictions we have build a
local broker so that users can nd their local desktop les.
The interaction between the users, the search client and the
multiple brokers is illustrated in Figure 1.
      </p>
      <p>Figure 1 illustrates how the search client is used to
contact all of the di erent brokers. It enables users to search
in di erent network areas. In each of the network,
multiple repositories are queried by the responsible broker. The
desktop from the user also has a dedicated broker because
local les should not be queried by an external broker.</p>
    </sec>
    <sec id="sec-6">
      <title>OPEN PROBLEMS</title>
      <p>In this section we describe following open problems
derived from our investigations on the distributed enterprise
search system for the administration of Berlin. We also
suggest how these open problem can be further investigated in
DIR research in enterprise context.
4.1</p>
    </sec>
    <sec id="sec-7">
      <title>Utilization of security features in enterprise repositories</title>
      <p>
        Until recently the research of DIR is based on the
assumption that all documents are accessible to all users [
        <xref ref-type="bibr" rid="ref10 ref11">11,
10</xref>
        ]. However, in enterprise environments access to the
documents is secured [
        <xref ref-type="bibr" rid="ref1 ref13 ref5">1, 5, 13</xref>
        ]. Depending on the access right
each user may see di erent search results for the same search
query. We argue that this property can be utilized as an
essential feature for improving distributed enterprise search.
Utilization means not only to comply with access
restrictions, but also to improve resource selection and result
merging algorithm by using the security information.
      </p>
      <p>In order to accomplish this, DIR algorithms for di erent
steps have to integrate the security information as a feature.
For example, in DIR literature resource selection is
responsible to select the resources with the most relevant documents.
However, when a relevant repository is restricted for
particular user or user group, these documents will not be shown.
This can be mitigated when repositories, with more
accessible documents, are higher prioritized even though they have
less relevancy. In this case in navigational search tasks
employees may get a better recall about the subject. To the
best of our knowledge such behavior in a DIR setting is not
yet researched. Creating suitable test collections for this
purpose is needed to evaluate this essential feature.
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>Adaptation of resource description and resource selection for enterprise model</title>
      <p>
        Li et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] highlighted how repositories in an enterprise
environment unique to each other are. Based on this fact Li
et al. evaluate how these uniqueness may in uence the result
merging performance. The concern about the exploitation
of unique features found in enterprise repositories should
also be considered in the other sub-tasks of DIR: resource
description and resource selection. Research in this area is
needed in order to improve the application of DIR in
Enterprise Search. Thus, it will help narrow the gap between
these two research elds.
      </p>
      <p>As an example use case of such exploitation, in an
enterprise environment a common repository type is le servers.
Documents from le servers, which are relevant for search
queries, are stored in a hierarchy of directories. We can
include these directory names as a part of the resource
description of a repository. This means content of a resource
description includes not only sampled documents but also
directory names. These directory information can be used
as important terms that improve the resource selection
algorithm. Such exploitation is yet to be investigated in
distributed enterprise search context.
4.3</p>
    </sec>
    <sec id="sec-9">
      <title>Repositories grouping</title>
      <p>
        One of the challenges from our enterprise environment
setting is the need of multiple brokers to handle di erent
networks. Even though the multiple brokers can be seen as a
technical feature, it introduces a new perspective in handling
multiple repositories. The ability of repositories grouping
open the possibility to boost repositories with similar types
as a group. The application of boosting for theme speci c
repositories is being actively researched in distributed web
search task context [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this paper, the authors
investigated how news sources can be ranked and placed in the
web search result. In our deployment scenario we found
that search queries about speci c law and regulations are
common. In this case, having a theme speci c group of
repositories means documents comes from the law-themed
repositories receive higher rank than documents from non
law-themed ones.
      </p>
      <p>By having groups of repositories, result merging
techniques may rank not only based on the repository rank but
also on repositories-group rank. To accomplish this the
subtasks of DIR must be adapted to broker context, namely
broker description and broker selection. When a particular
group of repositories is highly relevant for a search query,
the gained broker ranking may be used to highlight a group
of repository in the search result page.</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION</title>
      <p>
        Due to the availability of heterogeneous repositories,
previous works have proposed the application of DIR in
enterprise search [
        <xref ref-type="bibr" rid="ref11 ref13 ref5">5, 11, 13</xref>
        ]. However, the improvements achieved
in DIR research are rarely investigated in real-world
scenarios, especially in enterprise environments [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Recent works
show that further research in the application of DIR in
realworld settings is needed in order to close the gap between
DIR research and its' real-world application in enterprises.
In this paper, we introduced three issues based on our
experience from a real distributed enterprise setting.
      </p>
      <p>
        For future works, we need to investigate how security
information can be utilized for the di erent DIR sub-tasks.
This also applies for the unique features found in enterprise
repositories, such as le paths and directories. It permits
us to better understand the application of DIR in enterprise
search, thus, bridges the gap between these two research
areas. Also more e ort building appropriate test collections
for a real-world DIR use case, like the work from Nguyen et
al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], has to be done for evaluating distributed enterprise
search. This helps to adapt available techniques for di erent
tasks in DIR (resource description, resource selection and
result merging) for enterprise use cases. Another investigation
is needed on how to accommodate repositories grouping by
building multiple brokers. Similar to news integration in web
search [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], group of repositories can be di erently ranked for
incoming queries. The ranking result may then be used to
highlight a group of repository in the presentation of search
result page (SERP) in enterprise search.
6.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawking</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Matson</surname>
          </string-name>
          .
          <article-title>Secure search in enterprise webs: tradeo s in e cient implementation for document level security</article-title>
          .
          <source>CIKM '06 Proceedings of the 15th ACM international conference on Information and knowledge management</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Callan</surname>
          </string-name>
          .
          <article-title>Distributed information retrieval</article-title>
          .
          <source>In Advances in Information Retrieval</source>
          , pages
          <volume>127</volume>
          {
          <fpage>150</fpage>
          . Kluwer Academic Publishers,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Gunadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Meder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Plumbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Scheel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Albayrak</surname>
          </string-name>
          .
          <article-title>Distributed enterprise search using software agents</article-title>
          .
          <source>In Proceeding AAMAS '14</source>
          , pages
          <fpage>1623</fpage>
          {
          <fpage>1624</fpage>
          , Paris, France,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawking</surname>
          </string-name>
          .
          <article-title>Challenges in enterprise search</article-title>
          .
          <source>In ADC '04 Proceedings of the 15th Australasian database conference</source>
          , volume
          <volume>27</volume>
          , pages
          <fpage>15</fpage>
          {
          <fpage>24</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawking</surname>
          </string-name>
          . Enterprise Search. In R. Baeza-Yates and
          <string-name>
            <given-names>B.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          -Neto, editors,
          <source>Modern Information Retrieval</source>
          , pages
          <volume>641</volume>
          {
          <fpage>684</fpage>
          .
          <string-name>
            <surname>Addison-Wesley</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Jie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lamkhede</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sapra</surname>
          </string-name>
          , E. Hsu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Song</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <article-title>A uni ed search federation system based on online user feedback</article-title>
          .
          <source>In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining - KDD '13</source>
          ,
          <string-name>
            <surname>page</surname>
            <given-names>1195</given-names>
          </string-name>
          , New York, New York, USA,
          <year>2013</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P. V.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Thomas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Hawking</surname>
          </string-name>
          .
          <article-title>Merging algorithms for enterprise search</article-title>
          .
          <source>Proceedings of the 18th Australasian Document Computing Symposium on - ADCS '13</source>
          , pages
          <fpage>42</fpage>
          {
          <fpage>49</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>McCreadie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Ounis.</surname>
          </string-name>
          <article-title>News vertical search: when and what to display to users</article-title>
          .
          <source>In SIGIR '13: 36th international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>253</volume>
          {
          <fpage>262</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Mao</surname>
          </string-name>
          .
          <article-title>Enterprise search: Tough stu</article-title>
          .
          <source>Queue</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <volume>36</volume>
          {
          <issue>46</issue>
          , 4
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Trieschnigg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>Federated search in the wild: the combined power of over a hundred search engines</article-title>
          .
          <source>In Proceedings of the 21st ACM international conference on Information and knowledge management</source>
          , pages
          <year>1874</year>
          {
          <year>1878</year>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shokouhi</surname>
          </string-name>
          and
          <string-name>
            <given-names>L. Si. Federated</given-names>
            <surname>Search</surname>
          </string-name>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):1{
          <fpage>102</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Thomas</surname>
          </string-name>
          .
          <article-title>To what problem is distributed information retrieval the solution</article-title>
          ?
          <source>Journal of the American Society for Information Science and Technology</source>
          ,
          <volume>63</volume>
          (
          <issue>7</issue>
          ):
          <volume>1471</volume>
          {
          <fpage>1476</fpage>
          ,
          <year>July 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M. White. Enterprise</given-names>
            <surname>Search. O'Reilly Media</surname>
          </string-name>
          , Inc.,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhou</surname>
          </string-name>
          <article-title>. Multi-agent based distributed secure information retrieval</article-title>
          .
          <source>In CMC'10</source>
          , pages
          <fpage>76</fpage>
          {
          <fpage>79</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>