<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>In notitia i confide - Enterprise search and information quality</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mar n White</string-name>
          <email>martinswhite@outlook.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ISKO UK Conference</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Visiting Professor, Information School, University of Sheffield</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Enterprise search implementations first started to be undertaken in the early 1970s. Although there may well be around 100,000 research papers published on information retrieval there is only one that provides a detailed enterprise search case study. Related research indicates that content quality, along with technology limitations and a lack of training, contribute to a significant lack of satisfaction with the performance of enterprise search applications. This paper is based on experience gained by the author from enterprise search projects undertaken by the author between 2009 and 2019, highlighting the impact of information quality on enterprise search performance and satisfaction.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Enterprise Search</kwd>
        <kwd>Information Quality</kwd>
        <kwd>Information Management</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This paper considers the reasons for this lack of satisfaction, and the extent to which information
quality is a factor. It is based on the author’s personal experience with around 40 enterprise search
projects between 2009 and 2019.</p>
    </sec>
    <sec id="sec-2">
      <title>2. What do we mean by ‘search’?</title>
      <p>Given the title of this paper it is important to understand the diversity of ‘search’ applications and
processes. The recent arrival of ChatGPT has led to vendors and observers of search applications to
announce either that search is now dead or that AI will significantly improve the performance of search
applications. The use of ‘search’ in this context is similar to using ‘car’ to describe everything from a Fiat
500 up to a Maclaren supercar. In some situations (the design of car parks) this may be a valid use but not
in terms of fitness for purpose (taking the family on holiday).</p>
      <p>It is important to understand that there is a difference between ‘fitness to specification’ and ‘fitness to
purpose’ and any gap between them will almost inevitably result in a workaround by the employee to
enable employees to meet both their personal objectives and those of their organisation. There are eight
categories of ‘search’.</p>
      <p>
        1. Web search – Publicly available information with effort being paid to quality, metadata and links.
2. Website search - Publicly available information with effort being paid to quality, metadata and
links.
3. Intranet search - Highly curated web-server based enterprise-specific information searchable with
an internal search application. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
4. Academic search - Research services for academic users with highly curated content on
specialpurpose commercial and open-source applications. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
5. E-commerce search - Highly curated content accessed through a specialist website or a
vendorspecific search application. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
6. Professional search – Specialist collections of curated content for lawyers, clinicians, patent
agents and other professional groups who use search intensively. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
7. Systematic search - Highly curated content used by experienced researchers where a very high
degree of recall and reproducibility are important. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
8. Enterprise search - Structured and unstructured content, often in multiple languages and very little
of which is curated. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The characteristics of enterprise content</title>
      <p>
        Enterprise content is invariably written for a defined audience that the author is either familiar with
personally or has a good knowledge of the potential readership of the document. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] It will almost
certainly contain (for example) internal project names, trade and internal names for products,
alphanumeric product tags and short-hand expressions for offices (‘the team in Boston report…’).
      </p>
      <p>In many cases there will be no author identified, just that the document has been prepared by HR.
There will be multiple versions of similar documents which may vary in title and scope. Date tagging is
very important. Knowing the date of the most recent modification can be very misleading if a document
written in 2019 has now had a spelling correction made.</p>
      <p>
        An important issue for multinational organisations is that content can be in multiple languages.
Employees searching for information may be doing so in their second or even third language, which may
mean that they do not have an appreciation of synonyms that could be used to improve search quality and
that the content items retrieved could be in a language of which they have only a limited ability to
read.[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Content quality</title>
      <p>Just because a document has been deemed relevant by an algorithm does not mean to say that it is useful.
Content quality is a major issue in search implementation because it is often very obvious that the most
‘relevant’ references listed on the SERPs (Search Engine Results Pages) are widely different in terms of
quality and value to an employee.</p>
      <p>
        Research on the use of Enterprise Content Management applications (ECM) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] (which invariably have
good search functionality) strongly indicates that that the problems lie not with the functionality of the
ECM with regards to either adding content or finding it but with the quality of the content that is
retrieved. The research suggests that there are two aspects to enterprise content quality assessment.
To quote Laumerm Maier and Weitzel:
“The first is representational information quality. Our analysis of the interviews indicates that the format
of information is an important influencing factor for user satisfaction and a unique dimension in our
additional analysis. This dimension reflects the way information is presented to the user and subsumes
related characteristics of information including conciseness, presentation, and understandability.
Conciseness reflects the rigor and the sententiousness of information, presentation refers to the format and
the way information is designed to make it understandable to users, and understandability is the extent to
which information is clear, unambiguous, and easily comprehensible. All these characteristics have in
common that they focus on the way information is presented to the user and reflect the requirement that
information needs to be represented in an appropriate format that accentuates its meaning. They are
independent of the use of information in a specific context.
      </p>
      <p>The second dimension we identified in our interview analysis was contextual information quality, an
important influencing factor for user satisfaction. This dimension reflects the extent to which information
fits the needs of the task the information is used in. In our analysis, we identified completeness, relevance,
timeliness, and usability as information characteristics which we subsumed into the contextual
information quality dimension”.</p>
      <p>Some examples of poor information quality the author has encountered include:
•
•
•
•
•
•
•
•
•
missing versions of documents.
not being able to be sure that the version found is the current version.
no specified date of initial authorship.
authorship attributed to a department and not to an individual, making verification very difficult.
no context about why a document has been prepared and any restrictions on its scope.
references in a document to related documents but without the information needed to be able to
locate and obtain them.
no information about when a PowerPoint presentation was given and whether it has been
modified following presentation to correct errors
the scope of Excel spreadsheets and a lack of a ‘last updated’ comment
titles on documents that bear no relationship with the content, a particular issue with PowerPoint
presentations.</p>
      <p>The fundamental issue is that few organisations have developed a set of information quality policies and
even fewer have implemented an effective governance structure to achieve conformance.
As a result, employees have to place their trust (and reputation) in information that they cannot validate.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Enterprise searching – why and how?</title>
      <p>The image of the lone employee faced with a challenging problem and having to rely on a search
application to find an expert (as proposed by many search application vendors) is questionable.
Enterprises were full of supportive teams even before the advent of wide-scale remote working as a
reaction to the Covid pandemic. Moreover, employees are in receipt of data and information from many
database applications, email and social media. When confronted with the need to locate information in
order to make a decision, employees will invariably have a collection of information to hand but need to
verify, update and expand this collection of information. The result is that the search query is often along
the lines of ‘More like this’ and the employee already has a good vocabulary of query terms.</p>
      <p>That has implications for relevance, precision and performance metrics. The outcomes of a search
could include a significant number of relevant documents on the first two SERPs but these may well
duplicate information the employee has already acquired. Relevance and value are not synonyms. Efforts
by an enterprise search team to improve the click-through on the first two SERPS may do nothing more
than increase the effort involved in doing so with no visible benefit to the employee.</p>
      <p>For a significant number of enterprise queries a search application will return a substantial number of
results defined broadly by the business scopes of the enterprise. Research on professional searching
suggests that different communities of professionals make use of different aspects of the user interface to
filter the results and this pattern of use is similar across employees in the enterprise as each becomes a
‘professional searcher’ in their various roles, responsibilities and teams.</p>
      <p>
        This is a good point to consider the balance in enterprise search between precision and recall. Many
(too many!) search vendors claim that their applications deliver high precision results on the initial SERP.
In the enterprise environment there are many situations where a reasonable degree of recall is of value in
validating the initial query and perhaps an initial collection of documents. The employee may then wish
to either narrow the scope to improve precision or expand the scope to improve recall. This links into the
issue of stopping strategies, which is of significant importance in enterprise search but lies outside the
scope of this paper. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. The role of snippets</title>
      <p>In enterprise search situation relevance is obviously important but equally so is the quality of each result.
This cannot be directly assessed but it is likely that the user will be considering a range of clues from the
result snippet. These might include:
• the quality (especially clarity) of the title
• the name of the author
• their position in the organisation
• the department for which they were working when they authored the document
• the origination date of the document
• the file format where it might have an impact on accessibility
• the language
• the size of the document and therefore the challenge of locating the position in the document of the
information satisfying the query terms.</p>
      <p>In effect the employee is seeing the extent to which there is an audit trail that enables them to check on
the quality of the information they have found. This is a contributory factor to the very high level of
search queries for people in the organisation. Some might be to find an ‘expert’ but many more will be to
check out the authority of an unknown employee as a means of assessing the veracity, value and quality
of the information.</p>
      <p>
        Although there has been research undertaken into the value of various snippet formats [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] none of
this research has been on enterprise search use.
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Search dissatisfaction</title>
      <p>
        Despite over almost sixty years of enterprise search deployment a number of surveys conducted over the
last two decades indicate that perhaps in only 1 out of 5 organisations are employees very satisfied with
the performance of enterprise search. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
      </p>
      <p>Cleverley and Burnett [16] indicate that issues around technical performance, content quality and
training are the root causes of search dissatisfaction.</p>
      <p>A fundamental problem is that enterprise search is implemented on the basis that it can be used (in
principle) by any employee without the need for training and support. All too often the assumption of
both senior IT and business managers is that search is intuitive. Research indicates that search training
makes a substantial difference to employee search performance. [17]</p>
      <p>The reality is that very few employees have the expertise and experience to construct effective queries
and to assess the value of the results that are delivered. Each employee has their own domain knowledge
and expectations and has multiple information seeking options of which search is just one of many.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Impact of AI</title>
      <p>Despite the claims that Large Language Model (LLM)-based applications mark the end of ‘search’, no
result has yet been carried out on information discovery in the enterprise. Much attention has been paid to
the use of generative artificial intelligence (AIGC) in the form of summaries of documents and machine
translations, and what I refer to as ‘faux-search’ when a short summary is given in response to a
promptbased query. [18]</p>
      <p>With (at this stage) very little focus on the use of private LLMs in the enterprise, it is difficult to do
more than highlight the extent to which an employee is going to be able to validate the content of any
AIGC outputs. This challenge will be even greater when the multiple languages prevalent in an enterprise
are taken into account in the design of training sets and the modification of these to reflect changes in the
scope of the enterprise.</p>
      <p>It is also important to accept that just changing the search technology is not going to make any
significant improvement to employee satisfaction with enterprise search.</p>
    </sec>
    <sec id="sec-9">
      <title>9. Five steps to achieving enterprise search satisfaction</title>
      <p>There are four steps that need to be taken to ensure that employees can use enterprise search to find
information of the highest quality in order to make decisions of the highest quality.</p>
      <p>1. Adopt an information management strategy and related policies within a pragmatic governance
structure.
2. Integrate the information management strategy with the AI strategy.
3. Select search technology software on the basis of both functional and non-functional
requirements.
4. Identify content categories where the risk to the organisation from consistently poor information
quality puts the organisation at risk and take remedial action that is then carried forward as
examples of good practice.
5. Develop training and mentoring schemes for all employees (but especially for newcomers) that is
specific to the technology and the use cases of the content.
10.
[16] P.H. Cleverley, S. Burnett, L. Muir, Exploratory information searching in the enterprise: A study of
user satisfaction and task performance, Journal of the Association for Information Science and
Technology, 68(1) (2017) 77-96.
[17] Y.-L Lee, E.A. Chu, Chu, S. K.-W., Lee, M. M.-L. Chiu, R. C. H. Chan, Scaffolding in information
search: Effects on less experienced searchers, Journal of Librarianship and Information
Science, 48(2) (2015) 177-190.
[19] N. F. Liu, T. Zhang, P. Liang, Evaluating verifiability in generative search engines, 2023, arXiv
preprint arXiv:2304.09848. URL: https://doi.org/10.48550/arXiv.2304.09848</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <source>A History of Enterprise Search</source>
          <year>1938</year>
          -2022, University of Sheffield,
          <year>2022</year>
          . URL: https://sheffield.pressbooks.pub/eshistory1/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>IR</given-names>
            <surname>Anthology</surname>
          </string-name>
          , URL: https://ir.webis.de/anthology/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Russell-Rose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <article-title>Information retrieval in the workplace: A comparison of professional search practices</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>54</volume>
          (
          <issue>6</issue>
          ) (
          <year>2018</year>
          )
          <fpage>1042</fpage>
          -
          <lpage>1057</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lykke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bygholm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.B.</given-names>
            <surname>Søndergaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Byström</surname>
          </string-name>
          ,
          <article-title>The role of historical and contextual knowledge in enterprise search</article-title>
          ,
          <source>Journal of Documentation</source>
          , (
          <year>2022</year>
          )
          <volume>78</volume>
          (
          <issue>5</issue>
          ),
          <fpage>1053</fpage>
          -
          <lpage>1074</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <source>Achieving Enterprise Search Satisfaction</source>
          ,
          <year>2023</year>
          . URL : https://searchresearch.online/wpcontent/uploads/2023/01/Achieving-enterprise
          <article-title>-search-satisfaction</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kayley</surname>
          </string-name>
          , Intranet Search Essentials, Nielsen Norman Group,
          <year>2022</year>
          . URL: https://www.nngroup.com/articles/intranet-search/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O.</given-names>
            <surname>Hoeber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Storie</surname>
          </string-name>
          ,
          <article-title>A study of academic search scenarios and information seeking behaviour</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Human Information Interaction and Retrieval</source>
          , pp.
          <fpage>231</fpage>
          -
          <lpage>235</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tsagkias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.H</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kallumadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Murdock</surname>
          </string-name>
          , M.de Rijke,
          <article-title>Challenges and research opportunities in ecommerce search and recommendations</article-title>
          .
          <source>ACM SIGIR Forum</source>
          <volume>54</volume>
          (
          <issue>1</issue>
          ), (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Levay</surname>
          </string-name>
          , J. Craven (Eds)
          <article-title>Systematic Searching</article-title>
          .
          <source>Facet Publishing</source>
          <year>2019</year>
          ,
          <source>ISBN 978-1-78330-374-8</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.H.</given-names>
            <surname>Cleverley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Burnett</surname>
          </string-name>
          ,
          <article-title>Enterprise search: A state of the art</article-title>
          ,
          <source>Business Information Review</source>
          ,
          <volume>36</volume>
          (
          <issue>2</issue>
          ) (
          <year>2019</year>
          )
          <fpage>60</fpage>
          -
          <lpage>69</lpage>
          . URL: https://doi.org/10.1177/0266382119851880
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Mathieu</surname>
          </string-name>
          ,
          <article-title>Defining knowledge workers' creation, description, and storage practices as impact on enterprise content management strategy</article-title>
          ,
          <source>Journal of the Association for Information Science and Technology</source>
          ,
          <volume>73</volume>
          (
          <issue>3</issue>
          ) (
          <year>2022</year>
          )
          <fpage>472</fpage>
          -
          <lpage>484</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Harvey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brazier</surname>
          </string-name>
          ,
          <article-title>E-government information search by English-as-a Second Language speakers: The effects of language proficiency and document reading level</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          ,
          <volume>59</volume>
          (
          <issue>4</issue>
          ) (
          <year>2022</year>
          )
          <fpage>102985</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Laumer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Maier</surname>
          </string-name>
          , T. Weitzel,
          <article-title>Information quality, user satisfaction, and the manifestation of workarounds: A qualitative and quantitative study of enterprise content management system users</article-title>
          .
          <source>European Journal of Information Systems</source>
          <volume>26</volume>
          (
          <year>2017</year>
          )
          <fpage>333</fpage>
          -
          <lpage>360</lpage>
          . URL: https://www.tandfonline.com/doi/full/10.1057/s41303-016-0029-7
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>D.M. Maxwell</surname>
          </string-name>
          ,
          <article-title>Modelling search and stopping in interactive information retrieval, Doctoral dissertation</article-title>
          , University of Glasgow,
          <year>2019</year>
          . URL: https://theses.gla.ac.uk/41132/
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zimmerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Elsweiler</surname>
          </string-name>
          ,
          <article-title>Featured Snippets and their Influence on Users' Credibility Judgements</article-title>
          , in: CHIIR '
          <fpage>22</fpage>
          ,
          <string-name>
            <surname>Proceedings</surname>
            <given-names>of</given-names>
          </string-name>
          <source>the 2022 conference on Human Information Interaction and Retrieval, March</source>
          <volume>14</volume>
          -18, Regensburg, Germany,
          <year>2022</year>
          . URL: https://doi.org/10.1145/3498366.3505766
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>