<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information Reputation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Peter Davis</string-name>
          <email>peter.davis@neustar.biz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salman Haq</string-name>
          <email>salman.haq@neustar.biz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Distinguished Engineer</institution>
          ,
          <addr-line>Neustar, 21575 Ridgetop Circle, Sterling, Virginia</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Research Architect</institution>
          ,
          <addr-line>Neustar, 21575 Ridgetop Circle, Sterling, VA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>5</lpage>
      <abstract>
        <p>In this paper we describe the design of a reputation framework for an information management system under active development. The integration of a reputation framework with an IMS is a novel combination that can produce a distinctly more e↵ective business intelligence tool. claim One or more assertions made of a datum. data source A computer system that stores data such as a database or file system.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Neustar is a data analytics and intelligence services company that operates
several large database systems. To eciently manage these numerous, disparate
systems, we are developing an Information Management System (IMS) that maps
technical data models using a standard set of ontologies. The IMS is an online
community for employees where they can share, classify and discover metadata
about various Neustar data sources. Its main purpose is to assist users in
achieving two main objectives: a) reducing costs by utilizing existing information and
b) increasing revenues by creating new information [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>With these objectives in mind, users must have the ability to make value
judgments about data sources relative to one another. Such data sources may
number in the hundreds and the datum contained therein may number in the tens
of thousands. The majority of these entities will lack significant value for data
science, and those that are valuable will risk being lost in a deluge of information.
Therefore it is imperative that the system establish a bias towards meaningful
datum by highlighting interestingness. A well-crafted reputation framework can
excel at doing exactly this.
Glossary
data steward A individual or group of individuals holding domain-specific
knowledge of an information system.
datum An instance of metadata mapped to an atomic data field. This includes,
for example, columns in a relational database or entities defined in an XML
schema.
interestingness A scalar value indicating the suitability for inclusion in further
analysis.
reputation A qualitative measure that informs a value judgment about a datum
or user.</p>
    </sec>
    <sec id="sec-2">
      <title>Acronyms</title>
      <p>IMS Information Management System.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Framework Description</title>
      <p>The framework is comprised of several reputation models, each of which
computes one or more scores for a resource type. A fixed set of claims serve as
inputs to each model which assigns numerical values to them and passes them
through a series of mathematical filter processes. Models are distinguished by
their input selection, process configuration, and output scores. The IMS utilizes
a fixed ontology to define claims that include appropriate business and technical
classifications for data within the subject systems. The essential claims of the
datum model are summarized in Table 1. The IMS also incorporates techniques
to simplify crowd sourcing the classification of datum by data steward s.
However, we have concluded from early usage, that a simple classification process
is insucient. Classifications can be subjective, and classification sparseness
results in under utilization of the system. As a consequence, methods to encourage
accurate and complete classification will be implemented to enrich the overall
ecacy of the system.
Name Description
classified Data steward classified datum from the business domain ontology
described Data steward entered a description
discussed User participated in a discussion topic about datum
emailed User emailed the link to the datum page to another user
flagged User informed the data steward about insucient or inaccurate details
watched User will be notified of future updates by other users
wanted User requested access to the datum from the data steward</p>
    </sec>
    <sec id="sec-4">
      <title>Interestingness Reputation Model for Datum</title>
      <p>In the IMS, datum is an atomic unit of data. Its’ classification results in queriable
metadata, and can relate to a column in a relational database or an element,
attribute or phrase in a document. At the time of writing, the system had over
7,000 fields from merely four data sources. Even at this watermark, the task of
finding interesting datum is impractical for any user community. As more data
sources are imported into the system, this task will be become impossible even
if the datum population grows sub-linearly. Therefore it is imperative that the
system is capable of identifying and highlighting interesting datum to facilitate
user objectives.</p>
      <p>
        In Figure 1 we describe the simplified model for calculating the
interestingness reputation score for datum. Our approach is informed by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] which applies
a similar methodology for surfacing interesting media objects. The score is an
indicator of the likelihood that a particular datum has potential value. User
interactions with the IMS are interpreted as claims from Table 1. The figure shows
claims as they are consumed by various processes. The intermediate processes
(boxes 7, 8, 9, and 10) compute normalized counts of the claim interactions.
These counts are fed to the terminal process, InterestingnessCustomMixer (box
12) which scales and reduces the values into the scalar interestingness score. This
score can be used as a predictor for search and recommendation systems.
Omitted from this simplified model are lag and decay filters necessary to counteract
volatility and freshness bias respectively [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion And Future Work</title>
      <p>
        We have described a realistic blue print for a reputation system that is on the
roadmap of our IMS. Once implemented we think that it will dramatically
improve the quality of information that is retrievable by users, thus increasing its’
e↵ectiveness as a platform for information management and data science. We
have left outcome analysis of the approach and results for a future paper. Also
on the roadmap is a meaningful gamification system inspired by [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to further
enhance user engagement.
4
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Butterfield, Daniel, et al.
          <article-title>Interestingness Ranking of Media Objects</article-title>
          . Yahoo! Inc., assignee.
          <source>Patent US20060242139A1. 8 Feb</source>
          .
          <year>2006</year>
          . Print.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Farmer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Randall</surname>
            ., and
            <given-names>Bryce</given-names>
          </string-name>
          <string-name>
            <surname>Glass</surname>
          </string-name>
          .
          <source>Building Web Reputation Systems</source>
          . Sebastopol, CA:
          <string-name>
            <given-names>O</given-names>
            <surname>'Reilly</surname>
          </string-name>
          ,
          <year>2010</year>
          . Print.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Sveiby</surname>
            ,
            <given-names>K.E.</given-names>
          </string-name>
          <string-name>
            <surname>Knowledge</surname>
          </string-name>
          <article-title>Management: Lessons from the Pioneers</article-title>
          .
          <source>Tech. Sveiby Knowledge Associates</source>
          ,
          <year>2001</year>
          . Web.
          <volume>12</volume>
          Mar.
          <year>2013</year>
          . http://www.sveiby.com/articles/KM-lessons.doc
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Nicholson</surname>
            ,
            <given-names>Scott.</given-names>
          </string-name>
          <article-title>Strategies for Meaningful Gamification: Concepts behind Transformative Play and Participatory Museums</article-title>
          .
          <source>Tech. Meaningful Play 2012 Web. 13 Mar</source>
          .
          <year>2013</year>
          . http://scottnicholson.com/pubs/meaningfulstrategies.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>