<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Author pro le pages in zbMATH { improving accuracy through user interaction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Helena Mihaljevic-Brandt</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabian Muller</string-name>
          <email>fabian@zentralblatt-math.org</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>y Nicolas Royz</string-name>
          <email>roy@zentralblatt-math.org</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <volume>13</volume>
      <issue>2014</issue>
      <abstract>
        <p>A solid and distinctive online record of a scientist's research achievements is nowadays a must, in particular for a researcher at an early career stage looking for a better university position, applying for a grant or a promotion, or trying to establish new contacts and build partnerships and projects. In the global world of science there are already numerous services o ering research pro le sites, such as Google Scholar, Scopus or ORCID. In the eld of mathematics, however, the community still relies a lot on the traditional services MathSciNet and zbMATH, which stand for a high level of quality control regarding data inclusion and assignment. However, the enormously growing number of mathematical publications and new mathematicians begs the question of how to design author pro le pages for mathematicians which t the new trends of communication in the era of social networks, while keeping (or better increasing) the existing level of reliability. In our contribution we plan to present some ideas from the development team of zbMATH addressing this topic.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>A highly visible online presence of a scientist's research activities has become almost
inevitable to advance one's own academic career, even for a eld like mathematics with
communities of manageable size and a rather traditional communication structure. For
researchers at an early career stage, a solid record of scholarly contributions is
undoubtedly essential for getting a new job, a promotion or the funding for a new project. A
distinguishable pro le increases the chance of correct attribution, recognition and
citation of a scientist's impact and provides the possibility to nd others working on a
similar topic, establish new contacts, and build partnerships and projects.</p>
      <p>It is therefore not surprising that various online providers o er platforms for
researchers with the goal to increase the visibility of their scienti c activities. Many search
engines have recognized the potential of such services, Google Scholar being probably
the most successful example. Also the social media hype of making a pro le of oneself
publicly available has been brought into the scienti c world with enormous success as
shown, e.g., by ResearchGate. Last but not least, seeing the need for a global and
sustainable authorship administration, the initiative ORCID1 has the promising potential
of becoming the standard author identi er.</p>
      <p>In the eld of mathematics, there are two large traditional indexing and reviewing
services, namely MathSciNet and zbMATH, as well as some regional bibliographic services
such as Math-Net.ru for Russian literature, which o er pro les for the authors indexed
in their databases.</p>
      <p>Author pro les in zbMATH
zbMATH, known to many under its former name Zentralblatt MATH, provides access to
bibliographic data and content information of publications from all areas of mathematics
and its applications. Currently, about 3.4 millions publications are indexed, going back
to 1755 and with an annual growth of 120;000 new records.</p>
      <p>zbMATH indexes only peer-reviewed literature. All items undergo an internal
editorial process, and many of them, in particular in the core elds of mathematics, are
reviewed additionally by an independent expert in the corresponding eld. This editorial
procedure ensures a high integrity level of the mathematical content, helping, e.g., to
detect male cent behaviour such as \plagiarism" or \non-sense" papers2.</p>
      <p>The authors of a new publication record in zbMATH are either identi ed with existing
pro les, or new pro les are being created. The rst step of this procedure is solved
algorithmically by analysing the name string, but also additional information such as the
co-authors. The post-processing involves also manual correction which is often initialized
through user requests. The zbMATH author identi cation team receives such requests
on a daily basis, showing an increasing interest of the mathematical community in the
availability of reliable author pro les.</p>
      <p>The quality control of the indexed documents and of the authorship identi cation
explains why zbMATH (and similar services) enjoy high con dence of the mathematical
community, despite the numerous alternative services mentioned before. Furthermore,
since zbMATH aims at a complete coverage of all literature, and hence all authors, in
mathematics, this is the right place to obtain an overview of the entire publishing activity
in this eld, where all mathematicians since 1868 are visible and not only those able or
1The Open Researcher and Contributor ID (ORCID) is a recent non-pro t initiative to assign unique
identi ers to authors of research publications (http://orcid.org).
2For an enlightening review in zbMATH see, e.g., https://zbmath.org/?q=an:1202.51019
willing to use \social-like networks". The alternative services are not able to o er such
a broad coverage for the eld of mathematics, and presumably this won't change in the
near future. Even ORCID, the service with the most promising potential to become the
standard author identi er, has yet only 146;064 non-empty pro les across all scienti c
elds 3, among which only a few thousand are actually mathematicians 4.</p>
      <sec id="sec-1-1">
        <title>Why is author disambiguation di cult?</title>
        <p>The analysis of authorship has to be achieved with an enormously growing number of
publications and people: currently, there are more than 3:4 millions publications
corresponding to 5:5 millions authorships in zbMATH, for an estimated number of 860;000
3This corresponds to the o cial ORCID statistics from June 6 2014 (https://orcid.org/
statistics).
4We found only 3;000 mathematicians in ORCID by matching DOIs from the publications available
in the ORCID Public Data File from October 2013 with zbMATH.
author identities. The trend of maintaining pro le pages and registering at various
scienti c networking platforms (ResearchGate, ORCID, Google Scholar, . . . ) is quite recent,
and still not intensively picked up by mathematicians. Although the identi ers or links
to these services would obviously be a great help for the identi cation of authorship,
they are at the moment almost never available in the publications. Information like
a liation, email, etc. is a rather standard part of a publication's metadata, but these
tend to change, leaving practically no other information but the name for the author
identi cation5</p>
        <p>The task of attributing an exact set of publications to a certain author based only on
person name is far from trivial, due to several reasons:</p>
        <p>Incompleteness: The available data may be incomplete; some parts of the name
may be missing (e.g., middle name, second familiy name) or abreviated (use of
initials). In earlier years, it was also not uncommon to publish under the family
name only.</p>
        <p>Synonyms: For a single author, one can face a great variability in names, in
particular due to di erent spellings and transliterations, but also through name
changes such as after marriage.</p>
        <p>Homonyms: The same name may refer to multiple individuals. This is
particularly acute in the case of certain Eastern Asian names; for example, in China
the top three surnames (Wang, Li and Zhang) account for more than 20% of the
population.</p>
        <p>
          This explains why author name disambiguation is a longstanding research topic with
high relevance for bibliometric studies and publication retrieval, and why the general
eld named entity recognition has gained a lot of interest recently (see, e.g., [
          <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
          ]).
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2 Community input</title>
      <p>As outlined above, the problem of author disambiguation is too di cult to be tackled
completely algorithmically. On the other hand, the amount of publications incorporated
every year into zbMATH and other information services is too big to solve this task
manually. We thus decided to open up this process at least partially to community
input.</p>
      <p>
        Experience shows that mathematicians take an active interest in grooming and
optimizing their research pro les, both the ones they create for themselves (ORCID,
ResearchGate, Google Scholar, ResearcherID) and those that were initiated automatically
(Scopus Author Identi er)6. This correlates with results from the recent study [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
indicating that mathematicians like to be empowered when using mathematical search
5zbMATH is also involved in a project on text mining using techniques from NLP aiming, among
others, at the extraction of mathematical vocabulary, which could also be used as a tool for author
disambiguation.
6Some institutions even publish guidelines on how to increase visibility via research pro les, e.g.
      </p>
      <p>Utrecht University Library at http://libguides.library.uu.nl/profiles.
interfaces, i.e., they like having full control over what is included in a search result and
the way it is presented.</p>
      <p>At present, requests to zbMATH to clean up author pro les or correct mistakes in
attribution of publications are usually presented by researchers via email, an approach
that is evidently error-prone and not very scalable. Providing an interface to perform
these tasks in a de ned way is an obvious solution.</p>
      <p>On the other hand, in the interest of keeping reliability as high as possible such a
solution cannot be fully automated. Hence it would be desirable to include user input,
but nd a good combination between algorithmic attribution, user correction and manual
curation by trustworthy personnel. The latter is already facilitated by an internal author
disambiguation interface used by the zbMATH sta .</p>
      <sec id="sec-2-1">
        <title>A public interface for author disambiguation</title>
        <p>In order to include community input, we are faced with the task of designing a publicly
accessible interface that allows users to improve the quality of publication data in author
pro les. In particular, when dealing with a single author's list of publications they should
be able to
con rm correctly assigned publications,
exclude incorrectly assigned publications,
resolve ambiguities in algorithmic assignment,
merge pro les that use di erent names (or name variants) for the same person,
and
supply links to pro les in other services and personal homepages.</p>
        <p>Possible users of this service would be
zbMATH authors themselves, i.e. researchers in mathematics at every stage,
their associates (colleagues, students, coworkers etc.), as well as
people generally interested in quality of information content (e.g. librarians).
As this list implies, we do not want nor expect users to solely care for their own author
pro les, but rather see this enhancement of information as a service for science itself and
the community as a whole.</p>
        <p>Moreover, there are di erent usage scenarios that come to mind and that should be
supported by the interface: A casual user might notice a mistake in some author's pro le
and just want to quickly correct it. On the other hand, an author or one of her assistants
might want to completely clean up a pro le, i.e., ensure that all of the author's works
(and hopefully only those) are attributed to her and only to her, all name variants are
correctly recognized, and as much metadata as possible is associated with the pro le. As
this can be a somewhat daunting task in the case of established researchers, the interface
should allow for doing some partial work and picking up later where one left o .</p>
        <p>To make the interface accessible and easy to use, our goal was to provide a clear
overview of the data and the possible ways of action that the user could take. In
particular, some delicate questions concerned the outward presentation of the internal
states the data could have. For example, if algorithmic assignment of a paper turned up
multiple possible authorships, should we present that fact to the user when browsing a
single author's publication list? If some item has been manually assigned via another
user's input before, should we allow reassignment (or recon rmation) of the same?</p>
        <p>Furthermore, while the user should get some immediate feedback after taking an
action, it must be clearly communicated that the changes will not be incorporated
automatically, but rather are subject to an internal review process.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Implications for our design</title>
        <p>As shown in Figure 2 (which shows a preliminary version of the interface), we decided to
place an unobtrusive box at the top of the page with an overview of the possible actions
and links to the relevant sections of the page. There is no real authentication as of now7,
but we do ask for an email address. This in principle allows for further enquiries by our
author identi cation team, and may be turned into a real (but still lightweight) system
of authentication later on.</p>
        <p>Articles that have been manually assigned before are distinguished from others by
the lack of a green Confirm button. They still have a red Exclude button as we are
indeed interested in cases where users think that someone got it wrong before. Every
click sends o an immediate AJAX request queuing the action for internal review, and
the respective button afterwards appears pressed. We still provide an explicit submit
button, which does nothing but show a reassuring message that the input has been
received and will be reviewed internally. At a later stage we plan to introduce lter and
sorting capabilities, e.g., to show only non-manually assigned publications or only those
with ambiguities.</p>
        <p>Finally, we provide the option to write a free text message to our author identi cation
team, as there are bound to be some requests that cannot be reasonably handled using
an all-purpose interface. In this way we at least receive this information in a de ned
way and together with some identifying metadata.</p>
        <p>The user input subsequently passes through an internal checking procedure
involving algorithmical and manual analyses before the corresponding data modi cations are
executed in the database. This step ensures that the quality standard of the data and
hence zbMATH's authoritative status is preserved.</p>
        <p>7This is primarily to lower the hurdle for people to actually use the interface. On the other hand, the
question of nding a reliable authentication system allowing a mathematician to identify herself as
the author of a given set of publications is far from trivial.
For a couple of months we have been testing a prototype of the author disambiguation
interface among di erent user groups (zbMATH editors, mathematics students,
librarians). Based on the very positive feedback we are con dent that the public version of
the interface will enjoy high usage in the mathematical community.</p>
        <p>Provided the expected success of our interface, we envisage to enhance it by additional
features, which would require reliable authentication. Among the planned extensions
would be the possibility to add comments or to tag the content of publications, or to
upload and match additional metadata and sources such as references or photos.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ferreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Laender</surname>
          </string-name>
          ,
          <article-title>A brief survey of automatic methods for author name disambiguation</article-title>
          .
          <source>SIGMOD Record</source>
          ,
          <volume>41</volume>
          , No.
          <volume>2</volume>
          ,
          <issue>15</issue>
          {
          <fpage>26</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fahrni</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Strube</surname>
          </string-name>
          ,
          <article-title>Jointly disambiguating and clustering concepts and entities with Markov logic</article-title>
          .
          <source>in: Proceedings of COLING-2012</source>
          .
          <volume>815</volume>
          {
          <issue>831</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zheng</surname>
          </string-name>
          and
          <string-name>
            <surname>Z. Nie,</surname>
          </string-name>
          <article-title>Combining machine learning and human judgment in author disambiguation</article-title>
          .
          <source>in: Proceedings of CIKM-2011</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kohlhase</surname>
          </string-name>
          .
          <article-title>Search interfaces for mathematicians</article-title>
          .
          <source>in: Proceedings of CICM-2014. Lecture Notes in Computer Science</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>