<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Conference Indexing in Digital Libraries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Schloss Dagstuhl LZI</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>dblp christopher.michels@dagstuhl.de</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TH Koln - University of Applied Sciences</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Trier University</institution>
          ,
          <addr-line>dblp</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>30</fpage>
      <lpage>41</lpage>
      <abstract>
        <p>Digital library curators make relevance decisions in their daily work to prioritize the most urgent metadata updates. In this work, we propose a complex relevance and ranking model to support the decision and prioritization process of digital library curators. Our approach incorporates di erent aspects of relevance decisions into a framework for feasible data quality management in digital libraries. A case study demonstrates the e ects of the factors we use to model these aspects.</p>
      </abstract>
      <kwd-group>
        <kwd>bibliometrics</kwd>
        <kwd>conferences</kwd>
        <kwd>database curation</kwd>
        <kwd>dblp</kwd>
        <kwd>dig- ital library</kwd>
        <kwd>ranking models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Digital libraries4 need to retrieve and keep up with the most recent relevant
research to ful ll the information needs of their target audience in the scienti c
community. Their workload is dominated by indexing the tables of content with
missing old or expected new publications. With limited time and workforce,
curators have to tackle an extensive, complex, and dynamic pool of heterogeneous
data sources, ranging from numerous individual hints in e-mails on a small scale
to harvesting websites and data feeds of publishers on a larger scale. This
bottleneck requires bibliographic prioritization: The most important indexing updates
for missing or upcoming publications need to be identi ed and addressed rst.</p>
      <p>In terms of information retrieval, we can model this prioritization process as a
ranking problem. A given set of data sources has to be ranked according to
speci ed criteria. The features used for ranking should re ect the curators' relevance</p>
      <sec id="sec-1-1">
        <title>Copyright c 2020 for this paper by its authors. Use permitted under Creative Com</title>
        <p>mons License Attribution 4.0 International (CC BY 4.0). BIR 2020, 14 April 2020,</p>
      </sec>
      <sec id="sec-1-2">
        <title>Lisbon, Portugal.</title>
        <sec id="sec-1-2-1">
          <title>4 We use the term digital libraries to consolidate di erent systems like reference</title>
          <p>databases or online bibliographies.
decisions. These relevance decisions of digital library curators are governed by
higher standards of quality than those for other information access systems [1].
In contrast to web search engines such as Google Scholar, a digital library such
as dblp5 is expected to provide structured authority data as well as more
consistent and coherent system responses, relying on a higher understanding of the
data, tasks, communities, and the speci c information needs involved.</p>
          <p>In computer science and related disciplines, conferences constitute the main
channel of sharing results with the research community [5]. Dynamic,
communitydriven lifecycles and event structures are essential hallmarks of conferences, in
comparison with other publication venues such as journals or book series.</p>
          <p>Bibliometric analyses are a source of context knowledge which is of particular
importance to library curators. Lee [3] use several conference-related factors to
predict citation rates on conference papers. Some factors they investigate are
name, age, size, and internationality of conference series. Size is operationalized
in terms of the number of papers presented, which is a commonly available factor
for conferences as opposed to other possible operationalizations like the number
of submissions, attendees, sponsors, or conference pro ts. To operationalize
internationality, they use the degree of international collaborations in papers. They
also look at the age of conferences, trying to answer the question if longer
running conference series are able to gather more citations on papers than shorter
running ones. They nd that internationality is one of the factors that signi
cantly contribute to citation rates. Size has a negative correlation with citation:
the fewer papers are presented at a conference, the more they are cited.</p>
          <p>Besides conference-related metadata such as size, and external bibliometric
data such as citation numbers, archive-internal factors also in uence curator
decisions. For example, the prominence of an author within the archive might
also in uence which pending metadata updates are addressed rst.</p>
          <p>In this work, we propose a complex ranking and relevance model to support
the decision and prioritization process of digital library curators. Our approach
incorporates di erent aspects of curation decisions into a framework for feasible
data quality management in digital libraries. It incorporates both bibliometric
and retrieval-related elements. While the factors itself are motivated through
bibliometric research and ndings, all proposed methods can easily be integrated
into an actual ranking model of a retrieval or recommendation system. This work
is based on the prioritization mechanism for conference metadata updates from
our previous work presented at JCDL 2018 [7]. We extend our previous model
and describe the di erent ranking factors and the corresponding data sources
at a much higher level of detail and include a case study that illustrates the
feasibility of the approach. Our main goal is to thoroughly describe the ranking
factors to allow the readers to understand the mechanics behind them entirely.</p>
          <p>In Section 2, we describe the components we employ to rank conferences with
pending metadata updates. A case study on ranking conference updates in dblp
with these components is presented in Section 3. The work closes with a short
discussion and an outlook on future work in Section 4.
5 https://dblp.org
2</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Components of the Ranking Model</title>
      <p>The task at hand can be modeled as a ranking problem. All conferences listed
in the archive have to be ranked according to an information need inherent in
archive curation. The curators need to have the conferences ranked highest for
which an update is expected. In case of multiple conferences being due
simultaneously, their ranking should re ect their priority for the archive. Usually,
curators are guided by several criteria when deciding which conference to index
next. In describing and evaluating a set of such rank-establishing factors, the
following notation applies.</p>
      <p>Each conference c (like CIKM, ECIR, JCDL, etc.) of the conference set C
is constituted of conference events E(c) := (en; : : : ; e1). In bibliographic terms,
an event e groups all volume-level event members V (e) := (v1; : : : ; vn), e.g.,
proceedings, workshops, or other named parts of published content6. Modeling
the variety in conference event members is simpli ed here. Event members are
simply attributed to a conference by their date date(v). The date date(v)
consists of the year year(date(v)) and the month month(date(v)) of the last
day of the event, i.e., date(e) = date(v) 8v 2 V (e).</p>
      <p>The set Vf (c) := (v j 9e : e 2 E(c); v 2 V (e)) contains all event members of
a given conference, regardless of the event to which they belong, for a partial
function f . When f depends on external metadata sources, e.g., citation links,
no value might be provided for a given event member v 62 Vf (c).
2.1</p>
      <p>Base delay score
The primary criterion for ranking conferences by urgency is the delay between
the expected next indexing date of a new record and the current date NOW.
Conference events usually occur at regular intervals and at roughly the same
time of the year. In the archive, there is a delay between the event and archiving
date indexed(v) := indexed(v) date(v) for each proceedings record v.7 Given
these regularities and the edit history of an archive, we can estimate the next
event and when it is expected to be indexed.</p>
      <p>We assume that the events E(c) of a conference c are in decreasing order
of their date date(e), with the most recent event being denoted by en. The
limited set of up to u most recent events is referred to as recent u(c).We use
recent 6(c) to determine the characteristic interval between the last ve events
in months (unit M ), event(c), with a default of 12M . Furthermore, the archive
delay indexed(c) is the median of indexed(v) for recent 5(c). Finally, to determine
the expected event month expmonth(c), we take the mode of month(date(e))
of recent 5(c)8.</p>
      <p>
        Then, the expected date of the next conference event is date(en+1) =
(expmonth(c); year(date(en))) + event(c). The recording delay of c is
approximated based on this estimated next event and the archive delay in Eq. (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ).
      </p>
      <sec id="sec-2-1">
        <title>6 In the following referred to as proceedings.</title>
      </sec>
      <sec id="sec-2-2">
        <title>7 indexed(v) may be negative in the case of pre-proceedings.</title>
      </sec>
      <sec id="sec-2-3">
        <title>8 In case of a multi-modal distribution, we take the most recent most frequent month.</title>
        <p>
          If no new entry is expected yet (date(en+1) &gt; NOW), the conference scores 0
in the ranking. Otherwise, delay (c) is log-smoothed and inverted in Eq. (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) to
compute the delay scoring factor. Conferences for which new entries are expected
promptly then rank highest, while extremely high delays are practically ignored.
The base score and the factors below have the range [1; 2].
        </p>
        <p>wdelay (c) = 1 +
2.2</p>
        <p>Ranking Factors
delay (c) = NOW
(date(en+1) + indexed(c))</p>
        <p>1
1 + log2(delay (c) + 1)
; delay (c) &gt; 0
We de ne the following ranking factors, which re ect the extent to which there
is a need to index a conference and help determine a corresponding score for
the ranking. The estimated delay of the expected next event constitutes a basic
factor that is combined with a boosting factor for each of the remaining criteria
of a given conference. Each combination results in a corresponding ranking score:
score (c) = wdelay (c)</p>
        <p>w (c);
2</p>
        <p>
          = factive; rate; size; intl ; a l ; cite; promg
Activity. A conference might cease to be organized without any notice reaching
the archive. Conferences that are likely to be discontinued should thus receive
a lower score than active ones. The following scoring takes the relation between
the time since the last entry and the regular event interval into account (Eq.
4). The activity scoring factor in Eq. (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) then boosts conferences that have a
recently active life-cycle according to their archive history.
        </p>
        <p>
          Ratings. External ratings usually guide update cycles if they coincide with the
set of conferences that are relevant to an archive. If such ratings can be integrated
into a single list for a given conference c, a list of corresponding numeric rating
values attributed to c is yielded by rated (c) in Eq. (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ). With the average numeric
rating rate(c) the conference rating weighting factor is given in Eq. (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ).
age(c) =
        </p>
        <p>NOW
wactive (c) = 1 +</p>
        <p>date(en)
event(c)</p>
        <p>1
(1 + age(c)2)
rate(c) =
wrate (c) = 1 +</p>
        <p>P
r2rated(c) r
jrated (c)j</p>
        <p>
          rate(c)
maxrate C
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
(
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
Size. Similar to Lee [3], we de ne the size of a conference in terms of the average
number of papers for each event e(c). The set of papers for a given event member
v is denoted by papers(v). The size factor is normalized by the maximum value
present in the data set (Eq. 9).
        </p>
        <p>size(c) =</p>
        <p>Pv2Vsize (c) jpapers(v)j
wsize (c) = 1 +
jVsize (c)j</p>
        <p>size(c)
maxsize C
Internationality. The number of locations where events are organized can be an
indicator of the internationality of a conference. The multiset located(c) contains
all locations of all event members of a given conference, whereas locations(c)
contains all distinct elements in located(c). The internationality of a conference
intl (c) in Eq. (10) is de ned as the number of distinct event countries, divided
by the total of their occurrences. Normalization for the corresponding scoring
factor is trivial since there cannot be more locations than event members for a
single conference.</p>
        <p>wintl (c) = 1 +
intl (c) = jlocations(c)j</p>
        <p>jlocated(c)j
intl (c) intl (c)</p>
        <p>= 1 +
maxintl C 1
= 1 + intl (c)
A liations. Lee [3] measures the internationality of a conference via the a
liation countries of attendees rather than event locations. We introduce a factor to
approximate the internationality of a conference audience by the a liations of
its published authors as they are known to the archive. For each event member,
given the set of authors with a liation information a liated (v), and the set of
distinct a liation countries of its authors' a liations (v), we compute the
internationality of a liation histories a l (v). The rst quotient in Eq. (12) serves
as a weight, considering that the set of authors with a known a liation history
a liated (v) might only constitute a small part of all distinct authors of an event
member authors(v). The second quotient then models the actual audience
internationality, dividing the distinct locations of the a liation histories of the event
member v by all a liation locations known to the archive (countries) .</p>
        <p>a l (c) =
a l (v) = ja liated (v)j ja liations (v)j
jauthors (v)j countries</p>
        <p>Pv2Va l (c) a l (v)
wa l (c) = 1 +
jVa l (c)j</p>
        <p>
          a l (c)
maxa l C
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
(9)
(10)
(11)
(12)
(13)
(14)
Citations. Incoming citations, by analogy to incoming links on the web [8],
are commonly employed to quantify the scienti c impact of a publication. For
each v, given its publication year published(v), we get the number of incoming
citations cited(v; y) from records published in year y, and cited(y) as the total
of all incoming citations from the publication year y regardless of their target.
The set Yv in Eq. (15) contains all publication years for all citation origins known
to the archive for an event member v. Eq. (17) describes the weighted average
of the number of incoming citations across Yv for a given event member v and
each publication year yi 2 Yv, relative to the number of papers in v and to the
total of incoming citations in yi; it uses the normalized weights described in
Eq. (16). With the oldest publication year given by y1 2 Yv, incoming citations
receive higher weights the closer in time the publication year of their origin is
to the event year of their target. Thus, we consider power-law e ects immanent
to online networks: Similar to the in uence of older, established, and frequently
linked web documents, older, well-known events potentially accumulate more
citations than more recent events [2].
        </p>
        <p>Yv = fy j 9t : t cites v; published(t) = yg; jYvj = m
m m i + 1
X wi = 1; wi = Pm
i=1 i=1 i
cite(v) =
m</p>
        <p>cited(v; yi)
X wi jpapers(v)j cited(yi)
i=1
cite(c) =
wcite (c) = 1 +</p>
        <p>Pv2Vcite (c) cite(v)
jVcite (c)j</p>
        <p>cite(c)
maxcite C
Author Prominence. The total number of published works of an author a is
considered an indicator of their prominence in this scienti c eld. In Eq. (20)
the number of publications known to an archive jpapers(a)j is summed for all
distinct authors of an event member authors(v) and put in proportion to the
number of distinct authors. The average prominence value of all event members
of a conference in Eq. (21) forms the basis of the prominence scoring factor.
prom(v) =
prom(c) =</p>
        <p>Pa2authors(v) jpapers(a)j</p>
        <p>jauthors (v)j</p>
        <p>Pv2Vprom (c) prom(v)
wprom (c) = 1 +
jVprom (c)j</p>
        <p>prom(c)
maxprom C
(15)
(16)
(17)
(18)
(19)
(20)
(21)
(22)
3
3.1</p>
        <p>Case Study on dblp Conference Updates</p>
        <p>Data Sets
Most of the functions described in Section 2 are based on elds of dblp records.
Others rely on external data sets. Both types of data sources are described in this
section. The subset of covered conferences varies among the computed weighting
factors since the required data is sometimes not available for all conferences. If
there is no data available for some conference in a speci c factor, it contributes
the neutral weight 1 to the combined score of that conference. Augmenting
metadata in any case entails an existing mapping from the identi ers of
thirdparty data to the identi ers inherent in the target literature database. In addition
to external ratings and a citation graph, signatures marked up with a liation
locations are integrated into the ranking factors introduced above.</p>
        <p>dblp. The data set we use for our case study is the dblp collection9 as of
2018-12-17. Of more than 4.4 million distinct records in total (excluding author
homepages), this data set contains about 40,000 records of proceedings of about
4,600 di erent conferences to be considered.</p>
        <p>Since the relation between conference event dates and some other date is
essential to our approach, we need to make sure that this information is available
for the records under consideration. Exact date information is available in fewer
titles than month and year information. Therefore, we parse only the event year
and month values for date() from event member titles with simple pattern
matching. There are 4,395 conferences in the set that have at least one event
member for which these data elds could be parsed. Thus, we only take this
subset into account for our analysis. The next important date is the creation date
of a record. Whenever a record is modi ed in dblp, its timestamp is updated.
The creation date of a record thus corresponds to its earliest modi cation date.
The date of publication does not necessarily coincide with the date of the event
but may be several days or weeks in advance or even distinctly later. We use the
creation date as an approximation of the publication date.</p>
        <p>Some of the proposed methods rely on aggregation over simple elds, such
as keys for conferences, event members, publications, author pro les, or record
creation dates, to determine sets of distinct papers or authors and their
respective sizes for the citation-, prominence-, and size-related factors. Other, more
complex elds of dblp records require parsing for the computation of the factors
discussed above. Geographical information10 for located(), for example, are
extracted from the title eld of proceedings records, if possible. Of all conferences
suitable for evaluation, 4,000 have at least one country information available
(91%). A size score &gt; 0 is present for 4,153 of the evaluated conference streams
(94:4%). Prominence scores are available for 4,149 conferences (94:4%).</p>
        <p>Ratings. The rating factor is based on several local, external conference
ratings. The rating CORE originates from Australia and its ratings from 2008
9 https://doi.org/10.5281/zenodo.3051910
10 For parsing geographical information, the Python library geotext (https://pypi.
python.org/pypi/geotext) by Yaser Martinez Palenzuela was used.
and 2017 have been mapped to dblp. A similar, compatible mapping local to the
Brazilian computer science research community [4] has been integrated as well.
The integration process is checked for rare instances of disparate
conferencesubstructure modeling. For example, if dblp attributes two separate conference
identi ers from one of the ratings to the same conference, the rating values
are ignored. The alphabetical rating ranks are mapped to numerical ranks to
compute rate(c) for 791 conferences (18%).</p>
        <p>Citations. The Open Academic Graph (OAG) includes the Microsoft
Academic Graph (MAG) enhanced by AMiner11, comprising approx. 166 million
records. These records are mapped to dblp based on their DOIs, if possible,
falling back to matching the record titles otherwise. About 3.1 million incoming
citation edges from this graph have their citation targets and origins in dblp.
This set of incoming citations is used to compute cited(v; yi) and cited(yi)
above. With the OAG as of 2017-06-09, a citation-based score is computed for
3,900 evaluated conferences (88; 7%).</p>
        <p>A liations. To mark up the author-publication links in dblp with the
country labels corresponding to the author's a liation location at that time,
integrated data sets from OAG as of 2017-06-09 and a set of institutions derived from
Wikidata are used. The institution country labels from Wikidata are matched
based on unnormalized a liation strings of publication authors in MAG and
sanitized by checking if the parsed city exists in the parsed country. The country
labels are attributed to 3.3 million signatures in dblp based on the DOIs and
titles of the publications. For each conference event, the ratio of distinct country
labels which distinct event authors have had up to the year of aggregation to the
number of distinct country labels in the entire data set is computed, resulting
in a liation-based scores for 3,734 evaluated conferences (85%).
3.2</p>
        <p>Detailed Example
In this section, we provide an example in the form of a case study on how to
calculate a ranking score for a speci c conference at a given point in time. The
example describes the components from the perspective of their use in archive
curation: How do the individual factors incorporate latent features for each
scoring? How can they assist curators in di erentiating the various relevant aspects
of the archive update process? In the following, a comparison of ve di erent
conferences answers these questions while also illustrating how the applied
models cope with arising di culties and necessary adaptations. We de ne NOW as
2018-07.</p>
        <p>In Table 1, the estimation of the next event entry as well as the delay, in
relation to an evaluation date, for the International Conference on Data
Engineering (ICDE, unique stream key: conf/icde) is exempli ed. It appears to be
an annual conference ( event(c) = 12M ) with records being added to the archive
with a rather small delay ( indexed(c) = 2M ). The expected next entry of 2018
(date(en+1) = 2018-06) is one month overdue.
11 https://aminer.org/open-academic-graph</p>
        <sec id="sec-2-3-1">
          <title>Function</title>
          <p>limited 6(c)
event(c)
indexed(c)
date(en+1)
delay(c)</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Input / Body</title>
          <p>E(c)
f11; 13; 12; 12; 12g
f1; 1; 1; 2; 2; 1; 5; 2; 3g
2017-04 + 12M + 2M
2018-07 2018-06</p>
        </sec>
        <sec id="sec-2-3-3">
          <title>Output</title>
          <p>In Table 2, raw values and nal scores for each factor de ned above are
demonstrated in comparison to four other conferences { the Joint Conference
on Digital Libraries (JCDL), the conference on Automata Theory and Formal
Languages (Automata), the International Conference on Web-Age Information
Management (WAIM), and the Symposium on Principles of Database Systems
(PODS). Since the year of the date of consideration is 2018, all conference records
that have been added to dblp up to and including 2017 are taken into account.
Values for weighting factors are also calculated using data available up to this
point in time.</p>
          <p>One can see that all conferences share the same delay of one month, as all
of them were expected to be updated in June 2018. Thus, the base score is the
same for all of them12.</p>
          <p>Examining the di erences in scores for the proposed ranking factors
exempli es how they will be ranked in relation to each other in di erent settings.</p>
          <p>First of all, there is no big di erence in the activity factor between the
conferences, since all of them are active in a sense where records are added regularly
to the archive. ICDE gets the lowest score because the last indexed event has
already taken place 15 months before NOW, whereas for WAIM it is only 12
months.</p>
          <p>A rating-based ranking will put ICDE and PODS slightly higher than JCDL
as they received one more A rating in the past. WAIM has a much lower score
here since it has received only one C-level rating. There is no rating data
available at all for Automata; thus wrate (c) is 1:0, the lowest score in this example.
ICDE also scores highest in size with over 100 papers per proceeding on average,
while Automata has only about 21 papers per proceeding. On the other hand,
Automata gets the highest citation score in our example since papers from this
conference are being cited more than the others.</p>
          <p>When it comes to the internationality of a conference, our example highlights
that this notion depends on its de nition. When de ned in terms of diversity
of event locations, Automata outscores all the other conferences, with all eight
event venues having taken place in di erent countries. In this setting, WAIM
receives the lowest score, as all conference events of the past have taken place
in one country. On the other hand, when looking at the a liation countries of
12 To break ties, conferences with the same score are sorted alphabetically.
NOW</p>
          <p>c
delay(c)
wdelay (c)
event(c)</p>
          <p>date(en)
age(c)
wactive (c)
rated abc(c)
rated num (c)
rate(c)
wrate (c)
jpapers(v)j
jVsize (c)j
size(c)
maxsize C
wsize (c)
jlocations(c)j
jlocated(c)j</p>
          <p>intl (c)
maxintl C
wintl (c)
a l (c)
wa l (c)
cite(c)
maxcite C
wcite (c)
prom(c)
maxprom C
wprom (c)
1M
1:5
conf/icde</p>
          <p>conf/jcdl conf/automata conf/waim conf/pods
4:566E
publishing authors, JCDL scores highest by far, followed by ICDE and WAIM
that also have a rather diverse set of publishing authors in terms of a liations.</p>
          <p>In a scenario where the prominence of publishing authors w. r. t. the archive
guides the prioritization of indexing, the PODS conference will be on top of the
ranking just before ICDE. JCDL and Automata, on the other hand, appear to
have less prominent authors and will thus be ranked lower.</p>
          <p>In this case study, we looked at the in uence of the di erent factors by
calculating them on a given set of example conferences. This is not a formal
evaluation of the ranking factors, and the combination of these factors remains
as future work. However, we see the general feasibility and plausibility within
this case study, which was the primary concern in this work.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Discussion and Future Work</title>
      <p>In this paper, we have presented a complex set of ranking factors to support the
decision and prioritization process of digital library curators. Building upon our
previous work, we have elaborated on the de nition of several factors that in
uence the ordering of the pending metadata updates for conferences. A detailed
example in the form of a case study on dblp data demonstrated the e ects of
each of these factors. It exempli es how di erent factors in uence the ranking
of conferences with the same base score, i. e. that are due at the same time.</p>
      <p>Our example also replicates a nding of the bibliometric analysis of Lee [3]:
The conference with the fewest number of papers per proceeding got the highest
citation score in the example set. While this is no result of a formal evaluation,
it still shows the general plausibility of this factor.</p>
      <p>While our ranking factors might look very over-speci ed and very much
tailored to dblp, we believe that our work can be of use to other digital libraries
and use cases. This is due to the fact that we solely rely on publicly available
metadata like MAG, CORE, or Wikidata. The indexing data needed to compute
the base delay score should be available in other digital libraries as well.</p>
      <p>In this work, we focus on the use case of recommending most urgent
conferences to database curators. Of course, these factors can be used in other use
cases, such as retrieval tasks when searching for conference-related resources.
Another use case might be performing bibliometric studies on conferences, such
as nding the most in uential or prestigious conferences in a eld.</p>
      <p>There are some limitations to our work. Data aggregation needed for most
of the factors may su er from awed metadata, like unresolved ambiguous
author names. Even though derived data sets for author-name disambiguation have
proven high standards of quality for bibliographies such as dblp [6], errors are
unavoidable. And even in the case of dblp, this issue may remain despite
several methods of daily on- and o ine curation methods from automated tests to
consulting human experts. Not only the quality of metadata is a limiting factor,
but also the availability of data to calculate the di erent factors. As most of the
used metadata is imported from external sources, this issue might be neglectable
as these data sets are publicly available. Nevertheless, in our case, we could
aggregate only about 20% rating information for conferences. The other data had
much better coverage, but an exhaustive matching could not be achieved.</p>
      <p>Future work consists of the evaluation of our reworked factors against a gold
standard built from actual indexing times, as we did in [7]. This way, we will see
which of our factors (or which combination) models the relevance decisions of
curators best. We also want to focus more on the usefulness of the digital library
for the users, e. g. by incorporating latent signals of unful lled information needs
from the web logs of digital libraries. In the end, this work could lead to a system
that produces a ranking based on this model on demand and guides digital library
curators into updating the most pressing archive de cits.</p>
      <p>Acknowledgments
This work was supported by the German Research Foundation (Deutsche
Forschungsgemeinschaft - DFG, project no. 217852844).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agosti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crivellari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nunzio</surname>
            ,
            <given-names>G.M.D.</given-names>
          </string-name>
          :
          <article-title>Web log analysis: a review of a decade of studies about information acquisition, inspection and interpretation of user interaction</article-title>
          .
          <source>Data Min. Knowl. Discov</source>
          .
          <volume>24</volume>
          (
          <issue>3</issue>
          ),
          <volume>663</volume>
          {
          <fpage>696</fpage>
          (
          <year>2012</year>
          ). https://doi.org/10.1007/s10618-011-0228-8
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davison</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Freshness matters: In owers, food, and web authority</article-title>
          .
          <source>In: Proceeding of the 33rd SIGIR</source>
          <year>2010</year>
          , Geneva, Switzerland,
          <source>July 19-23</source>
          ,
          <year>2010</year>
          . pp.
          <volume>114</volume>
          {
          <fpage>121</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2010</year>
          ). https://doi.org/10.1145/1835449.1835471
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          :
          <article-title>Predictive power of conference-related factors on citation rates of conference papers</article-title>
          .
          <source>Scientometrics</source>
          <volume>118</volume>
          ,
          <issue>281</issue>
          {
          <fpage>304</fpage>
          (
          <year>2018</year>
          ). https://doi.org/10.1007/s11192- 018-2943-z
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Martins</surname>
            ,
            <given-names>W.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalves</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laender</surname>
            ,
            <given-names>A.H.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pappa</surname>
            ,
            <given-names>G.L.</given-names>
          </string-name>
          :
          <article-title>Learning to assess the quality of scienti c conferences: a case study in computer science</article-title>
          .
          <source>In: Proceedings of the JCDL</source>
          <year>2009</year>
          , Austin, TX, USA, June 15-19,
          <year>2009</year>
          . pp.
          <volume>193</volume>
          {
          <fpage>202</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2009</year>
          ). https://doi.org/10.1145/1555400.1555431
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Michels</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fu</surname>
          </string-name>
          , J.:
          <source>Systematic analysis of coverage and usage of conference proceedings in web of science. Scientometrics</source>
          <volume>100</volume>
          ,
          <issue>307</issue>
          {
          <fpage>327</fpage>
          (
          <year>2014</year>
          ). https://doi.org/10.1007/s11192-014-1309-4
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Muller,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Reitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          :
          <article-title>Data sets for author name disambiguation: an empirical analysis and a new resource</article-title>
          .
          <source>Scientometrics</source>
          <volume>111</volume>
          (
          <issue>3</issue>
          ),
          <volume>1467</volume>
          {
          <fpage>1500</fpage>
          (
          <year>2017</year>
          ). https://doi.org/10.1007/s11192-017-2363-5
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michels</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schenkel</surname>
          </string-name>
          , R.:
          <article-title>Prioritizing and scheduling conferences for metadata harvesting in dblp</article-title>
          .
          <source>In: Proceedings of the 18th JCDL</source>
          <year>2018</year>
          ,
          <string-name>
            <surname>Fort</surname>
            <given-names>Worth</given-names>
          </string-name>
          , TX, USA, June 03-07,
          <year>2018</year>
          . pp.
          <volume>45</volume>
          {
          <fpage>48</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2018</year>
          ). https://doi.org/10.1145/3197026.3197069
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gauch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Incorporating quality metrics in centralized/distributed information retrieval on the world wide web</article-title>
          .
          <source>In: Proceedings of the 23rd SIGIR</source>
          <year>2000</year>
          , Athens, Greece,
          <source>July 24-28</source>
          ,
          <year>2000</year>
          . pp.
          <volume>288</volume>
          {
          <fpage>295</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2000</year>
          ). https://doi.org/10.1145/345508.345602
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>