<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identifying collaborations among researchers: a pattern-based approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luca Cagliero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Garza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohammad Reza Kavoosifar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Baralis</string-name>
          <email>elena.baralisg@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Torino - Dipartimento di Automatica e Informatica - Torino</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In recent years a huge amount of publications and scienti c reports has become available through digital libraries and online databases. Digital libraries commonly provide advanced search interfaces, through which researchers can nd and explore the most related scienti c studies. Even though the publications of a single author can be easily retrieved and explored, understanding how authors have collaborated with each other on speci c research topics and to what extent their collaboration have been fruitful is, in general, a challenging task. This paper proposes a new pattern-based approach to analyzing the correlations among the authors of most in uential research studies. To this purpose, it analyzes publication data retrieved from digital libraries and online databases by means of an itemset-based data mining algorithm. It automatically extracts patterns representing the most relevant collaborations among authors on speci c research topics. Patterns are evaluated and ranked according to the number of citations received by the corresponding publications. The proposed approach was validated in a real case study, i.e., the analysis of scienti c literature on genomics. Speci cally, we rst analyzed scienti c studies on genomics acquired from the OMIM database to discover correlations between authors and genes or genetic disorders. Then, the reliability of the discovered patterns was assessed using the PubMed search engine. The results show that, for the majority of the mined patterns, the most in uential (top ranked) studies retrieved by performing author-driven PubMed queries range over the same gene/genetic disorder indicated by the top ranked pattern.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Weighted itemset mining</kwd>
        <kwd>data mining and knowledge discovery</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Plenty of scienti c studies have been published on scienti c journal, books, and
conference proceedings. To deepen their knowledge on speci c research topics
researchers commonly explore in uential studies written by the most renowned
authors. Digital libraries and online databases (e.g., PubMed [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], OMIM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ])
play a fundamental role in supporting researchers in their studies. For example,
topic- and author-driven searches are supported by most of the renowned digital
libraries. Typically, searches are manually performed to retrieve the publications
of interest.
      </p>
      <p>
        In literature many studies have analyzed the correlation between the authors
of scienti c papers and the topics covered by the scienti c literature (e.g., [
        <xref ref-type="bibr" rid="ref15 ref19 ref3 ref7">3, 7,
15, 19</xref>
        ]). For example, existing approaches allow us to identify the most relevant
publications of an author on a given topic. An established way to measure the
relevance of a publication in the research community is to count the number
of received citations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Hence, a large body of work addressed the problem
of identifying most in uential research works covering a given topic by means
of citation content analysis [
        <xref ref-type="bibr" rid="ref19 ref3 ref7">3, 7, 19</xref>
        ]. Citations not only indicate the relevance
of a publication but can be exploited also to assess the in uence/reputation of
individual researchers in their community [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>Since in several domains research works are often fruit of joint e orts of
many researchers, a parallel research issue is the study of the e ectiveness of
the research collaborations among multiple authors. Manually identifying the
most fruitful collaborations on a given research topic is, in general, a challenging
task, because it requires correlating the contribution of multiple authors on a
speci c topic by evaluating the signi cance of their joint research studies with
respect to the existing literature. Hence, automated solutions aimed at analyzing
publication data and automatically discovering fruitful research collaborations
would be desirable.</p>
      <p>
        To address the aforesaid issue, we propose an automated data mining
strategy based on the analysis of publication data acquired from digital libraries
and online databases. The aim is to identify interesting correlations between the
authors of most in uential research studies on speci c research topics. To
analyze publication data we apply a variant of an FP-Growth-like weighted itemset
mining algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Weighted itemset mining is an established data mining
technique that focuses on discovering recurrent combinations of items,
characterized by di erent importance levels, from transactional data [
        <xref ref-type="bibr" rid="ref14 ref16 ref18">14, 16, 18</xref>
        ]. In our
context, items represent authors or research topics. We discover a new type of
pattern, namely the Authors-Topic Pattern (ATP), which represents
combinations of authors and topics that frequently co-occur in the analyzed data (i.e.,
they are associated with a large number of publications). To consider also the
impact of the research collaboration on the research community each
publication in the source data is enriched with the current number of received citations.
Then, pattern relevance is measured as a weighted frequency of occurrence in
the analyzed data (hereafter denoted as in uence). To pinpoint for each
collaboration among multiple authors the research topics that have produced most
authoritative studies, the corresponding patterns are ranked by decreasing
inuence. Since the extracted patterns are easily interpretable, users may easily
explore the top ranked patterns generated by the automatic data mining process.
      </p>
      <p>
        We experimentally evaluated the e ectiveness of the proposed methodology
in a real case study, i.e., the analysis of scienti c studies on genomics and
genetics acquired from two independent libraries (i.e., OMIM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and PubMed [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
In the context under analysis, the ATPs mined from OMIM data represent
combinations of researchers working together on a speci c genetic disorder or gene.
We assessed the reliability of the discovered patterns by exploiting the search
engine of the PubMed library. Speci cally, for each ATP mined from OMIM
data we performed an author-driven query on PubMed to nd the most related
publications co-authored by the same researchers indicated in the pattern. The
query returned a ranked list of related publications. The results show that, for
the majority of the top ranked (automatically generated) ATPs, the most
inuential (top ranked) studies returned by author-driven PubMed queries range
over the same gene/genetic disorder indicated by the pattern. Hence, ATPs
compactly represent salient information about research collaborations on genomics.
The manual retrieval of the same information would entail performing many
PubMed search queries and then combining the results, which can be a
nontrivial and potentially time-consuming task.
      </p>
      <p>The rest of the paper is organized as follows. Section 2 compares the proposed
approach with existing studies. Section 3 thoroughly describes the proposed
methodology, while Section 4 experimentally evaluates its e ectiveness on real
data. Finally, Section 5 draws conclusions and discusses future developments of
the proposed work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        Studying the impact of scientists' research based on the citations received by
their academic publications is a known research problem (e.g., [
        <xref ref-type="bibr" rid="ref19 ref3 ref7">3, 7, 19</xref>
        ]).
Citation content analysis is a common way to tackle this problem. It focuses on
analyzing the semantics, syntax, and position in the text of the paper of the
citations to reveal the in uence of both authors and scienti c papers. For
example, in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] the authors analyzed the sentences including citation expressions to
identify interesting characteristics of scholarly communication. The works
presented in [
        <xref ref-type="bibr" rid="ref19 ref3">3, 19</xref>
        ] classi ed citations based on their semantics to gain insights
into the relationships between authors and topics. A social network of academic
researchers has been proposed in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The authors automatically extracted
researcher pro les from the Web and integrated publication data into the network
from existing digital libraries. The proposed academic search system, called
ArnetMiner, adopted a prede ned Author-Topic model to relate major research
topics to most in uential authors. However, to the best of our knowledge, it can
not be trivially adopted to automatically identify fruitful collaborations among
multiple authors on a speci c topic. Therefore, the goal of this work is
complementary to the above-mentioned approaches.
      </p>
      <p>To assist editors in the peer review of scienti c papers a signi cant research
e ort has been devoted to proposing new methodologies for matching paper
topics with researchers' expertise. For example, in [8{10] the authors addressed
the problem of choosing a pool of reviewers for a given article based on the
expertise of a potentially large set of candidate reviewers and on the main topics
covered by the paper under review. They tackled the optimization problem to
assign each paper to at least three independent reviewers with complementary
expertise so that the pool of reviewers assigned to each paper covers most of
the major topics of the paper and each candidate reviewer has a reasonable
number of reviews to do. Conversely, the problem addressed in this paper is not
an optimization issue. Even though discovering fruitful collaborations among
authors of scienti c studies may help conference chairs and journal editors to
e ectively plan reviewer assignments, the patterns extracted by our methodology
are general and they have not been speci cally designed to address the reviewer
assignment problem.</p>
      <p>
        A parallel research e ort has been devoted to e ciently extracting itemsets
and association rules from weighted data [
        <xref ref-type="bibr" rid="ref14 ref16 ref18 ref2">2, 14, 18, 16</xref>
        ]. This problem extends the
traditional association rule mining task, which was rst introduced in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in the
context of market basket analysis, to the case in which data items are no longer
considered as equally relevant within the analyzed data. For example, in the
context of market basket analysis the goal is to nd sets of products frequently
purchased together by taking into account not only the list of products that
customers have put into their market basket but also the purchased amount
and unitary price of each purchased product. In [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] the authors proposed to
extract weighted association rules, i.e., rules including weights denoting item
signi cance are extracted. In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] weights are used to drive the frequent
and infrequent itemset mining processes, respectively, while in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] weights are
automatically generated by means of graph indexing techniques. This paper
applies a variant of a weighted itemset mining algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to discover a new
type of patterns, which represents signi cant correlations between authors and
research topics.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Scienti c Collaboration Analyzer</title>
      <p>Scienti c Collaboration Analyzer (SCA) is a new data mining-oriented
methodology to analyze the scienti c literature accessible through digital libraries and
online database. The goal is to identify sets of authors (of arbitrary size) whose
joint research on speci c topics has produced publications with signi cantly high
impact. The methodology consists of two main steps: (i) Data collection and
preparation. Publication data and citations are acquired from online sources,
collected into a unique repository and tailored to the next mining process (see
Section 3.1). (ii) Pattern discovery, evaluation, and ranking. Patterns that
represent combinations of authors and topics are extracted from the prepared data
and ranked according to ad hoc evaluation metrics (see Section 3.2). A more
thorough description of each step follows.
3.1</p>
      <sec id="sec-3-1">
        <title>Data preparation</title>
        <p>Publication data are acquired from digital libraries and online databases and
stored in a unique repository. For our purposes, for each publication we acquire
the following information: (i) the Digital Object Identi er (DOI) of the
publication, (ii) the list of authors, (iii) the list of the research topics that are mainly
discussed in the publication, and (iv) the number of citations received.</p>
        <p>
          The current number of citations is considered because it has largely been
used to assess the in uence/popularity of a publication in the research
community [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. However, since the proposed methodology is general, di erent measures
can be easily integrated as well (e.g., the Hirsch index [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]).
        </p>
        <p>To allow pattern mining, publication data are collected into a weighted
transactional dataset. A weighted transactional dataset is a set of pairs htransaction,
weighti, where each transaction corresponds to a di erent scienti c publication,
while weight is the number of citations received by the corresponding publication.
We consider as transaction weight, denoting the relevance of each publication,
the number of citations.</p>
        <p>Transactions consist of sets of items, where items are publication authors
(e.g., Smith, L.), or research topics (e.g., topic X). Items are represented in
the form (feature:value), where feature is Author or Topic, while value is the
corresponding feature value.</p>
        <p>A more formal de nition of weighted transactional dataset is given below.</p>
      </sec>
      <sec id="sec-3-2">
        <title>De nition 1. Weighted transactional dataset. Let A be the set of authors</title>
        <p>and T be the set of topics. Let P be the set of all scienti c publications and let
C(pi) (pi 2 P ) be the number of citations received by publication pi. An item
ik is a pair feature:vq, where vq 2 A if feature is equal to Author or vq 2 T if
feature is equal to Topic. A transaction tj is a set of items related to publication
pj . A weighted transactional dataset D is a set of weighted transactions, where
each weighted transaction twj 2 D corresponds to a di erent publication pj 2 P
and it consists of a pair htj , C(pj )i.</p>
        <p>For instance, Table 1 reports an example of dataset consisting of six weighted
transactions, each one corresponding to a di erent scienti c publication. Each
publication, identi ed by the respective id, is weighted by the corresponding
number of citations (see Column # cit.). For each publication the list of authors
(see Column Authors) and the covered topics (see Column Topics) are known.
Publications can be co-authored, and can be related to many topics. For example,
publication with pub. id 1 received 10 citations (i.e., transaction weight equal
to 10). Its corresponding transaction consists of the following items: Author :
Brown; J:, Author : Smith; L:, T opic : A and T opic : X. The transaction
refers to a publication that was co-authored by Brown J. and Smith L. and that
relates to topics A and X.
3.2</p>
      </sec>
      <sec id="sec-3-3">
        <title>Pattern discovery, evaluation, and ranking</title>
        <p>This step entails discovering a new type of pattern from the prepared weighted
transactional dataset, namely the Authors-Topic Pattern (ATP). It represents a
potentially interesting correlation between a set of authors and a research topic.</p>
        <p>
          Preliminaries. Pattern extraction relies on itemset mining techniques.
Frequent itemset mining [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] is an established data mining technique to discover
recurrent correlations among data items hidden in large datasets. A k-itemset is
a set of k distinct items in a transactional dataset. It indicates the co-occurrence
of the corresponding items in the analyzed dataset. In our context of analysis,
an item represents either an author or a topic (see De nition 1). Hence,
itemsets may represent co-occurrences of multiple authors and topics in the analyzed
dataset. A more formal de nition of itemset is given below.
        </p>
        <p>De nition 2. Itemset. Let D be a weighted transactional dataset and let I
be the set of distinct items in the form feature:vq contained in any weighted
transaction twj 2 D. A k-itemset (i.e., an itemset of length k) is a set of k
distinct items in I.</p>
        <p>Note that each itemset may contain an arbitrary number of items belonging
to any feature.</p>
        <p>
          Since generating all the possible itemsets is computationally intractable even
on medium-size datasets, itemset mining is commonly driven by a minimum
support threshold [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. More speci cally, frequent itemset mining entails extracting
all the itemsets that frequently occur in the source dataset D, i.e., all itemsets
whose frequency of occurrence (support) in the source dataset is above a given
threshold minsup. The support threshold prevents the extraction of less relevant
or misleading itemsets. Thus, it allows us to consider only the most recurrent
and thus potentially reliable patterns.
        </p>
        <p>For example, itemset f(Author : Brown; J:),(T opic : X)g occurs three times
in the dataset in Table 1 (publications with ids 1, 2, and 5). Hence, by enforcing a
minimum support threshold minsup=2 the itemset would be extracted because
its frequency of occurrence (3) is above the minimum (user-provided) threshold.</p>
        <p>
          Unfortunately, the number of frequent itemsets can be very large. To prevent
the generation of redundant patterns, thus simplifying the manual inspection
of the result, a more compact subset of frequent itemsets, called the closed
itemsets [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], can be extracted. An itemset is closed if there exists no superset
that has the same support as this original itemset.
        </p>
        <p>Pattern de nition. For our purposes, we are interesting in mining a
particular type of closed itemsets: the combinations of authors with a speci c research
topic. Hereafter, we will denote it as Authors-Topic Patterns (ATP).</p>
      </sec>
      <sec id="sec-3-4">
        <title>De nition 3. Authors-Topic Patterns (ATP). Let I be a closed k-itemset.</title>
        <p>I is an ATP if (i) it contains one or more items feature:vq such that
feature=Author, and (ii) it contains exactly one item such that feature=Topic.</p>
        <p>Recalling the running example, Table 2 reports some examples of ATPs mined
from the dataset in Table 1. For example, f(Author : Brown; J:),(Author :
Smith; L:), (T opic : X)g indicates that researchers Brown J. and Smith L. have
co-authored many research publications on topic X.</p>
        <p>
          Pattern evaluation. The support quality index of an itemset [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] does not
consider the relative importance of each transaction in the source dataset. More
speci cally, in our context of analysis, each publication may have a di erent
impact on the research community. Some publication can be highly in uential,
whereas others may have a limited scope. Hence, to evaluate pattern signi cance,
pattern occurrences in each publication are weighted according to its impact on
the research community.
        </p>
        <p>
          Since our goal is to generate only the combinations of authors and topic
that have achieved a high impact, we extended the standard itemset mining
problem by integrating item weights [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Speci cally, item occurrences within
each transaction (publication) are weighted by the corresponding number of
citations. Therefore, the co-authorship of publications with a large number of
citations is rewarded, whereas co-authorship of publications with few citations
are penalized. To formalize this step, we introduce the concept of in uence of
an ATP as a weighted frequency of occurrence of the itemset in the weighted
transactional dataset.
        </p>
        <p>De nition 4. ATP in uence. Let D be a weighted transactional dataset and
I be an ATP. Let twj : htj , C(pj )i be an arbitrary weighted transaction in D.
The in uence of ATP I in D, hereafter denoted by inf(I), is de ned as follows:
inf(I) =</p>
        <p>X
twj2DjI tj</p>
        <p>C(pj )</p>
        <p>Recalling the previous example, ATP f(Author : Brown; J:),(Author :
Smith; L:), (Disorder : X)g has an in uence equal to 25 because it covers
the weighted transactions with publication ids 1 (weight 10), 2 (weight 5), and
5 (weight 10), respectively.</p>
        <p>Pattern ltering and ranking. To lter out less interesting patterns a
minimum in uence threshold is enforced. Speci cally, given a weighted
transactional dataset and a user-speci ed minimum in uence threshold mininf , we
extract all the ATPs whose in uence value is above or equal to the threshold.
Patterns can be sorted by decreasing in uence to quickly retrieve the most
relevant ones.
3.3</p>
      </sec>
      <sec id="sec-3-5">
        <title>The extraction algorithm</title>
        <p>
          Many frequent weighted itemset mining algorithms have already been proposed
in literature (e.g., [
          <xref ref-type="bibr" rid="ref14 ref16 ref18 ref2">2, 14, 16, 18</xref>
          ]). To accomplish the ATP mining task from
weighted transactional data, we adopt a variant of the algorithm rst proposed
in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], which performs FP-Growth-like closed itemset mining. FP-Growth [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
relies on an FP-tree data model, i.e., a compact, tree-based representation of
the original dataset residing in main memory. To e ciently generate ATPs on
top of closed itemsets, we separately extracted closed itemsets for each topic by
recursively visiting the FP-tree structure in a depth- rst manner.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Case study</title>
      <p>
        We investigated the applicability of the proposed methodology in a real case
study, i.e., the analysis of the research collaborations on genomics and genetics.
To perform our experiments we analyzed publication data that were acquired
from the Online Mendelian Inheritance in Man (OMIM) catalog of genetic
disorders [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The Online Mendelian Inheritance in Man (OMIM) database [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is one of the
most comprehensive and authoritative compendia of human genes and genetic
phenotypes. OMIM is part of the National Center for Biotechnology Information
(NCBI) system of databases [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and it is freely available on the Web. OMIM
collects information on all known mendelian disorders and over 12,000 genes.
Speci cally, it thoroughly describes the relationships between phenotypes and
genotypes by providing full-text, referenced overviews on genetic disorders. The
database is updated daily and thus its content is continuously evolving over
time. OMIM exposes public Application Programming Interfaces (APIs) for
genetic data crawling and download. Speci cally, it allows users to acquire the list
of all known disorders and a set of related annotations. Disorder annotations
consist of (i) the set of genes correlated with the disorder, (ii) a list of
scienti c publications ranging over the disorder (for each publication the complete
bibliographic information is known), (iii) a textual description of the disorder
including references, and (iv) links to other genetics resources.
      </p>
      <p>
        Our study is focused on discovering from OMIM data sets of researchers that
have conducted in uential studies on genomics or genetics. To tailor OMIM data
to our context of analysis, we considered as topic categories the genetic disorders
and the genes discussed in each publication. Speci cally, the weighted
transactional datasets contains items belonging to three di erent features: Author, Gene,
and Genetic Disorder. To crawl data from the online OMIM database, we
exploited the exposed APIs [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Instead, to retrieve the number of citations received
by each publication we exploited the APIs of the PubMed digital library [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
The integrated dataset, which were obtained by integrating publication data
from OMIM and citation data from PubMed, contains 8825 articles, 34555
authors, 302 disorders, and 1076 genes. The experiments were performed on a 2.67
GHz Intel Xeon workstation with 32 GB of RAM, running Ubuntu Linux 12.04
LTS. The data crawler and the data preparation steps are written in Java, while
the pattern mining algorithm is written in C.
0 0 200 400 600 800 1000 1200 1400 1600 1800 2000
      </p>
      <p>Minimum support threshold
We analyzed the characteristics of the patterns mined by setting di erent values
of minimum in uence threshold (i.e., the minimum number of received citations).
Figure 1 plots the number of mined ATPs related to genes, hereafter denotes
as Author-Gene Patterns (AGPs) for the sake of brevity, and the number of
mined ATPs related to disorders, denoted as Author-Disorder Patterns (ADPs),
by varying the mininf threshold value.</p>
      <p>By setting the in uence threshold we may discover research collaborations
whose production has had a rather di erent impact in their community. For
example, by setting mininf to 400 ATPs represent research teams whose scienti c
studies on a speci c gene/disorder have produced at least 400 citations. By
increasing mininf the constraint on the minimum number of citations becomes
more selective. Hence, as expected, the number of mined patterns decreases more
than linearly while increasing the mininf value. The two curves (those related
to AGPs and to ADPs) show similar trends.</p>
      <p>In Table 3 we categorized the extracted ADPs and AGPs according to the
number of authors appearing in each pattern. The reported categorization
approximately indicates the average impact of the research groups with a given
size. Notably, the in uence is not proportional to the number of authors.
Smalland medium-size groups (e.g., from 3 to 5 persons) with high research in uence
are quite frequent. However, a signi cant number of larger groups (7-8 persons)
have produced in uential studies as well1. To reduce the bias due to large
research teams, groups of few researchers should be analyzed rst. Alternatively,
pattern occurrences can be weighted by the group size beyond the number of
received citations.
4.2</p>
      <sec id="sec-4-1">
        <title>Pattern analysis</title>
        <p>We empirically analyzed the strength of the correlations between the research
teams and the topic identi ed by the pattern. For each topic we considered the
top-5 ranked patterns, mined from OMIM data, in order of decreasing in uence.
Each of the selected patterns indicates the most important topic addressed by
the research team. The research questions we would like to address in this section
are the following:
(A) Are the research team and the topic really correlated with each other?
(B) Among the topics addressed by the team, is the topic indicated in the pattern
the most in uential one?</p>
        <p>To address the above research questions, we assessed the quality of the mined
patterns by performing author-driven queries on the PubMed digital library.
Speci cally, for each pattern we picked the research team (i.e., the author names
occurring in the pattern) and we performed author-driven query on PubMed.
The PubMed query returns a list of publications ranked by decreasing relevance.
If a publication in the top-3 PubMed ranking covers the topic we can conclude
that research team and topic are, to some extent, correlated with each other
(question (A)). If the publication covering the topic is at the top of the PubMed
ranking, we can conclude that the addressed topic is the most in uential among
those addressed by the research team (question (B)).</p>
        <p>We performed the above-mentioned comparison separately for genes and
genetic disorders. Speci cally, Table 4 summarizes the results of the manual
validation performed on the top-5 Authors-Disorder Patterns (i.e., the ATPs related
to genetic disorders), while Table 5 reports similar results for the top-5
AuthorsGene Patterns (i.e., the ATPs related to gene). For each pattern we reported the
title and the Digital Object Identi er of the top ranked publication returned by
PubMed. In most cases, the selected publication matches the topic indicated by
the pattern. To show the correlation between genes/disorders and publication
content, in Column PubMed publication title we highlighted the title keywords
that recall, to some extent, the gene name or the genetic disorder indicated in the
pattern. For example, the top ranked Author-Gene Pattern in Table 5 concerns
gene AUTS1, whose mutations are strongly correlated with the autism disorder.
In the top ranked publication returned by PubMed the title made explicit
reference to the correlated disorder (Recurrent de novo mutations implicate novel
genes underlying simplex autism risk.).
1 Note that genomic and genetic studies are likely to be co-authored by many
researchers.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and future work</title>
      <p>
        This paper presents an itemset-based approach to analyzing publication data and
to discovering fruitful collaboration among researchers. The proposed
methodology generates interpretable patterns that compactly represent collaborations
among researchers that have produced the most in uential studies. The
applicability and e ectiveness of the proposed methodology has been experimentally
evaluated in real case study, i.e., the analysis of the publications related to
genomic and genetics studies. As future work, we aim at (i) testing di erent
weighting functions (e.g., weighting research group size beyond the number of
citations), (ii) applying the proposed methodology to support reviewer
assignment in the process of paper peer reviewer, and (iii) address community
detection and organization nding in the context of project applications. We aim at
integrating Authors-Topic associations into existing optimization-based
strategies (e.g., [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]). For example, considering correlations between multiple authors
and a topic would help editors to diversify assignments across researchers with
complementary expertise.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Imielinski</surname>
          </string-name>
          , and
          <string-name>
            <surname>Swami</surname>
          </string-name>
          .
          <article-title>Mining association rules between sets of items in large databases</article-title>
          .
          <source>In ACM SIGMOD</source>
          <year>1993</year>
          , pages
          <fpage>207</fpage>
          {
          <fpage>216</fpage>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>L.</given-names>
            <surname>Cagliero</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Garza</surname>
          </string-name>
          .
          <article-title>Infrequent weighted itemset mining using frequent pattern growth</article-title>
          .
          <source>IEEE Trans. Knowl</source>
          . Data Eng.,
          <volume>26</volume>
          (
          <issue>4</issue>
          ):
          <volume>903</volume>
          {
          <fpage>915</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ding</surname>
          </string-name>
          , G. Zhang, T. Chambers,
          <string-name>
            <given-names>M.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          .
          <article-title>Content-based citation analysis: The next generation of citation analysis</article-title>
          .
          <source>JASIST</source>
          ,
          <volume>65</volume>
          :
          <year>1820</year>
          {
          <year>1833</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Hamosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Scott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Amberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Valle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>McKusick</surname>
          </string-name>
          .
          <article-title>Online mendelian inheritance in man (omim)</article-title>
          .
          <source>Human Mutation</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <volume>57</volume>
          {
          <fpage>61</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. J. Han,
          <string-name>
            <surname>J</surname>
          </string-name>
          . Pei, and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yin</surname>
          </string-name>
          .
          <article-title>Mining frequent patterns without candidate generation</article-title>
          .
          <source>In SIGMOD'00</source>
          , Dallas, TX, May
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>J. E. Hirsch.</surname>
          </string-name>
          <article-title>An index to quantify an individual's scienti c research output that takes into account the e ect of multiple coauthorship</article-title>
          .
          <source>Scientometrics</source>
          ,
          <volume>85</volume>
          (
          <issue>3</issue>
          ):
          <volume>741</volume>
          {
          <fpage>754</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. K.</given-names>
            <surname>Jeong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Song</surname>
          </string-name>
          .
          <article-title>Exploring the leading authors and journals in major topics by citation sentences and topic modeling</article-title>
          .
          <source>In BIRNDL@JCDL</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Kou</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. H. U.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mamoulis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gong</surname>
          </string-name>
          .
          <article-title>Weighted coverage based reviewer assignment</article-title>
          .
          <source>In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, SIGMOD '15</source>
          , pages
          <year>2031</year>
          {
          <year>2046</year>
          , New York, NY, USA,
          <year>2015</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Kou</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. H. U</surname>
          </string-name>
          , N. Mamoulis,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gong</surname>
          </string-name>
          .
          <article-title>A topic-based reviewer assignment system</article-title>
          .
          <source>Proc. VLDB Endow</source>
          .,
          <volume>8</volume>
          (
          <issue>12</issue>
          ):
          <year>1852</year>
          {1855, Aug.
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y. T.</given-names>
            <surname>Hou</surname>
          </string-name>
          .
          <article-title>The new automated ieee infocom review assignment system</article-title>
          .
          <source>IEEE Network</source>
          ,
          <volume>30</volume>
          (
          <issue>5</issue>
          ):
          <volume>18</volume>
          {
          <fpage>24</fpage>
          ,
          <year>September 2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>C. Lu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ma</surname>
          </string-name>
          .
          <article-title>How does citing behavior for a scienti c article change over time?: A preliminary study</article-title>
          .
          <source>In Proceedings of the 78th ASIS&amp;T Annual</source>
          Meeting:
          <article-title>Information Science with Impact: Research in and for the Community</article-title>
          ,
          <source>ASIST '15</source>
          , pages
          <issue>97:1</issue>
          {
          <issue>97</issue>
          :4,
          <string-name>
            <surname>Silver</surname>
            <given-names>Springs</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MD</surname>
          </string-name>
          , USA,
          <year>2015</year>
          . American Society for Information Science.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>NCBI</surname>
          </string-name>
          .
          <article-title>National Center for Biotechnology Information Website</article-title>
          . Available at http://www.ncbi.nlm.nih.gov/ Last access:
          <source>May</source>
          <year>2017</year>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>A.</given-names>
            <surname>Sidiropoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Katsaros</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Manolopoulos</surname>
          </string-name>
          .
          <article-title>Generalized hirsch h-index for disclosing latent facts in citation networks</article-title>
          .
          <source>Scientometrics</source>
          ,
          <volume>72</volume>
          (
          <issue>2</issue>
          ):
          <volume>253</volume>
          {
          <fpage>280</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>K.</given-names>
            <surname>Sun</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Bai</surname>
          </string-name>
          .
          <article-title>Mining weighted association rules without preassigned weights</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ):
          <volume>489</volume>
          {
          <fpage>495</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>J. Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            , L. Yao,
            <given-names>J.-Z.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            , and
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Su</surname>
          </string-name>
          .
          <article-title>Arnetminer: extraction and mining of academic social networks</article-title>
          .
          <source>In KDD</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>F.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Murtagh</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Farid</surname>
          </string-name>
          .
          <article-title>Weighted association rule mining using weighted support and signi cance framework</article-title>
          .
          <source>In Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          ,
          <source>KDD'03</source>
          , pages
          <fpage>661</fpage>
          {
          <fpage>666</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. J.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
            Han, and J.
          </string-name>
          <string-name>
            <surname>Pei</surname>
          </string-name>
          . Closet+
          <article-title>: searching for the best strategies for mining frequent closed itemsets</article-title>
          . In L. Getoor,
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Senator</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Domingos</surname>
          </string-name>
          , and C. Faloutsos, editors,
          <source>Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , pages
          <volume>236</volume>
          {
          <fpage>245</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>E cient mining of weighted association rules (WAR)</article-title>
          .
          <source>In Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          ,
          <source>KDD'00</source>
          , pages
          <fpage>270</fpage>
          {
          <fpage>274</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. G. Zhang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ding</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Milojevic</surname>
          </string-name>
          .
          <article-title>Citation content analysis (cca): A framework for syntactic and semantic analysis of citation content</article-title>
          .
          <source>JASIST</source>
          ,
          <volume>64</volume>
          :
          <fpage>1490</fpage>
          {
          <fpage>1503</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>