<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Supporting Scholarly Awareness and Researchers' Social Interactions using PUSHPIN</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wolfgang Reinhardt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pranav Kadam</string-name>
          <email>pdkadam@mail.upb.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tobias Varlemann</string-name>
          <email>tobiashv@upb.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junaid Surve</string-name>
          <email>jsurve@mail.upb.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muneeb I. Ahmad</string-name>
          <email>muneeb06@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johannes Magenheim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Paderborn Department of Computer Science Computer Science Education Group Fuerstenallee 11</institution>
          ,
          <addr-line>33102 Paderborn</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>31</fpage>
      <lpage>46</lpage>
      <abstract>
        <p>With the advent of Research 2.0, the way research is conducted has significantly changed. New tools and methodologies have emerged and an increasing amount of research is conducted in networked communities including the use of social networking tools. Apart from the well-known social networks, smaller and tailored social networks for researchers have emerged that are geared towards the specific needs of researchers. As more and more potentially relevant information is being made available, many researchers feel the need for awareness support in order to cope with the available amount of data. In this article we introduce the PUSHPIN application that aims at supporting researchers' awareness of publications, peers and research trends. The application is based on an eResearch infrastructure that analyzes large corpora of scientific publications and combines the extracted data with the social interactions in an active social network.</p>
      </abstract>
      <kwd-group>
        <kwd>research 2</kwd>
        <kwd>0</kwd>
        <kwd>eResearch infrastructure</kwd>
        <kwd>scholarly communication</kwd>
        <kwd>social networking</kwd>
        <kwd>hadoop</kwd>
        <kwd>storm</kwd>
        <kwd>big data analysis</kwd>
        <kwd>near-copy detection</kwd>
        <kwd>object-centered sociality</kwd>
        <kwd>bibliometrics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the early days, the Internet was mostly a top-down information distribution
system in which only few people provided information. Users of the Internet
merely consumed the information without being enabled to interact with or
create own information easily. With the rise of Web 2.0, Internet usage has
been revolutionized. It has enabled mankind to more easily participate in the
spread of information and the participation in global discourse [
        <xref ref-type="bibr" rid="ref20 ref21 ref9">9,20,21</xref>
        ]. The
different developments in Web 2.0 have resulted in a wide range of new tools
and methodologies, which reshaped social interaction, distribution of news and
other content as well as it fostered user participation. Applications like Facebook
and Twitter not only have had impact on the worldwide social system but also
influenced researchers to make applications that modernized how research is
done.
      </p>
      <p>
        The usage of Web 2.0 tools, practices and methodologies in the context of
scholarly communication has been recently labeled as Science 2.0 or Research
2.0 [
        <xref ref-type="bibr" rid="ref27 ref29">27,29</xref>
        ]. Similarly, the term eResearch is used when the talk is about
technologies and infrastructures to support Research 2.0, big data analysis and data
sharing on a large scale. Scholarly communication is generally referred to as the
publication and peer review of scientific publications. In line with [
        <xref ref-type="bibr" rid="ref22 ref23">22,23</xref>
        ] we
consider scholarly communication in a broader scope and consider each social
interactions and communicative activities, which is part of research cycle.
      </p>
      <p>Thus, we especially consider the joint developing of ideas and the exchange
of short texts, like in tweets or status updates as potentially relevant research
information. Moreover, the use of social networks is considered as very relevant
part of the modern research methodology. Despite the fact that Facebook has
evolved to be the de-facto standard in social networking sites (SNS), there are
several SNS that are tailored to the use by researchers and that help them in
connecting to like-minded researcher, publications and other content.</p>
      <p>Applications like Mendeley1, ResearchGate2, Academia.edu3 or iamResearcher4
compete with the top dog Facebook by providing features that cannot be found
in the general-purpose social network. Mendeley for example focuses on the
sharing and annotation of scientific documents in private or public groups. Moreover,
it supports researchers in generating bibliographies and recommending
publications that the research might be interested in.</p>
      <p>However, the new way of conducting research, communicating research ideas
and findings and sharing data also results in a very scattered network of
potentially relevant information. Researchers are in urgent need of awareness support
tools and techniques that provide detailed recommendations and hints for
possible collaborators. Many of the existing approaches seem to be based on first-level
metadata and collaborative filtering approaches only and this is where
PUSHPIN (Supporting Scholarly Awareness in Publications and Social Networks ) will
enhance the state-of-the-art. Through the application of in-depth publication
and citation analysis combined with the immense power of the social graph,
PUSHPIN aims to provide better awareness support for researchers than the
existing tools.</p>
      <p>In the following sections, we present our new application called PUSHPIN
and its approach for awareness support for researchers (Section 2). In Section 3,
we present the implementation details for PUSHPIN and present the underlying
eResearch infrastructure. We also discuss the three user interfaces for web, mobile
and tabletops that PUSHPIN provides for its users. Finally in Section 4, we
give an outlook on future research opportunities and present our evaluation and
public release plans.</p>
    </sec>
    <sec id="sec-2">
      <title>1 http://www.mendeley.com/ 2 http://www.researchgate.net/ 3 http://academia.edu/ 4 http://www.iamresearcher.com/</title>
      <p>2</p>
      <p>The PUSHPIN approach for awareness support for
researchers
PUSHPIN is an ongoing research project at the University of Paderborn
(Germany) that aims to provide awareness support for researchers through the
integration of social networking and big data analysis features. While many features
of the whole approach have already been implemented and can be used, other
features are not yet realized and are currently under development.</p>
      <p>In this section we give an introduction to how PUSHPIN will help researchers
to become and stay aware of their connections to other researchers and
publications. In particular we describe how the social layer and the available social
networking features contribute to the overall awareness of researchers (Section
2.1) and discuss the power of email notifications to keep the users engaged to
visit the platform (Section 2.6). In Section 2.3, we describe how the automatic
analysis of big data sets of publications is supporting object-centered sociality
in PUSHPIN and how it gives insight to the relations of people and objects in
PUSHPIN. Moreover, we present visualizations (Section 2.5) and
recommendations (Section 2.4) that support researchers’ awareness and discuss how we use
mobile devices and interactive displays to access data in our ecosystem (Section
2.7).
2.1</p>
      <sec id="sec-2-1">
        <title>The Social Layer of PUSHPIN</title>
        <p>
          To raise awareness of an idea and to create a circle of supporters of the same,
it is essential for any research idea to reach a wide audience. Social networking
makes it possible to connect to potential collaborators thereby supporting the
start of an incipient Research Network. Where social networking tools are often
based on the people element, on the other hand, social awareness tools tell us
a story using various data associated with people and helps us build a network
based on such data. Often, we also find social networks that assemble around
specific objects, which become the hub for social interactions [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. In PUSHPIN,
the objects that realize this object-centered sociality [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] are scientific
publications. While PUSHPIN can identify that there is a direct connection between
two researchers as they follow each other, we also provide social awareness by
stating that there are x publications that both of them have cited in their own
writings. This way, the system may make the researchers aware of their shared
interest and common knowledge in a certain research area and may trigger a
user action.
        </p>
        <p>
          The social layer of PUSHPIN aims to support users in creating an active
social network that is created by the users themselves through social interactions
and conscious activities. The other parts of PUSHPIN rather contribute to a
passive social network that is automatically generated by the system and that is
built based on abstract information and activities such as collaboratively writing
publications, working at the same institution or citing similar works [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Both,
social networking features and social awareness support together can provide a
powerful framework to support research [
          <xref ref-type="bibr" rid="ref14 ref23">14,23</xref>
          ]. The following points describe
how PUSHPIN support object-centered sociality and active network constructs.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Sign-up and sign-in using existing accounts To ease the sign-up and sign</title>
        <p>in process for users and to support them to reuse their existing social profiles
as login, we enable login via Facebook, Twitter and Mendeley. Moreover,
PUSHPIN gets access to the respective social graphs and can recommend
friends from the other social networks that already use PUSHPIN.
User profile updates A user’s profile plays an important role in getting to
know the user. To name a few, it consists of information about the user’s
affiliations, research interests and research disciplines, which highlights the
user’s research areas. This information has impact on the engagement in
social networks as they reflect the personality of a user. Any changes to
the user’s profile are presented to the followers of that user in their activity
stream.</p>
        <p>Following a user Users can follow other users to get an account of all their
activities. The number of followers (users following the current user) and
followees (users followed by the current user) are shown on the dashboard of
a user as well as any user’s profile. The number of followers of a user can be
taken as quantification of the popularity and networking efforts of that user
on PUSHPIN.</p>
        <p>Status updates, likes and comments Sharing status updates is a common
construct in social networking applications which allow users to share their
current thoughts or their work progress. In PUSHPIN, the status updates
could be used not only for sharing ideas or current readings but also for
requesting help or simply sharing some news. Also, all followers of the user
can like and comment on a status message, which may eventually result in
a discussion of the content shared. Moreover, if a user has connected other
social media accounts to his PUSHPIN account, she can automatically share
the status update with all of her other accounts.</p>
        <p>Private messaging To support non-public information exchange, all users on
PUSHPIN can exchange private messages with each other. Messages are
stored in conversations that multiple users can be part of. Any member of a
conversation can add additional users to the conversation and each user can
leave a conversation at any time.</p>
        <p>User’s activities When PUSHPIN users successfully sign in, they are
redirected to their personal dashboard. A significant part of the dashboard
consists of an activity stream, which is a sorted summary of activities. These
activities consist of stories such as status updates of users, likes and
comments on statuses, changes in profile information, users following and tagging
other users, users uploading, bookmarking, rating and tagging publications,
etc. In short, it tells stories of the users’ interaction with other users and
publications. Users can only see updates of other users, whom they follow.
Apart from the dashboard, users can also see activities of a particular user
on their user profile. This kind of feature is common with most of the social
networking platforms including Twitter and Facebook and hence, most of
the users are already familiar with it.
Uploading publications Since scientific publications are the central hub for
object-centered sociality in PUSHPIN, users can upload publications to the
service5. This may be done by selecting publication from the local computer
and uploading them, or by connecting their Mendeley account to PUSHPIN.
In the latter case, all the PDFs in the user’s Mendeley collections are
automatically imported in the PUSHPIN infrastructure. All the publications
that have been uploaded to the system, are then automatically analyzed and
information is extracted from them (see Section 2.3 for a detailed description
of this process).</p>
        <p>Interacting with publications All the users have access to the dedicated
profiles of all the publications in PUSHPIN. On the profile, users can rate the
publication and share it on other social networking sites. Moreover, users
can recommend the publication to other PUSHPIN users or send the
recommendation via email. Finally, users can bookmark the publication and put
it in one of their collections on PUSHPIN.</p>
        <p>
          Tagging objects Social tagging is one of the most prominent features of Web
2.0 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and is available for all kinds of objects in PUSHPIN. Users can
tag publications and institutions and to classify other users they can also
tag users (this is commonly referred to as people tagging [
          <xref ref-type="bibr" rid="ref19 ref3 ref7">3,7,19</xref>
          ]). When
someone explores a keyword, all the users tagged with that keyword form a
part of search results in researchers’ list.
2.2
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Publications analysis</title>
        <p>All scientific publications that are uploaded to the PUSHPIN infrastructure 6
are automatically analyzed according to several aspects. This automatic analysis
represents a series of processing steps that are executed after a publication is
uploaded the PUSHPIN system.</p>
        <p>
          The first and foremost step taken is to check if the publication is already in
the publication corpus and/or if a full analysis has to be started. If this is not
the case, the uploaded publication is inserted into HBase7. After that, Storm8
is triggered for further analysis of the publication. This analysis by Storm
involves activating the metadata extraction and reference extraction modules to
obtain the metadata and references from the publication. The metadata,
being referred to, can be the title, the author(s) and their email addresses, the
authors’ institutions, abstract, and keywords. For each of the references that
have been cited in the publication, the reference extraction module looks for
title, author(s), year of publication and publication outlet. The two modules
5 Due to potential copyright infringements, we will only process the uploaded data in
order to extract metadata from the publications. We will not, however, allow the
public download of the PDFs shared with the PUSHPIN system.
6 Currently we only process articles in PDF format. In particular, we do not process
books or theses.
7 http://hbase.apache.org
8 http://storm-project.net
use GROBID9 and ParsCit10 as key software tools. If additional metadata in
BibTEX or PLoS XML format is available, the modules make use of this
information as well. The extracted data is then compared and combined to get
the most exact metadata (similar to our approach in [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]). Alongside metadata
extraction, Storm also triggers a module that creates thumbnails of each page
of the uploaded publication.
2.3
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Near-copy detection and publication similarities</title>
        <p>
          A problem of modern science is the rising amount of plagiarism. In the digital age
it has become much easier to access scientific publications and to copy content. In
order to detect conscious or unconscious plagiarism we introduced algorithms to
PUSHPIN, which are capable of doing near-copy detection (NCD). NCD means
that correctly cited paragraphs will also be detected. To distinguish between
full-text quotes and plagiarism, additional algorithms have to be used to detect
plagiarism indicators. This could be done in future projects. The NCD algorithm
used in PUSHPIN are inspired by the fuzzy string similarity detection algorithm
described in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>Each uploaded paper first goes through initial text preprocessing steps before
it can be analyzed by our NCD algorithm. These initial steps are used to remove
irrelevant and uninteresting parts of the text and to make the different text
better comparable:
Text extraction The papers are uploaded as PDF files. From these files, the
text, along with the information about its position in the PDF file are
extracted. This gives the exact location of a copied text in the documents it
appears in.</p>
        <p>Text cleaning The extracted text contains – for the NCD algorithm –
uninteresting information, like headers and footers of the document. These lines
are removed and hyphenated words are joined again.</p>
        <p>Language detection Some algorithms need to know the language of the text
as they work with trained models that are specific for one language.
Part-of-speech tagging The "Part-of-Speech" (POS) tagging determines the
grammatical meaning of a word in a sentence. This information is necessary
for detecting synonym groups of words later on. Moreover, POS tagging is
also useful in combination with lemmatization for calculating word clouds.
Lemmatization and stemming For comparing words in our NCD algorithm,
it is necessary to bring all words to the principal form, which is the same for
all tenses and plural and singular forms. Lemmatization transforms words
in the principal form using a dictionary algorithm. This algorithm is
expensive in time and memory but the results are real words, which also can be
displayed in word clouds. Stemming is an algorithmic transformation of the
input word that will transform it to the stem. The stem, however, does not</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>9 http://grobid.no-ip.org 10 http://aye.comp.nus.edu.sg/parsCit</title>
      <p>need to be a real word and thus should not be used in word clouds or the
like but its calculation is very fast.</p>
      <p>Number and stop word removal In this step, we remove unimportant
elements from the text in order to reduce the complexity of the NCD algorithm
computation.</p>
      <p>
        Synonym detection Often, copiers try to conceal the copies by replacing words
with synonyms of the word. This makes it harder to detect certain parts of
a text as copied. This makes it necessary to detect synonym groups that a
given word belongs to and to check all synonyms of the word for potential
copies. In this step we make use of the WordNet project [
        <xref ref-type="bibr" rid="ref16 ref8">16,8</xref>
        ] and a modified
Lesk algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for distinguishing the different meanings of a word.
      </p>
      <p>
        After the text preprocessing is finished, the NCD algorithm can calculate the
similarities between all sentences of the publication and the preprocessed
background corpus. This procedure is inspired by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] but additionally incorporates
the similarity between two synonym groups. Whereas the original algorithm
uses a similarity of 1 if two words are equal, a similarity of 0.5 if they are in the
same WordNet synonym groups and 0 in all other cases, we calculate the Wu
and Palmer WordNet similarity [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] between two words if they are not equal.
Additionally to the sentence-level calculation of similarities, we also compute
several text-based similarity measures on a fulltext-level of all publications in
the PUSHPIN corpus with respect to each other.
      </p>
      <p>This computation needs very large computational power and produces a lot
of similarity data. We rely on the Apache Hadoop framework to scale the
computation to a cluster of computers (see Section 3 for a detailed inspection of the
PUSHPIN eResearch infrastructure).
2.4</p>
      <sec id="sec-3-1">
        <title>Recommendations</title>
        <p>In PUSHPIN we use an ample number of recommender algorithms due to the
following reasons:
1. The system has to take into consideration the networks that result from the
extracted co-authorship information as well as the co-citation and
bibliographic coupling data of publications.
2. For item-based recommendations, the system also has to employ the use of
textual similarities, clustering results, author-assigned and extracted
keywords as well as user tags.
3. Also, the system is capable of tracking user activity on the PUSHPIN web
application, store the user activity, and based on these, be able to
recommend resources (e.g., users who bookmarked publication X also bookmarked
publication Y; mutual followers; you might also assign these tags to the
resource because others did so; people who visited this resource also visited
that resource).</p>
        <p>To sum up, the recommender system takes into account all the above
information for recommendation. In addition, the recommendations will be textual
and visual, and also can be explained to the user.
2.5</p>
      </sec>
      <sec id="sec-3-2">
        <title>Visualizations</title>
        <p>
          Visualizations prove very useful in presenting and understanding large and
complex sets of data and mining for hidden patterns within them. They serve as a
very useful decision support tool in research networks and help researchers to
become and stay aware of large data sets [
          <xref ref-type="bibr" rid="ref18 ref23">18,23</xref>
          ]. Sometimes, they also allow
interaction with the data in order to enhance the understanding [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]. In
PUSHPIN, visualizations play an important part to support social awareness using a
set of aesthetic visualizations of data related to researchers, affiliations and
publications. We will have a brief look at some of the visualizations that we have or
plan to have in PUSHPIN.
        </p>
        <p>Usage and statistical visualizations This category of visualizations will be
prevalent throughout PUSHPIN. For researchers, there will be a simple chart
depicting the development of followers, co-authors, publications, etc.
Similarly, there will be charts for a publication how the number of citations
and bookmarks developed over time. Besides, visualizations based on
general statistical data like typical co-authorship network sizes, most referenced
articles, top research disciplines, etc. will have a place in PUSHPIN.
Trend-based visualizations This category will include trends using numbers
as well as trends in usage of text over time. Trending citations, authors,
topics and keywords will be visualized in appropriate manner.</p>
        <p>
          Similarity-based visualizations Details of textual similarity between papers
and bibliographic coupling similarity between papers will be explored here.
Moreover, appropriate visualization of paragraphs that have been found
during the near-copy detection will be developed and provided in PUSHPIN.
Map-based visualizations Geo-spatial visualizations show us the
geographical location of researchers and institutions and help us understand the widely
spread co-authorship networks and the associations of different institutions
(inspired by the works of [
          <xref ref-type="bibr" rid="ref17 ref18">17,18</xref>
          ]). Particularly, we have interactive
visualizations that show and link us to various information related to a researcher
or an institution and relations between them.
        </p>
        <p>Co-authorship visualizations For a researcher, there will be a circular
visualization with the researcher at center and his co-authors around him in
circles. This give us a chance to explore the co-authors of this researcher.
When a user explores a discipline, a research interest, an institution or a
tag, there can be sets of co-authorship networks related to the explore query
which may not be connected. Hence, we do not use a radial layout here,
instead build a graph comprising of different networks(not connected) to show
various sets effectively.</p>
        <p>Besides the above categories, we will also have tag-based visualizations like
word clouds, spark lines, etc. and also circle-based visualizations
2.6</p>
      </sec>
      <sec id="sec-3-3">
        <title>Email notifications</title>
        <p>
          As Fred Wilson points out “ if you want to drive retention and repeat usage [of
your service], there isn’t a better way to do it than email ” [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. Instead of making
email disappear, social media has created new application fields for email and
makes heavy use of them in all kind of domains. In PUSHPIN, we also use the
power of email notifications to keep the users of the system up-to-date what is
going on in PUSHPIN. Users will receive emails when they have new followers or
someone comments on their publications. PUSHPIN will send alerts if it found
new publications of an author or if someone tagged an author’s publication. If
users do not want to be bothered with emails, they can deactivate them or set
adjust their granularity and frequency levels.
2.7
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Access on mobile devices and interactive displays</title>
        <p>
          In our previous research we found that mobile access to research information,
together with context-awareness and push notification of relevant information is
very relevant for researchers overall awareness of their research networks [
          <xref ref-type="bibr" rid="ref23 ref25">25,23</xref>
          ].
Moreover, research conducted by Nagel et al. [
          <xref ref-type="bibr" rid="ref17 ref18">17,18</xref>
          ] and Vandeputte et al. [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]
shows that interactive tabletop applications are useful for sensemaking of
publication data and co-authorship networks. Moreover, most of the existing social
networks and Research 2.0 applications make allowance for the immense
pervasion of mobile devices among all social classes by providing dedicated mobile
applications the resemble the features of their web-based counterparts. Often,
the mobile applications even make extensive use of the specific technical
characteristics of the mobile devices such as camera, microphone, GPS positioning.
Against this background, we decided to provide a mobile application, which
could be used by all PUSHPIN users and a multitouch application that should
be used for special occasions such as conferences.
        </p>
        <p>The PUSHPINmobile application resembles a significant part of the features
of the web application. Making use of specific mobile interface patterns such as
dashboards and multitouch gestures, researchers are enabled to access all the
information from the social layer and to engage in social interactions with their
peers. Researchers are also able to view their own and other researchers’ profiles,
search nearby researchers depending on their physical location, and also explore
the different research disciplines, institutions and publications in the system.
Moreover, researcher will also be able to tag other researchers and communicate
with each other through private messages.</p>
        <p>Beyond that, users of the mobile application will be enabled to
authenticate and exchange data with the multitouch table application (PUSHPINMT ).
Therefore, researchers can connect to PUSHPINMT using either Bluetooth or
NFC. Additionally, the mobile application can bring up QR codes that can be
scanned by the multitouch application. The QR codes can contain information
about the researcher’ own or other researchers’ profile, institutions or
publications. On PUSHPINMT , users will be able to explore their relations to other
researchers and publications based on several scientometric measures. Moreover,
they can explore the publications in PUSHPIN based on tags and other
classifications. Finally, they can scan QR codes of any PUSHPIN object and get a
virtual representation of the object on the tabletop.
3</p>
        <p>PUSHPIN’s eResearch infrastructure implementation
In this section we describe the technological underpinning of PUSHPIN’s
eResearch and big data analysis infrastructure and relevant technologies we employ
in the realization of the PUSHPIN user interfaces.</p>
        <p>
          The Hadoop Distributed File System is an open source implementation of a
fault tolerant, self-healing, distributed filesystem for large datasets inspired by
the Google filesystem (GFS)[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. It is designed to store large file, which are split
and distributed over several nodes of a cluster, and to achieve high performance,
while serving the data to computing processes. The processing methodology
of Hadoop is an implementation of the MapReduce paradigm [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], which is
designed to handle large amounts of data by splitting the input stream into chunks,
which are computed on several nodes of a cluster. The MapReduce paradigm
divides the processing into two stages to reduce the complexity. The first stage
(map) processes several input key/value pairs and outputs a set of intermediate
key/value pairs, which are sorted and transferred to the second stage (reduce).
The reducer, eventually, merges all intermediate values, which are associated to
the same key and outputs results for that key.
        </p>
        <p>Hadoop provides batch processing function, which perfectly scales with the
number of nodes in a cluster. This functionality excellently supports parallelism
to a wide range of algorithms especially in data mining and information retrieval.</p>
        <p>In PUSHPIN, Hadoop is used for several algorithms, which need large
computational performance and that process big data. Amongst others, these
algorithms compute the similarity of texts, clusters the papers, builds recommender
models or run near-copy detection algorithms. Moreover, we use Apache
Mahout14 for the calculation of text-based similarities, text clustering, classification
and recommender algorithms based on Hadoop MapReduce.
3.2</p>
      </sec>
      <sec id="sec-3-5">
        <title>Text preprocessing</title>
        <p>As described in Section 2.3, we perform several text preprocessing steps before a
paper can be analyzed by the near-copy detection algorithm. The text extraction
and thumbnail generation is done using Apache PDFBox15. Since many
algorithms need to have knowledge about the language of a text, we use a Java-based
language detection library16 for that. The Part-of-speech tagging is realized
using Apache OpenNLP17. Stemming and lemmatization of the extracted texts
is implemented on top of the Mate Tools natural language analysis toolkit18.
Finally, we make use of Apache Lucene19 in the process of removing numbers
and stop words that we consider as being not relevant for text similarities or
near-copy detection.
3.3</p>
      </sec>
      <sec id="sec-3-6">
        <title>Metadata and reference extraction</title>
        <p>During the metadata and reference extraction processes we are trying to
accurately detect a publication’s title, author(s), contact information, like emails and
14 http://mahout.apache.org
15 http://pdfbox.apache.org
16 http://code.google.com/p/language-detection
17 http://opennlp.apache.org
18 http://code.google.com/p/mate-tools
19 http://lucene.apache.org
address data as well as author-provided keywords and the publication’s abstract.
Moreover, we are interested in the list of references and all the relevant data from
each of the references. This metadata is extracted for different purposes, e.g.,
the attribution of publications to PUSHPIN users, the creation of co-authorship
graphs, the calculation of recommendations and for detecting reference and
research trends.</p>
        <p>
          Once a publication has been uploaded to PUSHPIN and inserted into HBase,
the metadata and reference extraction modules get triggered by Storm. The
process involves triggering ParsCit and GROBID in parallel threads. GROBID
(GeneRatiOn of BIbliographic Data) employs the concept of Conditional
Random Fields (CRFs) for pattern recognition and data extraction [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ]. Using this,
"GROBID extracts the bibliographical data corresponding to the header
information (title, authors, abstract, etc.) and to each reference (title, authors, journal
title, issue, number, etc.). The references are associated to their respective
citation contexts " [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. ParsCit also employs the use of CRF model at its core for
metadata extraction by locating reference strings, parsing them and retrieving
their citation contexts. It employs state-of-the-art machine learning models to
achieve its high accuracy in reference string segmentation, and heuristic rules to
locate and delimit the reference strings and to locate citation contexts. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Each tool does an independent metadata and reference extraction and the
two results, obtained at the end, are then combined with potentially available
other metadata like BibTEX data or PLoS XMLs. This merging is necessary
as sometimes the metadata extracted from both tools differs, and also at times
either of the tool misses out on some important metadata. If available, the data
available in BibTEX or PLoS XML format are the most accurate source of
information since they have been manually created by people knowledgeable of the
publication.
3.4</p>
      </sec>
      <sec id="sec-3-7">
        <title>Sign-up and sign-in using OAuth</title>
        <p>
          In PUSHPIN, we use the Open Authorization (OAuth) protocol20 to allow users
to login to PUSHPIN using their Facebook, Twitter or Mendeley accounts.
OAuth “ is a security protocol that enables users to grant third-party access
to their web resources without sharing their passwords ” [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Apart from this,
PUSHPIN also serves as an OAuth service provider, which implies that websites
can use PUSHPIN for the sign-up and sign-in of users. OAuth is also used to
connect the three PUSHPIN user interfaces to the backend.
3.5
        </p>
      </sec>
      <sec id="sec-3-8">
        <title>The PUSHPIN API</title>
        <p>In PUSHPIN, we use provide a REST (REpresentational State Transfer) API
(Application Programming Interface) to communicate between the frontends
(web-based application, mobile application and multitouch table) and the Java
backend. The frontend sends/requests data to the backend using the REST API,
20 http://oauth.net
e.g., information about a certain resource such as a publication. The backend
in turn returns a representation of the resource in JSON notation. The reasons
for using REST (over other available web services such as SOAP) are that it is
light-weight, simple, very popular among web applications and that it provides
better performance and scalability.
3.6</p>
      </sec>
      <sec id="sec-3-9">
        <title>PUSHPIN user interfaces</title>
        <p>PUSHPIN currently provides three user interfaces for its users. The web-based
application serves as the main interface to our service and will be used by the
average user. Moreover, we provide a mobile application for Android smartphones
that allows the anytime-anywhere access to PUSHPIN’s main features. Finally,
we also provide a multitouch application for tabletop-displays that supports
users in exploring the PUSHPIN data in new ways.</p>
        <p>Web-based application The web-based PUSHPIN front-end is a self-contained
application and serves as the primary application to most of the users (see
Figure 1). This application is written in PHP5 and builds on the state-of-the-art
in HTML5 and CSS3 development. It also involves extensive use of JavaScript
that enhances the user experience. Also, various Javascript frameworks are used
for different visualizations.</p>
        <p>
          Mobile application The PUSHPINmobile application is developed using the
Android 4 SDK and supports all smartphones running Android OS 4.0 and
higher. PUSHPINmobile currently provides users an interface to the social layer
of PUSHPIN and lets them flip through their activity stream, like and comment
entries and post new status updates. The application can scan QR codes of any
PUSHPIN object and present the data related to that object. Moreover, the
users can locate themselves and see relevant researchers around them.
Multitouch application The main purpose of the PUSHPINMT application
is to provide different interactions with the data in PUSHPIN. In [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] we discern
four basic modes of data exploration on PUSHPINMT : the 1) people-based,
2) topic-based, 3) event-based and 4) trend-based approach. Users can use the
search to bring up researcher or publication profiles or authenticate themselves
using PUSHPINmobile or QR codes. Moreover, they can explore the relations
between publications, which can be related by common references or authors,
textual similarity or even by copied/cited paragraphs. Finally, users can explore
the trends in reference and publication data as well as exploring the authorship
patterns found during the automatic analysis of the publications.
4
        </p>
        <p>Conclusion and future research opportunities
In this paper we have introduced the PUSHPIN approach for awareness
support in research networks. In PUSHPIN we combine the best of two worlds:
classic features of Facebook-like social networking sites and those of innovative
eResearch infrastructures. The integration of these features results in enhanced
awareness support for researchers on both a social and a content layer. The
recommender systems in PUSHPIN will not only recommend publications based on
collaborative filtering but also on the actual content and reference data within
the publications. Thus, PUSHPIN goes beyond the state-of-the-art and might
help overcoming unwanted fragmentation in research networks and connecting
researchers that otherwise would have stayed unknown to each other. In the
coming months we will continue to improve the implementation of the
analytical backend and further enhance the three user interfaces. We will invite selected
users to an alpha test of the PUSHPIN web-based application in August and
evaluate the existing features with them. The feedback on early versions of the
software will help shaping the further development. We plan to release the system
to public beta in early October 2012.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Salha</given-names>
            <surname>Alzahrani</surname>
          </string-name>
          and
          <string-name>
            <given-names>Naomie</given-names>
            <surname>Salim</surname>
          </string-name>
          .
          <article-title>Fuzzy Semantic-Based String Similarity for Extrinsic Plagiarism Detection</article-title>
          .
          <source>Lab report</source>
          , Taif University Saudi Arabia and Universiti Teknologi Malaysia,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Satanjeev</given-names>
            <surname>Banerjee</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ted</given-names>
            <surname>Pedersen</surname>
          </string-name>
          .
          <article-title>An adapted lesk algorithm for word sense disambiguation using wordnet</article-title>
          . In Alexander Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing</source>
          , volume
          <volume>2276</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>117</fpage>
          -
          <lpage>171</lpage>
          . Springer Berlin / Heidelberg,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Simone</given-names>
            <surname>Braun</surname>
          </string-name>
          , Christine Kunzmann, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Schmidt</surname>
          </string-name>
          .
          <source>People Tagging &amp; Ontology Maturing: Towards Collaborative Competence Management</source>
          , pages
          <fpage>133</fpage>
          -
          <lpage>154</lpage>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Isaac</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Councill</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lee Giles</surname>
          </string-name>
          , and
          <article-title>Min yen Kan. Parscit: An open-source crf reference string parsing package</article-title>
          .
          <source>In INTERNATIONAL LANGUAGE RESOURCES AND EVALUATION. European Language Resources Association</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sanjay</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          .
          <article-title>Mapreduce: simplified data processing on large clusters</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>51</volume>
          (
          <issue>1</issue>
          ):
          <fpage>107</fpage>
          -
          <lpage>113</lpage>
          ,
          <year>January 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Jyri</given-names>
            <surname>Engeström</surname>
          </string-name>
          .
          <article-title>Why some social network services work and others don't - or: the case for object-centered sociality</article-title>
          . Available online http://bit.ly/eJA7OQ (accessed
          <issue>31 December 2010</issue>
          ),
          <year>April 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Farrell</surname>
          </string-name>
          , Tessa Lau, Stefan Nusser, Eric Wilcox, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Muller</surname>
          </string-name>
          .
          <article-title>Socially augmenting employee profiles with people-tagging</article-title>
          .
          <source>In Proceedings of the 20th annual ACM symposium on User interface software and technology, UIST '07</source>
          , pages
          <fpage>91</fpage>
          -
          <lpage>100</lpage>
          , New York, NY, USA,
          <year>2007</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Christiane</given-names>
            <surname>Fellbaum</surname>
          </string-name>
          . Wordnet. In Roberto Poli,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Healy</surname>
          </string-name>
          , and Achilles Kameas, editors,
          <source>Theory and Applications of Ontology: Computer Applications</source>
          , pages
          <fpage>231</fpage>
          -
          <lpage>243</lpage>
          . Springer Netherlands,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Fuchs</surname>
          </string-name>
          .
          <source>Handbook of Research on Web 2.0</source>
          ,
          <issue>3</issue>
          .0,
          <string-name>
            <surname>and X.</surname>
          </string-name>
          <article-title>0: Technologies, Business, and Social Applications</article-title>
          , volume II,
          <source>chapter Social Software and Web 2.0: Their Sociological Foundations and Implications</source>
          , pages
          <fpage>764</fpage>
          -
          <lpage>789</lpage>
          . IGI-Global, Hershey, PA,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Sanjay</surname>
            <given-names>Ghemawat</given-names>
          </string-name>
          , Howard Gobioff, and
          <string-name>
            <surname>Shun-Tak Leung</surname>
          </string-name>
          .
          <article-title>The google file system</article-title>
          .
          <source>SIGOPS Oper. Syst. Rev.</source>
          ,
          <volume>37</volume>
          (
          <issue>5</issue>
          ):
          <fpage>29</fpage>
          -
          <lpage>43</lpage>
          ,
          <year>October 2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Eran</given-names>
            <surname>Hammer</surname>
          </string-name>
          .
          <source>Introducing oauth 2</source>
          .0, May
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>K. Knorr Cetina</surname>
          </string-name>
          .
          <article-title>Sociality with Objects: Social Relations in Postsocial Knowledge Societies</article-title>
          .
          <source>Theory Culture Society</source>
          ,
          <volume>14</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Patrice</given-names>
            <surname>Lopez</surname>
          </string-name>
          .
          <article-title>Grobid: combining automatic bibliographic data recognition and term extraction for scholarship publications</article-title>
          .
          <source>In Proceedings of the 13th European conference on Research and advanced technology for digital libraries</source>
          ,
          <source>ECDL'09</source>
          , pages
          <fpage>473</fpage>
          -
          <lpage>474</lpage>
          , Berlin, Heidelberg,
          <year>2009</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tamara M. McMahon</surname>
            ,
            <given-names>James E.</given-names>
          </string-name>
          <string-name>
            <surname>Powell</surname>
            , Matthew Hopkins, Daniel A. Alcazar,
            <given-names>Laniece E.</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
          </string-name>
          , Linn Collins, and
          <string-name>
            <surname>Ketan</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Mane</surname>
          </string-name>
          .
          <article-title>Social awareness tools for science research</article-title>
          .
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <volume>18</volume>
          (
          <issue>3</issue>
          /4),
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>David</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Millen</surname>
            , Jonathan Feinberg, and
            <given-names>Bernard</given-names>
          </string-name>
          <string-name>
            <surname>Kerr</surname>
          </string-name>
          . Dogear:
          <article-title>Social bookmarking in the enterprise</article-title>
          .
          <source>In Proceedings of the SIGCHI conference on Human Factors in computing systems, CHI '06</source>
          , pages
          <fpage>111</fpage>
          -
          <lpage>120</lpage>
          , New York, NY, USA,
          <year>2006</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>George</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Wordnet: a lexical database for english</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ):
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>November 1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Till</given-names>
            <surname>Nagel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Erik</given-names>
            <surname>Duval</surname>
          </string-name>
          .
          <article-title>Muse: Visualizing the origins and connections of institutions on co-authorship of publications</article-title>
          .
          <source>In Proceedings of the Science 2.0 for Technology Enhanced Learning Workshop</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Till</surname>
            <given-names>Nagel</given-names>
          </string-name>
          , Erik Duval, and
          <string-name>
            <given-names>Frank</given-names>
            <surname>Heidmann</surname>
          </string-name>
          .
          <article-title>Visualizing geospatial co-authorship data on a multitouch tabletop</article-title>
          .
          <source>In Proceedings of the 11th international conference on Smart graphics, SG'11</source>
          , pages
          <fpage>134</fpage>
          -
          <lpage>137</lpage>
          , Berlin, Heidelberg,
          <year>2011</year>
          . SpringerVerlag.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Peyman</surname>
            <given-names>Nasirifard</given-names>
          </string-name>
          , Sheila Kinsella, Krystian Samp, and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Decker</surname>
          </string-name>
          .
          <article-title>Social people-tagging vs. social bookmark-tagging</article-title>
          .
          <source>In Philipp Cimiano and H</source>
          . Pinto, editors,
          <source>Knowledge Engineering and Management by the Masses</source>
          , volume
          <volume>6317</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>150</fpage>
          -
          <lpage>162</lpage>
          . Springer Berlin / Heidelberg,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Tim O'Reilly</surname>
          </string-name>
          .
          <source>What is web 2</source>
          .0. Available online http://oreilly.com/web2/ archive/what-is-web-
          <volume>20</volume>
          .html,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Tim O'Reilly and John Battelle</surname>
          </string-name>
          .
          <source>Web Squared: Web</source>
          <volume>2</volume>
          .0 Five
          <string-name>
            <given-names>Years</given-names>
            <surname>On. Whitepaper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Reilly Media Inc</surname>
          </string-name>
          .,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Rob</surname>
            <given-names>Procter</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Robin</given-names>
            <surname>Williams</surname>
          </string-name>
          , James Stewart, Meik Poschen, Helene Snee, Alex Voss, and
          <string-name>
            <surname>Marzieh</surname>
          </string-name>
          Asgari-Targhi.
          <article-title>Adoption and use of Web 2.0 in scholarly communications</article-title>
          .
          <source>Phil. Trans. R. Soc. A</source>
          ,
          <volume>368</volume>
          :
          <fpage>4039</fpage>
          -
          <lpage>4056</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>Wolfgang</given-names>
            <surname>Reinhardt</surname>
          </string-name>
          .
          <article-title>Awareness Support for Knowledge Workers in Research Networks</article-title>
          . Available online at http: // bit. ly/ PhD-Reinhardt .
          <source>PhD thesis</source>
          , Open University of the Netherlands,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Wolfgang</surname>
            <given-names>Reinhardt</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muneeb I. Ahmad</surname>
          </string-name>
          , Pranav Kadam, Ksenia Kharadzhieva, Jan Petertonkoker, Amit Shrestha, Pragati Sureka, Junaid Surve, Kaleem Ullah, Tobias Varlemann, and
          <string-name>
            <given-names>Vitali</given-names>
            <surname>Voth</surname>
          </string-name>
          .
          <article-title>Exploration wissenschaftlicher Netzwerke und Publikationen mittels einer Multitouch-Anwendung [Exploration of Research Networks and Publications using a Multitouch Application]</article-title>
          . In Florian Klompmaker, Karten Nebe, and Nils Jeners, editors,
          <source>Proceedings of the 3rd Workshop Kollaboratives Arbeiten an interaktiven Displays [Collaborative Work on interactive displays] at the Mensch &amp; Computer Konferenz</source>
          <year>2012</year>
          ,
          <year>September 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Wolfgang</surname>
            <given-names>Reinhardt</given-names>
          </string-name>
          , Christian Mletzko, Hendrik Drachsler, and
          <string-name>
            <surname>Peter</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Sloep</surname>
          </string-name>
          .
          <article-title>Design and evaluation of a widget-based dashboard for awareness support in Research Networks</article-title>
          .
          <source>Interactive Learning Environments</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Wolfgang</surname>
            <given-names>Reinhardt</given-names>
          </string-name>
          , Christian Mletzko, Benedikt Schmidt, Johannes Magenheim, and
          <string-name>
            <given-names>Tobias</given-names>
            <surname>Schauerte</surname>
          </string-name>
          .
          <article-title>Knowledge Processing and Contextualisation by Automatical Metadata Extraction and Semantic Analysis</article-title>
          .
          <source>In Pierre Dillenbourg and Marcus Specht</source>
          , editors,
          <source>Proceedings of the 3rd European Conference on Technology Enhanced Learning (EC-TEL</source>
          <year>2008</year>
          ), Maastricht, The Netherlands,, volume
          <volume>5192</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>378</fpage>
          -
          <lpage>383</lpage>
          . Springer Berlin,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>Ben</given-names>
            <surname>Shneiderman</surname>
          </string-name>
          .
          <source>Science 2.0. Science</source>
          ,
          <volume>319</volume>
          (
          <issue>5868</issue>
          ):
          <fpage>1349</fpage>
          -
          <lpage>1350</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Bram</surname>
            <given-names>Vandeputte</given-names>
          </string-name>
          , Erik Duval, and
          <string-name>
            <given-names>Joris</given-names>
            <surname>Klerkx</surname>
          </string-name>
          .
          <article-title>Interactive sensemaking in authorship networks</article-title>
          .
          <source>In Proceedings of the 2011 ACM International Conference on Interactive Tabletops and Surfaces</source>
          , Kobe, Japan,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>M.M. Waldrop</surname>
          </string-name>
          .
          <source>Science 2.0. Scientific American</source>
          ,
          <volume>298</volume>
          (
          <issue>5</issue>
          ):
          <fpage>68</fpage>
          -
          <lpage>73</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Hanna</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Wallach</surname>
          </string-name>
          .
          <article-title>Conditional random fields: An introduction</article-title>
          .
          <source>CIS MS-CIS-04- 21</source>
          , University of Pennsylvania,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Matthew O. Ward</surname>
          </string-name>
          , Georges Grinstein, and Daniel Keim.
          <article-title>Interactive Data Visualization: Foundations, Techniques, and Applications</article-title>
          . Taylor &amp; Francis,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32. Fred Wilson. Email:
          <article-title>Social Media's Secret Weapon</article-title>
          . Available online http://articles.businessinsider.com/2011-05-15/tech/30100968_
          <article-title>1_ return-path-matt-blumberg-</article-title>
          <string-name>
            <surname>facebook</surname>
          </string-name>
          ,
          <year>May 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <given-names>Zhibiao</given-names>
            <surname>Wu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Martha</given-names>
            <surname>Palmer</surname>
          </string-name>
          .
          <article-title>Verbs semantics and lexical selection</article-title>
          .
          <source>In Proceedings of the 32nd annual meeting on Association for Computational Linguistics, ACL '94</source>
          , pages
          <fpage>133</fpage>
          -
          <lpage>138</lpage>
          , Stroudsburg, PA, USA,
          <year>1994</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>