<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The AIDA Dashboard: Analysing Conferences with Semantic Technologies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simone Angioni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angelo Salatino</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Osborne</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diego Reforgiato Recupero</string-name>
          <email>diego.reforgiatog@unica.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrico Motta</string-name>
          <email>enrico.mottag@open.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Mathematics and Computer Science, University of Cagliari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Knowledge Media Institute, The Open University</institution>
          ,
          <addr-line>Milton Keynes</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Scienti c conferences play a crucial role in the eld of Computer Science by promoting the cross-pollination of ideas and technologies, fostering new collaborations, shaping scienti c communities, and connecting research e orts from academia and industry. However, current systems for analysing research data do not provide a good representation of conferences. Speci cally, these solutions do not allow to track research trends, to compare conferences in similar elds, and to analyse the involvement of industrial sectors. In order to address these limitations, we developed the AIDA Dashboard, a tool for exploring and making sense of scienti c conferences which integrates statistical analysis, semantic technologies, and visual analytics.</p>
      </abstract>
      <kwd-group>
        <kwd>Scholarly Data</kwd>
        <kwd>Knowledge Graphs</kwd>
        <kwd>Topic Detection</kwd>
        <kwd>Bib- liographic Data</kwd>
        <kwd>Scholarly Ontologies</kwd>
        <kwd>Research Dynamics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Scienti c conferences play a crucial role in the eld of Computer Science by
promoting the cross-pollination of ideas and technologies, fostering new
collaborations, shaping scienti c communities, and connecting research e orts from
academia and industry. For this reason, every signi cant eld is usually
associated with multiple conferences that help de ning its challenges and paradigms
and to coordinate the e ort of all the interested stakeholders.</p>
      <p>Therefore, understanding and monitoring Computer Science conferences is
an important task for editors, researchers, companies, research policy makers
and other users working in this space. Several applications and services already
provide a wide variety of functionalities to support the exploration of research
data and produce various kinds of analytics. These include Microsoft Academic
Graph, Semantic Scholar, Scopus, Web of Science, OpenCitations, and many
others. However, these systems tend to neglect conferences and o er only a very
Copyright c 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
limited set of relevant analytics, such as the number of papers or citations. In
the rst instance, they do not allow users to examine the trends of the relevant
research topics. It is thus di cult to assess what are the research challenges
that a conference is actually addressing and how its focus changed in time.
Secondly, there is poor support in comparing conferences to determine the best
performing ones in speci c elds. For instance, we would like to know which are
the main conferences in Semantic Web, how they compare in terms of average
citations or other metrics, and how their performance changes in the last few
years. A third limitation is that current systems do not report any analytics
about the industry involvement. Conversely, it can be argued that conferences
are the premium public venues in which industry and academia interact, hence
monitoring these dynamics is critical for assessing a conference.</p>
      <p>
        In order to address these limitations, we developed the AIDA Dashboard,
a tool for exploring and making sense of scienti c conferences which integrates
statistical analysis, semantic technologies, and visual analytics. The AIDA
Dashboard was developed in collaboration with Springer Nature for assisting editors
in assessing conferences, but it also supports several other use cases. It introduces
three novel features that state-of-the-art systems are currently lacking. First, it
associates to conferences a very granular representation of their topics from the
Computer Science Ontology (CSO)[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] 3 and uses it to produce several analytics
about its research trends over time. Second, it enables to easily compare and
rank conferences according to several metrics within speci c elds (e.g.,
Semantic Web) and time-frames (e.g., last ve years). Finally, the AIDA Dashboard
o ers several features for assessing the involvement of industry in a conference.
This includes the ability to focus on companies and their performance when
assessing organizations, to report the ratio of publications and citations from
academia, industry, collaborative e orts, and to distinguish industrial
contributions according to 66 industrial sectors (e.g., automotive, nancial, energy,
electronics) from the Industrial Sectors Ontology (INDUSO)4. A demo of AIDA
Dashboard is currently available at http://w3id.org/aida/dashboard.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The AIDA Dashboard</title>
      <p>The AIDA Dashboard is a web application that allows users to visualize several
kind of analytics about a speci c conference (see Figure 1). The backend is
developed in Python, while the frontend is in HTML5 and Javascript.</p>
      <p>
        The AIDA Dashboard builds on the Academia/Industry DynAmics [
        <xref ref-type="bibr" rid="ref1 ref2">2,1</xref>
        ]
knowledge graph (AIDA)5, a large knowledge base describing 14M articles and
8M patents in the eld of Computer Science according to the research topics
drawn from CSO. 4M articles and 5M patents are also classi ed according to
the type of the author's a liations (academy, industry, or collaborative) and
66 industrial sectors drawn from INDUSO, which was speci cally designed to
      </p>
      <sec id="sec-2-1">
        <title>3 CSO - https://cso.kmi.open.ac.uk/ 4 INDUSO - http://w3id.org/aida/downloads/induso.ttl 5 AIDA - http://w3id.org/aida</title>
        <p>support AIDA. AIDA was generated by integrating several knowledge graphs
and bibliographic corpora, including Microsoft Academic Graph (MAG),
Dimensions, DBpedia, CSO, and the Global Research Identi er Database (GRID).</p>
        <p>
          The research papers were annotated with CSO topics using the CSO
Classi er [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]6, which is a tool that uses part-of-speech tagging to identify promising
terms and then exploits word embeddings to infer semantically related topics
from CSO. In addition, to extract further relevant topics, the classi er includes
also all their super topics according to the CSO. For instance, a paper tagged
with Neural Networks would be assigned the topic Arti cial Intelligence. This
solution enables identifying high level topics that are not typically mentioned in
the documents. The CSO Classi er powers the current version of the Smart Topic
Miner [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], which is the application used by Springer Nature to semi-automatically
annotate Proceedings books in the eld of Computer Science. Since CSO is often
updated, the set of topics used by AIDA is also evolving, constantly including
new emerging topics. As an example, topics can be extended by using hyperlinks
present in papers that might become Semantic Web entities or properties [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Each research article was also linked to the industrial sectors described in
INDUSO by mapping the a liations of the authors to their DBpedia entities,
which in turn are mapped to INDUSO. For instance, an article that was written
by authors who have Toyota as a liation would be associated to the industrial
sector Automotive.</p>
        <p>
          AIDA is available at http://w3id.org/aida under the CC-BY 4.0 license.
It was recently used for supporting the generation of adavanced analytics about
research dynamics and forecasting the impact of research topics on industry [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
However, using these data was not easy for less technical-savvy users. AIDA
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>6 CSO Classi er - https://pypi.org/project/cso-classifier/</title>
        <p>Dashboard is the rst step in allowing users to access AIDA through a
userfriendly but comprehensive interface.</p>
        <p>In order to support the AIDA Dashboard, we pre-computed a full set of
analytics for each conference from AIDA-KG and store it in a JSON le that
will be loaded by the web interface. This solution allows AIDA Dashboard to
be extremely scalable, since for a given conference it needs to query the server
only once, to retrieve its associated le. Every other operation is handled by the
front-end.</p>
        <p>AIDA Dashboard is highly scalable and allows to browse the di erent facets of
a conference according to seven tabs: Overview, Citation Analysis, Organizations,
Authors, Topics, Similar Conferences, and Industry.</p>
        <p>Figure 1 shows the Overview tab. This is the main view of a conference
that provides introductory information about its performance, the main authors
and organization, and the conference rank in its main elds in terms of average
citations for paper during the last ve years.</p>
        <p>The Citation Analysis tab reports the evolution in time of several
citationbased metrics such as the impact factor and the average citations for paper. It
also shows the evolution of the rank and the percentile of the conference in
di erent elds. For instance, the Conference on Neural Information Processing
Systems (NeurIPS) is currently the second conference in terms of average
citations in Neural Network, the third in Machine Learning, and the twelfth in
Arti cial Intelligence. This visualization is typically used by Springer Nature
editors to assess the performance of conferences within di erent communities and
to identify emerging conferences.</p>
        <p>The Organizations and Authors tabs show several analytics about the
main institutions and researchers active in the conference. Organizations can be
ltered according to their type (academia or industry) and are associated with
their number of publications, citations, and average citations for paper. The
researchers are associated with similar analytics, but also with their H-index and
H5-index, in order to quickly identify high impact researchers. Editors use this
information to understand the quality of researchers and organizations attracted
by the conferences. This is particularly important for assessing relatively young
conferences that may not have developed yet a strong citation record.</p>
        <p>The Topic tab allows users to analyse the topic trends in time. Speci cally it
shows two selections of topics: frequent topics and ngerprint topics. The rst is
the set of topics which appear more frequently in the conference. The second is
the set of most distinctive topics of the conference. It is obtained by computing
the di erence between the topic distribution of the conference and the one of the
full dataset. Preliminary analyses revealed that this second set is usually able to
better represent the topics considered central to the conference.</p>
        <p>The Similar Conferences tab compares the conference under analysis with
all the other conferences in the same elds according to their number of
publications, citations, and average citations for paper. The user can contextualise the
comparison to di erent elds. For example, ISWC can be compared with all the
other conferences in the elds of Semantic Web, Internet, or Computer Science.</p>
        <p>Finally, the Industry tab reports the percentage of publications and
citations from academia, industry, and collaborative e orts as well as the industrial
sectors analysis. The latter shows the percentage of produced publications and
citations received by companies in di erent sectors. For instance, the main
industrial sectors of ISWC are Computing and IT, Information Technology,
Management, Telecommunication, and Health Care.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>The current version of AIDA Dashboard already provides an array of interesting
functionalities, many of which go beyond what is available in other current tools.
Nevertheless, we are still at a relatively early stage and we are planning to
introduce new ones. As rst step, we plan to add a geographical tab for analysing
the distribution of countries active in a conference. We also want to expand
the set of entities that could be analysed by the dashboard, producing similar
analytics also for journals, organizations, and scienti c communities. Finally, we
plan to perform a comprehensive user study with editors and researchers from
di erent communities in order to assess the system and collect useful feedback.
For such a purpose, to generalize the presented dashboard we only need to replace
the CSO ontology with others within the domain under study. As such, we
have already started working with the MeSH ontology within the bio-informatics
domain to have our dashboard working in that domain as well.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Angioni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salatino</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Recupero</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Integrating knowledge graphs for comparing the scienti c output of academia and industry</article-title>
          .
          <source>In: Proc. of the ISWC 2019 Satellite Tracks. CEUR Workshop Proceedings</source>
          , vol.
          <volume>2456</volume>
          , pp.
          <volume>85</volume>
          {
          <issue>88</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Angioni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salatino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reforgiato Recupero</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Integrating knowledge graphs for analysing academia and industry dynamics</article-title>
          .
          <source>In: ADBIS, TPDL and EDA 2020 Common Workshops and Doctoral Consortium</source>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Presutti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nuzzolese</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Consoli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gangemi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Recupero</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>From hyperlinks to semantic web properties using open knowledge extraction</article-title>
          .
          <source>Semantic Web</source>
          <volume>7</volume>
          (
          <issue>4</issue>
          ),
          <volume>351</volume>
          {
          <fpage>378</fpage>
          (
          <year>2016</year>
          ). https://doi.org/10.3233/SW-160221
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Salatino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.:
          <article-title>Research ow: Understanding the knowledge ow between academia and industry</article-title>
          .
          <source>In: Knowledge Engineering and Knowledge Management</source>
          . Springer International Publishing (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Salatino</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birukou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.:
          <article-title>Improving editorial work ow and metadata quality at springer nature</article-title>
          .
          <source>In: The Semantic Web { ISWC 2019</source>
          . pp.
          <volume>507</volume>
          {
          <fpage>525</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Salatino</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thanapalasingam</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The cso classi er: Ontology-driven detection of research topics in scholarly articles</article-title>
          .
          <source>In: Digital Libraries for Open Knowledge</source>
          . pp.
          <volume>296</volume>
          {
          <fpage>311</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Salatino</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thanapalasingam</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mannocci</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The computer science ontology: a large-scale taxonomy of research areas</article-title>
          . In: International Semantic Web Conference. pp.
          <volume>187</volume>
          {
          <fpage>205</fpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>