<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>BibSLEIGH: Bibliography of Software (Language) Engineering in Generated Hypertext</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vadim Zaytsev vadim@grammarware.net</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Copyright c 2015 by the paper's authors. Copying permitted for private and academic purposes. This volume is published and copyrighted by its editors. In: A.H. Bagge, T. Mens (eds.): Postproceedings of SATToSE 2015 Seminar on Advanced Techniques and Tools for Software Evolution, University of Mons, Belgium</institution>
          ,
          <addr-line>6-8</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universiteit van Amsterdam</institution>
          ,
          <addr-line>The Netherlands Raincode</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The body of research contributions is vast and full of papers. Existing projects help us navigate through it and relate authors to papers and papers to venues. In this paper we list features missing from those projects and propose a solution in the form of BibSLEIGH a work in progress on facilitated browsing of scientic knowledge objects. Through leveraging domain focus, by actively employing automated data collection and scraping tools, and with automated annotating of the corpus, we are able to gain and provide insights into scientic communities and topics, as well as surface potential interdisciplinary opportunities.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>BibSLEIGH has started in 2014 as a project to scratch
some personal itches and solve problems that were
eating away from the authors’ time as well as anyone
else’s. These issues can be broadly categorised into
four categories. In 1.1, we will discuss in some detail
problems with the bibTEX format and the unnecessary
diversity of conventions for equivalent items, which has
a chance of making academic publications look
unprofessional and can also lead to confusion and mistakes.
However, consistency enforcing is very time
consuming. In 1.2, the focus will be on domain specicity,
which is a specialisation, and just as any specialisation,
can lead to signicant optimisation. We have collected
some features in 1.3 that are missing from the
current widespread remedies (we refuse to call them
solutions). Each of the features is missing for a good
reason: each requires research, development and
domain focus. This makes them both attractive to invest
eort in and dangerous because most are non-trivial.
Finally, in 1.4 the most obvious point will be raised
about information that is interesting in bibliographical
context, being distributed over various unconnected
sources of not that structured data.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>BibTEX non-uniformity across sources</title>
      <p>If we attempt to download .bib les for the same
publication from various sources, they will all look
dierently, sometimes drastically so. Many
publishers do not curate their data, rely on automatic text
recognition and only occasionally and serendipitously
x misspellings. BibT EX providers are often volatile
when it comes to conference naming. IEEE and
ACM are obviously inclined to include their
aliation (the IEEE/ACM international conference on...),
sometimes in favour of more useful information like
the number of the conference in the series. DBLP
has changed their policy on abbreviating venue names
during the period of writing this paper (between
SATToSE in July 2015 and post-proceedings in November).</p>
      <p>When information is available, bibTEX providers
usually decide to include it yet what was the last
time someone cared about whether ESOP 1986 took
place in Saarbrcken or in Passau? This information
can be leveraged for other purposes, like tracking
country and continent preferences and their shifting over
the years, or investigating the impact of location on
the number, quality and aliation of papers. However,
it is not used for any of those purposes, yet included in
the bibliographical entry. Nevertheless, many details
about in which hotel near which city on which exact
days the conference has taken place, nd their way into
bibTEX, even though they were important only for the
briefest of times, and only to immediate attendees of
the event.</p>
      <p>So, on one hand, there is too much information
in the bibTEX entries supplied by publishers and
accumulators like DBLP and Google Scholar:
addresses, dates, timestamps, keywords, sometimes
entire abstracts. On the other hand, however, some
of more useful information is routinely missed.
Frequent omissions concern editor names and
hyperlinks that can be used to access the actual content
of the publication. Editor names play exactly the
same role in events and journal special issues as
author names play in individual publications: they help
to identify the item but also establish community
links across dierently named and formally unrelated
events. Hyperlinks are not always entirely missing,
but oftentimes hidden behind non-standard elds like
ee or acmid; not curated in a way that a doi eld
sometimes starts with http://; and even outdated
most if not all links like http://www.computer.
org/proceedings/csmr/0546/05460161abs.htm
being provided by DBLP have been dead (HTTP Status
404) for several years since the redesign of the IEEE
Computer Society website made them obsolete.</p>
      <p>Time lost in reformatting is only a part of this
side of the problem. Inconsistencies lead to
unprofessional look of those papers whose authors have
decided against wasting time on bibliography
beautication; and worse yet to duplicate entries appearing
within the same paper with slight variations in spelling
and data details provided, which made searching for
the right entry harder and clone detection impossible
within a typical textual editor.
1.2</p>
      <sec id="sec-2-1">
        <title>Lack of domain focus</title>
        <p>Academic researchers tend to specialise but never limit
themselves overly to one particular series of events.
Yet, when we look at sources of information we have
at our disposal, they come in two sizes only. On one
extreme we have websites devoted to individual
conferences. They usually contain a lot of information
that is not immediately required for a decent bibTEX
entry, but can be quite useful in the long run for
community recognition: after all, one is much more likely
to submit to a conference chaired by someone whose
name they recognise and whose work they can relate
to that of their own. Organisation committee details
and programme committee members provide
refreshingly large foundation for automation of this process,
as demonstrated by the recent work of Vasilescu et
al. [VSM13, VSM+14] that harvested PC members of
several top conferences and cross-checked them with
authors publishing there to measure academic
inbreeding. However, the focus of such a website is limited to
one event, or in some lucky cases to a series of events,
and such websites are very prone to disappearing
forever once their organisers retire or change employers.</p>
        <p>As the other extreme we have services that make an
endeavour to collect information over a broad choice of
conferences on all kinds of topics, and put them in one
place for display and consumption. The most famous
ones are DBLP with its 6500+ venues, Google Scholar
which is based on web crawling and Microsoft
Academic Search that contains ranking tables sorting
conferences of one eld by the number of citations their
articles enjoyed over the years. Such services try to be as
general and comprehensive as possible, and this is
exactly where they fail short. Broad generalisations are
impossible without compromises on metadata models,
on information representation, on clone detection. A
website of one particular conference typically shows
very clearly which volume of which journal contains its
post-proceedings special issue while DBLP
habitually gives you all issues of the conference and all issues
of all journals and leaves the search for a match in your
own hands. University libraries fall into the same
category: while limiting their databases to material
available physically or through subscriptions, they do not
dierentiate among domains, so searching for
mutation will likely result in many items unrelated to
mutation testing; and searching for graph, while more
productive, will still yield results from graph
transformation research as well as from general graph theory.</p>
        <p>The quest for broad coverage makes the project
vulnerable. For instance, DBLP covers millions of authors
and thus has to be extremely careful about not
confusing authors with similar names however, many
researchers, especially in the pre-google era, did not
write their names always in the same fashion. This
would have been known to domain experts who are
familiar with key authors in their eld, but domain
knowledge does not scale up. Similarly, Google Scholar
relies on its web crawler, and so it is not uncommon for
it to point you to papers that are no longer available or
are in fact no papers at all, no matter what their
authors claim. Microsoft Academic Search is based on
citation information and as a result of dierent people
citing the same venue in dierent ways (e.g., with
International Conference or without it), the same venue
appears several times in the ranking, both positioned
much lower than they deserve.
1.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Missing features</title>
        <p>When we like a paper, we often begin investigating
its authors to see if they have contributed to
similar lines of research before or after. DBLP lookup
has become a part of a routine check in many cases
from research exploration to job candidate evaluation.
However, a graph transformation researcher that
occasionally published a model transformation paper, or a
grammarware engineer masquerading as a metamodel
evolution contributor, will have dierent styles across
other of their papers, and might not be as fruitful to
investigate if your interest is particular and your time
budget is limited. What could have helped here is
visualisation beyond textual : instead of browsing
through a multi-page wall of text prole on DBLP,
some of us would have wanted to take a quick look
at a diagram depicting community contribution in a
concise and illustrative manner.</p>
        <p>Natural language processing techniques have a
powerful arsenal: even the simplest analyses like
stemming and lemmatisation can provide great aid in
surfing through the ocean of papers to pick the right ones
to read and cite. It is common knowledge that the
names of conferences do not always completely
represent their intentions: having languages in the name
can mean one or two of a dozen of entirely dierent
research directions; venues with engineering in their
name can get quite science-y and theoretical, just as
a name starting with trends does not mean all
papers are surveys, overviews and vision statements. To
the best of our knowledge, no currently existing
bibliographic website currently provides a lot of NLP-based
features, although ACM Digital Library has recently
started collaborating with IBM Watson to pursue that.</p>
        <p>Scraping older sources from document scans to
websites that fell apart decades ago and have their
ruins exposed though the Wayback Machine, is
usually beyond the goals and capabilities of bibliographic
websites. Armed with domain knowledge and the
interest seriously linked to that domain, we can gather
enough eort to complete such endeavours and ask
senior and emeritus colleagues directly about that
one long-forgotten obscure workshop that a reputable
conference has grown from.</p>
        <p>Grouping and clustering of conferences is
usually either manual work, or done though event
colocation, or not done at all. The rst option is
labour-intensive, error-prone, vulnerable to biases and
prejudice. The second option delivers complications
for roaming venues like BX (deliberately co-locating
each year with a dierent community: ETAPS, STAF,
VLDB, etc) and for diverging venues that stopped
colocating deliberately to emphasize pursuing a
divergent path. The third option is not an option at all,
since even fairly focused researchers will nd
themselves contemplating submission to a dozen or two
reasonable venues. There is quite some space for
automated clustering.</p>
        <p>Topic-driven grouping is not the only kind of
classication that would be sensible for a bibliographic
portal: some venues are linked by a subcommunity of
people who strongly contribute to both. For instance,
there are many people who publish regularly both at
MoDELS and ICSME/SCAM, even though they
cannot attend both within the same year (they happen
simultaneously). Having linked data about people’s
contributions, we can surface such relations and
some RDF frontends to DBLP let you do that with a
couple of medium-size SPARQL queries.</p>
        <p>All that being said in 1.1 about the state of
bibTEX entries obtainable from available sources, we
still want to have some freedom in formatting :
everyone in computer science research knows what LNCS
is; in a paper submitted to SLE one does not need
to explain this abbreviation; editor names are nice to
have but sacriceable under pressing space constraints,
etc. We want exible bibT EX formatting: DBLP
provides you with some very limited options (crossref or
no crossref); IEEE Xplore and Elsevier as well
(abstract or no abstract); but BibSLEIGH even in its very
beginning stage provides its users with more freedom.</p>
        <p>Desktop software for managing bibliographies like
Mendeley has tagging functionality that can help its
users to annotate the papers they read into dierent
categories or add brief descriptions to them. However,
there is a huge gap between doing that and providing a
comprehensive annotated bibliography on the subject:
in fact, such contributions are rare and properly
treasured, for it takes a lot of expertise and work to craft
them. Unfortunately, there are much many topics and
subtopics than there will even be annotated
bibliographies. We need some semi-automatic way of providing
us with at least bundles of related papers if we
indicate the selection criteria.
1.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Distributed information</title>
        <p>It was already pointed out above that participating in
event organisation and serving in programme
committees can be seen as community binding and is therefore
metadata of interest. Yet, to the best of our
knowledge, there is no project currently dedicated to
collecting this kind of information, and it remains scattered
half over the internet and half in the Way Back
Machine.</p>
        <p>Mathematics Genealogy Project [C +] is a totally
disconnected project dedicated to documenting
topics of doctoral dissertations (and occasionally
habilitations) and supervisorship information. It certainly
has a merit of its own, but we believe it can also be
coupled with other kinds of metadata in a sensible way.</p>
        <p>Aliation information very occasionally nd its
way into DBLP as well as into Google Scholar where
academics can log in and update it (unfortunately,
some choose to log in and prohibit Google from ever
showing information about them), but there is no easy
way of tracking and leveraging it. However, it is not
outrageous to think of research dedicated to tracking
research centres of activities on particular topics over
the years.</p>
        <p>Finally, citation information it is available on
publishers’ websites in limited form (because they are
not big fans of sharing it among themselves) and on
Google Scholar (where it is heavily guarded against
any form of automated scraping). While
acknowledging some interest in it, we choose to avoid this aspect
for now, because it is not static by nature: citation
information available today can be totally out of date by
tomorrow. However, there is a lot of potential research
here that goes way beyond traditional bibliometrics:
for instance, we can identify canonical sources (which
often will be books, like the Dragon Book [ASU85])
that are used throughout a large fraction of papers in a
specic conference, and nd other venues in a dierent
language that have the tendency to cite translations of
this book.</p>
        <p>Additionally, academic articles also contain links to
web resources such as additional documentation, wikis
and tool repositories, and such links have a half life
of 4 years on average [Spi03]. The Software Heritage
Project was recently proposed by Roberto Di Cosmo as
a project to organise, preserve and share all
academically produced software to provide much desired
availability, traceability and uniformity. Unfortunately the
project seems to be in early stages, its call to action is
available on SlideShare [Cos15] but the project itself
is yet unknown to public search engines. It will be
interesting to see if the corpus of BibSLEIGH can be
automatically mined for references to tools and
clustered by technological space.
2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>BibSLEIGH to the rescue!</title>
      <p>BibSLEIGH is a work in progress. Keeping that in
mind, we would like to sketch preliminary
requirements and architecture decisions in 2.1, point out
some related work in 2.2 and describe the state of
the project as it is by the time of submission in 2.3.
Next, 3 will draft some possible future directions we
might decide to explore.
2.1</p>
      <sec id="sec-3-1">
        <title>Proposed solution In the centre of BibSLEIGH there is one centralised repository containing all its data in JSON format 4</title>
        <p>we call it LRJ, short for Lexically Reliable JSON,
because we store all key-value pairs one per line sorted
by keys. This was chosen over a more classic database
setup in order to allow individual traceable edits of
each piece of data and at the same time to guarantee
user responsiveness. Data is imported to this central
place through any of the existing importers, which are
usually implemented as iterative parsers (to process
the DBLP dump which is around 2 GB) or
webscrapers (at this moment we have those for individual DBLP
pages, CEUR and EasyChair). JSON les can also
obviously be added manually. There is also an ad-hoc
importer that creates appropriate JSON entities from
a list it reads from a textual le this helps to
properly add ancient entries.</p>
        <p>Once the data is in the repository, it can be
further curated, normalised, improved, enhanced and
crosschecked with other sources. Typical maintenance
activities include adding a fresh issue of an already
known conference or a journal issue known to be
related to one of the known conferences (automated:
one just needs to run an incremental updater),
improving the name of the proceedings booktitle
(semiautomated: changed manually at the top and
automatically propagated downwards), removing
nonacademic clutter such as forewords and panel
summaries (manually or heuristic-based). As an example
of crosschecking we can talk about adding PC
members and organisers: this information is never found on
DBLP, but can be harvested elsewhere and integrated
into the same system.</p>
        <p>Once normalisation reaches a point of being a valid
input for analysis, we enrich the data by stemming all
titles and tagging them by predened tags following
the spirit of the rest of the project, each tag has its own
denition stored in a separated JSON le which can
be accessed, inspected and changed right on GitHub.
Stemming provides fully automated foundation to
naturally link papers to their conceptual neighbours, tags
play the same role for previously known manually
dened concepts (so that -lifting falls under the same
tag as -calculus, but -kernel is kept away from
calculus, even though the characters look similar 1).
Each tag denition can contain links to Wikipedia,
Wikidata and other places that are displayed on the
tag’s webpage. Stems can only rely on automatically
derivable information, so their webpages display
neighbours stems that are commonly used together with
them.</p>
        <p>1As a side remark, in Unicode these are dierent symbols:
kernel is read as microkernel and therefore uses the micro sign
character (U+00B5), while -calculus is read as mu-calculus
and is thus represented by the Greek small letter mu (U+03BC).
BibSLEIGH is the only website that gets it right in all places,
the readers are welcome to check.</p>
        <p>Whenever the central dataset of BibSLEIGH is
needed for inspection, it is formatted as a collection of
almost-static XHTML pages: the only dynamic part of
them is the pretty-printing of bibTEX itself. The
outlook of BibSLEIGH is less austere than that of DBLP,
it makes full use of a palette of colours and a collection
of icons for each covered brand of conferences.
2.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Related work</title>
        <p>In the eld of High-Energy Physics there has been a
movement concerning long time preservation of
publications, datasets, repositories and relations between
them [GMH+09, GMB10, AAA+12, Sou13], and there
is a prospering project called INSPIRE-HEP at http:
//inspirehep.net . It covers a dierent domain than
software (language) engineering, but otherwise partly
addresses the same problems we have pointed out. It
does oer additional functionality such as job listings
and does not intend to cover some of our goals such as
visualisations.</p>
        <p>ACM Digital Library in recent collaboration with
IBM Watson has started to provide feature called
Concept Insights. For each paper, two things can
be explored: concepts in this article that links
glossary terms mined from the full text of the
paper, to their denitions on Wikipedia and recent
authors with related interests that visualises people
who recently published something that share these
concepts. This functionality is certainly welcome,
even though it remains to be seen how such
automated concept matching can compete with and
complement manual research eorts in taxonomies that
try to identify key publications and tie them with
key concepts and relations between them:
examples exist for taxonomies of domain specic aspect
languages [FDNT15], reverse engineering [CC90],
reverse architecting [PDP +07], (un)parsing [ZB14],
algorithm animated visualisation [KKM06], security
topics [KLS09]. Information retrieval research has also
demonstrated promising results in helping to select
features for automated induction [YC09, LWT08] and
renement [HZL06, Nov07] of taxonomies, which we
have not yet explored.</p>
        <p>One step farther from bibliographical
repositories there are model repositories such as FMI (Free
Model Initiative) [SHK14], ReMoDD (Repository
for Model Driven Development) [FBM +12], CDO
(Connected Data Objects) [Ecl09], Atlantic
Metamodel Zoo [Atl05], Grammar Zoo [Zay15],
GenMyModel [Gen14], that are on a quest of collecting
models for various purposes. There are quite a
number of initiatives related specically to community
management and facilitation: DBLP [Ley02],
Reengineering wiki [vDV02], Researchr [VVvC09], Research
2.0 [ABFM09], SL(E)BOK, etc. They usually combine
requirements elicitation with experience reports with
calls to arms. One of those very similar to ours is
MetaScience [CCCB14] unlike BibSLEIGH that mainly
aims at cross-referencing various information sources
and using domain knowledge, MetaScience is focused
exclusively on automatically deriving metadata such
as coauthor graphs and pages published per year, and
contains impressive interactive visualisations of it.</p>
        <p>Linked data is an initiative that started in the
semantic web community and has gained a lot of
attention over the decade of its existence. The idea
revolves around uniform identication of entities by
URIs and uniform encoding of a graph of their
relations as a collection of subject-predicate-object triples.
They have standard formats for specifying the triples
(mostly RDF or Turtle), languages for querying them
(nowadays mostly SPARQL) and over half a thousand
open datasets containing up to several billion of such
triples [CJ14]. There is research evidence backed up
by operational prototypes, that points to usefulness
of linked data for many related tasks from
connecting community heritage [WNB +15] to mining software
repositories [KFH+12].
2.3</p>
        <p>Terminology and current state of
BibSLEIGH
By domain we mean a top group of conferences: the
front page of BibSLEIGH displays logos of its domains.
Right now they are dened ad-hoc with the help of
some domain knowledge; in the future we will use
automated clustering techniques to form such domains.
A brand is a series of events with continuing numbering
and, more often than not, the same name. One event
can belong in several brands: a brand of MoDELS
covers the UML series because they kept the numbering,
but events of the brand LDTA and ATEM belong only
to the domain of SLE, but not to the brand SLE. Each
proceedings entity is called an issue : usually it is
regular conference proceedings issue, but it can also be a
journal special issue. Multi-volume proceedings have
one issue per volume because bibTEX entries for such
volumes are dierent. A tag is a predened term such
as context-free grammar or visual notation
specied as a set of matching rules covering spelling variants
and synonyms (so a paper with graphical notation in
the title will be tagged with visual notation). There
are several style-dening tags like question (the
title ends in a question, like Can Programming Be
Liberated from the Von Neumann Style?), towards
(like Towards Incremental Execution of ATL
Transformations), considered harmful, past, present and
future, etc. Interestingly, one of the most popular
tags (covering around 7.2% of all papers) is named,
which corresponds to the pattern of starting the title
with a word followed by a colon or an em-dash like
Lilith: A Personal Computer for the Software
Engineer, or Miranda: A Non-Strict Functional language
with Polymorphic Types, or GHC: Operational
Semantics, Problems, and Relationships with CP (#; j).
Currently tags are created based on titles only, because
that information is indisputably in the public domain
and can be used fairly; there is an ongoing discussion
about fair use of abstracts and keywords, but
technically they can be harvested as well, so we plan to do so
(perhaps not committing the results of such harvest to
public repositories to avoid copyright claims). A word
is what we call a stem obtained from a classic Snowball
stemmer for English. We use our own lexer that tries
to split camelcased words properly: not just
CamelCase to Camel and Case, but also APIExplorer
to API and Explorer and XSDtoMOF to XSD,
to and MOF (it also leaves JavaScript intact!).
Figure 1 shows a typical use of a word link. A role is
some facilitating role a person has played in an issue:
being an editor, a keynote speaker, a PC member, etc.,
are roles.</p>
        <p>By the time of submission of this paper,
BibSLEIGH covered 166 brands in 26 domains,
summarised on Table 1. There are 2726 issues of these
brands with 144589 papers in total. There are
currently 684 tags with 354720 markings. The total
vocabulary is 24359 stems derived from 1183492 words.</p>
        <p>The oldest entry so far is the First International
LISP Conference held in 1963 in MØxico, with
attendees like John McCarthy and Marvin Minsky. It has
mostly historical value, but a nice part was that it was
possible to surface most of the papers and reconstruct
metadata by googling and scraping. This issue is not
present on DBLP.</p>
        <p>Many mistakes in DBLP data (and sometimes in
publishers’ data) were corrected because they were
becoming quite apparent once automated processing
began: the longest stems were words erroneously
glued together; matching heuristics work
reasonably well to equate dierent spellings of diacritical
names, etc. An example of DBLP mismatch could
be seen by comparing http://dblp.uni-trier.
de/db/conf/edoc/edoc2007.html to http:
//bibtex.github.io/EDOC-2007.html : except for
10.1109/EDOC.2007.42 and 10.1109/EDOC.2007.44,
all DOIs at DBLP are incorrect but xed at
BibSLEIGH. This was spotted automatically by reporting
that some entries in this issue had no page
information; an attempt to x it revealed a mismatch
between DBLP and IEEE Xplore. DOI information
is usually reliable; we know of only one
counterexample: http://doi.ieeecomputersociety.org/10.
1109/ICSM.1997.624246 resolves successfully, but
http://dx.doi.org/10.1109/ICSM.1997.624246
does not.</p>
        <p>BibSLEIGH contains proles on 150454 people,
some of them might erroneously view several
namesakes as one person no noticeable attention was
devoted to this issue so far. Some scraping for roles
has begun, so far we have 4154 roles, which is
almost 10 times the size of the dataset of Vasilescu et
al. [VSM13], but still around 5% of total work if we
optimistically estimate 10 organisers and 20 PC members
on average per issue. Figure 3 and Figure 4 show two
examples of person proles, with corresponding
narrations in the captions. Notice how the prole is
interpreted without the usual bibliometric remarks about
the number of papers!</p>
        <p>Exploring the rest is left as an exercise to the reader:
http://bibtex.github.io web front end
http://github.com/slebok/bibsleigh
tially curated JSON data
http://github.com/bibtex/bibsleigh
JSON refactorings and visualisations
parWhat makes BibSLEIGH become more than a gloried
wrapper for DBLP is harvesting its domain specicity
and community specicity. While keeping the
automated, semi-automated and heuristic-based
transformations as maintenance activities, we can continue
ingraining the bibliographic entities and their groups
with information relating them to one another, as
well as to concepts, methods, frameworks, approaches,
toolkits, datasets. Implementing various distance
metrics, as well as annotating them manually or
automatically with topic information can aid clustering and
linking beyond traditional methods depending on the
citation information. We see this as another step
towards the construction of a body of knowledge for the
domain of software language engineering (SLEBoK).</p>
        <p>Expansion of the BibSLEIGH data set will continue,
but not far: most interesting next steps involve
strategically adding special issues and role annotations to
already imported conferences. We are afraid that overly
eager expansion will deprive us of the main advantage
of being domain-specic. However, if we could nd a
way to eventually hide irrelevant parts from sight so
that a user can productively focus on a reasonable
subset, that could solve the problem and open the door
wider for interdisciplinary growth of this project.</p>
        <p>Navigational support at the current stage of
development is already quite strong: domains, brands, tags
and words let you browse through thousands of papers
quite easily to nd that dozen that you are interested
in. However, we believe this can be improved further
through adding annotations, leveraging metadata,
proper visualisations, ground-based ranking and
clustering, etc.</p>
        <p>At BibSLEIGH’s webpage the project is called
facilitated browsing of scientic knowledge .
Indeed, providing interactive access to the curated
annotated corpus of academic papers on programming
language theory, compiler construction,
metaprogramming, software evolution and analytics, refactoring and
other related topics can serve as an entrance point into
the research domain as well as the foundation for some
metaresearch activities. Software engineering Master
students at the University of Amsterdam have already
started using BibSLEIGH actively in their studies.</p>
        <p>It remains to be seen which open problems of
software language engineering can this project contribute
to solving [BZ15]. SLE, besides being a subdomain of
software engineering, is known to be a bridging area
of research, where a fair share of activities is devoted
to seeking similarities between technologies and
technical spaces, and to developing techniques with wide
and cross-space applicability. However, even within
one space reaching a point of soundly relating concepts
can take substantial time and eort consider laying
relations between attribute grammars and ax
grammars [Kos91] or between object algebras to attribute
grammars [RBO14]. We will try to push BibSLEIGH
towards facilitating this, and any help is welcome.
[ASU85]
[Atl05]
[BZ15]
[C+]
[CC90]</p>
      </sec>
      <sec id="sec-3-3">
        <title>A. V. Aho, R. Sethi, and J. D. Ullman. Compilers: Principles, Techniques and Tools. Addison-Wesley, 1985.</title>
      </sec>
      <sec id="sec-3-4">
        <title>AtlanMod. Atlantic Metamodel Zoo, 2005. http://www.emn.fr/z-info/ atlanmod/index.php/Zoos .</title>
      </sec>
      <sec id="sec-3-5">
        <title>Anya Helene Bagge and Vadim Zaytsev.</title>
        <p>Open and Original Problems in Software
Language Engineering 2015 Workshop
Report. SIGSOFT Software Engineering
Notes, 40:3237, May 2015.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Harry Coonce et al. Mathematics Genealogy Project. http://www.genealogy. ams.org.</title>
      </sec>
      <sec id="sec-3-7">
        <title>Elliot J. Chikofsky and James H. Cross II. Reverse Engineering and Design Recovery: A Taxonomy. IEEE Software, 7(1):1317, 1990.</title>
        <p>[CJ14]
[Cos15]
[Ecl09]</p>
      </sec>
      <sec id="sec-3-8">
        <title>Richard Cyganiak and Anja Jentzsch. The Linking Open Data Cloud Diagram, 2014. http://lod-cloud.net .</title>
      </sec>
      <sec id="sec-3-9">
        <title>Roberto Di Cosmo. Ten Years Analysing Large Code Bases: A Perspective. http: //tinyurl.com/z44ydlw , 2015. EvoLille 2015.</title>
      </sec>
      <sec id="sec-3-10">
        <title>Eclipse. CDO (Connected Data Objects) Model Repository, 2009. https: //eclipse.org/cdo/ .</title>
        <p>[FBM+12] Robert B. France, James M. Bieman,
Sai Pradeep Mandalaparty, Betty H. C.</p>
        <p>Cheng, and Adam C. Jensen.
Repository for Model Driven Development
(ReMoDD). In Martin Glinz, Gail C. Murphy,
and Mauro PezzŁ, editors, Proceedings of
the 34th International Conference on
Software Engineering, pages 14711472. IEEE,
2012.
[FDNT15] Johan Fabry, Tom Dinkelaker, Jacques
NoyØ, and ric Tanter. A Taxonomy
of Domain-Specic Aspect Languages.</p>
        <p>ACM Computing Surveys , 47(3):40:1
40:44, February 2015.
[Gen14]
[GMB10]</p>
      </sec>
      <sec id="sec-3-11">
        <title>GenMyModel, 2014. //repository.genmymodel.com . https:</title>
      </sec>
      <sec id="sec-3-12">
        <title>Anne Gentil-Beccot, Salvatore Mele, and Travis C. Brooks. Citing and Reading Behaviours in High-energy Physics. Scientometrics, 84(2):345355, 2010.</title>
        <p>[GMH+09] Anne Gentil-Beccot, Salvatore Mele,
Annette Holtkamp, Heath B. O’Connell, and
Travis C. Brooks. Information resources
in high-energy physics: Surveying the
present landscape and charting the future
course. JASIST, 60(1):150160, 2009.
[HZL06]</p>
      </sec>
      <sec id="sec-3-13">
        <title>Ruizhang Huang, Zhigang Zhang, and</title>
        <p>Wai Lam. Rening hierarchical
taxonomy structure via semi-supervised
learning. In Proceedings of the 29th
International ACM SIGIR Conference on
Research and Development in Information</p>
        <p>Retrieval, pages 653654. ACM, 2006.
[KFH+12] Iman Keivanloo, Christopher Forbes,
Aseel Hmood, Mostafa Erfani,
Christopher Neal, George Peristerakis, and
Juergen Rilling. A Linked Data Platform
for Mining Software Repositories. In
Proceedings of the Ninth IEEE Working
Conference on Mining Software
Repositories, pages 3235. IEEE Computer
Society, 2012.
[KKM06]
[KLS09]
[Kos91]
[Ley02]
[LWT08]
[Nov07]</p>
      </sec>
      <sec id="sec-3-14">
        <title>Ville Karavirta, Ari Korhonen, and Lauri</title>
        <p>Malmi. Taxonomy of Algorithm
Animation Languages. In Proceedings of the
ACM Symposium on Software
Visualization, pages 7785. ACM, 2006.</p>
      </sec>
      <sec id="sec-3-15">
        <title>Justin King, Kiran Lakkaraju, and</title>
        <p>Adam J. Slagell. A Taxonomy and
Adversarial Model for Attacks Against
Network Log Anonymization. In Sung Y. Shin
and Sascha Ossowski, editors, Proceedings
of the 24th Symposium on Applied
Computing, pages 12861293. ACM, 2009.</p>
      </sec>
      <sec id="sec-3-16">
        <title>C. H. A. Koster. Ax Grammars for Pro</title>
        <p>gramming Languages. In H. Alblas and
B. Melichar, editors, Attribute Grammars,
Applications and Systems , volume 545 of
LNCS, pages 358373. Springer, 1991.</p>
      </sec>
      <sec id="sec-3-17">
        <title>Michael Ley. The DBLP Computer Sci</title>
        <p>ence Bibliography: Evolution, Research
Issues, Perspectives. In Alberto H. F.
Laender and Arlindo L. Oliveira, editors,
Proceedings of the 9th International
Symposium on String Processing and
Information Retrieval, volume 2476 of LNCS,
pages 110. Springer, 2002.</p>
      </sec>
      <sec id="sec-3-18">
        <title>Yuefeng Li, Sheng-Tang Wu, and Xiao</title>
        <p>hui Tao. Eective Pattern Taxonomy
Mining in Text Documents. In
Proceedings of the 17th ACM International
Conference on Conference on Information
and Knowledge Management , pages 1509
1510. ACM, 2008.</p>
      </sec>
      <sec id="sec-3-19">
        <title>Vt NovÆcek. Imprecise Empirical Ontol</title>
        <p>ogy Renement Application to
Taxonomy Acquisition. In Jorge Cardoso,
JosØ Cordeiro, and Joaquim Filipe,
editors, Proceedings of the Ninth
International Conference on Enterprise
Informa</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [AAA+12]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Akopov</surname>
          </string-name>
          , Silvia Amerio, David Asner,
          <string-name>
            <given-names>Eduard</given-names>
            <surname>Avetisyan</surname>
          </string-name>
          , Olof Brring, James Beacham, Matthew Bellis, Gregorio Bernardi, Siegfried Bethke, Amber Boehnlein, Travis Brooks, Thomas Browder, Rene Brun, Concetta Cartaro, Marco Cattaneo, Gang Chen, David Corney,
          <string-name>
            <given-names>Kyle</given-names>
            <surname>Cranmer</surname>
          </string-name>
          , Ray Culbertson, Suenje Dallmeier-Tiessen, Dmitri Denisov, Cristinel Diaconu, Vitaliy Dodonov, Tony Doyle,
          <string-name>
            <given-names>Gregory P.</given-names>
            <surname>Dubois-Felsmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Ernst</surname>
          </string-name>
          , Martin Gasthuber, Achim Geiser, Fabiola Gianotti, Paolo Giubellino, Andrey Golutvin, John Gordon, Volker Guelzow, Takanori Hara, Hisaki Hayashii, Andreas Heiss, Frederic Hemmer, Fabio Hernandez, Graham Heyes, AndrØ G. Holzner,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Igo-Kemenes</surname>
          </string-name>
          , Toru Iijima, Joe Incandela, Roger Jones, Yves Kemp, Kerstin Kleese van Dam,
          <string-name>
            <surname>Juergen Knobloch</surname>
          </string-name>
          , David Kreincik,
          <article-title>Kati Lassila-Perini, and Francois Le Diberder</article-title>
          .
          <source>Status Report of the DPHEP</source>
          Study Group:
          <article-title>Towards a Global Eort for Sustainable Data Preservation in High Energy Physics</article-title>
          . CoRR, abs/1205.4667,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [ABFM09]
          <string-name>
            <given-names>Denis</given-names>
            <surname>Avrilionis</surname>
          </string-name>
          , Grady Booch, JeanMarie Favre, and
          <string-name>
            <surname>Hausi</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mller</surname>
          </string-name>
          .
          <source>Software Engineering 2.0 &amp; Research</source>
          <volume>2</volume>
          .0. In Patrick Martin,
          <string-name>
            <given-names>Anatol W.</given-names>
            <surname>Kark</surname>
          </string-name>
          , and Darlene A. Stewart, editors,
          <source>Proceedings of the conference of the Centre for Advanced Studies on Collaborative Research (CASCON)</source>
          , pages
          <fpage>353355</fpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>A</given-names>
            <surname>Taxonomy</surname>
          </string-name>
          .
          <source>IEEE Software</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1317</fpage>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [CCCB14]
          <string-name>
            <given-names>Javier</given-names>
            <surname>Canovas</surname>
          </string-name>
          , Valerio Cosentino, Jordi Cabot, and Robin Boncorps.
          <source>MetaScience: Analyzing the Research Prole of Authors, Conferences and Journals</source>
          ,
          <year>2014</year>
          . http://som-research.uoc.edu/ tools/metaScience .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>tion Systems</source>
          , Volume
          <volume>2</volume>
          : AIDSS , pages
          <fpage>3138</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [PDP+07]
          <string-name>
            <surname>Damien</surname>
            <given-names>Pollet</given-names>
          </string-name>
          , StØphane Ducasse, Loc Poyet, Ilham Alloui, Sorana Cmpan, and
          <string-name>
            <given-names>HervØ</given-names>
            <surname>Verjus. Towards A ProcessOriented Software Architecture Reconstruction Taxonomy. In RenØ L. Krikhaar</surname>
          </string-name>
          , Chris Verhoef, and Giuseppe Antonio Di Lucca, editors,
          <source>Proceedings of the 11th European Conference on Software Maintenance and Reengineering</source>
          , pages
          <fpage>137148</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>IEEE Computer Society</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [VVvC09]
          <string-name>
            <given-names>Eelco</given-names>
            <surname>Visser</surname>
          </string-name>
          , Sander Vermolen, and Elmer van Chastelet.
          <source>Researchr</source>
          ,
          <year>2009</year>
          . http: //researchr.org .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [WNB+15]
          <string-name>
            <surname>Gemma</surname>
            <given-names>Webster</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hai H. Nguyen</surname>
            ,
            <given-names>David E.</given-names>
          </string-name>
          <string-name>
            <surname>Beel</surname>
            , Chris Mellish,
            <given-names>Claire D.</given-names>
          </string-name>
          <string-name>
            <surname>Wallace</surname>
            , and
            <given-names>Je Z.</given-names>
          </string-name>
          <string-name>
            <surname>Pan</surname>
          </string-name>
          . CURIOS:
          <article-title>Connecting Community Heritage through Linked Data</article-title>
          .
          <source>In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work &amp; Social Computing</source>
          , pages
          <fpage>639648</fpage>
          . ACM,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [YC09]
          <article-title>[Zay15] [ZB14] Hui Yang and Jamie Callan. Feature Selection for Automatic Taxonomy Induction</article-title>
          .
          <source>In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , pages
          <fpage>684685</fpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Vadim</given-names>
            <surname>Zaytsev. Grammar Zoo</surname>
          </string-name>
          :
          <article-title>A Corpus of Experimental Grammarware</article-title>
          .
          <source>Fifth Special issue on Experimental Software and Toolkits of Science of Computer Programming (SCP EST5)</source>
          ,
          <volume>98</volume>
          :
          <fpage>2851</fpage>
          ,
          <year>February 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>Parsing in a Broad Sense</article-title>
          . In Jrgen Dingel, Wolfram Schulte, Isidro Ramos, Silvia Abrahªo, and Emilio InsfrÆn, editors,
          <source>Proceedings of the 17th International Conference on Model Driven Engineering Languages and Systems</source>
          , volume
          <volume>8767</volume>
          <source>of LNCS</source>
          , pages
          <fpage>5067</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[RBO14] [SHK14] [Sou13] [Spi03] [vDV02] [VSM13] Tillmann Rendel</source>
          , Jonathan Immanuel Brachthuser, and Klaus Ostermann.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>From</given-names>
            <surname>Object</surname>
          </string-name>
          <article-title>Algebras to Attribute Grammars</article-title>
          .
          <source>In Proceedings of the 29th International Conference on Object Oriented Programming Systems Languages and Applications</source>
          , pages
          <fpage>377395</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Harald</given-names>
            <surname>Strrle</surname>
          </string-name>
          , Regina Hebig, and
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Knapp</surname>
          </string-name>
          .
          <article-title>An Index for Software Engineering Models</article-title>
          . In Stefan Sauer and Manuel Wimmer, editors,
          <source>Poster Session of MoDELS</source>
          <year>2014</year>
          , volume
          <volume>1258</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>36</fpage>
          <lpage>40</lpage>
          . CEUR-WS.org,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>David M.</given-names>
            <surname>South</surname>
          </string-name>
          . The DPHEP Study Group:
          <article-title>Data Preservation in High Energy Physics</article-title>
          . CoRR, abs/1302.3379,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Diomidis</given-names>
            <surname>Spinellis</surname>
          </string-name>
          .
          <article-title>The decay and failures of web references</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>46</volume>
          (
          <issue>1</issue>
          ):
          <fpage>7177</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Arie van Deursen</surname>
            and
            <given-names>Eelco</given-names>
          </string-name>
          <string-name>
            <surname>Visser</surname>
          </string-name>
          .
          <article-title>The Reengineering Wiki</article-title>
          .
          <source>In Proceedings of the Sixth European Conference on Software Maintenance and Reengineering</source>
          , pages
          <fpage>217220</fpage>
          . IEEE Computer Society,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Bogdan</given-names>
            <surname>Vasilescu</surname>
          </string-name>
          , Alexander Serebrenik, and
          <string-name>
            <given-names>Tom</given-names>
            <surname>Mens</surname>
          </string-name>
          .
          <article-title>A Historical Dataset of Software Engineering Conferences</article-title>
          .
          <source>In Proceedings of the 10th Working Conference on Mining Software Repositories</source>
          , pages
          <fpage>373376</fpage>
          . IEEE Computer Society,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [VSM+14]
          <string-name>
            <surname>Bogdan</surname>
            <given-names>Vasilescu</given-names>
          </string-name>
          , Alexander Serebrenik, Tom Mens, Mark G. J. van den Brand, and
          <string-name>
            <given-names>Ekaterina</given-names>
            <surname>Pek</surname>
          </string-name>
          .
          <source>How Healthy are Software Engineering Conferences? Science of Computer Programming</source>
          ,
          <volume>89</volume>
          :
          <fpage>251</fpage>
          272,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>