<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Automatic Classi cation of EU Projects for Supporting Open Fiscal Data Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ondrej Zamazal</string-name>
          <email>ondrej.zamazal@vse.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information and Knowledge Engineering, University of Economics</institution>
          ,
          <addr-line>W. Churchill Sq.4, 130 67 Prague 3</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The European union funding for job creation and a sustainable and healthy EU economy and environment is channelled through the ve EU structural and investment funds. Although there is EU categorization system for EU projects, EU countries apply their own di erent categorization systems. Some EU countries already apply European categorization system, but many do not. As a result, many projects in available datasets are not categorized using the European categorization system which hinders straightforward scal analyses. The long-term goal of this work is to support an open scal data analysis by an automatic classi cation of EU projects using a machine learning classi er.</p>
      </abstract>
      <kwd-group>
        <kwd>EU projects</kwd>
        <kwd>Open Fiscal Data</kwd>
        <kwd>RDF Code list</kwd>
        <kwd>Classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The European union (EU) funding for job creation and a sustainable and healthy
European economy and environment is channelled through the 5 European
structural and investment funds (ESIF). Although there is the European
categorization system for EU projects, EU countries apply their own di erent
categorization systems. Some EU countries already apply European categorization system,
but many do not.</p>
      <p>
        The Open Knowledge Foundation Deutschland (OKFD)1 gathers data about
EU projects published by responsible authorities (often in PDF), clean and share
them via the GitHub in di erent formats such as CSV, XSLX or JSON. During
data processing information about EU projects are aligned with one scal data
model of Open Spending.2 Recently, OKFD published the full dataset described
in documents [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for period 2000 to 2020 having 2.7 milion of projects.3 In
all, there are 113.446 projects from the latest period 2014-2020 out of which only
1 https://www.okfn.de/en/
2 https://github.com/os-data/eu-structural-funds/blob/master/
specifications/fiscal.schema.yaml
3 For this work we used the dataset from the 21st of April 2017. Up-to-date numbers
are available at the supplementary web page, http://owl.vse.cz:8080/ISWC2017/
7989 projects (7%) have been categorized into the EU categorization system
(intervention code) valid for 2014-2020 period. Although the analysis of the dataset
is already enabled by the OS Viewer tool,4 additional categorization information
would enhance straightforward scal comparative analyses. The motivation of
this work is to support an open scal data analysis by EU project classi cation
into the EU categorization system 2014-2020 using a machine learning classi er.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Gathered Data and the Approach</title>
      <p>The EU categorization system, having 123 categories, is available in the o cial
EU documents as well as in the tabular form from the Data for research web
page.5 We extracted the categorization system for 2014-2020 as the RDF code
list6 where the items contain English labels and are structured in the taxonomy.</p>
      <p>
        The extent to which the European projects in the full data set are described
varies a lot: project name, funding amount, funding period, bene ciary name,
project description etc. In this work we only focus on a lexical description of the
projects; particularly we consider attributes such as bene ciary name, project
name and project description. For projects representation we use the bag of
words approach where a vector has as many attributes as a number of all unique
terms (after removing of stop words) and we count a frequency of a term within
the given project. As a result we have data of a high dimensionality. In order
to cope with a high dimensionality we chose the machine learning algorithm
suitable for such a data [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. According to our initial testing (see the web page),
we decided for the implementation of the classi er based on SVM (Support
Vector Machines) with a linear kernel which uses the LibSVM Java library.7
      </p>
      <p>
        In order to gather the training data we rst found projects having an
intervention code in the dataset. Second, since we needed to unify the semantics
of projects' lexical description, we translated all words into the one natural
language. We applied the translation into English since translation into this
language has usually the best performance and labels in the RDF code list are also
in English. We used the Microsoft Translator API 8 since this API is available
for free up to 2 millions characters per month and there is a su cient coverage
of European languages. As an alternative to natural language translation we
experimented with a word sense disambiguation based on Babelfy [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Third,
in order to obtain only unique projects we deduplicated them based on their
lexical description and intervention code. For translated data we arrived at 5269
unique project descriptions where countries are distributed as follows: 3238
Germany, 1430 France, 8 Malta, 198 UK and 395 Greece. For disambiguated data
4 http://subsidystories.eu
5 http://ec.europa.eu/regional_policy/sources/docgener/evaluation/data/
categorisation_2014_2020_mapping.xls
6 https://github.com/openbudgets/Code-lists/blob/master/EUcategorization/
2014-2020/2014_2020_intervention_fields.trig
7 https://www.csie.ntu.edu.tw/~cjlin/libsvm
8 https://www.microsoft.com/en-us/translator/translatorapi.aspx
we arrived at 3631 unique project descriptions where countries are distributed as
follows: 1864 Germany, 1187 France, 8 Malta, 195 UK and 376 Greece. Training
data are unequally distributed to target classes (intervention codes).
Distributions for translated and for disambiguated data are available on the
supplementary web page.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Preliminary Experiments</title>
      <p>In order to evaluate the SVM classi er we have experimented with three di erent
settings with regard to the preparation of training and testing data. Due to the
fact that target classes are distributed unequally in the gathered data (and there
is no any instance for 32 target classes out of 123), during preprocessing we
performed oversampling, i.e. some randomly selected instances are duplicated,
and undersampling, i.e. some randomly selected instances are removed in order
to receive 100 training instances per target class.9</p>
      <p>For training data the experiment A considers only target classes having more
than 30 instances, the experiment B considers all target classes having at least
one instance, nally the experiment C considers all target classes. For target
classes which do not have any training instance we used words in labels from
the RDF code list of the categotization system. For testing data the experiments
A, B and C include a half of instances from target classes having at least one
instance to be used as testing data. Training and testing data are always disjoint.</p>
      <p>Results for the translated and disambiguated data are in Table 1.
Regarding translated data from the results we can see that the classi er performance
slightly increases from the experiment A to the experiments B and C (from the
precision of .727 to the precision of .762).10 While we could expect that the
classi er model which classi es to more target classes (ceteris paribus) would lead to
9 For training we experimentally used 100 instances, but we want to inspect the in
uence of this paramater on the classi cation performance in future.
10 The precision was averaged from three runs. The demonstration program,
EUProjectsClassi er, for each variant is available at the supplementary web page.
lower performance (since it is more complex task) the experiment results show
that the performance is better. The classi er in the experiment A provides the
classi cation into 22 most frequent target classes and thus instances of those 22
target classes dominate in the testing data. Although this setting is in favour
of the experiment A, the performance in the experiments B and C show that
a growing number of target classes is successfully associated with a modest
increase of the precision thanks to the fact that the classi er in the experiments
B and C can classify additional instances11 compared with the experiment A.
By comparing results for translated and disambiguated data we can see that the
performance of classi ers using translated data is better by approximately 20 %
than classi ers using disambiguated data.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>This work aims at classi cation of EU projects in order to support
straightforward scal analyses. Performed experiments showed that the approach with
natural language translation overperformed the approach with disambiguation
using Babelfy. Although the results of preliminary experiments are promising,
the testing data represent a limited portion of EU projects. Thus, for our near
future work we plan to further evaluate our approach using randomly selected
and newly annotated testing data. In our ongoing work, we randomly selected
EU projects (209.037 projects regardless the funding period) having a non-zero
length of the project description and not having the EU categorization. We
separated those EU projects into di erent groups according to the length of their
project descriptions and we assigned them to domain experts to annotate them.
The initial annotations point out the di culty of such a task since
preliminary inter-annotator agreement is about 25%. Besides this ongoing work on the
preparation of the EU projects classi cation gold standard and the subsequent
evaluation of our approach on top of that we further plan to build a web based
application for an online classi cation of given EU project based on a lexical
description where the top-n predicted categories will be graphically indicated in
the RDF code list visualization.</p>
      <p>The work has been supported by the H2020 project no. 645833 (OpenBudgets).
11 This can be seen from the precision per target class at the supplementary web page
where are also available the results from other experiments with di erent settings.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. SubsidyStories.eu. Dataset Descriptions. https://tinyurl.com/yc subg.
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. SubsidyStories.eu. Methodology &amp; Variables. https://tinyurl.com/y7errfx5.
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Moro</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raganato</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Entity</surname>
          </string-name>
          <article-title>Linking meets Word Sense Disambiguation: a Uni ed Approach</article-title>
          . In:
          <article-title>Transactions of the Association for Computational Linguistics</article-title>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wang</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>J. Mining</given-names>
          </string-name>
          <string-name>
            <surname>High-Dimensional Data</surname>
          </string-name>
          .
          <source>In: Data Mining and Knowledge Discovery Handbook</source>
          . Springer.
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>