<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modeling Semantic Relations from a Dependency-Based Graph: A Corpus-Based Network Analysis of Croatian Parliamentary Debates</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benedikt Perak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Rijeka Rijeka</institution>
          ,
          <country country="HR">Croatia</country>
        </aff>
      </contrib-group>
      <fpage>172</fpage>
      <lpage>192</lpage>
      <abstract>
        <p>The following paper examines the application of graph technologies to parliamentary data. Using the debates that took place in the Croatian Parliament between 2003 and 2017 as an example, it demonstrates how natural language processing (NLP) tools, graph databases, and network algorithms can be used to conduct corpus statistical, stylometric, and semantic analysis. Special attention will be paid to the structure of morpho-syntactically tagged corpora, which are embedded in a property graph database that enables the exploration of corpus-specific semantic relations. As will be shown, such a structure allows parliamentary data to be empirically analyzed with regard to the communication, conceptualization, and framing of social identities, interactions, institutions, and cultural models.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This paper examines the application of graph technologies to parliamentary
data by means of natural language processing (NLP) techniques, the graph
database Neo4j, and a Python implementation of igraph algorithms. It
describes the embedding of a dependency-tagged corpus in a graph database
model, and presents an innovative approach to the empirical linguistic
analysis of the social identities, interactions, institutions, and cultural practices
recorded in parliamentary data. The goal of the project is to establish an
extensive knowledge base with a clearly structured ontology, which in turn
enables data integration, data enrichment, and quantitative-qualitative
analysis.</p>
      <p>
        In addition to the application of standard corpus statistical methods,
semantic and stylometric analysis of the syntactic dependency structures
was carried out using the Universal Dependencies NLP parser
        <xref ref-type="bibr" rid="ref18">(Straka and
Straková, 2017)</xref>
        . Conceptual profiling of parliamentary discourse was
undertaken based on the semantic features of the syntactic dependency lexical
matrix with the aid of graph algorithms involving network centrality and
community measures.
      </p>
      <p>By way of a case study, this paper considers records detailing Croatian
parliamentary debates between 2003 and 2017. The following section (Section
2) discusses parliamentary data in general, while Section 3 elaborates on the
genesis of the Croatian Parliamentary Corpus (CroParl). Section 4 presents
a number of specific examples drawn from the said corpus, whereas the fifth
and final section contains concluding remarks and some thoughts on the
work that lies ahead.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Parliamentary Data</title>
      <p>Parliamentary data constitutes a rich and publicly available resource that is
inherently endowed with politically and historically significant information
concerning the discourse between the political representatives of democratic
systems of government and the socio-cultural interactions in which they
participate.</p>
      <p>
        In recent years, the improved accessibility of parliamentary data has made
the democratic process increasingly transparent
        <xref ref-type="bibr" rid="ref12 ref3 ref8">(Janssen, 2011; Andrews
and da Silva, 2013; Granickas, 2014)</xref>
        . Digital records of parliamentary
debates have become an indispensable resource for computational data
analysis in the humanities and social sciences
        <xref ref-type="bibr" rid="ref10 ref4">(Glavaš et al., 2019; Berntzen et al.,
2019; Hofmann et al., 2020)</xref>
        . Moreover, there are a number of reasons why
companies, private citizens, and public organizations may wish to engage
with parliamentary data, including but not limited to creating business value,
enabling local citizen value, addressing global societal challenges, and
advocating the open data agenda
        <xref ref-type="bibr" rid="ref13">(Lassinantti et al., 2019)</xref>
        .
      </p>
      <p>Yet the proper archiving, structuring, synchronization, and visualization
of the treasure trove of multimodal data that is generated over the course of
the socially and institutionally highly complex interactions that characterize
parliamentary debates are not without their pitfalls. For one, the diversity
of the data management approaches of national and regional parliaments
with respect to issues such as language, rhetorical strategies, actor properties,
social features, and political context poses a significant obstacle to systematic
scientific inquiry.</p>
      <p>
        Such standards for the processing of textual data and the creation of
structured corpora of parliamentary and other governmental debates as currently
exist – along with a broad variety of metadata formalizations – are the result
of the eofrts of a small number of scholars based out of various European
research centers. At the European level, the Digital Corpus of the European
Parliament (DCEP)
        <xref ref-type="bibr" rid="ref9">(Hajlaoui et al., 2014)</xref>
        is perhaps the most significant
digital collection, while the CLARIN ERIC consortium assembles the
available resources from a number of European national parliaments.1 Within
the ParlaCLARIN initiative, Erjavec and Pancur have laid out Text
Encoding Initiative (TEI) guidelines for corpora of parliamentary proceedings
        <xref ref-type="bibr" rid="ref6">(Erjavec and Pancur, 2019)</xref>
        , with recommendations for the structure of the
corpus, the encoding of metadata (including, for example, the speakers and the
political parties to which they belong), speeches, and notes, and guidelines
for linguistic annotation and integration of multimedia content.2 However,
despite these eofrts at standardization, the community of researchers
working with parliamentary data remains fragmented and in need of a unified
platform.
      </p>
      <p>Keeping these various factors in mind, the goal of this paper is to
showcase a flexible information framework based on graph technologies that is
capable of processing and integrating regional, national, and European
parliamentary data. In so doing, I hope to demonstrate how current advances in
natural language processing and data management can be harnessed to build
a socio-linguistic parliamentary data network that is accessible not only to
specialized scholars, but also to journalists, NGOs, and private citizens.</p>
      <p>The working prototype for the web implementation of Croatian
parliamentary data is available on the website hosted by the University of Rijeka’s
EmoCNet project.3
1https://www.clarin.eu
2https://github.com/clarin-eric/parla-clarin
3http://emocnet.uniri.hr/croparl</p>
      <p>Croatian Parliamentary Corpus
A significant portion of parliamentary data is made up of texts that detail
political debates among representatives. This corpus of speech acts allows
researchers, particularly linguists and political scientists, to engage in
diefrent kinds of semantic and pragmatic analysis.</p>
      <p>
        Here I am concerned with the debates that took place in the Croatian
Parliament between 2003 and 2017, during the fifth to ninth parliamentary
assemblies. As can be seen in Figure 1, the process of corpus creation in this
instance consisted of data gathering, NLP parsing, data integration, and data
management in the property graph database
        <xref ref-type="bibr" rid="ref15">(Perak and Rodik, 2018)</xref>
        .
3.1
      </p>
      <sec id="sec-2-1">
        <title>Data Gathering and Integration</title>
        <p>The data was gathered using a Selenium scraper4 from the Croatian
Parliament web repository.5 Data pertaining to the debates that took place
during the fifth to the ninth parliamentary assemblies is to be found in two
separate datasets: one containing values that relate to the assemblies, sessions,
and topics, and another containing values that concern persons, debate
transcripts, topic IDs, and announcement metadata.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>NLP Parsing</title>
        <p>
          Extracted from 390,000 transcripts, the various speech acts of the
representatives in question were processed using the UDPipe Natural Language
Processing Toolkit
          <xref ref-type="bibr" rid="ref17">(Straka et al., 2016)</xref>
          , which includes features such as
tokenization, part-of-speech tagging, lemmatization, and dependency parsing.
Parsing was done using the Universal Dependencies 2.0 Model (UDPipe
repository)6 developed in Croatia
          <xref ref-type="bibr" rid="ref1">(Agić and Ljubešić, 2015)</xref>
          .
        </p>
        <p>The NLP parsing created over 4 million sentences and 70 million word
tokens in the CoNLL-U output format (Figure 2). The morpho-syntactic
metadata appears in 10 tab-separated value fields:
1. ID: word index
2. FORM: word form or punctuation symbol
3. LEMMA: lemma or stem of word form
4. UPOS: universal part-of-speech tag
5. XPOS: language-specific part-of-speech tag
6. FEATS: list of morphological features from the universal feature
inventory
4https://github.com/ropensci/RSelenium
5The data gathering process is published under the GitHub handle https://github.com/
rodik/Sabor. The Croation Parliament web repository can be accessed via http://edoc.sabor.hr/
6https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-2364</p>
        <sec id="sec-2-2-1">
          <title>7. HEAD: head of the current word 8. DEPREL: universal dependency relation to the HEAD 9. DEPS: enhanced dependency graph 10. MISC: any other annotation</title>
          <p>3.3</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Data Integration</title>
        <p>
          The data obtained from the Croatian Parliament was then integrated with
the CoNLL-U formatted results
          <xref ref-type="bibr" rid="ref17">Straka et al. (2016)</xref>
          in the property graph
database Neo4j
          <xref ref-type="bibr" rid="ref20">(Webber, 2012)</xref>
          . Using the existing data categories of the
Croatian parliamentary records in conjunction with parsed linguistic
structures, an informational graph structure was created, in which labels, nodes,
node properties, relationships, and relationship properties form an
ontology of the various entities involved in the parliamentary debates (Figure
3). For example, labels represent instances of ontologically similar entities,
while links represent their inter-structural connections. Labels also have as
instances nodes that can store their relational properties in the form of
keyvalue metadata structures, which is one of the major advantages of
employing a property graph database.
        </p>
        <p>Structurally complex entities are metonymically related to ontologically
less complex entities: Assembly can be narrowed down to Session, which
can be narrowed down to Topic, which can be narrowed down to Utterance,
which can be narrowed down to Sentence, which can be narrowed down
to Token. Utterances are related to a Representative, who is socially related
to a Party and a Parliamentary Club. Some labels have inter-structural
relationships: tokens, sentences, and utterances are sequential in nature, while
tokens also have syntactic dependency relations.7</p>
        <p>
          Arranging the data in question in a property graph has a number of
advantages. First, it enables the integration of various datasets into a single
framework. Second, labels, nodes, relations, and their properties can be
easily updated or remodeled according to newly adopted standards, and the
structure of the graph can be enriched with additional knowledge resources.
Last but not least, the user-friendly graph representation of the information
ontology enables digital humanities scholars to intuitively develop new
approaches to their material based on the interrelation of data structures.8
7For details concerning the data integration and NLP parsing process, see
          <xref ref-type="bibr" rid="ref15">Perak and
Rodik (2018)</xref>
          .
        </p>
        <p>8The tagged CroParl graph database’s Neo4j dump is a case in point, https://drive.google.
com/file/d/1zRy3EmwPrb4r3vGM3J5N5bpeKwCLEaHF/view?usp=sharing
4.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Corpus Analysis</title>
      <p>Statistical Summarization of Lexical Usage
The CroParl corpus contains data from five parliamentary assemblies and
covers 5,599 topics, which were broached by 895 members of parliament
belonging to 42 political parties. In the process of NLP parsing, 390,078
utterances were related to 4 million sentences and 70 million tokens.
4.1.1</p>
      <sec id="sec-3-1">
        <title>Lexical Summarization</title>
        <p>The graph-based lexical store thus enables standard lexical summarizations
based on tokens, lemmas, or part-of-speech counts, and their respective
proportions. For instance, Table 1 represents the lemma count for nouns in the
corpus.</p>
        <p>The corpus can also be used for the representation of a keyword in
context (KWIC). The KWIC solution oefred on the oficial site
(https://edoc.sabor.hr/) relies on a string-based query search. This type of
search is relatively easy to implement from a technical standpoint, and is
suitable for morphologically lean languages. However, this is not the case
for Croatian, which has a complex declination and inflection system.
Consequently, a lemma-based query yields much more accurate results in the
KWIC concordance. Table 2 presents an example of sentence level results
for the noun ljubav (‘love’) using a lemma-based KWIC query. The results
are based on the lemma=‘ljubav’ and part-of-speech pos=‘noun’ properties
assigned to the tokens.
4.1.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Stylometric Summarization</title>
        <p>
          Stylometric summarization based on the linguistic behaviour of
parliamentary representatives can provide us with valuable insights into their individual
linguistic profiles, stylistic idiosyncrasies, and similarities compared to other
speakers
          <xref ref-type="bibr" rid="ref2">(Amancio, 2015)</xref>
          . As I will demonstrate in this section, utterance
analysis and word/concept counts are especially illuminating in this regard.
        </p>
        <p>The number of utterances of a speaker, for example, corresponds fairly
directly to their prominence within the parliament. Table 3 shows a list of
the 15 representatives with the highest number of utterances in the CroParl
corpus, with the overwhelming majority belonging to Vladimir Šeks, who
served as President of the Croatian Parliament from 2003 to 2008. A quick
examination of Table 3 reveals that the two runners-up, Bebić and Leko, also
served in this role, which, by its formal nature, involves numerous speech acts
with low token counts per utterance.</p>
        <p>Prominence can also be measured by other lexical features, such as
lemmas per representative, part-of-speech count per representative, etc., and
the data management structure enables various types of filtering (e.g.
utterances/tokens/lemmas per date/assembly/session/topic).</p>
        <p>Another application of corpus-based stylometric analysis is the frequency
with which a certain lexeme is used by a specific representative. For instance,
we might be interested in how often emotionally charged lexemes were used,
and by whom. By way of example, Tables 4 and 5 show a portion of the data
for two such terms: love and peace.</p>
        <p>Here we see that the lexeme mir (‘peace’) was used almost five times more
often than the lexeme ljubav (‘love’). Column “F” represents the frequency
with which the respective lexeme occurred in the representative’s utterances,
column “p in All” represents the overall proportion of the lemma in the
corpus, column “Auth” the number of tokens per representative, and column
“p in Auth” the proportion in which the lemma occurred in speeches by the
representative. While the value “p in All” can be interpreted as an indicator
for the representative’s conceptual influence on the debates, the value “p in
Auth” represents the importance assigned to a particular concept by a given
speaker.</p>
        <p>This type of analysis opens up a whole range of possibilities when it comes
to further studies concerning the significance of specific concepts to
individual representatives. Stylometric linguistic analysis could also be used to
interrogate the socio-, pragma-, and cross-linguistic profiles of a given
political party or parliamentary club based on how often and in which contexts
lexemes such as ‘love’ and ‘peace’ are used by its members.
4.2</p>
        <p>Dependency-Based Semantic Network Analysis
The NLP parsing of the texts in question created 74,968,809 syntactic
dependency relations between tokens. The structure of the relations tagged by
the Universal Dependencies parser is shown in Table 6.</p>
        <p>The graph structure of these relations enables a dependency-based type of
semantic analysis that is conducive to a number of NLP tasks, including the
discovery of corpus-specific word senses.
4.2.1</p>
        <p>
          Semantic Domain Induction Based on “Conj” Dependency
Word sense disambiguation and word sense identification have been
important tasks from the earliest days of natural language processing research
          <xref ref-type="bibr" rid="ref11 ref14">(Schütze, 1998; Ide and Véronis, 1998; Navigli, 2009)</xref>
          . NLP studies have
repeatedly demonstrated the usefulness of dependency-tagged corpora when
it comes to the unsupervised assembling of semantic knowledge. The
coordination structure labeled as conj in the UD framework, which expresses
an asymmetrical dependency relation between two elements that are
connected by a coordinating conjunction, such as and, or, etc., has proved
particularly useful for distinguishing between associated classes of concepts
          <xref ref-type="bibr" rid="ref21 ref21 ref22 ref5 ref5">(Widdows and Dorow, 2002; Cederberg and Widdows, 2003; Widdows, 2003)</xref>
          .
        </p>
        <p>In the method being outlined here, the iterating graph algorithm is used
to produce an undirected graph from all the nouns collocated with the
construction dependency, operating as a kind of lexical embedding that
represents the semantic structure of the concepts found in the Croatian
Parliamentary corpus. The coordination dependency graph can be used to
identify ontologically similar lexemes and related semantic domains for a
given source lexeme:
• choose a source lexeme and identify n number of the most frequently
collocated target lexemes
• construct a lexical network with associated lexemes as nodes and
coordination-based weighted relations
• detect lexically coherent communities with sub-graphs as a
representation of the semantic domains, and use parametrizable granularity as
a measure of the level of categorical consistency vs continuity and
abstraction
• analyze the distribution of prototypical association patterns as an
indication of cross-cultural framing and cross-linguistic variations
Graph analysis, meanwhile, is performed with the help of the Python
implementation of igraph, which is used to identify, measure, and visualize the
conceptual framing. The procedure starts with extracting collocations of an
arbitrary source lexeme. For instance, the first 50 most common collocates
for the lexeme mir (‘peace’) are represented in Figure 4.</p>
        <p>For the source lexeme mir with n = 50 first order lexemes, the structural
function of a friend coordination network is enhanced by identifying n=50
second order lexemes in a friend-of-a-friend (FoF) pattern. By analyzing the
subgraphs created by the second order coordinated dependency noun
collocates we can distinguish the semantic clusters that indicate the sense
association of the source lexeme.</p>
        <p>
          The FoF network with 945 nodes was pruned with regards to the node
degree measure degree&gt;4 in order to filter out the less interconnected
lexemes from the semantic network. As can be seen in Figure 5, the resulting
FoF network for mir contained 120 nodes, which were then clustered using
the Leiden clustering algorithm
          <xref ref-type="bibr" rid="ref19">(Traag et al., 2019)</xref>
          . The resulting network
with six clusters reveals the corpus-specific structure of the semantically
associated concepts and domains, as can be seen in Table 7. If needed,
modiifcations of the resolution parameter can render more fine-grained or more
robust communities.
        </p>
        <p>Finally, the prominence of the associative conceptual profiling of the
source lexeme mir can be discerned from a speaker’s weighted contribution
to the coordination dependency graph. The weighted graph was constructed
using the collocates that are in a coordination dependency with the source
lexeme mir used by each representative.</p>
        <p>The resulting graph with 461 nodes was pruned to highlight the more
prominent associations using weighted degree &gt; 4 settings. The weighted
degree represents the sum of weights assigned to the node’s connections. By
ifltering out the nodes by weighted degree, we were able to focus on the
lexemes and representatives that saliently contribute to the association matrix
of the lexeme mir by their occurrence. As is shown in Figure 6, the pruned
associative network with 97 elements was then clustered using the cpm
resolution 0.34, which yielded 50 clusters. The 12 most prominent association
clusters of the lexeme mir can be found in Table 8.</p>
        <p>Clearly, then, graph measures and network representation of socio-lexical
phenomenon can be used to glean an empirical, yet intuitive,
understanding of the complex weighted structural relations between lexical content and
the agents who generate this content. Of particular interest here is the
question of which lexemes exhibit the highest number of associations, and who
introduced these associations. Figure 6 shows that stabilnost (‘stability’),
sigurnost (‘safety’), and red (‘order’) are central for understanding the
associative conceptual content of the lexeme mir. The clusters, represented with
diefrent colors in Figure 6 and transcribed in Table 8, reveal the agents who
most frequently made use of the lexemes in question. Interestingly, most
of the agents connected with the concepts ‘safety’ and ‘order’ have served as
members of institutions responsible for the nation’s defense. For instance,
both Ivan Šantek and Tomislav Čuljak were members of the committee on
Internal Policy and National Security. Šantek was also a member of the
Defense Committee from 2013, while Berislav Rončević served as its vice
chairman.</p>
        <p>Conversely, the concept ‘peace’ has a diefrent connotation when coupled
with the lexemes ‘stability’ and ‘cooperation’ – the agents in this cluster seem
to be more concerned with international relations. Jozo Radoš, for example,
was vice-chairman of the Committee on European Aafirs and a member
of the Committee on European Integration, while Davor Ivo Stier was a
member of the Committee on Interparliamentary Cooperation, the Foreign
Policy Committee, and the Committee on European Aafirs. The same
holds true for Stier’s fellow part member, Andrej Plenković, who would go
on to become prime minister. Similarly, the aspect of ‘prosperity’ in
connection with ‘peace’ was promoted with particular vigor by the deputy chairman
of the Committee on Croats outside the Republic of Croatia, Boro
Grubišić, and by a member of the Committee on Regional Development and
European Union Funds, Petar Baranović. The antonym ‘war,’ meanwhile,
was prominently associated with ‘peace’ by the president of the Delegation
of the Croatian Parliament to the NATO Parliamentary Assembly, Krešimir
Ćosić.</p>
        <p>Although these results were acquired from a complex set of queries and
graph algorithms, they intuitively represent the conceptual associative
dimension of the lexeme mir, and as such can help us to begin to understand
how conceptual relations are influenced and motivated by the political views
of the speakers in question and their respective institutional functions. The
possibility of such an insight demonstrates the true value of graph-based
socio-linguistically and morpho-syntactically tagged corpus analysis: it
allows us not only to simply count lexical items or allocate frequencies of
syntactical relations, but to explore the socio-cognitive aspects of
conceptualization by revealing a dynamic pragmatic context that is determined by the
graph structure of the nodes and the relationships that link them together.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In this essay, I have introduced a graph-based data management and
analysis framework that seeks to integrate parliamentary data with a NLP
dependency tagged corpus. The proposed framework can ingest multiple data
formats with disparate internal structures, while producing flexible and
intuitive information structures that are highly conducive to comparative data
analysis. Making use of the readily available Neo4j database, NLP tools, and
a Python implementation of graph algorithms, it can easily be adapted to
a wide variety of discipline-specific research approaches and customized to
meet the requirements of individual researchers.</p>
      <p>In our continued work on parliamentary data analysis, we plan to: a)
harvest recent parliamentary data, b) implement new NLP tools for the
Croatian language provided by the CLASSLA initiative, and include
namedentity recognition (NER) data in the data tagging, c) carry out further
research on stylometric analysis, particularly with regard to its diachronic
dimension, d) integrate data from external sources of parliamentary data
(https://edoc.sabor.hr, wikipedia, etc.), and e) extend existing semantic graph
research to other dependency relations.</p>
      <p>Special attention will be paid to further enhancing the web
infrastructure for Croatian parliamentary data,9 which has recently been established
with the help of funding from the University of Rijeka and the University of
Zagreb University Computing Centre (SRCE), as such resources are
essential for establishing a platform for collaborative scholarship on parliamentary
data that encourages the exchange of information at an international level.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work has been jointly sponsored by the Croatian Science Foundation
under the project number UIP-05-2017-9219 and the University of Rijeka
under the project number UNIRI-human-18-243.
O
_ E
ER TA E
B :D T</p>
      <p>A
EM rom :oD
M f-
t</p>
      <p>N
I: R
D
e TR :TS
I
v
i S e
ta :e m
t
n m na
e a e
s
re -n r</p>
      <p>u
p -s
e
r</p>
      <p>T
X
E
N
N</p>
      <p>A
T</p>
      <p>E
e T I:N R L</p>
      <p>N r T R O
c :ID be t:S T O
n I m xe :gS t:B</p>
      <p>tcene ceuN ceneT irco em
e n n
t d e
n</p>
      <p>n en t
e -se tsne -sen -re cun
S o
- nn
a
S
A
H</p>
      <p>H
S
A
S
A
H</p>
      <p>S
A
H</p>
      <p>S
A
H</p>
      <p>Y
B
D
E
R
E
V
I
L
E
D
S
A
H
S
A
H
S
A
H
y
t
r
a</p>
      <p>P
F
O
_ E
ER TA E
B :D T</p>
      <p>A
EM rom :oD
M f-
te
v
i
t
a
t
n
e
s
e
r
p
e
R</p>
      <p>N
:I R
D
e TR :TS
I
v
i S e
ta :e m
t
n m na
e a e
s
re -n r</p>
      <p>u
p -s
e
r</p>
      <p>T
N
I: R
D
e TR :TS
I
v
i S e
ta :e m
t
n m na
e a e
s
re -n r</p>
      <p>u
p -s
e
r</p>
      <p>T
X
E</p>
      <p>N
e T</p>
      <p>N R R O
c I: E T T O
na IcenD :TAD i:trpS i:gnS t:nB
r e c d e
e ra ta sn rco em
tt tteu -d t-ra r-e cnu</p>
      <p>U</p>
      <p>N
A
E
L
o
n
n
a</p>
      <p>R</p>
      <p>TN TR :TS
c I
i :</p>
      <p>S g
p ID : n
c e i
i m rd
o
T t-op -na co
e
r</p>
      <p>T
n</p>
      <p>IN E
o : T E
i ID A T
s on :D A</p>
      <p>i D
s s m :o</p>
      <p>s ro
te se
fS
y
l
e
s
s
A
b</p>
      <p>D
I
y
m l
b
T
N
I:
m
e
s
s
a
their key:value properties: (Assembly) -[HAS]-&gt; (Session) -[HAS]-&gt; (Topic) -[HAS]-&gt;
((Utterance) -[HAS]-&gt; (Sentence) -[HAS]-&gt; (Token) -[SEQUENCE]-&gt; (Token) -[DEPENDENCY]-&gt;
(Token)) -[DELIVERED BY]-&gt; ((Representative) -[MEMBER OF]-&gt; (Club) ,-[MEMBER OF]-&gt;
(Political party)).</p>
      <p>z ic
za it t
i a c
n
N to lemm tsyna ENR
P e
L k
END l:reS
P p</p>
      <p>ED -ed
R
T
o
t
:IInDN IEUQ :TSR :aTSR :gaTSR :tgaTSR :egTSR I:TN l:eTSR :sTSR
T :S
e N rm m ts s ua ade rp ep
tko nU f-o le p p n -h -ed -d</p>
      <p>m o o g
- ke - -u -x l-a</p>
      <sec id="sec-5-1">
        <title>Example</title>
        <p>Htio bih odmah na početku kazati pošto govorimo
o djeci, a ponekad i u nekim raspravama ima dosta
politiziranosti, htio bih kazati da se djeca ne rađaju
ni radi politike, ni radi države, ni radi klase, ni radi
nacije, da se rađaju radi ljubavi kao što je netko
već jutros rekao i da je to posljedica možda jednog
od posljednjih prirodnih odnosa između muškarca i
žene, a koji su opet posljedica ljubavi.</p>
        <p>Dakle, ja mislim da je žena kao majka, žena kao dio
obitelji i ne znam zašto su se ljudi smijali kada se
govorilo o ljubavi pa ja onda ne moram ni ljubav
spomenuti ali dakle, žena koja izabere da će sa svojim
suprugom barem u prijateljstvu imati određeni broj
djece a usput završi fakultet, usput doktorira, usput i
radi nazadna ja mislim da je ona vrlo slobodna osoba,
da je ona suvremena i da je ona ravnopravnija od
kolega muškaraca jer emancipacija se živi a ne priča...
...
Count
100,684
94,195
91,425
89,971
87,245
85,868
85,413
85,325
80,573
80,302</p>
        <p>Name
1 Šeks, Vladimir
2 Tafra, Višnja
3 Kajin, Damir
4 Đakić, Josip
5 Brkić, Milijan
6 Jandroković, Gordan
7 Pernar, Ivan
8 Špoljar, Dunja
9 Ćosić, Krešimir
10 Zgrebec, Dragica
p in All
7.13e-06
2.852e-06
1.617e-06
1.592e-06
1.401e-06
1.095e-06
1.082e-06
1.07e-06
1.057e-06
1.006e-06</p>
        <p>Auth
1,247,153</p>
        <p>95,950
2,171,645
501,417
141,999
404,767
461,554
60,456
242,607
699,250</p>
        <p>Nodes
mir, stabilnost, vrijednost, demokracija,
napredak, tolerancija, oprost,
proizvodnja, blagdan, život, tišina, prosperitet,
blagostanje, rast, sustav, gospodarstvo,
zaštita, kvaliteta, aktivnost, povećanje,
uvjet, društvo, novac, imovina, načelo,
uloga, rješenje
red, sigurnost, sloboda, pravda,
pozivanje, govor, sjednica, zastupnik,
rasprava, tema, prijedlog, pitanje,
izvješće, standard, zdravlje, pravo,
zakonodavstvo, građanin, odbor,
odgovornost, odluka
materijal, država, akcija, desetljeće,
općina, grad, ministarstvo, zakon,
proračun, prihod, samouprava,
županija, razlog, zajednica, broj, republika,
sredstvo, uprava, jedinica, turizam,
udruga
suradnja, interes, fokus, razvoj, plan,
dobar, cilj, područje, reforma, izmjena,
politika, strategija, program, potpora,
prioritet, rezultat, potreba, stvar, mjera,
dokument
žena, mjesto, ugovor, posao, plaća,
doprinos, radnik, rad, način, uvođenje,
odnos, čovjek, dio, prostor, tisuća,
obveza, projekt
rat, stanje, oružje, prijetnja, slučaj,
kod, problem, vrijeme, situacija,
godina, primjer, djelo, mogućnost, podatak
material, state, action, decade,
municipality, city, ministry, law, budget,
revenue, self-government, county, reason,
community, number, republic, means,
administration, unit, tourism,
association
cooperation, interest, focus,
development, plan, good, goal, area, reform,
change, policy, strategy, program,
support, priority, result, need, thing,
measure, document
woman, place, contract, job, salary,
contribution, worker, work, manner,
introduction, relationship, man, part, space,
thousand, obligation, project
war, condition, weapon, threat, case,
code, problem, time, situation, year,
example, work, possibility, data</p>
        <p>C
1
2</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Agić</surname>
          </string-name>
          , Ž. and
          <string-name>
            <surname>Ljubešić</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Universal Dependencies for Croatian (that work for Serbian, too)</article-title>
          .
          <source>In The 5th Workshop on Balto-Slavic Natural Language Processing (BSNLP</source>
          <year>2015</year>
          ), pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <string-name>
            <given-names>Hissar. INCOMA</given-names>
            <surname>Ltd</surname>
          </string-name>
          . Shoumen, https://aclanthology.org/W15-5301.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Amancio</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A Complex Network Approach to Stylometry</article-title>
          .
          <source>PLOS ONE</source>
          ,
          <volume>10</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          , DOI: 10.1371/journal.pone.
          <volume>0136076</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Andrews</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>da Silva</surname>
            ,
            <given-names>F. S. C.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Using Parliamentary Open Data to Improve Participation</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Theory and Practice of Electronic Governance</source>
          , pages
          <fpage>242</fpage>
          -
          <lpage>249</lpage>
          . DOI:
          <volume>10</volume>
          .1145/2591888.2591933.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Berntzen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johannessen</surname>
            ,
            <given-names>M. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andersen</surname>
            ,
            <given-names>K. N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Crusoe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Parliamentary Open Data in Scandinavia</article-title>
          . Computers,
          <volume>8</volume>
          (
          <issue>3</issue>
          ):
          <fpage>65</fpage>
          , DOI: 10.3390/computers8030065.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Cederberg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Widdows</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Using LSA and Noun Coordination Information to Improve the Recall and Precision of Automatic Hyponymy Extraction</article-title>
          .
          <source>In Proceedings of the seventh conference on Natural language learning at HLT-NAACL</source>
          <year>2003</year>
          , volume
          <volume>4</volume>
          , pages
          <fpage>111</fpage>
          -
          <lpage>118</lpage>
          . DOI:
          <volume>10</volume>
          .3115/1119176.1119191.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Erjavec</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Pancur</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Parla-CLARIN: TEI Guidelines for Corpora of Parliamentary Proceedings</article-title>
          .
          <article-title>In Book of Abstracts of the TEI2019: What is text, really? TEI and beyond</article-title>
          , number
          <volume>157</volume>
          . https://gams.uni-graz. at/o:tei2019.bookofabstracts.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Glavaš</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nanni</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Computational Analysis of Political Texts: Bridging Research Eofrts Across Communities. In Proceedings of the 57th annual meeting of the association for computational linguistics: Tutorial abstracts</article-title>
          , pages
          <fpage>18</fpage>
          -
          <lpage>23</lpage>
          , Florence, Italy. Association for Computational Linguistics, DOI: 10.18653/v1/
          <fpage>P19</fpage>
          -4004.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Granickas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Open Data as a Tool to Fight Corruption</article-title>
          .
          <source>Technical Report 4</source>
          ,
          <string-name>
            <given-names>European</given-names>
            <surname>Public</surname>
          </string-name>
          Sector Information Platform.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Hajlaoui</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolovratnik</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Väyrynen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinberger</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , et al. (
          <year>2014</year>
          ).
          <article-title>DCEP-Digital Corpus of the European Parliament</article-title>
          .
          <source>In LREC 2014 (Language Resources and Evaluation Conference</source>
          , pages
          <fpage>3164</fpage>
          -
          <lpage>3171</lpage>
          . http://www. lrec-conf.org/proceedings/lrec2014/pdf/943_Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Hofmann</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marakasova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumann</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neidhardt</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al. (
          <year>2020</year>
          ).
          <article-title>Comparing Lexical Usage in Political Discourse Across Diachronic Corpora</article-title>
          .
          <source>In Proceedings of the Second ParlaCLARIN Workshop</source>
          , pages
          <fpage>58</fpage>
          -
          <lpage>65</lpage>
          , Marseille, France. European Language Resources Association, https: //aclanthology.org/
          <year>2020</year>
          .parlaclarin-
          <volume>1</volume>
          .
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Véronis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Introduction to the Special Issue on Word Sense Disambiguation: The State of the Art</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>24</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          , DOI: 10.5555/972719.972721.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Janssen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>The Influence of the PSI Directive on Open Government Data: An Overview of Recent Developments</article-title>
          .
          <source>Government Information Quarterly</source>
          ,
          <volume>28</volume>
          (
          <issue>4</issue>
          ):
          <fpage>446</fpage>
          -
          <lpage>456</lpage>
          , DOI: 10.1016/j.giq.
          <year>2011</year>
          .
          <volume>01</volume>
          .004.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Lassinantti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ståhlbröst</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Runardotter</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Relevant Social Groups for Open Data Use and Engagement</article-title>
          .
          <source>Government Information Quarterly</source>
          ,
          <volume>36</volume>
          (
          <issue>1</issue>
          ):
          <fpage>98</fpage>
          -
          <lpage>111</lpage>
          , DOI: 10.1016/j.giq.
          <year>2018</year>
          .
          <volume>11</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Word Sense Disambiguation: A Survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>41</volume>
          (
          <issue>2</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>69</lpage>
          , DOI: 10.1145/1459352.1459355.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Perak</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rodik</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Building a Corpus of the Croatian Parliamentary Debates Using UDPipe Open Source NLP Tools and Neo4j Graph Database for Creation of Social Ontology Model, Text Classification and Extraction of Semantic Information</article-title>
          . In Fišer, D. and
          <string-name>
            <surname>Pančur</surname>
          </string-name>
          , A., editors,
          <source>Conference on Language Technologies &amp; Digital Humanities</source>
          . https://www.bib.irb.hr/960280.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Schütze</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Automatic Word Sense Discrimination</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>24</volume>
          (
          <issue>1</issue>
          ):
          <fpage>97</fpage>
          -
          <lpage>123</lpage>
          , DOI: 10.5555/972719.972724.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Straka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajič</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Straková</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>UDPipe: Trainable Pipeline for Processing CoNLL-U Files Performing Tokenization, Morphological Analysis, POS Tagging and Parsing</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)</source>
          , pages
          <fpage>4290</fpage>
          -
          <lpage>4297</lpage>
          , Portorož, Slovenia.
          <source>European Language Resources Association (ELRA)</source>
          , https://aclanthology.org/L16-1680.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Straka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Straková</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2017</year>
          ). Tokenizing, POS Tagging,
          <article-title>Lemmatizing and Parsing UD 2.0 With UDPipe</article-title>
          .
          <source>In Proceedings of the CoNLL</source>
          <year>2017</year>
          <article-title>Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies</article-title>
          , pages
          <fpage>88</fpage>
          -
          <lpage>99</lpage>
          , Vancouver. Association for Computational Linguistics, DOI: 10.18653/v1/
          <fpage>K17</fpage>
          -3009.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Traag</surname>
            ,
            <given-names>V. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waltman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and van Eck,
          <string-name>
            <surname>N. J.</surname>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>From Louvain to Leiden: Guaranteeing Well-Connected Communities</article-title>
          .
          <source>Scientific Reports</source>
          ,
          <volume>9</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          , DOI: 10.1038/s41598-019-41695-z.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Webber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>A Programmatic Introduction to Neo4j</article-title>
          .
          <source>In Proceedings of the 3rd Annual Conference on Systems, Programming, and Applications: Software for Humanity</source>
          , pages
          <fpage>217</fpage>
          -
          <lpage>218</lpage>
          , New York, NY. Association for Computing Machinery, DOI: 10.1145/2384716.2384777.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Widdows</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Unsupervised Methods for Developing Taxonomies by Combining Syntactic and Statistical Information</article-title>
          .
          <source>In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics</source>
          , pages
          <fpage>276</fpage>
          -
          <lpage>283</lpage>
          . https://aclanthology.org/N03-1036.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Widdows</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Dorow</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>A Graph Model for Unsupervised Lexical Acquisition</article-title>
          .
          <source>In COLING 2002: The 19th International Conference on Computational Linguistics</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          . DOI:
          <volume>10</volume>
          .3115/1072228.1072342.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>