<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Analysis of Contractual Agreements to Support End-User Interpretation</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Smart Data Analytics (SDA) - University of Bonn and Fraunhofer Intelligent Analysis and Information Systems (IAIS)</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ubiquitous availability of the Internet results in a huge number of apps, software and online services with accompanying contractual agreements in the form of `terms of use' and `privacy policy'. Although everyone is exposed to such consent forms, the majority tend to ignore them due to their length and complexity. In this thesis, we focus on interpretation of contractual agreements for the bene t of endusers. By applying text mining and semantic technologies, we develop an approach that extracts important information and retrieves helpful links and resources for the better comprehension. Our approach is based on ontology-based information extraction and machine learning and delivers the unpleasant consent form in a user friendly and visualized format. The evaluation shows that although semi-automatic approaches lead to information loss, they save time and e ort by producing instant results and facilitate the end-users' understanding of legal texts.</p>
      </abstract>
      <kwd-group>
        <kwd>Contractual agreement</kwd>
        <kwd>Ontology-based information extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The increasing availability of online services and mobile apps has led to a huge
proliferation of terms and conditions regulating their use. In the digital age
everyone is exposed to such terms, and in their majority this constitutes ordinary
people with limited to no knowledge of legal terms. The problem arises when
people ignore the consent forms due to their length and complex terminology.
In a recent study, \The biggest lie on the Internet", 543 students were asked
to agree to a privacy policy and terms of use in order to join a ctitious social
network [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Although 26% did not choose the 'quick join', the average time of
reading was only 73 seconds. Ignoring these terms is a risk, taken by most users.
According to Skandia1, 10% were bound by a longer contract than they expected
and 5% lost money by not being able to cancel or amend their bookings.
      </p>
      <p>In order to facilitate the process of digesting terms and conditions for
regular end-users, we consider applying text mining and use of domain ontologies
1
http://www.prnewswire.co.uk/news-releases/skandia-takes-the-terminal-out-ofterms-and-conditions-145280565.html
to provide visualized summaries. The approach considered is broadly applicable
to other forms of text-based contractual agreements. However, in this thesis we
speci cally focus on terms of use (aka. End-User License Agreement or EULA)
and privacy policies, since they have the broadest impact and a ect everyone.
Our research questions which will be answered in the sequel, speci cally
include: 1) Dose text mining techniques for extracting and summarizing
important information from consent forms lead to information loss?
and 2) Does our approach need less time and e ort for contractual
agreements comprehension?
2</p>
    </sec>
    <sec id="sec-2">
      <title>State of the Art</title>
      <p>
        Consent forms such as terms of use and privacy policies have clear structure and
terminologies. Therefore, OBIE is a tting method for processing such texts,
since the mappings between natural language text and machine-understandable
conceptualizations is more straightforward. OBIE uses an ontology to guide the
IE pipeline and annotates the text with the ontology concepts. In recent years,
along with the increasing emergence of domain ontologies, OBIE has gained a
lot of interest. According to a survey of OBIE applications [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the most
widelyused tools for OBIE are GATE2, sProUT3 and the Stanford CoreNLP4. We have
chosen GATE due to its excellent support for OBIE.
      </p>
      <p>
        Our literature review covers speci cally license agreements and privacy policy
studies. Tl;drLegal5 is an online service that uses a manual and crowdsourced
way to present a summary of popular EULAs. Furthermore, NLL2RDF is a rst
attempt which employs NLP and ML techniques to generate RDF expressions
of license agreements [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The framework is evaluated against a goldstandard
which was created manually using Open Digital Right Language (ODRL)6 and
CC REL7 ontologies. However, NLL2RDF is able to generate only a few number
of rights and conditions due to the incomplete training data.
      </p>
      <p>
        Some e orts have speci cally studied privacy policies [
        <xref ref-type="bibr" rid="ref2 ref4 ref5">2, 4, 5</xref>
        ]. A common
approach is to use prede ned categories and supervised ML to assign classes
to policy paragraphs. Furthermore these categories are helpful for assessing the
completeness of privacy policies. The primary limitation of these studies is a lack
of su cient training data. The only proper dataset was created by the Usable
Privacy Policy Project8. OPP-115 contains 115 privacy policies from American
companies and was annotated by 3 experts into 10 categories [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Polisis9 exploits
OPP-115 to process the privacy policies and presents them in a visualized format.
2 https://gate.ac.uk/
3 http://sprout.dfki.de/
4 https://stanfordnlp.github.io/CoreNLP/
5 https://tldrlegal.com/
6 https://www.w3.org/community/odrl/
7 https://creativecommons.org/ns
8 https://usableprivacy.org/
9 https://pribot.org/polisis
To the best of our knowledge, the missing chain in Polisis is analyzing a policy's
risk factor and its compliance with the law. In future, we plan to apply ML
using OPP-115 to assign a risk score to privacy policies and identify potential
mappings between a speci c policy and data protection legislation.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Proposed Approach</title>
      <p>After a thorough literature review covering terms of use (or EULAs) and privacy
policies, we considered di erent approaches to identify the ones that are more
suitable for each type of agreement and the in/availability of training corpora.
For privacy policy analysis, we rely on the OPP-115 dataset to apply supervised
ML and train a reliable model. In contrast, there is no annotated corpus for
terms of use, and the creation of one poses a major challenge because they don't
follow a speci c structure and their scope is generally broader. Depending on the
type of an asset (software, website, digital products, etc.), the terms of use di er
signi cantly. They may contain copyright conditions, speci c rules on accessing
the service, intellectual property rights and various other content. However they
all share a common characteristic: they are written using legal terminology |
from which it is able to extract a common structure. Based on this assumption,
we apply OBIE for extracting pre-de ned classes of information.</p>
      <p>
        Having investigated the existing EULA ontologies and vocabularies, we have
chosen ODRL as the main ontology for our OBIE pipeline. It is speci ed in
W3C recommendations and has also demonstrated the highest community
endorsement. Although the focus of ODLR is digital content, it is broad enough
to cover di erent types of resources. Bene ting from Permission, Prohibition
and Duty classes and their properties (e.g., hasAction), we de ne tailored rules
utilizing GATE JAPE grammar [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Based on repeated observations and
consultation with legal experts, we enhanced the ontology to expand the coverage
of our rules, e.g., some instances are added to the Action class (which is the
`range' of hasAction property). Our nal framework extracts Permissions,
Prohibitions and Duties from an EULA.
      </p>
      <p>Although OBIE is a standard approach, there has been no prior study
utilizing OBIE for EULAs. From this point of view our application is new.
Furthermore, since we are bene ting from a standard `model' of the domain, there is a
huge potential to better structure similar documents along the same taxonomy.
Moreover, having a vocabulary to cover such legal texts can become a standard
for structuring also new documents (and not just the existing ones).
4</p>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <p>Our framework consists of two separate modules: EULAide is responsible for
processing license agreements (or terms of use) and is based on OBIE; and
KnIGHT assigns pre-de ned categories to a privacy policy paragraphs and is
built upon a supervised ML approach. Figure 1 shows the high-level architecture
of our framework and each module is presented in the following subsections.
EULAide extracts important excerpts from EULAs. As shown in the picture, a
pre-processing module performs common NLP tasks: tokenizer, sentence
splitter, POS tagger, root nder and a text- le gazetteer which contains important
keywords from EULAs (e.g., di erent synonyms for the terms `license' and
`asset'). The pre-processed EULA will be ingested to the OBIE pipeline, which
contains JAPE hand-coded grammar rules based on ODRL community speci
cation documentation10. Till now, we have implemented 15 rules, some of which
are:
i) Ontology-based annotation: separates all ontology-derived Action instances
into DutyAction &amp; PermProhAction based on the ontology speci cation;
ii) ExtractPermWords : identi es the important keywords for permissions
detection, e.g., may, can, grant, permit, etc.;
iii) ExtractPermission: extracts the whole sentence, if the pattern is matched.
Table 1 shows the steps towards extracting of a sample permission. After the
pre-processing phase, rst the text- le gazetteer produces two annotations:
License &amp; Asset. Second, the ontology-based annotation generates PermAction
annotation. Third, the extractPermWords rule res and PermWord annotation is
created. Finally, extractPermission detects the whole sentence as a Permission.
10 https://www.w3.org/TR/odrl-vocab/</p>
      <p>The summarization component clusters the similar extracted excerpts and
creates a short description for each cluster. Figure 2 shows an example of
EULAide output. The number of extracted excerpts by OBIE pipeline is 14, whereas
the summarization module has reduced the number of clusters to 9.</p>
      <p>In order to evaluate the e ciency of EULAide we conducted an experiment
to identify if the solution enables end-users to invest less time and e ort to
su ciently comprehend it. At the same time, we wanted to identify the
tradeo between the added support and the information loss expected when
applying semi-automatic IE and summarization. As a rst step, a corpus containing
twenty EULAs in their natural language texts was compiled. Two annotators
familiar with EULA texts annotated the corpus independently following an
introduction to the relevant ODRL concepts. The Inter-Annotator Agreement (IAA)
between two annotators is 90%, which indicates the production of a reliable
gold standard. To identify the cost of IE-in icted information loss, a legal
expert designed 5 multiple choice questions for four EULAs (e.g., 20 in total). All
questions are related to Permission, Prohibition &amp; Duty. In the last step 6
volunteers from the university campus (postgraduate students and sta ) were
required to answer these questions using two methods: i) reading the EULA in
full text and ii) utilizing EULAide. The results are brie y presented in section 5.
Privacy policies are legal documents stipulating how companies will gather,
manage and process customer data. They are legally required for any service that
uses, maintains or discloses data that can be used to identify an individual. In
contrast to EULAs, privacy policies must comply with a smaller set of legislation,
i.e., data protection laws. This focus enables us to perform more speci c analysis
and check compliance against speci c data protection regulation. For such
contractual agreements, we employ a deep learning approach utilizing an existing
corpus. The OPP-115 dataset is divided into paragraphs, each of which includes
annotations from three legal experts. There are two types of annotations: at the
top level each paragraph is labeled with one (or more) pre-de ned classes; and
at a lower level a class may contain speci c attributes. For example, the top level
category \User Choice/Control" can be narrowed down to: choice type, choice
scope, personal information type, purpose. We are working on training KnIGHT
with OPP-115 and extract top level classes and lower level attributes from a
policy. Having this structured information from privacy policies, we are able to:
1) check their completeness according to the legislation; 2) measure their risk
factor by analyzing the values of attributes; 3) map the top level categories to
data protection laws for more advanced users (legal experts, data o cers, etc.).</p>
      <p>The rst two goals target regular users as the intended audience, whereas
the third one is more suitable for experts. Although the OPP-115 consists
of policies de ned by American companies, the top level categories can be
mapped to GDPR. For instance, the category \First Party Collection/Use" is
related to Article 13, `Information to be provided where personal data
are collected', or \User Access, Edit &amp; Deletion\ category can be linked to
Article 16 &amp; 17 (`Right to Rectification/Erasure'). The mappings can be
as general as a whole article or as detailed as a speci c paragraph. For the
evaluation of KnIGHT, the rst two directions will be assessed using the OPP-115
as a gold standard, while the mapping accuracy should be assessed by experts.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>In this section, we will only discuss results from experiments seeking to evaluate
the support provided by EULAide to make sense of legal agreements (terms of
use). The evaluation of our privacy-policy (training data-based) approach is still
in planning stage. In order to measure the performance of the OBIE pipeline,
we have used our compiled gold standard based on the manually-annotated
examples (excluding the 10% disagreement in the IAA exercise). The evaluation
results are shown in tables 2a &amp; 2b. The ontology enhancement was a feedback
cycle during which we improved domain-speci c coverage by adding additional
instances (around 50), with the support of a legal expert. Considering the
complexity of EULAs and the 90% agreement observed between human annotators,
the results indicate that our OBIE method yields useful results and is feasible.
The (90 - 72)% information loss comes from the incomplete set of grammar rules
and ODRL instances coverage. Expanding the ontology with more concepts will
allow us to de ne more rules and will eventually increase the system accuracy.</p>
      <p>F1</p>
      <p>Tables 3a &amp; 3b present results from the previously described EULAide
usability experiment. Phase1 and Phase2 indicate di erent phases when answering
the multiple-choice questions. In the rst phase participants read EULAs (either
in their full text or utilizing EULAide) and answered the questions using their
memory. In the second phase they were allowed to use search tools for
unanswered questions. The rationale behind this setting was to recreate the baseline
method for users to check and read policies without any tool. Thus, we sought
to identify how well regular people can remember policies and how fast they can
search for information in an EULA. In practice, when one is agreeing with terms,
this process should be followed so as to be aware of the rights and regulations.
Our results verify our initial hypotheses, i.e., even though EULAide is e ected
by a (12 - 1.5 = 10.5)% information loss, it considerably saves time and e ort
spent by users to arrive to a similar level of understanding. Finally it should be
stated that although due to funding restrictions the number of selected EULAs
and participants was the bare minimum required for an experiment of this kind,
the results were su cient to indicate the value in extending and improving our
approach.</p>
      <p>Reading Phase1 Phase2
EULA
Full Text
EULAide
1185
315
75
72
152
77
(a) Average time (In Sec.)</p>
      <p>Correct Incorrect
EULA
Full Text
EULAide
67
62
8
15</p>
      <p>Unanswered in Phase1
Phase2 Phase2 Phase2
Correct Incorrect Unanswered
18.5 5 1.5
6.5
4.5
12
(b) Average percentage of questions results (%)</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>This thesis tackles the important issue of di cult-to-read legal documents and
investigates automated methods for the bene t of end-users. The experiments
conducted con rm the complexity of the task and the subjectivity of human
judgment. Somewhat counter-intuitively, we observe that agreement between
the experts is generally harder to achieve than between average users. This is
probably due to the experts' higher understanding and ability for a more critical
inspection of legal texts. A constant non-technical challenge in our e orts is to
attain commitment from legal experts on a voluntary basis. Despite the
challenges and di culties, our results so far indicate that NLP techniques combined
with OBIE and ML can be very useful to support legal text comprehension and
that with su cient funding broader experiments can be carried out.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>I would like to thank Simon Scerri, Soren Auer and Jens Lehmann for their
guidance and fruitful discussions during the development of this doctoral work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cabrio</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmero</surname>
            <given-names>Aprosio</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Villata</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>These are your rights</article-title>
          .
          <source>In: The Semantic Web: Trends and Challenges</source>
          . pp.
          <volume>255</volume>
          {
          <fpage>269</fpage>
          . Springer International Publishing (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Costante</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petkovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , den Hartog,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>A machine learning solution to assess privacy policy completeness: (short paper)</article-title>
          .
          <source>In: Proceedings of the 2012 ACM Workshop on Privacy in the Electronic Society</source>
          . pp.
          <volume>91</volume>
          {
          <fpage>96</fpage>
          . WPES '12,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2012</year>
          ). https://doi.org/10.1145/2381966.2381979
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tablan</surname>
          </string-name>
          , V.:
          <article-title>JAPE: a Java Annotation Patterns Engine (Second Edition)</article-title>
          . Research Memorandum CS{
          <volume>00</volume>
          {
          <fpage>10</fpage>
          , Department of Computer Science, University of She eld (
          <year>November 2000</year>
          ), http://www.dcs.shef.ac. uk/~diana/Papers/jape.ps
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Guntamukkala</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dara</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grewal</surname>
            ,
            <given-names>G.W.:</given-names>
          </string-name>
          <article-title>A machine-learning based approach for measuring the completeness of online privacy policies</article-title>
          .
          <source>2015 IEEE 14th International Conference on Machine Learning and Applications</source>
          (ICMLA) pp.
          <volume>289</volume>
          {
          <issue>294</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Harkous</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fawaz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lebret</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaub</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shin</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aberer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Polisis: Automated analysis and presentation of privacy policies using deep learning</article-title>
          .
          <source>CoRR abs/1802</source>
          .02561 (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Obar</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oeldorf-Hirsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The biggest lie on the internet: Ignoring the privacy policies and terms of service policies of social networking services</article-title>
          . Information, Communication &amp; Society pp.
          <volume>1</volume>
          {
          <issue>20</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Wilson,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Schaub</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Dara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.A.</given-names>
            ,
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Cherivirala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Leon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.G.</given-names>
            ,
            <surname>Andersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Zimmeck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sathyendra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.M.</given-names>
            ,
            <surname>Russell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.C.</given-names>
            ,
            <surname>Norton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.B.</given-names>
            ,
            <surname>Hovy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.H.</given-names>
            ,
            <surname>Reidenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.R.</given-names>
            ,
            <surname>Sadeh</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.M.:</surname>
          </string-name>
          <article-title>The creation and analysis of a website privacy policy corpus</article-title>
          .
          <source>In: ACL</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wimalasuriya</surname>
            ,
            <given-names>D.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Ontology-based information extraction: An introduction and a survey of current approaches</article-title>
          .
          <source>J. Inf. Sci</source>
          .
          <volume>36</volume>
          (
          <issue>3</issue>
          ),
          <volume>306</volume>
          {323 (Jun
          <year>2010</year>
          ). https://doi.org/10.1177/0165551509360123, http://dx.doi.org/10.1177/ 0165551509360123
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>