<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Automated Directive Extraction from Policy Texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alex Lyte The MITRE Corporation Bedford</string-name>
          <email>cpfeifer@mitre.org</email>
          <email>spetersen@mitre.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>USA alyte@mitre.org</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carlos Balhana Language Technology Lab University of Cambridge Cambridge</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Craig Pfeifer The MITRE Corporation Ann Arbor</institution>
          ,
          <addr-line>MI</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Karl Branting Jim Finegan David Shin Stacy Petersen The MITRE Corporation McLean</institution>
          ,
          <addr-line>VA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>case Letters</institution>
          ,
          <addr-line>Lowercase Letters, Number Digits</addr-line>
          ,
          <country>Solid Bullet Points</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>17</volume>
      <issue>2019</issue>
      <abstract>
        <p>Federal agencies must comply with directives expressed in documents issued by authoritative sources elsewhere in the government. To automate identification of these directives, the ADEPT (Automated Directive Extraction from Policy Texts) system exploits the observation that directive sentences are usually characterized by deontic modality (e.g. “must”, “shall”, etc.) permitting the open-ended task of summarizing obligations to be reduced to a well-defined and circumscribed linguistic analysis task. ADEPT comprises a linearizer, which converts deeply nested sentences into a form that can be handled by standard parsers, a deontic sentence classifier trained on an annotated corpus of sentences drawn from representative policy documents, a semantic role analyzer, and other analytic tools for extracting and analyzing the deontic content of policy documents.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Modern administrative states are regulated by statutes, regulations,
and other authoritative legal sources that are expressed in complex,
interconnected texts. Compliance with these rules is challenging
for agencies, citizens, rule-drafters, and attorneys alike. For
agencies, compliance requires understanding changes in federal laws,
executive orders, and authoritative directives, policies, regulations,
and standards. Simply identifying and summarizing these changes,
which often originate from a multitude of sources, can be a
burdensome drain on staf resources. The diversity of authoritative
sources imposing requirements of a given nature is typified by the
proliferation of cybersecurity requirements on U.S. federal agencies.
Directives can be expressed in Executive Orders, Ofice of
Management and Budget (OMB) circulars and memoranda, Department of
Homeland Security (DHS) Binding Operational Directives (BODs),
National Institute of Standards and Technology (NIST) Federal
Information Processing Standards (FIPS), and Special Publications
(SPs). Each agency must devote staf to monitor and review multiple
streams of publications to identify changes afecting their
cybersecurity profile (i.e., policies, practices, procedures, standards, and/or
guidance).</p>
      <p>A similar monitoring task is required for all other areas within
an agency where compliance is compulsory, such as privacy, health
policy, and processing of sensitive information. An algorithmic
process that automated the identification of sentences expressing
obligations incumbent upon a given agency could significantly
reduce the burden on staf having to review a large stream of
documents. Such automated processes could provide agencies with early
warnings of pending obligations, enabling them to better plan for
implementation once the obligation is finalized.</p>
      <p>A key observation of human performance on the
documentmonitoring task is that the summaries produced by staf typically
focus on sentences that express obligations, i.e., that are
characterized by deontic modality. This suggests that the tasks of monitoring
and extracting directive sentences depend critically on the
identification of such deontic sentences. We hypothesize that exploiting
this observation will permit an important portion of the open-ended
task of summarizing obligations to be reduced to a well-defined
and circumscribed linguistic analysis task.</p>
      <p>The remainder of this paper describes the design of a system for
automated extraction of directives, ADEPT, and the evaluation of
the critical deontic-sentence classification component. Section 2
presents examples of directives and describes the characteristics
that distinguish directives from non-directives and diferent types
of directives from one another. Section 3 discusses prior related
work on modality classification, and the handling of nested
directives, that is, sentences where dependent clauses or sentential
complements share a common root clause is discussed in Section 4.
Section 5 sets forth ADEPT’s approach to identifying and classifying
directive sentences, and Section 6 describes the use of semantic role
labeling and frame instantiation to extract structured knowledge
from sentences identified as directives. The implemented ADEPT
architecture is described in Section 7, and Section 8 summarizes
and outlines future eforts.
2</p>
    </sec>
    <sec id="sec-2">
      <title>DIRECTIVE SENTENCES IN POLICY</title>
    </sec>
    <sec id="sec-3">
      <title>DOCUMENTS</title>
      <p>ADEPT is based on an analysis of the work products of subject
matter experts engaged in monitoring federal policy documents
originating from the authoritative sources such as those listed in
Section 1. Analysis of these sentences revealed that directives
typically consist of expressions of obligations on the part of an agency
or other government entity to perform or refrain from some
speciifed actions, such as:
(1)
(2)</p>
      <sec id="sec-3-1">
        <title>Agencies must establish performance goals.</title>
        <p>Agencies are required to provide narrative responses
regarding their risk management decision process.
(3) Each agency business owner is directed to ensure that 3DES
and RC4 ciphers are disabled on mail servers.
(4)</p>
        <p>Chief Information Oficers are to submit a report within 180
days.</p>
        <p>
          These directive sentences can be viewed as illocutionary [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] or
performative texts [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] that make a given action compulsory for a
given government entity (i.e., the agency or a holder of a role within
the agency). Frequently, as in sentence 1 above, directive sentences
use modal verbs, such as “must”, “shall”, “may”, and “should”, as
auxiliaries [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. However, sentences 2–4 illustrate that obligations
can be expressed without the use of modal verbs.
        </p>
        <p>In addition to these absolute, i.e., unqualified , sentences, there
are two other types of sentences that are important for some, but
not all, applications.</p>
        <p>First, some directives are qualified in the sense of expressing
either permission or weak necessity, as in the following two sentences:
(5) Senior executives may consider delaying awarding new
ifnancial assistance obligations (permission).
(6)</p>
        <p>Agencies should establish and report other meaningful
performance indicators and goals (weak necessity).</p>
        <p>Second, some sentences merely report an obligation created by a
diferent document, rather than creating an obligation themselves,
such as:
(7) Section 1 of the Executive Order requires agency heads to
ensure appropriate risk management.</p>
        <sec id="sec-3-1-1">
          <title>We term such sentences indirect obligation sentences.</title>
          <p>We exclude sentences from our set of directive sentences those
that specify the details of an obligation created in a diferent
sentence, e.g., by elaborating on the requirements of a work product
obligation:</p>
          <p>(8) Reports must enumerate performance goals.</p>
          <p>We treat these sentences as non-directives because they provide
details of obligatory actions but do not in themselves create an
obligation for an agency or other government entity. We defer
handling of these sentences to future applications.</p>
          <p>In summary, we found that directive summaries extracted from
policy documents by human experts typically have deontic force,
which may be absolute, qualified, or indirect, depending on the
construction of the sentence. We hypothesize that summaries
consisting of these deontic sentences closely match existing work
products by agency personnel who currently monitor such documents
and that summaries of this type could benefit agencies by enabling
agency personnel to quickly identify the impact of new
obligations, improving an agency’s capability for complete and timely
compliance.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>RELATED WORK</title>
      <p>
        Providing assistance to agencies in complying with complex
regulatory and policy constraints is increasingly recognized as an
important AI application. Typical examples include development of
knowledge acquisition techniques to increase the agility in public
administration [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and information retrieval techniques optimized
for regulatory texts [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Research in this area has addressed both
cross-document relationships among regulatory and statutory texts,
such as network structure [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], and within-document analysis, such
as discourse analysis of regulatory paragraphs [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and parsing
statutory and regulatory rule texts into a computer-interpretable form
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The work most closely related to the objectives of the current
work is [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], which addressed sentential modality classification of
sentences in financial regulation texts.
      </p>
      <p>
        A number of previous research projects have addressed the
general task of modal sense disambiguation in legal and government
texts. Marasović and Frank [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] developed a classifier for epistemic,
deontic, and dynamic modal categories in English and German using
a one-layer convolutional neural network (CNN) with feature maps
and semantic feature detectors, reporting better results than with
MaxEnt or a one-layer neural network. O’Neill et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] combined
a neural network with both legal-specific and more general
distributional semantic model representations to distinguish among the
deontic modalities obligation, prohibition, and permission. Wyners
and Peters [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] used a rule-based approach to extract conditional
and deontic rules from the U.S. Federal Code of Regulations. They
found that this approach worked well for a specific set of regulatory
texts, but its generality is unclear. Maat et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] compared machine
learning approaches to knowledge-based approaches for legal text
classification in Dutch legislation, finding that while machine
learning classifiers performed as well as the pattern-based model, the
pattern-based approach generalized better than the machine
learning model to new texts.
      </p>
      <p>The modality classification task addressed by ADEPT difers
from this prior work in that it focuses on the deontic distinctions
relevant specifically for the task of extracting and summarizing the
directives from administrative and policy documents, e.g.,
distinguishing deontic from non-deontic sentences and distinguishing
among the categories of deontic sentences relevant to a particular
application (e.g., absolute and qualified obligations). As discussed
below, ADEPT additionally addresses tasks both upstream from
deontic sentence detection, such as linearization of nested
directive sentences, and downstream, such as instantiation of obligation
frames and conversion of instantiated frames into a structured form
useful to agency personnel.</p>
    </sec>
    <sec id="sec-5">
      <title>HANDLING NESTED DIRECTIVES</title>
      <p>Authoritative administrative texts, including directives, regulations,
and statutes, are often expressed in the form of nested
enumerations, such as the directive set forth in Figure 1. Nested structures
are characterized by multiple dependent clauses or sentential
complements to common superordinate clauses. Such structures are
intended to express complex rules and directives in a compact and
comprehensible style by reducing textual redundancy. Human
readers can easily understand the logical structure of such sentences
because the relationships among clauses are signaled by
hierarchical relations between varying levels of enumeration symbols,
punctuation marks, and varying indentation depths.</p>
      <p>
        Unfortunately, parsers trained on standard treebanks, which
are generally based on articles from news sources such as the
Wall Street Journal, are often unable to process sentences with
nested enumerations [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Thus, until domain-specific treebanks
have been developed for legal texts which include nested sentences,
it will remain necessary to convert such sentences into a
logicallyequivalent representations that are more amenable to conventional
parsers.
      </p>
      <p>
        One approach to simplifying the syntactic structure of nested
enumerations is to convert them into a series of unnested sentences
“by starting from the root of the tree and by concatenating, for each
possible path, the phrases found until the leaves are reached” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Each depth-first traversal of this tree is a simple (non-compound)
sentence. We refer to this process as linearization. For example, the
ifrst sentence in a linearization of the nested sentence shown in
Figure 1 is:
(9)
      </p>
      <p>All agencies are required to within 30 calendar days after
issuance of this directive, develop and provide to DHS an
agency Plan of Action for BOD 18-01 to enhance email
security by within 90 days after issuance of this directive
configuring all internet-facing mail servers to ofer
STRTTLS.</p>
      <p>
        Linearization of regulatory and statutory text can be complicated
by ambiguity in the scope of logical connectives that can arise from
inconsistencies in expressing conjunction and disjunction in legal
texts [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Nested directives, on the other hand, appear to generally
be implicitly conjunctive, so linearization into a set of separate
individual directives, each corresponding to a path in the depth-first
traversal of the tree representing the logical form of the sentence,
is generally consistent with the intended semantics of the original
nested form.
      </p>
      <p>
        As a practical matter, the greatest challenge in documents
published in PDF (the primary format used by the agencies that we
support) is determining the nesting level of each constituent clause
with respect to surrounding clauses. Text extracted using standard
tools, such as Apache Tika [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Tesseract [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], does not
reliably retain the indentation depths of the original PDF. Punctuation
marks often signal the nesting level, e.g., a clause that ends with a
colon is to be followed by one or more subordinate (more deeply
nested) clauses, and a period usually indicates a leaf node. However,
there is an inherent ambiguity in sentences that follow a leaf node,
such as the sentence in the box in Figure 1: “Within 120 days after
issuance of this directive, ensuring:”. Without either an
unambiguous indication of indentation depth relative to surrounding clauses
or an enumeration mark signaling a clear relationship to other
lines of enumerated text, it is impossible to determine whether this
sentence is (1) at the level of the sentence that starts “Within 90
days”, (2) at the level of the sentence that starts “Enhance email
security by:”, or (3) the start of a new nested expression.
      </p>
      <p>The lack of accurate indentation depths in text extracted from
PDF documents and the ambiguity of the typical punctuation
conventions suggest that the enumeration and bullet symbols and
punctuation must be the source of nesting information. After all,
these are generally unambiguous for human readers. Unfortunately,
there is no canonical hierarchical practice of enumerations and
bullets; document conventions vary not just among agencies but
often within the same issuing agency as well from one document to
the next. Enumeration and bulleting formats are sometimes applied
inconsistently even within the same document. Our strategy is
therefore to make an initial traversal of each document, recording
the order of occurrence of each of a standard set of possible
enumeration styles and conventions to establish a given document’s
hierarchical structure in each section. Each nested expression is
then replaced with its linearized equivalent as determined from the
hierarchy determined in the initial pass. The Appendix sets forth
this procedure in more detail.</p>
      <p>
        Our approach difers from Dragoni et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which mapped
enumerated propositions onto a legal ontology to define the domain
of directives and their constituent subparts, in using a
conceptagnostic approach that may be better suited for domains in which
directives are frequently revised, rescinded, or recontextualized in
ways that may not be amenable to previous ontologies.
      </p>
      <p>The extraction tools described below are intended to remove
reference footnotes, HTML links, page numbers, and other
extraneous information from within the span of single extracted sentences,
but remaining bits of extraneous text create challenges for NLP
processes downstream in our pipeline, such as POS and
dependency parsing, event extraction, and modality detection. The last
step of the linearization component therefore attempts to push
these remaining items to the bottom of the linearized document as
standardized endnotes.</p>
    </sec>
    <sec id="sec-6">
      <title>DIRECTIVE SENTENCE CLASSIFICATION</title>
      <p>Our working hypothesis is that policy-document summaries
consisting of some or all categories of directive sentences described
above can be a proxy for, assist in the creation of, or supplement
manually-created compliance summaries. Thus, we focus on
classifying sentences with respect to these directive sentence categories.
5.1</p>
    </sec>
    <sec id="sec-7">
      <title>Directive-Sentence Corpus</title>
      <p>Unfortunately, none of the models or corpora developed in the
prior work on sentence modality classification described above are
directly applicable to our task. We therefore found it necessary to
develop a new annotated directive sentence corpus based on U.S.
executive-branch policy directives. Our initial focus was on OMB
Memoranda and DHS Binding Operational Directives, for which
we had examples of agency work products. We downloaded 5 years
of OMB directives from the White House website.1</p>
      <p>
        Each of the documents in the corpus was originally published
in PDF format, usually with the first page scanned and signed.
Each document was converted to plain text using the Apache Tika
software package [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In parallel, each document was processed
with Grobid [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] to identify elements such as headers and footers
that can interrupt text that spans from one page to the next. The
elements identified using Grobid were disinterleaved from the main
text and concatenated at the end of each document.2
      </p>
      <p>As described in Section 4, policy documents often contain
complex sentences, including bullet-pointed lists and enumerations, that
establish multiple distinct obligations. Accordingly, each nested
sentence in the corpus was converted into a set of simple sentences
using the linearization process described in Section 4. Each of the
resulting sentences was then annotated according to the categories
set forth in Section 2 by several annotators, including a
subjectmatter expert and several linguists.</p>
      <p>
        The resulting set of 2,582 labeled sentences served as ground
truth in the construction of the machine learning-based models
described below. The mean length of these sentences was 38 tokens.
Table 1 shows the proportion of sentences of each of the 3 directive
types that have a modal auxiliary.3 These ratios illustrate that the
presence of modal auxiliaries is neither necessary nor suficient for
directives in this domain.4
1https://www.whitehouse.gov/omb/information-for-agencies/memoranda/
2Footnote texts must be retained because they sometimes contain directives.
3Modal verbs included can, could, may, might, must, shall, should, will, or would
4This annotated corpus will be made available to researchers in 2019 at
http://matannotation.sourceforge.net/.
We converted each sentence of our corpus into a vector of semantic
role values using AllenNLP [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. These vectors were converted to
ARFF format5 and evaluated in 10-fold cross-validation using the
Weka [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] implementation support vector machine (SVM) (Platt’s
algorithm for sequential minimal optimization [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]). As shown
in Table 2, a mean F-score of 0.812 was achieved across all four
categories. A mean F-score of 0.846 (with ROC Area of 0.689) was
obtained for the binary task of distinguishing non-directives from
any of the 3 types of directive sentences.
      </p>
      <p>This experiment indicates that the deontic categories of relevance
to our task can be distinguished by a model trained on a corpus
of modest size. We anticipate that this accuracy can be improved
by expanding the annotated data set size and refining the text
extraction and linearization processes that provide input into the
classifier.
6</p>
    </sec>
    <sec id="sec-8">
      <title>SEMANTIC ROLE LABELING AND</title>
    </sec>
    <sec id="sec-9">
      <title>TEMPLATE INSTANTIATION</title>
      <p>For many agency applications, the most useful representation of
directives is often in the form of structured tables or spreadsheets
summarizing multiple sentences. Analysis of representative work
products indicated that the information of interest from each
sentence includes the following:
• Actor - the agency or ofice to which the obligation applies
• Activity - the activity that is required of the Actor
• Object - the work product to be produced by the Activity, if
any
• Time - any time-related qualification of the directed activity
• Manner - any non-time-related qualification of the directed
activity
• Modal - whether the activity is obligatory, permitted, or
suggested, as indicated by the particular modal or other verb
used to convey the deontic character of the expression, i.e.,
“must” vs. “may.”
For each directive, we instantiate a frame containing argument slots
for each of the types of information above. For example, the
instantiated frame shown in Table 3 summarizes the key information
from the following directive sentence:
(10)</p>
      <p>Within 60 days of this Memorandum’s publication agencies
must update their list of non-governmental URLs.
5https://www.cs.waikato.ac.nz/ ml/weka/arf.html</p>
      <p>
        The slots in the directive frame are a domain-specific adaptation
of standard semantic roles. We use the Semantic Role Labeling
model of AllenNLP [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to assign Propbank semantic role labels
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] to directive sentences. We then use a set of simple heuristic
rules for mapping these SRLs to the slots of our frames, e.g., a
Propbank “ARG0” is generally the Actor, “ARG1” is generally the
Object, and “Temporal” corresponds to the Time slot. Directives
expressed without a modal verb ("All agencies are required to ...")
will have no entry in the "Modal" field.
      </p>
    </sec>
    <sec id="sec-10">
      <title>7 SYSTEM ARCHITECTURE</title>
      <p>As illustrated in Figure 2, ADEPT’s directive extraction and
analysis tasks require a series of processing steps. We have adopted a
modular architecture that can accommodate a variety of alternative
components.</p>
      <p>
        The first stage of the pipeline consists of concurrent calls to the
APIs of the Tika and Grobid services ofered by their respective
Docker [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] containers. Tika outputs the PDF extraction as plain
text whereas Grobid outputs the footnotes embedded in XML. The
merge stage integrates this content and outputs a text file
consisting of disinterleaved page content followed by all footnotes. The
linearizer takes this text as input and outputs a text file containing
one linearized sentence per line. The linguistic feature
extractor converts each sentence into a feature vector of n-grams and
features derived from a dependency parse.
      </p>
      <p>An API call to the Docker container of the AllenNLP service
is then made with a JSON file containing all sentences identified
as being of the target deontic type or types (e.g., absolute). The
AllenNLP output is passed to the template instantiation stage. The
ifnal output consists of CSV and HTML files that can be loaded into
a spreadsheet or viewed through a web browser.</p>
    </sec>
    <sec id="sec-11">
      <title>8 DISCUSSION AND FUTURE WORK</title>
      <p>ADEPT illustrates how a document analysis task that imposes a
significant burden to a wide range of agencies—directive extraction—
can be addressed by deontic sentence classification in combination
with nested sentence disambiguation and semantic role labeling.
We anticipate that an ADEPT directive-extraction pilot will take
place in mid-2019 with a representative U.S. federal agency.</p>
      <p>Future work will relax ADEPT’s current simplifying assumption
that the directive content of policy documents can be determined
by analyzing individual sentences divorced from their surrounding
context. For within-document contextual information, we plan
to introduce entity resolution and link connecting sentences that
elaborate on an obligation with the obligation sentence to which
they apply. To improve cross-document contextual information,
we plan to develop techniques to detect and classify references to
other documents, particularly statements that the current document
rescinds directives from other policy documents.</p>
      <p>Automated analysis of policy documents presents a rich set of
text-analytic tasks but promises very significant rewards to both
agencies and citizens. ADEPT represents an initial realization of this
approach to improving the administrative state through modern
computational linguistics techniques.</p>
    </sec>
    <sec id="sec-12">
      <title>ACKNOWLEDGMENTS</title>
      <p>The MITRE Corporation is a not-for-profit company, chartered in
the public interest, that operates multiple federally funded research
and development centers. This document is approved for Public
Release; Distribution Unlimited. Case Number 18-4602.</p>
      <sec id="sec-12-1">
        <title>EXTRACT strings matching footnote format</title>
      </sec>
      <sec id="sec-12-2">
        <title>STORE matching strings in References array</title>
      </sec>
      <sec id="sec-12-3">
        <title>DELETE matching strings in their original positions DELETE all multiple (n-1) vertical and horizontal spacing</title>
        <sec id="sec-12-3-1">
          <title>Detect Document Section Boundaries: Identify positions of each document section to prevent enumerated elements from spanning multiple distinct lists.</title>
        </sec>
      </sec>
      <sec id="sec-12-4">
        <title>MATCH list of known section headers STORE matches in partition along with starting ofset position for each section in index READ any enumerated lists in between section boundaries</title>
        <p>Parse and Concatenate Enumerations: Map document
hierarchical enumeration conventions against diferent
symbol sets. Concatenate all directly subordinated
sentence fragments with their subordinating fragments to
form full (flat) sentences from the enumerated elements
for downstream processing later in the classification
pipeline.</p>
        <p>MATCH lines in each enumerated list within each section against
enumeration symbol style list delimited by punctuation cues
(Uppercase Roman Numerals, Lowercase Roman Numerals,
Upper</p>
      </sec>
      <sec id="sec-12-5">
        <title>Hollow Bullet Points)</title>
        <p>STORE the sequential order (i.e., layers) of enumeration styles
encountered to set document convention, where each layer begins
with its own closet set of enumeration symbols</p>
      </sec>
      <sec id="sec-12-6">
        <title>FOR lower-order layers</title>
        <p>CONCATENATE lines recursively with all parent layers
TERMINATE upon reaching new paragraph with no enumeration
symbol at the start of the line</p>
      </sec>
      <sec id="sec-12-7">
        <title>ITERATE over all sections</title>
      </sec>
      <sec id="sec-12-8">
        <title>WRITE to [FILENAME]_paths.txt file</title>
        <sec id="sec-12-8-1">
          <title>Standardize Global Enumeration: Rewrite enumeration</title>
          <p>conventions to standard format (e.g. I.iii.B.a. → 1.3.2.1.)</p>
        </sec>
      </sec>
      <sec id="sec-12-9">
        <title>FOR all enumerated lists,</title>
        <p>REWRITE each line’s enumeration symbol with its corresponding
digit based on the layer order and within-layer order</p>
      </sec>
      <sec id="sec-12-10">
        <title>WRITE to [FILENAME]_trees.txt file</title>
        <sec id="sec-12-10-1">
          <title>Post-Process Footnotes: Add previously extracted footnotes to the bottom of document</title>
          <p>APPEND footnote elements to bottom of the [FILENAME]_paths.txt
ifle under the new section header “Footnotes”</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Allen</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Saxon</surname>
          </string-name>
          .
          <article-title>More IA needed in AI: Interpetation assistance for coping with the problem of multiple structural interpetations</article-title>
          .
          <source>In Proceedings of the Third International Conference on Artificial Intelligence and Law</source>
          , pages
          <fpage>53</fpage>
          -
          <lpage>61</lpage>
          , Oxford, England, June 25-28
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Apache</surname>
          </string-name>
          tika
          <article-title>- a content analysis toolkit</article-title>
          . https://tika.apache.org/. Accessed:
          <fpage>2018</fpage>
          - 11-16.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Austin</surname>
          </string-name>
          .
          <article-title>How to do things with words</article-title>
          . Oxford U. Press, New York,
          <year>1962</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Boer</surname>
          </string-name>
          and
          <string-name>
            <surname>T. van Engers.</surname>
          </string-name>
          <article-title>An agent-based legal knowledge acquisition methodology for agile public administration</article-title>
          .
          <source>In Proceedings of the 13th International Conference on Artificial Intelligence and Law</source>
          ,
          <source>ICAIL '11</source>
          , pages
          <fpage>171</fpage>
          -
          <lpage>180</lpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Buabuchachart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Metcalf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Charness</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Morgenstern</surname>
          </string-name>
          .
          <article-title>Classification of regulatory paragraphs by discourse structure, reference structure, and regulation type</article-title>
          .
          <source>In Proceedings of the 26th International Conference on Legal Knowledge-Based Systems JURIX</source>
          , University of Bologna, Bologna, Italy,
          <year>November 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Collarana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heuss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          , I. Lytra, G. Maheshwari,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nedelchev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Trivedi</surname>
          </string-name>
          .
          <article-title>A question answering system on regulatory documents</article-title>
          .
          <source>In Proceedings of the 31st international conference on Legal Knowledge and Information Systems (JURIX)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>E. de Maat</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Krabben</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Winkels</surname>
          </string-name>
          .
          <article-title>Machine learning versus knowledge based classification of legal texts</article-title>
          .
          <source>In Proceedings of the 2010 Conference on Legal Knowledge and Information Systems: JURIX</source>
          <year>2010</year>
          :
          <article-title>The Twenty-Third Annual Conference</article-title>
          , pages
          <fpage>87</fpage>
          -
          <lpage>96</lpage>
          , Amsterdam, The Netherlands, The Netherlands,
          <year>2010</year>
          . IOS Press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8] DOCKER. https://www.docker.com/. Accessed:
          <fpage>2019</fpage>
          -01-24.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dragoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Villata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Rizzi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Governatori. Combining NLP</surname>
          </string-name>
          <article-title>Approaches for Rule Extraction from Legal Documents</article-title>
          .
          <source>In 1st Workshop on MIning and REasoning with Legal texts (MIREL</source>
          <year>2016</year>
          ), Sophia Antipolis, France, Dec.
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gardner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Grus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Tafjord</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dasigi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmitz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <article-title>Allennlp: A deep semantic natural language processing platform</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1803</year>
          .07640,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Grobid</surname>
          </string-name>
          <article-title>(or grobid) means GeneRation of BIbliographic data</article-title>
          . https://grobid. readthedocs.io/en/latest/. Accessed:
          <fpage>2018</fpage>
          -12-18.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hall</surname>
          </string-name>
          , E. Frank,
          <string-name>
            <given-names>G.</given-names>
            <surname>Holmes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pfahringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Reutemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>The weka data mining software: An update</article-title>
          .
          <source>SIGKDD Explorations</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Keerthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shevade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Murthy</surname>
          </string-name>
          .
          <article-title>Improvements to platt's smo algorithm for svm classifier design</article-title>
          .
          <source>Neural Computation</source>
          ,
          <volume>13</volume>
          (
          <issue>3</issue>
          ):
          <fpage>637</fpage>
          -
          <lpage>649</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Koniaris</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Anagnostopoulos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Vassiliou</surname>
          </string-name>
          .
          <article-title>Network analysis in the legal domain: a complex model for european union legal sources</article-title>
          .
          <source>Journal of Complex Networks</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <fpage>243</fpage>
          -
          <lpage>268</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Marasović</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Frank</surname>
          </string-name>
          .
          <article-title>Multilingual modal sense classification using a convolutional neural network</article-title>
          . In P. Blunsom,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grefenstette</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Hermann</surname>
            , L. Rimell,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Weston</surname>
          </string-name>
          , and S. W. Yih, editors,
          <source>Proceedings of the 1st Workshop on Representation Learning for NLP, Rep4NLP@ACL</source>
          <year>2016</year>
          , Berlin, Germany,
          <year>August 11</year>
          ,
          <year>2016</year>
          , pages
          <fpage>111</fpage>
          -
          <lpage>120</lpage>
          . Association for Computational Linguistics,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Morgenstern</surname>
          </string-name>
          .
          <article-title>Toward automated international law compliance monitoring (tailcm)</article-title>
          .
          <source>Technical report, LEIDOS</source>
          , INC,
          <year>2014</year>
          .
          <article-title>AFRL-RI-</article-title>
          <string-name>
            <surname>RS-TR-</surname>
          </string-name>
          2014-206.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>J. O'Neill</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Robin</surname>
          </string-name>
          , and L.
          <string-name>
            <surname>O'Brien.</surname>
          </string-name>
          <article-title>Classifying sentential modality in legal language: a use case in financial regulations, acts and directives</article-title>
          .
          <source>In Proceedings of the 16th edition of the International Conference on Articial Intelligence and Law</source>
          ,
          <string-name>
            <surname>ICAIL</surname>
          </string-name>
          <year>2017</year>
          , London, United Kingdom, June 12-16,
          <year>2017</year>
          , pages
          <fpage>159</fpage>
          -
          <lpage>168</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gildea</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kingsbury</surname>
          </string-name>
          .
          <article-title>The proposition bank: An annotated corpus of semantic roles</article-title>
          .
          <source>Comput. Linguist.</source>
          ,
          <volume>31</volume>
          (
          <issue>1</issue>
          ):
          <fpage>71</fpage>
          -
          <lpage>106</lpage>
          , Mar.
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>W.</given-names>
            <surname>Peters</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. Z.</given-names>
            <surname>Wyner</surname>
          </string-name>
          .
          <article-title>Legal text interpretation: Identifying hohfeldian relations from text</article-title>
          . In N. Calzolari,
          <string-name>
            <given-names>K.</given-names>
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Goggi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Grobelnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mazo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Odijk</surname>
          </string-name>
          , and S. Piperidis, editors,
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation LREC</source>
          <year>2016</year>
          , Portorož, Slovenia, May
          <volume>23</volume>
          -28,
          <year>2016</year>
          .
          <source>European Language Resources Association (ELRA)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[20] The Plain Writing Act of</source>
          <year>2010</year>
          ,
          <year>2010</year>
          . 111th
          <string-name>
            <surname>Congress</surname>
            <given-names>H.R.</given-names>
          </string-name>
          <year>946</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Platt</surname>
          </string-name>
          .
          <article-title>Fast training of support vector machines using sequential minimal optimization</article-title>
          . In B.
          <string-name>
            <surname>Schölkopf</surname>
            ,
            <given-names>C. J. C.</given-names>
          </string-name>
          <string-name>
            <surname>Burges</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. J</surname>
          </string-name>
          . Smola, editors,
          <source>Advances in Kernel Methods</source>
          , pages
          <fpage>185</fpage>
          -
          <lpage>208</lpage>
          . MIT Press, Cambridge, MA, USA,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Searle</surname>
          </string-name>
          .
          <article-title>Speech Acts: An Essay in the Philosophy of Language</article-title>
          . Cambridge University Press, Cambridge,
          <year>1969</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <article-title>Tesseract ocr</article-title>
          . https://opensource.google.com/projects/tesseract. Accessed:
          <fpage>2018</fpage>
          - 11-16.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wyner</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Peters</surname>
          </string-name>
          .
          <article-title>On rule extraction from regulations</article-title>
          .
          <source>Frontiers in Artificial Intelligence and Applications</source>
          , (
          <volume>235</volume>
          ),
          <year>January 2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>