<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mapping landmark cases in the U.S. legal system</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rilder S. Pires</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erneson A. Oliveira</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlos G. O. Fernandes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>João A. Monteiro Neto</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasco Furtado</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Banco do Nordeste do Brasil S.A.</institution>
          ,
          <addr-line>60743-902 Fortaleza, Ceará</addr-line>
          ,
          <country country="BR">Brasil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centro de Ciências Jurídicas, Universidade de Fortaleza</institution>
          ,
          <addr-line>60811-905 Fortaleza, Ceará</addr-line>
          ,
          <country country="BR">Brasil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Empresa de Tecnologia da Informação do Ceará</institution>
          ,
          <addr-line>Governo do Estado do Ceará, 60130-240 Fortaleza, Ceará</addr-line>
          ,
          <country country="BR">Brasil</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Laboratório de Ciência de Dados e Inteligência Artificial, Universidade de Fortaleza</institution>
          ,
          <addr-line>60811-905 Fortaleza, Ceará</addr-line>
          ,
          <country country="BR">Brasil</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Mestrado Profissional em Ciências da Cidade, Universidade de Fortaleza</institution>
          ,
          <addr-line>60811-905 Fortaleza, Ceará</addr-line>
          ,
          <country country="BR">Brasil</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Programa de Pós Graduação em Informática Aplicada, Universidade de Fortaleza</institution>
          ,
          <addr-line>60811-905 Fortaleza, Ceará</addr-line>
          ,
          <country country="BR">Brasil</country>
        </aff>
      </contrib-group>
      <fpage>97</fpage>
      <lpage>103</lpage>
      <abstract>
        <p>Court decisions and emblematic legal cases are central elements of Law. They influence practitioners, scholars and public oficers at the same time they define and shape the legal reality and its boundaries. Despite being an overlooked area, understanding how legal cases establishes connections and relationships can provide important insights not only about how influence and impact are built, but also to identify influential cases that are not listed as landmarks. Here, we explore data from the U.S. legal system by modeling it as a citation network fed with 360 years of legal cases. We characterize the probability distributions for degree-centrality measures of the network and find a power-law behavior for the in-degree probability distribution with an exponent  ≈ 2.66. We also obtain the probability distribution for landmarks according to their in-degree and out-degree in order to find the region in the in-degree× out-degree space where landmarks are more likely to be found. Finally, we highlight some extreme special cases and make some considerations about the ratio between the number of landmarks and the total number of legal cases in a given spot of the in-degree× out-degree space.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;U</kwd>
        <kwd>S</kwd>
        <kwd>legal system</kwd>
        <kwd>Legal cases</kwd>
        <kwd>Landmark cases</kwd>
        <kwd>Complex networks</kwd>
        <kwd>Citation networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Precedents are fundamental to understanding the United States (U.S.) legal system since they
are the pillars of Common Law, the judicial system in Anglo-Saxon countries. Unlike Civil Law
(where statutes are the foundations of the judicial apparatus), Common Law is based on the
previous judgments of the courts, which establish the rules to be followed, persuading or binding
judges to that series of decisions [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. By legal point of view, there are some special cases,
namely “landmark cases”, that become relevant by setting key legal concepts or interpretations
and in doing so they influence a great number of other cases along the years. In recent years,
an increasing number of studies have been proposed to characterize legal networks through
mathematical models [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ]. Despite the eforts of such studies, properly defining the
properties of a landmark case through quantitative approaches remain an open problem in law
research areas.
      </p>
      <p>
        Here, we use a model based on concepts of complex networks fed with 360 years of data in
order to characterize landmark cases in the U.S. legal system. We know that vertices with high
degree-centrality measures play an important role in the information dynamics of the network,
especially in citation networks [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Such fact suggests that "landmarks" could be easily identified
as "outliers" with the highest in-degrees and that out-degree may have some importance for
landmarks since their degree of "innovation" could be measured through it. For these reasons,
the main contribution of our study is to perform a topological map suggesting the location of
landmark cases. We observe that the most cited cases in U.S. legal system are not landmarks. In
addition, we find a well-located region with all landmark cases and, even within this region, the
top cited landmark cases are exceptions. This allows us to shed some light on understanding
the structure and evolution of U.S. law.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Datasets</title>
      <p>
        We use the Caselaw Access Project (CAP) and the Historic Supreme Court Decisions (HSCD)
open datasets [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], which are compiled and made available by the Harvard and the Cornell
Law Schools, respectively. The former is composed of citation records and other retrieved
information for each digitalized court decision (legal case) in the U.S. legal system. In this work,
we only use the citation records from the CAP dataset, which is a CSV file with ≈ 5 million
records (≈ 380 MB), where each row has a legal case ID and all other legal cases IDs that cite
it. The later is a list of referral court decisions (landmark cases) from the Supreme Court. The
HSCD dataset is also a CVS file with 538 records (≈ 4 KB).
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The Model</title>
      <p>
        In order to characterize the U.S. legal system, we use the citation network model, where the
vertices  are legal cases and the directed edges  = (,  ) are citations in the CAP dataset,
from a newer legal case  to an older legal case  . In fact, our model is a network similar
to those used to describe paper citations [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Furthermore, we remove a few duplicated edges
in order to have a perfect acyclic network. The numbers of vertices and edges of the citation
network are  ≈ 5 million and  = 43 million, respectively. We emphasize that such amount
of citations unveils an additional computational challenge in our modeling. Mathematically, a
network is characterized by its adjacency matrix A. This matrix is defined in such way that each
of its elements , are either “1” if there is an edge between  and  or “0” otherwise [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ].
Some measures, namely vertex-centrality measures, can be derived from A in order to define
some kind of global importance of a given vertex in the network. Two basic examples of these
measures are the degree-centrality measures [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], defined as
      </p>
      <p>in() = ∑︁ , and out() = ∑︁ , ,
 
(1)
where in() and out() are usually called in-degree and out-degree of vertex , respectively.
In practical terms, in() represents the number of legal cases that cite the legal case  and
out() corresponds to the number of legal cases that is cited by the legal case .</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>In order to understand the structure of citations in the U.S. legal system, we perform topological
measures on the citation network. Precisely, we focus our analysis in the in and out centrality
measures due the large size of the network. These measures have the advantage of being simple,
i.e., they are easily computed even for large networks. Despite the simplicity of these measures,
they allows us to understand basic aspects of this network.</p>
      <p>
        We show both in and out probability distributions for the U.S. legal system (Fig. 2a). We
ifnd that the probability distribution for in is described by a power law with exponent  ≈ 2.66
characteristic of other scale-free networks [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. The  exponent was obtained using
Maximum Likelihood Estimation [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and assuming that a power-law distribution describe
the data. For the out probability distribution, we are able to identify a long-tailed behavior,
however we have no strong evidence to support a power-law hypothesis. In Figure 2b, we
show both in and out probability distributions for the landmark cases. Such distributions are
drastically diferent from those shown in Fig. 2a. Numerically, landmark cases have in average
in ≈ 1, 252.5 and out ≈ 44.6. In contrast with the general case, where in ≈ out ≈ 8.96,
landmark cases have in and out, respectively, ≈ 140 times and ≈ 5 times higher than usual
legal cases.
(a)
10−2
10−4
)
(K 10−6
P
10−8
10−10
10−12
(b)
out
      </p>
      <p>In Figure 3a, we show the probability distribution of cases of the U.S. legal system as function
of in and out. We observe that, in general, the legal cases are more likely to be in a region
with small values of in and out. The landmark cases, however, have a higher probability of
being in regions of large in and out as shown in Fig. 3b. The black line, shown in Fig. 3a,
highlights the areas where the probability distribution of landmark cases, shown in Fig. 3b, is
greater than zero. In Figure 3a, we also observe that there are regions, where the probability
distribution is higher than zero, beyond the limiting line defined by landmark cases.</p>
      <p>In Table 1, we show the special legal cases highlighted in Fig. 3. These cases were chosen
according to the following criteria: case “A” is the case with highest in, case “B” is the case with
highest out, and cases “C” and “D” are the legal and landmark cases, respectively, closer to the
(1, 252.5, 44.6) point. We choose the point (1, 252.5, 44.6) as a reference, since it corresponds
to coordinates of in and out for the landmark cases.</p>
      <p>The case Henry v. New Jersey Department of Human Services, 204 N.J. 320, 9 A.3d 882 (2010),
103
tou 102
K
101
100100
101
102
103
104</p>
      <p>105
Kin
10−5
10−6
10−7
10−8
10−9
Kin
10−2
10−4
10−6
10−8
10−10
10−12
cites more than 2, 900 cases. While it deals primarily with a claim related to the New Jersey Law
Against Discrimination, the incredible high number of citations is connected to a secondary
issue raised during the judgment. During the case, the Supreme Court of New Jersey debated
an issue related to the composition of the Court as the Honorable in charge of delivering the
opinion in this case was nominated to the court in a temporary assignment. This secondary
issue debate led Justice Rivera-Soto to cite more than 2, 000 cases supporting his argument that
the court was unconstitutionally constituted. The identification of anomalous cases can be an
interesting staring point to qualitative approach investigating the relevant secondary issues
that normally are shadowed by the initial discussion and do not receive the attention they need.</p>
      <p>On the opposite side, Anderson v. Liberty Lobby, Inc. (1986) is a very important case where
the U.S. Supreme Court set the standards guiding the acceptance or not of a summary judgment
request. A summary judgment will happen when there are no factual issues to discuss and
the trial court would analyze only matters of law. This case is highly influential, being cited
more than 65, 000 times because it set the mandatory rules that every case pledging a motion of
summary judgment needs to fulfill. Knowing better this network can reveal not only the level
of influence of the case, but also indications of in which jurisdictions the case is more often
cited and even if there is a rise of new cases questioning or reinforcing its standards.</p>
      <p>The ratio land/all in the in-degree× out-degree space is shown in Fig. 4. Here, land and
all are the total number of legal and landmark cases, respectively, inside of a given bin defined
by the same divisions used in Fig. 3. We observe that such ratio tends to be greater in regions of
high in and out, i.e., in regions of very low probability even for landmark cases.
104
103
ou 102
t
K
101
100</p>
      <p>Nland/Nall100
10−1
10−2
10−3
10−4
100
101
102
103
104</p>
      <p>105</p>
      <p>Kin</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In summary, we performed a numerical data analysis in order to establish a topological map
characterizing the location of landmark cases in the U.S. legal system. Precisely, we modeled the
U.S. legal system as a citation network fed with 360 years of digitalized documents and focused
our analysis on the in-degree and out-degree centrality measures. By evaluating in and out
for all legal cases, we obtained the probability distributions for these measures and found a
power-law decay for the probability distribution of in with an exponent  ≈ 2.66. We also
characterized the landmark cases and evaluated their in and out. Probability distributions of
in and out for the landmark cases were found to be drastically diferent from those found
for all legal cases, showing that landmark cases have in and out ≈ 140 times and ≈ 5 times
higher than usual legal cases, respectively. Moreover, we showed that there is an area in the
in × out space where landmarks are more likely to be found. Surprisingly, there are regions
beyond of the limiting area of landmark cases where the probability distribution for all legal
cases is higher than zero. This result could be confirmed by identifying some special legal cases
where in or out were greater than the limiting values of the landmark region. Using a similar
approach, we also observed that it is possible to find a landmark case and an usual legal case
with very similar in and out. In order to understand the relative occurrence of landmark
cases inside the region where they are more likely to be, we studied the land/all ratio in this
region and found that it is greater in regions of very low probability for both landmarks and
usual legal cases. As perspective for future works, we will propose a temporal analysis of the
citation network as well as the introduction of other centrality measures in order to better
characterize legal cases.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We gratefully acknowledge CNPq, CAPES, FUNCAP, BNB, and the Edson Queiroz Foundation
for financial support.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. V.</given-names>
            <surname>Calvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Coleman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Coleman</surname>
          </string-name>
          , S. Coleman,
          <article-title>American law and legal systems</article-title>
          , Prentice Hall,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Schauer</surname>
          </string-name>
          ,
          <article-title>Is the common law law</article-title>
          ,
          <source>HeinOnline</source>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fortunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. T.</given-names>
            <surname>Bergstrom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Börner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Helbing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Milojević</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Petersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Radicchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sinatra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Uzzi</surname>
          </string-name>
          , et al.,
          <source>Science of science, Science</source>
          <volume>359</volume>
          (
          <year>2018</year>
          ). URL: https://doi.org/10.1126/science.aao0185.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Whalen</surname>
          </string-name>
          ,
          <article-title>Legal networks: The promises and challenges of legal network analysis, Mich.</article-title>
          <string-name>
            <surname>St. L. Rev.</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <fpage>539</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>I.</given-names>
            <surname>Carmichael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wudel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jushchuk</surname>
          </string-name>
          ,
          <article-title>Examining the evolution of legal precedent through citation network analysis</article-title>
          ,
          <source>NCL Rev</source>
          .
          <volume>96</volume>
          (
          <year>2017</year>
          )
          <fpage>227</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Koppaka,
          <article-title>Semantics-based legal citation network</article-title>
          ,
          <source>in: Proceedings of the 11th international conference on Artificial intelligence and law</source>
          ,
          <year>2007</year>
          , pp.
          <fpage>123</fpage>
          -
          <lpage>130</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Newman</surname>
          </string-name>
          , Networks, Oxford university press,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H. L.</given-names>
            <surname>School</surname>
          </string-name>
          , Caselaw access project,
          <year>2021</year>
          . URL: https://api.case.law/v1/citations/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C. L.</given-names>
            <surname>School</surname>
          </string-name>
          , Historic supreme court decisions,
          <year>2021</year>
          . URL: https://www.law.cornell.edu/ supct/cases/name.htm.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Fundamentals of complex networks: models, structures and dynamics</article-title>
          , John Wiley &amp; Sons,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T. H.</given-names>
            <surname>Cormen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Leiserson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Rivest</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stein</surname>
          </string-name>
          , Introduction to algorithms, MIT press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Redner</surname>
          </string-name>
          ,
          <article-title>How popular is your paper? an empirical study of the citation distribution</article-title>
          ,
          <source>The European Physical Journal B-Condensed Matter and Complex Systems</source>
          <volume>4</volume>
          (
          <year>1998</year>
          )
          <fpage>131</fpage>
          -
          <lpage>134</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Roco</surname>
          </string-name>
          ,
          <article-title>Patent citation network in nanotechnology (1976- 2004)</article-title>
          ,
          <source>Journal of Nanoparticle Research</source>
          <volume>9</volume>
          (
          <year>2007</year>
          )
          <fpage>337</fpage>
          -
          <lpage>352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Eom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fortunato</surname>
          </string-name>
          ,
          <article-title>Characterizing and modeling citation dynamics</article-title>
          ,
          <source>PloS one 6</source>
          (
          <year>2011</year>
          )
          <article-title>e24926</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Clauset</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Shalizi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <article-title>Power-law distributions in empirical data</article-title>
          ,
          <source>SIAM review 51</source>
          (
          <year>2009</year>
          )
          <fpage>661</fpage>
          -
          <lpage>703</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>