<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linked Data and Holocaust Era Art Markets: Gaps and Dysfunctions in the Knowledge Supply Chain</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Laurel Zuckerman</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Independent Researcher</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bry sur Marne</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>FRANCE Laurel.zuckerman@gmail.com</string-name>
        </contrib>
      </contrib-group>
      <fpage>13</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>Without robust quality control, insights gained by exploiting linked data will be incomplete, unreliable and misleading. Much of what passes for quality control, however, is not designed to verify complete end-to-end real world processes but only certain isolated steps, leaving gaps for errors which too often are not only tolerated but considered “normal”. This paper presents the alarming results of a series of entity reconciliation tests in the domain of Nazi looted art. Only 6% of plundered Jewish collectors were correctly matched to LOD entities. The causes were multiple: missing or confusing Wikimedia content, LOD or authority files, dysfunctions across languages, redirects, problematic labelling, and a bias towards fictional characters over real people. It is not clear who, if anyone, is responsible for verifying and correcting these complex errors stemming from the intersection of multiple jurisdictions, raising the question of governance, in particular where marginalized communities are concerned. Conceived as a first step in an ongoing, iterative process, the paper concludes with some suggestions on how to monitor and improve the usefulness and reliability of linked data, as well as the urgency of doing so.</p>
      </abstract>
      <kwd-group>
        <kwd>linked data</kwd>
        <kwd>NER</kwd>
        <kwd>LOD</kwd>
        <kwd>best practices</kwd>
        <kwd>LOD quality</kwd>
        <kwd>art market</kwd>
        <kwd>Holocaust</kwd>
        <kwd>Nazi looted art</kwd>
        <kwd>erasure</kwd>
        <kwd>marginalized people</kwd>
        <kwd>Jewish art collectors</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Linked open data as imagined by Tim Berners-Lee in his seminal paper is a
decentralized network knitted together by a common respect of certain conventions.[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
However, if linked data is to fulfill its considerable promise, it needs, as Verborgh and
Sande so convincingly argued, to deal with “trivial” problems [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], such as providing
the end user with 100% accurate and complete information. To this end, identifying
and fixing errors in the overall chain of information is essential. A potentially helpful
approach might be end-to-end quality control of an industrial supply chain, with
quality monitoring and continuous improvement by domain. This is a particular
challenge for crowdsourced information that has no “owner”. The temptation to dismiss
specific pockets of poor quality as unimportant because statistically small is a
dangerous one, as the impact of even a tiny percentage of false information in a specific
domain can have such an outsized impact on trust that it casts a cloud of suspicion
Copyright © 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
over the whole.[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] The Holocaust is one such domain where accurate and complete
information is considered so important to the well-being of society that laws have
been enacted criminalizing the dissemination of certain false information.[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] What a
pity it would be then, if the semantic web were, through sloppiness, lack of quality
control and diffuse dysfunction, to contribute not to greater knowledge but to greater
ignorance.
      </p>
      <p>
        This paper looks at the small but high stakes domain of linked data for Nazi looted
art. The objective is to use successive sets of increasingly targeted tests to identify
errors in the linked data knowledge supply chain and to trace them back to their
origins so that they may be corrected. The term supply chain, borrowed from industry, is
employed to emphasize the focus on real world end-to-end processes as opposed to a
theoretical testing environment, as well as a concern for the challenges of quality
control, error detection and correction across multiple jurisdictions, involving
numerous actors, technologies, and procedures. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Method</title>
      <p>
        Named-entity linking (NEL) tests were executed from April to June 2019 and in
September 2020 using public data from two sources: “lootedart.com/news”[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a site
specialized in Holocaust –related art cases, and the German Wikipedia Liste von
Restitutionsfällen, which lists Holocaust-related art restitution cases involving 120
spoliated Jewish collectors.[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
      </p>
      <p>
        IBM Bluemix NLP [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and TextRazor [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], selected from a much larger pool of
NEL tools for ease of use, replication and interpretability, were not optimized for the
tests; they were untrained, unenhanced with no additional specific domain
information. The purpose was not to compare the NLP/NEL systems but rather to see how
the underlying linked data resources used in default mode did or did not contribute to
providing identification and context and why.
      </p>
      <p>Error verification was manual, involving multiple searches to check whether all
entities had been identified and correctly matched and, if not, what had gone wrong.
Results from each round of tests determined the research questions and the datasets
for the next, narrowing successively like a funnel.</p>
    </sec>
    <sec id="sec-3">
      <title>Entity Linking Tests</title>
      <sec id="sec-3-1">
        <title>Lootedart.com news articles</title>
        <p>
          Lootedart.com/news is published by the Central Registry of Information on Looted
Cultural Property 1933-45. Created in 2001 to provide a “central repository of
information on Nazi looting and contemporary efforts to research and resolve all
outstanding issues”, the Central Registry website aggregates news about Holocaust-related art
cases on lootedart.com/news. Lootedart.com/news articles have four characteristics
that deserve special mention in the context of these entity recognition and linking
tests. First, though the news articles themselves are recent, they contain references to
people, organizations, places, and events situated in the past. Second the deceased
individuals mentioned in the articles tend to appear in a variety of other sources,
including other news articles, books, catalogues, legal cases, scholarly documents,
museum websites, government archives and authority files. Third, the
lootedart.com/news articles concern the Holocaust and Holocaust research, which
European authorities have declared exempt from both GDPR and Right to Be
Forgotten.[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] And fourth, while lootedart.com is an English language site, it aggregates
news from other languages, which means the people mentioned may be considered
notable in languages other than English.
        </p>
        <p>Method: News articles from lootedart.com/news were pasted into the demos for
IBM BlueNLP- entities and Textrazor entities, and the output was analysed to see
what had been identified, how it was classified (“persons”, “organizations”, “places”,
etc) and which LOD entities it had been reconciled to. IBM Bluemix linked to
DBpedia while Textrazor linked to Wikipedia, Freebase and Wikidata.</p>
        <p>For each entity in the news article, the following questions were asked: Was the
entity recognized? Was it correctly classified as person/organization/place/event/etc?
Was it linked to DBpedia (Bluemix; DBpedia Spotlight), or
Wikipedia/Wikidata/Freebase (TextRazor) or anything else? Was this matching correct? If
the LOD match was incorrect, which wrong entity was returned? What was the likely
cause of the error? Are there errors for which the cause cannot be identified?
Could any patterns be detected? What (and where) might be the fix?</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results: Missing entities and mismatches for Jewish art collectors</title>
        <p>Nearly all visual artists, media outlets, museums, legal terms (including arcane
vocabularies) and most Holocaust-related institutions mentioned in the news articles
were successfully recognized and reconciled by both Bluemix and Textrazor.
However, results were very poor for art collectors and dealers mentioned in the same articles.
Textrazor, which links entities not only to Wikipedia but also to Wikidata, did better
than Bluemix which linked only to DBpedia, but both tended to omit the persecuted
Jewish collectors despite the hundreds of news article published about them.</p>
        <p>In several cases, the names of persecuted Jewish art collectors were incorrectly
classified as “places”, even though the first and last names were in the text.</p>
        <p>Books and films about the Holocaust victims or claimants, and the actors and
directors involved in created them tended to be successfully recognized and linked to
Dbpedia (Bluemix) or Wikipedia/Freebase/Wikidata (Textrazor) at high rates, while
in the very same texts, the people whose stories are dramatized tend NOT to be
recognized in DBpedia (Bluemix) with some links to Wikipedia (TextRazor: Maria
Altman).</p>
      </sec>
      <sec id="sec-3-3">
        <title>Questions raised by the differences in entity linking and selection of a data source for a second round of entity linking tests</title>
        <p>Why did the entity linking software fail to find so many persecuted Jewish art
collector and art dealers? Even when Wikipedia, Wikidata or Viaf entities exist? Why
are published results for entity linking so much better than the results observed here?
Is it a question of the texts used or the method or something else?</p>
        <p>Differences in spelling (ü vs ue; ö vs oe), languages other than English, as well as
missing, badly created or confusing authority files or Wikimedia entrees coincided
with many of the dysfunctions observed. (Emil Georg Bührle was linked to DBpedia
and Wikipedia, but Emil Buehrle was not; Wikpedia labels that included parentheses
performed worse than labels without parentheses). The failure of entity linking for
entities that existed in Wikipedia or Wikidata was of particular concern.</p>
        <p>
          The verification had to be manual because it required detailed domain knowledge
backed up by extensive research – which poses the question: on what is automated
quality evaluation of entity linking based? [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]
        </p>
        <p>NEL results for persecuted Jewish collectors in the news articles were so poor
(Bluemix made more erroneous links than correct links to persons involved in Nazi
art looting) and possible causes so numerous that the lootedart.news tests were cut
short and testing shifted to a targeted subset representing these entities to try to
identify the sources of error.
3.2</p>
      </sec>
      <sec id="sec-3-4">
        <title>Liste von Restitutionsfällen: Named-Entity Linking for spoliated art collectors</title>
        <p>The German Wikipedia page Liste Von Restitutionsfällen was initially created on
March 23, 2009 and, at the time of testing, had been updated by 60 separate users,
most recently on April 27, 2019. A corresponding page exists in the French Wikipedia
but not in the English language Wikipedia.</p>
        <p>The data used for the entity linking tests were drawn from column three of the
Wikipedia list, which contains, in German, both the name of the Jewish collector and
the name of the institution which received the claim. There are 223 rows with 120
names. Some names appear more than once because the families have filed multiple
claims for different works of art. The assumption is that these 120 names reflect a
crowdsourced consensus of German Wikipedia contributors concerning the most
notable Jewish art collectors who were victims of Nazi persecution. There are, in
addition, 156 footnotes, most of which identify news articles or books that mention the
Jewish collectors and their cases.</p>
        <p>All 120 names had been the subject of numerous news articles, in the New York
Times, BBC, NPR, Wall Street Journal, Guardian, Spiegel, Süddeutsche, Le Monde,
Le Figaro, and other contributing publications. Some had Wikipedia entries in several
languages, some only in one language, some no Wikipedia page at all. More than half
were referenced in both Viaf and Wikidata.</p>
        <p>Method: The names of 120 Jewish collectors were copied and pasted, as one list,
into the demo for the entity linking tools IBM Bluemix and Textrazor or via the API.
The results on screen and in JSON were analysed, using Google searches and
authority files to verify whether each match was correct.1</p>
      </sec>
      <sec id="sec-3-5">
        <title>Results: IBM Bluemix failed to link 94% of the spoliated art collectors</title>
        <p>The results confirm the observations made in the initial round of tests on
lootedart.com/news items. Less than 6% of the names were correctly matched to
DBpedia. Of 120 Jewish collectors mentioned in the German Wikipedia list of
Restitution cases, only seven were linked to the correct person in DBpedia by Bluemix.
These were Jacques Goudstikker, Alfred Hess, Max Silberberg, Sophie
LissitzkuyKüppers, Richard Semmel, Max Emdem, Eduard Einschlag, and Leo Bendel. Due to
technical problems in the German DBpedia, however, none of the correct DBpedia
links actually functioned; the identity was verified by consulting the corresponding
German Wikipedia pages.</p>
        <p>Bluemix failed to match the other 113 Jewish collectors to DBpedia. Additional
tests and verifications were needed to understand why.</p>
        <p>In three cases – for Paul Stern, Arthur Goldschmidt, and Arthur Feldmann –
Bluemix found a DBpedia entity, but it was for the wrong person. (The correct person had
no Wikipedia page.)</p>
        <sec id="sec-3-5-1">
          <title>Quantity</title>
          <p>7
3
113
data
IBM
Bluemix NLP</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>IBM BLUEMIX NEL results for the 120 spoliated art collectors in the</title>
          <p>German Wikipedia Liste von Restitutionsfällen
Correctly linked to Dbpedia entities
Incorrectly linked to wrong Dbpedia entities
NOT correctly linked by IBM Bluemix to DBpedia
https://de.wikipedia.org/wiki/Liste_von_Restitutionsf%C3%A4llen
https://natural-language-understanding-demo.ng.bluemix.net/
%
5,83%
2,50%
94,17%
1https://docs.google.com/spreadsheets/d/1ua3lv2KljSEipg1oauS9bEEDYY4AP8OcnMCtrTxZ
qSw/edit?ts=5f649400#gid=0
Questions raised: What was the cause of the 94% Named-Entity Linking failure in
Bluemix?</p>
          <p>Was the problem due to a glitch in disambiguation? A confusion between several
similar-looking entities?</p>
          <p>Was the problem due to language? Issues with spelling differences had been noted
in the first round of tests. What else might be at work?</p>
          <p>Was the problem due to a technical issue, either with the NPL product used or with
the entity knowledge base?</p>
          <p>Was the entity knowledge base outdated, so that information about the names,
while present live, were not in the old download used?</p>
          <p>Was there a problem due to recent data protection laws?
Did the Google Knowledge Graph identify them?</p>
          <p>Bluemix reconciles entities to DBpedia. Was there a specific issue with DBpedia
for these datasets?
3.3</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>Targeted tests of components to pinpoint defects</title>
        <p>There are many places for failure in the long and complex supply chain that underlies
entity linking. Bluemix and Textrazor, like many NLP tools, rely on linked data
assembled by DBpedia and the Wikimedia foundations, as well as many other
components.</p>
        <p>Method: To try to diagnose the problem – to see where the “break” had occurred, a
new round of small, targeted tests were performed using four tools that frequently
play an important role in the knowledge supply chains of cultural heritage and news
organizations.</p>
        <p>
          • OpenRefine: Wikidata [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ],
• OpenRefine ULAN [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ],
• DBpedia Lookup2[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
• Google Knowledge Graph API [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ],
OpenRefine is widely used by cultural heritage institutions for data cleaning and
entity linking, notably with Wikidata and more recently, the Union List of Artists Names
(ULAN). TMS, the leading museum management software in the USA, offers
integration with ULAN for entity lookup for collections management. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]
Tests focused on identifying patterns in the errors of the 120 names and identifying a
small, representative subset that could be used for quality testing and monitoring.
2 DBpedia graciously offered assistance in understanding how datasets are indexed and filtered.
        </p>
        <p>Proposed solution
Who can</p>
        <p>fix?
content - Create Wikipedia in English
non tech for spoliated collector
content - Create Wikipedia in English
non tech for spoliated collector
content, improve description labels,
non tech add birth, death
content
Content</p>
        <p>Getty
tech or
content
tech</p>
        <p>Delete redirect
Create missing entities in
ULAN, correct type errors
improve NLP NEL tech
content, create Wikipedia entity in
non tech English
Old datafile used by NEL
tool, omits newer entities
selection tech- Show dates and filters of
of dataset procedure datafile used by NEL tool
3 See results
https://docs.google.com/spreadsheets/d/e/2PACX-1vSexBfiDOrW4Ce8I6EVgFMfHNCQYnDflJDtVvXh0sxCgt4mIZx7t7cLaKBOgxOd0jaHdggjh_lJm62
/pubhtml</p>
        <p>Homonym creates confusion, late - NLP
prevents selection of entity
Wikidata exists but has been
filtered out of LOD dataset</p>
        <p>middle, tech, pro- Don't filter Wikidata. Better
late - NLP cedure integrate with NLP NEL tools
late - NLP
tech</p>
        <p>Index and use all languages,
in NLP tools (DBpedia)
content, If no tech fix, create in
Engnon tech lish Wikipedia
tech
tech
content,
non tech
tech</p>
        <p>
          Verify, improve NEL
improve NLP NEL tech
content redirect in Wikipedia
NLP improvement, using
additional info about entity
Discussion: The table above lists dysfunctions observed. It is a mix of technical, user
content and integration problems in different parts of the process for the specific
domain of spoliated collectors. Experience in complex logistics supply chains suggests
that these issues can be resolved and quality improved in this domain by applying an
end-to-end approach involving all the actors in the process.[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]
        </p>
        <p>Legacy choices concerning which language an entity was created in (content) and
how it was indexed or filtered by the LOD data aggregators like DBpedia used by
NEL tools (tech) had a major impact on end results. Entities created in the German
Wikipedia but not in the English Wikipedia were not picked up by DBpedia Lookup,
or Google KG. The spoliated collector Walter Westfeld, for example, is referenced in
the German Wikipedia, Wikidata, VIAF, GND and ISNI, as well as hundreds if not
thousands of newspaper articles, books and scholarly papers, but Westfeld was still
invisible to IBM Bluemix, Textrazor, Google Knowledge Graph, and ULAN
OpenRefine reconciliation.</p>
        <p>More than half (68) of the 120 spoliated collectors (including Walter Westfeld)
were successfully matched to Wikidata entities using OpenRefine reconciliation
suggestions. However these links were not picked up in earlier tests with Textrazor which
performed entity linking using Wikidata, which suggests that Textrazor’s NEL demo
may be using a smaller or less recent subset of Wikidata.</p>
        <p>
          OpenRefine ULAN linking was successful for only nine out of 120 names, less
than 10%. Manual verification confirmed that names of the spoliated collectors
simply did not exist in ULAN. This has major implications both cultural heritage
institutions that perform data cleaning and reconciliation using OpenRefine, as well as users
of the TMS collections management system and Europeana, both of which use ULAN
for entity linking. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]
        </p>
        <p>
          Including languages other than English in key DBpedia indexes and extractions
could solve a significant portion of the problems.. Performance may be improved by
translating information contained in foreign languages into English.[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]
Alternatively, a content based approach could ensure that all Wikipedia pages exist also in
English – perhaps in the framework of a Wikipedia project like Women Scientists.[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Lessons Learned and Recommendations</title>
      <sec id="sec-4-1">
        <title>On the cumulative impact of English-language bias in a multilingual knowledge supply chain for linked data</title>
        <p>
          Concern about multilingual entity linking is not new [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]; however it would appear
that the cumulative effect of a bias favoring English in all stages of the knowledge
supply chain may be underappreciated.
        </p>
        <p>
          At every stage – the creation of Wikipedia content [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], data cleaning4, data
selection, filtering, indexing, NEL development and product marketing – English is treated
differently from other languages, which can find themselves excluded from datasets.
Even inclusive, multilingual knowledge platforms like Wikidata can be subject to
retroactive filtering according to criteria that favor English, such as existence in
Wikipedia or number of links.[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]
        </p>
        <p>Each small action, taken independently and with no malicious intent, results in a
winnowing of knowledge stored in languages other than English. The impact of such
an exclusionary process becomes problematic in dealing with knowledge, like that
related to the Holocaust, which is stored in languages other than English, because it
occurred in places where people spoke languages other than English.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Knowledge is specific to a domain, and NEL testing practices ignore this fact</title>
        <p>Global recall scores are meaningless when it comes to assessing the quality of
entity linking in a specific domain. The mass of sports and celebrity and artist-related
information, which can be correctly linked, obscures, when aggregated, a devastating
failure to recognize and link entities in a specialized domain like art collectors
spoliated during the Holocaust.</p>
      </sec>
      <sec id="sec-4-3">
        <title>The Essential Role of Human Domain Experts</title>
        <p>
          It should be emphasized that a link does not mean a correct link. Human domain
experts are needed to identify such errors.[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] The question of how data quality is
evaluated is also posed, as no automated system can catch these errors. There is
arguably no “acceptable” fail rate given the sensitivity of this data. There is a growing
awareness of the importance of data quality in linked data [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], too often quality is
thought of as an isolated content management issue not as a supply chain challenge.
4 How many developers or data wranglers think it is “normal” to remove foreign characters as
part of “cleaning”?
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Recommendations</title>
      </sec>
      <sec id="sec-4-5">
        <title>Open the Black Box (and know the contents)</title>
        <p>NLP NEL tools are often plugged in, as a kind of blackbox component, by cultural
heritage institutions, news organizations, developers and scholars. Yet the
completeness and accuracy of NEL depends on filters, selections and rules made in the
laboratories of these components which neither developers nor end users know nor
understand. In the case of spoliated art collectors, many NEL failures were traced back to a
language filtering choice made by DBpedia: some products index only English with
the result that entities created only in German are not linked.</p>
        <p>Awareness about exactly what is in the black boxes (dates of dataset, filters
applied, operations) and how this impacts end-to-end quality in NEL should be
encouraged as it makes it easier to avoid problems and find solutions.</p>
        <p>For DBpedia Lookup, for example, a solution could be to include German and
other key languages in the index (tech solution at the DBpedia level); or users could
create the missing entities in the English Wikipedia (user content solution). Another
option would be for developers of NEL tools to shift to the multilingual Wikidata.</p>
        <p>Domain specific Data Quality Dashboard
Quality Monitoring should be regular because even names which have been
successfully reconciled in past tests can suddenly fail to reconcile due to the creation of
homonyms, duplicate records or other changes. Likewise, every element in this knowledge
chain is constantly evolving - the content, the tech, the procedures, and the tools used
to test; it is essential to verify that they continue to work together correctly and to spot
quickly any new defects in the process. This suggests that NEL should be tested and
validated: 1) by specific micro-domain, 2) by end users who know and care about the
accuracy and completeness of the information being transmitted, 3) at regular
intervals, 4) with a regular test set, and 5) with results published on a public Data Quality
Dashboard.</p>
      </sec>
      <sec id="sec-4-6">
        <title>Domain specific Test Datasets</title>
        <p>A standard domain specific test set, like that for spoliated collectors, makes results
easy to replicate and interpret, facilitating error correction. Test sets could include:
• Names previously linked correctly
• Names previously linked incorrectly or not linked at all
• Information about the status of these names in Wikipedia, Wikidata, Viaf etc
• Texts (news article, provenances, legal documents...) that are known to
contain these names.5
5 Due to link rot it is recommended to retrieve the text as well as an internet archive link</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Concluding Remarks</title>
      <p>The entity reconciliation results in this paper point to serious defects in linked data
in the domain of looted art. If NEL fails on the best documented names- publicly
known spoliated Jewish collectors drawn from a crowdsourced Wikipedia page
what are the implications for the rest of the spoliated Jewish collectors in linked data?</p>
      <p>They, like so many marginalized populations, will be invisible in linked data, as if
they had never existed at all.</p>
      <p>It is at the intersections of content, process and tech, over multiple jurisdictions and
long periods of time, that the battle for end-to-end quality is won or lost.</p>
      <p>A solid methodology for monitoring LOD erasure will make it easier to identify
defects and take corrective action, in a virtuous cycle that continuously improves NEL
results and reduces the number of errors that need correction.</p>
      <p>The methods explored here can be applied to any domain, in particular those that
concern marginalized populations documented in languages other than English.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] CERN, Tim Berners-Lee's proposal</article-title>
          , http://info.cern.ch/Proposal.html
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Verborgh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Vander</given-names>
            <surname>Sande</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          (
          <year>2020</year>
          ),
          <article-title>The Semantic Web identity crisis: in search of the trivialities that never were, Semantic Web Journal</article-title>
          , IOS Press, Vol.
          <volume>11</volume>
          No.
          <issue>1</issue>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>27</lpage>
          . https://ruben.verborgh.org/articles/the-semantic
          <article-title>-web-identity-crisis/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Färber</surname>
          </string-name>
          , Michael et all,
          <source>Linked Data Quality of DBpedia</source>
          , Freebase, OpenCyc, Wikidata, and
          <string-name>
            <surname>YAGO</surname>
          </string-name>
          ,
          <source>Semantic Web</source>
          <volume>1</volume>
          (
          <year>2016</year>
          )
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          , IOS Press, http://www.semantic-webjournal.net/system/files/swj1366.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Yad</given-names>
            <surname>Vashem</surname>
          </string-name>
          ,
          <source>Holocaust Denial Laws and Other Legislation Criminalizing Promotion of Nazism</source>
          , https://www.yadvashem.org/holocaust/holocaust-antisemitism/holocaust-deniallaws.html
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Robinson</surname>
          </string-name>
          , Carol, Malhotra, v.,
          <article-title>Defining the concept of supply chain quality management and its relevance to academic and industrial practice</article-title>
          ,
          <source>International Journal of Production Economics</source>
          , Volume
          <volume>96</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>3</given-names>
          </string-name>
          ,
          <issue>18</issue>
          <year>June 2005</year>
          , Pages
          <fpage>315</fpage>
          -337, https://doi.org/10.1016/j.ijpe.
          <year>2004</year>
          .
          <volume>06</volume>
          .055
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Central</given-names>
            <surname>Registry</surname>
          </string-name>
          for Information on Looted Cultural Property, https://lootedart.com/news
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>German</given-names>
            <surname>Wikipedia</surname>
          </string-name>
          , Liste von Restitutionsfällen, https://de.wikipedia.org/wiki/Liste_von_Restitutionsf%C3%
          <fpage>A4llen</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>IBM</given-names>
            <surname>Bluemix</surname>
          </string-name>
          , https://www.ibm.com/demos/live/natural-language-understanding/selfservice/home
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Textrazor</surname>
          </string-name>
          , https://www.textrazor.com/demo
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>The</given-names>
            <surname>International Holocaust Remembrance Alliance</surname>
          </string-name>
          (IHRA), https://www.holocaustremembrance.com/stories/reference-holocaust-gdpr
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Gruetze</surname>
          </string-name>
          ,
          <string-name>
            <surname>Toni</surname>
          </string-name>
          , et all,
          <article-title>CohEEL: Coherent and efficient named entity linking through random walks</article-title>
          ,
          <source>Journal of Web Semantics, Volumes</source>
          <volume>37</volume>
          -38,
          <year>March 2016</year>
          , Pages
          <fpage>75</fpage>
          -89, https://doi.org/10.1016/j.websem.
          <year>2016</year>
          .
          <volume>03</volume>
          .001
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>[12] OpenRefine, https://www.wikidata.org/wiki/Wikidata:Tools/OpenRefine</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>[13] Getty Research Institute, Getty Vocabularies OpenRefine Tutorial, https://www.getty.edu/research/tools/vocabularies/obtain/getty_vocabularies_openrefine_tutori al.pdf</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>DBpedia</given-names>
            <surname>Lookup</surname>
          </string-name>
          , https://wiki.dbpedia.org/lookup
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Google</given-names>
            <surname>Knowledge Graph</surname>
          </string-name>
          <string-name>
            <surname>API</surname>
          </string-name>
          , https://developers.google.
          <article-title>com/knowledge-graph Dataset file with Wikidata, Viaf and birth, death dates for spoliated collectors</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <article-title>TMS Collections and eMuseum by Gallery Systems</article-title>
          , https://www.canada.ca/en/heritageinformation-network/
          <article-title>services/collections-management-systems/collections-managementsoftware-vendor-profiles/tms-emuseum-gallery-systems-profile</article-title>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>[17] see attached Datasets https://docs.google.com/spreadsheets/d/e/2PACX1vTh2segiyBrwwIGi53198QFPj1xgwhwlwfMDM1cV9Z1IHI8K6E1BjA9P_riodQKJKcFDaLimAhMnIu/pubhtml?gid=1954686210&amp;single=true</mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Understanding</given-names>
            <surname>Supply Chain Excellence - Best Practices</surname>
          </string-name>
          and
          <source>Case Studies | MSU Online</source>
          , https://www.michiganstateuniversityonline.com/resources/supply-chain/supply-chainexcellence/
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19] https://www.slideshare.net/antoineisaac/designing
          <article-title>-a-multilingual-knowledge-graphdcmi2018</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Gesese</surname>
            ,
            <given-names>Genet</given-names>
          </string-name>
          <string-name>
            <surname>Asefa</surname>
          </string-name>
          , and
          <string-name>
            <surname>Alam</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sack</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <article-title>Semantic Entity Enrichment by Leveraging Multilingual Descriptions for Link Prediction</article-title>
          , ArXiv, abs/
          <year>2004</year>
          .10640 (
          <year>2020</year>
          ), https://arxiv.org/abs/
          <year>2004</year>
          .10640
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>[21] https://en.wikipedia.org/wiki/Wikipedia:WikiProject_Women_scientists.</mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <article-title>Muhao et all, Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment</article-title>
          ,
          <source>Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-17)</source>
          , (
          <year>2017</year>
          ) https://www.ijcai.org/Proceedings/2017/0209.pdf
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>[23] https://en.wikipedia.org/wiki/Wikipedia:Systemic_bias</mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Ismayilov</surname>
          </string-name>
          ,
          <article-title>Ali et all, Wikidata through the Eyes of DBpedia, Editor(s): Aidan Hogan</article-title>
          ,
          <source>Semantic Web</source>
          <volume>0</volume>
          (
          <issue>0</issue>
          )
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          IOS Press
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Fink</surname>
            ,
            <given-names>Eleanor E.</given-names>
          </string-name>
          ,
          <article-title>Overview and Recommendations for Good Practices, American Art Collaborative (AAC) Linked Open Data (LOD) Initiative (</article-title>
          <year>2018</year>
          ), https://s3.amazonaws.com/assets.saam.media/files/documents/2020- 07/OverviewandRecommendationsAccessible.pdf
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Aroyo</surname>
          </string-name>
          ,
          <article-title>Lora, Data excellence: Better data for better AI</article-title>
          ,
          <year>ODSC 2020</year>
          (
          <year>2020</year>
          ) https://www.slideshare.
          <article-title>net/laroyo/data-excellence-better-data-for-better-ai</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>