<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Don't Stop Thinking about Tomorrow: Use Cases Demonstrating the Asymmetric Impact of Contextual Temporal Links in Knowledge Graph Evolution &amp; Retrieval1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>K. Kr</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Decentralized Information Group, Computer Science and Artificial Intelligence Lab, Massachusetts Institute of Technology 02139</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This short paper presents use cases to prompt consideration of the asynchronous and asymmetric nature of context updates when devising schemes and standards for managing and preserving decentralized knowledge graphs. As data are increasingly connected in knowledge graphs that evidence the relationships among them, an open challenge is how to manage and preserve decentralized data so that a graph updates, and a query returns, data that correctly evidences the contextual relationship. Much of the focus on managing and preserving the evolution of data has been about preserving the internal (internal to a dataset or source) history, where preservation and retrieval are synchronous. But, as demonstrated here, in many real-world use cases the correct linkage and, therefore, preservation and retrieval, is neither a temporal match nor related version match.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge graph</kwd>
        <kwd>knowledge graph evolution</kwd>
        <kwd>temporal nodes</kwd>
        <kwd>temporal relationships</kwd>
        <kwd>web standards</kwd>
        <kwd>data management</kwd>
        <kwd>data preservation</kwd>
        <kwd>data context</kwd>
        <kwd>context mapping</kwd>
        <kwd>linked data</kwd>
        <kwd>semantic web</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Context tells us about the environment in which information exists. Context educates
us as to how information is relevant, by its relation to other things - from the most
common comparators of time and location to little known events that impact or are
impacted by our data. From the inception of this workshop on Managing the Evolution
and Preservation of the Data Web (“MEPDaW”), there has been a recognition of the
importance of temporal relationships in the update, recording, storing, and retrieval of
linked data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref13 ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. More recent work has expanded on the abilities to work at scale
and with ever increasing numbers of versions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Overwhelmingly, work has
operated on the presumption that linked data should be preserved or analyzed at a moment
in time, past or present. But decentralized data providing context in an evolving graph
is not necessarily a temporal match nor an exact versioning match. This paper provides
use cases in which searching for historical versions of a knowledge graph will require
the ability to identify and retrieve data which does not share the same archival date
and/or requires the retrieval of more than one version of some but not all nodes, and
possibly edges, of the graph.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Temporal Context &amp; Graph Evolution</title>
      <p>In the initial design of a knowledge graph, the relationships are often defined based
upon knowledge or theory of a use case as a snapshot. As graphs are deployed in
production, and the data flows, it becomes apparent that an additional sort of descriptor is
required. What should be defined and where – when the impact of updates to temporal
context on the evolution of the graph is known?
2.1</p>
      <p>
        Simultaneous Context
Circumstances where a simultaneous set of facts provides context are perhaps the
easiest to call to mind. For example, periods of rain readily provide the context for many
traffic accidents [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In such a case, a decentralized graph might tie the exact time and
geo-perimeter of meteorologic data about a phenomenon [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], highway data about the
number of vehicles in the vicinity at that time from EZpass or traffic cam counts [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
and law enforcement data about accidents from published police reports [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In this
case, it is straight-forward to retrieve the data by querying event_date. Even if, as so
often occurs, smaller accidents are reported and entered on later days, retrieval of all
the graph’s data based on event_date will still be effective.
There are many circumstances in which there is a time lag between one set of facts
which provides the context for another set of facts. A common example is the
relationship between national testing scores of a local school and changes in house prices in
the district [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In such a case, the testing scores are typically released and ranked once
in a year [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ][13][14], while house prices, averages, etc. are updated at least monthly
by real estate companies and governmental agencies [15]. In these cases, the relevant
temporal nodes share neither the same name nor date. For example, the relevant score
date is not the event_date (regardless of whether that is defined as test_date or
scoring_date), but the pub_date – the first date that the scores could have been known by
others; the relevant price date is not the pub_date but arguably the offer_date – the first
date there may be evidence of the market response; and the timing of the relationship
begins at pub_date+N – the number of days after publication that it could have reached
a real estate agent or a buyer. Statistics on either side of the graph can be corrected or
updated for a particular date. For these cases, it is important to remember that, even
though there is not an exact temporal match, the modification of either set should not
break the graphed relationship. And, retrieval should be of the final corrected versions
only.
Conversely, there are instances in which retrieval should pull all the iterations, not just
the final. Consider this circumstance, where data in one store has a predictive link to
data in another store. For example, over the summer of 2021, there was a record number
of dogs surrendered to a local animal control agency and this appeared to be a predictor
of the number of households to be in distressed circumstances at the end of
Covidrelated eviction moratoria. Figure 4 shows one possible graph in which the
Surrendered_Dog_Count node is a separate and distinct daily report, but it causes only
intermittent updates to the versioning of the singular node for Updated_Evictions_Forecast.
It is possible that the iterations of the edges (and resulting node versions) may not be
consistently temporally spaced, for reasons ranging from testing and refining the
forecasting model to additional forecasting when there is a significant influx of dogs. What
then is the appropriate query to restore the history?
Another sort of context which would require the reevaluation of the relationship based
upon knowledge at a particular time, is when the change can be prompted by any node.
For example, consider advances in knowledge about human reactions to substances and
changes in grocery contents. There may be new medical practice or research reporting
– for example, the impact that Sucralose has on blood sugar [16] – which changes the
graphed labels between diabetes and numerous foods. Sucralose was recently the
leading ingredient globally for new foods and beverages with sugar-related claims [17], an
example of the constant changes to the contents of groceries [18][19] – foods, toiletries,
cleaning supplies – which can also change the nature of the label between an item and
a medical condition (e.g., allergy, celiac, diabetes).
      </p>
      <p>This particular example is complicated by the fact that, for most purposes, the
primary users of each dataset would prefer different outcomes from the updates. The
medical researcher more likely would wish to see the medical conclusion mapped to
each version of a product’s ingredient list, requiring the retention of each as a separate
node and a separate edge. While the consumer would likely prefer to see the medical
conclusion mapped only to the form of the product currently stocked on shelves,
requiring an overwrite that retains only one node and one edge.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Discussion</title>
      <p>The provided use cases show instances in which searching for historical versions of a
knowledge graph will require the ability to identify and retrieve data which does not
share the same archival date and/or requires the retrieval of more than one version of
some but not all nodes of the graph. These challenges are complicated by the data being
owned by different parties, in different subjects, who may not even be aware of the use
to which their data is put. Generally, these are not challenges that can be solved by
simply using date+n_days, as the number of days and versions may be inconsistent.
For knowledge graphs to evolve appropriately, there must be a notation to indicate, and
a mechanism to produce, the desired impact of a contextual temporal relationship. As
shown with the temporal context examples of simultaneity, lag, prediction, and
bi-directionality, graph creators need to be able to express whether a temporal update to the
data in a node should create a new node or overwrite the existing one, and whether it
can result in a change to an edge or create a new one. To facilitate historical retrieval,
there should be a standard for data owners to describe not only the data and time
produced, but also versioning methodology – for example, metadata indicating whether
data is overwritten or new date named versions produced; whether there is a marker for
the final version of iterated data; and whether there is a graphed relationship that causes
changes to this data.
13. See, e.g., http://www.globalreportcard.org/about.html (downloadable global school district
data).
14. See, e.g., https://infohub.nyced.org/reports/school-quality/information-and-data-overview
(New York City open data on education, including testing scores).
15. See, e.g., https://www1.nyc.gov/site/finance/taxes/property-rolling-sales-data.page (NYC
rolling sales data for residential real estate).
16. Pepino, Y.M., Tiemann, C.D., et al, Sucralose Affects Glycemic and Hormonal Responses
to an Oral Glucose Load, Diabetes Care, Vol. 36(9), pp. 2530-2535 (American Diabetes
Association, Sept. 2013) (https://care.diabetesjournals.org/content/36/9/2530 ).
17. “Sugar Reduction Innovation,” Aug. 2021
(https://www.foodingredientsfirst.com/analysispopup/sugar_reduction_aug_2021.html).
18. See, e.g., https://www.foodingredientsfirst.com/ (a website for the food industry with focus
on rising and declining ingredient trends).
19. See, e.g., USDA Branded Food Products Database
(https://data.nal.usda.gov/dataset/usdabranded-food-products-database/resource/cfceb689-7dab-498f-8762-707cd299646b)
(providing ingredients for branded foods).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Taelman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , et al,
          <article-title>Continuously Updating Query Results over Real-Time Linked Data</article-title>
          ,
          <source>In Proceedings of the 2nd Workshop on Managing the Evolution and Preservation of the Data Web (MEPDaW</source>
          <year>2016</year>
          )
          <article-title>co-located with 13th European Semantic Web Conference (ESWC</article-title>
          <year>2016</year>
          )
          <article-title>CEUR-WS</article-title>
          , vol.
          <volume>1585</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          , Heraklion, Crete, Greece (
          <year>2016</year>
          )
          <article-title>(http://ceurws</article-title>
          .org/Vol1585/mepdaw2016_paper_01.pdf).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Bendiken</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <article-title>Transaction-Time Queries in Dydra (Industry Paper)</article-title>
          ,
          <source>In Proceedings of the 2nd Workshop on Managing the Evolution and Preservation of the Data Web (MEPDaW</source>
          <year>2016</year>
          )
          <article-title>co-located with 13th European Semantic Web Conference (ESWC</article-title>
          <year>2016</year>
          )
          <article-title>CEUR-WS</article-title>
          , vol.
          <volume>1585</volume>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>19</lpage>
          , Heraklion, Crete, Greece (
          <year>2016</year>
          )
          <article-title>(retaining past and current state as separately addressable stores) (http://ceur-ws</article-title>
          .
          <source>org/</source>
          Vol-
          <volume>1585</volume>
          /mepdaw2016_paper_02.pdf).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Fernandez</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polleres</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Umbrich</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <article-title>Towards Efficient Archiving of Dynamic Linked Open Data</article-title>
          ,
          <source>In Proceedings of the First DIACHRON Workshop on Managing the Evolution</source>
          and
          <article-title>Preservation of the Data Web co-located with 12th European Semantic Web Conference (ESWC 2015), CEUR-WS</article-title>
          , vol.
          <volume>1377</volume>
          , pp.
          <fpage>34</fpage>
          -
          <lpage>49</lpage>
          , Portorož, Slovenia (
          <year>2015</year>
          )
          <article-title>(http://ceur-ws</article-title>
          .
          <source>org/</source>
          Vol-
          <volume>1377</volume>
          /paper6.pdf).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Quevas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Hogan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <source>Versioned Queries over RDF Archives: All You Need is SPARQL? In Proceedings of the 6th Workshop on Managing the Evolution</source>
          and
          <article-title>Preservation of the Data Web (MEPDaW) co-located with the 19th International Semantic Web Conference (ISWC 2020), CEUR-WS</article-title>
          , vol.
          <volume>2821</volume>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          , Virtual (instead of Athens, Greece) (
          <year>2020</year>
          )
          <article-title>(exploring querying in and across massive versioned archives) (http://ceur-ws</article-title>
          .org/Vol2821/paper6.pdf).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gleim</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Decker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <article-title>Open Challenges for the Management and Preservation of Evolving Data on the Web</article-title>
          ,
          <source>In Proceedings of the 6th Workshop on Managing the Evolution</source>
          and
          <article-title>Preservation of the Data Web (MEPDaW) co-located with the 19th International Semantic Web Conference (ISWC 2020), CEUR-WS</article-title>
          , vol.
          <volume>2821</volume>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>16</lpage>
          , Virtual (instead of Athens, Greece) (
          <year>2020</year>
          )
          <article-title>(referring to the resolution of synchronization possibly through TimeMaps) (http://ceur-ws</article-title>
          .
          <source>org/</source>
          Vol-
          <volume>2821</volume>
          /paper9.pdf).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>6. See, e.g., https://ops.fhwa.dot.gov/weather/q1_roadimpact.htm.</mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. See, e.g., https://www.weather.gov/media/aly/Past_Events/2015/PNS_Microburst_
          <article-title>Jun_9_2015.pdf (example of open web data re: a microburst, showing time, longitude</article-title>
          and latitude).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. See, e.g., NY Open Data https://data.ny.gov/Transportation/Annual-Average-
          <article-title>Daily-TrafficAADT-Beginning-1977/6amx-2pbv (providing average daily vehicle usage per stretch of roadway).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. https://catalog.data.gov/dataset/e-zpass
          <string-name>
            <surname>-</surname>
          </string-name>
          usage-statistics-beginning
          <article-title>-2008 (providing EZ Pass usage by year by toll plaza).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>See</surname>
          </string-name>
          , e.g., https://data.ny.gov/Transportation/Motor-Vehicle-
          <article-title>Crashes-Case-InformationThree-Year-/e8ky-4vqe (providing timestamp and DOT mileage marker for location of accidents).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>See</surname>
          </string-name>
          , e.g., https://www.opendoor.com/w/blog/how
          <article-title>-school-ratings-impact-home-prices</article-title>
          and https://www.niche.com/k12/search/best-school
          <article-title>-districts/ (offering houses for sale tied to each school district ranking).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>See</surname>
          </string-name>
          , e.g., https://nces.ed.gov/programs/digest/mrt_tables.
          <source>asp (annual release of education statistics).</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <article-title>2 All web citations are as of September 5, 2021 and are not listed separately in each reference</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>