<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Refining predicates for relation extraction through thesaurus integration (abstract)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ewan Hannaford</string-name>
          <email>ewan.hannaford@glasgow.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Youcef Benkhedda</string-name>
          <email>youcef.benkhedda@manchester.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marc Alexander</string-name>
          <email>marc.alexander@glasgow.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Goran Nenadic</string-name>
          <email>gnenadic@manchester.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Riza Batista-Navarro</string-name>
          <email>riza.batista@manchester.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Manchester</institution>
          ,
          <addr-line>Oxford Road, Manchester M13 9PL</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Critical Studies, University of Glasgow</institution>
          ,
          <addr-line>Glasgow, G11 6EW</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Using natural language processing (NLP) approaches combined with humanities expertise, the Our Heritage, Our Stories project seeks to connect community archives from across the UK - archives run for and by local communities that contain unique datasets of community-generated digital content (CGDC) - with materials in The National Archives (TNA) of the UK. Foundational to this research is the refinement of NLP methods for their application to CGDC, in order to extract entities appearing in materials (people, places, dates, organisations, etc) and relations between them, and link these across collections. A key step in this process is the definition of predicates (i.e., relation types) that form the basis of the automated extraction of relationships between named entities. As there may exist many diferent linguistic formulations of the same relationship, normalisation of these heterogeneous, yet synonymous, forms to a prototypical relation is often necessary to adequately capture similar relationships appearing in diverse materials. This ensures that materials expressing the same meaning in diferent ways are recognised as doing so, and that connections can be drawn across materials as a result. Wikidata, a canonical knowledge base for linked data approaches, contains an extensive list of relation types, which researchers can map relations appearing in their materials to as a means of relation normalisation. In the Our Heritage, Our Stories project, a set of ~30 key relations relevant to CGDC were identified, which covered the core relationships between entities typically presented in community archive materials. However, these relations, as prototypical expressions, do not, and were not intended to, comprehensively cover the diverse ways in which such relations can be expressed in CGDC. As a result, in order to prevent relations that appeared in CGDC but not in lists of canonical relations going unrecognised and unrepresented, a broader range of expressions was required to capture the full range of linguistic manifestations of relationships between entities in materials. Using the Historical Thesaurus of English - the world's largest thesaurus of English - the project team identified and enriched the set of predicates selected from Wikidata with synonymous terms from similar semantic fields, subsequently integrating these expanded relation types into annotation and processing. This talk discusses this work, explaining how key relations were selected, how synonymous terms for these relations were identified from the Historical Thesaurus, and how these were integrated into the project's NLP approaches to improve the automated interpretation of relationships appearing in CGDC. In doing so, it delineates how linguistic thesauri may be used as a means of refining relation extraction and linking, enabling more comprehensive capture of relations whilst maintaining normalisation necessary for linked data approaches. Consequently, it proposes a hybrid approach to relation extraction, demonstrating how NLP methods can be supported by manual integration of linguistic expertise and resources.</p>
      </abstract>
    </article-meta>
  </front>
  <body />
  <back>
    <ref-list />
  </back>
</article>