<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>T e xt mi ni n g g e n o m e -wi d e a s s o ci ati o n st u d y lit er at ur e</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>T h o m a s R o wl a n d s a n d Ti m B e c k</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>U ni v ersit y of L ei c est er , U ni v ersit y R o a d, L ei c est er , L E 1 7 R H , U K</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>A b str a ct G e n o m e -wi d e ass o ci ati o n st u di es ( G W A S) h a v e i m pr o v e d o ur u n d erst a n di n g of dis e as e a eti ol o g y b y i d e ntif yi n g g e n eti c v ari a nts ass o ci at e d wit h c o m pl e x h u m a n tr aits a n d dis e as e p h e n ot y p es. T h e G W A S C e ntr al r es o ur c e e n a bl es br o a d a n d c o n v e ni e nt i nt e gr ati v e vis u alis ati o n of, a n d a c c ess t o, s u m m ar y -l e v el G W A S d at a. G W A S r e p ort e d i n t h e bi o m e di c al lit er at ur e ar e m a n u all y c ur at e d a n d i m p ort e d i nt o t h e d at a b as e. T o a c c el er at e t h e i m p ort pr o c ess, at a ti m e w h e n t h e si z e of G W A S ar e i n cr e asi n g , w e ar e usi n g n at ur al l a n g u a g e pr o c essi n g t o st a n d ar dis e p u bli c ati o n f ull t e xt a n d t a bl es, a n d t e xt mi ni n g t o e xtr a ct G W A S d at a. W e ar e als o d e v el o pi n g a n a n n ot at e d f ull -t e xt G W A S c or p us t o e v al u at e t h e a c c ur a c y of t h e t e xt mi ni n g al g orit h m . K e y w or d s 1 G e n o m e -wi d e ass o ci ati o n st u di es , G W A S C e ntr al, t e xt mi ni n g, bi o-o nt ol o gi es 1. I ntr o d u cti o n H e alt h r es e ar c h is a d v a n c e d t hr o u g h g e n o m e -wi d e ass o ci ati o n st u di es ( G W A S) w hi c h pr o vi d e a d e e p er u n d erst a n di n g of dis e as e a eti ol o g y b y d et e cti n g ass o ci ati o ns b et w e e n g e n eti c m ar k ers a n d dis e as e p h e n ot y p es i n p o p ul ati o n s a m pl es. G W A S C e ntr al ( w w w. g w as c e ntr al. or g) is a c o m pr e h e nsi v e c oll e cti o n of s u m m ar y -l e v el G W A S d at a [ 1]. G e n eti c m ar k er d at a ar e st a n d ar dis e d wit h d b S N P i d e ntifi ers a n d p h e n ot y p es ar e a n n ot at e d usi n g M e di c al S u bj e ct H e a di n gs ( M e S H) an d H u m a n P h e n ot y p e O nt ol o g y ( H P O) t er ms. B uil di n g o n t h e ri c h s e m a nti c p h e n ot y p e a n n ot ati o n l a y er, a c or e s u bs et of G W A S d at a is m a d e a v ail a bl e as R D F n a n o p u bli c ati o ns [ 2]. R e c e ntl y, t h er e h as b e e n a l ar g e i n cr e as e i n t h e a m o u nt of d at a r e p ort e d b y i n di vi d u al st u di es t h at i n v esti g at e g e n eti c ass o ci ati o ns wit h li pi d o m e, m et a b ol o m e a n d mi cr o bi o m e p h e n ot y p es. T his , i n t ur n, r e q uir es eff orts fr o m bi o c ur at ors t o i nt er pr et a n d e xtr a ct i n cr e as e d a m o u nts of ass o ci ati o n d at a fr o m p u bli c ati o ns. As t h e si z e of G W A S i ncr e as e, r e q uiri n g e v er l ar g er d at as ets t o b e e xtr a ct e d fr o m t h e lit er at ur e a n d i m p ort e d i nt o d at a b as es, t h er e is a n e e d f or n e w bi oi nf or m ati cs c a p a biliti es t o s u p p ort s c al a bl e d at a c ur ati o n. W e h a v e d e v el o p e d a n at ur al l a n g u a g e pr o c essi n g ( N L P) t o ol kit t o e xtr a ct G W A S e ntiti es a n d r el ati o ns fr o m t h e s ci e ntifi c lit er at ur e. T h e t o ol kit li n ks p h e n ot y p e e ntiti es wit h bi o-o nt ol o g y t er ms a n d us es t h e t e xt mi ni n g Bi o C st a n d ar d f or m at f or s h ari n g t e xt d o c u m e nts a n d a n n ot ati o ns. T h e t o ol kit c a n b e us e d t o pr o c ess n e w ' u ns e e n' p u bli c ati o ns as w ell as p u bli c ati o ns t h at h a v e pr e vi o usl y b e e n s e e n a n d c ur at e d. E ntiti es t h at h a v e pr e vi o usl y b e e n m a n u all y e xtr a ct e d b y d at a b as e c ur at ors c a n b e m a p p e d b a c k o n t o p u bli c ati o ns t o g e n er at e n e w a n n ot at e d c or p or a i n t h e Bi o C st a n d ar d. H er e, w e pr es e nt A ut o C O R P us f or st a n d ar disi n g p u bli c ati o n t e xts a n d t a bl es, G W A S -T a g g er f or a n n ot ati n g c ur at e d G W A S p u bli c ati o ns, a n d G W A S -Mi n er f or e xtr a cti n g G W A S e ntiti es a n d r el ati o ns fr o m u ns e e n p u bli c ati o ns. 2. I m pl e m e nt ati o n of a n at ur al l a n g u a g e pr o c e s si n g t o ol kit 2. 1. A ut o -C O R P u s : p u bli c ati o n t e xt a n d t a bl e st a n d ar di s ati o n</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>converted to the BioC JSON format with each publication section (e.g., introduction, methods, results)
annotated using the Information Artifact Ontology (IAO). A directed graph formed from s2e1c,t8i4o9n
headers from 2,441 open access publications is used by a section prediction algorithm to assign the IAO
classification. Since BioC does not provide native support for tabular data, the publication tables are
converted to a BioC compliant ta-bJlSeON structure. Allabbreviations declared withinapublication
are converted to JaSON output that relateesach abbreviation with the full definitioTn.he abbreviation
collection supports text mining tasks such as named entity recogni(tNioEnR) by including
abbreviations unique to individual publications that are not contained withi-nonbtoi ologies.
2.2.</p>
    </sec>
    <sec id="sec-2">
      <title>GWAS-Tagger: annotation of seen publications</title>
      <p>The GWAS-Tagger application searches BioC fu-tlelxt and tabl-eJSON files for GWAS entities and
relations that have previously been extreadct from a publication because of, for example, manual
biocuration or text mining. Entities and relations are written to the BioC an-JdSOtaNblefiles using the
BioC annotation format. The following GWAS entities, and relations between them, can be provided
to GWAS-Tagger to find within publications:
• Genetic marker given as a dbSNP identifier.
• Phenotype given as a MeSHor HPO term. Synonyms and variations of text such as
hyphenation, roman numeral usage and plurals are used if an exact term match is not found.
• P-value association of the genetic marker and phenotype given as a number with optional
exponential notation.
2.3.</p>
    </sec>
    <sec id="sec-3">
      <title>GWAS-Miner: text mining of unseen publications</title>
      <p>GWAS-Miner is a hybri-dbased NLP applicationthat usesthe SciSpaCy data model for procedures
such as tokenisation, pa-rotf-sentence tagging, dependency parsing, together witrhules for NERusing
a combination of bio-ntologies, vocabularies and pattern matchingB. ioC full-text files are processed
to identify key sentences which include relevdaanta. Sentencestructure and contexatre analysed using
a dependency tree alongside the use of shortest dependency path to isolate relationGsWhipAsS.-Miner
is designed to be easily updated with new ontology data as soon as ontologies are updated.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Discussion and future work</title>
      <p>We have used GWA-STagger to annotate 700 publications that have been curated by GWAS
Central. We aer manually evaluating these annotations, and editing them where necessary, to build a
gold standard ful-ltext GWAS corpus. The corpus will be used to assess the precision, recall and F1
score of GWASM-iner. The supplementary materials of GWAS publicatiocnans contain large amounts
of GWAS data, so we will extenAduto-CORPus to process supplementary files. Converting
heterogenous supplementary formats to the BioC standard will optimise the text and table contents for
text mining. Finally, we will incorporaGteWAS-Miner into the GWAS Central data import pipeline.</p>
    </sec>
    <sec id="sec-5">
      <title>4. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rowlands</surname>
          </string-name>
          , T. Shortearn,d
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Brookes</surname>
          </string-name>
          , GWAS Central:
          <article-title>an expanding resource for finding and visualising genotype and phenotype data from gen-owmidee association studies</article-title>
          .
          <source>Nucleic Acids Res</source>
          . (
          <year>2022</year>
          ) doi: 10.1093/nar/gkac1017.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Free</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Thorissaonnd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Brookes</surname>
          </string-name>
          ,
          <article-title>eSmantically enabling a genom-ewide association study database</article-title>
          .
          <source>J Biomed Semantics</source>
          . (
          <year>2012</year>
          )
          <article-title>3(1):9</article-title>
          . doi:
          <volume>10</volume>
          .1186/-21044810-3-9.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shorter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Popovici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A. R.</given-names>
            <surname>McQuibban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Makraduli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Yeung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rowlands</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Posma</surname>
          </string-name>
          ,
          <article-title>-ACuOtoRPus: A Natural Language Processing Tool for Standardizing and Reusing Biomedical Literature</article-title>
          . Front Digit Health. (
          <year>2022</year>
          )
          <volume>4</volume>
          :
          <fpage>788124</fpage>
          . doi:
          <volume>10</volume>
          .3389/fdgth.
          <year>2022</year>
          .
          <volume>788124</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>