<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The psycho-env corpus: research articles annotated for knowledge discovery on correlating mental diseases and environmental factors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hui Wang</string-name>
          <email>hui.1.wang@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quan Sun</string-name>
          <email>quan.sun@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anika Oellrich</string-name>
          <email>anika.oellrich@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Honghan Wu</string-name>
          <email>honghan.wu@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard Dobson</string-name>
          <email>richard.j.dobson@kcl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Biostatistics and Medical Informatics, King's College London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Informatics, King's College London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Psychiatry, Psychology &amp; Neuroscience, King's College London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <fpage>36</fpage>
      <lpage>40</lpage>
      <abstract>
        <p>While the published scientific literature is used in a biomedical context such as building gene networks for disease gene discovery, it seems to be an undervalued resource with respect to mental illnesses. It has been rarely explored for the purpose of gaining psychopathology insights. This limits our capability of better understanding the underlying mechanisms of mental disorders. In this paper we describe the psycho-env corpus, which aims at annotating published studies for facilitating knowledge discovery on pathologies of mental diseases. Specifically, this corpus focuses on the correlations between mental diseases and environmental factors. We report the first preliminary work of psycho-env on annotating 20 articles about two mental illnesses (bipolar disorder and depression) and two particular environmental factors - light and sunlight. The corpus is available at https://github.com/ KHP-Informatics/psycho-env.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The success stories of cognitive computing (e.g., IBM
Watson’s Jeopardy game) and deep learning (e.g., DeepMind’s
AlphaGo) have sparked a massive wave of using artificial
intelligence (AI) to improve numerous aspects of our daily
life. Not surprisingly, healthcare is among the hottest areas.
For example, IBM Watson is now utilised in decision
support for lung cancer at the Memorial Sloan Kettering Cancer
Center. However, AI models require data to derive better
understanding of the underlying mechanisms of diseases before
they can really improve existing treatments or increase the
recovery rate. Unfortunately, the lack of data is a major hurdle
in many areas of the clinical domain, such as understanding
the pathologies of mental illnesses.</p>
      <p>As with other diseases, it has been established that mental
illnesses are influenced in their origins and pathology by
environmental factors. For example, it has been found that higher
rates of schizophrenia occur in people of Caribbean origin
than general population living in the UK [Fung et al., 2006].
To date, no complete list of environmental factors for all
existing mental illnesses has been compiled that can be used for
patient screening and planning treatment strategies [Rutter,
2005].</p>
      <p>While the published scientific literature is used in a
biomedical context such as building gene networks for
disease gene discovery [Lage et al., 2007] or symptom
networks of inheritable human disorders [Zhou et al., 2014], it
seems to be an undervalued resource with respect to
mental illnesses. It has been rarely explored for the purpose
of gaining psychopathology insights. The potential of this
resource lies within the amount and variety of data
available: all journals that publish scientific results are covered
mostly since 1966, though some even date back to 1809.
Although there is a body of work trying to identify
“extended” phenotypes [Oellrich et al., 2016; Groza et al., 2015;
Collier et al., 2015], however, none of these efforts included
environmental factors, which are necessary to understand
gene-phenotype relationships. In order to make use of this
tremendous resource for finding potential environmental
factors that (i) cause, (ii) contribute to and (iii) influence the
origin and pathology of mental illnesses, (AI backed) automated
methods are needed to digest the large quantities of existing
data.</p>
      <p>In order to facilitate this endeavour, data collection and
annotation would be required to identify relevant studies and
the representation of environmental factors in the published
literature. In this paper we describe the psycho-env corpus1,
which is a manually curated dataset from the abstracts of
20 published studies on associations between two mental
illnesses (bipolar disorder and depression) and one particular
environmental factor - light. We believe this is the first effort
to produce curated corpus for knowledge discovery on
associations between mental illness and environmental factors.</p>
      <p>In the next section, we introduce the article selection,
annotation process, annotation tool used and data format of
annotations. In section 3, we describe the psycho-env corpus
and discuss the limitation of this work. Finally, we conclude
our work in section 4.</p>
      <p>1https://github.com/KHP-Informatics/
psycho-env
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Materials and methods</title>
      <sec id="sec-2-1">
        <title>Article selection</title>
        <p>In this preliminary study, we limited our scope on two
types of mental disorders (i.e., bipolar and depression) and
one particular environmental factor - light (including
sunlight and light in general). A manual retrieval method was
adopted to search and select articles from various
bibliographic databases and search engines. This was to ensure that
we were able to identify the most relevant and representative
investigations in this domain for the pilot study. The search
and selection process are briefly described in the following.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Literature search</title>
        <p>The bibliographic databases and search engines used were
MEDLINE (accessed via PubMed search engine), Web of
Science and Google Scholar. The aim was to look for
relevant and representative research articles including clinical
studies, case reports and clinical trials published during the
period from May 1877 to May 2017.</p>
        <p>The terms used for searching disorders included: bipolar,
manic and depression, while terms for environmental factors
included sunlight, “light therapy” and phototherapy. In some
situations, extra constrains were added to narrow down the
search results, e.g., clinical trial, case reports and etc.</p>
        <p>In general, we found PubMed combined with Google
Scholar can produce the most comprehensive list for our
searches. For example, when searching sunlight and bipolar
disorder, PubMed results contained 7 relevant hits, Google
Scholar had 6, and Web of Science gave 5. All combined,
there were 8 distinct relevant hits. The overlap between
PubMed and Google Scholar was 5 - PubMed brought in 2
new results and Google Scholar added 1, while all results
from Web of Science were covered by other two search
services.</p>
        <p>Also, we found the terminologies used in the literature are
quite heterogenous. For example, when denoting the usage of
light in the therapy, many different terms were used -
brightlight therapy, light therapy and phototherapy. Therefore, we
found it necessary to follow the reference graph of articles to
check and include more articles or search terms.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Article selection</title>
        <p>The studies were selected based on the following inclusion
criteria:
• published as an original article in a peer-reviewed
journal
• designed as a clinical trial, pilot study or case report
• used light or sunlight as one of the investigation aspects
or treatment alternatives
• subjects were diagnosed as bipolar disorder or
depression
When reviewing the articles, curators were asked to extract
the following information to create a correlation between
mental illnesses and environmental factors. When combined
together, the annotated items should be able to a) capture the
most important aspects for deriving the correlations and b)
form a concise description of the study. For well-defined
clinical concepts like disorders, phenotypes and clinical
measurements, the curators were asked to map them to UMLS
(Unified Medical Language System)2 concepts using a UMLS
search tool.</p>
        <p>1. The most important finding(s) of the study (e.g.,
Bipolar inpatients in E rooms (exposed to direct sunlight in
the morning) had a mean 3.67-day shorter hospital stay
than patients in W rooms [Benedetti et al., 2001]).
2. Environmental factors. Although this preliminary study
focused on light only, other types of environmental
factors might need to be annotated as well because they
were used in the study to derive or measure light factors,
such as “latitudes 6.3 to 63.4 degrees from the
equator”. Type of environmental factors including, but not
limited to: sunlight exposure, seasonal pattern, sunlight
in springtime, natural light, 36 collection sites from 23
countries, and monthly climate variables.
3. Environmental factor classification or measurement.</p>
        <p>This type of information includes the conceptual
classification or quantity metrics for environmental factors
investigated in the study, such as meteorological data on
light intensity, the amount of sunlight exposure (i.e.
insolation), maximum monthly increase in solar insolation
and etc.
4. Mental disorders. As mentioned earlier, two types of
diseases were to be curated in this work: bipolar and
depression disorders. Any diseases that are specific types
of these diseases need to be annotated, which include,
but not limited to, bipolar I disorder, recurrent
depression, non-seasonal depression, and rapid cycling bipolar.
5. Investigation aspects of disorders - the aspects of
disease pathologies or phenotypes that were investigated in
the study, such as the onset of bipolar disorder, mood
swings, length of hospitalization and plasma melatonin
levels.
6. Diagnosis methods (if available), such as Young Mania</p>
        <p>Rating Scale (YMRS).
7. Patient cohort information including number of patients,
patient demographic information, and control/case
settings.</p>
        <p>8. Data collection methods and data sources, such as
patient records and/or direct interviews and NASA Surface
Meteorology and Solar Energy (SSE) database.
9. Data analysis methodologies, such as Autoregressive
Integrated Moving Average (ARIMA) method.</p>
        <p>To the best of our knowledge, this is the first attempt to
curate literature in this particular domain. A large part of
the curation is unknown to us, for example, what aspects of
diseases were studied and how they were quantified, what
terminologies were used to describe both clinical and
environmental concepts, how environmental factors were measured
and etc. Considering this underdeveloped nature, we adopted
an agile curation process, which was designed to be adaptive
and able to achieve continuous improvement. The idea was
borrowed from the agile software development. Technically,
articles were partitioned into several subsets and curations
were conducted on each subset at a time. After each curation
step, a curator meetup would be arranged to discuss problems
encountered and the lessons learned, and subsequently
propose amendments on the curation guidelines for improving
the next rounds. We found this iterative process and efficient
inter-curator communications very helpful and effective.
2.3</p>
      </sec>
      <sec id="sec-2-4">
        <title>Annotation tool and annotation data format</title>
        <p>A browser based annotation tool, PsychoEnv annotator,
was used for annotating articles. The tool is backed
with an automated article highlighting service described
in [Wu et al., 2017]. PsychoEnv annotator is available on
Github: https://github.com/KHP-Informatics/
psycho-env. Figure 1 is a screenshot of PsychoEnv
annotator being used for annotating a PubMed article. Features of
the tool include:
• Easy to setup: the annotation tool is a Chrome
extension and the backend service is cloud based. Any article
with an online XHTML version (e.g., PubMed article
abstracts) is available for annotating immediately
without the need of any preprocessing.
• Easy to use: all curation operations are browser based,
which minimises the learning curve of curation process.
In addition, the free text labelling allows project-wise
acronyms, which speeds up the process.
• Easy to share: associating annotations with
webaddressed articles makes the annotations directly
retrievable either for the browser visualisation by using
PsychoEnv annotator or for software agents by RESTful
API calls.
Annotation node locator
1. a jQuery3 selector that locates the
parent element of the text node,
where the annotation appears;
2. an integer number that indicates
the index of the text node within
its parent’s children list.</p>
        <p>For example: a locator can be {
Selector: ABSTRACTTEXT:eq(1), Index: 0 }
The offsets have two integer
components: start offset and end offset, where
start offset indicates the start position
of the annotated text in its annotation
node’s text content and end offset
indicates the end position.</p>
        <p>The text content of the annotation</p>
        <p>The type of the annotation
• Structure preserving: compared to most existing
annotation tools, PsychoEnv annotator is featured by its unique
capability of locating annotations on the XHTML DOM
tree of the articles’ web pages (see annotation node
locator in table 2). This associates the annotations with
semi-structured DOM trees and, in turn, brings these tree
structures as additional and easy-to-consume features to
software models.</p>
        <p>The annotation data format is a 5-element tuple as
described in Table 2.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and discussion</title>
      <sec id="sec-3-1">
        <title>Corpus description</title>
        <p>The psycho-env corpus resulted in 27 annotated text nodes
that mark mental disorder mentions, 30 annotated text nodes
that mark environmental factors, 25 annotated text nodes that
mark environmental factor classifications/measurements and
23 annotated sentences marked as important findings. These
numbers are summarized in Table 3 which also shows the
average number of annotations and range of annotations per
article in the 20 articles in the corpus.</p>
        <p>The psycho-env corpus was selected to represent bipolar
and depression disorders associated with two environmental
factors - sunlight and (general) light. The aim was to have a
similar coverage on each of the four sub-domains (as shown
in Table 1) so that we could cover relatively diverse topics
within a preliminary study. We summarised the major types
of annotations in table 4. Duplicated instances have been
removed using a syntax approach - string comparison . The
first observation is that the environmental concepts seem to
be very heterogeneous (1.4 per article for light factors and
2.05 per article for light measurements) even when we
limited the scope on light only. However, a close inspection on
the list of instances revealed that many different terms might
mean the same concepts. This suggests the necessity of
having a consistent terminology so that different mentions of the
same instances can be mapped. The second interesting
observation is that the numbers of specific disorders, phenotypes
and their measurements are relative large considering only 2
disorders were selected. This suggests very little overlaps
between studies, which might imply that the curation could be
very efficient in terms delivering new knowledge.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Discussion</title>
        <p>The main purpose of this preliminary study is to conduct a
small scale case study on limited types of mental illnesses
and environmental factors. Therefore, the number of
documents annotated is rather small. But it has resulted with a
very valuable experience, which gave us a good
understanding about the quality and representation of environmental
factors and their associations with mental disorders. Particularly,
the typed annotations as summarised in table 4 can be used
to populate controlled vocabularies or ontologies to represent
knowledge in this domain.</p>
        <p>The corpus covers four subdomains of associations of
mental disorders and environmental factors as depicted in table 1.
The authors are confident that they have covered the most
representative studies in the top 3 subdomains. However,
regarding the last subdomain - Light to depression, due to a
relatively large body of available studies, the selected four
articles might not cover the most representative studies.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In order to facilitate knowledge discovery on the pathologies
of mental disorders, we initiated work on psycho-env corpus,
which is dedicated to curating the associations between
mental illnesses and environmental factors from published
literature. The first version reported in this paper focused on
bipolar and depression disorders associated with lights, and was
curated from abstracts of 20 articles. Both the annotation tool
and the corpus are open source and publicly available.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>The work was supported by NIHR Biomedical Research
Centre for Mental Health, the Biomedical Research Unit for
Dementia at the South London, the Maudsley NHS
Foundation Trust and Kings College London, and European Union’s
Horizon 2020 research and innovation programme under
grant agreement No 644753(KConnect).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Benedetti et al.,
          <year>2001</year>
          ]
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Benedetti</surname>
          </string-name>
          , Cristina Colombo, Barbara Barbini, Euridice Campori, and
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Smeraldi</surname>
          </string-name>
          .
          <article-title>Morning sunlight reduces length of hospitalization in bipolar depression</article-title>
          .
          <source>Journal of affective disorders</source>
          ,
          <volume>62</volume>
          (
          <issue>3</issue>
          ):
          <fpage>221</fpage>
          -
          <lpage>223</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Collier et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Nigel</given-names>
            <surname>Collier</surname>
          </string-name>
          , Anika Oellrich, and
          <string-name>
            <given-names>Tudor</given-names>
            <surname>Groza</surname>
          </string-name>
          .
          <article-title>Concept selection for phenotypes and diseases using learn to rank</article-title>
          .
          <source>Journal of biomedical semantics</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>24</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Fung et al.,
          <year>2006</year>
          ]
          <string-name>
            <given-names>WL</given-names>
            <surname>Alan Fung</surname>
          </string-name>
          , Dinesh Bhugra, and
          <string-name>
            <surname>Peter B Jones</surname>
          </string-name>
          .
          <article-title>Ethnicity and mental health: the example of schizophrenia in migrant populations across europe</article-title>
          .
          <source>Psychiatry</source>
          ,
          <volume>5</volume>
          (
          <issue>11</issue>
          ):
          <fpage>396</fpage>
          -
          <lpage>401</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Groza et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Tudor</given-names>
            <surname>Groza</surname>
          </string-name>
          , Sebastian Ko¨hler, Sandra Doelken, Nigel Collier, Anika Oellrich, Damian Smedley, Francisco M Couto,
          <string-name>
            <given-names>Gareth</given-names>
            <surname>Baynam</surname>
          </string-name>
          , Andreas Zankl, and
          <string-name>
            <surname>Peter N Robinson</surname>
          </string-name>
          .
          <article-title>Automatic concept recognition using the human phenotype ontology reference and test suite corpora</article-title>
          .
          <source>Database</source>
          ,
          <year>2015</year>
          :bav005,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Lage et al.,
          <year>2007</year>
          ]
          <string-name>
            <given-names>Kasper</given-names>
            <surname>Lage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E Olof</given-names>
            <surname>Karlberg</surname>
          </string-name>
          , Zenia M Størling,
          <source>Pa´ll I Olason</source>
          , Anders G Pedersen,
          <article-title>Olga Rigina</article-title>
          , Anders M Hinsby, Zeynep Tu¨mer, Flemming Pociot,
          <string-name>
            <given-names>Niels</given-names>
            <surname>Tommerup</surname>
          </string-name>
          , et al.
          <article-title>A human phenome-interactome network of protein complexes implicated in genetic disorders</article-title>
          .
          <source>Nature biotechnology</source>
          ,
          <volume>25</volume>
          (
          <issue>3</issue>
          ):
          <fpage>309</fpage>
          -
          <lpage>316</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Oellrich et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Anika</given-names>
            <surname>Oellrich</surname>
          </string-name>
          , Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam Shah, Olivier Bodenreider, Mary Regina Boland, Ivo Georgiev, Hongfang Liu,
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Livingston</surname>
          </string-name>
          , et al.
          <article-title>The digital revolution in phenotyping</article-title>
          . Briefings in bioinformatics,
          <volume>17</volume>
          (
          <issue>5</issue>
          ):
          <fpage>819</fpage>
          -
          <lpage>830</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Rutter</source>
          , 2005]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Rutter</surname>
          </string-name>
          .
          <article-title>How the environment affects mental health</article-title>
          .
          <source>The British Journal of Psychiatry</source>
          ,
          <volume>186</volume>
          (
          <issue>1</issue>
          ):
          <fpage>4</fpage>
          -
          <lpage>6</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
          </string-name>
          et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Honghan</given-names>
            <surname>Wu</surname>
          </string-name>
          , Anika Oellrich, Christine Girges, Bernard de Bono,
          <string-name>
            <given-names>Tim J.P.</given-names>
            <surname>Hubbard</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Richard J.B.</given-names>
            <surname>Dobson</surname>
          </string-name>
          .
          <article-title>Automated PDF highlighting to support faster curation of literature for Parkinson's and Alzheimer's disease</article-title>
          .
          <source>Database</source>
          ,
          <year>2017</year>
          (1):bax027,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>[Zhou</surname>
          </string-name>
          et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>XueZhong</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Jo¨rg Menche,
          <article-title>AlbertLa´szlo´ Baraba´si, and Amitabh Sharma. Human symptoms-disease network</article-title>
          .
          <source>Nature communications, 5</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>